// Orden Works

sial

AI Operating Platform

Idea

AI is doing the work.
You’re still holding it together.

Agents write code, reconcile invoices, draft campaigns, and make a prototype look easy. Production is still a long road: integrations, permissions, failures, and the daily work of keeping it all running.

You still carry the context between the tools, chase the next step, and turn yesterday’s mistakes into tomorrow’s fixes. The tasks are automated. The responsibility still lands on you.

It builds. It runs. It improves.

Sial carries work from prototype into production, then manages the day-to-day: planning, scheduling, delegating, and following up. Shared context keeps it connected; permissions, budgets, and approvals keep it within the limits you set.

It reviews outcomes, remembers what it learns, and proposes better procedures for your approval. Operation feeds improvement. You set the direction; Sial manages the work and learns from doing it.

1.09×

the work quality of Claude Code

Same model on both sides — the harness is the only variable.

55.8×

cheaper than Claude Code on Opus 5

And still judged higher on the quality of the work produced.

89.2%

LoCoMo, strict

1,540 questions across 10 long conversations.

29.8 ms

to recall everything it knows

Median, against a 50 ms budget, at ten times the tested scale.

A harness, a memory, and an operating system.

The harness plans and executes. The memory carries knowledge forward. The operating system runs and governs the work. Together, they give Sial the means to build, operate, and improve in one persistent workspace.

The resident dispatches. Agents create, maintain, and operate your apps: finance tools, marketing operations, software delivery, engineering workflows, and whatever your business needs next.

orden.worksSIAL
sial.monitorv17
SIAL APP
sial://work1 Sial · 3 executing · 4 open goals · month $18.42

NOW — 3 EXECUTING

GOALS — ALL PROJECTS · 4

UPCOMING · SCHEDULED WAKES · 48H · 2

sat 12daily operating reviewAlvaro Matawake
sat 12editorial calendar checkAlvaro Matawake

Technology

Harness

A model thinks. The harness decides what it thinks about.

The loop that turns a model into an agent which plans, delegates and finishes. Four decisions produce the difference in the numbers below: how context is kept, how work is divided, how your conventions are recalled, and how conduct adapts while the conversation moves.

Sial vs Claude Code- DeepSeek v4 Flash Vision in both harnesses

quality1.09×speed1.25×savings2.15×

Sial vs Claude Code- DeepSeek v4 Flash Vision vs Opus 5

quality1.02×speed0.79×savings55.8×

Sial vs Pi- DeepSeek v4 Flash Vision in both harnesses

quality1.06×speed1.11×savings0.89×

Sial vs Kimi Code- Kimi K3 in both harnesses

quality1.05×speed0.74×savings4.76×

quality — Judged blind by codex:gpt-5.6-sol and cross-checked with an Opus 5 rejudge of the best cells — the two rulers correlate, panel totals within one point. Hidden test oracles the agents never see, every pairing scored in both A/B positions (4 samples × 2 arrangements) to cancel position bias. The multiplier is Sial’s panel total over the reference’s.

speed — Wall-clock time for the full panel, one draw per cell, same machine and gateway for both harnesses. The multiplier is the reference’s total time over Sial’s.

savings — Metered token usage priced at each provider’s list, cache-read at 0.1×; the Opus 5 cell is self-metered by the Claude Code CLI. The multiplier is the reference’s panel cost over Sial’s.

Context as a living ledger

The Lead Agent holds an infinite conversation — effortlessly, and cheap. Facts are distilled continuously, raw turns evicted without loss, everything restored on resume. The context window stops being the limit.

Division of labour via mods

The leader agent never executes — it orchestrates specialist mods, coder, explorers, writers, analysts, as looped workflows over a shared memory. Flatten the cascade onto one strong model and scores drop; the split is load-bearing.

Procedures, recalled when relevant

How you or your team work — authored once, recalled exactly when the task matches and never loaded when it doesn’t. When the agent spots a new convention in a session, it proposes it as a draft; a human approves. Nothing publishes itself.

Adaptive behavior

A behavior window shifts the agent’s conduct as the conversation moves — topics it cares about warm or cool its tone, forbidden ground arms guardrails on first mention, and your programs steer live sessions from outside. Tool limits are enforced at dispatch, not suggested.

Memory

To improve, you need to remember.

Everything the agent has ever learned — and everything you and your team know — recalled in under 50 milliseconds. An agent that starts every session from nothing cannot get better at your work. This one recalls the right thing at the right moment, and would rather arrive empty than drown the turn in maybes.

LoCoMo

n=1,540 questions across 10 long conversations · open-domain 94.9% · temporal 85.7%

strict89.2%
answer84.9%

LongMemEval

n=100 · SE ±4.4

strict74%
answer79%

Document retrieval

n=20 documents · 50+ bench runs

P@398
R@585
R@1094

Retrieval at scale

ambient recall at 300 conversations + 100 docs · latency p50 against the 50 ms budget

ambient recall97.5% ± 2.5
cross-session100%
ambient p5029.8 ms
deep p5035.9 ms

Adaptive memory window

An ambient lane rides with every agent turn; a deep-recall lane escalates only when the agent demands deep understanding. Fast when it can be, thorough when it must be.

The window knows when to shut up

It grows when the turn touches what memory knows and empties itself on off-topic chatter — and that silence is deliberate. Injected a hundred turns an hour, it would rather arrive empty than distract the agent with maybes.

Every substrate, one memory

Facts retrieved from any prose like documents, webs, tasks, emails, past conversations, or tabular records alike — team knowledge, user facts, and derived entities resolved onto shared relational nodes.

Embeddings for meaning, term statistics for identity

Dense vectors put every same-shaped sentence about a topic in one band — the one carrying the exact name gets no advantage. Term statistics let that evidence force its way in. Two different questions, both answered; neither is a model call.

OS

An agent you cannot govern is one you cannot deploy.

Your apps, their work, and the rules they follow, in one place. Each app brings together programs, services, and screens, with clear ownership, permissions, and a record of what happened.

Programs

Repeatable work, from reconciling invoices to preparing a campaign. Each program combines automated steps, AI judgment, and human decisions as needed.

Libraries

Shared building blocks, such as a calculation or a data check. Build them once and reuse them across your apps.

Services

Start the right program when it is needed: on a schedule, when an invoice arrives, or when someone places an order.

Screens

Dashboards, forms, and controls for following the work and taking action. Each person sees what their role allows.

Approvals

Pause work for a human decision. Requests reach the responsible role, with deadlines and a clear next step if nobody responds.

Exec

Start an existing program or describe what you need in plain language. The same permissions and safeguards apply either way.

Sources & Targets

Connections to your existing tools. Sources bring in information; targets send messages, publish content, or update records, within the permissions you set.

Roles & Principals

People and agents get defined responsibilities and permissions. You control who can do what, and can trace each action back to who performed it.