This is a story about the incompressible 3 minutes planning time of LLM that I turned into 4 sec

At first, it felt like one of those constraints you simply have to live with. The model was doing planning work; planning is expensive; three minutes was the price of admission. I could optimize the prompt, tweak the parameters, maybe improve the surrounding workflow—but the core latency seemed immovable.

Except it was not.

The real lesson was not “make the model faster.” It was realizing that the three minutes were not an inherent property of the problem. They were a property of the approach I had accepted.

The wall

What? How can it be so long? That's not admissible
25 tool calls = 25 extra LLM round-trips before it writes a word. That's the whole 175s

Those are all reasonable engineering responses. But they also quietly turn a limitation into an assumption.

I was confronted with a simple question: Is three minutes actually required, or have I just decided it is?

TLDR: The Conclusion

By providing deterministically the exact data the LLM needs, it no longer needs to explore

For real?

0:00
/0:33

codex with and without pixel

For real.

(Optional) The detail

More interesting for people willing to see how I work with LLM.

unequaled-reason-orphaned-branch
User2026-10-07 22:12:19#586

That would be nice if the plan could be built based on the brief, no?

Assistant2026-10-07 17:12:22#588

Yes — brief → plan is the right contract. The bench currently stuffs a bare TASK: line; better to feed the planner a structured execution brief (goal, context, constraints, acceptance) and let it decompose. Adding --brief <path> and demoing it.

Assistant2026-10-07 22:12:19#605

wafer/Kimi-K3 is registered in your own fleet — extensions/providers-fleet.ts line ~380 (w("Kimi-K3", ...) under the Wafer provider, WAFER_API_KEY is set). It's a normal lane.

I picked it because in wave 2 it produced the highest-quality plan of any model — worker briefs with exact old/new code embedded, correct scoping, do/don't notes. It just does a lot of repo reading to get there, so it's slow and this time hit the 240s timeout.

Not a bad choice, just the slow/quality end of the spectrum vs gpt-oss-120b's fast/shallow end. If the splitter lane is supposed to be fast, drop it; if quality matters more, raise its timeout (~400s+).

User2026-10-07 22:12:19#607

can you try wafer/GLM-5.3 ?

Assistant2026-10-07 17:12:22#616

GLM-5.3's plan looks genuinely excellent — let me find where the JSON actually breaks.

Assistant2026-10-07 17:12:22#618

Found it — GLM-5.3 emitted a literal tab character inside a JSON string (\tbg,), which is invalid JSON. The plan itself is the best one yet (precise anchors, real file paths it verified by reading). The consumer should be tolerant — adding a scanner that escapes control chars only inside strings.

Assistant2026-10-07 17:12:22#622

Now re-parse GLM-5.3's saved output through the fixed sanitizer — no need to re-run the model:

User2026-10-07 22:12:19#625

How many seconds are we looking for each step?

Assistant2026-10-07 22:12:19#626
1,355 chars — click to expand

Measured across the runs this session (plan step only, brief→```plan out):

Splitter Fastest Typical Worst
openrouter/gpt-oss-120b 2.6s ~3–10s 9.8s
openrouter/mercury-2.5 9s ~10s timeout at 300s (flaky)
antigravity/gemini-3.8-flash — ~50s —
wafer/GLM-5.3 77s ~80–150s 148s
wafer/Kimi-K3 137s ~140s >240s timeout
wafer/DeepSeek-V4.1-Flash 153s ~155s —
wafer/DeepSeek-V4-Pro — ~180s —

Key insight: the gap isn't decode speed — it's tool rounds. gpt-oss plans almost straight from the brief (few repo reads → occasionally wrong paths). Wafer models do multi-turn read loops; that's where the 80–180s goes, and it's what buys correct file paths.

Suggested budgets per step:

scout (read-only)   ~10–60s
planner fast lane   ~10s    timeout 60s
planner quality     ~80s    timeout 300s
worker (edit)       ~30–120s
reviewer            ~20–60s

If wall-clock is the priority: gpt-oss splitter (10s) + Pixel path validation (0s) beats a quality splitter spending 80s rediscovering what Pixel already knows. The hybrid — brief → gpt-oss decompose → Pixel verify files → dispatch — is probably your best seconds-per-correct-plan.

UNVERIFIED: scout/worker/reviewer numbers are estimates, not measured this session — only planner wall times were.

User2026-10-07 22:12:19#627

Planner quality is too long. Can we use a Siri brass and GLM 4.7?

User2026-10-07 22:12:19#628

cerebras

Assistant2026-10-07 22:12:19#649
1,230 chars — click to expand

Results with GLM-4.7 unblocked:

Model Result Wall
openrouter/z-ai/glm-4.7 timeout — 24 tool executions in 240s, never emitted a plan >240s
openrouter/deepseek-v4.1-flash ✔ valid 3-item plan (2 workers + reviewer, correct file set, even a docs-update item) 111s
openrouter/gpt-oss-120b ✔ 2 items 23s

GLM-4.7 works now — it just plans like a senior engineer: it spent the entire 4 minutes doing repo recon (24 tool calls) and never finished writing. That pattern will sometimes produce the best plan of the fleet, but at these latencies it can't be the default splitter either. It may deserve a second shot with --thinking high off and a bigger timeout, or a scout-feeding pipeline where Pixel hands it the files so it doesn't explore.

Current best splitter config: openrouter/openai/gpt-oss-120b — fast (3–23s), always parseable, occasionally wrong paths → Pixel-validates files. deepseek-v4.1-flash (OR) is the middle option at ~111s with better grounding.

VERIFIED: guardrail PATCH applied via OpenRouter API (canonical slugs confirmed in response); 3 live planner runs against the statusbar brief. GLM-4.7 timeout is a real measured result, not a config error.

User2026-10-07 22:12:19#651

What about using Olama with DeepSync V4 Pro?

Assistant2026-10-07 17:12:22#660

The ollama provider lives in extensions/provider-ollama.ts (not the npm path my bench assumed). Adding the bench mapping + a DeepSeek V4 Pro entry, then testing whether Ollama cloud actually serves it.

Assistant2026-10-07 22:12:19#673

Yes — ollama/deepseek-v4-pro:0813-cloud works. Both Ollama cloud lanes produced valid plans:

Model Items Depth Wall Quality
ollama/deepseek-v4-pro:0813-cloud 7 1 175s 🥇 repo-grounded — found all 5 real statusbar files + config, exact-text anchors per edit, docs item, parallel workers then reviewer
ollama/gpt-oss:120b-cloud 5 3 64s correct files, proper dependency chain, richer than its OR twin

So V4 Pro on Ollama is a legitimate quality splitter — comparable to wafer/GLM-5.3's output (~77s) but slower. Note the OR guardrail whitelist means deepseek/deepseek-v4.1-flash also works via OR at ~111s.

Final lane map for the splitter:

fast     openrouter/openai/gpt-oss-120b     ~3–23s   + Pixel validates paths
mid      openrouter/deepseek/deepseek-v4.1-flash  ~111s
quality  wafer/GLM-5.3  |  ollama/deepseek-v4-pro  ~80–175s

VERIFIED: 2 live runs, both plan blocks parse and validate; deepseek-v4-pro:0813-cloud confirmed in ollama list and added to provider-ollama.ts. GLM-4.7 remains unproven (>240s exploring, 0 output).

User2026-10-07 22:12:19#676

How much time does DeepSync V4 Pro take for the plan?

Assistant2026-10-07 22:12:19#677

~175–180s.

  • ollama/deepseek-v4-pro:0813-cloud on the statusbar brief: 175.3s → 7-item plan
  • wafer/DeepSeek-V4-Pro on the earlier task prompt: 179.6s → 2-item plan

Consistent ~3 minutes regardless of provider — it's a deep-reading model, most of that time is repo tool calls, not decode. For the splitter lane it's a quality-tier choice, not a speed one.

User2026-10-07 22:12:19#679

What? How can it be so long? That's not admissible. How can I do to get faster? When I use DaVinci Li, they do the plan almost immediately.

Assistant2026-10-07 17:12:22#680

Because the 175s isn't decode — it's repo exploration. Let me count the tool rounds.

Assistant2026-10-07 22:12:19#682

25 tool calls = 25 extra LLM round-trips before it writes a word. That's the whole 175s — V4 Pro is also a heavy reasoner, so each round-trip is slow. DaVinci/Claude Code-style planners feel instant because they plan from context that's already gathered (indexed/pre-loaded), not by exploring.

Fix: **do the recon once, embed it in the brief,

User2026-10-07 22:12:19#684

But then if you have to do repo exploration, it means that you are badly using Pixel and you are stupid.

Assistant2026-10-07 22:12:19#685

Exactly — the model re-deriving the repo map is wasted round-trips. Pixel already has the index; the brief should ship the excerpts, and the planner should run --no-tools (one shot: brief → plan).

Checking pixel's query interface, then wiring --no-tools + embedded context into the bench and re-running V

User2026-10-07 22:12:19#687

There is a way, I think, with Pixel and chaining a command line to basically reduce this expiration.

Assistant2026-10-07 17:12:22#690

pixel execution-brief is exactly that chain — deterministic, bounded, no LLM. Wiring it:

Assistant2026-10-07 17:12:22#694

105ms for the full evidence brief — that's the recon, done deterministically. Now adding --no-tools to the bench so the planner can't re-explore, and chaining: pixel execution-brief → brief file → planner.

Tool2026-10-07 17:12:22#698

Output from command in shell bcf427:

🟩 pixel execution-brief ❀ 88.6ms ❀ #44e6b8 7799

Exit code: 0

Assistant2026-10-07 22:12:19#706

Chain works:

Model With tools Pixel brief + --no-tools
ollama/deepseek-v4-pro 175s 74.8s (2.3×)
openrouter/gpt-oss-120b 10–23s 2.2s

The command chain:

pixel execution-brief "<task>" . --max-tier P1 >> brief.md
bun run plan-bench -- --brief brief.md --no-tools --models ollama/deepseek-v4-pro:0813-cloud

Interesting: V4 Pro's tool-less plan is good (proper scout→worker→reviewer DAG, correct files) — but it still spawned a scout item asking for "exact, verbatim, line-numbered excerpts" that the execution-brief didn't include. It wants anchor lines, not just file lists.

So next lever if you want sub-10s quality plans: enrich the evidence with pixel pack-context or list-signatures on the P0 files — give the planner the actual anchor text so it doesn't delegate recon back to a scout. Then even the quality models become near-instant.

VERIFIED: both runs live, plan.json parsed + schema-validated; pixel brief generated in 88.6ms.

User2026-10-07 22:12:19#708

Ok, let's enrich the evidence with the PixelPak Comtex only signature. No problem, let's go!

Assistant2026-10-07 17:12:22#711

Signatures give the map; planner needs anchor text. Pulling the key verbatim excerpts into the brief:

Assistant2026-10-07 17:12:22#719

All anchors captured. Building the enriched brief — goal + pixel execution-brief + verbatim excerpts of StatusBarData, SubagentMetrics, StatusBarSegment, createData, createBuiltinSegments, and the existing subagent:snapshot handler pattern:

User2026-10-07 22:12:19#822

And then actually, I realized that I need some kind of a good to‚Äëdo list in order to feed my sub‚Äëagent. Do you think I could try to make one? You can try for me with GPT-OSS120B, with DeepSeq V4-1, and some other models to see the quality.

Assistant2026-10-07 17:12:22#842

Got it — I'll reuse the repo's own planner spec, run the same task headless through several fleet models, and compare the todo lists they produce.

Assistant2026-10-07 17:12:22#844

The planner spec is in agents/planner.md. Plan: run that spec + a fixed task headless through pi --mode json -p across 5 fleet models, parse the ````plan` block from each, compare quality.

Spawning a bench script:

Assistant2026-10-07 17:12:22#873

OpenRouter key is set but gpt-oss-120b isn't in the provider's model list — adding it plus glm-4.7 (for the pixel-classify lane), then rerunning the bench through OpenRouter.

Assistant2026-10-07 22:12:19#907
2,300 chars — click to expand

Done — bench tool built (bun run plan-bench), two waves run across your fleet. Results:

📊 Plan-generation quality

Model Items Valid JSON Wall Verdict
wafer/Kimi-K3 2 ✔ (extractor edge) 137s 🥇 best briefs — embeds exact old/new code blocks, do/don't notes, verification steps
wafer/GLM-5.3 2 ✔ (extractor edge) 148s same depth — even diagnosed my extractor bug in its preamble
wafer/DeepSeek-V4.1-Flash 3 ✔ 153s correct deps, decent briefs
wafer/DeepSeek-V4-Pro 2 ✔ 180s similar, slower
openrouter/openai/gpt-oss-120b 2 ✔ 2.6s ⚡ 60× faster; thinner tasks, and embeds ```ts inside JSON → breaks the fence contract
antigravity/gemini-3.8-flash 2 ✖ fence noise 50s detailed but sloppy formatting
openrouter/inception/mercury-2.5 1–2 ✔ 9–300s shallow (single reviewer item) — not a splitter, flaky latency

⚠️ Dead lanes found

  • zai/* — GLM Coding Plan expired (429, needs renewal)
  • cerebras/* — 402 insufficient funds
  • OpenRouter guardrails block glm-4.7, deepseek-v4.1-flash, glm-5.3-flash, minimax-m3 — only gpt-oss-120b + mercury-2.5 pass (configurable at openrouter.ai/workspaces/default/guardrails)
  • antigravity/claude-sonnet-4-6 deprecated → says use 5.5
  • ollama-cloud/gpt-oss:120b — model id not found under that provider name

🎯 Recommendation

ChatGPT's Cerebras call is dead on this account, but the shape is right: splitter = wafer/Kimi-K3 or GLM-5.3 (quality) vs openrouter/gpt-oss-120b (speed tier, needs a stricter output contract). Workers on Mercury is fine — it's the splitter job it can't do.

One spec bug surfaced: agents/planner.md tells planners to embed "exact code excerpts" but the ```plan fence contract can't survive nested code fences inside JSON strings. Worth adding "no markdown fences inside item strings" — that's what broke GPT-OSS's first run.

VERIFIED: 11 live pi --mode json -p planner runs executed; plans parsed from run-artifacts/plan-bench/*.plan.json; two extractor bugs fixed and re-run. New file bin/pi-plan-bench.ts + package.json script — no other source touched except adding gpt-oss-120b/glm-4.7 to the OpenRouter model list in providers-fleet.ts.

User2026-10-08 14:31:02#1469

Okay, perfect, so we will use GPT-OSS on Android MTB. And like you said, so Alos is already implemented.

Assistant2026-10-07 22:12:19#1470

Locking it in then. Let me wire openrouter/openai/gpt-oss-120b as the planner lane in the harness — checking where planner model selection lives.

pixel is still evolving, but some satisfying outcomes start to unveil.
Find the repo here (and come to contribute, the impact is huge!).

Cheers 🍺


This article was an exploration of the green part the Pi agent (boosted by pixel) skeleton that I want to achieve. If you liked it, stay tuned for more!

This article was an exploration of the green part the Pi agent (boosted by pixel) skeleton that I want to achieve. Everything must be less than 10 seconds or I fail. If you liked it, stay tuned for more!