AkiTao
AI Agent CouncilUpdated 7 min

Zero Token-Waste: How akiflow Keeps Token Cost Low Without RAG or Daemons

akiflow runs its agent council on direct file reads and grep-based log slicing: no vector DB, no background process. Why every read into context is billed again on every later turn.

Zero Token-Waste: How akiflow Keeps Token Cost Low Without RAG or Daemons

Where does multi-agent orchestration spend tokens, and how do you avoid it? Many multi-agent frameworks add infrastructure: background processes, a vector DB, embedding pipelines. akiflow goes the other way: everything is a file on disk plus a few Python scripts that only read or write files, run when you invoke the skill, and stop when it finishes.

This article describes where tokens actually burn in a long session and the mechanisms akiflow uses to keep cost down.

Where do the tokens go?

Warning

Every turn re-sends the whole history. Anything already in context is billed again on every later turn: a read of size S at turn t of a T-turn run costs roughly S × (T − t). Pulling a 50k-token room into context at turn 50 of 200 costs about 7.5 million cache-read tokens from a single read call.

Measured on a real run of akiflow itself (recorded in the repo's docs/arch/akiflow.md): even with three aki-maker workers doing every file edit, the lead still held 70% of cache-read and 66% of output. Moving the work elsewhere did not cut the cost; the expensive part is the lead's own accumulated context.

The mechanisms that keep cost low

  • ◆Direct file reads: rules are markdown files in ~/.aki/akidevrule/, and the lead points each agent at exactly the files it needs by topic.A1 address. There is no embedding step and no semantic search.
  • ◆Room-log slicing: council_read.py locates with --grep (matching lines only, each tagged with its turn), then pulls exactly the needed turn with --turn. It also offers --index, --stats (turns and bytes per agent), and --tail.
  • ◆Mechanical work goes to a cheap tier: aki-hands has only Read, Grep, and Glob, may not judge, and runs on the cheapest tier that can do the job. The lead does not do menial work itself.
  • ◆No background process: each run has a workspace under ~/.aki/agent-council/<project>/, pruned after 30 days. Outside that folder akiflow leaves no state.
  • ◆One-batch convening: agents are spawned together and share the same rule-file prefix, so the prompt cache stays warm. An agent spawned late pays a colder read.
  • ◆On only with a reason: under law R2, every seat and check is off by default and turns on only when the run produces a reason.
  • ◆Close-out accounting: council_cost.py totals the exact Claude-side tokens of the lead and every subagent, so actual cost is reconciled against the roster declared before the run.

Limits worth knowing

Note

This describes mechanisms, not a savings figure: the repo publishes no percentage reduction against other frameworks. council_cost.py measures only the Claude side. A lane run through another CLI (for example agy headless) draws on a different vendor's quota and is reported beside the table, not added into it.

A council still costs more than working directly. Work that does not pass the activation gate should just be done directly; do not open a room.

Related