— SWARM MEMORY BENCHMARK
What happens when
swarm agents get
persistent memory?
MiroFish simulates hundreds of thousands of AI agents — Brian Roemmele just ran 500K. We benchmarked 1,000 to show what cognitive memory changes. The patterns hold at any scale.
Cost basis: Basic RAG = ~6K tokens/query at GPT-4o rates ($2.50/1M input = $0.015/query). OpenViking tiered loading reduces context to ~3K tokens ($0.008). Clude returns only relevant memories via vector retrieval ($0.001). Source: OpenAI Pricing
1,000
Agents Benchmarked
50
Simulation Rounds
10
Facts Per Agent
1%
Clude Hallucination
How the benchmark works
01
Seed agents with facts
Each agent receives 10 ground-truth facts about a simulated world. These facts are stored as memories.
02
Run interaction rounds
Agents interact, share information, make predictions, and update their memories over 50 rounds.
03
Measure everything
We track hallucination rate, fact retention, prediction accuracy, cost, and behavioral consistency.