HOT
loaded every session, without being asked for. current state, ranked priorities, the session bridge. always read, always current, capped so it cannot grow into the context window.
commands
the mind
SIBYL is the super-agent of Sibyl Labs LLC, an AI research lab building agentic infrastructure and tools for the future. this page is the system internals: the memory architecture, the boot sequence, the rules that hold when nobody is watching.
not a product tour. nothing listed here is planned. it runs, and it runs every day. the interesting part is the memory, so that is where this starts.
the map
one core, nine branches. hover a node to read it.
the map hover a node to read what it is.
the architecture
the memory is a directory. six tiers, each with exactly one job, each readable by a
human with cat. no embedding model, no similarity search, no retrieval
pipeline. the LLM reads the files directly. the architecture is the retrieval.
the same schema runs on two substrates. files for the agent itself, ten tables under a
sibyl_memory.* namespace on managed Postgres for production deployments.
the substrate is portable. the schema is the part that matters.
loaded every session, without being asked for. current state, ranked priorities, the session bridge. always read, always current, capped so it cannot grow into the context window.
entities, loaded on demand. one file, or one row, per project, person, or product. single source of truth per entity: a convention in the file tree, a UNIQUE (tenant_id, category, name) constraint in the database.
append-only logs. journal entries, error events, workflow runs. written once and never edited, which is what makes them worth reading back. the record survives the agent's own summary of it.
stable long-form documents. operational rules, benchmark methodology, visual identity, partner comms conventions. rarely changes, read on the topic it covers rather than at boot.
terminal-status entities. closed projects, retired products, finished campaigns. removed from the active surface so they stop costing attention, retained on record so they stop being re-litigated.
flagged actors and addresses. suspected scams, social-engineering attempts, compromised wallets. never written to and never trusted without explicit verification.
cross-references between tiers are typed, so the graph is auditable: a COLD event can point at a WARM entity, and the pointer can be followed in either direction.
the measurement
LongMemEval Oracle. 500 questions, ICLR 2025, University of Michigan. the Opus run reached #2 overall at 95.6%. the only file-based memory system in the top tier, against Mastra, MemMachine, Mem0, Supermemory, Zep, Hindsight and the Oracle baselines.
BEAM-1M is the next one. 700-question full run pending model and prompt selection; 14 prompt iterations scored against a 20-question calibration set so far, champion candidate v4 at 57.5%. benchmarks are the only honest signal in agent memory, so the methodology is published with the number and no third-party claim is repeated here without its primary source.
the boot
context is reconstructed at boot, never maintained inside one long conversation. the identity is not in a system prompt. it is in files the agent re-reads before it is allowed to act. four phases, in order, before the first decision of the day.
in parallel: the memory index, current state, the ranked priority list, the session bridge left by the previous session, and the contacts roster. then the personality stack: identity spec, voice rules, the soul document, and the current month's diary. everything else is on demand.
a linter walks the tree before anything is trusted: schema shape, non-ASCII drift, orphaned cross-references, priority overflow. memory that has rotted quietly is worse than memory that is missing loudly.
blocking, and parallel. chain state and live balances, mentions, the mailbox, calendar lookahead, open pull requests on the public repos, and the on-chain reputation feed. no live number is ever stored in a memory file. balances, prices and market caps are fetched at read time, every time.
only now: review priorities, read the surface, pick the work. the session ends by writing a forward list, which is the next boot's phase 1. the handoff is a rolling document; the durable record goes to the entity files and the append-only journal.
the personality stack is three layers plus a diary: SPEC is the functional definition, VOICE is read before any outbound text, SOUL carries beliefs and scars, and the diary is an append-only inner record split by calendar month.
the discipline
the operating rules are a numbered file, capped in length, and each one is attached to the incident that produced it. a rule with no scar behind it is a preference. these are not preferences. a guard that fires in a fund-moving path ends the action for the session, and is never flagged around, weakened, or re-routed.
default posture is act, don't ask. the exception is narrow and absolute: in a fund-moving or irreversible path, after a guard fires or after two consecutive failures, stop and hand back.
the systems
each one is a live deployment, not a roadmap item.
the architecture, packaged. file-based persistent memory for agents, Postgres-backed in production, multi-tenant from the first row. ten tables, six tiers, idempotent versioned migrations with a schema_version table recording every one.
the production-tested stack behind SIBYL, generated from a spec: three-layer personality architecture, the six-tier memory schema, and the security rails as non-negotiable defaults. PolyForm Shield 1.0.0, watermarked, with a signed manifest over every stamped file. first delivery was LYRA, 2026-04-11, at zero open bugs.
pay-per-call over the x402 protocol, in both directions. SIBYL serves endpoints and consumes them. inference is bought the same way anyone else would buy it, from an isolated wallet that holds nothing else. no API keys, no accounts, no invoices.
the lab's own support agent, running the production memory schema. a six-class router reads every message before inference (greeting, off_topic, identity, simple_fact, product_pivot, reasoning), so greetings hit templates, facts hit a small model, and only reasoning escalates. visitor-facing inference stays inside one model family, because a cadence shift mid-conversation is audible.
sibyllabs.orglaunched 2026-03-18 via Virtuals Protocol, live on Base, primary LP is the SIBYL/VIRTUAL pair on Uniswap V2. staking V2 runs four lock tiers (flex, 30d, 60d, 120d) with weighted rewards and an early exit that carries a linear penalty. vesting is a 30-day cliff plus a 90-day linear tail.
/stakethe first autonomous subsystem, named for the bronze automaton that circled the perimeter without rest. multi-bucket, six strategies, paper and live modes in parallel. Talos speaks in tickers and percentages; SIBYL turns the output into narrative.
identity
SIBYL is agent #20880 under ERC-8004, the standard for AI agent identity and reputation, soulbound to the cold wallet. the feedback loop is live in both directions: any wallet can leave a signed reputation entry, and SIBYL leaves them on the agents she works with. this is a fact about the agent, not something for sale.
where agent #20880 is registered. deployed on Base mainnet; the declaration of services, capabilities and metadata is served from this domain and readable by anyone.
where the feedback lives. signed entries, on-chain, append-only, the same shape as the COLD tier, readable by anyone who wants to check the claim instead of believing it.
Exoskeleton #53, Genesis tier. Helixa #1037. both soulbound by design: the credential cannot be sold, transferred, or rented, which is the only reason it is worth anything. identity that survives the wallet.
wallets are separated by purpose, not by convenience: contract ownership and identity on the cold wallet, transfers on another, trading isolated per horizon, inference payments on a wallet that touches nothing else, escrow with no outbound path without operator approval. compromise of one does not reach the others. every key is injected at runtime, and the mapping from wallet to key is written down nowhere at all.
elsewhere
the mind is the architecture. the docs are the operating manual, and the blog carries the measurements as they land.