Introduction

ARES (Agentic Runtime Extensible Server) is an agent server written in Rust. It runs chat agents with multi-provider large language model (LLM) support, tool calling, and retrieval-augmented generation (RAG). It also speaks the Model Context Protocol (MCP) and serves a web UI.

The Cordis kernel

Every ARES component is a service on a typed Context. Services declare dependencies, and the Cordis kernel resolves them at run time. You can add, replace, or intercept services while the server runs.

ARES × Cordis — the live service graph, zero-downtime provider replacement, guarded operations, and fail-closed policy gates

How the pieces fit

ARES has four layers. Each layer depends only on the layer below it.

  1. Kernel. The Cordis kernel (ares-cordis) owns one typed Context graph, fibers, events, and intercepts. It knows nothing about agents or LLMs.
  2. Capability crates. Crates under crates/ register services on that context: storage (ares-store), tools (ares-tools), providers (ares-llm), agents (ares-agent), retrieval (ares-rag), protocol glue (ares-mcp), and HTTP routes (ares-http). They talk to each other through services on the context, not through direct wiring.
  3. Server facade. The ares-server package holds the binary and the library facade. It boots the context, registers factories, applies the entries program, and binds the listener.
  4. Surfaces. The CLI, the HTTP API, the MCP server, and the embedded web UI are entry points into the same running context graph.
flowchart TD
    CLI[CLI] --> Facade
    HTTP[HTTP API] --> Facade
    MCP[MCP server] --> Facade
    UI[Web UI] --> Facade
    subgraph Facade [ares-server]
        Boot[Boot + bind]
    end
    Boot --> Kernel[cordis kernel: Context, fibers, events]
    Kernel --> Capabilities[capability crates]

The layering gives you two properties. First, you can build a smaller server by dropping capability crates, because nothing above the kernel hard-depends on a capability. Second, you can replace a capability at run time through the loader, because consumers resolve it through the context at call time.

What you can do with ARES

You use ARES in three ways:

  • Command line interface (CLI): the ares-server binary starts the server. It also scaffolds projects, inspects configuration, and drives RAG collections from your terminal. See the Command Line Interface chapter.
  • HTTP server: the same binary exposes an HTTP API on Axum. Clients send chat requests over POST /v1/chat with an API key. Streaming uses Server-Sent Events (SSE). See the HTTP API chapter.
  • Rust library: the ares-server crate is also a library facade. You inject Execute, Tools, and Llm on a Cordis Context and run an agent with no HTTP stack. See ARES as a Library.

New to ARES? Follow Installation, then build your First Server.

Key capabilities

  • Multi-provider LLM routing through one API: OpenAI-compatible endpoints, Azure AI Foundry, AWS Bedrock, and Ollama.
  • Tool calling: define tools in configuration. ARES runs the tool-call loop and assembles the response.
  • Retrieval-augmented generation: ingest documents into collections and ground responses in them.
  • Workflows: chain agents into multi-step flows with entry and fallback agents.
  • Multi-tenancy: tenant isolation, scoped API keys, quotas, and usage metering.
  • Hot reload: edits to ares.toml and TOON files apply without a restart.
  • Supervision: the built-in supervisor restarts the server after hot-restart exits.
  • MCP integration: bridge external MCP servers as agent-callable tools.

Where to go next

ChapterPurpose
InstallationInstall or build the ares-server binary
First ServerScaffold, validate, run, and call the server
Command Line InterfaceEvery subcommand and flag

Reading paths

Pick the row that matches your goal. Each path lists chapters in reading order.

I want to...Read, in order
Run a server todayInstallation, First Server
Call the API from my applicationHTTP API; skim First Server for tenant and key setup
Embed ARES in my own Rust binaryARES as a Library, then Kernel Patterns in Rust
Understand what happens at bootSystem Overview, then Ideas and Map
Replace services while the server runsFiber Lifecycle, Interception Points, Configuration and Deployment
Define agents, tools, or workflowsAgents and Skills; Command Line Interface for scaffolding commands
Ground answers in my documentsRetrieval

Installation

This chapter installs the ares-server binary. Pick one method:

  • Install from crates.io with cargo install.
  • Build from source with cargo build.

Prerequisites

You need a Rust toolchain. The crate declares rust-version = "1.98", so use Rust 1.98 or newer. Check your version:

$ rustc --version

Install Rust through rustup if you do not have it.

Install from crates.io

The crate is published as ares-server. Install it without pinning a version:

cargo install ares-server

To include the embedded web UI in the build, add the ui feature:

cargo install ares-server --features ui

The install puts the binary in $HOME/.cargo/bin. Make sure that directory is on your PATH.

Build from source

Clone the repository and build the release binary:

git clone https://github.com/dirmacs/ares
cd ares
cargo build --release

The binary lands at target/release/ares-server. Copy it to a directory on your PATH, or call it by path.

Feature flags

Features select LLM providers, database backends, and vector stores. The table lists every feature of the ares-server package.

FeatureWhat it enables
defaultpostgres, openai, ares-vector, mcp, inventory, rhai-policy
openaiOpenAI API and compatible endpoints such as NVIDIA NIM
azureAzure AI Foundry chat completions
bedrockClaude on AWS Bedrock
postgresPostgreSQL tenant database through sqlx (default)
tursoTurso/libSQL, an edge-native SQLite-compatible store
ares-vectorEmbedded pure-Rust vector store with HNSW (default)
lancedbLanceDB embedded vector store; needs protoc
qdrantQdrant vector database client
pgvectorpgvector, a PostgreSQL extension for vectors
chromadbChromaDB vector database client
pineconePinecone managed vector database (alpha)
mcpMCP protocol glue, client, auth, and registry
inventoryCordis static registration at compile time (default)
rhai-policyRhai policy scripts on kernel events (default)
eruka-contextPer-agent context injection from Eruka
local-embeddingsONNX local embedding models; not on Windows MSVC
hmrHot swap of compiled plugins through dlopen
skillsSKILL.md discovery and loading
emailEmail sending over SMTP
search-toolsWeb search and scraping tools
uiEmbedded Leptos web UI served by the backend
swagger-uiInteractive API documentation pages

Feature bundles combine several flags:

BundleContents
all-llmopenai, azure, bedrock
all-dbpostgres
all-vectorstoresares-vector, qdrant, pgvector, chromadb, pinecone
local-vectorstoresares-vector only
fullAll LLM providers, postgres, qdrant, ares-vector, mcp, swagger-ui
full-uifull plus ui
minimalNothing optional

Choose feature combinations

Features compose along three independent axes. Pick one option per axis:

  1. LLM providers (openai, azure, bedrock, or none for Ollama). These add provider clients to ares-llm. They do not interact with each other, so all-llm is safe when you want runtime choice.
  2. Database backend (postgres or turso). The server binary requires the postgres feature. A binary built without it prints a rebuild hint and exits with code 1 at startup (src/main.rs compiles a stub main without it). Keep postgres unless you embed the library and run no HTTP server.
  3. Vector store (ares-vector, qdrant, pgvector, chromadb, pinecone, lancedb). Clients are additive. local-vectorstores keeps the build small because only the embedded store compiles.

Cross-axis rules worth knowing:

  • postgres also gates sqlx code paths in ares-store, ares-agent, ares-mcp, ares-tools, and ares-http through feature forwarding.
  • mcp, inventory, and rhai-policy ride in default; dropping default drops all three. Re-add them explicitly if you build with --no-default-features plus your own picks.
  • swagger-ui needs nothing extra, but the OpenAPI document includes RAG paths only when both local-embeddings and ares-vector are on (see the #[cfg(all(...))] gate around the OpenApi derive in src/main.rs).

Some features cost real compile time or native dependencies:

FeatureCost
lancedbNeeds the protoc compiler on PATH at build time
local-embeddingsPulls the ONNX Runtime; unsupported on Windows MSVC; slow link step
uiBuilds the embedded Leptos UI as part of the crate; longest cold build of any single feature
full-uiEverything above together; budget several minutes on a modest machine

For a first install, stay on defaults plus what you actually call. Defaults already give you postgres, openai, ares-vector, mcp, inventory, and rhai-policy.

Build offline or air-gapped

The repository ships no vendor directory. For an air-gapped machine, vendor dependencies on a connected machine first:

cd ares
cargo vendor vendor

Copy the whole tree, including vendor/, to the target machine. Then point Cargo at it through .cargo/config.toml next to Cargo.toml:

[source.crates-io]
replace-with = "vendored-sources"

[source.vendored-sources]
directory = "vendor"

Build with cargo build --release --offline. Two notes apply:

  • SQL migrations live inside the ares-store crate and ship inside the published package, so an offline build needs no external migration files.
  • The default TLS stack uses rustls, so you need no system OpenSSL headers. If a non-default feature drags in OpenSSL on a host without pkg-config/libssl-dev, enable its vendored form in Cargo.toml (see the commented vendored example near the end of the dependency list) instead of installing system packages.

Troubleshoot installation

SymptomCauseFix
package \ares-server v0.10.0` cannot be built because it requires rustc 1.98 or newer`Toolchain older than the declared rust-versionRun rustup update stable, then retry
Installed binary prints requires the \postgres` feature` and exits 1Built or installed with --no-default-features or without postgresReinstall with --features postgres, or keep default
error: failed to run custom build command naming protoclancedb enabled without Protocol Buffers compilerInstall protoc, or drop lancedb from --features
Link errors mentioning ONNXRuntime under local-embeddingsMissing ONNX Runtime library, or Windows MSVC hostInstall ONNX Runtime, or use a remote embeddings endpoint without the feature
ares-server: command not found after install$HOME/.cargo/bin missing from PATHAdd export PATH="$HOME/.cargo/bin:$PATH" to your shell profile
Build succeeds but /ui returns 404ui feature absent from this binaryRebuild with --features ui

Compile-time versus run time matters here. Features such as openai, azure, or bedrock decide which provider code exists inside the binary. A provider that is absent at compile time cannot appear at run time by editing ares.toml. Configuration selects among compiled-in options; it never adds new ones.

Verify the install

Print the version:

$ ares-server --version
ares-server 0.10.0

First Server

This chapter runs your first ARES server. You scaffold a project, validate the configuration, start the server, and send one chat request.

The commands use ares-server from your PATH. Substitute your binary path if you built from source.

Scaffold a project

Run init in an empty directory:

mkdir my-ares && cd my-ares
ares-server init

init accepts a target path, --force to overwrite, and these switches:

  • --no-examples skips the example TOON files under config/. Verified against 0.10.0: the run then creates only ares.toml, .env.example, .gitignore, and empty directories.
  • --provider openai|both selects which provider template lands in ares.toml; the default is ollama.
  • --minimal asks for a smaller configuration. In 0.10.0 both modes write the same file set; prefer plain init.

The command creates this layout:

.
├── ares.toml          # Main configuration file
├── .env.example       # Template for environment variables
├── .gitignore
├── data/              # Local database and RAG data
└── config/
    ├── agents/        # router.toon, orchestrator.toon
    ├── models/        # fast.toon, balanced.toon, powerful.toon
    ├── tools/         # calculator.toon, web_search.toon
    ├── workflows/     # default.toon, research.toon
    └── mcps/          # TOON MCP server definitions (empty)

Inspect the generated configuration

Open ares.toml. The generated file contains these sections (comments stripped):

[server]
host = "127.0.0.1"
port = 3000
log_level = "info"

[auth]
jwt_secret_env = "JWT_SECRET"   # Name of the env var that holds the JWT secret
jwt_access_expiry = 900         # Access token lifetime in seconds
jwt_refresh_expiry = 604800     # Refresh token lifetime in seconds
api_key_env = "API_KEY"         # Name of the env var that holds the service API key

[database]
url = "./data/ares.db"          # Local SQLite-compatible store; or a postgres:// URL

[providers.ollama-local]
type = "ollama"
base_url = "http://localhost:11434"
default_model = "ministral-3:3b"

[tools.calculator]
enabled = true
description = "Performs basic arithmetic operations (+, -, *, /)"
timeout_secs = 10

[agents.router]
model = "meta/llama-3.3-70b-instruct"
tools = []
max_tool_iterations = 1
parallel_tools = false
system_prompt = """..."""        # Routing prompt; answers with one agent name

[agents.orchestrator]
model = "meta/llama-3.3-70b-instruct"
tools = ["calculator", "web_search"]
max_tool_iterations = 10

[workflows.default]
entry_agent = "router"
fallback_agent = "orchestrator"
max_depth = 3

[config]
agents_dir = "config/agents"    # TOON directories, watched for hot reload
hot_reload = true
watch_interval_ms = 1000

A [rag] block also appears (embedding_model = "BAAI/bge-small-en-v1.5", chunk size and overlap). Retrieval reads it when you add collections.

What each block does

[server] controls the listener and logging. host and port feed the single TcpListener::bind call in run_server (src/main.rs). Keep 127.0.0.1 while you test; use 0.0.0.0 only when another machine must connect. Two more fields exist beyond the template: cors_origins (default ["http://localhost:3000"], set explicit origins in production) and the rate-limit pair below.

rate_limit_per_second = 100   # Requests per second per IP; 0 disables limiting
rate_limit_burst = 10         # Bucket size admitted above the steady rate

The limiter is tower_governor (src/main.rs, rate-limit layer build): it admits bursts up to rate_limit_burst and refills one slot every \(1/\text{rate_limit_per_second}\) seconds. Responses carry x-ratelimit-* headers.

[auth] names the secret environment variables and token lifetimes. jwt_secret_env = "JWT_SECRET" means: read the signing key from the environment variable called JWT_SECRET. The same pattern applies to api_key_env. Lifetimes are seconds: 900 gives 15-minute access tokens; 604800 gives 7-day refresh tokens.

[database] holds one connection URL. The template writes a local file path for an embedded store. A production deployment points at postgres://user:pass@host/db; the store then runs embedded migrations from ares_store::MIGRATOR on first connect.

[providers.<name>] defines one LLM provider. <name> (ollama-local here) is your label for it; agents reference this label. type picks the client implementation. base_url is where the client sends requests. default_model applies when an agent names no model of its own.

[tools.<name>] declares one tool in the main file: an enabled switch, a description the model sees, and a timeout_secs cap. The matching TOON file under config/tools/ carries the same fields for hot reload.

[agents.<name>] defines one agent. model selects the model (a provider default when omitted). tools lists callable tools; empty means the agent answers without side effects. max_tool_iterations caps the tool-call loop rounds. parallel_tools runs independent calls in one round together. system_prompt sets the persona text sent with every request. The scaffold defines two agents: router classifies a request and answers with one agent name; orchestrator does the real work and owns the tools.

[workflows.default] chains agents. A workflow enters at entry_agent and falls back to fallback_agent; max_depth bounds the hops.

[config] points at the TOON directories. With hot_reload = true, a watcher re-reads them every watch_interval_ms. Edits apply without a restart.

Set environment variables

Copy the template and fill in the values:

cp .env.example .env

Set at least these variables before you start:

export JWT_SECRET="$(openssl rand -base64 32)"   # At least 32 characters
export API_KEY="change-me-service-key"

With the default Ollama provider, also start Ollama and pull a model:

ollama serve
ollama pull ministral-3:3b

Validate the configuration

Check the file before starting:

ares-server config --validate

The command reports valid configuration or names the problem:

$ ares-server config --validate
  ✓ Configuration is valid!

  Configuration Summary

    Config file: ares.toml
    Server: 127.0.0.1:3000
    Log level: info

Validation checks structure and cross-references: unknown providers in agent blocks and malformed TOML fail here. Verified against 0.10.0: validation passes even while JWT_SECRET and API_KEY are unset, so export them anyway before you start.

Start the server

Run without arguments to start:

ares-server

What happens on boot

run_server in src/main.rs runs one ordered pass. Each step gates the next:

  1. Load environment and tracing (src/main.rs:528-532). .env is read if present, then the log filter starts at info.
  2. Create the root context (src/main.rs:534-554). A Cordis root Context appears, plus a ReflectService. The service registers notifiers for the Tools and Llm types so later changes fan out immediately.
  3. Register loader factories (src/main.rs:557-558). Built-in factories (Store, Llm, Tools, CalculatorService, and others) enter the PluginRegistry, either through explicit chains or inventory collection.
  4. Boot the entries program (src/main.rs:560-568). The loader parses config/cordis-entries.toml and instantiates entries in file order. When it reaches the Overlay entry, empty entry configs fill from ares.toml. The Store factory connects to the database, runs migrations, and seeds default agents here. Any boot failure logs Cordis Loader: boot failed and exits with code 1.
  5. Guard configuration presence (src/main.rs:570-585). No loaded config means a friendly error that points at ares-server init, then exit 1.
  6. Start the entries watcher (src/main.rs:592-625). File events re-compose the program and apply diffs through the loader journal. If the watcher cannot start, a 30-second poll takes over.
  7. Preload runtime providers and snapshot agent configs (src/main.rs:630-707). Runtime provider registrations load; current agent definitions land in the version history table.
  8. Build HTTP layers (src/main.rs:895-944). CORS applies from cors_origins; the rate-limit layer builds only when rate_limit_per_second > 0.
  9. Bind and serve (src/main.rs:949-964). The listener binds host:port, and Axum serves with graceful shutdown on Ctrl+C or SIGTERM.

If step 4 or 9 fails, you see the reason in the log before the process exits. Nothing listens before step 9, so a failure never leaves a half-open port.

Send a chat request

The external chat route is POST /v1/chat. It authenticates with a tenant API key of the form ares_....

Create a tenant and its first key. The admin routes need the ADMIN_API_KEY environment variable on the server process and the matching x-admin-secret header:

# Start the server with admin routes enabled
ADMIN_API_KEY="local-admin-secret" ares-server &

# Create a tenant
curl -s -X POST http://localhost:3000/admin/tenants \
  -H "x-admin-secret: $ADMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "local", "tier": "free"}'

# Create an API key for that tenant (use the tenant id from the response)
curl -s -X POST http://localhost:3000/admin/tenants/TENANT_ID/api-keys \
  -H "x-admin-secret: $ADMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "first-key"}'

The second response carries raw_key. Store it now; you cannot read it again.

Export the raw key for later requests:

export ARES_TENANT_API_KEY="ares_paste-the-raw-key-here"

Send the chat request with that key:

curl -s -X POST http://localhost:3000/v1/chat \
  -H "Authorization: Bearer $ARES_TENANT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"message": "What is 21 times 2?"}'

The response echoes the answering agent and the token usage:

{
  "response": "42",
  "agent": "orchestrator",
  "source": "registry",
  "model": "ministral-3:3b",
  "provider": "ollama-local",
  "usage": {
    "input_tokens": 12,
    "output_tokens": 3
  }
}

Exercise a tool through chat

The scaffold already wires one tool-capable agent: orchestrator lists calculator and web_search in its tools field, and both tool files exist under config/tools/. The router agent carries tools = [] on purpose.

To exercise the calculator directly, ask through the orchestrator path:

curl -s -X POST http://localhost:3000/v1/chat \
  -H "Authorization: Bearer $ARES_TENANT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"message": "Use the calculator to compute 21 * 2."}'

What happens inside:

  1. Execute::run receives the message.
  2. The model answers with a tool call instead of text.
  3. ARES resolves calculator through the Tools service and executes it.
  4. ARES sends the tool result back to the model, up to max_tool_iterations rounds.
  5. The final response arrives in the same JSON shape as before. Token usage now covers every round.

You can also give the router agent tools: set tools = ["calculator"] in ares.toml, or add tools[0]: calculator to config/agents/router.toon. Save the file; hot reload applies it within about one second.

If the response says the tool is unknown, check three things: the tool file exists under config/tools/, enabled: true is set, and the agent's tools list matches the tool name exactly.

Common first-run errors

SymptomCauseFix
Bind error such as Address already in use (os error 98)Another process holds host:port; often a previous server that never stoppedStop the old process, or change port in [server]
Chat returns an authentication errorMissing, malformed, or revoked tenant API keyConfirm the header reads Bearer ares_... with the raw key from creation time
Cordis Loader: boot failed in the log, then exit code 1A loader entry failed during boot; most often the database URL points at an unreachable serverCheck [database].url; for postgres:// URLs confirm the server accepts connections, then retry
Friendly banner naming the missing config file, exit code 1ares.toml absent from the working directoryRun ares-server init in that directory, or start from a directory that has one
Chat fails with a provider connection errorOllama (or your provider) is down, or base_url is wrongStart ollama serve, pull the model named in the provider block, and re-check base_url
Admin routes answer 401Server started without ADMIN_API_KEY setRestart the server process with ADMIN_API_KEY exported

Stop the server

Press Ctrl+C in the server terminal. The server shuts down cleanly.

For daemon-style operation, run under the built-in supervisor:

ares-server --supervise

The supervisor respawns the child after a hot-restart exit (code 51) and stops on clean exits. See Command Line Interface for details.

Architecture

ARES is a Cargo workspace. One root package, ares-server, holds the binary and the library facade. Ten crates under crates/ hold the capabilities. The kernel crate is published as ares-cordis; the workspace imports it as cordis.

Crate map

CrateRoleDepends on
ares-serverBinary and library facade. Boots the server, registers factories, binds HTTP.all capability crates, cordis
cordis (ares-cordis)Kernel. Typed Context, plugins, loader, fibers, events, intercepts.none
ares-typesShared types and errors (AppError, TenantContext).cordis
ares-vectorEmbedded vector store with HNSW. Standalone; no workspace dependencies.none
ares-ragRetrieval-augmented generation pipeline and embeddings.ares-types, cordis
ares-storePersistence. PostgreSQL through sqlx, Turso/libSQL behind a feature. Embeds migrations.ares-types, cordis; optional sqlx, libsql, ares-vector
ares-toolsTool trait, static and runtime tool registry, calculator.ares-types, cordis; optional ares-store, ares-mcp
ares-llmProvider clients, factory, pool, circuit breaker, Llm service.ares-types, ares-tools, cordis; optional ares-store
ares-mcpModel Context Protocol glue: client, auth, registry, server.ares-types, cordis; optional ares-store
ares-agentAgents: registry, router, orchestrator, Execute service, tenant scoping.ares-types, ares-llm, ares-tools, cordis; optional ares-store
ares-httpAxum routes, middleware, auth, overlay config, admin handlers.ares-agent, ares-tools, ares-llm, ares-store, ares-rag, ares-types, cordis; optional ares-mcp, ares-vector

The same crates with one-line data-flow notes each:

CrateData that flows through it
ares-serverBoot order and wiring only: it pushes factories into the context and hands the context to the HTTP listener; no domain data crosses it at request time.
cordisService handles and dependency epochs. Every get::<T>() in the server resolves here; versions bump when a store entry changes.
ares-typesStructs every crate shares: requests, responses, errors, TenantContext. Pure data; no I/O.
ares-vectorRaw f32 vectors plus HNSW graph nodes in process memory; queries enter as vectors and leave as neighbor ids and distances.
ares-ragText in, embeddings out: chunks go to an embedder, vectors to a store, and query results back through scoring and reranking.
ares-storeRows over sqlx: tenants, API keys, usage, agent versions, runs. Migrations flow outward from the embedded migrator at boot.
ares-toolsJSON tool calls in, JSON results out. The registry maps names to Tool impls; per-tenant allowlists gate resolution.
ares-llmChat completions to provider HTTP APIs; token counts and latency back. The circuit breaker wraps each provider client.
ares-mcpMCP protocol frames both ways: external servers become agent tools; the built-in MCP server exposes ARES agents outward.
ares-agentThe run pipeline: admitted request, tool-call rounds against Tools/Llm, assembled final response, run records toward storage.
ares-httpHTTP in, responses out. Middleware attaches tenant identity and usage; handlers open realms and call Execute.

Feature flags forward down the chain: postgres enables sqlx in store, agent, mcp, tools, http; openai, azure, and bedrock add providers in ares-llm.

Process lifecycle

Startup follows one ordered pass in run_server (src/main.rs:524). Each step gates the next:

  1. Load .env and start tracing (src/main.rs:528-532).
  2. Create the root Cordis Context and the ReflectService (src/main.rs:534-554). The service registers notifiers for the Tools and Llm types and fires an initial notify, so dependents reconcile from the first moment.
  3. Register loader factories through register_loader_factories (src/main.rs:328, called at src/main.rs:557-558). Explicit chains run without the inventory feature; inventory collection runs with it.
  4. Boot the entries program (src/main.rs:560-568): parse config/cordis-entries.toml, compose includes, instantiate entries in file order. The Overlay entry runs early and fills empty entry configs from ares.toml. A boot failure logs Cordis Loader: boot failed and exits with code 1.
  5. Guard configuration presence (src/main.rs:570-585). A missing config prints an ares-server init hint and exits with code 1.
  6. Start the entries watcher (src/main.rs:592-625). File events re-compose the program and apply diffs through the loader journal. When the watcher cannot start, a 30-second modified-time poll replaces it.
  7. Preload runtime providers (src/main.rs:630-632) and snapshot current agent definitions into the version history (src/main.rs:673-707).
  8. Build CORS and rate-limit layers (src/main.rs:895-944), bind the TCP listener (src/main.rs:949-950), and serve the Axum router with graceful shutdown on Ctrl+C and SIGTERM (src/main.rs:957-964).

Step 8 is the first step that touches the network for serving. A failure in any earlier step exits before any port opens.

The rate-limit layer

When rate_limit_per_second > 0, startup wraps the router in tower_governor (src/main.rs:900-913). The limiter is a GCRA bucket per client IP. It admits up to rate_limit_burst requests immediately and then admits one more request every \(1/\text{rate_limit_per_second}\) seconds. A background task prunes idle per-IP buckets every 60 seconds (src/main.rs:917-930). Responses carry x-ratelimit-* headers. Setting rate_limit_per_second = 0 removes the layer entirely and logs a warning.

Request path

request / job
  -> TenantRealms.open then intercept (HTTP/MCP/JWT) or isolate only (background)
  -> agent.admit (Execute, JWT chat, API-key middleware, MCP)
  -> Execute::run
  -> Tools / Llm / skills via EventsService waterfalls
  -> response
flowchart LR
    A[Request] --> B{Route class}
    B -->|protected| C[JWT auth]
    B -->|admin| D[Admin secret]
    B -->|/v1| E[API key]
    C --> F[TenantDb inject]
    E --> F
    F --> G[Usage tracking]
    D --> H[Handler]
    G --> H
    H --> I[Realm open + intercept]
    I --> J[agent.admit]
    J --> K[Execute.run]
    K --> L[Tools / Llm waterfalls]
    L --> M[Response]

Middleware order in create_router (crates/ares-http/src/api/routes.rs) matches this flow. On protected routes the outermost layer validates the JWT, the next layer injects TenantDb into extensions, and the innermost layer records token usage after the handler returns. Admin routes sit behind admin_middleware, which checks the X-Admin-Secret header. The /v1 routes use API key authentication, plus context injection when the eruka-context feature is enabled.

Inside a handler, request_tenant_ctx opens the tenant realm and then applies the TenantContext intercept. Realms are cached child contexts; the same tenant gets the same realm. Background jobs call tenant_scope instead. They isolate without an intercept. The agent.admit event is the shared quota gate. A deny maps to HTTP 429 or an MCP tool error. After the gate, Execute::run drives tools, model calls, and skills through EventsService waterfalls.

Cordis concepts as server concerns

  • Realms equal isolates. TenantRealms keeps one child context per tenant id, each backed by one fiber. Data-bearing services such as Tools resolve inside the realm. Tenant delete calls dispose, which undoes that realm's provides in last-in-first-out order.
  • Providers equal LLM clients. The Llm service wraps the provider registry. Its circuit breaker feeds Service::check. When the breaker opens, dependent fibers deactivate through guarded withdrawal.
  • Fibers own lifecycle. Each provided service has a fiber. Fiber::refresh recomputes the dependency epoch. Losing a dependency moves a fiber to Pending, not Failed. Hot swap replaces a service in place; retire withdraws it under guard while consumers exist.

Peer-dependency versions

Provider versions are plain u64 values with a packed layout (crates/cordis/src/context.rs:227-247, constant VERSION_MAJOR_SCALE = 100_000). The major lives in the high digits and the minimum compatible floor in the remainder:

$$\text{major}(v) = \left\lfloor \frac{v}{100,000} \right\rfloor, \qquad \text{floor}(v) = v \bmod 100,000$$

An inject that requires \(M \cdot 100,000 + f\) is satisfied by provider version \(p\) exactly when the provider exists and is available, \(\text{major}(p) = M\), and \(p \geq M \cdot 100,000 + f\). Any mismatch leaves the inject unsatisfied and the dependent fiber Inactive; it never binds across a major boundary. Legacy provide() installs version 0, which satisfies only unconstrained injects.

Storage model

PostgreSQL is the default backend, through sqlx with rustls TLS. Turso/libSQL compiles behind the turso feature for edge deployment. Vector data lives in ares-vector by default, or in Qdrant, pgvector, ChromaDB, Pinecone, or LanceDB behind features. SQL migrations are embedded in the ares-store crate at compile time. The migrator ships inside the published package, so cargo install ares-server needs no external files.

Tenant realms

TenantRealms (crates/ares-store/src/realms.rs) is the cache that turns one row per tenant into one live child context per tenant.

  • Open once, reuse forever. open(root, tenant_id) returns the cached child, or creates it on first call with a double-checked write (realms.rs:35-47). The child is root.extend().isolate_type(tools_tid, tenant_id).
  • Only data-bearing services isolate. The realm isolates on the Tools TypeId, so each tenant sees its own tool state and allowlists. Execute stays shared on purpose: it is a stateless engine, and its tenancy comes from the context handed to it (realms.rs:29-34). No TenantContext intercept lands inside the cached realm; handlers add it per request.
  • Disposal unwinds LIFO. dispose(tenant_id) removes the child from the map and disposes its fiber (realms.rs:49-57). Fiber disposal undoes that realm's provides in last-in-first-out order, so teardown runs in exact reverse registration order. Tenant delete (crates/ares-http/src/api/handlers/admin/tenants.rs) calls this before dropping rows.
sequenceDiagram
    participant R as Request
    participant TR as TenantRealms
    participant C as Cached child ctx
    R->>TR: open(root, "acme")
    TR->>C: first open: extend + isolate Tools
    C-->>R: realm (same Arc for every later request)
    R->>R: intercept TenantContext per request
    Note over TR: delete_tenant -> dispose("acme")
    TR->>C: fiber dispose, undos unwind LIFO

Failure philosophy

ARES fails closed at every trust boundary:

  • Missing or invalid JWT, API key, or admin secret rejects the request before any handler runs.
  • Quota denial through agent.admit blocks execution.
  • An availability predicate rejection marks a service Failed. Failed is terminal.
  • Guarded withdrawal deactivates dependents instead of binding to an incompatible provider.
  • Missing tenant database access during tenant resolution fails closed.
  • Loader composition stays fail-open for bad includes so one broken file cannot brick a running reload, but a failed boot pass still exits with code 1.
  • The emergency stop switch halts all agent execution.

Retrieval scoring

Embedding similarity uses cosine similarity over dense vectors (crates/ares-rag/src/embeddings.rs:128-149):

$$\cos(a, b) = \frac{a \cdot b}{\lVert a \rVert,\lVert b \rVert}$$

The implementation clamps the result to \([-1, 1]\) and returns \(0\) for mismatched lengths or zero magnitude instead of dividing by zero. Search results report distance, not similarity; distance_to_similarity maps cosine distance back with \(1 - d\), L2 with \(1/(1+d)\), and inner product with \(-d\) (crates/ares-rag/src/search.rs:296-302). The hybrid reranker then min-max normalizes retrieval scores across the candidate set and blends: \(\text{final} = (1 - w),\text{norm}_{\text{retrieval}} + w,\text{rerank}\) (crates/ares-rag/src/reranker.rs:362-371).

Design rules

Three rules explain most ARES behavior at failure boundaries. Each rule names one concrete place you can observe it.

Fail closed

When identity or availability cannot be proven, reject; never guess. JWT scope resolution is the sharpest example: if claims name a tenant but no matching TenantContext exists, or the tenant store is absent, the request fails instead of falling back to an unscoped context (crates/ares-http/src/auth/middleware.rs:31-32). The same rule puts auth middleware outside every handler and makes quota denial through agent.admit block execution.

Guarded withdrawal

A provider with live consumers does not vanish underneath them. Retiring a service through POST /admin/cordis/services/{name}/retire answers 409 {"retired": false, "reason": "guarded", "consumers": N} while N active consumer fibers still rely on it; the removal happens only after the last consumer drops its reliance (crates/ares-http/src/api/handlers/admin/cordis.rs:71-80). Dependent fibers deactivate rather than bind to an incompatible replacement.

Eager reconciliation

State changes propagate immediately, not at next use. Two mechanisms carry this rule:

  • Boot wires ReflectService notifiers for Tools and Llm and fires an initial notify, so dependent fibers recompute their dependency epoch right away (src/main.rs:534-554).
  • Editing config/cordis-entries.toml triggers parse, compose, diff-apply, and classify in one pass (cordis::reload_entries_from_disk, called from src/main.rs:488-493); consumers see the new provider set without a restart.

Together the three rules produce the observed lifecycle semantics: Pending means waiting, Failed means rejected permanently, and a guarded retire never leaves a consumer resolving nothing.

Command Line Interface

The ares-server binary is the single entry point. Run it without a subcommand to start the server, or use a subcommand to manage a project.

$ ares-server --help
A production-grade agentic chatbot server with multi-provider LLM support,
tool calling, RAG (Retrieval Augmented Generation), and MCP integration.

Run without arguments to start the server, or use 'init' to scaffold a new project.

Usage: ares-server [OPTIONS] [COMMAND]

Commands:
  init    Initialize a new A.R.E.S project with configuration files
  config  Show configuration information
  agent   Manage agents
  rag     Ingest and search RAG collections through the ARES API
  help    Print this message or the help of the given subcommand(s)

Global options

These options work on every command:

OptionEffect
-c, --config <CONFIG>Path to the configuration file. Default: ares.toml
-v, --verboseEnable verbose output
--no-colorDisable colored output
--mcpStart in MCP server mode over stdio transport
--superviseRun under the built-in supervisor
-h, --helpPrint help. Use -h for a short form and --help for details
-V, --versionPrint the version

Global flag placement

Every option carries global = true (src/cli/mod.rs). Place a global flag before or after the subcommand. These two lines parse identically:

ares-server --no-color config --validate
ares-server config --validate --no-color

-v, --verbose changes the log filter of a server run from info to debug,ares=trace (run_server, src/main.rs). Subcommands print their diagnostics to standard error regardless of this flag.

--mcp needs the mcp build feature. Without it, the binary prints a rebuild hint and exits 1.

Supervisor semantics

--supervise runs the real server in a child copy of the same executable. The child signals the parent through its exit code:

Exit codeMeaningParent action
51Hot-restart requestRespawn a fresh child
52Clean shutdownStop the loop
53Boot failureStop and mirror the non-zero code to the service manager

The constants live in src/supervisor.rs. The child carries the CORDIS_SUPERVISED environment variable. The daemon holds the write end of the child's standard input. Dropping it closes the pipe, so the child sees end-of-file and tears down gracefully.

Example: run a daemon that survives hot restarts:

ares-server --supervise

Pair it with a service manager such as systemd. Boot failures still surface as non-zero exits.

Restart safeguards

Four safeguards keep the loop responsive (src/supervisor.rs):

  • Rapid-restart guard. Five exits inside any 30-second window stop the loop. This bounds a crash loop that a misbehaving child cannot out-wait.
  • Health ladder reset. A child that ran for at least 10 minutes before exiting counts as healthy. The next crash sequence starts its backoff from zero, not from stale strikes.
  • Exponential backoff. A child that exited within 10 seconds never proved it could serve. The daemon delays the respawn: 100 ms first, then double each consecutive unhealthy run, capped at 5 s.
  • Stop grace. A stopped worker gets 10 seconds to exit after its standard input closes. A worker that ignores the request is force-killed.

Exit codes

Use these codes in scripts and service units:

CodeProducerMeaning
0every subcommandSuccess. config --validate reports a valid file
1initTarget files already exist (--force overwrites), or scaffolding failed
1server bootMissing config file, missing Overlay entry, failed entries program, or --mcp without the feature
51supervised childHot-restart request; the daemon respawns
52supervised childClean shutdown; the daemon stops
53supervised childBoot failure; the daemon stops and mirrors 53

Two details matter for wrappers:

  • Without --supervise, a failing boot returns 1. With --supervise, the parent mirrors the real child code (exit(last_code & 0xff)), so a boot failure surfaces as 53.

init

Scaffold a new project. Creates ares.toml, a .env.example template, the data/ directory, and the config/ directory tree with example TOON files.

$ ares-server init [OPTIONS] [PATH]
OptionEffect
[PATH]Directory to initialize. Default: current directory
-f, --forceOverwrite existing files without prompting
-m, --minimalCreate fewer agents and tools
--no-examplesSkip example TOON files in config/
--provider <PROVIDER>LLM provider to configure: ollama, openai, or both. Default: ollama
--host <HOST>Server host address. Default: 127.0.0.1
--port <PORT>Server port. Default: 3000

Example session:

$ ares-server init --minimal
   Agentic Runtime Extensible Server v0.10.0

  Initializing A.R.E.S Project

  Creating directories
  ✓ directory data
  ✓ directory config/agents
  ✓ directory config/models

  Creating configuration files
  ✓ config ares.toml
  ✓ env .env.example

  A.R.E.S project initialized successfully!

config

Show configuration information.

$ ares-server config [OPTIONS]
OptionEffect
-f, --fullShow the full configuration instead of the summary
--validateValidate the configuration file and report problems

Example summary:

$ ares-server config
  Configuration Summary

    Config file: ares.toml
    Server: 127.0.0.1:3000
    Log level: info

  Providers
    • ollama-local

  Agents
    • orchestrator
    • router

Validation fails when a referenced environment variable is missing, or when an agent references an unknown provider. A zero exit code means the file is valid:

$ ares-server config --validate
  ✓ Configuration is valid!

agent

Manage configured agents.

$ ares-server agent <COMMAND>

Subcommands:

  • list — list all configured agents.
  • show <NAME> — show details for one agent.

Example listing:

$ ares-server agent list
  Configured Agents

    Name            Model                        Tools
    ────────────────────────────────────────────────
    router          meta/llama-3.3-70b-instruct  -
    orchestrator    meta/llama-3.3-70b-instruct  calculator, web_search

Show one agent by name:

ares-server agent show orchestrator

Both commands read the same sources as the server: static agents from ares.toml plus TOON files in the configured directories.

rag

Ingest documents into RAG collections and search them. These commands are API clients: they call a running ARES server over HTTP, so start the server first.

Authenticate with --user and --password, or skip login with --token. Both forms accept --host for a remote server. Default host: http://localhost:3000.

rag ingest-dir

Recursively ingest local text documents into a collection.

$ ares-server rag ingest-dir [OPTIONS] --collection <COLLECTION> --docs-path <DOCS_PATH>

All options (--help output):

OptionEffect
--collection <NAME>Collection to ingest into. Required
--docs-path <DIR>Directory with the documents. Required
--chunking-strategy <KIND>word (default), semantic, or character
--tag <TAG>Attach a tag. Repeat for multiple tags
--dry-runList the files that would be ingested. Send no requests
--host <URL>ARES server base URL. Default: http://localhost:3000
--user / --passwordLogin credentials
--token <TOKEN>Bearer token; skips login

Preview an ingest before running it:

ares-server rag ingest-dir \
  --collection docs \
  --docs-path ./handbook \
  --tag handbook \
  --dry-run

Remove --dry-run to perform the ingest.

Search a collection.

$ ares-server rag search [OPTIONS] --collection <COLLECTION> --query <QUERY>

All options (--help output):

OptionEffect
--collection <NAME>Collection to search. Required
--query <TEXT>Search query. Required
--top-k <N>Maximum number of results. Default: 10
--strategy <KIND>semantic (default), bm25, fuzzy, or hybrid
--host <URL>ARES server base URL. Default: http://localhost:3000
--user / --passwordLogin credentials
--token <TOKEN>Bearer token; skips login

Example:

ares-server rag search \
  --collection docs \
  --query "how do I rotate API keys" \
  --top-k 5 \
  --strategy hybrid

Scripting and composition

Every command reports its result through the exit code. Chain commands with && so a later step runs only after an earlier step succeeds.

Validate, then start under the supervisor:

ares-server config --validate && ares-server --supervise

Preview an ingest, then run it for real:

ares-server rag ingest-dir --collection docs --docs-path ./handbook --dry-run \
  && ares-server rag ingest-dir --collection docs --docs-path ./handbook \
       --tag handbook --user me@example.com --password "$PASS"

Scripting guidance:

  • Check $? after each call. 0 means success; see the exit-code table above for failure codes.
  • --verbose raises server log verbosity to debug,ares=trace. It does not change subcommand output.
  • Progress lines go to standard output; failures go to standard error. Redirect them separately: 2>err.log.
  • A failed ingest still processes every document first. The summary line on stdout reads summary<TAB>documents=N succeeded=S failed=F chunks=C; parse it to decide whether to retry.
  • A dry run prints one tab-separated line per file: <path><TAB><title><TAB><N bytes>, then dry_run=true documents=<count>. Both formats are stable for parsing.

A guarded ingest in shell:

if ares-server rag ingest-dir --collection docs --docs-path ./handbook \
     --token "$ARES_TOKEN" > out.log 2> err.log; then
  grep '^summary' out.log
else
  cat err.log
  exit 1
fi

Where to go next

  • Scaffold and run your first project: First Server.
  • Call the server over HTTP instead of the CLI: HTTP API.

HTTP API

ARES exposes one Axum HTTP service. The main router mounts every application route under the /api prefix (crates/ares-http/src/lib.rs). The server also answers GET /health outside /api. This chapter documents the routes as implemented in crates/ares-http/src/api/routes.rs.

Base URL

The server binds to server.host and server.port from its configuration file. Defaults are 127.0.0.1 and port 3000 (crates/ares-http/src/config.rs). All paths below are relative to http://localhost:3000/api.

curl -s http://localhost:3000/health

The /health route returns the plain text OK. The server binary adds GET /health/detailed and GET /config/info next to it.

Authentication

Three schemes exist. Pick the scheme that matches the route group.

JWT bearer tokens (user routes)

Register or log in to get a token pair:

curl -s -X POST http://localhost:3000/api/auth/register \
  -H "Content-Type: application/json" \
  -d '{"email": "user@example.com", "password": "secret", "name": "User"}'
{
  "access_token": "<jwt>",
  "refresh_token": "<jwt>",
  "expires_in": 3600
}

Send the access token as a bearer token:

Authorization: Bearer <access_token>

Token anatomy

Claims lives in crates/ares-types/src/types/mod.rs; the signing logic lives in crates/ares-http/src/auth/jwt.rs. One decoded access token:

{
  "sub": "9b2f3c58-4b1e-4a7d-9c11-0f5a6b8c2d10",
  "email": "user@example.com",
  "exp": 1756221600,
  "iat": 1756220700
}
ClaimPresenceMeaning
subalwaysUser id. Refresh tokens must match the session row by this field.
emailalwaysAccount email.
exp, iatalwaysExpiry and issue time as Unix seconds. Validation allows a 60-second clock-skew leeway.
jtirefresh tokens onlyRandom UUID that identifies one refresh session. Access tokens omit it.
tenant_idtenant-scoped tokens onlyTenant that issued or owns the session.

Defaults from AuthConfig (crates/ares-http/src/config.rs): access tokens live 900 seconds (15 minutes), refresh tokens 604800 seconds (7 days). expires_in in the response echoes the configured access expiry, so read it instead of hard-coding 900.

Register, login, refresh, logout

Validation runs before any database call:

RouteFailureResponse
POST /auth/registerEmpty email or password under 8 characters400 {"error":"Email required and password must be at least 8 characters"}
POST /auth/registerEmail already registered400 {"error":"User already exists"}
POST /auth/loginUnknown email or wrong password401 {"error":"Invalid credentials"}

Passwords hash with Argon2id; refresh tokens are stored only as SHA-256 hashes.

The refresh flow rotates sessions — each refresh token works exactly once (refresh_token, crates/ares-http/src/api/handlers/auth.rs):

  1. Verify the refresh token's HS256 signature and expiry.
  2. Hash it and look up the session row. No row answers 401 {"error":"Refresh token has been revoked or expired"}.
  3. Compare the session's user id with the sub claim. A mismatch answers 401 {"error":"Token mismatch"}.
  4. Delete the old session row.
  5. Issue and return a fresh pair; the new refresh token lands in its own session.
curl -s -X POST http://localhost:3000/api/auth/refresh \
  -H "Content-Type: application/json" \
  -d '{"refresh_token": "<refresh_token>"}'

The response is a full TokenResponse. Reuse of an already-rotated refresh token fails at step 2, which is why clients must persist the newest pair after every call.

POST /auth/logout takes {"refresh_token": "..."}, deletes the matching session by hash, and returns {"message":"Logged out successfully"} even when the session is already gone.

Admin secret (admin routes)

Admin routes check the X-Admin-Secret header against the ADMIN_API_KEY environment variable. As an alternative, admin routes accept a JWT with an admin role claim.

curl -s http://localhost:3000/api/admin/stats \
  -H "X-Admin-Secret: $ADMIN_API_KEY"

A rejected request answers 401 with:

{"error":"Admin access requires X-Admin-Secret header or JWT with admin role"}

Tenant API keys (/v1 routes)

Routes under /v1 authenticate machine clients with tenant API keys. Keys start with ares_ and travel in the same bearer header:

Authorization: Bearer ares_<key>

Scheme matrix

PropertyJWT bearerAdmin secretTenant API key
CredentialAccess token from login/registerStatic value of ADMIN_API_KEY env varKey created via POST /v1/api-keys, prefix ares_
HeaderAuthorization: Bearer <access_token> or ?token=X-Admin-Secret: <value>; a JWT with an admin role claim also worksAuthorization: Bearer ares_<key>
Route group/chat, /research, /user/agents, /conversations, /workflows, .../admin/*/v1/*
IdentityUser id in sub claimNone (operator)Tenant resolved from the key row
MeteringNo quota gate at the middlewareNot meteredMonthly and daily quota checks run before the handler
RevocationRefresh rotation plus logout deletes the sessionRotate the environment variable and restartRevoke with DELETE /v1/api-keys/{id}

The middleware rejects malformed /v1 credentials before touching the database (crates/ares-http/src/middleware/api_key_auth.rs). All format failures answer 401 {"error": "<message>"}:

ConditionMessage
No Authorization headerMissing Authorization header
Header not valid ASCIIInvalid Authorization header
Value does not start with Bearer (case-sensitive)Invalid Authorization format. Expected: Bearer ares_...
Key does not start with ares_Invalid API key format. Must start with ares_
Key is well formed but unknownInvalid API key

Quota breaches answer 429: Monthly request quota exceeded or Daily rate limit exceeded. The monthly check wins when both are exhausted. Tier limits come from the tenant's quota row; the unit tests pin examples — a Free-tier tenant blocks at 1,000 requests per month or 50 per day, a Dev-tier tenant at 2,000 per day, Enterprise tiers allow large volumes. Infrastructure faults answer 500 with messages such as Tenant database not configured, Failed to verify API key, Failed to check usage, or Failed to check rate limit.

Response Envelope

Successful handlers return the documented payload directly. Errors return one consistent shape with two fields, error and code (crates/ares-http/src/error.rs):

{
  "error": "agent my-agent not found",
  "code": "NOT_FOUND"
}

Error catalog

Handlers return HttpError, which wraps AppError (crates/ares-http/src/error.rs). The status comes from AppError::status_code() and the code from AppError::code(), both in crates/ares-types/src/types/mod.rs. The mapping is fixed:

AppError variantHTTP statuscodeExample message prefix
Database500DATABASE_ERRORDatabase error:
LLM500LLM_ERRORLLM error:
Auth401AUTHENTICATION_FAILEDAuthentication error:
NotFound404NOT_FOUNDNot found:
InvalidInput400INVALID_INPUTInvalid input:
Configuration500CONFIGURATION_ERRORConfiguration error:
External502EXTERNAL_SERVICE_ERRORExternal service error:
Internal500INTERNAL_ERRORInternal error:
Unavailable503INTERNAL_ERRORService unavailable:
RateLimited429INTERNAL_ERRORRate limited:
FeatureDisabled400INTERNAL_ERRORFeature disabled:

Three variants carry status codes that do not match their code class: Unavailable answers 503 but reports INTERNAL_ERROR, and RateLimited answers 429 while FeatureDisabled answers 400, both also reporting INTERNAL_ERROR. Match on the status plus the message prefix, not on code alone.

Chat

POST /chat

Runs one agent turn. Requires a JWT bearer token.

Request fields (ChatRequest, crates/ares-types/src/types/mod.rs):

FieldTypeNotes
messagestringRequired. The user message.
agent_typestringOptional. Defaults to the router agent.
context_idstringOptional. Continues a conversation.
workspace_idstringOptional. Eruka workspace scope.
modelstringOptional per-request model override.
curl -s -X POST http://localhost:3000/api/chat \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"message": "Summarize my notes", "agent_type": "researcher"}'

Response (ChatResponse):

{
  "response": "Here is the summary...",
  "agent": "researcher",
  "context_id": "8f14e45f-ea9b-4d2a-9c3b-1f6a2b7c9d01",
  "sources": [
    {"title": "Meeting notes", "url": null, "relevance_score": 0.87}
  ]
}

POST /chat/stream and GET /chat/stream

Streams Server-Sent Events with the same request body. The GET variant reads the fields from query parameters for EventSource clients. Each event is a StreamEvent object with fields event, content, agent, context_id, and error; absent optional fields are omitted:

data: {"event":"start","agent":"researcher","context_id":"8f14e45f-..."}

data: {"event":"token","content":"Here "}

data: {"event":"done","agent":"researcher","context_id":"8f14e45f-..."}

Event anatomy

StreamEvent and its four constructors live in crates/ares-http/src/api/handlers/chat.rs. Absent optional fields are omitted from the JSON, never sent as null:

EventFields setProducer behavior
startagent ("<name> (system)"), context_idSent once before any model output, after agent resolution succeeds.
tokencontentOne per streamed token chunk. No agent or context_id.
doneagent ("{AgentType:?} ({source})", for example "Sales (system)"), context_idFinal event of a successful run.
errorerror; context_id when knownTerminal. Failures before a context exists (admission denial, missing Llm service) omit context_id entirely.

An admission failure yields an error event with no other fields:

data: {"event":"error","error":"monthly quota exceeded"}

A mid-stream failure carries the conversation scope:

data: {"event":"start","agent":"product (system)","context_id":"8f14e45f-..."}

data: {"event":"token","content":"Here "}

data: {"event":"error","context_id":"8f14e45f-...","error":"Stream error: provider closed connection"}

The endpoint attaches an SSE keep-alive comment every 15 seconds (Sse::keep_alive in chat_stream_response). Idle connections therefore never time out silently; clients should ignore comment frames.

The GET variant takes the request fields as query parameters (ChatStreamQuery): message (required), plus optional agent_type, context_id, and workspace_id. Authenticate it with Authorization: Bearer or the ?token= fallback:

curl -N -s "http://localhost:3000/api/chat/stream?message=Summarize%20my%20notes&agent_type=researcher&token=$ACCESS_TOKEN"

POST /research

Runs deep research. Body is {"query": "...", "depth": 3, "max_iterations": 10}; both limits are optional.

GET /memory

Returns stored facts and preferences for the authenticated user:

{
  "user_id": "42",
  "preferences": [
    {"category": "communication", "key": "style", "value": "concise", "confidence": 0.9}
  ],
  "facts": [
    {
      "id": "f-1", "user_id": "42", "category": "work",
      "fact_key": "timezone", "fact_value": "UTC+1", "confidence": 0.95
    }
  ]
}

An empty memory returns no body content.

Agents

User agents (JWT)

MethodPathPurpose
GET/agentsPublic list of shared agents.
GET/user/agentsList the caller's agents.
POST/user/agentsCreate an agent.
GET/user/agents/{name}Read one agent.
PUT/user/agents/{name}Update one agent.
DELETE/user/agents/{name}Delete one agent.
POST/user/agents/importImport an agent from TOON format.
GET/user/agents/{name}/exportExport an agent to TOON format.

Create body (CreateUserAgentReq):

{
  "name": "my-agent",
  "display_name": "My Agent",
  "description": "Answers billing questions",
  "model": "gpt-4o-mini",
  "system_prompt": "You are a billing assistant.",
  "tools": ["calculator"],
  "max_tool_iterations": 10,
  "parallel_tools": false,
  "is_public": false,
  "extra": {}
}

Responses carry id, usage_count, average_rating, created_at, and updated_at alongside the input fields.

Loop-mode agents (JWT)

POST /loops/start starts a loop run, GET /loops lists loops, DELETE /loops/{id} stops one.

Conversations (JWT)

GET /conversations lists conversations. GET, PUT, and DELETE on /conversations/{id} read, rename, and delete one.

Workflows, Skills, Tools

Workflows require a JWT:

  • GET /workflows lists available workflows.
  • POST /workflows/{workflow_name} executes one.

With the skills feature enabled:

  • GET /skills lists skills.
  • GET /skills/{name} reads one skill.

Admin surfaces manage runtime tools and skills with the X-Admin-Secret header:

MethodPathPurpose
GET / POST/admin/runtime-toolsList or create tools.
GET/admin/runtime-tools/capabilitiesList tool capability descriptors.
GET / PUT / DELETE/admin/runtime-tools/{id}Manage one tool.
POST/admin/runtime-tools/{id}/testExecute a tool with sample input.
GET/admin/runtime-tools/{id}/versionsList versions.
POST/admin/runtime-tools/{id}/rollback/{version}Roll back.
GET / POST/admin/skillsList or create skills.
POST/admin/skills/runRun a skill.
GET / PUT / DELETE/admin/skills/{id}Manage one skill.

Tool test example. The body field input_args holds the JSON arguments passed to the tool's execute method:

curl -s -X POST http://localhost:3000/api/admin/runtime-tools/7/test \
  -H "X-Admin-Secret: $ADMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"input_args": {"x": 2, "y": 3}}'
{"ok": true, "output": {"sum": 5}, "error": null, "latency_ms": 4}

Skill run example. tenant_id must name an existing tenant:

curl -s -X POST http://localhost:3000/api/admin/skills/run \
  -H "X-Admin-Secret: $ADMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"skill_id": "summarize", "tenant_id": "tenant-a", "input": {"text": "..."}}'

RAG

These routes need the local-embeddings and ares-vector features at build time. They require a JWT. Collections are scoped per user; the server prefixes your collection name with your user id internally.

POST /rag/ingest

Body fields come from RagIngestRequest: collection, content, plus optional title, source, tags, and chunking_strategy.

curl -s -X POST http://localhost:3000/api/rag/ingest \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "collection": "notes",
    "content": "Quarterly review text...",
    "title": "Q3 review",
    "tags": ["finance"]
  }'
{
  "chunks_created": 4,
  "document_ids": ["d1", "d2", "d3", "d4"],
  "collection": "notes"
}

POST /rag/search

Strategy is one of semantic, bm25, fuzzy, or hybrid. Defaults: limit 10, threshold 0.0, rerank false.

curl -s -X POST http://localhost:3000/api/rag/search \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"collection": "notes", "query": "budget owners", "limit": 5, "strategy": "hybrid"}'
{
  "results": [
    {
      "id": "d1",
      "content": "Budget owners meet on Mondays.",
      "score": 0.91,
      "metadata": {}
    }
  ],
  "total": 5,
  "strategy": "hybrid",
  "reranked": false,
  "duration_ms": 23
}

Collection management

  • GET /rag/collections lists collections as CollectionInfo objects.
  • DELETE /rag/collection deletes one. Body: {"collection": "notes"}. Response: {"success": true, "collection": "notes", "documents_deleted": 12}.

MCP

Model Context Protocol surface is read-only today:

curl -s http://localhost:3000/api/mcp/runtime_tool_capabilities \
  -H "X-Admin-Secret: $ADMIN_API_KEY"

The route lives behind the admin middleware because it merges into the admin router set (build_routes).

/v1 External API

Machine clients use tenant API keys. Metered routes record usage per call:

MethodPathPurpose
POST/v1/chatChat completion.
POST/v1/researchDeep research run.
POST/v1/agents/{name}/runRun a named agent.
POST/v1/agents/{name}/sandbox-runSandbox execution.
GET/v1/agentsList agents visible to the tenant.
GET/v1/agents/{name}Read one agent.
GET/v1/agents/{name}/runsList run history.
GET/v1/agents/{name}/logsList run logs.
GET/v1/usageTenant usage summary.
GET / POST/v1/api-keysList or create API keys.
DELETE/v1/api-keys/{id}Revoke a key.
POST/v1/search/semanticSemantic search (feature-gated).
DELETE/v1/tenant/dataDelete all tenant data.

Example:

curl -s -X POST http://localhost:3000/api/v1/chat \
  -H "Authorization: Bearer ares_$API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"message": "Hello"}'

Quota breaches answer with a quota-exceeded error body.

Pagination and Filtering

List endpoints use two different parameter conventions.

/v1 page-based pagination

GET /v1/agents, GET /v1/agents/{name}/runs, and GET /v1/agents/{name}/logs take page and per_page query parameters and return a Paginated<T> envelope (crates/ares-http/src/api/handlers/v1/shared.rs):

{
  "items": [],
  "total": 0,
  "page": 1,
  "per_page": 20,
  "total_pages": 0
}
ParameterNormalizationNotes
pageDefaults to 1; values under 1 clamp to 1
per_pageDefaults to 20 for agents, 25 for runs; caps at 100Logs default to 50

Example:

curl -s "http://localhost:3000/api/v1/agents?page=2&per_page=50" \
  -H "Authorization: Bearer ares_$API_KEY"

Admin limit/offset pagination

Admin list endpoints in crates/ares-http/src/api/handlers/admin/audit.rs and siblings take limit and offset. The handler clamps the values before querying:

Route groupParametersClamping
GET /admin/alertslimit, severity, resolvedDefault limit 50, cap 200; filter by severity string and resolved flag
GET /admin/audit-loglimit, offsetDefault limit 50, cap 200
GET .../tenants/{tenant_id}/usage/dailydaysDefault 30, cap 90
Tenant agent runs (.../agents/{name}/runs)limit, offsetDefault 50, cap 200
Feedback summary (.../{agent_name}/feedback/summary)daysDefault 30, clamped to 1..366
Missed runs (GET .../schedules/{id}/missed-runs)limitDefault 10, clamped to 1..100
Run history costs (POST /admin/run-history/costs)limit, offset in bodyLimit clamped to 1..10000

Tenant-scoped list routes such as /admin/triggers, /admin/pipelines, and /admin/schedules require ?tenant_id=. An empty value answers 400 {"error":"tenant_id query param is required"}.

Webhooks, OAuth, Events

Public routes without authentication:

  • POST /webhooks/{trigger_id} — webhook receiver for triggers.
  • GET /oauth/authorize and GET /oauth/callback — connector OAuth flow.
  • POST /events/document-upload and POST /events/field-change — event ingestion.

Admin Surfaces

All admin routes take the X-Admin-Secret header. Route groups in routes.rs:

GroupExample routes
TenantsPOST/GET /admin/tenants, GET /admin/tenants/{tenant_id}, POST/GET /admin/tenants/{tenant_id}/api-keys, GET .../usage, PUT .../quota, GET .../usage/daily
ProvisioningPOST /admin/provision-client
Tenant agentsGET/POST /admin/tenants/{tenant_id}/agents, PUT/DELETE .../agents/{agent_name}, .../versions, .../rollback/{version}, .../test, .../runs, .../stats, .../feedback/*
Cross-tenant agentsGET/POST /admin/agents, GET/PUT/DELETE /admin/agents/{tenant_id}/{agent_name}, .../versions, .../rollback/{version}, GET/POST /admin/agents/emergency-stop
Templates and modelsGET/POST /admin/agent-templates, DELETE /admin/agent-templates/{id}, GET /admin/models
Alerts and auditGET /admin/alerts, POST /admin/alerts/{alert_id}/resolve, GET /admin/audit-log
DeploymentPOST /admin/deploy, GET /admin/deploy/{deploy_id}, GET /admin/deploys, GET /admin/services, GET /admin/services/{service_name}/logs
Model tiersGET/POST /admin/tenants/{tenant_id}/model-tiers, GET/PUT/DELETE .../{tier_name}
AllowlistsGET/POST .../allowed-tools, .../allowed-models, .../allowed-rag-sources, each with DELETE .../{name}
Triggers and pipelinesGET/POST .../triggers, PUT/DELETE .../triggers/{id}, same shape for pipelines and platform-wide /admin/triggers, /admin/pipelines
Fleet providersGET /admin/fleet-providers, GET .../capabilities, PUT/DELETE .../{provider_name}, POST .../verify
SchedulesGET/POST /admin/schedules, PUT/DELETE /admin/schedules/{id}, tenant variants and GET .../missed-runs
ConnectorsGET/POST /admin/connectors, PUT/DELETE /admin/connectors/{id}, tenant connectors and oauth-creds
BillingGET .../billing/summary, GET .../billing/line-items, GET /admin/billing/model-rates, GET /admin/billing/unit-rates
BudgetsGET/PUT/DELETE /admin/run-history/budgets/{tenant_id}, GET/PUT /admin/token-budgets/{tenant_id}, GET .../status, POST .../reset, GET .../usage, GET /admin/run-history/alerts, POST /admin/run-history/alerts/{id}/acknowledge
Run historyGET/POST /admin/run-history/llm-calls, GET .../llm-calls/{id}, same shape for tool-calls, GET /admin/run-history/costs/{run_id}, GET .../costs, GET/POST .../health-metrics, GET .../model-metrics, GET /admin/runs/live (active-run stream)
Runtime providersGET/POST /admin/runtime_providers, GET/DELETE /admin/runtime_providers/{name}
Platform statsGET /admin/stats

Cordis Service Lifecycle

These routes manage the plugin runtime. Unknown loader state answers 503.

Retire removes a service; provide re-registers a known direct service:

curl -s -X POST http://localhost:3000/api/admin/cordis/services/events_service/provide \
  -H "X-Admin-Secret: $ADMIN_API_KEY"
{"provided": true, "service": "events_service", "type": "cordis::events::EventsService"}

POST /admin/cordis/services/{name}/retire answers 200 {"retired": true, ...} on removal and 200 {"retired": false, ...} when the service was already absent. Guarded withdrawal refuses the removal while active consumer fibers still rely on the provider; it answers 409 {"retired": false, "reason": "guarded", "consumers": <N>}. Names that are not direct Cordis services answer 409 as well — wrapper types are not supported today (crates/ares-http/src/api/handlers/admin/cordis.rs, retire_cordis_service).

Two read-only routes help interpret those outcomes:

  • GET /admin/cordis/services lists every tracked fiber with fiber_id, state (the debug form of FiberState: Active, Inactive, Loading, Failed, Reloading, Unloading), error when the fiber rests in a terminal state with a message, disposed, and pending_undo_count.
  • GET /admin/cordis/undo lists the labeled undo closures still pending per fiber, in registration order. Only labeled undos surface; anonymous ones count toward pending_undo_count only.

Both answer 503 {"error":"RegistryService is not provided on this context"} on library deployments without a registry.

POST /admin/cordis/services/{name}/replace

Rolling drain-and-shift replacement of a journaled provider. The body must be {"config": <value>} carrying the new configuration. Success answers:

{"replaced": true, "plugin": "calc", "fiber_id": 17}

A refusal (unknown plugin label, untracked provider, failing trial) leaves the old provider serving untouched and answers 409 {"replaced": false, "service": "calc", "reason": "..."}. A missing config field answers 400.

Cordis Entries

Entries live in a TOML program file. Routes:

MethodPathPurpose
GET/admin/cordis/entriesList the entry tree.
PUT/admin/cordis/entriesUpsert an entry.
PATCH/admin/cordis/entries/{id}Partial update.
DELETE/admin/cordis/entries/{id}Remove an entry.
POST/admin/cordis/entries/{id}/toggleEnable or disable.
POST/admin/cordis/entries/reloadReload from disk.
POST/admin/cordis/entries/{id}/moveRelocate an entry.
GET/admin/cordis/eventsPer-event dispatch counters.
GET/admin/cordis/undoPending undo labels per fiber.

PATCH /admin/cordis/entries/

Applies only the present fields: config, disabled, isolate, intercept. An empty body is a validated no-op that still persists and re-applies the tree. Present parent or position fields move the entry first. Invalid moves answer 409; unknown ids answer 404.

When the new configuration fails the factory pre-flight, the response carries a structured issues array next to the legacy error string. Each issue has a message and a path. The error string is the loader's marker plus the rendered error ("config pre-flight failed: {error}", crates/cordis/src/loader.rs; a validation issue renders as - <message> (at <path>)):

{
  "applied": [],
  "patched": false,
  "reloaded": false,
  "error": "config pre-flight failed: invalid config: - missing url (at calc.url)",
  "issues": [{"message": "missing url", "path": ["calc", "url"]}]
}

Move-then-update in one call

A body with a present parent or position field relocates the entry first, then applies the remaining fields. One request can rename a subtree and reconfigure its root:

curl -s -X PATCH http://localhost:3000/api/admin/cordis/entries/calc \
  -H "X-Admin-Secret: $ADMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"parent": "tools-group", "config": {"precision": 2}}'

The response carries the post-patch entry plus renamed old-to-new id pairs from the move phase:

{
  "applied": [],
  "patched": true,
  "renamed": [["calc", "tools-group:calc"]],
  "entry": {"id": "tools-group:calc"}
}

The live fiber keeps its identity across the structural move. The journal re-keys the record to the new id while preserving the fiber id, so consumers never observe a restart. The config update then lands under the new id. This sequence comes from the test patch_endpoint_moves_entry in crates/ares-http/src/api/handlers/admin/cordis.rs.

A position-only body reorders within the current parent; an explicit "parent": null moves to the tree root.

Failed pre-flight: the issues array

When the patched config fails the factory trial pre-flight, the loader stashes machine-readable issues for the entry and the handler attaches them to the failure body shown above. The status is 422 (patch_endpoint_returns_structured_issues_on_bad_config).

The failed trial leaves nothing behind: a follow-up well-formed patch succeeds with 200 and no issues field, instead of tripping stale issues from the earlier attempt.

POST /admin/cordis/entries/{id}/move

Relocates the entry and its whole {id}:* descendant namespace under a new parent. Body fields:

  • parent — string entry id, or null to move to the tree root.
  • position — non-negative integer child index. Omit it to append after the target's existing children.
curl -s -X POST http://localhost:3000/api/admin/cordis/entries/calc/move \
  -H "X-Admin-Secret: $ADMIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"parent": "tools-group", "position": 0}'

Pure structural moves preserve fiber identity; running fibers never restart. Renamed descendant ids appear as old-to-new pairs:

{
  "moved": true,
  "noop": false,
  "renamed": [["calc", "tools-group:calc"]],
  "applied": []
}

Parent semantics in detail (move_cordis_entry and EntryTree::move_entry):

  • "parent": "<id>" moves under that entry. Every descendant id prefixed {moved-id}: renames mechanically; ids without the prefix stay untouched.
  • "parent": null moves the entry to the tree root and strips any parent prefix from it and its descendants.
  • Omitting the parent field behaves like null: the entry moves to the tree root. To reorder within the current parent without relocating, use PATCH with a position only.
  • "position" must be a non-negative integer. A wrong type answers 400 {"error":"\"position\" must be a non-negative integer"}. A non-string, non-null parent answers 400 {"error":"\"parent\" must be a string or null"}.

Pure structural moves never restart fibers: the loader detects that plugins, configs, disabled flags, and isolates are identical on both sides, takes the noop path, re-keys journal records while keeping fiber ids, and reports "noop": true when nothing but placement changed.

ARES as a Library

ARES runs as an embedded library. The ares-server package publishes a library facade next to the server binary. You build one Cordis context, put services on it, and call them. No HTTP port opens unless you start the router yourself.

This chapter shows every step with code that mirrors real tests in the repository. Each section names the source it adapts.

Add the dependency

[dependencies]
ares-server = "0.10"
ares-cordis = "0.10"   # kernel types beyond the facade re-exports
tokio = { version = "1", features = ["rt-multi-thread", "macros"] }
serde_json = "1"

The default features are the server defaults: postgres, openai, ares-vector, mcp, inventory, and rhai-policy. For a lean embed build, turn them off:

ares-server = { version = "0.10", default-features = false }

The Axum HTTP stack stays a compiled dependency of the package either way. It never binds a socket unless the Http service runs.

The public surface

ares_server re-exports this set (see src/lib.rs):

ExportPurpose
ContextThe Cordis context; holds every service
ExecuteUnified agent execution service
ToolsTool listing, resolution, and dispatch
LlmLarge language model client coordination
StoreTenant database (feature postgres)
Plugin, ServiceFactory and service traits from the kernel
Loader, PluginRegistryEntries-file loader and its factory table
DispatchEvent dispatch modes
register_pluginsRegisters all capability-crate factories

Types such as AgentRequest, ExecutionResult, LLMClient, LLMResponse, Tool, ToolDefinition, TenantContext, and AppError ride the same re-export list. Deeper types live in the capability crates: ares-tools, ares-llm, ares-agent, and ares-store.

End-to-end in one file

This complete program builds a context, registers services, dispatches one tool call, and runs one agent turn. Every piece is adapted from code that compiles and passes in this repository:

  • The static tool list follows tools_with_probe() in crates/ares-agent/src/execution.rs (Tools::from_static).
  • The calculator arguments and result shape come from Calculator in crates/ares-tools/src/tools/calculator.rs, re-exported by the facade.
  • The event bus provide follows the middleware tests in crates/ares-http/src/middleware/api_key_auth.rs (ctx.provide(cordis::EventsService::new())).
  • The agent turn uses Execute::run; with no Llm on the context it takes the documented echo fallback in crates/ares-agent/src/execution.rs.
[dependencies]
ares-server = "0.10"
ares-cordis = "0.10"
tokio = { version = "1", features = ["rt-multi-thread", "macros"] }
serde_json = "1"
use std::sync::Arc;

use ares_server::{AgentRequest, Context, Execute, Tool, Tools};

#[tokio::main(flavor = "multi_thread", worker_threads = 2)]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // One context is the whole graph.
    let ctx = Context::new_root();

    // The event bus. Tool calls fan their arguments through a
    // `tools.execute` waterfall when this service is present.
    ctx.provide(cordis::EventsService::new());

    // A tool set with one real tool. Calculator answers
    // {"result": <number>} for basic arithmetic.
    ctx.provide(Tools::from_static(
        [Arc::new(ares_server::Calculator) as Arc<dyn Tool>],
    ));

    let tools = ctx.get::<Tools>().expect("tools on context");
    let out = tools
        .execute(&ctx, "calculator", serde_json::json!({
            "operation": "add", "a": 2, "b": 3
        }))
        .await?;
    println!("calculator -> {}", out["result"]); // calculator -> 5.0

    // One agent turn through the shared engine. No Llm is provided, so
    // run() returns the message itself through the echo fallback path.
    let execute = Execute::new();
    let req = AgentRequest {
        agent_name: "echo".to_string(),
        message: "hello".to_string(),
        ..Default::default()
    };
    let result = execute.run(&req, &ctx).await?;
    println!("agent -> {}", result.response.content); // agent -> hello

    Ok(())
}

Run it with cargo run. This exact program was compiled and executed against the workspace; its output:

calculator -> 5.0
agent -> hello

Three details keep this working:

  • Use a multi-threaded Tokio flavor. Kernel plugin activation calls tokio::task::block_in_place, which current-thread runtimes reject.
  • Execute is stateless; construct or provide one instance anywhere. Do not isolate it per tenant — isolating it hides the root instance from request paths (the regression guard test request_tenant_ctx_keeps_root_execute_resolvable pins this).
  • Add an Llm::from_client(...) provider to make the agent turn produce real model output instead of the echo fallback.

Minimal embed through the loader

The server boots as one ordered pass over an entries file. A library embed can run the same pass with two calls. This example adapts the in-tree test calculator_entry_loads_factory_and_executes_tool (src/main.rs) and the factory-registration helper register_loader_factories in the same file:

use ares_server::{Context, Loader, PluginRegistry, register_plugins};
use cordis::loader::Entry;

#[tokio::main(flavor = "multi_thread", worker_threads = 2)]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // One context is the whole graph.
    let ctx = Context::new_root();

    // Fill the string-keyed factory table, then put it on the context.
    ctx.provide(PluginRegistry::new());
    let registry = ctx.get::<PluginRegistry>().expect("registry");
    register_plugins(&registry);

    // One entry describes one service instance.
    let entry = Entry {
        id: "calc".to_string(),
        plugin: "CalculatorService".to_string(),
        config: serde_json::json!({}),
        ..Default::default()
    };
    Loader::instantiate(&ctx, &entry.plugin, &entry.config, &entry.id)?;

    // The factory provided a live service. Call it.
    let calc = ctx.get::<ares_tools::CalculatorService>().expect("calculator");
    let output = calc
        .execute(serde_json::json!({ "operation": "add", "a": 2.0, "b": 3.0 }))
        .await?;
    println!("{}", output["result"]); // 5

    Ok(())
}

Three facts matter here:

  • Use a multi-threaded Tokio runtime. Factories run tokio::task::block_in_place, which current-thread runtimes reject (crates/ares-tools/src/plugins.rs).
  • register_plugins registers the manual fallback chain. With the default inventory feature, compile-time collected nodes add the same keys; both paths land identical factory names (tests/server_inventory_probe.rs).
  • A second instantiation of the same service type fails with a duplicate-provider error. Single-source discipline is active in libraries too (src/main.rs, same test).

Pulling facades at call time

Services are values behind their types. Code anywhere in your process pulls one with ctx.get::<T>(). The kernel walks parent contexts, honors tenant realms, and returns None when nothing provides the type. This mirrors execute_lists_tools_using_tenant_context_intercept in crates/ares-agent/src/execution.rs:

#![allow(unused)]
fn main() {
use ares_server::{AgentRequest, Context, Execute, Tools};

async fn run_agent(ctx: &std::sync::Arc<Context>) -> Result<String, ares_server::AppError> {
    // Shared engine, pulled at call time like the tests do.
    let execute = ctx.get::<Execute>().expect("Execute on context");

    if let Some(tools) = ctx.get::<Tools>() {
        for def in tools.list(ctx) {
            println!("available tool: {}", def.name);
        }
    }

    let req = AgentRequest {
        agent_name: "echo".to_string(),
        message: "hi".to_string(),
        ..Default::default()
    };
    let result = execute.run(&req, ctx).await?;
    Ok(result.response.content)
}
}

The same pattern reaches the other facades:

#![allow(unused)]
fn main() {
// Llm: complete a prompt (crates/ares-llm/src/llm_service.rs, `complete`).
if let Some(llm) = ctx.get::<ares_server::Llm>() {
    let reply = llm.complete(ctx, "Summarize this document.").await?;
}

// Store: tenant database handle. Requires the postgres feature.
#[cfg(feature = "postgres")]
if let Some(store) = ctx.get::<ares_server::Store>() {
    let pool = store.pool();
    let _ = pool;
}
}

For tests and proofs, pin an in-process client instead of a network provider. Llm::from_client exists for exactly this (crates/ares-llm/src/llm_service.rs). The echo client below adapts the crate's own test double:

#![allow(unused)]
fn main() {
use ares_server::{AppError, ConversationMessage, LLMClient, LLMResponse, ToolDefinition};

struct EchoClient;

#[async_trait::async_trait]
impl LLMClient for EchoClient {
    async fn generate(&self, prompt: &str) -> Result<String, AppError> {
        Ok(format!("echo:{prompt}"))
    }

    async fn generate_with_system(&self, _system: &str, prompt: &str) -> Result<String, AppError> {
        self.generate(prompt).await
    }

    async fn generate_with_history(
        &self,
        _messages: &[(String, String)],
    ) -> Result<LLMResponse, AppError> {
        Ok(LLMResponse { content: String::new(), tool_calls: vec![], finish_reason: "stop".into(), usage: None })
    }

    async fn generate_with_tools(
        &self,
        _prompt: &str,
        _tools: &[ToolDefinition],
    ) -> Result<LLMResponse, AppError> {
        Ok(LLMResponse { content: String::new(), tool_calls: vec![], finish_reason: "stop".into(), usage: None })
    }

    async fn generate_with_tools_and_history(
        &self,
        _messages: &[ConversationMessage],
        _tools: &[ToolDefinition],
    ) -> Result<LLMResponse, AppError> {
        Ok(LLMResponse { content: String::new(), tool_calls: vec![], finish_reason: "stop".into(), usage: None })
    }

    async fn stream(
        &self,
        _prompt: &str,
    ) -> Result<Box<dyn futures::Stream<Item = Result<String, AppError>> + Send + Unpin>, AppError> {
        Err(AppError::Internal("echo stream not implemented".into()))
    }

    async fn stream_with_system(
        &self,
        _system: &str,
        _prompt: &str,
    ) -> Result<Box<dyn futures::Stream<Item = Result<String, AppError>> + Send + Unpin>, AppError> {
        Err(AppError::Internal("echo stream not implemented".into()))
    }

    async fn stream_with_history(
        &self,
        _messages: &[(String, String)],
    ) -> Result<Box<dyn futures::Stream<Item = Result<String, AppError>> + Send + Unpin>, AppError> {
        Err(AppError::Internal("echo stream not implemented".into()))
    }

    fn model_name(&self) -> &str {
        "echo"
    }
}

let llm = ares_server::Llm::from_client(std::sync::Arc::new(EchoClient));
let reply = llm.complete(&ctx, "hi").await?;
assert_eq!(reply, "echo:hi");
}

Register a custom plugin

A plugin is a typed factory. It declares a config type and the service it provides; the kernel calls apply once, inserts the value, and tracks the fiber for lifecycle and hot reload. The trait lives in crates/cordis/src/registry.rs; the shape below follows PipelinePlugin (crates/ares-agent/src/pipeline.rs) and the kernel's own registry tests:

#![allow(unused)]
fn main() {
use std::sync::Arc;

use ares_server::{Context, Plugin, Service};
use cordis::{CordisError, RegistryService};

pub struct GreetingService {
    prefix: String,
}

impl GreetingService {
    pub fn greet(&self, name: &str) -> String {
        format!("{}{name}", self.prefix)
    }
}

impl Service for GreetingService {}

#[derive(Default, serde::Serialize, serde::Deserialize)]
pub struct GreetingConfig {
    pub prefix: String,
}

pub struct GreetingPlugin;

impl Plugin for GreetingPlugin {
    type Config = GreetingConfig;
    type Provides = GreetingService;

    fn apply(
        &self,
        _ctx: &Arc<Context>,
        config: Self::Config,
    ) -> Result<Arc<Self::Provides>, CordisError> {
        Ok(Arc::new(GreetingService { prefix: config.prefix }))
    }
}
}

Register through the kernel registry and pull the service back:

#![allow(unused)]
fn main() {
let ctx = Context::new_root();
let registry = RegistryService::new();

registry.register(
    &ctx,
    GreetingPlugin,
    GreetingConfig { prefix: "hello, ".to_string() },
)?;

let greeting = ctx.get::<GreetingService>().expect("provided by apply");
assert_eq!(greeting.greet("ares"), "hello, ares");
}

To expose the same plugin to the entries loader, add a string-keyed factory. The closure-free function form follows the capability-crate factories in crates/ares-tools/src/plugins.rs:

#![allow(unused)]
fn main() {
fn factory_greeting(
    ctx: &Arc<Context>,
    config: &serde_json::Value,
) -> Result<cordis::FiberId, cordis::CordisError> {
    let cfg: GreetingConfig = serde_json::from_value(config.clone())
        .map_err(|e| CordisError::Configuration(e.to_string()))?;
    let svc = GreetingService { prefix: cfg.prefix };
    tokio::task::block_in_place(|| {
        tokio::runtime::Handle::current().block_on(ctx.plugin(svc))
    })
}


// Pull the factory table from the context, then add the key
// after register_plugins(&registry).
let factories = ctx.get::<ares_server::PluginRegistry>().expect("registry");
factories.register("Greeting", std::sync::Arc::new(factory_greeting));
}

An entry can now name it:

[[entry]]
id = "greeter"
plugin = "Greeting"
disabled = false

[entry.config]
prefix = "hello, "

Lifecycle hooks every service can implement

The Service trait (crates/cordis/src/service.rs) has three methods with working defaults. Override only what you need:

HookDefaultWhen the kernel calls it
name()The Rust type nameDiagnostics and duplicate-provider messages
init(ctx)Returns Ok(None)Once per activation, after apply produces the value
check()Returns trueEvery time a freshly built instance meets the graph

init returns an optional cleanup handle. The boxed value is any closure with the Disposable shape; the kernel pushes it onto the owning fiber's undo accumulator. Use it to close connections or cancel background work:

#![allow(unused)]
fn main() {
use cordis::effect::Disposable;
use cordis::ServiceInitFuture;

impl Service for PoolService {
    fn name(&self) -> &'static str { "pool" }

    fn init(&self, _ctx: &Arc<Context>) -> ServiceInitFuture<'_> {
        Box::pin(async move {
            println!("pool online");
            // Returned cleanup runs once when the fiber disposes.
            let guard: Box<dyn Disposable> = Box::new(|| println!("pool closed"));
            Ok(Some(guard))
        })
    }
}
}

check() is an availability predicate evaluated at build-and-met points, such as after RegistryService::register runs your factory and before the value is provided. A false verdict is terminal and inspectable: the fiber rests in state Failed, never Pending (crates/cordis/src/service.rs, check documentation). Services whose availability changes later — circuit breakers, feature gates — must re-provide or notify instead of relying on spontaneous re-checks. The kernel holds no downcasting machinery for per-read checks.

Dispose ordering

One rule covers every teardown path: effects run in reverse registration order, last-in first-out (docs/src/kernel/runtime.md; Fiber::dispose). Concretely, for a plugin that registered a listener, then a timer, then returned an init cleanup handle:

  1. The init cleanup handle runs.
  2. The timer cancels.
  3. The listener detaches.

The same LIFO pass runs during reactive loss (Unloading), during loader rollback (newest-first across applied steps), and at process shutdown. Nothing observes a half-torn configuration: by the time an undo starts, every effect registered after it is already gone. Design init cleanups to depend on nothing registered later than themselves.

Error handling patterns

Kernel calls return Result<_, CordisError> (crates/cordis/src/service.rs). The enum has ten variants. Match the ones you can act on and let the rest bubble:

#![allow(unused)]
fn main() {
use cordis::{CordisError, ValidationIssue};

fn describe(err: &CordisError) -> String {
    match err {
        // Config failed to deserialize into the plugin's Config type.
        CordisError::Configuration(msg) => format!("bad wiring: {msg}"),
        // Structured pre-flight failures carry placed issues.
        CordisError::Validation(err) => err
            .issues
            .iter()
            .map(|i: &ValidationIssue| format!("{} at {}", i.message, i.path.join(".")))
            .collect::<Vec<_>>()
            .join("; "),
        // Two plugins provide the same service type.
        CordisError::DuplicateProvider { name, owner } => {
            format!("{name} provided twice; second by {owner}")
        }
        // A transition lease timed out after 10 s of contention.
        CordisError::TransitionStuck { fiber, waited_ms } => {
            format!("fiber {fiber} stuck for {waited_ms} ms")
        }
        // Typed property reads that fail to downcast.
        CordisError::PropertyTypeMismatch { name, expected } => {
            format!("property {name} is not a {expected}")
        }
        other => other.message(), // Fiber, ServiceNotFound, InvalidConfig,
                                  // Internal, ReadOnlyProperty
    }
}
}

Guidance per variant, grounded in kernel behavior:

  • ServiceNotFound means nothing provides the type on this context or its parents. Check isolate labels first: get_isolated::<T>("other-realm") fails even when another realm serves the type.
  • DuplicateProvider is single-source discipline firing. Drop one factory or move it to a realm.
  • InvalidConfig is stringly; Validation is structured. Only Validation exposes machine-readable issues through validation_error(). The loader lifts structured issues with CordisError::validation(...), keeping the "invalid config: ..." Display prefix.
  • Apply errors recorded on a fiber are terminal (Failed). Retrying the same registration never recovers it; re-register with a fresh fiber instead (docs/src/kernel/lifecycle.md).

At HTTP edges, convert with the existing adapter: HttpError::from(app_err) wraps an AppError, and app_error_into_response renders the {"error", "code"} body (crates/ares-http/src/error.rs).

Loader entries versus pure-library composition

Pick one composition style per process:

AspectLoader entriesDirect provides
Declarationconfig/cordis-entries.tomlRust code
Service shape[entry]: id, plugin, JSON config, disabled, optional isolate and interceptTyped values
Hot reloadYes, through the admin surface and journalNo; rebuild and restart
Server bootRequired; the binary exits without a working Overlay entryNot applicable
Best fitDeployment-time wiring, operator editsTests, embedded agents, fixed topologies

The server treats the entries file as the program: it loads the tree, composes includes, instantiates each enabled entry in file order, and fills empty configs after the Overlay entry lands (boot_loader_program, src/main.rs; entry fields in crates/cordis/src/loader.rs). A library embed can reuse that machinery with Loader::load_from_file and Loader::instantiate_entry, or skip it and call ctx.provide(...) directly. Direct provides are what every kernel and agent test in the repository does.

Choosing a composition style

Answer two questions. Who changes the wiring? How often does it change?

SituationStyleWhy
Wiring is fixed at build time; you own all call sitesDirect providesOne less file format; the compiler checks every reference
Operators re-wire services between releases without rebuildsLoader entriesEditing TOML and reloading beats shipping a binary
You need admin-surface retire/replace/patch per serviceLoader entriesThe journal tracks each entry's fiber for lifecycle routes
Unit tests and examplesDirect providesEvery kernel and agent test composes this way
Product boots from entries; tests exercise one pluginHybridBoth styles feed the same context and coexist

The hybrid form registers your string-keyed factory beside register_plugins, then names it from an entry. The factory table and the typed store meet inside the kernel: an entry instantiates through the factory, and the produced value lands as a typed provider like any direct provide. Use this when library users should wire your plugin declaratively while your own tests keep constructing it directly.

Tenant isolation

Multi-tenant hosts keep shared services on a root context and scope per-request data into child contexts. Two mechanisms exist:

  • Isolate labels a service type with a realm name. get on the labeled child resolves only matching realms through get_isolated.
  • Intercept shadows one service type with a request-scoped value, such as a TenantContext.

Both come from the kernel context (crates/cordis/src/context.rs). The production helper is tenant_scope in crates/ares-agent/src/execution.rs; its behavior is pinned by the test execute_isolate_label_wins_over_intercept_for_tools:

#![allow(unused)]

fn main() {
use ares_server::{Context, TenantContext, TenantTier, Tool, Tools};

// Root context holds shared services.
let root = Context::new_root();

// Scope to one tenant: realm label on Tools, then provide inside the realm.
let scoped = root.isolate::<Tools>("acme");
scoped.provide(Tools::from_static(Vec::<Arc<dyn Tool>>::new()));

// Add request data on top. Isolate wins over intercept for resolution.
let request = scoped.with_intercept(TenantContext::new(
    "acme".to_string(),
    TenantTier::Pro,
));

// The realm's tool set resolves on the request context...
assert!(request.get::<Tools>().is_some());
assert!(request.get_isolated::<Tools>("acme").is_some());

// ...and stays invisible under a different label.
assert!(request.get_isolated::<Tools>("other").is_none());
}

Rules the tests enforce:

  • An isolate label beats a TenantContext intercept during agent resolution (user_id_from_ctx_isolate_label_wins_over_intercept, resolver tests).
  • Execute stays shared. It is a stateless engine; isolating it hid the root instance and broke request paths, so tenant_scope isolates Tools only.
  • Background jobs scope with isolate alone; HTTP requests add the intercept afterward (request_tenant_ctx, same module).

Where to go next

Kernel concepts

The Cordis kernel is a typed service graph with a lifecycle engine. This chapter explains the ideas. The how-to chapters cover the state machine and interception in detail.

Ideas

Typed service graph. A Context stores services keyed by Rust TypeId. A child context walks its parent chain like a prototype chain. Isolate labels mark realm boundaries. Two realms can provide the same service type. Multi-tenant isolation uses this.

Fibers own lifecycles. Each registration creates a fiber. The fiber holds the state machine, the dependency epoch, and an undo accumulator. When a provider changes, the kernel refreshes dependent fibers. Disposing a fiber runs every undo in reverse order (LIFO).

Effects are undos. Every provide, timer, and listener pushes an undo closure onto its fiber. There is no separate effect tree. Undo metadata carries a label and a timestamp for inspection.

Event-first middleware. Product code talks over an event bus with five dispatch modes: emit, parallel, serial, bail, and waterfall. Waterfall handlers wrap downstream work through a next continuation.

Interception. Five kernel meta-events act as veto points around reads, writes, config resolution, restarts, and listener registration. A sixth meta-event observes every non-internal dispatch. With no listener registered, each point costs two map lookups.

Layered overrides. An intercept override shadows one service type. Overrides stack in layers. The innermost layer wins. Accessors are name-keyed computed properties beside the TypeId store; they bypass interception on purpose.

Declarative management. A loader applies entry trees against a journal. It stages changes in two phases and rolls back on failure. A reflect service fans change notifications out through a dependency graph.

Why a kernel

Applications wire objects together by hand. Hand wiring creates four recurring problems. The kernel solves each one with one mechanism.

Problem 1: wiring. Components reference each other by concrete type. Every change to a component ripples through every construction site. The kernel replaces hand wiring with typed lookup: provide once, resolve anywhere through the parent chain. Consumers name what they need; nobody names who builds it.

Problem 2: lifecycle. Services start in the wrong order and stop in the wrong order. Timers, listeners, and connections leak when shutdown misses one. The kernel gives every registration a fiber. The fiber tracks every undo closure in registration order. Teardown runs them in reverse, always, even during reactive dependency churn.

Problem 3: observation. Once components talk directly, nothing sees the traffic. Debugging relies on print statements at every call site. The kernel routes reads, writes, config resolution, restarts, and listener registration through six meta-events. One subscription observes every sensitive operation. With no subscriber, each point costs two map lookups.

Problem 4: replacement. Swapping an implementation means touching call sites or restarting the process. The kernel separates declaration from use. A loader stages a new entry tree in two phases and rolls back on failure. Layered intercept overrides shadow one service type for one subtree without touching the parent context. Hot swap of native plugins rides the same fibers.

The result is one graph that owns construction, teardown, inspection, and substitution. Product code states facts; the kernel keeps the facts true.

Glossary

TermDefinition
FiberLifecycle owner of one registration. Holds the state machine, the dependency epoch, and the undo accumulator. Type: cordis::fiber::Fiber.
ServiceAny value stored by Rust TypeId. Implement the empty Service trait to participate.
EffectOne labeled undo closure on a fiber. Every provide, timer, and listener pushes one. Disposal pops effects last-in, first-out.
EventNamed message on the event bus. Dispatch modes: emit, parallel, serial, bail, waterfall.
Meta-eventKernel-owned veto point such as internal/get or internal/set. Distinct from product events.
IsolateRealm label marking a boundary. Contexts beyond an isolate do not see layers or services across it. Multi-tenant isolation uses this.
LoaderDeclarative engine applying entry trees against a journal in two phases, with rollback.
EpochString encoding every declared dependency version. A changed epoch triggers a refresh pass.
Readiness gatePredicate holding a fiber in reversible Pending until it reports ready.
AccessorName-keyed computed property beside the TypeId store. Bypasses interception by design.

Map: idea to module

IdeaImplemented in
Typed store, parent walk, isolate realms, accessors, layered overridescordis::context
Fibers, states, epochs, undo accumulatorcordis::fiber
Single-provider registry, plugins, readiness gatescordis::registry
Disposable effectscordis::effect
Event bus, dispatch modes, meta-eventscordis::events
Declared event contracts and typed payloadscordis::events_catalog, cordis::events_payload
Errorscordis::service (CordisError)
Declarative loader, staged batches, journalcordis::loader; LoaderJournal in the crate root
Change fan-out and BFS refreshReflectService in the crate root
File-watch hot reloadcordis::watcher, cordis::reload, cordis::stamp
Plugin-module graphcordis::module_graph
Native-code hot swapcordis::hmr
Dependency-cycle detectioncordis::cycles
Peer-dependency versionsContext::VERSION_MAJOR_SCALE (context), constraints in fiber
Executable kernel guaranteescordis::metatheory
Timer and logger primitivescordis::timer, cordis::logger
Entry composition (@include, @group, Rhai interpolation)cordis::compose, cordis::rhai_service
Supervised worker exit protocolcordis::worker

Read Lifecycle for the full state machine. Read Interception for the exact veto semantics.

Fiber lifecycle

A fiber is the lifecycle owner of one registration. Its state machine has seven states. The type is cordis::fiber::FiberState.

States

StateMeaning
Inactive { error }Not serving. Pristine, waiting for a first apply, or resting after an unsatisfied declaration.
LoadingA plugin activation is in flight.
Active { epoch }Serving. epoch encodes every declared dependency version.
ReloadingA refresh pass runs. Effects may be undone and re-applied.
Unloading { error }Effects are being disposed in reverse order (LIFO).
PendingReactive waiting. Effects were disposed, but the fiber is not disposed. It waits for dependencies or its readiness gate.
Failed { error }Terminal after an apply error, or inspectable after an availability-predicate rejection.

Pending carries no error. Reactive waiting is not a failure.

State diagram

stateDiagram-v2
    [*] --> Inactive
    Inactive --> Loading : dependencies satisfied and gate open
    Loading --> Active : runner returns Ok(true)
    Loading --> Inactive : runner returns Ok(false)
    Loading --> Failed : runner returns Err
    Active --> Reloading : dependency epoch changed
    Active --> Unloading : reactive dependency loss or closed gate
    Reloading --> Unloading : effects undone before re-apply
    Unloading --> Pending : fiber ever activated
    Unloading --> Inactive : dispose completes
    Pending --> Loading : gate opens or provider returns
    Pending --> Pending : still-closed gate (no-op)
    Failed --> Failed : terminal; refreshes refuse to revive

Observers see the exact documented sequence across one reactive cycle: Active, Reloading, Unloading, Pending, Loading, Active. Subscribe with Fiber::subscribe_state. Test anchor: observer_sees_unloading_pending_active_sequence.

One fiber's life, narrated

Follow one registration through its whole life. Every step names the region inside Fiber::refresh (crates/cordis/src/fiber.rs) that owns it.

1. Registration. RegistryService::register runs the factory once, records the raw config on the fiber, and rests it Inactive until its dependencies exist. No undo exists yet.

2. Dependencies appear. A provider registers elsewhere. The kernel re-kicks our fiber. refresh acquires the transition lease, then checks the terminal short-circuit region: no apply error is recorded, so the pass continues.

3. Activation. The pass reaches the runner region. The runner re-runs and pushes effects — timers, listeners, provided values — onto the undo accumulator. Ok(true) stores the new epoch string and rests the fiber Active. Observers saw Inactive, Loading, Active.

4. Provider withdrawal. Another module disposes the dependency's provider. A refresh computes a new epoch that differs from the stored one. The genuine-loss guard fires: the provider reads as unavailable with version 0, the fiber has a reload runner, and it activated at least once.

5. Unloading. The reactive-loss region sets Unloading { error: None } and pops every effect in LIFO order. Timers cancel. Listeners detach. The provided value leaves the store. This is why consumers never observe a half-torn configuration.

6. Pending. Effects are gone but the fiber survives. It rests Pending with no error field. Registry pruning skips it. Strict get refuses its value because none exists; nothing lies about availability.

7. Return. The provider re-registers. ReflectService fans the change out and the fiber re-kicks. The Pending fast-path sees an open path this time and falls through to one full pass. The runner region executes again under Loading, effects are rebuilt, and a new epoch lands.

8. Reactivated. The fiber serves again with fresh effects. Consumers that waited on strict get now resolve it.

Two details matter in production:

  • If the plugin factory had errored during any pass, region 3 would have set the terminal marker instead. Region 2's short-circuit then returns forever. Recovery means explicit re-registration with a fresh fiber id.
  • A constraint refusal over a live provider never reaches region 4. The provider still reports available, so the fiber rests Inactive, not Pending. See the version rules below.

Test anchors covering this walk: dependent_reactivates_when_provider_returns, fiber_refresh_passes_through_reloading, observer_sees_unloading_pending_active_sequence.

Reactive Pending reversal rules

The kernel converts a working configuration into Pending only under narrow conditions:

  • The fiber lost a dependency genuinely: the provider is unavailable and reads provider version 0 again after its undo ran.
  • The fiber has a reload runner and completed at least one fully-satisfied application (ever_activated). Fibers that never activated rest Inactive instead.
  • A peer-version constraint refusal over a live provider is policy, not loss. The provider is still available, so those fibers stay Inactive.

Apply errors are different. A factory error sets a terminal marker (apply_failed). Refreshes never revive such a fiber. Recovery means explicit re-registration with a fresh fiber id. Test anchors: dependent_reactivates_when_provider_returns, failed_stays_failed_on_dep_return.

Late inject declarations never rest Pending. An unsatisfied declaration deactivates the fiber immediately with the note "missing or inactive dependency". The next reactive refresh converts a genuine loss to Pending.

A re-kick on a Pending fiber with a still-closed gate is a no-op. An open gate falls through to one full pass: Pending, Loading, Active. RegistryService::prune_disposed keeps Pending fibers alive so they can reactivate. Tests: pending_fiber_survives_prune_disposed, prune_disposed_drops_disposed_but_keeps_failed.

Every transition acquires the fiber inertia guard through a bounded wait (TRANSITION_WAIT, 10 seconds). Same-thread reentrancy fails immediately. Cross-task contention times out with an error naming the stuck fiber id. Tests: reentrant_transition_detected_fast, contention_times_out_named.

Peer-version compatibility

A provider can publish a version, and an inject can demand one. The scheme lives in Context::VERSION_MAJOR_SCALE (crates/cordis/src/context.rs). One scale constant splits a plain u64 into two fields:

$$S = 100,000, \qquad \operatorname{major}(v) = \left\lfloor \frac{v}{S} \right\rfloor, \qquad \operatorname{floor}(v) = v - S \cdot \operatorname{major}(v)$$

An inject constrained with requirement \(r = M \cdot S + f\) accepts a provider of version \(p\) if and only if:

$$\text{satisfied}(p, r) \iff \bigl(\operatorname{major}(p) = \operatorname{major}(r)\bigr) ;\wedge; (p \ge r)$$

Read the two clauses as separate guarantees:

  • Same major means compatible. Peer dependencies never bind across a breaking boundary. A provider at major \(M+1\) leaves the dependent unsatisfied even when its remainder is huge.
  • At or above the floor means feature-complete. Under equal majors, \(p \ge r\) reduces to \(\operatorname{floor}(p) \ge f\). The provider has every capability the consumer asked for.

Any mismatch keeps the inject unsatisfied. The dependent fiber rests Inactive rather than silently binding a wrong version. This is the "constraint refusal over a live provider" rule from the section above.

Legacy provide installs version 0. Version 0 satisfies only unconstrained injects. Migrate providers to provide_versioned to opt into matching. Declare the constraint side with Fiber::declare_inject_versioned.

Worked example: a provider publishes \(p = 200,003\) (major 2, floor 3).

Requirement \(r\)major match\(p \ge r\)Verdict
\(2 \cdot S + 3 = 200,003\)yesyesSatisfied
\(2 \cdot S + 7 = 200,007\)yesnoUnsatisfied; floor too high
\(1 \cdot S + 1 = 100,001\)noyesUnsatisfied; cross-major
unconstrainedn/an/aSatisfied

Source: VERSION_MAJOR_SCALE documentation in context.rs; constraint matching in cordis::fiber. Test anchor: refresh_reruns_provider_after_dependency_version_change in registry.rs.

Readiness barriers versus availability predicates

The two mechanisms answer different questions.

Readiness barrier — "not yet". Install one with RegistryService::register_with_readiness and compose predicates with ReadinessBarrier::new, ReadinessBarrier::and, and with_readiness. While the predicate reports false:

  • The fiber rests in inspectable Pending. This is quiet, reversible waiting.
  • The factory still runs once at registration, so config errors surface early.
  • Strict ctx.get refuses values owned by non-Active fibers, so consumers never see the half-ready service.
  • Declare watch keys with ReadinessBarrier::watching. Any settlement on those types fans out through ReflectService and re-kicks the gated fiber.

Test anchors: ready_when_holds_pending_until_true_then_activates, readiness_composes_and_semantics, external_rekick_reactivates_waiting_fiber.

Availability predicate — "never, as built". Implement Service::check. The registry consults it before the value reaches consumers. A rejection rests the fiber as Failed { error: "availability predicate rejected service" }. Registration itself stays non-throwing; later refreshes converge once the instance reports healthy.

Test anchors: availability_predicate_rejection_registers_failed, predicate_passing_reregistration_activates_dependents.

In short: a closed readiness gate parks the fiber quietly as reversible Pending. An availability rejection fails loudly to Failed{error}. They are complements, not alternatives.

Two-phase staged reload

The loader applies a desired entry tree against the current tree in two phases:

  1. Verify without mutating. Phase one resolves entries and pre-flights config trials against scratch candidates. On the first failed verification, the batch aborts. Nothing changed, so no rollback is needed.
  2. Apply in dependency order. Phase two applies verified candidates. Begin and rebuild steps run first, then updates, then retire steps. On a failure during phase two, the loader rolls back every applied change newest-first: configs restore, rebuilt fibers dispose. The live tree serves the originals.

On any failure the loader leaves current unchanged, so a retry re-diffs cleanly. Config-only patches on active fibers go through the update path with a pre-flight trial, so factory apply counts stay flat.

Rollback scenarios

The table lists what each phase-two failure undoes. "Applied" means the step already ran when the failure happened; later steps never start.

Failure duringRollback actionLive-tree result
A config update stepRestore that fiber's prior effective config from the staged candidateOld config keeps serving
A rebuild (dispose + re-apply) stepDispose the rebuilt fiber's effects in LIFO orderOriginal fiber's value is gone only if it was retired earlier in the batch — retire steps run after rebuilds
A begin stepNothing to undo; batch stops before any mutation of existing fibersTree untouched
A retire stepRe-apply is not attempted; earlier applied updates and rebuilds roll back newest-firstOriginals serve again
Verification (phase one)No rollback needed; nothing mutatedIdentical tree

Newest-first ordering matters. Suppose a batch updates A, then rebuilds B which injects A. If B's re-apply fails, rollback disposes B' first and restores A second — the reverse of application order. Consumers of A see its original value throughout.

Test anchors: staged_batch_rolls_back_on_first_failure, staged_batch_applies_in_order_on_success.

Cascade batching

When a provider fiber sits mid-config-update, dependent refreshes defer instead of churning. Loader::cascade_defer_needed probes an in-flight ledger. A deferred dependent rests Pending quietly. One post-settle re-kick converges every deferred dependent at once, instead of running one full cascade wave per concurrent patch. Without a RegistryService on the context, the probe always reports false and legacy behavior holds.

The timeline below shows two dependents deferring behind one provider patch and converging together:

sequenceDiagram
    participant L as Loader
    participant P as Provider fiber (db)
    participant D1 as Dependent (api)
    participant D2 as Dependent (worker)
    L->>P: mark in-flight, stage new config
    L->>P: apply config update
    D1->>P: refresh() re-kick
    P-->>D1: cascade_defer_needed = true
    D1->>D1: Unloading -> undo effects -> rest Pending
    D2->>P: refresh() re-kick
    P-->>D2: cascade_defer_needed = true
    D2->>D2: Unloading -> undo effects -> rest Pending
    L->>L: settle batch, clear in-flight marker
    L->>D1: single post-settle re-kick
    L->>D2: single post-settle re-kick
    D1->>D1: Pending -> Loading -> Active
    D2->>D2: Pending -> Loading -> Active

Without batching, each concurrent patch would trigger one full cascade wave through every dependent. With three patches landing together, that is three teardown-rebuild cycles per dependent. Batching collapses them into one.

Interception

The kernel routes its own sensitive operations through the event bus. Five meta-events are veto points. A sixth observes dispatches. The constants live in cordis::events:

Meta-eventGuards
internal/getStrict service reads (Context::get)
internal/setService writes (provide paths)
internal/configConfig resolution before a plugin apply
internal/updateRestart scheduling on config change
internal/listenerListener registration
internal/dispatchObservation only; no veto power

These meta-events do not join the product event catalog, so catalog validation skips them.

Decision guide: which meta-event for which need

Pick by the operation you want to guard, not by the mechanism.

You need to...Subscribe toVeto power
Hide or rewrite one service from consumersinternal/getRefuse the read, redirect to the parent frame, or substitute the value
Freeze writes or audit every provideinternal/setBlock the write; old value stays intact
Transform or validate plugin configurationinternal/configReplace the effective config, or fail activation
Defer a restart during a change windowinternal/updatePark the proposed config in vetoed_config, keep serving
Gate listener registration per realminternal/listenerCancel registration with an inert handle
Trace every product dispatchinternal/dispatchNone — observation only

Rules of thumb:

  • One need, one meta-event. Do not emulate internal/update vetoes with config rewriting; update preserves the running application untouched, while config changes force a new application.
  • For multi-tenant scoping, register listeners inside the isolate realm. Meta-event listeners are ordinary listeners; they obey isolate boundaries like everything else on that context.
  • Prefer internal/dispatch over wrapping product code in logging. The observer sees every dispatch without touching call sites.

Zero-cost gates

Every interception point calls EventsService::listener_count first. The count is one pair of map lookups over the flat and waterfall registries. With zero listeners, the operation proceeds on its historical path with no chain work. This gate makes the whole feature free until someone subscribes.

Veto semantics

All veto chains run as bail chains: handlers run in registration order, and the first non-null result terminates the chain.

Ordering guarantees within one chain

Order inside a chain is deterministic:

  • Default registrations append to the back of that event's list. First registered runs first.
  • A registration with EventOptions { prepend: true } inserts at the front. It runs before every default registration, including ones made earlier.
  • Handlers are async but the chain awaits them sequentially. Handler N+1 never starts before handler N returns. There is no concurrency inside a bail chain.
  • The first non-null result wins. Later handlers do not run at all — not even for observation. Put audit listeners on internal/dispatch instead if they must always run.
  • Disposing a listener handle removes exactly that slot. Remaining handlers keep their relative order.

Test anchors: prepend_ordering_observed in events.rs.

internal/get

Payload: { "service", "ctx" }. Rules for intercept_get:

  • No listeners: Ok(None) at map-lookup cost.
  • A chain terminal of null passes the read through untouched.
  • A non-null result replaces what the consumer sees.
  • A chain error vetoes the read.

The synchronous bridge in Context::get maps verdicts further:

  • Non-null JSON with "refuse": true refuses the read outright (ReadVerdict::Refuse, lookup returns None).
  • Any other non-null value redirects the read to the parent frame (ReadVerdict::RedirectFrame). This frame's store and intercept bindings are skipped, so a parent binding serves the read.
  • Null passes normally.

Test anchors: get_interceptor_rewrites_read, accessor_bypasses_intercept_waterfalls.

internal/set

Payload: { "service", "ctx" }. A chain error vetoes the write. The previous value stays fully intact; no store, owner, or version mutation happens. Null and pass-through allow the write unchanged.

Test anchor: set_interceptor_vetoes_write_leaves_old_value.

internal/config

Input: the raw config value. The chain's non-null terminal is the effective configuration. A null terminal passes the raw config through. A chain error fails the activation or update that was resolving config; the fiber rests Failed.

The registry captures raw config at registration time. Each refresh pass re-resolves the effective config from that single source through Fiber::resolve_effective_config, stages it once, and the runner consumes it via effective_config_override. With no interception, the path is byte-identical to the legacy runner.

Test anchors: config_interceptor_rewrites_effective_config, config_waterfall_covers_activation_path, interceptor_error_fails_fiber_activation.

internal/update

Payload: { "service" }. Three outcomes inside Fiber::update:

  • Proceed: the chain passes (null counts as proceed here), so the restart runs.
  • Veto: a bail with explicit JSON false parks the proposed config in vetoed_config and returns Ok(()). No restart runs; the fiber keeps serving its current application. Operators can inspect what was deferred through Fiber::vetoed_config.
  • Error: a chain error propagates out of Fiber::update as an Err. The fiber stays Active on its old configuration. Nothing was applied and nothing was deferred.

Test anchors: update_interceptor_veto_skips_restart_keeps_config, update_veto_defers_config_and_returns_ok, update_error_stays_active_old_config.

internal/listener

Payload: { "event" }. Registrations are synchronous APIs, so the veto chain runs to completion on the current thread through block_in_place. Rules:

  • Ok(true) or a null terminal lets the registration proceed.
  • A bail (non-null, non-true) or a chain error cancels it. Fail-closed.
  • A cancelled registration returns an inert handle. Disposing it flips nothing.
  • On runtimes that cannot park a worker (single-thread flavors), the registration falls open and a warning records the skipped veto.

Test anchor: listener_interceptor_bail_cancels_registration_inert_handle.

internal/dispatch

Observer, not veto. Before every non-internal dispatch, listeners receive { mode, name, args } fire-and-forget. Handler results and errors drop by design: observability must never break or delay the observed operation. Internal meta-events are exempt, so observation cannot recurse into itself.

Test anchor: internal_dispatch_observes_non_internal_only.

Worked mini-scenarios

Each scenario shows the payload shape and one verdict path. Payloads are JSON objects; shapes come from the bridge implementations in events.rs and its tests.

Scenario: tenant A may not read the billing service. A listener on internal/get receives:

{ "service": "BillingStore", "ctx": "<frame id>" }

The handler returns { "refuse": true } when "service" names a billing type and the dispatch belongs to tenant A. The synchronous bridge maps that to ReadVerdict::Refuse, so Context::get returns None. Tenant B's handler returns null for the same event; null passes the read through untouched.

Scenario: freeze all writes during a migration. One listener on internal/set receives { "service": "SessionCache", "ctx": "<frame id>" } and returns an Err(CordisError). The chain aborts with that error and vetoes the write. The previous value stays fully intact: no store slot, owner, or version changes. Callers see the error; nothing else moves.

Scenario: inject feature flags into plugin config. The registry captured raw config {"pool": 4} at registration. A listener on internal/config receives that raw value and returns

{ "pool": 4, "feature_x": true }

That non-null terminal IS the effective configuration for every apply and refresh pass of this fiber. Removing the listener reverts the next refresh to the raw value.

Scenario: hold a restart during peak hours. An operator gate listens on internal/update. It receives { "service": "SearchIndex" } and returns JSON false outside the maintenance window. The proposed config parks in Fiber::vetoed_config and update returns Ok(()); the old application keeps serving. Returning any other non-null value — including an object — counts as proceed. Only explicit false is a veto.

Scenario: block a debug-only listener in production. A listener on internal/listener receives { "event": "tools/execute" }. For events matching a deny list it returns Ok(json!(false)). The registration cancels fail-closed and the caller gets an inert handle whose dispose flips nothing.

Scenario: trace latency per dispatch mode. A listener on internal/dispatch receives

{ "mode": "emit", "name": "agent/step", "args": { "n": 1 } }

before every non-internal emit. The handler records the timestamp and returns anything; results drop by design. Internal meta-events never trigger it, so tracing cannot recurse into itself.

Layered intercept chains

Intercept overrides stack per TypeId. Every set appends a layer. Layers order outermost to innermost across ancestor frames; the innermost layer is the effective value every getter returns. Context::intercept_chain returns the full chain. An isolate label is a realm boundary: layers beyond it do not leak in.

Structural comparison uses shared-instance identity (Arc::ptr_eq per layer) via Context::chains_structurally_equal. Freshly built values compare unequal by design.

Test anchors: chained_layers_append_innermost_effective, intercept_chain_returns_all_layers_in_order, inject_appends_layer.

Accessors bypass the waterfalls

Name-keyed accessors live beside the TypeId store. Register them with Context::register_accessor; alias names share one slot and dispose together. Reads use read_property (typed variant: read_property_typed); writes use write_property. Accessor reads and writes never consult or re-enter the internal/get or internal/set waterfalls. An interceptor that refuses all strict reads cannot block accessor traffic.

Test anchor: accessor_bypasses_intercept_waterfalls.

Target-carrying filtered dispatches

The *_from variants carry a target filter alongside the payload:

  • bail_from(event, payload, Option<ListenerFilter>)
  • waterfall_from(event, payload, Option<ListenerFilter>)
  • waterfall_async_from(event, payload, filter, core)

Non-global listeners whose registration options fail the filter skip this dispatch but stay registered. Global listeners bypass the filter and always participate. Kernel meta-events ride bail_from for their veto chains. A filter-empty snapshot still leaves live registrations intact; only cancelled slots prune.

Test anchors: target_carrying_dispatches_filter_per_dispatch, filter_excludes_nonmatching_contexts, global_bypasses_filter.

Re-entrancy protection

A thread-local fence (InterceptFence) marks the thread while a synchronous bridge drives its chain. Nested operations on the same thread pass through unintercepted, so an internal/get listener that reads services does not recurse into its own veto. Bridges also require a multi-thread runtime; without one they fall open and log a warning.

Why the fence exists

Consider an internal/get listener that inspects other services to make its verdict. The inspection calls ctx.get. Without protection that read re-enters the same veto chain, which reads services again. The stack grows until something breaks.

The fence is one boolean in thread-local storage:

  1. InterceptFence::enter sets it and returns a guard.
  2. The bridge drives its chain while the guard lives.
  3. Any nested operation on this thread sees the flag set. Its bridge short-circuits to "allow" before consulting listeners.
  4. Dropping the guard clears the flag.

The scope is deliberately per-thread, not global. Two worker threads can each drive their own chain at the same time. Only true nesting — one bridge inside another on the same call stack — passes through unintercepted.

The same reasoning protects internal/set, internal/config, internal/update, and internal/listener: every synchronous bridge enters the fence first. A fenced nested write proceeds unintercepted, so an interceptor can record state without tripping its own gate.

Test anchor: accessor_bypasses_intercept_waterfalls shows the related design bypass for accessors; the fence covers same-thread recursion for every bridge.

Fall-open on single-thread runtimes

Synchronous bridges park the current worker with block_in_place. Parking requires spare workers, so it needs a multi-thread tokio runtime. On a current-thread flavor there are no spare workers, and block_in_place panics.

Each bridge checks the flavor before parking:

listener_count == 0        -> historical path (zero cost)
fence already held         -> allow, no recursion
single-thread runtime      -> allow + warning: skipped veto
otherwise                  -> park and run the bail chain

Fall-open means fail-open by design choice. Registration and reads are core APIs; refusing them because of a runtime flavor would break every test harness and embedded use. The warning names the skipped veto so deployments notice when they expected enforcement. Product deployments run multi-thread flavors, where the chain always runs.

Runtime Services

This chapter describes the runtime services that live beside the kernel core: effect ownership, the fiber-scoped timer suite, the logger service, the module graph, and the tenant file fence. Every signature here comes from crates/cordis/src and crates/ares-tools/src/fence.rs.

Effect Ownership

The kernel models cleanup as effects. An effect is anything that implements one method:

#![allow(unused)]
fn main() {
pub trait Disposable: Send + 'static {
    fn dispose(self: Box<Self>);
}
}

Every closure with a compatible signature is a Disposable. A fiber holds its effects as labeled undo entries. Fiber::dispose pops them in last-in, first-out (LIFO) order and runs each undo once. A reactive pass through Unloading runs the same stack. This gives one rule: teardown order is the reverse of registration order, always.

EffectHandle dispose semantics

The timer suite returns an EffectHandle per registration. Its rules:

  • Dropping the handle does NOT cancel the effect. Callers must dispose it explicitly or let the owning fiber do it.
  • Clones share one cancellation flag. Disposing any clone cancels the registration.
  • Disposal is idempotent: the flag flips once and the teardown hook runs once.
  • handle.is_cancelled() reports the state at any time.
  • Each handle pushes a labeled undo (prefix timer:) onto the current fiber scope. Fiber disposal therefore cancels timers without caller action.

A registration made outside a fiber scope logs a warning and returns an orphan handle. An orphan still works when you dispose it by hand; no fiber cancels it automatically.

Inert handles

Two APIs return handles that may be inert: disposing them flips nothing.

  • Event listener registration rides the internal/listener veto point. When that chain bails or errors, the registration is cancelled before it enters either registry. The caller receives an inert handle whose dispose does nothing. The failure is fail-closed: an erroring veto chain cancels too.
  • Context::register_accessor returns an accessor EffectHandle. handle.dispose() removes the declaration and every alias bound to it, and returns true only when the declaration was still live. After removal, reads resolve None.

Fiber-Scoped Timer Suite

cordis::timer provides six primitives. All of them run on one shared timer thread named cordis-timer, never on the owning task. The thread sleeps until the nearest deadline in a shared wheel, drains all due entries under one short lock, then runs callbacks outside the lock. Callbacks must be cheap and non-blocking. A panicking callback is caught and logged; the thread survives.

Wheel mechanics

The wheel is a min-heap of entries ordered by (deadline, seq). The sequence number breaks ties so equal deadlines fire in registration order. One process-wide instance serves every fiber; it lives behind a LazyLock<Mutex<Wheel>>.

The driver loop has two phases:

  1. Sleep. Read the nearest deadline while holding the lock, release, then park_timeout for that long. A park with no timeout waits for the next insert. Any schedule call pushes its entry and unparks the thread, so an earlier deadline preempts a long sleep immediately.
  2. Drain. Re-acquire the lock, pop every entry whose deadline is at or before now, release, then run each job. Callbacks run outside the lock on purpose: an interval callback re-arms itself by calling schedule, which needs the lock. Holding it across callbacks would deadlock every self-re-arming pattern, including debounce and throttle emits.

Cancellation is cooperative through EffectHandle. Disposal flips the shared flag but does not remove the heap entry. Two paths make that safe:

  • The timeout job checks the flag inside the job body. A disposal racing the drain still prevents the callback from running.
  • The future-based sleep resolves early and silently when disposed while pending; nothing observes a cancelled wake-up.

A panicking callback lands in catch_unwind. The driver logs a warning and moves to the next job. One bad callback cannot starve the rest of the wheel.

PrimitiveShape
timeout(delay, callback)One-shot delay, then callback. Returns EffectHandle.
sleep(delay)One-shot delay as a future. Returns (EffectHandle, impl Future).
interval(delay, callback)Repeating callback. Returns EffectHandle.
interval_stream(delay)Repeating ticks as a pollable stream. Returns Interval.
debounce(delay)Trailing-edge burst collapse. Returns Scheduled<T>.
throttle(delay, no_trailing)Leading edge plus optional trailing edge. Returns Scheduled<T>.

Register inside with_current_fiber(&fiber, || ..) to attach the effect to that fiber. Key behaviors:

  • timeout checks the cancellation flag inside the job, so a disposal that races the drain still prevents the callback from running.
  • interval re-arms the next tick from the moment each tick fires. The cadence never runs ahead of the callback.
  • sleep resolves early and silently when its handle is disposed while the future is pending.
  • interval_stream queues ticks in a channel while nobody polls. After disposal the stream yields exactly ONE final Err(InactiveEffect) item, then closes. Ticks queued before the disposal are discarded, so teardown is always the final observation.
  • debounce keeps only the last value of a burst and delivers it after a quiet window. throttle delivers the first value immediately and, unless no_trailing is set, the last value of the window at close.

Scheduled<T> pairs a submit side (call) with a consumer side (receive, receive_timeout). Cancellation drops pending values; later receives return None.

LoggerService

LoggerService is a bounded ring of recent messages plus exporter fan-out. Provide it once on the root context:

#![allow(unused)]
fn main() {
ctx.provide(LoggerService::new());
}

The default capacity is 1000 messages. Use LoggerService::with_capacity to change it. The oldest message leaves first at capacity. snapshot returns detached clones, oldest first.

Write path

Every write follows four steps:

  1. Resolve the effective threshold for the logger name.
  2. Bail BEFORE argument assembly when the kind fails the gate.
  3. Append the record to the ring.
  4. Fan out to every exporter that accepts (name, kind).

Step 2 matters for cost. Prefer log_with with a closure so disabled paths never build arguments:

#![allow(unused)]
fn main() {
logger.log_with(&ctx, "db", LogKind::Debug, || {
    vec!["rows".into(), count.into()]
});
}

Convenience methods error, warn, info, and debug take pre-built arguments and still pass the gate. Facade methods on Context (ctx.info(..), and so on) no-op when no LoggerService is provided.

Levels

Severity ranks are numeric: Error=0, Warn=1, Info=2, Debug=3. Lower means more severe. A kind passes a threshold when its rank is less than or equal to the threshold value.

  • set_default_level pins the threshold for unlisted names. The default is DEBUG, which passes everything.
  • set_level(name, level) pins one name. It wins over the default.
  • clear_level(name) removes a pin.

Exporters

An exporter implements export(&self, message: &Message, text: &str). It runs inline on the writer's thread and must not panic. Registration takes an ExporterConfig with two fields:

  • levels: per-name thresholds for this sink. Unlisted names pass.
  • max_length: character cap on rendered text. Default 4096.

register returns a Box<dyn Disposable>. Disposing it removes the sink. The buffer keeps recording after a sink leaves.

Printf placeholders

When the leading argument is a string containing %, Message::render treats it as a format string:

SpecifierMeaning
%sString
%d, %iInteger
%fFloat
%oCompact JSON object
%OPretty JSON object
%cColorized with the stable palette slot for this name
%CBold colorized variant of %c
%%Literal percent

Unknown specifiers and exhausted arguments stay literal. Unconsumed arguments join at the end with spaces. Without a format head, arguments join with single spaces.

Stable color slots

%c and %C pick one of sixteen ANSI colors from the logger name. The slot must be stable across processes and platforms, so it comes from a hash, not a counter. name_color_code in logger.rs computes FNV-1a over the name bytes:

$$h_0 = \texttt{0xcbf29ce484222325}, \qquad h_{i+1} = (h_i \oplus b_i) \cdot \texttt{0x00000100000001b3} ;\bmod; 2^{64}$$

where \(b_i\) is the \(i\)-th name byte. The palette index is then

$$\text{color}(name) = \text{ANSI16}[,h \bmod 16,]$$

with ANSI16 = [30..=37, 90..=97]: eight normal foregrounds followed by their bright variants. The multiplication wraps (wrapping_mul), so no input can overflow-panic. The same name always renders in the same color — in tests, in production logs, and across restarts. Test anchor: colorization_is_stable_per_name.

LoggerIntercept override

LoggerIntercept rides the normal intercept channel. Install it with ctx.intercept(..). Writes through that context handle resolve it at write time:

#![allow(unused)]
fn main() {
let child = root.intercept(LoggerIntercept {
    name: Some("svc".into()),   // None matches every logger
    level: Some(LogLevel::ERROR),
});
}

level: Some(l) replaces the effective threshold for matching writes, over both pins and defaults. Names that do not match keep the ambient configuration.

Derived names

hyphenate turns CamelCase into kebab-case and handles acronym heads (HTTPServer becomes http-server). derived_name::<T>() applies it to the short type name. Use it for logger naming: ctx.info(&derived_name::<Self>(), ..).

Module Graph Transactional Reloads

The watcher fans file changes out to service-level dependents by TypeId. That layer cannot answer "which plugin must reload because this file changed?" because file edges carry no TypeId. ModuleGraph is that missing layer.

Callers register every dynamic module under a key, usually the watched file stem:

#![allow(unused)]
fn main() {
graph.register_module("foo", vec!["shared".into()], "FooPlugin");
}

Each entry carries its declared dependencies and the plugin that implements it.

Transaction shape

ModuleGraph::change_many(ctx, keys) runs one settled batch in two phases:

  1. Compute (read-only): walk the transitive dependent set across ALL input keys with a shared visited set. Cycles terminate. A plugin reachable from several inputs appears exactly once. If no input key matches a registered module, nothing is computed.
  2. Apply (sequential): reload each affected plugin through the ModuleReload seam in breadth-first propagation order. The FIRST failure rolls that plugin back to its previous state and stops the batch. Earlier successes stay active.

The classified result is a ChangeOutcome:

  • Ignored: no key matched a registered module. Nothing changed.
  • Reloaded(plugins): every affected plugin reloaded, deduped.
  • RolledBack { reloaded, failed_plugin, error }: names what applied, what failed, and the error text. The text also reports a rollback failure when the restore itself failed.

The default seam, NoopReload, never fails. Deployments wire their own reload / rollback pair, or swap one in later with set_reloader.

One transaction, narrated

Watch two shared files change at once: routes.toml and auth.toml. Three modules depend on them. foo depends on both; bar depends on shared; baz is independent.

  1. The watcher's debounce settles with both paths. Each path maps to its file stem, and the watcher hands ["routes", "auth"] to change_many.
  2. The compute phase walks the transitive dependent set across BOTH keys with one shared visited set. It reaches foo through either key but records it once — dedup happens during the walk, not after.
  3. The apply phase reloads affected plugins in breadth-first order: dependencies before their dependents, so each plugin reloads into a kernel where what it injects already exists.
  4. Suppose bar's reload fails on its new code. The seam rolls bar back to its previous state and the batch stops there. baz never ran — it matched no input key.

The result is RolledBack { reloaded: ["foo"], failed_plugin: "bar", error: .. }.

Sibling survival. Plugins that reloaded before the failure stay active on their NEW code. The batch does not unwind earlier successes. This guarantee shapes how you write reloaders:

  • A reload must leave the kernel consistent on its own. Earlier siblings will not be reverted for you.
  • Order failures so cheap ones fail first when possible; breadth-first order plus early failure minimizes divergence between old and new.
  • The rollback text reports a restore failure separately. Rolling back bar can itself error; the error field says so. Surface that case as an operator alert: the running state then matches no recorded state.

Contrast with the loader's two-phase reload (see Lifecycle): the loader rolls back everything newest-first; the module graph deliberately keeps successful siblings. The loader owns declarative trees. The module graph owns native-code hot swap, where a reloaded .so cannot always be unloaded again safely.

Watcher integration

When a debounced watcher batch settles, the watcher maps each changed path to its file stem and hands those stems to change_many — but only when a ModuleGraph is provided on the context. No graph registered means zero cost; the TypeId path stays unchanged. The HMR dynamic library fingerprint gate is untouched by this layer; neither consults the other.

File Fence Layers L0-L3

The tenant filesystem permission fence lives in crates/ares-tools/src/fence.rs. One Fence instance serves one session. Its policy value is pure and shareable; the observed-set ledger and audit ring sit behind a mutex.

Layers run in fixed order. A path passes only when every active layer passes:

  • L0 mode: FenceMode::ReadOnly denies every write. Reads still pass L1 and L2.
  • L1 boundary: the resolved path must stay inside workspace_root. FenceMode::Full waives this layer.
  • L2 blocklist: a blocked name denies reads and writes in every mode.
  • L3 write guards: session-level enforcement over the policy.

check_read and check_write on FencePolicy stay pure path checks (L0-L2). Only the Fence methods touch file contents.

Layer matrix

The matrix lists, for each layer, what it guarantees and which stable code reports its failure. Layers run top to bottom; the first failure decides the code.

LayerGuaranteeFails withApplies to
L0 modeReadOnly denies every writeFS_FENCE_DENIEDWrites only
L1 boundaryResolved path stays inside workspace_root (Full waives)FS_FENCE_DENIEDReads and writes
L2 blocklistBlocked names denied in every modeFS_FENCE_DENIEDReads and writes
L3 observationCanonical path observed before any guarded write in non-blind modes; missing paths record version 0FS_NOT_OBSERVEDWrites only
L3 contractGuard matches observed state: absent path for create, unchanged version for replaceFS_EXISTS, FS_VERSION_CONFLICTWrites only
L3 I/OAtomic sibling-temp-plus-rename writeFS_IOWrites only

Reading the table as an operator:

  • Three different denials all report FS_FENCE_DENIED; the audit ring entry records the reason text that separates them.
  • FS_NOT_OBSERVED is a protocol error, not a permission error. The agent forgot to read before writing. A read of a missing path counts, so creating a new file needs one prior failed-or-absent read.
  • FS_VERSION_CONFLICT means someone changed the file after your read. Re-read and re-apply the edit.
  • FS_IO covers everything underneath the policy: permission bits at the OS level, full disks, vanished parents. The reason text carries the OS message.

Determinism is the point. The same path, mode, guard, and observed state always produce the same code. Agent-facing retry logic branches on codes, not on parsed prose.

L3 write guards

Every write names a guard contract (WriteGuard):

  • Unconditional: overwrite whatever is there.
  • CreateIfAbsent: fails with FS_EXISTS when the path already exists.
  • ReplaceIfVersion { version }: fails with FS_VERSION_CONFLICT when the file is gone or changed since observation.

In modes without blind-write allowance, the canonical path must have been observed through Fence::fence_read first. Otherwise the write fails with FS_NOT_OBSERVED. This covers every contract, including creating new files. A read records a version fingerprint; a missing path records version 0, so a later create can prove absence.

Writes land through a sibling temporary file and an atomic rename, so an interrupted write leaves no torn file behind. A successful write becomes the new observed version, so chained guarded writes work against your own output.

Errors carry stable FS_* codes: FS_NOT_OBSERVED, FS_VERSION_CONFLICT, FS_EXISTS, FS_FENCE_DENIED, and FS_IO. The first failing layer determines the code, so agent-facing errors stay deterministic. Every operation lands in a bounded audit ring (audit_log()); the oldest entry leaves at capacity 200.

Kernel Patterns in Rust

This chapter is a cookbook for the Cordis kernel application programming interface (API). Every snippet is a light adaptation of a real in-tree test. The citation names the test and file, so you can read the full context there.

Provide and Consume Typed Services

Provide a service on a context; read it back typed with get. Child contexts inherit through the parent walk. intercept creates a child where one type resolves to an override; the parent stays untouched.

#![allow(unused)]
fn main() {
use cordis::{Context, Service};

struct Greeting(String);
impl Service for Greeting {}

let root = Context::new_root();
root.provide(Greeting("hello".into()));
assert_eq!(root.get::<Greeting>().unwrap().0, "hello");

// Per-request override: innermost intercept wins.
let req = root.intercept(Greeting("override".into()));
assert_eq!(req.get::<Greeting>().unwrap().0, "override");
assert_eq!(root.get::<Greeting>().unwrap().0, "hello");
}

Source: isolate_and_intercept in crates/cordis/src/lib.rs.

Listeners: on_with, once_with, and emit_filtered

on and once delegate to on_with / once_with with default options. The explicit forms take EventOptions: prepend: true inserts at the front of the dispatch order, global: true exempts the listener from context filters.

#![allow(unused)]
fn main() {
use cordis::events::{EventOptions, EventsService};
use std::sync::Arc;

let svc = EventsService::new();
let order = Arc::new(parking_lot::Mutex::<Vec<String>>::new(Vec::new()));

for name in ["first", "second"] {
    let slot = order.clone();
    svc.on("prepend.test".into(), move |_p| {
        let slot = slot.clone();
        async move {
            slot.lock().push(name.to_string());
            Ok(serde_json::Value::Null)
        }
    });
}

// Prepend lands in front of both default registrations.
let prepended = order.clone();
svc.on_with(
    "prepend.test".into(),
    EventOptions { prepend: true, global: false },
    move |_p| {
        let slot = prepended.clone();
        async move {
            slot.lock().push("prepended".to_string());
            Ok(serde_json::Value::Null)
        }
    },
);
}

Sources: prepend_ordering_observed and once_with usage in crates/cordis/src/events.rs.

emit_filtered dispatches to listeners whose registration options pass the filter. Exclusion is per dispatch: nobody is unregistered, and a later unfiltered emit runs everyone again. Global listeners bypass the filter entirely.

#![allow(unused)]
fn main() {
use std::sync::atomic::{AtomicUsize, Ordering};

let ran_a = Arc::new(AtomicUsize::new(0));
let a = ran_a.clone();
svc.on_with("filtered.test".into(), EventOptions::default(), move |p| {
    let a = a.clone();
    async move {
        a.fetch_add(1, Ordering::SeqCst);
        Ok(p)
    }
});

// A rejecting filter excludes non-global listeners for THIS dispatch.
svc.emit_filtered(
    "filtered.test".into(),
    serde_json::json!({ "tenant": "a" }),
    Box::new(|_opts| false),
)
.unwrap();

// An unfiltered dispatch runs both listeners again.
svc.dispatch("filtered.test".into(), serde_json::json!({}), cordis::Dispatch::Emit)
    .await
    .unwrap();
}

Source: filter_excludes_nonmatching_contexts in crates/cordis/src/events.rs.

Interceptor on internal/set

The kernel exposes six meta-events as veto points. internal/set runs before every service write. No listener means zero cost. Null or pass-through allows the write. A chain error vetoes the write and the previous value stays.

#![allow(unused)]
fn main() {
use cordis::events::{EventsService, INTERNAL_SET_EVENT};
use cordis::CordisError;

let svc = EventsService::new();
assert!(svc.intercept_set("Svc", None).await.is_ok());

// Freeze writes while this gate lives.
let d = svc.on(INTERNAL_SET_EVENT.into(), |_payload| async move {
    Err::<serde_json::Value, CordisError>(CordisError::Configuration(
        "writes are frozen".into(),
    ))
});
assert!(svc.intercept_set("Svc", None).await.is_err());

d.dispose();
assert!(svc.intercept_set("Svc", None).await.is_ok());
}

Source: set_interceptor_vetoes_write_leaves_old_value in crates/cordis/src/events.rs.

Readiness Gates: register_with_readiness and .watching

A readiness barrier holds a fiber out of service while it reports false. The fiber rests in an inspectable Pending state; this is quiet waiting, not failure. .watching(..) declares which service types re-kick the gate when their providers settle, so external provides and withdrawals re-evaluate it without polling.

#![allow(unused)]
fn main() {
use cordis::{Context, RegistryService, Service};
use cordis::registry::ReadinessBarrier;
use std::any::TypeId;

struct Dependency;
impl Service for Dependency {}

struct ConsumerPlugin;
impl cordis::Plugin for ConsumerPlugin {
    type Config = ();
    type Provides = Dependency;
    fn apply(
        &self,
        _ctx: &std::sync::Arc<Context>,
        _cfg: (),
    ) -> Result<std::sync::Arc<Dependency>, CordisError> {
        Ok(std::sync::Arc::new(Dependency))
    }
}

let ctx = Context::new_root();
let registry = RegistryService::new();

// Gate observes a plain context fact: is Dependency provided?
let fid = registry
    .register_with_readiness(
        &ctx,
        ConsumerPlugin,
        (),
        ReadinessBarrier::new(|ctx: &std::sync::Arc<Context>| {
            ctx.get::<Dependency>().is_some()
        })
        .watching([TypeId::of::<Dependency>()]),
    )
    .expect("fact-gated registration");
}

The factory still runs once at registration, so configuration errors surface immediately. Strict ctx.get refuses values owned by non-Active fibers, so consumers never see a gated service early.

Source: the fact-gated leg of the readiness tests in crates/cordis/src/registry.rs (barrier over ctx.get::<Dependency>() plus .watching([TypeId::of::<Dependency>()])).

Accessors and Aliases

An accessor is a name-keyed computed property. Reads and writes bypass the internal/get / internal/set waterfalls by design. Disposing the handle removes the declaration AND every alias bound to it.

#![allow(unused)]
fn main() {
use cordis::{Accessor, Context, CordisError};
use parking_lot::Mutex;
use std::any::Any;
use std::sync::Arc;

#[derive(Debug, PartialEq)]
struct PropValue(pub u64);

let ctx = Context::new_root();
let cell = Arc::new(Mutex::new(5u64));
let read_cell = cell.clone();
let write_cell = cell.clone();

let handle = ctx
    .register_accessor(
        "primary",
        Accessor::read_write(
            move |_ctx| {
                Ok(Some(Arc::new(PropValue(*read_cell.lock()))
                    as Arc<dyn Any + Send + Sync>))
            },
            move |_ctx, value: Arc<dyn Any + Send + Sync>| {
                *write_cell.lock() =
                    value.downcast::<PropValue>().unwrap().0;
                Ok(())
            },
        ),
    )
    .unwrap();

// Bind an alternate name resolving through the SAME registration.
ctx.alias("nick", "primary").expect("alias binds");
ctx.write_property("nick", Arc::new(PropValue(6))).unwrap();
assert_eq!(*cell.lock(), 6);

// Disposal removes both names at once.
assert!(handle.dispose());
assert!(ctx.read_property("nick").unwrap().is_none());
}

Collisions and unknown alias targets return errors: DuplicateProvider and ServiceNotFound. A typed read that fails to downcast returns PropertyTypeMismatch, never a silent None.

Source: alias_resolves_same_value and accessor_read_write_roundtrip in crates/cordis/src/context.rs.

Programmatic Tree Moves: Loader::move_entry

Loader::move_entry relocates a subtree in the live entry tree and makes the running kernel agree with it. Validation happens first; a refusal leaves the tree untouched. Renamed descendants follow the {parent}: id convention (svc under grp becomes grp:svc).

For a pure structural move — same plugins, configs, disabled flags, and isolates on both sides — the contexts-equivalence gate takes the noop path. Every journaled record re-keys old to new while KEEPING its fiber id. Consumers keep resolving the same live instances; nothing disposes or re-creates.

#![allow(unused)]
fn main() {
use cordis::{Context, EntryTree, Loader};

// current: tree loaded from config; journal: the LoaderJournal service.
// Both come from the normal loader bootstrap.
let outcome = Loader::move_entry(&ctx, &mut current, &journal, "svc", Some("grp"), 0)
    .await
    .expect("move succeeds");

assert!(outcome.noop, "pure structural move takes the noop path");
assert_eq!(
    outcome.renamed,
    vec![("svc".to_string(), "grp:svc".to_string())]
);
}

Refusals include unknown ids, moving an entry under itself, and moves under its own descendant. When mixed edits rode along so composition differs, the call falls back to a full reconcile apply instead of the noop path.

Source: the structural-move test in crates/cordis/src/loader.rs (Loader::move_entry with out.noop asserted against a two-entry tree).

Consuming interval_stream

Poll the stream yourself; ticks queue while nobody polls. After the owning fiber disposes, the stream yields exactly ONE final Err(InactiveEffect), then terminates. Live ticks queued before the disposal are discarded.

#![allow(unused)]
fn main() {
use cordis::timer::{interval_stream, with_current_fiber, InactiveEffect};
use cordis::{Fiber, timer::Stream};
use std::pin::Pin;
use std::sync::Arc;
use std::task::Poll;
use std::time::{Duration, Instant};

let fiber = Arc::new(Fiber::new());
let mut stream =
    with_current_fiber(&fiber, || interval_stream(Duration::from_millis(10)));

// Collect two live ticks (Ok items).
let mut live_ticks = 0u32;
while live_ticks < 2 {
    // poll_stream_once wraps Stream::poll_next with a noop waker.
    match poll_stream_once(&mut stream) {
        Some(Ok(())) => live_ticks += 1,
        Some(Err(_)) => unreachable!("not disposed yet"),
        None => std::thread::sleep(Duration::from_millis(2)),
    }
}

// Dispose through the owning fiber.
fiber.dispose().await.expect("dispose ok");

// Exactly ONE final error, then end-of-stream.
assert_eq!(poll_stream_once(&mut stream), Some(Err(InactiveEffect)));
assert_eq!(poll_stream_once(&mut stream), None);
}

Where poll_stream_once is the small helper from the source test:

#![allow(unused)]
fn main() {
fn poll_stream_once(stream: &mut cordis::timer::Interval) 
    -> Option<cordis::timer::TickResult> 
{
    let waker = std::task::Waker::noop();
    let mut cx = std::task::Context::from_waker(&waker);
    match Stream::poll_next(Pin::new(stream), &mut cx) {
        Poll::Ready(item) => item,
        Poll::Pending => None,
    }
}
}

Dropping the Interval also stops scheduling, so ownership without a fiber stays safe.

Source: interval_stream_final_err_on_dispose and interval_stream_discards_stale_live_ticks_before_final_err in crates/cordis/src/timer.rs.

emit_filtered with a Global Bypass

Register an audit listener with global: true when it must observe every dispatch, even filtered ones. The filter excludes only non-global listeners; the global one always runs. Exclusion is per dispatch: nobody is unregistered, and a later unfiltered emit runs everyone again.

#![allow(unused)]
fn main() {
use cordis::events::{EventOptions, EventsService};
use std::sync::Arc;
use std::sync::atomic::{AtomicUsize, Ordering};

let svc = EventsService::new();
let ran = Arc::new(AtomicUsize::new(0));

// Non-global listener for tenant "b".
let b = ran.clone();
svc.on_with("global.test".into(), EventOptions::default(), move |payload| {
    let b = b.clone();
    async move {
        if payload["tenant"] == "b" {
            b.fetch_add(1, Ordering::SeqCst);
        }
        Ok(payload)
    }
});

// Global listener: exempt from every filter verdict.
let g = ran.clone();
svc.on_with(
    "global.test".into(),
    EventOptions { prepend: false, global: true },
    move |payload| {
        let g = g.clone();
        async move {
            if payload["tenant"] == "b" {
                g.fetch_add(10, Ordering::SeqCst);
            }
            Ok(payload)
        }
    },
);

// A filter that admits NOTHING still lets the global listener run.
svc.emit_filtered(
    "global.test".into(),
    serde_json::json!({ "tenant": "b" }),
    Box::new(|_opts| false),
)
.unwrap();
assert_eq!(ran.load(Ordering::SeqCst), 10);
}

Use this for security audit trails, metrics, and tracing sinks. Anything that must not miss an event rides global: true.

Source: global_bypasses_filter in crates/cordis/src/events.rs.

Inspecting Values Mid-Transition with get_relaxed

Strict Context::get refuses values owned by fibers resting in Loading, Reloading, Unloading, or reactive Pending. Lifecycle and observer code needs exactly those values. get_relaxed resolves a locally-owned value while its owner transitions; terminal Failed and disposed owners stay refused in relaxed mode too.

#![allow(unused)]
fn main() {
use cordis::{Context, Fiber, Service};
use std::sync::Arc;

struct TransitionProbe(u32);
impl Service for TransitionProbe {}

let ctx = Context::new_root();
let fiber = Arc::new(Fiber::new());
fiber.set_reload_context(&ctx);
fiber.set_id(96_001);

// Provide ON the registration fiber so the owner link exists.
ctx.provide_on_fiber(Arc::new(TransitionProbe(7)), &fiber);

// Strict get refuses non-Active owners...
assert!(ctx.get::<TransitionProbe>().is_none());

// ...but relaxed reads serve the transitioning value itself.
fiber.set_state(cordis::FiberState::Pending);
assert_eq!(ctx.get_relaxed::<TransitionProbe>().unwrap().0, 7);
}

The setup calls set_reload_context, set_id, provide_on_fiber, and set_state are crate-internal. The source test uses them to place the owner fiber into each transitioning state directly. Product code reaches those states through a readiness gate or a dependency loss instead; only get_relaxed is public API.

Reach for this in diagnostics endpoints, state inspectors, and tests — never in ordinary consumers. Consumers keep strict get so they never observe half-torn configurations.

Source: relaxed_read_succeeds_while_provider_transitioning in crates/cordis/src/context.rs.

Inspecting a Deferred Config After an Update Veto

An internal/update listener returning JSON false vetoes the restart. The proposed config parks in Fiber::vetoed_config, the fiber stays Active on its old application, and update returns Ok(()). Only explicit false is a veto; any other non-null value proceeds. Operators can read what was deferred and apply it later inside the window.

#![allow(unused)]
fn main() {
use cordis::{Context, EventsService, Fiber};
use std::sync::Arc;

let ctx = Context::new_root();
let events = Arc::new(EventsService::new());
ctx.provide_arc(events.clone());
let fiber = Arc::new(Fiber::new());
fiber.set_reload_context(&ctx);   // crate-internal, see note below
fiber.set_id(70_300);             // crate-internal
// ... install a reload runner and satisfy its declared injects ...

fiber.set_raw_config(serde_json::json!({ "deferred": true })); // crate-internal

// An explicit JSON `false` bail verdict IS the veto.
let gate = events.on(
    cordis::events::INTERNAL_UPDATE_EVENT.into(),
    |_p| async move { Ok(serde_json::json!(false)) },
);
fiber.update(&ctx).await.expect("veto is Ok, not an error");
gate.dispose();

assert!(matches!(fiber.state(), cordis::fiber::FiberState::Active { .. }));
assert_eq!(
    fiber.vetoed_config(),
    Some(serde_json::json!({ "deferred": true })),
);
}

The setup calls set_reload_context, set_id, and set_raw_config are crate-internal. The source test uses them to stage a minimal fiber. Product code gets the same state from a normal RegistryService registration; only the veto listener, Fiber::update, and Fiber::vetoed_config are public surface.

Pair the gate with a maintenance-window check. During the window return null (proceed); outside it return false (defer). A later update call with a fresh proposed config overwrites vetoed_config.

Source: update_veto_defers_config_and_returns_ok in crates/cordis/src/fiber.rs.

Per-Fiber Log Level Override via LoggerIntercept

Install a LoggerIntercept on a child context to quiet one noisy logger for one subtree. Writes through that context handle resolve the override at write time; writes through other handles keep the ambient configuration. name: None matches every logger.

#![allow(unused)]
fn main() {
use cordis::logger::{LogLevel, LoggerIntercept, LoggerService};

let root = Context::new_root();
root.provide(LoggerService::new());

// Fiber-scoped override: only "svc" drops to Error-only on the child.
let child = root.intercept(LoggerIntercept {
    name: Some("svc".into()),
    level: Some(LogLevel::ERROR),
});
child.debug("svc", vec!["suppressed".into()]);
child.error("svc", vec!["survives".into()]);
// Non-matching names keep the ambient configuration.
child.debug("other", vec!["other-passes".into()]);

// Wildcard intercept: name=None forces the level for every logger.
let wild = root.intercept(LoggerIntercept {
    name: None,
    level: Some(LogLevel::ERROR),
});
wild.info("anything", vec!["blocked".into()]);
}

level: Some(l) replaces the effective threshold for matching writes, over both per-name pins and the default. Stack two intercepts and the innermost matching layer wins, like every layered override. Use the scoped form for request-scoped suppression; use the wildcard form for a temporary global mute during a hot path benchmark.

Source: logger_intercept_overrides_level in crates/cordis/src/logger.rs.

Agents

This chapter describes the agent configuration model in ARES v0.10.0. Every field here is read from the source files named in each section.

Agent behavior comes from TOML configuration. The root file ares.toml carries an [agents] table. You can also keep agents in TOON files under config/agents/. The config.agents_dir key in ares.toml sets that directory (crates/ares-http/src/overlay.rs, DynamicConfigPaths).

Agent fields

Each agent entry deserializes into AgentConfig (crates/ares-agent/src/config.rs):

FieldTypeDefaultMeaning
modelstringrequiredModel name defined under [models]. Missing values fail deserialization.
system_promptstringnonePersonality and instructions for the agent, per the source doc comment.
toolslist of stringsemptyTool names this agent uses.
allowed_toolslist of stringsall toolsWhitelist of permitted tool names. Absent means all tools are allowed.
max_tool_iterationsinteger10Maximum tool-calling rounds before the agent stops.
parallel_toolsbooleanfalseRun independent tool calls in parallel when possible.
compaction_enabledbooleanfalseTurn on per-session history compaction through the LLM Compactor. Long conversations stay a bounded working set instead of a last-5 history slice.
(extra keys)tableemptyUnknown keys pass through unchanged via #[serde(flatten)].

Deserialization details worth knowing when you write TOML:

  • Only model lacks a serde default. Every other field tolerates absence.
  • allowed_tools carries skip_serializing_if = "Option::is_none". An absent key stays absent on round-trip, so "all tools" survives config rewrites.
  • compaction_enabled is an Option<bool>. The engine reads unwrap_or(false), so omitting the key equals false.
  • Extra keys land in extra: HashMap<String, toml::Value> (#[serde(flatten)]). They never fail parsing. Runtime consumers read them by name.

Agent resolution order

The resolver (crates/ares-agent/src/resolver.rs) picks one agent definition from three tiers:

  1. Tenant database row (tenant_db tenant_agents).
  2. Community public agent.
  3. System agent from static TOML/TOON configuration.

The first tier that holds the name wins. The scope comes from the request isolate label, then a tenant context intercept, then a fallback id.

Skills configuration

The [skills] group deserializes into SkillsTomlConfig (crates/ares-agent/src/workflows_config.rs):

  • project_dir — project skills directory, for example ./.claude/skills/.
  • personal_dir — personal skills directory, for example ~/.claude/skills/.
  • plugin_dirs — extra directories to scan for SKILL.md files.

All three keys are optional.

Workflows

A workflow defines how agents work together. Each entry under [workflows.<name>] deserializes into WorkflowConfig (crates/ares-agent/src/workflows_config.rs):

FieldTypeDefaultMeaning
entry_agentstringrequiredAgent that receives the initial request.
fallback_agentstringnoneAgent used when routing fails or no match exists.
max_depthinteger3Maximum depth for nested workflows.
max_iterationsinteger5Maximum iterations for iterative workflows.
parallel_subagentsbooleanfalseExecute sub-agent calls in parallel.

The engine (crates/ares-agent/src/workflows/engine.rs) runs steps through the unified executor and records a WorkflowStep per hop: agent_name, input, output, timestamp, duration_ms. The result is a WorkflowOutput with final_response, steps_executed, agents_used, and reasoning_path. Workflows live on the context as a WorkflowSet; the HTTP overlay copies [workflows] from TOML into it.

Delegation flags

Skill delegation arguments parse into DelegationArgs (crates/ares-agent/src/skills/mod.rs). Parsing runs in two steps: tokenizing, then flag assignment.

Tokenizer rules

tokenize_delegation_args scans characters once:

  1. Whitespace outside quotes splits tokens. Empty tokens never appear.
  2. A double-quoted segment forms one token. It may contain spaces and bare |.
  3. Inside quotes only, a backslash escapes the next character. Outside quotes a backslash is an ordinary character.
  4. Quotes themselves disappear from the output token. "delta echo" becomes delta echo.

Flag assignment

parse_delegation_args walks the token list left to right:

  • --parallel latches split-per-token mode. After the flag, every plain token starts its own task and | separators are ignored. Tokens before the flag stay one sequential task.
  • --model <value> consumes exactly one following token as the model override. A missing value returns the error --model consumes exactly one value; none given.
  • --tools enables the inner tool loop for delegated tasks. It takes no value.
  • Before the latch, bare | separates sequential tasks. After the latch, bare | is dropped.

Parallel task cap

MAX_PARALLEL_DELEGATED_TASKS equals MAX_SKILL_CALL_DEPTH, which is 8 (crates/ares-agent/src/skills/mod.rs). Each parallel task becomes its own delegated skill call one level deeper, so the recursion depth is the natural ceiling. Excess tasks are truncated, not rejected.

Model resolution precedence

DelegationArgs::resolved_model(profile_default, global_default) applies this order:

  1. Explicit --model flag.
  2. Profile default model, when non-empty.
  3. Global default model.

An empty profile default falls through to the global default.

Worked examples

Both examples come verbatim from the test suite (delegation_flag_tests, same file).

Example 1 — quoted pipe inside one sequential task:

translate "hello | world" --model fast-model --tools

Result: one task with tokens ["translate", "hello | world"], model = "fast-model", tools = true, parallel = false. The quoted pipe does not split the task because quoting happens at tokenize time, before any | handling.

Example 2 — latch mid-stream:

alpha | beta --parallel gamma "delta echo" | epsilon

Result: five tasks — alpha, beta, gamma, delta echo, epsilon. alpha and beta split on the first pipe. After --parallel latches, gamma starts a task, the quoted "delta echo" forms its own task despite the space, and the trailing | is ignored.

Review gate for delegated results

SkillsPluginConfig.review_delegated_results (crates/ares-agent/src/skills/mod.rs) opts nested skill results into a review micro-step before integration:

  • Off by default. A missing or failing reviewer passes the original result through.
  • On acceptance, the original result integrates unchanged.
  • On rejection, the parent receives a structured rejection with review notes; the original stays intact in metadata for re-dispatch.

Verdict protocol

The reviewer prompt uses a fixed preamble (DELEGATED_REVIEW_TEMPLATE, same file). The preamble asks the reviewer to judge consistency between result and requested input. It requires the verdict word on the first line, followed by one sentence of notes. The preamble is byte-stable across calls so provider-side prompt caches hit on it.

parse_review_verdict applies these rules:

  1. Take the first non-empty line of the reply.
  2. Skip leading non-alphabetic characters, then read the first alphabetic run as the verdict word.
  3. Accept exactly ACCEPT or REJECT. Anything else is a parse error.
  4. Notes are the rest of the first line plus any further lines.

An empty reply or an unparsable verdict returns an error. Callers treat that error like a reviewer outage: the original result integrates unchanged. The gate never guesses.

Rejection payload shape

On rejection the parent receives this JSON structure instead of the original result:

{
  "status": "rejected",
  "review": { "accepted": false, "notes": "<reviewer text>" },
  "metadata": {
    "delegated_skill_id": "<nested skill id>",
    "original_result": { "...": "unchanged original" }
  }
}

The gate snapshots the delegated input before running the sub-execution. With the gate off, no snapshot allocation happens.

Without the postgres feature there is no reviewer, so the gate is permanently off and every result passes through.

Self-check rounds

SkillEngine::with_self_check_rounds(n) (crates/ares-agent/src/skills/engine.rs) adds critique rounds over each nested skill result before integration. One LLM call runs per round. Off by default; zero behaves identically to off.

Round flow

Each round follows this sequence:

  1. Extract the answer text. Structured results contribute their content field; bare strings contribute themselves. Results with neither carry nothing checkable, so the loop skips them.
  2. Build the prompt from a byte-stable template (SELF_CHECK_TEMPLATE, same file). The template asks two questions: did the answer address the requested task, and are there obvious errors or omissions. The tail names the delegated skill id, the requested input, and the current answer.
  3. Send one LLM call.
  4. On success, compare the reply with the current answer. A verbatim reply means no corrections were found, so the loop exits early.
  5. On error, stop the loop. The last good answer stays in place. No retry happens inside the loop.
  6. After the rounds end, fold the answer back. Structured results keep their shape; only content moves. Bare strings are replaced whole. An unchanged answer leaves the result untouched.
flowchart TD
    A[Nested skill result] --> B{Self-check on?}
    B -- off --> H[Integrate result]
    B -- on --> C[Extract answer]
    C -- none found --> H
    C --> D[Round: LLM call over stable template]
    D -- error --> G[Keep last good answer]
    G --> F
    D -- verbatim reply --> F[Fold answer back]
    D -- corrected reply --> E{Rounds left?}
    E -- yes --> D
    E -- no --> F
    F --> H

The template prefix stays identical across rounds, so provider-side prompt caches hit on every round after the first.

Delegated sub-workflows accept only allowlisted step kinds. Nested tool rounds hard-cap at three (check_tool_round_cap). Anti-recursion blocks delegation inside a delegated sub-workflow (validate_delegated_step). Slash-command chatter lines are stripped from delegated text before it enters the parent context (sanitize_result_value).

Ambient enrichment

AmbientEnrichmentConfig.enabled (crates/ares-agent/src/skills/engine.rs) is off by default. When enabled, each LLM step completion fires two parallel micro calls after the answer: intent classification and keyword tagging. Outcomes attach as session metadata on the skill-step record under the ambient_enrichment key. Enrichment input truncates at 4000 characters. Failures log at debug level and never delay or fail the completion.

Emergency stop

EmergencyStop (crates/ares-agent/src/emergency_stop.rs) is a Cordis service holding one atomic boolean. set_active(true) turns the kill switch on; set_active(false) clears it. Startup provides a fresh instance set to inactive when none exists yet (crates/ares-agent/src/plugins.rs).

Propagation path

The switch reaches execution through the shared context:

  1. Plugins provide EmergencyStop on the root context.
  2. Skill execution resolves it with ctx.get::<EmergencyStop>().
  3. ensure_execution_active (crates/ares-agent/src/skills/engine.rs) checks the flag before every nested skill call. The check runs at call boundaries, so a stop request lands at the next boundary without cancelling in-flight work mid-call.
  4. An active flag aborts with the stable marker subtask_cancelled prefixed to the error text, so callers classify aborts without parsing prose.
  5. The aborted subtask integrates nothing into the parent context.

Per-subtask cancel tokens work alongside the global switch. Tokens are sticky cancel flags keyed "{run_id}/{skill_id}"; registration is idempotent and there is no un-cancel. Execution reads the registry fresh at every boundary, so an external trigger racing a starting subtask still stops it at its first call boundary. Token aborts reuse the same subtask_cancelled marker.

Worked example

Every key below appears in the structures cited above. The file passes ares-server config --validate on the v0.10 tree:

[server]
host = "127.0.0.1"
port = 3000
log_level = "info"

# Required tables; every field inside defaults, so empty tables work.
[auth]
jwt_secret_env = "JWT_SECRET"
api_key_env = "API_KEY"

[database]

[providers.local]
type = "ollama"
base_url = "http://localhost:11434"
default_model = "llama3.1:8b"

[models.default]
provider = "local"
model = "llama3.1:8b"
temperature = 0.7
max_tokens = 512

[agents.helper]
model = "default"
system_prompt = "You answer questions about internal documents."
tools = ["calculator"]
allowed_tools = ["calculator"]
max_tool_iterations = 10
parallel_tools = false
compaction_enabled = true

[workflows.support]
entry_agent = "helper"
fallback_agent = "helper"
max_depth = 3
max_iterations = 5
parallel_subagents = false

[skills]
project_dir = "./.claude/skills/"

Validation output, captured from a real run (exercised):

$ ares-server config --validate
  ✓ Configuration is valid!

  Configuration Summary

    Config file: ares.toml
    Server: 127.0.0.1:3000
    Log level: info

  Providers
    • local

  Models
    • default

  Agents
    • helper

Omitting either [auth] or [database] fails with missing field 'auth' or missing field 'database'. The tables themselves are required even though every field inside them carries serde defaults.

Retrieval Augmented Generation

ARES ships a RAG (Retrieval Augmented Generation) system: document ingestion, chunking, embeddings, and search. This chapter describes only behavior verified in the source.

Enabling RAG

The [rag] group in ares.toml deserializes into RagConfig (crates/ares-rag/src/config.rs). It has four sub-groups: vector, chunking, search, and rerank. Every field defaults, so an empty [rag] table is valid.

Vector group ([rag.vector])

FieldDefaultMeaning
enabledfalseMaster switch for the RAG feature.
embedding_model"bge-small-en-v1.5"Local embedding model name.
sparse_embeddingsfalseEnable sparse embeddings for hybrid search.
sparse_model"splade-pp-en-v1"Sparse model name.
vector_path"./data/vectors"Path for vector data on disk.

The default model emits 384-dimension vectors (BgeSmallEnV15, crates/ares-rag/src/embeddings.rs). Quantized variants of several models exist under separate names.

Chunking group ([rag.chunking])

FieldDefaultMeaning
chunking_strategy"word"One of "word", "semantic", "character".
chunk_size200Chunk size; words for word chunking, characters otherwise.
chunk_overlap50Overlap between consecutive chunks.
min_chunk_size20Smallest chunk size to keep, in characters.

The chunker (crates/ares-rag/src/chunker.rs) exposes TextChunker with word, semantic, and character modes. chunk_with_metadata returns chunks that carry offsets. Each Chunk records index, content, start_offset, and end_offset.

Strategy names accept aliases when parsed from strings (FromStr for ChunkingStrategy): words; sentence or paragraph for semantic; char or chars for character. Unknown names fail with an error naming the valid set.

Note one asymmetry. The standalone ChunkerConfig struct defaults chunk_size to 512. The HTTP ingest handler builds chunkers with its own fixed sizes per strategy (see below).

Ingestion paths

Two paths reach the same ingest handler:

Command line

ares-server rag ingest-dir walks a local directory and posts each supported UTF-8 text file (.md, .txt, .json, .jsonl) to the server (src/cli/rag.rs). Flags:

FlagMeaning
--hostServer base URL. Defaults to http://localhost:3000.
--collectionCollection name to ingest into. Required.
--docs-pathDirectory with documents. Required.
--user / --passwordLogin credentials. Used when --token is absent.
--tokenBearer token; skips login when present.
--chunking-strategyword, semantic, or character. Defaults to word.
--tagTag to attach. Repeat for multiple tags.
--dry-runList files without sending API requests.

The command fails closed: a bad path or a failed login stops before any network write. Each document reports its own audit trail, so a partial run leaves no half-ingested files unrecorded.

The following outputs come from running this binary against a real directory (exercised, ARES v0.10 tree):

$ ares-server rag ingest-dir --collection demo --docs-path /tmp/ragdemo --dry-run
/tmp/ragdemo/alpha.md	alpha	21 bytes
/tmp/ragdemo/beta.txt	beta	12 bytes
dry_run=true documents=2

$ ares-server rag ingest-dir --collection demo --docs-path /tmp/no-such-dir
Error: "docs path is not a directory: /tmp/no-such-dir"

The dry-run prints one path<TAB>title<TAB>size line per candidate file and a summary line. The bad path aborts with an error before any login or HTTP call.

HTTP

POST /api/rag/ingest accepts a JSON body with collection, content, title, source, tags, and chunking_strategy; it returns chunks_created plus document_ids and echoes the unscoped collection name (crates/ares-http/src/api/handlers/rag.rs, src/cli/rag.rs). The endpoint requires authentication and checks tenant-level RAG source allowlists before ingesting. Routes register only when both the local-embeddings and ares-vector features compile in (crates/ares-http/src/api/routes.rs).

Related endpoints: POST /api/rag/search, collection listing under /api/rag/collections, and DELETE /api/rag/collection.

Ingest handler stages

The handler (crates/ares-http/src/api/handlers/rag.rs, ingest) runs these stages in order:

  1. Validate input. An empty collection or empty content returns an error before any work starts.
  2. Check the tenant allowlist. TenantAllowlistStore::is_rag_source_allowed consults PostgreSQL. A denied collection fails with an auth error naming the collection.
  3. Scope the collection name. user_scoped_collection combines the user id and requested name, giving each user an isolated namespace.
  4. Resolve services. The embedding service comes from the context, or gets constructed on first use. Construction pre-downloads ONNX model files through lancor, because fastembed's own client fails on HuggingFace CDN redirects. The vector store opens at [rag.vector] vector_path.
  5. Parse the chunking strategy. Absent keys fall back to Word. The handler then fixes chunker parameters per strategy: word uses size 200 with overlap 50, semantic uses size 500 with no overlap, character uses size 500 with overlap 100.
  6. Chunk the content. Zero chunks fail with Content too small to chunk. Short trailing remainders below min_chunk_size drop out silently instead.
  7. Create the collection on demand. If the scoped collection does not exist, the handler creates it with the model's dimension count.
  8. Embed all chunk texts. The batch passes through the dedup planner described below.
  9. Build documents. Each chunk becomes a Document with id {base_uuid}_{index} and metadata carrying title, source, created_at, and tags.
  10. Upsert into the vector store and return chunks_created with the document id list. A structured log records user, collections, chunk count, and duration.
flowchart TD
    A[POST /api/rag/ingest] --> B{Auth valid?}
    B -- no --> X[401]
    B -- yes --> C{Collection allowlisted?}
    C -- no --> Y[Auth error]
    C -- yes --> D[Scope collection to user]
    D --> E[Chunk content by strategy]
    E -- zero chunks --> Z[Input error]
    E --> F[Embed chunks with dedup]
    F --> G[Upsert documents to vector store]
    G --> H[chunks_created + document_ids]

Chunking algorithms

Each mode behaves differently (crates/ares-rag/src/chunker.rs):

  • Word. Split text on whitespace. Walk forward by step = chunk_size - chunk_overlap, never less than 1. Join each window back with single spaces. Character offsets are approximated by summing word lengths plus one separator per earlier word, so offsets drift when original spacing differed.
  • Semantic. Hand the text to TextSplitter::new(chunk_size) from the text-splitter crate. It splits on sentence and paragraph boundaries up to the size cap. Overlap does not apply in this mode. Offsets come from searching each produced slice in the remaining text.
  • Character. Collect chars and walk by the same step formula as word mode. Offsets are exact char positions.

All three modes drop candidates shorter than min_chunk_size.

Embeddings and deduplication

Embeddings live in crates/ares-rag/src/embeddings.rs. Two facts matter:

  • Dedup per request: identical inputs collapse by a whitespace-normalized SHA-256 content hash (normalize_for_dedup plus content_hash_hex) before the backend call. Computed vectors fan back to every duplicate slot, so callers receive full-length results while identical texts cost one backend call.
  • Vectors are L2-normalized to unit length after computation (normalize_embedding).

Dedup walkthrough

Three functions cooperate:

  1. normalize_for_dedup(" hello world \n\t again ") returns "hello world again". Trimming and collapsing whitespace makes spacing differences invisible to the hash.
  2. content_hash_hex hashes the normalized text with SHA-256 and formats lowercase hex.
  3. DedupPlan::plan(texts) walks inputs once. First occurrences record their index in unique_indices; every duplicate records which unique slot it maps to in sources.

After the backend embeds only the unique texts, fan_out(&vectors) rebuilds the full-length list where result[i] == vectors[sources[i]]. The seen-set lives only for the duration of one call, so memory stays bounded. Both the local fastembed path and the HTTP batched path use the same plan (embed_texts, embed_texts_batched).

Worked example, given inputs ["alpha", " alpha\n", "beta"]:

  • Hashes: slot 0 covers alpha; the second entry normalizes to alpha too, so it maps to slot 0; beta takes slot 1.
  • Backend call receives ["alpha", "beta"] — two texts, not three.
  • Fan-out returns [v0, v0, v1].

Model initialization serializes through global per-model locks, so parallel first-use requests cannot race a model download.

Similarity math

Two dense vectors compare through cosine similarity (cosine_similarity, crates/ares-rag/src/embeddings.rs):

$$\cos(\mathbf{a},\mathbf{b}) ;=; \frac{\sum_i a_i b_i}{\sqrt{\sum_i a_i^2},\sqrt{\sum_i b_i^2}}$$

Mismatched lengths or a zero magnitude return 0.0, and the result clamps into \([-1, 1]\). When both vectors have unit length — which normalize_embedding produces — the denominator equals 1 and the whole expression reduces to the dot product:

$$|\mathbf{a}| = |\mathbf{b}| = 1 \quad\Longrightarrow\quad \cos(\mathbf{a},\mathbf{b}) = \sum_i a_i b_i$$

Cosine distance stored by pgvector-style comparisons is 1 - cos(a,b) (cosine_distance). At query time, raw distances convert back to similarity scores where higher means better (distance_to_similarity, crates/ares-rag/src/search.rs):

$$\text{sim}{\text{cosine}} = 1 - d \qquad \text{sim}{L2} = \frac{1}{1+d} \qquad \text{sim}_{\text{inner}} = -d$$

The local-embeddings feature enables the fastembed backend (crates/ares-rag/Cargo.toml). Without it, embedding calls go over HTTP.

Search flow

SearchStrategy (crates/ares-rag/src/search.rs) supports four strategies, also accepted by the CLI as --strategy:

StrategyAliasesMatches onIndex usedPersisted
semanticdense, vectorMeaning, via dense embeddingsVector store collectionYes (vector data)
bm25lexical, sparseExact terms, TF-IDF weightedInverted indexYes (bm25_index.json)
fuzzyapproximateTypos, via edit distanceVocabulary plus document mapYes (fuzzy_index.json)
hybridcombined, rrfUnion of all threeAll indicesVia component indices

Unknown names fail with an error naming the four valid sets.

Semantic scoring

The engine validates query embeddings first: non-empty, all finite, and dimension-matched when a dimension is known (validate_embedding). SQL orders by the metric operator and computes the score expression inline; collection names must pass validation before they can reach a table name ({prefix}_{collection}).

Result post-processing follows one shared path (rank_results): drop scores under the threshold, keep the highest-scoring copy of duplicate ids, sort descending, truncate to top_k.

BM25 scoring

Bm25Index keeps an inverted index and document frequencies. Tokenization lowercases, splits on non-alphanumeric characters, and drops tokens of length 1 or less. Term rarity enters through IDF (Bm25Index::idf):

$$\mathrm{idf}(t) ;=; \ln!\left(\frac{N - df_t + 0.5}{df_t + 0.5} ;+; 1\right)$$

Defaults follow standard BM25 saturation and length normalization: \(k_1 = 1.2\), \(b = 0.75\) (Bm25Index::new). Search collects candidate documents containing any query term, scores them, sorts descending, and truncates.

Fuzzy matching

FuzzyIndex stores a vocabulary of indexed words. Query correction finds the closest vocabulary word within max_distance edits (default 2, Levenshtein). Exact hits short-circuit at distance 0. correct_query rewrites each misspelled word and reports every correction as an original/corrected/distance triple. Match score per document is one minus normalized distance, averaged over matched query words.

Hybrid fusion

Hybrid search fuses three ranked lists through reciprocal rank fusion (RrfFusion::fuse). Raw scores are ignored; only positions count. With lists \(\ell\) and weights \(w_\ell\), a document's fused score is:

$$\mathrm{score}(d) ;=; \sum_{\ell} \frac{w_\ell}{k + \mathrm{rank}_\ell(d)} \qquad k = 60 \text{ by default}$$

Ranks start at 1 in this implementation (search.rs, RrfFusion::fuse). Default weights are semantic = 0.6, bm25 = 0.3, fuzzy = 0.1 (HybridWeights::default); the sum should be 1.0. Hybrid fetches top_k * 2 candidates from BM25 and fuzzy before fusion, then truncates the fused list to top_k. A typo-corrected variant corrects the query first, then reruns the whole fusion.

Worked fusion, with unit weights and \(k = 60\) (adapted from the passing test test_rrf_fusion_ranking_prefers_shared_top_ranks). The semantic list ranks doc_a first and doc_b second. The BM25 list ranks doc_b first and doc_c second:

$$\mathrm{score}(\text{doc_a}) = \tfrac{1}{61} = 0.0164$$ $$\mathrm{score}(\text{doc_b}) = \tfrac{1}{62} + \tfrac{1}{61} = 0.0325$$ $$\mathrm{score}(\text{doc_c}) = \tfrac{1}{62} = 0.0161$$

The document appearing in both lists wins even though it never held rank one anywhere.

The BM25 and fuzzy indices support save() and load(), so they survive restarts without re-indexing. SearchEngine::save writes bm25_index.json and fuzzy_index.json into one directory; load_or_new falls back to fresh empty indices when the directory is missing. ares-server rag search --collection <name> --query <text> [--top-k N] [--strategy S] drives the flow from the terminal.

Vector stores

Two backends exist behind features:

  • ares-vector (crates/ares-vector/) — a pure-Rust embedded store built on an HNSW index. Two construction modes differ only in persistence (Config, crates/ares-vector/src/config.rs):

    ModeConstructorBehavior
    MemoryConfig::memory()No data_path; data dies with the process.
    PersistentConfig::persistent(path)Sets data_path and turns on periodic snapshotting; collections reload on startup.

    Tuning knobs include max_vectors per collection and the HnswConfig fields (m, m_max, ef_construction, ef_search, thread counts). A memory_efficient() preset trades accuracy for footprint (m = 8, single-threaded construction). Collections are typed by dimension and distance metric at creation; searches reject wrong-dimension queries with DimensionMismatch. This backend backs the default [rag.vector] path.

  • Qdrant — optional external store configured through database.qdrant in ares.toml (QdrantConfig in crates/ares-store/src/config.rs). Fields: url (default http://localhost:6334) and api_key_env, the environment variable holding the API key. ARES treats it as an external service and holds no local state for it.

Operations

This chapter covers configuration, supervision, observability, security posture, CI gates, and backup notes for running an ARES server. All keys come from the config structs cited per section.

Configuration file

The server reads ares.toml from the working directory. ares-server --config my-config.toml selects another file. ares-server init scaffolds a project with ares.toml, .env.example, and a config/ directory tree (src/cli/init.rs). The root schema is AresConfig (crates/ares-http/src/overlay.rs). Check any config file before deploy with ares-server config --validate. The [auth] and [database] tables are required, but an empty table satisfies them because every field inside carries a default.

[server] group

ServerConfig (crates/ares-http/src/config.rs):

FieldDefaultMeaning
host"127.0.0.1"Bind address.
port3000Listen port.
log_level"info"One of trace, debug, info, warn, error.
cors_origins["http://localhost:3000"]Allowed CORS origins. Set explicit origins in production.
rate_limit_per_second100Requests per second per IP; 0 disables limiting.
rate_limit_burst10Rate limiter burst size.

CORS behavior

build_cors_layer (src/main.rs) maps the origin list to three modes:

  1. Single "*" entry, or an empty list: any origin is allowed, credentials are disabled, and a warning logs at startup. Browsers reject credentials together with wildcard origins, so this mode never sends them.
  2. Any other list: only listed origins pass. Credentials are enabled. Origins that fail to parse drop out silently.

Allowed methods cover GET, POST, PUT, PATCH, DELETE, and OPTIONS. Allowed request headers are Authorization, Content-Type, Accept, Origin, and x-admin-secret.

Rate limit algorithm

The two rate_limit_* keys drive one middleware built at startup (src/main.rs, around the layer assembly):

  • When rate_limit_per_second > 0, the app wraps in tower_governor::GovernorLayer. The governor runs the Generic Cell Rate Algorithm (GCRA) per client IP: per_second sets the sustained refill rate, and burst_size permits short bursts above it before rejections start.
  • use_headers() adds standard x-ratelimit-* headers to responses, so clients can observe their budget.
  • A background task prunes stale per-IP state every 60 seconds to bound memory.
  • Setting rate_limit_per_second = 0 skips the layer entirely. Startup then logs a warning that disabling limiting is not recommended for production.

A second, independent limiter lives in the API-key auth middleware: tenant daily usage accumulates in the PostgreSQL daily_rate_limits table keyed by (tenant_id, usage_date) and caches per tenant in memory (crates/ares-store/src/tenants.rs). If the database check itself fails, requests fail closed with HTTP 500 Failed to check rate limit.

[auth] group

AuthConfig:

FieldDefaultMeaning
jwt_secret_env"JWT_SECRET"Environment variable name holding the JWT (JSON Web Token) secret.
jwt_access_expiry900Access token lifetime in seconds (15 minutes). Short-lived by design; refresh tokens renew sessions.
jwt_refresh_expiry604800Refresh token lifetime in seconds (7 days).
api_key_env"API_KEY"Environment variable name holding the API key.

Secrets live in environment variables by name, never in ares.toml. Expiry values are plain second counts; there is no separate unit key.

[database] group

DatabaseConfig (crates/ares-store/src/config.rs):

FieldDefaultMeaning
urlpostgres://postgres:postgres@localhost:5432/aresPostgreSQL connection string. Holds tenants, agents, skills, run history, billing, and compaction snapshots.
qdrantnoneOptional QdrantConfig table for an external vector store. See RAG.

[providers.*] group

Each named provider deserializes into ProviderConfig (crates/ares-llm/src/config.rs), tagged by type = "...". Variants: openai (fields api_key_env, api_base, default_model; also serves NVIDIA NIM and compatible endpoints), azure (api_key_env, base_url_env, default_model), anthropic (api_key_env, default_model), bedrock (api_key_env, region_env, default_model), and ollama (base_url, default_model). A missing environment variable fails at client creation with a clear error.

An optional [nvidia] group (NvidiaConfig) adds catalog settings: api_key_env, api_base, models_url, catalog_refresh_seconds, and default_model. When absent, the registry synthesizes one NVIDIA provider from defaults.

[models.*] group

ModelConfig (crates/ares-llm/src/config.rs): provider (name under [providers]), model (identifier sent to the provider), temperature (default 0.7), max_tokens (default 512).

Provider pool

The client pool takes a PoolConfig (crates/ares-llm/src/pool.rs) with these fields:

  • max_in_flight — maximum simultaneous dispatches admitted per provider. Absent means unlimited and no governor installs.
  • governor_acquire_timeout — how long a dispatch waits for an in-flight slot before failing closed. Default 30 seconds.
  • max_connections_per_provider — default 10.
  • min_idle_connections — default 2.
  • idle_timeout — default 300 seconds.
  • max_lifetime — default 1800 seconds.
  • health_check_interval — default 60 seconds.
  • acquire_timeout — wait budget for borrowing a pooled client. Default 30 seconds.
  • enable_health_check — default true.

A permit spans the whole call including streams; saturation fails closed. This admission model rejects excess callers instead of queueing them onto the backend, so a saturated provider degrades loudly and early rather than silently stretching latencies.

Supervised operation runbook

Start the server with --supervise (src/main.rs, src/supervisor.rs). The daemon runs the real server as a child copy marked by the CORDIS_SUPERVISED environment variable. Dropping the child's standard-input handle is the stop request; the child watches for end-of-file and tears down gracefully.

Exit codes drive the loop:

Exit codeConstantEffect
51EXIT_RESTARTStart a fresh child. Rapid loops back off exponentially.
52EXIT_QUITEnd supervision; shut down for good.
53EXIT_BOOTBoot failed; report and do not restart.

Any other terminal status also ends the loop, including death by signal. Hot restarts use code 51 so configuration changes apply without dropping the daemon.

Restart loop guard rails

The loop carries four safeguards (src/supervisor.rs):

  • Rapid-restart cap. Five exits inside a 30-second window stop the loop with an error instead of spinning. A plugin that crashes at boot trips this cap.
  • Health reset. A child that ran at least 10 minutes counts as healthy. Its exit clears the strike ladder, so old crashes never doom a fresh process.
  • Backoff ladder. A child that exited within 10 seconds never proved health. The next respawn delays 100 ms, doubling per consecutive unhealthy run: 100 ms, 200 ms, 400 ms, 800 ms, 1.6 s, 3.2 s, capped at 5 s.
  • Shutdown grace. After a stop request, the child gets 10 seconds to exit on its own; past that, the daemon force-kills it. The grace bounds the goodbye, never the working lifetime.

Nested supervision refuses to start: a child marked CORDIS_SUPERVISED never spawns its own daemon. Exit code 53 mirrors the child's real code to the daemon's process exit, so systemd or another service manager still observes the failure.

flowchart TD
    A[Child exits] --> B{Code}
    B -- 51 RESTART --> C{5 restarts in 30 s?}
    C -- yes --> H[Stop loop with error]
    C -- no --> D{Run under 10 s?}
    D -- yes --> E[Delay backoff ladder]
    D -- no --> F[Reset ladder]
    E --> G[Spawn fresh child]
    F --> G
    B -- 52 QUIT --> I[Shut down daemon]
    B -- 53 BOOT --> J[Report failure, do not restart]
    B -- other / signal --> I

Observability

  • Health endpoints: /health (Http plugin) and /health/detailed (src/main.rs). Admin routes expose health metrics and model metrics under /health/list_health_metrics and /health/list_model_metrics (crates/ares-http/src/api/routes.rs). Observed responses from a running v0.10 server (exercised):

    $ curl -s http://localhost:3000/health
    OK
    
    $ curl -s http://localhost:3000/health/detailed
    {"status":"healthy","version":"0.1.0","checks":{},"agents":[],"latency_ms":1}
    

    /health answers plain text for cheap probes. /health/detailed returns JSON with per-check status, registered agents, and measured latency.

  • Telemetry records: every LLM call produces an LlmCallRecord that carries cached_tokens and total_time_ms alongside token counts (crates/ares-llm/src/observability.rs). Micro-call cache hits report latency_ms: 0 and carry a cache_hit flag.

Telemetry field semantics

Both new columns are optional integers (Option<i64>), and their absence carries meaning:

  • cached_tokens reports tokens served from the provider-side prompt cache. It is None when unknown or unreported. When present it is always zero or more and forms a subset of prompt_tokens.
  • total_time_ms measures end-to-end wall-clock time for the whole call, including retries and queueing. Callers commonly mirror latency_ms into it when they cannot measure the two separately. None means not measured.

Exporter routing

Log exporters fan records out through the ExporterRouter (crates/ares-llm/src/exporter.rs):

  • Each registered sink declares which records it accepts through a RecordLevel gate: Debug, Info, Warn, or Error. The level is metadata about the record; it does not change the record.
  • The built-in stdout formatter emits a tracing info event for successful calls and a warn event for failures. Both include cached_tokens and total_time_ms; an absent Option emits nothing rather than a placeholder.
  • Sinks include stdout formatters, database writers, OTLP (OpenTelemetry Protocol) forwarders, and test captures.
  • An exporter failure logs a warning inside the exporter and never fails inference.

There is no Prometheus endpoint; scrape-style monitoring must read the admin endpoints above. Skill-step records attach ambient enrichment metadata under the ambient_enrichment key when enabled (see Agents).

Security posture

ARES fails closed on its trust boundaries:

  • RAG ingestion and search require authentication and check tenant allowlists before touching data.
  • Delegation arguments, review gates, and tool allowlists restrict what agents may call; unknown tools stay blocked when allowed_tools names a set.
  • Provider governors cap concurrent dispatches and reject excess callers instead of queueing them onto the backend.
  • Secrets resolve from named environment variables at use time; config files hold only variable names.

Rotate credentials with this procedure:

  1. Add the new value under a fresh environment variable name.
  2. Update the matching *_env key in ares.toml.
  3. Restart or hot-restart (exit code 51) the supervised process.
  4. Remove the old environment variable after the new one proves active.

Never place secret values in ares.toml, TOON files, or version control.

Backup and restore runbook

Storage choice comes from [database] (DatabaseConfig, crates/ares-store/src/config.rs):

  • url — PostgreSQL connection string, default postgres://postgres:postgres@localhost:5432/ares. PostgreSQL holds tenants, agents, skills, run history, billing, and compaction snapshots.
  • qdrant — optional external vector store. Back up collections through Qdrant's own snapshot mechanism; ARES treats it as an external service.
  • The embedded ares-vector store persists under [rag.vector] vector_path (default ./data/vectors). Include that directory in file-level backups.

Backup procedure:

  1. Quiesce RAG ingestion first. Stop writes or pause ingest traffic so chunk files stay consistent during the copy.
  2. Dump PostgreSQL with standard tooling: pg_dump "$DATABASE_URL" > ares-backup.sql. Schedule dumps to match your recovery point objective.
  3. Copy the vector directory while writes are quiesced: cp -a ./data/vectors /backups/vectors-$(date +%F)/.
  4. For Qdrant deployments, take its snapshot through Qdrant's API or tooling. Do not copy its files behind its back.
  5. Verify each artifact restores before trusting the schedule: load the dump into a scratch database and open the copied vector directory read-only.

Restore procedure, in this order:

  1. Stop the ARES server.
  2. Restore PostgreSQL first: create the database, apply ares-backup.sql, then let migrations reconcile schema state on next boot.
  3. Restore the vector directory second, back to the exact [rag.vector] vector_path the restored config names.
  4. Restore Qdrant snapshots third, if used.
  5. Start the server so migrations run against the restored database before traffic arrives.

Order matters: vector data references documents whose metadata lives in PostgreSQL, so restoring the database first keeps ids consistent.

CI quality gates

The repository pins one automated quality gate in GitHub Actions (.github/workflows/ci.yml, job crap, named "CRAP Score Gate"):

cargo install cargo-crap
cargo crap --workspace --format json --threshold 30 --fail-above

CRAP (Change Risk Anti-Patterns) combines cyclomatic complexity with test coverage per function. The upstream metric definition (Savoia & Evans, 2007; implemented by cargo-crap) is:

$$\mathrm{CRAP}(m) ;=; \mathrm{comp}(m)^2 \times \left(1 - \frac{\mathrm{cov}(m)}{100}\right)^{3} ;+; \mathrm{comp}(m)$$

Here \(\mathrm{comp}(m)\) is the function's cyclomatic complexity and \(\mathrm{cov}(m)\) its line coverage percentage. Three properties explain the gate's shape:

  • A trivial, fully covered function scores exactly \(1\).
  • At full coverage the cubic term collapses, so \(\mathrm{CRAP}\) equals complexity. Tests cap the risk; they do not remove it.
  • Above complexity \(\approx 30\), no coverage level brings the score under the threshold of 30. Oversized functions fail regardless of tests.

Two concrete scores, taken from the upstream tool's own example table: a function with complexity 12 and zero coverage scores \(12^2 \times (1-0)^3 + 12 = 156\). A function with complexity 4 at roughly 44% coverage scores \(16 \times (1-0.444)^3 + 4 \approx 6.7\).

The gate fails when any workspace function exceeds the threshold of 30. Treat a red gate as work, not noise. Split the flagged function or raise its test coverage.

Cordis → rust mapping (Phase 0, step 5)

Source: DeepSeek Cordis paper ("A Programming Paradigm for Spatiotemporal Composability", Aug 2026, cordiverse/cordis + cordiverse/paper) and DeepSeek Harness (deepseek-ai/deepseek-harness, TS, ~60 packages, 12 layers). Target: ARES /opt/ares (dirmacs/ares v0.7.3, 11 crates + ares-server root, Rust 1.98, Tokio/Axum). This doc is strategy only, no code changes. Spike crate crates/cordis (Phase 1) must prove the theorems before adoption.

Phase 2 crate graph (crates.io already has cordis 0.0.0, so the Cargo package is ares-cordis with [lib] name = "cordis"; workspace dep key stays cordis):

PathPackageRust crate
crates/cordisares-cordiscordis
crates/ares-agentares-agentares_agent
crates/ares-storeares-storeares_store

Root package stays ares-server with lib name ares. Domain config lives in plugin crates; Overlay (src/overlay.rs) owns ares.toml after Phase 4.


1. core equation

Cordis: Γ^∞ = μΓ. Γ × (Γ → Γ) × Σ, unified context that lifts effect systems (revertible mutations) and coeffect systems (typed dependency declarations) to runtime.

Rust mapping:

#![allow(unused)]
fn main() {
pub struct Context {
    // Γ — value environment (store)
    store: RwLock<HashMap<TypeId, Arc<dyn Any + Send + Sync>>>, // Σ impl, see §3
    // Isolate table: TypeId → Symbol (scoped identity)
    isolate: RwLock<HashMap<TypeId, Symbol>>,
    // Intercept table: TypeId → override (prototype-chain)
    intercept: RwLock<HashMap<TypeId, Arc<dyn Any + Send + Sync>>>,
    // Fiber that owns this context's lifecycle
    fiber: Arc<Fiber>,
    // Parent for hierarchical lookup (prototype chain)
    parent: Option<Arc<Context>>,
    // Root for epoch-computation reachability
    root: Weak<Context>,
}
}
  • Context::new_root() -> Arc<Context> creates Fiber::Inactive.
  • Context::extend(&self) -> Arc<Context> creates child with parent = Some(self) (lexical scope / request scope).
  • Context::provide::<T: Service>(&self, svc: T) inserts TypeId::of::<T>() → Arc<T> into store; witnessed by effect.
  • Context::get::<T: Service>(&self) -> Option<Arc<T>> walks store → intercept → parent.store (coeFFECT lookup).
  • Context::get_relaxed::<T: Service>(&self) -> Option<Arc<T>> — same walk, but a locally-owned provider whose owner fiber rests mid-transition (Active, Loading, Reloading, Unloading, reactive Pending) still resolves. Strict get refuses those so consumers never observe mid-transition values; disposed owners (undos already ran) and terminal Failed{error} owners stay refused even relaxed.
  • Context::isolate::<T>(&self, label: &str) -> Arc<Context> creates child whose isolate[TypeId::of::<T>()] = Symbol(label).
  • Context::intercept::<T>(&self, override: T) -> Arc<Context> creates child whose intercept[TypeId::of::<T>()] = override.

No unsafe. No libloading in spike (stubbed); YAGNI HMR deferred behind #[cfg(feature = "hmr")].


2. witnessed effects, temporal composability

Cordis witnessed effect function: (Γ → Γ) × (Γ → Γ) pair (do + undo) with LIFO accumulator for revertible mutations. Guarantees: if fiber disposes, all effects it applied are reverted in reverse order.

Rust:

#![allow(unused)]
fn main() {
pub trait Disposable: Send + 'static {
    fn dispose(self: Box<Self>);
}

pub struct EffectGuard {
    // LIFO accumulator — Box<dyn FnOnce() + Send>
    acc: Vec<Box<dyn FnOnce() + Send>>,
}

impl Drop for EffectGuard {
    fn drop(&mut self) {
        while let Some(undo) = self.acc.pop() { undo(); } // reverse order
    }
}

pub trait Effect: Send + Sync + 'static {
    fn apply(&self, ctx: &Context) -> Box<dyn Disposable>;
}

// Helper on Context — mirrors Cordis Context::effect
impl Context {
    pub fn effect<E: Effect>(&self, eff: E) -> Box<dyn Disposable> {
        let guard = eff.apply(self);
        self.fiber.accumulator.lock().push({
            let ptr = /* capture undo closure */;
            Box::new(move || { /* undo */ })
        });
        guard
    }
}
}

Spike verification (Phase 1, §8): temporal composability test

#![allow(unused)]
fn main() {
#[tokio::test]
async fn temporal_composability() {
    let ctx = Context::new_root();
    let fiber = ctx.plugin(FooService::new(), FooConfig::default()).await;
    ctx.provide(BarService(42));
    assert_eq!(ctx.get::<BarService>().unwrap().0, 42);
    fiber.dispose().await;
    assert!(ctx.get::<BarService>().is_none());
    assert_eq!(ctx.snapshot(), pre_plugin_snapshot);
}
}

Must hold before Phase 2.


3. coeffect table Σ, spatial composability

Cordis Σ is a TypeId-keyed table of dependency declarations (inject = ["foo"]). Fiber recomputes epoch from dependency UIDs; if epoch unchanged, no reload.

Rust anymap/typemap equivalent (hand-rolled to avoid extra dep in spike):

#![allow(unused)]
fn main() {
pub struct CoeffectTable {
    // TypeId → (TypeId, Symbol, UID)
    injects: HashMap<TypeId, (Symbol, String)>, // String = epoch fragment ":uid"
}

impl CoeffectTable {
    pub fn declare<T: Service>(&mut self, label: Symbol) {
        self.injects.insert(TypeId::of::<T>(), (label, uid_for::<T>(label)));
    }
    pub fn epoch(&self) -> String {
        // Monoid over concatenation, per Cordis: ":uid1:uid2:..."
        let mut frags: Vec<_> = self.injects.values().map(|(_, uid)| uid).cloned().collect();
        frags.sort();
        frags.join(":")
    }
}
}
  • aranymap crate is not needed; HashMap<TypeId, Box<dyn Any + Send + Sync>> with TypeId::of::<T>() suffices.
  • Symbol is Arc<str> or &'static str for isolate labels (e.g., tenant:abc).
  • Handlers currently take State<AppState> (17,22 fields, src/lib.rs:230). New handlers: State<Arc<Context>> + ctx.get::<T>() where T is declared as inject. Example:
#![allow(unused)]
fn main() {
 // Before (P0 god-struct):
 async fn chat(State(state): State<AppState>, ...) -> Response
 
 // After (decomposed Context):
 async fn chat(State(ctx): State<Arc<Context>>, ...) -> Response {
     let exec = ctx.get::<dyn AgentExecutionService>().expect("no execution service");
     exec.execute(req, &ctx).await
 }
}

Spike verification: spatial composability

#![allow(unused)]
fn main() {
#[tokio::test]
async fn spatial_composability() {
    let ctx = Context::new_root();
    let consumer = ConsumerService::new(inject: vec![TypeId::of::<FooService>()]);
    let fid = ctx.plugin(consumer, Config::default()).await;
    assert_eq!(ctx.fiber_state(fid), FiberState::Inactive); // dep missing
    ctx.provide(FooService);
    assert_eq!(ctx.fiber_state(fid), FiberState::Active); // auto-reload
    ctx.provide(FooService::v2()); // re-provide
    assert_eq!(ctx.fiber_epoch(fid).prev, ":foo_v1");
    assert_eq!(ctx.fiber_epoch(fid).current, ":foo_v2"); // reload triggered
}
}

4. isolate & intercept

Isolate (spatial scoping)

Cordis isolate("name", label) creates realm where provide/inject are scoped.

Rust:

#![allow(unused)]
fn main() {
impl Context {
    pub fn isolate<T: Service>(&self, label: impl Into<Symbol>) -> Arc<Context> {
        let child = self.extend();
        child.isolate.write().insert(TypeId::of::<T>(), label.into());
        child
    }
}

// Usage: per-tenant tool isolation (P10, Phase 3)
let tenant_ctx = root_ctx.isolate::<dyn ToolService>("tenant:acme");
tenant_ctx.provide(TenantToolService::new(tenant_id));
// tenant_ctx.get::<dyn ToolService>() returns tenant-scoped service
// root_ctx.get::<dyn ToolService>() still returns fleet service
}

Intercept (prototype-chain override)

Cordis intercept("key", config) overrides a coeffect without mutating the provider.

Rust:

#![allow(unused)]
fn main() {
impl Context {
    pub fn intercept<T: Service>(&self, override_val: T) -> Arc<Context> {
        let child = self.extend();
        child.intercept.write().insert(TypeId::of::<T>(), Arc::new(override_val) as Arc<dyn Any + Send + Sync>);
        child
    }
}

// Usage: per-request model pinning (P10, Phase 5)
let req_ctx = root_ctx.intercept(ModelOverride { model: "gpt-4o-mini".into() });
let llm = req_ctx.get::<dyn LlmService>().unwrap(); // sees override via prototype walk
}

Kernel intercept meta-events (beyond the prototype chain)

Five kernel operations expose listener-driven veto points on reserved events (internal/get, internal/set, internal/config, internal/update, internal/listener), plus an internal/dispatch observer. These sit outside the product event catalog on purpose.

  • EventsService::intercept_get/set/config/update/listener implement the semantics; internal/dispatch reports every NON-internal dispatch as (mode, name, args) and exempts itself from observation.
  • internal/get consults a Bail chain on every strict Context::get: a non-null terminal replaces the returned value, {"refuse": true} fails the lookup, a redirect verdict continues at the parent frame, null passes through, and a chain error refuses the read.
  • internal/set errors veto the provider write; the previous binding stays fully intact.
  • internal/config's non-null terminal IS the effective configuration for one apply pass (staged on the fiber and consumed by the registry runner); a chain error rests the fiber terminal Failed.
  • internal/update bails skip the scheduled restart entirely — the fiber keeps serving its current application and the deferred config stays readable via Fiber::vetoed_config.
  • internal/listener bails cancel the registration; the caller receives an inert handle and neither registry sees the listener.

Zero-cost gating: every helper short-circuits on listener_count == 0 before doing anything else, so the default path is two map lookups. Synchronous call sites bridge through block_in_place on multi-thread tokio runtimes; single-thread flavors fall OPEN (warning + allow), matching historical no-listener behavior. A thread-local re-entrancy fence keeps operations made inside a chain un-intercepted. The *_from dispatch family — bail_from, waterfall_from, waterfall_async_from — exposes the same chains for product code, adding an optional per-dispatch ListenerFilter; filtered listeners skip one dispatch and remain registered.


5. fiber lifecycle (with inertial lock)

Cordis Fiber states: Inactive → Reloading → Active → Unloading plus Inertia lock to serialize transitions (Thm 63, guarded withdrawal: provider does not withdraw until dependents deactivate).

Rust:

#![allow(unused)]
fn main() {
pub enum FiberState {
    Inactive { error: Option<AppError> },
    Reloading { iter: Box<dyn EffectIterator>, acc: EffectAcc, committed: CommittedView },
    Active { acc: EffectAcc, committed: CommittedView },
    Unloading { acc: EffectAcc, committed: CommittedView, outcome: Option<AppError> },
}

pub struct Fiber {
    state: RwLock<FiberState>,
    inertia: Arc<tokio::sync::Mutex<()>>, // serialize transitions
    acc: Mutex<EffectAcc>, // Vec<Box<dyn FnOnce() + Send>>
    epoch: RwLock<String>, // computed epoch
    injects: CoeffectTable,
    committed: CommittedView, // snapshot for rollback
}

impl Fiber {
    pub async fn refresh(&self) {
        let _guard = self.inertia.lock().await; // Thm 63
        let new_epoch = compute_epoch(&self.injects);
        if *self.epoch.read() == new_epoch { return; } // no change
        self.reload().await;
    }

    async fn reload(&self) { /* iterate effects, recompute, commit or rollback */ }
    async fn dispose(self: Arc<Self>) { /* LIFO undo, state → Inactive */ }
}
}
  • EffectIter is Box<dyn Iterator<Item = Box<dyn Effect>> + Send>, each Service::init yields effects.
  • CommittedView is HashMap<TypeId, Arc<dyn Any>> snapshot taken at Active entry; used for rollback on failure.
  • notify (see §7) triggers Fiber::refresh() via BFS over dependent fibers.

File placement: crates/cordis/src/fiber.rs (spike) → later crates/ares-context/src/fiber.rs.

The fiber lifecycle adds Loading and Failed: Loading marks a fiber mid-instantiation and Failed records a terminal error from a plugin activation. A failed fiber remains observable via the registry so a loader or admin tool can report why a registration did not become Active; a later successful re-registration starts a fresh fiber that reaches Active. In 0.9.0, Fiber::refresh still compares epochs, then undoes prior effects and reruns the registered plugin apply when a reload is required.

Pending rest state (reversible withdrawal)

A fifth state completes the machine: an Active runner whose dependency is genuinely withdrawn disposes its effects LIFO under Unloading, then rests Pending instead of dying or going Inactive. While Pending, the fiber keeps its registry key (it survives prune_disposed) and reactivates through Loading when the provider returns.

  • Pending is reserved for reactive waiting only: apply errors still rest terminal Failed{error}, and a peer-version constraint refusal over a live provider rests Inactive (the provider is still available; the refusal is policy, not loss).
  • Eligibility needs one fully-satisfied refresh pass first — registration alone cannot mark a fiber eligible because its declares may still be unserved.

Readiness barriers vs availability predicates

RegistryService::register_with_readiness(ctx, plugin, config, ready_when) installs a ReadinessBarrier consulted before every activation pass. While the gate reports not-ready the fiber rests inspectable Pending — quiet waiting that never becomes Failed — with the factory run once up front and strict get refusing the non-Active owner, so the service stays out of consumer reach until the observed environment turns ready.

  • ReadinessBarrier::new(pred) wraps one Fn(&Arc<Context>) -> bool; .and(other) AND-composes; with_readiness([a, b, c]) folds any number of barriers (an empty list is vacuously ready).
  • .watching([TypeId]) unions the provider keys whose settlements re-kick the gated fiber through the ReflectService fan-out: an external provide or withdrawal re-evaluates the gate without anyone touching the fiber.

This complements rather than replaces availability predicates (Service::check): a rejected availability predicate rests Failed{error: "availability predicate rejected service"} — loud and terminal per the rules above — while a closed readiness gate is quiet, reversible waiting. Use the predicate when the factory cannot produce a valid service; use the barrier when production succeeds but serving should hold.


6. epoch, hash of dependency UIDs

Cordis epoch is monoid ":uid1:uid2:..." (concatenation). Fiber skips reload if epoch unchanged.

Rust:

#![allow(unused)]
fn main() {
pub fn compute_epoch(injects: &CoeffectTable) -> String {
    // Monoid over concatenation per paper §4.3
    let mut frags: Vec<String> = injects.uids_sorted();
    if frags.is_empty() { return ":".into(); }
    format!(":{}", frags.join(":"))
}

// uid_for<T> = format!("{}:{}", std::any::type_name::<T>(), label)
// Example: epoch = ":ares_llm::LlmService:tenant_acme:ares_tools::ToolService:tenant_acme"
}
  • The epoch type is String, not a hash. The paper uses concatenation for debuggability; if perf matters, switch to sha2 hash later.
  • Fiber::refresh compares self.epoch.read() vs compute_epoch(&self.injects); logs diff via tracing::debug!.

7. notify, reactive recomputation (tokio::sync::watch)

Cordis notify is fan-out to dependent fibers. In the TS Harness, notify uses EventEmitter; in Rust, it uses tokio::sync::watch.

#![allow(unused)]
fn main() {
pub struct ReflectService {
    // TypeId of changed service → watch channel sender
    notifiers: RwLock<HashMap<TypeId, watch::Sender<()>>>,
    // dependency graph: provider TypeId → [dependent FiberId]
    dependents: RwLock<HashMap<TypeId, Vec<FiberId>>>,
}

impl ReflectService {
    pub fn notify(&self, changed: TypeId) {
        // BFS walk dependents
        let deps = self.dependents.read().get(&changed).cloned().unwrap_or_default();
        for fid in deps {
            if let Some(sender) = self.notifiers.read().get(&changed) {
                let _ = sender.send(()); // fan-out, ignore closed receivers
            }
            // also trigger Fiber::refresh via task
            tokio::spawn({
                let fiber = self.fiber_for(fid);
                async move { fiber.refresh().await }
            });
        }
    }
}

// DB-backed source example (replaces 60s poll in runtime_registry.rs etc.):
// On Postgres NOTIFY (or polling fallback every 60s if no NOTIFY), call:
// ctx.get::<ReflectService>().unwrap().notify(TypeId::of::<RuntimeToolService>());
}

Replaces: RuntimeToolRegistry::start_background_reload (60s poll, crates/ares-tools/src/runtime_registry.rs), ProviderRegistry poll (crates/ares-llm/src/provider_registry.rs), NvidiaCatalogCache::start_background_refresh (crates/ares-llm/src/nvidia_catalog.rs).


8. events, 5 dispatch modes

Cordis Events: emit / parallel / serial / bail / waterfall typed bus.

Rust:

#![allow(unused)]
fn main() {
#[derive(Clone, Copy, Debug)]
pub enum Dispatch {
    Emit,      // fire-and-forget, no return, no error propagation
    Parallel,  // tokio::JoinSet, collect all, fail-open (one handler error doesn't cancel others)
    Serial,    // sequential, fail-open
    Bail,      // sequential, fail-fast (first error aborts)
    Waterfall, // sequential, each handler receives previous handler's output (chained)
}

pub struct EventsService {
    handlers: RwLock<HashMap<EventId, Vec<Handler>>>, // Handler = Box<dyn Fn(Value) -> Future<Output=Result<Value>> + Send>
    bus: broadcast::Sender<EventEnvelope>, // tokio::sync::broadcast for cross-task fan-out
}

impl EventsService {
    pub fn on(&self, event: EventId, handler: Handler) -> Box<dyn Disposable> {
        self.handlers.write().entry(event).or_default().push(handler);
        // return Disposable that removes handler on dispose (LIFO undo)
        Box::new(RemoveHandler { event, idx: len - 1 })
    }

    pub async fn dispatch(&self, event: EventId, payload: Value, mode: Dispatch) -> Result<Value> {
        match mode {
            Dispatch::Emit => { self.bus.send(envelope(payload)); Ok(Value::Null) }
            Dispatch::Parallel => { /* JoinSet */ }
            Dispatch::Serial => { /* loop */ }
            Dispatch::Bail => { /* loop with bail */ }
            Dispatch::Waterfall => { /* chain payload through handlers */ }
        }
    }
}
}

Mapping from TS Harness (12 layers, ~60 packages), in Rust, one crate suffices; do not replicate layering ceremony. EventsService is a Service itself (ctx.provide(EventsService::new())), so any fiber can ctx.get::<EventsService>().unwrap().on(...).

Dispatcher parity (shipped)

The dispatch implementation follows the five modes exactly:

  • Emit: every handler is spawned and not awaited (fire-and-forget); dispatch returns JSON null after broadcasting the event and payload on the bus. A caller that needs completion can listen on the bus or use a oneshot channel, not await this call.
  • Parallel: handlers run concurrently via tokio::task::JoinSet; successful dispatch returns JSON null (handler values are discarded). The first error observed is propagated (a joined panic surfaces as CordisError::Fiber).
  • Serial: handlers run in registration order with the original payload; the first non-null result bails and is returned. An all-null chain returns the original payload. Serial and Bail share this path.
  • Bail: stops at the first handler that returns a non-null result and returns that value without running later handlers; a null result means not bailing and the chain continues with the original payload.
  • Waterfall: each handler transforms the payload and passes the result to the next; a handler short-circuits by returning an object whose waterfall_stop field is true. This is the Rust static-dispatch analogue of the TS next() closure: instead of passing a next function, a handler opts out by returning the sentinel.

Dispatch participation knobs (EventOptions / emit_filtered)

Flat listeners can register through on_with / once_with with EventOptions { prepend, global }:

  • prepend: true inserts the listener at the FRONT of the dispatch-order list, so it runs before previously registered listeners of the same event.
  • global: true marks the listener realm-agnostic: emit_filtered(event, args, filter) offers every non-global listener to the filter predicate first and excludes it from that one dispatch on a false verdict — without unregistering it. Global listeners bypass the filter entirely.
  • The historical on / once / emit signatures delegate with default options (false/false), so existing registrations are byte-compatible. The broadcast bus fan-out is not filtered; only registered handlers participate.

Event-first skill execution (0.9.0)

Each skill receives the request Context. Tool steps run on a tenant-scoped context created with ctx.isolate::<Tools>(tenant_id) and invoke Tools::execute through tools.execute. Every Skill LlmCall step uses strict Llm::complete through llm.complete; SkillEngine and SkillsService do not call providers directly or fall back to generate_with_history. When EventsService is present, waterfall_around wraps these capability calls. Tools, Llm, and Execute public methods stay on the same event-first path.


9. loader & config reconciliation (Declarative)

Cordis Loader: Entry { id, plugin, config, disabled, isolate, intercept } + EntryTree(Vec<Entry>) persisted to config/entries.json (or config/cordis-entries.toon via toon-format 0.4.1). Loader::reconcile(current, desired) diffs incrementally.

Rust (Phase 3, crates/cordis/src/loader.rs):

#![allow(unused)]
fn main() {
#[derive(Serialize, Deserialize, Clone)]
pub struct Entry {
    pub id: String,               // fiber id, e.g. "tool:calculator"
    pub plugin: PluginId,         // e.g. "ares_tools::CalculatorService"
    pub config: serde_json::Value,// Plugin::Config serialized
    pub disabled: bool,
    pub isolate: Option<String>,  // e.g. Some("tenant:acme")
    pub intercept: HashMap<String, Value>,
}
pub struct EntryTree(pub Vec<Entry>);

impl Loader {
    pub fn reconcile(&self, current: &EntryTree, desired: &EntryTree) {
        // per-field dispatch (paper §5):
        // id/plugin change → rebuild fiber (dispose + new)
        // config change → fiber.update(new_config)
        // disabled toggle → fiber.retire() / fiber.begin()
    }
}
}

Persistence: config/entries.json (or config/cordis-entries.toon) separate from ares.toml symlink (/opt/ares-config/ares.toml), do not conflict (see Assumptions in plan). Reuse toon-format serialization.

Loader journal (shipped)

LoaderJournal makes the UpdateConfig and Retire arms real. It stores a JournalRecord per entry: the plugin label owning the entry, the last applied config, the live fiber id when known, and a monotonically increasing generation counter (every mutation bumps it). It is the single source of truth for "is this entry live, with which fiber, at what config/version".

  • Loader::instantiate and the RebuildFiber arm call journal.upsert(id, plugin, config, Some(fid)) after a successful factory invocation.
  • UpdateConfig reads the recorded fiber id, resolves it via RegistryService::get_fiber, and calls Fiber::update when a live fiber is known (running block_in_place on a multi-thread runtime, journal-only on a current-thread runtime or no runtime); it then calls journal.update_config(id, new_config, recorded), bumping generation.
  • Retire calls journal.retire(id), clearing the record.

It is provided as a service with ctx.provide(LoaderJournal::new()), so Context::get::<LoaderJournal> returns the shared handle. When absent, Loader::execute_action and Loader::instantiate degrade to log-only.


10. Plugin & RegistryService

Cordis plugins are FnOnce(&Context, Config) -> Result<Disposable> or struct with apply. Registry enforces single-source discipline.

Rust (Phase 2, crates/cordis/src/registry.rs):

#![allow(unused)]
fn main() {
pub trait Plugin: Send + Sync + 'static {
    type Config: Serialize + DeserializeOwned + Send + Sync;
    fn apply(&self, ctx: &Context, config: Self::Config) -> Result<Box<dyn Disposable>>;
}

pub struct RegistryService {
    fibers: RwLock<HashMap<FiberId, Arc<Fiber>>>,
}

impl RegistryService {
    pub fn plugin<P: Plugin>(&self, plugin: P, config: P::Config) -> Result<FiberId> {
        // check duplicate provider: no two fibers may provide same TypeId in same isolate realm
        // if violation: return Err(AppError::Configuration("duplicate provider for TypeId"))
        // else: create Fiber, store, return FiberId
    }
}
}

Static registration (preferred production): inventory/linkme (compile-time plugin set), real surface is RegistryService::plugin (single-source discipline); inventory::submit! / linkme::distributed_slice of fn(&Arc<Context>) -> Result<FiberId, CordisError> is the future shortcut once crate count stabilizes (see crates/cordis/src/lib.rs HMR section and Wiring task). Dynamic HMR (dev only, behind #[cfg(feature = "hmr")] off by default): libloading path that dlopens .so and calls Plugin::apply via extern "C"; if libloading ABI fragility blocks (Rust 1.98 toolchain coupling, unsafe soundness), fall back to file-watch + full fiber reload (re-read config/TOON), 90% of value per plan, see watcher fallback below.

HMR YAGNI decision (plan Assumptions §Contingencies)

Decision: DEFER libloading HMR, keep file-watch + Fiber::reload as production path.

  • Rationale: libloading::Library::new + Symbol<extern "C"> requires unsafe, a stable repr(C) ABI boundary, and the .so to be built with the exact same Rust toolchain (1.98). ABI drift across patches, rust-doctor soundness flags, and Box::leak ownership hazards make libloading too brittle for a generic runtime. As plan contingency states: "If libloading HMR proves too complex for Rust (dynamic library ABI fragility, unsafe surface), fall back to file-watch + full fiber reload without dynamic code swapping … file watcher still triggers Fiber::reload() by re-reading config/TOON, which already covers 90% of self-evolution value. Dynamic code HMR can be deferred to a later phase behind #[cfg(feature = "hmr")] without blocking the core redesign."
  • Fallback implemented: crates/cordis/src/watcher.rs (watch_many / watch_cordis_entries) uses notify::RecommendedWatcher (debounced 500 ms + 100 ms settle, same as AresConfigManager::start_watching) to watch config/agents/*.toon (recursive) and config/entries.json (or config/cordis-entries.toon parent dir). On Modify/Create it calls ReflectService::notify(tid) which BFS-walks dependents and spawns Fiber::refresh (epoch recompute via compute_epoch). No restart, no libloading. Logs Configuration hot-reloaded successfully via Cordis watch (generalizes AresConfigManager's Configuration hot-reloaded successfully which is already proven on random-port E2E 39476/39120, see docs/cordis-redesign.md §9/9b).
  • Stub preserved: crates/cordis/src/hmr.rs is #[cfg(feature = "hmr")] (Cargo feature hmr = ["dep:libloading"], off by default). It shows libloading::Library::new + get::<HmrEntryFn> + owned HmrLibrary holder (RAII, no Box::leak) calling cordis_plugin_apply (extern "C"). Enable with cargo build --features hmr and a .so built with the same toolchain. Not invoked by src/main.rs or ReflectService, watcher is the production path.
  • Cargo.toml: [features] hmr = ["dep:libloading"] (libloading 0.8 optional, notify 8.2.0 always for watcher), default = [].

11. what is explicitly not ported in spike (updated)

Per YAGNI (Phase 1, §8) + HMR deferral above:

  • ❌ libloading HMR DEFERRED, file-watch fallback crates/cordis/src/watcher.rs (notify → ReflectService::notify → Fiber::reload via epoch) covers 90% value without dynamic code. Dynamic code swap remains as crates/cordis/src/hmr.rs stub behind #[cfg(feature = "hmr")] (off by default, libloading 0.8 optional). See HMR decision above and lib.rs HMR section.
  • ❌ WASM, deferred.
  • ❌ Visual layer package (~60 TS packages, 12 layers), in Rust, one crate; do not replicate ceremony.
  • ❌ ares.toml symlink handling, keep AresConfigManager::start_watching() as-is for Phase 2; Loader is additive. watcher generalizes it to Cordis entries/TOON without touching ares.toml symlink (/opt/ares-config/ares.toml).

12. critical anchors (Reread before phases 2/4/5)

  • /opt/ares/src/main.rs run_server (lines 296,889, 17 steps) → becomes root_ctx.plugin(...).plugin(...).await (5,8 lines). Every registry/pool/cache must migrate to a Service.
  • /opt/ares/src/lib.rs AppState (230,274, 17,22 fields) + base_router() → Arc<Context>, build_router(ctx: Arc<Context>).
  • /opt/ares/crates/ares-tools/src/runtime_registry.rs start_background_reload (60s poll, ArcSwap) → epoch-driven notify.
  • /opt/ares/crates/ares-llm/src/provider_registry.rs ArcSwap<HashMap> + NvidiaCatalogCache → LlmService with circuit breaker.
  • /opt/ares/src/api/handlers/admin.rs 190 KB, 5,946 lines, split by domain in Phase 6.

13. consequences & alternatives

  • If async fn in trait causes dyn issues, use async_trait only for that trait and document why (Rust 1.98 floor, 1.75+ stable for async fn in trait, but a dyn Service can require async_trait; prefer an impl Future return).
  • If TypeId + HashMap proves too coarse (downcasting ergonomics), evaluate anymap/typemap crates, but hand-rolled HashMap<TypeId, Box<dyn Any>> is sufficient for spike.
  • If epoch String concatenation bloats logs, switch to sha2 digest and keep :uid1:uid2 only in tracing::debug!.

14. Current architecture (0.9.0)

This section describes the shipped 0.9.0 runtime. Earlier sections remain the Phase 0 mapping and 0.8 spike record.

Fiber refresh

Fiber::refresh recomputes the dependency epoch. When the epoch changes, or the fiber is not already Active with dependencies satisfied, it undoes prior effects and reruns the registered plugin apply. Dispose still LIFO-undoes and passes through Unloading.

Wave-1 refinements: a working runner whose dependency genuinely vanishes rests reversible Pending (see §5) instead of dying, and reactivates when the provider returns; Context::get_relaxed lets lifecycle code read locally-owned values during those transitions while strict get stays conservative; Fiber::subscribe_state exposes every transition to panic-contained synchronous observers.

EventsService dispatch

Shipped behavior in crates/cordis/src/events.rs:

  • Emit: fire-and-forget; dispatch returns JSON null.
  • Parallel: handlers run on a JoinSet; successful dispatch returns JSON null. Handler return values are not collected. The first join or handler error is propagated.
  • Serial: handlers run in registration order with the original payload. The first non-null result bails and is returned. An all-null chain returns the original payload. Serial and Bail share this path.
  • Waterfall: around-middleware with next; skipping next skips later handlers and core.

Event-first product path

Tools, Llm, Execute, and skills stay event-first. Public methods go through EventsService::waterfall_around when the bus is on ctx (tools.execute, llm.complete, agent.run). Skill tool and LLM steps use those same events on a tenant isolate. They do not call providers or the tool registry directly.

agent.admit

agent.admit (Dispatch::Bail) is the shared quota gate. Execute::run, JWT chat, API-key middleware, and MCP all dispatch it before work. The default handler denies monthly or daily quota for non-enterprise tenants. HTTP maps deny to 429; MCP maps deny to a tool error.

Store, Overlay, TOON

The Store loader factory connects, runs sqlx migrations, and seeds default agent templates, then provides Store. Overlay fills empty loader entry.config from ares.toml sections and leaves non-empty configs unchanged. TOON file changes notify TypeId::of::<Tools>() and TypeId::of::<Execute>() so those fibers refresh.

Isolate vs intercept

For the same TypeId, an isolate label wins: get skips intercept and walks the isolated store/parent. Unlabeled types still intercept, so request TenantContext and ModelOverride keep working on a Tools/Execute realm.

TenantRealms

Request paths open the tenant realm then intercept TenantContext (open then with_intercept), including JWT middleware (tenant claims), JWT chat/research, and v1 chat/stream/agents. JWT user: isolate does not invent a dummy Free tenant. Background jobs open or isolate only and do not attach a request intercept. Admin tenant delete calls TenantRealms::dispose before SQL delete.

Facade

The default ares crate has no axum. Context, Execute, Tools, Llm, and register_plugins are enough to run an agent. Enable feature http to pull ares-http. ProviderRegistry is not re-exported from ares; construction tests import it from ares-llm.

Honest residuals

  • ProviderRegistry still exists on ares-llm because Llm::new / AgentRegistry::from_config still take one during construction.
  • run_server still instantiates Overlay first, fills empty loader configs, then instantiates remaining config/cordis-entries.toml entries.
  • Scheduler, pipeline, and trigger domain loops remain native ARES engines. They inject Execute and run behind it; they are not a second public agent API.
  • Root ares-server is a binary. Overlay lives in crates/ares-http/src/overlay.rs; the server still registers the Overlay factory.
  • Overlay / optional ServerRuntime provide host extras (ActiveRuns, SkillEngine, MCP). Execute is registered once, by ares-agent.

15. Round 4 (0.9.x): policies, HMR finish, static factories, metrics, atomic saves

Typed event payloads

All 22 catalog events in crates/cordis/src/events_catalog.rs (CONTRACTS, contract_for) carry typed payload structs in crates/cordis/src/events_payload.rs. Each implements the TypedEvent trait (type Payload: Serialize + DeserializeOwned + Clone + Send + Sync, plus NAME/MODE/AROUND const-linked to the catalog). EventsService::dispatch_typed<E> serializes and dispatches with the declared mode; on_typed / on_typed_waterfall deliver deserialized payloads to listeners, skipping malformed ones with a warning. A consistency test asserts every binding matches CONTRACTS and that no event is unbound. The raw Value API remains for kernel and dynamic cases.

Engine choreography

Scheduler admission is scriptable through two waterfall events: scheduler.before_run output overrides agent_name/message on the executed request (audit identity stays schedule-tied); a Bail denial emits {ok:false, denied:true} and returns without running — due-pass callers solely advance next_run. Boundary events SCHEDULER_TICK, SCHEDULER_SCHEDULE_DISPATCHED, PIPELINE_STEP_STARTED/FINISHED, PIPELINE_FANOUT_COMPLETED, and TRIGGER_FIRED are emitted from scheduler/pipeline/trigger engines, completing full-catalog adoption.

RhaiPolicy scripting (default-on)

RhaiServiceConfig.listen: Vec<RhaiListenerConfig { event, fn_name }> lets a declarative entry attach sandboxed Rhai functions to any catalog event:

[[entry]]
id = "policy-example"
plugin = "RhaiPolicy"
[entry.config]
script = "fn gate(p) { if p.tenant_id == \"banned\" { #{deny: \"banned\"} } else { () } }"
[[entry.config.listen]]
event = "scheduler.admit"
fn_name = "gate"

Semantics: the payload arrives as an object map (plain property access). Returning ()/null passes through (Bail) or delegates via next (waterfall); any other value becomes the dispatch result — deny marker or short-circuit. Script runtime errors log a warning and pass through/delegate. Init validates each event against the catalog (unknown ⇒ Configuration error, recorded per-entry at boot). All listener disposables combine with the init guard so fiber dispose unregisters them. Enabled by default through root feature rhai-policy; factory key "RhaiPolicy" registered under the rhai feature.

Static registration (inventory-collected factories)

cordis::CordisPluginFactory { name, make } (make is a plain fn(&Arc<Context>, &Value) -> Result<FiberId, CordisError>) is inventory-collected; cordis::register_inventory_factories(reg) installs every submission and is the server's primary boot path (src/main.rs). The hand-written register_plugins chains remain compiled as the --no-default-features fallback and for tests. Each capability crate declares an optional inventory feature forwarded from root/facade; because inventory nodes live in linker sections, a crate with no otherwise-referenced code loses its submissions — tests force linkage by calling the manual chains before collecting. Parity proofs: crates/ares/tests/inventory_parity.rs (library crates) and tests/server_inventory_probe.…

Admin surface

  • POST /admin/cordis/services/{name}/retire / provide — runtime retire/re-provide.

  • GET|PUT /admin/cordis/entries, DELETE /admin/cordis/entries/{id}, POST /admin/cordis/entries/{id}/toggle — declarative entries management (Null configs normalize to {}), 503 when loader state is absent.

  • PATCH /admin/cordis/entries/{id} — typed partial update via cordis::loader::EntryUpdate (config, disabled, isolate, intercept; id/plugin deliberately not patchable). Only present fields change, {} is a validated no-op that still persists and re-applies; replies with the post-patch entry and per-action outcomes, 404 on unknown ids. A failed config pre-flight attaches a machine-readable issues array ([{message, path}, …]) beside the legacy error string.

  • POST /admin/cordis/entries/reload — reload from disk through the shared apply flow.

  • GET /admin/cordis/events — per-event dispatch counters {total_dispatched, by_event} from EventsService::dispatch_snapshot() (counts every mode via the single dispatch choke point).

Atomic entries persistence

EntryTree::save_to_toml_file writes <name>.tmp-<pid> then renames over the target (atomic on POSIX), preserving leading comment headers verbatim; the temp file is removed on failure so crashes mid-write leave either the previous or the new file, never a truncated one.

HMR resolution

The §10/§11 "defer" decision above is superseded. The dylib path is finished and correct, still opt-in: apply_plugin_so copies the library to a process-unique sibling (<stem>.<pid>.<seq>.hmr-load) before dlopen — glibc caches handles by path, so a rebuilt .so never swaps without this copy — and HmrLibrary::drop removes the copy after dlclose (best effort). Watcher hook apply_plugin_so_if_dylib benefits automatically. Production reload stays watcher + TOML reconcile; enable dylib loading with --features hmr only with same-toolchain cdylibs.


16. Round 5–6 (0.9.x): policy activation, effect removal, lifecycle hardening, composition

Round 5 — RhaiPolicy production activation

RhaiListenerConfig gained on_error: passthrough | deny (serde default = passthrough, byte-compatible with round-4 configs). deny is fail-closed: a flat-listener script error returns the built-in deny-marker shape instead of the payload; a waterfall-listener error vetoes without calling next. Multi-instance policies must use isolate realms (a pass-through handler on a shared Bail chain terminates every dispatch); each realm keeps its own EventsService. config/cordis-entries.toml ships an ACTIVE policy-admission-audit entry logging scheduler.admit agent names. The paper's witnessed-effect primitive (Effect, EffectGuard, Context::effect) was DELETED in round 5 — zero users ever; revertibility lives entirely in Disposable + Fiber LIFO a…

Round 6 — Guarded withdrawal (§4.3.1 reliedₙ)

RegistryService::reliance_count(&(TypeId, label)) derives (never seeds) how many ACTIVE fibers other than the provider resolve a key at its isolate label and declare an inject on it — always current under late declare_inject, no stale rows. A per-fiber realms ledger (FiberId -> Weak<Context>) captured at registration supplies context for label resolution. Context::remove<T> refuses while consumers remain ("guarded withdrawal: N active consumer(s)…"); internal rollback uses remove_forced (undo never blocks). Admin retire maps refusal to 409 {retired:false, reason:"guarded", consumers:N}.

Round 6 — Verified hot-swap

LoaderAction::RebuildFiber applies the new plugin OUT-OF-BAND on a scratch child context first: failure ⇒ old provider untouched. Success bridges via intercept (new values win instantly), disposes old while the bridge covers stale store rows, promotes peeked values store-first, then drops the bridge — get::<T> resolves at every instant (test probes from a concurrent task). Falls back to classic dispose+instantiate (reported verified=false) for untracked old fibers, isolated entries, or side-effectful factories (Store migrations run twice across trial+promotion). AppliedAction carries verified through admin PUT responses.

Round 6 — Typed listeners adopted

Scheduler's three prod listeners (agent.completed observability, agent.failed runtime-control, legacy shim) use on_typed::<AgentCompletedEvent/AgentFailedEvent>. rhai_service intentionally stays on the raw Value API (dynamic script maps are its purpose).

Round 6 — Dependency-cycle detection

cordis::cycles::{find_dependency_cycle, DependencyGraph} — colored-DFS over the fiber inject graph with deterministic node ordering; CycleLedger maps (TypeId,label)→provider-fiber and entry-id↔fiber so loader-side callers can reconstruct edges without registry internals. Unit-tested for self-loops, 2/3-cycles, nested-behind-prefix, disconnected components, cross-edge false positives.

Round 6 — Intercept meta-events, readiness barriers, cascade batching, logger, timers

Intercept meta-events + *_from dispatch. Five veto points (internal/get, internal/set, internal/config, internal/update, internal/listener) plus the internal/dispatch observer (see §4) ride EventsService::bail_from; product code gets the same chains through bail_from / waterfall_from / waterfall_async_from with an optional per-dispatch ListenerFilter. Veto semantics: a config-rewrite terminal becomes the effective config for that apply pass; an erroring config chain rests the fiber Failed; an update veto skips the restart and parks the deferred config in Fiber::vetoed_config; a listener bail cancels registration and returns an inert handle. Every gate is zero-cost when unregistered, and the synchronous bridges fall open on single-thread tokio flavors.

ReadinessBarrier. ReadinessBarrier::new(pred) + .and(..) / with_readiness([...]) + .watching([TypeId]), installed via RegistryService::register_with_readiness (see §5): closed gates hold fibers at inspectable Pending and re-kick through the ReflectService fan-out; availability predicates stay the loud Failed path.

Cascade batching. An in-flight ledger around loader re-applies collapses concurrent provider config updates to ONE dependent convergence wave per settled batch (CASCADE_INFLIGHT, consulted from the kernel refresh path).

LoggerService (cordis::logger). Bounded ring (default 1000) of Messages; effect-owned Exporter sinks (registration returns the removing Disposable) with per-sink ExporterConfig (per-name levels, max_length truncation); per-name thresholds via set_level with set_default_level fallback; zero-cost enabled gate before argument assembly; printf rendering %s %d %i %f %o %O %c %C %% with ANSI16 name-hashed colors (%c) and bold (%C); hyphenate/derived_name kebab-case logger names; LoggerIntercept per-fiber threshold overrides resolved through the relaxed read channel; Context::log/info/warn/debug/error facade no-ops when no logger is provided.

Timers (cordis::timer). Six primitives — timeout, sleep, interval, interval_stream, debounce, throttle — std-only, fiber-scoped via with_current_fiber labeled undos, running on one shared wheel thread; dispose yields exactly one final Err(InactiveEffect) on streams then closes them; dropping handles never cancels.


17. Round 7 (0.9.x): wiring composition and detection, drain-and-shift

Cycle detection wired

Loader::apply now reconstructs the post-apply inject graph via the CycleLedger (fresh-provide diff at instantiate_entry records providers; fiber.injected_type_ids() + ctx.isolate_label() resolve edges) and runs find_dependency_cycle. Detection never fails the apply — a found ring logs a warning naming entry ids; Loader::detect_cycle_entry_ids(&ctx) queries it, and GET /admin/cordis/entries carries an additive dependency_cycles key (entry-id rings, empty when healthy).

Composition wired at boot + reload

Both entry-load sites (boot_loader_program parse, watcher/poll reload via reload_current(ctx, path, &mut current, desired_composed, &journal) — callers own parsing+composition now) run cordis::compose_all (@include splice → @group flatten → ${rhai: …} interpolation) with the entries file's parent as base dir. Fail-open: composition errors log loudly naming path+error and proceed with the RAW entries. Interpolation scope variable is entry (rhai reserves $). Commented examples live in config/cordis-entries.toml.

Drain-and-shift provider replacement

Loader::replace_provider(ctx, plugin_name, config, journal): trial new provider out-of-band → intercept bridge → dispose old → promote store-first → fresh Active fiber journaled. Zero absence window (concurrent-get probe test). Shared tail extracted into SwapPromotion used by both verified rebuild and replace; fixing it closed a latent double-swap bug (promoted values had no undo, blocking any second swap). replace_provider bypasses Context::remove's guard legitimately — the bridge guarantees continuous resolution, which is exactly what the guard protects; genuine retire keeps the guarded path. Root-realm only; no dispose-then-rebuild fallback.


18. Round 8 (0.9.x): composed program, operational replace, metatheory smoke

Entries program split (composition load-bearing)

The production config/cordis-entries.toml now exercises the round-6 compose pipeline for real: events/overlay/store inline at the head, tools..probe spliced verbatim from config/cordis-entries-shared.toml via @include, http inline. Parity-proven: simulated composition of the new layout deep-equals the old flat 15-entry program, boot order unchanged. A second ACTIVE policy joined the program: policy-emergency-gate (isolate policy-emergency) denies any scheduled agent named EMERGENCY-HALT with on_error="deny" — a fail-closed kill-switch; policy-admission-audit remains last.

Operational provider replacement

POST /admin/cordis/services/{name}/replace (body {"config": <Value>}) exposes round-7's Loader::replace_provider: trial out-of-band → bridge → dispose old → promote → fresh fiber journaled. 200 {replaced:true, plugin, fiber_id} on success; 409 {replaced:false, reason} when replacement is refused (unknown label, untracked/isolated/failing trial — old provider untouched by design); 400 malformed body; 503 without journal. Providers can now be swapped under live traffic from outside the process.

Metatheory property smoke tests

cordis::metatheory encodes paper guarantees as executable checks over the public API:

  • quiescence_after_every_op (Thm 66 progress): 24-op deterministic schedule; no fiber rests in transitional states; Active fibers always have injects available. HELD — inertia keeps transitions unobservable at rest.
  • order_confluence_of_registrations (Cor 21 / Thm 73): consumer-first vs providers-first converge to identical epochs and projections. HELD.
  • dependent_never_active_without_provider (§spatial reactive invariant): HELD with two documented deltas — (1) a declare_inject can land on an Active fiber outside the state machine until refresh; (2) factories whose apply() errors become Failed without ReflectService wiring, where the paper expects permanently-Inactive dependents. Both are future-round deliverables.
  • lifo_dispose_restores_store (Thm 16): exact LIFO undo order, disposables fire exactly once, store restored.

19. Round 9 (0.9.x): eager reconciliation, failed-factory wiring, peer versioning

Metatheory deltas resolved (round-8 follow-ups)

Both documented deltas are now resolved behavior:

  1. Eager declaration: declare_inject on a fiber resting Active reconciles immediately — inertia try-lock fast path updates epoch+state in place (satisfied) or drives the refresh transition to Inactive{missing-dependency} (unsatisfied). A pending_declare flag folded into refresh's recompute loop makes declare-vs-refresh races lossless.
  2. Failed-factory wiring: both RegistryService::register failure paths wire the fiber into ReflectService (register_fiber + notify on the type key), so dependents observe provider-loss and rest Inactive at quiescence. Failed{error} remains the terminal visible state (deliberate divergence from the paper's permanently-Inactive expectation — operationally more useful); the provided slot stays vacant and re-registration supersedes with a fresh fiber id.

Bonus fix: a latent deadlock in register itself — an if-let scrutinee held the provided-map read guard across stale-slot cleanup that took the write lock. Exposed by metatheory leg F; fixed by deciding staleness under the read guard before any write.

Peer-dependency versioning

Addresses the paper's open problem (§discussion):

  • Provide side: Context::provide_versioned<T>(value, version); legacy provide() = version 0. Versions live in a parallel map with LIFO-undo restore (disposal restores the prior version exactly).
  • Inject side: Fiber::declare_inject_versioned<T>(min_compatible: Option<u64>). Satisfaction rule: provider exists AND major(provider) == major(requirement) AND provider >= requirement, where major(v) = v / 100_000. Mismatch ⇒ dependent stays Inactive — never binds incompatible versions.
  • Reactivity: constrained epoch fragments fold in "v@:", so same-major upgrades flip the epoch and reactively reactivate dependents (test: v200_001 → Inactive → compatible v100_005 re-provide → Active with new projection). Unconstrained fragments byte-identical to the prior scheme.
  • Deliberate open point kept: structural interface compatibility (paper's full problem) is NOT attempted — majors-only buckets with explicit floors cover the practical case; deeper structural checks remain future work.

Status vs paper guarantees

Quiescence (Thm 66): held. Order confluence (Cor 21/Thm 73): held. LIFO disposal (Thm 16): held. Reactive invariant (§spatial): now fully held — eager declarations close the last observable gap. Terminal-state divergence (Failed vs Inactive) is documented and deliberate.

20. Round 10 (0.9.x): accessors, layered chains, lifecycle riders, module graph, entry moves

Accessor registry (name-keyed computed properties)

Context::register_accessor(name, Accessor::{read_only, read_write, setter_only}) installs a computed property beside the TypeId service store and returns an EffectHandle whose disposal removes the declaration AND every alias. Context::alias(alias, target) binds an alternate name through the same registration slot; duplicates (including alias collisions) are rejected with DuplicateProvider. Typed reads surface PropertyTypeMismatch instead of a silent None; writes against a read-only property are refused with ReadOnlyProperty. Crucially, accessor traffic BYPASSES the internal/get / internal/set intercept waterfalls entirely — resolving an accessor never consults or re-enters a veto chain (accessors are policy plumbing, not provider state). Anchors: crates/cordis/src/context.rs.

Layered intercept chains

Intercept layers per TypeId are an ordered outermost..innermost sequence; NEW registrations APPEND rather than replace, so the innermost layer stays effective for all existing getters — no caller breakage when another policy layer joins. Context::intercept_chain(tid) returns every layer in dispatch order; Context::chains_structurally_equal(a, b) compares two chains by shared-instance identity (Arc::ptr_eq per layer pair) — erased values carry no comparable contract, so freshly-built values compare unequal by design, which is the honest test for restart decisions.

Lifecycle riders

Fiber::update now returns Result<(), CordisError>: a restart-path error propagates to the caller while the fiber stays Active serving its OLD configuration (effects never unwind on a failed restart). An internal/update veto parks the deferred config in Fiber::vetoed_config() and returns Ok — the skip is observable without being destructive. The internal/config waterfall consult now also covers the ACTIVATION path in the registry, so config rewrites apply on first activation identically to re-applies.

Module graph fan-out (opt-in)

cordis::module_graph::ModuleGraph maps module keys to their dependencies. With a ModuleReload implementation installed (with_reloader / set_reloader), change_many(ctx, keys) runs in two phases: phase 1 computes the TRANSITIVE affected plugin set READ-ONLY across the dependency closure; phase 2 reloads each affected plugin EXACTLY ONCE per transaction through ModuleReload, classified into a ChangeOutcome. A failing reload rolls back that plugin only — successfully reloaded siblings stay Active. The file watcher's debounced batch fans through a registered ModuleGraph when one is provided on the context; WITHOUT one, watcher behavior is byte-identical to before. Anchors: crates/cordis/src/module_graph.rs, wiring in crates/cordis/src/watcher.rs.

Entry moves (move_entry surface)

Entries gain hierarchy: EntryPosition { parent, position } rides on Entry. PATCH /admin/cordis/entries/{id} accepts optional parent / position applied move-THEN-update — an invalid placement (e.g. moving an entry into its own descendant) answers 409 before ANY mutation of file or live tree. POST /admin/cordis/entries/{id}/move relocates an entry together with its whole {id}:* subtree in one rename cascade. A valid move preserves fiber identity: the live instance refreshes IN PLACE under the new key, so consumers never observe a dispose/recreate window. Disabled groups move suppressed-then-restored. Anchors: EntryPosition, EntryTree::move_entry, subtree_ids, Loader::move_entry in crates/cordis/src/loader.rs; handlers in crates/ares-http/src/api/handlers/admin/cordis.rs.

Upstream Parity Ledger

What this document is

This document records every deliberate divergence between our kernel and the upstream model. The paper states the reference model. Our kernel sometimes chooses a different behavior on purpose. This file is the audit ledger for those decisions.

  • Date: 2026-08-24
  • Scope: crates/cordis versus the model in the paper
  • Method: every divergence lists the claim, the rationale, and the enforcement or test point

Divergence table

#DivergenceDecision
1Failed registrations stay visibleFailed{error} is a terminal rest state with reflective wiring
2Peer-dependency compatibilityMajors-only version buckets, no structural checks
3Late inject declarationsEager reconciliation on Active fibers
4Hot swap mechanicsOut-of-band trial then promote, honest swap_mode reporting
5Dynamic library loadingStrictly opt-in behind hmr, exact fingerprint handshake
6Factory collectionInventory primary, manual chains as fallback
7Serial dispatchDirect alias of Bail, waterfall uses real next continuations
8Worker supervisionReserved exit codes plus stdin-EOF death detection
9Log routingOne exporter router fans records to every gated sink
10Dependency withdrawalGenuine loss rests working fibers Pending (reversible); apply errors stay terminal Failed
11Dispatch participation knobsEventOptions{prepend,global} + emit_filtered; filters never exclude global listeners
12Reads during transitionsStrict get refuses transitioning owners; get_relaxed is the explicit opt-in

| 14 | Kernel operations are interceptable | Five meta-events veto or rewrite kernel ops; unregistered path stays zero-cost; sync bridges fall open | | 15 | Readiness gates wait quietly | Closed ready_when barriers rest fibers Pending (never Failed); availability predicates remain the loud path | | 16 | Config cascades batch | Concurrent provider updates collapse to one dependent convergence wave | | 17 | Validation errors carry paths | Pre-flight failures surface message + path issues beside the legacy string | | 18 | Logger adopted natively | Ring buffer, effect-owned exporters, per-name routing live in the kernel crate | | 19 | Timers adopted natively | Six fiber-scoped primitives share one std-only wheel thread | | 20 | Accessor traffic bypasses interception | Name-keyed computed properties resolve outside the internal/get / internal/set waterfalls entirely | | 21 | Intercept layers are ordered and inspectable | Append-on-set keeps the innermost layer effective; intercept_chain returns outermost..innermost | | 22 | Restart errors keep the old application | Fiber::update propagates errors with the fiber still Active; a veto parks vetoed_config and returns Ok | | 23 | Config interception covers activation | The internal/config waterfall consults on first activation too, not only re-applies | | 24 | Module changes fan out through one graph | change_many computes the affected set read-only first, reloads each plugin once per transaction, and rollback keeps siblings Active (EXCEEDS upstream: no module-graph concept) | | 25 | Entries relocate without losing identity | PATCH move-then-update answers 409 on conflict; /move renames the {id}:* subtree; in-place refresh preserves the fiber (EXCEEDS upstream: flat key list only) | | 26 | Subtask cancellation is wired end to end | Sticky cancel tokens honored at step boundaries plus an external trigger (EXCEEDS upstream: reference defines the hook but never wires it) | | 27 | Deterministic micro calls are cached | LRU+TTL keyed (model, system, input) with cache_hit telemetry; salvage/retry results never cached (EXCEEDS upstream: roadmap prose only) | | 28 | Guided grammars ride typed hints | Schema-shaped values become json_schema everywhere; raw GBNF stays a provider extension; absent hint = byte-identical wire | | 29 | Duplicate embeddings cost one backend call | Content-hash dedup on local and HTTP paths fans vectors back to duplicate slots (EXCEEDS upstream: roadmap prose only) |

Each row expands below with the claim, the rationale, and the evidence.

1. Failed registrations stay visible

  • Upstream expectation: a factory error leaves the fiber permanently Inactive and unreachable.
  • Claim: Failed{error} is a terminal VISIBLE rest state.
  • Detail: RegistryService::register returns Err.
  • Detail: the fiber enters the bookkeeping graph through RegistryService::wire_failed_registration.
  • Detail: ReflectService wiring registers the fiber against the attempted provider key.
  • Detail: notify fans out, so dependents observe the provider loss reactively and rest Inactive.
  • Detail: a later successful registration allocates a fresh fiber id and supersedes the failed fiber.
  • Rationale: operators inspect failures directly.
  • Rationale: dependents get a real notification instead of silence.
  • Rationale: fresh-id supersession keeps the provided slot free for retry.
  • Source: crates/cordis/src/registry.rs, wire_failed_registration
  • Test: metatheory property 3 legs D-H, metatheory_dependent_never_active_without_provider

2. Peer dependencies use majors-only buckets

  • Upstream expectation: compatibility needs full structural interface checks.
  • Claim: compatibility compares majors only.
  • Detail: versions are plain u64 values.
  • Detail: major(v) = v / VERSION_MAJOR_SCALE and floor(v) = v % VERSION_MAJOR_SCALE.
  • Detail: VERSION_MAJOR_SCALE equals 100_000.
  • Detail: a requirement binds when the major matches AND the provider reaches the floor.
  • Detail: any mismatch leaves the dependent fiber Inactive.
  • Detail: structural interface compatibility is deliberately NOT attempted.
  • Rationale: majors-only buckets cover practical drift between builds.
  • Rationale: full structural compatibility remains the open problem the paper defers.
  • Source: Context::provide_versioned and VERSION_MAJOR_SCALE in crates/cordis/src/context.rs
  • Source: Fiber::declare_inject_versioned in crates/cordis/src/fiber.rs
  • Tests: the version_conformance module in crates/cordis/src/metatheory.rs

3. Late inject declarations reconcile eagerly

  • Upstream expectation: a late declaration waits for an external refresh trigger.
  • Claim: a declaration landing on a fiber resting Active reconciles eagerly.
  • Detail: Fiber::reconcile_after_declare runs with the same transition shape as refresh.
  • Detail: satisfied declarations update the epoch in place.
  • Detail: unsatisfied declarations undo effects and rest the fiber Inactive.
  • Detail: a declaration that races an in-flight refresh folds into that refresh through a pending flag.
  • Detail: declarations on Inactive and Failed fibers wait for the next transition.
  • Rationale: eager recompute loses no racing declaration.
  • Rationale: the quiescence invariant survives every declaration path.
  • Source: Fiber::declare_inject and Fiber::reconcile_after_declare in crates/cordis/src/fiber.rs
  • Source: register paths in crates/cordis/src/registry.rs
  • Test: reactive invariant leg of dependent_never_active_without_provider

4. Hot swap drains then shifts

  • Upstream expectation: swap mutates providers in place.
  • Claim: swap builds out-of-band, then promotes.
  • Detail: new instances build inside a scratch context.
  • Detail: SwapPromotion bridges the new values through intercept bindings.
  • Detail: the old fiber disposes while the bridge keeps serving consumers.
  • Detail: promotion moves bridge values into the store before intercept removal.
  • Detail: consumers never observe an absence window.
  • Detail: an unverifiable swap reports swap_mode = "unverified" instead of a fake success state.
  • Rationale: drain-and-shift proves a zero absence window under concurrency.
  • Rationale: honest reporting lets operators tell verified swaps from unverified ones.
  • Source: Loader::replace_provider and SwapPromotion in crates/cordis/src/loader.rs
  • Tests: replace_provider_zero_absence_window
  • Tests: rebuild_same_type_verified_swap probes resolution from a concurrent task during the swap

5. Dylib loading is strictly opt-in with a fingerprint handshake

  • Upstream expectation: dynamic library loading runs as a first-class default.
  • Claim: dylib loading requires the hmr cargo feature.
  • Detail: the default production path is file-watch plus Fiber::reload through watcher::watch_many.
  • Detail: as of this change, every dylib load performs an exact ABI fingerprint handshake.
  • Detail: the plugin must export cordis_plugin_fingerprint.
  • Detail: the returned string must equal the host fingerprint exactly.
  • Detail: a missing symbol refuses the load.
  • Detail: a mismatched string refuses the load and names both fingerprints.
  • Rationale: stale libraries fail fast at load time.
  • Rationale: unchecked dylibs corrupt the process across the FFI boundary.
  • Source: load_plugin_so and FINGERPRINT_SYMBOL in crates/cordis/src/hmr.rs
  • Tests: load_plugin_so_rejects_missing_fingerprint
  • Tests: load_plugin_so_rejects_mismatched_fingerprint

6. Inventory collection is the primary registration path

  • Upstream expectation: registration walks an explicit hand-written factory list.
  • Claim: inventory collection gathers factories automatically as the primary path.
  • Detail: hand-written register_plugins chains remain as the fallback without the inventory feature.
  • Detail: the linker drops inventory nodes from crates that nothing references.
  • Detail: parity tests force-link every contributing crate before collection.
  • Rationale: collection deletes a hand-maintained list and its drift bugs.
  • Rationale: the linker failure is silent, so it needs a written warning.
  • Source: register_inventory_factories in crates/cordis/src/lib.rs
  • Test: inventory_registry_matches_expected_factory_set in tests/inventory_parity.rs

7. Serial dispatch aliases Bail, waterfall composes handlers

  • Upstream expectation: serial dispatch differs from bail semantics.
  • Claim: Dispatch::Serial is a direct alias of Dispatch::Bail.
  • Detail: both variants run the same run_bail_handlers code path.
  • Detail: waterfall is around-middleware.
  • Detail: every waterfall handler receives a real next continuation.
  • Detail: the terminal next runs the core operation.
  • Detail: no sentinel value ever stops a chain.
  • Rationale: one shared bail mode removes a near-duplicate implementation.
  • Rationale: around-middleware composition matches the paper shape directly.
  • Source: EventsService::dispatch in crates/cordis/src/events.rs
  • Tests: serial_stops_at_first_non_null_result
  • Tests: waterfall_around_short_circuit_skips_core

8. Worker supervision uses reserved exit codes plus stdin EOF

  • Upstream expectation: the paper defines no process supervision model.
  • Claim: supervised workers terminate through a fixed exit-code protocol.
  • Detail: EXIT_RESTART (51) asks the daemon for a fresh worker.
  • Detail: EXIT_QUIT (52) stops without a restart.
  • Detail: EXIT_BOOT (53) surfaces boot failure non-zero to the manager.
  • Detail: codes sit in the 51-53 band, clear of shell (1-2) and panic (101) codes.
  • Detail: workers set CORDIS_SUPERVISED watch stdin; EOF means the daemon died.
  • Rationale: pipe EOF is the only loss-free signal that survives daemon SIGKILL.
  • Source: crates/cordis/src/worker.rs, src/supervisor.rs
  • Tests: exit_codes_are_distinct
  • Tests: child_exit_codes_drive_loop

9. Log routing fans out through one exporter router

  • Upstream expectation: the paper defines no observability surface.
  • Claim: one router fans call records out to every registered exporter.
  • Detail: per-exporter level gates filter records before delivery.
  • Detail: exporter failures stay contained; inference never fails on them.
  • Detail: registration validates once; duplicate registrations are skipped.
  • Rationale: a single fan-out point replaces ad hoc sink plumbing per consumer.
  • Source: ExporterRouter in crates/ares-llm/src/exporter.rs
  • Tests: router_fans_out_to_all_exporters
  • Tests: accepts_gate_filters_records

10. Dependency withdrawal is reversible for working fibers

  • Upstream expectation: a fiber whose provider disappears rests Inactive (or is disposed) and never comes back on its own.
  • Claim: a previously-working runner fiber whose dependency genuinely vanished disposes its effects LIFO under Unloading and rests a new Pending state; when the provider returns it reactivates through Loading.
  • Detail: Pending is reserved for reactive waiting only — an apply error still rests terminal Failed{error} (row 1), and a peer-version constraint refusal over an existing-but-incompatible provider still rests Inactive because the provider remains available.
  • Detail: eligibility requires one fully-satisfied refresh pass first; registration cannot mark a fiber eligible.
  • Detail: Pending fibers reserve their registry key and survive prune_disposed, so reactivation needs no re-registration.
  • Rationale: the paper's permanently-Inactive outcome discards a healthy instance that only waits for its dependency; keeping it reversible preserves work.
  • Source: FiberState::Pending, the reactive-loss branch of Fiber::refresh in crates/cordis/src/fiber.rs
  • Tests: dependent_reactivates_when_provider_returns, failed_stays_failed_on_dep_return, pending_fiber_survives_prune_disposed

11. Dispatch participation knobs and filtered emits

  • Upstream expectation: listener registration has fixed semantics with no ordering or participation control.
  • Claim: flat listeners register through on_with / once_with with EventOptions { prepend, global }; emit_filtered runs a per-dispatch predicate over non-global listeners.
  • Detail: prepend: true inserts at the front of the dispatch-order list; global: true marks the listener realm-agnostic, and filters never exclude it.
  • Detail: a filter exclusion skips one dispatch without unregistering the listener.
  • Detail: the historical on / once / emit signatures delegate unchanged, and the broadcast bus fan-out is not filtered.
  • Rationale: per-realm policies need ordered, selectively-participating listeners without duplicating the bus.
  • Source: EventOptions, EventsService::on_with / once_with / emit_filtered in crates/cordis/src/events.rs
  • Tests: prepend_ordering_observed, filter_excludes_nonmatching_contexts, global_bypasses_filter

12. Reads during transitions are explicit and relaxed

  • Upstream expectation: every read either resolves an Active value or fails; mid-transition values are unreachable by construction.
  • Claim: strict Context::get keeps refusing providers resting in transitional states; Context::get_relaxed serves locally-owned values while their owner sits in Loading / Reloading / Unloading / reactive Pending.
  • Detail: terminal rest states stay refused even relaxed — disposed owners (undos already ran) and Failed{error} owners return nothing.
  • Rationale: lifecycle and observer code must inspect the value that is about to serve or was just retracted; making that a distinct method keeps the default read conservative.
  • Source: Context::get_relaxed in crates/cordis/src/context.rs
  • Test: relaxed_read_succeeds_while_provider_transitioning

13. Fiber state observers

  • Upstream expectation: the model defines no notification surface for individual fiber state changes.
  • Claim: Fiber::subscribe_state fans every lifecycle transition out to synchronous observers.
  • Detail: observers run inline under the short state-lock critical section and MUST NOT call back into the fiber.
  • Detail: observer panics are caught, so one broken observer cannot corrupt a transition; cancelled subscriptions are pruned on the next event.
  • Rationale: tooling (admin surfaces, tests, supervision) needs transitions as they happen, not just polling after quiescence.

14. Kernel operations are interceptable through meta-events

  • Upstream expectation: reads, writes, config resolution, restart schedules, and listener registration are fixed kernel behavior with no override points.
  • Claim: five veto meta-events (internal/get, internal/set, internal/config, internal/update, internal/listener) wrap those operations and an internal/dispatch observer reports every non-internal dispatch with (mode, name, args); the un-intercepted path is a zero-cost gate, and synchronous bridges FALL OPEN on runtimes that cannot park the worker.
  • Detail: internal/get — a non-null terminal replaces the value a strict read returns, {"refuse": true} fails the lookup outright, null passes, a chain error refuses the read; a redirect verdict continues the lookup at the parent frame.
  • Detail: internal/set — a chain error vetoes THIS write; the previous binding stays fully intact (no store/owners/version mutation).
  • Detail: internal/config — the chain's non-null terminal IS the effective config staged for that apply pass; a chain error rests the fiber terminal Failed{error} (row 1 semantics unchanged).
  • Detail: internal/update — a bail or explicit JSON false skips the restart; the fiber keeps serving its current application and the deferred config stays visible via vetoed_config.
  • Detail: internal/listener — a bail or chain error cancels the registration and the caller receives an INERT handle; neither registry ever sees the listener (fail-closed).
  • Detail: every consult checks listener_count == 0 first (map-lookup cost); a thread-local fence keeps operations made inside a chain un-intercepted; single-thread tokio flavors log a warning and fall open, matching historical behavior.
  • Detail: bail_from / waterfall_from / waterfall_async_from carry the operating context through an optional per-dispatch ListenerFilter; exclusions skip one dispatch without unregistering.
  • Rationale: policy layers need to observe and veto kernel operations without duplicating them; the zero-cost gate keeps the default path byte-identical for every existing caller.
  • Source: INTERNAL_*_EVENT constants, intercept_get/set/config/update/listener, bail_from / waterfall_from / waterfall_async_from, and the synchronous bridges in crates/cordis/src/events.rs
  • Source: consult points in Context::get / the provider-write path (crates/cordis/src/context.rs); config staging and the update veto in crates/cordis/src/fiber.rs
  • Tests: get_interceptor_rewrites_read, set_interceptor_vetoes_write_leaves_old_value, config_interceptor_rewrites_effective_config, update_interceptor_veto_skips_restart_keeps_config, listener_interceptor_bail_cancels_registration_inert_handle, internal_dispatch_observes_non_internal_only, interceptor_error_fails_fiber_activation, target_carrying_dispatches_filter_per_dispatch

15. Readiness gates wait quietly; availability predicates fail loudly

  • Upstream expectation: the model defines no way to hold a produced service out of rotation while its environment warms up (and nothing distinguishes that from failure).
  • Claim: register_with_readiness installs a composable ReadinessBarrier consulted before every activation pass; while it reports not-ready the fiber rests inspectable Pending — quiet waiting that NEVER becomes Failed — while availability predicates (Service::check) remain the loud complement resting Failed{error: "availability predicate rejected service"}.
  • Detail: ReadinessBarrier::new(pred) wraps one Fn(&Arc<Context>) -> bool; .and(other) AND-composes; with_readiness([a, b, c]) folds any number of barriers (an empty list is vacuously ready).
  • Detail: .watching([TypeId]) unions the provider keys whose settlements re-kick the gated fiber through the ReflectService fan-out — an external provide or withdrawal re-evaluates the gate without touching the fiber.
  • Detail: the factory runs once at registration (config errors still surface immediately); a closed gate only keeps the produced service OUT of consumer reach because strict get refuses non-Active owners; opening the gate activates without re-running the factory.
  • Rationale: not-ready-yet (warming caches, absent external system) differs fundamentally from broken; conflating them buries healthy waiting fibers under failure noise, and separating them lets operators read intent from state alone.
  • Source: ReadinessBarrier, with_readiness, register_with_readiness, and the re-kick wiring in crates/cordis/src/registry.rs
  • Source: the readiness consult in Fiber::refresh (crates/cordis/src/fiber.rs)
  • Tests: ready_when_holds_pending_until_true_then_activates, readiness_composes_and_semantics, external_rekick_reactivates_waiting_fiber

16. Concurrent config updates collapse to one cascade wave

  • Upstream expectation: every provider settle triggers its own full dependent refresh wave.
  • Claim: an in-flight ledger marks providers mid-reapply; dependents defer during the window and converge EXACTLY ONCE per settled batch.
  • Detail: CASCADE_INFLIGHT maps fiber id to open-window count (reentrant-safe); Loader::drive_fiber_update opens and closes windows around one live re-apply.
  • Detail: the kernel refresh path consults cascade_any_inflight, so a storm of racing patches costs one dependent apply pass and ends Active with the final config.
  • Rationale: N concurrent patches against one provider must not cost N dependent convergence waves.
  • Source: CASCADE_INFLIGHT, cascade_begin / cascade_end / cascade_any_inflight in crates/cordis/src/loader.rs
  • Test: concurrent_config_updates_collapse_to_single_cascade

17. Config pre-flight failures carry structured issues

  • Upstream expectation: configuration errors are lossy prose strings.
  • Claim: plugins reject configs with ValidationIssue { message, path } items aggregated in a ValidationError; CordisError::validation lifts the aggregate into the existing invalid config: class, and the loader trial stashes per-entry failures so the admin PATCH answers 4xx with a machine-readable issues array beside the legacy error string.
  • Detail: stash slots mirror the LATEST trial outcome; recording a non-validation error clears the entry and consumption removes it, so a later successful patch carries no issues.
  • Rationale: API consumers need to render field-level feedback, not parse sentences.
  • Source: ValidationIssue / ValidationError / trial stash in crates/cordis/src/error.rs; CordisError::validation in crates/cordis/src/service.rs; issue attachment in crates/ares-http/src/api/handlers/admin/cordis.rs
  • Test: patch_endpoint_returns_structured_issues_on_bad_config

18. The logger lives in the kernel crate

  • Upstream expectation: logging ships as a satellite console package beside the kernel.
  • Claim: the logger is adopted NATIVELY (cordis::logger) with upstream-style semantics: bounded ring, effect-owned exporter sinks, per-name level routing, printf rendering.
  • Detail: LoggerService keeps the last 1000 Messages (monotonic sequence, timestamp, name, kind, numeric level, args, fiber label) and snapshots without copying payloads.
  • Detail: exporters are effect-owned — register returns a Disposable whose disposal removes the sink; ExporterConfig gates per name and truncates rendered text (default cap 4096 chars, char-boundary safe).
  • Detail: thresholds resolve per-name pin, then the LoggerIntercept override (read through the relaxed channel, so per-fiber overrides apply on child contexts), then the default level (Debug); enabled bails BEFORE argument assembly.
  • Detail: rendering supports %s %d %i %f %o %O %c %C %%; unknown specifiers and exhausted arguments stay literal; %c picks a stable ANSI16 slot by FNV-1a hash of the logger name, %C adds bold; hyphenate / derived_name yield kebab-case logger names.
  • Detail: the Context facade (ctx.log/info/warn/debug/error/log_with) is a no-op when no logger is provided.
  • Rationale: observability belongs where fibers dispose, so sink lifetimes tie to effects instead of a satellite package boundary; the multi-package layering ceremony is deliberately not replicated.
  • Source: LoggerService, Exporter, ExporterConfig, LoggerIntercept, Message::render, hyphenate in crates/cordis/src/logger.rs
  • Tests: buffer_bounded_at_capacity_snapshot_reads, level_routing_per_name_with_default_fallback, printf_placeholders_format_correctly, logger_intercept_overrides_level, hyphenate_and_derived_names, exporter_disposal_removes_sink

19. Timer primitives are fiber-scoped and std-only

  • Upstream expectation: timing ships as a dedicated satellite package with its own runtime assumptions.
  • Claim: six primitives (timeout, sleep, interval, interval_stream, debounce, throttle) live natively in cordis::timer, run on ONE shared wheel thread, and attach to the owning fiber through labeled undos.
  • Detail: the wheel is a min-heap on a dedicated cordis-timer thread; due entries drain under one short critical section and callbacks run outside the lock; panics are caught and the thread survives.
  • Detail: registrations made under with_current_fiber push timer:-labeled undos, so Fiber::dispose (or a reactive unload) cancels them; dropping a handle does NOT cancel; out-of-scope registrations degrade to warned orphan handles that stay explicitly disposable.
  • Detail: a disposed Interval stream yields exactly ONE final Err(InactiveEffect) then closes; queued live ticks are discarded so teardown is the final observation.
  • Detail: debounce collapses a burst into one trailing delivery after the last call; throttle delivers leading-edge plus optional trailing in a fixed window.
  • Rationale: timers must die with the fiber that owns them or they leak firings past teardown; a shared thread keeps thousands of registrations at one thread's cost with no async runtime dependency.
  • Source: timeout / sleep / interval / interval_stream / debounce / throttle, with_current_fiber, Scheduled, Interval in crates/cordis/src/timer.rs
  • Tests: timeout_fires_once_and_disposes_with_fiber, timeout_dispose_before_deadline_prevents_fire, interval_ticks_repeatedly_and_stops_on_dispose, interval_stream_final_err_on_dispose, debounce_collapses_bursts, throttle_trailing_edge_respected

20. Accessor traffic bypasses interception

  • Upstream expectation: every value read or write consults the internal/get / internal/set veto waterfalls.
  • Claim: name-keyed computed properties (register_accessor) resolve OUTSIDE both waterfalls — resolving an accessor never consults or re-enters a veto chain.
  • Detail: Accessor::{read_only, read_write, setter_only} installs a getter/setter pair beside the TypeId service store; registration returns an EffectHandle whose disposal removes the declaration and every alias.
  • Detail: Context::alias binds an alternate name through the SAME registration; duplicate declarations (including alias collisions) are rejected with DuplicateProvider.
  • Detail: typed reads surface CordisError::PropertyTypeMismatch instead of a silent None; writes to a read-only property are refused with CordisError::ReadOnlyProperty.
  • Rationale: computed properties are policy plumbing, not provider state — vetoing them would let an interceptor break accessor invariants it cannot see.
  • Source: Context::register_accessor, Accessor, Context::alias, EffectHandle in crates/cordis/src/context.rs
  • Tests: accessor_read_write_roundtrip, duplicate_accessor_declaration_rejected, readonly_property_rejects_set, dispose_accessor_resolves_none, alias_resolves_same_value, accessor_bypasses_intercept_waterfalls

21. Intercept layers are ordered, append-on-set, and inspectable

  • Upstream expectation: one intercept binding per key; later registrations replace earlier ones (last-write-wins).
  • Claim: intercept layers per TypeId form an ordered outermost..innermost sequence; NEW registrations APPEND, so the innermost layer stays effective for all existing getters.
  • Detail: append-on-set means no existing caller observes a behavior change when another layer joins.
  • Detail: Context::intercept_chain(tid) returns every layer outermost..innermost for inspection and restart-decision logic.
  • Detail: Context::chains_structurally_equal compares two chains by shared-instance identity (Arc::ptr_eq) per layer pair; erased values carry no comparable contract, so freshly-built values compare unequal by design.
  • Rationale: layered policies need composition without clobbering, and restart decisions need an honest equality test over opaque layers.
  • Source: Context::intercept_chain, Context::chains_structurally_equal, layer storage in crates/cordis/src/context.rs
  • Tests: chained_layers_append_innermost_effective, intercept_chain_returns_all_layers_in_order

22. Restart errors keep the old application; vetoes defer loudly

  • Upstream expectation: an update-pass failure leaves the fiber in an unspecified mid-transition state.
  • Claim: Fiber::update returns Result<(), CordisError> — a restart-path error propagates to the caller and the fiber stays Active serving its OLD configuration; an internal/update veto parks the deferred config in Fiber::vetoed_config and returns Ok.
  • Detail: error propagation never disposes effects of the still-running application.
  • Detail: the vetoed config remains inspectable through vetoed_config() so operators can see what was declined and why.
  • Rationale: a failed restart must not destroy the working instance, and a silent skip must still be observable.
  • Source: Fiber::update, vetoed_config in crates/cordis/src/fiber.rs
  • Tests: update_error_stays_active_old_config, update_veto_defers_config_and_returns_ok

23. Config interception covers the activation path

  • Upstream expectation: config rewriting applies only on later re-applies; first activation runs the raw config.
  • Claim: the internal/config waterfall is consulted on the ACTIVATION path too, so rewrites apply on first activation identically to re-applies.
  • Detail: the same non-null-terminal-becomes-effective-config semantics hold on both paths (row 14).
  • Rationale: activation-time-only rewrites would make a policy's effect depend on whether the fiber happened to start fresh.
  • Source: config-waterfall consult in the register/activation path in crates/cordis/src/registry.rs
  • Test: config_waterfall_covers_activation_path

24. Module changes fan out through one dependency graph

  • Upstream expectation: no module-level change propagation exists; each watcher event maps to at most one reload target.
  • Claim: ModuleGraph maps module keys to dependencies; change_many computes the TRANSITIVE affected plugin set read-only FIRST, then reloads each affected plugin EXACTLY ONCE per transaction; a failing reload rolls back that plugin while successfully reloaded siblings stay Active.
  • Detail: ModuleReload implementations perform the reloads; ChangeOutcome classifies the transaction result.
  • Detail: the file watcher's debounced batch fans through a registered ModuleGraph when one is provided on the context; WITHOUT registration the watcher path is unchanged (opt-in).
  • Rationale: batched filesystem events must not reload shared dependents N times or leave siblings dead because one peer failed.
  • Source: ModuleGraph, ModuleEntry, ModuleReload, ChangeOutcome, change_many in crates/cordis/src/module_graph.rs; fan-out wiring in crates/cordis/src/watcher.rs
  • Tests: dependency_change_reloads_dependents_transitively, batched_changes_reload_each_plugin_once, rollback_keeps_successful_siblings_active, watcher_module_graph_fan_out_reloads_dependents, module_graph_without_registration_is_ignored

25. Entries relocate without losing fiber identity (we exceed upstream)

  • Upstream expectation: entries form a flat id-keyed list; relocation means delete-plus-recreate with a fresh fiber.
  • Claim: PATCH accepts optional parent / position (EntryPosition) applied move-THEN-update (invalid placements answer 409 before any mutation), POST /admin/cordis/entries/{id}/move relocates an entry with its whole {id}:* subtree in one rename cascade, and a valid move preserves fiber identity via in-place refresh.
  • Detail: moving into a descendant is refused; disabled groups move suppressed-then-restored.
  • Detail: invalid moves touch neither the entries file nor the live tree; unknown ids answer 404.
  • Rationale: hierarchical entry organization must not cost consumer-visible dispose/recreate windows.
  • Source: EntryPosition, EntryTree::move_entry, subtree_ids, Loader::move_entry in crates/cordis/src/loader.rs; patch / move_cordis_entry handlers in crates/ares-http/src/api/handlers/admin/cordis.rs
  • Tests: move_preserves_fiber_identity_and_lands_update_in_new_parent, descendant_move_refused, subtree_rename_cascades_descendants, disabled_group_move_suppresses_start_then_restores, patch_endpoint_moves_entry, patch_endpoint_invalid_move_conflicts_without_mutating

26. Subtask cancellation is wired end to end (we exceed upstream)

  • Upstream expectation: the reference client defines cancellation hooks but never wires them into its delegation loop.
  • Claim: delegated subtasks register sticky cancel tokens keyed by run/skill id; SkillEngine::cancel_subtask() flips a token exactly once and is honored at step boundaries alongside the EmergencyStop hook; an aborted subtask integrates nothing into the parent.
  • Detail: quote-aware delegation argument parsing: double-quoted segments are single tokens (backslash escapes inside quotes); --parallel latches split-per-token mode with | separators ignored, --model consumes exactly one token, --tools enables the inner tool loop; precedence is flags > profile > global.
  • Rationale: long-running delegated work needs an external off switch that lands between model rounds, not only at process exit.
  • Source: SubtaskCancelToken, SkillEngine::cancel_subtask, registered_cancel_token, step-boundary checks in crates/ares-agent/src/skills/engine.rs; parse_flags tokenizer in crates/ares-agent/src/skills/mod.rs
  • Tests: cancel_token_aborts_subtask_between_rounds, parse_flags_quote_aware_tokens

27. Deterministic micro calls are cached; repaired answers are not (we exceed upstream)

  • Upstream expectation: the reference roadmap describes response caching but ships none; every identical micro call re-hits the network.
  • Claim: deterministic-class micro outcomes serve from a bounded LRU map keyed by a content hash over (model, system template, input); answers reached through retries or salvage fallback are NEVER cached.
  • Detail: default 256 entries, 15-minute TTL, master switch via MicroCacheConfig.
  • Detail: hits skip the network entirely, report latency_ms: 0, and carry the cache_hit telemetry flag.
  • Rationale: classify/tag-style enrichment calls dominate micro traffic and their answers are stable; a repeated or repaired request proves the call was NOT deterministic-class, so caching it would pin a bad answer.
  • Source: MicroCacheConfig, MicroOutcome::cache_hit, cache_key, LruOutcomeCache in crates/ares-llm/src/micro.rs
  • Tests: identical_inputs_serve_cached_outcome, retries_exhausted_falls_back_to_salvage (salvage path stays uncached by construction)

28. Guided grammars ride typed hints with byte-identical absence

  • Upstream expectation: constrained output requires per-provider request surgery with no portable hint channel.
  • Claim: GenerationHints::guided_grammar carries a schema-shaped JSON value as response_format json_schema on every OpenAI-compatible path; raw GBNF/EBNF text rides the provider-specific guided_grammar extension field on NON-streaming OpenAI-compatible requests; providers without a channel silently ignore the hint; an ABSENT hint leaves the wire byte-identical.
  • Detail: classification is structural — a JSON object with an object root is a schema; anything else is raw grammar text.
  • Rationale: one opt-in hint field covers structured outputs where supported and vendor grammar extensions where they are not, without changing default requests.
  • Source: GenerationHints::guided_grammar in crates/ares-llm/src/client.rs; classification and GUIDED_GRAMMAR_EXTENSION in crates/ares-llm/src/openai.rs
  • Test: grammar_hint_present_reaches_request_builder

29. Duplicate embeddings cost exactly one backend call (we exceed upstream)

  • Upstream expectation: the reference roadmap mentions embedding dedup but implements none; every input in a batch hits the backend.
  • Claim: per-request content-hash dedup collapses duplicate inputs (whitespace-normalized SHA-256) BEFORE the backend call on BOTH local and HTTP embedding paths; computed vectors fan back to every duplicate slot.
  • Detail: DedupPlan maps duplicates to the first occurrence's slot; callers receive full-length results.
  • Rationale: identical texts in one batch are common (templates, retries) and each costs a paid embedding call.
  • Source: DedupPlan, content_hash_hex, normalize_for_dedup in crates/ares-rag/src/embeddings.rs
  • Tests: dedup_plan_maps_duplicates_to_first_occurrence_slot, duplicate_texts_single_backend_call_vectors_fanned_back, content_hash_hex_ignores_whitespace_differences_only

Properties we prove beyond the paper

crates/cordis/src/metatheory.rs proves five properties as executable checks. These hold regardless of the divergence choices above.

  1. Quiescence after every operation (quiescence_after_every_op). Every fiber rests in a well-defined state between operations. Transitional states appear only mid-await, never at rest. Allowed rest states include the reversible Pending (row 10) and terminal Failed{error} (row 1); only Active fibers must hold all declared injects available.
  2. Registration confluence (order_confluence_of_registrations). Registration order does not change the final graph.
  3. Reactive spatial invariant (dependent_never_active_without_provider). A dependent never activates while its provider is absent. It activates reactively when the provider appears.
  4. LIFO dispose restores the store (lifo_dispose_restores_store). Disposal unwinds effects in strict LIFO order. The store returns to its pre-registration contents.
  5. Version-conformance flips (version_conformance module). A compatible upgrade flips the dependent back to Active. A mismatch holds it at Inactive.

Maintenance note

Any pull request that changes kernel semantics MUST add a row here or update an existing row. State the claim, the rationale, and the enforcement point in the entry. Reviewers reject semantic kernel changes without a ledger entry. Before merge, re-run the cited tests.

Changelog

All notable changes to ARES are documented here. This project follows Semantic Versioning.


0.10.0 - 2026-08-26

Reactive fiber lifecycle, kernel interception points, guarded file writes, and opt-in intelligence controls.

Added

Kernel (ares-cordis)

  • Reactive Pending fiber state: an Active fiber whose dependency is genuinely withdrawn disposes its effects (LIFO) and rests Pending; it reactivates through Loading when the provider returns. Apply errors stay terminal Failed; peer-version refusals over a live provider still rest Inactive. Pending fibers reserve their registry key and survive prune_disposed
  • EventOptions { prepend, global } on the new on_with / once_with listener registrations, plus emit_filtered: per-dispatch filtering where non-global listeners are offered to a filter predicate and global listeners always join. Existing on / once / emit signatures are unchanged
  • Context::get_relaxed::<T>(): like get, but serves a locally-owned value while its provider fiber transitions (Active / Loading / Reloading / Unloading / Pending); disposed and Failed owners stay refused
  • Fiber state observers: Fiber::subscribe_state delivers every lifecycle transition to synchronous observers with panic isolation; the returned handle cancels the subscription
  • Kernel intercept meta-events: listeners on internal/get, internal/set, internal/config, internal/update, and internal/listener veto or rewrite the matching kernel operation, and internal/dispatch observes every non-internal dispatch with its (mode, name, args). internal/get can replace a strict read's value, refuse the lookup (refuse: true), or redirect it to the parent frame; an erroring chain refuses the read. internal/set errors veto the provider write with the previous binding left fully intact. internal/config's non-null terminal IS the effective config for that apply pass; an erroring chain rests the fiber terminal Failed (unchanged semantics). An internal/update bail skips the restart and keeps the current application, with the deferred config visible as vetoed_config. An internal/listener bail (or chain error) cancels the registration and returns an inert handle. Every consult short-circuits at map-lookup cost when no listener is registered, a re-entrancy fence keeps reads inside a chain un-intercepted, and the synchronous bridges fall open (warn + allow) on tokio flavors that cannot block_in_place
  • Target-carrying dispatch family: bail_from / waterfall_from / waterfall_async_from run the Bail / Waterfall / around-waterfall chains with an optional per-dispatch ListenerFilter; a filtered-out listener skips that one dispatch and stays registered
  • Readiness barriers: register_with_readiness takes a composable ReadinessBarrier — ReadinessBarrier::new(pred), .and(..) / with_readiness([a, b]) AND-composition (empty is vacuously ready), .watching([TypeId]) re-kicks the gated fiber when those providers settle through the ReflectService fan-out. While the gate is closed the fiber rests inspectable Pending — quiet waiting that never becomes Failed — with the factory run once up front and strict get keeping the service out of consumer reach. Complements (does not replace) availability predicates, which still fail loudly to Failed
  • Cascade batching: concurrent config updates against one provider collapse to a single dependent convergence wave. Providers mid-reapply are marked in an in-flight ledger; dependents defer during the window and converge once per settled batch instead of once per patch
  • Name-keyed computed properties: Context::register_accessor(name, Accessor::{read_only, read_write, setter_only}) installs a computed property beside the TypeId service store and returns an EffectHandle whose disposal removes the declaration and every alias. Context::alias binds an alternate name through the same registration. Typed reads surface PropertyTypeMismatch instead of a silent None; writes to a read-only property are refused with ReadOnlyProperty; duplicates (including alias collisions) are rejected. Accessor traffic BYPASSES the internal/get / internal/set intercept waterfalls entirely
  • Layered intercept chains: intercept layers per TypeId form an ordered outermost..innermost sequence; new registrations APPEND, so the innermost layer stays effective for all existing getters (no caller breakage). Context::intercept_chain returns every layer in dispatch order, and Context::chains_structurally_equal compares two chains by shared-instance identity for restart-decision checks
  • Lifecycle riders: Fiber::update returns Result<(), CordisError> — an error on the restart path propagates to the caller and the fiber stays Active serving its OLD configuration. An internal/update veto parks the deferred config in Fiber::vetoed_config and returns Ok. The internal/config waterfall now also covers the activation path, so rewrites apply on first activation, not only on re-applies
  • Module graph fan-out (opt-in): cordis::module_graph::ModuleGraph maps module keys to their dependencies and, given a ModuleReload implementation, change_many computes the TRANSITIVE affected plugin set read-only FIRST, then reloads each affected plugin exactly once per transaction; a failing reload rolls back that plugin while successfully reloaded siblings stay Active. When a ModuleGraph is registered on the context, the file watcher's debounced batch fans through it; without one, watcher behavior is unchanged

Logger

  • In-kernel LoggerService: bounded ring (default 1000 records) with monotonic sequences, fan-out to effect-owned Exporter sinks (registration returns a Disposable that removes the sink), and per-name level routing (set_level) with a service-wide fallback (set_default_level, default Debug). enabled gates writes before argument assembly. Message::render applies printf placeholders %s %d %i %f %o %O %c %C %% (%o/%O render compact/pretty JSON; unknown specifiers stay literal); %c colorizes over the ANSI16 palette by an FNV-1a hash of the logger name, %C adds bold. hyphenate / derived_name turn type names into kebab-case logger names (HTTPServer → http-server). LoggerIntercept overrides thresholds per fiber through ctx.intercept (resolved via the relaxed read); the Context facade (ctx.info, …) is a no-op when no logger is provided

Timers

  • cordis::timer: six fiber-scoped primitives — timeout, sleep, interval, interval_stream, debounce, throttle — std-only on a shared wheel thread (min-heap deadlines drained under one short critical section, callbacks outside it, panics caught). Registrations under with_current_fiber push labeled undos onto the owning fiber, so dispose or a reactive unload cancels them; dropping a handle does NOT cancel. A disposed Interval stream yields exactly one final Err(InactiveEffect) and closes; debounce collapses bursts to one trailing delivery, throttle

Admin HTTP

  • Structured validation errors on the same PATCH endpoint: a config pre-flight can reject with ValidationIssue { message, path } items aggregated in a ValidationError; the 4xx body then carries a machine-readable issues array beside the legacy error string (success carries none, and a failed trial leaves no stale slot behind)
  • Entry moves: PATCH accepts optional parent / position (EntryPosition) applied move-THEN-update — an invalid placement answers 409 without touching the file or the live tree. POST /admin/cordis/entries/{id}/move relocates an entry together with its whole {id}:* subtree in one rename cascade. A valid move preserves fiber identity through in-place refresh, so consumers never observe a dispose/recreate window

Tools fence

  • Layer 3 write guards on Fence: writes require a prior fence_read observation unless the mode allows blind writes (FS_NOT_OBSERVED); CreateIfAbsent refuses existing paths (FS_EXISTS); ReplaceIfVersion compares the mtime ^ size fingerprint captured at read time (FS_VERSION_CONFLICT). Bytes land through a sibling temp file renamed into place, new files get 0600 on unix, errors carry structured FS_* codes, and a bounded 200-entry audit ring is readable through audit_log. Layers 0-2 behave exactly as before

LLM

  • Retry-before-salvage JSON policy: the micro engine re-requests a malformed-JSON answer identically up to json_retries times (default 2) before the substring-salvage fallback runs
  • Per-provider concurrency governor: optional pool setting max_in_flight caps simultaneous dispatches per provider, with governor_acquire_timeout (default 30 s) bounding the wait. Permits release only at terminal stream items, and saturation fails closed. Without the setting, behavior is unchanged and no wrappers install
  • Model-profile catalog: one cross-provider ModelProfile table (capabilities, context window, speed tier, cost) merging the static tables with runtime catalog entries. lean_hint renders the whole catalog for prompt injection in well under 50 tokens, describe_full prints one record, and route picks the cheapest capable model for a modality. The catalog is opt-in; nothing wires it into default model selection
  • Guided-output grammar hints: GenerationHints::guided_grammar carries a schema-shaped value (JSON object with a "type": "object" root) as response_format json_schema on every OpenAI-compatible path; raw GBNF/EBNF-style text rides the provider-specific guided_grammar extension field on non-streaming OpenAI-compatible requests instead. Providers without a channel silently ignore the hint, and an ABSENT hint leaves the wire byte-identical
  • Micro-call response cache: deterministic-class micro outcomes are served from a bounded least-recently-used map keyed by a content hash over (model, system template, input) — default 256 entries with a 15-minute TTL and a master switch via MicroCacheConfig. Hits skip the network entirely, report latency_ms: 0, and carry a cache_hit telemetry flag. Answers reached through retries or the salvage fallback are NEVER cached

Agent

  • Per-subtask cancellation: delegated subtasks register sticky cancel tokens keyed by run/skill id; SkillEngine::cancel_subtask() flips a token exactly once and is honored at step boundaries alongside the existing EmergencyStop hook. An aborted subtask integrates nothing into the parent context
  • Quote-aware delegation arguments: double-quoted segments parse as single tokens that may contain spaces and | separators; --parallel latches split-per-token mode (separators ignored), --model consumes exactly one token, and --tools enables the inner tool loop for delegated tasks. Precedence is flags > profile > global

RAG

  • Embedding dedup per request: duplicate inputs collapse by whitespace-normalized SHA-256 content hash before the backend call on both the local and HTTP embedding paths; computed vectors fan back to every duplicate slot, so callers receive full-length results while identical texts cost exactly one backend call

Skills

  • Delegated-result review gate (opt-in review_delegated_results): nested SkillCall results pass a fixed-template consistency and task-fit review before integration. A rejection replaces the result with a structured rejection that keeps the original for re-dispatch; a reviewer outage passes the result through unchanged. Off by default
  • Self-check critique rounds (opt-in SkillEngine::with_self_check_rounds): nested SkillCall results pass up to N LLM critique rounds over a cache-stable template before integration; a verbatim reply ends the loop, and an LLM failure keeps the last good answer silently. Off by default
  • Delegation hygiene: delegated sub-workflows accept only allowlisted step kinds (delegated_step_not_allowed:), nested tool rounds hard-cap at three (tool_round_cap_exceeded: aborts after exactly three rounds), and slash-command chatter lines are stripped from delegated result text before it enters the parent context

0.9.0, 2026-08-22

Service architecture, unified execution, wrapper removal.

Added

  • Execute::run with the full resolve-create-execute pipeline and RunTracker observability
  • EventsService::waterfall_around: around-middleware waterfall that runs core at the end; a skip of next skips core
  • Context::inject waits on the ReflectService TypeId notifier (ensure_notifier + changed); if ReflectService is absent or the sender is dropped, it falls through to a 5ms poll
  • Product events: tools.list / tools.resolve / tools.execute, llm.get_client / llm.complete, llm.generate / llm.generate_tools (ConfigurableAgent), agent.run (waterfall), agent.admit (Dispatch::Bail), agent.started (Dispatch::Parallel)
  • Skills isolate the request ctx (isolate::<Tools>(tenant_id)) instead of opening a new root
  • Skill LlmCall steps strictly run Llm::complete through the llm.complete waterfall; SkillEngine and SkillsService have no direct provider generate_with_history fallback
  • Skill ToolCall steps run Tools::execute (tools.execute waterfall) on the tenant isolate
  • ExecutionResult return type with resolution metadata (source tier, run ID)
  • RunTracker trait extracted to ares-agent for decoupled run observability
  • Service impl directly on AgentRegistry and ConfigBasedLLMFactory (no wrappers needed)
  • agent_config_from_user_agent helper in ares-agent::configurable
  • Fiber::refresh reruns registered plugin apply after epoch recompute
  • EventsService Parallel returns JSON null; Serial bails on the first non-null handler result
  • Store loader factory runs SQL migrations and seeds agent templates
  • Overlay fills empty loader entry.config from ares.toml; TOON reloads notify Tools and Execute
  • TenantRealms open-then-intercept on request paths; dispose on admin tenant delete
  • JWT research plus remaining v1 stream/agent handlers open the tenant realm before intercept
  • Isolate labels win over intercept for the same TypeId; unlabeled types still intercept
  • Leftover execution_stack dual Execute installer removed
  • Default ares-server library build has no axum (http is optional) and no longer re-exports ProviderRegistry
  • Single Execute loader key (ares-agent); Overlay/ServerRuntime provide host extras
  • JWT middleware looks up tenant claims in Store, fail-closes 401 when the tenant does not exist, then opens TenantRealms and intercepts TenantContext; user claims isolate with no dummy Free tenant
  • Llm::from_client is the public test constructor; no_http no longer builds ProviderRegistry
  • Root ares-server package keeps its binary; the library target serves embedders; integration tests depend on ares-server / ares-http

Changed

  • All 5 execution sites (chat, v1, scheduler, trigger, pipeline) now delegate to Execute
  • Tools, Llm, and Execute public methods run through Cordis waterfall_around when EventsService is on ctx
  • SkillsService and SkillEngine LLM/tool steps use those same events instead of calling the tool or client directly
  • resolve_agent delegates to crate-private Resolver when available (legacy fallback retained)
  • Removed AgentRegistryService and LlmFactoryService wrappers (consumers use types directly)
  • Deleted deprecated start_background_reload function and its test
  • Calculator tool registered via ctx.plugin(CalculatorService) in addition to legacy path
  • Version bump to 0.9.0
  • Tools, Llm, Execute, and skills remain event-first on EventsService waterfalls
  • run_server still instantiates Overlay first, then remaining loader entries
  • Scheduler, pipeline, and trigger domain loops remain native ARES engines behind Execute
  • ProviderRegistry remains on ares-llm for Llm::new / AgentRegistry::from_config
  • Overlay lives in crates/ares-http/src/overlay.rs; the server still registers the Overlay factory

0.8.0, 2026-08-21

Service-based architecture with dependency injection.

Added

  • cordis crate: typed Context container, Fiber lifecycle, Service trait, RegistryService with plugin pattern, Loader with config reconciliation, EventsService with 5 dispatch modes, ReflectService for hot-reload coordination.
  • Unified services: UnifiedToolService (merges static, runtime, and MCP tools), LlmService (circuit breaker with failover), AgentResolverService (3-tier resolution: tenant, community, system).
  • Handler migration: 177 handlers moved from State<AppState> to State<Arc<Context>> with ctx.get::<T>().
  • Admin API split from single 190KB file into 15 domain-specific modules.
  • V1 API split from 73KB file into 5 modules.
  • File-watch hot-reload with 500ms debounce (replaces 60s polling).
  • Rhai scripting support for custom tools and services.

Changed

  • AppState god-struct (17-22 fields) replaced by pub type AppState = Arc<Context>.
  • build_router(ctx) is the primary router constructor; base_router remains as deprecated shim.
  • Rust toolchain updated to 1.98.
  • All docs humanized (removed AI-sounding prose patterns).
  • docs/src/SUMMARY.md now includes Cordis chapters (mapping, remedies, capabilities, baseline, YAGNI, redesign) plus Architecture.
  • mdBook GH-pages rebuilt for 0.8.0 (gh-pages branch docs: rebuild gh-pages book for 0.8.0 Cordis).

0.7.3

Previous release line (see git tags). Changes tracked in git history before changelog formalization.

0.6.3

Multi-provider LLM, tenant agents, and enterprise metering.

This release transforms ARES from a single-provider system into a full multi-provider LLM platform with enterprise-grade tenant management.

Added

  • Multi-provider LLM routing: support for 4 providers (Groq, Anthropic, NVIDIA DeepSeek, Ollama) and 11 models through a unified API.
  • Model tier system: fast, balanced, powerful, deepseek, and local tiers with automatic provider routing.
  • Tenant agent system: agents stored in the database per tenant. Template-based provisioning with full CRUD via admin API.
  • Agent templates: seed templates applied automatically on startup. New tenants receive a default agent set.
  • Usage metering: usage_events table, monthly_usage_cache, and daily_rate_limits for tracking tokens, requests, and costs per tenant.
  • API key authentication: Authorization: Bearer ares_xxx on /v1/* routes with tenant scoping.
  • Enterprise agent templates: 4 specialized agent templates (trade-classifier, trade-risk, trade-monitor, trade-reporter) for the first enterprise deployment.
  • Tenant-scoped API routes: both JWT-protected (/api/trading/*) and API-key (/v1/trading/*) endpoints.
  • Admin provisioning API: atomic tenant creation: schema + agents + API key in a single operation.

Changed

  • Chat handler now resolves tenant_id from authentication context instead of hardcoded values.
  • Provider configuration moved from code to ares.toml for runtime flexibility.
  • Rate limit enforcement now operates at both the provider and tenant level.

Fixed

  • Chat handler tenant_id resolution for multi-tenant requests.

0.6.2

Streaming and SSE support.

Added

  • Server-Sent Events streaming: POST /v1/chat/stream endpoint for real-time token-by-token responses.
  • Stream handler: unified streaming across all providers with consistent SSE format.
  • Context continuation: context_id parameter for maintaining conversation history across requests.

Changed

  • Response format standardized to {"response", "agent", "context_id"} across all endpoints.

0.6.1

Tool calling and RAG foundations.

Added

  • Tool calling framework: define tools per agent. ARES manages the tool-call loop, execution, and response assembly.
  • RAG pipeline: retrieval-augmented generation with pluggable document stores.
  • Workflow engine: chain multiple agents into multi-step workflows with deterministic execution.

Changed

  • Agent configuration schema extended to support tool definitions and RAG settings.

0.5.0

JWT authentication and user management.

Added

  • User registration and login: POST /api/auth/register, POST /api/auth/login.
  • JWT token lifecycle: 15-minute access tokens, refresh token rotation, logout/invalidation.
  • Role-based access: user roles with permission checks on protected routes.
  • Admin authentication: X-Admin-Secret header for internal administration endpoints.

Changed

  • All /api/* routes now require JWT authentication.
  • Error responses standardized with error and code fields.

0.4.0

PostgreSQL backend and multi-tenant schema.

Added

  • PostgreSQL integration: full migration from in-memory storage to PostgreSQL with sqlx.
  • Auto-migration: sqlx::migrate!() runs on startup. No manual SQL required.
  • Tenant schema: tenants, tenant_agents, and api_keys tables with foreign key relationships.
  • Tenant tiers: Free, Dev, Pro, and Enterprise tiers with configurable limits.

Changed

  • All state persistence moved from in-memory structures to PostgreSQL.
  • Connection pooling via sqlx::PgPool with configurable pool size.

For the complete commit history, see the ARES repository on GitHub.