The Talos mark: a crystalline T split by a vertical spine, inside concentric construction circles

TALOS

An agent that writes the tools it's missing, tests them, and keeps the ones that work.

Five steps, one loop

Each step is its own node in the graph, so you can watch every run happen.

vault PLAN FORGE TEST EXECUTE LEARN retry up to 3×
  1. PLAN

    Splits your request into sub-tasks and checks the vault before building anything.

  2. FORGE

    Writes a typed Python function for the gap, researching free APIs first.

  3. TEST

    Runs generated tests in a subprocess. Failures go back to Forge.

  4. EXECUTE

    Runs the tool on your real input.

  5. LEARN

    Saves it to the vault so next time it skips straight to Execute.

Ask once, it builds. Ask again, it remembers.

The same kind of question, twice. The first run writes three tools. The second writes none.

First run3 tools forged
talos › compare GitHub stars for langgraph and crewai
plan3 sub-tasks, 3 need new tools
vaultno match
forgefetch_github_star_counts()
testpassed
forgepercentage_difference()
testfailed, ZeroDivisionError, retrying
testpassed
forgeformat_repo_comparison()
testpassed
learn3 tools saved
executedone
Next day0 tools forged
talos › compare GitHub stars for fastapi and flask
plan3 sub-tasks, all in the vault
vaultfetch_github_star_counts
vaultpercentage_difference
vaultformat_repo_comparison
executedone

Illustrative replay of how a session looks.

Built as a state graph

Talos runs on LangGraph. Every agent below is a node with typed state, and every run is a trace you can open in LangSmith.

primitive already in the vault fails: retry, max 3 needs an API key all done next sub-task Forge sub-graph Orchestrator Planner Primitive Vault tool Forger Tester Learn Human check Executor Answer

Calls a model. Everything else is plain Python.

Planner

Labels every sub-task as a primitive, a vault tool, or something to forge, before anything runs.

Forger and Tester

Code and tests come out of one structured call. The tester runs them in a 10-second subprocess and hands back the trace.

Researcher

A ReAct agent the Forger calls to read API docs with Tavily search and Jina reader before writing web code.

Human check

If a tool needs an API key, the graph pauses, asks you once, and saves it for next time.

Executor

Fills in arguments from your query and earlier results, then runs the tool.

Skill vault

A JSON manifest and a folder of .py files. No database, nothing hidden.

45

tools in the vault, every one written by Talos

None were written by hand. They came out of test runs, passed their own tests, and got reused.

Reaches the web through a free API it found itself

Evaluated on 62 end-to-end queries

The same suite, before and after the latest round of changes.

Suite score 56/62 was 45/62
Search and reading questions 5/5 was 0/5
Queries that crashed 1 was 3
Median time per query 17.1s was 13.9s