TALOS
An agent that writes the tools it's missing, tests them, and keeps the ones that work.
Five steps, one loop
Each step is its own node in the graph, so you can watch every run happen.
- PLAN
Splits your request into sub-tasks and checks the vault before building anything.
- FORGE
Writes a typed Python function for the gap, researching free APIs first.
- TEST
Runs generated tests in a subprocess. Failures go back to Forge.
- EXECUTE
Runs the tool on your real input.
- LEARN
Saves it to the vault so next time it skips straight to Execute.
Ask once, it builds. Ask again, it remembers.
The same kind of question, twice. The first run writes three tools. The second writes none.
Illustrative replay of how a session looks.
Built as a state graph
Talos runs on LangGraph. Every agent below is a node with typed state, and every run is a trace you can open in LangSmith.
Primitives and vault tools go straight to the Executor.
needs a new toolA failing test sends the trace back to the Forger, up to 3 times.
Human check pauses for a missing API key. Learn saves the tool to the vault.
Another sub-task left? Back to the Planner.
Calls a model. Everything else is plain Python.
Planner
Labels every sub-task as a primitive, a vault tool, or something to forge, before anything runs.
Forger and Tester
Code and tests come out of one structured call. The tester runs them in a 10-second subprocess and hands back the trace.
Researcher
A ReAct agent the Forger calls to read API docs with Tavily search and Jina reader before writing web code.
Human check
If a tool needs an API key, the graph pauses, asks you once, and saves it for next time.
Executor
Fills in arguments from your query and earlier results, then runs the tool.
Skill vault
A JSON manifest and a folder of .py files. No database, nothing hidden.
- LANGGRAPH
- LANGCHAIN
- LANGSMITH
- TAVILY
- JINA READER
- OPENROUTER
tools in the vault, every one written by Talos
None were written by hand. They came out of test runs, passed their own tests, and got reused.
- fetch_github_star_counts(repo1, repo2)
- levenshtein_distance(s1, s2)
- haversine_distance(lat1, lon1, lat2, lon2)
- json_to_yaml(json_str)
- fetch_latest_python_version()
- count_http_status_codes(log_text)
- prime_factorization(n)
- normalize_dates(dates)
- fetch_bitcoin_price_usd()
- score_password_strength(password)
- flatten_json(data)
- extract_style_rules(url)
- caesar_cipher(text, shift, mode)
- add_revenue_column(input_path, output_path)
- fetch_top_trending_repo()
- hex_to_rgb(hex_code)
- slugify(title)
- validate_email(email)
Reaches the web through a free API it found itself
Evaluated on 62 end-to-end queries
The same suite, before and after the latest round of changes.