Artemis and your coding agent
Coding agents such as Claude Code and Codex write, run and test code well. Artemis works alongside them. You stay in your local agent, and it hands work to Artemis when the job needs measurement on real hardware, many attempts, or a record that lasts after the chat.
What Artemis adds
Measured verification on your hardware
When your agent says a change is faster, Artemis can build, test and benchmark the Baseline and the change on your own Runner, repeat the benchmark, and report each Metric with its spread. The Script with the benchmark and tests is a Project setting, and the agent does not change it during a run.
Many candidates
A Discovery Run generates and scores many Versions with several models, builds on the promising ones, and keeps the failed ones visible.
Evidence that is kept
Experiments, scores, diffs and agent traces stay with the Project in Artemis after the conversation ends, and each Pull Request comes with its numbers.
How they connect
Your agent talks to Artemis through the Artemis CLI. Any agent that can run shell commands can use it to start a Discovery Run, run a Validation on a Runner, run a Maintain Scan or read results. Neither the CLI nor your agent is the Runner: the Runner is where Artemis runs the code, and it can be on another machine.
For Claude Code there is also the Artemis skills plugin, which teaches Claude Code how to use the CLI for these jobs. With Codex and other agents, use the CLI directly. From Maintain, Copy for local agent gives your agent an Issue to work on. The CLI and the skills are versioned separately, so pin both if you script against them.
Four ways to use it
- Verify a change your agent made, on your hardware.
- Delegate a Discovery and keep coding while it runs.
- Find what's worth fixing with a Maintain Scan.
- Review a plan before your agent builds it.