What is a Discovery?
A Discovery Run is a session in which Artemis explores many possible changes to your code toward one Objective and keeps the evidence for each one. You can start one from the Discover tab, from the Artemis CLI, or from your coding agent.
How it works
You describe the Objective, for example "make the simulation step faster without changing its output". Artemis proposes Experiments, each an approach worth trying, and several models generate Discovery Versions of the code for them. Models also review the ideas before budget is spent on them. Each Version is scored against the Baseline, the unchanged code, and each Experiment ends as Validated, Refuted or Inconclusive.
Promising Versions are built on in later Experiments. The rest stay visible with their scores, so you can see what was tried and why it was dropped. Many Versions fail or show no useful gain, and that is expected.
You can steer while it runs: approve or dismiss Experiments, add your own ideas, redirect it, or add budget. Effort and Version limits bound how much work it does. When it finishes you review the best Versions, read the diffs, and decide what to merge.
Measured runs and model-judged runs
A Discovery Run works in one of two ways, and the difference matters when you read the numbers.
With a Runner and a benchmark, every Version is built, tested and benchmarked on your Runner. The scores are measurements from your hardware.
With no Runner, for work that cannot be executed such as a Markdown Plan, Artemis turns your Objective into scoring criteria and a panel of judge models scores each Version against the Baseline. These are model-judged scores. They are useful for comparing ideas, but they are not measurements, so never read them as a speed-up. Measure code on a Runner before you rely on a performance claim.