Skip to main content

What is a benchmark?

A benchmark is a Command that runs your code on a workload that looks like real use and reports numbers. Artemis runs it on the Baseline, your unchanged code, and on each changed Version, on the same Runner, and compares the results.

The workload and the fixed checks​

The workload is the part you own. If it is too small or unlike real use, a Version can look faster on the benchmark and be no faster in production. A few different inputs usually tell you more than one.

The benchmark sits in your Script next to the build and test Commands. The Script is a Project setting: the agent that searches for improvements writes code against it and does not change it during a Discovery Run. A Version counts only if it builds and passes the tests as well as scoring better on the benchmark. Passing tests shows the Version behaves correctly on what the tests cover. It cannot prove the code is right for every possible input.

Metrics, direction and importance​

Artemis records runtime, CPU and memory for each benchmark run by itself. Anything else, such as throughput, latency or accuracy, your benchmark writes to a small results file (artemis_results.json or .csv).

Each Metric has a direction. Throughput is one to maximise (Higher is better in the UI); latency and memory are ones to minimise (Lower is better). When there are several Metrics, each gets an importance tier of High, Medium, Low or None, and Artemis combines them into a Composite score against the Baseline to rank Versions. The Composite score is a summary for ranking, not a percentage speed-up. A Metric set to None is still measured and shown, but it does not count toward the score.

Improvements often come with trade-offs. A Version can be faster and use more memory. The per-Metric results show both, so you can decide whether the trade is worth it.

Why run it more than once​

Timings vary between runs of the same code. Other processes, caches, CPU frequency and the operating system all add noise, so a single run of each Version can make a small difference look real.

Artemis can repeat the benchmark Command, either a fixed number of times or until the results settle, and reports each Metric's mean and spread along with the change against the Baseline. Only the benchmark repeats; setup Commands such as the build run once. If the difference is smaller than the noise, treat the result as inconclusive and run more repeats. For Versions that are close, Discovery recommends how many repeats would settle the comparison.