Run and gate evals from your terminal, CI, or agentic harness in minutes.
No credit card · bring your own key · read the docs
Your harness + observability + eval + gating glue.
Your harness + Autoeval.
evaluate · measure reliability · gateUpstream · your stack
e.g. DeepSeek Harness, coding agents
The agent does the work and the run is already the artifact: scenarios, conversations, and agent traces come straight out of the harness. No logging pipeline to wire up first.
Local · open source
The CLI. Author, run, and gate a run from your terminal, in CI, or driven by a coding agent over MCP. Bring your own OpenAI, Anthropic, or Together.ai key.
Server · the reliability engine
Calibrated judges grade to your experts' bar, several judges score each answer independently, and scenarios get statistical multi-run evaluations with adaptive stopping and confidence intervals.
The engine behind Plumloom Autoeval is EvalCore. It turns a raw model score into a result you can act on.
Judges grade against anchors set by your own experts, so the standard stays consistent from run to run and does not drift over time. A 4.2 means the same thing every time.
The agentic harness already gives you the operational loop: the agent acts, the result is judged, and the harness decides what to keep and what to improve next. It is fast, and it is autonomous.
Every decision in that loop rests on the judgment underneath it. Judge with a single run and the same case can pass one moment and fail the next, so the loop moves quickly on a signal that is part noise.
The harness keeps, promotes, or discards based on a verdict that may not hold. One moment a change looks like an improvement; the next, it is a regression you cannot see yet.
Plumloom makes the judgment step reliable: calibrated to your experts and scored by multiple judges. For scenarios, it adds statistical multi-run evaluations with adaptive stopping and confidence intervals. The loop keeps its speed, and its decisions now rest on evidence that holds.
Plumloom makes the judgment step reliable without slowing the loop.
Frameworks run your evaluation and give the runs back, leaving the judgment to you. Plumloom returns the judgment: calibrated, scored by multiple judges, and for scenarios using statistical multi-run evaluations with adaptive stopping and confidence intervals.
Observability shows what happened after a response ships, and needs your app instrumented first. Plumloom tells you whether to ship at all, with a provider key and nothing to install.
Run one evaluation from your terminal and see whether the number holds before you ship it.
No credit card · bring your own key · read the docs