Ablitron / HomeABLITRON RESEARCH / 001

Evidence.
Before superlatives.

A coding agent should be judged by the work it completes. Here is our proposed evaluation framework, published before the results.

Methodology draft · No results published

We have not run or published an Ablitron benchmark. The graphic on the homepage is illustrative and represents no measurements.

The question

Does an abliterated model complete legitimate coding tasks more reliably than its original checkpoint when both run in the same agent environment?

We want to distinguish changes in model behavior from changes in the agent, tools, prompts, and inference settings. Comparisons should hold those conditions constant whenever possible.

What we plan to measure

MeasureEvidence
Task completionTask-specific checks against the final repository state.
Instruction adherenceA rubric for explicit constraints, scope, and output requirements.
Unnecessary refusalHuman-reviewed refusal labels on clearly legitimate tasks.
Unrequested changesDiff review for edits outside the requested scope.
Cost and timeRecorded usage, price assumptions, and end-to-end runtime.
RecoveryWhether the agent corrects a failed check within the run budget.

A reproducible run

  1. Freeze the task, repository commit, environment, and acceptance criteria.
  2. Record the model checkpoint, quantization, prompt, tools, and sampling settings.
  3. Use the same time and tool budgets for paired comparisons.
  4. Repeat runs to measure variability; report failures and timeouts.
  5. Publish permitted task data, run logs, costs, and the scoring method.

How we will report results

We plan to show absolute task counts, denominators, uncertainty, and per-task outcomes. We will identify who built and funded the evaluation, including our interest in Ablitron, and explain limitations rather than presenting one universal “best model” score.

What this will not prove

A narrow coding test does not establish overall model safety, capability, or suitability for every repository. Public task contamination, small samples, judging errors, and provider changes can all affect results.

See the build roadmap