Case Studies
MethodMethods note 0.3

Evaluation Before Eloquence

Why North separates internal targets from independent benchmarks and keeps caveats attached to every claim.

By North MLAugust 4, 20266 minute read
NORTH / RESEARCH RECORD140
Abstract

Why North separates internal targets from independent benchmarks and keeps caveats attached to every claim.

01

Separate the dimensions

Research quality contains several different problems: source discovery, claim extraction, factual support, synthesis, calibration, latency, and cost. One preference score cannot explain which part improved.

North evaluates these dimensions separately and treats prose style as a presentation layer. A system should not receive factual credit because its answer is fluent.

02

Reproducible evaluation records

Every reported result should identify the model version, prompt or policy version, retrieval configuration, evaluator, sample set, and date. Where a judge model is used, its identity and known biases belong in the record.

Small hand-audited sets are used to catch rubric failures before scaling automated evaluation. Disagreements between human and automated graders are retained rather than averaged away.

03

Publication rules

Internal targets are labeled as targets. Internal test results are labeled as internal. Third-party benchmarks are linked to their source and are not reformatted into stronger claims than the source supports.

A result is publishable when another person can understand what was tested, what was excluded, and what could invalidate the conclusion. Eloquence comes after that threshold, not before it.

Continue exploring North ML.

Return to Case StudiesOpen Horizon ↗