Why North separates internal targets from independent benchmarks and keeps caveats attached to every claim.
Separate the dimensions
Research quality contains several different problems: source discovery, claim extraction, factual support, synthesis, calibration, latency, and cost. One preference score cannot explain which part improved.
North evaluates these dimensions separately and treats prose style as a presentation layer. A system should not receive factual credit because its answer is fluent.
Reproducible evaluation records
Every reported result should identify the model version, prompt or policy version, retrieval configuration, evaluator, sample set, and date. Where a judge model is used, its identity and known biases belong in the record.
Small hand-audited sets are used to catch rubric failures before scaling automated evaluation. Disagreements between human and automated graders are retained rather than averaged away.
Publication rules
Internal targets are labeled as targets. Internal test results are labeled as internal. Third-party benchmarks are linked to their source and are not reformatted into stronger claims than the source supports.
A result is publishable when another person can understand what was tested, what was excluded, and what could invalidate the conclusion. Eloquence comes after that threshold, not before it.