How good is Bifrost at finding usages?
IMMUTABLE EVIDENCE MAP · USAGEBENCH
The immutable v1 result and its generated comparison page are the authority for v1 scores. The independent v2 result is published from v0.3.0; its generated report pages and hashes remain separately identified.
The reviewed legacy core remains pending a separately verified release.Published historical release v0.2.0; 12 per profile
Immutable result v0.3.0
10 × 11 languages; retrospective only; report pending
42 overflow · 6 controls · 2 semantic-pack cases
The generated evidence map shows every profile and language denominator, the frozen manifest checksums, and the safeguards that keep v1, v2, the reviewed legacy core, controls, overflow, and development cases from being pooled. The v1 score below is historical release evidence; v2 and legacy status/provenance are derived from immutable bundles when supplied; all score tables must be generated from checksum-verified reports.
What this result means
Section titled “What this result means”UsageBench is Bifrost’s recurring conformance and regression suite for source usages and navigation. Established language servers provide strong comparison evidence, but the reviewed source contract remains the authority: Bifrost should match a reference server where that behavior satisfies the contract and may preserve a narrower or broader result when the source evidence requires it.
Language servers primarily support interactive development inside configured editor workspaces. Bifrost is a code-analysis and navigation substrate for repository-scale consumers, especially coding agents and static-analysis tools. UsageBench measures their overlapping References and Definition surface; it does not compare completion, diagnostics, refactoring, or either product’s complete feature set.
This result is the independently reviewed, source-locked
real-project-v1 partition published in immutable release
v0.2.0. Its
bounded claim covers only descriptive, per-profile comparisons for 12 sampled
public repositories and 36 preregistered cases. It does not estimate
language-wide accuracy, ecosystem-wide superiority, latency, memory, or
cold-start performance.
The generated v1 result page and case comparison are the score authority. They are regenerated and provenance-checked from the immutable release bundle during publication; this overview intentionally does not duplicate score totals.
Why this result is publication-qualified
Section titled “Why this result is publication-qualified”The cases were selected before analyzer execution, checked through blinded OpenAI and Anthropic review sessions, adjudicated by an accountable human, and bound to immutable source archives. The release audit hashes the protocol, selection, review records, adjudication, source lock, runner profiles, reports, and exact UsageBench revision.
The broader 158-case fixture suite remains valuable development and regression evidence, but it was selected differently and is not pooled with this evaluation. Its corrected 24 July comparison remains available as an explicitly historical result.
The newest frozen evaluation is published as the current result, with its own profiles and denominators; read it for location-level precision and recall, versions, exclusions, replacements, and artifact hashes. Its case comparison lists every strict result unique to one side. Each frozen slice keeps a separate denominator and none of them pool. For scoring and claim boundaries, see the methodology; for rerunning the evidence, use the reproduction guide.