Overview

Which setup should I run?

Next action

Loading workspace state

InferGrade is checking account, runner, and evidence readiness.

Runner

Get ready to run

One path: account, runner, first benchmark.

Overview

Find the setup to run next, then inspect the evidence behind it.

Start with Choose for an answer-first setup choice. Compare public evidence or queue a benchmark when you already know what you want to inspect.

How evidence works
Evidence and setup status Checking evidence and runner readiness.
Active runs
—
syncing
Verified results
—
usable in decisions
Open blockers
—
checking
Sign in
Account attached
Pair a runner
Local execution ready
Choose evidence
Recommendation ready
Run or compare
Next action

Choose

Which setup should I run?

Recent runs

Tracked execution

More tools

Community

View contributor activity See who is expanding the public evidence corpus.

Top contributors

Community evidence stays cumulative and inspectable.

Overview

What should I run?

Choose your machine and task

InferGrade ranks model + quant setups that fit.

Already running a local model? Paste one localhost OpenAI-compatible URL. InferGrade runs a short check, then tells you what to verify next.
More evidence and candidates Candidate table, caveats, source notes, and the next benchmark. Evidence details ready Open for the complete candidate and proof trail.
Advanced filters Runtime, trust, family, quant, axes, and evidence scope.
Download data

Compare

Compare local models

Capability, speed, memory, and evidence—scoped to practical local setups.

Evidence browser

Recent benchmark evidence

Model Time / task Hardware Capability Trust

Compare

Choose between families, variants, and quants

Preset views

Start from a useful model-choice stance, then refine the exact variants or inspect individual runs.

Individual run comparison

Result

Result

Family comparison

Branches, quants, and nearby matches

Download data

Run

Plan a benchmark run

Why run this benchmark

Run the benchmark that would change the answer.

Start from Choose when possible; otherwise select a model and benchmark goal below.

1 Model 2 Benchmarks 3 Queue
Model

Choose the benchmark goal, then use a suggested model or paste your own.

Use public artifacts without connecting Hugging Face.

Benchmark scope

Choose the evidence this run should produce.

Customize benchmark checks
Benchmark groups

Adjust related checks together.

Individual checks

Exact checks for this run.

Run details

Optional context for history.

Advanced overrides
Artifact and runtime

Only adjust these if you need an override.

Ontology hints

Most users should keep the inferred values.

Run plan JSON Inspect or export the prepared plan.

Run plan

Ready to queue after preparation.

No run plan prepared yet.

Run Status

Active and recent runs

Recent runs

Live timeline

Saved plans

Reusable runs

My results

Contributor activity