Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Open models · sources checked 22 Sept 2026

AI you can download, run and keep.

Open models are published as files anyone can download: you run them on your own hardware instead of sending your work to someone else’s. Mutinai tracks the models, the machines they fit and the tools that run them, and labels every figure with where it came from.

Open-weight models tracked
25
Releases & launches in the last 30 days
8
Best open GPQA Diamond score · DeepSeek-R1
71.5%
Tools & runtimes tracked
13
  • Run models yourself

    Download the weights and run them on your own machine — no account, no per-token bill.

    What can I run?
  • Choose your hardware

    A laptop, one graphics card, a Mac or a server. You decide what the model runs on.

    Compare hardware
  • Control your data

    Prompts, code and documents stay on the machine you ran them on.

    How local models work
  • Customize the stack

    Swap runtimes, shrink a model to fit your memory, or fine-tune one on your own data.

    Tools & runtimes
  • Avoid lock-in

    The weights are files you keep. Nobody can withdraw, reprice or quietly change them.

    What “open” means

What’s happening right now

Top story
runtime release ·

vLLM v0.30.0

v0.30.0 Highlights This release features 762 commits from 315 contributors (104 new)! New models : DeepSeek-V4.1-Flash ( #56214 , #56228 , #56208 ) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 ( #56893 ), DeepGEMM Mega-mHC ( #56962 ), and async…

Latest

  1. Toolruntime release
  2. Event
    Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
    announcement
  3. Eventannouncement
  4. Toolruntime release
  5. Event
    Your Agent Aced the Task. Will It Do It Again?
    announcement
  6. Toolruntime release
  7. Toolruntime release
  8. Eventannouncement

What the community is talking about

Community Runs and reviews written by members will appear here, kept separate from the sourced facts everywhere else on this page and never merged into them.

Not collected yet

Contributions are not open yet

This preview has no sign-in, so nothing here could have been written by a member: the accounts, runs and reviews in its database are fictional sample content, and are left out rather than dressed up as activity. Real runs, reviews and discussion appear here once contributions open.

Explore the ecosystem

Start with a question

Models worth knowing

Open models arrive in families. Start with the family and who builds it, then open a release to see its sizes and variants.

  • Qwen

    Qwen Team (Alibaba Cloud)
    Known forCoding, reasoning and tool use
    596M – 235B parametersHow many numbers the model learned during training — the usual rough measure of its size.Dense and mixture of expertsA model split into parts where only a few run for each word, so it answers faster than its size suggests.Newest release 5 Aug 2026
  • Gemma

    Google DeepMind
    Known forImages, long documents and many languages
    9.24B – 27.4B parametersHow many numbers the model learned during training — the usual rough measure of its size.DenseNewest release 12 Mar 2025
  • Mistral

    Mistral AI
    Known forTool use and many languages
    7.25B – 23.6B parametersHow many numbers the model learned during training — the usual rough measure of its size.DenseNewest release 30 Jan 2025
  • DeepSeek-R1

    DeepSeek
    Known forCoding, reasoning and long documents
    671B parametersHow many numbers the model learned during training — the usual rough measure of its size.Mixture of expertsA model split into parts where only a few run for each word, so it answers faster than its size suggests.Newest release 20 Jan 2025
  • Phi

    Microsoft
    Known forCoding and reasoning
    14.7B parametersHow many numbers the model learned during training — the usual rough measure of its size.DenseNewest release 12 Dec 2024
  • Llama

    Meta
    Known forCoding, tool use and long documents
    8.03B – 70.6B parametersHow many numbers the model learned during training — the usual rough measure of its size.DenseNewest release 6 Dec 2024

How they compare

A benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is a fixed set of questions every model answers, so scores can be compared. Positions below are relative to the best open result Mutinai holds for that test.

Coding best scores on HumanEval

  1. Qwen2.5-Coder 32B92.7
  2. Llama 3.3 70B88.4
  3. Qwen2.5 32B88.4

developer-reported what it measures →

Reasoning best scores on GPQA Diamond

  1. DeepSeek-R1 671B71.5
  2. Qwen3 30B-A3B65.8
  3. Qwen2.5 32B62.1

developer-reported what it measures →

Knowledge best scores on MMLU-Pro

  1. DeepSeek-R1 671B84.0
  2. Phi-4 14B70.4
  3. Qwen2.5 32B69.0

developer-reported what it measures →

Instruction following best scores on IFEval

  1. Llama 3.3 70B92.1
  2. Gemma 3 27B90.4
  3. Llama 3.1 8B80.4

developer-reported what it measures →

Runs on your own machine

  1. Qwen2.5-Coder 32B~21 GB to runruns well on 11 of 13 reference systems
  2. Qwen2.5-Math 7B~6 GB to runruns well on 12 of 13 reference systems
  3. Qwen3 30B-A3B~18 GB to runruns well on 11 of 13 reference systems

estimated smallest download at 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each.

Images and multimodal

1 tracked model accepts images, but Mutinai holds no multimodal benchmark results yet, so there is nothing to rank here. See the models →

Best open score over time · GPQA Diamond

Each step is a release that beat the previous best open result. Developer-reported scores.

3080Llama 3.1 8B Instruct: 30.4%Qwen2.5 32B Instruct: 49.5%Llama 3.3 70B Instruct: 50.5%Phi-4: 56.1%DeepSeek-R1: 71.5%DeepSeek-R1 · 71.5Jul 24Jan 25

Hardware watch

Memory decides what fits; memory speed decides how fast it answers. Prices are shown only with what they mean and when they were checked.

Graphics memoryThe memory on a graphics card. A model has to fit in it to run at full speed. tiers

Mutinai records launch prices (MSRPThe price the manufacturer set at launch. What a shop charges today can be very different.) and, where a retail price has been checked, the date it was checked. It keeps no price history yet, so it shows no price trends.

Unified memoryOne pool of memory shared by the processor and graphics, as on Apple silicon, so large models can fit.

These chips share one pool of memory with the processor, so what fits is chosen when the machine is bought, not by the chip.

What can I run?

Start from the machine you have.

Pick the closest setup — we recommend a download, a runtime, and show memory and speed.

RTX 4090 workstation (64 GB DDR5)

Largest models that fit entirely in its 24 GB GPU, at 8K context.

  1. Start hereWhat open models are and why running them yourself matters.
  2. Run locallyMemory, quantization and runtimes decide what you can run.
  3. Understand modelsFamilies, releases, variants and benchmarks, decoded.
  4. Build with themAPIs, coding assistants, agents and fine-tuning.