On 6 October 2026, Mistral released a public preview of Mistral Large 4, its biggest model so far. In Mistral’s own words: “Unofficially ML4, very officially: le Chonk.” The launch post comes with around twenty benchmark charts. We read them all and kept what matters if you run a business rather than an AI lab.
The short version:
- ML4 is the strongest open-weight model built outside China, not the strongest open model overall.
- It stands out on legal work, resistance to manipulation, and locating details in images.
- On business automation it is level with the best open models. In finance it trails the leaders.
- It is a preview. Mistral expects the scores to improve before the weights are published at the end of October.
What Mistral released
ML4 has 1 trillion parameters, of which 49 billion work on any given answer. It reads text and images, handles up to 1 million tokens of context (about 750,000 words), and was trained on more than 160 languages, including every official EU language. The weights will be open.
Two details matter more than the size. First, it was trained and is served in Europe, on Nvidia chips in Mistral’s own data centres, with a European deployment run “under European law”. That answers the jurisdiction question from our guide to sovereign AI. Second, it costs more than Large 3: $1.36 per million tokens read and $4.18 written at list price, currently discounted by half. For a typical chatbot answer, that is still a fraction of a cent.
The benchmarks
A benchmark is a fixed set of tasks, scored the same way for every model. Below are the seven that map to real business work, each against the best rival on Mistral’s own chart. Most were measured by outside evaluators (Artificial Analysis and Vals.ai).
ML4 wins one of these seven outright and ties another. GLM-5.3 and Kimi K3, both Chinese open-weight models, beat it on charts Mistral chose to publish. Credit to Mistral for showing them. But “best open model” needs the qualifier Mistral itself uses: “developed in the US or Europe”.
What each result means in practice:
- Business automation. AutomationBench runs 657 workflows across Gmail, Sheets, Slack and Salesforce, close to what we build in email automation and TicketFlow. ML4 scores 59.9%, in the same band as the best open models. Mistral’s previous all-rounder scored 6.3%. A European model is no longer the weak option here.
- Legal and finance. ML4 leads every open model on Harvey’s legal benchmark (15.8%, nearly three times GPT-6 Astra). The low score shows how hard the test is: no model does a lawyer’s job unsupervised. In finance it comes within about a point of the leader on one test, and trails Kimi K3 by ten points on spreadsheet work.
- Documents and images. It edges out GPT-6 Astra on locating objects in crowded images, and reads complex PDFs better than most rivals shown, though Kimi K3 does better. Useful for invoices, delivery notes and technical drawings.
- Resistance to manipulation. ML4 resists 93.3% of attacks designed to hijack an AI agent, such as instructions hidden in a web page or document, joint best. For any assistant that reads outside content and can act, this is the number to watch. We explain why in connecting AI to your company tools.
- Coding. Good, not first. Kimi K3 is ahead, and per VentureBeat the top closed models score around 74% on the public leaderboard, against ML4’s 62%.
How much to trust these numbers
- It is still training. Mistral expects “large and rapid improvements”. Scores will move.
- Mistral picked the charts. The closed leaders appear mostly where ML4 beats them or where they refused the task.
- Benchmarks are not your workload. A model that tops a legal test can still fail on your contracts, in your language. The test that counts is fifty of your real cases, which is what a proof of concept is for. We wrote more on measuring this in AI you can trust.
What it means for your business
Until now, a company wanting a top-tier open model it could run itself mostly ended up with a Chinese one, which is harder to explain to a compliance officer. ML4 gives Europe that option. Once the weights are out, and subject to a licence Mistral has not detailed yet, it can run on Mistral’s platform, on a European cloud, or in a data centre you choose. Not in your office, though: a trillion parameters needs a server with several high-end GPUs.
Every Flowful assistant runs on Vectoria, which is not tied to a single model, and most of our packages already work with Mistral models. Trying ML4 is a configuration change, not a project. The obvious first candidates are the Internal Chatbot, which reads contracts and procedures, and the Web Chatbot, which reads what your visitors send it. How we handle your data, whichever model you pick, is on our security page.
Do not switch models because of a launch post. Tell us which tasks you would hand to an assistant, and we will tell you whether ML4 is the right model for them, or whether a smaller one does the job for less.
Checked on 6 October 2026, the day of the preview. Figures come from Mistral’s launch materials unless another source is linked.
