Open Doesn't Mean Free

Open Doesn't Mean Free

Auf Deutsch lesen →

Listen to this article
Speed

AI narration · automatically generated

Subscribe as a podcast
Personal note

Stored only locally in your browser — nothing is sent.

Two reading levels: the running text is the thread — for everyone, no prior knowledge required. The specially marked deep dive boxes are elaborations for specialists and can be skipped without losing the thread.


Price and performance it was not — those were the first two parts. That leaves the third reason, and it is the most emotional: freedom. An open model, so the promise goes, belongs to you. No corporation talking into your decisions, no subscription that can be cancelled, no terms of use that change overnight. You download the file, and it is yours.

The file, yes. The right to work with it, not necessarily. Between “I have the weights” and “I may build a product with them” stands a document that few ever open — the licence. And in 2026 that document holds more than the two words on the download button let on.

Open source, open weights — and why this is no hair-splitting

Let us start with the label. “Open source” has a fixed meaning in the software world: you get not only the finished product, but everything you need to rebuild and understand it yourself — the source code. Carried over to AI models, the organisation that has guarded the term for decades (the Open Source Initiative) set a benchmark in 2024: a model is only open source when, alongside the weights, sufficient information about the training data and the complete training code are disclosed — detailed enough that a knowledgeable person could build an equivalent model.

Measured against that benchmark, practically no commercially relevant model is open source. Almost all ship only the weights — the finished block of numbers — but neither the data nor the blueprint. That is open weights: downloadable, but not reproducible. You get the car, not the engineering drawings.

The finest proof of this comes not from a critic but from a provider himself. The research institute Ai2 splits the comparison table in the model card of its own model into two expressly named blocks: “Open-weight Models” and “Fully-open Models” — and the big names otherwise advertised as “open source” sit in the first block. When even the providers who mean it seriously draw that dividing line, the label has been hijacked on the market.

Deep dive · for specialists A note for the sake of honesty: the few models that plausibly reach the strict benchmark — the research models of the OLMo, Pythia and Apertus families — are the only group that also lay open their training data and their blueprint, but in performance they lag noticeably behind. In the comparison table of that same provider they land within reach of models that are themselves already two generations old. Full openness currently costs roughly one model generation — that is the real price one rarely mentions when demanding “real open source”. And one further subtlety: the Open Source Initiative certifies no individual models. Every statement “model X is open source” is therefore an outside assessment, not an official confirmation.

The licence is the fine print — and it has migrated into the mid-sized business

Because almost everything is open weights, the enclosed licence decides everything practical: whether you may use the model commercially, whether you may resell it, what you have to write on the wall. And here something shifted in 2026 that you need to know about.

The most famous threshold comes from Meta’s Llama: above 700 million monthly active users you need an extra licence. It sounds threatening and is not — that number is reached by perhaps five companies worldwide. For everyone else Llama was practically free. It is precisely that reassuring order of magnitude that fell in 2026. Two examples, both readable word for word in the licence texts:

  • Mistral, the French flagship lab, places its middle model line under a self-written licence that strips all rights above 20 million US dollars of consolidated monthly revenue — above that a commercial licence exists only at the provider’s discretion.
  • The Chinese MiniMax M3 grants its rights at all only “for non-commercial purposes”; commercial use above 20 million US dollars of annual revenue requires prior written permission.

With MiniMax, twenty million in annual revenue is already enough — that is a solid mid-sized German firm; with Mistral it is twenty million a month, that is to say a mid-size corporation. Both lie far below Meta’s “five biggest in the world”. The threshold has thus not been tightened in the detail, it has changed category: from “concerns the biggest corporations in the world” to “may well concern you”. And this does not stand in an analysis, it stands in the licence text you click away at download.

Where the decisive clause hides

It gets worse: some restrictions do not even stand where you would look for them. The most instructive example is again Llama. In the actual licence file there is not a single hit on “European Union”. So whoever reads the LICENSE and is reassured has missed half the game. For the licence pulls in a second document by reference, the Acceptable Use Policy — and there stands the sentence that has it in it:

For all multimodal models in Llama 4, the rights granted are not granted to individuals domiciled in the EU or companies headquartered in the EU.

Since all Llama 4 models are multimodal, the clause hits the entire family. A company based in Germany is therefore never even granted the rights to use these models — and that not through a European law, but through the terms of use of the American provider himself. (End users of a finished product that contains such a model are expressly exempt — which is why Llama-based products are available in Europe all the same.)

The lesson is uncomfortable and simple: with these models it is not enough to read the file with “LICENSE” in its name. You must also read every document it refers to — and with some licences the provider may even change those documents unilaterally, while the duty to notice lies with you.

Who is liable for what it was trained on?

Now the question that can turn expensive, and that none of these licences answers in your favour: what if copyright-protected material sits inside the model — and comes back out again?

That this is no theoretical problem was shown by a German court in November 2025: the Regional Court of Munich I held it proven that a well-known language model could reproduce whole protected song lyrics nearly verbatim, and rated this memorisation of training data as a copyright infringement where no licence exists. The judgment is not final, the appeal is running — but it marks the risk. In that case it struck the provider of the model; when self-hosting, though, you are the provider.

And that risk you bear, with open weights, entirely yourself. None of the market-standard open-weight licences contains an indemnity in your favour — that is, the provider’s promise to stand in for such claims. The weights come “as is”, expressly without any assurance that they infringe no third-party rights. Meta’s licence even turns the tables: whoever sues Meta over intellectual property automatically loses the licence. That is the economic core of this chapter: the legal uncertainty over the training data is shifted onto the user — unlike some commercial paid offering that supplies exactly this as a contractual shield. Downloaded “for free” does not mean deployed “risk-free”.

The AI Act rewards real openness — the market term is not enough for it

A ray of light, with a catch. The European AI Act exempts providers of “free and open-source” licensed models from certain obligations. It sounds as if “open source” were an advantage here — but the legislator means the strict term, not the marketing. Models with research-only or non-commercial clauses fall out of the exception, user-number thresholds likewise, and even those who enjoy the exception must still keep a copyright policy on hand and publish a summary of the training content. The exception is thus narrower than the term on the download button leads you to believe.

More important for most readers is the worry that thereby dissolves: do I become a regulated provider myself if I host an open model and post-train it on my documents? Almost certainly not.

Deep dive · for specialists The Commission names a tangible guiding measure for this: you become an independent provider of a general-purpose model only when the compute you deploy for your own adaptation exceeds one third of the model’s original training compute (an indicative criterion; where you do not know the starting figure, the fallback threshold is around 3.3 × 10²² compute operations for models without systemic risk). A typical fine-tune — a so-called LoRA over a few graphics-card days — lies many orders of magnitude below that. The mid-sized firm that runs an open-weight model itself and specialises it on its own files does not thereby become a model provider in the sense of the law. It may very well be a deployer of an AI system, with its own obligations — but that is a different role, and the two are constantly confused.

The German punchline

That leaves the most uncomfortable finding, and it is a home game with a bad ending. Among the most permissive licences on the market — the ones with which you can go least wrong — are, of all things, the Chinese: DeepSeek and the GLM provider place their weights under the plain MIT licence, Alibaba’s main Qwen line under Apache 2.0. (A US model such as the open gpt-oss is likewise under Apache — only it is weak; the strongest unrestrictedly usable open model currently comes from China.) These are the licences a lawyer likes to see: short, commercial, without a revenue threshold, without a naming requirement.

And the German flagship models? The publicly funded Teuken, in its current freely downloadable version, stands under a non-commercial licence (an earlier, expressly commercial variant existed but was not carried on); the openly published models of Aleph Alpha stand under a licence that, beside commercial use, even excludes administrative use. One should not inflate this into the headline “Germany locks up its own models” — both houses offer commercial paths, just not via the freely downloaded file. But the sober, provable version stays uncomfortable enough, and next to the Llama finding from earlier it grows sharper: a US corporation shuts EU firms out completely — and even our own flagship models still put a hurdle in front of it where China puts none. Whoever wants the strongest unrestrictedly usable open model currently reaches for one from Hangzhou.

What this means for us

The licence is thereby exposed as the silent passenger — “open” simply does not mean “free”. And unlike price and performance, the circle closes here even more clearly: the licence restrictions apply wherever the model runs — in the cloud as on your card. The licence is therefore never an argument for running things locally. Where it turns into a risk — with liability for the training data — self-hosting is even the worse deal, because you are then the only one left to answer for it. Three thoughts to take away:

  • Do not read the label, read the licence — and everything it refers to. “Open source” on the button is marketing. What counts stands in the licence text and in the documents it pulls in by reference — right up to clauses that shut EU firms out completely.
  • Check the threshold against your own revenue. The limits at which rights lapse have migrated into the mid-sized business in 2026. Twenty million in revenue is reached faster than 700 million users.
  • “Free” is not “risk-free”. For the training data you are liable with open weights, not the provider. Self-hosting almost never turns you into a regulated model provider — but it does make you the only one still standing in the room in a legal dispute.

So if the freest choice is, of all things, a Chinese model — what are you bringing into the house with it? It is not the licence that then gives cause for worry, but the provenance: censorship baked into the weights, the supply chain, back doors. What of that is proven, what of it is panic, and what you can concretely do about it, we take on in Part 4.


If this piece gave you something to think about, feel free to share it — and at the next “open source” model, before you build it in, actually open the licence file and every document it refers to. This series has two more parts — provenance, and what remains at the end. If you do not want to miss any of them, subscribe to the newsletter: a short note the moment a new piece appears, no promotional newsletter, unsubscribe at any time with one click.


Sources (selection):