Open source in the AI age: open weights vs open source AI, training data, and the community

One of six themes that emerged from the EOLE 2026 online kick-off workshop (25 June). Background: see the kick-off topic.

Why it matters for sovereignty. AI was called the elephant in the room: most systems will soon embed AI, and if sovereignty is about autonomy, control and the capacity to decide, dependence on third-party AI cuts directly against it. AI also strains the classic open source promise (read, audit, rewrite) when models are opaque.

What came up at the workshop. A key distinction: open weights are not the same as open source AI. Open-weights models are free to use but give no visibility on training data, which is fine for a large share of applications; truly open source models remain few (Apertus, Allen Institute’s OLMo, Hugging Face initiatives) and are not yet seen as frontier. Open-weights models, including Chinese ones, were credited with breaking the proprietary stranglehold. Two “no-no’s” were named from the OSAID experience: use restrictions (Llama-style conditions, including bans on generating synthetic data to prevent forks), and AI as a service, meaning dependence on hyperscalers: you only use open stuff when you control it and it runs on your machines. Attention was also drawn to the heavy subsidies behind proprietary frontier models, and to a gap Europe could fill: an open, continuously extended corpus of training data, notably for post-training and reinforcement learning. A deeper question was left open: what is the role of the community when a large share of code is AI-generated (quality assurance? data curation? and what future for junior developers)?

What already exists to build on. OSAID (Open Source AI Definition); EuroLLM and EU open-model initiatives; open and open-weights models (Mistral, DeepSeek, Llama with caveats, Apertus, OLMo, Hugging Face); the open data agenda.

Open questions. How do sustainability, access and transparency need to be rethought for AI-based stacks? Can Europe build an open training-data commons, especially for post-training? What are the concrete consequences of Llama-style licence limits? What is the community for in an AI-assisted world?

How to contribute. Reply below with analyses, model reviews and initiatives, ideally by end of September 2026. Rapporteurs welcome for the Barcelona event (November 2026): volunteer by replying below or by direct message on this forum. Possible outputs: an “open weights vs open source AI” explainer and a mapping of EU open models and open training-data initiatives.