---
title: "Open Models Tack Toward the Frontier"
description: "Open-weight models keep catching closed frontier systems on Chatbot Arena Elo, 2023-2026, but have not materially led."
categories: ["AI","open source","startups"]
keywords: ["open-source AI","closed AI models","DeepSeek R1","GLM-5.2","Kimi K3","Chatbot Arena Elo"]
ai_summary: "Open-weight models have repeatedly reached equivalency with closed frontier models, first with DeepSeek R1, then GLM-4.6, GLM-5.2, \u0026 Kimi K3. But open source has not materially led; closed models such as GPT-5.2-class systems \u0026 Opus continue to create step-change moments."
date: 2026-07-20
lastmod: 2026-07-20
canonical_url: https://www.tomtunguz.com/open-models-tack-toward-the-frontier/
author: "Tomasz Tunguz"
---


{{< email_image src="hctnco7jfmuvm7ercweu" alt="Line chart comparing best open-weight & closed-weight LLM Chatbot Arena Elo from 2023 to 2026" width="540" height="357" >}}

Like two sailboats in a marathon race, open & closed AI labs are tacking & jibing in San Francisco Bay.

In 2023, closed source models led by an enormous margin on Chatbot Arena Elo[^1]. Two years later, the DeepSeek R1 moment arrived, the open-source answer to the ChatGPT moment. The two boats raced side-by-side for nearly a year.

Architectural improvements & the first Blackwell-trained models brought a step change with GPT-5.2 & Fable 5 starting in 2026.

The next few weeks will see another flurry of open-source releases. Moonshot [shipped Kimi K3](https://simonwillison.net/2026/Jul/16/kimi-k3/), a 2.8T parameter open-weight model, on July 16. Alibaba [previewed Qwen 3.8](https://www.marktechpost.com/2026/07/19/alibaba-previews-qwen3-8-max-a-2-4-trillion-parameter-multimodal-model-days-after-moonshots-kimi-k3-open-weight-launch/), a 2.4T model, on July 19. DeepSeek V4 [graduates from preview](https://technode.com/2026/06/30/deepseek-to-launch-v4-in-mid-july-with-new-peak-time-api-pricing/) in mid-July. These follow [Thinking Machines' Inkling](https://thinkingmachines.ai/news/introducing-inkling/), a 975B Apache-2.0 multimodal model released July 15, & Meta Superintelligence Labs' Muse Spark in April.

Open-source models have never taken an open-water lead, but that may not be necessary. Blend prices at a 90/10 input-to-output ratio & the median open-weight frontier model runs about 15% cheaper than GPT-5.2. The cheapest open model, DeepSeek V4 Flash, is roughly 90% cheaper.

{{< email_image src="ujreabprnfwmdlc3ozfi" alt="Bar chart: blended API pricing per million tokens, open-weight vs closed frontier models, July 2026" width="540" height="299" >}}

We may have a dynamic where the closed models drive the industry forward & open-source rapidly copies to commoditize. Will that slow down innovation?

Competition tends to do the opposite. OpenAI has [cut inference costs by 50%](https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half). Kimi shipped a [new attention architecture, KDA](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart). Fable's step function has an entire industry redoubling to catch up.

The major question put to the industry is what will happen to margins. Anthropic is [about to post its first profitable quarter](https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-propel-anthropic-into-its-first-profitable-quarter-7edbf2f4). Bezos said your margin is my opportunity. Open source's competitive dynamics keep margins & pricing competitive.

The AI wave will be among the largest infrastructure projects[^2] ever for the US & likely one of the greatest contributors to faster economic growth. Competition is essential to keeping the race fast.

The frontier is no longer a one-way race. It is a repeating cycle: closed models pull ahead, open models catch up, & the whole market moves faster.

[^2]: See [The GDP Impact of LLMs](https://tomtunguz.com/llm-impact-gdp/) for the scale estimate & the growth channel it flows through.

[^1]: Chatbot Arena Elo is a rating system borrowed from chess. Users see responses from two anonymous models side-by-side & vote for the better one. Each model starts at 1000. Winning against a stronger model earns more points than winning against a weaker one; the gap in ratings predicts the probability of winning a matchup. A 100-point Elo gap implies the higher-rated model wins about 64% of the time. The score reflects human preference on open-ended chat, not reasoning, coding, or agentic benchmarks.
