Data-Driven Analysis

AI Insights

Trends, patterns, and analysis derived from 236 models, 53 papers, and 29 hardware milestones.

#1
Insight #1

The Cambrian Explosion

From 2 models in 2018 to 80 in 2024 — a 40× increase

With 236 models now tracked, the AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. landscape went from a handful of research projects to an industry producing 80 new models per year. 2023 was the inflection point — the year ChatGPT's success triggered an industry-wide arms race. Every major tech company, from Apple to Amazon, scrambled to release their own models. 67 organizations across 10+ countries are now competing.

#2
Insight #2

The Open Source Revolution

Open models went from 0% in 2021 to 59% in 2024
Closed Models
Open Models

In 2021, every major model was closed and proprietary. By 2024, open-weightopen-weightModel weights are publicly released but training data/code may not be. Enables fine-tuning but not full reproduction. and open-source models made up the majority of releases. Meta's LLaMA leak in 2023 was the spark — once researchers could study and fine-tune frontierfrontierThe absolute leading edge of AI model capability, representing state-of-the-art parameters, compute, and benchmark performance.-class models, the community produced an explosion of derivatives. This democratization may be the most consequential trend in AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. history.

#3
Insight #3

The Deepest Family Trees

GPT lineage reaches 18 generations deep

OpenAI's GPT family is the deepest evolutionary tree in AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence., with 14 generations from GPT-1 to GPT-5.5. This isn'tTTemperature — a hyperparameter dividing the logits to scale next-token selection randomness. just version numbering — each generation represents genuine architectural or training breakthroughs. Anthropic's Claude lineage, while shorter at 10 generations, shows the fastest iteration pace.

#4
Insight #4

The MoE Takeover

56 models now use Mixture-of-Experts architecture

Mixture-of-Expertsmixture-of-expertsArchitecture where only a fraction of the model's parameters are active for each input, allowing massive scale with lower compute. started as a 2017 paper and was mostly ignored. Mixtral 8x7B's viral success in late 2023 proved MoEMoEMixture of Experts — architecture where only a fraction of parameters activate per input, enabling massive scale at lower compute cost. could deliver GPT-4-class quality at a fraction of the inferenceInferenceUsing a trained model to generate predictions or outputs (as opposed to training it). cost. Within 12 months, MoEMoEMixture of Experts — architecture where only a fraction of parameters activate per input, enabling massive scale at lower compute cost. became the default architecture for any model over 100B parameters — adopted by DeepSeek V2/V3, Grok, DBRX, Arctic, and Qwen.

#5
Insight #5

The Context Window Explosion

Context windows grew from 512 to 10M+ tokens — a 19,000× increase

In 2018, GPT-1 had a 512-tokentokenA basic unit of text (such as a word or subword fragment) processed by a language model. context windowContext windowThe maximum number of tokens a model can process in a single input. Ranges from 2K to 10M+.. By 2025, Gemini offered 10 million tokens — enough to process entire codebases or dozens of novels at once. This wasn'tTTemperature — a hyperparameter dividing the logits to scale next-token selection randomness. just incremental improvement; it required fundamental innovations like Flash AttentionFlash AttentionAn IO-aware exact attention algorithm that's 2-4× faster by minimizing GPU memory reads/writes. and rotary position embeddings. The practical impact: AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. went from answering single questions to analyzing entire projects.

#6
Insight #6

67 Companies, One Race

Models from 67 organizations across 10+ countries
33
OpenAI
22
Google DeepMind
16
Anthropic
15
Meta
12
Google
9
DeepSeek
9
NVIDIA
8
Community
7
Mistral AI
7
xAI
7
Alibaba Cloud
5
Cohere

What started as a two-horse race between OpenAI and Google has become a global competition spanning 67 organizations. 21 models come from Chinese labs — DeepSeek alone has 7 entries, while Alibaba's Qwen and Zhipu AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence.'s GLM families are rapidly expanding. Community contributors and startups (Nous Research, Eric Hartford) punch far above their weight through fine-tuningFine-tuningAdapting a pre-trained model to a specific task or domain by training on additional data. and abliterationabliterationRemoving safety guardrails from a model through targeted fine-tuning or weight manipulation. Controversial but popular in open-source community. techniques.

#7
Insight #7

Innovation Pipeline: Paper to Product

Average ~2 years from paper to widespread model adoption
Attention Is All You Need (2017)
GPT-1 (2018)
1 year
RLHF (2017)
InstructGPT (2022)
5 years
Chain-of-Thought (2022)
o1 (2024)
2 years
MoE (Shazeer et al.) (2017)
Mixtral 8x7B (2023)
6 years
LoRA (2021)
Widespread use (2023)
2 years
Flash Attention (2022)
Default everywhere (2023)
1 year
RoPE (2021)
Default everywhere (2023)
~2 years
GQA (2023)
LLaMA 2 (2023)
<1 year
Mamba (SSM) (2023)
Jamba (2024)
~4 months

The pipeline from research paper to production model has dramatically accelerated. Early innovations like RLHFrlhfReinforcement Learning from Human Feedback — training models to align with human preferences by having humans rank outputs. took 5+ years to go mainstream. Now, architectures like Mamba go from paper to production in months. This compression is both exciting (faster progress) and concerning (less time for safety evaluation).

#8
Insight #8

The Modality Matrix

Text dropped from 100% to just 20% of new models
text
multimodal
code
image
audio
video

Early AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. was text-only. Now, over half of new models handle multiple modalities — images, audio, video, and code. The trend is unmistakable: the future of AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. is models that can see, hear, speak, code, and reason simultaneously. The arrival of models like GPT-4o (text+image+audio) and Gemini 2.0 (text+image+video) marks the beginning of truly general-purpose AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence..

#9
Insight #9

The Reasoning Revolution

53 models now claim reasoning capabilities
Our Take (Early 2026)
36 models

ReasoningreasoningStructured step-by-step problem solving, often using chain-of-thought or tree-of-thought approaches. has exploded from a niche capability (Chain-of-Thoughtchain-of-thoughtPrompting technique where the model 'thinks out loud' step by step before giving a final answer. prompting in 2022) to the most sought-after feature in AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence.. OpenAI's o1 proved that 'thinking longer' (test-time compute) could dramatically improve performance on hard problems. Now every major lab — Anthropic, Google, DeepSeek — is racing to build models that don'tTTemperature — a hyperparameter dividing the logits to scale next-token selection randomness. just pattern-match but actually reason step-by-step.

Our Take (Mid 2026)
48 models

In mid 2026, reasoningreasoningStructured step-by-step problem solving, often using chain-of-thought or tree-of-thought approaches. evolved into continuous self-correcting loopsloopsIterative processes where a model performs actions, observes the results, and reflects on them repeatedly to solve a task. where models automatically backtrack and fix their own logic during inferenceInferenceUsing a trained model to generate predictions or outputs (as opposed to training it)., seen in GPT-5.6 and Claude 5 Opus. The paradigm has shifted from discrete 'thinking steps' to fluid, adaptive problem-solving that mimics human reflection.

Our Take (Fall 2026)
53 models

By late September 2026, reasoningreasoningStructured step-by-step problem solving, often using chain-of-thought or tree-of-thought approaches. crossed an inflection point with GPT-6 Astra, Claude Mythos 5.1, and Claude Opus 5.5. The frontierfrontierThe absolute leading edge of AI model capability, representing state-of-the-art parameters, compute, and benchmark performance. is now proactive credit assignmentProactive Credit AssignmentAn advanced reinforcement learning technique where models anticipate future reward bottlenecks during long-horizon planning and assign credit across branching actions., automated formal verification, and dealing with 'behavioral shadowsBehavioral ShadowsPersistent, unintended behavioral patterns and reasoning tendencies imprinted on a model during intense reinforcement learning post-training.' (arXiv:2609.29233)—where intense post-training RLVR transfers reasoningreasoningStructured step-by-step problem solving, often using chain-of-thought or tree-of-thought approaches. reflexes and policy biases into completely unrelated decision domains.

#10
Insight #10

The China Factor

21 Chinese models tracked, up from 0 in 2022
United States
China
Our Take (Early 2026)
28 models

China has emerged as the world's second AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. superpower. DeepSeek, Alibaba, and Baidu are rapidly closing the gap with Western frontierfrontierThe absolute leading edge of AI model capability, representing state-of-the-art parameters, compute, and benchmark performance. models. The US-China AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. race is now the defining dynamic of the industry, with implications for regulation, export controls, and the future of open research.

Our Take (Mid 2026)
38 models

Companies like DeepSeek proved that innovative architecture (MLAMLAMulti-head Latent Attention — DeepSeek's innovation that compresses key-value caches into a low-rank latent space., multi-head latent attentionattentionA mechanism in Transformers that determines how much focus to place on other words in a sequence when processing the current word.) can compete with brute-force scaling. Moonshot AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence.'s Kimi K3 (open-weightopen-weightModel weights are publicly released but training data/code may not be. Enables fine-tuning but not full reproduction.) and MiniMax show that Chinese labs are no longer just following — they're leading on specific frontiers.

Our Take (Fall 2026)
21 models

In September 2026, Chinese AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. labs demonstrated unprecedented architectural efficiency: DeepSeek's V4.1-Flash achieved a 4x reduction in KV-cache memory overhead, ByteDance scaled Seed 2.1 Turbo for ultra-low-latency enterprise inferenceInferenceUsing a trained model to generate predictions or outputs (as opposed to training it)., Zhipu expanded the GLM-5.3 series, and Xiaomi's MiMo-V2.6 pioneered edge-to-cloud multimodalmultimodalProcessing multiple types of input (text, images, audio, video) in a single model. IoT orchestration.

#11
Insight #11

The Efficiency Revolution

77 models use efficiency innovations (distillation, MoE, parameter sharing)

The AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. industry hit a wall: training ever-larger models became prohibitively expensive. The responseresponseThe output text generated by a language model in reply to a prompt. was an efficiency revolution. DistilBERT showed you could compress BERT to 60% of its size while keeping 97% of its capability. ALBERT proved parameterparameterA variable or weight inside a neural network that is adjusted during training to store knowledge. sharing could slash model size 18×. Mixture-of-Expertsmixture-of-expertsArchitecture where only a fraction of the model's parameters are active for each input, allowing massive scale with lower compute. architectures (Mixtral, DeepSeek-V2) activate only a fraction of parameters per queryqueryIn self-attention, the vector representing the current token seeking context from other parts of the sequence.. LoRALoRALow-Rank Adaptation — an efficient fine-tuning technique that adds small trainable matrices to frozen model weights. made fine-tuningFine-tuningAdapting a pre-trained model to a specific task or domain by training on additional data. accessible on consumer GPUs. The Chinchilla paper proved most models were undertrained relative to their size. The new mantra: smaller, smarter, cheaper.

#12
Insight #12

The Safety Imperative

3 dedicated safety models + 7 RLHF-aligned models
Our Take (Early 2026)
Alignment & Guardrails

Safety went from an academic afterthought to an industry imperative. The RLHFrlhfReinforcement Learning from Human Feedback — training models to align with human preferences by having humans rank outputs. paper (2017) took 5 years to become standard practice. Constitutional AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. gave Anthropic a principled framework for self-improvement. Meta's Llama Guard and Google's ShieldGemma created accessible safety classifiers, while open-source abliterationabliterationRemoving safety guardrails from a model through targeted fine-tuning or weight manipulation. Controversial but popular in open-source community. challenged whether guardrails should be removable.

Our Take (Fall 2026)
Critical-Cyber Safeguards

With the September 2026 releases of GPT-6 Astra and Gemini 3.8 Flash Cyber, models for the first time crossed official 'critical-cyberCritical-CyberA frontier risk classification tier triggered by models demonstrating automated vulnerability discovery, binary analysis, and exploit synthesis.' capability thresholds. FrontierfrontierThe absolute leading edge of AI model capability, representing state-of-the-art parameters, compute, and benchmark performance. labs now mandate gated defender programs and runtime policy enforcement frameworks (like the NVIDIA Open Agent Safety Platform), treating autonomous multi-agent systems as critical infrastructure with tangible systemic risk.

#13
Insight #13

From Text to Everything

6 modalities × 8 specialized verticals = 30 specialist models
9
Image Gen
7
Coding
3
Search
3
Safety
2
Music
2
Speech
2
Embedding
2
Robotics

AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. has fragmented from one thing (text prediction) into a dozen specialized disciplines. EmbeddingembeddingA high-dimensional vector representing the semantic meaning of a token, grouping similar words close together in vector space. models (text-embeddingembeddingA high-dimensional vector representing the semantic meaning of a token, grouping similar words close together in vector space.-3, BGE) power every search engine and RAGRAGRetrieval-Augmented Generation — combining a language model with a search/retrieval system to ground responses in external knowledge. pipeline. Safety models (Llama Guard, ShieldGemma) act as AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. immune systems. Robotics models (RT-2, PaLM-E) bridge language and physical action. Music generators (Suno, Udio), speech synthesizers (VALL-E, ElevenLabs), and coding agents (Devin, Cursor, SWE-Agent) each represent billion-dollar verticals. The 'foundation model' era is giving way to an era of specialized, deeply integrated AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. products.

#14
Insight #14

The Era of Autonomous Agents

38 models specifically designed for autonomous, long-horizon workflows
Our Take (Mid 2026)
14 models

The paradigm has shifted from conversational chatbots to autonomous agents capable of extended task planning. Triggered by innovations like Google DeepMind's 'Prospective Credit Assignment' and open alternatives like Nemotron Lightning, models are now increasingly designed to run continuously, correcting their own mistakes and managing multi-step workflows over hours or days.

Our Take (Fall 2026)
38 models

By October 2026, autonomous systems operate as multi-agent swarms. Models like Grok 4.6 run parallel agent sub-routines, while ContextcontextThe text or prompt preceding a generation that the model uses to understand what it is currently processing. Language Models (CLMs, arXiv:2609.37725) treat live contextcontextThe text or prompt preceding a generation that the model uses to understand what it is currently processing. as an editable persistent workspace. Agent development has shifted from simple prompting loopsloopsIterative processes where a model performs actions, observes the results, and reflects on them repeatedly to solve a task. to robust long-horizon execution platforms with formal tool verification and critical-cyberCritical-CyberA frontier risk classification tier triggered by models demonstrating automated vulnerability discovery, binary analysis, and exploit synthesis. guardrails.

#15
Insight #15

Hardware Specialization

Scaling from chips to gigawatt racks and model-specific ASICs
🖥️
General GPU
Flexible, High-Cost
🏢
Rack-Scale
Dense, Agentic Workloads
🧠
Silicon ASICs
Embedded Weights
Our Take (Mid 2026)
Gigawatt Racks & ASICs

We are witnessing the rapid diversification of AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. hardware. It is no longer just about buying faster general-purpose GPUs. In mid 2026, companies introduced massive rack-scale deployment solutions (AMD Helios) and hyper-specialized ASICs (Taalas) designed to embed specific model weights directly into silicon. This solves the memory-bandwidth bottleneck, enabling dense agenticagenticModels that can autonomously plan, execute multi-step tasks, use tools, and self-correct without human intervention. workflows at scale.

Our Take (Fall 2026)
Vera Rubin & Wafer-Scale

In September 2026, hardware scaling branched into two dominant trajectories: NVIDIA began volume shipping of the Vera RubinVera RubinNVIDIA's next-generation GPU architecture featuring HBM4 memory and extreme compute density per gigawatt of datacenter infrastructure. (NV-VR200) platform with 288GB HBM4 to power massive multi-trillion MoEMoEMixture of Experts — architecture where only a fraction of parameters activate per input, enabling massive scale at lower compute cost. clusters like GPT-6, while Cerebras launched CS-4 Nexus, delivering up to 30x faster inferenceInferenceUsing a trained model to generate predictions or outputs (as opposed to training it). tokentokenA basic unit of text (such as a word or subword fragment) processed by a language model. generation to eliminate latency bottlenecks in agenticagenticModels that can autonomously plan, execute multi-step tasks, use tools, and self-correct without human intervention. thinking loopsloopsIterative processes where a model performs actions, observes the results, and reflects on them repeatedly to solve a task..

Insight #16

What's Next: Emerging Patterns

Based on the trajectories we've tracked, here are the patterns most likely to define AI's next chapter.

🤖
Agent-first

38 models already have agenticagenticModels that can autonomously plan, execute multi-step tasks, use tools, and self-correct without human intervention. capabilities — expect this to become the default interaction model.

⚡
Efficiency over scale

ParameterparameterA variable or weight inside a neural network that is adjusted during training to store knowledge. growth is plateauing; efficiency (MoEMoEMixture of Experts — architecture where only a fraction of parameters activate per input, enabling massive scale at lower compute cost., Mamba, MLAMLAMulti-head Latent Attention — DeepSeek's innovation that compresses key-value caches into a low-rank latent space.) is the new frontierfrontierThe absolute leading edge of AI model capability, representing state-of-the-art parameters, compute, and benchmark performance.. Smaller, smarter models are winning.

🔓
Open is winning… for now

Open-source share peaked at 62% but dropped to 31% in 2026 as frontierfrontierThe absolute leading edge of AI model capability, representing state-of-the-art parameters, compute, and benchmark performance. labs restrict access to their most powerful models.

🖥️
Hardware is the bottleneck

The 29 hardware milestones show compute doubling every ~18 months, but model demands are growing faster.

🎯
Specialization

Coding tools, search engines, music generators — AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. is fragmenting into specialized, deeply integrated products that do one thing exceptionally well.

🌏
The US-China race

The US-China AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. race will intensify — Chinese open-weightopen-weightModel weights are publicly released but training data/code may not be. Enables fine-tuning but not full reproduction. models (DeepSeek, Kimi K2) are already matching Western closed models on keykeyIn self-attention, the vector representing what information a token contains, matched against queries to compute attention weights. benchmarks.

🏥
Vertical foundation models

Every industry vertical will have its own foundation model — legal AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence., medical AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence., financial AIAIArtificial Intelligence — the broad field of computer science focused on building systems capable of performing tasks that typically require human intelligence. — each trained on domain-specific data at scale.

Analysis based on 236 models, 53 papers, and 29 hardware milestones tracked in the LLM Tree of Life.