See Aiera in Action
Whether you’re looking for a quick walkthrough or enterprise pricing, our team is ready to show you how Aiera can fit your needs.
Whether you’re looking for a quick walkthrough or enterprise pricing, our team is ready to show you how Aiera can fit your needs.
We may request cookies to be set on your device. We use cookies to let us know when you visit our websites, how you interact with us, to enrich your user experience, and to customize your relationship with our website.
Click on the different category headings to find out more. You can also change some of your preferences. Note that blocking some types of cookies may impact your experience on our websites and the services we are able to offer.
These cookies are strictly necessary to provide you with services available through our website and to use some of its features.
Because these cookies are strictly necessary to deliver the website, refusing them will have impact how our site functions. You always can block or delete cookies by changing your browser settings and force blocking all cookies on this website. But this will always prompt you to accept/refuse cookies when revisiting our site.
We fully respect if you want to refuse cookies but to avoid asking you again and again kindly allow us to store a cookie for that. You are free to opt out any time or opt in for other cookies to get a better experience. If you refuse cookies we will remove all set cookies in our domain.
We provide you with a list of stored cookies on your computer in our domain so you can check what we stored. Due to security reasons we are not able to show or modify cookies from other domains. You can check these in your browser security settings.
We also use different external services like Google Webfonts, Google Maps, and external Video providers. Since these providers may collect personal data like your IP address we allow you to block them here. Please be aware that this might heavily reduce the functionality and appearance of our site. Changes will take effect once you reload the page.
Google Webfont Settings:
Google Map Settings:
Google reCaptcha Settings:
Vimeo and Youtube video embeds:
You can read about our cookies and privacy settings in detail on our Privacy Policy Page.
Privacy Policy
Nearly Everything is Converging
This article was originally posted on LinkedIn.
How the 2026 model landscape looks from a research-weighted perspective
If you plot every model on the Aiera leaderboard by its composite score, you don’t get a smooth slope. Instead, you get: two models sitting alone at the top, a long flat plateau, and then a cliff. Looking closely, those stages tell a very interesting story about 2026 so far.
The plateau
Let’s start with the top half of the leaderboard: starting with Claude Fable 5 at #3 (64.9) and moving down to Claude Sonnet 4.6 at #16 (58.1), fourteen models are separated by just 6.85 points. Most of these models were released in just the last three months; on the axis of “which new model is best overall,” they are, for practical purposes, a near-tie, with scores differences that could be argued are mostly noise.
Meanwhile, below Sonnet 4.6 the floor drops out: about 8 points are lost by #20, and a further 8 points by #24. The bottom of the board (single-digit research scores) isn’t even in the same conversation as the plateau. The field isn’t gradually spreading out; it’s being compressed into a dense band at the top, with a sharp descent.
The slate of models released in 2026 have mostly started to converge, and selecting one over the other has become a pricing and throughput exercise (with some specific task performance to be measured and considered). This is at the heart of the Kimi K3 freak-out, as the frontier closed-source providers realize their moat might be shallow.
The breakaway
The exceptions to this phenomenon (for now) are gpt-5.6-sol (69.6) and gpt-5.5 (68.9). The single step from #3 up to #2 is 3.95 points, which is more than half the entire 14-model plateau internal spread. These two aren’t leading the pack; they’ve left it.
What separates them is almost entirely one thing: research. Our composite weights four tasks: research (60%), Q&A (24%), summarization (10%), sentiment (6%); and research is the only task measured by connecting the model to live, proprietary data and asking it to answer questions the open web can’t. The two leaders average 61.7 on research; the plateau averages 53.5. Run the decomposition and ~70% of the gap from #3 to #2 is in the research axis alone. For our clients, that is quite significant.
Why the middle is so flat
The plateau isn’t a coincidence; it’s what happens when scoring axes saturate.
Three of the four core tasks we measure are things that modern models are simply good at now. When a capability saturates, it stops discriminating: nearly everyone scores well, so that slice of the composite becomes roughly constant across the field. Forty percent of the Aiera Score is now close to flat from #3 to #16. That compression is the plateau.
The remaining 60% of the score (research) is the one axis that hasn’t yet saturated, because it isn’t a test of what the model knows but of whether it can drive tools, retrieve the right proprietary data, and reason against that data to create a correct, well-sourced answer. That’s still hard, and it’s where the real spread survives.
The cliff below #16 is also telling: models that can’t reliably drive the tools or effectively manage the data have research scores that collapse toward zero. Notably, which side of this cliff a model lands on isn’t predicted by size or lab prestige: open-weight models intermix with frontier labs on the plateau, and some very large models fall off it.
What it means if you’re choosing a model
For practitioners, the takeaway is that “best overall model” is becoming a less useful question than “best at the axis you care about.” If your workload is general finance NLP (extraction, summarization, etc.) the field has converged. A dozen models are effectively interchangeable; choose on cost, latency, context window, or open weight, not leaderboard rank. But if your workload is agentic, deep financial analysis (reasoning over live transcripts, filings, and institutional research) a genuine gap remains, it’s fairly wide, and it’s where you should spend your evaluation budget.
The signal
The tidy version of the 2026 story is “models are plateauing.” The board tells a sharper one: the common axes have plateaued, and model differentiation has migrated almost entirely onto the hard, data-connected one, where a small ceiling still stands and most of the field is bunched a few points below it. This is precisely what a research-weighted benchmark is built to surface, and why we pay such close attention to it.