top of page
davydov consulting logo

DAVYDOV CONSULTING BLOG

HOME  >  NEWS  >  POST

Open-Weight AI Models Now Do Most of the Work: Should Your Business Switch?

1 minute ago
5 min read

For the past three years, the AI conversation in most companies has started and ended with the same few names. You picked OpenAI, Anthropic or Google, connected the API, and paid whatever the invoice said at the end of the month. Open-source alternatives existed, of course, but they were mostly treated as research toys, something your team might look at later, when things calm down.

Well, later has arrived. According to Vercel's AI Gateway production index, open-weight AI models processed 56% of all tokens that went through the platform in August 2026. In December 2025 the same figure was 7%. In eight months, open models went from a rounding error to the majority of production AI traffic on one of the largest model gateways on the market. On the busiest day, 22 August, they peaked at 62%.

For business owners and managers this is not an academic milestone, it is a pricing signal. In this article we look at what actually changed, why companies are moving routine AI workloads to open models, and how to decide whether your product should follow. Let's sink in!

A business team discussing AI model costs and usage dashboards during a strategy meeting

The month open models took the lead

First, a quick definition. Open-weight models are AI models whose trained weights are published, so anyone can download them, run them on their own infrastructure or through an inference provider, and often fine-tune them for their own tasks. DeepSeek, Meta's Llama family, Alibaba's Qwen, Moonshot AI and Z.ai's GLM series are the names you will see most often.

The growth curve on Vercel's gateway is unusually steep even by AI standards. Open models handled 7% of tokens in December 2025, 13% in April, 36% in July and 56% in August. This is not one viral product skewing the numbers either; the volume is spread across many teams and many models. Z.ai's GLM-5.3-Flash, for example, overtook its own predecessor within a single day of launch.

Quality is the reason this became possible. A recent Mozilla report puts the capability gap between the leading open and proprietary models at around 3.3%. For summarising documents, drafting product descriptions, classifying support tickets or powering a chatbot, most users would honestly never notice the difference.


Cheap tokens, expensive answers

Here is the twist that makes the story interesting. While open models now carry 56% of the tokens, they only account for about 14% of the money spent on the gateway. Anthropic alone still takes roughly 64 cents of every dollar, a share that has not dropped below 61% since December 2025.

The explanation is simple: open-weight tokens are dramatically cheaper. Closed-model tokens cost about 7.8 times more on average, which means companies route the heavy, repetitive workloads to open models and reserve the expensive frontier models for the tasks that genuinely need them. High volume goes one way, high value goes the other.

Prices are falling across the board too. The average price per token on the gateway dropped 23.2% in August alone, the third monthly decline in a row. Among teams processing more than ten million tokens, the median price per token fell 7.6% in a single month. Whatever you agreed to pay for AI six months ago, the fair price today is probably lower.


Why companies are switching to open-weight AI models

The first reason is obvious: budget. If open models give you comparable output for routine work at a fraction of the price, the same monthly spend suddenly buys several times more inference. For products where AI features run on every user action, such as search, recommendations or content generation, that difference lands directly on your margins.

The second reason is control. With an open model you can choose where it runs, keep sensitive data inside your own infrastructure, fine-tune the model on your domain and switch providers without rewriting your product. For companies in regulated industries, that flexibility is sometimes worth more than the savings themselves.

And the third reason is that buyers have stopped being loyal. Vercel's data shows that when Anthropic released Opus 5 at roughly half the price of its Fable 5 model, nine out of ten Fable customers cut their spending on it, and most of them simply moved down the price ladder within the same lab. Users follow value, not brand. Open models are the logical end point of that behaviour.

Server racks in a data center of the kind that hosts open-weight AI models

What open models still will not give you

Before you cancel every proprietary subscription, some honesty about the limits. The hardest tasks, complex multi-step reasoning, agents that plan and execute long workflows, high-stakes code generation, are still dominated by frontier models, and that is exactly why they keep 86% of the spending. The 3.3% capability gap is an average; on the toughest problems it is much wider.

Running open models is also your responsibility in a way an API subscription never was. Someone has to choose an inference provider or manage GPUs, monitor quality, apply updates and handle failures. Managed inference platforms remove most of this pain, but they also take back part of the cost advantage.

One caveat about the data itself is worth mentioning. Vercel's index covers traffic on its own gateway, so it does not see direct API contracts, private enterprise deals or fully self-hosted deployments, and its spend figures are based on list prices. The trend is real, but the exact numbers describe one platform, not the whole market.


What this means for businesses building digital products

The practical conclusion is not "switch everything to open models tomorrow". It is: stop building products that are welded to a single AI vendor. If your application talks to one hardcoded API today, you are one price change away from a margin problem and one outage away from downtime you cannot control.

When we design AI features for clients, model routing has become a standard part of the architecture. A thin routing layer, whether a commercial gateway or a simple internal service, lets you send routine requests to a cheap open model and difficult ones to a frontier model, and lets you change that mapping without touching product code.

It also changes how you should negotiate. With per-token prices falling every month, long commitments priced at today's rates are a bad deal. Shorter contracts, usage-based pricing and a tested fallback model give you room to benefit from the next price drop instead of watching it from inside a lock-in.

Developers comparing outputs of an open-weight AI model and a proprietary model at their workstations

How to run the numbers before you switch

Start by measuring what you actually spend and on what. Break your AI bill down by feature, not by invoice. In most products we audit, two or three routine features generate the bulk of the tokens, and those are the natural candidates for an open model.

Then test quality on your own data, not on benchmarks. Take a sample of real requests, run them through both your current model and one or two open alternatives, and have people who know the product review the outputs blind. If the open model wins or ties on your routine workload, the migration maths is usually easy.

Finally, count the full cost. Include the inference provider's fees, the engineering time to set up routing and evaluation, and the frontier model you will still keep for the hard 10% of requests. Even with all of that included, teams that make the move report cutting the cost of migrated workloads by half or more.


Final notes

August 2026 will likely be remembered as the month open-weight AI stopped being an experiment. The majority of production tokens on a major gateway now run on models anyone can download, while the money still flows to frontier models for the work that justifies them. That split, cheap open models for volume, premium models for difficulty, is quickly becoming the default architecture for AI products.

For your business the takeaway is simple. Know your workloads, keep your product model-agnostic and re-check your AI costs every quarter, because the market is moving faster than annual budgets do. The companies that treat model choice as an ongoing decision rather than a one-time pick will quietly end up with the same features at a fraction of the cost.

 
 
 

Comments


​Thanks for reaching out. Some one will reach out to you shortly.

CONTACT US

bottom of page