IT Brief India - Technology news for CIOs & IT decision-makers
India
Open-weight AI models take majority of Vercel traffic

Open-weight AI models take majority of Vercel traffic

Fri, 2nd Oct 2026 (Today)
Joseph Gabriel Lagonsin
JOSEPH GABRIEL LAGONSIN News Editor

Vercel has published data showing that open-weight AI models handled most tokens on its AI Gateway in August. The figures also show that average token prices fell sharply during the month.

According to Vercel's September AI Gateway Production Index, open-weight models accounted for 56% of token volume in August, up from 13% in April and 7% in December 2025. August was the first month in which open-weight models made up the majority of token traffic on the gateway.

The data covers anonymised aggregate traffic routed through Vercel's AI Gateway, which moves tens of trillions of tokens each month between production applications and AI model providers. Token volume in the report includes input, output, reasoning, cached-input and cache-creation tokens, while spending is estimated from published list prices.

The shift in usage came with lower prices. The average price per token across the gateway fell 23.2% in August, the third straight monthly decline and the steepest drop since April 2026.

For teams running more than 10 million tokens in both July and August, the median team paid 7.6% less per token, compared with a 2.9% decline in July. The average token now costs less than half what it did five months earlier.

Spend split

While open-weight models accounted for most token volume, they represented a much smaller share of spending. In August, they made up 14% of AI Gateway spend, indicating that higher-priced closed models still captured most revenue.

Frontier model providers continued to dominate spending even as leadership between individual models shifted quickly. Half of all tokens in August went to models that had been available for less than three months, while almost nine in 10 tokens ran on models that had arrived on the gateway within the previous four months.

That pace of change points to rapid turnover in model adoption. Model leadership on the gateway has become unusually fragile, even as the largest providers retain most token spending.

Anthropic lead

Anthropic held 64% of gateway spend in August. It has taken at least 61 cents of every dollar spent through the gateway each month since December, and its models have held the top two positions by spend throughout that period.

Within Anthropic's model lineup, Fable 5 lost ground quickly. Its share of AI Gateway spend fell from 13.2% in July to 4.9% in August, while Opus 5 rose to 22.5%.

Opus 5 costs roughly half as much per token as Fable 5. Nine in 10 teams that had run Fable reduced their usage, with more of those workloads moving to Opus 5 than to any other model.

As a result, spend stayed with Anthropic even as customers moved away from one of its more expensive models. The figures suggest some production users are moving to lower-cost models within the same provider rather than switching labs altogether.

OpenAI launch

The report also highlighted strong early take-up for OpenAI's GPT-6 Astra. The model reached one in every three dollars spent on OpenAI models through the gateway within two days of launch, and its share of OpenAI spending then remained between 28% and 39%.

Within OpenAI's lineup, Astra and GPT-5.6 Sol together processed 27% of OpenAI's tokens but accounted for 71% of its spending over the period from 4 to 16 September. Astra launched at the same price as Fable 5.1 and at two and a half times the price of GPT-5.6 Sol.

Vercel compared Astra's early performance with Anthropic's Fable 5.1, which arrived on the gateway two days earlier. In each model's first 12 days on the gateway, Astra took 7.7% of all gateway spend, more than double Fable 5.1's 3.7%, and was used by twice as many teams.

Price pressure

The figures indicate that cheaper models are taking a larger share of production traffic, putting downward pressure on average token prices. At the same time, premium models continue to account for a disproportionate share of total spending.

Half of tokens in August went to models that had been available for less than three months, underlining how quickly usage patterns are changing across the market.