IT Brief India - Technology news for CIOs & IT decision-makers
India
Leiolai launches AI model that runs on user devices

Leiolai launches AI model that runs on user devices

Thu, 27th Aug 2026 (Today)
Mark Tarre
MARK TARRE News Chief

Leiolai has launched its leiolai-1 artificial intelligence model, along with a consumer app and developer API.

The model runs inference across users' existing devices rather than through data centres. That approach is central to Leiolai's pitch to developers and consumers. The company says the system uses no water for inference and pays users when their devices contribute computation, with the app showing estimated water savings and user earnings after each answer.

Leiolai is entering a crowded large language model market with a claim that sets it apart on architecture as much as price. Rather than tying growth to new server build-outs, it argues that capacity can expand as more people add their phones, tablets and computers to the network.

The launch version includes what Leiolai describes as the world's largest publicly available context window at 11 million tokens. It also says the model can continue generating without fixed output limits because runtime scales linearly as context increases.

Both features are available in the consumer product and the API. Developers can access the model through an OpenAI-compatible Chat Completions API, lowering switching friction for software teams already using that format.

Pricing model

Price is a key part of the launch. The Fast tier starts at USD $0.01 per million input tokens and USD $0.02 per million output tokens.

At that rate, one billion output tokens cost USD $20, 10 billion cost USD $200 and 100 billion cost USD $2,000. Leiolai says Fast is aimed at tasks that generate very large volumes of text, including synthetic datasets, evaluation corpora, document transformation, catalogue generation and product features that rely heavily on output volume.

Developers can choose how much computation to use for a request by setting response depth. According to the company, that lets teams apply more compute only where needed while using the lower-cost Fast option for large-scale generation workloads.

Leiolai is also offering Research and Private modes. Research starts at the same base price as Fast, with input at USD $0.01 per million tokens and output at USD $0.02 per million tokens.

Private starts at USD $20 per million input tokens and USD $90 per million output tokens. Plus Ultra, a higher-depth option available in either Research or Private, costs USD $160 per million input tokens and USD $700 per million output tokens.

Different modes

Research mode is intended for public problems such as protein design, mathematics, algorithm discovery and work related to artificial intelligence itself. Private mode, by contrast, is designed to process requests confidentially and exclusively on trusted hardware and the user's own devices.

These distinctions suggest Leiolai is trying to appeal to several groups at once: developers seeking low-cost bulk generation, researchers needing distributed compute, and customers wanting tighter control over where sensitive requests are handled.

In the consumer market, the app will be free to use and will not require a payment card. Users can also earn money by contributing device computation, and access to higher levels of intelligence increases as users progress, Leiolai says.

The launch also replaces Leiolai's Early Access model with leiolai-1. By releasing the consumer app and developer API together, the company is trying to build demand on both sides of its network at the same time: people who use the service and people whose devices supply compute.

Infrastructure bet

Leiolai's broader argument is that AI growth does not need to depend on another wave of data-centre construction. It says every participating device adds computation and claims that, with 10 million devices, it could deliver more long-context throughput than all major frontier labs combined.

That claim is likely to attract scrutiny because distributed inference across consumer hardware presents technical and operational challenges, including reliability, coordination, latency and trust. Even so, the model reflects a wider push in the AI sector to find alternatives to the rising cost, power demand and resource use associated with centralised compute infrastructure.

The consumer app also includes design elements tailored for Apple devices, including dynamic backgrounds, selected colour palettes, motion tied to device tilt and animated particles within glass-style interface components. Those features sit alongside functional displays of earnings and estimated water savings.

"Growth does not have to begin with another data center," Leiolai said.