Cosine AI Review 2026: Features, Pricing & Verdict
Most AI coding tools are fighting over the same customer: a developer building a React app who wants faster autocomplete. Cosine AI walked away from that fight.
Instead, it trains its own coding models, supports languages like COBOL and Verilog, and will install itself inside an air-gapped server room. It is now building the UK’s first sovereign frontier AI model, with HSBC, Lloyds and BAE Systems signed on as partners.
That is a very different product from Cursor. This review covers what Cosine AI actually does, what the Lumen models are, what it really costs in 2026, and who should skip it.
Looking for cosine similarity? If you searched “cosine AI” expecting the vector-maths formula used in embeddings and semantic search, that is a different topic entirely. Start with cosine similarity instead. This article is about Cosine, the company.
What Is Cosine AI?
Cosine describes itself as “the AI engineer for production software teams.” It is not a chat wrapper sitting on top of someone else’s model. Cosine post-trains its own family of coding models on production codebases.
The company was founded in 2022 by Alistair Pullen (CEO), Sam Stenner and Yang Li. It launched as Buildt, an AI codebase search tool, went through Y Combinator, and rebranded to Cosine in 2024.
The pitch is narrow and deliberate. Cosine is not trying to be the fastest autocomplete. It is trying to be the AI engineer that a bank, a defence contractor or a telco can legally and safely install.
Cosine AI at a glance
| Founded | 2022 (as Buildt) |
| Founders | Alistair Pullen, Sam Stenner, Yang Li |
| Backing | Y Combinator, Lakestar |
| Products | Lumen model family, Cosine CLI, Cosine Cloud, Red Team |
| Starting price | $19/month |
| Best for | Regulated industries, legacy codebases, air-gapped deployments |
| Not for | Solo devs on a standard web stack |
From Buildt to Genie to Lumen
Cosine’s story explains its strategy better than any feature list.
In August 2024, Cosine launched Genie and briefly held the AI coding crown. Genie scored 30.1% on SWE-bench Full and 50.7% on SWE-bench Lite, beating Amazon Q’s 19.75% by a wide margin, according to DeepLearning.AI’s reporting. Genie was trained on a proprietary dataset covering six software engineering tasks across 15 programming languages.
In 2025, Genie 2 added agentic workflows: read a Jira ticket, write the code, run the tests, open the pull request.
Then in 2026, Cosine retired the Genie branding and shipped the Lumen model family instead.
Here is why that matters. By 2026, the frontier labs had won the benchmark race outright. Chasing them on SWE-bench was a losing game for an eight-million-dollar startup. So Cosine changed the question. Instead of “which model solves the most tasks,” it started asking “which model produces code a senior engineer will actually merge, in a place where you’re not allowed to send code to a US cloud.”
That pivot is the whole company now.
The Lumen Model Family Explained
Cosine runs three models, each aimed at a different job.
Lumen Scout is post-trained from Devstral 123B. It is the fast, cheap option, built for on-device work where latency matters more than depth.
Lumen Outpost is the workhorse, announced on 13 May 2026. It is post-trained from Kimi K2.6 using reinforcement learning — the same Moonshot AI model line that has been quietly beating far more expensive Western models. Cosine claims Outpost delivers three times the cost-efficiency of GPT-5.5 on everyday production tasks.
Lumen Sovereign is the ambitious one. Announced at London Tech Week on 8 June 2026, it is Britain’s first sovereign frontier AI model — built from scratch rather than fine-tuned, using proprietary datasets drawn from more than 30 regulated workflows. It trains on Isambard-AI, the Nvidia-powered supercomputer in Bristol, under the UK government’s £500 million Sovereign AI programme.
The partner list is the real signal: HSBC, Lloyds, NatWest, BT, Telefónica Tech, BAE Systems, Babcock, Thales UK, Leonardo UK, LSEG and PwC. Target deployment is the end of 2026.
Pullen’s argument is that vendor lock-in creates “security risk, dependency risk, and cost escalation risks.” For a UK bank, that is not a marketing line. It is a board-level concern.
The benchmark numbers
Cosine built three of its own benchmarks, because the public ones do not measure what it sells:
| Benchmark | What it measures | Lumen Outpost |
|---|---|---|
| Niche-Bench | Legacy and rare language accuracy | 53.9% |
| Slop-Bench | Code maintainability, duplicate logic | 25.4% |
| Vibe-Bench | Collaborative agent behaviour | 29.4% (up from 22.7%) |
Cosine says Outpost beats GPT-5.5, GPT-5.4 and Gemini 3.1 Pro on niche language tasks.
Read those numbers with care. These are self-reported benchmarks on tests Cosine designed itself. That does not make them false, but it does mean nobody independent has reproduced them yet. Treat them as a claim, not a fact.
What Cosine AI Actually Does
1. Code built for review, not volume
Cosine optimises for readable, maintainable diffs rather than raw output. Its stated target is “slop” — the sprawling patches and duplicated logic that make AI-generated code expensive to review. Anyone who has read a 900-line AI pull request knows the problem. If you’re new to how machine-written code differs from hand-written code, our guide to basic AI code covers the fundamentals.
2. Legacy and niche language support
This is the moat. Cosine handles COBOL, Fortran, Verilog, ABAP, Rust and complex SQL.
That sounds unglamorous until you remember who runs those languages. Banks run COBOL. Chip designers run Verilog. Enterprise resource systems run ABAP. Mainstream models saw very little of this code in training, so they guess. A model deliberately post-trained on it does not.
No other major coding assistant is competing seriously here.
3. Workflow and deployment
Cosine ships two surfaces — a CLI and a collaborative Cloud platform — plus multi-agent orchestration, parallel task execution, and external tool access via MCP. It integrates with Git, GitHub, GitLab, Slack, Jira and Linear. There is also a separate Red Team application for security work.
Deployment is where it separates from the pack: public cloud, managed single-tenant, or fully air-gapped on-premise. If your code cannot leave the building, Cosine is one of very few options. This is part of a wider shift we track in our AI dev tools trends coverage.
Cosine AI Pricing in 2026
Cosine moved to credit-based billing. Here are the current plans:
| Plan | Price/month | Credits included | Extra credits | Built for |
|---|---|---|---|---|
| Starter | $19 | 4M | $6.50 per 1M | Solo devs, side projects |
| Team | $199 | 47M | $5.00 per 1M | Growing teams |
| Enterprise | $999 | 240M | $4.50 per 1M | Regulated industries |
Credits cover agent work, model calls and cloud execution. Burn rate depends on task size, model choice and runtime.
Three things worth knowing before you buy.
First, the Team plan is not per-seat. Spread across a six-person team, $199 works out at roughly $33 per developer per month. Compared to per-seat rivals, that is the single strongest financial argument for Cosine.
Second, credit billing is harder to forecast than seats. You cannot know your monthly cost until you have run a real sprint through it. Start on Starter, measure your actual burn, then upgrade. Credit models reward heavy users and punish casual ones — the same pattern we found reviewing 1min AI.
Third, there is no free tier in 2026. Several tool directories still list an old free plan with “80 tasks” and a per-seat Professional tier. That pricing is gone, along with the Genie branding it belonged to. If a review quotes it, that review has not been updated in a year.
Cosine AI vs Cursor, Devin and Claude Code
| Cosine | Cursor | Devin | Claude Code | |
|---|---|---|---|---|
| Billing | Credits, flat team price | Per seat | Per seat / ACU | Usage or subscription |
| Best codebase | Legacy, regulated | Modern web | Modern web | Broad |
| Niche languages | Strong | Weak | Weak | Moderate |
| Air-gapped install | Yes | No | No | No |
| Own models | Yes | No | Partly | Yes |
| Ecosystem maturity | Small | Large | Medium | Large |
- Choose Cursor if you want the fastest in-editor loop on a modern web stack.
- Choose Claude Code if you want a terminal-native agent with a deep plugin ecosystem.
- Choose Cosine if you have regulated data, legacy code, or a compliance team that will not approve a US public cloud.
These are not really the same product. They only look similar from the outside.
Who Cosine AI Is For
Good fit:
- Enterprises in finance, defence, healthcare or telecoms
- Teams maintaining COBOL, Fortran, ABAP or mainframe systems
- Organisations with data residency rules that block US cloud processing
- Teams tired of paying per seat for licences half the team barely uses
Poor fit:
- Solo developers on a standard React or Node stack — Cursor is cheaper and faster for you
- Anyone who needs a mature plugin and extension ecosystem today
- Anyone who needs a free tier to evaluate
- Anyone hoping AI will replace their engineers outright, which our testing on that question suggests is still some way off
Honest Limitations
No review on aiera.blog ends without the downside, and Cosine has real ones.
The team is small next to Cursor, Anthropic and Cognition. Independent validation is thin — there are no G2 reviews yet, and the benchmarks are Cosine’s own. Credit forecasting is genuinely difficult in month one. There is no free tier, so evaluation costs money.
Most importantly, Lumen Sovereign has not shipped. It is targeted for late 2026. It is the most exciting thing Cosine is doing and it is also a promise, not a product. Do not buy today for a capability that arrives next year.
Frequently Asked Questions
Is Cosine AI free? No. There is no free tier in 2026. Entry is the $19/month Starter plan with 4M credits.
What happened to Cosine Genie? Genie was Cosine’s flagship agent from 2024 to 2025. It has been replaced by the Lumen model family — Scout, Outpost and Sovereign.
Is Cosine AI better than Cursor? It depends on your stack. Cosine wins on legacy languages and air-gapped deployment. Cursor wins on speed, ecosystem and everyday web development.
Can Cosine AI run on-premise? Yes. Cosine offers managed single-tenant deployments and fully air-gapped on-premise installations.
Is Cosine AI a UK or US company? It is registered in San Francisco but runs its sovereign AI programme on UK infrastructure with UK institutional partners.
Verdict
Cosine AI is a specialist tool, and it is honest about that.
If you write TypeScript in a startup, buy Cursor and move on. Cosine costs more, offers less tooling, and solves problems you do not have.
If you maintain a COBOL banking system, work under data residency rules, or need an AI engineer that installs behind your own firewall, Cosine is one of very few credible options — and the flat team pricing makes it cheaper than it first looks.
The risk is straightforward: it is a small company selling to some of the most conservative buyers on earth. The reward is that almost nobody else is trying. Cosine stopped chasing the leaderboard and started selling deployability, and in 2026 that looks like the smarter bet.
Rating: 4/5 for regulated and legacy teams. 2/5 for everyone else.
For more tested reviews of AI tools, browse the AI tool reviews section on aiera.blog, or read our companion breakdown of Qwen AI for another look at open-weight models built outside the US.