A sovereign AI lab: the Lumen model family and specialised coding agents, trained for organisations that need frontier capability inside air-gapped or regulated environments.
Last updated
Cosine has pivoted, and most of what you will read about it elsewhere describes a product that no longer exists. It was known for Genie, an autonomous coding agent in roughly the same category as Devin: describe a bug, the agent searches the repository, plans, edits the files the change touches, and hands back a diff. That is not how the company presents itself now. cosine.sh describes Cosine as the Sovereign AI Lab, and the one-sentence version of the business is that it trains AI models and specialised coding agents for organisations that need frontier capability inside secure environments.
The distinction that matters is who that sentence is addressed to. The reader it describes is not a developer choosing an assistant; it is an organisation with a constraint — a classification boundary, an air gap, a regulator, a codebase written in something the general-purpose models barely saw during training — that rules out the normal options before the evaluation starts. Worth noting that the pricing page has not entirely followed the front page there: it still sells a $19 Starter plan described as being for solo developers and side projects. So the self-serve door is open, but the product behind it is being built for the constrained buyer, and that is the gap to keep in mind when you read a feature list written for someone whose problems are not yours.
The product line is a set of models under the Lumen name — Lumen Scout, Lumen Outpost and Lumen Sovereign — rather than a single subscription tier list, and the naming tracks the deployment story rather than a good, better, best ladder. Around them the site organises its solutions along four axes: sovereign AI, air-gapped operation, cybersecurity, and niche languages. Applications are narrower and more specific than a general coding assistant: a red-team application is offered, and a training application is listed as coming soon. You reach the system through a CLI or through Cosine's cloud rather than as an editor plugin, which tells you something about the assumed user — this is tooling for an engineering or security team with an environment to integrate with, not a tab-completion experience.
Of those four axes, niche languages is the one most worth pausing on, because it is the only one that is about capability rather than deployment. A general-purpose assistant's competence in a language tracks how much of that language it saw in training, which is why they are fluent in TypeScript and vague in whatever your industrial control system is written in. A lab that trains its own models can attack that directly. Whether it succeeds is not something this page can tell you, and it is not something a vendor's own numbers can tell you either — Cosine publishes benchmark results of its own devising, and vendor-run benchmarks are evidence about the vendor's confidence rather than about your codebase.
If your code can go to a commercial cloud API, this is not your tool, and the pivot is the reason: you would be buying a sovereignty and security story you do not need, from a small vendor, in place of mature products. A developer who wants an assistant should be reading about Cursor or GitHub Copilot; a team that wants tickets worked autonomously should be reading about Devin and the broader shift from assistants to agents. Cosine only becomes the interesting answer at the point where those have already been eliminated by a rule somebody else wrote.
The evaluation advice for the old Genie still applies to the coding-agent half of this, and it is the part vendors are least keen to have you do: take ten tickets your team actually closed last quarter, replay them, and count how many produced a diff you would have merged with light edits. That number is the one that predicts anything. Everything else — the leaderboard, the demo, this page — is a proxy for it. And because deployment is the whole premise here, get the deployment questions answered in writing before a pilot: where the weights live, what leaves the boundary during a run, what is retained, and what happens to your workflows if the company is acquired.
The premise of the whole product. Where an environment has no route to a commercial model API, the choice is not between vendors but between having AI assistance at all and not having it, and that is the only situation in which a specialist lab beats an incumbent on anything other than price.
Defence, industrial, telecoms and financial systems are full of languages and dialects that general-purpose assistants saw little of in training, and their output degrades accordingly. A vendor that trains its own models can target that gap directly, which is a different proposition from a thinner wrapper around someone else's frontier model.
Cosine lists a red-team application alongside its coding agents, which puts it in front of a security team rather than a product team. Treat that as a separate evaluation with separate success criteria — the question is what it finds on a system you already know the answers for, not whether it writes pleasant code.
There is no free tier, so any description of Cosine as freemium — including an earlier version of this page — is out of date. The published plans are self-serve and credit-based rather than per-seat: Starter at $19/month with 4M credits, which the pricing page addresses to solo developers and side projects; Team at $199/month with 47M credits; and Enterprise at $999/month with 240M credits, aimed at highly regulated industries that prioritise data privacy. Extra credits are sold by the million and get cheaper as the plan gets larger — roughly $6.50, $5.00 and $4.50 per million across the three tiers — which is worth an extra second of arithmetic, because it means the entry plan carries the worst unit price at exactly the point a heavy user starts overrunning its allowance. The credit mechanic is the real budgeting risk, and the vendor is straightforward about why: its own FAQ defines credits as usage across agent work, model calls and cloud execution, with actual consumption depending on task size, model choice and runtime. That is an honest answer and an unforecastable one. The bill tracks how much work the agents do rather than how many people you employ, so a team of four can outspend a team of forty, and nobody — the vendor included — can tell you what one of your tickets costs until you have run a few. Meter a representative sample before committing to a tier. Private and air-gapped deployments sit outside the table entirely: the same FAQ says enterprise and private deployments are scoped with sales because infrastructure, support and security requirements differ, so if a deployment constraint is why you are reading this, treat the published prices as context rather than as your quote. These figures come from the vendor's own pricing page and this is a market where they move — confirm them there before budgeting.
It is the same company, but the positioning has moved. Cosine now presents itself as a sovereign AI lab training models and specialised coding agents for organisations working inside secure environments, with a model line called Lumen. If you came here looking for a general-purpose autonomous coding agent to point at a normal repository, that is no longer what the front door describes.
Organisations whose environment disqualifies the mainstream tools before the comparison starts — air-gapped networks, classified or sovereignty-constrained deployments, security teams, and codebases written in languages that general-purpose models handle poorly. Anyone can buy the $19 Starter plan, and the pricing page still names solo developers on it, so nothing stops an individual trying it. But if none of those constraints describe you, you are evaluating a specialist against incumbents with far more maturity and a much larger support surface, and the entry price is close enough to theirs that it will not be what decides it.
The surfaces the company lists are a CLI and its cloud, not an IDE plugin. That is consistent with the buyer: a team integrating a capability into an existing secure environment and its pipelines, rather than an individual developer installing something into VS Code over lunch.
Treat them as vendor-run, because they are. A benchmark designed and executed by the company whose product it scores is a statement of what the company believes it is good at, which is genuinely informative and is not independent evidence. For a specialist claim like competence in an unusual language, the only test that settles it is running the thing against your own code.
No, and that is not a limitation of this particular vendor. No autonomous coding agent available today should merge unreviewed into a codebase anyone depends on. The specific risk is that a wrong change is internally consistent rather than obviously broken: an agent that misidentifies where a change belongs does not stop, it makes every subsequent edit consistent with the mistake, and the result is harder to catch in review than code that is simply broken.
Only if a rule prevents you from buying those. They solve the everyday problem — making a developer faster at work they are already doing — with far more maturity and a much larger support surface. Cosine is aimed at the case where that option is off the table. The cost of picking it anyway is not really the sticker price, which starts in the same range as a Cursor or Copilot seat; it is a smaller vendor, a thinner ecosystem, and a credit-metered bill that moves with how hard the agents work rather than with headcount.
Full review coming soon.