
By BarathVector Editorial — 2026-08-15
India now has enough artificial-intelligence announcements to fill a scoreboard. It still needs the scoreboard.
As of July 2026, the IndiaAI Mission said it was supporting 20 foundation-model proposals: 12 large language models and eight small language models. The programme has also assembled public compute, datasets and safety projects. This is real industrial policy, not a slogan.
But a count of funded models is an input. Sovereignty is an outcome. India will not become AI-sovereign because twenty teams complete training runs, any more than it becomes energy-secure by counting power plants without asking whether they deliver electricity when needed.
The country needs a public sovereign-AI scorecard that measures what each supported system can do, for whom, at what cost and under whose control.
Sovereignty is the ability to choose
The word is often used as if it means domestic ownership. Ownership matters, but it is only one dependency. A model trained by an Indian company on rented foreign accelerators, using imported software and licensed data, may still be useful. A fully domestic model that cannot serve Indian users at an affordable price may be sovereign on paper and irrelevant in practice.
Operational sovereignty means retaining credible choices across the stack: compute access, data rights, model weights or dependable interfaces, skilled people, evaluation capacity and the ability to deploy under Indian law. It also means being able to change suppliers without rebuilding an essential public service from zero.
That standard does not require autarky. India has always advanced through selective partnerships. It requires knowing where dependence sits and ensuring that no single external actor can withdraw a critical capability without a practical substitute.
The mission has built useful foundations
The government's July 2026 update reported 20 supported model proposals, 237 approved AI projects and about 9.3 million GPU hours allocated through the national compute programme. AIKosh is intended to give researchers and developers access to datasets, models and tools.
These investments attack genuine constraints. Indian teams need compute without negotiating every experiment from scratch. They need lawful, well-documented data in languages and domains that global products neglect. They need evaluation and safety capacity that is independent of a vendor's marketing page.
Several projects report ambitious coverage. BharatGen, for example, says its text models support all 22 scheduled Indian languages and that its Param2 system was trained on 22 trillion tokens. Those are the developer's claims and should be treated as claims until reproduced under public tests. That is not suspicion unique to BharatGen. It is the discipline every publicly supported model should face.
What the scorecard must measure
First, language quality must be tested outside translation-friendly benchmarks. A model should handle dialect, code-switching, script variation, government vocabulary and culturally specific questions. Average performance can conceal severe failure in a smaller language. Results should therefore be published separately for every claimed language and task.
Second, capability must include Indian use cases rather than a collection of foreign academic tests. Can the system explain a welfare notice without changing its legal meaning? Can it retrieve the correct agricultural advisory for a district? Can it refuse a dangerous medical instruction while preserving useful general information? Can it cite the source rather than invent one?
Third, the public needs the economics. Publish training compute, serving cost, latency, energy use and the hardware required to run the model. A smaller model that performs a narrow public task cheaply on available infrastructure may contribute more sovereignty than a prestige system that requires a scarce cluster for every query.
Fourth, disclose access and control. Are weights available? Under what licence? Can a government agency deploy the model in its own environment? Who can change the terms? What happens if the original developer fails? A model accessible only through an opaque endpoint can be a good service and a poor sovereign asset.
Fifth, safety tests must be repeatable. Measure data leakage, cyber misuse, caste and religious bias, harmful advice and performance after adversarial prompting. Publish both failures and mitigations. A state-funded model should not be protected from scrutiny by the claim that disclosure would embarrass the mission.
Benchmarks can lie politely
A scorecard can also become propaganda. Teams can train against known test sets, choose favourable averages or exclude invalid answers. The Stanford AI Index 2026 notes that some frontier evaluations have reported invalid-response rates as high as 42 percent and that the gap between open- and closed-weight systems varies sharply by task.
India should therefore use independent evaluators, hidden test sets and periodic refreshes. Every score needs confidence intervals, sample sizes and an explanation of what counted as invalid. Human evaluation should include speakers from the communities whose language is being claimed, not merely bilingual annotators in a major city.
The scorecard should be a public service, not a league table. Different models will be best for different jobs. Its purpose is to reveal trade-offs and help buyers choose, not crown one national champion.
The strongest countercase
Young model teams cannot disclose every architectural and cost detail without giving away commercial advantage. Early public rankings may punish experimentation and push developers toward safe, benchmark-friendly systems. Sovereign capacity also includes companies strong enough to survive, so policy should not turn grants into compulsory self-exposure.
That concern is legitimate. Publication can be staged. Security-sensitive data, proprietary training recipes and exact infrastructure vulnerabilities may remain confidential. But a team accepting public compute or grants owes the public verifiable performance, material limitations, licensing terms and total support received. Commercial secrecy can protect a method; it cannot substitute for evidence that the promised result exists.
Fund milestones, not announcements
Future tranches of public support should be tied to measured milestones: demonstrated language coverage, reproducible task performance, a serving-cost target, a deployment outside the developer's laboratory and a documented path to continuity. Failed experiments should not be treated as scandal if their results are reported honestly. Research policy becomes wiser when negative evidence survives.
India should also publish aggregate dependencies: which accelerators, cloud regions, software components and foreign licences the supported portfolio requires. That view will identify the next useful investment far better than another race to announce the largest parameter count.
Twenty model projects can seed an ecosystem. They cannot by themselves prove independence, public value or strategic resilience. The proof begins when the models are tested in Indian languages, priced for Indian institutions and deployed under terms India can control.
Count the models if the number is useful. Then open the scorecard. Sovereignty starts where the announcement ends.