Discover key insights from Lumen Sovereign's RFI: what critical sectors like finance, defence, and healthcare demand from secure, sovereign AI models.
We have spent the last three months running a public Request for Input (RFI) for Lumen Sovereign, Britain’s first fully sovereign frontier AI model. Before committing half a million GPU hours on Isambard-AI, we wanted to hear from the institutions that will be using this capability.
The volume and depth of response exceeded our expectations. We received detailed submissions from defence primes, banks, energy operators, consultancies, and research bodies from across the UK. This report summarises what they shared.¹
What our RFI asked
The RFI for Lumen Sovereign covered six questions:
- Capabilities: what high-value tasks should the model prioritise?
- Shortfalls: where do current frontier models fall short in workflows?
- Safety & governance: what assurance, red lines, and oversight mechanisms must be built in?
- Datasets & benchmarks: what private or sector-specific evaluation approaches should we use?
- Tools: what enterprise systems must the model integrate with?
- Anything else: open feedback on deployment, cost, or ecosystem design.
Responses we heard back
1. The capability brief: software engineering is critical, but too narrowly focused
Every respondent agreed that strong coding and secure software engineering are essential. However, a consistent theme across submissions was that a sovereign model optimised only for software engineering would be too narrow for the operational value these institutions need.
Multiple respondents argued that the same core problem-solving steps being developed for coding – long-context evidence use, decomposition, iterative refinement, recovery from failure, tool use, and verification before claiming success – are directly transferable to adjacent high-value workflows, but respondents want a model able to go further. The suggested expansion areas mapped directly to critical sector operations included:
- Financial services & banking: deep regulatory comprehension for KYC customer onboarding, AML alert triage, and automated financial crime investigations. Broad themes included actuarial pricing, credit risk reporting, regulatory change impact assessments, and using LLM-as-a-judge patterns to verify compliance against strict national regulations.
- Defence & systems engineering: support across full systems engineering lifecycles, safety and assurance case co-piloting, multi-source intelligence synthesis, and data-driven “what-if” operational readiness support. Respondents also wanted small-footprint models that could deploy at the edge in severely constrained environments.
- Energy & critical national infrastructure: navigating highly regulated environments to manage field safety, operational risks, and automate predictive maintenance across large physical estates.
- Cyber triage & defensive security: secure configuration review, incident summarisation, and vulnerability remediation, without the over-refusal often seen in commercial models applied to defensive contexts.
- Healthcare & life sciences: structured multi-agent coordination for complex scientific workflows, biomedical information extraction, and medical reasoning governed by strict clinical safety rubrics.
- Telecommunications: automated integration with internal business and operational support systems (BSS/OSS) and infrastructure scripts.
- Legacy estate comprehension: modernising brownfield codebases in older and niche languages (such as Ada, COBOL, Delphi, and Assembler) where modern frontier models – typically tuned for Python and TypeScript – struggle to provide value.
One respondent put it directly: “Coding should remain the proving ground, but not the entire identity of the system.”
A related point, raised by several respondents, is that these use cases deliver value not from the model in isolation but from the harness surrounding it. A model that can reason well but cannot reliably call the right tool (e.g. query a core legacy database, integrate with enterprise project management tools, or navigate a secure CI/CD pipeline) will struggle to deliver operational value regardless of its raw capability.
2. Shortfalls in current frontier models go beyond accuracy
The most common shortfall cited was trust. Respondents repeatedly described current models as fluent but opaque. Defence and other regulated-sector users need to know which sources were used, whether those sources are current versus superseded, what assumptions were made, what evidence supports the answer, and where uncertainty remains. Without traceability and the ability to challenge outputs, model-generated answers are difficult to use in safety, assurance, engineering, operational, or programme decision-making.
A second shortfall was sovereignty and supply-chain clarity. Multiple respondents cited concerns about overseas infrastructure dependencies, opaque data handling, proprietary APIs, and sub-processor chains that cannot be enumerated.
Third was long-context and multi-document workflow bottlenecks. Regulated tasks frequently require reasoning across requirements, architecture descriptions, safety cases, trial reports, risk registers, standards, contractual artefacts, code repositories, and configuration records. Current models were described as struggling to preserve provenance, distinguish authoritative from draft material, and maintain a clear line of reasoning across long evidence chains.
Fourth were cost transparency and deployment concerns. Consumption-based pricing, rate limits, vendor lock-in, and limited air-gapped deployment flexibility were all cited as adoption barriers. One respondent noted that if a system uses millions of tokens to solve a task, cost transparency is needed, and that upfront capital expenditure for owned infrastructure can be preferable to ongoing API dependency.
Fifth was meta-cognition and reasoning transparency. Standard chain-of-thought was described as insufficient for realistic meta-cognition, which refers to a model’s ability to evaluate its own reasoning while executing a task. Respondents argued that reasoning processes must be shown, not hidden, so that if the system goes wrong, operators can see exactly where it failed and add guardrails or modify prompts accordingly.
3. Safety and governance, assurance-by-design
Governance expectations were consistently stringent and specific.
Human oversight should be risk-based rather than blanket. Several respondents argued that bounded, low-risk workflows may be appropriate for autonomous operation within an agreed risk appetite, but that any actions affecting safety, mission outcome, legal position, personnel, or material commercial commitment must require human review.
Air-gapped and high-security deployment was treated as a design principle. Requirements included customer-managed keys, strict identity and access management, compartmentalisation, protective monitoring, and clear separation between customer environments with no external data transfer.
Auditability and provenance must be systematic and domain-relevant. Suggested testing areas included: hallucination; data leakage; unsafe cyber enablement; prompt injection; over- and under-classification; mishandling of sensitive material; tool misuse; behaviour in edge cases. Several respondents stressed that automated benchmark scoring is necessary but insufficient; expert blind review, source-grounded answer checking, adversarial testing, and longitudinal monitoring after deployment are essential.
Multiple defence-facing respondents raised safety-case support. The model should help draft, structure, and challenge assurance evidence, but its output must not be treated as the evidence itself. The value comes from accelerating the assurance workflow while preserving independent scrutiny and accountable human judgement.
4. Datasets and benchmarks
Public benchmarks were almost uniformly described as table stakes but increasingly uninformative. Concerns included saturation, contamination risk, and the fact that standard benchmarks do not measure how often the model is confidently wrong in a workflow where nobody checks.
The consensus alternative was layered evaluation that combined:
- Public benchmarks for generic capability baselines
- Synthetic sector tasks tailored to the relevant operational context
- Private or controlled workflow evaluations using internal documents and tools
- Expert human assessment, including blind review and adversarial testing
Specific evaluation areas proposed included:
- Assurance and systems engineering: requirements decomposition, inconsistency detection, ambiguity identification, change-impact analysis, and verification mapping over long, multi-document contexts.
- Mission data exploitation and operational analysis: scenario construction, assumptions capture, course-of-action comparison, and sensitivity analysis with clear distinction between evidence, interpretation, and conjecture.
- Defensive cyber and resilience: vulnerability impact analysis, secure configuration review, incident summarisation, and remediation planning, with red-team controls to ensure the model supports defensive outcomes without enabling misuse.
- Secure software engineering: code comprehension, vulnerability identification, secure refactoring, and test generation over legacy code with incomplete documentation.
- Faithful long context reasoning: contradiction detection, evidence reconciliation, uncertainty calibration, and robustness to hidden distractors.
- Scientific and clinical safety: evaluation governed by medical safety rubrics and controlled dual-use risk testing.
Generally, respondents noted that the strongest evaluation design would mirror the philosophy used for coding evaluation: not only testing whether outputs sound plausible, but whether they use context faithfully, recover from bad intermediate states, and produce outcomes useful in real-world workflows.
5. Tools
Respondents were clear that the model must operate inside the systems where their work already happens rather than replacing them. Tool integration requirements spanned:
- Document and knowledge repositories: controlled read, summarise, compare, and draft across Word, PDF, PowerPoint, Excel, Markdown, and structured document sets with version awareness, citation-level provenance, and strict respect for user permissions.
- Engineering lifecycle environments: requirements management, architecture modelling, test and evaluation repositories, safety and assurance case tools, configuration management, and issue trackers with traceability across the lifecycle.
- Code and cyber workflows: access approved code repositories, CI/CD pipelines, static analysis, security scanners, test frameworks, and vulnerability management systems, with the ability to explain legacy code, generate tests, and propose secure refactoring within governed boundaries.
- Data and analytics: query approved databases, analyse spreadsheets and structured datasets, generate reproducible reports, and produce visualisations and evidence packs with data-classification awareness and robust auditability.
- Operational and simulation environments: scenario generation, parameter explanation, output interpretation, and structured after-action review for wargaming, mission planning, and autonomy support where security controls allow.
- Governance and audit tooling: native review of model-use logs, prompt and response history, retrieval records, tool calls, policy enforcement events, and human approval steps.
- Enterprise platforms: internal knowledge bases, case management tools, link-analysis systems, and ticketing and workflow orchestration.
6. Ecosystem and deployment design
Several respondents raised structural points about the model ecosystem:
SME and supply chain enablement was highlighted by one respondent as a major sovereign advantage. A UK sovereign model should be designed for adoption across the whole industrial base. It should work for specialist engineering organisations, small primes, SMEs, and secure supply chains as well as the largest primes. Alongside model capability, this requires education, training, deployment pattern workshops, and structured feedback mechanisms.
Multiple respondents requested appropriately sized deployment tiers. Not every use case needs a frontier-scale model. One respondent asked for a “right-sized, cost-effective deployment tier” with predictable costs, immutable version control, and clear exit and continuity provisions. Another requested small language models for edge deployment targeting 16GB VRAM or less, given Size, Weight, and Power (SWaP) constraints in on-edge autonomous systems.
Model continuity and provenance were stressed as essential. Respondents want clear provenance of training data to mitigate vulnerabilities and bias, immutable version control, a Software Bill of Materials for supply chain assurance, and the ability to self-assess output quality and flag risks to users.
What we are doing in response
These are just some of the directional signals we’re using to guide our progress. Whether regulated organisations deem Lumen Sovereign to be useful, safe, and worth deploying is the ultimate test.
The consistent message is that a sovereign frontier model must be judged on operational fit rather than generalised benchmark performance. The UK institutions that responded want:
- Strong coding as an anchor, but reasoning capability beyond that
- Transparent, challengeable outputs with explicit provenance
- Air-gapped deployment
- Deep integration with existing enterprise stacks
- Assurance-by-design
- Evaluation that mirrors real workflows
- Ecosystem support that enables adoption across the full industrial base
These are not easy requirements, but that is the point. Lumen Sovereign is being built for the workflows where sensitivity and complexity have historically made AI adoption impossible. Delivering on this mandate aligns with Cosine’s DNA. With a proven track record of building models that can navigate complex software engineering tasks, we have already tackled the hardest parts of autonomous reasoning: orchestrating multi-step tool use, maintaining accuracy over long contexts, and recovering from failure. By translating our deep code expertise into broader operational workflows, we are uniquely positioned to deliver a sovereign model that reliably executes complex tasks.
Since forming the coalition and opening the public input process, we have moved through several concrete stages of our training on Isambard-AI. The feedback we received is actively shaping our training curriculum, deployment targets, and evaluation strategies. Based on these responses, we are making active adjustments to Lumen Sovereign’s training:
- Exploring a 200B checkpoint release: to meet the demand for long-context reasoning across specific, complex enterprise workflows, we are currently exploring the release of a 200B parameter checkpoint for Lumen Sovereign.
- Deep-domain post-training: we are including highly specific, coalition-assisted data into our post-training pipeline. This includes complex regulatory reasoning (such as KYC and AML workflows) and biomedical information extraction.
- Targeting edge hardware: multiple critical sectors highlighted the need for autonomous, on-edge deployments operating under strict SWaP constraints. We are actively adapting our training strategies to ensure Lumen Sovereign can operate effectively in these constrained edge environments. More on this soon.
- Adopting sector-specific evaluations: we are moving beyond saturated public benchmarks to produce specific, workflow-driven evals. Based on the RFI, we are building evaluations around defence principles, incorporating telecom-specific frameworks, testing against clinical safety rubrics, and more.
Keeping up with Lumen Sovereign
We will continue to publish updates on how the design and training are evolving in response to this input. This process is ongoing, and if your organisation operates in a highly regulated UK sector and you want to contribute to our evaluation phases and future training runs, submit your input here.
If you want to follow the build as it happens – from data and architecture decisions through to post-training and evaluation, along with exclusive early previews – the best place is the Lumen Sovereign newsletter, where we share progress every two weeks.
Thank you to all our RFI respondents, and especially our Lumen Sovereign coalition partners.
__________________
¹ Several respondents gave permission to quote their feedback anonymously. This report is built entirely from that authorised material. No identifying details are included, and no organisation names are attached to specific quotes.