Governments are being asked to choose and deploy AI for cyber defense, critical infrastructure, software assurance, and other public missions. In many cases, procurement is moving faster than the evidence needed to support it.
A model demonstration can look convincing while leaving the important questions unanswered. Did the system complete the task? Did the result hold under hidden checks? What did it cost? How did it fail? Could an independent reviewer reproduce the conclusion?
THE QUESTIONHow well does this system perform the security task the mission actually requires?
01 / WHAT IT IS
An independent evaluation practice for public missions.
OrbitCurve for Government is a focused evaluation service for agencies and public sector teams considering AI systems for cybersecurity work. We do not sell the models being evaluated. We design the task population, control the execution conditions, define success before the run, and preserve the evidence behind the result.
The output is not a generic score. It is a decision record: what the system was asked to do, the conditions it operated under, the actions it took, the artifacts it produced, the checks it passed, the resources it consumed, and the failure modes we observed.
02 / WHY SECURITY IS DIFFERENT
Cybersecurity performance lives in actions and artifacts.
Security work is not well measured by whether an answer sounds plausible. A binary is recovered or it is not. A crash is reproduced or missed. A patch survives hidden malformed inputs or fails. The evidence is usually found in the environment, not in the final paragraph.
Our evaluations therefore grade outcomes wherever possible. Capability is reported beside cost, latency, time, tokens, and compute so that a high score cannot hide an impractical operating profile.
03 / BENCHMARK PROGRAMS
Real security work with objective consequences.
Our benchmark programs are built around tasks where the outcome can be checked without relying on a model to judge another model. They are designed to expose long-horizon behavior, brittle shortcuts, and the difference between producing an answer and completing the work.
DeobfuscationBench
Measures whether models and agents can recover independently usable code and technical artifacts from heavily obfuscated x86-64 binaries over long-horizon tasks.
APEX CyberMicro Bench
Measures native-code vulnerability triage and patch validation through sanitizer evidence, deterministic milestones, regression checks, and hidden malformed inputs.
View the public repository ↗04 / THE ORCHESTRATION LAYER
Akula turns an evaluation plan into an evidence record.
Akula is OrbitCurve's proprietary orchestration layer. It binds the benchmark, model, agent, environment, budget, and execution policy into one controlled plan. It then captures the run identity, trajectories, artifacts, grader output, and resource use needed to review the result.
Deployments are scoped for each customer environment and security requirement. The objective is consistent: make every recommendation traceable to the run that produced it.
05 / WORKING TOGETHER
Start with one consequential mission task.
An engagement can begin with a focused evaluation of a single system and task. From there, the work can expand into a private benchmark suite, a procurement study, or recurring assurance as models, agents, and operating conditions change.
The first step is to define the decision you need to make and the evidence that would make that decision defensible.
