Option 1
General-purpose AI
Fast and cheap — but incomplete, not executable, and blind to the science. And it needs you to already be the expert who knows how to solve the problem.
Discovers the best prior art, builds on it, and verifies every result against the governing math — then signs a proof you own. It sharpens itself every run and generates what comes next: code, models, agentic systems. Days, not months, at a fraction of the cost.
The core idea
Automation follows a script someone else wrote. A self-generating system runs its own loop: it sets its own goals, acts autonomously, acquires its own resources, and improves itself recursively — each capability feeding back into the core, so it keeps producing what comes next.
The pipeline
The problem
Today's AI is plausible by construction — not correct by construction. That gap is where projects go wrong.
looks right is not is right — we diff every number.
Option 1
Fast and cheap — but incomplete, not executable, and blind to the science. And it needs you to already be the expert who knows how to solve the problem.
Option 2
Correct, eventually. Months of work. $350K–$750K per problem — and no proof you can hand to an auditor.
Option 3 · SingulusAI
Complete, verified, executable solutions in days — correct by construction, token-lean, at an economic cost. The system knows how to solve it for you.
Where we stand
Copilots suggest. Agents ship PRs. Primae discovers, builds, verifies, signs, and remembers — autonomous and domain-correct, token-lean and science-grounded. Most AI is one or the other. We're both.
Autonomous
Manual
speaks your stack — real repos, real toolchains
How it works
The task: move a robot arm to its goal without colliding with the obstacles around it — as directly and quickly as possible.
One plan was hand-built by a team of domain experts over months. The other, our system produced in a few hours. Same scene, same robot — watch the path each one takes.
Built by a team · months
SingulusAI · hours
Straight to the goal, not a detour.
The arm reaches its goal faster.
To compute the plan.
The robot arm is one result. Below is the deeper proof — how we check: three real research problems, the generated code diffed against the reference answer, number by number.
Proof · verification
Execution-based differential testing — not "did the tests pass," but every number the code produces, diffed against the reference. Four steps:
The task: predict a depth map from a single image, measured against the published Depth-Anything-V2 reference. All three codebases are Python/PyTorch, so we feed a byte-identical input through each one's real functions — the metrics, losses, transforms, and forward pass — and diff the numbers.
The radar plots all nineteen output-affecting components on a continuous log-precision scale, one ring per decade — from exact at the edge to 100% error at the centre. Hover any point for its measured error.
The eight modules of the DINOv2 backbone are the network. Since none of these models is trained, we copy the reference's weights into each system's counterpart, feed both an identical input, and diff. If the layer is the same mathematics, the transplanted weights must reproduce the reference exactly — which isolates the architecture from its initialization.
| Reference layer | SingulusAI | Claude | Codex |
|---|---|---|---|
| Attention | identical | 2.2e‑7 | 2.2e‑7 |
| Block (full) | identical | 6.8e‑8 | 6.8e‑8 |
| Mlp | identical | exact · 0.0 | exact · 0.0 |
| PatchEmbed | identical | exact · 0.0 | exact · 0.0 |
| LayerScale ★ | identical | ABSENT | ABSENT |
| DropPath | identical | ABSENT | ABSENT |
| SwiGLUFFN | identical | ABSENT | ABSENT |
| MemEffAttention | identical | ABSENT | ABSENT |
Simulate two-phase gas→oil flow with species transfer across a moving interface, measured against the MassTransferFoam reference. Here the method has to change: the reference is OpenFOAM C++ while two systems are Python, so the code can't be diffed across frameworks — and it is a running simulation.
So we compiled the real solvers, wrote unit tests that call the actual compiled functions, and — first — recorded whether each codebase builds and runs at all: the decisive gate for simulation software.
Detect and classify six damage types with an attention-diversified ensemble and test-time augmentation, measured against the yolov5_SODRv1 reference. All codebases are Python/PyTorch, so this uses the same procedure as Problem 1 — byte-identical inputs through each one's real functions (the box criteria, IoU, non-max suppression, letterbox, scale-back, anchor grid, and average precision), diffed on the same log-precision scale.
Proof · attestation
A verification battery runs, standards clauses cite in, and the result is stamped — a proof you can hand to a reviewer.
V&V battery
convergence · benchmark · invariant · symmetry · MMS · dimensional
Attestation
PASSED · 5 pass / 0 fail
attestation id att-3f9c7a
artifact sha256:9b1e…44c7
capability shown with illustrative values — every run can produce a signed, replayable attestation.
Always on
works on a shadow branch · surfaces only measurable improvements · you accept — never auto-applied.
Where we can help
Domain 01
A wrong sign costs a tape-out, a launch window, or worse.
CFD/FEA at hypersonic regimes — CFL, BCs, turbulence models.
TCAD, lithography, advanced packaging. One wrong Maxwell stencil burns a tape-out.
Battery thermal/electrochemical, crash sim, aero CFD, NVH, autonomy sim.
Seismic, wind, fluid-structure interaction, BIM-coupled analysis.
Domain 02
Force-field conventions, ensemble corrections, convergence — the stuff agents always botch.
MD, DFT, free-energy, PK/PD. Force-field conventions, ensemble corrections, converged DFT setups.
Multi-scale modeling, polymer dynamics, catalyst design, process simulation.
FDA V&V40 makes physics-based sim mandatory — and that means audited, reproducible code.
Domain 03
PDEs, conservation laws, and reactors of every scale.
One product, four reactors of accuracy.
PDE-heavy, grid-based, conservation-critical — from numerical weather to cat-risk reinsurance.
Crop-growth models, soil physics, hydrology.
Domain 04
Where AI literally writes the science — and where every bug compounds.
PINNs, neural operators (FNO · DeepONet), equivariant GNNs, surrogate-coupled solvers — implemented correctly.
Stochastic PDEs, Monte Carlo, high-precision numerics — code that cannot silently lose accuracy.
SAR processing, hyperspectral, atmospheric correction — numerically delicate.
About Primae
Any field. No AI expertise required — bring the problem; keep the signed proof.
Early access Primae is working with an initial group of design partners.
Early access
We're working with a small group of early collaborators. Leave your details and we'll be in touch.