Your Models Work in a Notebook. Production Is a Different Problem.
Senior engineering ownership for machine learning in regulated finance: a real deployment path, monitoring that measures realized outcomes, and validation evidence that survives a sponsor bank questionnaire or an examiner asking who signed off.
When to call us
The models work in a notebook and nowhere else
Your data scientists have models that perform well offline. Nothing has a deployment path, a retraining trigger, or an owner on call. The gap is engineering, and hiring more data scientists widens it.
Someone asked for model documentation and there is none
A sponsor bank, an auditor, or a board risk committee has asked how a model was validated and who signed off. The answer currently lives in a notebook, a Slack thread, and whatever one engineer happens to remember.
Nothing watches the model after deployment
No drift detection, no performance monitoring against realized outcomes, no threshold that triggers revalidation. The model is quietly getting worse and the first signal will be a loss or a complaint.
An AI feature is blocked on governance nobody owns
The build is ready and cannot ship because there is no framework for approving it. Legal will not sign off on a process that does not exist, and engineering is not going to invent a control framework unprompted.
Where the big firms and we differ
Ask who does model risk work in finance and you get the large consultancies. They are genuinely good at governance frameworks. The gap is narrower than it looks.
| What you need | A large consultancy | Us |
|---|---|---|
| An enterprise-wide model risk framework and policy set | Better fit | We would decline |
| An independent validation opinion for a regulator | Better fit | We cannot, we sit on the build side |
| Someone to write the deployment pipeline and the monitoring | Offered, with more program overhead | Better fit |
| One model in production with monitoring and validation evidence, inside two months | Higher overhead at this scope | Better fit |
| Answering a sponsor bank diligence questionnaire next month | Depends on bench availability | Better fit |
If your problem is a policy framework for a large bank, hire one of them. Most of the teams that call us have the opposite problem: working models, no production path, and a diligence deadline.
What you actually get
- A written model inventory: every model in use, its owner, its purpose, its inputs, its last validation date, and its risk tier
- Validation evidence covering the three things any reviewer asks: whether the approach is sound, whether it still performs, and how predictions compare against realized outcomes
- Production deployment path: versioned artifacts, reproducible training, staged rollout, rollback, and a retraining trigger that is a rule rather than a decision
- Monitoring that measures the thing that matters: realized outcomes against predictions, input drift, segment-level performance, and alert thresholds that map to revalidation
- A governance process light enough to survive contact with a shipping team: who approves what, at what risk tier, with what evidence
- Escalation and change control, so a model change does not silently become an unreviewed policy change
Named proof
Engagements where the models reached production and stayed there.
Latent Sciences
Boston, MAServerless ML infrastructure connecting experiments to on-demand GPUs. Cut idle infrastructure cost 98% and raised GPU utilization 4x.
AI Trader
AnonymizedMulti-model agentic platform across four LLM providers with quality-gated iteration and schema-validated outputs. 3x research throughput, unattended operation.
Global Advertising Leader
AnonymizedKnowledge graph and AI platform replacing weeks of manual audience research. 28,000+ segments across 9+ markets, in production.
How we work
30-min discovery call
What models exist, who is asking questions about them, and what is actually blocking production. We will tell you if the honest answer is that you need a validator rather than an engineer.
Model risk and production assessment
2-3 weeks, fixed fee. Model inventory, gap analysis against supervisory expectations, and a written plan ordered by risk rather than by ease. You keep the document whether or not we do the build.
Embedded build
We pair with your team on the deployment path, the monitoring, and the first two validation packages, so the pattern is one your engineers can repeat without us.
Frequently asked
Are you a model validation firm?
No. We sit on the build side: we get models into production, instrument them, and produce the evidence a validator or an examiner will ask for. On independence the 2026 guidance is less rigid than most people expect, saying the quality of the validation process depends on the rigor and effectiveness of the review rather than on the organizational structure of the risk management function. We would still not validate our own build and call it independent. If what you need is an independent validation opinion, you want a separate firm, and we will tell you that on the first call.
Which supervisory guidance does this map to?
US interagency model risk guidance was revised on 17 April 2026. Federal Reserve SR 26-2 supersedes SR 11-7, which had been the reference point since 2011, with companion issuances from the OCC and FDIC. Anything published before April cites the 2011 letter, and plenty of internal model risk policies still do too. If yours does, that is a reasonable first thing to look at. What we will not do is tell you what the revision means for your specific models before reading your policy set against the current text, and you should be wary of anyone who will.
We are a fintech, not a bank. Does model risk guidance apply to us?
Not directly. Federal model risk guidance is addressed to supervised banking organizations. It reaches a non-bank in practice through the partner: a sponsor bank is accountable for the models used in its programs, and it discharges that through third-party risk management. So the expectations arrive in your diligence questionnaire rather than from a regulator, and arguing about the technicality does not help you answer the questionnaire.
Does an LLM count as a model?
Under the April 2026 interagency guidance, no. Footnote 3 places generative and agentic AI outside its scope, on the stated grounds that both are novel and rapidly evolving. That is narrower than most people assume, and it is not the relief it sounds like: out of scope for this guidance is not the same as unreviewed, and a sponsor bank will still ask what controls you run before it lets a language model near a customer decision. So the work does not go away, it just stops being a model validation exercise. We validate the system rather than the weights: retrieval, guardrails, the evaluation harness, refusal behavior, and which human is accountable when it is wrong.
Our models are simple rules and gradient boosting, not AI. Is that in scope?
Gradient boosting, yes. Simple rules, probably not. The April 2026 guidance defines a model as a complex quantitative method applying statistical, economic, or financial theories, and it explicitly excludes deterministic rule-based processes and software where no such theory underpins the design or use. So a hand-tuned rules engine can sit outside the definition even while it makes consequential decisions. That catches people out, and it cuts both ways: outside the definition means nobody is required to validate it, which is how a rules engine quietly becomes unreviewed policy. We would put it in the inventory anyway.
How long until we can ship?
For a single model with a clear owner, 4-8 weeks from kickoff to a production deployment with monitoring and a validation package. A portfolio of undocumented models is a longer job, and the first pass should triage by risk rather than attempt everything. We would rather get your three highest-risk models properly governed than produce thin paperwork for thirty.
What if our data science team does not want this?
Reasonable of them, if governance has previously meant a spreadsheet they fill in for someone else. The version worth having gives them a deployment path they do not currently have, and takes the audit correspondence off their desk. If the process does not make shipping easier, it will be circumvented, and a circumvented control is worse than an acknowledged gap.
Related
AI/ML engineering for finance
The broader build: inference infrastructure, LLM integration, multi-model orchestration, cost discipline.
Why FinTech ML models stall before production
The specific engineering gaps between a model that scores well offline and one that runs.
PCI DSS and SOC 2 readiness
The other compliance workstream that usually arrives with the same diligence questionnaire.
Ready to talk?
30 minutes, free, no pitch. Bring one model you cannot ship and we will tell you what is actually in the way.