Skip to content

Service-disabled veteran-owned

AI for public institutions, built to be audited, not just demoed.

Flint yard builds applied AI systems for government, and the evaluation work that proves they do what they claim. Commercial-scale engineering, held to a standard the public can inspect.

Ownership
Veteran-owned
Primary NAICS
541511
Based in
Sacramento, CA
Engagement
Prime or subcontract
§01

What we do

01

Applied AI systems

Document processing, eligibility triage, case summarization, and retrieval over records that were never designed to be searched. Built with human review in the loop and a clear record of why the system produced what it produced.

Document processingRetrievalHuman-in-the-loop
02

Evaluation and assurance

The part almost nobody does. Before an AI system touches a benefit determination or a public-facing decision, there has to be a way to measure whether it is right, and to explain months later why it did what it did. We build that measurement first.

Eval harnessesProvenanceAudit trailsFailure analysis
03

Software people actually use

Public-facing web applications, mobile apps, and the internal tools that caseworkers and program staff sit in front of all day. Most of the value in a program is decided by whether the people operating it can do their jobs without fighting the software.

Web applicationsMobileInternal toolsAccessibility
04

The systems underneath

AI and applications are only as good as the data and infrastructure beneath them. Legacy integration, data migration, and load engineering: the unglamorous work that decides whether anything above it survives contact with real volume.

Legacy integrationData migrationLoad engineering
§02

Where to start

Four engagements that are small enough to be awarded quickly and specific enough to be judged on their results. Each produces something written that you can circulate, act on, or use to justify what comes next.

Assessment

AI readiness review

Two to three weeks inside a program, looking at the actual workflows, data, and staff capacity. You get a written assessment of where AI would genuinely help, where it would not, and what it would cost to find out.

Written assessment · 2–3 weeks
Measurement

Evaluation harness

For a program that already has an AI system, bought or built, and cannot tell whether it works. We build the test set, the scoring, and the reporting, so the question stops being a matter of opinion.

Working eval suite · 3–6 weeks
Pilot

Document processing pilot

One narrow, high-volume document workflow, automated end to end with human review and a full audit trail. Scoped so that a failure is cheap and a success is obvious.

Running pilot · 6–10 weeks
Review

Technical due diligence

An independent read on a troubled program, a vendor proposal, or an architecture you have been asked to approve. Written plainly enough to hand to a non-technical decision maker.

Written review · 2–4 weeks
§03

How we work

The people who scope the work are the people who write the code. No account manager in between, no hand-off to a different team after the award, and no pyramid where the engineers who won the contract are not the ones who show up to build it.

01

Discovery

Two weeks with the people who use the system and the staff who maintain it. What gets written down is what actually happens, not what the process document says happens.

02

Working software

Something real in a staging environment inside the first month, and every sprint after that. Progress is demonstrated by software you can click, not by a status deck.

03

Handoff

Your team owns it when the engagement ends. Documentation, runbooks, and pairing throughout, not a knowledge-transfer meeting scheduled for the final week.

Flint yard takes on a small number of engagements at a time and staffs them deeply rather than broadly. That is the right trade for programs where the binding constraint is judgment rather than headcount: a modernization that needs a plan before it needs forty contractors.

As Flint yard grows, every hire will be a veteran or the family member of one. That is a hiring commitment, not a marketing position.

§04

The experience behind Flint yard

Flint yard is a veteran-owned firm. The engineering experience behind it was built in the commercial technology industry, across three kinds of work that turn out to matter for public programs.

Building AI at scale. Machine learning and AI systems in production, serving real traffic, where the failure modes are measured rather than guessed at. That is a different discipline from building a demo, and it is the one government programs actually need.

Founding and funding startups. Founding a venture-backed company through Y Combinator, and the years of shipping under real constraint that go with it. This is where the conviction about small scope and short feedback loops comes from. Those habits are not a methodology preference. They are what survives contact with a deadline and a real user.

Consulting with Fortune 500 companies. Years of working inside large organizations with entrenched systems, competing priorities, and staff who have watched consultants come and go. Government programs are more like that than they are like a startup, and knowing how to be useful in that environment is its own skill.

Flint yard exists because the gap between how commercial systems get built and how government systems get built is not a talent gap. It is a structure gap: incentives, contract shape, and feedback loops. Most of what makes public software fail is decided before anyone writes a line of code.

Prior engineering experience at

NetflixCLEARIBMGoogle

Former employers, listed as engineering background. They are not Flint yard clients and imply no endorsement.

§05

How to contract with us

Micro-purchase

Under $15,000

Awardable on a government purchase card without competitive quotes. The fastest way to start a working relationship, typically a scoped assessment or a two-week discovery engagement.

Simplified acquisition

$15,000 to $350,000

Simplified acquisition procedures, reserved for small business. This is where most of our work sits: a defined problem, a fixed scope, and working software at the end of it.

Subcontract

Under a prime

We subcontract to primes carrying small business subcontracting plans, and we team under FAR 9.6 arrangements on larger pursuits where the fit is genuine.

Flint yard is structured to be easy to buy from. Most first engagements are scoped as a discovery sprint or an assessment with a concrete written deliverable, which keeps the initial commitment small and lets the work prove itself before anything larger is contemplated.

Most government software fails for reasons that have nothing to do with technology. Flint yard works on the parts that actually break.

Have a program that needs to actually ship?

Flint yard responds to sources sought notices, RFIs, and market research requests, and is glad to talk well before any of that.