Public-Sector AI Use-Case Scorecard
Self-initiated · 2026 · Applied AI / public sector
Open the live tool →Most public-sector AI never makes it out of the demo. A model does something impressive in a workshop, everyone's excited, and then it dies in the gap between that demo and a system a ministry can actually trust and run.
The hard part isn't the model. It's deciding which use cases are worth doing, being honest about the risk, and knowing what "reliable enough to ship" concretely means before anyone builds.
I built a working tool that models the front of that funnel: the part a forward-deployed product person owns before engineering starts. You scope a candidate use case (entity, problem, users, solution pattern), score it across five levers (mission impact, feasibility, data readiness, trust and risk exposure, time to value), and it computes a priority and plots it on a value-versus-feasibility quadrant.
The part I care most about is what happens next. It generates an evaluation plan (the specific metrics and the eval set) that changes by solution pattern, because evals for document extraction look nothing like evals for an agentic workflow or a forecasting model. And it produces a production-readiness gate that gets stricter as risk rises, up to human-in-the-loop review, red-teaming, and data residency for sovereign contexts.
I prototyped the whole thing hands-on in Claude Code. The scoring is deliberately transparent and opinionated. It's not a validated model. What it does is force an honest conversation about reliability before build, which is exactly the conversation these projects tend to skip.
A self-contained, working demonstration of how I'd take public-sector AI from pilot to production, and of a way of working I think is fairly rare in product: building the thing rather than just describing it. Open the live tool to try it.