The most advanced AI model is not automatically the right model for a federal mission. Selection should begin with the work being performed, the information involved, the operating environment, and the consequences of error.
A strong recommendation connects model capability to mission evidence. It also considers the surrounding application, controls, deployment approach, and operational ownership rather than a benchmark score in isolation.
Model landscape
Start with the problem, then choose the model family
“AI model” covers several different capabilities. The intended task determines which family belongs in the conversation.
| Model family | Best suited for | Primary consideration |
|---|---|---|
| Predictive | Forecasting, risk scoring, and probability estimates | Quality and relevance of historical data |
| Classification | Routing, categorization, and prioritization | Accuracy across important groups and edge cases |
| Computer vision | Images, video, inspections, and detection | Environmental and image variability |
| Speech and language | Transcription, translation, search, and extraction | Language, accessibility, and context |
| Generative | Drafting, summarization, code, and assistance | Grounding, verification, and permitted actions |
| Specialized | Narrow mission workflows and domain tasks | Training evidence and lifecycle ownership |
The right model family is determined by the work to be performed, not by the visibility of the technology.
System thinking
A model is only one part of the AI system
Capability, control, and accountability are distributed across five connected layers. Leaders should evaluate the complete operating arrangement.
Produces a prediction, classification, representation, or generated output.
Provides prompts, documents, examples, instructions, and runtime data.
Connects users, workflow logic, integrations, permissions, and rules.
Applies security, privacy, testing, human review, logging, and approvals.
Owns monitoring, incidents, cost, continuity, changes, and retirement.
A capable model can still produce an unreliable mission service when the surrounding data, workflow, controls, or operations are weak.
Tradeoffs
General-purpose versus specialized
The decision is not “large or small.” It is which option meets the mission threshold with acceptable risk, cost, latency, and operational burden.
| Decision factor | General-purpose | Specialized or smaller |
|---|---|---|
| Strength | Broad capability and fast experimentation | Focused performance and tighter control |
| Best fit | Multiple evolving tasks | Bounded, repeatable workflows |
| Typical tradeoff | More grounding, cost, or provider dependence | More engineering and lifecycle ownership |
| Evidence needed | Task-by-task tests and guardrails | Domain coverage and change controls |
Deployment choice
How the model is accessed changes the risk
Data exposure, security responsibilities, portability, cost, and change control depend on the delivery arrangement.
Managed service or API
Fast access to vendor-operated capabilities.
- Confirm data use and retention.
- Define version-change notice.
- Plan for availability and exit.
Private deployment
Additional isolation in an approved environment.
- Define shared responsibilities.
- Validate monitoring needs.
- Test cost at mission scale.
Self-hosted
Greater configuration and deployment control.
- Own patching and updates.
- Assess licensing and supply chain.
- Fund full lifecycle support.
Selection framework
Use six lenses to compare candidates
Vendor benchmarks can inform market research. They do not replace tests built around representative agency scenarios.
Task fit
Does the model perform the exact function better than rules, search, analytics, or process redesign?
Information fit
Are the data sources authoritative, permitted, traceable, and suitable for the intended use?
Performance fit
Do mission measures, thresholds, and unacceptable failures support the recommendation?
Operational fit
Can it meet integration, latency, availability, accessibility, and continuity needs?
Control fit
Can the agency constrain access, trace outputs, review decisions, and respond to incidents?
Economic fit
Does the full lifecycle cost include usage, hosting, testing, monitoring, support, and exit?
Minimum evidence package
Require evidence that can survive the demonstration
A repeatable package makes the selection explainable and establishes the baseline for monitoring after deployment.
- Current-state baseline
Compare the model with the existing process and simpler alternatives.
- Representative scenarios
Include normal, difficult, rare, ambiguous, and adversarial cases.
- Acceptance thresholds
Define success measures and failures that cannot be accepted.
- Workflow test
Exercise retrieval, integrations, permissions, human review, and fallback.
- Configuration record
Capture model version, prompts, parameters, sources, and evaluation set.
- Monitoring plan
Set measures, owners, review triggers, rollback, and replacement criteria.
Generative AI
The interaction pattern matters
Similar chat interfaces can hide materially different sources, permissions, and operational responsibilities.
| Pattern | What it does | Control focus |
|---|---|---|
| Closed-book | Responds mainly from learned parameters | Do not treat general knowledge as an authoritative agency source |
| Grounded | Uses approved documents or data at runtime | Source quality, access, citations, and freshness |
| Tool-enabled | Searches, calculates, queries, or initiates steps | Permissions, confirmation, logging, and recovery |
| Fine-tuned | Adapts behavior for a task or domain | Training data, evaluation, documentation, and maintenance |
Acquisition lifecycle
Keep the model decision visible from requirement through operations
Outcome-focused requirements preserve the agency’s ability to compare, govern, and replace the solution.
Define the mission threshold
Set outcomes, authorized AI role, data conditions, evaluation scenarios, documentation, and portability expectations.
Test under consistent conditions
Validate capabilities, limitations, controls, costs, provider claims, and lock-in risk using the same evidence standard.
Manage change as an event
Monitor agreed measures and require review when models, data, configurations, or mission uses change.
Leadership approval
Six questions before approving the recommendation
The answers should be concise, evidence-based, and understandable in mission terms.
- 01What exact task will the model perform?
Name the input, output, user, workflow, and decision supported.
- 02Why this model family and deployment?
Connect the recommendation to the mission and operating constraints.
- 03What evidence shows it meets the threshold?
Use representative tests instead of general rankings or a polished demonstration.
- 04How are outputs verified and actions controlled?
Define sources, review, permissions, logging, escalation, and correction.
- 05What happens when the model changes?
Require notice, regression testing, approval, rollback, and documentation.
- 06Can the agency sustain or replace it?
Confirm lifecycle cost, portability, operating ownership, and exit readiness.
TDG perspective
Model selection is a mission architecture decision
The federal enterprise does not need one model for every problem. It needs a repeatable way to select, test, integrate, and operate the right capability for each mission context.
That discipline produces clearer requirements, proportionate controls, stronger acquisition decisions, and a credible path to replace the model when the mission, evidence, technology, or risk changes.
Selected federal AI references
- OMB M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust
- OMB M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government
- OMB M-26-04, Increasing Public Trust in Artificial Intelligence Through Unbiased AI Principles
- NIST Artificial Intelligence Risk Management Framework
- NIST AI 600-1, Generative Artificial Intelligence Profile
- NIST AI 700-1, Text-to-Text Evaluation Overview and Results
- GAO-26-107859, Artificial Intelligence Acquisitions



