The prompt is no longer the whole job. As AI systems move from answering questions to carrying out multi-step work, organizations are not only delegating to a model—they are depending on a vendor, a toolchain, permission model, data policy, integration surface, and operating environment.
That changes due diligence.
A capable demo is evidence that a tool can do something. It is not evidence that the tool is appropriate for your data, risk, workflow, or customers.
The operating contract for AI therefore has two sides: the instructions you give the agent and the conditions under which you are willing to trust the vendor that runs it.
Why a prompt is not enough for delegated work
A prompt tells a system what you want. An operating contract tells it what counts as acceptable work.
If the task is “research this topic,” the contract defines acceptable sources, date range, citation rules, contradiction handling, and stop conditions. If the task is “update this product page,” the contract defines canonical templates, fields that may change, fields that must not change, the controlling source material, validation steps, and the release gate.
That is the difference between assistance and accountable execution.
Vendor due diligence starts before the workflow
Before an AI tool receives real business data or permissions, evaluate the environment it creates around the model.
NIST’s Generative AI Profile extends the AI Risk Management Framework with considerations specific to generative AI and frames risk management across the AI lifecycle. For a small business, the practical translation is not a 100-page procurement process. It is a disciplined set of questions that prevents convenience from silently becoming exposure.
The nine-part AI vendor due-diligence record
-
Business purpose.
What exact job will the tool perform, and what value justifies introducing another vendor or data flow? -
Data scope.
What information will be sent to the tool—public content, customer records, financial data, source code, proprietary research, credentials, or internal strategy? Classify the data before connecting it. -
Training and retention.
Document whether submitted business data is used to train models by default, what retention controls exist, and what deletion or export options are available. Verify the vendor’s current policy rather than assuming all plans behave the same way. -
Permissions and actions.
List what the system can read, create, edit, publish, send, delete, purchase, or otherwise execute. Prefer the least privilege that still allows the workflow to work. -
Security and access controls.
Review authentication options, role-based controls, admin visibility, audit logging, encryption claims, and relevant compliance documentation for the risk level involved. -
Model and service change risk.
Understand whether the underlying model, feature behavior, connector, or pricing can change without your workflow changing. Critical workflows need regression tests, not blind trust in yesterday’s behavior. -
Output validation.
Define how important results will be checked. A fluent answer is not a control. Use deterministic tests, source reconciliation, human approval, or downstream verification appropriate to the consequence. -
Incident and exit path.
Know how to revoke access, disable integrations, rotate credentials, export needed data, and continue the workflow if the vendor has an outage or no longer fits the business. -
Commercial terms.
Record material limits, usage costs, ownership terms, support level, rate limits, and any dependency that could change the economics of the workflow.
Separate model capability from vendor suitability
A model can be excellent at reasoning and still be the wrong deployment choice for a particular workflow. Conversely, a tool with a slightly less capable model may be the better business fit if its access controls, integrations, observability, and operating terms match the task.
Evaluate at least three layers separately:
- Model layer: capability, reliability, context limits, tool use, structured output, and known failure modes.
- Platform layer: permissions, connectors, logging, orchestration, administration, storage, and deployment controls.
- Business layer: cost, support, portability, policy fit, and operational dependency.
That separation makes comparisons more honest.
Permission design is part of procurement
Once an agent can use tools, the question is not merely whether the vendor is trustworthy. The question is what a compromised, confused, or simply incorrect workflow would be allowed to do.
A strong default is read before write, reversible before irreversible, narrow scope before broad scope. Drafting an email is different from sending it. Preparing a theme change is different from publishing it. Reading a ledger is different from moving money.
Least privilege reduces the size of a mistake.
Require evidence for consequential claims
Agents should not declare a task complete because the output looks plausible. Build evidence into the workflow:
- Source citations for factual research.
- Readback after a mutation.
- Tests after a code change.
- Reconciliation after financial analysis.
- Live-URL validation after publication.
- Permission checks before irreversible actions.
The evidence should match the claim.
Plan for vendor change
AI products are moving quickly. A workflow that depends on one undocumented behavior or one proprietary prompt format can become fragile.
Preserve the parts you control: canonical instructions, test cases, source data, schemas, acceptance criteria, and an inventory of integrations. Where practical, keep business logic outside the vendor-specific interface so a tool can be replaced without rebuilding the operating doctrine from memory.
This does not mean avoiding managed AI services. It means preventing unnecessary lock-in.
Write the completion criteria before execution
For every consequential agent workflow, define:
- Outcome: what tangible result must exist?
- Authority: which sources control?
- Boundaries: what may and may not change?
- Evidence: what proves the work is correct?
- Release gate: what must be true before completion?
- Failure path: when must the workflow stop rather than improvise?
Those six lines make the vendor evaluation actionable because they show what the tool must support in the real operating environment.
What not to automate blindly
Agentic capability does not mean every action should become autonomous. Higher-consequence work benefits from stronger gates, especially where actions are difficult to reverse, affect customers, change money or permissions, publish publicly, or depend on incomplete evidence.
Good automation removes unnecessary human effort. It does not remove accountability.
The practical takeaway
AI vendor selection should move beyond feature comparison. Evaluate the full operating surface: data, retention, permissions, security, change risk, validation, exit strategy, commercial terms, and the exact job the system is allowed to perform.
Then give the agent an operating contract that defines authority, boundaries, evidence, completion criteria, and a failure path.
For the data side of the same problem, read The Data Agent Era. For a broader system view, see The Lean AI Stack and the Kairos overview.
Sources
- NIST — Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)
- NIST — AI Risk Management Framework
- OpenAI — Introducing the Agents API
Related resources