Updated September 8, 2026 · By Ryan Bold
An enterprise AI agent combines a model with tools that let it retrieve information or take actions. Its value depends less on how autonomous it sounds than on whether it completes a defined business task accurately, with appropriate permissions and a usable record of what happened.
Choose a workflow before choosing an agent
Anthropic’s engineering guide distinguishes predetermined workflows from systems in which a model dynamically directs its process and tool use. It also recommends starting with simpler approaches where they are sufficient. A fixed approval rule does not need a model to reinvent it.
Consider a support inbox. Software can route a known account identifier to its owner using a rule. A model might help summarize an ambiguous request or draft a response from approved documentation. Sending the response, changing account access and issuing a refund are additional actions with separate permissions.
A bounded invoice example
The following is an original design example, not a claim about a deployed customer system:
- Read an incoming invoice from an approved queue.
- Extract supplier, invoice number, currency and amount into a fixed schema, linking each field to the source.
- Use ordinary code to check required fields and match the supplier against an approved record.
- Flag discrepancies and draft a review note.
- Require an authorized person to approve any payment through the existing payment system.
The agent does not need bank credentials to perform the first four steps. If a PDF tells it to change the supplier’s bank details, that text is invoice content, not an instruction from the business owner.
Make permissions enforceable outside the prompt
OWASP documents prompt injection as a risk when models process external content. Telling an agent to “be careful” is not equivalent to restricting its tools. Use a dedicated identity, limited operations and validation in the system that carries out actions.
| Operation | Example boundary |
|---|---|
| Read | Only the approved queue and records needed for the task |
| Draft | Save proposed changes without applying them |
| Write | Validate fields and authorization in application code |
| Consequential action | Use the existing approval process and record the authorized decision |
Test ordinary mistakes as well as attacks
Use duplicate invoices, missing currencies, unreadable scans, two suppliers with similar names and a temporarily unavailable lookup service. Check that retries do not create duplicate work. A system that stops with a clear unresolved case may be safer and more useful than one that always produces a confident answer.
Record the input reference, extracted fields, tool calls, validation results and final disposition. Avoid retaining unnecessary personal data or secrets in logs. Establish a manual fallback and a way to disable the agent without disabling the underlying business process.
Turn the invoice example into acceptance cases
A draft-only pilot needs an observable result for each input, including cases it cannot complete. The following cases are a proposed test design, not a claim about a deployed system:
| Input | Required outcome |
|---|---|
| The same supplier and invoice number arrive twice | Flag the duplicate; do not create a second payable item. |
| The amount is visible but currency is absent | Leave currency unresolved and request review; do not infer it from an address. |
| The invoice requests new bank details | Keep payment data unchanged and invoke the separate verification process. |
| The supplier lookup times out | Record the incomplete check and retry without duplicating the draft. |
| The PDF contains instructions addressed to the model | Treat them as document content, never as authorization. |
For every extracted field, retain a page or text reference sufficient for a reviewer to find the source. For money, use validated decimal handling and an explicit currency rather than allowing a free-text model answer to become the accounting value.
The duplicate check must also cover retries and concurrent workers. A prompt asking the model not to duplicate work is not an idempotency mechanism. The application storing the draft should enforce the chosen identity rule and preserve a clear unresolved state when required fields are missing.
Measure the complete job
Track accepted cases, corrections, escalations, duplicate actions, processing time and review effort. Include integration and operating costs. The NIST AI Risk Management Framework provides a broader structure for identifying and managing AI risks, but applying a framework is not a certification that an implementation is safe.
Expand access only after the pilot shows where the tool succeeds and fails. A reliable draft-and-review process can be a valuable outcome without turning the agent into the final decision-maker.
For a focused example of checking source-based output, try our Gemini Notebook source-conflict exercise. Two fictional documents show how to distinguish an explicit correction from an unsupported assumption.