An AI agent RFP should make every supplier price and prove the same workflow. State the trigger, current baseline, authorised sources, decisions the agent may make, actions it may take, human approvals, accepted final state and first three months of operating cost. Then require evidence against representative cases rather than a generic capability presentation.

Procurement can now compare like with like. The response also shows where a polished demo depends on clean sample data, manual work, broad permissions or implementation tasks left outside the quote.

Start with a case suppliers cannot redefine

“Deploy an enterprise AI agent” leaves the supplier free to choose the easiest interpretation. One proposal may cover a conversational interface over documents. Another may include identity, source preparation, system actions, evaluation, monitoring and user rollout. Their prices describe different products.

Anchor the RFP to one recurring case:

  • who starts the work and under what condition;
  • the inputs and source systems used today;
  • current volume, handling time and failure cost;
  • where judgement enters the process;
  • the record, artefact or action that completes the case;
  • who accepts that result;
  • the baseline the implementation must improve.

A Product × GTM planning case, for example, may collect product evidence from Slack and SharePoint, reconcile claims, draft a planning artefact and route unresolved points to named owners. The accepted outcome is a reviewed artefact with usable evidence, rather than an answer that merely sounds plausible.

The AI agents versus workflow automation guide helps split deterministic transport from model-led judgement before suppliers add expensive autonomy to stable steps.

Define evidence and authority before the interface

Internal agents produce answers from company material that varies in freshness, ownership and sensitivity. The RFP needs a source schedule that names each system, the data required, the owner and the access rule.

Ask suppliers to explain:

  1. how source-system permissions survive retrieval and output;
  2. what happens when two authorised sources disagree;
  3. how stale, deleted or revoked material stops influencing later work;
  4. which evidence appears beside an important answer;
  5. where corrections go and who can approve them;
  6. how the implementation behaves when a source or permission service is unavailable.

Resolve this before integration. Otherwise, the buyer learns too late that “connected” means searchable, while operators still reconcile conflicting policies and customer records by hand.

Microsoft’s agent architecture checklist asks teams to define acquisition and governance, processing and data flow, input and output specifications, control, accountability and return on investment before detailed design. Those categories belong in the buying brief because each one changes scope.

Put action boundaries into the requirements

Retrieval carries one risk profile. An agent that updates CRM stages, sends external messages or changes a commercial record carries another.

List each tool and action separately. For every action, state:

  • eligible users and cases;
  • read or write scope;
  • required evidence;
  • approval threshold;
  • retry behaviour;
  • duplicate-action protection;
  • audit record;
  • stop, rollback or reconciliation route;
  • person accountable when recovery fails.

Require the supplier to map the proposed identity model across the user, agent, connected tool and destination system. “Role-based access” is too broad to show whether the agent acts as itself, on behalf of a user or through a shared service identity.

A live vendor demo should exercise one consequential action and one partial failure from this list. Procurement then sees the operating path behind the presentation.

Turn acceptance into a contractual evidence pack

Accuracy without a case definition gives the buyer little control. Agent output varies across runs, and a correct sentence can still accompany the wrong tool call or destination state.

Attach a representative case set to the RFP. Include normal work, missing information, conflicting sources, restricted information, tool failure, an approval rejection and an attempted duplicate action. Define the graders and release threshold before the supplier tunes the system.

Anthropic’s agent evaluation guidance separates tasks, repeated trials, graders, traces and final outcomes. That distinction matters in a contract. A transcript can show reasonable behaviour while the destination system contains the wrong state. Acceptance should verify both the path and the result.

Ask each supplier to return:

  • baseline results on the buyer’s cases;
  • the number of trials per case;
  • deterministic, model-based and human grading used;
  • hard-stop failures that block release;
  • accepted outcome rate and material revision rate;
  • latency and full cost per accepted case;
  • regression method for model, prompt, tool and source changes;
  • evidence retained for audit and dispute resolution.

The existing pilot success criteria can turn those results into a scale, revise or stop decision after the first bounded deployment.

Make suppliers expose the whole operating boundary

A credible response separates what exists today from what will be built. Ask suppliers to label every requirement as standard, configured, custom, third-party, buyer-provided or excluded.

Use this response matrix:

RequirementBuyer constraintSupplier responseEvidence available nowWork before acceptanceOwner after launchOne-time costMonthly or usage cost
Workflow and accepted stateNamed case and baseline
Sources and permissionsSystems, roles, conflict rule
Models and orchestrationApproved providers and change controls
Tools and actionsRead/write scope and approvals
EvaluationCases, trials, graders and thresholds
ObservabilityTraces, alerts and retention
RecoveryStop, rollback, reconciliation and owner
AdoptionTraining, support and correction route
Handover and exitArtefacts, access and export

The “evidence available now” column prevents roadmap claims from receiving the same score as demonstrated capability. “Owner after launch” exposes permanent buyer workload before the contract converts it into an operational surprise.

The AI agent implementation cost guide provides the companion budget worksheet across build, company context, controls, adoption and recurring operation.

Price change, dependence and exit

The first release will change. Sources gain new fields, policies move, models update and workflow owners tighten acceptance after seeing real cases.

Require rates and response terms for source changes, new test cases, model migration, prompt or policy updates, incidents and material workflow expansion. Set usage caps and state who approves spend above the agreed range.

The handover schedule should identify configuration, prompts, source maps, permission rules, test cases, traces, runbooks, known limitations and deployment records the buyer receives. Ask how data, evaluations and operating artefacts can be exported at exit, plus the work required to remove supplier access and verify deletion under the contract.

The resourcing choice becomes concrete in that handover schedule. The AI consultancy versus in-house guide defines the decision rights the company should retain and the live recovery test that proves handover.

Add the UAE procurement questions where they apply

The UAE Ministry of Finance’s guidance workshop on AI procurement gives buyers a useful discipline: define the problem clearly, test whether AI is the right technology, involve final users, use proof-of-concept stages, assess data-protection and AI-ethics obligations, and audit implementation against the purchase contract.

For a Dubai or UAE buyer, translate applicable policy and legal requirements into named acceptance evidence with qualified counsel and internal owners. A broad promise to “comply with UAE regulation” cannot tell procurement which data moves, where it is processed, who can retrieve it or how the organisation will verify the control.

The same guidance leaves room for a non-AI solution when it addresses the problem more efficiently. That commercial test belongs near the front of the tender. Suppliers should identify deterministic steps and existing platform capabilities before pricing an agent around the entire process.

Score the proposal against the consequence

Weight the matrix according to the workflow. A drafting assistant can place more weight on output quality and adoption. An agent that changes external or financial state needs higher weights for identity, approvals, duplicate protection, traceability and recovery.

Keep commercial scoring connected to accepted outcomes. Compare the three-month total across implementation, licences, model and tool usage, buyer labour, review, rework and support. A lower platform fee loses its advantage when missing controls create custom work or every result needs specialist reconstruction.

Model Operator uses this procurement logic in bounded AI implementation: one recurring workflow, named sources, explicit decision rights and evidence-backed acceptance. For AI-active companies assessing an internal agent, the first useful conversation is the case suppliers will otherwise define for you. Bring Model Operator that workflow and the proposal can be structured around its evidence, authority and operating consequence.