Every question here is agent-specific: memory isolation between tenants, prompt injection defenses, model deprecation history, what happens to a half-finished task when a rate limit hits. It assumes the vendors across the table are running agents in production, not chatbots with a better UI. Score each answer 1 to 5 against the rubric on the last page.
Companion: filled sample. A complete response from a fictional vendor (‘AgentFlow AI’) including a candid ‘Known Limitations & Gaps’ section.