A vendor-led product tour answers the question the seller prepared, not necessarily the question the buyer needs to decide. A useful software demonstration is a controlled test: every viable option receives the same scenario, roles, data, time boundary and evidence rules.
The purpose is not to see every feature. It is to verify a small number of important workflows, expose exceptions and identify what still requires proof. The runbook below keeps the session disciplined without preventing a supplier from showing a better method.
Choose decisions the demo must inform
Start with three to five requirements that could change the shortlist, implementation approach or risk decision. For each, write the business scenario, actors, starting data, expected outcome, material exception and evidence status that would count as verified. Avoid feature names unless they are the actual constraint.
For example, replace "show approval workflow" with: "A requester submits an order above a delegated limit; the first approver is unavailable; the service routes to the authorised substitute, preserves the reason and produces an audit record." This lets different products demonstrate different designs while answering the same need.
Send a pre-demo pack
Provide the script several working days in advance. Include roles, anonymised or synthetic data, integration assumptions, permitted configuration, time allocation and rules for follow-up. Ask the supplier to identify anything that will be simulated, unavailable in the proposed edition or dependent on third-party work.
State the evaluation boundary: the demonstrated environment and version, configuration performed before the session, included and excluded services, and whether the people presenting would support implementation. A polished prototype can be relevant evidence, but it is not the same as a generally available function in the proposed service.
Assign roles inside the buyer team
| Role | Responsibility during the session |
|---|---|
| Facilitator | controls sequence, time and interruption rules |
| Scenario owner | confirms whether the work and exception are represented correctly |
| Operator or administrator | asks how rules, access, monitoring and change are managed |
| Evidence recorder | captures observations, status, conditions and follow-up owner |
| Evaluators | score independently after the evidence segment, not during sales discussion |
Do not make one person facilitate, take notes and decide. The cognitive load produces missing evidence and inconsistent treatment.
Use a fixed 75-minute run order
- 0–10 minutes — prove the boundary: name the product, edition, environment, roles, data and preconfiguration. Record any simulation.
- 10–30 — normal workflow: the supplier completes the buyer's scenario without substituting a prepared story.
- 30–45 — exception: introduce the unavailable approver, invalid data, integration delay or correction. Observe recovery and audit behavior.
- 45–60 — operation: change a rule, inspect access, monitoring, logs, configuration history or report lineage.
- 60–75 — unresolved evidence: ask targeted questions and assign dated follow-ups.
Adapt the duration, but preserve the order. Normal-path confidence before exception testing often produces false assurance.
Define interruption rules
The facilitator may pause when the demonstration leaves the scenario, uses different data, hides a prerequisite or makes a claim that needs an evidence note. Use neutral prompts: "Which step in our scenario does this answer?"; "Is that included in the proposed edition?"; "Can you show the resulting record?"; "What happens when that dependency is unavailable?"
Do not debate product design during the run. Record a gap or alternative, then continue. Give each supplier equal time and equivalent opportunities to clarify.
Capture observations before scores
For each requirement, write what happened, who performed it, the resulting record and any condition. Use four statuses:
- Verified: the relevant behavior and output were observed in the defined boundary.
- Conditional: it worked with a named configuration, service, edition, dependency or future proof.
- Gap: the scenario or control was not supported within the proposed boundary.
- Not tested: the session did not produce enough evidence.
Not tested is not zero and not verified. It creates a decision about follow-up, proof or risk. NASA's systems-engineering material distinguishes verification from validation; the same discipline helps buyers ask both whether the product behaved as specified and whether that behavior supports the intended use.
Force an exception and a correction
Happy paths hide the cost of operating software. Introduce one predictable exception: a duplicate record, rejected request, changed approver, late data, unavailable integration or user mistake. Then ask the presenter to correct it without deleting the history.
Observe who can act, which permissions are required, whether downstream data changes, how users are notified and what the audit trail retains. A workaround may be acceptable, but its effort and ownership must be visible.
Test administration as part of product fit
Ask who performs common changes after launch: adding a field, updating a rule, changing access, investigating a failed integration, restoring data or producing evidence for an audit. Have the supplier demonstrate one representative change and its approval or history.
This prevents selection based only on the end-user surface. A fast workflow that requires specialist intervention for every change can create an operating bottleneck.
Keep questions tied to missing proof
Do not spend the final segment on a generic question list. Review the evidence sheet and ask only what would change a status or decision. Convert every deferred answer into a deliverable: exact question, acceptable evidence, named owner and deadline.
Acquisition.gov's source-selection framework emphasises evaluating proposals against stated factors and documenting the rationale. A controlled demo applies the same principle to observed product behavior.
Score independently, then reconcile evidence
Evaluators should submit their status and score before group discussion. Investigate large differences by comparing evidence notes. Do not average a verified observation and an untested assumption. Resolve the boundary or retain uncertainty.
After all demonstrations, compare by requirement rather than replaying each vendor's presentation. The decision record should distinguish observed capability, conditional capability, gaps and unanswered evidence. A successful demo makes the next decision clearer; it does not merely make the product look impressive.
Assemble the evidence pack the same day
Store the completed scenario sheet, environment and version, screenshots or recordings permitted by the session, evaluator notes, conditions, scores and follow-up requests under one identifier. Record who attended and which supplier statements were later corrected. Freeze the initial observations before new documents arrive.
Then issue a concise factual clarification: missing evidence, acceptable form, owner and deadline. Keep late evidence visibly separate from what was demonstrated. This preserves a fair comparison and lets the implementation team trace every accepted condition back to the test that created it.