A weighted scorecard should expose judgment, not hide it behind decimals. Start by removing options that fail mandatory gates. Then set a small number of criteria and weights before demonstrations, score only observed evidence, and test whether reasonable changes in disputed weights reverse the result.
This guide includes a worked three-option calculation. The numbers are fictional and illustrate the method; they are not ratings of any product.
The baseline tie and changed ranking show why a scorecard must report uncertainty, not manufacture a winner.
Gate first, score second
Define conditions that no weighted strength can compensate for: a critical workflow, mandatory security or privacy control, required deployment boundary, data portability need or affordability ceiling. Mark each option pass, fail or unresolved. A failed option leaves the comparison unless an authorised decision maker records an explicit exception.
This distinction is not unusual in formal evaluation. Public Acquisition.gov source-selection guidance separates stated weighted factors from factors such as go/no-go acceptability. BBS uses the principle as a practical decision control, not as legal procurement advice.
Choose criteria that represent different decisions
Use five to eight criteria that trace to the outcome, user needs and major risks. Avoid scoring the same idea several times—for example usability, ease of use and intuitive interface—because duplication gives it hidden weight. A criterion should have one owner, one definition and evidence that can change the score.
In the worked example, all three options pass the mandatory gate. The team compares four distinct criteria: workflow evidence, implementation effort, operating fit and commercial resilience. The weights total 100%: 40, 25, 20 and 15.
Anchor the evidence scale
Define the scale before evaluating products:
- 0: no evidence or unacceptable;
- 1: material gap, workaround or uncommitted future dependency;
- 2: acceptable with known effort or limitation;
- 3: strong fit demonstrated with relevant evidence.
Every score links to a demonstration note, test, document or contractual response. Evaluators record unknown rather than guessing. They score independently first, then resolve differences by reviewing evidence—not by averaging incompatible interpretations.
Calculate the baseline result
For each option, multiply each evidence score by its decimal weight and add the results:
Weighted total = Σ (criterion weight × evidence score).
| Criterion | Weight | A | B | C |
|---|---|---|---|---|
| Workflow evidence | 0.40 | 3 | 2 | 2 |
| Implementation effort | 0.25 | 1 | 3 | 2 |
| Operating fit | 0.20 | 2 | 2 | 3 |
| Commercial resilience | 0.15 | 2 | 2 | 3 |
| Total | 1.00 | 2.15 | 2.35 | 2.35 |
The baseline does not identify one winner. B and C tie, while A is 0.20 behind. That is a valid result: the current evidence and preferences do not justify a unique choice.
Run a sensitivity test on the disputed judgment
The team believes workflow fit may deserve more weight and implementation effort less. It changes only those two assumptions: workflow rises from 40% to 50%; implementation falls from 25% to 15%. Other weights remain 20% and 15%.
| Scenario | A | B | C |
|---|---|---|---|
| Baseline 40/25/20/15 | 2.15 | 2.35 | 2.35 |
| Sensitivity 50/15/20/15 | 2.35 | 2.25 | 2.35 |
The ranking changes. A joins C at the top and B falls behind. The decision is sensitive to a debatable preference, so the sheet should not claim a robust winner. The team can obtain better workflow or implementation evidence, negotiate a condition, or ask the accountable decision maker to choose the trade-off explicitly.
Test score uncertainty as well as weight uncertainty
A demonstration score may also be uncertain. If Option C's operating-fit score of 3 depends on a roadmap item, recalculate it as 1 or mark the criterion unresolved. Report a range where evidence cannot yet support one point. A product that wins only when every unknown receives the optimistic value is not the evidence leader.
Keep price or total cost visible using the same horizon and assumptions. Depending on the decision method, cost may be a separate affordability gate, a transparent factor or part of a value assessment. Do not hide a different cost basis inside a technical score.
Prevent six common scoring failures
- Weights after demos: lock them before preferences form.
- Duplicate criteria: map each criterion to one decision.
- Vague scales: anchor each number to evidence quality and fit.
- False precision: do not treat a 2.31 as meaningfully exact.
- Unknown equals zero—or three: keep uncertainty explicit.
- No reversal test: always challenge the most disputed weights and scores.
The decision record beside the spreadsheet
The scorecard needs a short narrative: mandatory-gate results; criteria and why they matter; evidence behind decisive scores; sensitivity scenarios; unresolved risks; total-cost assumptions; and the trade-off accepted by the owner. Acquisition.gov's source-selection planning guidance similarly connects evaluation factors to user requirements, objectives, market research and risk analysis.
A scorecard is successful when another person can reconstruct the choice and see what new evidence might change it. If the spreadsheet merely produces the answer the team already wanted, remove the decimals and return to the evidence.
Calibrate evaluators before the real demonstrations
Give all evaluators the same short practice scenario and ask them to score it independently. Compare both the number and the evidence note. A disagreement between 1 and 3 usually means the scale, criterion or expected proof is unclear. Resolve that ambiguity before product names and preferences enter the room.
During live evaluation, keep individual scores private until everyone has submitted them. The group then reviews large differences criterion by criterion. The goal is not to negotiate toward an average; it is to determine whether people saw different evidence, applied different assumptions or interpreted the scale differently. Change a score only with a recorded reason.
Set decision thresholds before reading the totals
Define what the result can authorize. For example, a product may need to pass every mandatory gate, score at least 2 on each high-risk criterion, stay within the approved three-year cost range and lead the runner-up by more than the score uncertainty. If no option meets those conditions, the correct outcome is not necessarily to select the highest total. The team may need to clarify requirements, obtain stronger evidence, negotiate safeguards, expand the market scan or stop the purchase.
Also decide who owns a close call. The accountable decision maker should receive the baseline, sensitivity results, unresolved risks and recommendation—not just a colour-coded ranking. A tie can be a useful finding because it reveals where commercial terms, implementation capacity or further testing should decide the next step.
Keep an auditable evidence ledger
Store one row per score with the evaluator, date, evidence link, assumption, confidence and any condition attached to the score. Lock the calculation cells and version changes to criteria or weights. When new evidence arrives, update the ledger and rerun the same scenarios rather than editing the final total directly.
This makes the scorecard useful after contract signature. Implementation teams can see which capabilities were demonstrated, which relied on configuration or roadmap commitments, and which risks were accepted. The buying decision becomes a set of testable promises instead of a spreadsheet that loses meaning once the preferred vendor is chosen.