
Two analysts score the same vendor questionnaire and land 15 points apart. Which one is right? If your methodology cannot answer that question, you do not have a scoring model; you have opinions with decimals attached. A defensible model has five components: domain weighting, evidence pricing, partial credit rules, decision thresholds, and a governed override path. Here is how to build each one.
Principle: score the risk, not the questionnaire
The purpose of a score is to drive one of three decisions: approve, approve with conditions, or reject. Every design choice below should be tested against a single question: does this make the resulting decision more defensible to the person who will challenge it later, whether that is an auditor, an examiner, or your own leadership after an incident?
Step 1: weight by control domain, not by question count
The naive model scores every question equally, which means a vendor can fail encryption and access control yet score well by acing forty easy governance questions. Weight at the domain level first, so the shape of the instrument does not determine the shape of the risk. An illustrative weighting for a vendor handling sensitive data:
| Control domain | Weight | Why |
|---|---|---|
| Access control and identity | 20% | Most common initial failure point in third-party incidents |
| Data protection and encryption | 20% | Directly governs the harm if compromise occurs |
| Vulnerability and patch management | 15% | Determines exposure window to known flaws |
| Incident detection and response | 15% | Determines blast radius and your notification timeline |
| Business continuity and resilience | 10% | Availability harm is real harm |
| Governance, policy, and personnel | 10% | Leading indicator, weak direct control |
| Subprocessor and supply chain management | 10% | Your fourth-party exposure lives here |
Two rules make weighting defensible. First, weights are set per vendor category, not per vendor: a data processor and an on-site services firm get different weight profiles, but two similar data processors get the same one, or your comparisons are meaningless. Second, document the rationale column. "Why is encryption 20 percent?" is the first question a skeptical reviewer asks, and the answer should already be written down.
Step 2: price evidence above attestation
A "yes" with a screenshot, a policy excerpt, or an independent audit artifact is not the same answer as a bare "yes." A scoring model that treats them identically teaches vendors that evidence is optional. Apply an evidence multiplier to each answer's base score:
- Independently verified (1.0): supported by an independent audit report, certification with relevant scope, or your own testing.
- Evidence-backed (0.9): vendor-supplied artifact consistent with the claim: configuration export, policy document, dated screenshot.
- Self-attested (0.7): a bare assertion with nothing behind it.
- Attested but contradicted (0.0 plus a flag): the claim conflicts with observable reality, such as external monitoring showing the opposite. This is never a scoring discount; it is a finding, because it undermines the credibility of the whole response. This is the questionnaire-level version of the assessed versus live posture problem.
The multiplier values are illustrative; the structural point is that the gap between attested and verified must be priced, visible, and consistent.
Step 3: partial credit rules, written down
Real answers are rarely clean yes/no. "MFA is enforced for administrators, rollout to all staff completes next quarter" deserves something between full credit and zero, and the something must be a rule, not a mood. A workable four-level rubric per question:
- Full credit (100%): control fully implemented across the relevant scope.
- Substantial (67%): implemented for the majority of scope, including all critical assets, with a dated plan for the remainder.
- Partial (33%): implemented in a limited scope, or compensating controls only.
- None (0%): not implemented, no credible plan, or answer not provided. Unanswered questions score zero, never "excluded," because excluding them silently inflates the score of the least cooperative vendors.
Add one non-negotiable overlay: critical control gates. Designate a short list of controls (encryption of your data class at rest and in transit, MFA on administrative access, breach notification capability) where a "None" answer caps the total score below the approval threshold regardless of arithmetic. Weighted averages exist to be gamed by strength elsewhere; gates exist to stop that.
Step 4: thresholds that map to decisions
A score without a decision boundary is trivia. Set explicit bands, and scale them by the vendor's tiering framework tier, because the same score should not buy the same outcome at every criticality level. An illustrative matrix on a 0 to 100 scale:
| Decision | Tier 1 vendors | Tier 2 | Tier 3 |
|---|---|---|---|
| Approve | 85+ | 75+ | 65+ |
| Conditional approval | 70 to 84 | 60 to 74 | 50 to 64 |
| Reject / do not proceed | Below 70, or any gate failure | Below 60, or any gate failure | Below 50 |
Conditional approval is the band that does real work, so define its mechanics precisely: every condition gets a named owner on the vendor side, a remediation deadline (30/60/90 days by severity is a sane default), and an automatic consequence for missing it, which is escalation to reject or to formal risk acceptance by someone senior enough to own the residual risk. A conditional approval without deadlines is an approval wearing a disclaimer.
Step 5: reviewer override, governed not forbidden
No model captures everything. A vendor can score 88 while their answers reveal an architecture that concentrates exactly the risk you care about; a vendor can score 68 with gaps that are irrelevant to your actual use of them. Reviewers must be able to override the computed outcome, and the override must be governed:
- Reviewer judgment always outranks the computed score, in both directions. A model that cannot be overridden becomes the thing everyone works around.
- Every override carries a written rationale recorded with the assessment. "Reviewer discretion" is not a rationale.
- Overrides are logged and periodically analyzed. If one domain is overridden in a fifth of assessments, the model is mispricing that domain; fix the weights instead of accumulating exceptions.
- Upgrades demand more than downgrades. Overriding toward approval on a gate failure should require sign-off one level up, because that is a risk acceptance, not a scoring correction.
This human-final principle extends to any automation you apply. AI-assisted scoring can meaningfully reduce assessment burnout by drafting scores and flagging weak evidence, and it belongs under exactly the same rule; ThirdSentry applies it natively, with AI-generated scores always subordinate to the reviewer's recorded decision.
Make it auditable end to end
The final property of a defensible model is reconstructability. For any historical score, you should be able to produce: the instrument and version used, the domain weights in force, each answer with its evidence tier and partial credit level, the computed score, any override with rationale and approver, and the resulting decision with conditions and their outcomes. If any link in that chain lives in a spreadsheet formula someone edited since, the chain is broken.
ThirdSentry runs this methodology as workflow rather than spreadsheet: weighted domain scoring across SIG, CAIQ, and custom instruments, evidence attached at the answer level, and reviewer decisions that override and outlive the computed score, with the full trail retained for the audit that will eventually ask.

