Articles

Why refining an operational due diligence rubric makes it weaker

Averaging cannot express “this issue is disqualifying”. Adding domains halves the weight of the one that could kill the fund, and re-weighting changes nothing at all.

Twelve managers are scored across eight operational risk domains and the scores are averaged. Three of the twelve carry a flaw that is terminal on its own — no independent custody, an unmanageable conflict, self-marked valuations with no oversight. The averaging model gives all three its top rating.

Worse than that. The manager with no independent custody of client assets scores 1.500. A manager with no fatal flaw at all scores 2.250. The model does not merely fail to catch the disqualifying issue — it ranks it 0.75 of a point higher.

Every operational due diligence framework carries a warning against averaging your way to comfort. The warning is correct and almost always unquantified, which is a strange thing for a warning about arithmetic. Here is what it is worth in decisions.

The model

Risk domainCan it be terminal on its own?
Governance
Structure & terms
Finance & controls
Admin, custody, cashyes
Service providers
Compliance & conflictsyes
Technology & resilience
Valuationyes
Eight domains, scored 1 (low concern) to 5 (high). Approve at 2.0 or below, approve with conditions at 3.0, not yet at 3.5, reject above.

Three of the eight domains can be terminal on their own. That is not a controversial claim — most frameworks name the same three. The question is whether the scoring model does anything with the fact, and an average does not: it has no way to express “this one is disqualifying regardless of the rest”.

Twelve managers, two decision rules

ManagerWorst veto domainAverageDecision on the averageDecision with veto rules
Alderman Credit51.500ApproveReject (veto)
Brightmoor Partners11.250ApproveApprove
Caldwell Special Sits22.250Approve with conditionsApprove with conditions
Draycott Infrastructure51.625ApproveReject (veto)
Ellerby Growth33.000Approve with conditionsApprove with conditions
Fenwick Real Assets51.500ApproveReject (veto)
Garrow Mid-Market32.250Approve with conditionsApprove with conditions
Halstead Ventures33.500Not yetNot yet
Inverleith Secondaries11.000ApproveApprove
Jarrow Opportunistic43.625RejectNot yet (veto)
Kelsall Private Credit22.125Approve with conditionsApprove with conditions
Larkhall Buyout42.250Approve with conditionsNot yet (veto)
Twelve managers, one scoring sheet, two decision rules.

Five decisions out of twelve change when the same scores are read with veto rules instead of an average. Three of those five are managers the average approves and the veto rules reject outright. Nothing about the underlying diligence changed; only the arithmetic did.

The comparison that should end the argument

Alderman CreditCaldwell Special Sits
Admin, custody, cash52
Every other domain12 or 3
Average1.5002.250
Decision on the averageApproveApprove
Which one has a terminal flaw?this oneneither
One has no independent custody of client assets. It scores three quarters of a point better.

How good does the rest have to be?

Solve it directly. With n domains, one scored 5 and an approval threshold of 2.0, the others may average up to (2n − 5)/(n − 1) and the fatal flaw still disappears.

Domains in the rubricAverage the others may carryWeight of the fatal domain
51.25020.0%
81.57112.5%
101.66710.0%
121.7278.3%
161.8006.2%
201.8425.0%
A single domain scored 5, and the overall score still lands on Approve.

Read that table twice. At eight domains the other seven need to average 1.571 — not perfection, just good. At sixteen domains they need only 1.800. Refining the rubric made it easier to hide a fatal flaw, by 0.229 of a point, because splitting eight domains into sixteen halves the weight of the one domain that could kill the fund.

Nobody running that exercise would describe it that way. A team that adds granularity is trying to capture nuance and believes it is tightening the model. It is doing the opposite, and the direction of the error is the dangerous one: the more careful the framework looks, the more comfortably a disqualifying issue rides through it.

And the debate committees actually have does not change anything

Weightings are where the argument goes. Should custody carry three times governance? Should valuation be doubled? Two plausible schemes, applied to the same twelve managers.

Change made to the modelDecisions it changes, out of 12
Re-weight: controls-heavy0
Re-weight: governance-heavy0
Replace averaging with veto rules5
Zero from re-weighting. Five from restructuring.

Zero decisions out of twelve change under either re-weighting. Five change when averaging is replaced by veto rules. Committee time spent debating weights is time that cannot change an outcome — while the design decision that determines every outcome, whether the model may say “this issue is disqualifying”, is usually taken in one sentence and never revisited.

How close is each answer to a different one?

One more diagnostic worth running on your own rubric. Measure each manager’s distance to the nearest threshold in analyst-points — one notch of disagreement on one domain. It tells you which conclusions survive a reasonable difference of opinion between two people who did the same work.

ManagerAverageNearest thresholdAnalyst-points to a different answer
Alderman Credit1.5002.04
Brightmoor Partners1.2502.06
Caldwell Special Sits2.2502.02
Draycott Infrastructure1.6252.03
Ellerby Growth3.0003.01
Fenwick Real Assets1.5002.04
Garrow Mid-Market2.2502.02
Halstead Ventures3.5003.51
One analyst-point is one notch of disagreement on one domain.

A rating one analyst-point from a different answer is not a rating. It is a coin flip with a decimal place, and it should be reported as such to whoever relies on it.

What to change on Monday

Three things, none of which requires rebuilding anything. First, name the domains that are terminal on their own and give the model a rule that says so — a veto, not a weight. Second, before adding domains to a rubric, compute what the addition does to the weight of the terminal ones; if the answer is that it dilutes them, add the veto instead of the nuance. Third, publish the analyst-point distance next to every rating, so a committee can see which of its conclusions are robust and which are arithmetic.

The warning against averaging your way to comfort is right. It just needs a number attached before anyone acts on it, and the number is five decisions in twelve.

The workbook behind this article

Every figure above is a live formula in the companion file for Operational Due Diligence in Private Equity, which holds the full twelve-manager sheet, the granularity table, the three weighting schemes and the fragility measure. It is free, and it needs no account and no email address.

Open the companion file →

Also on this site

This note is drawn from Operational Due Diligence in Private Equity. The book is on Amazon.

If this book helped — or didn’t — a few lines on Amazon are worth more than they look: they are what the next reader goes on. Write a review. The workbook stays free either way.