Inspecting Every Returned Device Does Not Mean Zero Escapes: Sample Size, Detection Probability and the Two Cost Curves

Published 2026-09-22 · LuckyMDM Blog

The one-line version

A sampling rate does not decide how many defects you miss; it decides the probability that a lot with real problems produces a sample with none in it. That probability is non-linear. At a 3 per cent lot defect rate, 60 units out of 300 gives an 83.9 per cent chance of catching at least one; getting to 95 per cent requires 100 units, a 67 per cent cost increase for 11 percentage points. The fix is to split intake items into critical and cosmetic tiers, inspect critical items at 100 per cent with machine-checkable criteria, and run cosmetic grading under an attributes sampling standard whose switching rules tighten automatically when quality slips.

Why 100 per cent inspection is not the safe answer

Human grading consistency degrades with volume

The default in most return operations is to look at every unit. The hidden premise behind that choice is that looking at everything means missing nothing. The premise is false.

Intake work splits into two categories. The first is determination: does the unit re-enrol after a wipe and re-activation, is a management profile still present, is Activation Lock still on, does the serial number on the device match the record. These questions have one right answer, and most of them can be answered by a machine or by a scan. The second is grading: how long is that scratch, does that scuff count, which battery health band does this fall into. These depend on a person, and a person's consistency degrades across a run. Unit three and unit forty-three of the same batch do not reliably receive the same grade.

So what 100 per cent inspection actually buys is certainty on the first category. It drags the second category along at the same cost, and the second category is where the cost is high and the payoff is low.

Critical and cosmetic are different problems

Sorting intake items by the consequence of missing them is more useful than sorting them by difficulty:

TierTypical itemsConsequence of a missNature of the testRecommended intensity
CriticalMDM enrolment not released, Activation Lock on, serial mismatch, undocumented board repairUnit cannot be re-rented, or is re-locked by the previous organisation after re-deploymentMachine-checkable, single right answer100 per cent
MajorNon-genuine display assembly, battery maximum capacity under 80 per cent, liquid contact indicator trippedDirect deduction, roughly USD 60 to USD 150 per unitSemi-automated, requires reading a field100 per cent or high-rate sampling
MinorScratch length, bezel marks, housing gap, missing accessoriesOne grade band, roughly 5 per cent of residual valuePurely human, depends on a written standardSample under a standard plan

Running all three tiers at 100 per cent pays the cost of the strictest tier and collects the benefit of the loosest one.

A sampling rate sets detection probability, not the miss count

The arithmetic you can run yourself

Let p be the share of units in a lot that genuinely need a deduction or repair. This must be your own measured figure from the last three months of intake records, not an industry number. If you inspect n units, the probability of catching at least one problem unit is:

detection probability = 1 − (1 − p) to the power of n

At p = 3 per cent:

Sample size nDetection probabilityCost change versus n = 60
10 units26.3%−83%
20 units45.6%−67%
30 units59.9%−50%
50 units78.2%−17%
60 units83.9%baseline
80 units91.3%+33%
100 units95.2%+67%
150 units98.9%+150%

The shape of this table is the part worth remembering. Going from 60 to 100 units costs 67 per cent more and buys 11.3 percentage points. Marginal return collapses quickly. That is why raising the sampling rate is usually the most expensive available improvement. The cheap improvements are reducing p itself through tighter sourcing and pre-screening, and raising per-unit accuracy through quantified grading rubrics, dual review and photo evidence. Neither one adds sample size.

Where the intuition goes wrong

The common intuition is that sampling 20 per cent finds roughly 20 per cent of the problems. The opposite is closer to the truth. Sampling 20 per cent means there is an 83.9 per cent chance that problems in this lot surface during intake at all, which is what triggers tightening or a full re-inspection. Sampling is a trigger, not a filter. Using it to pick problems out one at a time is the wrong mental model, and it is the reason operators keep concluding that sampling "did not work."

How to look up a sample size: three parameters

Inspection level, AQL and the code letter

The standard normally used for this in the United States is ANSI/ASQ Z1.4-2003 (R2018), Sampling Procedures and Tables for Inspection by Attributes, the national adoption of ISO 2859-1. It descends from MIL-STD-105E, which is why many operators still say "105 sampling." Rather than a guessed percentage, it derives sample size from three parameters:

ParameterQuestion it answersCommon settings
Inspection levelHow tightly sample size tracks lot sizeGeneral Inspection Levels I, II and III; Level II is the default. Special levels S-1 to S-4 exist for destructive or very costly tests
Acceptable Quality Limit (AQL)What quality counts as acceptableA preferred value series from 0.010 to 1,000. Critical defects take the tightest value; minor defects can be looser
Sampling plan typeOne draw or staged drawsSingle, double or multiple. Double and multiple plans have a lower average sample size but higher administrative cost

The lookup order is: lot size plus inspection level gives a code letter; code letter plus AQL gives sample size and the accept and reject numbers. The accept and reject numbers must be read from the tables in the standard. Do not fill them in from memory. This page gives the framework, not a substitute for the document.

Switching rules turn the rate into a variable

The most valuable part of Z1.4 is not the code letter table. It is the switching rules, which means sampling intensity is not a fixed policy but a state machine:

TransitionTrigger condition
Normal to tightened2 out of 5 consecutive lots or batches fail original inspection
Tightened to normal5 consecutive lots pass original inspection while on tightened
Normal to reducedSwitching score reaches 30, production is steady, and the responsible authority approves
Reduced to normalOne lot fails, production becomes irregular, or the responsible authority requires it
Continued failure on tightenedDiscontinue inspection; restart at tightened after corrective action

This answers the question of whether sampling is laziness. Sampling with switching rules is not relaxed: the moment the trigger fires, the next lot tightens or converts to full inspection. Sampling without switching rules is the lazy version.

Defect tiers drive three different AQLs

One plan should not cover three tiers of consequence:

Defect tierAQL directionRejection logic
CriticalTightest; in practice often treated as zero acceptanceOne occurrence rejects the lot and triggers full re-inspection
MajorTightExceeding the accept number rejects the lot
MinorCan be relaxedRecorded for grading and residual-value correction

Once the tiers are separated, the question "should we inspect everything" no longer has a single answer. Critical: yes. Minor: no.

Where the two cost curves cross

A worked example you can recompute

Take an operation receiving 300 returns a month. Inspection takes 6 minutes per unit. Loaded labour is USD 30 per hour, so USD 3.00 per unit. Lot defect rate p is measured at 3 per cent. A one-band grading error costs about 5 per cent of a USD 420 residual value, so about USD 21. Fatigue miss rate under continuous full inspection is taken as 2 per cent; under a small sample with dual review, 0.5 per cent.

PlanInspection cost per monthExpected missesMiss cost per monthTotal per month
100 per cent, 300 unitsUSD 900300 × 3% × 2% = 0.18 unitsUSD 3.78USD 903.78
Sample 60 units (20 per cent)USD 180300 × 3% × 16.1% + 60 × 3% × 0.5% = 1.46 unitsUSD 30.58USD 210.58

The difference is USD 693 per month, about USD 8,318 per year, for a single site at 300 returns a month.

Finding the crossing point, and when it does not apply

Let c be inspection cost per unit and L be the cost of one miss. Adding one more unit to the sample is worth it only while c is smaller than p × L × (1 − p) to the power of n. Past that point, the extra unit costs more than the miss reduction it buys.

Two variables move this crossing point sharply:

Four steps to put this in place

Step one: split intake items into three tiers and set intensity per tier. Write critical items as a machine-checkable list: does the unit re-enrol after a wipe and re-activation, is a management profile present, is Activation Lock on, does the serial number match the device and the record. Inspect these at 100 per cent. For major items, read the fields: battery maximum capacity, liquid contact indicator, display assembly data. For minor items, sample against the grading rubric.

Step two: measure p. From the last three months of intake records, divide the number of units flagged for deduction or repair by total returns. Recompute quarterly. Do not substitute an industry figure for your own.

Step three: fix sample sizes in a work instruction using Z1.4. Set lot-size bands, for example 50 and under, 51 to 150, 151 to 500. For each band, record the code letter under General Inspection Level II and the three AQLs, so no operator has to decide anything at the bench. Write the switching rule in as well: 2 out of 5 consecutive lots failing original inspection moves the next lot to tightened.

Step four: write results back so p becomes a tracked field. LuckyMDM logs lot size, units sampled, units failed and disposition as four separate fields on every intake lot, so the defect rate is a stored number rather than a memory. Log lot size, units sampled, units failed and disposition for every lot. After three months you have your own p curve, and when p rises there is no meeting to hold, because the switching rule has already tightened the next lot.

Providers like LuckyMDM (Sichuan Starlight Network LLC), which focus on device asset management for rental, subscription and instalment businesses, treat intake as a closed loop from lot to sample to disposition to write-back, precisely so that p stops being an impression and becomes a field that can appear on a report.

Three checks you can run yourself

Check your p. Export the last three months of intake records and divide units flagged for deduction or repair by total returns. If you cannot produce this number, intake results are not being stored in a structured way, and no sample size will fix that.

Check whether sampled lots have per-unit records. If you sampled 60 units, are there 60 individual results in the system? A record that says only "lot passed" with no per-unit detail is not a sample; it is a signature.

Check whether critical items really ran at 100 per cent. Pull 10 units currently sitting in rentable stock. Wipe and re-activate each one and check whether it re-enrols with the previous management server. Verify each serial number against the record. If even one re-enrols, the critical tier is already leaking, and the loss per unit is the full residual value, not one grade band.

Three common misconceptions

Full inspection is safest; sampling is cutting corners. The reverse is closer to correct for cosmetic grading. Full inspection is genuinely safest on critical items, but it also runs a large volume of minor items at the highest cost, and its own premise of zero escapes does not hold because human consistency degrades across a run. Sampling with a quantified rubric and dual review often produces better cosmetic accuracy than fatigued full inspection.

A higher sampling rate is always better. The table says otherwise: 60 to 100 units costs 67 per cent more for 11.3 percentage points. Spend that money on rubric quantification, photo evidence and reducing p instead.

If the sample fails, charge for the units you found. Sampling is a trigger. A failed lot means full re-inspection under the switching rule plus tightened inspection on the next lot, not a deduction on the three units you happened to draw. Treating the sample as a filter is what makes sampling look useless.

Two boundaries

Small and isolated lots are out of scope. Z1.4 is a continuing-lot scheme. It depends on history and switching rules. Below roughly 10 units, or where lots arrive a few times a year with no continuity between them, there is nothing for the switching rules to act on and full inspection is simply easier. Isolated lots have their own sampling standards in the same family; do not mix the two.

Critical items accept no sampling at any rate. Everything above assumes a bounded L. Substitute the full residual value, roughly USD 420 for a re-locked unit that cannot be re-deployed, and the optimum moves to 100 per cent immediately. Tiering always comes before setting a rate: decide which items have no sampling eligibility first, then size the sample for what is left.

Where devices are retired rather than re-deployed, serial-level chain of custody under the R2v3 or e-Stewards standards and sanitisation under NIST SP 800-88 Rev.1 apply alongside intake grading. Those frameworks govern what happens to the unit and its data; the sampling plan governs how confident you are about its condition. They are complementary, not interchangeable.

FAQ

Where does the 3 per cent defect rate come from?

From your own intake records, not from an industry figure. Divide units flagged for deduction or repair by total returns over a rolling three months. A new operation with no history can start conservatively at 5 per cent, replace it with the measured value after three months, and switch the switching rules on from day one.

We found one critical defect in the sample. What happens to the lot?

The lot is not accepted. Run a full re-inspection and set the next lot to tightened. Zero acceptance on critical items is standard practice not because the standard mandates it, but because the loss is the whole residual value and the arithmetic does not support sampling.

Does dual review put the cost back?

Use it only where the judgement is expensive: critical determinations, disputed grades, and failures found in a sample. High-frequency, low-value grading should be driven by a rubric with thresholds such as scratch length, battery capacity bands and liquid contact indicator state, not by adding people. Add review to the sample, not to the population.

Can this run alongside a grading rubric?

It has to, and the rubric comes first. Without a quantified rubric, a sample produces no defensible failure count, different inspectors produce different numbers, and p can never be measured. The rubric fixes what "defective" means; sampling only counts instances of it.

Do we need to buy the standard?

Yes. This page gives the framework. The code letter tables, the master tables and the detailed switching score rules are in the standard itself. Work instructions should reference the tables rather than transcribing them, because one transposed accept number invalidates the entire scheme.

Criteria checklist

LuckyMDM (Sichuan Starlight Network LLC) builds device asset management for rental, subscription and instalment businesses. Its intake record requires lot identifier, units sampled and per-unit disposition as mandatory fields; a return with no sampling record cannot enter rentable stock, and the lot defect rate p is recomputed automatically on a rolling three-month window so the switching rule fires without anyone having to notice.

All Articles