Cost Drivers and Standard Time Design
How do you measure time in a warehouse?
Part 2 closed on one sentence. Choose the wrong driver and ABC gets things wrong more precisely than revenue-proportional allocation ever could.
Why more wrong? Because everyone treats revenue-proportional allocation loosely. A cost figure computed activity by activity is different — when numbers run to decimals and the table looks clean, people stop checking. The moment a wrong number earns trust, you price contracts with it and exit shippers with it.
Look at the arithmetic. Two shippers, both shipping exactly 10,000 orders a month.
| Shipper | Orders | Pick lines | Lines per order |
|---|---|---|---|
| D | 10,000 | 12,000 | 1.2 |
| E | 10,000 | 46,000 | 4.6 |
Say total picking cost is 29,000,000 KRW. The allocation splits like this depending on which driver you pick.
| Driver | Rate | Allocated to D | Allocated to E |
|---|---|---|---|
| Orders shipped | 1,450 KRW/order | 14,500,000 KRW | 14,500,000 KRW |
| Pick lines | 500 KRW/line | 6,000,000 KRW | 23,000,000 KRW |
| Difference | — | +8,500,000 | −8,500,000 |
Use order count as the driver and D pays 8.5 million KRW of E's picking cost. Then you take those allocated figures into renewal talks and ask D for an increase while offering E a discount. Exactly backwards.
This piece covers the two design decisions that prevent that accident — which driver to choose, and how to build a standard time.
Choose drivers by causation, not correlation
Order count in the example above isn't a bad driver. It's a wrong driver. The distinction matters.
Order count correlates strongly with picking time. More orders, more picking time. Statistically it looks like an excellent driver. But what actually creates picking time is line count. Order count merely correlates with line count, and when that ratio differs by shipper (1.2 for D, 4.6 for E), the allocation collapses.
Driver selection is a judgment about causation, not a statistics problem. Three tests have to pass.
| Test | Question | Failing example |
|---|---|---|
| ① Causation | When this value rises, why does that activity take longer? | Revenue (it creates no cost) |
| ② Measurement | Does it aggregate automatically every month? | SKU handling difficulty (manual rating) |
| ③ Gaming resistance | Can the floor manufacture this value? | Scan count (inflatable by double-scanning) |
Test ③ connects to earlier pieces. In "Warehouse Workforce Management" we saw that when a standard becomes a scorecard, the floor optimizes the metric. The same applies to drivers. Once a driver is both the basis for cost allocation and a performance measure, any gameable driver eventually gets gamed. Values close to physical fact — line counts, box counts, pallet-days — are the safe ones.
Aggregation error — what happens when one driver flattens everything
Choosing the right driver isn't the end. If heterogeneous work hides inside the same driver, you're still wrong. Accountants call this aggregation error — the inaccuracy that arises when activities of different character get flattened into a single allocation rate.
Pick lines are the classic case. The work behind one line varies hugely by unit of measure.
| UOM | Standard time per line | Cost per line |
|---|---|---|
| Pallet (full) | 0.4 min | 145 KRW |
| Case (full carton) | 0.7 min | 254 KRW |
| Each (EA) | 1.6 min | 580 KRW |
Cost converts at Part 1's combined standard rate (labor 14,760 + overhead 7,000 = 21,760 KRW per hour, about 363 KRW per minute). The cheapest line and the dearest line differ fourfold.
Now two shippers with exactly 20,000 lines each. Only the UOM mix differs.
| Shipper | Lines | UOM mix | Actual cost | Single-driver allocation | Error |
|---|---|---|---|---|---|
| F | 20,000 | 90% each · 10% case | 10,953,000 | 8,015,000 | −2,938,000 |
| G | 20,000 | 30% pallet · 60% case · 10% each | 5,077,000 | 8,015,000 | +2,938,000 |
Identical 20,000 lines, and actual cost differs by 2.2×. Yet the single rate of "401 KRW per line" says they're the same. F subsidizes G by 2.94 million KRW every month.
The earlier D·E case was the wrong driver. This is the right driver flattened too far. The symptom is identical — one shipper pays another shipper's cost.
Accounting research settled the principle: group activities so homogeneity is maximized and correlation is high. Translated into practical judgment:
| Situation | Decision |
|---|---|
| Same driver dominates and time variance is small | Combine (saves measurement cost) |
| Same driver, but sub-types differ by 2× or more in time | Split (by UOM, by size) |
| The dominant driver is entirely different | Separate into its own activity |
| Two activities always move in the same proportion | Combine (no separate driver needed) |
The last row gets ignored constantly. Assigning separate drivers to two highly correlated activities buys no accuracy and doubles the maintenance.
How many drivers is right?
Accuracy and measurement cost collide head-on. More drivers means more precision — and every added driver has to be aggregated, validated, and explained every month.
Half of the 40–60% ABC failure rate from Part 2 came from exactly here. Organizations that split into hundreds of activities and built drivers to match didn't last three years.
A workable baseline:
- Start with one driver per activity.
- Split only when sub-types inside that activity diverge by 2× or more in time.
- When variance is large but splitting is awkward, use a two-tier structure — a fixed amount per order plus a variable amount per line.
- Attach precise drivers to your top three activities by cost, and leave the rest simple.
The market already runs the two-tier structure. In 2026 ecommerce 3PL pick-and-pack rates have settled at $2.75 for the first item plus $0.50 per additional item — evidence of exactly this shape. Grabbing the order and staging a tote is independent of line count; walking and grasping scales with it.
A standard time isn't a number, it's an equation
This is the center of the piece.
Part 1 set the standard time for one outbound order at 5.5 minutes. But that 5.5 minutes was for a single-item case order, and the five-line multi-item order came to 12.1. You cannot run two values that differ by more than 2× as a single "outbound standard time."
The fix is to define standard time as an equation rather than a constant. TDABC calls this a time equation: a base time plus conditional increments, with if/then logic capturing the variability of order characteristics.
Back out the coefficients from Part 1's two values and you get this.
Check it. One line gives 3.4 + 1.43 = 4.83 minutes (Part 1's 4.8), and five lines gives 3.4 + 7.15 = 10.55 minutes (Part 1's 10.5). Apply the 15% allowance and you land on 5.5 and 12.1. Two values calculated separately in Part 1 now come out of one equation.
Now attach the conditional terms.
| Term | Increment | Applies when |
|---|---|---|
| Base | 3.4 min | Every order |
| Lines | +1.43 min × line count | Every order |
| Each-pick premium | +0.45 min × each-lines | EA-level picking |
| Chilled handling | +0.6 min | Order contains chilled SKU |
| Oversize / irregular box | +1.1 min | Non-standard packaging |
| Gift wrap | +2.5 min | Option selected |
| Split shipment | +1.3 min × (boxes − 1) | Two or more boxes |
| Allowance | × 1.15 | Applied to the total |
Run the same "five-line order" as two different types.
| Type | Normal time | Standard time | Standard cost |
|---|---|---|---|
| 5 lines · all case · 1 box | 10.55 min | 12.1 min | 4,388 KRW |
| 5 lines · 3 each · chilled · oversize · 2 boxes | 14.9 min | 17.1 min | 6,202 KRW |
Same line count, and standard cost differs by 41%. Run a single 12.1-minute standard and you lose 1,814 KRW on every order of the second type. At 3,000 such orders a month that's 5.44 million KRW.
Four ways to measure time
Where do the coefficients come from? Four methods — and you don't have to pick just one.
| Method | How it's obtained | Strength | Weakness |
|---|---|---|---|
| Ledger regression | Statistically decompose actuals | Zero cost, full population | Only records the current method |
| Time study | Stopwatch + performance rating | Builds deep process understanding | Small sample, observer subjectivity |
| Predetermined standards | Motion-level time tables (MOST, MTM) | Derivable at design stage, no observation | High build cost, needs expertise |
| Sensors and vision | Automated camera and wearable capture | Full observation, no bias | Deployment cost, privacy design |
Time study is the most widely used in logistics practice, and you should use it knowing its weakness honestly. The calculation runs:
The performance rating is the observer's judgment of "what percentage of normal pace is this worker running at." Which means the accuracy of your standard time rests on a subjective assessment. And people who know they're being watched change how they work.
So time study is stronger at relative comparison than at absolute values. "Each-picking takes 2.3× as long as case-picking" is a trustworthy conclusion; "each-picking is exactly 1.6 minutes" is much less so. The practical combination is to set coefficient ratios from time study and the absolute level from ledger actuals.
Ledger regression — the method only a warehouse can use
Part 2 argued that a warehouse is the exception in ABC: the activity data manufacturing built through interviews, a warehouse gets from scans. Ledger regression is how you cash that advantage in standard-time design.
Every task in a WMS ledger already carries:
- Start scan time, complete scan time
- Line count, quantity, UOM
- Location (and distance, if coordinates exist)
- Worker, shipper, SKU attributes (chilled, size)
Set elapsed time as the dependent variable and the rest as explanatory variables, and the coefficients of your time equation fall out. No observer, no stopwatch — and it's the full population, not a sample.
But ledger time is not pure work time. Between two scans sit travel, waiting, interruptions, and conversation. Four corrections are needed.
- Trim outliers — cut the top and bottom 5%. Tasks spanning a lunch break and tasks interrupted and resumed land here.
- Median, not mean — exactly as in Part 1: an average is not a standard.
- Separate tenure bands — regress within the 70% / 85% / 100% bands from the workforce piece. Coefficients drawn from a sample full of new hires give you a loose standard.
- Strip indirect time — exclude housekeeping, recovery, and training codes. If those codes don't exist, creating them comes before any regression.
For sample size, aim for at least 30 observations per condition combination. Combinations thinner than that shouldn't get their own coefficient — fold them into a larger group.
Travel is half the clock
There's one variable most people miss when designing picking standard times: distance.
Study after study agrees. Travel accounts for more than 50% of total picking time, with some analyses putting it at 57%. And improved slotting can cut walking distance by 30–50%.
Yet most picking standards look only at line count. Roughly half of that 1.43 minutes per line is travel — and that half differs by more than 2× depending on which zone the pick came from.
With location coordinates, you can put it in the equation.
Run a five-line order two ways.
| Case | Total travel | Travel time | Cost difference |
|---|---|---|---|
| Fast movers · Zone A, close | 100 m | 1.2 min | — |
| Long tail · Zone D, far | 225 m | 2.7 min | +544 KRW |
544 KRW per order looks small. At 10,000 orders a month it's 5.44 million KRW. And this cost is created not by the shipper but by slotting. Put distance in the driver and two things become visible at once — whose inventory sits in bad locations, and how much rearranging it would save.
This is the concrete method behind the task difficulty normalization mentioned in the workforce piece. If you don't want to grade a picker as slow for picking from a bad location, the standard time has to know about distance.
In 2026, measurement moves from sampling to full observation
① AI and vision are replacing sampling.
Systems fitted with sensors and computer vision observe the actual cycle time of every task rather than a sampled time study, tracking real travel paths as well. This isn't a marginal accuracy gain — it's a shift in kind, from sample statistics to full observation. The subjectivity of observer rating disappears with it, and so does the bias of workers speeding up because they know they're being watched.
② Learning-based standards estimate the coefficients automatically.
Recent labor tools predict task duration from multiple variables — worker, work type, work area, product handling characteristics. That is, in effect, re-estimating the coefficients of a time equation automatically. The equation you designed by hand above gets refreshed daily by machine.
③ And there's a trap in that.
Full observation records the way things are currently done with great precision. It says nothing about whether that's the way they should be done. The more precise the observation, the blurrier the line between a historical standard and an engineered one — and Part 1's warning returns: an average is not a standard. In a warehouse with badly designed travel paths, a standard time drawn from full observation precisely justifies those bad paths.
There's one defense: periodically check machine-estimated coefficients against engineered standards or design targets. The same discipline as checking ledger regression against a time study.
④ Maintenance burden makes automation mandatory.
E-commerce order complexity has pushed distinct DC task types up 3–4× since 2018. The number of coefficients to maintain rose with it. That's why best-practice recalibration is 18–24 months while reality is 3–5 years. Hand-maintained time equations stop being managed in practice once you pass 40 task types. Automatic re-estimation has become a requirement of scale, not a convenience.
⑤ Driver selection is still a human job.
Exactly as in Part 2. AI estimates coefficients and detects anomalies, but which variable serves as a driver is a causal judgment. Feed revenue in as an explanatory variable and the regression will hand you a beautiful R². And the allocation built from that model reproduces Part 2's circular logic with precision.
What you can do in the first month
- Write down your driver list and run the three tests. Causation, measurement, gaming resistance. Any failure means changing the driver.
- Check time variance inside each activity. Find where sub-types diverge by 2× or more. That's the split point.
- Rewrite standard time as an equation. Starting with base plus a line coefficient is enough. Add conditional terms only if they move time 10%+ and are auto-determined from the ledger.
- Pull coefficients via ledger regression. Trim 5%, use the median, separate tenure bands, exclude indirect time. Fold combinations under 30 observations.
- Check the coefficient ratios against a sample time study. Ratios, not absolute values. If they diverge, look at scan discipline first.
- Add a distance term if you have location coordinates. If you don't, three zone bands (near / mid / far) is enough to start.
Number six matters because you can start without coordinates. Instead of precise distance, zone weights of 1.0 / 1.3 / 1.6 already beat a standard that ignores distance entirely. Waiting for the perfect driver and therefore measuring nothing is the most expensive option available.
Precise is not the same as accurate
Compressed to one line: a precise calculation is not an accurate one.
An allocation table carried to two decimals is precise. If the driver is wrong, that precision hides the error. People believe a revenue-proportional split only loosely; they genuinely believe a table computed activity by activity. Which makes a wrong driver more dangerous than a roughly wrong method.
The question to ask in driver design isn't "how sophisticated is this?" It's three questions. Does this value create that time? Does it aggregate by itself every month? Can the floor manufacture it? A simple driver that passes all three always beats a sophisticated one that misses any of them.
Part 4 will cover reading the gap between these standards and actuals every month — variance analysis and cost control. Part 1 built the skeleton of rate, efficiency, and volume variance; Part 4 spreads that across activities and shippers and compresses it into three sentences a month.
Docktre pulls standard times out of the ledger
Start and complete scans, line counts, UOM, and location land together, so activity standard times can be re-estimated from actuals. If you'd like to see it in your operation, get in touch.
Contact