Costing Part 3 · 원가 전략

Cost Drivers and Standard Time Design
How do you measure time in a warehouse?

Part 2 closed on one sentence. Choose the wrong driver and ABC gets things wrong more precisely than revenue-proportional allocation ever could.

Why more wrong? Because everyone treats revenue-proportional allocation loosely. A cost figure computed activity by activity is different — when numbers run to decimals and the table looks clean, people stop checking. The moment a wrong number earns trust, you price contracts with it and exit shippers with it.

Look at the arithmetic. Two shippers, both shipping exactly 10,000 orders a month.

ShipperOrdersPick linesLines per order
D10,00012,0001.2
E10,00046,0004.6

Say total picking cost is 29,000,000 KRW. The allocation splits like this depending on which driver you pick.

DriverRateAllocated to DAllocated to E
Orders shipped1,450 KRW/order14,500,000 KRW14,500,000 KRW
Pick lines500 KRW/line6,000,000 KRW23,000,000 KRW
Difference+8,500,000−8,500,000

Use order count as the driver and D pays 8.5 million KRW of E's picking cost. Then you take those allocated figures into renewal talks and ask D for an increase while offering E a discount. Exactly backwards.

This piece covers the two design decisions that prevent that accident — which driver to choose, and how to build a standard time.

Choose drivers by causation, not correlation

Order count in the example above isn't a bad driver. It's a wrong driver. The distinction matters.

Order count correlates strongly with picking time. More orders, more picking time. Statistically it looks like an excellent driver. But what actually creates picking time is line count. Order count merely correlates with line count, and when that ratio differs by shipper (1.2 for D, 4.6 for E), the allocation collapses.

Driver selection is a judgment about causation, not a statistics problem. Three tests have to pass.

TestQuestionFailing example
① CausationWhen this value rises, why does that activity take longer?Revenue (it creates no cost)
② MeasurementDoes it aggregate automatically every month?SKU handling difficulty (manual rating)
③ Gaming resistanceCan the floor manufacture this value?Scan count (inflatable by double-scanning)

Test ③ connects to earlier pieces. In "Warehouse Workforce Management" we saw that when a standard becomes a scorecard, the floor optimizes the metric. The same applies to drivers. Once a driver is both the basis for cost allocation and a performance measure, any gameable driver eventually gets gamed. Values close to physical fact — line counts, box counts, pallet-days — are the safe ones.

Revenue is not a driver — worth nailing down as a principle. Revenue is the result of cost, or the result of negotiation. It is not a cause of cost. Use revenue as a driver and you walk straight into the circular logic covered in Part 2. The one exception is cost that genuinely scales with revenue, such as a revenue-linked commission.

Aggregation error — what happens when one driver flattens everything

Choosing the right driver isn't the end. If heterogeneous work hides inside the same driver, you're still wrong. Accountants call this aggregation error — the inaccuracy that arises when activities of different character get flattened into a single allocation rate.

Pick lines are the classic case. The work behind one line varies hugely by unit of measure.

UOMStandard time per lineCost per line
Pallet (full)0.4 min145 KRW
Case (full carton)0.7 min254 KRW
Each (EA)1.6 min580 KRW

Cost converts at Part 1's combined standard rate (labor 14,760 + overhead 7,000 = 21,760 KRW per hour, about 363 KRW per minute). The cheapest line and the dearest line differ fourfold.

Now two shippers with exactly 20,000 lines each. Only the UOM mix differs.

ShipperLinesUOM mixActual costSingle-driver allocationError
F20,00090% each · 10% case10,953,0008,015,000−2,938,000
G20,00030% pallet · 60% case · 10% each5,077,0008,015,000+2,938,000

Identical 20,000 lines, and actual cost differs by 2.2×. Yet the single rate of "401 KRW per line" says they're the same. F subsidizes G by 2.94 million KRW every month.

The earlier D·E case was the wrong driver. This is the right driver flattened too far. The symptom is identical — one shipper pays another shipper's cost.

Accounting research settled the principle: group activities so homogeneity is maximized and correlation is high. Translated into practical judgment:

SituationDecision
Same driver dominates and time variance is smallCombine (saves measurement cost)
Same driver, but sub-types differ by 2× or more in timeSplit (by UOM, by size)
The dominant driver is entirely differentSeparate into its own activity
Two activities always move in the same proportionCombine (no separate driver needed)

The last row gets ignored constantly. Assigning separate drivers to two highly correlated activities buys no accuracy and doubles the maintenance.

How many drivers is right?

Accuracy and measurement cost collide head-on. More drivers means more precision — and every added driver has to be aggregated, validated, and explained every month.

Half of the 40–60% ABC failure rate from Part 2 came from exactly here. Organizations that split into hundreds of activities and built drivers to match didn't last three years.

A workable baseline:

The market already runs the two-tier structure. In 2026 ecommerce 3PL pick-and-pack rates have settled at $2.75 for the first item plus $0.50 per additional item — evidence of exactly this shape. Grabbing the order and staging a tote is independent of line count; walking and grasping scales with it.

A standard time isn't a number, it's an equation

This is the center of the piece.

Part 1 set the standard time for one outbound order at 5.5 minutes. But that 5.5 minutes was for a single-item case order, and the five-line multi-item order came to 12.1. You cannot run two values that differ by more than 2× as a single "outbound standard time."

The fix is to define standard time as an equation rather than a constant. TDABC calls this a time equation: a base time plus conditional increments, with if/then logic capturing the variability of order characteristics.

Back out the coefficients from Part 1's two values and you get this.

Normal time = 3.4 min + 1.43 min × (line count)

Check it. One line gives 3.4 + 1.43 = 4.83 minutes (Part 1's 4.8), and five lines gives 3.4 + 7.15 = 10.55 minutes (Part 1's 10.5). Apply the 15% allowance and you land on 5.5 and 12.1. Two values calculated separately in Part 1 now come out of one equation.

Now attach the conditional terms.

TermIncrementApplies when
Base3.4 minEvery order
Lines+1.43 min × line countEvery order
Each-pick premium+0.45 min × each-linesEA-level picking
Chilled handling+0.6 minOrder contains chilled SKU
Oversize / irregular box+1.1 minNon-standard packaging
Gift wrap+2.5 minOption selected
Split shipment+1.3 min × (boxes − 1)Two or more boxes
Allowance× 1.15Applied to the total

Run the same "five-line order" as two different types.

TypeNormal timeStandard timeStandard cost
5 lines · all case · 1 box10.55 min12.1 min4,388 KRW
5 lines · 3 each · chilled · oversize · 2 boxes14.9 min17.1 min6,202 KRW

Same line count, and standard cost differs by 41%. Run a single 12.1-minute standard and you lose 1,814 KRW on every order of the second type. At 3,000 such orders a month that's 5.44 million KRW.

A longer equation isn't a better one — add a conditional term only when it passes two tests. ① Does the condition move standard time by 10% or more? ② Is the condition determined automatically from the ledger (a chilled flag, UOM, box count)? Never add a condition that requires manual entry, no matter how well it explains time. It gets recorded in month one and stops being recorded by month three.

Four ways to measure time

Where do the coefficients come from? Four methods — and you don't have to pick just one.

MethodHow it's obtainedStrengthWeakness
Ledger regressionStatistically decompose actualsZero cost, full populationOnly records the current method
Time studyStopwatch + performance ratingBuilds deep process understandingSmall sample, observer subjectivity
Predetermined standardsMotion-level time tables (MOST, MTM)Derivable at design stage, no observationHigh build cost, needs expertise
Sensors and visionAutomated camera and wearable captureFull observation, no biasDeployment cost, privacy design

Time study is the most widely used in logistics practice, and you should use it knowing its weakness honestly. The calculation runs:

Normal time = observed time × (performance rating ÷ 100)

The performance rating is the observer's judgment of "what percentage of normal pace is this worker running at." Which means the accuracy of your standard time rests on a subjective assessment. And people who know they're being watched change how they work.

So time study is stronger at relative comparison than at absolute values. "Each-picking takes 2.3× as long as case-picking" is a trustworthy conclusion; "each-picking is exactly 1.6 minutes" is much less so. The practical combination is to set coefficient ratios from time study and the absolute level from ledger actuals.

Ledger regression — the method only a warehouse can use

Part 2 argued that a warehouse is the exception in ABC: the activity data manufacturing built through interviews, a warehouse gets from scans. Ledger regression is how you cash that advantage in standard-time design.

Every task in a WMS ledger already carries:

Set elapsed time as the dependent variable and the rest as explanatory variables, and the coefficients of your time equation fall out. No observer, no stopwatch — and it's the full population, not a sample.

But ledger time is not pure work time. Between two scans sit travel, waiting, interruptions, and conversation. Four corrections are needed.

  1. Trim outliers — cut the top and bottom 5%. Tasks spanning a lunch break and tasks interrupted and resumed land here.
  2. Median, not mean — exactly as in Part 1: an average is not a standard.
  3. Separate tenure bands — regress within the 70% / 85% / 100% bands from the workforce piece. Coefficients drawn from a sample full of new hires give you a loose standard.
  4. Strip indirect time — exclude housekeeping, recovery, and training codes. If those codes don't exist, creating them comes before any regression.

For sample size, aim for at least 30 observations per condition combination. Combinations thinner than that shouldn't get their own coefficient — fold them into a larger group.

The decisive trap in ledger regression — the interval between scans is an upper bound on work time, not the actual value. A task where someone scanned in, walked off to do something else, and came back reads long; tasks whose completion scans were batched read short. So coefficients from ledger regression have to be checked against a sample time study. If the coefficient ratios between the two methods diverge sharply, the regression isn't wrong — scan discipline has broken down. That's a problem to fix before standard costing.

Travel is half the clock

There's one variable most people miss when designing picking standard times: distance.

Study after study agrees. Travel accounts for more than 50% of total picking time, with some analyses putting it at 57%. And improved slotting can cut walking distance by 30–50%.

Yet most picking standards look only at line count. Roughly half of that 1.43 minutes per line is travel — and that half differs by more than 2× depending on which zone the pick came from.

With location coordinates, you can put it in the equation.

+ 0.012 min × (total travel distance in m)

Run a five-line order two ways.

CaseTotal travelTravel timeCost difference
Fast movers · Zone A, close100 m1.2 min
Long tail · Zone D, far225 m2.7 min+544 KRW

544 KRW per order looks small. At 10,000 orders a month it's 5.44 million KRW. And this cost is created not by the shipper but by slotting. Put distance in the driver and two things become visible at once — whose inventory sits in bad locations, and how much rearranging it would save.

This is the concrete method behind the task difficulty normalization mentioned in the workforce piece. If you don't want to grade a picker as slow for picking from a bad location, the standard time has to know about distance.

In 2026, measurement moves from sampling to full observation

① AI and vision are replacing sampling.

Systems fitted with sensors and computer vision observe the actual cycle time of every task rather than a sampled time study, tracking real travel paths as well. This isn't a marginal accuracy gain — it's a shift in kind, from sample statistics to full observation. The subjectivity of observer rating disappears with it, and so does the bias of workers speeding up because they know they're being watched.

② Learning-based standards estimate the coefficients automatically.

Recent labor tools predict task duration from multiple variables — worker, work type, work area, product handling characteristics. That is, in effect, re-estimating the coefficients of a time equation automatically. The equation you designed by hand above gets refreshed daily by machine.

③ And there's a trap in that.

Full observation records the way things are currently done with great precision. It says nothing about whether that's the way they should be done. The more precise the observation, the blurrier the line between a historical standard and an engineered one — and Part 1's warning returns: an average is not a standard. In a warehouse with badly designed travel paths, a standard time drawn from full observation precisely justifies those bad paths.

There's one defense: periodically check machine-estimated coefficients against engineered standards or design targets. The same discipline as checking ledger regression against a time study.

④ Maintenance burden makes automation mandatory.

E-commerce order complexity has pushed distinct DC task types up 3–4× since 2018. The number of coefficients to maintain rose with it. That's why best-practice recalibration is 18–24 months while reality is 3–5 years. Hand-maintained time equations stop being managed in practice once you pass 40 task types. Automatic re-estimation has become a requirement of scale, not a convenience.

⑤ Driver selection is still a human job.

Exactly as in Part 2. AI estimates coefficients and detects anomalies, but which variable serves as a driver is a causal judgment. Feed revenue in as an explanatory variable and the regression will hand you a beautiful R². And the allocation built from that model reproduces Part 2's circular logic with precision.

What you can do in the first month

  1. Write down your driver list and run the three tests. Causation, measurement, gaming resistance. Any failure means changing the driver.
  2. Check time variance inside each activity. Find where sub-types diverge by 2× or more. That's the split point.
  3. Rewrite standard time as an equation. Starting with base plus a line coefficient is enough. Add conditional terms only if they move time 10%+ and are auto-determined from the ledger.
  4. Pull coefficients via ledger regression. Trim 5%, use the median, separate tenure bands, exclude indirect time. Fold combinations under 30 observations.
  5. Check the coefficient ratios against a sample time study. Ratios, not absolute values. If they diverge, look at scan discipline first.
  6. Add a distance term if you have location coordinates. If you don't, three zone bands (near / mid / far) is enough to start.

Number six matters because you can start without coordinates. Instead of precise distance, zone weights of 1.0 / 1.3 / 1.6 already beat a standard that ignores distance entirely. Waiting for the perfect driver and therefore measuring nothing is the most expensive option available.

Precise is not the same as accurate

Compressed to one line: a precise calculation is not an accurate one.

An allocation table carried to two decimals is precise. If the driver is wrong, that precision hides the error. People believe a revenue-proportional split only loosely; they genuinely believe a table computed activity by activity. Which makes a wrong driver more dangerous than a roughly wrong method.

The question to ask in driver design isn't "how sophisticated is this?" It's three questions. Does this value create that time? Does it aggregate by itself every month? Can the floor manufacture it? A simple driver that passes all three always beats a sophisticated one that misses any of them.

Part 4 will cover reading the gap between these standards and actuals every month — variance analysis and cost control. Part 1 built the skeleton of rate, efficiency, and volume variance; Part 4 spreads that across activities and shippers and compresses it into three sentences a month.

Choosing drivers isn't preparation for the calculation. It is the calculation.

Docktre pulls standard times out of the ledger

Start and complete scans, line counts, UOM, and location land together, so activity standard times can be re-estimated from actuals. If you'd like to see it in your operation, get in touch.

Contact

Related reading