Analytics & machine learning

Models built for the decision, not the benchmark.

Causal inference, forecasting and segmentation, delivered into the meeting where somebody has to commit budget. The test is whether the recommendation survives contact with the executive who has to act on it.

Return on marketing spend by channel, from a multi-variate regression - a method I have run for a decade, modernized. It is live and synthetic: drag any channel's spend and watch the modeled return, the blended ROAS, and where the next dollar earns the most.

$1.42M
Modeled return
1.77x
Blended ROAS
Paid search
Best next dollar
$0$250k$500k$750k$1.00M0$100k$200k$300k$400k$500k
Paid searchSocialVideoDisplay
$100knext $1 → 2.87
$200knext $1 → 1.45
$250knext $1 → 1.33
$250knext $1 → 0.59

Each channel's effect is estimated by a multi-variate regression that holds seasonality, price and promotions constant, so a channel's ROI is isolated from the others - in the real method each coefficient is reported with its confidence interval. The decision is the marginal column: move the next dollar to Paid search, where it returns the most, not to whoever spent the most last quarter. Synthetic data, illustrative of the method.

The carryover decay, price elasticity, the method, and the real record - below.

Machine learning

The question. You have budget to reach a fraction of your list. Which customers should get the campaign, and what do you hand the team so they can run it next quarter without you in the room?

Why a tree and not clustering. Clustering groups customers by how similar they are to each other, and the groups it finds often differ on nothing you care about. A decision tree groups them by what predicts the outcome. That is supervised segmentation, and what comes out is a rule rather than a centroid.

How to read it

  • Left, the customers. Each dot is one person: how often they visit, and how much of your email they open. Green responded to the last campaign, grey did not. The model has never been told which rule to look for.
  • The sweep. The dashed line is the split being tested right now. The trace underneath is how mixed the two resulting groups would be, tested at every threshold on both features in turn. The tree keeps the lowest point and draws it in solid.
  • Right, the tree. Each box is a group, with its size and the share of it that responded. Every box below is that group cut in two by the rule above it.
  • The two controls. How deep the tree may go, and how few customers a group may contain. Those are the only two decisions a person makes here, and they are what separates a rule that holds next quarter from one that memorized this quarter.
Synthetic data, seeded

1 · The customers, and the split being tested

036912050100sessions / weekemail open rate %sessionsopen rateimpurity

Press play to watch the tree search every candidate split.

2 · The groups it splits them into

31% · n=210

Hover or tab to any node to read it.

3 · The rule it hands yousessions > 5.8 and open rate > 80% → 87% respond (n=31), against a 32% base rate.

How much detail to allow. Loosen both and the tree will describe this quarter perfectly and next quarter badly.

95%55%max depth

training rowsheld-out rows

Depth 3: 83% on training, 78% held out. Pull min customers per group down to 2 and the two lines come apart.

A real CART fit, run in the browser on 300 synthetic customers (210 training, 90 held out). Illustrative only, not a measured result.

What you do with the output

The rule in the blue box is the deliverable. It names a segment in terms the team already has in the CRM, so it can be turned into an audience the same afternoon, and it carries the two numbers that decide whether to bother: how much better that segment responds than the list as a whole, and how many people are in it. A lift that big on a segment too small to matter is not a campaign.

It is also auditable, which is the real argument. Anyone can read why a customer was included and disagree with it on the merits. A model that cannot be argued with does not survive a room of people who have to sign off on the spend.

So why not just use XGBoost?

The honest answer is to go and measure it rather than assert it. Same 300 customers, same hold-out, the real library. Boosting grows many shallow trees instead of one deep one, each fitted to what the previous ones got wrong.

Real XGBoost 3.4.1 · same 300 customers120 trees · depth 2 · lr 0.12

Boosting does not grow one deep tree. It grows many shallow ones, and each is fitted to what the ones before it got wrong. On its own, no single tree here decides anything: each contributes a small nudge up or down, and the prediction is the running total.

tree 1
sessions ≤ 6.0open rate ≤ 94%-0.21-0.04open rate ≤ 40%-0.12+0.12
tree 2
sessions ≤ 6.0open rate ≤ 94%-0.18-0.04open rate ≤ 40%-0.11+0.11
tree 3
sessions ≤ 6.0open rate ≤ 94%-0.17-0.03open rate ≤ 40%-0.10+0.10
tree 4
sessions ≤ 6.0open rate ≤ 94%-0.15-0.03open rate ≤ 59%-0.06+0.11
tree 5
sessions ≤ 6.0open rate ≤ 94%-0.14-0.03open rate ≤ 40%-0.08+0.08
tree 6
sessions ≤ 6.0open rate ≤ 94%-0.13-0.03open rate ≤ 70%-0.04+0.10
tree 7
sessions ≤ 6.0sessions ≤ 2.6-0.15-0.08open rate ≤ 40%-0.07+0.07
tree 8
sessions ≤ 6.0open rate ≤ 94%-0.12-0.02open rate ≤ 70%-0.04+0.09
+112more

The first 8 of 120, exactly as the model dumped them. Bars are each leaf's log-odds nudge: green pushes toward responding, grey away. Notice how alike the early ones are, and how small the numbers get.

76%80%84%88%one tree, held out (80.0%)peak 75trees in the ensemble

training rowsheld-out rows

At 75 trees: 88.1% on training, 83.3% held out. Past the peak, the extra trees are fitting noise.

What it bought, and what it cost

  • One tree: 80.0% held out, 7 groups, and a rule a person can read out loud.
  • XGBoost at its best: 83.3% held out at 75 trees. 3.3 points better, and no rule at all.
  • XGBoost run to 120: 78.9% held out. Past the peak it overfits, and ends up worse than the single tree while looking better on training.
Fitted by scripts/portfolio-site/fit_xgboost.py on the same synthetic customers, 210 training and 90 held out. Real measurements from that fit, on synthetic data.

Three points of accuracy at the peak, in exchange for a model nobody in the room can argue with, and only if you know where to stop. Run it to the end and it lands below the single tree. That is the actual trade, and on two clean features with 210 training rows it is a much closer call than the reflex to reach for the stronger model suggests. On a wider, messier problem the ensemble would win by more, and I would use it.

A tree is the answer when a person has to read the model, defend it, and act on it. Its limits are real: it is unstable to small changes in the data, it approximates smooth relationships with staircases, and it will happily carve noise into rules if you let it, which is what the depth control above demonstrates.

How long a dollar keeps working (adstock)

A dollar of advertising does not spend its whole effect the day it runs, it decays over the weeks after. Modeling that carryover is what tells you how far apart to space campaigns before the returns erode.

Finding the optimal price (elasticity)

Every product answers a price change differently. Estimating elasticity - how much volume you lose for a given price increase - is how you find the point where the next price move starts costing more than it earns.

The method is chosen by the decision, and the record
  • Causal inference and experimental design when the question is whether this actually caused that, and a correlation would send real money in the wrong direction.
  • Forecasting when the decision is a commitment: inventory, headcount, revenue.
  • Segmentation when the decision is who to talk to and what to say. A segmentation that produces seven elegant clusters and no change in how a single customer is contacted has produced nothing.

Hi-tech manufacturing.I built a division's data-science and reporting ecosystem from scratch, then moved it to a cloud lakehouse - recruiting and training internal partners so the capability outlived my involvement. Machine-learning forecasting and segmentation markedly sharpened quote-conversion prediction across global markets.

Solar and cleantech. Econometric and regression forecasting that significantly improved financial forecast accuracy, replacing a set of disparate, separately maintained forecasts with a single global one, presented with its assumptions to executive leadership.

Forecasting: the shape of the improvement

time →
ActualModelPrior approach

Illustrative shape only, no figures quoted. The accuracy figures for this work stay on my resume rather than on a public page. What the chart shows is the shape: the model tracking actuals closely where the prior approach swung wide.

What forecasting does not do

It does not tell you what will happen. It narrows the range and makes the assumptions arguable, which is worth considerably more than a point estimate delivered with false confidence. The forecasts I have built earned trust by being explicit about which drivers they were sensitive to, not by being right every quarter. A number without its caveats is not a finding, it is a liability - which is why the last mile is the meeting: the recommendation, the assumptions behind it, and the conditions under which it stops being true, in that order.

Causal inferenceExperimental design and A/B testingRegression and econometric forecastingCustomer segmentationCloud lakehousePythonSQL