Siya ZanwarSelected work / Product · Strategy · AI
AI evaluation · Marketing analytics · Product concept

Marketing Budget
Review Copilot

I extended a UCL marketing simulation into an exploratory comparison of AI budget recommendations, then used the reasoning differences to design a review screen bringing allocation, evidence and customer value into one decision.

Six-person simulationIndependent AI comparison and Copilot conceptAI allocations were not executed
01

The budget-review problem

ExerciseMinder was a nine-year UCL simulation allocating a $10m annual marketing budget across five channels. In Year 4 our allocation produced more new customers and lower customer value at once.

ROMI is return on marketing investment, CLV customer lifetime value. A recommendation reporting one without the other can look like success and be a loss.

A trade-off hidden by a single metric

More customers did not mean higher customer value

Year 4 → Year 5 · Recorded simulation outcomes

New customers
+13.4%
606,868 → 688,263
Customer lifetime value
−27.4%
351 → 255
ROMI
+17.2%
2.03× → 2.38×

A reviewer needs to see this trade-off before accepting a growth recommendation.

Each pair indexed to Year 4 = 100. Simulated outcomes. The divergence does not establish its cause.
02

Year 4 comparison

I revisited five checkpoints, asking the same model for an allocation twice: plainly, then instructed to rank evidence by strength first.

Baseline response · illustrative paired summary

Prioritised direct performers

Used visit and conversion data as decisive signals. Allocated 0% to TV.

With evidence ranking · illustrative paired summary

Considered the wider funnel

Ranked A/B evidence above visit data, noting a low direct-conversion channel may contribute earlier in the funnel. Kept 8% in TV.

One response per condition. Allocations were not executed and 8% is not established as better. One response named the evidence it weighed.

03

The review screen

A reviewer needs the evidence and the value picture beside the numbers. Three design decisions follow.

Evidence beside allocation

Each share carries its evidence, with tests, associations and observed metrics distinguished rather than weighted equally.

A review flag where removal lacks support

Proposed rule: flag a channel set to zero when no supporting evidence is recorded. Zero is not an error, and the flag does not force spend back.

Customer value beside acquisition

Acquisition and CLV appear together, so growth cannot be read without its effect on value.

Concept screen using the team’s Year 4 allocation and retrospective Year 4 to 5 simulation outcomes. The allocation is the human team’s, not an AI proposal. Static mockup, no evidence attached.
04

MVP and what I deferred

First build: the allocation with its evidence attached, a recorded rationale, and one untuned rule that flags a channel set to zero when no evidence was recorded. Three other specified features are deliberately not in it.

DeferredAutomated evidence scoringScoring evidence automatically needs an agreed standard for what counts as strong. Nobody has written one, so a score would manufacture confidence rather than check it.
DeferredThresholds on the removal flagA threshold needs a false-positive rate to tune against, which only exists once the flag has run against real allocations.
DeferredSearchable decision historyValuable once there is a volume of decisions to search. On day one there are none.

The rule each time: build what makes a judgement visible, defer what makes it automatic until there is evidence about what a good judgement looks like. A proposed prioritisation, not a roadmap that existed.

Next test

Ten runs per condition against a frozen checkpoint, scored on the five dimensions in the protocol below, then review tasks with marketing leads. Neither has been run.

05

Evidence and method

Project scope and comparison method

The original ExerciseMinder simulation was a six-person UCL MSc Management team project. I contributed to collective decisions and co-owned Real-World Application & Conclusion with Armaan. The independent comparison, risk framework and Copilot concept are my extension.

Claude, September 2026. Five checkpoints, one baseline and one guardrailed response at each. Original self-scoring was unblinded and has not been independently reviewed. Exact original prompts, full outputs and model version are not published here.

Prior exposure to the source report may have affected the model. The direction and size of that effect are unknown. Changing several prompt checks together does not identify which individual check caused a difference. Single pairs cannot distinguish a prompt effect from normal variability. Historical simulated outcomes do not establish the performance of unexecuted alternative allocations.

Nine-year simulation trend

Efficiency and customer value told different stories

Marketing Budget Review Copilot · Siya Zanwar0×1×2×3×2.0 target1.42Y12.57Y22.58Y32.03Y42.38Y52.43Y62.68Y72.98Y83.14Y9Marketing Budget Review Copilot · Siya Zanwar020040090Y182Y283Y3351Y4255Y5257Y6258Y7360Y8333Y9
Separate scales, shared Year axis. CLV shown in source units. Year 4 highlighted. All figures are simulation outputs.
Other checkpoints

Year 2: the guardrailed response retained a small TV allocation and proposed a causal test. Year 5: it emphasised retention and a CLV check. Year 8: it questioned causation while largely retaining the allocation. Year 9: it favoured retaining a working strategy. These are descriptions of single pairs, not estimates of a prompt effect.

The original prompts, the full five-channel outputs and the model version string were not retained. These are recorded observations, not transcripts, and nothing has been reconstructed.

Annual simulation data and campaign history
YearROMICLVNew customers
Y11.42×90454,534
Y22.57×82638,783
Y32.58×83675,905
Y42.03×351606,868
Y52.38×255688,263
Y62.43×257727,229
Y72.68×258767,623
Y82.98×360827,295
Y93.14×333879,065

Years 2 and 3: New Member Price Discount. Year 4: Message Customers with Tips and Ideas. Years 5 to 7: Sign Up a Friend. Year 8: Goals With Friends. Year 9: Loyalty Points for Usage.

Years 2 to 7 and Year 9 used 30% Facebook, 35% Branded Search and 35% Email. Year 8 used 25% / 37% / 38% according to the appendix execution record. The report narrative differs. TV and Unbranded Search were 0% in these allocations.

Final customer base: 7,184,869. Cumulative simulated revenue: $2,134,441,426. Average ROMI across Years 2 to 9: 2.59875, rounded to 2.6×. These are team simulation outcomes, not real-world impact.

Proposed replication protocol

Status: a prospective protocol, not a transcription of the original experiment. No repeated-trial results are claimed.

Shared prompt template

You are advising the CMO of a simulated wearable subscription business. You have a $10m annual marketing budget. Use only the pre-decision evidence below. Choose one campaign and allocate exactly 100% across TV, Facebook, Branded Search, Unbranded Search and Email. Explain the evidence for each choice, distinguish causal findings from associations and assumptions, identify the main risk, and say what new evidence would change your recommendation. Do not use future-year outcomes. [Insert frozen checkpoint evidence verbatim.]

Guardrail addition

Before finalising, rank evidence by strength, considering test design, relevance and uncertainty. Check whether an awareness channel supports later conversions. Inspect new customers, exits and CLV together. Flag channel concentration and unsupported assumptions. If evidence conflicts, state the uncertainty and require human review. Preserve the same budget and output format.

Proposed scoring

Five dimensions: numerical validity, evidence labelling, funnel and channel interactions, retention and CLV, and uncertainty discipline. Score 0 for absent or incorrect, 1 for mentioned without evidence or decision consequence, and 2 for supported and used in the decision. Record supporting output excerpts. Validate allocation totals separately. Report distributions and reviewer agreement, without inferring business uplift.

Run in fresh contexts with the same model version, settings, evidence and campaign choices. Randomise condition order. Preserve all outputs, including failures. Remove condition labels before independent review, while acknowledging that guardrail wording may still reveal the condition.

Original team-report evidence
Key Simulation Metrics table
Key Simulation Metrics table · Original team report
Customer Data table
Customer Data table · Original team report
Human campaign allocation
Human campaign allocation · Original team report
Final simulation results
Final simulation results · Original team report