AI evaluation · Financial analysis · Product concept

AI-Assisted
Financial Review

I extended a UCL comparison of Moonpig and M&S into a five-case exploration of AI financial commentary. The work examined ratio definitions, business context and source support, then informed a review workflow for checking interpretations before publication.

Four-person coursework project Independent AI comparison and review concept Exploratory comparison · five cases, one response per condition
2 firms
Contrasting operating models
3 years
FY22 to FY24 source evidence
5 cases
10 captured retrospective outputs
1 run
Per condition. Single run, exploratory
01

The interpretation problem

The coursework compared Moonpig, a digital-first asset-light gifting platform, with Marks & Spencer, a large retailer with stores and inventory. Ratios were calculated by hand, then interpreted with ChatGPT-4 and Claude.

The failures were not arithmetic. They appeared when a model had to explain why a number looked unusual. A low current ratio reads as distress unless you know how the business collects cash, and a ratio built on an unresolved definition cannot support a conclusion at all. Both need a source check before publication.

02

Worked example: liquidity

Moonpig’s current ratio fell from 2.34 to 0.24. I ran the same liquidity case with and without a set of review checks in the prompt.

Baseline response

Investigate before declaring a crisis

The newer baseline model did not call it distress. It asked for cash flow, liability composition, customer payment timing and financing access before judging.

Response with review checks

An explicit decision gate

The checked version named what had to be confirmed first: customer prepayments, supplier timing, operating cash generation and liability composition, and said no recommendation could be made until they were.

Five cases, one response per condition, same session, source report already read. Exploratory, not a benchmark. The baseline was already cautious here. The difference is that the checked version named which facts were missing.

03

Review workflow

The concept is a review layer for one proposed user: a junior analyst checking AI-assisted commentary before it is published. It is not an autonomous analyst, and investment sign-off stays out of scope.

01

Definition

Confirm what each ratio is built from before any comparison. Different bases are not compared.

02

Context

Check the business model and sector before mapping a ratio to a conclusion.

03

Evidence

Label every claim as supported fact, inference, or unresolved. Link each supported claim to its source.

04

Human decision

If a definition is unresolved or a source is missing, the system abstains and asks for a source check before sign-off.

A prompt checklist is cheapest, but records nothing. Spreadsheet QA handles formulas and definitions, less so a narrative claim with no source. A review layer keeps the claim, its label and its evidence together.

04

MVP and next test

Designed a review workflow linking commentary to definitions, context and evidence. The MVP would carry source-linked ratio commentary, definition checks, evidence labels and abstention.

Three proposed measures, none yet collected: unsupported claims correctly flagged, material errors missed, analyst review time. Validation needs to be external: fresh cases, repeated runs in clean contexts, independent scoring.

05

Evidence and method

Project scope and authorship

Four-person UCL group project comparing Moonpig Group plc and Marks & Spencer Group plc, with hand-calculated ratios and LLM-assisted interpretation. The problem reframing, the five-case retrospective evaluation, the review workflow and the product concept on this page are my independent extension of it.

Source: UCL MSIN0211 Financial Frameworks group report and its cited company filings.

Evaluation method and its limits

Model: GPT-5.6 Sol. Date: 29 Sep 2026. Five cases, two conditions (baseline and with review checks), one response per condition. Scoring was a same-session qualitative rubric applied by me, not independently reviewed.

Both conditions ran in the same conversation, after the source report had already been read. That creates leakage and scorer bias. The comparison is exploratory and no effectiveness claim is made from it.

I scored the responses against a five-part rubric at the time. Those scores are not reproduced here, because a single self-scored run cannot support a performance claim.

Rubric used

DimensionQuestion
Numerical accuracyDoes the answer use the supplied figures correctly?
Accounting disciplineDoes it test whether the ratio/definition is economically appropriate?
Business-model contextDoes the interpretation fit how the company operates?
Industry contextDoes it avoid applying generic thresholds mechanically?
Evidence disciplineDoes it distinguish fact, inference and unresolved evidence?
Recommendation qualityAre actions proportionate to the evidence supplied?
Example: liquidity baseline vs guardrail prompt
Baseline
Moonpig: Current Ratio 2.34→0.24; Quick Ratio 2.13→0.17. M&S: Current Ratio 0.92→0.86; Quick Ratio approximately 0.57→0.52. Interpret the liquidity position of Moonpig and Marks & Spencer. Assess the financial risk implied by these ratios and recommend management actions.
Guardrail addition
Before treating a liquidity ratio as a distress signal: - identify the company operating model; - consider customer cash-collection timing; - examine supplier/payment cycles; - distinguish operational working-capital structure from financing distress; - state what additional evidence is required before recommending emergency financing or restructuring.
Interpretation risks found in the original outputs
  • Asset-story hallucination. The model linked Moonpig’s non-current asset shift to warehouses and physical expansion, which the filings do not support.
  • ROE without denominator context. A reported −291.7% ROE was read as deteriorating profitability without examining what the negative equity base was doing.
  • Liquidity threshold blindness. Low current ratios were read as generic distress without checking how the business collects and pays cash.
  • Method inconsistency. Inventory, payables and the cash conversion cycle changed materially depending on closing versus average balances. The coursework cash conversion cycle of −228.85 days depends on which payables definition is used, so it is not treated as ground truth.

Reported margins also sit on different bases. Moonpig’s 22.9% coursework figure reflects an adjusted operating margin and is not directly comparable with the M&S figure. Reconciling these definitions from the filings is unfinished work, and the review checks above exist because of it.

Original coursework outputs and calculations
Manual ratio calculations, Appendix 1.5
Original three-year profitability, efficiency, liquidity and solvency calculations used in the coursework.
Manual ratio calculations from original report
ROE interpretation, original Claude output
Source evidence for the original negative-ROE interpretation later challenged in the report.
Original Claude profitability transcript
Liquidity interpretation, original ChatGPT output
Source evidence for the original liquidity-distress framing.
Original ChatGPT liquidity transcript
Efficiency / CCC methodology, original ChatGPT output
Shows the competing inventory and payables calculation approaches.
Original ChatGPT efficiency transcript