AI-Assisted
Financial Review
I extended a UCL comparison of Moonpig and M&S into a five-case exploration of AI financial commentary. The work examined ratio definitions, business context and source support, then informed a review workflow for checking interpretations before publication.
The interpretation problem
The coursework compared Moonpig, a digital-first asset-light gifting platform, with Marks & Spencer, a large retailer with stores and inventory. Ratios were calculated by hand, then interpreted with ChatGPT-4 and Claude.
The failures were not arithmetic. They appeared when a model had to explain why a number looked unusual. A low current ratio reads as distress unless you know how the business collects cash, and a ratio built on an unresolved definition cannot support a conclusion at all. Both need a source check before publication.
Worked example: liquidity
Moonpig’s current ratio fell from 2.34 to 0.24. I ran the same liquidity case with and without a set of review checks in the prompt.
Investigate before declaring a crisis
The newer baseline model did not call it distress. It asked for cash flow, liability composition, customer payment timing and financing access before judging.
An explicit decision gate
The checked version named what had to be confirmed first: customer prepayments, supplier timing, operating cash generation and liability composition, and said no recommendation could be made until they were.
Five cases, one response per condition, same session, source report already read. Exploratory, not a benchmark. The baseline was already cautious here. The difference is that the checked version named which facts were missing.
Review workflow
The concept is a review layer for one proposed user: a junior analyst checking AI-assisted commentary before it is published. It is not an autonomous analyst, and investment sign-off stays out of scope.
Definition
Confirm what each ratio is built from before any comparison. Different bases are not compared.
Context
Check the business model and sector before mapping a ratio to a conclusion.
Evidence
Label every claim as supported fact, inference, or unresolved. Link each supported claim to its source.
Human decision
If a definition is unresolved or a source is missing, the system abstains and asks for a source check before sign-off.
A prompt checklist is cheapest, but records nothing. Spreadsheet QA handles formulas and definitions, less so a narrative claim with no source. A review layer keeps the claim, its label and its evidence together.
MVP and next test
Designed a review workflow linking commentary to definitions, context and evidence. The MVP would carry source-linked ratio commentary, definition checks, evidence labels and abstention.
Three proposed measures, none yet collected: unsupported claims correctly flagged, material errors missed, analyst review time. Validation needs to be external: fresh cases, repeated runs in clean contexts, independent scoring.
Evidence and method
Project scope and authorship
Four-person UCL group project comparing Moonpig Group plc and Marks & Spencer Group plc, with hand-calculated ratios and LLM-assisted interpretation. The problem reframing, the five-case retrospective evaluation, the review workflow and the product concept on this page are my independent extension of it.
Source: UCL MSIN0211 Financial Frameworks group report and its cited company filings.
Evaluation method and its limits
Model: GPT-5.6 Sol. Date: 29 Sep 2026. Five cases, two conditions (baseline and with review checks), one response per condition. Scoring was a same-session qualitative rubric applied by me, not independently reviewed.
Both conditions ran in the same conversation, after the source report had already been read. That creates leakage and scorer bias. The comparison is exploratory and no effectiveness claim is made from it.
I scored the responses against a five-part rubric at the time. Those scores are not reproduced here, because a single self-scored run cannot support a performance claim.
Rubric used
| Dimension | Question |
|---|---|
| Numerical accuracy | Does the answer use the supplied figures correctly? |
| Accounting discipline | Does it test whether the ratio/definition is economically appropriate? |
| Business-model context | Does the interpretation fit how the company operates? |
| Industry context | Does it avoid applying generic thresholds mechanically? |
| Evidence discipline | Does it distinguish fact, inference and unresolved evidence? |
| Recommendation quality | Are actions proportionate to the evidence supplied? |
Example: liquidity baseline vs guardrail prompt
Interpretation risks found in the original outputs
- Asset-story hallucination. The model linked Moonpig’s non-current asset shift to warehouses and physical expansion, which the filings do not support.
- ROE without denominator context. A reported −291.7% ROE was read as deteriorating profitability without examining what the negative equity base was doing.
- Liquidity threshold blindness. Low current ratios were read as generic distress without checking how the business collects and pays cash.
- Method inconsistency. Inventory, payables and the cash conversion cycle changed materially depending on closing versus average balances. The coursework cash conversion cycle of −228.85 days depends on which payables definition is used, so it is not treated as ground truth.
Reported margins also sit on different bases. Moonpig’s 22.9% coursework figure reflects an adjusted operating margin and is not directly comparable with the M&S figure. Reconciling these definitions from the filings is unfinished work, and the review checks above exist because of it.