Case study

Day 46 — from a WhatsApp chat to a payment claim

Freelancers get paid late, and the deal usually lives in a chat. I built the tool that turns that chat into a dated case, and kept the AI on a short leash, because in a document someone signs, a confident mistake is worse than no help at all.

Built and shipped September 2026 · Python · FastAPI · Claude API · 1,707 tests · live

The problem

A designer delivers the logo, sends the invoice, and waits. The order, the delivery, the invoice, the "will pay by Friday" are all in a WhatsApp chat, spread over weeks. India's MSMED Act gives a registered micro or small supplier a hard deadline and compound interest at three times the RBI Bank Rate after it, but using that means knowing the day the work was accepted, the rate for each stretch of time, and how to write it all down without overstating anything.

The project takes its name from that deadline. The sentence that explains it is fixed wording, used the same way everywhere:

Day 46 is the first day on which a payment to a registered micro or small supplier is late under section 15 of the MSMED Act whatever the contract says, because no written credit period may run beyond 45 days from acceptance. Where nothing was agreed in writing, the payment was already late on day 16.

A small survey (nine freelancers) changed the design early: none of them were registered on Udyam, the registration that route depends on. So the ordinary contract-law route had to be built as an equal, not a footnote. Six eligibility questions sort each case onto a route, or say they cannot tell, with the reasons printed in the check's own words.

The rule I built around: the AI only points

A model that writes a date will one day write the wrong one, confidently, into a letter someone signs. So the AI in Day 46 has exactly one job, and every other step belongs to code or to the person:

stepwho does the work
Read the export: every message gets an id, a time, a sendercode, tested
Show exactly what the AI will see, masked, before anything is sentcode, tested
Point at the messages that record an order, a delivery, an invoice, a promise to paythe AI, and only this
Confirm, correct or reject every event; type what a chat cannot saythe person; checked again by code
Due date, interest, part-paymentscode, tested
Reminder, demand letter, case file for an advocatefixed templates

The AI step is one request to Claude under a fixed JSON schema. Its reply can only be a list of type, message, quote, confidence across twelve event types: no dates, no figures, no sentences. Code keeps an entry only if its quote is word for word inside the message it cites, then takes the date and the speaker from that message and reads any amount from the quoted words itself. A buyer's "sent 10,000" is recorded as the buyer saying so, never as money that arrived.

Two smaller guards matter as much. The chat goes in as evidence inside a fence named after its own hash, so text in the chat cannot close the fence and start giving instructions. And the server rebuilds the masked text and compares its SHA-256 with what the page showed, so what is sent is byte for byte what the person read.

Masking, on screen first

Before any request, the page shows the exact text the AI would receive. The names of the people in the chat become SUPPLIER and BUYER; phone numbers, e-mail and UPI addresses, PAN, GSTIN, Aadhaar, bank and card numbers and links become tokens. Amounts, dates and invoice numbers stay, because the events cannot be recognised without them. The person can leave out any message or hide any word. A leak test plants 1,200 private details in a chat and fails the build if a single one survives into what the AI would see. Masking is still a best effort, and the page says so.

The server has no database and no accounts, writes no upload to disk and logs no chat text. The masked excerpt the person approves is sent to the AI provider, and the page says so beside the button.

Measuring the AI step

Twelve synthetic Hinglish chats, 362 messages, 262 events. The answers came first: for every message the events it records were fixed as labels, then the words were written, then the export files were rendered in three phone formats from the same list. Four of the chats were labelled a second time, blind (115 of 119 pairs agreed; the four disagreements were ruled on before the run). No model saw the chats before it was measured, and the prompt was not tuned on them.

modelfoundrighthallucinatedper chatmedian
opus-5 (the app's)99%94%0 of 276USD 0.05416.6 s
sonnet-597%95%0 of 266USD 0.02917.8 s
haiku-4.578%95%0 of 215USD 0.0065.0 s
sonnet-5, one-line prompt92%85%0 of 286USD 0.02817.7 s

Found is the share of labelled events the model pointed at with the right type; right is the share of what it pointed at that is in the labels; hallucinated counts entries whose quote is not in the message they cite, which code drops before anyone sees them.

Same model, same chats: the one-sentence prompt was right 85% of the time, the app's prompt 95%. The instructions are worth ten points.

The caveat, stated before anyone asks: these chats are synthetic. An AI assistant wrote them, not any of the models that read them, and no real chat is in the set. Real exports, with consent, are the next test.

Getting the arithmetic right without a person checking it

With no one to check the sums by hand, every figure is checked three ways. The interest engine must agree to the paisa with a second implementation written from the spec alone, without sight of the engine, on 2,500 random cases, and on a set of golden cases with an Excel workbook of live formulas recalculated by Excel itself. The same idea runs elsewhere: a blind second reading of the confirmation rules is compared on 3,000 generated cases, and the chat reader on 3,000 generated chats every run. Every finding from the adversarial reviews stays in the suite as a test, and the whole suite, 1,707 tests, runs on every push.

Where the law is silent, the engine takes the reading that lowers the claim, prints it, and shows the other reading once as an upper figure. The RBI Bank Rate lives in a dated table with a source for every row; a table not checked for 75 days refuses to produce a figure at all. And wording that overstates the law, from "eligible" to a promised outcome, fails the build.

Shipping it

It runs on Render. The server checks itself before it starts (page files, templates, the saved sample answer), so a broken deploy never answers its health check. Every reply carries a content security policy and no-framing headers; requests from other sites are refused before their body is read; heavy work is counted so that slow uploads cannot shut others out. One command checks the live site end to end, including a whole sample case calculated there and compared with the same code run locally:

$ python -m day46.check https://day46.onrender.com PASS the site answers: after 22 s, AI live PASS the RBI rate table is fresh PASS the law check is fresh PASS the security headers are on PASS /, /chat, /accuracy open: status 200 PASS the sample chat replays: 21 events PASS the sample case is calculated: MSME Act route, interest Rs 2626, the same as this code computes READY: all 12 checks passed.

Decisions I'd defend

What's next

Real chat exports with consent, measured the same way; reading screenshots, for chats that cannot be exported, with a clear line that the images go to the provider unmasked; opening a saved case file again; and a cheaper model once the table says it can be trusted with the promises to pay. Each is one more row in a measurement that already exists.

Try the sample case ↗ The accuracy page ↗ Back to the site →

David Jeremie Anand · New Delhi · every figure on this page comes from the live accuracy page, the committed test suite or a real run of the check above. Day 46 is not a law firm, and nothing it produces is legal advice.