Methods and decisions
How the revived 4399 CRM works, and how well
Where the data comes from, how the 2021 back-end maps onto the 2026 app, how the numbers on the Insights page are calculated, how the AI assistant and its redaction step were evaluated (including the weak results), what AI is and is not used for, and the decisions behind it all.
Data provenance
- Original code: the 2021 submission is preserved in
coursework/; its results and behaviour are not altered. Ports live inweb/src/lib/legacyand are parity-tested against the original functions. - Demo and guest data: generated deterministically by
web/src/lib/sample-data.ts. Names are invented, e-mails use the reservedexample.*domains and phone numbers come from ACMA's ranges reserved for fiction. Meeting places are about 100 Melbourne suburbs and landmarks looked up once through Photon (data © OpenStreetMap contributors, ODbL); the offline basemap is Natural Earth (public domain). - Evaluation corpora: 34 notes with labelled personal details for the redactor, and 32 notes with labelled follow-ups for the assistant. All synthetic and written for this project.
- Visitor data: whatever visitors type into the shared demo account or their own guest sandbox. No course-provided data, real people's details or employer data are used anywhere.
Functional parity with the 2021 back-end
All 40 REST endpoints of the original Express API, and what replaced each one: 15 implemented with the same behaviour, 24 changed (same capability, different mechanism), 1 dropped. A unit test checks this table against the original router files and checks that every replacement it names is really exported.
POST
implemented/contact/ createContact createNewContact
Now:
createContactActionDuplicate check and account linking ported; scoped to the signed-in owner.
POST
implemented/contact/ createContactByUserName createContactbyUserName
Now:
addByUserNameActionAlso the target of the QR scanner and /contacts/add?u= links.
GET
changed/contact/ showContact showAllContact
Now:
listContactsRead inside the /contacts Server Component instead of a JSON endpoint.
POST
changed/contact/ showOneContact showOneContact
Now:
getContactRead inside /contacts/[id]; the original returned any contact whose id the client sent.
GET
changed/contact/ deleteOneContact/ :userName/ :contact_id deleteOneContact
Now:
deleteContactActionA POST Server Action with confirmation; a GET that deletes data can be triggered by a link or prefetch.
POST
changed/contact/ uploadContactImage contactPhotoUpload
Now:
createContactAction,updateContactActionPhotos are a validated field of the contact form (PNG/JPEG/WebP data URL, at most 180 KB), not a multer upload to disk.
POST
implemented/contact/ updateContactInfo updateContactInfo
Now:
updateContactActionSame validation rules (zod schema shared by form and server).
POST
changed/contact/ searchContact searchContact
Now:
searchContactsThe 2021 client never called this endpoint; it filtered the loaded list. That client-side filter is ported 1:1 and parity-tested.
POST
implemented/contact/ synchronizationContactInfo synchronizationContactInfo
Now:
syncContactActionReplays the team's Jest fixture; only the changed fields are copied.
POST
changed/contact/ connectContactToAccount linkToAccount
Now:
confirmInviteActionNo public endpoint: linking happens on the server when an invitee confirms. The original let any signed-in user link any contact to any account.
POST
implemented/contact/ createContactOneStep createContactOneStep
Now:
createContactActionCreate-with-photo in one step is now the only create path.
POST
changed/profile/ addPhone addPhone
Now:
updateProfileActionFolded into one validated profile update (the 2021 UI saved the whole form anyway).
POST
changed/profile/ delPhone delPhone
Now:
updateProfileActionFolded into one validated profile update (the 2021 UI saved the whole form anyway).
POST
changed/profile/ addEmail addEmail
Now:
updateProfileActionFolded into one validated profile update (the 2021 UI saved the whole form anyway).
POST
changed/profile/ delEmail delEmail
Now:
updateProfileActionFolded into one validated profile update (the 2021 UI saved the whole form anyway).
POST
changed/profile/ editFirstName editFirstName
Now:
updateProfileActionFolded into one validated profile update (the 2021 UI saved the whole form anyway).
POST
changed/profile/ editLastName editLastName
Now:
updateProfileActionFolded into one validated profile update (the 2021 UI saved the whole form anyway).
POST
changed/profile/ editOccupation editOccupation
Now:
updateProfileActionFolded into one validated profile update (the 2021 UI saved the whole form anyway).
POST
changed/profile/ editStatus editStatus
Now:
updateProfileActionFolded into one validated profile update (the 2021 UI saved the whole form anyway).
POST
implemented/profile/ editProfile editProfile
Now:
updateProfileActionFree-text status is stored separately from the account state it used to overload.
GET
changed/profile/ showProfile showProfile
Now:
requireUserThe /profile Server Component reads the signed-in user directly.
POST
implemented/profile/ uploadUserImage uploadPhoto
Now:
setPortraitActionSize- and type-checked data URL instead of a file on the server's disk.
GET
dropped/profile/ displayImage displayImage
Portraits are stored as small data URLs and rendered inline, so no image-serving endpoint is needed.
POST
implemented/record/ createRecord createRecord
Now:
saveRecordActionWho, when, where (with map coordinates), notes and custom fields; contact must be the owner's.
GET
changed/record/ showRecord showAllRecords
Now:
listRecordsRead inside /records, /map, /calendar and /home Server Components.
GET
changed/record/ searchRecord searchRecord
Now:
searchRecordsThe original handler used an undefined expressValidator (it threw on every call) and the client never called it; the client-side record search is ported instead.
POST
implemented/record/ deleteOneRecord deleteOneRecord
Now:
deleteRecordActionOwner-scoped: you can only delete your own meetings (the original deleted any id it was sent).
POST
implemented/record/ editRecord editRecord
Now:
saveRecordActionSame action as create, with the record id.
GET
changed/user/ jwtTest isAuth
Now:
getCurrentUserSigned httpOnly session cookie verified on every request (proxy.ts redirect + requireUser), not a JWT in localStorage.
POST
implemented/user/ login handleLogin
Now:
loginActionbcrypt cost 10 as before; case-insensitive user name.
POST
implemented/user/ signup emailCodeVerify + register
Now:
registerActionThe e-mailed code is verified on the server (single use, 5 attempts).
POST
changed/user/ sendEmailcode emailAuthSend
Now:
sendSignupCodeActionThe code goes to the on-screen demo inbox instead of Gmail SMTP (DR-002).
POST
changed/user/ emailVerify emailCodeVerify
Now:
registerActionNo separate verify call: the code is checked when the account is created, so it cannot be verified once and reused.
POST
changed/user/ fastRegisterPrepare emailFastRegister + emailRegisterCodeSend
Now:
inviteContactActionInvitation e-mail with the 10-digit link lands in the inviter's demo inbox (DR-002).
POST
implemented/user/ fastRegisterConfirm emailRegisterVerify + emailFastRegisterConfirm
Now:
confirmInviteActionFixes the inverted check in the original verifier.
POST
implemented/user/ changePassword emailCodeVerify + updatePassword
Now:
sendChangePasswordCodeAction,changePasswordActionCode from the demo inbox; a new password equal to the old one is now refused, which the original only attempted in resetPassword.
POST
changed/user/ sendResetCode sendResetCode
Now:
sendResetCodeActionThe code is only shown in the demo inbox of a browser that has signed in to that account before.
POST
changed/user/ codeValidation userCodeVerify
Now:
verifyResetCodeActionA valid code issues a signed, 10-minute reset ticket (httpOnly cookie).
POST
changed/user/ resetPassword resetPassword
Now:
resetPasswordActionRequires the reset ticket; the original accepted a constant codeVerified: "4399CRMVerified" from any client, and its "same as the old password" check compared two salted bcrypt hashes with ===, so it never fired (now bcrypt.compare).
POST
implemented/user/ checkUserName checkUserDuplicate
Now:
checkUserNameActionSame messages, plus a user-name format rule.
Not counted as API endpoints: /api/* (Served a 'testing' ping and the generated JSDoc HTML pages - not part of the API; not ported) /test/* (Served a generated test report - not ported (CI runs the tests instead)) GET /contact/, GET /record/ (Declared in the routers without a handler, so requests fell through to the 404 catch-all)
Statistics on the Insights page
- Meetings per week. Meetings that already happened, counted per complete Monday-Sunday week in Melbourne time, from the first week with a meeting (at most 26 weeks back). The current week is left out because it is unfinished. The line is the mean of the last 8 weeks; the band is a 95% percentile bootstrap interval from 2,000 resamples of those weeks with the fixed seed 4399. The headline rate uses the same method over all weeks shown, and states the number of weeks.
- Assumption: the weeks in a window are treated as exchangeable (no trend, no week-to-week correlation). Real meeting habits have both, so the band describes variability rather than forecasting it, and with only 8 weeks a percentile interval runs a little too narrow.
- When you meet. A plain count of every meeting by weekday and starting hour. One hue, darker means more, with a table view for screen readers.
- Contacts by recency. Days since the last meeting that already happened. This is a census of the whole address book, not a sample, so it is shown without confidence intervals on purpose.
- Verified helpers. The statistics code in
web/src/lib/stats(normal quantile, Wilson interval, exact sign test, type-7 quantiles, percentile and paired bootstrap) is unit-tested against values from scipy and R, produced byscripts/stats_reference.py.
Evaluation design and results
Two things are evaluated: the redaction step that runs before any AI call, and the follow-up suggestions the assistant makes, against a rule-based baseline. Everything below is recomputed from the code on every build.
Redaction
34 synthetic notes with 49 labelled personal details. A detail counts as caught only if every character of it is removed (any category); a redaction counts as correct if it overlaps a labelled detail. 95% Wilson intervals, which treat details in the same note as independent and so run a little narrow. The redactor runs as the app runs it since DR-006: the meeting contact's and the user's names, plus the full names in the user's address book. For the evaluation that address book is fixed by a rule: every meeting contact in the corpus plus the four directory accounts (19 people).
| Category | Caught | Recall (95% CI) | Precision (95% CI) |
|---|---|---|---|
| E-mail addresses | 7 of 8 (1 partly) | 88% (53%-98%) | 100% (65%-100%) |
| Phone numbers | 14 of 15 | 93% (70%-99%) | 93% (70%-99%) |
| Street addresses | 8 of 9 | 89% (57%-98%) | 89% (57%-98%) |
| Known names | 13 of 17 (1 partly) | 76% (53%-90%) | 94% (73%-99%) |
| All | 42 of 49 | 86% (73%-93%) | 94% (83%-98%) |
Notes fully cleaned: 21 of 28, 75% (57%-87%). Missed or partly missed: "Leila", "mateo dot lopez at example dot com", "0491 five seven zero 313", "Sienna", "corner of Lygon and Elgin streets, number 9", "Oliver", "Grace Okafor". False positives: "3 Collins Street", "1300 4471", "priya".
With only the meeting contact's and the user's names (the DR-003 setting), names were 12 of 17 and all details 41 of 49, 84% (71%-91%). The address book adds one name, “Sam Patel”, and only because he is one of the directory accounts; the intervals overlap almost entirely, so this is a fix for a known leak, not a measured improvement. People outside the address book and lone first names are still missed. These numbers are optimistic because the same person wrote the rules and the corpus.
Follow-up suggestions: LLM against a rule-based baseline
32 synthetic notes with hand-labelled follow-ups, in two splits of 16. A suggestion matches a labelled follow-up when it contains a keyword from each of its keyword groups (“e-mail” is read as “email”; a test checks that every labelled follow-up matches its own keywords). Metrics: mean recall per note and mean F1 per note with seeded percentile bootstrap intervals (4,000 resamples, seed 4399), pooled recall and precision with Wilson intervals, and suggestions made on notes with nothing to do. F1 is defined on every note (on a note with nothing to do it is 1 for no suggestions and 0 for any), so padding the list with guesses does not pay. The two methods are compared note by note on the same notes, for recall and for F1: paired bootstrap interval of the difference, win / tie / loss counts and an exact sign test. A model answer that is invalid, refused or cut off is scored as an empty answer; only infrastructure failures (network, rate limits, provider errors) are left out, and they are counted in the results and the exports.
| Rule-based baseline | Mean recall per note | Mean F1 per note | Pooled recall | Precision | On no-action notes |
|---|---|---|---|---|---|
| Development split (rules written on it) | 100% (no interval), n = 14 | 99% (96%-100%), n = 16 | 24 of 24, 100% (86%-100%) | 96% (80%-99%) of 25 | 0 |
| Held-out split (rules frozen first) | 31% (12%-54%), n = 13 | 42% (21%-63%), n = 16 | 6 of 23, 26% (13%-46%) | 75% (41%-93%) of 8 | 1 |
The baseline is perfect on the notes it was written against and finds roughly a third of the follow-ups in notes it has not seen. That gap is the reason for the split, and the held-out row is the one to quote. On the development split every note scored 100%, so the bootstrap has no spread and no interval is shown for it; the pooled Wilson interval next to it is the honest range.
The LLM side is not published. The site has no AI budget, so the comparison runs in a signed-in visitor's browser with their own key, in the evaluation harness. It uses the assistant's exact prompt (meeting-assist/v1), logs every call to the visitor's AI log and exports results as JSON or CSV. Keyword matching under-credits paraphrases, so recall is a lower bound for both methods.
Parity of the ported 2021 logic is checked separately: the tests load the original functions from coursework/ and compare outputs, and replay fixtures from the team's Jest suites.
AI use statement
The app works fully without AI. One optional feature uses a language model: the meeting-note assistant on a meeting page, plus the evaluation harness that tests it.
What AI does
- Summarises one meeting note in one to three sentences.
- Suggests follow-up actions for the note's author, with timing as written.
What AI never does
- Make or influence any decision about a person, rank contacts or infer anything sensitive.
- Save anything by itself: every answer is a draft until a person accepts, edits or rejects it.
- Send messages, create meetings or change contacts.
- Run without a visitor choosing to use it with their own key.
Data sent to the provider
- Fixed instructions, the meeting date and the note after redaction (e-mails, phone numbers, addresses, the contact's and user's names, and the full names of everyone in the user's contacts removed). The visitor sees the exact text before sending.
- Sent from the visitor's browser straight to the provider they chose: Anthropic (Claude Haiku 4.5 by default, or Claude Sonnet 5.5) or OpenAI (model id of their choice, default gpt-5-mini). The API key stays in the browser (sessionStorage unless they opt in to remembering it) and never reaches this site.
- A Content Security Policy (report-only for now) lists the only places the browser may connect to: this site, the two AI providers and the map tiles. Violations are reported back, so a script sending the key anywhere else would show up.
Human in the loop and audit trail
- Every output is labelled “AI-generated”, including after it is accepted.
- Every call is written to the AI log: the redacted text sent, the answer, the requested and served model, latency, token counts and the human decision, which can be recorded once. The log is viewable and exportable at
/ai-log. The administrator sees these rows in/admin/recordswith the text masked, except on the shared demo account. - “Accepted” means the model's logged answer, unchanged: the server saves the answer already in the log and ignores any text sent with an accept. Any change is recorded as “edited”, together with the final text.
- Limit: entries are reported by the visitor's browser, because the call never passes through this server. The server cannot verify what the provider actually returned; it can only check the shape of the entry and refuse anything that looks like an API key.
This design is informed by the Australian Government's (DTA) policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is not a claim of compliance with any of them.
Privacy by design
- Your data page. Signed-in users can download everything stored about them (JSON, plus CSV per table) and hard-delete their account. Deletion removes the account and every contact, meeting, link, invitation, demo-inbox e-mail, pending code, activity entry and AI-log entry that belongs to it, in one database transaction (all of it or none of it). Other people's contacts that were linked to the account keep their own copy, unlinked. One anonymous row (counts only) records that a deletion happened.
- Activity log. Append-only: views of contact and meeting pages; creates, changes, deletes and exports; sign-ins, sign-outs, sign-ups and password resets; profile changes and AI actions. List, search, map and Insights pages are not logged as views. It stores ids, field names and counts, never the contents. Entries past the retention period are never shown or exported, and are deleted at server start and on routine clean-ups.
- Ownership. Every read and write is scoped to the signed-in owner on the server.
- The public demo admin is untrusted (DR-005). Anyone can sign in as it, so
/admin/recordsalways masks password hashes, e-mail codes, invitation links and the inbox browser key, and shows names, contact details, notes and AI text only for the seeded demo accounts. Guest and registered visitors' rows show ids, timestamps and counts only, and search does not look inside them. Every admin view and export is written to the admin account's activity log.
| What | Why | Kept for |
|---|---|---|
| Account: user name, password hash, name, occupation, status line, phones, e-mails, optional photo | Sign-in, and the details other people copy when they add you by user name or QR code. | Until you delete the account. Guest sandboxes are deleted after 24 hours. |
| Contacts and meetings (who, when, where, notes, custom fields, accepted AI summaries) | The address book and meeting log you keep - the purpose of the app. | Until you delete them or the account. Deleting a contact deletes its meetings. |
| Activity log (contact and meeting pages viewed; anything created, changed, deleted or exported; sign-ins and password resets; AI actions) | Accountability: what was done through your account, and when. List, search and map pages are not logged as views. | 180 days (older entries are never shown or exported, and are deleted at server start); or with the account. |
| Administrator access | The public demo admin sees your rows in /admin/records with names, contact details, notes and AI text masked, and never sees codes, links or password hashes. Each admin view and export is logged under the admin account. | Admin log entries follow the activity-log retention above. |
| AI audit log (redacted text sent to the provider, the answer, model, timing, tokens, your decision) | Transparency about every AI call made with your key, and the human decision on it. | Until you delete the account. Never contains your API key. |
| Demo inbox e-mails (codes, invitations) | Replaces real e-mail in this demo so sign-up and password reset work. | Until the account is deleted. |
| E-mail codes and invitation links | Verifying an e-mail address, password reset, inviting a contact. | Codes 5 minutes, invitations 15 minutes (unconfirmed invitee accounts 16 minutes). |
- No analytics, tracking pixels or third-party cookies.
- No IP addresses or user agents in the database (rate limits are in memory only).
- Your AI API key: it stays in your browser (sessionStorage by default) and goes only to the AI provider you chose.
- Background location: a meeting stores the place you pick on the map, or your current position only when you press "Use my location".
Assumptions and limitations
- The hosted database was attached late (DR-007). The live demo now runs on Turso, so writes persist and every serverless instance sees them, but the privacy and AI features were built and demonstrated on a local production build while production still used a per-instance copy (DR-004). The functions now run in Sydney (DR-008) but the database is in Tokyo, so every query still crosses to Japan and back.
- The shared
demoaccount is shared: other visitors see its activity log, AI log and inbox. Guest sandboxes are private (the demo admin sees only masked rows for them) and deleted after 24 hours. - Until DR-005, the demo admin could read live password-reset codes and every visitor's notes. That was found in review before this upgrade was merged and fixed by masking; it is recorded rather than quietly removed.
- E-mail verification is simulated by the demo inbox (DR-002); it does not prove that anyone owns an address.
- Evaluation corpora are small, synthetic and written by one person; intervals are wide and the redaction results are optimistic.
- Insights treat weeks as exchangeable and the seed data is synthetic, so the demo charts show the method, not real behaviour.
Decision records
Each record follows the same shape: context, the decision (stated first), the options considered, why, what happened (including the weak numbers) and what I would change. Past records are never edited; a new record supersedes an old one.
DR-001 · Accepted
Merge the React front-end and the Express back-end into one Next.js app
Port both halves into a single Next.js App Router application in TypeScript (strict), with Server Components for reads, Server Actions for writes, a handful of Route Handlers for the demo inbox, geocoding and exports, and SQLite through libSQL and Drizzle.
ReadDR-002 · Partly superseded by DR-005
Deliver e-mail to an on-screen demo inbox instead of sending real e-mail
Store every e-mail the app would send in an email_outbox table and show it in an on-screen demo inbox, with the code or link pulled out so the flow can be completed in a click.
ReadDR-003 · Partly superseded by DR-006
Redact meeting notes in the browser before any language-model call
Before any call, the browser removes e-mail addresses, phone numbers, street addresses and PO boxes, and the names the app already knows (the meeting contact and the signed-in user), replacing each with a token such as [EMAIL].
ReadDR-004 · Partly superseded by DR-007
Use Turso when configured, and fall back to a /tmp copy of the seed database
Use Turso (hosted libSQL) whenever a database URL is configured, and otherwise fall back to a per-instance copy of the committed seed database in /tmp, with a visible notice that demo storage resets.
ReadDR-005 · Accepted (supersedes the reset-code claim in DR-002)
Treat the public demo admin as untrusted: mask secrets and visitors' data, and log every admin read
The admin view masks secrets in every row and shows personal details and free text only for rows that belong to the seeded demo accounts, and every admin view and export is logged.
ReadDR-006 · Accepted (supersedes the known-names scope of DR-003; the rest of DR-003 stands)
Redact every full name in the user's address book before a language-model call
The meeting page passes the full names of everyone in the user's address book to the redactor, which replaces a name with [NAME] when the whole "First Last" appears with the same capitalisation.
ReadDR-007 · Accepted
Run the live demo on the hosted Turso database
Production uses the Turso database comp30022-personal-crm (libSQL, hosted in AWS ap-northeast-1), connected through DATABASE_URL and DATABASE_AUTH_TOKEN in the Vercel Production environment only.
ReadDR-008 · Accepted (follows DR-007
Run the server functions in Sydney (syd1), close to the visitors
Pin the functions to Vercel's syd1 (Sydney) region with "regions": ["syd1"] in web/vercel.json, and leave the database in Tokyo.
ReadModel card
Meeting-note assistant and redactor
Intended use, data, evaluation with intervals, failure modes and ethical considerations.
What I'd change
- Provision the hosted database first, fail a production deployment that has none, and choose the function region by an interleaved comparison that includes the database's own region.
- Detect names of people outside the address book in the browser before sending, and have someone else write a held-out set of notes to evaluate it.
- Replace the keyword matcher with a pre-registered rubric scored by two people, so the follow-up evaluation credits paraphrases fairly.
- Add browser end-to-end tests of the main journeys to CI.
- Expire demo-inbox e-mails after a fixed period, like the activity log.