Skip to content

Methods and decisions

How the revived 4399 CRM works, and how well

Where the data comes from, how the 2021 back-end maps onto the 2026 app, how the numbers on the Insights page are calculated, how the AI assistant and its redaction step were evaluated (including the weak results), what AI is and is not used for, and the decisions behind it all.

Data provenance

  • Original code: the 2021 submission is preserved in coursework/; its results and behaviour are not altered. Ports live in web/src/lib/legacy and are parity-tested against the original functions.
  • Demo and guest data: generated deterministically by web/src/lib/sample-data.ts. Names are invented, e-mails use the reserved example.* domains and phone numbers come from ACMA's ranges reserved for fiction. Meeting places are about 100 Melbourne suburbs and landmarks looked up once through Photon (data © OpenStreetMap contributors, ODbL); the offline basemap is Natural Earth (public domain).
  • Evaluation corpora: 34 notes with labelled personal details for the redactor, and 32 notes with labelled follow-ups for the assistant. All synthetic and written for this project.
  • Visitor data: whatever visitors type into the shared demo account or their own guest sandbox. No course-provided data, real people's details or employer data are used anywhere.

Functional parity with the 2021 back-end

All 40 REST endpoints of the original Express API, and what replaced each one: 15 implemented with the same behaviour, 24 changed (same capability, different mechanism), 1 dropped. A unit test checks this table against the original router files and checks that every replacement it names is really exported.

  1. POST /contact/createContact

    implemented

    createNewContact

    Now: createContactAction

    Duplicate check and account linking ported; scoped to the signed-in owner.

  2. POST /contact/createContactByUserName

    implemented

    createContactbyUserName

    Now: addByUserNameAction

    Also the target of the QR scanner and /contacts/add?u= links.

  3. GET /contact/showContact

    changed

    showAllContact

    Now: listContacts

    Read inside the /contacts Server Component instead of a JSON endpoint.

  4. POST /contact/showOneContact

    changed

    showOneContact

    Now: getContact

    Read inside /contacts/[id]; the original returned any contact whose id the client sent.

  5. GET /contact/deleteOneContact/:userName/:contact_id

    changed

    deleteOneContact

    Now: deleteContactAction

    A POST Server Action with confirmation; a GET that deletes data can be triggered by a link or prefetch.

  6. POST /contact/uploadContactImage

    changed

    contactPhotoUpload

    Now: createContactAction, updateContactAction

    Photos are a validated field of the contact form (PNG/JPEG/WebP data URL, at most 180 KB), not a multer upload to disk.

  7. POST /contact/updateContactInfo

    implemented

    updateContactInfo

    Now: updateContactAction

    Same validation rules (zod schema shared by form and server).

  8. POST /contact/searchContact

    changed

    searchContact

    Now: searchContacts

    The 2021 client never called this endpoint; it filtered the loaded list. That client-side filter is ported 1:1 and parity-tested.

  9. POST /contact/synchronizationContactInfo

    implemented

    synchronizationContactInfo

    Now: syncContactAction

    Replays the team's Jest fixture; only the changed fields are copied.

  10. POST /contact/connectContactToAccount

    changed

    linkToAccount

    Now: confirmInviteAction

    No public endpoint: linking happens on the server when an invitee confirms. The original let any signed-in user link any contact to any account.

  11. POST /contact/createContactOneStep

    implemented

    createContactOneStep

    Now: createContactAction

    Create-with-photo in one step is now the only create path.

  12. POST /profile/addPhone

    changed

    addPhone

    Now: updateProfileAction

    Folded into one validated profile update (the 2021 UI saved the whole form anyway).

  13. POST /profile/delPhone

    changed

    delPhone

    Now: updateProfileAction

    Folded into one validated profile update (the 2021 UI saved the whole form anyway).

  14. POST /profile/addEmail

    changed

    addEmail

    Now: updateProfileAction

    Folded into one validated profile update (the 2021 UI saved the whole form anyway).

  15. POST /profile/delEmail

    changed

    delEmail

    Now: updateProfileAction

    Folded into one validated profile update (the 2021 UI saved the whole form anyway).

  16. POST /profile/editFirstName

    changed

    editFirstName

    Now: updateProfileAction

    Folded into one validated profile update (the 2021 UI saved the whole form anyway).

  17. POST /profile/editLastName

    changed

    editLastName

    Now: updateProfileAction

    Folded into one validated profile update (the 2021 UI saved the whole form anyway).

  18. POST /profile/editOccupation

    changed

    editOccupation

    Now: updateProfileAction

    Folded into one validated profile update (the 2021 UI saved the whole form anyway).

  19. POST /profile/editStatus

    changed

    editStatus

    Now: updateProfileAction

    Folded into one validated profile update (the 2021 UI saved the whole form anyway).

  20. POST /profile/editProfile

    implemented

    editProfile

    Now: updateProfileAction

    Free-text status is stored separately from the account state it used to overload.

  21. GET /profile/showProfile

    changed

    showProfile

    Now: requireUser

    The /profile Server Component reads the signed-in user directly.

  22. POST /profile/uploadUserImage

    implemented

    uploadPhoto

    Now: setPortraitAction

    Size- and type-checked data URL instead of a file on the server's disk.

  23. GET /profile/displayImage

    dropped

    displayImage

    Portraits are stored as small data URLs and rendered inline, so no image-serving endpoint is needed.

  24. POST /record/createRecord

    implemented

    createRecord

    Now: saveRecordAction

    Who, when, where (with map coordinates), notes and custom fields; contact must be the owner's.

  25. GET /record/showRecord

    changed

    showAllRecords

    Now: listRecords

    Read inside /records, /map, /calendar and /home Server Components.

  26. GET /record/searchRecord

    changed

    searchRecord

    Now: searchRecords

    The original handler used an undefined expressValidator (it threw on every call) and the client never called it; the client-side record search is ported instead.

  27. POST /record/deleteOneRecord

    implemented

    deleteOneRecord

    Now: deleteRecordAction

    Owner-scoped: you can only delete your own meetings (the original deleted any id it was sent).

  28. POST /record/editRecord

    implemented

    editRecord

    Now: saveRecordAction

    Same action as create, with the record id.

  29. GET /user/jwtTest

    changed

    isAuth

    Now: getCurrentUser

    Signed httpOnly session cookie verified on every request (proxy.ts redirect + requireUser), not a JWT in localStorage.

  30. POST /user/login

    implemented

    handleLogin

    Now: loginAction

    bcrypt cost 10 as before; case-insensitive user name.

  31. POST /user/signup

    implemented

    emailCodeVerify + register

    Now: registerAction

    The e-mailed code is verified on the server (single use, 5 attempts).

  32. POST /user/sendEmailcode

    changed

    emailAuthSend

    Now: sendSignupCodeAction

    The code goes to the on-screen demo inbox instead of Gmail SMTP (DR-002).

  33. POST /user/emailVerify

    changed

    emailCodeVerify

    Now: registerAction

    No separate verify call: the code is checked when the account is created, so it cannot be verified once and reused.

  34. POST /user/fastRegisterPrepare

    changed

    emailFastRegister + emailRegisterCodeSend

    Now: inviteContactAction

    Invitation e-mail with the 10-digit link lands in the inviter's demo inbox (DR-002).

  35. POST /user/fastRegisterConfirm

    implemented

    emailRegisterVerify + emailFastRegisterConfirm

    Now: confirmInviteAction

    Fixes the inverted check in the original verifier.

  36. POST /user/changePassword

    implemented

    emailCodeVerify + updatePassword

    Now: sendChangePasswordCodeAction, changePasswordAction

    Code from the demo inbox; a new password equal to the old one is now refused, which the original only attempted in resetPassword.

  37. POST /user/sendResetCode

    changed

    sendResetCode

    Now: sendResetCodeAction

    The code is only shown in the demo inbox of a browser that has signed in to that account before.

  38. POST /user/codeValidation

    changed

    userCodeVerify

    Now: verifyResetCodeAction

    A valid code issues a signed, 10-minute reset ticket (httpOnly cookie).

  39. POST /user/resetPassword

    changed

    resetPassword

    Now: resetPasswordAction

    Requires the reset ticket; the original accepted a constant codeVerified: "4399CRMVerified" from any client, and its "same as the old password" check compared two salted bcrypt hashes with ===, so it never fired (now bcrypt.compare).

  40. POST /user/checkUserName

    implemented

    checkUserDuplicate

    Now: checkUserNameAction

    Same messages, plus a user-name format rule.

Not counted as API endpoints: /api/* (Served a 'testing' ping and the generated JSDoc HTML pages - not part of the API; not ported) /test/* (Served a generated test report - not ported (CI runs the tests instead)) GET /contact/, GET /record/ (Declared in the routers without a handler, so requests fell through to the 404 catch-all)

Statistics on the Insights page

  • Meetings per week. Meetings that already happened, counted per complete Monday-Sunday week in Melbourne time, from the first week with a meeting (at most 26 weeks back). The current week is left out because it is unfinished. The line is the mean of the last 8 weeks; the band is a 95% percentile bootstrap interval from 2,000 resamples of those weeks with the fixed seed 4399. The headline rate uses the same method over all weeks shown, and states the number of weeks.
  • Assumption: the weeks in a window are treated as exchangeable (no trend, no week-to-week correlation). Real meeting habits have both, so the band describes variability rather than forecasting it, and with only 8 weeks a percentile interval runs a little too narrow.
  • When you meet. A plain count of every meeting by weekday and starting hour. One hue, darker means more, with a table view for screen readers.
  • Contacts by recency. Days since the last meeting that already happened. This is a census of the whole address book, not a sample, so it is shown without confidence intervals on purpose.
  • Verified helpers. The statistics code in web/src/lib/stats (normal quantile, Wilson interval, exact sign test, type-7 quantiles, percentile and paired bootstrap) is unit-tested against values from scipy and R, produced by scripts/stats_reference.py.

Evaluation design and results

Two things are evaluated: the redaction step that runs before any AI call, and the follow-up suggestions the assistant makes, against a rule-based baseline. Everything below is recomputed from the code on every build.

Redaction

34 synthetic notes with 49 labelled personal details. A detail counts as caught only if every character of it is removed (any category); a redaction counts as correct if it overlaps a labelled detail. 95% Wilson intervals, which treat details in the same note as independent and so run a little narrow. The redactor runs as the app runs it since DR-006: the meeting contact's and the user's names, plus the full names in the user's address book. For the evaluation that address book is fixed by a rule: every meeting contact in the corpus plus the four directory accounts (19 people).

CategoryCaughtRecall (95% CI)Precision (95% CI)
E-mail addresses7 of 8 (1 partly)88% (53%-98%)100% (65%-100%)
Phone numbers14 of 1593% (70%-99%)93% (70%-99%)
Street addresses8 of 989% (57%-98%)89% (57%-98%)
Known names13 of 17 (1 partly)76% (53%-90%)94% (73%-99%)
All42 of 4986% (73%-93%)94% (83%-98%)

Notes fully cleaned: 21 of 28, 75% (57%-87%). Missed or partly missed: "Leila", "mateo dot lopez at example dot com", "0491 five seven zero 313", "Sienna", "corner of Lygon and Elgin streets, number 9", "Oliver", "Grace Okafor". False positives: "3 Collins Street", "1300 4471", "priya".

With only the meeting contact's and the user's names (the DR-003 setting), names were 12 of 17 and all details 41 of 49, 84% (71%-91%). The address book adds one name, “Sam Patel”, and only because he is one of the directory accounts; the intervals overlap almost entirely, so this is a fix for a known leak, not a measured improvement. People outside the address book and lone first names are still missed. These numbers are optimistic because the same person wrote the rules and the corpus.

Follow-up suggestions: LLM against a rule-based baseline

32 synthetic notes with hand-labelled follow-ups, in two splits of 16. A suggestion matches a labelled follow-up when it contains a keyword from each of its keyword groups (“e-mail” is read as “email”; a test checks that every labelled follow-up matches its own keywords). Metrics: mean recall per note and mean F1 per note with seeded percentile bootstrap intervals (4,000 resamples, seed 4399), pooled recall and precision with Wilson intervals, and suggestions made on notes with nothing to do. F1 is defined on every note (on a note with nothing to do it is 1 for no suggestions and 0 for any), so padding the list with guesses does not pay. The two methods are compared note by note on the same notes, for recall and for F1: paired bootstrap interval of the difference, win / tie / loss counts and an exact sign test. A model answer that is invalid, refused or cut off is scored as an empty answer; only infrastructure failures (network, rate limits, provider errors) are left out, and they are counted in the results and the exports.

Rule-based baselineMean recall per noteMean F1 per notePooled recallPrecisionOn no-action notes
Development split (rules written on it)100% (no interval), n = 1499% (96%-100%), n = 1624 of 24, 100% (86%-100%)96% (80%-99%) of 250
Held-out split (rules frozen first)31% (12%-54%), n = 1342% (21%-63%), n = 166 of 23, 26% (13%-46%)75% (41%-93%) of 81

The baseline is perfect on the notes it was written against and finds roughly a third of the follow-ups in notes it has not seen. That gap is the reason for the split, and the held-out row is the one to quote. On the development split every note scored 100%, so the bootstrap has no spread and no interval is shown for it; the pooled Wilson interval next to it is the honest range.

The LLM side is not published. The site has no AI budget, so the comparison runs in a signed-in visitor's browser with their own key, in the evaluation harness. It uses the assistant's exact prompt (meeting-assist/v1), logs every call to the visitor's AI log and exports results as JSON or CSV. Keyword matching under-credits paraphrases, so recall is a lower bound for both methods.

Parity of the ported 2021 logic is checked separately: the tests load the original functions from coursework/ and compare outputs, and replay fixtures from the team's Jest suites.

AI use statement

The app works fully without AI. One optional feature uses a language model: the meeting-note assistant on a meeting page, plus the evaluation harness that tests it.

What AI does

  • Summarises one meeting note in one to three sentences.
  • Suggests follow-up actions for the note's author, with timing as written.

What AI never does

  • Make or influence any decision about a person, rank contacts or infer anything sensitive.
  • Save anything by itself: every answer is a draft until a person accepts, edits or rejects it.
  • Send messages, create meetings or change contacts.
  • Run without a visitor choosing to use it with their own key.

Data sent to the provider

  • Fixed instructions, the meeting date and the note after redaction (e-mails, phone numbers, addresses, the contact's and user's names, and the full names of everyone in the user's contacts removed). The visitor sees the exact text before sending.
  • Sent from the visitor's browser straight to the provider they chose: Anthropic (Claude Haiku 4.5 by default, or Claude Sonnet 5.5) or OpenAI (model id of their choice, default gpt-5-mini). The API key stays in the browser (sessionStorage unless they opt in to remembering it) and never reaches this site.
  • A Content Security Policy (report-only for now) lists the only places the browser may connect to: this site, the two AI providers and the map tiles. Violations are reported back, so a script sending the key anywhere else would show up.

Human in the loop and audit trail

  • Every output is labelled “AI-generated”, including after it is accepted.
  • Every call is written to the AI log: the redacted text sent, the answer, the requested and served model, latency, token counts and the human decision, which can be recorded once. The log is viewable and exportable at /ai-log. The administrator sees these rows in /admin/records with the text masked, except on the shared demo account.
  • “Accepted” means the model's logged answer, unchanged: the server saves the answer already in the log and ignores any text sent with an accept. Any change is recorded as “edited”, together with the final text.
  • Limit: entries are reported by the visitor's browser, because the call never passes through this server. The server cannot verify what the provider actually returned; it can only check the shape of the entry and refuse anything that looks like an API key.

This design is informed by the Australian Government's (DTA) policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is not a claim of compliance with any of them.

Privacy by design

  • Your data page. Signed-in users can download everything stored about them (JSON, plus CSV per table) and hard-delete their account. Deletion removes the account and every contact, meeting, link, invitation, demo-inbox e-mail, pending code, activity entry and AI-log entry that belongs to it, in one database transaction (all of it or none of it). Other people's contacts that were linked to the account keep their own copy, unlinked. One anonymous row (counts only) records that a deletion happened.
  • Activity log. Append-only: views of contact and meeting pages; creates, changes, deletes and exports; sign-ins, sign-outs, sign-ups and password resets; profile changes and AI actions. List, search, map and Insights pages are not logged as views. It stores ids, field names and counts, never the contents. Entries past the retention period are never shown or exported, and are deleted at server start and on routine clean-ups.
  • Ownership. Every read and write is scoped to the signed-in owner on the server.
  • The public demo admin is untrusted (DR-005). Anyone can sign in as it, so /admin/records always masks password hashes, e-mail codes, invitation links and the inbox browser key, and shows names, contact details, notes and AI text only for the seeded demo accounts. Guest and registered visitors' rows show ids, timestamps and counts only, and search does not look inside them. Every admin view and export is written to the admin account's activity log.
WhatWhyKept for
Account: user name, password hash, name, occupation, status line, phones, e-mails, optional photoSign-in, and the details other people copy when they add you by user name or QR code.Until you delete the account. Guest sandboxes are deleted after 24 hours.
Contacts and meetings (who, when, where, notes, custom fields, accepted AI summaries)The address book and meeting log you keep - the purpose of the app.Until you delete them or the account. Deleting a contact deletes its meetings.
Activity log (contact and meeting pages viewed; anything created, changed, deleted or exported; sign-ins and password resets; AI actions)Accountability: what was done through your account, and when. List, search and map pages are not logged as views.180 days (older entries are never shown or exported, and are deleted at server start); or with the account.
Administrator accessThe public demo admin sees your rows in /admin/records with names, contact details, notes and AI text masked, and never sees codes, links or password hashes. Each admin view and export is logged under the admin account.Admin log entries follow the activity-log retention above.
AI audit log (redacted text sent to the provider, the answer, model, timing, tokens, your decision)Transparency about every AI call made with your key, and the human decision on it.Until you delete the account. Never contains your API key.
Demo inbox e-mails (codes, invitations)Replaces real e-mail in this demo so sign-up and password reset work.Until the account is deleted.
E-mail codes and invitation linksVerifying an e-mail address, password reset, inviting a contact.Codes 5 minutes, invitations 15 minutes (unconfirmed invitee accounts 16 minutes).
  • No analytics, tracking pixels or third-party cookies.
  • No IP addresses or user agents in the database (rate limits are in memory only).
  • Your AI API key: it stays in your browser (sessionStorage by default) and goes only to the AI provider you chose.
  • Background location: a meeting stores the place you pick on the map, or your current position only when you press "Use my location".

Assumptions and limitations

  • The hosted database was attached late (DR-007). The live demo now runs on Turso, so writes persist and every serverless instance sees them, but the privacy and AI features were built and demonstrated on a local production build while production still used a per-instance copy (DR-004). The functions now run in Sydney (DR-008) but the database is in Tokyo, so every query still crosses to Japan and back.
  • The shared demo account is shared: other visitors see its activity log, AI log and inbox. Guest sandboxes are private (the demo admin sees only masked rows for them) and deleted after 24 hours.
  • Until DR-005, the demo admin could read live password-reset codes and every visitor's notes. That was found in review before this upgrade was merged and fixed by masking; it is recorded rather than quietly removed.
  • E-mail verification is simulated by the demo inbox (DR-002); it does not prove that anyone owns an address.
  • Evaluation corpora are small, synthetic and written by one person; intervals are wide and the redaction results are optimistic.
  • Insights treat weeks as exchangeable and the seed data is synthetic, so the demo charts show the method, not real behaviour.

Decision records

Each record follows the same shape: context, the decision (stated first), the options considered, why, what happened (including the weak numbers) and what I would change. Past records are never edited; a new record supersedes an old one.

What I'd change

  • Provision the hosted database first, fail a production deployment that has none, and choose the function region by an interleaved comparison that includes the database's own region.
  • Detect names of people outside the address book in the browser before sending, and have someone else write a held-out set of notes to evaluate it.
  • Replace the keyword matcher with a pre-registered rubric scored by two people, so the follow-up evaluation credits paraphrases fairly.
  • Add browser end-to-end tests of the main journeys to CI.
  • Expire demo-inbox e-mails after a fixed period, like the activity log.