# Booking study — y = f(x) on the records we have

**Leads → Booking study** (`/manage/leads/booking-study`, the button in the Leads toolbar). A re-runnable
research page, not a dashboard. Read this before touching it, and before quoting any number from it.

## What it does

The founder's question (2026-09-19): *"y = total booking — covering Loyal, Repeat, Client, Dropped, Booked. We
do closing first; conversion can be considered later. The rest of the parameters are my x. I want a study based
on current data."*

So: **y = a first booking** (whether it later converted or dropped — a booking is what sales controls, and
`bookings.booking_date` is the only purchase event with a trustworthy date, from 2024), **x = what the person
did in a fixed window BEFORE it**, and the page prints, for each x, the booking rate per group with its 95%
range and a p-value, then every x at once in a logistic regression.

It also keeps the record the NEXT study needs: a weekly, append-only snapshot of every lead.

## How it works

`Src\Lead\Study\BookingStudy::run($horizon)` is the whole study. The design is the study; the arithmetic after
it is routine. Five decisions, **each forced by a way the study was wrong first**:

1. **x is never the Leads list's totals.** Those are lifetime figures "as of today", and 96% of buyers were
   imported into the CRM after they bought — their engagement is SPA and loan chasing. Everybody gets the SAME
   window: the first `WINDOW_DAYS` (30) after **day 0 = their first webinar on record**. Anyone who had already
   booked by the end of it is excluded.
2. **Everybody must have been WATCHED for window + horizon.** Otherwise a recent attendee who has not had time
   to book is counted as somebody who did not want to. `HORIZONS` = 180 / 360; **360 is the default because the
   median booking comes 187 days after the first webinar**, and 60 of 119 come after day 180.
3. **Day 0 is the first webinar because nothing else is old enough.** A 360-day horizon needs people first
   seen 390+ days ago — and then the CRM recorded webinars, memberships and funnels, nothing else. WhatsApp
   starts 2025-07, calls 2026-05, the portal 2026-07, AI calls 2026-08 (`BookingStudy::SOURCES` /
   `sourceTimeline()`, printed on the page). **A long y and a rich x do not overlap yet.**
4. ⚠️ **`member_subscriptions.paid_at` is a DATE.** Compared as TEXT against a datetime it sorts before
   midnight of its own day, which silently dropped every membership bought ON the webinar day — 87 of the 151
   real ones — from the first version. Dates are parsed (`Carbon`), never compared as strings.
5. ⚠️ **TWO POPULATIONS, NEVER POOLED — the finding that reshaped the study.** Of 2,520 people first seen before
   mid-2025, 426 had NO lead row at the time; 369 of those were created in the **July 2026 bulk import of
   members and buyers**, with their old webinar attendance attached afterwards. They are in the CRM *because
   they later paid*. `STRATA`:
     - **`known`** — lead row existed by the end of the window. The honest population: **2,088 people, 1
       booking (0.05%)**. Nothing can be concluded from it yet.
     - **`linked_later`** — **362 people, 36 bookings (9.94%)**. Community members who never bought were not
       imported, so its RATES are inflated by an unknown amount; only comparisons WITHIN it mean anything
       (3+ webinars 20% vs 7%; bought a membership 23% vs 6.6%; in the joint model only membership stays clear
       of 1 — OR 1.46, 1.05–2.03).
   Pooled, the study "found" that people who attend more webinars book 13× as often, and that a tracked funnel
   lowers booking forty-fold. What it had found was the import.

**What it may claim:** association, with its size and its uncertainty. **Not cause** — people who come to
three webinars may simply be the ones who already wanted to buy. Only a randomised experiment separates those.
The page says so beside the numbers.

**The statistics are hand-written** ([`Statistics`](/src/Lead/Study/Statistics.php): Wilson interval,
chi-square via the incomplete gamma function, logistic regression by IRLS) because there is no stats package in
this codebase and the page must re-run on production by a button. Every one was checked against scipy /
scikit-learn on the same cohort (chi-square p = 1.752e-15 both ways) and
[StudyStatisticsTest](/tests/Unit/Lead/StudyStatisticsTest.php) pins those values. A chi-square with an expected
cell under 5 is flagged "too few bookings to test" rather than printed as if it were valid; a regression with
fewer than 10 events is not fitted at all.

**Machine learning is deliberately absent.** It answers "can I predict people I have not seen", judged on
LATER people — and the later people here have had no time to book (a forward-in-time test set held 4 bookings).
It belongs on the snapshot data, once each outcome has a few hundred events.

**Access:** `LeadVisibility::LEVEL_ALL` only (company-wide aggregates). Cached 6h per horizon; `?fresh=1`
recomputes. Counts and dates only — no name, phone or message text is read.

### The weekly snapshot — `leads:snapshot-weekly`

Today's tables only say what is true NOW (readiness is overwritten hourly; the list's counts are lifetime
totals). `lead_weekly_snapshots` is the missing half: **one row per lead per week** — stage, y (bookings /
converted / dropped), membership, every cumulative behaviour count the list shows, the seven readiness states,
the reply tier, action items open/done, the account manager and the three Follow Up roles.

- [`WeeklySnapshotBuilder`](/src/Lead/Study/WeeklySnapshotBuilder.php) — ~15 grouped reads for ~12,700 leads
  (12s), **re-using the list's own definitions** (LeadStages' SQL twin, ReplyOwed's tier, Zoom meetings that
  HAPPENED, non-ignored calls, no sandbox / CEO WhatsApp lines): a snapshot that disagreed with the screen on the
  day it was taken would be a second version of the truth.
- ⚠️ **APPEND-ONLY.** Unique `(lead_id, snapshot_on)`, `insertOrIgnore`, no update path anywhere. A second run
  in the same week adds nothing, even if the figures moved — a corrected history is not a history.
- Stamped with the week's **Monday** whichever day it runs; scheduled Mondays 02:30. ⚠️ **A week that is not
  recorded can never be reconstructed** — if the scheduler was down, run it by hand that same week.
- Counts are CUMULATIVE; a window is the difference of two rows. ~660,000 rows a year.
- First snapshot: week of 2026-09-14 (12,692 rows; it matched the list exactly — 480 bookings on 361 leads).
  At ~26 weeks the page should gain a second study: every x, against a booking in the following 60–180 days,
  for the `known` population.

## Related files

- `src/Lead/Study/BookingStudy.php`, `Statistics.php`, `WeeklySnapshotBuilder.php`
- `src/Lead/LeadWeeklySnapshot.php`, `src/Lead/Repositories/LeadWeeklySnapshotRepository.php`
- `app/Console/Commands/SnapshotLeadsWeekly.php` (+ `app/Console/Kernel.php`)
- `app/Http/Controllers/Manage/Leads/BookingStudyController.php`, route `manage.leads.booking-study`
  (a literal segment, declared before `GET {id}`)
- `resources/js/Pages/Manage/Leads/BookingStudy.vue`
- `database/migrations/2026_09_19_230000_create_lead_weekly_snapshots.php`
- `tests/Feature/Manage/Leads/BookingStudyTest.php`, `tests/Unit/Lead/StudyStatisticsTest.php`
