# What entreprenoid collects

> ⚠️ **Generated from `packages/event-schema` by `pnpm fields:build`. Do not edit by hand** --
> `collected-fields.test.ts` fails if this file and the schema disagree. The specification
> (§18.4) requires documentation that accurately enumerates collected fields, and generating
> it is the only way that stays true.

There are **two** channels, and both are listed here. The first is what crosses from a
customer's SERVER to our ingest; the second is what the optional page tag sends from a
VISITOR's BROWSER. Nothing else is sent by either.

## Which requests are stored

A request is stored when a known AI client makes it (every such request, whatever it asked for),
when it asks for markdown, when it fetches a discovery file such as /llms.txt, or when it is
served an HTML page. Everything else -- scripts, form posts, redirects, missing pages, assets,
and the site's own admin and scheduled traffic -- is counted by reason and not stored.

A request addressed to a server's bare public IP address, rather than to the site by name, is
counted by reason and not stored, whichever client made it: visitors ask for a site by its name,
and scanners walking a server ask for the machine.

A request addressed to a local name such as localhost, at a site whose traffic otherwise arrives
under its own domain, is counted by reason and not stored, whichever client made it: it was made
to the server directly, not to the site. A site whose server never sees its own name, behind a
proxy that does not pass it on, keeps these requests.

A batch may also carry `notRecorded`: a count per reason of the requests the SDK saw and did not
send. Counts only -- no path, address, header or time finer than the hour it arrives in.

## The 39 fields a server sends

| Field | Type | Always sent | What it is, and why we need it |
|---|---|---|---|
| `eventId` | `string` | yes | A unique id the SDK mints for this observation. It is what makes a safe retry idempotent -- the same event sent twice is stored once. |
| `siteId` | `string` | yes | Which site the event belongs to. ⚠️ On arrival this is OVERWRITTEN with the site the credential belongs to; a value in the body is never trusted, so one customer's key cannot write into another's data. |
| `observedAt` | `string` | yes | When the customer's server saw the request. Displayed, never sorted on: it is client-supplied and therefore clock-skewed. Ordering uses our own receive time. |
| `requestId` | `string` | no | Correlates two observations of one request. The Next.js proxy runs before the response exists, so it emits what it knows and carries this id. ⚠️ NOTHING MERGES IT YET: the next-proxy adapter emits it, and no consumer reads it. Next offers no general post-response hook, so the second observation needs a per-route wrapper that does not exist. The field is here so that merge can arrive without a schema change or a customer upgrade -- it is not evidence that it happened. |
| `pathRedacted` | `boolean` | no | True when at least one path or route segment was replaced with [redacted] before sending, by a customer pattern or by the SDK's default detector. Shown to the operator so a redacted path reads as a setting rather than a bug. |
| `method` | `string` | yes | The HTTP method. |
| `path` | `string` | yes | The request path, normalised and WITHOUT its query string. Segments matching the site's configured secret patterns are redacted before the event is created. |
| `route` | `string` | no | The route template, such as `/users/:id`, when the customer configured templating. Collapsing identifiers out of paths is what keeps per-user values out of analytics and keeps cardinality finite. |
| `host` | `string` | no | The `Host` header, for sites serving several domains. |
| `protocol` | `"http" \| "https"` | no | The request scheme. |
| `userAgent` | `string` | no | The `User-Agent` header, truncated at 512 characters. This is the primary evidence for classifying a client, and the dashboard shows the exact string a rule matched -- an explanation a site owner cannot check is not an explanation. A User-Agent is a claim, never proof. |
| `referrerOrigin` | `string` | no | The ORIGIN of the referrer -- scheme and host only, never the full URL, so the page someone came from is not collected. Used to attribute a visit to an AI assistant that sent it. |
| `campaign` | `object` | no | Allowlisted campaign parameters only. Arbitrary query-string values are never collected. |
| `campaign.source` | `string` | no | The `utm_source` value, when present. Needed to attribute a human visit to an AI assistant that tagged the link. Only the allowlisted campaign parameters are ever read; no other query value is collected. |
| `campaign.medium` | `string` | no | The `utm_medium` value. Distinguishes a referral from a paid placement. |
| `campaign.campaign` | `string` | no | The `utm_campaign` value. Groups visits belonging to one campaign. |
| `campaign.content` | `string` | no | The `utm_content` value. Distinguishes variants within one campaign. |
| `clickIdType` | `"gclid" \| "fbclid" \| "msclkid"` | no | WHICH advertising click-id was present, never its value. The value identifies a person; its presence identifies a traffic source. |
| `response` | `object` | no | Facts about the completed response. Absent when the adapter could not observe one. |
| `response.status` | `number` | no | The HTTP status code the customer's server returned. |
| `response.contentType` | `string` | no | The response `Content-Type` header. This is what tells a site owner whether an AI client received HTML where markdown would have served it better -- the measurement the product exists to make. |
| `response.contentLength` | `number` | no | The response `Content-Length` header, when the server set one. Read from the header only; response bytes are never counted or buffered, because counting means wrapping the stream and being on the critical path. |
| `response.latencyMs` | `number` | no | How long the customer's server took to produce the response, in milliseconds. |
| `response.observation` | `"measured" \| "inferred" \| "unknown"` | yes | How the response facts were obtained. `measured` means we read the completed response; `inferred` means we derived it from something weaker, such as a path suffix; `unknown` means we looked and could not tell. The `response` object being ABSENT is a fourth, different fact: we never had a chance to look, which is the Next.js proxy case. |
| `serve` | `object` | no | The markdown-twin decision for this request. Present only when the serve half is configured. |
| `serve.decision` | `"served" \| "advertised" \| "fell_through" \| "error"` | yes | What the serve half did. `served` means a markdown twin was returned; `advertised` means the HTML went out carrying a `rel=alternate` link to the twin; `fell_through` means the request was not ours to answer; `error` means something failed and the customer's own handler ran unchanged. |
| `serve.reason` | `string` | no | Why the decision went the way it did. `no_signal` is the ordinary case: the client did not ask for markdown, so it received the page. ⚠️ `no_twin` means a twin was looked for and NOT found; it must never appear beside `served` or `advertised`, which both prove one was found. Values outside `KNOWN_SERVE_REASONS` are ACCEPTED rather than rejected: an SDK in a customer's process outlives any list this schema can hold, and rejecting one unknown string would throw away every event in the batch beside it. |
| `serve.format` | `string` | no | The content type actually served, when we served one. |
| `network` | `object` | no | Network facts about the end client, and which setting produced them. |
| `network.ip` | `string` | no | The END CLIENT's IP address, as the customer's server saw it. ⚠️ It must come from the SDK: our ingest's socket peer is the customer's SERVER, not their visitor, so unlike a browser beacon we cannot observe it ourselves. The address is RETAINED on the stored request and is deleted with it, on the site's own retention schedule -- except for a browser we classify as a person, whose address is discarded as soon as its network type has been recorded, normally as the request is stored. It is also used at the ingest boundary for coarse country, rate limiting and crawler verification, and a separate 24-hour hold exists for that purpose. It is also matched, on our own servers, against public network data to record what kind of network it belongs to; no third-party service is consulted. Each request additionally carries a site-scoped, daily-rotating HMAC pseudonym and a two-letter country code, which are what the aggregates group on. |
| `network.countryCode` | `string` | no | An ISO-3166-1 alpha-2 country code, when the customer's platform already resolved one (Cloudflare and Vercel both do). Saves us deriving it, and is coarse by construction. |
| `network.ipSource` | `"platform" \| "forwarded" \| "off" \| "custom"` | no | WHICH SETTING produced the address, or produced none: the operator's `clientIp` choice, never the header that answered. ⚠️ It is the difference between an install that was never finished and a customer who chose to send nothing, and without it the only thing a dashboard can tell the second one is that their install is broken. `off` is an answer; `platform` with no `ip` beside it is a self-hosted install behind a proxy that writes no platform header, where crawler verification can never run. An ABSENT field is the fourth state and means an SDK older than this one -- it is never read as `off`. |
| `internal` | `boolean` | no | Whether the customer's own filters, or our own asset path, marked this as internal traffic. Since Phase 55 an SDK does not send internal traffic at all -- it is counted by reason in the batch's `notRecorded` -- UNLESS an AI client made the request, in which case the event is sent with this flag set and is excluded from reporting and never billed. Older SDK versions still send it; ingest drops and counts those, whether or not the adapter observed a response. |
| `sdk` | `object` | yes | Which SDK produced this event. |
| `sdk.name` | `string` | yes | The SDK package name. |
| `sdk.version` | `string` | yes | The SDK version, so an event can be traced to a release. |
| `sdk.adapter` | `string` | yes | Which integration produced the event -- `express`, `web`, `next-proxy` and so on. Different adapters can observe different things, and the dashboard must not present an adapter's blind spot as a fact about the traffic. |
| `sdk.runtime` | `string` | yes | The JavaScript runtime and version the SDK is running on. |
| `sdk.dropped` | `number` | no | How many events this SDK instance discarded since the last successful send, because its bounded buffer filled during an outage. Reported so the dashboard can say the count is short rather than silently under-reporting. |

## What is never collected

The wire schema rejects unknown properties at every level, so these cannot be sent even by a
modified SDK -- the ingest boundary refuses the event rather than storing what it does not
recognise. `event.test.ts` asserts each one is rejected.

- `body`
- `requestBody`
- `responseBody`
- `authorization`
- `cookie`
- `cookies`
- `headers`
- `query`
- `queryString`
- `search`
- `fragment`
- `hash`
- `formData`
- `console`
- `email`
- `name`
- `userId`
- `username`
- `password`
- `token`
- `apiKey`
- `serverKey`

## How an IP address is handled

⚠️ **The end client's IP must be sent by the SDK.** Our ingest's socket peer is the customer's
*server*, not their visitor, so unlike a browser beacon we cannot observe the visitor's address
ourselves.

⚠️ **The address is STORED on the request, as of 2026-09-17.** Until that date this page
promised the opposite -- that it survived only as a pseudonym -- and a promise a product has
stopped keeping is worse than one it never made.

The address is RETAINED on the stored request and is deleted with it, on the site's own retention
schedule -- except for a browser we classify as a person, whose address is discarded as soon as
its network type has been recorded, normally as the request is stored. It is also used at the
ingest boundary for coarse country, rate limiting and crawler verification, and a separate 24-hour
hold exists for that purpose. It is also matched, on our own servers, against public network data
to record what kind of network it belongs to; no third-party service is consulted. Each request
additionally carries a site-scoped, daily-rotating HMAC pseudonym and a two-letter country code,
which are what the aggregates group on.

A request from a browser we classify as a person is kept as its own record for seven days, then
folded into hourly counts that keep no address, user-agent, visitor pseudonym or full path, and
the record deleted -- unless an AI assistant referred that visit, or the record predates referral
tracking, in which case it is kept for the site's retention window. The counts are deleted on the
same schedule.

## What the page tag sends from a browser

The page tag is **optional and off until a site enables it**. No metric depends on it; its
purpose is generating markdown twins from pages a server-side fetch cannot read.

On each page view it sends only a **content hash**, a path, and the site's public key,
and is told whether that content is already held. It sends the content itself only when
the answer is no -- in practice, once per version of a page.

**Each of those questions is kept as a vote** (Phase 129): which version of which path a
visitor rendered, because agreement across visitors is what makes content usable. The
visitor's address waits in a queue only until the worker turns it into the same site-scoped,
daily-rotating pseudonyms an event carries -- normally seconds, and a queued row older than a
day is removed unconverted. The vote itself is kept seven days.

| Field | What it is |
|---|---|
| `k` | The site's public key. Published by design; it cannot write events. |
| `path` | `location.pathname` only. The query string and fragment are never read. |
| `contentHash` | SHA-256 of the pruned content. |
| `title` | The page title, already redacted in the browser. |
| `textLength` | How many characters of visible text the pruned content held. |
| `html` | The pruned content subtree. Converted to markdown on receipt; the HTML is deleted in the same transaction. |

**Removed in the browser, before anything is sent:** every form value, anything marked
`data-entreprenoid-private`, hidden and `aria-hidden` subtrees, scripts and styles, every
attribute except `href`, `src`, `alt` and `title`, and text matching email, card and
long-digit patterns. The tag reads no cookie and writes no storage of any kind.

⚠️ **Stripping is not the guarantee.** Harvested content is used only when several DISTINCT
anonymous visitors produced the same content, or a cookie-less fetch of your verified domain
reproduced it -- and it is SERVED only in the second case. A page that differs per visitor --
personalised or behind a login -- does neither, so it is never used or served, and its copies
are deleted after seven days. A page the
site marks `noindex` is never harvested at all.
