Skip to content

Experiment analytics

Experiment analytics connects a declared flag allocation to actual feature use and committed business events. Console shows mature and immature cohorts, delivery coverage, binary outcomes, guardrails, and the evidence used for a winner decision. It measures feature consumption or delivery, not human attention.

Signed-in subjects are supported in ordinary Apps and tenant Apps. An ordinary App’s public form can attribute events to an explicitly established visitor. Visitor conversion attribution in tenant public intake or platform triage is not supported: those placement journals do not support the necessary event provenance. Suite startup and source provisioning refuse tenant visitor experiments. Do not substitute a tenant member, service actor, or payload field for a visitor.

Each metric comes from a declared committed command event. Arbitrary SQL changes, generic CRUD without an event, historical backfill, continuous metrics, and identity stitching are outside this contract. Each metric counts at most one conversion per exposed unit.

The App’s experiments entry has the same key as its flag. Declare a version, subject or visitor unit, control variant, fixed UTC start/end timestamps, conversion and delivery-grace windows, eligibility, allocation, primary conversion, and optional binary guardrails. For example:

experiments: {
offerCopy: {
version: 'offer-copy-1', unit: 'subject', control: 'standard',
startsAt: '2027-01-01T00:00:00.000Z',
endsAt: '2027-01-15T00:00:00.000Z',
conversionWindowHours: 48, deliveryGraceHours: 24,
eligibility: { rollout: 100 },
allocation: [
{ value: 'standard', weight: 50 },
{ value: 'clearer', weight: 50 },
],
primary: { conversion: 'applicationSubmitted' },
},
}

Export experimentConversions beside the default defineCommands registry in the server commands module. Reuse the actual defineEvent handle that the command emits; a matching string or payload field is insufficient:

import { defineExperimentConversion } from 'tablewalk/commands';
export const experimentConversions = [
defineExperimentConversion({
id: 'applicationSubmitted', on: applicationSubmittedEvent,
unit: { kind: 'subject', source: 'actor' },
}),
];

For an ordinary public form, use unit: { kind: 'visitor', source: 'public-intake' } and unit: 'visitor' in the App epoch. The command must be public. Its verified visitor provenance is written with the receipt, audit, and event in the business transaction. The public actor and idempotency namespace remain unchanged. An intake without visitor consent emits its business event normally but creates no experimental conversion.

Changing the allocation, variants, eligibility, dates, windows, or metrics requires a new version. Active experiments cannot use targeting rules that force a specific variant. Pause/kill preserves assignments; a fixed winner ends enrollment. Flags and experiments retain their separate scoped role capabilities.

The suite needs shared identity, a flag control database, and a separate dedicated PostgreSQL evidence database owned by its configured role. Neither database may be reused as the business or identity database. Set a dedicated canonical base64url 32-byte assignmentKeyEnv secret; do not reuse identity or placement key material.

{
"experiments": {
"store": "postgres://experiment_owner@localhost/tablewalk_experiments",
"deploymentId": "experiment-evidence",
"assignmentKeyEnv": "EXPERIMENT_ASSIGNMENT_KEY",
"sources": [{
"appId": "lending",
"flagId": "offerCopy",
"tenantId": null,
"sourceId": "business",
"storeId": "app:lending",
"commandIds": ["submitApplication"],
"maxClockSkewMs": 1000
}]
}
}

Each source selects a deployed App connection through sourceId; it does not accept a request-supplied URL. An ordinary App’s event storeId is app:<appId>. A tenant source uses its existing tenant store ID and an explicit tenant ID. Its native service authority rechecks current tenant/platform anchors, independently of the historical user who emitted the event. maxClockSkewMs is an explicit operator bound on emitter clock skew against the source database, from 0 to 120000 milliseconds.

For global-control observations evaluated separately in a tenant, declare that tenant as tenantId and set controlTenantId: null. Such evidence is not pooled into a global result and cannot authorize a global winner from one tenant’s sample. Omitting controlTenantId uses the selected tenant as the control scope.

Provision the business command journal through its normal maintenance workflow, then run these explicit operator commands with the server stopped for SQLite sources:

Terminal window
tablewalk experiments provision --suite suite.json
tablewalk experiments provision-sources --suite suite.json
tablewalk suite suite.json

provision creates the evidence layout. provision-sources registers declared epochs and provisions their event store and dedicated outbox destination. It does not activate flag allocation. Serving opens and attests existing stores only. After adding an epoch or destination, run provision-sources again before serving. A verified version-1 evidence database uses the explicit migration:

Terminal window
tablewalk experiments upgrade-v1-to-v2 --suite suite.json

The migration preserves old observations/evidence. Their missing exposure coverage remains unknown; re-registering an old epoch cannot manufacture complete history.

A flag evaluation alone creates no exposure. Browser components report consumption using short-lived server-issued assignment receipts tied to the current native session or verified visitor proof. Browser requests never choose the unit or arm. Receipt delivery first creates a durable obligation. A confirmed use resolves it; repeated renders and retries count once per epoch/unit. Lost receipts or reports remain visible as pending/unknown coverage, including after process restart.

For an ordinary public surface, establish visitor identity only after an explicit interaction using useExperimentVisitor().start(). A form can use start({ form: 'apply', address: '/apply' }); the address is interpreted by the native public-intake owner, not as a browser-selected tenant. Reads and renders do not mint visitor cookies. Identity is opaque, scoped, and HttpOnly; signing in does not stitch visitor history to the account.

Server command flag reads are recorded only while the admitted handler uses them; preparation and availability checks remain passive. A server command cannot commit untracked experimental use: enrollment failure rolls its transaction back. Once the obligation is durable, a later exposure-ingestion outage may leave pending coverage while business work commits. Conversions use the existing business outbox: a rollback emits none, an ingestion outage preserves committed business work, and a crash after evidence ingestion but before acknowledgement is safe to replay.

The worker serially drains bounded batches and publishes immutable evidence. Changed observations are reflected on the next successful tick; otherwise maturity/coverage is refreshed at most every 30 seconds until complete final results exist. Database outages never turn into zero conversions or a success claim. The source owner examines bounded receipt coverage; undiscovered work, failed delivery, retention loss, in-flight source writers, and unacknowledged exposure obligations block complete evidence. This is a bounded experiment system, not an unbounded analytics pipeline.

Conversions must occur after first exposure and within that unit’s conversion window. Reordered delivery reconciles against bounded pending facts. Estimates use mature cohorts; immature subjects remain separately visible. Finalization requires the end date, conversion window, delivery grace, source watermark, and complete coverage. Statistical evidence retains method/version metadata, allocation diagnostics, multiplicity adjustment, and binary guardrail results.

A winner is an explicit authorized Console operation requiring complete final evidence, acceptable allocation diagnostics, guardrails, and current control and evidence revisions. A decision writes a flag audit entry referencing the immutable evidence revision. Late corrections create new evidence and do not rewrite the old decision. Flag and evidence databases have no cross-database atomic commit: only the flag’s decide audit row proves the winner was applied. On an unknown outcome, read control history before attempting another decision.