# Controlled pilot checklist

Ayonix Video Analytics · Version 1.0 · Reviewed 11 September 2026
Author: Gabriel Bamola, Chief Marketing Officer · Technical review: Dr Sadi Vural, Founder and CEO

---

## The purpose of a pilot

A pilot should end in a decision, not a debate about whose numbers to believe.
That requires acceptance criteria agreed **before** it starts, and ground truth
that does not come from the system being evaluated.

## Phase 1 — Camera readiness review

Before any rule is configured.

- [ ] Every camera in scope assessed against the camera planning guide
- [ ] Pixel-on-target measured, not estimated, at the furthest point of interest
- [ ] Outcome recorded per camera: pass, re-aim or unsuitable
- [ ] Remediation for re-aim cameras completed and re-verified
- [ ] Unsuitable cameras removed from scope, and the decision documented
- [ ] Clock synchronisation verified across cameras, analytics and VMS

## Phase 2 — Baseline (log only)

**No alerting is enabled during this phase.** This is the most frequently skipped
step and the most frequently regretted.

- [ ] Rules configured in log-only mode
- [ ] Minimum two weeks, covering weekday and weekend patterns
- [ ] At least one period of poor weather, for outdoor cameras
- [ ] Baseline event volume recorded per camera per hour
- [ ] Current nuisance-alarm rate on the existing system recorded, for comparison
- [ ] Stream availability recorded per camera

Without a baseline you cannot demonstrate improvement, only assert it.

## Phase 3 — Ground truth

Ground truth must be independent of the system being measured. The system's own
confidence scores are not ground truth.

- [ ] **Scripted tests** at defined positions, times and conditions
- [ ] **Manual counts** for any counting metric, at quiet, normal and peak periods
- [ ] **Stopwatch timing** for any duration metric
- [ ] Tests repeated in the worst realistic condition, not only the best
- [ ] Scripted edge cases: group entry abreast, occlusion, a person stepping aside
- [ ] Who collects ground truth, and when, agreed in writing

## Phase 4 — Tuning

- [ ] Confidence thresholds set **per camera**, from observed conditions
- [ ] Persistence windows set from observed nuisance patterns
- [ ] Schedules configured to match actual operating hours
- [ ] Zone geometry verified against the physical space
- [ ] Cooldowns set so a sustained presence does not flood the queue
- [ ] Every change logged with its rationale

Note the trade-off explicitly: tuning too aggressively converts false alerts into
missed events, which is usually the worse failure.

## Phase 5 — Measurement

Measure against the criteria agreed at the start, not criteria chosen afterwards
to fit the result.

| Metric                             | How measured                              | Target agreed? |
| ---------------------------------- | ----------------------------------------- | -------------- |
| Event recall                       | Scripted ground truth                     |                |
| False alerts per camera-hour       | Operator-marked, over a continuous window |                |
| Alert latency                      | Event time to operator queue              |                |
| Duplicate alerts per genuine event | Counted from the event log                |                |
| Stream availability                | Percentage over the pilot period          |                |
| Operator handling time             | Sampled                                   |                |
| Counting error                     | Against manual counts                     |                |
| System resource usage              | Against deployed hardware                 |                |

## Phase 6 — Decision

- [ ] Results compared against the agreed acceptance criteria
- [ ] Cameras that failed readiness listed, with remediation cost
- [ ] Residual false-alert rate stated plainly, not averaged away
- [ ] Decision recorded: expand, adjust or stop
- [ ] If expanding: which sites, in what order, and what changes first

## Setting acceptance criteria

There are no universal numbers, and any vendor offering them should be asked how
they were measured and on what cameras.

Reasonable criteria are:

- **Specific to a camera group** rather than site-wide
- **Expressed with their conditions** ("in dry conditions above 50 lux")
- **Paired**: recall and false-alert rate together, never one alone
- **Compared to the existing baseline**, not to an absolute ideal
- **Agreed by the person who will operate the system**, not only by the buyer

## Common pilot failures

1. No baseline, so improvement cannot be demonstrated.
2. Ground truth taken from the system's own output.
3. Testing only in good conditions.
4. Acceptance criteria set after seeing the results.
5. Cameras that failed readiness left in scope anyway.
6. Alerting enabled on day one, so operators lose trust before tuning.
7. No one nominated to make the decision at the end.

---

Questions: infojp@ayonix.com · https://videoanalytics.ayonix.com
