Skip to content

Worked Example: A Bayesian Case Assessment

Context shifts posteriors. This walkthrough carries two named incidents (the New Jersey December 2024 sightings and a Langley Air Force Base case) through the framework’s full prior-to-posterior sequence, then shows what changes when the analyst’s context changes.

Try it: prior to posterior

Adjust the evidence likelihoods, or switch the prior context, and watch the posterior shift, the same move the paper makes from New Jersey to Langley.

Prior context

Civilian context: no strong prior reason to favor any hypothesis. Near-uniform priors assign equal weight to all five possibilities before seeing evidence.

H1 Data or sensor artifact
20%
3%
3%
H2 Natural physical source
20%
5%
5%
H3 Human-made, unclassified
20%
36%
36%
H4 Human-made, classified
20%
49%
49%
H5 Unknown phenomenon
20%
7%
7%
Prior (text, read-only)Posterior

Every observation is weighed across five hypotheses, not sorted into “explained” or “unexplained.”

  1. Data or Sensor Artifact3%

    The apparent event is primarily a product of sensor behavior, processing error, or interpretation failure.

  2. Natural Physical Source5%

    A known, real physical phenomenon from atmospheric, astronomical, oceanic, or geophysical causes.

  3. Human-Made Physical Source (unclassified)36%

    A real object consistent with conventional or commercially accessible technology.

  4. Human-Made Physical Source (classified)49%

    A real object associated with restricted state or defense activity.

  5. Unknown Physical Phenomenon7%

    A real physical phenomenon not adequately explained by H1–H4: the open residual that holds probability for explanations not yet named.

Worked posterior: New Jersey civilian context

Bayesian reasoning is a method for reasoning under uncertainty. It states plainly what is believed before new evidence arrives, revises that belief as evidence comes in, and produces a range of possibilities each weighted by how likely it is, rather than a single verdict. For policymakers, the appeal is practical: it requires analysts to put their assumptions in writing, makes any disagreement easy to trace to its source, and allows an assessment to be updated as new facts emerge.

We apply here a simple probabilistic model to a toy example. See the following sources for treatments of structured probabilistic reasoning in intelligence analysis [151], its application to anomalous aerial observations [152], forensic evidence evaluation [142], and Bayesian network modeling for risk assessment [153].

Case Profile: Toy Example

The composite first draws on accounts during the New Jersey wave. Note: While this example is inspired by the NJ Wave, it is meant to be an illustrative toy example and not to draw conclusions about the real NJ-centered events. Trained observers reported low-altitude objects moving in coordinated grid patterns over residential areas. The objects were often described to produce “rotor slap,” a heavy rhythmic thumping distinct from the high-pitched whir of consumer drones, characteristic of large rotorcraft displacing air with heavy blades. A resident reported one object hovering over a truck at approximately two meters in diameter. The activity repeated on successive evenings, following consistent flight paths. No operator was identified despite investigation by the State Police, , and .

The evidence profile: large-scale aerial objects at low altitude; coordinated grid-pattern flight; acoustic signature consistent with heavy rotorcraft; repeated activity over multiple nights; no identified operator despite federal investigation.

We also include an additional hypothetical observation of rapid, anomalous acceleration of a ‘drone’ observed near a military base. This was not a something widely reported in the NJ wave itself but is commonly associated with sightings.

Key terms

  • Prior: what we believe before looking at this case's evidence.

  • Relative : for one piece of evidence, how well a given hypothesis predicts it, on a simple scale where 1.0 means the evidence tells us nothing about that hypothesis. A value above 1 means the evidence fits that hypothesis; below 1 means it argues against it.

  • Combined weight: a single running score for each hypothesis, found by multiplying together the relative likelihoods of all the evidence.

  • Posterior: the updated belief after the evidence, expressed as a percentage. The posteriors are the numbers a decision-maker actually reads.

Case analysis

We group the observations into a small number of independent evidence items. This grouping matters. Several widely reported facts, that the issued flight restrictions, that nothing was shot down, that the government did not attribute the activity, are not three independent clues; they are three faces of one underlying fact (a cautious, non-attributing official response). Counting them separately would multiply the same signal three times and artificially sharpen the result, so we treat them as one item.

We use three evidence items in total:

  • E1: the civilian observation profile. What witnesses saw and heard: large objects, low altitude, coordinated grid flight, rotor-slap acoustics, repeated over several nights.

  • E2: the investigation and attribution outcome. What the official response produced: flight restrictions, no object shot down, and no operator identified despite federal investigation.

  • E3: a single, uncorroborated anomalous- report (introduced only in Stage 2): one military sensor logging rapid, propulsion-less acceleration, with no second sensor confirming it.

The numbers below are one analyst's assignment. Different analysts would assign different values, and locating where they differ (a prior? a likelihood?) is exactly what this method is for.

How to read the relative likelihood weights

  • ≈ 1.0: evidence is uninformative for this hypothesis

  • 0.3–0.8: mild evidence against

  • 1.5–3: evidence for

  • ~4: strong evidence for

We deliberately cap the weights around 4: no single observation in this toy case is treated as decisive, a constraint whose importance becomes clear in Stage 2. One note on the uniform prior: starting every hypothesis at 20% is a teaching convenience, not a claim of true ignorance. A flat prior treats "unknown phenomenon" (, ✶) as a priori exactly as plausible as "sensor artifact" (, ⊞), which real-world base rates do not support. A genuine assessment would set context-dependent priors: the same evidence over an active military base, for instance, would start with far more weight on classified activity (, △).

Stage 1: the reported evidence

We first weigh the two real-world evidence items, the civilian observation profile (E1) and the investigation-and-attribution outcome (E2), to see where the conventional record alone points, before testing what a single anomalous report would do to that picture.

HypothesisE1: civilian observationE2: investigation outcomeCombined weight (E1 × E2)Posterior
⊞ H1 Sensor artifact0.31.00.303%
Natural source0.31.00.303%
Human-made (unclassified)3.00.82.4028%
△ H4 Human-made (classified)2.52.05.0058%
✶ H5 Unknown phenomenon0.61.00.607%

Table 16. Stage 1: weights and posterior for the civilian observation (E1) and investigation outcome (E2), from a uniform prior.

Following the arithmetic. Take H4: its combined weight is E1 × E2 = 2.5 × 2.0 = 5.00. Add up all five combined weights, 0.30 + 0.30

  • 2.40 + 5.00 + 0.60 = 8.60, and divide to get each posterior. For H4: 5.00 ÷ 8.60 = 0.58, i.e. 58%. For H1: 0.30 ÷ 8.60 = 0.035, i.e. about 3%. (Percentages are rounded and may not total 100%.)

Why these weights? The structured, repeating, rotorcraft-like signature is poorly explained by a sensor artifact or a natural source (both well below 1 on E1), and is only mild-evidence against an exotic unknown: a genuinely novel phenomenon need not sound like a helicopter. It fits human-made activity well, and the unattributed-despite-investigation outcome (E2) fits classified activity. The result concentrates roughly 86% of the probability on authorized human activity (H3 + H4), with the artifact, natural, and unknown hypotheses each small but never driven to zero.

Stage 2: adding one uncorroborated anomalous report

Stage 1's combined weights now become the starting point: carrying a previous result forward as the new prior is Bayesian updating. We multiply each by the relative likelihood for the new evidence item E3.

HypothesisWeight carried from Stage 1E3: single anomalous sensor observationNew weightPosterior
⊞ H1 Sensor artifact0.304.01.2012%
⩯ H2 Natural source0.300.50.152%
☒ H3 Human-made (unclassified)2.400.51.2012%
△ H4 Human-made (classified)5.001.05.0050%
✶ H5 Unknown phenomenon0.604.02.4024%

Table 17. Stage 2: the Stage 1 weights updated by one uncorroborated anomalous report (E3).

Following the arithmetic. Take H5: its new weight is the carried-forward 0.60 × 4.0 = 2.40. The five new weights total 1.20

  • 0.15 + 1.20 + 5.00 + 2.40 = 9.95, so H5's posterior is 2.40 ÷ 9.95 = 24%, up from 7%.

The instructive result is what does not happen: the unknown hypothesis does not take over. A single sensor reporting impossible kinematics is predicted equally well by a sensor artifact (H1) and by a genuine unknown (H5): one instrument alone cannot tell them apart. The distribution becomes more uncertain, not more exotic, and classified human technology remains the single leading explanation.

This is the appendix's main lesson, and it echoes the sensor shortcut argument (Figure 5): for the unknown hypothesis (H5) to rise above the artifact hypothesis (H1), one needs corroborating evidence that specifically rules the artifact out: independent multi-sensor confirmation. A dramatic single report does not resolve a case toward "unknown"; it widens the range of live possibilities and shows exactly what would narrow them again: independent corroboration from a second sensor.

The full derivation (with Appendix G and complete evidence matrices) lives in the working paper.

Read the working paper →