Skip to content

Part II · Detect & Identify

Identifying Unknowns

June 15, 2026 · 13 min read

Abstract

From binaries to probabilities: ranking five mutually exclusive hypotheses (H1–H5) under explicit uncertainty, AI-assisted anomaly resolution, and dual-use detection pathways.

identification has historically been approached as retrospective case resolution: identified or unidentified, resolved or unresolved. That binary model, inherited from (1952–1969) and continued through , is useful for archival reporting but poorly suited to three real-world demands.

FrameworkEraKey categoriesPrimary use
Project 1952–1969Identified, insufficient data, unidentified case sorting
system1972Nocturnal lights, daylight discs, radar-visual, close encounters (1st–3rd kind)Civilian and academic encounter classification
/ UAP Task Force2021Airborne clutter, natural phenomena, USG/industry, foreign adversary, otherIntelligence analytical classification
AARO2022–presentSame as ODNI plus resolved/unresolved labelsU.S. government UAP investigations
Confronting Unknowns Framework (proposed)2026Artifact, natural, unclassified tech, classified tech, unknownOperational decision-making and scientific sensemaking

Table 8. Historical identification frameworks compared.

First, many events develop over time and involve partial, multimodal evidence. A binary label applied at a single point discards the structure of the evidence. Second, operational decision-making cannot wait for definitive attribution. A probability distribution over explanatory hypotheses is more useful than a pending “unresolved” label. Third, cumulative scientific analysis requires structured representations of uncertainty that can be updated as new data, methods, or context become available. A case that resists resolution at first pass should carry a living record of uncertainty rather than a closed file.

The alternative: treat identification as a function. Given a flagged event record, produce a ranked list of explanatory hypotheses with current probability assignments, explicit about the assumptions behind them, the dominant uncertainties, and what additional evidence would most change the assessment. Rather than collapsing to a vote that terminates as "unresolved," the assessment updates across the hypothesis set as evidence arrives—a provisional distribution rather than a verdict.

Anomaly Resolution: From Binaries to Probabilities

proposes five explanatory hypotheses to form a set that is mutually exclusive and, by construction, collectively exhaustive, allocating probability for explanations not yet articulated. The set is designed to be revised; as evidence accumulates, a hypothesis may be split, merged, or added, and probability that once sat in may migrate to a newly named category. This is a feature and is how the framework represents that the space of possible explanations is not ‘mentally foreclosed.’

TypeIconHypothesisDescription
Data or Sensor ArtifactThe apparent event is primarily a product of sensor behavior, processing error, or interpretation failure. Examples: parallax, lens flare, radar
Natural Physical SourceThe event reflects a known, real physical phenomenon from atmospheric, astronomical, oceanic, or geophysical causes. Examples: birds, ball lightning, meteors, mirage
Human-Made Physical Source (unclassified)The event reflects a real object consistent with conventional or commercially accessible technology. Examples: space/air debris, satellite, balloon, crewed and uncrewed aircraft (“drone”)
Human-Made Physical Source (classified)The event reflects a real object associated with restricted state or defense activity. Examples: allied or adversary test program
H5Unknown Physical PhenomenonThe event involves a real physical phenomenon not adequately explained by H1–H4. H5 functions as the open residual: it holds probability for explanations the current hypothesis set does not yet name. Examples: novel natural phenomena, non-human intelligence, residual unknowns

Table 9. CUF explanatory hypotheses.

Three design choices distinguish this hypothesis space. Artifact (H1, ⊞) is an explicit top-level hypothesis, not merely an attribute of a report, because artifact classification is always in tension with the most exotic explanations: for H5 to rank high, H1 must rank very low. An event should be beyond reasonable doubt not a sensor artifact before extraordinary explanations are entertained.

Human-made technology is bisected into unclassified (H3, ☒) and classified (H4, △): military sensor systems disproportionately encounter classified programs: a collection bias with direct legal and operational implications for stable identification.

The space of hypotheses affords prosaic explanations a natural home: a potential UAP later recognized as a civilian aircraft falls squarely into H3, and categorizing known objects and phenomena is as valuable as categorizing unknowns, because it establishes a baseline for what “normal” looks like.

Two qualifications anchor the hypothesis space. First, anomalous and unknown are not the same thing. A phenomenon can be genuinely anomalous yet already understood by someone. A classified program (H4, △) is the clearest case, but the principle is general: knowledge is unevenly distributed, and what is unknown to a frontline operator may be known elsewhere. The operational goal is therefore to move cases toward resolution by whoever holds, or can develop, the relevant knowledge, resisting the treatment of 'unknown' as a permanent property of the event. Second, the boundaries between categories are crisp in principle but can be fuzzy in practice, because the line between 'natural' (H2, ⩯) and 'unknown' (H5, ✶) tracks the provisional state of physical science. A phenomenon classed as natural may later prove novel, or one classed as unknown may resolve to an unusual but conventional process. Reclassification across these boundaries is expected as understanding improves; the framework treats it as the system working, not failing.

Approaching identification with a probabilistic mindset means maintaining a distribution over all five hypotheses that reflects current evidence—not collapsing prematurely to a single answer. The starting distribution is not necessarily uniform: context carries information, and a prior that ignores it does not improve objectivity. In the ideal case, evidence raises one hypothesis decisively: the functional equivalent of a "Resolved" case. In practice, many incidents land in an intermediate state, with two or three hypotheses each holding 20–40% probability. These are the hardest cases to respond to and explain, but the ones that most benefit from honest probabilistic representation, as well as the ones where targeted collection has the highest expected payoff.

As a further example, Figure 7 illustrates our suggested approach, which assigns different probabilities to different causes. A UAP ‘trilemma’ is presented, in which there is a tension of supporting multiple competing explanatory hypotheses: H4 (classified human-made technology), H5 (unknown phenomenon), and even H2 (natural phenomenon) may explain the observation given the data, though none of these hypotheses alone can be assigned with high (e.g., >80%) confidence.

Typical likelihoods across the hypotheses H1 to H5 (illustrative).Figure by the authors of Confronting Unknowns, illustrated by Jonathan Miller.

Functioning as a loop, the CUF ensures that communication failures feed back into detection and identification, and institutional responses are architected in a manner to generate signals more consequential than the original events. Feedback paths matter as much as the forward path: identification feedback refines detection parameters; communication feedback refines the vocabulary and delivery of identification outputs; the accumulated record of cases builds institutional knowledge for all stages.

Five Steps from Data to Ranked Hypotheses

A disciplined identification workflow runs in five steps. Each step is procedural for the operator and auditable for a reviewer; the framework's probabilistic claims rest on the discipline of the earlier steps being done first.

  1. Integrity check. Before any inference, establish what kind of record is in front of you. Verify provenance, calibration, and synchronization; look for the signatures of processing artefacts. The goal at this step is not to confirm the data is anomalous. The goal is to confirm the data is what it claims to be.

  2. Kinematic and signature analysis. With clean data, estimate position, velocity, acceleration, and trajectory, each with explicit uncertainty bounds. Characterize additional observables—compositional, acoustic, electromagnetic, and magnetic signatures or environmental effects. Separate what the sensor measured from what an analyst inferred.

  3. Context correlation. Run the event against the conventional explanations: air traffic, weather, astronomical objects, satellite passes, launch records, and restricted-area operations. Most anomalous reports resolve at this step.

  4. Multi-sensor consistency. When the event survives context correlation, test whether independent sensors and observers agree on the basic facts. Agreement across modalities raises confidence that the signal is physical. Agreement alone does not make the signal anomalous.

  5. Hypothesis scoring. Score the hypotheses against the evidence, record the dominant uncertainties, and route the case: archive, monitor, re-task for additional collection, or escalate for restricted correlation.

To illustrate how context shapes the result: in a wargame applying this workflow to a simulated low-altitude incursion over a civilian area near a military installation, the two human-made explanations (unclassified and classified technology, H3 and H4) together dominate, with classified technology (H4) alone accounting for roughly half the total. Running the identical evidence against a restricted military-airspace scenario shifts the result further toward H4. Nothing the sensors recorded changed between the two runs; only the context (and, therefore, the starting assumptions) differed. Context carries information; ignoring it does not improve objectivity.

A Note on Method: Probability in Emergency Management

Bayesian statistics are a method for reasoning under uncertainty. It starts by stating what is believed before new evidence arrives, then revises that belief as evidence comes in, and rather than forcing a single yes-or-no verdict, it produces a range of possibilities each weighted by how likely it is. For policymakers, the appeal is practical: it requires analysts to put their assumptions in writing, makes any disagreement easy to trace to its source, and lets an assessment to be updated as new facts emerge instead of being scrapped and rebuilt from scratch.

Because of these qualities, the method is established across high-stakes fields where decisions cannot wait for complete data: emergency management situation awareness [21], early outbreak detection [22], wildfire [23] and earthquake forecasting, and intrusion detection in critical infrastructure networks [24], among others [25], [26].

The implications for disaster response are direct: responders rarely have clean or complete data, yet must commit resources before the picture resolves, and Bayesian methods let an assessment begin with whatever is known and sharpen as observations accumulate.

Critically, these same response systems are inherently dual-use and portable to new domains. The sensor-fusion and situation-awareness infrastructure built for emergency response and critical-infrastructure protection, discussed in earlier sections, already performs the core task UAP assessment requires, and would demand little fundamental modification: the analyst still weighs competing explanations, records the reasoning behind each, and updates as evidence accumulates.

The way this paper handles uncertainty is deliberately kept simple. We name a fixed set of explanations, assign each a rough starting probability based on context, and adjust them up or down as evidence comes in. We made this choice for transparency: a decision-maker, a congressional staffer, and an analyst can all look at the same numbers and pinpoint where they disagree, such as the starting assumptions or how well the evidence fits each explanation.

We want to be explicit about what the simple version leaves out. Two limitations matter most. The first: we treat the list of explanations as fixed during any single analysis, even though the framework is meant to be revised between analyses. More advanced methods handle this head-on: allowing the number of explanations to grow as evidence accumulates, holding probability in reserve for explanations no one has named yet and creating a new category only when the data genuinely call for one. Our open residual (H5) is the plain-language version of that idea, a deliberate reserve of probability set aside for the possibility that the current categories are incomplete, but a fuller treatment would let that reserve actually generate new, named explanations rather than leaving it to an analyst's judgment to spot when one is needed. The idea of recursive hypothesis generation has been studied extensively in the statistics literature. See [27], [28] for foundational sources.

The second limit: we commit to a single set of starting assumptions for each context. A different tradition works with a whole range of plausible initial assumptions at once and favors decisions that hold up well across the entire range—often the more honest approach when reasonable people disagree about where to start. We use the single-set approach because policy audiences need one legible picture, not a cloud of competing ones. But the confidence of any single result depends on starting assumptions that other careful analysts could reasonably dispute.

The practical takeaway is modest and important: the framework's outputs are inputs to human judgment, not replacements for it. The numbers are only as good as the explanations we thought to name, our reading of the evidence, and the assumptions we started from, all three of which are revisable. Future work should build the open-ended version of the hypothesis space formally, so that the reserve category does not merely sit as a placeholder but actively proposes new explanations as evidence grows.

Anomaly Analysis: Artificial Intelligence for Unknowns

Recent advances in artificial intelligence create an opportunity to augment expert-driven identification across the resolution workflow described above.

Large Language Models (LLMs) and multimodal models can assist by normalizing heterogeneous reports, extracting structured information from narrative text, correlating records against external databases, surfacing likely artifact modes, triaging cases, and summarizing uncertainty for different audiences. They can also generate a provisional distribution over the H1–H5 hypotheses. Critically, makes the probabilistic framework itself more accessible: Bayesian reasoning has traditionally required specialist statistical expertise, but AI tools can do the analytical heavy lifting of integrating evidence streams, computing likelihoods, and updating estimates, while presenting results as natural-language summaries and ranked hypothesis lists. This affords frontline practitioners to apply rigorous probabilistic reasoning without deep statistical training.

Multimodal LLMs offer a specific capability relevant to UAP analysis: processing visual data (images, plots of physical quantities) using few-shot, in-context learning. Where traditional machine learning requires hundreds of thousands of labeled examples, multimodal LLMs can identify anomalies from a handful of data points, a property well suited to a domain plagued by data scarcity. Over-simplifying: "given these 5 images of confirmed UAP and these 5 images of conventional drones, reason about the next image and tell me if it's likely to be a UAP or a drone, and why." The same few-shot approach is already being adopted in industry (e.g., manufacturing defect detection) and applies to any data that can be rendered as an image, including sensor-data plots. A mere 20-second military infrared video can be decomposed into timestamped multimodal observations: image frames, apparent motion vectors, thermal contrast patterns, and scene metadata. Even 'small' UAP datasets can be enriched along multiple dimensions (including visual, kinematic, thermal, temporal) to partially bridge this data gap.

Multimodal LLMs can identify anomalies from a handful of data points, a property well-suited to a domain plagued by data scarcity.

Testing these Multimodal LLMs for UAP analysis requires no new technology or large datasets: a small set (tens to hundreds) of well-selected, validated examples plus access to state-of-the-art models suffices for early experimentation. However, these systems are vulnerable to false confidence, inconsistent reasoning, hallucination, and non-transparent failure modes, and their inference cost makes real-time or large-scale use impractical today, so current results may be insufficient for operational deployment. Yet the pipelines are portable to better models as the technology evolves and could soon exceed the performance and explainability of traditional methods that require massive training datasets—and which do not yet exist for UAP.

Dual-use pathways using existing and emerging detection infrastructures

Much of the required detection infrastructure already exists; it is just not under that label as such. A mature ecosystem of , environmental monitoring, and industrial surveillance systems is already built to continuously detect, classify, and respond to small, low-altitude, and hard-to-identify aerial targets.

This sensing capability extends beyond the air domain. Offshore and industrial sites increasingly deploy sensor systems that include sonar for underwater detection (e.g. [29], [30]), identifying divers, submersibles, and anomalous objects near critical infrastructure. In maritime settings, combining these acoustic and sonar inputs with radar, optical, and sensing is a natural extension for UAP detection.

In the security sector, modern counter-drone radars (e.g. [31], [32]) explicitly target objects most relevant to UAP reporting: low-flying objects that are hard to detect by radar and move with irregular flight patterns. These systems increasingly read the fine motion signatures of an object’s moving parts using MTI/ techniques to distinguish birds, drones, and helicopters.

Environmental systems offer parallel capabilities. Bird-monitoring radar systems used in wind-farm operations (e.g. [33], [34], [35]) continuously track and classify small airborne objects under noisy, real-world conditions, integrating radar, optical sensors, and real-time analytics to track movement and trigger responses like turbine shutdowns.

Atmospheric platforms sensing such as methane detection and anomaly mapping tools (e.g. [36]) can help explain common misidentifications such as thermal inversions, gas emissions, or optical distortions. More broadly, the field is converging on large-scale sensor networks, drone integration, and AI-driven analysis for observing complex, dynamic systems (e.g. [37], [38]).

, , the U.S. Space Force, and private-sector space companies are already building out the sensing, cataloging, and traffic-coordination infrastructure needed to operate in an increasingly congested near-Earth environment (e.g. [38], [39], [40]). As satellite constellations, launch activity, and orbital debris grow, the technical problem is no longer simply detecting objects, but integrating heterogeneous measurements, maintaining custody, filtering known objects, and sharing safety-relevant information across civil, military, and commercial users. These efforts provide a useful dual-use model for spaceborne anomaly detection: not a dedicated UAP observatory, but a federated architecture that combines public data systems, debris measurements, commercial sensing, and operational space-domain-awareness capabilities.

Across these domains, the bottleneck is integration and standardization, not novel sensing. A credible detection framework for unknowns would emerge from federating existing industrial and environmental systems into an open, multi-modal pipeline that both flags anomalies and rules out known objects under operational conditions.

Try it: prior to posterior

Adjust the evidence likelihoods, or switch the prior context, and watch the posterior shift, the same move the paper makes from New Jersey to Langley.

Prior context

Civilian context: no strong prior reason to favor any hypothesis. Near-uniform priors assign equal weight to all five possibilities before seeing evidence.

H1 Data or sensor artifact
20%
3%
3%
H2 Natural physical source
20%
5%
5%
H3 Human-made, unclassified
20%
36%
36%
H4 Human-made, classified
20%
49%
49%
H5 Unknown phenomenon
20%
7%
7%
Prior (text, read-only)Posterior

Suggested citation

The Confronting Unknowns ’26 Program (2026). Identifying Unknowns. In Confronting Unknowns. Sensemaking. https://sensemaking.wtf/work/cu26-p01-ch06