Skip to content

Rating Observation Credibility: A Frontline Rubric

Not all observations carry equal weight. These two rating scales give frontline operators a shared vocabulary for tagging detection credibility before a report enters the analysis pipeline.

How to use these scales

While further analysis into dimensions of credibility could be developed to create a more detailed rating scale for use in scientific investigations [156], the Confronting Unknowns working paper proposes two scales of credibility for immediate use in active frontline operations and decision-making [157]. The first scale combines several aspects of a detection scenario, including the nature of the detection method (sensor, ), their number (single/multiple), and their grade (professional/consumer, trained/untrained). A more fine-grained scale specific to human observers follows.

These scales are one input among many and should not gate inclusion. They function as a shared vocabulary and an initial confidence tag for downstream human or analysis and are not intended to serve as filters for discarding reports.


Scale 1: Detection scenario credibility

This scale combines the nature of the detection method (sensor vs. human), the number of independent detections (single vs. multiple), and the grade of sensor or observer (professional/consumer, trained/untrained).

Higher ratings indicate stronger corroboration. Rating 5 marks the most strongly corroborated detections; rating 1 marks the weakest; and rating 0 marks an observation as an error, artefact, or hoax.

RatingDetection scenarioRelationship
5Multiple sensorsANDMultiple human observers
4Multiple sensorsANDSingle observer
3Single sensorAND/ORMultiple observers
2Single sensor (professional grade)ORSingle trained observer
1Single sensor (consumer grade)ORSingle untrained observer
0Error, artefact, hoax

Table 18. Credibility rating scale combining multiple aspects of a detection scenario.

Rating 5 example: Observation recorded by both radar and , confirmed from both cockpit and ship bridge.

Rating 4 example: Cockpit radar and visual observation made by one pilot only.

Rating 3 example: Radar contact tracked by both cockpit and ATC; or multiple pilots on separate aircraft visually confirm a sighting with no sensor recordings.

Rating 2 example: Radar contact from ATC only, with no cockpit confirmation or visual sighting.

Rating 1 example: Mobile phone footage captured by one civilian passenger on an aircraft.


Human observer challenges

Human observers present unique challenges [158] in the context of anomalous incident detection, including:

  • Vulnerability to perceptual error, memory distortion, and geometry misjudgment, particularly when they are not trained observers.
  • Multiple human observers experiencing different sensory input and processing, resulting in different memories when reporting the same anomalous incident.
  • Intentional deception including hoaxes, illusions of the senses, and fabrications.

These challenges may be mitigated by observational training and may be exacerbated by observing conditions, particularly for untrained observers. This leads to the following scale of credibility for human observations, to aid the analysis of eyewitness reports.


Scale 2: Human observer credibility

RatingHuman observer
5Military trained observer. Example: pilots, radar operators, naval lookouts, air traffic controllers, weapons systems operators
4Civil aviation trained observer. Example: commercial pilots, flight control room personnel
3Trained non-aviation professional. Example: law enforcement, coast guard, meteorologist, air traffic controller acting as visual observer
2Untrained observer, good conditions. Example: general public; daytime, clear skies, unobstructed view
1Untrained observer, poor conditions. Example: general public; night, adverse weather, obstructed or distant view
0Suspected or confirmed hoaxer. Example: known agent or individual with confirmed history of fabrication

Table 19. Types of human observers.


Observer category notes

The four types of human observers typically involved in aerial incidents are listed below, in descending order of observational training.

Military (rating 5). A witness report by military personnel is often made by observers specifically trained to detect and identify classified domestic and foreign military vehicles. However, data accessibility for military witnesses remains restricted due to lack of reporting (stigma or other forms of social pressure) or authorization to speak publicly [159].

Civil Aviation (rating 4). Commercial pilots are trained in air observation, including the ability to distinguish military and civilian craft and atmospheric phenomena. Pilot eyewitness descriptions include anomaly shape, size, and speed (captured in ATC transcripts). Observations of anomalies in civil aviation have been historically underreported due to stigma and professional consequences to pilots who report anomalous observations [160].

Law Enforcement (rating 3). State and local law enforcement officers are trained in general visual observation, investigation procedures, and reporting, including collection of eyewitness testimony on scene immediately after an incident, collection and analysis of nearby security camera footage, and collection of forensic analysis of physical evidence on scene including biological and material specimens. While they may not be specifically trained for air observation, they may be equipped with relevant methods and tools to reliably capture evidence of anomalous incidents.

General Public (ratings 1 and 2). The general population provides an abundance of visual eyewitness reports. Historically, most of these reports were made to civilian reporting centers such as , which has a database of 140,000 incidents dating back to 1969. Individual observations of the general public are the most susceptible to the human observer challenges mentioned above, due to the lack of any observational training, including to account for illusory perception due to poor conditions of observation. However, general public eyewitness reports are occasionally high-quality, or valuable in any case for aggregate data analysis to derive statistics and trends.


Human observers complement sensors

The relevance of human observers should not be underestimated even in contexts concurrently observed by multi-sensor suites. For example, the Artemis II mission, which went around the moon in April 2026, provided unexpected and scientifically valuable eyewitness reports by the crew of micrometeor impacts (flashes) on the far side of the moon [161]. This known, natural phenomenon, which may be critical for future manned surface mission safety, is hard to image with even the capable camera systems available to the crew. And even when potentially within the capability of the imaging systems, it may be hard to get the timing of the observation right if observing such impact flashes is not the main or only purpose of a mission. The astronauts were also asked to provide information on color hues of the lunar surfaces — subtle variations of color are not easily captured by imaging sensors. Artemis II was a mission equipped with state-of-the-art, calibrated sensors based on the mission profile. Yet human observers filled the gap of phenomena unobservable or unobserved by even some of the most capable sensors.

The full scale, with scoring guidance and Appendix I, lives in the working paper.

Read the working paper →