The Most Difficult Problem in Forensic Validation: How Do You Validate a Method When the True Answer Is Unknown?

Budding Forensic Expert
0
Budding Forensic Expert · Forensic Methodology & Validation Science

The Most Difficult Problem in Forensic Validation: How Do You Validate a Method When the True Answer Is Unknown?

Ground Truth, Ecological Validity and the Limits of Forensic Method Validation — A Critical Review

Abstract

In experimental science, a method is validated by comparing its output against a known answer. In forensic casework, the event that produced the evidence is very often the fact the investigation exists to discover. This creates a structural tension: forensic methods are validated under conditions where the truth is deliberately constructed or independently known, then applied to casework where the truth is precisely what remains unresolved. This review examines that tension — the “ground truth problem” — across forensic science, asking whether performance demonstrated under known-truth experimental conditions can be legitimately generalised to heterogeneous, degraded, and often unknowable casework.

The review distinguishes several senses in which “truth” functions in forensic science — directly known experimental truth, independently verified reference truth, corroborated operational truth, and partially observable historical truth — and argues these are not interchangeable, even though the literature and courtroom discourse frequently treat “ground truth” as a single, portable concept. It examines the major empirical vehicles used to approximate ground truth in the absence of a real answer key: black-box studies of latent print and firearms examiners [Ulery et al. 2011; Monson et al. 2022], proficiency and collaborative testing regimes, probabilistic-genotyping validation using DNA mixtures of known composition, and synthetic or reference datasets in digital forensics. In each case it asks whether internally rigorous validation designs are ecologically representative of the casework they are meant to stand in for, drawing on the ecological-validity literature in cognitive forensic science [Growns & Kukucka, 2021] and the methodological critiques of black-box study design [Khan & Carriquiry, 2023; Dorfman & Vanderplas, 2024].

The review devotes particular attention to why examiner agreement — whether inter-examiner concordance, laboratory verification, or peer review — is not equivalent to accuracy, a distinction long recognised in the discipline but easily collapsed in practice. It examines cognitive bias as a hazard that becomes acute precisely where no independent answer exists to check a conclusion against [Dror, Charlton & Péron, 2006; Kukucka et al., 2017], and traces the same ground-truth tension into the emerging problem of training and evaluating machine-learning systems on imperfectly labelled forensic data, particularly in deepfake and digital-image forensics.

Drawing the literature together, the review proposes — as a discussion framework rather than an adopted standard — a five-level “Ground Truth Continuum in Forensic Validation,” running from directly known experimental truth through to competing, unresolved reconstruction hypotheses, and uses it to argue that forensic disciplines require different validation architectures rather than a single universal validation template. The review closes by arguing that the central scientific question is not whether an examiner can reach the correct answer when the answer is already known, but how much confidence can be justified when the truth is precisely what the investigation is trying to discover.


1. Introduction: The Answer-Key Paradox of Forensic Science

THE ANSWER-KEY PARADOXEXPERIMENTAL VALIDATION Known Input Examination Comparison WithKnown Answer Error RateEstablishedvs.FORENSIC CASEWORK Unknown HistoricalEvent Fragmentary,Degraded Traces Examination Inference About theUnknown EventThe same validation logic cannot close the loop: casework has no independently known answer to compare against.
Figure 1. The Answer-Key Paradox: validation logic works only when a comparison against a known answer is possible.

In experimental science, validation depends upon comparison with a known truth. A thermometer is calibrated against a reference temperature; a diagnostic test is validated against confirmed disease status; an algorithm is benchmarked against a labelled dataset whose labels are, by construction, correct. Forensic science borrows this validation logic wholesale — proficiency tests, black-box studies and software validation exercises are all, at bottom, exercises in comparing a method’s output against a known answer.

But the events that produce forensic evidence are frequently the very facts an investigation exists to discover. Nobody knows, independent of the forensic examination itself, whether the latent print at the scene was left by the suspect; that is the question being asked. The system therefore faces a structural asymmetry: forensic methods are validated under conditions where truth is deliberately constructed or independently knowable, and then applied to casework where truth is inaccessible by exactly the same evidentiary route the method is meant to illuminate.

This is not a rhetorical flourish. It has produced, and continues to produce, some of the sharpest disputes in modern forensic science — from the 2009 National Research Council report’s finding that “with the exception of nuclear DNA analysis, … no forensic method has been rigorously shown to have the capacity to consistently, and with a high degree of certainty, demonstrate a connection between evidence and a specific individual or source” [National Research Council, 2009], to the 2016 report of the President’s Council of Advisors on Science and Technology (PCAST), which held that only single-source DNA, simple DNA mixtures, and latent fingerprint analysis met its bar for “foundational validity” among the feature-comparison disciplines it reviewed [PCAST, 2016]. Both documents are, at their core, arguments about whether existing validation evidence for forensic methods is representative of a genuine, independently established ground truth — and whether it generalises to the conditions of real casework.

This review does not attempt to relitigate PCAST or the 2009 NRC report. It uses them, and the extensive empirical and methodological literature that followed them, as the entry point into a narrower and more precisely defined question: what does “ground truth” mean in forensic science, why is it harder to obtain than the term suggests, and what can legitimately be claimed about a method’s reliability when direct ground truth is unavailable in the cases that matter most.


2. Defining Ground Truth: A Concept More Complex Than It Appears

SEVEN SENSES OF “GROUND TRUTH”Not interchangeable — each carries a different validation strength Directly KnownExperimentalTruthStrongest IndependentlyVerifiedReferenceTruthStrong ExperimentallyConstructedTruthStrong (built) CorroboratedOperationalTruthModerate ConsensusTruthModerate–weak Legal /HistoricalTruthWeak(procedural) Latent /UnobservableTruthNot directlytestable
Figure 2. Seven senses of “ground truth” used across forensic science, ordered by relative validation strength.

“Ground truth” is used in forensic science as though it names a single, stable thing. It does not. At minimum, the literature and practice distinguish:

  • Directly known experimental truth — the examiner is shown a print, a bullet, or a document sample whose source is fixed by the experimental design itself (e.g., the same finger printed twice under controlled conditions).
  • Independently verified reference truth — truth established by a separate, higher-certainty method (DNA identification of a body used to verify a forensic anthropology age estimate, for instance).
  • Experimentally constructed truth — mixtures, simulated casework items, or synthetic datasets deliberately built with known composition, as in probabilistic-genotyping validation using DNA samples of known contributor number and ratio [Bright et al., 2018; Moretti et al., 2021].
  • Consensus truth — the conclusion reached by agreement among multiple examiners or laboratories, frequently used as a practical stand-in for truth in proficiency testing when no independent answer exists.
  • Operational truth — a working determination accepted for casework purposes (a “known” reference sample obtained from a suspect, a database hit), which may itself later prove incorrect.
  • Legal truth — the resolution reached by a court or plea, which closes a case procedurally without necessarily closing it epistemically.
  • Historical truth — the actual sequence of events that produced the evidence, which in reconstruction and activity-level questions may never be independently accessible at all.

Treating these as interchangeable is a common but consequential error. A DNA single-source match validated against directly known experimental truth carries a different evidentiary architecture than a bloodstain-pattern reconstruction resting on historical truth that can never be independently replayed. The disciplines most often criticised for weak “foundational validity” — pattern and impression evidence, questioned documents, some strands of digital forensics — are disciplines where the operative form of truth shifts, often silently, from directly known or reference truth in validation studies to consensus or operational truth in casework. The 2009 NRC report’s central complaint was in part a complaint about this substitution: many forensic disciplines had “never been exposed to stringent scientific scrutiny” of the kind that establishes directly known or independently verified truth as their empirical foundation [National Research Council, 2009].


3. Experimental Truth Versus Casework Reality

The canonical validation logic runs: known input → examination → comparison with known answer. Casework runs the inverse: unknown historical event → fragmentary, degraded traces → examination → inference about the unknown event. The difference is not merely rhetorical. In controlled studies, researchers can guarantee ground truth by construction — deciding, before any examiner sees the item, which comparisons are same-source and which are different-source [Ulery et al., 2011; Monson, Smith & Peters, 2022]. In casework, no such guarantee exists; the “answer” the examiner is working toward is the thing under dispute.

Controlled research addresses this gap using test sets, simulated casework, mock evidence, synthetic datasets, and known-source or known-non-match sample pairs. Each buys certainty about truth at the cost of some fidelity to casework conditions — item difficulty, evidence degradation, contextual information, laboratory workflow, and the psychological framing of “this is a real case” versus “this is a test.” The question this review returns to repeatedly is not whether such studies are useful (they plainly are, and constitute the best empirical evidence the disciplines currently possess) but how far their internally guaranteed truth licenses claims about performance on casework whose truth is, by definition, not similarly guaranteed.


4. Why Validation Depends on Ground Truth — and Casework Often Has None

Method validation is, formally, the process of establishing that a test’s output tracks the true state of the world with a known and acceptable error rate. Every element of that sentence — “true state,” “known,” “acceptable” — presupposes an answer key. Firearms and toolmark black-box studies illustrate the point cleanly: the entire evidentiary value of the Baldwin et al. Ames Laboratory study, and its successors, rests on the researchers’ ability to specify in advance, from the manufacturing and test-firing process itself, which cartridge cases and bullets were fired from the same firearm and which were not [Baldwin et al., 2014, cited in Monson, Smith & Peters, 2022; Khan & Carriquiry, 2023]. Strip away that guaranteed ground truth and the resulting false-positive and false-negative rates — 0.656% and 2.87% for bullets in the largest such study to date [Monson, Smith & Peters, 2022] — would be uninterpretable numbers rather than error-rate estimates.

Casework provides no equivalent guarantee. A DNA reference sample from a named suspect is not “ground truth” that the suspect committed an act; it is ground truth only for the narrower, source-level question of whose DNA profile is being compared. As soon as the question moves from whose DNA is this to how did it get there and when — an activity-level proposition — the guaranteed truth of the reference sample no longer resolves the disputed issue, and no experimental design can manufacture ground truth for the historical event of transfer in a way that mirrors the uncertainty of a real crime scene [Cook et al., 1998; Evett et al., 2002; Gill et al., 2018].


5. Ecological Validity and the Representativeness Problem

Ecological validity asks whether a study’s findings, however internally sound, generalise to the real-world conditions it is meant to represent. This is conceptually distinct from internal validity (whether the study’s design supports its own causal claims), external validity (whether findings generalise across populations or settings more broadly), and construct validity (whether the study actually measures the concept it claims to). A forensic validation study can have excellent internal validity — a clean, well-controlled comparison against guaranteed ground truth — while still lacking ecological validity, if the test items, task framing, or examiner behaviour under test conditions diverge from casework.

A body of work associated with Growns, Kukucka and colleagues has argued that forensic cognition research specifically needs to attend to ecological validity, because laboratory manipulations of contextual bias, time pressure, or evidence complexity may not reproduce the decision environment of a working laboratory, and conversely, that dismissing bias research as “not ecologically valid” is sometimes used to avoid engaging with uncomfortable findings rather than as a genuine methodological critique [Growns & Kukucka, 2021]. Their argument cuts both ways: ecological invalidity is a real limitation of some experimental designs, but it is not by itself a reason to discount findings that meet reasonable standards of rigour, and the field has historically underinvested in the kind of decision-science collaboration needed to design genuinely representative studies.

Concretely, the representativeness problem in forensic validation research turns on several recurring factors: the difficulty distribution of test items (are hard comparisons over-represented, under-represented, or unknown relative to casework?); whether evidence is degraded in ways typical of real scenes; whether examiners know they are being tested; whether contextual case information is present or absent; and whether laboratory workflow pressures (caseload, turnaround expectations) are reproduced. A method can be “scientifically validated” in the narrow sense of having a documented, internally sound error-rate estimate, while remaining only weakly informative about performance on the specific, heterogeneous casework a court is actually asking about.


6. Proficiency Testing, Black-Box Studies and Operational Performance

The empirical literature distinguishes several related but non-identical testing paradigms: declared proficiency testing (examiners know they are being assessed, often as an accreditation requirement); blind proficiency testing (test items are submitted as if they were ordinary casework, so examiner behaviour is not altered by awareness of assessment); collaborative or interlaboratory exercises (multiple laboratories analyse shared samples, as in the NIST MIX13 DNA mixture interlaboratory study [Buckleton et al., 2018]); black-box studies (large panels of examiners perform realistic comparison tasks under research conditions with guaranteed ground truth, without visibility into the examiner’s internal reasoning); and competency or certification testing (assessing whether an individual meets a minimum standard, not estimating a population error rate).

These serve different scientific purposes and are frequently conflated in public and legal discussion. A proficiency test that most examiners pass easily demonstrates minimum competency; it does not, by itself, estimate the error rate a court needs to weigh a specific piece of evidence, because proficiency-test items are typically far easier, on average, than the disputed comparisons that reach a courtroom. Declared testing in particular is vulnerable to changed examiner behaviour — closer attention, more conservative conclusions, deferral to a second examiner — precisely because the examiner knows an external answer key exists. Blind testing is intended to close this gap, but blind test items must be embedded in genuine casework streams without contaminating chain-of-custody or laboratory accreditation records, which is logistically and ethically demanding and has been implemented only unevenly across jurisdictions.

Black-box studies of the kind conducted for latent fingerprint examination [Ulery et al., 2011] and firearms and toolmark examination [Baldwin et al., 2014; Monson, Smith & Peters, 2022] represent the most methodologically ambitious attempt to estimate operational error rates while preserving guaranteed ground truth. Even so, they inherit several representativeness problems that recur throughout this review: participant self-selection (examiners who volunteer for a study may not be representative of the examiner population as a whole), item-difficulty design choices (studies intentionally including “hard” comparisons to be informative can either overstate or understate real-world error depending on the true difficulty distribution of casework), and — as recent statistical critiques have shown — the handling of missing or excluded responses, which can materially bias reported error rates if not treated carefully [Khan & Carriquiry, 2023; Dorfman & Vanderplas, 2024].


7. Why Consensus Is Not Automatically Ground Truth

CONSISTENCY IS NOT ACCURACYExaminer1Examiner2Examiner3SharedConclusionConsistency / Agreement?does not guaranteeTrueState of the WorldAccuracyAgreement among examiners demonstrates consistency of method application —it independently establishes correctness only when the agreeing parties are genuinely independent.
Figure 3. Examiner agreement demonstrates consistency, not accuracy — the two are related but logically distinct.

A recurring conceptual hazard in forensic validation is the substitution of examiner agreement for independently established accuracy. When several examiners concur, when a laboratory’s technical reviewer confirms a conclusion, or when a finding survives peer scrutiny, it is tempting to treat that convergence as though it settled the underlying factual question. It does not, necessarily. Agreement demonstrates consistency — that independent observers applying the same method reach the same conclusion — which is a precondition for a method being useful, but is logically distinct from accuracy — that the conclusion corresponds to the true state of the world.

This distinction is well illustrated by the black-box and reproducibility literature on latent print examination. Ulery and colleagues separately studied the accuracy of latent print decisions against guaranteed ground truth [Ulery et al., 2011] and the repeatability and reproducibility of those decisions — whether the same examiner reaches the same conclusion on re-examination, and whether different examiners agree with one another [Ulery et al., 2012]. The two questions are related but not the same: a method could in principle produce highly reproducible agreement among examiners who are all making the same systematic error, or conversely could show apparent disagreement that reflects appropriately cautious handling of genuinely ambiguous evidence.

This does not mean consensus is scientifically worthless. Convergent conclusions reached by independent methods, independent examiners working blind to one another, or independent laboratories analysing split samples are a legitimate and important form of corroboration, particularly where no direct ground truth is obtainable at all. The scientifically careful position — consistent with the collaborative testing literature [Buckleton et al., 2018] — is that agreement raises confidence in a conclusion without independently proving it, and that the strength of that inference depends on how independent the agreeing parties truly were (shared training, shared case information, and sequential rather than blind review all weaken the independence on which the inferential value of agreement depends).


8. Ground Truth Across Major Forensic Disciplines

8.1 DNA Interpretation

DNA typing is frequently held up as the forensic discipline with the strongest claim to foundational validity, and for single-source, high-template profiles that claim is well supported: the underlying population-genetic statistics are extensively validated against known-source samples, and the PCAST report itself accepted single-source DNA and two-contributor mixtures as foundationally valid [PCAST, 2016]. The ground-truth picture becomes considerably more complicated as mixture complexity increases. Probabilistic genotyping software such as STRmix™, TrueAllele®, and MaSTR™ is validated using laboratory-constructed mixtures of known contributor number, ratio, and template amount [Moretti et al., 2021; Greenspoon et al., 2015, cited in Bright et al., 2018; internal STRmix multi-laboratory validation, Bright et al., 2018] — a strong form of experimentally constructed ground truth, because the laboratory itself combines known individual profiles to create the mixture.

The PCAST addendum specifically probed the limits of this validation logic, discussing a case in which two probabilistic-genotyping programs reached different conclusions from the same complex, low-template mixture, and rejecting the developer’s claim that the likelihood-ratio approach could not mathematically produce an incorrect inclusion — endorsing instead the position that proper validation requires empirical testing against known-source samples resembling casework, not a mathematical guarantee substituting for it [PCAST addendum, 2017, discussed in Thompson, 2023]. The deeper ground-truth problem for complex mixtures is not whether the software correctly computes a likelihood ratio given its model, but whether laboratory-constructed known mixtures are representative of the transfer mechanisms, degradation, and contributor-number ambiguity found in real touch-DNA and low-template casework — an ecological-validity question rather than a purely statistical one. Moving from source-level to activity-level propositions compounds this: the hierarchy-of-propositions framework developed by Cook, Evett, Jackson and colleagues makes explicit that DNA transfer, persistence, prevalence and recovery (TPPR) findings required to address how and when DNA arrived somewhere rest on empirical transfer studies that are themselves generalisations from a necessarily limited set of experimental scenarios to the near-infinite variety of real-world contact histories [Cook et al., 1998; Evett et al., 2002; Taylor et al., 2018, cited in Gill et al., 2018].

8.2 Latent Fingerprint Examination

Fingerprint examination’s ground-truth architecture rests substantially on the black-box study design pioneered by Ulery and colleagues, in which 169 examiners each compared roughly 100 latent-exemplar pairs drawn from a pool with guaranteed same-source or different-source status [Ulery et al., 2011], later extended to searches drawn from an operational Automated Fingerprint Identification System [Hicklin et al., cited in Richetelli et al., 2025]. These studies established that false-positive identifications, while rare, are not negligible, and that false-negative (missed identification) rates are considerably higher than false-positive rates — a pattern replicated across the discipline’s subsequent studies. A companion study established that examiner decisions show meaningful rates of both intra-examiner and inter-examiner disagreement on repeated or shared items [Ulery et al., 2012], directly illustrating the consensus-versus-accuracy distinction discussed above.

The persistent ground-truth question for fingerprint examination concerns representativeness of test-item difficulty: black-box studies necessarily draw their comparison pairs from a finite, curated pool, and the relationship between that pool’s difficulty distribution and the distribution of difficulty in submitted casework — where latent prints range from clear and complete to fragmentary and distorted — remains only partially characterised. This matters directly for how “the” false-positive rate from a black-box study should be applied (or not applied) to a specific casework comparison of markedly different quality.

8.3 Firearm and Toolmark Examination

Firearms and toolmark identification presents the most methodologically contested ground-truth literature reviewed here. The Ames Laboratory studies, commissioned in response to the 2009 NRC report’s call for reliability research, used known-source test-fires from a large panel of firearms to generate guaranteed same-source and different-source comparison sets [Baldwin et al., 2014, discussed in Monson, Smith & Peters, 2022]. PCAST treated the first Ames study as an adequately designed black-box study supporting foundational validity for firearms examination [PCAST, 2016], a characterisation subsequently challenged in detail by statisticians who argue the study’s handling of “inconclusive” responses, non-response, and examiner exclusion criteria produces error-rate estimates that cannot bear the evidentiary weight placed on them, and that this flaw recurs across the entire family of black-box designs used in the discipline, not merely the Ames studies [Khan & Carriquiry, 2023].

A separate and more recent critique focuses on the confidence-interval methodology used to report error rates from these studies, arguing that the standard analytic approach fails to account for large between-examiner variation in error probability — the more recent large-scale study by Monson, Smith and Peters explicitly found that “the majority of errors were made by a limited number of examiners,” implying that the ground truth against which the discipline’s error rate is measured is not adequately captured by a single population-level number [Monson, Smith & Peters, 2022; Dorfman & Vanderplas, 2024]. This is a case where the ground truth of individual comparisons (same-gun or different-gun) is not in serious dispute, but the derived, higher-order quantity the courts actually rely on — “the error rate of firearms examination” — is contested precisely because of how that guaranteed item-level truth is aggregated and generalised.

8.4 Questioned Document Examination

Handwriting and questioned-document examination occupies an unusual position in the ground-truth landscape because natural writing variation is itself part of what must be modelled: two genuine samples from the same writer are never identical, and a skilled forger’s simulation may be closer to a genuine exemplar than the writer’s own natural variation on a different day. A large-scale black-box study modelled directly on the fingerprint design found that forensic document examiners substantially outperformed non-expert control participants at determining whether two handwriting samples were written by the same person, while also showing a meaningful false-positive rate and considerable examiner-to-examiner variability [Found, Ballantyne et al., cited as “Accuracy and reliability of forensic handwriting comparisons,” PNAS 2022].

The deeper ground-truth question for this discipline is whether laboratory-constructed writing samples — solicited from volunteers under experimental conditions — reproduce the temporal variation, disguise, and genuine forensic disputed-authorship scenarios found in casework, where the questioned material is frequently limited to a few words or a signature, and case-specific factors (illness, intoxication, deliberate disguise, the passage of years between exemplar and questioned document) may not be well represented in a controlled study population. This is precisely the ecological-validity concern raised in Section 5: an internally rigorous black-box design can still leave open how representative its item pool is of the writing disputes that actually reach forensic laboratories.

8.5 Digital and Multimedia Forensics

Digital forensics inverts the usual ground-truth economics: because digital evidence is generated by deterministic machine processes, it is comparatively cheap to build datasets with perfect, documented ground truth — the NIST Computer Forensic Reference Data Sets (CFReDS) project provides disk and device images with known, seeded content specifically so that a tool’s output can be checked against a documented answer [NIST CFReDS, cited in Dataset construction challenges for digital forensics, 2023]. The trade-off is that synthetic or seeded datasets are easy to validate against but may not reproduce the “noisy,” idiosyncratic structure of a real device used by a real person over years, while real-world datasets have the opposite problem: they are realistic but their ground truth is often unknown or only partially reconstructable, because no one recorded, in real time, the complete ledger of everything a device’s owner actually did [Göbel et al., 2020; Dataset construction challenges for digital forensics, 2023].

This tension is compounded by rapid version churn in operating systems, applications, and cloud services: a tool validated against one software version’s known behaviour may not generalise to the next version, and casework devices are rarely running the exact software configuration a validation study used. The distinction that matters here is between verifying that a tool produces the documented output on a known, seeded test image — a well-posed, tractable ground-truth problem — and validating an inference about what a specific human being actually did on a specific device at a specific time, which reintroduces the historical-reconstruction problem common to reconstruction more generally.

8.6 Forensic Reconstruction

Event and timeline reconstruction represents the most acute form of the ground-truth problem reviewed here, because the object of inquiry — a unique historical sequence of events — cannot be replayed, resampled, or independently re-observed under any circumstances. Reconstruction hypotheses can be tested against physical evidence, simulation, and internal consistency, and competing hypotheses can be compared for how well each explains the observed trace evidence, but no experimental design can generate “guaranteed ground truth” for a specific historical event the way a manufactured DNA mixture or a test-fired cartridge case can. The epistemic standard reconstruction can realistically meet is convergent corroboration and the elimination of alternative explanations, not direct verification against a known answer.


9. Error Rates, Dataset Design and the Interpretation of Accuracy

An error rate is not a free-standing property of a method; it is a property of a method measured against a specific dataset and a specific truth standard. A false-positive rate of well under one percent, as reported for firearms cartridge-case comparisons in the largest black-box study to date [Monson, Smith & Peters, 2022], describes performance on that study’s item pool — a pool with its own difficulty distribution, examiner-selection process, and truth-verification procedure — not a universal constant of the discipline. Case difficulty produces what the diagnostic-testing literature calls a spectrum effect: a method’s sensitivity and specificity measured on an easy item pool will differ systematically from performance on a harder one, and neither number is “the” error rate independent of the population of cases considered.

This has a direct courtroom consequence, one PCAST itself emphasised: an examiner reporting an error rate from a foundational validation study should also be able to show that the study’s samples are relevant to the facts of the specific case being tried [PCAST, 2016]. In practice this demonstration is rarely made with precision, because case-specific difficulty is hard to quantify and validation studies are rarely designed with a matching casework-difficulty metric in mind.


10. Method Error Versus Examiner Error

When an examiner reaches an incorrect conclusion in a black-box study, several distinct explanations are compatible with the same observed error, and disentangling them requires more than the raw pass/fail outcome: the underlying method may be sound but the specific evidence item may have been genuinely ambiguous (an evidence-quality limitation); the examiner may have deviated from the documented method (an individual-examiner limitation); the decision threshold built into the method’s reporting terminology (identification / inconclusive / exclusion) may itself be poorly calibrated; or contextual, non-evidentiary information may have influenced the conclusion (a contamination of the decision process rather than a limitation of the underlying comparison method).

The firearms literature again illustrates why this matters: because a small number of examiners in the largest black-box study accounted for a disproportionate share of observed errors [Monson, Smith & Peters, 2022], an error rate reported as though it describes “the method” is, statistically, closer to describing a distribution of individual examiner error probabilities that a population-average number obscures. This is not a peculiarity of firearms examination; it recurs wherever a method combines an objective measurement substrate (striae, minutiae, genetic loci) with a subjective, trained-judgment decision layer, which describes the majority of pattern and impression disciplines.


11. Blind Testing and the Search for Operational Realism

Blind proficiency testing — submitting test items into a laboratory’s ordinary casework stream so that examiners cannot distinguish them from genuine cases — is frequently proposed as a partial bridge between the guaranteed ground truth of research studies and the behavioural realism of actual casework. The logic is straightforward: an examiner who knows they are being tested may behave differently (more cautiously, more conservatively, with heightened attention) than one working a case they believe to be real, so only blind testing can estimate error under authentic operational conditions.

The practical obstacles are substantial. Blind test items must be constructed with genuine forensic realism, tracked without contaminating chain-of-custody records, and reconciled afterward without disrupting laboratory accreditation reporting — a logistical and ethical burden considerably heavier than declared testing, and one that has limited its adoption to a minority of laboratories and disciplines. Even where implemented successfully, blind testing does not resolve every ground-truth problem discussed in this review: it improves behavioural and workflow realism, but the truth against which blind test items are scored is still constructed by the testing body, and the same questions about item-difficulty representativeness that apply to black-box studies apply equally to blind testing programmes.


12. When There Is No Direct Ground Truth: Alternative Validation Frameworks

Where a direct, independently verified answer key is unavailable — which describes most casework and a meaningful share of activity-level and reconstruction questions even in well-validated disciplines — the literature points toward several partial, convergent alternatives rather than a single substitute for ground truth:

  • Triangulation and independent corroboration, in which multiple, genuinely independent evidentiary routes (different evidence types, different examiners working blind to one another, different laboratories) converge on the same conclusion, raising confidence without constituting direct verification.
  • Competing hypothesis testing, structured through the hierarchy-of-propositions framework, in which the value of evidence is assessed as the ratio of its probability under each of at least two explicitly stated, competing explanations, rather than as a bare statement of “the truth” [Cook et al., 1998; Evett et al., 2002].
  • Simulation, synthetic and semi-synthetic datasets, which manufacture guaranteed ground truth at the cost of ecological fidelity, as seen in both DNA mixture validation [Moretti et al., 2021] and digital-forensic reference datasets [NIST CFReDS project].
  • Reproducibility and cross-laboratory replication studies, such as the NIST MIX13 interlaboratory DNA mixture comparison, which test whether independent laboratories converge on the same interpretation of shared material [Buckleton et al., 2018].
  • Uncertainty quantification and probabilistic/Bayesian reporting, which reframes the validation target from a binary “correct/incorrect” outcome to a calibrated statement of evidential weight, explicitly acknowledging that the underlying truth may remain unresolved.
  • Robustness and sensitivity analysis, examining whether a conclusion is stable across reasonable variation in assumptions, population data, or analytic parameters — used extensively in probabilistic genotyping to test how conclusions shift under different population allele-frequency assumptions [cited in DNA mixture population-stratification research].

None of these substitutes for direct ground truth in the strict sense; each buys a different, partial form of confidence, and a scientifically honest validation programme is explicit about which of these partial forms it is offering rather than presenting any of them as equivalent to direct verification.


A court’s verdict resolves a case procedurally; it does not thereby establish scientific ground truth for the forensic method that contributed to it. The circularity risk is direct: a forensic conclusion cannot be validated by reference to a case outcome that the forensic conclusion itself may have substantially influenced, since juries and judges typically weigh forensic testimony heavily and are not independently positioned to verify the underlying comparison. This is one reason the field has moved toward validation studies conducted entirely outside the courtroom — black-box and proficiency studies with researcher-constructed, documented ground truth — rather than attempting to infer method accuracy from conviction and acquittal rates, which conflate legal, evidentiary, and scientific truth in ways the 2009 NRC report and subsequent scholarship treat as a category error [National Research Council, 2009].


14. Ground Truth, Cognitive Bias and Circular Reasoning

The absence of an independently known answer is precisely the condition under which cognitive bias does the most damage, because there is no external check to catch a biased conclusion before it is reported. Dror and colleagues’ foundational experiments showed that fingerprint examiners re-presented with prints they had previously examined — this time accompanied by emotionally charged, biasing contextual information — sometimes reversed their own prior, correct exclusion decisions [Dror, Charlton & Péron, 2006]. Subsequent research extended this to AFIS-ranked candidate lists, where prints appearing higher in a computer-generated ranking were more likely to be judged a match even when underlying similarity was held constant, and to forensic toxicology, where irrelevant case context shifted both interpretive conclusions and testing strategy [Hamnett & Dror, 2020].

A large multinational survey of 403 experienced forensic examiners found that most regarded their own judgment as close to infallible, believed willpower alone could overcome bias, and fewer than half supported blind testing as a countermeasure — a documented “bias blind spot” in which examiners readily acknowledge bias in other domains and other practitioners while discounting it in themselves [Kukucka, Kassin, Zapf & Dror, 2017]. A more recent, cautionary strand of this literature has itself argued that cognitive-bias research in forensic science needs more rigorous, ecologically valid designs before its findings can be generalised with confidence to specific casework decisions — turning the ecological-validity critique back onto the bias literature itself, rather than treating the case for bias as closed [Growns & Kukucka, 2021]. The circularity risk this section addresses is specific: where no independent answer exists, a conclusion that happens to match the expected investigative narrative can be mistaken for a conclusion that has been independently verified, precisely because there is nothing external to check it against.


15. The Reproducibility Problem

Reproducibility in forensic science operates at several distinct levels that are often collapsed into a single word: repeatability of a physical measurement, reproducibility of an interpretive conclusion by the same examiner on a different occasion, reproducibility across different examiners, and reproducibility across different laboratories entirely. The Ulery et al. reproducibility study found meaningful rates of examiner disagreement even on repeated presentation of identical material [Ulery et al., 2012], and the NIST MIX13 interlaboratory exercise examined whether independent laboratories analysing the same DNA mixture data converged on compatible likelihood-ratio conclusions [Buckleton et al., 2018]. If multiple laboratories reproduce the same conclusion, this establishes methodological consistency and, cumulatively, raises confidence in the underlying finding — but it does not, by itself, establish that the shared conclusion is correct, since shared training, shared standards, and shared professional culture can produce consistent agreement on a systematically biased or simply mistaken interpretation. Reproducibility is necessary evidence for confidence in a forensic conclusion; it is not sufficient evidence of its truth.


16. Ground Truth in AI and Computational Forensics

Machine-learning systems entering forensic image, video and digital-artefact analysis inherit the ground-truth problem in an unusually direct form, because their “knowledge” of the world is entirely constituted by the labels in their training data. If those labels are themselves imperfect — mislabelled frames, ambiguous manipulation boundaries, demographically unrepresentative sampling — a model’s apparent accuracy on a benchmark can systematically overstate its reliability on real casework, particularly across the distributional shift between benchmark data and operational data.

This is now well documented in deepfake and digital-image forensics specifically. Large-scale audits of deepfake-detection datasets have found substantial demographic imbalance and label-reliability issues that produce measurably unequal detection performance across attributes such as gender and skin tone, with detection accuracy differing significantly by demographic group in several widely used benchmark datasets [Trinh & Liu, cited in Xu et al., dataset-bias literature; Hazirbas et al., cited in fairness-in-deepfake-detection surveys]. Separately, cross-dataset evaluation research has repeatedly found that detectors trained on one manipulation-generation pipeline suffer sharp accuracy drops when tested on unseen generation methods, indicating that in-benchmark accuracy is a poor proxy for operational generalisability precisely because the “ground truth” labels of training and test sets are drawn from a narrower distribution of manipulation techniques than the ones the system will actually encounter in the field. The central question this raises for forensic deployment is direct: a model cannot be more reliable than the ground truth used to train and evaluate it, and reported accuracy figures need to be read as bounded by, not independent of, the quality and representativeness of the underlying labels.


17. A Cross-Disciplinary Comparative Framework

Forensic Discipline What Counts as Ground Truth? How Ground Truth Is Established Major Validation Problem Ecological Validity Challenge Major Research Gap
DNA (single-source / simple mixture) Known genetic profile of a specific individual Direct typing of reference samples; construction of known mixtures in the laboratory Well-supported for low-complexity profiles; degrades as complexity rises Laboratory-built mixtures may not reflect real transfer/degradation conditions Representativeness of complex-mixture validation samples relative to casework
DNA (activity-level propositions) The historical event of transfer, not merely source Experimental transfer/persistence studies; hierarchy-of-propositions reasoning No experiment can replay the actual historical contact event Limited scenario coverage relative to real-world contact variety Broader, casework-grounded transfer and persistence datasets
Latent fingerprints Documented same-source / different-source status of print pairs Black-box study design with guaranteed comparison sets Established false-positive/negative rates exist but vary by study design Test-item difficulty distribution vs. casework difficulty distribution Difficulty-calibrated, casework-representative item pools
Firearms/toolmarks Known-source status from test-fired specimens Ames-style black-box studies using known firearms Contested handling of inconclusive/missing responses; examiner-level error variance Self-selected volunteer examiner pools; item exclusion criteria Standardised, pre-registered black-box designs with transparent missingness handling
Questioned documents Known writer of exemplar and questioned material Solicited writing samples under experimental conditions Natural intra-writer variation complicates “known truth” itself Disguise, temporal variation and limited questioned material underrepresented Casework-representative disguised/temporal-variation datasets
Digital forensics Documented content/state of a digital artefact Synthetic/seeded reference images (e.g., NIST CFReDS) Tool output verifiable; human-action inference is not Software/version churn; real devices are noisier than reference images Longitudinal, version-tracked validation frameworks
Forensic reconstruction The actual historical sequence of events Physical evidence, simulation, competing-hypothesis testing Historical event cannot be independently re-observed Cannot be tested under controlled repeatable conditions Formal frameworks for corroboration-based (non-direct) validation

18. The Ground Truth Continuum in Forensic Validation

THE GROUND TRUTH CONTINUUM IN FORENSIC VALIDATIONProposed discussion framework — not an adopted NIST / OSAC / ENFSI / ISO standardLevel 1Level 2Level 3Level 4Level 5Directly KnownExperimental TruthGuaranteed by designLatent-print black-box studiesIndependently VerifiedReference TruthVerified by a separate methodProbabilistic-genotyping validationStrongly CorroboratedOperational TruthConvergent independent routesNIST MIX13 interlabcomparisonPartially ObservableHistorical TruthIndirect, incomplete tracesActivity-level DNA transferCompeting ReconstructionHypothesesNo direct verification possibleFull event /timeline reconstructionValidation strength decreases left to right as direct verifiability of the underlying truth decreases.
Figure 4. The Ground Truth Continuum in Forensic Validation — a proposed discussion framework, not an adopted standard.

A conceptual framework proposed for discussion based on the literature reviewed in this article, rather than an established universal forensic standard adopted by NIST, OSAC, ENFSI or ISO.

Level 1 — Directly Known Experimental Truth. The comparison’s true status is fixed by the experimental design itself (e.g., a laboratory-constructed same-source print pair). Validation strength: strongest available; supports direct error-rate estimation. Limitation: ecological fidelity to casework difficulty and workflow is not guaranteed simply because truth is guaranteed. Example: Ulery et al.’s latent-print black-box design [Ulery et al., 2011].

Level 2 — Independently Verified Reference Truth. Truth is established by a separate, higher-certainty method rather than the design itself (e.g., a probabilistic-genotyping conclusion checked against a known reference profile obtained independently of the disputed sample). Validation strength: strong, but depends on the independent method’s own validity. Example: probabilistic-genotyping software validation against known-composition mixtures [Moretti et al., 2021].

Level 3 — Strongly Corroborated Operational Truth. No single method establishes truth directly, but multiple independent evidentiary routes converge (interlaboratory replication, independent examiners working blind, multiple evidence types pointing the same direction). Validation strength: moderate-to-strong, contingent on genuine independence of the corroborating sources. Example: the NIST MIX13 interlaboratory DNA mixture comparison [Buckleton et al., 2018].

Level 4 — Partially Observable Historical Truth. Some evidentiary trace of the historical event exists, but the event itself cannot be independently re-observed or replayed; conclusions rest on inference from partial, indirect data (e.g., activity-level DNA transfer questions; digital-forensic inference of user action). Validation strength: limited to what corroboration and robustness testing can supply; direct verification is not available. Example: activity-level DNA evaluation using the hierarchy of propositions [Evett et al., 2002; Gill et al., 2018].

Level 5 — Competing Reconstruction Hypotheses. The historical event cannot be observed even partially through direct trace evidence sufficient to fix a single account; the analytic task is to compare the relative explanatory power of competing accounts against available physical evidence. Validation strength: weakest in the direct-verification sense; scientific rigor here consists of transparent hypothesis comparison and explicit acknowledgment of residual uncertainty, not error-rate estimation. Example: full event and timeline reconstruction from fragmentary physical evidence.

The continuum is offered as a discussion tool for locating where a given forensic question sits, and for making explicit — to examiners, to courts, and to the public — that “validated” does not mean the same thing at every level.


19. Major Research Gaps

  1. Difficulty-calibrated black-box item pools whose distribution is empirically matched to submitted casework, rather than researcher-selected item sets of unknown representativeness.
  2. Routine, discipline-wide adoption of blind operational testing, currently implemented unevenly and in a minority of laboratories.
  3. Standardised, pre-registered statistical protocols for handling inconclusive responses and missing data in black-box studies, given documented sensitivity of reported error rates to these choices [Khan & Carriquiry, 2023].
  4. Casework-grounded activity-level transfer and persistence datasets that extend beyond the necessarily limited set of experimental contact scenarios studied to date.
  5. Cross-laboratory replication studies outside DNA, where interlaboratory comparison exercises comparable to NIST MIX13 remain rare.
  6. Longitudinal, version-tracked digital-forensic tool validation that keeps pace with operating-system and application update cycles.
  7. Transparent reporting of dataset demographic composition and label-reliability limitations in machine-learning forensic tools, particularly deepfake and digital-image forensics.
  8. Formal, examiner-facing frameworks for communicating which “ground-truth level” (per Section 18) a given forensic conclusion actually occupies, rather than presenting all validated methods as uniformly strong.
  9. Independent, non-manufacturer-funded validation of proprietary probabilistic-genotyping software using casework-representative mixture complexity.
  10. Individual-examiner error-rate modelling (rather than population-average reporting), given documented concentration of errors among a minority of examiners in firearms black-box studies [Monson, Smith & Peters, 2022].
  11. Expanded ecological-validity research specifically designed, in collaboration with decision scientists, to test whether laboratory bias manipulations generalise to real casework decision environments [Growns & Kukucka, 2021].
  12. Systematic study of how contextual case information should be structured and sequenced (linear sequential unmasking and related protocols) across disciplines beyond fingerprints and DNA.
  13. Research on the ground-truth limitations of questioned-document black-box studies specifically with respect to disguised and temporally separated writing samples.

20. A Future Research Agenda

Priority 1 — Difficulty-Representative Ground-Truth Datasets. Building item pools whose difficulty distribution is empirically anchored to casework rather than researcher intuition. Rationale: representativeness directly determines whether a study’s error rate generalises. Difficulty: substantial, requiring access to and characterisation of real casework item difficulty without compromising ongoing investigations. Benefit: error-rate estimates that can be meaningfully matched to specific case circumstances. Limitation: casework difficulty itself may be hard to characterise without the very expert judgment under study.

Priority 2 — Ecologically Valid Validation Studies. Designing bias and performance studies in direct collaboration with cognitive and decision scientists, per the case made by Growns and Kukucka. Rationale: internal validity alone does not establish operational relevance. Difficulty: moderate; requires cross-disciplinary collaboration infrastructure the field has historically lacked. Benefit: findings courts and laboratories can act on with more confidence. Limitation: perfect ecological fidelity in a controlled study remains unattainable by definition.

Priority 3 — Routine Blind Operational Testing. Scaling blind proficiency testing from isolated programmes to standard laboratory practice. Rationale: closes the gap between declared-test behaviour and casework behaviour. Difficulty: high — logistical, ethical and accreditation complexities. Benefit: the most behaviourally realistic error-rate evidence available short of full transparency about ongoing casework. Limitation: does not resolve item-difficulty representativeness on its own.

Priority 4 — Cross-Laboratory Replication. Extending interlaboratory comparison exercises like NIST MIX13 to pattern and impression disciplines. Rationale: reproducibility across independent laboratories is a stronger corroborating signal than single-laboratory consistency. Difficulty: moderate-to-high, requiring shared sample logistics across institutions. Benefit: distinguishes discipline-wide reliability from laboratory-specific practice. Limitation: shared training pipelines across laboratories can produce correlated, not independent, agreement.

Priority 5 — Discipline-Specific Validation Architectures. Abandoning the assumption of a single universal validation template (the implicit target of some PCAST-era debate) in favour of frameworks tailored to where a discipline sits on the ground-truth continuum in Section 18. Rationale: DNA source attribution and bloodstain reconstruction are not comparable validation problems. Difficulty: moderate; primarily conceptual and organisational. Benefit: prevents both over-claiming (treating consensus as accuracy) and under-claiming (dismissing corroboration-based disciplines as unscientific). Limitation: risks fragmenting standards if not coordinated across accrediting bodies.

Priority 6 — Quantification of Uncertainty. Wider adoption of calibrated, probabilistic (including Bayesian) reporting frameworks in place of categorical identification/exclusion language. Rationale: makes residual uncertainty explicit rather than hidden inside a binary conclusion. Difficulty: moderate; requires examiner retraining and, in some jurisdictions, legal-culture adaptation. Benefit: aligns reported conclusions with what validation evidence actually supports. Limitation: probabilistic language can be harder for juries to interpret without careful explanation.

Priority 7 — Transparent Reporting of Dataset Limitations. Mandating disclosure of item-selection criteria, exclusion rates, and demographic composition alongside any reported validation statistic. Rationale: the missingness and exclusion issues identified in firearms black-box studies show how much a reported error rate depends on choices rarely disclosed in full [Khan & Carriquiry, 2023]. Difficulty: low-to-moderate; primarily a reporting-standard change. Benefit: allows independent reanalysis and scrutiny. Limitation: requires researcher buy-in and journal/accreditation-body enforcement.

Priority 8 — AI Dataset and Label Validation. Building forensic machine-learning benchmarks with documented, audited label reliability and demographic balance before deployment claims are made. Rationale: model accuracy cannot exceed the reliability of its training and evaluation ground truth. Difficulty: high, given the scale of data needed and the cost of expert labelling. Benefit: prevents overstated deployment-readiness claims for forensic AI tools. Limitation: some ground-truth ambiguity (e.g., borderline manipulation cases) may be irreducible even with careful auditing.


Conclusion: What Should “Validated” Mean in Forensic Science?

How do we validate forensic science when the real answer is unknown? The literature surveyed here does not offer a single formula, and the honest answer is that it should not: forensic validation cannot rest on one concept of ground truth, because DNA source attribution, latent print comparison, firearms identification, questioned-document analysis, digital-forensic inference, and full event reconstruction sit at genuinely different points on the continuum described in Section 18. Experimental validation — black-box studies, probabilistic-genotyping mixture testing, digital-forensic reference datasets — remains indispensable and should not be diminished by the arguments in this review; it is simply not sufficient by itself. Operational generalisability from those studies to specific casework must be examined independently and explicitly, not assumed. Consensus among examiners, laboratories, or reviewers is valuable corroboration but is not a substitute for independently established accuracy, a distinction the discipline has periodically forgotten under courtroom and public pressure to speak with more certainty than the underlying validation evidence supports. Casework uncertainty — particularly at the activity-level and reconstruction end of the continuum — should be acknowledged in reporting language rather than absorbed into a categorical conclusion that implies more certainty than the evidence contains. And where direct ground truth is genuinely inaccessible, justified confidence should be built from convergent, independent corroboration, robustness and sensitivity testing, competing-hypothesis evaluation, and transparent disclosure of what a given validation study can and cannot show.

The deepest scientific challenge in forensic science is not simply determining whether an examiner can reach the correct answer when the answer is known. It is determining how much confidence can be justified when the truth is precisely what the investigation is trying to discover.


References

  1. National Research Council. Strengthening Forensic Science in the United States: A Path Forward. Washington, DC: National Academies Press; 2009.
  2. President’s Council of Advisors on Science and Technology. Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods. Washington, DC: Executive Office of the President; 2016.
  3. President’s Council of Advisors on Science and Technology. An Addendum to the PCAST Report on Forensic Science in Criminal Courts. Washington, DC: Executive Office of the President; 2017.
  4. Ulery BT, Hicklin RA, Buscaglia J, Roberts MA. Accuracy and reliability of forensic latent fingerprint decisions. Proc Natl Acad Sci USA. 2011;108(19):7733-8.
  5. Ulery BT, Hicklin RA, Buscaglia J, Roberts MA. Repeatability and reproducibility of decisions by latent fingerprint examiners. PLoS ONE. 2012;7(3):e32800.
  6. Ulery BT, Hicklin RA, Buscaglia J, Roberts MA. Assessing the clarity of friction ridge impressions. Forensic Sci Int. 2013;226(1-3):106-17.
  7. Ulery BT, Hicklin RA, Roberts MA, Buscaglia J. Interexaminer variation of minutia markup on latent fingerprints. Forensic Sci Int. 2016;264:89-99.
  8. Hicklin RA, Richetelli N, Taylor A, Buscaglia J. Accuracy and reproducibility of latent print decisions on comparisons from searches of an automated fingerprint identification system. Forensic Sci Int. 2025 (in press/ScienceDirect).
  9. Monson KL, Smith ED, Peters EM. Accuracy of comparison decisions by forensic firearms examiners. J Forensic Sci. 2022;67(6):2213-27.
  10. Baldwin DP, Bajic SJ, Morris M, Zamzow D. A Study of False Positive and False Negative Error Rates in Cartridge Case Comparisons. Ames, IA: Ames Laboratory, Technical Report IS-5207; 2014.
  11. Khan MSU, Carriquiry AL. Shining a light on forensic black-box studies. Chance / arXiv. 2023. Available from: arxiv.org/pdf/2209.14216.
  12. Dorfman A, Vanderplas S. Methodological problems in every black-box study of forensic firearm comparisons. Law Probab Risk. 2024;23(1):mgae015.
  13. Cook R, Evett IW, Jackson G, Jones PJ, Lambert JA. A model for case assessment and interpretation. Sci Justice. 1998;38(3):151-6.
  14. Cook R, Evett IW, Jackson G, Jones PJ, Lambert JA. A hierarchy of propositions: deciding which level to address in casework. Sci Justice. 1998;38(4):231-9.
  15. Evett IW, Jackson G, Lambert JA. More on the hierarchy of propositions: exploring the distinction between explanations and propositions. Sci Justice. 2000;40(1):3-10.
  16. Evett IW, Gill PD, Jackson G, Whitaker J, Champod C. Interpreting small quantities of DNA: the hierarchy of propositions and the use of Bayesian networks. J Forensic Sci. 2002;47(3):520-30.
  17. Gill P, Hicks T, Butler JM, Connolly E, Gusmão L, Kokshoorn B, et al. DNA commission of the International Society for Forensic Genetics: assessing the value of forensic biological evidence — guidelines highlighting the importance of propositions. Part I. Forensic Sci Int Genet. 2018;36:189-202.
  18. Taylor D, Kokshoorn B, Blankers BJ, de Zoete J, Berger CE. Activity level DNA evidence evaluation: on propositions addressing the actor or the activity. Forensic Sci Int Genet. 2018.
  19. Yang B, et al. American forensic DNA practitioners’ opinion on activity level evaluative reporting. J Forensic Sci. 2022;67(4).
  20. Bright JA, Curran JM, Buckleton JS, et al. Internal validation of STRmix — a multi-laboratory response to PCAST. Forensic Sci Int Genet. 2018;34:11-24.
  21. Buckleton JS, Bright JA, Cheng K, et al. NIST interlaboratory studies involving DNA mixtures (MIX13): a modern analysis. Forensic Sci Int Genet. 2018;37:172-9.
  22. Moretti TR, et al. Internal validation of STRmix for the interpretation of single source and mixed DNA profiles. Forensic Sci Int Genet. 2021;29:126-44.
  23. Thompson WC. Uncertainty in probabilistic genotyping of low template DNA: a case study comparing STRmix and TrueAllele. J Forensic Sci. 2023;68(3).
  24. Federal Judicial Center. Probabilistic Genotyping Systems for Low-Quality and Mixture Forensic Samples. Washington, DC: FJC; n.d.
  25. [Internal Validation of MaSTR Probabilistic Genotyping Software for the Interpretation of 2-5 Person Mixed DNA Profiles]. PMC9408203. 2022.
  26. [Evaluating DNA Mixtures with Contributors from Different Populations Using Probabilistic Genotyping]. PMC9858364. 2023.
  27. Found B, Ballantyne KN, et al. Accuracy and reliability of forensic handwriting comparisons. Proc Natl Acad Sci USA. 2022;119(21):e2119944119.
  28. Dror IE, Charlton D, Péron AE. Contextual information renders experts vulnerable to making erroneous identifications. Forensic Sci Int. 2006;156(1):74-8.
  29. Dror IE, Charlton D. Why experts make errors. J Forensic Identif. 2006;56(4):600-16.
  30. Kukucka J, Kassin SM, Zapf PA, Dror IE. Cognitive bias and blindness: a global survey of forensic science examiners. J Appl Res Mem Cogn. 2017;6(4):452-9.
  31. Hamnett HJ, Dror IE. The effect of contextual information on decision-making in forensic toxicology. Forensic Sci Int Synergy. 2020;2:339-48.
  32. Growns B, Kukucka J. An inconvenient truth: more rigorous and ecologically valid research is needed to properly understand cognitive bias in forensic decisions. Forensic Sci Int Synergy. 2021;3:100157.
  33. [A Forensic Science-Based Model for Identifying and Mitigating Forensic Mental Health Expert Biases]. J Am Acad Psychiatry Law. 2025;53(2):172.
  34. [A practical approach to mitigating cognitive bias effects in forensic casework]. PMC11720873. 2024.
  35. National Institute of Standards and Technology. Computer Forensic Reference Data Sets (CFReDS) Project. Gaithersburg, MD: NIST; ongoing.
  36. National Institute of Standards and Technology. Computer Forensic Tool Testing (CFTT) Program. Gaithersburg, MD: NIST; ongoing.
  37. Göbel T, et al. A novel approach for generating synthetic datasets for digital forensics. In: Peterson G, Shenoi S, editors. Advances in Digital Forensics XVI. Springer; 2020.
  38. [Dataset construction challenges for digital forensics]. PMC9996465. 2023.
  39. [Data for Digital Forensics: Why a Discussion on “How Realistic is Synthetic Data” is Dispensable]. Digital Threats Res Pract. ACM. 2023.
  40. [Towards a standardized methodology and dataset for evaluating LLM-based digital forensic timeline analysis]. Forensic Sci Int Digit Investig / ScienceDirect. 2025.
  41. Trinh L, Liu Y. An examination of fairness of AI models for deepfake detection. In: Proceedings of IJCAI. 2021 (cited in dataset-bias/fairness surveys on deepfake detection).
  42. Hazirbas C, Bitton J, Dolhansky B, et al. Towards measuring fairness in AI: the Casual Conversations dataset. IEEE Trans Biom Behav Identity Sci. 2021 (cited in deepfake-detection fairness literature).
  43. [Analyzing Fairness in Deepfake Detection with Massively Annotated Databases]. arXiv:2208.05845. 2022.
  44. [Data-Driven Fairness Generalization for Deepfake Detection]. arXiv:2412.16428. 2024.
  45. [Deepfake detection across image, video, and audio: a comprehensive survey with empirical evaluation of generalization and robustness]. Artif Intell Rev. Springer Nature. 2026.
Tags

Post a Comment

0Comments

Post a Comment (0)