Is Probabilistic Genotyping Making DNA Evidence Better—or More Difficult for Courts to Understand?
A forensic-science editorial examining the statistical power and the courtroom-comprehension cost of probabilistic genotyping in modern DNA mixture interpretation.
INTRODUCTION
A swab from a doorknob, a steering wheel, or the inside of a discarded mask rarely carries DNA from one person. It carries DNA from whoever touched it last, and often from several people before that — the owner, a passenger, a shopkeeper, a relative who borrowed the car. When that mixed sample is amplified and run through a genetic analyzer, the resulting electropherogram does not look like the clean, single-peak profile most people picture when they hear the words "DNA match." It looks like a crowd of overlapping peaks of different heights, some of them so short they could be genuine alleles or could be background noise. Deciding whose DNA is actually present, and how strongly the evidence points toward one explanation over another, is no longer a matter of reading a chart. It is a statistical inference problem.
Probabilistic genotyping software was built to handle exactly this problem. Over roughly the last decade, tools such as STRmix and TrueAllele have moved from research curiosities to standard casework instruments in many forensic DNA laboratories worldwide, precisely because they can extract usable information from mixtures that older, rule-based methods either excluded outright or interpreted only partially. That is a genuine scientific advance, and the literature reviewed below supports it.
But the same complexity that lets these tools interpret difficult evidence also makes their output harder to explain. A likelihood ratio of several billion, generated by a continuous probabilistic model running Markov chain Monte Carlo simulations across thousands of possible genotype combinations, is not self-explanatory to a jury, and arguably not fully self-explanatory to many of the lawyers and judges who must decide how much weight it deserves. This is the tension this article sets out to examine: probabilistic genotyping can pull more information out of difficult DNA evidence, but the more mathematically sophisticated that extraction becomes, the more work is required to make sure courts understand what the resulting number does and does not mean.
Single-source DNA profiles — a reference sample from a known individual, or a stain that clearly came from one person — are comparatively easy to interpret. Each locus shows one or two clean peaks, contributor genotypes can usually be read directly off the electropherogram, and the statistical question reduces to how rare that specific combination of alleles is in the relevant population.
Casework samples are frequently not like that. NIST's 2024 scientific foundation review on DNA mixture interpretation notes that improvements in the sensitivity of DNA testing have allowed profiles to be generated from a few skin cells, which has extended DNA analysis into new categories of casework but has also made complex mixtures — DNA from two, three, or more contributors — a routine rather than an exceptional occurrence. The report identifies several factors that make mixture interpretation inherently harder than single-source analysis: distinguishing one contributor's alleles from another's, estimating how many people actually contributed DNA, judging whether a trace amount is even relevant to the case or is the product of incidental contact or contamination, and detecting trace levels of a suspect's or victim's DNA sitting beneath a larger, unrelated contribution [NIST, 2024].
Several distinct phenomena compound the difficulty. Low-level or "low-template" DNA — amounts near the limit of detection — produces amplification that is inherently stochastic rather than proportional, so a true allele from one contributor can fail to amplify at all (drop-out) while an artefact can appear where no real allele exists (drop-in). Degraded DNA, common in older or environmentally exposed samples, preferentially loses larger fragments, distorting the expected relationship between peak height and contributor proportion. PCR stutter — a technical artefact of the amplification process — creates minor peaks one repeat unit away from a true allele, which can be mistaken for a genuine minor contributor if not modelled correctly. And heterozygote peak imbalance means that even the two alleles belonging to the same person at the same locus may not appear at equal heights, which complicates any attempt to infer mixture ratios by eye.
None of these effects are exotic edge cases. NIST IR 8351 catalogues them as ordinary features of casework-type samples, and the report's central argument is that if they are not properly modelled and communicated, they can distort both the strength and the perceived relevance of DNA evidence presented in court [NIST, 2024]. Traditional binary interpretation — where an analyst simply includes or excludes a genotype above a fixed stochastic threshold — was designed for a simpler era of DNA testing and struggles once several of these effects operate simultaneously in the same sample. That gap is what probabilistic genotyping software was built to close.
It is worth being precise about what probabilistic genotyping is, because loose language does real damage here. Probabilistic genotyping software does not "look at DNA and calculate who committed the crime." It is a statistical tool that evaluates DNA profiling results against two or more competing, explicitly stated propositions — for example, "the DNA came from the person of interest and two unknown individuals" versus "the DNA came from three unknown individuals unrelated to the person of interest" — and calculates how much more probable the observed electropherogram data is under one proposition than under the other.
Mechanically, the software builds a statistical model of the biological and analytical processes that could have produced the observed peak pattern: the number of contributors, their relative proportions, the probability of allele drop-out and drop-in at given DNA quantities, stutter ratios, and degradation. Continuous models — the category that includes STRmix and TrueAllele — use quantitative peak-height information directly; earlier semi-continuous models used only which alleles were present or absent, combined with drop-out and drop-in probabilities. The software then explores an enormous space of possible genotype combinations, typically using Markov chain Monte Carlo methods, to estimate how likely the observed data is under each proposition, and reports the ratio of those two likelihoods as the likelihood ratio (LR).
This is fundamentally different from a combined probability of inclusion or exclusion, the older statistic that simply asked whether a person's genotype could be present in a mixture at all, without weighing how well it explains the data relative to an alternative. NIST's review discusses the LR framework and probabilistic genotyping software together precisely because the LR is the conceptual product the software exists to compute [NIST, 2024]. The output is a number that expresses relative support for one proposition over another — not a probability of guilt, not a probability that a named person is the source, and not a standalone verdict on the evidence. That distinction is central enough to this entire debate that it needs its own section, further below.
The scientific case for probabilistic genotyping rests on a fairly narrow but genuine claim: for mixtures that binary or semi-continuous methods struggle with, continuous probabilistic models can use more of the information actually present in the data — full peak-height patterns rather than simple presence/absence calls — and can therefore produce a more informative, and often more discriminating, assessment of the evidence.
The clearest institutional endorsement of this point comes from the source that is often treated as the most skeptical: the 2016 report of the U.S. President's Council of Advisors on Science and Technology (PCAST). Even while raising concerns discussed later in this article, PCAST stated that these probabilistic genotyping programs "clearly represent a major improvement over purely subjective interpretation" of complex mixtures, while adding that they still require careful scrutiny to establish where their results are and are not reliable [PCAST, 2016].
Peer-reviewed developmental and internal validation studies support the same general conclusion for specific software packages, within defined bounds. The developmental validation of STRmix, conducted according to the Scientific Working Group on DNA Analysis Methods (SWGDAM) 2015 validation guidelines, reported sensitivity, specificity, and precision data supporting the software's suitability for interpreting single-source and mixed DNA profiles [Bright et al., Forensic Science International: Genetics]. A subsequent multi-laboratory internal validation study compiled data from thirty-one laboratories across 2,825 mixtures of three to six donors, explicitly framed as a response to PCAST's call for broader empirical testing, and found the software's discriminating power to be consistent with earlier single-laboratory findings, while also identifying the specific conditions — low template amounts and high contributor numbers — under which likelihood ratios become less discriminating for both true donors and non-donors [Bright et al., Forensic Science International: Genetics, 2018]. Independent validation of a different package, MaSTR, tested against two-to-five-person mixtures likewise reported findings consistent with current probabilistic genotyping validation standards [Adamowicz, Rambo & Clarke, Genes, 2022].
It is worth stressing what these studies do and do not show. They demonstrate that, within the mixture complexity and template ranges actually tested, a given software implementation performs in line with its design goals. They do not establish that probabilistic genotyping is "more accurate" than human interpretation in some universal sense, and none of the sources reviewed for this article make that claim. The more defensible statement, and the one this article adopts, is that probabilistic genotyping can produce a more informative and scientifically defensible interpretation of specific classes of mixture — provided the software has been properly validated for the mixture type in question and is applied within the bounds that validation established.
The advantage described above comes at a communication cost, and that cost is not incidental — it follows directly from what makes the method powerful. A binary inclusion/exclusion call can be explained to a jury in a sentence. A likelihood ratio generated by a continuous probabilistic model, built on assumptions about contributor number, degradation, stutter, and drop-out/drop-in probabilities, and computed through a stochastic simulation process, cannot be explained nearly as briefly without either oversimplifying it or losing the courtroom.
This is not a hypothetical concern. Empirical research on how people interpret DNA evidence framed as a likelihood ratio has found that mock jurors' verdicts can be influenced by the presence of strong DNA evidence in ways that are not fully sensitive to other case facts, such as the strength of a suspect's alibi, and the researchers note this pattern is consistent with — though not conclusively proof of — the kind of overweighting associated with the prosecutor's fallacy [Psychonomic Bulletin & Review, 2020]. A related line of experimental work compared how mock jurors respond to random match probabilities, likelihood ratios, and verbal equivalents of likelihood ratios, finding that verdicts tracked the strength of DNA evidence reasonably well regardless of the framing used, but also documenting that fallacious interpretations — consistent with the classic "source probability error" — occurred under more than one presentation format [Prosecutor's Fallacy and Expert Testimony study].
None of this means the underlying science is unsound. It means that a scientifically valid number can still be misread if it is not explained carefully, and that the more information-dense the statistical method becomes, the more explanatory work an expert witness has to do to prevent that misreading. That is the crux of the central question this article set out to examine, and the literature reviewed here does not resolve it neatly in either direction — it points to a genuine, ongoing communication burden that grows alongside the method's sophistication.
A likelihood ratio is the probability of observing the evidence given one proposition, divided by the probability of observing the same evidence given a competing proposition. In forensic DNA casework, a typical pairing might be: "the DNA mixture contains material from the person of interest and two unknown people" (proposition Hp, often aligned with the prosecution's position) against "the DNA mixture contains material from three unknown people unrelated to the person of interest" (proposition Hd, often aligned with the defence's position). The LR expresses how many times more probable the observed profiling data is under Hp than under Hd.
Key Distinction
A likelihood ratio measures the relative support the evidence provides for one proposition over another. It is not, and cannot logically be, the probability that the person of interest is guilty. Guilt is a legal conclusion reached by weighing all the evidence in a case, including but never limited to a DNA result.
This distinction is not editorial preference; it is basic probability theory, and its courtroom violation has a name: the prosecutor's fallacy, also called the fallacy of the transposed conditional. It occurs when a probability of the form "probability of this evidence, if the defendant is innocent" is silently converted into "probability that the defendant is innocent, given this evidence." These are different quantities, and conflating them can differ from the correct figure by orders of magnitude, because the second quantity depends on prior probability of guilt — information the DNA result alone does not and cannot supply [prosecutor's fallacy literature].
The fallacy is not a theoretical risk invented by academics. In the English case R v Deen, the Court of Appeal quashed a rape conviction after a DNA expert testified that the probability of a positive test from someone other than the defendant was 1 in 3 million, and that figure was then treated — by the expert and subsequently by the judge's summing-up — as though it were the probability that the DNA came from someone other than the defendant. A retrial was ordered [R v Deen, discussed in Open University teaching materials on expert evidence]. The case is now a standard teaching example of exactly the error that likelihood-ratio reporting, properly explained, is intended to prevent — because an LR, correctly presented, forces the fact-finder to keep the direction of the conditional probability straight, and explicitly leaves the assessment of prior probability, and therefore the ultimate question of guilt, to the court.
Large likelihood ratios and rarity statistics are common in modern forensic DNA reporting; figures in the billions are not unusual given the discriminating power of current STR multiplexes [see, e.g., commentary in Lexology, 2023, on the frequency of very large likelihood ratio figures in casework]. It is worth working through, in a clearly hypothetical example, what such a number does and does not mean.
Suppose an expert testifies that a DNA mixture recovered from a crime scene is one billion times more likely to occur if the person of interest and two unknown people are the contributors than if three unknown, unrelated people are the contributors. Two very different statements might be built from that number.
The first, incorrect statement: "there is only a one-in-a-billion chance that this person is innocent." This treats the likelihood ratio as if it were a posterior probability of innocence — exactly the transposition described in the prosecutor's fallacy above. It silently assumes a prior probability of guilt (in effect, that the person of interest was equally likely, before the DNA evidence, to be the source as any of billions of other people), and it hands the jury a false sense that the DNA result alone has settled the ultimate question.
The second, correct statement: "the observed DNA profiling data is a billion times more probable under the proposition that the person of interest contributed to this mixture than under the proposition that three unrelated, unknown people did." This is a statement about the relative explanatory power of two propositions with respect to one specific body of evidence — the DNA result — and nothing more. It says nothing on its own about whether the person of interest was at the scene, whether they had a lawful reason to be there, or whether some other piece of evidence entirely changes the picture. Converting this figure into an actual probability of guilt requires combining it, formally or informally, with everything else known about the case — precisely the reasoning process a jury, not a laboratory report, is supposed to perform.
These two statements are not stylistic variants of each other. They answer different questions, and treating them as interchangeable is the single most consequential communication failure this article's source material repeatedly identifies.
Commercial probabilistic genotyping software is, in most jurisdictions, proprietary. Developers argue that their algorithms, source code, and specific implementation details are trade secrets protected by intellectual property law, and that the underlying scientific methodology — published in peer-reviewed validation papers — is what matters for assessing reliability, not the literal source code [see, e.g., Cybergenetics' public position in coverage of the New Jersey v. Pickett litigation]. Defence advocates and some legal scholars argue the opposite: that a defendant's right to confront and meaningfully challenge evidence used against them requires access to the actual code that generated a number used to convict, because independent examination of comparable forensic software has, in some instances, revealed implementation errors that only became visible once the code was reviewed [Electronic Frontier Foundation, 2020; 2021].
This dispute has produced real, divergent court rulings — a useful reminder that "the courts have decided this" is not currently an accurate description of the state of the law. In Washington v. Fair, a King County Superior Court judge in Seattle denied a defence motion to compel disclosure of TrueAllele's source code, holding that the software's reliability could be established through published validation studies and expert testimony without reviewing the code itself, and that the demand for source code was accordingly not material to determining reliability [Washington v. Fair, King County Superior Court ruling, reported by Cybergenetics, 2017]. In New Jersey v. Pickett, by contrast, the Superior Court of New Jersey's Appellate Division reversed a trial court's denial of a similar motion, holding that an independent defence expert should be permitted to examine TrueAllele's source code in advance of an admissibility hearing — a decision welcomed by the Electronic Frontier Foundation and the ACLU of New Jersey as a due-process victory, and one the prosecution in that case sought (unsuccessfully, at the trial level) to have reconsidered [New Jersey v. Pickett, Superior Court of New Jersey Appellate Division, 2021; EFF, 2021; Cyberlaw Clinic, Harvard Law School, 2021].
It is worth being fair to both positions rather than treating "black box" as automatically synonymous with "bad science." Proprietary status and scientific opacity are not the same thing. A method can be commercially closed at the level of its specific code implementation while remaining scientifically transparent at the level of its published statistical model, its documented assumptions, and its externally verifiable validation performance — which is the position software developers and courts that have declined to compel source-code disclosure have generally taken. Conversely, a method can be fully open-source and still poorly validated, poorly documented, or poorly understood by the people relying on it. The genuinely contested question, on the evidence reviewed here, is not whether courts should care about transparency — everyone involved in this debate agrees they should — but at what level that transparency needs to operate: the published methodology and validation record, or the literal executable code.
It is a mistake — one this article deliberately avoids — to describe probabilistic genotyping as though the software reaches a conclusion on its own. A substantial amount of human judgement precedes and follows every likelihood ratio the software produces.
- Sample and peak assessment: an analyst decides which peaks in the electropherogram represent genuine alleles versus artefacts before any data enters the model.
- Estimating the number of contributors: this is frequently an analyst judgement call, informed by but not dictated by the software, and NIST's review specifically flags contributor-number estimation as a source of interpretive uncertainty [NIST, 2024].
- Formulating propositions: the Hp/Hd pair the software tests is chosen by the analyst (often in consultation with investigators or the court), and the framework for choosing propositions correctly — the "hierarchy of propositions" — is itself a distinct, decades-old area of forensic statistical scholarship, tracing to foundational work by Cook, Evett, Jackson and colleagues in the late 1990s and refined since [Cook et al., Science & Justice, 1998; Evett et al., Journal of Forensic Sciences, 2002].
- Software configuration and settings: parameters, reference population databases, and interpretation thresholds are chosen by the laboratory, generally under standard operating procedures set during validation, not by the software itself.
- Review, reporting, and testimony: results are typically subject to technical and administrative review before being reported, and the expert who testifies is responsible for explaining what the number means and does not mean to the court.
The likelihood ratio is therefore better described as the output of a human-configured statistical process applied to human-classified data, rather than as an independent algorithmic verdict. This does not eliminate the risk of error — human judgement introduces its own sources of variability, including the possibility of cognitive bias in peak calling or proposition framing — but it does mean that criticisms aimed purely at "the algorithm" often miss where a substantial share of the interpretive work, and the interpretive risk, actually sits.
Validation is the backbone of any claim that a probabilistic genotyping result is scientifically defensible, and the forensic DNA community has built a reasonably structured framework for it. SWGDAM's 2015 Guidelines for the Validation of Probabilistic Genotyping Systems established the standard against which developmental validation (testing by the software developer, establishing sensitivity, specificity, and precision across a defined range of mixture types) and internal validation (testing by each individual laboratory adopting the software, against its own instruments, protocols, and casework conditions) are conducted [referenced across multiple validation studies, including Bright et al. and Adamowicz et al.].
NIST's review usefully distinguishes reliability from relevance as separate axes of assessment: a method can produce a numerically well-behaved, reproducible likelihood ratio (reliability) while the underlying question of whether that number bears on the actual dispute in the case — for instance, how the DNA got onto an item, rather than merely whose DNA it is — remains a separate matter (relevance), addressed at the activity level of interpretation rather than the source level [NIST, 2024]. This distinction matters because a well-validated source-level LR does not, by itself, resolve activity-level questions such as how or when DNA was deposited — a point developed extensively in the DNA Commission of the International Society for Forensic Genetics' guidelines on evaluating findings at the activity level [DNA Commission of the ISFG, Forensic Science International: Genetics, 2018–2019].
Proficiency testing and interlaboratory comparison provide an additional, independent check. One study that specifically re-examined a set of proficiency test results referenced in NIST's draft foundation review found that, across the tests reviewed for the years 2018–2021, none of the false positive or false negative results identified could be attributed to the mixture interpretation strategy used, and certainly not specifically to the use of probabilistic genotyping software [Bille, Coble, Kalafut & Buckleton, Genes, 2022]. This is a narrow, specific finding about a specific dataset — it does not establish that probabilistic genotyping is error-free in general, and it should not be read that way — but it is a relevant, peer-reviewed data point in an area where the underlying NIST report itself acknowledged some false positives and false negatives occurred across the broader set of proficiency tests it examined [NIST IR 8351, discussed in Bille et al., 2022].
Version control and documentation round out the validation picture. Because probabilistic genotyping software is periodically updated, laboratories are expected to revalidate, or at minimum assess the impact of, new versions before deploying them on casework, and NIST's review situates this expectation within its broader discussion of software reliability and the practical realities of implementing continuously evolving tools in an accredited laboratory setting [NIST, 2024].
NIST Interagency Report 8351, DNA Mixture Interpretation: A NIST Scientific Foundation Review, published in final form in December 2024, is the most comprehensive single document available on this topic and deserves treatment on its own terms rather than as a source mined for supporting quotes. The report runs to six chapters plus two supplemental documents, cites 497 references, and was authored by John Butler, Hariharan Iyer, Richard Press, Melissa Taylor, Peter Vallone, and Sheila Willis [NIST, 2024].
Its stated purpose is to document and evaluate the scientific basis for the methods forensic laboratories use to interpret DNA mixtures — not to endorse or condemn any specific commercial product. The report explicitly discusses the likelihood ratio framework and probabilistic genotyping software together, describes the biological background needed to understand mixture measurement, and structures its core reliability and relevance analysis around the hierarchy of propositions, distinguishing sub-source-level conclusions (whose DNA is present) from activity-level conclusions (how it came to be present) [NIST, 2024]. A supplemental document, NISTIR 8351sup1, traces how the field's interpretation methods have evolved over roughly thirty-five years, and a second supplemental document, NISTIR 8351sup2, compiles summarized information from publicly accessible validation studies, proficiency test results, and interlaboratory comparison data [NIST IR 8351sup1 and 8351sup2, 2024].
What the report does not do is as important as what it does. It does not declare probabilistic genotyping software unreliable, and it does not declare any specific commercial package superior to another. Instead, consistent with the broader posture of NIST's scientific foundation review series, it catalogues where the science is well established, where practice still varies across laboratories, and where communication and documentation gaps create real risk of the evidence being misunderstood in court — precisely the balance this article has tried to reflect throughout.
Not every question in this field has a settled answer, and it would misrepresent the literature to pretend otherwise. Several genuine points of disagreement or unresolved practice recur across the sources reviewed for this article.
- Source-code transparency versus trade-secret protection: Washington v. Fair and New Jersey v. Pickett reached opposite conclusions on essentially the same underlying question, and no uniform national or international rule currently exists [cybgen.com, 2017; EFF/Cyberlaw Clinic, 2021].
- How best to communicate likelihood ratios to lay fact-finders: experimental studies comparing random match probabilities, likelihood ratios, and verbal equivalents have not converged on a single presentation format that reliably prevents fallacious reasoning across all evidence types [Prosecutor's Fallacy and Expert Testimony study; Psychonomic Bulletin & Review, 2020].
- The reach of activity-level propositions: the DNA Commission of the ISFG's own guidelines note ongoing practical obstacles — cost, time, the absence of default formulaic computations — to routinely extending probabilistic evaluation beyond source-level questions into activity-level questions of how DNA was transferred [DNA Commission of the ISFG, Forensic Science International: Genetics, 2019].
- Where validation coverage genuinely ends: PCAST's 2016 finding that foundational validity was established for mixtures of up to three contributors with a minor contributor above roughly 20 percent of the intact DNA was explicitly contested by STRmix's developers, who argued the published validation literature already supported a wider range, illustrating that even validation boundaries are argued rather than universally agreed [PCAST, 2016; STRmix developer response, 2016].
Where the literature does show convergence, it is worth naming plainly rather than manufacturing false balance: there is broad consensus among the primary sources reviewed here — NIST, PCAST, ISO 21043-4, and the peer-reviewed validation literature — that continuous probabilistic genotyping represents a genuine methodological advance over purely binary interpretation, that validation is a precondition for reliable use rather than an optional formality, and that the likelihood ratio must never be presented or received as a probability of guilt.
India's evidentiary framework for forensic DNA has recently been reshaped by statute. Under the Bharatiya Sakshya Adhiniyam (BSA), 2023, which replaced the Indian Evidence Act, 1872, forensic DNA analysis is treated as expert opinion evidence under Section 39 — admissible and capable of providing significant corroboration, but not "substantive evidence" that is conclusive in itself [Bharatiya Sakshya Adhiniyam, 2023, Section 39, as discussed in Indian legal commentary]. The Supreme Court of India reinforced this framing in Kattavellai @ Devakar v. State of Tamil Nadu (2025), issuing guidelines on the collection, transport, and chain-of-custody handling of DNA samples, and explicitly reiterating that DNA evidence is opinion evidence whose probative value must be assessed case by case rather than treated as automatically decisive [Kattavellai @ Devakar v. State of Tamil Nadu, Supreme Court of India, 2025].
On the specific question this article is built around — probabilistic genotyping — the available evidence does not support a claim that Indian forensic science laboratories have widely adopted probabilistic genotyping software as a standard casework tool, and this article does not make that claim. The most direct evidence available is critical rather than affirmative. The Forensic Science India Report, produced by Project 39A at the National Law University, Delhi in 2023, reviewed the state of forensic DNA analysis in Indian laboratories and found that the majority of samples received for DNA analysis are mixed samples requiring software-based, probabilistic interpretation — but that software- and statistics-based interpretation of this kind is currently lacking in Indian practice and is not mandated. The same report identified inadequate contamination-control measures, the absence of internal staff-elimination databases to rule out laboratory-personnel contamination, unvalidated internally developed procedure manuals, and significant variability in the DNA profiling kits used across different state laboratories [Forensic Science India Report, Project 39A, National Law University Delhi, 2023, as summarised in a 2024 peer-reviewed article on DNA profiling in India].
This is a materially different starting point from the jurisdictions — principally the United States, the United Kingdom, Australia, and New Zealand — where most of the validation literature, court rulings, and standards discussed earlier in this article were generated. It would be a mistake to assume that debates about the black-box status of STRmix or TrueAllele, or about jury comprehension of likelihood ratios running into the billions, map directly onto Indian courtrooms today, when the more immediate, evidenced gap identified in Indian forensic practice concerns basic mixture-handling infrastructure and the routine use of any probabilistic statistical interpretation at all, rather than the fine-grained transparency of a specific proprietary algorithm.
That said, the direction of India's evidentiary reform is relevant to where this debate is heading domestically. Legal commentary examining the BSA's treatment of forensic genomics has already begun asking whether evidentiary law, drafted with classical STR profiling in mind, is adequate for next-generation sequencing and more advanced statistical methods, and has flagged this as an open question requiring further legislative and judicial attention rather than a settled matter [Criminal Law Blog, National Law University Jodhpur, 2025]. As Indian forensic science laboratories modernise — a process the BSA and the Bharatiya Nagarik Suraksha Sanhita (BNSS) are explicitly intended to support through mandated forensic examination for serious offences — the courtroom-comprehension questions this article raises about likelihood ratios, activity-level propositions, and software transparency are likely to become live issues in India as well, but on the evidence currently available, they remain anticipatory rather than descriptive of present Indian practice.
Framed as a binary — does it work, yes or no — the probabilistic genotyping debate is close to unanswerable in a useful way, because the honest answer is "it depends on conditions that vary case by case and laboratory by laboratory." The more productive set of questions, consistent with how NIST, ISO 21043-4, and the validation literature all implicitly frame the issue, is narrower and more procedural.
- Was the specific software validated for the mixture complexity, contributor number, and template quantity actually present in this case, or is the case pushing beyond the software's demonstrated validation range?
- Were the competing propositions tested by the software appropriately formulated for the actual dispute in the case — source-level, or, where relevant, activity-level?
- Were the model's assumptions about degradation, stutter, and drop-in/drop-out reasonable for the sample type and condition in question?
- Was the resulting likelihood ratio, and the distinction between relative support and probability of guilt, communicated to the court in terms a lay fact-finder could correctly apply?
- Can the result, in principle, be independently scrutinised — through published validation data, expert testimony, or, where a court orders it, direct examination of the underlying methodology?
ISO 21043-4:2025 effectively codifies this procedural framing at an international standards level: it applies both when an opinion is based directly on human judgement and when it is based on a statistical model, and it is built around safeguarding the interpretation process itself — addressing alternative propositions, documenting methods, and ensuring interpretation is fit for its intended use — rather than certifying any particular statistical technique as inherently correct [ISO 21043-4:2025]. That is arguably the right level at which to ask the question: not whether probabilistic genotyping as a category is trustworthy, but whether the specific application, in the specific case, met the specific standards the field has already built for it.
Probabilistic genotyping is neither a scientific breakthrough that resolves the difficulty of interpreting DNA mixtures nor a statistical smokescreen that obscures more than it reveals. The peer-reviewed validation record, PCAST's own 2016 assessment, and NIST's 2024 scientific foundation review converge on a more modest and more defensible position: within its validated range, continuous probabilistic genotyping extracts more of the information present in complex DNA mixture data than earlier binary methods did, and it does so through a transparent, published statistical framework, even where specific commercial implementations remain proprietary at the code level.
That same statistical sophistication is precisely what makes the resulting likelihood ratio harder to explain correctly to a court than a simple inclusion or exclusion ever was. The prosecutor's fallacy predates probabilistic genotyping by decades, but the sheer scale of the numbers this newer generation of software produces — likelihood ratios running into the billions, generated by simulation processes most fact-finders will never see the internal workings of — raises the stakes of getting that explanation right every time an expert takes the stand.
The evidence reviewed here does not support concluding that probabilistic genotyping is simply "good" or simply "bad." It supports a narrower, evidence-anchored conclusion: the value of any specific probabilistic genotyping result depends on the quality of the underlying sample, the appropriateness of the statistical model to that sample, the rigor of the software's validation for the mixture type in question, the reasonableness of the propositions tested, and — perhaps most consequentially for the question this article set out to answer — whether the number that comes out the other end is explained to the court as what it actually is: a measure of relative support for one proposition over another, and never a verdict.
Primary Scientific & Standards Sources
- Butler, J., Iyer, H., Press, R., Taylor, M., Vallone, P. and Willis, S. (2024). DNA Mixture Interpretation: A NIST Scientific Foundation Review. NIST Interagency Report (NISTIR) 8351. National Institute of Standards and Technology, Gaithersburg, MD. https://doi.org/10.6028/NIST.IR.8351
- Butler, J. (2024). History of DNA Mixture Interpretation. NISTIR 8351sup1, Supplemental Document to NIST IR 8351. https://doi.org/10.6028/NIST.IR.8351sup1
- Butler, J. (2024). Summarized Information from Publicly Accessible Validation Studies, Proficiency Test Results, and Interlaboratory Comparison Data. NISTIR 8351sup2, Supplemental Document to NIST IR 8351. https://doi.org/10.6028/NIST.IR.8351sup2
- International Organization for Standardization. (2025). ISO 21043-4:2025 — Forensic sciences — Part 4: Interpretation. Edition 1. https://www.iso.org/standard/72039.html
- President's Council of Advisors on Science and Technology (PCAST). (2016). Report to the President: Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods. Executive Office of the President, Washington, DC.
- Scientific Working Group on DNA Analysis Methods (SWGDAM). (2015). Guidelines for the Validation of Probabilistic Genotyping Systems.
Peer-Reviewed Research
- Bille, T., Coble, M.D., Kalafut, T. and Buckleton, J. (2022). Study of CTS DNA Proficiency Tests with Regard to DNA Mixture Interpretation: A NIST Scientific Foundation Review. Genes, 13(11), 2171. https://doi.org/10.3390/genes13112171
- Bright, J.-A. et al. (2016). Developmental validation of STRmix™, expert software for the interpretation of forensic DNA profiles. Forensic Science International: Genetics.
- Bright, J.-A. et al. (2018). Internal validation of STRmix™ — A multi-laboratory response to PCAST. Forensic Science International: Genetics.
- Adamowicz, M.S., Rambo, T.N. and Clarke, J.L. (2022). Internal Validation of MaSTR™ Probabilistic Genotyping Software for the Interpretation of 2–5 Person Mixed DNA Profiles. Genes, 13(8), 1429. https://doi.org/10.3390/genes13081429
- Cook, R., Evett, I.W., Jackson, G., Jones, P.J. and Lambert, J.A. (1998). A hierarchy of propositions: deciding which level to address in casework. Science & Justice, 38, 231–240.
- Evett, I.W., Gill, P.D., Jackson, G., Whitaker, J. and Champod, C. (2002). Interpreting small quantities of DNA: the hierarchy of propositions and the use of Bayesian networks. Journal of Forensic Sciences, 47(3), 520–530.
- DNA Commission of the International Society for Forensic Genetics (ISFG). (2018–2019). Assessing the value of forensic biological evidence — Guidelines highlighting the importance of propositions, Parts I & II. Forensic Science International: Genetics.
- Psychonomic Bulletin & Review (2020). Does DNA evidence in the form of a likelihood ratio affect perceivers' sensitivity to the strength of a suspect's alibi? Springer Nature.
- Forensic Science International: Synergy (2025). Finally a really forensic worldwide standard: ISO 21043 Forensic sciences, Part 4, Interpretation. https://doi.org/10.1016/j.fsisyn.2025.100589
- Sciencedirect / Science & Justice (2024). DNA profiling in India: Addressing issues of sample preservation, databasing, marker selection, and statistical approaches.
Legal / Judicial / Evidence Sources
- Bharatiya Sakshya Adhiniyam, 2023, Section 39 (Government of India).
- Kattavellai @ Devakar v. State of Tamil Nadu (2025), Supreme Court of India — guidelines on DNA evidence collection, custody, and evaluation.
- R v Deen, Court of Appeal (England and Wales) — discussed in Open University, "Expert evidence and forensic science in the courtroom."
- Washington v. Fair, King County Superior Court (2017) — TrueAllele admissibility and source-code disclosure ruling, reported by Cybergenetics.
- New Jersey v. Pickett, Superior Court of New Jersey, Appellate Division (2021) — source-code disclosure ruling, reported by the Electronic Frontier Foundation and the Cyberlaw Clinic, Harvard Law School.
Additional Authoritative Sources
- Electronic Frontier Foundation (2017, 2020, 2021). Press releases and commentary on TrueAllele source-code disclosure litigation.
- Cyberlaw Clinic, Harvard Law School (2021). "Victory for Transparency in Probabilistic Genotyping Case."
- National Institute of Justice (NIJ). Post-PCAST Court Decisions Assessing the Admissibility of Forensic Science Evidence.
- Criminal Law Blog, National Law University Jodhpur (2025). Admissibility of DNA evidence beyond DNA profiling: Is evidence procured using next-generation sequencing admissible?
- Forensic Science India Report (2023). Project 39A, National Law University, Delhi.
The most load-bearing sources for this article’s claims were: NIST Interagency Report 8351, DNA Mixture Interpretation: A NIST Scientific Foundation Review (2024); ISO 21043-4:2025, Forensic sciences — Part 4: Interpretation; the 2016 PCAST Report to the President on forensic feature-comparison methods; peer-reviewed developmental and internal validation studies for STRmix and MaSTR published in Forensic Science International: Genetics and Genes; the foundational hierarchy-of-propositions literature (Cook, Evett, Jackson et al.; Evett et al.); court rulings in Washington v. Fair and New Jersey v. Pickett on probabilistic genotyping source-code disclosure; India’s Bharatiya Sakshya Adhiniyam, 2023 and the Supreme Court’s 2025 ruling in Kattavellai @ Devakar v. State of Tamil Nadu; and the 2023 Forensic Science India Report produced by Project 39A, National Law University, Delhi.

