Abruquah v. State
Kobina Ebo Abruquah v. State of Maryland, No. 10, September Term, 2022. Opinion by Fader, C.J. EVIDENCE – EXPERT EVIDENCE Firearms identification examiner testifying as an expert witness should not have been permitted to offer an unqualified opinion that crime scene bullets and a bullet fragment were fired from the petitioner’s gun. The reports, studies, and testimony presented to the circuit court demonstrate that the firearms identification methodology employed by the examiner in this case can support reliable conclusions that patterns and markings on bullets are consistent or inconsistent with those on bullets fired from a particular known firearm. Those reports, studies, and testimony do not, however, demonstrate that the methodology used can reliably support an unqualified conclusion that such bullets were fired from a particular firearm.
Circuit Court for Prince George’s County Case No. CT121375X Argued: October 4, 2022 IN THE SUPREME COURT OF MARYLAND* No. 10 September Term, 2022 ______________________________________ KOBINA EBO ABRUQUAH v. STATE OF MARYLAND ______________________________________ Fader, C.J., Watts, Hotten, Booth, Biran, Gould, Eaves, JJ. ______________________________________ Pursuant to the Maryland Uniform Electronic Legal Materials Act (§§ 10-1601 et seq. of the State Government Article) this document is authentic. Opinion by Fader, C.J. Hotten, Gould, and Eaves, JJ., dissent. 2023-07-14 09:00-04:00 ______________________________________ Filed: June 20, 2023 Gregory Hilton, Clerk * At the November 8, 2022 general election, the voters of Maryland ratified a constitutional amendment changing the name of the Court of Appeals of Maryland to the Supreme Court of Maryland. The name change took effect on December 14, 2022. Firearms identification, a subset of toolmark identification, is “the practice of investigating whether a bullet, cartridge case or other ammunition component or fragment can be traced to a particular suspect weapon.” Fleming v. State, 194 Md. App. 76, 100-01 (2010).
The basic idea is that (1) features unique to the interior of any particular firearm leave unique, microscopic patterns and marks on bullets and cartridge cases that are fired from that firearm, and so (2) by comparing patterns and marks left on bullets and cartridge cases found at a crime scene (“unknown samples”) to marks left on bullets and cartridge cases fired from a known firearm (“known samples”), firearms examiners can determine whether the unknown samples were or were not fired from the known firearm. At the trial of the petitioner, Kobina Ebo Abruquah, the Circuit Court for Prince George’s County permitted a firearms examiner to testify, without qualification, that bullets left at a murder scene were fired from a gun that Mr. Abruquah had acknowledged was his. Based on reports, studies, and testimony calling into question the reliability of firearms identification analysis, Mr. Abruquah contends that the circuit court abused its discretion in permitting the firearms examiner’s testimony. The State, relying on different studies and testimony, contends that the examiner’s opinion was properly admitted.
Applying the analysis required by Rochkind v. Stevenson, 471 Md. 1 (2020), we conclude that the examiner should not have been permitted to offer an unqualified opinion that the crime scene bullets were fired from Mr. Abruquah’s gun. The reports, studies, and testimony presented to the circuit court demonstrate that the firearms identification methodology employed in this case can support reliable conclusions that patterns and markings on bullets are consistent or inconsistent with those on bullets fired from a particular firearm. Those reports, studies, and testimony do not, however, demonstrate that that methodology can reliably support an unqualified conclusion that such bullets were fired from a particular firearm. The State also contends that any error in the circuit court’s admission of the examiner’s testimony was harmless.
Because we are not convinced “beyond a reasonable doubt, that the error in no way influenced the verdict,” Dionas v. State, 436 Md. 97, 108 (2013) (quoting Dorsey v. State, 276 Md. 638, 659 (1976)), we must reverse and remand for a new trial. BACKGROUND Factual Background On August 3, 2012, police responded to three separate calls complaining of disturbances at the house that Mr. Abruquah shared with his roommate, Ivan Aguirre- Herrera. On the third of these occasions, just before midnight, two officers arrived at the house. According to the officers, Mr. Abruquah appeared “agitated,” “very aggressive,” and uncooperative.
One of the officers testified that Mr. Aguirre-Herrera appeared to be terrified of Mr. Abruquah. Before leaving around 12:15 a.m., the officers told the men to stay away from each other. A neighbor of Messrs. Abruquah and Aguirre-Herrera testified that he heard multiple gunshots sometime between 11:30 p.m. on August 3 and 12:30 a.m. on August 4.
Four days later, officers discovered Mr. Aguirre-Herrera’s body decomposing in his bedroom. An autopsy revealed that he had been shot five times, including once in the back 2 of the head. The police recovered four bullets and two bullet fragments from the crime scene. During questioning, Mr. Abruquah told the police that he owned two firearms, both hidden in the ceiling of the basement of the residence he shared with Mr. Aguirre-Herrera.
The police recovered both firearms, a Glock pistol and a Taurus .38 Special revolver. A jailhouse informant testified that Mr. Abruquah had said that he had engaged in “a heated argument” with Mr. Aguirre-Herrera, “snapped,” and shot him with “a 38” that he kept in the ceiling of his basement.1 Procedural Background Mr. Abruquah was convicted by a jury of first-degree murder and related handgun offenses in December 2013. Abruquah v. State, No. 246, Sept. Term 2014, 2016 WL 7496174 , at 1 & n.1 (Md. App. Dec. 20, 2016). In an unreported opinion, the Appellate Court of Maryland (then named the Court of Special Appeals)2 reversed the judgment and remanded the case for a new trial on grounds that are not relevant to the current appeal.
Id. at 9. On remand, Mr. Abruquah filed a motion in limine to exclude firearms identification evidence the State intended to offer through its expert witness, Scott McVeigh, a senior firearms examiner with the Firearms Examination Unit of the Prince George’s County The jailhouse informant testified at Mr. Abruquah’s first trial in 2013. At his 1 second trial, in 2018, the State read into the record a transcript of that prior testimony. 2 At the November 8, 2022 general election, the voters of Maryland ratified a constitutional amendment changing the name of the Court of Special Appeals of Maryland to the Appellate Court of Maryland. The name change took effect on December 14, 2022. 3 Police Department, Forensic Science Division.
The circuit court held a four-day Frye- Reed hearing3 during which both parties introduced evidence and elicited testimony that we summarize below. Following the hearing, the circuit court largely denied, but partially granted, the motion. The court concluded that “firearm and toolmark identification is still generally accepted and sufficiently reliable under the Frye-Reed standard” and therefore should not be “excluded in its entirety.” Nonetheless, the court agreed with Mr. Abruquah that the subjective nature of the matching analysis made it inappropriate for an expert to “testify to any level of practical certainty/impossibility, ballistic certainty, or scientific certainty that a suspect weapon matches certain bullet or casing striations.” The court thus restricted the expert to opining whether the bullets and bullet fragment “recovered from the murder scene fall into any of” a particular set of five classifications, one of which is “[i]dentification” of the unknown bullet as a match to a known bullet. At trial, Mr. McVeigh testified about the process by which he eliminated the Glock pistol as a source of the unknown crime scene samples, created known samples from the Taurus revolver, and compared the microscopic patterns and markings on the two sets of samples.
Over defense objection, Mr. McVeigh opined that four bullets and one bullet 3 Prior to our decision in Rochkind v. Stevenson, 471 Md. 1 (2020), courts in Maryland determined the admissibility of expert testimony using the Frye-Reed evidentiary standard, which “turned on the ‘general acceptance’ of such evidence ‘in the particular field in which it belongs.’” Rochkind, 471 Md. at 4 (discussing Frye v. United States, 293 F. 1013 (D.C. Cir. 1923) and Reed v. State, 283 Md. 374 (1978)). 4 fragment recovered from the crime scene “at some point had been fired from [the Taurus revolver].”4 Mr. Abruquah was again convicted of first-degree murder and use of a handgun in the commission of a crime. His first appeal from that conviction resulted in a remand to the circuit court to consider whether it “would reach a different conclusion concerning the admission of firearm and toolmark identification testimony” applying our then-new decision in Rochkind v. Stevenson, 471 Md. 1, 27 (2020). In that decision, which was issued after Mr. Abruquah’s second conviction while his appeal was pending, we abandoned the Frye-Reed standard for admissibility of expert testimony in favor of the standard set forth in Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579 (1993), and its progeny. Abruquah v. State, 471 Md. 249, 250 (2020).
On remand, the circuit court held a hearing in which it once again received evidence from both sides, which is discussed further below. The court ultimately issued an opinion in which it reviewed each of the ten factors this Court set forth in Rochkind and concluded that the testimony remained admissible. The court noted that although Mr. Abruquah “ha[d] made a Herculean effort to demonstrate why the evidence should be heavily scrutinized, questioned and potentially impeached, the State has met the burden for admissibility of this evidence.” The court therefore sustained Mr. Abruquah’s prior conviction. 4 Four bullets and two bullet fragments were recovered from the crime scene but Mr. McVeigh found that one of the fragments was not suitable for comparison. As a result, his testimony was limited to the bullets and one of the fragments. 5 Mr. Abruquah filed another timely appeal to the intermediate appellate court and, while that appeal was pending, he filed a petition for writ of certiorari in this Court.
We granted that petition to address whether the firearms identification methodology employed by Mr. McVeigh is sufficiently reliable to allow a firearms examiner, without any qualification, to identify a specific firearm as the source of a questioned bullet or cartridge case found at a crime scene. See Abruquah v. State, 479 Md. 63 (2022). DISCUSSION We review a circuit court’s decision to admit expert testimony for an abuse of discretion. Rochkind, 471 Md. at 10 .
Under that standard, we will “not reverse simply because . . . [we] would not have made the same ruling.” State v. Matthews, 479 Md. 278, 305 (2022) (quoting Devincentz v. State, 460 Md. 518, 550 (2018)). In connection with the admission of expert testimony, where circuit courts are to act as gatekeepers in applying the factors set out by this Court in Rochkind, a circuit court abuses its discretion by, for example, admitting expert evidence where there is an analytical gap between the type of evidence the methodology can reliably support and the evidence offered.5 See Rochkind, 471 Md. at 26-27 . 5 This Court has frequently described an abuse of discretion as occurring when “no reasonable person would take the view adopted by the circuit court” or when a decision is “well removed from any center mark imagined by the reviewing court and beyond the fringe of what the court deems minimally acceptable.” Matthews, 479 Md. at 305 (first quoting Williams v. State, 457 Md. 551, 563 (2018), and next quoting Devincentz v. State, 460 Md. 518, 550 (2018)). In our view, the application of those descriptions to a trial court’s application of a newly adopted standard, such as that adopted by this Court in Rochkind as applicable to the admissibility of expert testimony, is somewhat unfair. In this case, in the absence of additional caselaw from this Court implementing the newly adopted standard, the circuit court acted deliberately and thoughtfully in approaching, analyzing, 6 Part I of our discussion sets forth the standard for the admissibility of expert testimony in Maryland following this Court’s decision in Rochkind v. Stevenson, 471 Md. 1 (2020).
In Part II, we discuss general background on the firearms identification methodology employed by the State’s expert witness, criticisms of that methodology, studies of the methodology, the testimony presented to the circuit court, and caselaw from other jurisdictions. In Part III, we apply the factors set forth in Rochkind to the evidence before the circuit court. I. THE ADMISSIBILITY OF EXPERT TESTIMONY The admissibility of expert testimony is governed by Rule 5-702, which provides: Expert testimony may be admitted, in the form of an opinion or otherwise, if the court determines that the testimony will assist the trier of fact to understand the evidence or to determine a fact in issue. In making that determination, the court shall determine (1) whether the witness is qualified as an expert by knowledge, skill, experience, training, or education, (2) the appropriateness of the expert testimony on the particular subject, and (3) whether a sufficient factual basis exists to support the expert testimony.
Trial courts analyzing the admissibility of evidence under Rule 5-702 are to consider the following non-exhaustive list of “factors in determining whether the proffered expert testimony is sufficiently reliable to be provided to the trier of facts,” Matthews, 479 Md. at 310 : and resolving the question before it. This Court’s majority has come to a different conclusion concerning the outer bounds of what is acceptable expert evidence in this area. 7 (1) whether a theory or technique can be (and has been) tested; (2) whether a theory or technique has been subjected to peer review and publication; (3) whether a particular scientific technique has a known or potential rate of error; (4) the existence and maintenance of standards and controls; . . . (5) whether a theory or technique is generally accepted[;] ... (6) whether experts are proposing to testify about matters growing naturally and directly out of research they have conducted independent of the litigation, or whether they have developed their opinions expressly for purposes of testifying; (7) whether the expert has unjustifiably extrapolated from an accepted premise to an unfounded conclusion; (8) whether the expert has adequately accounted for obvious alternative explanations; (9) whether the expert is being as careful as [the expert] would be in [the expert’s] regular professional work outside [the expert’s] paid litigation consulting; and (10) whether the field of expertise claimed by the expert is known to reach reliable results for the type of opinion the expert would give.
Rochkind, 471 Md. at 35-36 (first quoting Daubert, 509 U.S. at 593-94 (for factors 1-5) and next quoting Fed. R. Evid. 702 Advisory Committee Note (cleaned up) (for factors 6- 10)). In applying these “Daubert-Rochkind factors,” we have observed that the guidance provided by the United States Supreme Court in Daubert and its progeny, especially General Electric Co. v. Joiner, 522 U.S. 136 (1997), and Kumho Tire Co. v. Carmichael, 8 526 U.S. 137 (1999), “is critical to a trial court’s reliability analysis.” Rochkind, 471 Md. at 36 . In Matthews, we summarized that guidance in five principles: • “[T]he reliability inquiry is ‘a flexible one.’” Matthews, 479 Md. at 311 (quoting Rochkind, 471 Md. at 36 ). • “[T]he trial court must focus solely on principles and methodology, not on the conclusions that they generate. However, conclusions and methodology are not entirely distinct from one another.
Thus, [a] trial court . . . must consider the relationship between the methodology applied and conclusion reached.” Id. (internal citations and quotation marks omitted). • “[A] trial court need not admit opinion evidence that is connected to existing data only by the ipse dixit of the expert; rather, [a] court may conclude that there is simply too great an analytical gap between the data and the opinion proffered.” Id. (internal quotation marks omitted). • “[A]ll of the Daubert factors are relevant to determining the reliability of expert testimony, yet no single factor is dispositive in the analysis. A trial court may apply some, all, or none of the factors depending on the particular expert testimony at issue.” Id. at 37 . • “Rochkind did not upend [the] trial court’s gatekeeping function.
Vigorous cross-examination, presentation of contrary evidence, and careful instruction on the burden of proof are the traditional and appropriate means of attacking shaky but admissible evidence.” Id. at 38 (internal quotation marks omitted). The overarching criterion for the admission of relevant expert testimony under Rochkind, and the goal to which each of the ten Daubert-Rochkind factors and the five principles summarized in Matthews are all addressed, is reliability. The question for a trial court is not whether proposed expert testimony is right or wrong, but whether it meets a minimum threshold of reliability so that it may be presented to a jury, where it may then be questioned, tested, and attacked through means such as cross-examination or the submission of opposing expert testimony. Because we evaluate a trial court’s decision to admit or exclude expert testimony under an abuse of discretion standard, our review is necessarily limited to the information 9 that was before the trial court at the time it made the decision.
A trial court can hardly abuse its discretion in failing to consider evidence that was not before it.6 II. FIREARMS IDENTIFICATION EVIDENCE Through multiple submissions by the parties and two evidentiary hearings over the course of five days, the circuit court ultimately received the testimony of five witnesses (one twice); 18 reports or articles discussing firearms identification, describing studies testing firearms identification, or criticizing the theory or the results of the studies testing it; and a chart identifying dozens of additional or planned studies or reports. In section A of this Part II, we discuss firearms identification evidence generally. In sections B and C, we review criticisms and studies of the methodology, respectively.
In section D, we summarize the testimony presented to the circuit court. Finally, in section E, we discuss how some other courts have resolved challenges to the admissibility of firearms identification evidence. 6 On appeal, the State cited articles presenting the results of studies that were not presented to the circuit court and, in some cases, that were not even in existence at the time the circuit court ruled. See, e.g., Maddisen Neuman et al., Blind Testing in Firearms: Preliminary Results from a Blind Quality Control Program, 67 J. Forensic Scis. 964 (2022); Eric F. Law & Keith B. Morris, Evaluating Firearm Examiner Conclusion Variability Using Cartridge Case Reproductions, 66:5 J. Forensic Scis. 1704 (2021). We have not considered those studies in reaching our decision.
If any of those studies materially alters the analysis applicable to the reliability of the Association of Firearm and Tool Mark Examiners theory of firearms identification, they will need to be presented in another case. 10 A. Firearms Identification 1. The Theory Underlying Firearms Identification Generally Firearms identification is a subset of toolmark identification. A toolmark—literally, a mark left by a particular tool—is “generated when a hard object (tool) comes into contact with a relatively softer object,” such as the marks that result “when the internal parts of a firearm make contact with the brass and lead that comprise ammunition.” United States v. Willock, 696 F. Supp. 2d 536, 555 (D. Md. 2010) (quoting Nat’l Rsch. Council, Nat’l Acad. of Scis., Strengthening Forensic Science in the United States: A Path Forward 150 (2009)), aff’d sub nom.
United States v. Mouzone, 687 F.3d 207 (4th Cir. 2012). The marks are then viewable using a “comparison microscope,” which a firearms examiner uses “to compare ammunition test-fired from a recovered gun with spent ammunition from a crime scene[.]” United States v. Monteiro, 407 F. Supp. 2d 351, 359 (D. Mass. 2006). As a forensic technique to identify a particular firearm as the source of a particular ammunition component, firearms identification is based on the premise that no two firearms will make identical marks on a bullet or cartridge case. United States v. Natson, 469 F. Supp. 2d 1253, 1260 (M.D. Ga. 2007).
That, the theory goes, is because the method of manufacturing firearms results in the interior of each firearm being unique and, therefore, making unique imprints on ammunition components fired from it. Id. As the United States District Court for the District of Massachusetts explained: When a firearm is manufactured, the “process of cutting, drilling, grinding, hand-filing, and, very occasionally, hand-polishing . . . will leave individual characteristics” on the components of the firearm. See Brian J. Heard, Handbook of Firearms and Ballistics 127 (1997).
Although modern manufacturing methods have reduced the amount of handiwork performed 11 on an individual gun, the final step in production of most firearm parts requires some degree of hand-filing which imparts individual characteristics to the firearm part. See id. at 128. This process results in “randomly produced patterns of individual stria,” or thin grooves or markings, being left on firearm parts. Id.
These parts are assembled to compose the final firearm. When a round (a single “shot”) of ammunition is fired from a particular firearm, the various components of the ammunition come into contact with the firearm at very high pressures. As a result, the individual markings on the firearm parts are transferred to the ammunition. Id.
The ammunition is composed primarily of the bullet and the cartridge case. The bullet is the missile-like component of the ammunition that is actually projected from the firearm, through the barrel, toward the target. . . . The cartridge case is the part of the ammunition situated behind the bullet containing the primer and propellant, the explosive mixture of chemicals that causes the bullet to be projected through the barrel. Id. at 42.
Monteiro, 407 F. Supp. 2d at 359-60 . The patterns and marks left on bullets and cartridge cases are classified into three categories. First, “class characteristics” are common to all bullets and cartridge cases fired from “weapons of the make and model that fired the ammunition.” Willock, 696 F. Supp. 2d at 557-58 . “Examples of class characteristics include the bullet’s weight and caliber; number and width of the lands and grooves in the gun’s barrel; and the ‘twist’ (direction of turn, i.e., clockwise or counterclockwise, of the rifling in the barrel).”7 Id. at 558 . Second, “subclass characteristics” are common to “a group of guns within a certain make or model, such as those manufactured at a particular time and place.” Monteiro, 407 7 “Rifling” refers to “a pattern of channels that run the length of a firearm barrel, manufactured with a helical pattern, or twist,” which has raised areas called “lands,” and lowered areas called “grooves.” Ass’n of Firearms & Tool Mark Exam’rs, What Is Firearm and Tool Mark Identification?, available at https://afte.org/about-us/what-is-afte/what-is- firearm-and-tool-mark-identification (last accessed June 14, 2023), archived at https://perma.cc/UYA4-99CS. “The number and width of lands and grooves is determined by the manufacturer and will be the same for a large group of firearms.” Id. 12 F. Supp. 2d at 360. “An example would include imperfections ‘on a rifling tool that imparts similar toolmarks on a number of barrels before being modified either through use or refinishing.’” Willock, 696 F. Supp. 2d at 558 (quoting Ronald G. Nichols, Defending the Scientific Foundations of the Firearms and Tool Mark Identification Discipline: Responding to Recent Challenges, 52 J. Forensic Scis. 586, 587 (2007)).
Third, “individual characteristics” are those unique to an individual firearm that therefore “distinguish [the firearm] from all others.” Willock, 696 F. Supp. 2d at 558 (quoting Monteiro, 407 F. Supp. 2d at 360 ). Individual characteristics include “[r]andom imperfections produced during manufacture or caused by accidental damage.” Id. Notably, not all individual characteristics are unique, Willock, 696 F. Supp. 2d at 558 , and individual characteristics can change over the life of a firearm as a result of, for example, wear, polishing, or damage. As will be discussed further below, one dispute between proponents of firearms identification and its detractors is the degree to which firearms examiners can reliably identify the difference between subclass and individual characteristics when performing casework. 2.
The Association of Firearm and Tool Mark Examiners Methodology The leading methodology used by firearms examiners, and the methodology employed in this case by Mr. McVeigh, is the Association of Firearm and Tool Mark Examiners (“AFTE”) “Theory of Identification” (the “AFTE Theory”).8 See Committee 8 According to its website, the AFTE “is the international professional organization for practitioners of Firearm and/or Toolmark Identification and has been dedicated to the exchange of information, methods and best practices, and the furtherance of research since 13 for the Advancement of the Science of Firearm & Toolmark Identification, Theory of Identification as it Relates to Toolmarks: Revised, 43 AFTE J. 287 (2011). Examiners employing the AFTE Theory follow a two-step process. At step one, the examiner evaluates class characteristics of the unknown and known samples. See AFTE, Summary of the Examination Method, available at https://afte.org/resources/swggun-ark/summary- of-the-examination-method (last accessed June 14, 2023), archived at https://perma.cc/4D8W-UDW9.
If the class characteristics do not match—i.e., if the samples have different numbers of lands and grooves or a different twist direction—the firearm that produced the known sample is excluded as the source of the unknown sample. Id. If the class characteristics match, the second step involves “a comparative examination . . . utilizing a comparison microscope.” Id. At that step, the examiner engages in “pattern matching” “to determine: 1) if any marks present are subclass characteristics and/or individual characteristics, and 2) the level of correspondence of any individual characteristics.”9 Id. its creation in 1969.” AFTE, What is AFTE?, available at https://afte.org/about-us/what- is-afte (last accessed June 14, 2023), archived at https://perma.cc/4VKT-EZW7.
According to AFTE’s bylaws, individuals are eligible to become members if they are, among other things, “a practicing firearm and/or toolmark examiner,” which is defined to mean a person who “derives a substantial portion of their livelihood from the examination, identification, and evaluation of firearms and related materials and/or toolmarks; or an individual whose present livelihood is a direct result of the knowledge and experience gained from the examination, identification, and evaluation of firearms and related materials and/or toolmarks.” AFTE, AFTE Bylaws, Art. III, § 1, available at https://afte.org/about-us/bylaws (last accessed June 14, 2023), archived at https://perma.cc/Y2PF-XWUF. 9 An alternative to the AFTE method is the “consecutive matching striae method of toolmark analysis” (“CMS”). Fleming, 194 Md. App. at 105 . “The CMS method . . . calls 14 Based on that “pattern matching,” the examiner makes a determination in accordance with the “AFTE Range of Conclusions,” which presents the following options: 1. “Identification” occurs when there is “[a]greement of a combination of individual characteristics and all discernible class characteristics where the extent of agreement exceeds that which can occur in the comparison of toolmarks made by different tools and is consistent with the agreement demonstrated by toolmarks known to have been produced by the same tool.” 2. There are three categories of “Inconclusive,” all of which require full agreement of “all discernible class characteristics”: (a) when there is “[s]ome agreement of individual characteristics . . . but insufficient for an identification”; (b) when there is neither “agreement [n]or disagreement of individual characteristics”; and (c) when there is “disagreement of individual characteristics, but insufficient for an elimination.” 3. “Elimination” occurs when there is “[s]ignificant disagreement of discernible class characteristics and/or individual characteristics.” AFTE, Range of Conclusions, available at https://afte.org/about-us/what-is-afte/afte- range-of-conclusions (last accessed June 14, 2023), archived at https://perma.cc/WKF5- M6HD. According to the AFTE, a positive “Identification” can be made when there is “sufficient agreement” between “two or more sets of surface contour patterns” on samples.
AFTE, AFTE Theory of Identification as It Relates to Toolmarks, available at for the examiner to consider the number of consecutive matching striae, or ‘scratches’ appearing on a projectile fragment. The theory provides that a positive ‘match’ determination can be made only when a certain, statistically established number of striae match.” Id. Proponents of the CMS method argue that it has a “greater degree of objective certainty” than other methods. Id.
The CMS method was not used in this case. 15 https://afte.org/about-us/what-is-afte/afte-theory-of-identification (last accessed June 14, 2023), archived at https://perma.cc/E397-U8KM. “[S]ufficient agreement,” in turn: (1) occurs when the level of agreement “exceeds the best agreement demonstrated between toolmarks known to have been produced by different tools and is consistent with agreement demonstrated by toolmarks known to have been produced by the same tool”; and (2) means that “the agreement of individual characteristics is of a quantity and quality that the likelihood another tool could have made the mark is so remote as to be considered a practical impossibility.” Id. The AFTE acknowledges that “[c]urrently the interpretation of individualization/identification is subjective in nature[.]” Id. The AFTE Theory provides no objective criteria to determine what constitutes the “best agreement demonstrated” between toolmarks produced by different tools or what rises to the level of “quantity and quality” of agreement demonstrating a “practical impossibility” of a different tool having made the same mark. There are also no established standards for classifying a particular pattern or mark as a subclass versus an individual characteristic.
B. Critiques of Firearms Identification Firearms identification has existed as a field for more than a century.10 Throughout most of that time, it has been accepted by law enforcement organizations and courts without 10 The first prominent use of firearms identification in the United States is attributed to examinations made in the aftermath of the 1906 race-related incident in Brownsville, Texas, known as the “Brownsville Affair.” There, Army personnel matched 39 out of 45 cartridge cases to two types of rifles “through the use of only magnified photographs of firing pin impressions[.]” Kathryn E. Carso, Amending the Illinois Postconviction Statute to Include Ballistics Testing, 56 DePaul L. Rev. 695 , 700 n.43 (2007). 16 significant challenge. However, the advent of Daubert, work exposing the unreliability of other previously accepted forensic techniques,11 and recent reports questioning the foundations underlying firearms identification have led to greater skepticism. Reports issued since 2008 by two blue-ribbon groups of experts outside of the firearms and toolmark identification field have been critical of the AFTE Theory. In 2008, the National Research Council of the National Academies of Science (the “NRC”) published a report concerning the feasibility of developing a national database of ballistic images to aid in criminal investigations.
National Research Council, National Academy of Sciences, Committee to Assess the Feasibility, Accuracy, and Technical Capability of a National Ballistics Database, Ballistic Imaging 1-2 (2008), available at https://nap.nationalacademies.org/read/12162/chapter/1 (last accessed June 14, 2023), archived at https://perma.cc/X6NG-BNVN. In the report, the committee identified challenges that complicate firearms identifications, and ultimately determined that the creation of a national ballistic image database was not advisable at the time. Id. at 4-5. 11 For example, comparative bullet lead analysis was initially widely accepted within the scientific and legal community, and admitted successfully in criminal prosecutions nationwide, yet its validity was subsequently undermined and such evidence is now inadmissible. See Chesson v. Montgomery Mut.
Ins. Co., 434 Md. 346, 358-59 (2013) (stating that, despite the expert’s “use of th[e] technique for thirty years,” comparative bullet lead analysis evidence was inadmissible because its “general and underlying assumption . . . was no longer generally accepted by the relevant scientific community”); Clemons v. State, 392 Md. 339, 364-72 (2006) (comprehensively discussing comparative bullet lead analysis and holding that it does not satisfy Frye-Reed); Sissoko v. State, 236 Md. App. 676, 721-27 (2018) (discussing that the “methodology underlying [comparative bullet lead analysis], which was developed in the 1960s and became a widely accepted forensic tool by the 1980s[,] . . . [was] undermined by many in the relevant scientific community” and was “no longer . . . ‘valid and reliable’” (quoting Clemons v. State, 392 Md. 339, 359 (2006))). 17 Then, in 2009, the NRC published a report in which it addressed “pressing issues” within several forensic science disciplines, including firearms identification. National Research Council, National Academy of Sciences, Strengthening Forensic Science in the United States: A Path Forward 2-5 (2009) (the “2009 NRC Report”), available at https://www.ojp.gov/pdffiles1/nij/grants/228091.pdf (last accessed June 14, 2023), archived at https://perma.cc/RLT6-49C3.12 The NRC observed that advances in DNA evidence had revealed flaws in other forensic science disciplines that “may have contributed to wrongful convictions of innocent people,” id. at 4, and pointed especially to the relative “dearth of peer-reviewed, published studies establishing the scientific bases and validity of many forensic methods,” id. at 8. With respect to firearms identification specifically, the NRC criticized the AFTE Theory as lacking specificity in its protocols; producing results that are not shown to be accurate, repeatable, and reproducible; lacking databases and imaging that could improve the method; having deficiencies in proficiency training; and requiring examiners to offer opinions based on their own experiences without articulated standards.
Id. at 6, 63-64, 155. In particular, the lack of knowledge “about the variabilities among individual tools and guns” means that there is an inability of examiners “to specify how many points of similarity are necessary for a given level of confidence in the result.” Id. at 154. Indeed, the NRC noted, the AFTE’s guidance, which is the “best . . . available for the field of 12 The lead NRC “Committee” behind the report was the “Committee on Identifying the Needs of the Forensic Science Community.” The committee was co-chaired by Judge Harry T. Edwards of the United States Court of Appeals for the District of Columbia Circuit and included members from a variety of distinguished academic and scientific programs. 18 toolmark identification, does not even consider, let alone address, questions regarding variability, reliability, repeatability, or the number of correlations needed to achieve a given degree of confidence.” Id. at 155. The NRC concluded that “[t]he validity of the fundamental assumptions of uniqueness and reproducibility of firearms-related toolmarks has not yet been fully demonstrated.” Id. at 70, 80-81, 154-55 (citation omitted).
In 2016, the President’s Council of Advisors on Science and Technology (“PCAST”)13 issued a report identifying additional concerns about the scientific validity of, among other forensic techniques, firearms identification. See Executive Office of the President, President’s Council of Advisors on Science and Technology, REPORT TO THE PRESIDENT, Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods (2016) (the “PCAST Report”), available at https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/PCAST/pcast_fo 13 The PCAST Report provides the following description of PCAST’s role: The President’s Council of Advisors on Science and Technology (PCAST) is an advisory group of the Nation’s leading scientists and engineers, appointed by the President to augment the science and technology advice available to him from inside the White House and from cabinet departments and other Federal agencies. PCAST is consulted about, and often makes policy recommendations concerning, the full range of issues where understandings from the domains of science, technology, and innovation bear potentially on the policy choices before the President. PCAST Report at iv.
Members of PCAST included scholars and senior executives at institutions and firms including Harvard University; the University of Texas at Austin, Honeywell; Princeton University; the University of Maryland; the University of Michigan; the University of California, Berkeley; United Technologies Corporation; Washington University of St. Louis; Alphabet, Inc.; Northwestern University; and the University of California, San Diego. Id. at v-vi. PCAST also consulted with “Senior Advisors” including eight federal appellate and trial court judges, as well as law school and university professors. Id. at viii-ix. 19 rensic_science_report_final.pdf (last accessed June 14, 2023), archived at https://perma.cc/3QWJ-2DGR.
With respect to all six forensic disciplines addressed in the report, including firearms identification, PCAST focused on whether there had been a demonstration of both “foundational validity” and “validity as applied.” Id. at 4-5. Foundational validity, according to PCAST, requires that the method “be shown, based on empirical studies, to be repeatable, reproducible, and accurate, at levels that have been measured and are appropriate to the intended application.” Id. Validity as applied requires “that the method has been reliably applied in practice.” Id. at 5. With respect to firearms identification specifically, PCAST described the AFTE Theory as a “circular” method that lacks “foundational validity” because appropriate studies had not confirmed its accuracy, repeatability, and reproducibility.
Id. at 60, 104-05. PCAST concluded that the studies performed to that date, with one exception, were not properly designed, had severely underestimated the false positive and false negative error rates, or otherwise “differ[ed] in important ways from the problems faced in casework.” Id. at 106. Among other things, PCAST noted design flaws in existing studies, including: (1) many were not “black-box” studies,14 id. at 49; and (2) many were closed-set studies, 14 “A black box study assesses the accuracy of examiners’ conclusions without considering how the conclusions were reached. The examiner is treated as a ‘black-box’ and the researcher measures how the output of the ‘black-box’ (examiner’s conclusion) varies depending on the input (the test specimens presented for analysis).
To test examiner accuracy, the ‘ground truth’ regarding the type or source of the test specimens must be known with certainty.” Organization of Scientific Area Committees for Forensic Science, OSAC Draft Guidance on Testing the Performance of Forensic Examiners (2018), available at https://www.nist.gov/document/drafthfcguidancedocument-may8pdf (last accessed June 14, 2023), archived at https://perma.cc/3LH5-KURT. 20 in which comparisons are dependent upon each other and there is always a “correct” answer within the set, id. at 106. The sole exception to PCAST’s negative critique of study designs was a study performed by the United States Department of Energy’s Ames Laboratory (the “Ames I Study”), which PCAST called “the first appropriately designed black-box study of firearms [identification].” Id. at 11. Nonetheless, PCAST observed that that study, which we discuss below, was not published in a scientific journal, had not been subjected to peer review, and stood alone. Id.
PCAST therefore concluded that “firearms analysis currently falls short of the criteria for foundational validity” and called for additional testing. Id. at 111-14. C. Recent Studies of the AFTE Theory Numerous studies of the AFTE Theory have been performed over the course of several decades. The State contends that many of those studies are scientifically valid, reflect extremely low false positive error rates, and therefore support the reliability of the methodology.
Mr. Abruquah argues that the studies on which the State relies are flawed and were properly discounted by the NRC and PCAST, that even the best studies present artificially low error rates by treating inconclusive findings as correct, and that the most recent and authoritative study reveals “shockingly” low rates of repeatability and reproducibility. The State is correct that numerous studies have purported to validate the AFTE Theory, including by identifying relatively low false positive error rates. One of the State’s expert witnesses, Dr. James E. Hamby, is the lead author on one such study, in which 697 21 examiners inspected “over 240 test sets consisting of bullets fired through 10 consecutively rifled RUGER P-85 pistol barrels.” James E. Hamby et al., A Worldwide Study of Bullets Fired from 10 Consecutively Rifled 9MM Ruger Pistol Barrels—Analysis of Examiner Error Rate, 64:2 J. Forensic Scis. 551, 551 (Mar. 2019) (the “Hamby Study”). In that closed-set study, of 10,455 unknown bullets examined, 10,447 “were correctly identified by participants to the provided ‘known’ bullets,” examiners could not reach a definitive conclusion on eight bullets, and none were misidentified.15 Id. at 556.
The error rate, excluding inconclusive results, was thus 0.0%. See id. Examples of other studies on which the State relies, all of which identify relatively low error rates based on the study method employed, include: (1) Jamie A. Smith, Beretta barrel fired bullet validation study, 66 J. Forensic Scis. 547 (2021) (comparison testing of 30 consecutively manufactured pistol barrels, producing a 0.55% error rate); and (2) Tasha P. Smith et al., A Validation Study of Bullet and Cartridge Case Comparisons Using Samples Representative of Actual Casework, 61 J. Forensic Scis. 939 (2016) (within-set study of 31 examiners matching bullets and cartridge cases, yielding a 0.0% false-positive rate for bullet comparisons and a 0.14% false-positive error rate for cartridge cases). The NRC and PCAST both are critical of closed-set studies like the Hamby Study and others that provide examiners with multiple “unknown” bullets or cartridge cases and a corresponding number of “known” bullets or cartridge cases that the examiners are asked 15 Of the eight, the authors point out that three examiners “reported insufficient individual characteristics for two of the test bullets and two trainees could not associate five of the test bullets to their known counterpart bullets.” Hamby Study, at 556. 22 to match.
The NRC and PCAST criticize such studies as not being representative of casework because, among other reasons: (1) examiners are aware they are being tested; (2) a correct match exists within the set for every sample, which the examiners also know; and (3) the use of consecutively manufactured firearms (or barrels) in a closed-set study has the effect of eliminating any confusion concerning whether particular patterns or marks constitute subclass or individual characteristics. PCAST Report, at 32-33, 52-59, 107-09; 2009 NRC Report, at 154-55. The Ames I Study, which PCAST had identified as the only one that had been “appropriately designed” to that point, PCAST Report, at 111, was a 2014 open-set, black- box study designed to measure error rates in the comparison of “known” and “unknown” cartridge cases (the Ames I Study did not involve bullets). See David P. Baldwin et al., A Study of False-Positive and False-Negative Error Rate in Cartridge Case Comparisons, Defense Biometrics & Forensics Office, U.S. Dep’t of Energy (Apr. 2014).
In the Ames I Study, 15 sets of four cartridge cases fired from 25 new, same-model handguns using the same type of ammunition were sent to 218 examiners. Ames I Study, at 3. Each set included one unknown sample and three known samples fired from the same known gun, which might or might not have been the source of the unknown sample. Id. at 4.
Even though there was a known correct answer of either an identification or an elimination for every set, examiners were permitted to make “inconclusive” responses, which were “not counted as an error or as a non-answer[.]” Id. at 6. Of the 1,090 comparisons where the “known” and “unknown” cartridge cases were fired from the same source firearm, the examiners incorrectly excluded only four cartridge cases, yielding a false-negative rate of 23 0.367%. Id. at 15. Of the 2,180 comparisons where the “known” and “unknown” cartridge cases were fired from different firearms, the examiners incorrectly matched 22 cartridge cases, yielding a false-positive rate of 1.01%.16 Id. at 16.
However, of the non-matching comparison sets, 735, or 33.7%, were classified as inconclusive, id., a significantly higher percentage than in any closed-set study. The Ames Laboratory later conducted a second open-set, black-box study that was completed in 2020, in between the Frye-Reed and Daubert-Rochkind hearings in this case. See Stanley J. Bajic et al., Report: Validation Study of the Accuracy, Repeatability, and Reproducibility of Firearm Comparisons, U.S. Dep’t of Energy 1-2 (2020) (the “Ames II Study”). The Ames II Study, which was undertaken in direct response to PCAST’s call for further studies to demonstrate the foundational validity of firearms identification, id. at 12, enrolled 173 examiners for a three-phase study to test for all three elements PCAST had identified as necessary to support foundational validity: accuracy (in Phase I), repeatability (in Phase II), and reproducibility (in Phase III).
In each of three phases, each participating examiner received 15 comparison sets of known and unknown cartridge cases and 15 comparison sets of known and unknown bullets. Id. at 23. The firearms used for the bullet comparisons were either Beretta or Ruger handguns and the firearms used for the cartridge case comparisons were either Beretta or Jimenez handguns. Id.
Only the researchers knew the “ground truth” for each packet; that is, which “unknown” cartridges and bullets matched or did not match the included “known” cartridges and bullets. Id. As with the 16 The authors stressed that a significant majority of the false positive responses— 20 out of 22—came from just five of the 165 examiners. Ames I Study, at 16. 24 Ames I Study, although there was a “ground truth” correct answer for each sample set, examiners were permitted to pick from among the full array of the AFTE Range of Conclusions—identification, elimination, or one of the three levels of “inconclusive.” Id. at 12-13.
The first phase of testing was designed to assess accuracy of identification, “defined as the ability of an examiner to correctly identify a known match or eliminate a known nonmatch.” Id. at 33. In the second phase, each examiner was given the same test set examined in phase one, without being told it was the same, to test repeatability, “defined as the ability of an examiner, when confronted with the exact same comparison once again, to reach the same conclusion as when first examined.” Id. In the third phase, each examiner was given a test set that had previously been examined by one of the other examiners, to test reproducibility, “defined as the ability of a second examiner to evaluate a comparison set previously viewed by a different examiner and reach the same conclusion.” Id. In the first phase, the results, shown in percentages, were: 25 Id. at 35.
Treating inconclusive results as appropriate answers, the authors identified a false negative rate for bullets and cartridge cases of 2.92% and 1.76%, respectively, and a false positive rate for each of 0.7% and 0.92%, respectively. Id. Examiners selected one of the three categories of inconclusive for 20.5% of matching bullet sets and 65.3% of non- matching bullet sets. Id.
As reflected in the following table, the results overall varied based on the type of handgun that produced the bullet/cartridge, with examiners’ results reflecting much greater certainty and correctness in classifying bullets and cartridge cases fired from the Beretta handguns than from the Ruger (for bullets) and Jimenez (for cartridge cases) handguns:17 17 “Of the 27 Beretta handguns used in the study, 23 were from a single recent manufacturing run, and four were guns produced in separate earlier manufacturing runs.” Ames II Study, at 56. The Ames II Study does not identify similar information for the Ruger or Jimenez handguns. 26 Id. at 53. Comparing the results from the second phase of testing against the results from the first phase, intended to test repeatability, the outcomes, shown in percentages, were: Id. at 39. Thus, an examiner classifying the same matching bullet or cartridge case set a second time classified it in the same AFTE category 79% and 75.6% of the time, respectively, and an examiner classifying the same non-matching bullet or cartridge case set a second time did so 64.7% and 62.2% of the time, respectively.
Id. The authors viewed these percentages favorably, concluding that this level of “observed agreement” exceeded the level of their “expected agreement.”18 Id. at 39-41. They did so, however, based on an expected level of agreement reflecting the overall pattern of results from the first phase of 18 The study authors also produced alternate calculations in which they merged either (1) all inconclusive results together or (2) positive identifications with “Inconclusive A” results and eliminations with “Inconclusive B” results. Ames II Study, at 40.
As expected, those results produced greater agreement, although still ranging only from 71.3% agreement to 85.5% agreement. Id. at 42. 27 testing. Id. at 39-40. In other words, the metric against which the authors gauged repeatability was, in essence, random chance.
Comparing the results from the third phase of testing against the results of the first phase, intended to test reproducibility, the outcomes, shown in percentages, were: Id. at 47. Thus, an examiner classifying a matching bullet or cartridge case set previously classified by a different examiner classified it in the same AFTE category 67.8% and 63.6% of the time, respectively, and an examiner classifying a nonmatching bullet or cartridge case set previously classified by a different examiner classified it in the same AFTE category 30.9% and 40.3% of the time, respectively. Id. The authors again viewed these percentages largely favorably.
Id. at 47-49. Again, however, that conclusion was based on a level of expected agreement that was essentially random based on the overall results from the first phase of testing. Id. at 48-49. The State claims support from the Ames I and Ames II Studies based on what it calls their relatively low overall false positive rates.
The State contends that those results confirm the low false positive rates produced in every other study of firearms identification, 28 which are worthy of consideration even if they were not as robust in design as the Ames studies. By contrast, Mr. Abruquah claims that the high rates of inconclusive responses in both studies and the low rates of repeatability and reproducibility in the Ames II Study further support the concerns raised by NRC and PCAST about the lack of demonstrated foundational validity of firearms identification. D. Witness Testimony 1. The Frye-Reed Hearing Five witnesses testified at the two hearings conducted by the circuit court.
In the Frye-Reed hearing, Mr. Abruquah called William Tobin, a 27-year veteran of the Federal Bureau of Investigation with 24 years’ experience at the FBI Laboratory and an expert in forensic metallurgy. Mr. Tobin’s testimony was broadly critical of firearms identification generally and the AFTE Theory specifically. Citing support from multiple sources, he opined that: (1) firearms identification is “not a science,” does not follow the scientific method, and is circular; (2) the AFTE Theory is wholly subjective and lacks any guidance for examiners to determine the number of similarities needed to achieve an identification; (3) in the absence of standards, examiners ignore or “rationalize away” dissimilarities in samples; (4) examiners are incapable of distinguishing between subclass characteristics and individual characteristics—a phenomenon referred to as “subclass carryover”—thus undermining a fundamental premise of the AFTE Theory; (5) the studies on which the State had relied are flawed, do not reflect actual casework, and underestimate error rates; and (6) the AFTE Theory had not been subject to any “valid hypothesis testing” because the studies cited as support for it “lack any indicia of scientific reliability.” Mr. Tobin opined 29 that, in the absence of a pool of samples from all other possible firearms that might have fired the bullets at issue, the most a firearms examiner could accurately testify to in reliance on the AFTE Theory is whether it was possible that the recovered bullets were fired from Mr. Abruquah’s revolver. The State presented three witnesses.
It first presented Dr. James Hamby, an AFTE firearms examiner with a Ph.D. in forensic science who had been Chief of the Firearms Division for the United States Army Lab, authored dozens of articles and studies in the firearms examination field, trained firearms examiners domestically and internationally, and who, over the course of nearly 50 years in the field, managed his own forensic laboratory and two others. Dr. Hamby testified generally about the AFTE Theory, which he asserted had been accepted by the relevant scientific community and by courts, and proven by numerous studies, for more than a century. Dr. Hamby agreed with PCAST that to have foundational validity, a methodology dependent on subjective analysis must be subjected to empirical testing by multiple groups, be repeatable and reproducible, and provide valid estimates of the method’s accuracy. He opined that studies of firearms identification proved that the AFTE Theory meets all those criteria and has consistently low error rates.
Dr. Hamby acknowledged that false positives can result when similarities in subclass characteristics are mistaken for individual characteristics, but testified that trained examiners would not make that mistake. Dr. Hamby also discussed the controls and standards governing the work of firearms identification examiners, including internal laboratory procedures, the AFTE training manual, and periodic proficiency training required of every examiner. He testified that one 30 way forensic labs guard against the possibility of false positive results is by having a second examiner review all matches to ensure the correctness of the first examiner’s decision. In his decades of experience, Dr. Hamby was not personally aware of a second examiner ever having reached a different conclusion than the first in actual casework, which he seemed to view as a positive reflection on the reliability of the methodology.
The State’s second witness was Torin Suber, a forensic scientist manager with the Maryland State Police. Like Dr. Hamby, Mr. Suber testified about the low false-positive error rates identified in the Ames I and other studies. Mr. Suber agreed that some examiners could potentially mistake subclass characteristics for individual characteristics, but testified that such errors would be limited to novice examiners who “don’t actually have that eye or knack for identification yet.” The final witness presented at the Frye-Reed hearing was the State’s testifying expert, Mr. McVeigh, whom the court accepted as an expert in firearms and toolmark examinations generally, as well as “the specifics of the examination conducted in this matter.” Mr. McVeigh testified that 100% of his work is in firearms examinations and that firearms identification is generally accepted as reliable in the relevant scientific community. Mr. McVeigh acknowledged the subjective standards and procedures used in the AFTE methodology but claimed that it is “a forensic discipline with a fairly strict methodology and a lot of rules and accreditation standards to follow.” He also relied heavily on what he described as low error rates revealed by the Ames I Study and a separate 31 study out of Miami-Dade County.19 Although acknowledging the concern that examiners might mistake subclass characteristics for individual characteristics, Mr. McVeigh testified that possibility is “the number one thing[] that firearm examiners guard against.” He said that the “current thinking in the field” is that a trained examiner can overcome that concern.
With respect to the examination he conducted in Mr. Abruquah’s case, Mr. McVeigh testified that he received for analysis two firearms, a Glock pistol and a Taurus revolver, along with “six fired bullet items,” one of which was unsuitable for comparison. Based on class characteristics, he first eliminated the Glock pistol. He then fired two rounds from the Taurus revolver and compared markings on those bullets against the crime scene bullets using the comparative microscope. In doing so, he focused on the “land impressions,” rather than the “groove impressions[, which] are the most likely place where the subclass [characteristics] would occur[.]” Mr. McVeigh opined, without qualification, that, based on his analysis, “at some point each one of those five projectiles had been fired from the Taurus revolver.” He testified that his conclusion had been confirmed by another examiner in his lab. 19 Mr. McVeigh referred to the Miami-Dade Study as an open-set study.
Although neither party introduced a report of the Miami-Dade Study, PCAST described it as a “partly open” study. PCAST Report, at 109. According to PCAST, examiners were provided 15 questioned samples, 13 of which matched samples that were provided and two of which did not. Id.
Of the 330 non-matching samples that were provided, the examiners eliminated 188 of them, reached an inconclusive determination for 138 more, and made four false classifications. Id. The inconclusive rate for the non-matching samples was thus 41.8% with a false positive rate of 2.1%. Id.
PCAST observed that even in that “partly open” study, the inconclusive rate was “200-fold higher” and the false positive rate was “100-fold higher” than in closed set studies. Id. 32 On cross-examination, Mr. McVeigh admitted that he did not know how Taurus manufactured its .38 Special revolver, how many such revolvers had been consecutively manufactured and shipped to the Prince George’s County area, or how many in the area might show similar subclass characteristics. He also admitted that the proficiency testing he had undergone during his career is not blind testing and is “straight forward.” Indeed, to his knowledge, no one in his lab had ever failed a proficiency test. Mr. McVeigh asserted that bias is not a concern in firearms examinations because the examiners are not provided any details from the police investigation before conducting an examination. 2.
The Daubert-Rochkind Hearing At the Daubert-Rochkind hearing, each party presented only one witness to supplement the record that had been created at the Frye-Reed hearing. The State began with Dr. Hamby. In addition to reviewing many of the same points from his original testimony, Dr. Hamby testified that the AFTE Theory had been tested since 1907 and peer reviewed hundreds of times. He highlighted the low error rates produced in studies, including those in which examiners matched bullets fired from consecutively manufactured barrels.
He was also asked about the more recent Ames II Study, but seemed to have limited familiarity with it. Mr. Abruquah presented testimony and an extensive affidavit from David Faigman, Dean of the University of California Hastings College of Law, whom the court accepted as an expert in statistical and methodological bases for scientific evidence, including research design, scientific research, and methodology. Dean Faigman discussed several concerns with the validity of the AFTE Theory, which were principally premised on the subjective 33 nature of the methodology, including: (1) the difference in error rates between closed- and open-set tests; (2) potential biases in testing that might skew the results in studies, including (a) the “Hawthorne effect,” which theorizes that participants in a test who know they are being observed will try harder; and (b) a bias toward selecting “inconclusive” responses in testing when examiners know it will not be counted against them, but that an incorrect “ground truth” response will; (3) an absence of pre-testing and control groups; (4) the “prior probability problem,” in which examiners expect a certain result and so are more likely to find it; and (5) the lack of repeatability and reproducibility effects. Dean Faigman agreed with PCAST that the Ames I Study “generally . . . was the right approach to studying the subject.” He observed, however, that if inconclusives were counted as errors, the error rate from that study would “balloon[]” to over 30%.
In discussing the Ames II Study, he similarly opined that inconclusive responses should be counted as errors. By not doing so, he contended, the researchers had artificially reduced their error rates and allowed test participants to boost their scores. By his calculation, when accounting for inconclusive answers, the overall error rate of the Ames II Study was 53% for bullet comparisons and 44% for cartridge case comparisons—essentially the same as “flipping a coin.” Regarding the other two phases of the Ames II Study, Dean Faigman found the rates of repeatability and reproducibility “shockingly low.” E. The Evolving Caselaw Until the 2008 NRC Report, most courts seem to have accepted expert testimony on firearms identification without incident. See David H. Kaye, Firearm-Mark Evidence: Looking Back and Looking Ahead, 68 Case Western Reserve L. Rev. 723, 723-26 (2018); 34 see also, e.g., United States v. Davis, 103 F.3d 660, 672 (8th Cir. 1996); United States v. Natson, 469 F. Supp. 2d 1253, 1261 (M.D. Ga. 2007) (permitting an expert to testify “to a 100% degree of certainty”); United States v. Foster, 300 F. Supp. 2d 375 , 376 n.1, 377 (D. Md. 2004) (stating that “numerous cases have confirmed the reliability” of firearms and toolmark identification); United States v. Santiago, 199 F. Supp. 2d 101, 111 (S.D.N.Y. 2002); State v. Mack, 653 N.E.2d 329, 337 (Ohio 1995); Commonwealth v. Moore, 340 A.2d 447, 451 (Pa. 1975).
However, “[a]fter the NRC Report issued, some jurisdictions began to limit the scope of a ballistics expert’s testimony.” Gardner v. United States, 140 A.3d 1172, 1183 (D.C. 2016); see also Commonwealth v. Pytou Heang, 942 N.E.2d 927, 938 (Mass. 2011) (“Concerns about both the lack of a firm scientific basis for evaluating the reliability of forensic ballistics evidence and the subjective nature of forensic ballistics comparisons have prompted many courts to reexamine the admissibility of such evidence.”). Initially, those limitations consisted primarily of precluding experts from testifying that their opinions were offered with something approaching absolute certainty. In United States v. Willock, for example, Judge William D. Quarles, Jr. of the United States District Court for the District of Maryland, in adopting a report and recommendation by then-Chief Magistrate Judge, later Judge, Paul W. Grimm of that court, permitted an examiner to testify as to a “match” between a crime scene cartridge case and a particular firearm, but “without any characterization as to degree of certainty.” 696 F. Supp. 2d at 572, 574 ; see also United States v. Ashburn, 88 F. Supp. 3d 239, 250 (E.D.N.Y. 2015) (limiting an expert’s conclusions to those within a “reasonable degree of certainty in the ballistics field” 35 or a “reasonable degree of ballistics certainty”); Monteiro, 407 F. Supp. 2d at 372 (stating that the proper standard is a “reasonable degree of ballistic certainty”); United States v. Taylor, 663 F. Supp. 2d 1170, 1180 (D.N.M. 2009) (“[The expert] will be permitted to give . . . his expert opinion that there is a match . . . . [He] will not be permitted to testify that his methodology allows him to reach this conclusion as a matter of scientific certainty.”); United States v. Glynn, 578 F. Supp. 2d 567, 574-75 (S.D.N.Y. 2008) (allowing expert testimony that it was “more likely than not” that certain bullets or casings came from the same gun, “but nothing more”). Following issuance of the PCAST Report, some courts have imposed yet more stringent limitations on testimony.
One example of that evolution—notable because it involved the same judicial officer as Willock, Judge Grimm, as well as the same examiner as here, Mr. McVeigh—is in United States v. Medley, No. PWG-17-242 (D. Md. Apr. 24, 2018), ECF No. 111. In Medley, Judge Grimm thoroughly reviewed the state of knowledge at that time concerning firearms identification, including developments since his report and recommendation in Willock. Judge Grimm restricted Mr. McVeigh to testifying only “that the marks that were produced by the . . . cartridges are consistent with the marks that were found on the” recovered firearm, and precluded him from offering any opinion that the cartridges “were fired by the same gun” or expressing “any confidence level” in his opinion. Id. at 119.
Some other courts, although still a minority overall, have recently imposed similar or even more restrictive limitations. See United States v. Shipp, 422 F. Supp. 3d 762 , 783 (E.D.N.Y. 2019) (limiting expert’s testimony to opining that “the recovered firearm cannot 36 be excluded as the source of the recovered bullet fragment and shell casing”); Williams v. United States, 210 A.3d 734, 744 (D.C. 2019) (“[I]t is plainly error to allow a firearms and toolmark examiner to unqualifiedly opine, based on pattern matching, that a specific bullet was fired by a specific gun.”); United States v. Adams, 444 F. Supp. 3d 1248 , 1256, 1261, 1267 (D. Or. 2020) (precluding expert from offering testimony of a match but permitting testimony about “limited observational evidence”).20 III. ANALYSIS In granting in part Mr. Abruquah’s motion in limine to exclude firearms identification evidence, the circuit court ruled that Mr. McVeigh could not testify “to any level of practical certainty/impossibility, ballistic certainty, or scientific certainty that a suspect weapon matches certain bullet or casing striations.” However, the court ruled that Mr. McVeigh could opine the bullets and fragment “recovered from the murder scene fall into any of the AFTE Range of Conclusions[,]” i.e., identification, any of the three levels of inconclusive, or elimination. Accordingly, at trial, after explaining how he analyzed the samples and compared their features, Mr. McVeigh testified, over objection and separately with respect to each of the four bullets and the bullet fragment, that each “at some point” “had been fired” from or through “the Taurus revolver.” He testified neither that his 20 In United States v. Davis, citing Judge Grimm’s reasoning in Medley with approval, a federal district court judge in West Virginia also precluded Mr. McVeigh and other examiners from testifying that marks on a cartridge case indicated a “match” with a particular firearm, while permitting the examiners to testify that marks on the cartridges were “similar and consistent with each other.” 2019 WL 4306971 , at 7, Case No. 4:18- cr-00011 (W.D. Va. 2019). 37 opinion was offered to any particular level of certainty nor that it was subject to any qualifications or caveats.
In his appeal, Mr. Abruquah does not challenge all of Mr. McVeigh’s testimony or that firearms identification is sufficiently reliable to be admitted for some purposes. Instead, he contends that the methodology is insufficiently reliable to support testimony “identify[ing] a specific firearm as the source of a questioned bullet,” and argues that an examiner should be limited to opining, “at most, that a firearm cannot be excluded as the source of the questioned projectile[.]” In response, the State argues that firearms identification evidence has been accepted by courts applying the Daubert standard as reliable, has repeatedly been proven reliable in studies demonstrating very low false- positive rates, and that, “[a]t best, [Mr. Abruquah] has demonstrated that there are ongoing debates regarding how to assess the AFTE methodology[,]” not whether it is admissible. In light of the scope of Mr. Abruquah’s challenge, our task is to assess, based on the information presented to the circuit court, whether the AFTE Theory can reliably support an unqualified opinion that a particular firearm is the source of one or more particular bullets. Our analysis of the Daubert-Rochkind factors is thus tailored specifically to that issue, not to the reliability of the methodology more generally.
Before turning to the specific Daubert-Rochkind factors, we offer two preliminary observations. First, our analysis is not dependent on whether firearms identification is a “science.” “Daubert’s general holding,” adopted by this Court in Rochkind, “applies not only to testimony based on ‘scientific’ knowledge, but also to testimony based on ‘technical’ and ‘other specialized’ knowledge.” Rochkind, 471 Md. at 36 (quoting Kumho 38 Tire Co., 526 U.S. at 141 ). Second, it is also not dispositive that firearms identification is a subjective endeavor. See, e.g., United States v. Romero-Lobato, 379 F. Supp. 3d 1111, 1120 (D. Nev. 2019) (“The mere fact that an expert’s opinion is derived from subjective methodology does not render it unreliable.”); Ashburn, 88 F. Supp. 3d at 246-47 (stating that “the subjectivity of a methodology is not fatal under [Federal] Rule 702 and Daubert”).
The absence of objective criteria is a factor that we consider in our analysis of reliability, but it is not dispositive. We now turn to consider each of the ten Daubert-Rochkind factors. Of course, those factors “are neither exhaustive nor mandatory,” Matthews, 479 Md. at 314 , but they provide a helpful framework for our analysis in this case. A. Testability Although significant dispute surrounds many of the studies conducted on firearms identification to date, and especially their applicability to actual casework, it is undisputed that firearms identification can be tested.
Indeed, the bottom-line recommendation of the most significant critics of firearms identification to date, the authors of the 2009 NRC and PCAST Reports, was to call for more and better testing, not to question whether such testing is possible. B. Peer Review and Publication The second Daubert-Rochkind factor considers whether a methodology has been submitted “to the scrutiny of the scientific community,” under the belief that doing so “increases the likelihood that substantive flaws in methodology will be detected.” Daubert, 509 U.S. at 593 . The circuit court concluded that the State satisfied its burden to show that 39 the firearms and toolmark identification methodology has been peer reviewed and published. We think the evidence is more mixed.
The two most robust studies of firearms identification—Ames I and II—have not been peer reviewed or published in a journal. The record does not disclose why. Some of the articles on which the State and its witnesses rely have been published in the AFTE Journal, a publication of the primary trade group dedicated to advancing firearms identification. The required steps in the AFTE Journal’s peer review process involve a review by “a member of [AFTE’s] Editorial Review Panel” for “grammatical and technical correctness” and review by an AFTE “Assistant Editor[]” for “grammar and technical content.” See AFTE, Peer Review Process, available at https://afte.org/afte-journal/afte- journal-peer-review-process (last accessed June 14, 2023), archived at https://perma.cc/822Y-C7G8.
That process appears designed primarily to review articles and studies to determine their adherence to the AFTE Theory, not to test the methodology. Although a handful of other firearms identification studies have been published in other forensic journals, the record is devoid of any information about the extent or quality of peer review as concerns the validity of the methodology. Nonetheless, NRC’s and PCAST’s critiques of some of those same studies, and of the AFTE Theory more generally, have served many of the same purposes that might have been served by a robust peer review process. See Shipp, 422 F. Supp. 3d at 777 (concluding that the AFTE Theory had been adequately subjected to peer review and publication due in large part to “the scrutiny of PCAST and the flaws it perceived in the AFTE Theory”). 40 C. Known or Potential Rate of Error The circuit court found that the parties did not dispute “that a known or potential rate of error has been attributed to firearms identification evidence,” and treated that as favoring admission of Mr. McVeigh’s testimony.
(Emphasis removed). Neither party disputes that there is a potential rate of error for firearms identification or that a number of studies have purported to identify such an error rate. However, they do dispute whether the studies to date have identified a reliable error rate. On that issue, we glean several relevant points from the record.
First, the reported rates of “ground truth” errors—i.e., “identification” of a non- matching sample or “elimination” of a matching sample—from studies in the record are relatively low.21 Error rates in most closed-set studies hover close to zero and the overall error rates calculated in the Ames I and II Studies were in the low single digits.22 It thus 21 Most of the parties’ attention in this case is naturally focused on the “false positive” rate. Although false positives create the greatest risk of leading directly to an erroneous guilty verdict, an examiner’s erroneous failure to eliminate the possibility of a match could also contribute to an erroneous guilty verdict if the correct answer— elimination—would have led to an acquittal. To that extent, it is notable that in the first round of testing in the Ames II Study, examiners correctly eliminated only 33.8% of non- matching bullets and 48.5% of non-matching cartridge cases. See Ames II Study, at 35. 22 The Ames I Study identified a false negative rate of 0.367%, with a 95% confidence interval of up to 0.94%, a false-negative-plus-inconclusive rate of 1.376%, with a 95% confidence interval of up to 2.26%, and a false positive rate of 0.939%, with a 95% confidence interval of up to 2.26%.
Ames I Study, at 17. The Ames II Study reports its results for bullets as having a false positive error probability of 0.656%, with a 95% confidence interval of up to 1.42%, and a false negative error probability of 2.87%, with a 95% confidence interval of up to 4.26%. The Ames II Study results for cartridge cases showed a false positive error probability of 0.933%, with a 95% confidence interval of up to 1.57% and a false negative error probability of 1.87%, with a 95% confidence interval of up to 2.99%. Ames II Study, at 77. 41 appears that, at least in studies conducted thus far, it is relatively rare for an examiner in a study environment to identify a match between a firearm and a non-matching bullet.
Second, the low error rates from closed-set, matching studies utilizing bullets or cartridges fired from consecutively manufactured firearms or barrels, offer strong support for the propositions that: (1) firearms produce some unique collections of individual patterns and markings on bullets and cartridges they fire; and (2) such collections of individual patterns and markings can be reliably identified when subclass characteristics are removed from the equation.23 Third, the rate of “inconclusive” responses in closed-set studies is negligible to non- existent, see, e.g., Hamby Study, at 555-56 (finding that examiners classified eight out of 10,445 responses as inconclusive); but the rate of such responses in open-set studies is significant, see, e.g., Ames I Study, at 16 (finding that examiners classified 33.7% of “true different-source comparisons” as inconclusive); Ames II Study, at 35 (finding that examiners classified more than 20% of matching bullet sets and more than 65% of non- matching bullet sets as inconclusive), suggesting that examiners choose “inconclusive” even when it is not a “correct” response. The State, its witnesses, and the studies on which they rely suggest that responses of “inconclusive” are properly treated as appropriate 23 The use of bullets and cartridges from consecutively manufactured firearms or barrels, although more difficult in the sense that the markings in total can be expected to be more similar than those fired from non-consecutively manufactured firearms or barrels, also makes it easier to eliminate any confusion concerning whether marks or patterns are subclass or individual characteristics. See Tasha P. Smith et al., A Validation Study of Bullet and Cartridge Case Comparisons Using Samples Representative of Actual Casework, 61 J. Forensic Scis. 939 (2016) (noting that toolmarks on consecutively manufactured firearms may be identified “when subclass influence is excused”). 42 responses because, as stated in the Ames I Study, if “the examiner is unable to locate sufficient corresponding individual characteristics to either include or exclude an exhibit as having been fired in a particular firearm,” then “inconclusive” is the only appropriate response. Ames I Study, at 6.
That answer would be more convincing if rates of inconclusive findings were consistent as between closed-set and open-set studies or if the Ames II Study had produced higher levels of consistency in the repeatability or reproducibility portions of the study. Instead, whether an examiner chooses “inconclusive” in a study seems to depend on something other than just the “corresponding individual characteristics” themselves. Fourth, if at least some inconclusives should be treated as incorrect responses, then the rates of error in open-set studies performed to date are unreliable. Notably, if just the “Inconclusive-A” responses—those for which the examiner thought there was almost enough agreement to identify a match—for non-matching bullets in the Ames II Study were counted as incorrect matches, the “false positive” rate would balloon from 0.7% to 10.13%.
That is particularly noteworthy because in all the studies conducted to date, the participating examiners knew that (1) they were being studied and (2) an inconclusive response would not be counted as incorrect. There is no evidence in the record that examiners in a casework environment—when processing presumably less pristine samples than those included in studies and that were provided to them by law enforcement officers in the context of an investigation—select inconclusive at the same rate they do in an open- set testing environment. 43 Fifth, it is notable that the accuracy rate in the Ames II Study varied significantly between the two different types of firearms tested. Examiners correctly classified 89.7% of matching bullet sets fired from Beretta handguns but only 56.6% of those fired from Ruger handguns. Ames II Study, at 53.
They also correctly eliminated 38.7% of non- matching bullet sets fired from Beretta handguns and only 21.7% of those fired from Ruger handguns. Id. Given that variability, it is significant that the record provides scant information about where Taurus revolvers might fall on the error rate spectrum.24 Finally, we observe that even if the studies reflecting potential error rates of up to 2.6% reflected error rates in actual casework—a proposition for which this record provides no support—that rate must be assessed in the context of the evidence at issue. Not all expert witness testimony is created the same.
Unlike testimony that results in a determination that the perpetrator of a crime was of a certain height range, see Matthews, 479 Md. at 285 , a conclusion that a bullet found in a victim’s body was fired from the defendant’s gun is likely to lead much more directly to a conviction. That effect is compounded by the fact that a defendant is almost certain to lack access to the best evidence that could potentially contradict (or, of course, confirm) such testimony, which would be bullets fired from other firearms from the same production run. 24 During the Frye-Reed hearing, Dr. Hamby testified, using Glock as an example, that high-quality firearms would produce bullets and cartridge cases with very consistent patterns and markings, even across 10,000 cartridges, because the process of firing has little effect on the firearm. He also testified that, by contrast, an examiner might not be able to tell the difference between cartridge cases from rounds fired even consecutively from a low-quality firearm, because each bullet “just eats up the barrel.” Asked where a Taurus .38 revolver falls on the spectrum between a “cheap gun versus the most expensive,” Dr. Hamby offered that “it’s mid-level.” 44 The relatively low rate of “false positive” responses in studies conducted to date is by far the most persuasive piece of evidence in favor of admissibility of firearms identification evidence. On balance, however, the record does not demonstrate that that rate is reliable, especially when it comes to actual casework.
D. Existence and Maintenance of Standards and Controls The circuit court found the evidence with respect to the existence and maintenance of standards and controls to be “muddled” and so to weigh against admission. We mostly agree. On the one hand, to the extent that this factor encompasses operating procedures designed to ensure a consistency in process, see, e.g., Adams, 444 F. Supp. 3d at 1266 (discussing annual proficiency testing, second reviewer verification, technical review, and training as relevant to the analysis of standards and quality control), the State presented evidence of such standards and controls. That evidence includes the AFTE training manual, laboratory standard operating procedures, and laboratory accreditation standards.
Together, those sources provide standards and controls applicable to: (1) the training and certification of firearms examiners; (2) proficiency testing of firearms examiners; and (3) the mechanics of how examiners treat evidence and conduct examinations. Accord Willock, 696 F. Supp. 2d at 571-72 (finding the existence of “standards governing the methodology of firearms-related toolmark examination”). Notably, however, the record also contains evidence that severely undermines the value of some of those same standards and controls. For example, one control touted by advocates of firearms identification is a requirement that a second reviewer confirm every identification classification.
See Taylor, 663 F. Supp. 2d at 1176 (noting an expert’s 45 testimony that “industry standards require confirmation by at least one other examiner when the first examiner reaches an identification”). Indeed, Dr. Hamby testified that he believes error rates identified in firearms identification studies are overstated because those studies do not permit confirmatory review by a second examiner. However, Dr. Hamby also testified that the confirmatory review process is not blind, meaning that the second reviewer knows the conclusion reached by the first. Even more significantly, Dr. Hamby testified that in his decades of experience in firearms identification in multiple laboratories in multiple states, he was not aware of a single occasion in which a second reviewer had reached a different conclusion than the first.
In light of the findings in the reproducibility phase of the Ames II Study concerning how frequently examiners in the study environment come to different conclusions, Dr. Hamby’s testimony strongly suggests that study results do not, in fact, reliably represent what occurs in actual casework. As a second example, although advocates of firearms identification tout periodic proficiency testing by Collaborative Testing Services Inc. (“CTS”) as a method of ensuring the quality of firearms identification, the record contains no evidence supporting efficacy of that testing. To the contrary, the evidence suggests that examiners rarely, if ever, fail CTS proficiency tests. Dr. Hamby confirmed that the industry’s mandate to CTS with respect to proficiency tests “was to try to make them [as] inexpensive as possible.” To the extent that “standards and controls” encompasses standards applicable to the analysis itself, see, e.g., Shipp, 422 F. Supp. 3d at 779-81 (discussing the “circular and subjective” nature of the sufficient agreement standard and the inability of examiners “to protect against false positives” as an absence of “standards controlling the technique’s 46 operation” (quoting Daubert, 509 U.S. at 594 )), firearms identification faces an even greater challenge.
As noted, “sufficient agreement,” the threshold for reaching an “identification” classification, lacks any guiding standard other than the examiner’s own subjective judgment. The AFTE Theory states that: “sufficient agreement” is related to the significant duplication of random toolmarks as evidenced by the correspondence of a pattern or combination of patterns of surface contours. The theory then observes that: [a]greement is significant when the agreement in individual characteristics exceeds the best agreement demonstrated between toolmarks known to have been produced by different tools and is consistent with agreement demonstrated by toolmarks known to have been produced by the same tool. AFTE Theory (emphasis removed).
The theory offers no guidance as to the quality or quantity of shared individual characteristics—even assuming it is possible to reliably differentiate these from subclass characteristics—that should cause an examiner to determine that two bullets were fired from the same firearm or the quality or quantity of different individual characteristics that should cause an examiner to reach the opposite conclusion.25 See William A. Tobin & Peter J. Blau, Hypothesis Testing of the Critical Underlying Premise of Discernible Uniqueness in Firearms-Toolmarks Forensic Practice, 53 Jurimetrics J. 121 , 125 (2013); 2009 NRC Report, at 153-54; see also Itiel E. Dror, Commentary, The Error in “Error Rate”: Why Error Rates Are So Needed, Yet So Elusive, 25 On cross-examination, Mr. McVeigh answered that he could not identify the “least number of matching individual characteristics” that he had “ever used to make an identification[,]” declining to say even whether it may have been as low as two shared characteristics. 47 65 J. Forensic Scis. 1034, 1037 (2020) (stating that “forensic laboratories vary widely in what decisions are verified”). As explained in the findings of the authors of the Ames II Study, in defending the decision not to treat inconclusive results as errors: When confronted with a myriad of markings to be compared, a decision has to be made about whether the variations noted rise above a threshold level the examiner has unconsciously assigned for each examination. Ames II Study, at 75; see also id. (“[A]ll examiners must establish for themselves a threshold value for evaluation[.]”).
A “standard” for evaluation that is dependent on each individual examiner “unconsciously assign[ing]” a threshold level “for each examination” may not undermine the reliability of the methodology to support generalized testimony about the consistency of patterns and marks on ammunition fired from a particular firearm and crime scene bullets. It does not, however, support the reliability of the methodology to identify, without qualification, a particular crime scene bullet as having been fired from a particular firearm. On this issue, we find the results of phases two and three of the Ames II Study particularly enlightening. The PCAST Report identified accuracy, repeatability, and reproducibility as the key components of the foundational validity of any forensic technique.
PCAST Report, at 5. Dr. Hamby testified at the Frye-Reed hearing that he agreed with that as a general proposition. The Ames II Study, which was not available at the time of the Frye-Reed hearing, was designed specifically to test the repeatability and reproducibility of the AFTE Theory methodology. For purposes of reviewing the 48 reliability of firearms identification to support the admissibility of expert testimony of a “match,” the level of inconsistency identified through that study is troublesome.
Notably, at the Frye-Reed hearing, Mr. McVeigh rejected the notion that a firearms examiner looking at a bullet multiple times might come to different conclusions, stating that he believed that firearms identification’s “repeatability is not in question.” By the time of the Daubert hearing, however, the Ames II Study had been released, with data revealing that an examiner reviewing the same bullet set a second time classified it in the same AFTE category only 79% of time for matching sets and 65% of the time for non-matching sets. Ames II Study, at 39. In light of the black-box nature of the study, there is no explanation of this lack of consistency or of the lack of reproducibility shown in the same study.26 Nonetheless, it highlights both (1) the absence of any standards or controls to guide the analysis of examiners and (2) the importance of testing unverified (though undoubtedly genuinely held) claims about reliability. The lack of standards and controls is perhaps most acute in discerning whether a particular characteristic is a subclass or an individual characteristic.
As noted, subclass characteristics are those shared by a group of firearms made using the same tools, such as those made in the same production run at a facility. Individual characteristics are those 26 As noted above, the Ames II Study also found that an examiner reviewing a bullet set previously classified by a different examiner classified it in the same AFTE category 68% of the time for matching sets and 31% of the time for non-matching sets. Ames II Study, at 47. Even when the authors of the Ames II Study paired Identifications with Inconclusive-A responses and Eliminations with Inconclusive-C responses, second examiners still reached the same results as the first examiners looking at the same set of matching bullet sets only 77.4% of the time, and did so when looking at the same set of non-matching bullet sets only 49.0% of the time.
Ames II Study, at 49. 49 specific to a particular firearm. Both can result from aspects of the manufacturing process; individual characteristics can also result from later events, such as ordinary wear and cleaning and polishing. Currently, there are no published standards or controls to guide examiners in identifying whether any particular pattern or mark is a subclass or an individual characteristic. Mr. McVeigh testified that examiners attempt to “guard against” this “subclass carryover,” and that it is possible for a “trained examiner” to do so.27 However, neither he nor any other witness identified any industry standards or controls addressing that topic.
On balance, consideration of the existence and maintenance of standards and controls weighs against admission of testimony of a “match” between a particular firearm and a particular crime scene bullet. Accord Shipp, 422 F. Supp. 3d at 782 (“[T]he court finds that the subjective and circular nature of AFTE Theory weighs against finding that a firearms examiner can reliably identify when two bullets or shell casings were fired from the same gun.”). 27 Mr. Abruquah relies on a 2007 study published in the AFTE Journal that was designed to test the possibility that cartridge cases fired from two pistols that had been shipped to the same retailer on the same date would show similarities in subclass characteristics. See Gene C. Rivera, Subclass Characteristics in Smith & Wesson SW40VE Sigma Pistols, 39 AFTE J. 247 (2007) (the “Rivera Study”). The Rivera Study found “alarming similarities” among the marks from the two different pistols, which, the author concluded, “should raise further concern for the firearm and tool mark examiner who may rely only on one particular type of mark for identification purposes.” Id. at 250.
The Rivera Study suggested that the AFTE Theory’s “currently accepted standard for an identification” may need to be reconsidered as a result of very “significant” agreement between the two different pistols. Id. The AFTE seems to have responded by clarifying in the statement of its theory that an examiner’s decision should be based on individual characteristics, but it has not provided standards for distinguishing those from subclass characteristics. 50 E. General Acceptance Whether the AFTE Theory of firearms identification is generally accepted by the relevant community is largely dependent on what the relevant community is. Based on materials included in the record, as well as caselaw, the community of firearms identification examiners appears to be overwhelmingly accepting of the AFTE Theory.
See, e.g., Romero-Lobato, 379 F. Supp. 3d at 1122 (stating that “[t]he AFTE method certainly satisfies th[e general acceptance] element”); United States v. Otero, 849 F. Supp. 2d 425, 435 (D.N.J. 2012) (stating that the AFTE Theory is “widely accepted in the forensic community and, specifically, in the community of firearm and toolmark examiners”); Willock, 696 F. Supp. 2d at 571 (“[D]espite its inherent subjectivity, the AFTE theory . . . has been generally accepted within the field of toolmark examiners[.]”); Monteiro, 407 F. Supp. 2d at 372 (“[T]he community of toolmark examiners seems virtually united in their acceptance of the current technique.”). On the other hand, groups of eminent scientists and other academics have been critical of the absence of studies demonstrating the validity of firearms identification generally and the AFTE Theory specifically. See, e.g., 2009 NRC Report, at 155; PCAST Report, at 111. Indeed, the record does not divulge evidence of general acceptance of the methodology by any group outside of firearms identification examiners and law enforcement.
We conclude that the relevant community for the purpose of determining general acceptance consists of both firearms examiners and the broader scientific community that has weighed in on the reliability of the methodology. The widespread acceptance of the 51 methodology among those who have vast experience with it, study it, and devote their careers to it is of great significance. However, we would be remiss were we to rely exclusively on a community that, by definition, is dependent for its livelihood on the continued viability of a methodology to sustain it, while ignoring the relevant and persuasive input of a different, well-qualified, and disinterested segment of professionals.28 We consider this factor to be neutral. F. Whether Opinions Emerged Independently or Were Developed for Litigation The circuit court found that Mr. McVeigh’s testimony grew naturally out of research independent of the litigation because “the ultimate purpose of th[e firearms and toolmark] evidence is investigation [into the victim’s death], not litigation.” We disagree. “Historically, forensic science has been used primarily in two phases of the criminal-justice process: (1) investigation, which seeks to identify the likely perpetrator of a crime, and (2) prosecution, which seeks to prove the guilt of a defendant beyond a reasonable doubt.”29 See PCAST Report, at 4.
The use of firearms identification in a criminal prosecution is not independent of its investigative use. Nonetheless, the purpose of this factor is to determine whether there is reason for skepticism that the opinion reached might 28 In his dissent, Justice Gould takes Mr. Abruquah to task for not retaining his own firearms examiner to provide a different analysis of the bullets at issue. Dissenting Op. of Gould, J. at 42. In doing so, Justice Gould assumes that there are firearms examiners whose services were readily available to Mr. Abruquah, i.e., who are willing and able to take on work for criminal defendants in such cases.
The record contains no support for that proposition. 29 Here, for example, it appears that Mr. Abruquah was already identified as the likely perpetrator of the murder before Mr. McVeigh began his analysis of the Taurus revolver and the crime scene bullets. 52 be tailored to the preferred result for the litigation, rather than the expert’s considered, independent conclusion. Here, the circuit court lauded Mr. McVeigh’s integrity and forthrightness, and we have no reason to second-guess that view.30 Crediting the court’s findings about Mr. McVeigh’s testimony, we are confident that the court would not weigh this factor against admissibility and so we will not either. G. Unjustified Extrapolation from Accepted Premise Citing Mr. Abruquah’s “voluminous data indicating that firearms identification evidence is unjustifiably extrapolated from the toolmarks” and the State’s “credible and persuasive evidence that all extrapolations are justifiably calculated and well-reasoned[,]” the circuit court found “this factor to be in equipoise” and so to weigh against admission. In Rochkind, we explained that this factor invokes the concept of an analytical gap, as “[t]rained experts commonly extrapolate from existing data[,]” but a circuit court is not required “to admit opinion evidence that is connected to existing data only by the ipse dixit of the expert.” 471 Md. at 36 (quoting Joiner, 522 U.S. at 146 ). “An ‘analytical gap’ typically occurs as a result of ‘the failure by the expert witness to bridge the gap between [the expert’s] opinion and the empirical foundation on which the opinion was derived.’” Matthews, 479 Md. at 317 (quoting Savage v. State, 455 Md. 138, 163 (2017)). 30 We observe that another seasoned trial judge, even while limiting Mr. McVeigh’s testimony more than we do here, was equally profuse in his laudatory comments about Mr. McVeigh’s integrity.
See United States v. Medley, No. PWG-17-242 (D. Md. April 24, 2019), ECF No. 111, at 14 (“Mr. McVeigh, who was, for an expert witness, . . . remarkably forthcoming in his testimony and credible.”); id. at 53-54 (“I’ve seldom seen an expert who is as sincere and straightforward and no baloney and genuine in what he did as Mr. McVeigh.”). Nothing about our opinion or our conclusion in this case should be understood as contradicting that sentiment. 53 Although we do not preclude the possibility that the gap may be closed in the future, for the reasons already discussed, this case presents just such an analytical gap. That gap should have foreclosed Mr. McVeigh’s unqualified testimony that the crime scene bullets and bullet fragment were fired from Mr. Abruquah’s Taurus revolver. Although the court precluded Mr. McVeigh from testifying to his opinions to a “certainty,” an unqualified statement that the bullets were fired from Mr. Abruquah’s revolver is still more definitive than can be supported by the record.
To be sure, the AFTE Theory is intended to allow firearms examiners to reach conclusions linking particular firearms to particular unknown bullets. Mr. McVeigh’s testimony was thus not an unjustified departure from the methodology employed by those practicing in his field. We conclude, however, for reasons discussed above, that although the studies and other information in the record support the use of the AFTE Theory to reliably identify whether patterns and lines on bullets of unknown origin are consistent with those known to have been fired from a particular firearm, they do not support the use of that methodology to reliably opine without qualification that the bullets of unknown origin were fired from the particular firearm. H. Accounting for Obvious Alternative Explanations The court found this factor “definitively weighs in favor of admission” because Mr. McVeigh and Dr. Hamby “clearly and concisely addressed how alternative interpretations of toolmarks are generally accounted for in the field of firearms identification,” and Mr. Abruquah’s “counters in this area were ineffective.” We disagree.
For reasons already addressed, without the ability to examine other bullets fired from other firearms in the same production run as the firearm under examination, the record simply 54 does not support that firearms identification can reliably eliminate all alternative sources so as to permit unqualified testimony of a match between a particular firearm and a particular crime scene bullet. I. Level of Care Mr. McVeigh’s testimony here was given as part of his regular professional work, rendering this factor technically inapplicable. Nonetheless, to the extent this factor can be re-cast as a general inquiry into the level of care he exhibited, we have no qualms about accepting the circuit court’s determination that Mr. McVeigh is a “consummate professional in his field” and demonstrated a “level of care in this case” that was not “assailed in any convincing manner.” J. Relationship Between Reliability of Methodology and Opinion to Be Offered Based on the State’s evidence concerning the reliability of firearms examinations and “a dearth of real-life examples of erroneous examinations,” the circuit court concluded that “firearm and toolmark evidence is known to reach reliable results” and, therefore, that this final factor favors admission of the evidence. We do not question that firearms identification is generally reliable, and can be helpful to a jury, in identifying whether patterns and markings on “unknown” bullets or cartridges are consistent or inconsistent with those on bullets or cartridges known to have been fired from a particular firearm.
For that reason, to the extent Mr. Abruquah suggests that testimony about the consistency of 55 such patterns and markings should be excluded, we disagree.31 It is also possible that experts who are asked the right questions or have the benefit of additional studies and data may be able to offer opinions that drill down further on the level of consistency exhibited by samples or the likelihood that two bullets or cartridges fired from different firearms might exhibit such consistency. However, based on the record here, and particularly the lack of evidence that study results are reflective of actual casework, firearms identification has not been shown to reach reliable results linking a particular unknown bullet to a particular known firearm. For those reasons, we conclude that the methodology of firearms identification presented to the circuit court did not provide a reliable basis for Mr. McVeigh’s unqualified opinion that four bullets and one bullet fragment found at the crime scene in this case were fired from Mr. Abruquah’s Taurus revolver. In effect, there was an analytical gap between the type of opinion firearms identification can reliably support and the opinion Mr. McVeigh offered.32 Accordingly, the circuit court abused its discretion in permitting Mr. McVeigh to offer that opinion. 31 As noted, Mr. Abruquah argues that the testimony of a firearms identification examiner should be limited to opining, “at most, that a firearm cannot be excluded as the source of the questioned projectile[.]” It is not entirely clear to us whether Mr. Abruquah believes that testimony about the consistency of patterns and markings on bullets would be permissible—and, indeed, necessary to establish the basis for an opinion that a firearm cannot be excluded—or whether he believes that testimony about the consistency of such patterns and markings goes too far and should be excluded.
If the latter, we disagree for the reasons identified. 32 Both dissenting opinions contend that we have been insufficiently deferential to the trial court’s determination. Although they observe, quite correctly, that we do not ask trial judges to play the role of “amateur scientists,” Dissenting Op. of Hotten, J. at 4 56 IV. HARMLESS ERROR The State argues in the alternative that any error in admitting Mr. McVeigh’s testimony was harmless. We disagree. “The harmless error doctrine is grounded in the notion that a defendant has the right to a fair trial, but not a perfect one.” State v. Jordan, 480 Md. 490, 505 (2022).
The doctrine is strictly limited only to “error[s] in the trial process itself” that may warrant reversal. Id. at 506 (quoting Weaver v. Massachusetts, 137 S. Ct. 1899, 1907 (2017)). For an appellate court to conclude that the admission of expert testimony was harmless, the State must show “beyond a reasonable doubt, that the error in no way influenced the verdict.” Dionas, 436 Md. at 108 (quoting Dorsey, 276 Md. at 659 ). Upon our review of the record, we are not convinced beyond a reasonable doubt that the expert testimony in no way contributed to the guilty verdict.
The firearm and toolmark identification evidence was the only direct evidence before the jury linking Mr. Abruquah’s gun to the crime. Absent that evidence, the guilty verdict rested upon circumstantial evidence of a dispute between the men, a witness who heard gunfire around the time of the dispute, a firearm recovered from the residence, and testimony of a jailhouse (quoting Rochkind, 471 Md. at 33-34 ); Dissenting Op. of Gould, J. at 1, 50, we also do not provide increased deference simply because the subject matter of the expert testimony is scientific. The forensic technique under review was, until relatively recently, accepted almost entirely without critical analysis. See discussion above at 16-17.
Daubert and Rochkind demand more than adherence to an orthodoxy simply because it has long been accepted or because of the number of impressive-sounding statistics generated by studies that do not establish the reliability of the specific testimony offered. They require that the party proffering such evidence, whatever type of evidence it is, establish that it meets a minimum threshold of reliability. 57 informant. To be sure, that evidence is strong. But the burden of showing that an error was harmless is high and we cannot say, beyond a reasonable doubt, that the admission of the particular expert testimony at issue did not influence or contribute to the jury’s decision to convict Mr. Abruquah.
See Clemons v. State, 392 Md. 339, 372 (2006) (stating that “[l]ay jurors tend to give considerable weight to ‘scientific’ evidence when presented by ‘experts’ with impressive credentials” (quoting Reed v. State, 283 Md. 374, 386 (1978))). CONCLUSION Based on the evidence presented at the hearings, we hold that the circuit court did not abuse its discretion in ruling that Mr. McVeigh could testify about firearms identification generally, his examination of the bullets and bullet fragments found at the crime scene, his comparison of that evidence to bullets known to have been fired from Mr. Abruquah’s Taurus revolver, and whether the patterns and markings on the crime scene bullets are consistent or inconsistent with the patterns and markings on the known bullets. However, the circuit court should not have permitted the State’s expert witness to opine without qualification that the crime scene bullets were fired from Mr. Abruquah’s firearm. Because the court’s error was not harmless beyond a reasonable doubt, we will therefore 58 reverse the circuit court’s ruling on Mr. Abruquah’s motion in limine, vacate Mr. Abruquah’s convictions, and remand for a new trial.
RULING ON MOTION IN LIMINE CONCERNING EXPERT TESTIMONY REVERSED; JUDGMENT OF THE CIRCUIT COURT FOR PRINCE GEORGE’S COUNTY VACATED; CASE REMANDED FOR A NEW TRIAL. COSTS TO BE PAID BY PRINCE GEORGE’S COUNTY. 59 Circuit Court for Prince George’s County Case No. CT121375X Argued: October 4, 2022 IN THE SUPREME COURT OF MARYLAND* No. 10 September Term, 2022 __________________________________ KOBINA EBO ABRUQUAH v. STATE OF MARYLAND __________________________________ Fader, C.J., Watts, Hotten, Booth, Biran, Gould, Eaves, JJ. __________________________________ Dissenting Opinion by Hotten, J., which Eaves, J., joins. __________________________________ Filed: June 20, 2023 *During the November 8, 2022 general election, the voters of Maryland ratified a constitutional amendment changing the name of the Court of Appeals to the Supreme Court of Maryland. The name change took effect on December 14, 2022. Respectfully, I dissent.
I would hold that the Circuit Court for Prince George’s County did not abuse its discretion in admitting the State’s expert firearm and toolmark identification testimony and evidence, following its analysis and consideration of the factors outlined in Rochkind v. Stevenson, 471 Md. 1 , 236 A.3d 630 (2020). “When the basis of an expert’s opinion is challenged pursuant to Maryland Rule 5-702, the review is abuse of discretion.” Id. at 10 , 236 A.3d at 636 (citation omitted); State v. Matthews, 479 Md. 278, 305 , 277 A.3d 991, 1007 (2022) (citation omitted). We have declared it “the rare case in which a Maryland trial court’s exercise of discretion to admit or deny expert testimony will be overturned.” Matthews, 479 Md. at 286, 306 , 277 A.3d at 996, 1008 . This should not be one of those instances. The Circuit Court Did Not Abuse Its Discretion in Admitting the State’s Firearm Toolmark Identification Testimony Under Rochkind.
In Rochkind, this Court abandoned the Frye-Reed standard in favor of the more “flexible” analysis set forth in Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579 , 113 S. Ct. 2786 (1993), concerning the admissibility of expert testimony. Rochkind, 471 Md. at 29, 34 , 236 A.3d at 646, 650 . Rochkind prescribes ten factors for trial judges to consider when applying Maryland Rule 5-702.1 See id. at 35 , 236 A.3d at 650 (emphasis added). First, the trial court must consider the original five Daubert factors: 1 Rule 5-702 pertains to the admissibility of expert testimony and provides, in full: Expert testimony may be admitted, in the form of an opinion or otherwise, if the court determines that the testimony will assist the trier of fact to understand the evidence or to determine a fact in issue.
In making that determination, the court shall determine[:] (continued . . .) (1) whether a theory or technique can be (and has been) tested; (2) whether a theory or technique has been subjected to peer review and publication; (3) whether a particular scientific technique has a known or potential rate of error; (4) the existence and maintenance of standards and controls; and
This is a preview of Abruquah v. State. About 50% of the opinion remains. Read the complete opinion in RecordCite.