Mitochondrial haplotypes for contamination detection in low coverage whole genome sequencing

By detecting mtDNA haplotypes and grouping them into haplogroups, the method addresses the challenge of low-depth sequencing contamination, enhancing genetic screening accuracy and preventing diagnostic errors.

WO2025213075A1PCT designated stage Publication Date: 2025-10-09QUEST DIAGNOSTICS INVESTMENTS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/023235
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-05
Filing Date
2025-04-04
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current methods for detecting DNA sequencing contamination are ineffective at low sequencing depths, leading to incorrect genetic screening results and disease diagnosis, as they rely on deep sequencing and high coverage that is not feasible below 4x.

Method used

Detecting mitochondrial DNA (mtDNA) haplotypes and grouping them into haplogroups to identify contamination in low coverage sequencing, utilizing the higher genome coverage of mtDNA to distinguish between samples and detect cross-contamination.

Benefits of technology

Accurately detects contamination at sequencing depths as low as 5x, preventing errors in genetic screening and improving diagnostic accuracy by identifying sample and equipment contamination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025023235_09102025_PF_FP_ABST
    Figure US2025023235_09102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure is in the field of low depth whole genome sequence and genetic sequencing. In particular, method of detecting contamination based on haplotype detection of the mitochondrial DNA.
Need to check novelty before this filing date? Find Prior Art

Description

MITOCHONDRIAL HAPLOTYPES FOR CONTAMINATION DETECTION IN LOW COVERAGE WHOLE GENOME SEQUENCINGCROSS-REFERENCE TO RELATED APPLICATION[00011 This application claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 63 / 575,347, filed April 5, 2024, the entire contents of which is incorporated herein by reference in its entiretyTECHNICAL FIELD

[0002] The present disclosure is in the field of low depth whole genome sequence and genetic screening. Described are methods of identifying contamination of sequencing based on haplotype detection of the mitochondrial DNA.BACKGROUND

[0003] The following discussion is merely provided to aid the reader in understanding the disclosure and is not admitted to describe or constitute prior art thereto.J0004] Contamination of DNA sequencing assays, samples, and sequencing equipment can affect the output sequence reads and data potentially resulting in incorrect interpretation of downstream analysis and outcomes. One particular consequence can be incorrect conclusions and results of genetic screening which in turn can end in incorrect disease diagnosis. Thus, effectively detecting contamination is an important aspect for assays that utilize DNA sequencing.

[0005] Current approaches of detecting contamination in DNA sequencing assay depend on deep sequencing and high sequencing coverage of DNA samples, for example, to look for allelic imbalances in specific genomic coordinates where deviations from an expected allele frequencies represent contributions of DNA from exogenous sources. However, when sequencing depth becomes very low (<4 times) this approach is not possible because generally all sites will have too few sequencings reads to detect a signal of contamination.

[0006] Thus, there is a substantial need for methods for detecting contamination in sequencing assays and samples, particularly in low sequencing depth assays. Mitochondrial genomes compared to nuclear genomes in DNA extracts generally have a much higher genome coverage in low sequencing depth outputs. Methods utilizing phylogenetic incompatibility of mitochondrial haplotypes in a low coverage whole genome data would aid in providing useful quality control checks for DNA sequencing assays.SUMMARY

[0007] The present disclosure provides methods for detecting contamination in low depth DNA sequencing by detecting one or more haplotypes in mitochondrial DNA (mtDNA) and grouping of the haplotypes into one or more haplogroups. Additionally, the disclosed methods can detect the presence or absence of contamination between one or more sequencing samples in serial sequencing runs.

[0008] In one aspect, the present disclosure provides, a method of detecting contamination in DNA sequencing comprising: (a) preparing a sequencing library from DNA extracted from a biological sample obtained from a subject; (b) performing low depth sequencing (e.g., whole genome sequencing, whole exome sequencing, etc.) on the sequencing library, thereby obtaining a plurality of sequencing reads; (c) detecting in the plurality of sequencing reads one or more mitochondrial DNA (mtDNA) haplotypes and grouping the one or more mtDNA haplotypes into one or more haplogroups, wherein contamination is present during sequencing when two or more haplotypes are detected and grouped into two or more haplogroups and contamination is not present during sequencing when one haplotype is detected and grouped into one haplogroup.

[0009] In another aspect, the present disclosure provides, a method of detecting one or more haplotypes in mitochondrial DNA (mtDNA) comprising: (a) obtaining genomic DNA extracted from a sample obtained from a subject; (b) performing low pass next generation sequencing on the genomic DNA, thereby obtaining mtDNA sequence reads; (c) detecting one or more haplotypes in the mtDNA sequence reads.

[0010] In some embodiments, the DNA is total genomic DNA or comprises amplified regions of interest.[00111 In some embodiments, the biological sample is selected from a tissue sample or a fluid sample. In some further embodiments, the fluid sample is selected from plasma, serum, or whole blood.

[0012] In some embodiments, the low depth whole genome sequencing or the low pass next generation sequencing generates between about lOOx to about lOOOx mitochondrial genome coverage. In some embodiments, the low depth whole genome sequencing or the low pass next generation sequencing generates between about 0.2x to about 5x nuclear genome coverage (e.g., about 0.75x to about 4x nuclear genome coverage).

[0013] In some embodiments, the two or more haplogroups comprises a major haplogroup and one or more minor haplogroups. In some embodiments, the one or more minor haplogroups comprises 2, 3, 4, 5, or 6, minor haplogroups.

[0014] In some embodiments, each individual minor haplogroup represents a source of contamination.

[0015] In some embodiments, the one haplogroup is a major haplogroup. In some embodiments, the detecting the one or more haplotypes further comprises grouping the one of more haplotypes into (i) one major haplogroup, or (ii) one major haplogroup and one or more minor haplogroups.{0016] In some embodiments, wherein the contamination is sample to sample cross contamination.

[0017] In some embodiments, wherein the plurality of sequence reads is concurrently used for genetic screening.

[0018] In some embodiments, the low pass next generation sequencing is whole genome sequencing or whole exome sequencing.

[0019] In some embodiment, the sample obtained from the subject is concurrently sequenced for genetic screening.

[0020] In further aspects, the present disclosure provides for a method for detecting the presence or absence of contamination between one or more sequencing samples comprising: (a) preparing a first sequencing library from extracted DNA from a first subject; (b) performing low pass whole genome sequencing on the first sequencing library, therebyobtaining a plurality of sequencing reads from the first sequencing library; (c) preparing a second sequencing library from extracted DNA from a second subject; (d) performing low pass whole genome sequencing on the second sequencing library using the same sequencer used to perform low pass whole genome sequencing on the first sequencing library, thereby obtaining a plurality of sequencing reads from the second sequencing library; (e) detecting in the plurality of sequencing reads from the second sequencing library one or more mitochondrial DNA (mtDNA) haplotypes and grouping the one or more mtDNA haplotypes into one or more haplogroups, wherein contamination is present during sequencing when two or more haplotypes are detected and grouped into two or more haplogroups and contamination is not present during sequencing when one haplotype is detected and grouped into one haplogroup.

[0021] In some embodiments, further comprising detecting in the plurality of sequencing reads from the first sequencing library one or more mtDNA haplotypes.

[0022] In some embodiments, the extracted DNA from the first subject, the extracted DNA from the second subject, or the extracted DNA from both the first subject and the second subject is obtained from a biological sample independently selected from a tissue sample or a fluid sample.

[0023] In some embodiments, the fluid sample is selected from plasma, serum, or whole blood.

[0024] In some embodiments, the low depth whole genome sequencing generates between about lOOx to about lOOOx mitochondrial genome coverage. In some embodiments, the low depth whole genome sequencing generates between about 0.2x to about 5x nuclear genome coverage (e.g., about 0.75x to about 4x nuclear genome coverage).

[0025] In some embodiments, further comprising repeating (c)-(e) with extracted DNA from multiple further subjects. In some embodiments, multiple further comprises at least a third subject, at least a fourth subject, or at least a fifth subject.

[0026] In some embodiments, the contamination occurs during the sequencing sample preparation.

[0027] In some embodiments, the contamination occurs during the during the low pass whole genome sequencing.

[0028] In some embodiments, the two or more haplogroups comprises one major haplogroup and one or more minor haplogroups. In some embodiments, the one haplogroup is a major haplogroup.

[0029] In some embodiments, the contamination is between about 0.5% to about 5%. In some embodiments, the contamination is between about 1% to about 2.5%. The percent of contamination may be defined as the fraction of reads from an exogenous source.(0030] Both the foregoing summary and the following description of the drawings and detailed description are exemplary and explanatory. They are intended to provide further details of the disclosure, but are not to be construed as limiting. Other objects, advantages, and novel features will be readily apparent to those skilled in the art from the following detailed description of the disclosure.

[0031] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below are provided as being part of the inventive subject matter disclosed herein and may be employed in any combination to achieve the benefits described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0032] FIG. 1 depicts the precision and recall results of the contaminated sequence reads variant call format with a control data set variant call format. The X axis represents the percentage of contamination per sample and the Y axis represents the sequencing coverage of the nuclear DNA.

[0033] FIG. 2 depicts the precision as a function (left boxplot) and recall as a function (right boxplot). The X axis represent each tested sequenced sample label as (sequence depth, contamination %).

[0034] FIG. 3 depicts precision and recall for individual samples compared with Genome in a Bottle (GIAB) data. Samples are labeled according to the contamination percentage and genome coverage, clean / 0 represents no contamination, 01 represent 1% contamination, 25represents 2.5% contamination, 05 or 5 represent 5% contamination, and 10 represent 10% contamination. Genome coverage was tested for O.Olx, 0.015x, 0.02x, 0.04x, 0.075x, O. lx, 0.15x, 0.2x, 0.4x, 0.75x, lx, 1.5x, 2x, and 4x.

[0035] FIG. 4 depicts how haplotype-based methods disclosed herein are able to estimate contamination in samples sequenced at low depth when conventional autosomal method (e.g., DRAGEN (Illumina, 2022)) could not.DETAILED DESCRIPTION

[0036] Embodiments according to the present disclosure will be described more fully hereinafter. Aspects of the disclosure may, however, be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. The terminology used in the description herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0037] Unless the context indicates otherwise, it is specifically intended that the various features described herein can be used in any combination. Moreover, the disclosure also contemplates that in some embodiments, any feature or combination of features set forth herein can be excluded or omitted. To illustrate, if the specification states that a complex comprises components A, B, and C (or A, B, and / or C), it is specifically intended that any of A, B or C, or a combination thereof, can be omitted and disclaimed singularly or in any combination.

[0038] Unless explicitly indicated otherwise, all specified embodiments, features, and terms intend to include both the recited embodiment, feature, or term and biological equivalents thereof.I. Definitions

[0039] As used in the description of the invention and the appended claims, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0040] The terms “substantially” and “about” are used herein to describe and account for small variations. When used in conjunction with an event or circumstance, the terms can refer to instances in which the event or circumstance occurs precisely as well as instances in which the event or circumstance occurs to a close approximation. When used in conjunction with a numerical value, the terms can refer to a range of variation of less than or equal to ±10% of that numerical value, such as less than or equal to ±5%, less than or equal to ±4%, less than or equal to ±3%, less than or equal to ±2%, less than or equal to ±1%, less than or equal to ±0.5%, less than or equal to ±0.1%, or less than or equal to ±0.05%. When referring to a first numerical value as “substantially” or “about” the same as a second numerical value, the terms can refer to the first numerical value being within a range of variation of less than or equal to ±10% of the second numerical value, such as less than or equal to ±5%, less than or equal to ±4%, less than or equal to ±3%, less than or equal to ±2%, less than or equal to ±1%, less than or equal to ±0.5%, less than or equal to ±0.1%, or less than or equal to ±0.05%. The terms or “acceptable,” “effective,” or “sufficient” when used to describe the selection of any components, ranges, dose forms, etc. disclosed herein intend that said component, range, dose form, etc. is suitable for the disclosed purpose.10041 ] Additionally, amounts, ratios, and other numerical values are sometimes presented herein in a range format. It is to be understood that such range format is used for convenience and brevity and should be understood flexibly to include numerical values explicitly specified as limits of a range, but also to include all individual numerical values or sub-ranges encompassed within that range as if each numerical value and sub-range is explicitly specified. For example, a ratio in the range of about 1 to about 200 should be understood to include the explicitly recited limits of about 1 and about 200, but also to include individual ratios such as about 2, about 3, and about 4, and sub-ranges such as about 10 to about 50, about 20 to about 100, and so forth.

[0042] Also as used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).

[0043] The terms “individual,” “subject,” and “patient” are used interchangeably herein, and refer to an individual organism, vertebrate, mammal, or a human. In some embodiments, the individual, subject, or patient is a human.

[0044] The terms “sample” or “biological sample” refers to a substance that is being assayed in the methods as disclosed herein. A biological sample obtained or derived from a source of interest (i.e., a subject). Non-limiting examples of a biological sample include tissue or fluid.

[0045] The terms “library” or “sequencing library” refers to a collection of nucleic acid fragments, e.g., a collection of DNA fragments derived from whole genomic DNA or amplified regions of DNA, that is processed for sequencing. In one embodiment, a portion or all of the library nucleic acid fragments comprises an adapter sequence.100461 The terms “next generation sequencing” or “NGS” refers to sequencing methods that allow for massively parallel sequencing of clonally amplified and of single nucleic acid molecules during which a plurality, e.g., millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced in unison.

[0047] The terms “whole genome sequencing,” “WGS,” “full genome sequencing,” “complete genome sequencing,” or “entire genome sequencing” refers to the process of determining the complete DNA sequence of an organism’s genome at a single point in time by sequencing all of an organism’s chromosomal DNA / nuclear DNA as well as mitochondrial DNA (mtDNA).

[0048] The terms “sequence reads” or “read” refers to sequence information (e.g., sequence data) of a nucleic acid fragment obtained through a sequencing assay, such as a next generation sequencing (NGS) assay.

[0049] The terms “mapped sequence reads” or “aligned sequence reads” refers to the process of determining a sequence reads location or local of origin in a reference genome, which is based on the similarity of the nucleotide sequence of the reads and the genome sequence.

[0050] The terms “read depth” or “raw read depth” refers to the total amount of sequence data or sequence reads produced by the sequencing assay (i.e., machine used for sequencing).

[0051] The terms “mapped read depth” or “coverage” refers to the average number of reads that align to a known reference bases of a reference genome.

[0052] The terms “low depth” or “low pass” refers to a sequencing assay that produces a low read depth. In some embodiments, low depth or low pass sequencing can produce a low read depth and low coverage of the nuclear genome and higher coverage of the mtDNA genome. Non-limiting examples of a low coverage of the human nuclear genome include about 0.25x to about 5x coverage.

[0053] The terms “variant” or “mutation” refers to a change introduced into a reference sequence, including, but not limited to, substitutions, insertions, deletions (including truncations) relative to the reference sequence. Variants can involve large sections of DNA (e.g, copy number variation), whole chromosomes (e.g, aneuploidy) or small sections of DNA (e.g., point mutations or single nucleotide polymorphisms (SNPs), single nucleotide variants (SNVs), multiple nucleotide polymorphisms, insertions, multiple nucleotide changes, deletions, inversions, or a genomic rearrangements.

[0054] The terms “haplotype,” “mitochondrial DNA haplotype,” or “mtDNA haplotype” refers to a set of variants, such as single nuclear polymorphisms (SNPs), in the mtDNA that tend to be inherited together. mtDNA haplotypes cluster together and show the phylogenetic origins of maternal lineages. In some embodiments, haplotypes can be grouped into major haplotypes and minor haplotypes based on the frequency of the polymorphic variants.

[0055] The terms “haplogroup,” “mitochondrial DNA haplogroup,” or “mtDNA” refer to a grouping of similar haplotypes that can be defined by various variants in the mtDNA (e.g., SNPs) inherited from a common ancestor. Haplogroups can be used to present major branching points on the mitochondrial phylogenetic tree. In some embodiments, thehaplogroups can be grouped into major and minor haplogroups based on the haplotype frequency.II. Contamination detection in DNA sequencing

[0056] Identifying contamination in low depth and low coverage DNA sequencing assays (e.g., < 4x) has been challenging as previous methods (e.g., based on unexpected allele frequencies) require adequate coverage and sequencing depth of the nuclear genome. The present disclosure provides methods for detecting contamination in low depth DNA sequencing by identifying one or more haplotype groups in mtDNA. In general, mitochondrial genome coverage is much higher in low depth sequencing as compared to coverage of the nuclear genome (e.g., about lOOx to about lOOOx for mitochondrial genome compared with about 0.2x to about 5x). Mitochondrial variants (e.g., polymorphic variants) may be homoplasmic or heteroplasmic and are grouped into haplotypes based on variant allele frequency, with the haplotypes being subsequently classified into haplogroups. Heteroplasmy is the occurrence of at least two different haplotype profiles. However, contamination during sequencing (e.g., sample to sample cross contamination or sample swap contamination) can create heteroplasmic polymorphic sites that can be detected and grouped into one or more haplotype groups and one or more haplogroups profiles. Thus, based on haplotype mix and subsequence haplogroup classification, the mtDNA phylogeny can be used as a source of identifying contamination in low depth and low coverage DNA.

[0057] The present disclosure provides of detecting contamination in DNA sequencing comprising: (a) preparing a sequencing library from DNA extracted from a biological sample obtained from a subject; (b) performing low depth whole genome sequencing on the sequencing library, thereby obtaining a plurality of sequencing reads; (c) detecting in the plurality of sequencing reads one or more mitochondrial DNA (mtDNA) haplotypes and grouping the one or more mtDNA haplotypes into one or more haplogroups, wherein contamination is present during sequencing when two or more haplotypes are detected and grouped into two or more haplogroups and contamination is not present during sequencing when one haplotype is detected and grouped into one haplogroup.

[0058] Additionally, the present disclosure provides methods of detecting one or more haplotypes in mitochondrial DNA (mtDNA) comprising: (a) obtaining genomic DNAextracted from a sample obtained from a subject; (b) performing low pass next generation sequencing on the genomic DNA, thereby obtaining mtDNA sequence reads; (c) detecting one or more haplotypes in the mtDNA sequence reads.

[0059] In some embodiments, the methods as disclosed herein comprise a plurality of sequence reads concurrently used for genetic screening. In some embodiments, the disclosed methods can detect contamination is sequence reads used for genetic screening and testing. In some embodiments, the methods can prevent errors in genetic screening results and medical diagnosis.

[0060] The present disclosure and Examples provided herein show that the disclosed methods can be used to successfully detect contamination in sequencing libraries at cut-off depths that were previously unachievable with conventional autosomal approaches, such as DRAGEN. By relying on the comparatively high copy number of mitochondrial DNA in a sample, the disclosed methods utilize mitochondrial haplotypes at low sequencing depths (e.g., less than 5x) and across various sample types (e.g., blood, skin, tissue, etc.) to accurately detect contamination. a. Sample preparation and sequencing

[0061] In some embodiments, DNA is extracted from a biological sample from a subject. In some embodiments, the biological sample is a tissue sample or a fluid sample. In some embodiments, the fluid sample is selected from plasma, serum, or whole blood. In some embodiments, a biological sample may be or comprise bone marrow, blood cells, ascites, tissue or fine needle biopsy samples, cell-containing body fluids, sputum, saliva, urine, cerebrospinal fluid, peritoneal fluid, pleural fluid, feces, lymph, gynecological fluids, skin swabs, skin cells, vaginal swabs, oral swabs, nasal swabs, washings or lavages such as a ductal lavage or broncheoalveolar lavages, aspirates, scrapings, secretions or excretions. In some embodiments, a biological sample is or comprises cells obtained from an individual. In some embodiments, the biological sample comprises blood, plasma, or serum. In some embodiments, the biological sample is or comprises a tissue sample, such as a tissue swab (e.g., buccal swab, vaginal swab, etc.) or tissue biopsy (e.g., a skin biopsy, prostate biopsy, tumor biopsy, colon biopsy, stomach biopsy, esophageal biopsy, liver biopsy, pancreatic biopsy, etc.). In some embodiments, the biological sample may comprise hair or skin.

[0062] In some embodiments, a sequencing library is prepared from DNA extracted from the biological sample as disclosed herein. In some embodiments, the extracted DNA is total DNA or comprises amplified regions of interest. In some embodiments, the total DNA may comprise, for example, nucleic acids or proteins extracted from a sample. In some embodiments, the amplified regions of interest may be obtained by subjecting a primary sample to techniques such as amplification, isolation and / or purification of certain components. In some embodiments, the DNA may be amplified to obtain a population of multiple copies of processed DNA.

[0063] In some embodiments, the sequencing library comprises extracted DNA (e.g., genomic DNA). In some embodiments, the DNA can be fragmented, for example, by using a hydrodynamic shear or other mechanical force, or fragmented by chemical or enzymatic digestion, such as restriction digesting. In some embodiments, the fragments may be end repaired prior to ligation of adapters. In some embodiments, the DNA is subjected to additional modifications. Non-limiting examples include incorporate sequencing adapters onto the nucleic acids. In some embodiments, the fragments may be end repaired prior to ligation of adapters. In some embodiments, sample specific barcodes may be incorporated into the nucleic acids to allow for multiplexing of multiple samples.

[0064] In some embodiments, the library is subject to sequencing. In some embodiments, the sequencing is used to detect variants. In some embodiments, the sequencing library is sequenced to produce a plurality of sequence reads. In some embodiments, the sequence reads are about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, about 500 bp, or more than 500 bp in length. In some embodiments, the sequencing is whole genome sequencing, whole exome sequencing, or targeted sequencing. In any embodiment, the sequencing can comprise next-generation sequencing (NGS). In some embodiments, the sequencing is low pass or low depth. In some embodiments, the low pass or low depth sequencing has low sequence read depth and low coverage of the nuclear genome (e.g., 0.075x, O. lx, 0.15x, 0.2x, 0.4x, 0.75x, lx, 1.5x, 2x, 2.5x, 3x, 3.5x, or 4x). Insome embodiments, the nuclear genome coverage of the low depth whole genome sequencing or next generation sequencing is between about 0.2x to about 5x or about 0.75x to about 5x nuclear genome coverage. In some embodiments, the nuclear genome coverage of the low depth whole genome sequencing or next generation sequencing is about 0.2x, about 0.25x, about 0.3x, about 0.35x, about 0.4x, about 0.45x, about 0.5x, about 0.55x, about 0. 6x, about 0.65x, about 0.7x, about 0.75x, about lx, about 1.5x, about 2x, about 2.5x, about 3x, about 3.5x, about 4x, about 4.5x, or about 5x. In some embodiments, the nuclear genome coverage of the low depth whole genome sequencing or next generation sequencing is less than about 0.2x, less than about 0.25x, less than about 0.3x, less than about 0.35x, less than about 0.4x, less than about 0.45x, less than about 0.5x, less than about 0.55x, less than about 0. 6x, less than about 0.65x, less than about 0.7x, less than about 0.75x, less than about lx, less than about 1.5x, less than about 2x, less than about 2.5x, less than about 3x, less than about 3.5x, less than about 4x, less than about 4.5x, or less than about 5x. In some embodiments, the low pass or low depth sequencing has low sequence read depth and a high coverage of the mitochondrial genome (e.g.,100x, 200x, 300x, 500x, 600x, 800x, 900x, or lOOOx). In some embodiments, the mitochondrial genome coverage of the low depth whole genome sequencing or next generation sequencing is between about lOOx to about lOOOx mitochondrial genome coverage. In some embodiments, the mitochondrial genome coverage of the low depth whole genome sequencing or next generation sequencing is about lOOx, about 150x, about 200x, about 250x, about 300x, about 350x, about 400x, about 450x, about 500x, about 550x, about 600x, about 650x, about 700x, about 750x, about 800x, about 850x, about 900x, about 950x, or about lOOOx.

[0065] In some embodiments, the plurality of sequence reads are mapped to a reference genome (e.g., hgl8, hgl9, GRCh38.pl4, or GRCh37.pl3). In some embodiments sequence reads are mapped to the mitochondrial genome. In some embodiments, algorithms can be used to align the sequence reads. Non-limiting examples include BLAST (Altschul et al., 1990), BLITZ (MPsrch) (Sturrock & Collins, 1993), FASTA (Person & Lipman, 1988), BOWTIE (Langmead et al, Genome Biology 10:R25.1-R25.10

[2009] ), or ELAND (Illumina, Inc., San Diego, Calif., USA). In one embodiment, the sequencing data is processed by bioinformatic alignment analysis for the Illumina Genome Analyzer, which uses the Efficient Large-Scale Alignment of Nucleotide Databases (ELAND) software. Additionalsoftware includes SAMtools (SAMtools, Bioinformatics, 2009, 25(16):2078-9), theBurroughs-Wheeler block sorting compression procedure, and DRAGEN (Illumina 2022). b. Detection of Haplotypes

[0066] In general, one or more haplotypes are detected in the plurality of sequences. In one embodiment, the plurality of sequences are aligned to a genome as disclosed herein at a low coverage (e.g., 0.075x, O. lx, 0.15x, 0.2x, 0.4x, 0.75x, lx, 1.5x, 2x, 2.5x, 3x, 3.5x, or 4x). In some embodiments, the one or more haplotypes are determined by detection of mitochondrial variants (e.g., polymorphic variants) which are grouped into one or more haplotypes based on variant allele frequency. In some embodiments, the mitochondrial variants are SNPs. In some embodiments, variants with an allele frequency of 50% or more are classified into a first haplotype group. In some embodiments, the first haplotype group is called a major haplotype group. In some embodiments, any variants with an allele frequency of 50% or less are grouped into at least a second haplotype group. In some embodiments, any variants with an allele frequency of 50% or less are grouped into a second, a third, a four, or a fifth haplotype group. In some embodiments, the haplotype groups in the second, third, fourth, or fifth haplotype group are called minor haplotype groups. In some embodiments, alleles with an allele frequency of 40% or more, 30% or more, or 20% or more are classified into the first haplotype group. In some embodiments, the one or more haplotypes comprises 2, 3, 4, 5, 6, or 7 or more haplotypes. In some embodiments, the one or more haplotypes comprises one major haplotype and one or more minor haplotypes.

[0067] In general, haplogroups are based on, for example, mitochondrial phylogeny and determined by quantification of variants such as SNPs. In some embodiments, the one or more haplotypes are grouped into one or more haplogroups. In some embodiments, two or more haplotypes are grouped into two or more haplogroups. In some embodiments, one haplotype is grouped into one haplogroup. In some embodiments, the one haplogroup is called a major haplogroup. In some embodiments, a second, a third, a fourth, or a fifth haplotype is grouped into a second, a third, a fourth, or a fifth haplogroup. In some embodiments, a second, a third, a fourth, or a fifth haplogroup are minor haplogroups. In some embodiment, the two or more haplogroups comprise a major haplogroup and one ormore minor haplogroups. In some embodiments, each single haplotype is classified as a single haplogroup. c. Detection of Contamination based on haplotypes and haplogroups

[0068] In general, contamination is present during sequencing when two or more haplotypes or two or more haplogroups are detected as disclosed herein. In some embodiments, two or more minor haplogroups indicates a source of contamination. In some embodiments, each individual minor haplogroup represents DNA from a different subject as compared to the subject the biological sample was obtained. In some embodiments, each identified minor haplogroup represents a source of contamination. In some embodiments, the contamination is sample to sample cross contamination. In some embodiments, sample to sample cross contamination can occur during DNA extraction from a biological sample, during library preparation, or during sequencing. Sample to sample cross contamination can occur when more than one biological sample is obtained from a subject and the biological samples are subsequently pooled (e.g., one or more whole blood samples in separate tubes pooled into one). Errors in biological sample collection can occur resulting in biological samples from more than one subject being pooled (e.g., one whole blood sample from a subject being pool with another whole blood sample from a second subject). Sample to sample cross contamination can also occur when equipment is contaminated (e.g., such as sample splashes, spilled, or contaminated test tubes and pipettes) or during serial sequencing runs. A source of contamination can also be the laboratory scientists preparing the assays, with the laboratory scientists DNA representing one or more of the detected minor haplogroups.

[0069] In some embodiments, sample to sample cross contamination can occur during sample and library preparation. In some embodiments, where the contamination occurs during the sequencing sample preparation. In some embodiments, where the contamination occurs during the low pass whole genome sequencing.

[0070] In some embodiments, the contamination is between about 0.5% to about 5%. In some embodiments, the contaminations is between about 0.5%, about 1%, about 2.5%, about 5%, about 10% or more. In some embodiments, the contamination is between about 1% toabout 2.5%. In some embodiments, the contamination is between about 1%, about 1.5%, about 2%, about 2.5% or more.(0071 ] In some embodiments, the contamination analysis is based on the haplogroups of the one or more haplotypes. In some embodiments, haplogroups for the major haplotype and minor haplotypes are assigned as a major haplogroup and the minor haplogroup. In some embodiments, the one or more minor haplotypes and / or the one or more minor haplogroups indicate contamination. In some embodiments, Haplocheck (Weissensteiner et al. 2021) can be used as a computational model for determining haplotypes and hapolgroups in mtDNA in low depth sequencing. d. Detection of Contamination in serial sequencing runs

[0072] In general, sequencing facilities perform sequencing on many different sequence libraries at one given time. Additionally, sequencing runs can be performed serially with limited cleaning or contamination checks. The methods as disclosed herein can detect the presence or absence of contamination between one or more samples within one sequencing cycle or between several serial sequencing cycles serving as a method for quality control in the laboratory.

[0073] The present disclosure provides a method for detecting the presence or absence of contamination between one or more sequencing samples as disclosed herein. In some embodiments the method comprising: (a) preparing a first sequencing library from extracted DNA from a first subject; (b) performing low pass whole genome sequencing on the first sequencing library, thereby obtaining a plurality of sequencing reads from the first sequencing library; (c) preparing a second sequencing library from extracted DNA from a second subject; (d) performing low pass whole genome sequencing on the second sequencing library using the same sequencer used to perform low pass whole genome sequencing on the first sequencing library, thereby obtaining a plurality of sequencing reads from the second sequencing library; (e) detecting in the plurality of sequencing reads from the second sequencing library one or more mitochondrial DNA (mtDNA) haplotypes and grouping the one or more mtDNA haplotypes into one or more haplogroups, wherein contamination is present during sequencing when two or more haplotypes are detected and grouped into two ormore haplogroups and contamination is not present during sequencing when one haplotype is detected and grouped into one haplogroup. In some embodiments, further comprising detecting in the plurality of sequence reads from the first sequencing library one or more mtDNA haplotypes.100741 In some embodiments, the extracted DNA from the first subject, the extracted DNA from the second subject, or the extracted DNA from both the first subject and the second subject is obtained from a biological sample independently selected from a tissue sample or fluid sample. In some embodiments, the fluid sample is selected from plasma, serum, or whole blood.

[0075] In some embodiments, DNA is extracted from at least a third, at least a fourth, or at least a fifth subject. In some embodiments, DNA is extracted from 5 or more additional subjects. In some embodiments, DNA is extracted from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 or more additional subjects.(0076] In some embodiments, at least a third, at least a fourth, or at least a fifth sequencing library is prepared and low pass whole genome sequencing is performed using the same sequencer as used for the first and second sequencing library. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 or more additional sequencing libraries are prepared and low pass whole genome sequencing is performed using the same sequencer as used for first, second, third, fourth, or fifth sequencing library. In some embodiments, a plurality of sequence reads from the third, fourth, or a fifth sequencing library are obtained. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 or more additional sequence reads are obtained from the 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 or more additional sequencing libraries. In some embodiments, detecting in the plurality of the third, fourth, or a fifth sequencing reads and the 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 or more additional sequencing reads, one or more mtDNA haplotypes and grouping the one or more mtDNA haplotypes into one or more halopgroups, wherein contamination is present during sequencing when two or more haplotypes are detected and grouped into two or more haplogroups and contamination is not present during sequencing when one haplotype is detected and grouped into one haplogroup.

[0077] In some embodiments, the source of contamination can be identified in the first, second, third, fourth, fifth, or the 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 or more additional sequencing reads. For example, if one or more minor haplogroups identified in any of the sequence reads map with a major haplogroup in any of the additional sequence reads.EXAMPLES

[0078] These examples are provided for illustrative purposes only and not to limit the scope of the claims provided herein.Example 1. Detection of Contamination in mixed sequence reads.

[0079] Low depth whole genome sequencing reads from two difference sources were downsampled to create a contamination fastq file at known contamination values of 1%, 2.5%, 5%, and 10% at differing sequence depths. Using DRAGEN (Illumina, 2022) the fastq files were aligned to the human genome (e.g., hgl8, hg 19, hg 38, etc.) and BAM files and imputed variant cell format (VCF) files were prepared at a nuclear DNA coverage of 0.75x, lx, 1.5x, 2x, and 4x. The median mitochondrial DNA coverage was observed to be significantly higher than the nuclear DNA coverage. (Table 1) Haplocheck (Weissensteiner et al. 2021), which calls mtDNA variants, groups the variants into one or more haplotypes, assigns major and minor haplogroups based on the frequency of the haplotypes and predicts contamination estimate based on the frequency of the minor haplogroups when incompatibilities are observed, was run on the BAM files, include two control files that were not known to be contaminated. The expected verse observed contamination values were compared. The control samples were called correctly to have no contamination. 1% Contamination was reliably called in >lx coverage. 2.5% contamination was reliably called at >lx coverage and >2.5% coverage was reliably called at all samples.

[0080] The imputed VCF files were generated at nuclear DNA coverage of O.Olx, 0.015x, 0.02x, 0.04x, 0.075x, O. lx, 0.15x, 0.2x, 0.4x, 0.75x, lx, 1.5x, 2x, and 4x. Precision and recall of the called variants used for contamination determinations was assessed and compared to VCF files prepared from Genome in a Bottle (GIAB) data (Zook et al., Data Descriptor, 160025 (2016)) at the same sequence coverage to compare precision and recall of the calledvariants used for contamination determinations. Precision and recall is tolerated at 2.5% contamination. (FIG. 1. - FIG.3).(0081 ] Table 1: Median mitochondrial DNA coverage of low depth DNA sequencing compared to nuclear DNA coverage.Example 2. Detection of Contamination in human skin samples.

[0082] Sequencing reads spanning a large range of average whole genome depth (0.39X- 102.49X) were generated from whole genome libraries prepared from human skin samples. Bacterial contamination was removed by classifying reads with a k-mer based approach against the human genome, which produced fastq files containing reads only of human origin. Using DRAGEN (Illumina, 2022), the fastq files were aligned to the human genome, GRCh38, BAM files generated, and contamination estimated via autosomal allele frequency when able. The median mitochondrial DNA coverage was observed to be significantly higher than the nuclear DNA coverage (FIG. 4). Haplocheck (Weissensteiner et al. 2021), which calls mtDNA variants, groups the variants into one or more haplotypes, assigns major and minor haplogroups based on the frequency of the haplotypes and predicts contamination estimate based on the frequency of the minor haplogroups when incompatibilities are observed, was run on the BAM files. In samples with an average whole genome depth lower than 3.77X, DRAGEN was not able to produce an estimate of sample contamination via its allele frequency method.

[0083] However, the haplotype-based method was able to estimate contamination, which was found to be as high as 11.8%. More specifically, FIG. 4 shows that the disclosed haplotypebased methods were able to detect contamination at sequencing depths less than 5x (including specific data points at 3.77x, 3.75x, 3.74x, 2.73x, 1.06x, 0.94x, and 0.39x), whereas the conventional autosomal approach (exemplified by DRAGEN) was unable to detect thecontamination. Thus, the utility of haplotype-based methods for increased sensitivity at lower whole genome depths is highlighted and exemplified and, when both haplotype and allele frequency based methods produce estimates, the correlation between these methods was 0.82 (Pearson). sfe sfe sfe sfe

[0084] These examples are provided for illustrative purposes only and not to limit the scope of the claims provided herein. While certain embodiments have been illustrated and described, it should be understood that changes and modifications can be made therein in accordance with ordinary skill in the art without departing from the technology in its broader aspects as defined in the following claims.

[0085] The present disclosure is not to be limited in terms of the particular embodiments described in this application. Many modifications and variations can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and compositions within the scope of the disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the appended claims. The present disclosure is to be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled. It is to be understood that this disclosure is not limited to particular methods, reagents, compounds, or compositions, which can of course vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0086] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.]0087[ All publications, patent applications, issued patents, and other documents referred to in this specification are herein incorporated by reference as if each individual publication, patent application, issued patent, or other document was specifically and individually indicated to be incorporated by reference in its entirety. Definitions that are contained in textincorporated by reference are excluded to the extent that they contradict definitions in this disclosure.

Claims

WHAT IS CLAIMED IS:

1. A method of detecting contamination in DNA sequencing comprising:(a) preparing a sequencing library from DNA extracted from a biological sample obtained from a subject;(b) performing low depth whole genome sequencing on the sequencing library, thereby obtaining a plurality of sequencing reads;(c) detecting in the plurality of sequencing reads one or more mitochondrial DNA (mtDNA) haplotypes and grouping the one or more mtDNA haplotypes into one or more haplogroups, wherein contamination is present during sequencing when two or more haplotypes are detected and grouped into two or more haplogroups and contamination is not present during sequencing when one haplotype is detected and grouped into one haplogroup.

2. The method of claim 1, wherein the DNA is total genomic DNA or comprises amplified regions of interest.

3. The method of claim 1 or 2, wherein the biological sample is selected from a tissue sample or a fluid sample.

4. The method of claim 3, wherein the fluid sample is selected from plasma, serum, or whole blood.

5. The method of any one of claims 1-4, wherein the low depth whole genome sequencing generates between about lOOx to about lOOOx mitochondrial genome coverage.

6. The method of any one of claims 1-5, wherein the low depth whole genome sequencing generates between about 0.2x to about 5x nuclear genome coverage.

7. The method of any one of claims 1-6, wherein the two or more haplogroups comprises a major haplogroup and one or more minor haplogroups.

8. The method of claim 7, wherein the one or more minor haplogroups comprises 2, 3, 4, 5, or 6, minor haplogroups.

9. The method of claim 8, wherein each individual minor haplogroup represents a source of contamination.

10. The method of any one of claims 1-9, wherein the one haplogroup is a major haplogroup.

11. The method of any one of claims 1-10, wherein the contamination is sample to sample cross contamination.

12. The method of any one of claims 1-11, wherein the plurality of sequence reads is concurrently used for genetic screening.

13. A method of detecting one or more haplotypes in mitochondrial DNA (mtDNA) comprising:(a) obtaining genomic DNA extracted from a sample obtained from a subject;(b) performing low pass next generation sequencing on the genomic DNA, thereby obtaining mtDNA sequence reads;(c) detecting one or more haplotypes in the mtDNA sequence reads.

14. The method of claim 13, wherein the low pass next generation sequencing is whole genome sequencing or whole exome sequencing.

15. The method of claim 13 or 14, wherein the low pass next generation sequencing generates between about lOOx to about lOOOx mitochondrial genome coverage.

16. The method of any one of claims 13-15, wherein the low pass next generation sequencing generates between about 0.2x to about 5x nuclear genome coverage.

17. The method of any one of claims 13-16, wherein the detecting the one or more haplotypes further comprises grouping the one of more haplotypes into (i) one major haplogroup, or (ii) one major haplogroup and one or more minor haplogroups.

18. The method of any one of claims 13-17, wherein the sample obtained from the subject is concurrently sequenced for genetic screening.

19. A method for detecting the presence or absence of contamination between one or more sequencing samples comprising:(a) preparing a first sequencing library from extracted DNA from a first subject;(b) performing low pass whole genome sequencing on the first sequencing library, thereby obtaining a plurality of sequencing reads from the first sequencing library;(c) preparing a second sequencing library from extracted DNA from a second subject;(d) performing low pass whole genome sequencing on the second sequencing library using the same sequencer used to perform low pass whole genome sequencing on the first sequencing library, thereby obtaining a plurality of sequencing reads from the second sequencing library;(e) detecting in the plurality of sequencing reads from the second sequencing library one or more mitochondrial DNA (mtDNA) haplotypes and grouping the one or more mtDNA haplotypes into one or more haplogroups, wherein contamination is present during sequencing when two or more haplotypes are detected and grouped into two or more haplogroups and contamination is not present during sequencing when one haplotype is detected and grouped into one haplogroup.

20. The method of claim 19, further comprising detecting in the plurality of sequencing reads from the first sequencing library one or more mitochondrial DNA haplotypes.

21. The method of claim 19 or 20, wherein extracted DNA from the first subject, extracted DNA from the second subject, or extracted DNA from both the first subject and the second subject is obtains for a biological sample independently selected from a tissue sample or a fluid sample.

22. The method of claim 21, wherein the fluid sample is selected from plasma, serum, or whole blood.

23. The method of any one of claims 19-22, wherein the low depth whole genome sequencing generates between about lOOx to about lOOOx mitochondrial genome coverage.

24. The method of any one of claims 19-23, wherein the low depth whole genome sequencing generates between about 0.2x to about 5x nuclear genome coverage.

25. The method of any one of claims 19-24, further comprising repeating (c)-(e) with extracted DNA from multiple further subjects.

26. The method of claim 25, wherein multiple further subjects comprises at least a third subject, at least a fourth subject, or at least a fifth subject.

27. The method of any one of claims 19-26, wherein the contamination is sample to sample contamination.

28. The method of claim 27, wherein the contamination occurs during the sequencing sample preparation.

29. The method of claim 27, wherein the contamination occurs during the low pass whole genome sequencing.

30. The method of any one of claims 19-29, wherein the two or more haplogroups comprises one major haplogroup and one or more minor haplogroups.

31. The method of any one of claims 19-29, wherein the one haplogroup is a major haplogroup.

32. The method of any one of claims 1-13 and 19-31, wherein the contamination is between about 0.5% to about 5%.

33. The method of claim 32, wherein the contamination is between about 1% to about 2.5%.

Citation Information

Patent Citations

  • Mitochondrial DNA Quality Control

    US20220042091A1