Method for full-coverage amplification of DNA sample, rapid DNA library building method, sequencing method and application thereof
By employing ultra-high multiplex PCR technology and a simplified library construction method, we have achieved efficient and low-cost single-cell mtDNA amplification and library construction, solving the problem of detecting low-frequency mtDNA mutations in existing technologies and advancing our understanding of the clinical significance of mtDNA mutations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies struggle to achieve high-throughput, low-cost, and high-coverage single-cell mtDNA sequencing, and conventional methods are complex and limited to specific sequencing modes, making it impossible to accurately detect low-frequency mtDNA mutations.
Employing ultra-high multiplex PCR technology, multiple primer pairs are designed for full-coverage amplification. Magnetic bead purification is used as an alternative to magnetic bead purification, simplifying the library construction process. Efficient library construction is achieved through two PCR steps, making it suitable for various sequencing platforms.
This technology enables rapid and economical mtDNA amplification and library construction with full coverage, improves detection sensitivity, and allows for accurate detection of mtDNA mutations at the single-cell level. This advances our understanding of the role of mtDNA mutations in the aging process and provides a foundation for clinical treatment.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of DNA library preparation and sequencing, specifically to a method for full-coverage amplification of DNA samples, a rapid DNA library preparation method, a sequencing method, and their applications. Background Technology
[0002] As vital organelles in eukaryotes (including humans), mitochondria play a crucial role in maintaining system stability and survival. Mitochondria are the central hubs of cellular energy production and metabolism, primarily providing chemical energy for cellular activities through oxidative phosphorylation and adenosine triphosphate (ATP) synthesis. In addition to providing energy, mitochondria participate in processes such as cell differentiation, cell signaling, and apoptosis, and have the ability to regulate cell growth and the cell cycle. Recent studies have also revealed the role of mitochondria in regulating epigenetics and immune responses, playing a vital role in the homeostasis of organisms.
[0003] mtDNA is a unique genetic material found in the mitochondria of eukaryotic cells. Unlike nuclear DNA, mtDNA has a circular structure, and each cell contains multiple copies. mtDNA encodes an important component of the mitochondrial respiratory chain, responsible for producing cellular energy through oxidative phosphorylation. mtDNA has a high mutation rate, lacks effective repair mechanisms, and is unevenly distributed during cell division, leading to the easy generation and accumulation of mtDNA mutations within individual cells. These mutations can result in a high frequency of mtDNA mutations in individual cells, leading to mitochondrial dysfunction, which in turn affects cellular energy metabolism, oxidative stress, and other key biological processes. It is closely associated with the occurrence and progression of various diseases, including cancer, neurodegenerative diseases, autoimmune diseases, aging, and age-related diseases.
[0004] Most confirmed mtDNA mutations associated with human diseases are in a so-called heterozygous state, meaning that both mutated and wild-type mtDNA are present in cells or tissues. Whole-genome sequencing (WGS) with deep and uniform coverage of mtDNA can be used to measure low levels of heterozygosity. However, the cost-effectiveness of using large sample sizes for mtDNA studies is limited. Even with available genomic datasets, mtDNA analysis is often constrained by the limitations of the original studies, limiting the exploration of key features such as heterozygosity dynamics. A common approach in directly designing techniques for detecting mtDNA heterozygosity involves testing large samples, which is relatively straightforward. This approach is effective for identifying shared mutations but may struggle to accurately detect and assess mutations present in some cells but absent in most. Therefore, there is a risk of underestimating the frequency of such mutations, leading to misinterpretations of their potential impact on cellular function. Given the significant variability of mtDNA mutations, high-throughput single-cell mtDNA sequencing technologies are urgently needed.
[0005] Recent studies have shown that mtscATAC-seq based on the 10X Genomics platform can simultaneously detect mtDNA heterozygosity and chromatin accessibility at the single-cell level. However, this technology is costly, and the inherent allele loss problem associated with scATAC-seq is often overlooked when determining the mutation profile of the mitochondrial genome. Furthermore, it is estimated that only about 20% of the total mtscATAC-seq reads from a single cell are mtDNA fragments, highlighting the problem of low mtDNA coverage.
[0006] Another typical single-cell mtDNA sequencing technology is scSTAMP, which can achieve high sequencing depth and coverage of mtDNA at the single-cell level. However, it involves complex steps such as long fragment amplification, probe capture, enzyme linking, and amplification, making its protocol lengthy and complicated. In addition, the resulting library is only compatible with PE250 sequencing mode, limiting its widespread application.
[0007] Therefore, it is necessary to develop a cell-level DNA library construction method, and then to obtain a DNA sequencing method with deep and uniform coverage of the whole genome. Summary of the Invention
[0008] This application provides a method for full-coverage amplification of DNA samples, a rapid DNA library construction method, a sequencing method, and their applications. This application develops a novel mtDNA sequencing method, called 123-seq, aiming to provide a rapid, reliable, and cost-effective method for analyzing mtDNA heterozygosity at the level of large amounts of genomic DNA and even single cells. This application achieves this goal through ultra-high multiplex PCR technology, covering the entire mtDNA genome, using multiple primer pairs. Furthermore, this application creatively eliminates the magnetic bead purification process during targeted library construction after multiplex PCR, simplifying the library construction process to only two PCR steps requiring three consecutive reagent additions; this method maximizes efficiency while minimizing library construction time and sequencing costs. In addition, this method was further applied to evaluate single-cell mtDNA variations from samples from elderly individuals, revealing a large number of mtDNA mutations in T cells. Overall, this application provides a rapid and powerful mtDNA analysis method, offering new insights into the role of mtDNA heterozygosity in the aging process. By facilitating large-scale population studies, this technology is expected to advance the understanding of the clinical significance of mtDNA mutations in the field and provide information for future treatment strategies.
[0009] This application involves the following:
[0010] 1. A method for full-coverage amplification of a DNA sample, comprising the following steps:
[0011] Low-temperature denaturing PCR: DNA samples are contacted with multiple sets of pre-amplification primer pairs for PCR amplification. The amplification product of any one of the multiple sets of pre-amplification primer pairs overlaps with the amplification product of the pre-amplification primer pair adjacent to it in the 5' direction by at least 20 bases, at least 25 bases, at least 30 bases, at least 35 bases, at least 40 bases, at least 45 bases, at least 50 bases, at least 55 bases, or at least 60 bases.
[0012] Preferably, during low-temperature denaturing PCR amplification, the amplification product of any one of the multiple sets of pre-amplification primer pairs and the amplification product of the pre-amplification primer pair adjacent to it in the 5' direction have an overlap of 20-70 bases, 25-65 bases, 30-60 bases, 35-55 bases, or 40-50 bases.
[0013] 2. According to the full-coverage amplification method described in item 1, the target sequence length amplified by any one of the multiple sets of pre-amplification primer pairs is at most 450 bases, at most 400 bases, at most 350 bases, or at most 300 bases.
[0014] Preferably, the target sequence length amplified by any one of the multiple sets of pre-amplification primer pairs is 350-400 bases, 280-400 bases, 170-250 bases, 170-350 bases, or 100-170 bases.
[0015] 3. The method for full-coverage amplification according to item 1 or 2, during low-temperature denaturing PCR, includes the following steps:
[0016] Denaturation for 20-60 seconds, annealing for 0.2-8 minutes, extension at the denaturation temperature for 5-15 seconds, followed by extension;
[0017] Preferably, the denaturation process specifically includes the following steps: pre-denaturation for 15-45 seconds, and denaturation for 5-15 seconds.
[0018] 4. According to the full-coverage amplification method described in any one of items 1-3, the annealing time for low-temperature denaturing PCR is 0.2-6 min;
[0019] Preferably, the annealing time is 2-5 minutes.
[0020] 5. The method for full-coverage amplification according to any one of items 1-4, wherein the DNA sample is derived from cells, and the cell source includes any one or more of animals, plants, non-cellular organisms, prokaryotes, and fungi;
[0021] More preferably, the cells are derived from human cells, and the sources of the human cells include any one or more of blood cells, immune cells, liver cells, tumor cells, stem cells, nerve cells, muscle cells, and skin cells;
[0022] Preferably, the method for obtaining the cells includes any one or a combination of the following methods: flow cytometry cell sorting, cell suspension dilution, mechanical separation, micromanipulation, tablet pressing, microfluidics, and laser cutting.
[0023] 6. The method for full-coverage amplification according to any one of items 1-5, wherein the DNA sample is selected from any one or more of nuclear DNA, mitochondrial DNA (mtDNA), chloroplast DNA, cytoplasmic DNA, body fluid DNA, blood DNA, plasmid DNA, and DNA obtained by reverse transcription;
[0024] Preferably, the DNA sample is derived from mitochondrial DNA (mtDNA) from a single cell or multiple cells.
[0025] 7. The method for obtaining mitochondrial DNA according to the full-coverage amplification method described in item 6 includes the following steps: treating cells with cell lysis buffer to release mitochondrial DNA;
[0026] Preferably, the cell lysis buffer comprises a nonionic surfactant;
[0027] More preferably, the components of the cell lysis buffer include any one or more of a buffer, an inorganic salt, albumin or a substitute thereof, and a reducing agent;
[0028] More preferably, the volume percentage (v / v) of the nonionic surfactant in the cell lysis buffer is 0.05-0.3%, 0.05-0.15%, 0.1-0.2%, 0.1-0.15%, 0.15-0.2%, 0.15-0.25%, or 0.2-0.3%; and / or,
[0029] The reducing agent in the cell lysis buffer is 0.2-1.7 mM, 0.2-1.4 mM, 0.2-1 mM, 0.5-1.5 mM, 0.8-1.2 mM, 0.2-0.7 mM, or 1.2-1.7 mM; and / or,
[0030] The inorganic salt content in the cell lysis buffer is 5-20 mM, 5-15 mM, 8-10 mM, 10-18 mM, 10-15 mM, 12-18 mM, or 12-20 mM; and / or,
[0031] The volume percentage (v / v) of albumin or its substitute in the cell lysis buffer is 0.1-3%, 0.1-1%, 0.1-2%, 0.5-1%, 0.5-2%, 0.5-3%, or 0.8-1.5%.
[0032] 8. In the method of full-coverage amplification according to item 7, when the cells are treated with cell lysis buffer, the conditions include: heat treatment at 55-70°C for 5-15 min;
[0033] Preferably, the heat treatment is carried out at 60-68℃ for 8-12 minutes.
[0034] 9. The method for full-coverage amplification according to any one of items 1-8, wherein the DNA sample is mitochondrial DNA;
[0035] During low-temperature denaturing PCR amplification, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 170-250 bases. The pre-amplification primer pairs include pre-amplification primer pairs 1 to 123 as shown below:
[0036] Preamplification primer pairs F1, R1, F2, R2, F3, R3, ..., Fx, Rx; where x is a natural number from 1 to 123.
[0037] Alternatively, the preamplification primer pair may comprise a combination of preamplification primer pairs as described above, wherein at least one primer is replaced with a primer having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with its sequence.
[0038] 10. The method for full-coverage amplification according to any one of items 1-8, wherein the DNA sample is mitochondrial DNA;
[0039] During low-temperature denaturing PCR amplification, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 280-400 bases. The pre-amplification primer pairs include those shown below.
[0040] Preamplification primer pairs F1, R2, F2, R3, F3, R4, ..., F... n-1 R n Where n is a natural number from 2 to 123;
[0041] Alternatively, the preamplification primer pair may comprise a combination of preamplification primer pairs as described above, wherein at least one primer is replaced with a primer having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with its sequence.
[0042] 11. The method for full-coverage amplification according to any one of items 1-8, wherein the DNA sample is mitochondrial DNA;
[0043] During low-temperature denaturing PCR amplification, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 280-400 bases, wherein the pre-amplification primer pairs include;
[0044] Preamplification primer pairs F1, R2, F3, R4, F5, R6, ..., F... m-1 R m Where m is an even number from 2 to 123;
[0045] Alternatively, the preamplification primer pair may comprise a combination of preamplification primer pairs as described above, wherein at least one primer is replaced with a primer having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with its sequence.
[0046] 12. A rapid DNA library construction method for full sequence coverage amplification, comprising the following steps:
[0047] Full-coverage amplification was performed using any one of the methods described in items 1-11 to obtain PCR products;
[0048] PCR products were treated with single-stranded DNA exonuclease to remove residual primers after PCR.
[0049] Secondary PCR: This involves bringing the exonuclease treatment product into contact with the secondary PCR primer pair to perform PCR and obtain a library.
[0050] 13. According to the rapid DNA library construction method described in item 12, the 5' end of the forward primer and the 3' end of the reverse primer in any set of pre-amplification primer pairs also include a shared sequence;
[0051] Preferably, the 5' end of the forward primer and the 3' end of the reverse primer in the secondary PCR primer pair further include the shared sequence and the barcode sequence.
[0052] 14. According to the rapid DNA library construction method described in item 13, the shared sequence satisfies the following condition:
[0053] a) The shared sequence and the genome and / or mitochondrial genome from which the DNA sample is derived do not contain the binding sequence;
[0054] b) The annealing temperature of the shared sequence is 70-80℃;
[0055] c) The shared sequence does not end with a repeating sequence of at least 3, at least 5, or at least 8 consecutive bases;
[0056] d) The length of the shared sequence is 18-25, 18-20, 20-25, or 20-22 bases.
[0057] 15. The rapid DNA library construction method according to item 13 or 14, wherein the shared sequence is selected from any one of the following sequences: such as the sequence shown in SEQ ID NO:247, such as the sequence shown in SEQ ID NO:448, such as the sequence shown in SEQ ID NO:449, and such as the sequence shown in SEQ ID NO:450;
[0058] Alternatively, the shared sequence is a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with the sequence shown in SEQ ID NO:247, SEQ ID NO:448, SEQ ID NO:449, or SEQ ID NO:450.
[0059] 16. The rapid DNA library construction method according to any one of items 13-15, wherein the barcode sequence is selected from any one or more of SEQ ID NO:250 to SEQ ID NO:441;
[0060] Or the barcode sequence is a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with any of the sequences shown in SEQ ID NO:250 to SEQ ID NO:441.
[0061] 17. The rapid DNA library construction method according to any one of items 12-16, wherein the single-stranded DNA exonuclease is selected from any one or more of exonuclease I, exonuclease T, exonuclease VII, and Lambda exonuclease;
[0062] Preferably, the single-stranded DNA exonuclease is exonuclease I; more preferably, it is thermosensitive exonuclease I.
[0063] 18. According to any one of the rapid DNA library construction methods described in items 12-17, when treating PCR products with single-stranded DNA exonuclease, the treatment method includes the following steps: treating at 30-38℃ for 5-15 min.
[0064] 19. The rapid DNA library construction method according to any one of items 12-18, wherein the PCR product is treated with a single-stranded DNA exonuclease in the presence of a buffer, wherein the buffer is selected from any one of NEB exonuclease I buffer, Thermoscientific buffer, and TaKaLa buffer.
[0065] 20. A hybrid library, wherein the hybrid library is constructed using any one of the methods described in items 12-19.
[0066] 21. A sequencing method for sequencing a DNA sample, said DNA sample comprising at least one sequence fully covering an amplified DNA fragment;
[0067] The sequencing method for amplifying the DNA fragment with full sequence coverage includes the following steps:
[0068] A mixed library was obtained by using the rapid DNA library construction method described in any one of items 12-19.
[0069] Library purification;
[0070] Sequencing.
[0071] 22. According to the sequencing method described in item 21, the DNA sample is a sequence-completely amplified DNA fragment.
[0072] 23. According to the sequencing method described in item 21 or 22, when any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 170-250 bases, the resulting library is suitable for any one of the following sequencing platforms: HiSeq 3000, HiSeq 4000, HiSeq X10, NextSeq 1000, NextSeq 2000, MiniSeq, and NovaSeq 6000 in PE150 sequencing mode;
[0073] When amplifying a target sequence of 280-400 bases using any one of the multiple sets of pre-amplification primer pairs, the resulting library is suitable for any one of the following testing platforms: any one of the following sequencing platforms: HiSeq 3000, HiSeq 4000, HiSeq X10, NextSeq 1000, NextSeq 2000, MiniSeq, and NovaSeq 6000 in PE250 sequencing mode.
[0074] 24. An application of the full-coverage amplification method according to any one of items 1-11 and / or the rapid DNA library preparation method according to any one of items 12-19 and / or the sequencing method according to any one of items 21-23, the application including genomic mutation analysis.
[0075] Invention Effects
[0076] 1. This application develops a method for target sequence amplification based on a sequence full coverage concept. It utilizes a special primer design method to achieve full-coverage amplification of the target sequence requiring amplification. For example, for mtDNA, 60-130 pairs of pre-amplification primers are specially designed to achieve full-coverage amplification of mtDNA in single cells and multi-cell structures. This ultimately enables the sequencing of any gene mutation present on the mtDNA.
[0077] 2. Based on the sequence-wide amplification method of this application, a hybrid library is further constructed. This application develops a special library construction method. Specifically, it first solves the specificity problem between target sequences by designing shared sequences on pre-amplification primers and combining this with a longer extension time used in low-temperature PCR. After the first PCR, this application can remove impurities by a single digestion treatment with exonuclease I without performing conventional steps such as magnetic bead purification. The further second PCR uses shared sequences as primers and high-temperature PCR to achieve exponential amplification of the target sequences, thus solving the compatibility problem between target sequences. Finally, a hybrid library for sequencing is obtained. This method has the characteristics of low cost and short experimental cycle, and it is suitable for the construction of libraries of mixed cellular genomic DNA or single-cell samples. The hybrid library of this application has high detection sensitivity and can achieve high-precision single-cell level sequencing of mtDNA in different cell types (including HEK-293T cells, Quatt cell line M8 cells, and peripheral blood-derived T cells).
[0078] 3. The amplification and library construction technology developed in this application can be used to discover the widespread presence of mtDNA mutations in individual T cells of healthy elderly individuals, and mtDNA mutations in elderly individuals may be one of the causes of weakened T cell function. This suggests that intervening in the generation and accumulation of mtDNA mutations may become a potentially important approach and method for the prevention and treatment of age-related diseases; providing a new foundation and conditions for the prevention and treatment of such diseases. Attached Figure Description
[0079] Figure 1 A schematic diagram of a full-coverage (or overlap) design when multiple pre-amplified primer pairs bind to the target sequence.
[0080] Figure 2 This application includes a diagram illustrating the library construction and a flowchart illustrating the data processing.
[0081] Figure 3 Agarose gel electrophoresis showed that the library was constructed without exonuclease I purification between low-temperature denaturing PCR and secondary PCR. Among them, lane 1 of the genomic DNA of 293 cells was called 293-rep-1, lane 2 of the genomic DNA of 293 cells was called 293-rep-2, and NTC was the template-free control group. In the two rounds of PCR, the one with the Exo I treatment step was called Exo I, and the one without was called no Exo I.
[0082] Figure 4 Accurate detection of mtDNA variations in mixed samples; among which, Figure 4A is an agarose gel electrophoresis image of the 123plex PCR product, showing the molecular weight marker (Lane Marker), 293 cell genomic DNA repeat 1 (Lane 293-rep-1), 293 cell genomic DNA repeat 2 (Lane 293-rep-2), 293 cell genomic DNA repeat 3 (Lane 293-rep-3), and no genomic DNA (Lane NTC). Figure 4 B compares the mtDNA mapping ratio before and after Numts removal, no_RtN refers to the mtDNA mapping ratio before Numts removal, and RtN refers to the mtDNA mapping ratio after Numts removal. Figure 4 C is the Pearson correlation analysis of three repeated detections of mutations in the same sample after library construction using the 123-seq technique.
[0083] Figure 5 Heterogeneity determination of mtDNA variations in mixed samples; among which, Figure 5 A represents the theoretical and detected mtDNA heterogeneity in a mixture of GM19223 and GM12878 cells, with a Pearson correlation coefficient of 0.98 and a P-value (two-tailed) of <0.0001. Figure 5 B is a heatmap of the detected and theoretical mtDNA variation heterogeneity in a mixture of genomic DNA from GM19223 and GM12878 cells.
[0084] Figure 6 Analysis of the mapping distribution of sequencing reads on each chromosome after Numts removal; no_RtN refers to the number of sequencing reads before Numts removal that are mapped to mtDNA, and RtN refers to the number of sequencing reads after Numts removal that are mapped to mtDNA. Chr1 refers to chromosome 1; Chr2 refers to chromosome 2; Chr3 refers to chromosome 3; Chr4 refers to chromosome 4; Chr5 refers to chromosome 5; Chr6 refers to chromosome 6; Chr7 refers to chromosome 7; Chr8 refers to chromosome 8; Chr9 refers to chromosome 9; Chr10 refers to chromosome 10; Chr11 refers to chromosome 11; Chr12 refers to chromosome 12; Chr13 refers to chromosome 13; Chr14 refers to chromosome 14; Chr15 refers to chromosome 15; Chr16 refers to chromosome 16; Chr17 refers to chromosome 17; Chr18 refers to chromosome 18; Chr19 refers to chromosome 19; Chr20 refers to chromosome 20; Chr21 refers to chromosome 21; Chr22 refers to chromosome 22; ChrX refers to chromosome X; ChrY refers to chromosome Y; ChrM refers to the mitochondrial genome.
[0085] Figure 7The performance of single-cell mtDNA mutation analysis using the method of Example 1; wherein, Figure 7 A is to screen lysis buffers using qPCR to assess their compatibility with mtDNA amplification; Figure 7 B is the amplification of a 373bp fragment of the mitochondrial ND5 gene in a single 293T cell using lysis buffer 1. Here, 293T genomic DNA refers to 10 ng of 293T cell genomic DNA template, and NTC refers to a template-free control.
[0086] Figure 8 Single-cell mtDNA coverage and mutation distribution using the method of Example 1; wherein, Figure 8 A is the single-cell mtDNA coverage analysis of HEK-293T cells; Figure 8 B is an mtDNA mutation with an average VAF exceeding 20% in two different wells of the same cells from the Quatet cell line M8.
[0087] Figure 9 Correlation analysis of mtDNA mutations in HEK-293T single cells after sequencing using the method described in Example 1.
[0088] Figure 10 Flow cytometry sorting strategy for T cells.
[0089] Figure 11 Distribution of .mtDNA mutations (horizontal axis) and distribution of its VAFs (vertical axis).
[0090] Figure 12 Analysis of the proportion of mutation sites in specific functional categories.
[0091] Figure 13 .mtDNA mutation status; Figure 13 A represents the percentage of cells carrying different numbers of mtDNA mutations (vertical axis); Figure 13 B represents the percentage of mutation sites with a specific VAF range (vertical axis) among all mutation sites (horizontal axis). Specific implementation methods
[0092] It should be noted that certain terms are used in the specification and claims to refer to specific components. Those skilled in the art will understand that different terms may be used to refer to the same component. This specification and claims do not distinguish components based on differences in terminology, but rather on differences in their functions.
[0093] As used throughout the specification and claims, the terms "comprising" or "including" are open-ended and should be interpreted as "comprising but not limited to". The subsequent descriptions in the specification are preferred embodiments for carrying out this application; however, these descriptions are for the purpose of understanding the general principles of the specification and are not intended to limit the scope of this application. The scope of protection of this application shall be determined by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0094] The terms “preamplified primer pair” and “preamplified primer” are used interchangeably in this application. As used herein, “primer” generally refers to an oligonucleotide. It is used, for example, to initiate nucleotide extension, ligation, and / or synthesis; for example, in the synthetic step of a polymerase chain reaction or in primer extension techniques used in certain sequencing reactions. Primers can also be used in hybridization techniques as a means of providing complementarity between a locus and a capturing oligonucleotide for the detection of specific nucleic acid regions.
[0095] In this application, "low-temperature denaturing PCR" refers to a type of PCR with a long annealing time, specifically within the range of 0.2-8 minutes. Low-temperature denaturing PCR is used to amplify all target genes.
[0096] The terms “PCR” and “polymerase chain reaction” are used interchangeably in this application.
[0097] As used herein, “nucleic acid” includes at least two nucleotide monomers linked together. Examples include, but are not limited to, DNA, such as genome or cDNA; RNA, such as mRNA, sRNA, or rRNA; or hybrids of DNA and RNA. As will be understood by those skilled in the art, nucleic acids may have naturally occurring nucleic acid structures or non-naturally occurring nucleic acid analog structures. Nucleic acids may contain phosphodiester bonds; however, in some embodiments, nucleic acids may have other types of backbones, including, for example, phosphoramide, thiophosphate, dithiophosphate, O-methylphosphoramide, and peptide nucleic acid backbones and links. Nucleic acids may have a positive backbone; a nonionic backbone; and a non-ribosyl backbone. Nucleic acids may also contain one or more carbon-cyclic sugars. Nucleic acids may contain portions of double-stranded and single-stranded sequences, for example, as demonstrated by forked linkers. Nucleic acids can contain any combination of deoxyribonucleotides and ribonucleotides, as well as any combination of bases, including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, isoguanine, and base analogs such as nitropyrrole (including 3-nitropyrrole) and nitroindole (including 5-nitroindole).
[0098] As used herein, a “nucleotide sequence” includes the order and type of nucleotide monomers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of forms, including, for example, descriptions, images, electronic media, series of symbols, series of numbers, series of letters, series of colors, etc. Information can be represented, for example, at single-nucleotide resolution, at higher resolution (e.g., indicating the molecular structure of nucleotide subunits), or at lower resolution (e.g., indicating chromosomal regions, such as haplotype blocks). A series of letters “A,” “T,” “G,” and “C” is a well-known sequence that represents DNA that can be associated with the actual sequence of a DNA molecule at single-nucleotide resolution. Similar notation is used for RNA, except that “U” is used instead of “T” in the series.
[0099] The term "barcode" is short for "barcode". A "barcode" refers to a nucleic acid sequence that is unique and completely independent of the target nucleic acid sequence. As understood by those skilled in the art, barcodes can label molecules, such as nucleic acids and / or peptides. The material used for labeling can be associated with information. Barcodes may be referred to as sequence identifiers (i.e., sequence-based barcodes or sequence indexes). Barcodes can be specific nucleotide sequences, artificial sequences, or naturally occurring sequences. Barcodes can be introduced into any region of a polynucleotide; this region can be known or unknown. Barcodes can be added to any position on a nucleotide; for example, the 5' end of a nucleotide; or, for example, the 3' end of a nucleotide; or, for example, between the 5' and 3' ends of a nucleotide.
[0100] The terms "nucleic acid," "polynucleotide," and "nucleotide sequence" are used interchangeably and do not have a specific limit on the number of bases. They refer to polymeric forms of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogues. "Oligonucleotide" and "oligonucleotide" are used interchangeably and refer to short polynucleotides having no more than about 50 nucleotides.
[0101] As used herein, “complementarity” refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid via conventional Watson-Crick base pairing. The complementarity percentage indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (i.e., Watson-Crick base pairing) with a second nucleic acid (e.g., approximately 5, 6, 7, 8, 9, 10 / 10, representing approximately 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively). “Complete complementarity” means that all consecutive residues in the nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in the second nucleic acid sequence. As used herein, “substantially complementary” refers to a degree of complementarity of at least approximately 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of approximately 40, 50, 60, 70, 80, 100, 150, 200, 250 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0102] As used herein, “strict conditions” for hybridization refer to conditions under which nucleic acids complementary to the target sequence hybridize primarily with the target sequence and substantially with no hybridization with non-target sequences. Strict conditions are typically sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature at which it hybridizes specifically with its target sequence.
[0103] The term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex, which is stabilized by hydrogen bonds between the bases of the nucleotide residues. Hydrogen bonding can occur through Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. A sequence that can hybridize with a given sequence is called the "complementary sequence" of that given sequence.
[0104] The term “to bring into contact” means to place or bring into contact with something, to be in contact with something, or to make contact with something. As used herein, “contact” means a state or situation of touching or being in direct or partial proximity to something.
[0105] As used herein, “binding” should be understood as a non-covalent interaction between macromolecules (e.g., the non-covalent interaction between a primer and a target nucleic acid; the non-covalent interaction between a protein and guide RNA, etc.). When in a non-covalent interaction state, macromolecules are referred to as “associated,” “interacting,” or “binding” (e.g., when molecule X is said to interact with molecule Y, it means that molecule X binds to molecule Y in a non-covalent manner). Not all components of a binding interaction need to be sequence-specific (e.g., in contact with phosphate residues in the DNA backbone), but some parts of a binding interaction can be sequence-specific.
[0106] "Sequencing" generally refers to any and all biochemical methods that can be used to determine the sequence of nucleotide bases in nucleic acids.
[0107] The term "library" refers to a collection of nucleic acids. In some implementations, a library includes a target sequence and a barcode. In some implementations, the size of the library is approximately [number] bases.
[0108] The terms "mitochondrial DNA" and "mtDNA" are used interchangeably. mtDNA is a double-stranded, closed-circular DNA, with a heavy chain on the outer loop and a light chain on the inner loop. mtDNA is primarily found within cells; a single cell contains approximately 10 to 1000 mitochondria, or even more; each mitochondrion contains approximately 2 to 10 mtDNA molecules. The genes in the mitochondrial genome are arranged very compactly, resulting in high gene utilization. The intergenic spacers in human mtDNA total only 87 bp, accounting for 0.5% of the total mtDNA length. mtDNA generally does not contain introns, but some eukaryotic mtDNA does.
[0109] It should be understood that the embodiments of this application described herein include embodiments that are "composed of" and / or "substantially composed of". References to values or parameters of "about" herein include (and describe) variations of that value or parameter itself. For example, a reference to "about X" includes a description of "X".
[0110] As used herein, references to “not” values or parameters generally refer to and describe “except” values or parameters. For example, “The method is not used to treat type X cancer” means that the method is used to treat cancers other than type X.
[0111] As used in this article, the term “approximately XY” has the same meaning as “approximately X to approximately Y”.
[0112] As used herein and in the appended claims, the singular forms “a / an” and “the” include the plural objects unless the context clearly indicates otherwise. It should also be noted that claims may be drafted to exclude any optional elements. Therefore, this statement is intended as a preliminary basis for the use of exclusive terms such as “only” or “merely” in conjunction with the description of the elements of the claim, or for the use of the limitation of “no”.
[0113] As used herein, the term "and / or" in words such as "A and / or B" is intended to include both A and B; A or B; A (alone); and B (alone). Similarly, as used herein, the term "and / or" in words such as "A, B and / or C" is intended to include each of the following embodiments: A, B and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0114] In this application, "overlap" means that the amplification products of two adjacent pre-amplification primer pairs have a certain number of identical base sequences. For example, the 3' end of the amplification product of pre-amplification primer pair A has a partial sequence of "...AACCTGT"; the 5' end of the amplification product of pre-amplification primer pair B, which is adjacent to this pre-amplification primer pair in the 5' direction, also has a partial sequence of "AACCTGT..."; that is, "AACCTGT" is the overlapping part of the amplification products of the two adjacent pre-amplification primer pairs A and B. It can also be understood that a partial sequence at the 5' end of the amplification product of pre-amplification primer pair A has a base sequence that is completely identical to a partial sequence at the 3' end of the amplification product of another pre-amplification primer pair C (which is adjacent to pre-amplification primer pair A in the 3' direction).
[0115] This application provides a method for full-coverage amplification of a DNA sample, comprising the following steps:
[0116] Low-temperature denaturing PCR: DNA samples are contacted with multiple sets of pre-amplification primer pairs for PCR amplification. The amplification product of any one of the multiple sets of pre-amplification primer pairs overlaps with the amplification product of the pre-amplification primer pair adjacent to it in the 5' direction by at least 20 bases, at least 25 bases, at least 30 bases, at least 35 bases, at least 40 bases, at least 45 bases, at least 50 bases, at least 55 bases, or at least 60 bases.
[0117] Preferably, during low-temperature denaturing PCR amplification, the amplification product of any one of the multiple sets of pre-amplification primer pairs and the amplification product of any adjacent set of pre-amplification primer pairs overlap by 20-70 bases, 25-65 bases, 30-60 bases, 35-55 bases, or 40-50 bases.
[0118] "Full coverage amplification" refers to amplifying all sequences in the target sequence, meaning the amplified product contains 100% of the base composition of the target sequence. Those skilled in the art should understand that the amplification method of this application does not impose excessive restrictions on the DNA sample. In some embodiments, the DNA sample is obtained from any one or more of, including but not limited to, blood, serum, tissue fluid (including cerebrospinal fluid and human lymph), urine, feces, cells, or tissues. In some embodiments, the DNA sample is derived from cells. The cells can be single-celled or multi-celled. In some embodiments, the cells are derived from any one or more of, including but not limited to, animals, plants, non-cellular organisms (e.g., DNA viruses), prokaryotes (including bacteria and microalgae, such as cyanobacteria and green algae), and fungi. In some embodiments, the cells are human cells, derived from any one or more of, including but not limited to, blood cells (e.g., peripheral blood cells), immune cells, hepatocytes, tumor cells, stem cells, blood cells, nerve cells, zygotes, muscle cells (e.g., cardiomyocytes), and skin cells. Preferably, the cells are multicellular, consisting of 2-4 cells. In this application, no restrictions are placed on the method of obtaining cells. Any method that can obtain cells can be used to obtain the desired single or multiple cells. Exemplary methods include, but are not limited to, any one or more of the following: flow cytometry, cell suspension dilution, mechanical separation, micromanipulation, smearing, microfluidics, and laser cutting. In some embodiments, the DNA includes, but is not limited to, any one or more of the following: nuclear DNA, mitochondrial DNA (mtDNA), chloroplast DNA, cytoplasmic DNA, body fluid DNA, blood DNA, plasmid DNA, DNA obtained by reverse transcription, and viral DNA. Mitochondrial DNA (mtDNA) is particularly important because its genome contains virtually no introns; therefore, when studying mtDNA, it is necessary to sequence almost all gene fragments on the mtDNA. In some embodiments, when the DNA sample is mtDNA, multiple sets of pre-amplification primer pairs amplify the mtDNA with complete sequence coverage. Those skilled in the art should understand that the DNA sample can also be other DNA samples requiring complete sequence coverage amplification, such as nuclear DNA or chloroplast DNA. In some implementations, the DNA sample is one or more DNA strands; for example, the DNA sample is a complete mtDNA or a combination of multiple mtDNA strands; for example, the DNA sample is a nuclear DNA (requiring full sequence coverage amplification) or a combination of multiple nuclear DNA strands (requiring full sequence coverage amplification); for example, the DNA sample is any combination of one or more mtDNA strands and one or more nuclear DNA strands (requiring full sequence coverage amplification).Those skilled in the art should understand that the method of this application is also applicable to the amplification of DNA samples containing both gene fragments requiring full sequence coverage amplification and gene fragments not requiring full sequence coverage amplification; that is, the full coverage amplification in this application is determined according to the amplification requirements. For gene fragments in DNA samples that do not require full sequence coverage amplification, such as gene fragments containing at least one intron (without amplifying the intron sequence), primers can be designed using conventional methods and amplified using the method of this application or other methods.
[0119] When designing primer sets for "full coverage amplification," the amplification product of any one of the multiple pre-amplification primer pairs overlaps with the amplification product of the pre-amplification primer pair adjacent to it in the 5' direction by at least 20, 25, 30, 35, 40, 45, 50, 55, or 60 bases. The purpose is that if a mutation occurs in the primer-binding region, causing a primer to fail to bind to the target sequence, then... Adjacent primer pairs (with overlapping amplification products) will amplify this target sequence, resulting in the complete sequence of the DNA sample. This facilitates subsequent DNA genome assembly to obtain any gene mutation information within the DNA sample. Furthermore, when the DNA sample is mtDNA, this primer design effectively addresses the issues of low coverage and poor sequencing depth in current mtDNA sequencing technologies, enabling the analysis of mtDNA heterozygosity at the level of large amounts of genomic DNA and even single cells. An exemplary explanation is as follows... Figure 1 As shown, the process is illustrated using adjacent F1 / R1 and F2 / R2 pairs as examples among multiple pre-amplification primer pairs; when these two pre-amplification primer pairs are amplified simultaneously, three amplification products will appear, such as... Figure 1 Product 1 (obtained by amplification with F1 and R1 primers), product 2 (obtained by amplification with F2 and R2 primers) and product 3 (obtained by amplification with F2 and R1 primers); the overlapping part of product 1 and product 2 is product 3.
[0120] In some embodiments, the amplification product of any one of the multiple sets of pre-amplification primer pairs overlaps with the amplification product of any adjacent set of pre-amplification primer pairs by at least 20 bases, at least 25 bases, at least 30 bases, at least 35 bases, at least 40 bases, at least 45 bases, at least 50 bases, at least 55 bases, or at least 60 bases. In some embodiments, the amplification product of any one of the multiple sets of pre-amplification primer pairs overlaps with the amplification product of any adjacent set of pre-amplification primer pairs by any number of bases within the range of 20-70 bases, 25-65 bases, 30-60 bases, 35-55 bases, 40-50 bases, or 20-70 bases. In a specific embodiment, the amplification product of any one of the multiple sets of preamplification primer pairs and the amplification product of any adjacent set of preamplification primer pairs overlap by 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 bases, or overlap by any number of bases within the range of 20-70 bases.
[0121] In some embodiments, the annealing time for low-temperature denaturing PCR is 0.2-8 min; in further embodiments, the annealing time for low-temperature denaturing PCR is 0.2-1 min, 0.5-1.5 min, 1.2-2.3 min, 2-3.5 min, 3.2-4 min, 3.8-5 min, 4.5-6.9 min, 6.7-7.4 min, or 7.2-8 min. In some specific embodiments, the annealing time for low-temperature denaturing PCR is 0.2 min, 0.8 min, 2 min, 3 min, 4 min, 5 min, 6 min, 7 min, or 8 min. Figure 1 As shown, the theoretical primer amplification fragment for product 3 is small, resulting in high amplification efficiency. To prevent the final product from being predominantly product 3, a long annealing time was used in the PCR program, allowing products 1 and 2 to be amplified simultaneously in large quantities. In some implementation schemes, denaturation during low-temperature denaturing PCR can be a single denaturation step or a pre-denaturation followed by further denaturation. In some implementation schemes, the pre-denaturation is 15-45 seconds, and the denaturation is 5-15 seconds.
[0122] In some embodiments, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of up to 450 bases, up to 400 bases, up to 350 bases, or up to 300 bases. In some embodiments, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 350-400 bases, 280-400 bases, 170-250 bases, 170-350 bases, or 100-170 bases. In one specific embodiment, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, or 450 bases, or a target sequence of any length within the range of 170-450 bases.
[0123] In this application, target sequences of different lengths can be obtained by adjusting the primer sequence composition.
[0124] This application does not limit the method for obtaining mitochondrial DNA from cells, as long as it achieves the purpose of releasing mitochondrial DNA from cells. Exemplary methods for releasing mitochondrial DNA from cells include lysis.
[0125] In some embodiments, the method for obtaining mitochondrial DNA includes the following steps: treating cells with a cell lysis buffer to release mitochondrial DNA.
[0126] In some embodiments, the cell lysis buffer comprises a surfactant, preferably a nonionic surfactant; exemplary nonionic surfactants include Triton X-100 and IGEPACAL CA-630. In some embodiments, the volume percentage (v / v) of the surfactant in the cell lysis buffer is any range from 0.05-0.6%, 0.05-0.15%, 0.1-0.2%, 0.1-0.15%, 0.15-0.2%, 0.15-0.25%, 0.2-0.6%, 0.3-0.6%, 0.35-0.5%, 0.4-0.6%, or 0.05-0.6%. In one specific embodiment, the volume percentage (v / v) of the surfactant in the cell lysis buffer is any amount within the range of 0.05%, 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.4%, 0.5%, 0.6%, or 0.05-0.6%. In some embodiments, the cell lysis buffer comprises a reducing agent, which may be, for example, dithiothreitol (DTT) or citric acid. In some embodiments, the reducing agent in the cell lysis buffer is present in any concentration range within the range of 0.2-1.7 mM, 0.2-1.4 mM, 0.2-1 mM, 0.5-1.5 mM, 0.8-1.2 mM, 0.2-0.7 mM, 1.2-1.7 mM, or 0.2-1.7 mM; for example, it can be any concentration within the range of 0.2 mM, 0.4 mM, 0.8 mM, 1 mM, 1.2 mM, 1.7 mM, or 0.2-1.7 mM. In some embodiments, the cell lysis buffer comprises an inorganic salt selected from sodium chloride and magnesium chloride. In some embodiments, the inorganic salt content in the cell lysis buffer is any concentration range within the range of 5-20 mM, 5-15 mM, 8-10 mM, 10-18 mM, 10-15 mM, 12-18 mM, 12-20 mM, or 8-20 mM; for example, exemplary concentrations may be any concentration within the range of 8 mM, 10 mM, 13 mM, 15 mM, 18 mM, 20 mM, or 8-20 mM. In some embodiments, the cell lysis buffer comprises albumin or an albumin substitute, preferably bovine serum albumin (BSA). In some embodiments, the volume percentage (v / v) of albumin or an albumin substitute in the cell lysis buffer is any range within the range of 0.1-3%, 0.1-1%, 0.1-2%, 0.5-1%, 0.5-2%, 0.5-3%, 0.8-1.5%, or 0.1-3%. In one specific embodiment, the volume percentage (v / v) of albumin or albumin substitute in the cell lysis buffer is any amount within the range of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, or 0.3-3%.In some embodiments, the cell lysis buffer comprises 0.05-0.3% (v / v) of a nonionic surfactant; preferably, the nonionic surfactant is Triton X-100. In some embodiments, the cell lysis buffer comprises 0.2-0.6% (v / v) of a nonionic surfactant; preferably, the nonionic surfactant is IGEPEAL CA-630. In some embodiments, the cell lysis buffer comprises 0.2-1.7 mM of a reducing agent, 5-15 mM of a buffer, 5-18 mM of an inorganic salt, 0.05-0.3% (v / v) of a nonionic surfactant, and 0.1-3% (v / v) of albumin. In some embodiments, the cell lysis buffer comprises 5-15 mM of a buffer, 5-18 mM of an inorganic salt, 0.05-0.3% (v / v) of a nonionic surfactant, and 0.1-3% (v / v) of albumin.
[0127] In some embodiments, when the cells are treated with a cell lysis buffer, the conditions include heat treatment at 55-70°C for 5-15 min. In some embodiments, the lysis temperature is 55-60°C, 58-64°C, 62-68°C, or 65-70°C; in some specific embodiments, the lysis temperature is 55°C, 57°C, 60°C, 63°C, 67°C, or 70°C. In some embodiments, the lysis time is 5-9 min, 7-12 min, 9-13 min, or 11-15 min; in some specific embodiments, the lysis time is 5 min, 8 min, 11 min, 14 min, or 15 min.
[0128] In some embodiments, mitochondrial DNA from human peripheral blood cells is amplified using the DNA amplification method of this application; exemplaryly, mitochondrial DNA amplification is performed using 123 pre-amplification primer pairs. During low-temperature denaturing PCR amplification, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 170-250 bases, wherein the pre-amplification primer pairs are a combination including pre-amplification primer pairs 1 to 123 as shown below;
[0129] Among them, the pre-amplification primer pairs are F1 and R1, F2 and R2, F3 and R3, ..., and F... x R x Where x is a natural number from 1 to 123. The sequence and sequence number corresponding to each primer are detailed in Table 1.
[0130] It should be noted that the pre-amplification primer pairs and their combinations used for full-coverage amplification of mitochondrial mtDNA should not be unique and are not limited to the exemplary schemes above. For example, the sequence composition and / or length of one primer, the combination of pre-amplification primer pairs, etc., can be adjusted in conjunction with the length of the amplification product, the base composition, etc. Those skilled in the art should understand that the above primer design schemes are not unique. Depending on the requirements, such as the length of the target sequence to be amplified, the length of the overlapping sequence of the amplification products of any two adjacent pre-amplification primer pairs, and / or the selection of a suitable sequencing platform, the combination of pre-amplification primer pairs can be adjusted. It is understood that when the length of the target sequence to be amplified is longer, the number of pre-amplification primer pairs in the pre-amplification primer pair will decrease; conversely, the number of pre-amplification primer pairs in the pre-amplification primer pair will increase. It is also understood that when the length of the overlapping sequence of the target sequence binding to the forward primer of any one of the multiple sets of preamplification primer pairs and the target sequence binding to the reverse primer of any adjacent set of preamplification primer pairs increases, the length and / or base composition of any primer in the preamplification primer pair will change.
[0131] For the sequence composition of a primer: In some embodiments, the pre-amplified primer pair comprises a combination of pre-amplified primer pairs as described above, and at least one of the primers is replaced with a primer having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with its sequence. In one specific embodiment, the preamplification primer pair comprises a combination of preamplification primer pairs as described above, and at least one of the primers is replaced with a primer having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.99% identity with its sequence.
[0132] Those skilled in the art should understand that, for full-coverage amplification of the same DNA sample, if the length of the amplification product of each pre-amplification primer pair increases, the number of pre-amplification primer pairs required to achieve full-coverage amplification can be appropriately reduced; conversely, if the length of the amplification product of each pre-amplification primer pair decreases, the number of pre-amplification primer pairs required to achieve full-coverage amplification needs to be appropriately increased.
[0133] For example, when the length of the amplification product of any pair of preamplification primers varies from 280 to 400 bases, in some embodiments, the combination of the preamplification primers includes the preamplification primer pairs shown below;
[0134] Preamplification primer pairs F1, R2, F2, R3, F3, R4, ..., F... n-1 R n ; where n is a natural number from 2 to 123.
[0135] Alternatively, in some embodiments, the preamplification primer pair comprises a combination of preamplification primer pairs as described above, and at least one of the primers is replaced with a primer having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with its sequence. In one specific embodiment, the pre-amplification primer pair comprises a combination of pre-amplification primer pairs as described above, and at least one of the primers is replaced with a primer having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.99% sequence identity. In the above scheme, the number of pre-amplification primer pairs for achieving full coverage amplification of the DNA sample is reduced to 122 pairs.
[0136] "Pre-amplification primer pair F" n-1 R n The meaning of "" is: from upstream primer F n-1 and downstream primer R n The primer pair consists of: for example, the pre-amplification primer pair F3, R4 refers to the primer pair consisting of upstream primer F3 and downstream primer R4, wherein the sequence of F3 is shown in SEQ ID NO:5 and the sequence of R4 is shown in SEQ ID NO:8.
[0137] For example, when the length of the amplification product of any pair of preamplification primers varies from 280 to 400 bases, in some embodiments, the combination of the preamplification primers includes the preamplification primer pairs shown below;
[0138] Preamplification primer pairs F1, R2, F3, R4, F5, R6, ..., F... m-1 R m ; where m is an even number from 2 to 123.
[0139] Alternatively, in some embodiments, the preamplification primer pair comprises a combination of preamplification primer pairs as described above, and at least one of the primers is replaced with a primer having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with its sequence. In one specific embodiment, the pre-amplification primer pair comprises a combination of pre-amplification primer pairs as described above, and at least one of the primers is replaced with a primer having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.99% sequence identity. In the above scheme, the number of pre-amplification primer pairs for achieving full coverage amplification of the DNA sample is reduced to approximately 61 pairs.
[0140] Of course, the primer pair designs given above are merely exemplary and not limiting of the scope of protection of this application, based on the method for achieving full-coverage gene amplification of DNA samples in this application; other pre-amplification primer pairs and combinations thereof designed based on the method of this application should also be included within the scope of protection of this application.
[0141] This application provides a rapid DNA library preparation method with full sequence coverage amplification, which includes the following steps:
[0142] The method described above was used to perform full-coverage amplification to obtain PCR products;
[0143] PCR products were treated with single-stranded DNA exonuclease to remove residual primers after PCR.
[0144] Secondary PCR: This involves bringing the exonuclease treatment product into contact with the secondary PCR primer pair to perform PCR and obtain a library.
[0145] In some implementations, the 5' end of the forward primer and the 3' end of the reverse primer in any set of preamplification primer pairs also include a shared sequence.
[0146] In some implementations, the 5' end of the forward primer and the 3' end of the reverse primer in the secondary PCR primer pair further include a shared sequence and a barcode sequence.
[0147] The primers designed as described above are intended to: during low-temperature denaturing PCR, DNA amplification is performed using pre-amplification primer pairs as primers, and all target sequences obtained will include the shared sequence; during further secondary PCR, DNA amplification is performed using secondary PCR primer pairs containing the shared sequence, through complementary base pairing between the shared sequence on the secondary PCR primer pairs and the shared sequence on the target sequences, i.e., DNA amplification is performed using the shared sequence as primers.
[0148] The term "shared sequence" refers to a gene sequence contained in all sequences within a mixed sequence and that is identical or complementary to any two sequences. During library construction, the shared sequence is introduced at the 5' and 3' ends of any primers on multiple pre-amplification primer pairs so that all target sequences after PCR contain the shared sequence at both ends; during subsequent PCR, the shared sequence can be used as a primer pair for mass amplification of the target sequence in a new round of PCR. In some embodiments, the shared sequence is directly designed onto the primers and ligated to the target sequence via complementary base pairing; in other embodiments, the shared sequence is directly ligated to the target sequence using biological and / or chemical methods. In some embodiments, the shared sequence does not have a binding sequence with the genome or mitochondrial genome from which the DNA sample is derived. If the DNA sample is derived from humans, the shared sequence does not have a binding sequence with the human genome. In some embodiments, the annealing temperature of the shared sequence is 70-80°C. In some implementations, the annealing temperature of the shared sequence is any temperature range selected from 70-72°C, 71-74°C, 73-76°C, 75-78°C, 77-79°C, 78-80°C, or 70-80°C; in one specific embodiment, the annealing temperature of the shared sequence is any temperature range selected from 70°C, 71°C, 72°C, 73°C, 75°C, 76°C, 77°C, 78°C, 79°C, 80°C, or 70-80°C. This annealing temperature is to ensure the sequential binding of the shared sequence to the pre-amplified primer pair. In some embodiments, the shared sequence does not end in repeating sequences of at least 3, 5, or 8 consecutive bases; in one specific embodiment, the shared sequence does not end in repeating sequences of 3, 4, 5, 6, 7, or 8 consecutive bases; in some embodiments, the repeating sequences may be, for example, AAA, TTT, GGG, CCC, etc. This design aims to avoid inducing polyaddition reactions. In some embodiments, the shared sequence is 18-25 bases long; in some embodiments, the shared sequence is 18-21, 20-22, 21-24, or 23-25 bases long. In some specific embodiments, the shared sequence is 18, 19, 20, 21, 22, 23, 24, or 25 bases long. It should be noted that a shared sequence that is too short may lead to insufficient specificity, while a shared sequence that is too long may reduce the efficiency of PCR.
[0149] When the DNA sample is mtDNA, the shared sequence is selected from any one of the following sequences: the sequence shown in SEQ ID NO:247 (CAGCGTCCGCTCTTCCGATC), the sequence shown in SEQ ID NO:258 (AAGCAGTGGTATCAACGCAGAGT), the sequence shown in SEQ ID NO:259 (ACACTGACGACATGGTTCTACA), and the sequence shown in SEQ ID NO:260 (AGTCACGACGTTGTAATACG); or, the shared sequence is selected from the sequence shown in SEQ ID NO:247, the sequence shown in SEQ ID NO:258, SEQ ID NO:259, and SEQ ID NO:260. Any sequence shown in NO:260 that has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity. For example, a shared sequence is a sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.99% identity with the sequence shown in SEQ ID NO:247; for example, a shared sequence is a sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.99% identity with the sequence shown in SEQ ID NO:258; a shared sequence is a sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 99.99% identity with the sequence shown in SEQ ID NO:258; The sequence shown in NO:259 has 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.99% identity. Those skilled in the art should understand that the above are specific choices of shared sequences that are exemplary and available, but shared sequences that meet the above requirements can also be used in the construction of the library of this application.
[0150] The barcode sequence is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more consecutive nucleotides. In some embodiments, the barcode sequence contains at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more consecutive nucleotides. In some embodiments, at least a portion of the different barcode sequences in a nucleic acid population containing multiple barcode sequences are different. In some embodiments, at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99% of the barcode sequences linked to different nucleic acids are different. In several such embodiments, all barcode sequences are different. The diversity of different barcode sequences in a nucleic acid population containing barcode sequences can be randomly generated or non-randomly generated. In some embodiments, the library structure includes barcode sequences. In some implementations, the resulting library includes one or more barcode sequences. In some implementations, the multiple barcode sequences in a library are different. In some implementations, barcode sequences are combined for combined labeling. Exemplary base compositions of barcode sequences are shown in Table 2.
[0151] The applicant discovered that after low-temperature denaturing PCR, only an exonuclease treatment step is needed before proceeding to the next step of secondary PCR to obtain a sequenceable mixed library. That is, in this library construction method, there is no PCR product purification step after low-temperature denaturing PCR; this method can achieve subsequent library construction and accurate sequencing. However, those skilled in the art should know that, for the purposes of this application, purifying the product before the secondary PCR is also a feasible approach; this approach is also included within the scope of protection of this application. That is, in some embodiments, conventional PCR product purification is not performed before the secondary PCR; in other embodiments, conventional PCR product purification methods are used before the secondary PCR, and the purified product is then used for the secondary PCR.
[0152] In some embodiments, when treating PCR products with exonucleases, the treatment method includes the following steps: treating at 30-38°C for 5-15 min. In some embodiments, the treatment time for treating PCR products with exonucleases is any time range within the range of 5-8 min, 6-11 min, 9-13 min, 12-15 min, or 5-15 min; in one specific embodiment, the treatment time for treating PCR products with exonucleases is any time range within the range of 5 min, 6 min, 7 min, 8 min, 9 min, 10 min, 11 min, 12 min, 13 min, 14 min, 15 min, or 5-15 min. In some embodiments, the treatment temperature for exonuclease treatment of PCR products is any temperature range within the range of 30-33°C, 32-35°C, 34-38°C, or 30-38°C; in one specific embodiment, the treatment temperature for exonuclease treatment of PCR products is any temperature range within the range of 30°C, 31°C, 32°C, 33°C, 34°C, 35°C, 36°C, 37°C, 38°C, or 30-38°C.
[0153] In some embodiments, the PCR product is treated with a single-stranded DNA exonuclease in the presence of a buffer. It is understood that there are no particular limitations on the buffer in this application, and the choice of buffer includes, but is not limited to, those listed below; exemplary optional buffers may be NEB Exonuclease I Reaction Buffer, Thermoscientific Buffer, or TaKaLa Buffer. The components of NEB Exonuclease I Reaction Buffer may, exemplaryly, be: 67mM Glycine-KOH, 6.7mM MgCl2, 10mM β-ME; the components of Thermoscientific Buffer may, exemplaryly, be: 67mM glycine-KOH, 67mM MgCl2, 10mM DTT; the components of TaKaLa Buffer may, exemplaryly, be: 67mM Glycine-KOH, pH 9.5, 6.7mM MgCl2, 1mM DTT, 0.25mg / ml substrate DNA.
[0154] This application also provides a hybrid library constructed using the method described above. Each library in the hybrid library includes a target sequence and a shared sequence. In some embodiments, each library in the hybrid library further includes a barcode sequence.
[0155] This application also provides a sequencing method for sequencing a DNA sample, wherein the DNA sample comprises at least one sequence that fully covers the amplified DNA fragment;
[0156] The sequencing method for amplifying the DNA fragment with full sequence coverage includes the following steps:
[0157] The above-described rapid DNA library construction method was used to construct a mixed library;
[0158] Library purification;
[0159] Sequencing.
[0160] The specific methods for building the library are as described above and will not be repeated here.
[0161] Using the method of this application, sequencing libraries of different lengths can be obtained by adjusting the length of the target sequence that any pair of pre-amplified primers can amplify during low-temperature denaturing PCR; and sequencing libraries of different lengths can be adapted to different sequencing platforms.
[0162] More specifically, when amplifying a target sequence of 170-250 bases using any one of the pre-amplification primer pairs from multiple sets of pre-amplification primer pairs, the resulting library is suitable for any one of the following sequencing platforms: HiSeq 3000, HiSeq 4000, HiSeq X10, NextSeq 1000, NextSeq 2000, MiniSeq, and NovaSeq 6000 in PE150 sequencing mode;
[0163] When amplifying a target sequence of 280-400 bases using any one of the multiple sets of pre-amplification primer pairs, the resulting library is suitable for any one of the following testing platforms: any one of the following sequencing platforms: HiSeq 3000, HiSeq 4000, HiSeq X10, NextSeq 1000, NextSeq 2000, MiniSeq, and NovaSeq 6000 in PE250 sequencing mode.
[0164] This application creatively develops a novel mtDNA sequencing method utilizing 123 pre-amplified primer pairs. The aim is to provide a rapid, reliable, and cost-effective method for analyzing mtDNA heterozygosity at the level of large genomic DNA and even single cells. This application achieves this goal through ultra-high multiplex PCR technology, with the 123 pre-amplified primer pairs covering the entire mtDNA genome. Furthermore, this method creatively eliminates the magnetic bead purification process during targeted library construction after multiplex PCR, simplifying the library construction process to only three consecutive reagent additions, thus maximizing efficiency while minimizing library construction time and cost. In addition, this method was further applied to evaluate single-cell mtDNA variants from elderly samples, revealing a large number of mtDNA mutations in T cells. Overall, this research provides a rapid and powerful method for mtDNA analysis, offering new insights into the role of mtDNA heterozygosity in the aging process. By facilitating large-scale population studies, this method is expected to advance the understanding of the clinical significance of mtDNA mutations in those skilled in the art and provide reference information for future treatment strategies.
[0165] Currently, large-scale mtDNA sequencing provides a pathway for analyzing mtDNA mutations. However, current mtDNA sequencing technologies, whether using batch genomic DNA templates or single-cell templates, suffer from drawbacks such as high cost, complex experimental protocols, and lengthy experimental cycles. This application, based on the concept of full-coverage sequence amplification, further develops an mtDNA sequencing technology (referred to as 123-seq sequencing technology) specifically for mtDNA library construction. This method is characterized by low cost and short experimental cycle, and is suitable for library construction from mixed cellular genomic DNA or single-cell samples. At the single-cell level, the average library construction cost per cell is approximately $0.13, and the total library construction time is approximately 90 minutes. Furthermore, this study evaluated its detection sensitivity in mixed samples and demonstrated its ability to achieve high-precision single-cell mtDNA sequencing in different cell types, including HEK-293T cells, Quatet M8 cell line, and peripheral blood-derived T cells. The 123-seq technology represents a method that maximizes advantages in terms of time, cost, and library construction efficiency. Therefore, 123-seq technology is expected to become an important research tool for analyzing mtDNA heterogeneity in large population samples in the future.
[0166] Multiplex PCR targeted library preparation represents a strategic approach to library construction, in which multiple PCR primers are introduced to selectively amplify regions of interest or targets, followed by library construction from the amplified DNA fragments for sequencing analysis. This methodology offers unique advantages by efficiently amplifying multiple regions of interest (multiple targets), thereby simplifying sample processing time and reducing costs. However, a common challenge in targeted sequencing strategies is the limited number of targets, typically no more than 60, primarily due to the difficulty of optimizing PCR amplification efficiency and compatibility across different targets in ultramultiplex PCR. In this study, the development of the 123-plex ultramultiplex PCR library construction system represents a significant advancement, providing an important reference for establishing targeted amplification systems for non-mtDNA targets. For example, it allows for targeted amplification without purification after cell lysis; and it simplifies the experimental procedure by removing redundant primers using exonucleases (specifically exonuclease I) between the first round of low-temperature denaturing PCR amplification and the second round of secondary PCR exponential amplification. Furthermore, the introduction of universal primers (e.g., shared sequences) during target amplification solves the compatibility problem between primers and enhances amplification efficiency across multiple targets.
[0167] Furthermore, immune system aging is a physiological decline that occurs post-reproduction and accompanies the aging process. Mitochondrial dysfunction, leading to energy and metabolic imbalances, is widely considered one of the core causes of the gradual decline of the immune system. However, the exact causes of mitochondrial dysfunction during immune system aging remain unclear. Using the 123-seq technology developed in this study, we were able to detect ubiquitous mtDNA mutations in individual T cells of healthy elderly individuals, indicating impaired mtDNA purification and selection function in aged T cells. Therefore, mtDNA mutations in elderly individuals may be one of the causes of weakened T cell function, highlighting the potential importance of intervening in the generation and accumulation of mtDNA mutations for the prevention and treatment of age-related diseases. Example
[0168] Specific embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While specific embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0169] In this application, the DNA sample specifically uses mtDNA as an example to illustrate the scheme of this application.
[0170] 1. Sample study and source
[0171] Peripheral blood samples were collected from a 69-year-old male patient for analysis of single-cell mtDNA mutation profiles. Human samples were provided by the Institute of Human Phenotyping, Fudan University. The study was approved by the Institutional Review Board of Fudan University, and all participants provided written informed consent in accordance with the guidelines of the Declaration of Helsinki.
[0172] The Quartet M8 cell line was generously provided by Professor Shi Leming's laboratory at Fudan University; while the HEK-293T cells were obtained from ATCC.
[0173] 2. Cell Culture
[0174] Quartet cell line M8 cells were cultured in RPMI-1640 medium containing 10% fetal bovine serum and incubated under standard conditions at 5% CO2 and 37°C.
[0175] Similarly, HEK-293T cells were cultured under the same conditions in DMEM medium containing 10% fetal bovine serum. Cells were collected and washed once with 1×PBS before the experimental procedures.
[0176] 3. Flow cytometry sorting of single cells
[0177] Peripheral blood samples were rapidly thawed in a 37°C water bath and then processed using a red blood cell lysis kit (ThermoFisher, cat.no.00-4333-57) according to the manufacturer's instructions, followed by cell counting. Subsequently, one million cells were resuspended in 100 μL of 1×PBS containing 2% FBS and incubated on ice in the dark for 30 min with 5 μL of CD9-PE (ThermoFisher, cat.no.MA5-16861) and 5 μL of CD3-FITC antibody (ThermoFisher, cat.no.MA1-80640). After incubation, 400 μL of 1×PBS containing 2% FBS was added, and the mixture was centrifuged at 300 g for 3 min. The cells were then resuspended in 500 μL of 1×PBS containing 2% FBS and sorted by flow cytometry using a SONY MA900 flow cytometer. A sorting chip with a diameter of 100 μm was used to separate CD3-positive cell populations.
[0178] HEK-293T cells and Quartet M8 cell line were stained with Calcein (ThermoFisher, cat.no.C3099) and sorted according to Calcein-positive cell populations.
[0179] 4. Obtaining mtDNA
[0180] The sorted single cells were transferred to 96-well plates containing 5 μL of lysis buffer (final concentration 0.5% CA630). The plates were then heated to 65°C for 10 min to ensure optimal mtDNA release.
[0181] 5. Library sequencing
[0182] Cluster generation, sequencing, image processing, demultiplexing, and quality fraction calculation were performed on the Illumina NovaSeq 6000 platform. Batch genomic DNA sequencing followed a similar procedure to the steps described above, omitting the cell lysis step. Instead, 10–50 ng of total genomic DNA was used as template for the first round of PCR reactions.
[0183] 6. mtDNA variant identification and data analysis
[0184] Low-quality reads were cut using trim galore(44) and sequencing adapters were removed. The raw data were then aligned with the hg38 composite reference genome using bwa mem and RtN was used. [5] The software package eliminated Numts readings, and the MitoMutCall package was used to analyze mtDNA mutations. Statistical analysis and graphical representation were performed using R (version 4.3.1). All statistical comparisons between two samples were performed unpaired, and two-sided p-values were calculated. Visualization was performed using basic R plotting functions and ggplot2.
[0185] Example 1. Library construction and sequencing
[0186] A rapid DNA library construction method with full sequence coverage amplification, referring to... Figure 2 The specific steps are as follows:
[0187] Low-temperature denaturing PCR: 50 ng of the obtained single-cell genomic DNA (the starting template can be single-cell lysis buffer or pre-purified genomic DNA; the genomic DNA contains mtDNA, generally 0.1 wt% to 0.05 ng of the genomic DNA content) and multiple sets of pre-amplification primer pairs are mixed. The specific mixing system is as follows: 1×Q5 reaction buffer (New England Biolabs, cat.no. M0491), 200 μM 10 mM dNTPs (New England Biolabs, cat.no. N0447L), 0.02 U / μL Q5 high-fidelity DNA polymerase (New England Biolabs, cat.no. M0491), and 10 μM 1st PCR primer mixture (the 1st PCR primer mixture refers to 123 pairs of pre-amplification primers, the specific sequences of which are shown in Table 1; the shared sequence linked to the 5' end of each pre-amplification primer pair is as shown in SEQ ID). As shown in NO:247 (CAGCGTCCGCTCTTCCGATC), add nuclease to remove water to bring the final volume to 20 μL. Perform PCR amplification under the following cycling conditions: 98℃ pre-denaturation for 30 s, followed by 98℃ denaturation for 10 s to increase specificity, annealing at 60℃ for 5 min, then 98℃ denaturation for 10 s for extension, 72℃ extension for 20 s, and finally 72℃ extension for 2 min. The reaction is then maintained at 4℃.
[0188] To remove residual primers after PCR, the PCR products were treated with exonucleases. Specifically, after the first round of PCR, 1.25 μL of thermosensitive exonuclease I (New England Biolabs, cat.no.M0568L) and 2.5 μL of NEB buffer 3.1 (New England Biolabs, cat.no.M0568L) were added to the reaction mixture, and then incubated at 37°C for 10 min.
[0189] Table 1. 123 pairs of pre-amplification primers
[0190]
[0191]
[0192]
[0193]
[0194]
[0195]
[0196] Secondary PCR: This involves bringing the exonuclease treatment product into contact with the secondary PCR primer pair to perform PCR and obtain a library. The specific steps are as follows: Take 12.5 μL of the exonuclease-treated reaction mixture as a template and add it to a new tube containing 25 μL of 2X PhusionBlood PCR buffer (ThermoFisher, cat.no.M0568L), 1 μL of Phusion Blood II DNA polymerase (ThermoFisher, cat.no.M0568L), 2.5 μL of 10 μM index primer Fv and 2.5 μL of 10 μM index primer Rv (see Table 2 for details; index primer Fv refers to index primer Fv-1 to index primer Fv-96, a total of 96 primers; index primer Rv refers to index primer Rv-1 to index primer Rv-96, a total of 96 primers; both index primer Fv and index primer Rv contain barcode sequences and shared sequences, as well as primer sequences compatible with next-generation sequencers). Add nuclease to remove water to bring the final volume to 50 μL. The cycling conditions for the second round of PCR were as follows: pre-denaturation at 98°C for 5 min, followed by 30 cycles, each cycle consisting of denaturation at 98°C for 1 s, annealing at 60°C for 5 s, extension at 72°C for 8 s, and a final extension at 72°C for 2 min. The reaction was then maintained at 4°C. All second-round PCR products were combined and purified using 1.2 volumes of VAHTSDNA Clean Beads (Vazyme, cat.no.N411-01) to obtain a library suitable for sequencing.
[0197] In the second PCR, the upstream primers were index primers Fv1 to Fv96; the primer composition was: second-generation sequencer-compatible primer sequence + barcode sequence + shared sequence. The downstream primers were index primers Rv1 to Rv96; the primer composition was: second-generation sequencer-compatible primer sequence + barcode sequence + shared sequence.
[0198] The general formula for the primer sequences of specific index primers Fv-1 to Fv-96 is as follows:
[0199] AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGACGCTCTTCCGATCTCTG, SEQ ID NO:248; where NNNNNNNN refers to a barcode sequence consisting of 8 bases; the difference between index primers Fv-1 to Fv-96 lies in the different barcode sequences, and the barcode sequences on index primers Fv-1 to Fv-96 are shown in Table 2.
[0200] The general formula for the primer sequences of index primers Rv-1 to Rv-96 is as follows:
[0201] CAAGCAGAAGACGGCATACGAGATNNNNNNNNAGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGAC, SEQ ID NO:249; where NNNNNNNN refers to a barcode sequence consisting of 8 bases; the difference between index primers Rv-1 to Rv-96 lies in the different barcode sequences, and the barcode sequences on index primers Rv-1 to Rv-96 are shown in Table 2.
[0202] Table 2. Barcode sequences in index primers Fv and Rv
[0203]
[0204]
[0205] In the above method, 123 pairs of pre-amplified primers targeting mtDNA are used to selectively amplify the entire mitochondrial genome. Each primer pair contains a shared sequence of 20 nucleotides at its 5' end. The first round of PCR (i.e., low-temperature denaturing PCR) includes a low-temperature, low-cycle phase for targeted amplification of mtDNA to address target specificity issues, and a high-temperature PCR phase (i.e., secondary PCR) for exponential amplification of the shared sequence to address target compatibility issues. In the second step, thermosensitive exonuclease I (Exo I) is added to the reaction system to remove multiple primers remaining from the first round of PCR. In the third step, sequencing-adapted primers containing the shared sequence and cell barcode (sample index) are added to the reaction system, followed by a second round of PCR to complete library construction.
[0206] In computational analysis workflows, in order to obtain high-precision mtDNA mutation information, such as... Figure 1 As shown, this application integrates a method that eliminates low-level sequencing and nuclear-encoded mtDNA homologous sequence (Numts) reads to mitigate the impact of Numts read contamination on mtDNA mutation retrieval. Specifically, the raw library data is filtered to consider data quality, with Q≥30 as the criterion; then, the reads are made consistent with GRCh38 (GRCh38 is a genome sequence published in 2013); then, the RtN algorithm is used to identify and remove Numts; finally, mtDNA mutations are evaluated.
[0207] Table 1. 123 pairs of pre-amplification primers and their sequences
[0208] Example 2. Detection of mtDNA variations at the level of mixed cellular genomic DNA
[0209] To maximize experimental efficiency and ensure high yields of mtDNA library products, this application designed 123 pairs of pre-amplification primers, optimized the biochemical reaction of targeted mtDNA sequencing technology, and eliminated the need for purification after low-temperature denaturing PCR.
[0210] 1) The impact of Exo I processing steps on the library
[0211] In this embodiment, mtDNA was obtained from HEK-293T cells according to the method described above, and the library product was obtained according to the method in Example 1.
[0212] Simultaneously, mtDNA was obtained from HEK-293T cells, and a library was constructed using the mtDNA according to the following method (this method does not involve exonuclease I treatment):
[0213] Low-temperature denaturing PCR: Same as in Example 1.
[0214] Secondary PCR: The low-temperature denaturing PCR product is brought into contact with the secondary PCR primer pair for PCR to obtain a library. The specific steps are as follows: 12.5 μL of the low-temperature denaturing PCR product is used as a template, and the other components in the PCR system are the same as in Example 1; the PCR conditions for secondary PCR are the same as in Example 1; the purification steps are the same as in Example 1 to obtain a library that can be used for sequencing.
[0215] The results showed that, Figure 3 As shown, compared to directly using single-stranded DNA as a template for a secondary PCR reaction without Exo I digestion after low-temperature denaturation PCR (using the same input amount, i.e., 50 ng of 293T cell single-cell genomic DNA), the gel electrophoresis analysis of the final library product obtained by the method in Example 1 showed a brighter band for the target 336 bp fragment, indicating a higher library concentration. Therefore, this application ultimately identifies the Exo I-mediated single-stranded DNA digestion step as an intermediate step between the two PCR reactions.
[0216] 2) Validation of consistency between parallel samples
[0217] like Figure 4 As shown, analysis of the final library bands from three HEK-293T genomic DNA samples (the genomic DNA contained mtDNA; the three HEK-293T genomic DNA samples were from three parallel experiments, with 50 ng of genomic DNA as the template) clearly showed the target 336 bp band in all three replicate samples, while the template-free control group (i.e., the NTC in the attached figure) showed no band. Figure 4 A). In addition, such as Figure 4 B and Figure 6 As shown, after removing Numts readings, the data were aligned to the rCRS human mitochondrial genome Cambridge reference sequence at a rate of 100%. Correlation analysis of mtDNA mutations identified from three HEK-293T genomic DNA samples is also included. Figure 4 As shown in C, it exhibits a high correlation.
[0218] 3) Sensitivity and specificity of mtDNA heterogeneity and its fraction.
[0219] To evaluate the sensitivity and specificity of 123-seq technology in detecting mtDNA heterogeneity and its fraction in total genomic DNA, the technology was applied to a series of sample mixtures composed of total genomic DNA (including nuclear genomic DNA, mitochondrial DNA, etc.) from GM19223 and GM12878 cell lines in different proportions. The mtDNA sequences of these cell samples differed at 19 single nucleotide sites; these 19 single nucleotide sites are as follows: Figure 5 As shown in B, the values are: 182C>T, 185G>T, 189A>G, 195T>C, 247G>A, 357A>G, 1018G>A, 1738T>C, 2352T>C, 2706A>G, 2758G>A, 2768A>G, 3308T>C, 3594C>T, 3666G>A, 3693G>A, 5036A>G, 5046G>A, 5393T>C. The results show that... Figure 5 A and Figure 5 As shown in B, heterogeneity at all 19 polymorphic sites changed with the proportion of DNA used to generate these sample mixtures, consistent with this (Pearson correlation coefficient = 0.98, P < 0.0001).
[0220] Example 3. Detection of mtDNA variations at the single-cell genomic DNA level
[0221] To apply the library construction method of this application to the single-cell level, this embodiment exemplarily utilizes 123-seq technology and designs a set of lysis buffers. These buffers are theoretically capable of disrupting the mitochondrial membrane to release mtDNA while maintaining compatibility with PCR enzymes. More specifically, referring to the method in "4. Acquisition of mtDNA," mtDNA is obtained using lysis buffers with different components and ratios. Specifically, the sorted single cells are transferred to 96-well plates containing 5 μL of lysis buffer (specifically, lysis buffer 1, lysis buffer 2, lysis buffer 3, lysis buffer 4, and lysis buffer 5). The plates are then heated to 65°C for 10 min to ensure optimal mtDNA release. Specifically, the components and volumes of lysis buffer 1 are as follows: final concentration 0.5% (v / v) IGEPAL CA-630; the components and volumes of lysis buffer 2 are as follows: final concentration 10mM Tris-HCl pH 7.4, 10mM NaCl, 3mM MgCl2, 0.1% (v / v) IGEPAL CA-630, 1% (v / v) BSA; the components and volumes of lysis buffer 3 are as follows: final concentration 110mM Tris-HCl pH 7.4, 0.25% (v / v) IGEPAL CA-630, 150mM NaCl; the components and volumes of lysis buffer 4 are as follows: final concentration 1mM DTT, 10mM Tris-HCl pH 7.4, 10mM NaCl, 3mM MgCl2, 0.1% (v / v) IGEPAL CA-630. CA-630, 1% (v / v) BSA; the components and amounts of lysis buffer 5 are: final concentration of 0.2% (v / v) Triton X-100.
[0222] 1) Screening of lysis buffer
[0223] In this embodiment, qPCR experiments targeting three different mtDNA sites were performed at the single-cell level using different solubilization buffers to determine which solubilization buffer can maximally solubilize cells and maximize mtDNA amplification compatibility.
[0224] The specific experimental procedure was as follows: mtDNA from the same single cell was selected, and qPCR was performed on three sites in the mitochondrial genome: positions 8-199, 5258-5444, and 10626-10811. The amplification primers for the three sites were:
[0225] Positions 8-199 of the mitochondrial genome (mitochondrial segment I):
[0226] Mitochondrial fragment FI: GGTCTATCACCCTATTAACCAC, SEQ ID NO:250;
[0227] Mitochondrial fragment RI: AGTAAGTATGTTCGCCTGTAAT, SEQ ID NO:251;
[0228] Mitochondrial genome positions 5258-5444 (mitochondrial segment II):
[0229] Mitochondrial fragment FII: ATGGGCCATTATCGAAGAATT, SEQ ID NO: 252;
[0230] Mitochondrial fragment RII: GAATGGGGTGGGTTTTGTAT, SEQ ID NO: 253;
[0231] Positions 10626-10811 of the mitochondrial genome (mitochondrial segment III):
[0232] Mitochondrial fragment FIII: CCCTCTTAGCCAATATTGTG, SEQ ID NO:254;
[0233] Mitochondrial fragment RIII: AAAGTCATGTCAGTGGTAGTAA, SEQ ID NO:255.
[0234] The result is as follows Figure 7 As shown in Figure A, lysis buffer 1 (final concentration: 0.5% v / v IGEPAL CA-630; IGEPAL CA-630 is a nonionic, non-denaturing detergent, CAS number 9002-93-1) exhibited the lowest Ct value.
[0235] 2) Verification of compatibility between lysis buffer 1 and the mtDNA amplification system
[0236] After lysing single cells with lysis buffer 1, lysis products were obtained, and the mitochondrial ND5 fragment was amplified at the single-cell level.
[0237] The specific method is as follows: Single 293T cells were placed in 96-well plates for analysis using flow cytometry. 5 μL of 0.5% v / v IGEPAL CA-630 lysis buffer was added to each well of the 96-well plate before lysis at 65°C for 10 min. The mitochondrial ND5 fragment was amplified using the Hot Start High-Fidelity 2X Master Mix amplification enzyme system with F primer (ACTTATTACTCTCATCGCTACC, SEQ ID NO: 256) at a final concentration of 0.2 μM and R primer (TGTGATGCTAGGGTAGAATCCG, SEQ ID NO: 257) at a final concentration of 0.2 μM. The amplified mitochondrial ND5 fragment was analyzed by gel electrophoresis, and the results are as follows: Figure 7 As shown in B, parallel experiments with eight single cells all demonstrated successful amplification, confirming its compatibility with the mtDNA amplification system.
[0238] 3) Validation of sequencing depth of the sequencing method in Example 1
[0239] HEK-293T cells were subjected to 123-seq using lysis buffer 1. The specific experimental method was as follows: library construction was performed according to the method in Example 1; the three HEK-293T cells were lysed with lysis buffer 1 according to the method described in "4. mtDNA Acquisition," and the mtDNA genome of each of the three HEK-293T cells was ultimately detected intact; the results are as follows. Figure 8 As shown in Figure A, the entire mtDNA genome was detected completely from all three individual cells, with an average sequencing coverage depth of 43,056-fold. Furthermore, a high correlation was observed between the mtDNA samples detected in the three individual cells, highlighting the ability of the sequencing method exemplified in this embodiment to detect mtDNA mutations at the single-cell level.
[0240] 4) Technical variability of the library construction base sequencing method in Example 1
[0241] Considering the diversity of mtDNA molecules within a single cell and the dominance of heterogeneous mtDNA mutations, a validation experiment was conducted in this embodiment to evaluate the technical variability of the library construct sequencing method (also known as 123-seq technology) of Example 1.
[0242] The specific experimental method is as follows: Quaret M8 cell line was used to obtain single cells according to the method described in "3. Flow Cytometry Sorting of Single Cells". Then, following the method described in "4. mtDNA Acquisition", cell lysis buffer 1 was used to obtain cell lysis buffer (single cells were placed in 5 μL of cell lysis buffer, lysed, mixed thoroughly, and divided into two 2.5 μL portions). Each cell lysis buffer was evenly divided into two aliquots, and library construction (also known as 123-seq technology) was performed on each sample according to the method described in Example 1. Related studies have shown that mutations with high VAF are unlikely to be attributed to PCR errors. Therefore, to strictly control false-positive mutation detection, this experiment focused on single-cell mtDNA mutations with a VAF exceeding 20%. The experimental results were analyzed using a two-tailed t-test (P = 0.2374), and the results are as follows... Figure 8 As shown in B, no significant difference was found in the mean VAF between two samples of the same cells; this result confirms that the 123-seq technology can achieve high-precision detection of mtDNA mutations at the single-cell level.
[0243] Application examples
[0244] Application Example 1. Analysis of mtDNA mutation distribution in T cells of elderly individuals
[0245] Immunoresenescence is considered a major driver of systemic aging, and mitochondrial dysfunction is one of the prominent features of age-related T cell senescence. However, the exact mechanisms underlying mitochondrial dysfunction in these T cells remain unclear. To investigate mtDNA mutations in T cells of older individuals, this application example, using Example 1 as an example, illustrates the application and effectiveness of the library construction and sequencing methods based on this application in the analysis of mtDNA mutation distribution in T cells of older individuals.
[0246] The mitochondrial genome of single T lymphocytes from peripheral blood obtained from a healthy 90-year-old male donor was analyzed using the method described in Example 1 (i.e., 123-seq technique). CD3-positive T cells were sorted using the method described in "3. Flow Cytometry Sorting of Single Cells," specifically as follows... Figure 10 As shown: Peripheral blood cells were first incubated with CD3 and CD19 antibodies, then analyzed by flow cytometry using FSC-A and SSC-A methods to remove dead cells and cell debris. The master cell population was then plotted using FSC-A and FSC-H methods, and CD3-positive cells were then gated to generate 286 single-cell mtDNA libraries.
[0247] Subsequently, mtDNA mutation analysis was performed on sequencing sites with a depth exceeding 500-fold, and 2687 mutations with a VAF (mutation frequency) exceeding 20% were identified at 111 mtDNA sites in T lymphocytes. Specific results can be found in [link to results]. Figure 11 ;like Figure 12 As shown, further analysis revealed that 83.78% of the mutations were located in the coding region of mtDNA. In the cells studied, reference... Figure 13 A, 97.9% carry at least one mutation, and the VAF distribution is biased towards the high and low ends of the VAF spectrum; specifically, refer to Figure 13 B, 71.6% of the mutated VAFs were distributed in the range of 0.9–1, and 13.62% were distributed in the range of 0.2–0.3. In summary, these results indicate that widespread mtDNA mutations exist in T cells of older individuals, and these mutations may affect mtDNA encoding, thereby impacting mitochondrial function.
Claims
1. A method for full-coverage amplification of a DNA sample, comprising the following steps: Low-temperature denaturing PCR: DNA samples are contacted with multiple sets of pre-amplification primer pairs for PCR amplification. The amplification product of any one of the multiple sets of pre-amplification primer pairs overlaps with the amplification product of the pre-amplification primer pair adjacent to it in the 5' direction by at least 20 bases, at least 25 bases, at least 30 bases, at least 35 bases, at least 40 bases, at least 45 bases, at least 50 bases, at least 55 bases, or at least 60 bases. Preferably, during low-temperature denaturing PCR amplification, the amplification product of any one of the multiple sets of pre-amplification primer pairs and the amplification product of the pre-amplification primer pair adjacent to it in the 5' direction have an overlap of 20-70 bases, 25-65 bases, 30-60 bases, 35-55 bases, or 40-50 bases.
2. The method for full-coverage amplification according to claim 1, wherein the target sequence length amplified by any one of the multiple sets of pre-amplification primer pairs is at most 450 bases, at most 400 bases, at most 350 bases, or at most 300 bases; Preferably, the target sequence length amplified by any one of the multiple sets of pre-amplification primer pairs is 350-400 bases, 280-400 bases, 170-250 bases, 170-350 bases, or 100-170 bases.
3. The method for full-coverage amplification according to claim 1 or 2, wherein low-temperature denaturing PCR includes the following steps: Denaturation for 20-60 seconds, annealing for 0.2-8 minutes, extension at the denaturation temperature for 5-15 seconds, followed by extension; Preferably, the denaturation process specifically includes the following steps: pre-denaturation for 15-45 seconds, and denaturation for 5-15 seconds.
4. In the method for full-coverage amplification according to any one of claims 1-3, the annealing time for low-temperature denaturing PCR is 0.2-6 min; Preferably, the annealing time is 2-5 minutes.
5. The method for full-coverage amplification according to any one of claims 1-4, wherein the DNA sample is derived from cells, and the cell source includes any one or more of animals, plants, acellular organisms, prokaryotes, and fungi; More preferably, the cells are derived from human cells, and the sources of the human cells include any one or more of blood cells, immune cells, liver cells, tumor cells, stem cells, nerve cells, muscle cells, and skin cells; Preferably, the method for obtaining the cells includes any one or a combination of the following methods: flow cytometry cell sorting, cell suspension dilution, mechanical separation, micromanipulation, tablet pressing, microfluidics, and laser cutting.
6. The method for full-coverage amplification according to any one of claims 1-5, wherein the DNA sample is selected from any one or more of nuclear DNA, mitochondrial DNA (mtDNA), chloroplast DNA, cytoplasmic DNA, body fluid DNA, blood DNA, plasmid DNA, and DNA obtained by reverse transcription; Preferably, the DNA sample is derived from mitochondrial DNA (mtDNA) from a single cell or multiple cells.
7. The method for full-coverage amplification according to claim 6, wherein the method for obtaining mitochondrial DNA includes the following steps: Cells were treated with a cell lysis buffer to release mitochondrial DNA; Preferably, the cell lysis buffer comprises a nonionic surfactant; More preferably, the components of the cell lysis buffer include any one or more of a buffer, an inorganic salt, albumin or a substitute thereof, and a reducing agent; More preferably, the volume percentage (v / v) of the nonionic surfactant in the cell lysis buffer is 0.05-0.3%, 0.05-0.15%, 0.1-0.2%, 0.1-0.15%, 0.15-0.2%, 0.15-0.25%, or 0.2-0.3%; and / or, The reducing agent in the cell lysis buffer is 0.2-1.7 mM, 0.2-1.4 mM, 0.2-1 mM, 0.5-1.5 mM, 0.8-1.2 mM, 0.2-0.7 mM, or 1.2-1.7 mM; and / or, The inorganic salt content in the cell lysis buffer is 5-20 mM, 5-15 mM, 8-10 mM, 10-18 mM, 10-15 mM, 12-18 mM, or 12-20 mM; and / or, The volume percentage (v / v) of albumin or its substitute in the cell lysis buffer is 0.1-3%, 0.1-1%, 0.1-2%, 0.5-1%, 0.5-2%, 0.5-3%, or 0.8-1.5%.
8. The method for full-coverage amplification according to claim 7, wherein when the cells are treated with cell lysis buffer, the conditions include: Heat-treat at 55-70℃ for 5-15 minutes; Preferably, the heat treatment is carried out at 60-68℃ for 8-12 minutes.
9. The method for full-coverage amplification according to any one of claims 1-8, wherein the DNA sample is mitochondrial DNA; During low-temperature denaturing PCR amplification, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 170-250 bases. The pre-amplification primer pairs include pre-amplification primer pairs 1 to 123 as shown below: Preamplification primer pairs F1, R1, F2, R2, F3, R3, ..., Fx, Rx; among which, x is a natural number from 1 to 123; Alternatively, the preamplification primer pair may comprise a combination of preamplification primer pairs as described above, wherein at least one primer is replaced with a primer having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with its sequence.
10. The method for full-coverage amplification according to any one of claims 1-8, wherein the DNA sample is mitochondrial DNA; During low-temperature denaturing PCR amplification, any one of the multiple sets of pre-amplification primer pairs amplifies a target sequence of 280-400 bases. The pre-amplification primer pairs include those shown below. Preamplification primer pairs F1, R2, F2, R3, F3, R4, ..., F... n-1 R n ;in, n is a natural number from 2 to 123; Alternatively, the preamplification primer pair may comprise a combination of preamplification primer pairs as described above, wherein at least one primer is replaced with a primer having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity with its sequence.