Human mitochondrial gene mononucleotide mutation detection method
By combining UMI-labeled PCR and nanopore sequencing technologies, the problems of PCR duplication and nuclear genome interference in NGS sequencing have been solved, thereby improving the accuracy and sensitivity of human mitochondrial DNA single nucleotide mutation detection.
Patent Information
- Application Number
- CN202511023913.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-07
AI Technical Summary
When detecting single nucleotide mutations in human mitochondrial DNA, NGS sequencing technology suffers from distortion of mutation frequency due to PCR duplication and the unavoidable influence of nuclear genome sequence, affecting the accuracy of the test results.
mtDNA was amplified and purified using a UMI-labeled PCR reaction system to construct an Oxford Nanopore library, which was then analyzed using nanopore sequencing technology. PCR duplication and nuclear DNA interference were removed by combining UMI repeat count and nucleotide similarity processing.
It improves the accuracy and sensitivity of mitochondrial DNA single nucleotide mutation detection, avoids PCR duplication and nuclear DNA interference in sequencing data, and the results are closer to the real situation.
Smart Images

Figure BDA0005515375810000081
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of biomedical engineering, and relates to a human mitochondrial genome detection technology, in particular to a human mitochondrial gene (MtDNA) single nucleotide mutation detection method. BACKGROUND
[0002] Mitochondria is one of the most important organelles in cells, which is usually involved in various metabolic pathways such as oxidative phosphorylation and fatty acid oxidation, and the mitochondrial genome is independent of the nuclear chromosome. Each mitochondrial molecule is a double-stranded covalently closed circular DNA molecule, and the outer ring of the circular DNA molecule is the heavy chain (H chain) and the inner ring is the light chain (L chain). The sequence of human mitochondrial DNA (MtDNA) consists of 16569 base pairs, which contains 37 genes.
[0003] MtDNA mutation can cause mitochondria to fail to perform normal functions, and further cause the occurrence of various diseases. Generally, MtDNA mutations include MtDNA rearrangement (such as large fragment deletion, etc.) and protein synthesis gene mutation (single nucleotide mutation, SNV). Among them, there are various detection methods for mitochondrial single nucleotide mutation, and with the development of sequencing technology, Next Generation Sequencing (NGS) technology has become the mainstream method for detecting MTDNA single nucleotide mutation.
[0004] However, the NGS technology for detecting MTDNA single nucleotide mutation has the following technical limitations:
[0005] 1. The NGS library construction process will introduce PCR duplication, which will cause the mutation frequency to be distorted and cannot approach the true situation;
[0006] 2. MtDNA and nuclear genome sequences have similarity, which makes it difficult for NGS sequencing technology to avoid detecting nuclear genome sequences, which may have some impact on the detection results. SUMMARY
[0007] In order to solve the technical problems of mutation proportion distortion caused by the introduction of PCR duplication in the process of sequencing human mitochondrial DNA by NGS sequencing technology, and the influence on the judgment results caused by the short sequencing fragments of NGS sequencing and the high similarity of nuclear genome fragments to MTDNA sequence during sequencing, the present application discloses a human mitochondrial gene (MtDNA) single nucleotide mutation detection method, which comprises:
[0008] S1, extracting whole genome DNA from the sample to be detected, wherein the whole genome DNA includes mtDNA and nuclear DNA;
[0009] S2, labeling the mtDNA with the UMI to obtain UMI-mtDNA;
[0010] S3, performing targeted amplification and purification on the UMI-mtDNA to obtain a reaction product, and constructing an mtDNA library by the reaction product using a library construction technology of Oxford Nanopore;
[0011] S4, performing sequencing on the mtDNA library using a nanopore sequencing technology to obtain Nanopore sequencing data, and analyzing mtDNA single nucleotide mutation results from the Nanopore sequencing data.
[0012] Further, in step S2, the mtDNA is labeled with the UMI to obtain UMI-mtDNA, which comprises:
[0013] S21, designing a UMI-labeled PCR reaction system, wherein the UMI-labeled PCR reaction system comprises 2xPlatinum TM SuperFi TM PCR Master Mix, a DNA template, and a UMI primer;
[0014] S22, performing amplification and magnetic bead purification on the whole genome DNA using the UMI-labeled PCR reaction system to obtain UMI-mtDNA.
[0015] Further, in step S21, the reaction procedure for amplifying the whole genome DNA using the UMI-labeled PCR reaction system is 98℃-1min, 70℃-65℃ (gradient reduction) 5s, and 72℃-10min.
[0016] Further, in step S3, the UMI-mtDNA is subjected to targeted amplification and purification to obtain a reaction product, which comprises:
[0017] The UMI-mtDNA is subjected to amplification and purification using a PCR reaction system, wherein the PCR reaction system comprises PrimeSTAR GXL DNA Polymerase, 5xPrimeSTAR GXL Buffer, dNTP Mixture, a DNA template, an upstream primer, a downstream primer, and nuclease-free water, and the reaction procedure is 95℃-1min, 98℃-10s, 64-70℃-14min, 30 cycles, and 68℃-5min.
[0018] Further, in step S4, the Nanopore sequencing data is analyzed to obtain mtDNA single nucleotide mutation results, which comprises:
[0019] S41, sequentially performing the preprocessing operations of removing adapter sequences, removing fixed sequences and extracting sequences containing UMI on the Nanopore sequencing data to obtain sequences to be analyzed;
[0020] S42, classifying the sequences to be analyzed according to the number of UMI repeats to obtain UMI single-copy sequences and UMI multi-copy sequences;
[0021] S43, clustering all UMI multi-copy sequences according to UMI to form UMI clusters;
[0022] S44, calculating the average nucleotide similarity of each sequence in the UMI cluster with the rest of the sequences, and taking the sequence with the highest average nucleotide similarity as the representative sequence of the UMI cluster;
[0023] S45, aligning the UMI single-copy sequences and all the representative sequences to the reference genome, and analyzing the alignment results using the clair3 tool or the bamreadcount tool to obtain mtDNA single nucleotide mutation results.
[0024] Further, in step S41, the preprocessing operations of removing adapter sequences, removing fixed sequences and extracting sequences containing UMI are sequentially performed on the Nanopore sequencing data to obtain sequences to be analyzed, comprising:
[0025] S411, using the Porechop tool to remove the adapter sequences in the Nanopore sequencing data;
[0026] S412, using the cutadapt tool to remove the fixed sequences from the 5' end and the 3' end of the Nanopore sequencing data, respectively;
[0027] S413, using the umi-tools tool to search for sequences containing UMI from the 5' end and the 3' end of the Nanopore sequencing data, respectively, using a Python script to find sequences containing specific UMI from the sequences containing UMI to obtain sequences to be analyzed.
[0028] Further, in step S413, the sequence containing specific UMI is NNNNTGNNNN.
[0029] The human mitochondrial gene single nucleotide mutation detection method of the present application can avoid obtaining Nanopore sequencing data containing sequences other than mtDNA, such as nuclear DNA, through UMI library construction and Nanopore sequencing technology, so that the mtDNA single nucleotide mutation results obtained after analyzing the Nanopore sequencing data will not be distorted, and will be closer to the true situation. DETAILED DESCRIPTION
[0030] The embodiments of the present application will be described in detail below.
[0031] The above examples are merely used to illustrate the embodiments of the present application, and the skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in the present specification. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. The present application can also be implemented or applied through other different specific embodiments, and the details in the present specification can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features of the embodiments can be combined with each other without conflict. Based on the embodiments in the present application, all other embodiments obtained by the skilled in the art without creative labor are within the scope of protection of the present application.
[0032] The following explains the abbreviations and key terms mentioned in the present application:
[0033] Nanopore technology: Nanopore sequencing technology, also a kind of third-generation sequencing technology, can pass through the nanometer hole by the motor protein traction of the single-stranded DNA which has been dissociated, and the current on the resistance film will change due to the different bases passing through, so as to identify different bases. It has the characteristics of longer sequencing read length (the longest read length can reach 2 Mb) and slightly lower quality value of sequencing data compared with the second-generation sequencing.
[0034] Next-generation sequencing technology: Second-generation sequencing technology, abbreviated as NGS, which is a kind of sequencing-by-synthesis scheme. In the process of sequencing, the bases labeled by fluorescence are added, and the bases are determined according to the different fluorescence emitted by the bases. It has the characteristics of high throughput and short sequencing fragments (generally not more than 500 bp).
[0035] UMI (Unique Molecular Identifier) technology: A technology for improving the accuracy and sensitivity of high-throughput sequencing, by adding a short, unique sequence tag to each DNA molecule, allowing researchers to distinguish between true sequencing signals and PCR amplification bias in subsequent analysis.
[0036] Nucleotide similarity: Nucleotide similarity refers to a measure for evaluating the similarity between DNA or RNA sequences when comparing them.
[0037] PCR duplication: refers to the repeated sequences generated by PCR amplification in the library construction process of the second-generation sequencing. These repeated sequences appear as completely identical reads in the sequencing data, which not only do not provide new information, but also may affect downstream analysis such as variant detection, genome assembly, etc. Therefore, in the process of data analysis, it is often necessary to remove these PCR duplicates to improve data quality and analysis accuracy.
[0038] The embodiment of the present application discloses a human mitochondrial gene (MtDNA) single nucleotide mutation detection method, the method comprises:
[0039] S1, extracting whole genome DNA from the sample to be tested, wherein the whole genome DNA comprises mtDNA and nuclear DNA;
[0040] S2, marking UMI on the mtDNA to obtain UMI-mtDNA;
[0041] S3, performing targeted amplification and purification on the UMI-mtDNA to obtain a reaction product, and constructing an mtDNA library by the reaction product using the library construction technology of Oxford Nanopore;
[0042] S4, sequencing the mtDNA library by using the nanopore sequencing technology to obtain Nanopore sequencing data, and analyzing the mtDNA single nucleotide mutation result of the Nanopore sequencing data.
[0043] Further, in step S2, marking UMI on the mtDNA to obtain UMI-mtDNA comprises:
[0044] S21, designing a UMI-labeled PCR reaction system, wherein the UMI-labeled PCR reaction system comprises 2xPlatinum TM SuperFi TM PCR Master Mix, DNA template and UMI primer;
[0045] S22, amplifying and magnetically purifying the whole genome DNA by using the UMI-labeled PCR reaction system to obtain UMI-mtDNA.
[0046] Further, in step S21, the reaction program for amplifying the whole genome DNA by using the UMI-labeled PCR reaction system is: 98℃-1min, 70℃-65℃ (gradient reduction) 5s, 72℃-10min;
[0047] Further, in step S3, performing targeted amplification and purification on the UMI-mtDNA to obtain a reaction product comprises:
[0048] The UMI-mtDNA is amplified and purified by a PCR reaction system, which comprises PrimeSTAR GXL DNA Polymerase, 5x PrimeSTAR GXL Buffer, dNTP Mixture, DNA template, upstream primer, downstream primer and nuclease-free water, and the reaction procedure is 95℃-1min, 98℃-10s, 64-70℃-14min, 30 cycles, and 68℃-5min.
[0049] Further, in step S4, the Nanopore sequencing data is analyzed to obtain mtDNA single nucleotide mutation results, which comprises:
[0050] S41, the Nanopore sequencing data is sequentially subjected to the pretreatment operations of removing adapter sequences, removing fixed sequences and extracting UMI-containing sequences to obtain sequences to be analyzed;
[0051] S42, the sequences to be analyzed are classified according to the number of UMI repeats to obtain UMI single-copy sequences and UMI multi-copy sequences, specifically, a sequence in which an UMI appears only once is considered as a single-copy UMI sequence, and there is no PCR duplication in this sequencing, and the sequence is directly output. A sequence in which an UMI appears twice or more is considered as a multi-copy UMI sequence, and the sequence contains PCR duplication, and the UMI is clustered as a class;
[0052] S43, all UMI multi-copy sequences are clustered to form UMI clusters, and in the clustering process, the isONcorrect tool can also be used to correct each class of UMI sequences;
[0053] S44, the average nucleotide similarity of each sequence in the UMI cluster to the remaining sequences is calculated, and the sequence with the highest average nucleotide similarity is taken as the representative sequence of the UMI cluster;
[0054] S45, the UMI single-copy sequences and all the representative sequences are aligned with a reference genome, and the clair3 tool or bamreadcount tool is used to analyze the alignment results to obtain mtDNA single nucleotide mutation results and output a VCF file.
[0055] Further, in step S41, the Nanopore sequencing data is sequentially subjected to the pretreatment operations of removing adapter sequences, removing fixed sequences and extracting UMI-containing sequences to obtain sequences to be analyzed, which comprises:
[0056] S411, removing the adapter sequence in the Nanopore sequencing data by using the Porechop tool;
[0057] S412, removing the fixed sequence from the 5' end and the 3' end of the Nanopore sequencing data respectively by using the cutadapt tool;
[0058] S413, searching for the sequence containing UMI from the 5' end and the 3' end of the Nanopore sequencing data respectively by using the umi-tools tool, and finding the sequence containing specific UMI from the sequence containing UMI by using the Python script to obtain the sequence to be analyzed.
[0059] Further, in step S413, the sequence containing specific UMI is the sequence containing NNNNTGNNNN.
[0060] The above method of the application will be described in detail through specific examples as follows:
[0061] I. Extracting the whole genome DNA of the sample to be tested
[0062] The sample to be tested (the blood, hair, abortion product, and blastocyst cells of non-embryonic tissue of the donor) is subjected to the extraction of 0.2 ml of the DNA of the peripheral blood sample or 1*10^6 cells of the DNA according to the blood / cell genomic DNA extraction kit, and the concentration of the extracted DNA sample is determined by using Qubit, wherein the obtained DNA sample contains mtDNA and nuclear DNA. In this example, 3 positive patients are selected as the detection objects. The pathogenic sites of the mtDNA single nucleotide mutation of the 3 positive patients are as follows: the F1 sample: m.3250T>C (24.4%); the F2 sample: m.3243A>G (78.14%); and the F3 sample: m.9176T>C (<1%).
[0063] II. UMI labeling and purification of mtDNA to obtain UMI-mtDNA
[0064] According to the designed sequence structure, that is, the reads from 5' to 3' end are added with the UMI sequence at the 5' end, and the reads from 3' to 5' end are added with the UMI sequence at the 3' end, the mtDNA in the extracted DNA sample is subjected to UMI labeling. The 2*Platinum TM SuperFi TM PCR Master Mix is used for labeling the mtDNA reaction, and the PCR reaction system is 25 μl, including 2*Platinum TM SuperFi TMPCR Master Mix, 10 μl of DNA template (DNA content is 50-60 ng) and UMI primer (10 μM) 2.5 μl. The reaction conditions are 98℃-1 min, 70℃-65℃ (gradient reduction) 5 s, 72℃-10 min. After labeling, 0.8X magnetic beads are used for purification, and the purified product is eluted with 10 μl of nuclease-free water.
[0065] III. Targeted amplification and purification of UMI-mtDNA to obtain reaction products
[0066] Targeted amplification of UMI-mtDNA is performed using PrimeSTAR GXL DNA Polymerase (1.25 U), and the PCR reaction system is 50 μl, including PrimeSTAR GXL DNA Polymerase (1.25 U) 1 μl, 5x PrimeSTAR GXL Buffer 10 μl, dNTP Mixture 4 μl, DNA template 10 μl (DNA content is about 100 ng) and upstream primer and downstream primer (10 μM) each 1 μl, and nuclease-free water to 50 μl. The reaction conditions are 95℃-1 min, 98℃-10 s, 64-70℃-14 min, 30 cycles, 68℃-5 min. After amplification, agarose gel electrophoresis is used to determine the fragment size, and after confirmation, PrimeSTAR GXL is used for amplification again for 15-20 cycles to obtain a large amount of target fragment for sequencing. After the second amplification, 0.9X magnetic beads are used for purification, the concentration of the purified product of each pair of primers is determined, and the DNA samples of equal amount are mixed. The sample amount is not less than 1000 ng for three-generation library construction.
[0067] IV. mtDNA library construction
[0068] mtDNA library construction is performed using Ligation Sequencing Kit (SQK-LSK110) of Oxford Nanopore Technologies, and the specific operation is performed according to the operation procedure. The main process includes DNA end repair reaction, Barcode ligation reaction, and Adapter ligation reaction.
[0069] V. Sequencing and analysis of mtDNA library by nanopore sequencing technology
[0070] The sequencing operation is performed according to the Oxford Nanopore (nanopore sequencing technology) sequencer instruction manual. The Oxford Nanopore MinION Mk1C single molecule real-time sequencing system is used for sequencing to obtain Nanopore sequencing data, and the method of steps S41-S45 is used for mtDNA single nucleotide mutation analysis, wherein the data in the Nanopore sequencing data of the three samples is shown in Table 1.
[0071] Table 1: Sequencing data distribution
[0072] Sample MT bases Total bases MT ratio F1 3674807 4387272 0.837606 F2 2683998 3266780 0.821604 F3 4066930 4742539 0.857543
[0073] As shown in Table 1, more than 80% of the sequencing bases of the three samples obtained by the method of the present application are on the mitochondrial genome, and only a small part of the sequencing results is the nuclear genome. It can be proved that the method of the present application can well limit the sequencing region, thereby improving the accuracy and sensitivity of sequencing, and reducing the sequencing depth.
[0074] At the same time, since the MT DNA is maternally inherited, by using the library building sequencing and analysis method of the present application, 3 positive samples with known pathogenic sites are tested, and the results are shown in Table 2.
[0075] Table 2: Positive sample test results
[0076]
[0077]
[0078] As shown in Table 2, the method of the present application can have good detection performance for true positive sites, and the results are not different from those obtained by the NGS method, but the mutation frequency of the NGS method and the method of the present application is quite different. The main reason is that the method of the present application can remove the defect of PCR duplication that cannot be removed by the NGS method, and avoid distortion of mutation frequency.
[0079] The human mitochondrial gene single nucleotide mutation detection method of the present application can avoid obtaining Nanopore sequencing data containing sequences other than mtDNA, such as containing nuclear DNA, so that the mtDNA single nucleotide mutation result obtained after analyzing the Nanopore sequencing data will not be distorted, and is closer to the true situation.
[0080] Obviously, those skilled in the art should understand that the above description is only the preferred embodiment of the present application, and is not used to limit the present application, and the present application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for detecting single nucleotide mutations in human mitochondrial genes, characterized in that, The method comprises the following steps: extracting whole genome DNA from a sample to be tested, wherein the whole genome DNA comprises mtDNA and nuclear DNA; labeling the mtDNA with a UMI to obtain UMI-mtDNA; performing targeted amplification and purification on the UMI-mtDNA to obtain a reaction product, and constructing an mtDNA library by using a library construction technology of Oxford Nanopore through the reaction product; sequencing the mtDNA library by using a nanopore sequencing technology to obtain Nanopore sequencing data, and analyzing the Nanopore sequencing data to obtain an mtDNA single nucleotide mutation result.
2. The method of claim 1, wherein the method is for detecting a single nucleotide mutation in a human mitochondrial gene. labeling the mtDNA with a UMI to obtain UMI-mtDNA, comprising: A UMI-labeled PCR reaction system was designed, which includes 2x Platinum TM SuperFi TM PCR Master Mix, DNA template, and UMI primer; amplifying the whole genome DNA by using a UMI-labeled PCR reaction system and purifying the whole genome DNA by using a magnetic bead to obtain UMI-mtDNA.
3. The method for detecting single nucleotide mutations in human mitochondrial genes according to claim 2, characterized in that, The reaction procedure for amplifying the whole genome DNA by using the UMI-labeled PCR reaction system is as follows: 98℃-1min, 70℃-65℃ (gradient reduction) 5s, and 72℃-10min.
4. The method of claim 1, wherein the method is for detecting a single nucleotide mutation in a human mitochondrial gene. performing targeted amplification and purification on the UMI-mtDNA to obtain a reaction product, comprising: amplifying and purifying the UMI-mtDNA by using a PCR reaction system, wherein the PCR reaction system comprises PrimeSTARGXL DNA Polymerase, 5×PrimeSTAR GXL Buffer, dNTP Mixture, a DNA template, an upstream primer, a downstream primer and nuclease-free water, and the reaction procedure is as follows: 95℃-1min, 98℃-10s, 64-70℃-14min, 30 cycles, and 68℃-5min.
5. The method of claim 1, wherein the method is for detecting a single nucleotide mutation in a human mitochondrial gene. analyzing the Nanopore sequencing data to obtain an mtDNA single nucleotide mutation result, comprising: performing preprocessing operations of removing an adapter sequence, removing a fixed sequence and extracting a sequence containing a UMI on the Nanopore sequencing data in sequence to obtain a sequence to be analyzed; classifying the sequence to be analyzed according to the number of UMI repetitions to obtain a UMI single-copy sequence and a UMI multi-copy sequence; clustering all UMI multi-copy sequences according to UMIs to form an UMI cluster; calculating the average nucleotide similarity of each sequence in the UMI cluster with the rest of the sequences, and taking the sequence with the highest average nucleotide similarity as a representative sequence of the UMI cluster; aligning the UMI single-copy sequence and all the representative sequences with a reference genome, and analyzing the alignment result by using a clair3 tool or a bamreadcount tool to obtain an mtDNA single nucleotide mutation result.
6. The method for detecting single nucleotide mutations in human mitochondrial genes according to claim 5, characterized in that, performing preprocessing operations of removing an adapter sequence, removing a fixed sequence and extracting a sequence containing a UMI on the Nanopore sequencing data in sequence to obtain a sequence to be analyzed, comprising: removing the adapter sequence in the Nanopore sequencing data by using a Porechop tool; Removing the fixed sequence from the 5' end and 3' end of the Nanopore sequencing data respectively by using the cutadapt tool; Searching for the sequence containing UMI from the 5' end and 3' end of the Nanopore sequencing data respectively by using the umi-tools tool, and finding the sequence containing specific UMI from the sequence containing UMI by using the Python script written to obtain the sequence to be analyzed.
7. The method of claim 6, wherein the method is for detecting a single nucleotide mutation in a human mitochondrial gene. The sequence containing specific UMI is NNNNTGNNNN.