A method for rapidly obtaining vertebrate mitochondrial genome sequences

Through RNA probe enrichment and multiple annealing circular cycle amplification technology, combined with magnetic bead capture and assembly algorithms, the time-consuming, labor-intensive or costly problems of existing technologies are solved, and animal mitochondrial genome sequences can be obtained quickly and economically.

CN115992204BActive Publication Date: 2025-09-30SHANGHAI OCEAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210999993.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-09-30
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

Existing methods for obtaining animal mitochondrial genome sequences are time-consuming, labor-intensive, or costly, and it is difficult to achieve complete mitochondrial genome assembly, especially in the absence of a reference sequence for the target species.

Method used

Mitochondrial DNA was enriched using RNA probes, and a mitochondrial genomic library was constructed by combining multiple annealing circular amplification and magnetic bead capture technology. The complete genome was then assembled using Trinity and NOVOPlasty.

Benefits of technology

It significantly improves the proportion and coverage of mitochondrial genome data, reduces sequencing costs and time, and can quickly obtain a complete mitochondrial genome without a reference sequence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115992204B_ABST
    Figure CN115992204B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of biotechnology, and specifically relates to a method for establishing a mitochondrial DNA gene library and a method for obtaining a mitochondrial genome sequence. The method for establishing the library is as follows: (1) after hybridization of the total DNA of the sample with an RNA probe, the target sequence is captured; (2) the target sequence is collected and amplified to obtain the amplified mitochondrial DNA. The amplified DNA is fragmented, blunt-end repaired, connected to a sequencing adapter, and indexed by PCR to construct a library. The obtained gene library is sequenced, and the sequencing results are assembled. The scheme of the present invention enriches the original long-fragment mitochondrial DNA before constructing the sequencing library, so that more mitochondrial DNA data can be obtained, thereby improving the enrichment efficiency, increasing the proportion and coverage of the mitochondrial genome data, and achieving complete assembly; at the same time, the amount of data required for assembly is reduced, and the sequencing cost is greatly reduced. The present invention does not require a reference genome of the corresponding species, and can relatively quickly and economically obtain the complete mitochondrial genome of various vertebrates, providing a feasible technical route for establishing a vertebrate environmental DNA mitochondrial genome database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biotechnology, and in particular relates to a method for establishing a mitochondrial DNA gene library and a method for obtaining a mitochondrial genome sequence. Background Art

[0002] Animal mitochondrial gene sequences are characterized by maternal inheritance, high copy numbers, and smaller genome sequence data sets than nuclear genomes. While they generally mutate at a high rate, some genes encoding crucial life processes are relatively conservative. For this reason, animal mitochondrial gene sequences have become valuable tools in fields such as phylogenetics and ecology. For example, environmental DNA (ED) studies often use mitochondrial gene sequences to assess the species and abundance of organisms in an environment. However, researchers often use different mitochondrial gene markers, such as CO1 and 12S, making it difficult to standardize EDN databases. Furthermore, analysis results from a single or a few molecular markers are often insufficient to fully characterize a specific habitat. Furthermore, the methods used by researchers to collect, extract, amplify, enrich, and process EDN for these markers vary widely, creating significant challenges for scientists in selecting the right method for their research and increasing the cost of trial and error. Using complete mitochondrial genome sequences as EDN databases would address these challenges. Because the research on single-site mitochondrial genes or multi-site mitochondrial genes can use the mitochondrial genome database as a reference database without being restricted by the method, and the corresponding species information provided by the complete mitochondrial genome is far greater than that of the gene sequences of several sites, multi-site combined analysis can also solve the problem of species identification caused by insufficient sequence information.

[0003] However, using the mitochondrial genome as a database of environmental DNA requires obtaining a large amount of mitochondrial genome data, and the current methods for obtaining the mitochondrial genome are either time-consuming and labor-intensive or expensive, such as primer walking or long-range PCR amplification based on PCR and Sanger sequencing, and genome skimming based on high-throughput sequencing.

[0004] Mitochondrial copy numbers vary significantly between different tissues and cells. Within cells containing mitochondria, the number of mitochondria ranges from a few to hundreds. Generally speaking, animal liver and muscle tissues have the highest number of mitochondria and, accordingly, higher amounts of mitochondrial DNA. Furthermore, nuclear and mitochondrial DNA are difficult to separate after cell lysis. Therefore, there are currently no methods specifically designed for mitochondrial extraction. Instead, whole-genomic DNA is directly extracted, and the mitochondrial DNA is subsequently purified using various methods for subsequent analysis.

[0005] Although mitochondrial DNA exists in multiple copies within cells, it is still a tiny fraction compared to nuclear DNA, whose sequences can reach billions of bases in length. Therefore, no matter which sequencing method is used, mitochondrial DNA needs to be purified to a certain degree to increase its proportion, thereby increasing sequencing efficiency. There are usually two approaches to purifying a specific type of target DNA: one is to amplify the target DNA through methods such as PCR to increase the amount of effective data and thus its proportion; the other is to design probes that hybridize and capture the target DNA while washing away non-target DNA fragments to reduce the amount of non-target DNA and increase the proportion of target DNA.

[0006] To obtain a species' complete mitochondrial genome, in addition to the aforementioned DNA extraction, purification, and sequencing, the sequencing results must also be assembled and proofread. In this process, if we have a mitochondrial genome reference sequence for the target species, or a reference sequence for a closely related species, then mapping can quickly and easily yield relatively accurate results. However, during the research process, researchers often encounter difficulties in determining the specific type of target species, or discover that there is no reference sequence for the target species after morphological identification. This complicates the assembly process considerably.

[0007] Target Sequence Capture, also known as gene capture or gene enrichment, is a method used to capture target sequences of interest to scientists. It is usually based on pre-designed RNA or DNA baits (probes), which hybridize with target sequences in an established DNA library, and then use a series of methods to separate and purify the enriched sequences for sequencing or downstream experiments. It uses the principle of probe hybridization to enrich and purify nucleic acid fragments of interest to researchers, thereby increasing the proportion of effective data and reducing sequencing and analysis costs. The emergence of this type of method conforms to the development of high-throughput sequencing technology. In addition, there are also methods that use nucleic acid-protein complexes such as CRISPR / Cas systems to capture target sequences.

[0008] For example, Sevigny J et al. designed probes based on the mitochondrial genome sequences of all metazoans. They then used the probes to enrich high-throughput sequencing libraries of any metazoan, followed by sequencing after enrichment. However, because the probe design used the entire sequence, the synthesis cost increased significantly. Furthermore, since the enrichment was limited to short sequences in the library, the proportion of valid data was low, and information outside the loci was incomplete, resulting in high costs and low efficiency.

[0009] Therefore, it is necessary to establish a method that can quickly obtain mitochondrial genome sequences, increase the proportion of effective data, reduce sequencing costs, and obtain more comprehensive sequence information. Summary of the Invention

[0010] The present invention aims to provide a method for rapidly obtaining animal mitochondrial genome sequences.

[0011] An RNA probe for enriching mitochondrial DNA, wherein the nucleotide sequence of the probe contains any one or more sequences selected from SEQ ID No. 1 to No. 11.

[0012] Preferably, the nucleotide sequence of the RNA probe is selected from any one or more of SEQ ID No. 1-No. 11.

[0013] More preferably, the nucleotide sequence of the RNA probe is shown as SEQ ID No.1-No.11.

[0014] Furthermore, the RNA probe is modified with biotin.

[0015] The RNA probes described above can be used to enrich animal mitochondria, determine animal mitochondrial genome sequences, obtain animal mitochondrial genome sequences, or establish an animal environmental DNA mitochondrial genome database. In particular, they can be used to enrich vertebrate mitochondria, determine vertebrate mitochondrial genome sequences, obtain vertebrate mitochondrial genome sequences, or establish a vertebrate environmental DNA mitochondrial genome database.

[0016] Another embodiment of the present invention is a method for establishing a mitochondrial genome library, comprising the following steps:

[0017] (1) After the total DNA of the sample is hybridized with the above-mentioned RNA probe, the target sequence is captured;

[0018] (2) Collect the target sequence and amplify it to obtain the amplified mitochondrial DNA.

[0019] Preferably, in step (2), the target sequence is collected using magnetic beads modified with streptavidin.

[0020] In step (2), multiple annealing cycles are used for amplification, followed by washing to obtain the amplified mitochondrial DNA.

[0021] Furthermore, the method for establishing a mitochondrial genome library further comprises the following steps:

[0022] The amplified mitochondrial DNA was fragmented, blunt-end repaired, ligated with sequencing adapters, and subjected to index PCR to construct the library.

[0023] Another embodiment of the present invention is a method for rapidly obtaining a vertebrate mitochondrial genome sequence, characterized in that it comprises the following steps:

[0024] 1. Sequencing the gene library obtained by the above method;

[0025] II. Assemble the sequencing results.

[0026] Preferably, in step II, Trinity and NOVOPlasty are used for assembly.

[0027] Furthermore, the conserved site reads obtained by de novo assembly and screening were assembled using Trinity, and contigs of the conserved site sequences were assembled and extended by NOVOPlasty to obtain the complete mitochondrial genome.

[0028] Preferably, the sequencing results are preprocessed, and then the most similar sequences are screened using BLAST based on the conserved sequences of the enriched sites, followed by de novo assembly using Trinity. The assembly results are further screened using BLAST to obtain the most similar contigs, and the final contig is extended using NOVOPlasty to obtain the complete mitochondrial genome.

[0029] The present invention is a "enrich first, then build library" method based on long-fragment enrichment. Unlike traditional methods of enriching shorter fragment library DNA, this method enriches the original long-fragment mitochondrial DNA before constructing the sequencing library, aiming to obtain more mitochondrial DNA data, thereby improving enrichment efficiency and reducing sequencing and subsequent assembly costs.

[0030] Using the model organism zebrafish (Danio rerio) as an example, the results showed that, based on the same amount of sequencing data, this new method increased the proportion of mitochondrial genome data by over 180 times compared to traditional genome overviews. Even compared to traditional enrichment methods, the proportion of mitochondrial genome data increased by approximately three times. More importantly, traditional methods provide incomplete coverage of mitochondrial genome data, resulting in incomplete mitochondrial genome assembly. Under the same conditions, this method achieved 100% data coverage, indicating that this method reduces the possibility of failure to assemble a complete mitochondrial genome due to the absence of sequences from non-enriched sites, and can achieve complete assembly.

[0031] The authors also calculated the amount of data necessary to assemble a complete vertebrate mitochondrial genome using this method and compared it with other methods. The results showed that the improved method only requires 50MB of data to assemble a complete mitochondrial genome, less than 5% of the minimum required data volume, significantly reducing sequencing costs.

[0032] The method was also successfully assembled using the Chinese soft-shell turtle (Pelodiscus sinensis) and the pig (Sus scrofa), demonstrating that the method is also applicable to vertebrates of other classes. The probes and methods of the present invention are particularly suitable for detecting and assembling vertebrate mitochondrial genomes.

[0033] For a single sample, the entire experimental process takes only three days, and the total cost of obtaining a complete mitochondrial genome is less than 50% of traditional methods. The final results show that this method can relatively quickly and economically obtain complete mitochondrial genomes of various vertebrates without the need for a reference genome of the corresponding species. This method provides a feasible technical route for establishing a vertebrate environmental DNA mitochondrial genome database. It can also be used to obtain mitochondrial genome assemblies of uncharacterized vertebrates or those without reference sequences, contributing to the improvement of the vertebrate mitochondrial genome database. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 Depth distribution curves of the three methods in Example 1

[0035] Figure 2 The coverage of the three methods in Example 1 under different data amounts DETAILED DESCRIPTION

[0036] Main reagents:

[0037] (1) Buffer Tango (10x) (Thermo, Cat. No. BY5);

[0038] (2) dNTPs (10 mM each) (Invitrogen, Cat. No. 18427088);

[0039] (3) ATP (100 mM) (Thermo, Cat. No. R0441);

[0040] (4)T4 polynucleotide kinase (10U / μL) (Thermo, Cat. No.: EK0031);

[0041] (5) T4 DNA polymerase (5 U / μL) (Thermo, Catalog No.: EP0061);

[0042] (6) T4 DNA ligase (5 U / μL) (Invitrogen, Cat. No. 15224041) and T4 DNA ligase buffer and PEG-4000 (50%) supplied with the T4 DNA ligase kit;

[0043] (7) Bsm polymerase, large fragment (8 U / μL) (Thermo, Catalog No.: EP0691) and matching Bsm buffer;

[0044] (8) KAPA HiFi taq Ready Mix (2x) (KAPABIOSYSTEMS, catalog number: KK2602);

[0045] (9) UltraPure TM SSPE, 20x (Invitrogen, cat. no. 15591043);

[0046] (10) UltraPure TM 0.5 M EDTA, pH 8.0 (Invitrogen, Cat. No. 15575020);

[0047] (11) Denhardt's Solution (50x) (Invitrogen, Cat. No. 750018);

[0048] (12) 10% SDS solution (Sangon, catalog number: B548118-0100);

[0049] (13)SUPERase·In TM RNase Inhibitor (20 U / μL) (Invitrogen, cat. no. AM2696);

[0050] (14) DEPC treated water (Shenggong, catalog number: B501005-0500);

[0051] (15) Human Cot-1 DNA (Invitrogen, Cat. No. 15279011);

[0052] (16)Dynabeads TM MyOne TM Streptavidin C1 (Invitrogen, Cat. No. 65002);

[0053] (17) Tween-20 (Amresco, catalog number: 0777-1L);

[0054] (18) TE buffer (Sanggong, catalog number: B548106);

[0055] The magnetic beads used were modified with streptavidin.

[0056] RNA probe sequences for enrichment (SEQ ID No. 1 - 11) are as follows:

[0057] 1.ccgggtactacgagcactagcttaaaacccaaaggacttggcggtgctttagatccacctagaggagcctgttctagaaccgataacccccgttaaacctcaccctctcttgttcttcccg

[0058] 2.gactataagtttaacggccgcggtattttgaccgtgcaaaggtagcgcaatcacttgtcttttaaatgaagacctgtatgaatggcataacgagggcttaactgtctcctttttccagtcaatgaaattgatctccccgtgca

[0059] 3.tcgacaagagggtttacgacctcgatgttggatcaggacatcctaatggtgcagccgctattaagggttcgtttgttcaacgattaaagtcctacgtgatctgagttcagaccggagtaatccaggtcagtttctatctatgccacgatcttttct

[0060] 4.gctcgaaccctacctgaagagatcaaaactcttagtgcttccactacaccacttccctagtaaagtcagctaaataagcttttgggcccataccccaaacatgttggttaaactccttcctttgct

[0061] 5.cccacatcttctgcatgcaaaacagacattttaattaagctaaagccttactagacaggaaggcctcgatcctacaaactcttagttaacagctaagcgcttaaaccaacaagcatctgtctaacctttccccgc

[0062] 6.ccatcttacctgtggcaatcacacgttgatttttctcaactaatcacaaagacatcggcaccctctatctagtatttggtgcttgagccggaatagtaggaactgcattaagcctcctaattcgggca

[0063] 7.cacatgctttcgtaataattttctttatagtaatgccaattataattggaggttttggaaactgactagtgccactaatgattggtgcaccagacatggccttccctcgaataaataacatgagt

[0064] 8.gccggcatcacaatacttctaacagaccgaaacctaaacacaaccttctttgaccctgccggaggaggagaccccatcctttaccaacacttattctgattctttggacaccctgaagtttatattct

[0065] 9.agggtttattgtctgagcccatcacatgttcaccgtaggaatggacgtagatacacgggcttactttacttccgccacaataattattgccatcccaaccggagtaaaagtcttcagctg

[0066] 10.cagtagccataattcaggcctatgtctttgttcttcttttaagcctttacctacaagaaaacgtttaatggcccatcaagcacacgcatatcacatagttgaccccagcccatgacccctaaca

[0067] 11.tctttagccctcttctcccccaatctacttggtgatcctgacaacttcacccccgcaaaccctctagttacccctccccacattaaacccgaatggtacttcttatttgcctacgccatcctacgctcaat

[0068] The universal primer sequences for MALBAC amplification are as follows:

[0069] MALBAC+8N:gtgagtgatggttgaggtagtgtggagnnnnnnnn

[0070] MALBAC: gtgagtgatggttgaggtagtgtggag

[0071] Example 1

[0072] Taking zebrafish as the research object, total DNA was extracted from zebrafish muscle, and a mitochondrial genome library was established, and the mitochondrial genome was obtained after sequencing.

[0073] (1) Pre-hybridization PCR

[0074] (1) Set the following program on the PCR instrument: 95°C for 5 minutes, 65°C for 5 minutes, 65°C for 10 minutes, 60°C for 10 hours, and then hold at 60°C.

[0075] (2) Prepare hybridization mixture (Hyb Mix) on ice according to the number of samples and the following system:

[0076]

[0077] Thoroughly shake and mix evenly and collect the reagents on the tube wall to the bottom of the tube by short centrifugation. Take 5 μL of each sample into a pre-prepared empty PCR tube; temporarily store in an ice box or a 4°C refrigerator until step (6) before use;

[0078] (3) Prepare the DNA mixed solution (Lib Mix) for library preparation according to the number of samples on an ice box:

[0079]

[0080] Add 1.25 μL of reagent to each empty PCR tube with the corresponding sample number written on it, and add 6.25 μL of sample DNA (100 ng to 500 ng). Gently pipette to mix, and briefly centrifuge; temporarily store in an ice box or a 4°C refrigerator until step (5).

[0081] (4) Prepare the probe mixture (Baits Mix, SEQ ID No. 1-11) on an ice box according to the number of samples and the following system:

[0082]

[0083] Mix the mixture gently by pipetting; store it temporarily in an ice box or a 4°C refrigerator until step (6).

[0084] (5) Transfer the Lib Mix to the PCR instrument and start the program set in (1); the DNA will be denatured at 95°C for 5 minutes;

[0085] (6) When the program reaches the second step, that is, when the temperature drops from 95°C to 65°C, turn on the PCR instrument and add the Hyb Mix and Bait Mix (do not move the Lib Mix); all mixed solutions will be preheated to 65°C in the PCR instrument;

[0086] (7) When the program enters the third step (65°C for 10 minutes), transfer 5 μL of Hyb Mix and 3 μL of Baits Mix to Lib Mix and mix gently by pipetting. The entire operation process is performed on the PCR instrument. Do not remove the PCR tube.

[0087] (8) Allow the mixed solution to complete all steps of the program on the PCR instrument;

[0088] (II) Magnetic bead enrichment and elution

[0089] (1) Add n × 10 μL (n is the number of samples) of MyOne (Invitrogen cat#65002) or M270 (Invitrogen cat#:653-06) magnetic beads (streptavidin modified) to a new PCR tube (no more than 180 μL per tube);

[0090] (2) Place the test tube on a magnetic plate to collect the magnetic beads, and remove the supernatant after the solution becomes clear;

[0091] (3) Add 200 μL of Binding Buffer (see the Appendix for the recipe) to wash the magnetic beads; gently pipette the magnetic beads to suspend them, then place them back on the magnetic plate to collect the beads and aspirate the supernatant;

[0092] (4) Repeat step (3) twice;

[0093] (5) Add n×20 μL Binding Buffer and pipette to suspend the magnetic beads, then add 1 μL 10% Tween;

[0094] (6) Take empty PCR tubes corresponding to the number of samples, number them, and add 180 μL of Binding Buffer to each tube, followed by 20 μL of magnetic bead suspension;

[0095] (7) Preheat at 60°C for 5 minutes;

[0096] (8) Transfer all the hybridization solution to the magnetic bead solution and then incubate at 60°C on a hybridizer for 30 minutes. Then, collect the magnetic beads using a magnetic plate and discard the supernatant. Do this as quickly as possible, taking care not to terminate the PCR program and keep it at 60°C.

[0097] (9) During the incubation process, prepare 3 × n PCR tubes, add 190 μL of Wash Buffer 2 to each tube, and transfer to a PCR instrument to preheat at 60°C for at least 10 minutes;

[0098] (10) Add 186 μL of Wash Buffer 2 to the magnetic bead tube, gently pipette the magnetic beads to suspend them, and then transfer them to a PCR instrument and let them stand for 10 minutes; then collect the magnetic beads using a magnetic plate and discard the supernatant;

[0099] (11) Repeat step (10) twice, performing a total of three washes at 60°C;

[0100] (12) Finally, wash again with 186 μL of DEPC water at room temperature;

[0101] (13) Add 25 μL of DEPC water to the tube;

[0102] (3) Multiple Annealing and Looping-Based Amplification Cycles (MALBAC)

[0103] (1) Prepare the following mixed solutions on an ice box according to the number of samples:

[0104]

[0105]

[0106] Prepare empty PCR tubes corresponding to the number of samples, add 15 μL of the mixture to each tube, and add 10 μL of DNA with magnetic beads. At the same time, add a negative control and use DEPC water instead of the added 10 μL DNA with magnetic beads; mix thoroughly, centrifuge briefly, and finally amplify on a PCR instrument according to the following program: 92°C for 3 minutes, 14-18 cycles of: 10°C for 45 seconds, 20°C for 45 seconds, 30°C for 45 seconds, 40°C for 45 seconds, 50°C for 45 seconds, 68°C for 2 minutes, 92°C for 20 seconds, and 58°C for 20 seconds; maintain at 4°C after the cycle is completed;

[0107] After the above PCR program is completed, add 0.8 μL of primer 2: MALBAC to each tube, mix well, centrifuge briefly, and then perform the following program: 92°C for 2 minutes, 12-16 cycles of: 92°C for 20 seconds, 58°C for 30 seconds, and 68°C for 1 minute; store at 4°C after the cycle is completed;

[0108] (2) The DNA was washed using the magnetic bead method, and finally 30 μL of water was added to dissolve the DNA. The concentration was measured using Nanodrop3300 for subsequent construction of the sequencing library;

[0109] (4) DNA fragmentation

[0110] Transfer 0.3–1 μg of DNA to a new PCR tube. Add PCR-grade ultrapure water (hereinafter referred to as PCR water) to the tube until no bubbles are present after capping. Use a Covaris M220 Focused-ultrasonicator (Covaris, Woburn, USA) to fragment the DNA to approximately 300 bp (the fragment length can be detected by agarose gel electrophoresis). Set up a positive control and a negative control before library construction. Then, use the magnetic bead method developed by Rohland et al. to wash and collect the DNA (no new solution is added after the alcohol wash step).

[0111] (5) Flat end repair

[0112] (1) Prepare the mixed solution on an ice box according to the number of samples and the following system:

[0113]

[0114]

[0115] (2) Add 20 μL of the mixture directly to the PCR tube with magnetic beads in the previous step and mix well;

[0116] (3) Perform the following program on a PCR instrument: 25°C for 15 minutes, 12°C for five minutes;

[0117] (4) Wash the DNA using the magnetic bead method (keep the beads in the tube and do not separate them);

[0118] (6) Connecting sequencing adapters

[0119] (1) Prepare the mixed solution on an ice box according to the number of samples and the following system (without adding a connector, and then add a connector to each sample separately):

[0120]

[0121] (2) Add 38 μL of the mixture to each tube and mix well;

[0122] (3) Add 1 μL of IS1 and IS2 to each tube (each sample number corresponds to a set of IS1+IS2 to distinguish different samples during sequencing and subsequent analysis. The numbers of different sets of IS1+IS2 cannot be exactly the same).

[0123] (4) Incubate at 22°C for 30 minutes on a PCR instrument;

[0124] (5) Wash the DNA using the magnetic bead method (keep the magnetic beads in the tube and do not separate them);

[0125] (7) Filling

[0126] (1) Prepare the mixed solution on an ice box according to the number of samples and the following system:

[0127]

[0128] (2) Add 40 μL of the mixture to each tube and mix well;

[0129] (3) Incubate at 37°C for 20 minutes on a PCR instrument;

[0130] (4) Wash the DNA using the magnetic bead method and finally dissolve it in 35 μL of TE Buffer (keep the magnetic beads in the tube and do not separate them);

[0131] (8) Indexing PCR

[0132] (1) Prepare the mixed solution on an ice box according to the number of samples and the following system:

[0133]

[0134] (2) Add 13 μL of the mixture, 0.5 μL of P7 primer, and 11 μL of the DNA solution with magnetic beads to each tube, mix well, and then transfer to the PCR instrument;

[0135] (3) Perform the following program on a PCR instrument: 98°C for 45 seconds, 12 to 16 cycles of: 98°C for 15 seconds, 60°C for 30 seconds, and 72°C for 1 minute; hold at 4°C after the cycle is complete;

[0136] (4) Wash the DNA using the magnetic bead method (add new magnetic beads) and dissolve it in 25 μL of TE Buffer;

[0137] (5) Take 1 μL and use agarose gel electrophoresis to check the library construction.

[0138] (IX) Sequencing and assembly

[0139] The library was sequenced using the Illumina Novaseq-PE150 platform. Sequencing results were preprocessed, including decompression, data sorting according to the inline index, and removal of adapter sequences and low-quality sequences. Sequences were then identified based on conserved sequences within the enriched sites using BLAST fusion to identify the most similar sequences. De novo assembly was performed using Trinity v2.11.0. The assembly results were further filtered using BLAST fusion to identify the most similar contigs. The final contigs were then extended using NOVOPlasty. The assembly script is available at https: / / github.com / Checunmil y / mito_assemble.

[0140] The sequencing data (SRA: PRJNA796186) and assembly sequence results (GenBank numbers: OM236540, OM236541) have been uploaded to NCBI and are not significantly different from the sequences already in the database.

[0141] From the analysis of sequencing depth, it can be seen that after enrichment and amplification, the mitochondrial DNA fragments obtained are longer.

[0142] Result analysis:

[0143] 1. Assembly results

[0144] For each zebrafish sample tested, paired reads were aligned to the assembled zebrafish sample mitochondrial genome using BWA-MEM (v0.7.16a-r1181) with default parameters, and the coverage, average depth, and depth of each site were calculated using the view, coverage, and depth commands of Samtools (v1.10) with default parameters. All comparative analysis data were normalized to the number of reads before alignment, and the same number of reads (but with varying read lengths and therefore different absolute data volumes) was captured for analysis. The final results were presented using the number of reads. The enrichment factor was calculated by dividing the average depth of all samples for each enrichment strategy by the average depth of all samples for the standard library construction strategy.

[0145] We required 2GB of sequencing data for all zebrafish samples used to test the efficiency of our method, and after data preprocessing, we obtained 170MB to 500MB of data per sample. Taking the zebrafish genome size (approximately 1400MB) as a reference, the average depth of our data was less than 1x. However, for the zebrafish mitochondrial genome (approximately 16k), the data volume obtained after enrichment sequencing was much larger than the mitochondrial genome size.

[0146] In addition to the "enrichment first, then library construction" method of the present invention, the "direct library construction" and "library construction first, then enrichment" methods were also used for comparison.

[0147] Direct library construction involves obtaining total DNA, breaking it into fragments of approximately 300 bp using an ultrasonic disruptor, performing blunt-end repair, adding adapters, filling in, and indexing PCR, followed by cleaning using a magnetic bead wash method.

[0148] The library construction-before-enrichment method involves ultrasonication to fragment the total DNA to approximately 300 bp, followed by blunt-end repair, adapter addition, fill-in, and pre-hybridization PCR. Each of these steps is followed by DNA cleansing and collection using the magnetic bead method developed by Rohland et al. The enrichment steps are similar to those used for exon enrichment, but only one enrichment step is performed: hybridization, capture and elution using magnetic beads with streptavidin, and indexing PCR. After the indexing PCR, the DNA is cleaned using the magnetic bead method.

[0149] After the assembly process was completed, only the improved "enrichment first, then library construction" sample among the three strategies assembled a complete mitochondrial genome (GenBank number: OM236539). Both the direct library construction and "library construction first, then enrichment" samples could not assemble a complete mitochondrial genome.

[0150] During the assembly process, genome assembly software such as SOAPdenovo and SPAdes, as well as mitochondrial assembly software such as MITObim, failed to assemble a complete mitochondrial genome. Using Trinity to directly assemble all sequencing results also failed to obtain a complete mitochondrial genome.

[0151] The final assembly strategy was to use Trinity to assemble the conserved site reads obtained from scratch, assemble the contigs of the conserved site sequences, and then use these contigs as "seeds" to extend them to NOVOPlasty, ultimately obtaining a complete mitochondrial genome.

[0152] 2. Comparison of Coverage and Depth

[0153] After obtaining the complete mitochondrial genome of the zebrafish sample, we compared the effective data percentage, average depth, coverage, and enrichment factor of different library construction strategies. The statistical results are shown in Table 3-1.

[0154] Table 1. Comparison statistics of three methods

[0155]

[0156] From the analysis of sequencing depth, it can be seen that after enrichment and amplification, the mitochondrial DNA fragments obtained are longer.

[0157] T-tests were used to test the mean depth. P values ​​were < 0.0025 for library construction followed by enrichment versus direct library construction, and P values ​​were < 0.001 for enrichment followed by library construction versus direct library construction, and for enrichment followed by library construction versus library construction followed by enrichment. Significance tests revealed highly significant differences in the available data between the three methods. In the direct library construction method, mitochondrial data accounted for only approximately 0.07%, which is consistent with the initial experimental design, given the approximately 100,000-fold size difference between the zebrafish nuclear and mitochondrial genomes and the range of mitochondrial DNA copy numbers within a single cell (ranging from a few to hundreds). This data varied significantly across species and tissues, with a higher proportion of mitochondrial DNA data in species with relatively small nuclear genomes and tissues with high mitochondrial abundance, such as liver and muscle. This proportion was lower in tissues with larger nuclear genomes and lower mitochondrial abundance. Comparisons with two other enrichment strategies showed that mitochondrial DNA enrichment significantly increased the proportion of mitochondrial DNA data, with the improved "enrichment first, then library construction" approach achieving the best results, with an average depth of 186 times that of direct library construction. Even compared to traditional enrichment strategies, this approach yielded nearly three times as much effective data. Although both direct library construction and library construction followed by enrichment have a minimum depth of 0, meaning there are sites with no sequence information and a complete sequence cannot be assembled, the enrichment method still significantly improved coverage.

[0158] If the data volume is increased, the "build library first, enrich library later" approach can still relatively well assemble a complete mitochondrial genome. Even with this low data volume, the "enrich library first, build library later" approach can still maintain 100% coverage and a minimum depth of 8, confirming that the data measured by this method does indeed cover the complete mitochondrial genome.

[0159] 3. Sequencing Depth Distribution

[0160] The depth command of Samtools and the default parameters were used to call out the depth of each base of all sequencing samples, and the average depth of each base of different strategies was calculated. Then, the data was visualized using R, and a depth curve was drawn to illustrate its depth distribution, as shown in the figure below. Figure 1 As shown, the curves from top to bottom are E (enrichment first, then library construction), L (library construction first, then enrichment), and D (direct library construction). The X-axis represents the base sequence in the mitochondrial genome, and the Y-axis represents the corresponding depth of each base site.

[0161] The depth curve also shows that the "enrich first, then build library" strategy (E) has a significantly higher maximum depth and complete coverage than the other strategies. However, the position of its depth peak is different from the "build library first, then enrich" strategy (L). The depth curve of "build library first, then enrich" is completely consistent with the probe site, with a higher depth at the enrichment site and its vicinity, and a very low depth outside the enrichment site, close to the depth of direct library construction. The depth curve of "enrich first, then build library" is not completely consistent with the enrichment site, and is even very low at some sites where peaks should appear (such as around 7000 and 10000). At the same time, peaks appear in some places outside the enrichment site (such as around 11000 and 14000). This feature is still shown after the depth curve of a single sample is output. We speculate that the reason for this phenomenon is that the random primers used in the MALBAC amplification step after enrichment are competed for by the probes used for enrichment, and the RNA probe binds to the target site again during the renaturation process, hindering the binding of the random primers used for amplification, thereby reducing the amplification efficiency of the site. However, regions farther away from the enrichment site or regions with large sequence differences from the enrichment site are not affected and can be amplified normally, resulting in this phenomenon.

[0162] The results showed that this method was significantly superior to the original method in terms of maximum sequence depth and coverage.

[0163] IV. Verification of the minimum amount of data required to obtain a complete mitochondrial genome

[0164] In order to determine the minimum amount of data required to obtain a complete mitochondrial genome using the improved method and to avoid sequencing waste, we tested the data of the three samples of "enrichment first and then library construction". During the test, we reduced the number of sequences by 10,000 sequences (20,000 reads) each time, and aligned the intercepted data to the assembled mitochondrial genome again using BWA-MEM. We also used Samtools to obtain the coverage, output the number of sequences and coverage each time, and used R to draw the corresponding curve graph. The coverage situation under different data amounts is shown below. Figure 2 As shown in the figure. The X-axis represents the number of reads, and the Y-axis represents the coverage given the corresponding number of reads. The curves from top to bottom in the figure are enrichment first and then library construction (E), library construction first and then enrichment (L), and direct library construction (D).

[0165] As can be seen from the figure, the first time the coverage of the "enrich first, then build library" method fell below 100% was at 140,000 reads, after which the coverage began to decline. The coverage of the other two methods, even with smaller data volumes, was less than 100% from the start, and the downward trend was significantly faster than that of the "enrich first, then build library" method. It is worth noting that even with extremely low data volumes (20,000 reads), the improved new method can still maintain over 90% coverage, indicating that our approach to enriching long fragments can indeed obtain DNA fragments farther from the enrichment site, thereby maintaining higher coverage and more uniform depth.

[0166] 5. Cost Comparison of the “Enrichment-First, Library Construction” Approach and Existing Methods

[0167] According to the above experimental results, it can be concluded that for zebrafish muscle tissue, this method can obtain complete mitochondrial genome data of about 50MB. This minimum data volume also needs to take into account the species genome size and tissue sample differences, plus the impact of PCR duplication caused by the pre-sequencing amplification link, so a sequencing data volume of no less than 200MB is recommended. However, using traditional methods, such as first building a library and then enriching it, the amount of data required for high-throughput sequencing is more than 4G, which greatly saves sequencing costs. However, the method based on Sanger sequencing and PCR is more expensive because the sequences that can be obtained by a single amplification and sequencing are limited, so repeated PCR experiments and sequencing are required.

[0168] In terms of time cost, traditional Sanger sequencing-based methods require repeated PCR amplification, a process that often takes several days, and even longer if sequencing and analysis after each amplification are included. In contrast, the "enrichment first, library construction later" approach employed in this method reduces the total experimental process from DNA extraction to library submission to testing by only three days, and each sample only needs to undergo the entire experimental process once, making this method more time-efficient. However, it should be noted that the "library construction first, enrichment later" approach based on high-throughput sequencing is faster than this method, primarily due to the omission of the MALBAC whole-genome random amplification step. Furthermore, pre-library construction can also reduce some experimental steps. For example, multiple samples can be pooled and then enriched and sequenced together, and then separated during analysis. In contrast, the "enrichment first, library construction later" approach does not pre-add sequencing adapters and specific index sequences before enrichment, so sample pooling and subsequent enrichment are not possible. Each enrichment tube can only contain one sample. While this eliminates the influence of enrichment bias, it increases the number of experimental steps. The "direct library construction" protocol, which lacks enrichment and whole-genome amplification steps, is the fastest among high-throughput sequencing-based protocols. However, as shown in the previous results, this method also has the lowest efficiency. In summary, despite its advantages in enrichment efficiency and cost-effectiveness, the improved method is not as convenient as the "library construction followed by enrichment" or "direct library construction" protocols in terms of time, requiring anywhere from half a day to a day depending on the protocol. Therefore, this method can achieve optimal results in terms of experimental efficiency, cost-effectiveness, and time cost.

[0169] Example 2

[0170] Muscle samples from the Chinese soft-shelled turtle (Reptiles) and the domestic pig (Mammalians) were used as research objects, respectively. The method of Example 1 was used to enrich mitochondrial DNA, construct libraries, and sequence them. Both methods required the sequencing company to provide 4GB of data, and complete mitochondrial genomes were assembled.

[0171] The sequencing data (SRA: PRJNA796186) and the assembled sequence results (GenBank IDs: OM236540 and OM236541) have been uploaded to NCBI and do not differ significantly from existing sequences in the database. This shows that this method can be applied to the sequencing of vertebrate mitochondrial genomes, obtaining vertebrate mitochondrial genomes without the need for a reference sequence of the target species.

Claims

1. An RNA probe for enriching mitochondrial DNA, characterized in that It consists of the nucleotide sequence shown in SEQ ID No.1-No.

11.

2. The RNA probe according to claim 1, wherein The RNA probe is modified with biotin.

3. Use of the RNA probe according to claim 1 or 2 in enriching vertebrate mitochondria, obtaining animal mitochondrial genome sequences, or establishing an animal environmental DNA mitochondrial genome database.

4. A method for establishing a mitochondrial genome library, characterized in that: The following steps are involved: (1) After the total DNA of the sample is hybridized with the RNA probe described in claim 1 or 2, the target sequence is captured; (2) Collect the target sequence and amplify it to obtain the amplified mitochondrial DNA.

5. The method for establishing a mitochondrial genome library according to claim 4, characterized in that: In step (2), multiple annealing cycles are used for amplification.

6. The method for establishing a mitochondrial genome library according to claim 4, characterized in that: The following steps are also included: The amplified DNA was fragmented, blunt-end repaired, ligated with sequencing adapters, and subjected to index PCR to construct the library.

7. A method for rapidly obtaining vertebrate mitochondrial genome sequences, characterized in that: The following steps are involved:

1. Sequencing the gene library obtained by the method according to any one of claims 4 to 6; II. Assemble the sequencing results.

8. The method for rapidly obtaining vertebrate mitochondrial genome sequences according to claim 7, characterized in that: In step II, Trinity and NOVOPlasty were used for assembly.

9. The method for rapidly obtaining vertebrate mitochondrial genome sequences according to claim 8, characterized in that: The conserved site reads obtained by de novo assembly and screening were assembled using Trinity. After assembling the contigs of the conserved site sequences, they were extended by NOVOPlasty to obtain the complete mitochondrial genome.

Citation Information

Patent Citations

  • Capture probe set and kit for detecting human mitochondrial genes

    CN105779590A

  • Bovine mitochondrial genome capture probe kit

    CN111172159A