Sequencing method, device and equipment of mitochondrial genome sequence and storage medium
By comparing and assembling three generations of read-long sequences in the species read-long database, the extraction difficulties in mitochondrial sequencing process are solved, and efficient and cost-free mitochondrial genome sequence recognition and assembly are achieved.
Patent Information
- Application Number
- CN202311529102.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, sequencing and analysis of mitochondria alone requires isolating and purifying mitochondrial DNA, which is difficult and costly. How to use the whole genome sequencing data of species to mine mitochondrial characteristic sequences and reduce the mitochondrial genome.
By obtaining the three-generation read length sequences in the species read length database, and aligning them with the reference genomic sequences of the mitochondria of the target species one by one, it is determined that the read length sequencing sequences with alignment and coverage above the threshold are high-quality sequences, and trimming and assembled to obtain the measured genomic sequences of the mitochondria of the target species.
Without mitochondrial extraction and purification, high-quality long read sequences are directly screened from the whole genome sequencing data for assembly, improving the recognition rate of mitochondrial genome sequences and achieving large-scale acquisition of mitochondrial genome information.
Smart Images

Figure CN120015122A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of genome sequencing, and more specifically, to a sequencing method, device, equipment and storage medium for mitochondrial genome sequences. Background Art
[0002] With the innovation and application of sequencing technology, new high-throughput sequencing technologies continue to emerge. After more than ten years of development and accumulation, nucleic acid sequencing technology, which is based on short-read second-generation sequencing, faces many challenges in various applications. In recent years, genome sequencing has entered the era of long-read third-generation sequencing. Compared with second-generation sequencing, third-generation sequencing has a longer read length, providing new technical means for the analysis of complex repetitive regions in genome sequences and the assembly of high-quality and high-integrity genomes. Mitochondria, as a type of organelle with unique DNA molecules and a complete genetic information transmission and expression system, participate in many life activities such as biological energy metabolism, signal transduction, and cell apoptosis. Therefore, sequencing and research on the mitochondrial genome are also indispensable.
[0003] Currently, sequencing and analyzing mitochondria alone requires isolating and purifying mitochondrial DNA and investing in sequencing costs, and the extraction process is difficult.
[0004] How to use the whole genome sequencing data of species to mine mitochondrial characteristic sequences and restore and obtain mitochondrial genomes is an issue that needs attention. Summary of the invention
[0005] In view of the above problems, the present application is proposed to provide a sequencing method, device, equipment and storage medium for mitochondrial genome sequences, so as to utilize the whole genome sequencing data of species to mine mitochondrial characteristic sequences and restore and obtain the mitochondrial genome.
[0006] In order to achieve the above objectives, the specific plan is proposed as follows:
[0007] A method for sequencing a mitochondrial genome sequence, comprising:
[0008] Obtain various read length sequencing sequences in the existing species read length database, and various read length sequencing sequences are all third-generation read length sequences;
[0009] By comparing various read-length sequencing sequences one by one with the reference genome sequence of the mitochondria of the target species, the alignment rate and coverage of each read-length sequencing sequence are determined;
[0010] Among various read length sequencing sequences, a read length sequencing sequence having an alignment rate higher than a first preset threshold and a coverage rate higher than a second preset threshold is determined as a high-quality sequence;
[0011] The various high-quality sequences are trimmed and assembled in sequence to obtain the determined genome sequence of the mitochondria of the target species.
[0012] Optionally, the trimming and assembling of the various high-quality sequences in sequence to obtain the determined genome sequence of the mitochondria of the target species comprises:
[0013] If all the high-quality sequences are obtained by sequencing in the high-quality HiFI mode, these sequences are directly used; otherwise, the high-quality sequences are stacked and base-corrected to obtain various corrected sequences, and each corrected sequence is trimmed to obtain a trimmed sequence;
[0014] The various trimmed sequences are assembled to obtain the determined genomic sequence of the mitochondria of the target species.
[0015] Optionally, the various high-quality sequences are stacked and base corrected to obtain various corrected sequences, including:
[0016] By stacking various high-quality sequences, determining sequencing error sites in the stacked various high-quality sequences;
[0017] According to the sequencing error sites, base correction is performed on the various stacked high-quality sequences to obtain various corrected sequences.
[0018] Optionally, trimming each correction sequence to obtain a trimmed sequence includes:
[0019] Determine the sequence-irrelevant base segments in each corrected sequence;
[0020] From each corrected sequence, the corrected sequence is trimmed to remove the sequence-irrelevant base segments in the corrected sequence to obtain a trimmed sequence.
[0021] Optionally, the assembling of various trimmed sequences to obtain the determined genome sequence of the mitochondria of the target species comprises:
[0022] According to each base of each trimmed sequence, each overlapping base segment is determined, and the splicing relationship between the trimmed sequences is determined through the overlapping base segments;
[0023] According to the splicing relationship between the trimmed sequences, and by splicing the trimmed sequences to which each overlapping base segment belongs, various trimmed sequences are assembled to obtain the measured genome sequence of the mitochondria of the target species.
[0024] Optionally, the method determines the alignment rate and coverage of each read length sequencing sequence by aligning each read length sequencing sequence with the reference genome sequence of the target species mitochondria one by one, including:
[0025] For each sequencing read sequence, each base segment that matches the sequencing read sequence is determined by comparing each base on the sequencing read sequence with each base of the reference genome sequence of the mitochondria of the target species;
[0026] Counting the total number of bases in each base segment of each read length sequencing sequence;
[0027] Determine the number of sequence bases in each read length sequencing sequence that match the reference genome sequence, and the number of corresponding sequencing sequence bases;
[0028] The number of sequence bases of the long read sequence that matches each read length sequencing sequence with the reference genome sequence is divided by the number of bases of the long read sequence sequence to obtain the alignment rate of the long read sequence sequence;
[0029] The total number of bases that match the reference genome sequence for each sequencing read length is divided by the number of bases in the reference sequence to obtain the coverage of the sequencing read length on the reference genome sequence.
[0030] Optionally, after sequentially trimming and assembling the various high-quality sequences to obtain the determined genome sequence of the mitochondria of the target species, the method further includes:
[0031] The determined genome sequence and the reference genome sequence are subjected to colinearity analysis to determine the similarity between the determined genome sequence and the reference genome sequence, so as to evaluate the homology and consistency between the target species corresponding to the mitochondria of the target species and the existing species.
[0032] A sequencing device for a mitochondrial genome sequence, comprising:
[0033] A read length sequencing sequence acquisition unit is used to obtain various read length sequencing sequences in the existing species read length database, and various read length sequencing sequences are all third-generation read length sequences;
[0034] An alignment unit is used to determine the alignment rate and coverage of each read-length sequencing sequence by aligning each read-length sequencing sequence with the reference genome sequence of the mitochondria of the target species one by one;
[0035] A high-quality sequence determination unit, used to determine, among various read-length sequencing sequences, a read-length sequencing sequence having an alignment rate higher than a first preset threshold and a coverage rate higher than a second preset threshold as a high-quality sequence;
[0036] The trimming and assembly unit is used to trim and assemble the various high-quality sequences in sequence to obtain the determined genome sequence of the mitochondria of the target species.
[0037] Optionally, the trimming and assembling unit comprises:
[0038] A first trimming unit is used for directly using the various high-quality sequences if they are obtained by sequencing in a high-quality HiFI mode;
[0039] A correction and trimming unit, used for stacking and base-correcting the various high-quality sequences to obtain various corrected sequences if the various high-quality sequences are not obtained by sequencing in the high-quality HiFI mode;
[0040] A second trimming unit is used to trim each correction sequence to obtain a trimmed sequence;
[0041] The assembly unit is used to assemble various trimmed sequences to obtain the determined genome sequence of the mitochondria of the target species.
[0042] Optionally, the correction and trimming unit includes:
[0043] A sequencing error site determination unit, used for determining the sequencing error sites in the stacked high-quality sequences by stacking the high-quality sequences if the high-quality sequences are not obtained by sequencing in the high-quality HiFI mode;
[0044] The base correction unit is used to perform base correction on the various stacked high-quality sequences according to the sequencing error sites to obtain various corrected sequences.
[0045] Optionally, the second trimming unit includes:
[0046] An irrelevant base segment determination unit, used for determining sequence irrelevant base segments in each corrected sequence;
[0047] The irrelevant base segment elimination unit is used to trim each corrected sequence from the corrected sequence to eliminate the sequence irrelevant base segments in the corrected sequence to obtain a trimmed sequence.
[0048] Optionally, the assembly unit includes:
[0049] An overlapping segment determination unit, used to determine each overlapping base segment according to each base of each trimmed sequence, and determine the splicing relationship between the trimmed sequences through the overlapping base segments;
[0050] The overlap splicing unit is used to assemble various trimmed sequences according to the splicing relationship between the trimmed sequences and by splicing the trimmed sequences to which each overlapped base segment belongs, so as to obtain the measured genome sequence of the mitochondria of the target species.
[0051] Optionally, the comparison unit includes:
[0052] A first alignment subunit is used for determining, for each read-length sequencing sequence, each base segment aligned and matched on the read-length sequencing sequence by aligning each base on the read-length sequencing sequence with each base of a reference genome sequence of the mitochondria of the target species;
[0053] A second alignment subunit is used to count the total number of bases in each base segment of each read length sequencing sequence;
[0054] The third alignment subunit is used to determine the number of sequence bases in each read length sequencing sequence that match the reference genome sequence, and the number of corresponding sequencing sequence bases;
[0055] The fourth alignment subunit is used to divide the number of sequence bases of the long read sequence that matches the reference genome sequence for each read length sequencing sequence by the number of bases of the long read sequence to obtain the alignment rate of the long read sequence; the fifth alignment subunit is used to divide the total number of bases that match the reference genome sequence for each read length sequencing sequence by the number of bases of the reference sequence to obtain the coverage rate of the long read sequence on the reference genome sequence.
[0056] Optionally, the device further comprises:
[0057] A colinearity analysis unit is used to perform colinearity analysis on the measured genome sequence and the reference genome sequence after the various high-quality sequences are trimmed and assembled in sequence to obtain the measured genome sequence of the target species' mitochondria, so as to determine the similarity between the measured genome sequence and the reference genome sequence, so as to evaluate the homology and consistency between the target species corresponding to the target species' mitochondria and the existing species.
[0058] A sequencing device for a mitochondrial genome sequence, comprising a memory and a processor;
[0059] The memory is used to store programs;
[0060] The processor is used to execute the program to implement the various steps of the sequencing method of the mitochondrial genome sequence as described above.
[0061] A storage medium stores a computer program, which, when executed by a processor, implements the various steps of the sequencing method of the mitochondrial genome sequence as described above.
[0062] By means of the above technical scheme, the present application obtains various read-length sequencing sequences in the existing species read-length database, and various read-length sequencing sequences are all third-generation read-length sequences. By comparing various read-length sequencing sequences with the reference genome sequence of the target species' mitochondria one by one, the alignment rate and coverage rate of each read-length sequencing sequence are determined. Among the various read-length sequencing sequences, the read-length sequencing sequences with an alignment rate higher than the first preset threshold and a coverage rate higher than the second preset threshold are determined as high-quality sequences, and various high-quality sequences are trimmed and assembled in turn to obtain the measured genome sequence of the target species' mitochondria. It can be seen that by directly extracting data from the current massive species public data and screening out the third-generation high-quality long-read sequences for assembly, the recognition rate of mitochondrial genome sequences can be improved, and there is no need for mitochondrial extraction, purification, and sequencing, that is, the mitochondrial genome information of existing published species can be obtained in large quantities. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0064] Figure 1 A schematic diagram of a process for sequencing a mitochondrial genome sequence provided in an embodiment of the present application;
[0065] Figure 2 A schematic diagram of a flow chart for calculating the alignment rate and coverage rate of a read length sequencing sequence provided in an embodiment of the present application;
[0066] Figure 3 The colinearity analysis results of the measured gene sequence of the avian animal and the reference genome sequence of the avian animal provided in the embodiments of the present application;
[0067] Figure 4 A schematic diagram of the structure of a device for sequencing mitochondrial genome sequences provided in an embodiment of the present application;
[0068] Figure 5 A schematic diagram of the structure of a device for sequencing mitochondrial genome sequences provided in an embodiment of the present application. DETAILED DESCRIPTION
[0069] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0070] The present application solution can be implemented based on a terminal with data processing capabilities, which can be a computer, server, cloud, etc.
[0071] Next, combine Figure 1 The sequencing method of the mitochondrial genome sequence of the present application may include the following steps:
[0072] Step S110, obtaining various read length sequencing sequences in the existing species read length database.
[0073] Specifically, various read-length sequencing sequences can all be third-generation read-length sequences, and can be published read-length sequences of species downloaded from public databases.
[0074] Step S120: Determine the alignment rate and coverage of each read length sequencing sequence by aligning each read length sequencing sequence with the reference genome sequence of the target species mitochondria one by one.
[0075] The target species mitochondria can be mitochondria of plant species or animal species. The alignment rate can represent the proportion of aligned bases in the read sequencing sequence, and the coverage rate can represent the proportion of aligned bases in the reference genome sequence.
[0076] In addition, the comparison length of each read length sequencing sequence can be determined by comparing various read length sequencing sequences one by one with the reference genome sequence of the mitochondria of the target species.
[0077] Specifically, the alignment length may represent the number of consecutive bases that match the alignment between the sequencing read sequence and the reference genome sequence.
[0078] Step S130: Among various read length sequencing sequences, a read length sequencing sequence having an alignment rate higher than a first preset threshold and a coverage rate higher than a second preset threshold is determined as a high-quality sequence.
[0079] Specifically, a high-quality sequence can represent a sequencing sequence with a read length that roughly matches the reference genome sequence after alignment.
[0080] Among them, the first preset threshold can represent the alignment rate standard for determining the read-length sequencing sequence as a high-quality sequence, and the optimal value of the first preset threshold is 20%. The second preset threshold can represent the coverage rate standard for determining the read-length sequencing sequence as a high-quality sequence, and the optimal value of the second preset threshold is 20%.
[0081] For example, assuming that the first preset threshold and the second preset threshold are both 20%, when the alignment rate of a certain read length sequencing sequence is higher than 20% and the coverage rate of the read length sequencing sequence is higher than 20%, the read length sequencing sequence is determined to be a high-quality sequence.
[0082] Step S140: various high-quality sequences are trimmed and assembled in sequence to obtain the determined genome sequence of the mitochondria of the target species.
[0083] Specifically, before the high-quality sequences are trimmed, base correction can also be performed on the high-quality sequences to improve the assembly quality of various high-quality sequences, so that the assembled measured genome sequence has a higher recognition rate.
[0084] The sequencing method of the mitochondrial genome sequence provided in the present embodiment obtains various read-length sequencing sequences in the existing species read-length database, and the various read-length sequencing sequences are all third-generation read-length sequences. By comparing the various read-length sequencing sequences with the reference genome sequence of the mitochondria of the target species one by one, the alignment rate and coverage rate of each read-length sequencing sequence are determined. Among the various read-length sequencing sequences, the read-length sequencing sequences with an alignment rate higher than the first preset threshold and a coverage rate higher than the second preset threshold are determined as high-quality sequences, and the various high-quality sequences are trimmed and assembled in turn to obtain the measured genome sequence of the mitochondria of the target species. It can be seen that by directly extracting data from the current massive species public data and screening out the third-generation high-quality long-read sequences for assembly, the recognition rate of the mitochondrial genome sequence can be improved, and the mitochondrial genome information of the existing published species can be obtained in large quantities without the need for mitochondrial extraction, purification, and sequencing.
[0085] In some embodiments of the present application, the process of sequentially trimming and assembling various high-quality sequences in step S140 to obtain the determined genome sequence of the mitochondria of the target species is introduced, and the process may include:
[0086] S1. If all high-quality sequences are obtained by sequencing in high-quality HiFI mode, these sequences are used directly; otherwise, various high-quality sequences are stacked and base-corrected to obtain various corrected sequences, and each corrected sequence is trimmed to obtain a trimmed sequence.
[0087] It is understandable that the sequencing mode of the high-quality HiFi mode itself is to perform multiple rounds of sequencing on the same sequence. The sequencing process can be automatically corrected, and the quality value of the obtained sequencing read length sequence is very high. This read length does not require a correction process, and these sequences can be directly used for assembly analysis. For example, the reads of conventional Nanopore sequencing and PacBio sequencing both need to be corrected, but PacBio sequencing is high-quality HiFi mode sequencing, so the read length sequence obtained by PacBio sequencing can skip the correction and trimming process, while the read length sequence obtained by conventional Nanopore sequencing needs to be corrected before entering the trimming process.
[0088] S2. Assemble various trimmed sequences to obtain the determined genome sequence of the mitochondria of the target species.
[0089] Specifically, the process of assembling various trimmed sequences can be to find the splicing points of various trimmed sequences, and splice various trimmed sequences through the splicing points, so that multiple trimmed sequences are spliced into a longer sequence to obtain the measured genome sequence of the mitochondria of the target species.
[0090] The sequencing method of the mitochondrial genome sequence provided in this embodiment determines whether various high-quality sequences are obtained by testing in the high-quality HiFi mode, thereby determining whether to perform correction processing on the various high-quality sequences. If the process of correcting the high-quality sequences is omitted, the sequencing speed of the mitochondrial genome sequence can be increased.
[0091] In some embodiments of the present application, the process of stacking and base correcting various high-quality sequences mentioned in the above embodiments to obtain various corrected sequences is introduced, and the process may include:
[0092] S1. Determine sequencing error sites in the stacked high-quality sequences by stacking various high-quality sequences.
[0093] It can be understood that the alignment rate and coverage of the high-quality sequence are respectively greater than the first preset threshold and the second preset threshold, and have high accuracy. In order to further improve the accuracy of the high-quality sequence, the high-quality sequence can be corrected.
[0094] Specifically, since sequencing errors are generated randomly, all high-quality sequences can be stacked together first during the correction process, and the sequencing error sites in the high-quality sequence stack can be determined based on the principle of minority obeys majority.
[0095] S2. According to the sequencing error sites, base correction is performed on various stacked high-quality sequences to obtain various corrected sequences.
[0096] The sequencing method of the mitochondrial genome sequence provided in this embodiment stacks various high-quality sequences, determines the sequencing error sites in the stacked various high-quality sequences, and performs base correction on the stacked various high-quality sequences according to the sequencing error sites to obtain various corrected sequences, so as to improve the accuracy of the high-quality sequences.
[0097] In some embodiments of the present application, the process of trimming each correction sequence mentioned in the above embodiment to obtain a trimmed sequence is introduced, and the process may include:
[0098] S1. Determine the sequence-irrelevant base segments in each corrected sequence.
[0099] Specifically, in the corrected sequence obtained after the correction process, there may be sequence-irrelevant base segments or suspicious regions, such as remaining SMRTbell aptamers, and the sequence-irrelevant base segments may reduce the assembly quality of the corrected sequence.
[0100] S2. Trimming each corrected sequence to remove sequence-irrelevant base segments in the corrected sequence to obtain a trimmed sequence.
[0101] It is understandable that, since sequence-irrelevant base segments may reduce the assembly quality of the corrected sequence, the sequence-irrelevant base segments need to be removed from the corrected sequence. Specifically, each corrected sequence can be trimmed, that is, the high-quality sequence portion is retained and the suspicious region, such as the remaining SMRTbell aptamer, is removed to obtain a trimmed sequence.
[0102] The sequencing method of the mitochondrial genome sequence provided in this embodiment determines the sequence-irrelevant base segments in each corrected sequence, and removes the sequence-irrelevant base segments in the corrected sequence from each corrected sequence to obtain a trimmed sequence. Compared with the corrected sequence, the trimmed sequence contains less useless or suspicious information, so the trimmed sequence has a higher assembly quality than the corrected sequence.
[0103] In some embodiments of the present application, the process of assembling various trimmed sequences mentioned in the above embodiments to obtain the determined genome sequence of the mitochondria of the target species is introduced, and the process may include:
[0104] S1. Determine each overlapping base segment according to each base of each trimmed sequence, and determine the splicing relationship between the trimmed sequences through the overlapping base segments.
[0105] Specifically, the overlapping base segment may represent a base segment consisting of a number of consecutive bases (not less than a preset number of bases 500 bp), wherein the preset number of bases may be customized, such as the preset number of bases is 1000.
[0106] It is understandable that the assembly stage can be to sort the obtained high-quality trimmed sequences, connect the overlapping base segments between the trimmed sequences, generate longer sequences, and thus complete splicing and assembly.
[0107] For example, trimmed sequence No. 1 is AATGCA, trimmed sequence No. 2 is TGCATA, and trimmed sequence No. 3 is CATAGG, then it can be determined that the overlapping base segments are TGCA and CATA. For the overlapping base segment TGCA, the two trimmed sequences it belongs to are trimmed sequence No. 1 and trimmed sequence No. 2. For the overlapping base segment CATA, the two trimmed sequences it belongs to are trimmed sequence No. 2 and trimmed sequence No. 3.
[0108] S2. According to the splicing relationship between the trimmed sequences, various trimmed sequences are assembled by splicing the trimmed sequences to which each overlapping base segment belongs, so as to obtain the determined genome sequence of the mitochondria of the target species.
[0109] For example, the trimmed sequence No. 1 is AATGCA, the trimmed sequence No. 2 is TGCATA, and the trimmed sequence No. 3 is CATAGG. Then for the overlapping base segment TGCA, the trimmed sequence No. 1 and the trimmed sequence No. 2 can be spliced, and for the overlapping base segment CATA, the trimmed sequence No. 2 and the trimmed sequence No. 3 can be spliced, so that the final measured genome sequence is AATGCATAGG.
[0110] The sequencing method of the mitochondrial genome sequence provided in this embodiment determines each overlapping base segment and the trimmed sequence corresponding to each overlapping base segment, and assembles various trimmed sequences by splicing the trimmed sequences belonging to each overlapping base segment to obtain the measured genome sequence of the mitochondria of the target species, thereby assembling multiple short reads into a mitochondrial genome sequence with a longer read.
[0111] In some embodiments of the present application, the above step S120, the process of comparing various read length sequencing sequences with the reference genome sequence of the target species mitochondria one by one to determine the alignment rate and coverage rate of each read length sequencing sequence is introduced, such as Figure 2 As shown, the process may include:
[0112] Step S210: for each read sequencing sequence, each base on the read sequencing sequence is compared with each base of the reference genome sequence of the target species' mitochondria to determine each base segment that matches the read sequencing sequence.
[0113] Specifically, the number of bases included in the base segment is not less than the minimum number of bases in the base segment. The minimum number of bases in the base segment can be customized in advance, such as the minimum number of bases in the base segment is 3.
[0114] For example, assuming that the minimum number of bases in a base segment is 3, the sequencing sequence of read length 1 is AACGCTTTA, and the reference genome sequence of the mitochondria of the target species is TAACCGGCCTTTCA, then when the sequencing sequence of read length 1 is aligned with the reference genome sequence of the mitochondria of the target species, the aligned base segments are AAC and CTT.
[0115] Step S220: Count the total number of bases in each base segment of each read length sequencing sequence.
[0116] For example, assuming that the minimum number of bases in a base segment is 3, the sequencing sequence of read length 1 is AACGCTTTA, the reference genome sequence of the target species' mitochondria is TAACCGGCCTTTCA, and the available base segments are AAC and CTT, then the total number of bases in each base segment is 6.
[0117] Step S230, determining the number of sequence bases in each read length sequencing sequence that matches the reference genome sequence, and the number of corresponding sequencing sequence bases.
[0118] For example, if the sequencing sequence of read length 1 is AACGCTTTA, then the number of sequencing bases of the sequencing sequence of read length 1 is 8, and the reference genome sequence of the mitochondria of the target species is TAACCGGCCTTTCA, then the number of sequencing bases of the reference genome sequence of the mitochondria of the target species is 14. The sequence bases that match the two are AAC and CTT, so the number of sequence bases that match the two is 6.
[0119] Step S240: Divide the number of sequence bases of the long read sequence that matches the reference genome sequence by the number of bases of the long read sequence to obtain the alignment rate of the long read sequence.
[0120] For example, assuming that the minimum number of bases in a base segment is 3, the sequencing sequence of read length 1 is AACGCTTTA, and the reference genome sequence of the target species' mitochondria is TAACCGGCCTTTCA. Since the total number of bases in the sequencing sequence of read length 1 is 8 and the total number of bases in each base segment is 6, the alignment rate of the sequencing sequence of read length 1 is 6 / 8=75%.
[0121] Step S250: divide the total number of bases of each read length sequencing sequence that matches the reference genome sequence by the number of bases in the reference sequence to obtain the coverage of the read length sequencing sequence on the reference genome sequence.
[0122] For example, assuming that the minimum number of bases in a base segment is 3, the sequencing sequence of read length 1 is AACGCTTTA, and the reference genome sequence of the target species' mitochondria is TAACCGGCCTTTCA. Since the number of bases in the sequencing sequence is 14 and the total number of bases in each base segment is 6, the coverage of the sequencing sequence of read length 1 on the reference genome sequence is 6 / 14≈42.9%.
[0123] The animal and plant mitochondrial genomes assembled according to the method of the present invention all have a complete circular sequence, no holes, and no unknown sequences. Compared with many currently public mitochondrial genomes, the effect is better and can be obtained in batches quickly. It provides a good data basis for the study of species mitochondria. The specific assembly situation is described in the attached table.
[0124] Appendix Description
[0125] Table 1 shows the statistics of the assembly indicators of the obtained animal mitochondrial genomes;
[0126] Table 2 shows the assembly index statistics of the obtained plant mitochondrial genomes.
[0127] Table 1
[0128]
[0129] Note: The total length of the assembled mitochondria of this animal is 16,888bp, which is assembled into a complete circular sequence.
[0130] No N and gap. GC content is 47.47%.
[0131] Table 2
[0132]
[0133] Note: The total length of the plant mitochondrial assembly is 56,416bp, which is assembled into a complete circular sequence.
[0134] No N and gap. GC content is 27.90%.
[0135] Considering the quality evaluation of the determined genome sequence to analyze the gap with the reference genome sequence, in some embodiments of the present application, the sequencing method of the mitochondrial genome sequence provided may also include a process of analyzing the determined genome sequence. Specifically, the process may include:
[0136] The measured genome sequence and the reference genome sequence were subjected to colinearity analysis to determine the similarity between the measured genome sequence and the reference genome sequence, so as to evaluate the homology and consistency between the target species corresponding to the mitochondria of the target species and the existing species.
[0137] Specifically, each base in the measured genome sequence can be compared with each base in the reference genome sequence at their respective corresponding positions to determine the degree of similarity between the measured genome sequence and the reference genome sequence, thereby determining the similarity between the measured genome sequence and the reference genome sequence.
[0138] like Figure 3 As shown, Figure 3 The figure shows the colinearity between the mitochondrial genome sequence of birds and the reference mitochondrial genome sequence. The horizontal axis is the corresponding position of the reference mitochondrial genome sequence of birds, and the vertical axis is the sequence position of the measured mitochondrial genome of birds. It can be seen that the two genome sequences have good colinearity and a high base matching success rate, so the assembly accuracy and completeness are good, and the measured genome sequence for birds can be analyzed to have high homology and consistency with known birds.
[0139] The following is a description of the apparatus for sequencing mitochondrial genome sequences provided in an embodiment of the present application. The apparatus for sequencing mitochondrial genome sequences described below and the method for sequencing mitochondrial genome sequences described above can be referenced to each other.
[0140] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of a device for sequencing mitochondrial genome sequences disclosed in an embodiment of the present application.
[0141] like Figure 4 As shown, the device may include:
[0142] A read length sequencing sequence acquisition unit 11 is used to acquire various read length sequencing sequences in an existing species read length database, wherein the various read length sequencing sequences are all third-generation read length sequences;
[0143] The comparison unit 12 is used to determine the comparison rate and coverage rate of each read length sequencing sequence by comparing each read length sequencing sequence with the reference genome sequence of the mitochondria of the target species one by one;
[0144] A high-quality sequence determination unit 13 is used to determine, among various read-length sequencing sequences, a read-length sequencing sequence having an alignment rate higher than a first preset threshold and a coverage rate higher than a second preset threshold as a high-quality sequence;
[0145] The trimming and assembling unit 14 is used to trim and assemble the various high-quality sequences in sequence to obtain the determined genome sequence of the mitochondria of the target species.
[0146] Optionally, the trimming and assembling unit comprises:
[0147] A first trimming unit is used for trimming each of the high-quality sequences to obtain a trimmed sequence if all the high-quality sequences are obtained by sequencing in a high-quality HiFI mode;
[0148] A correction and trimming unit, for stacking and base-correcting the various high-quality sequences to obtain various corrected sequences if the various high-quality sequences are not obtained by sequencing in the high-quality HiFI mode;
[0149] A second trimming unit is used to trim each correction sequence to obtain a trimmed sequence;
[0150] The assembly unit is used to assemble various trimmed sequences to obtain the determined genome sequence of the mitochondria of the target species.
[0151] Optionally, the correction and trimming unit includes:
[0152] A sequencing error site determination unit, used for determining the sequencing error sites in the stacked high-quality sequences by stacking the high-quality sequences if the high-quality sequences are not obtained by sequencing in the high-quality HiFI mode;
[0153] The base correction unit is used to perform base correction on the various stacked high-quality sequences according to the sequencing error sites to obtain various corrected sequences.
[0154] Optionally, the second trimming unit includes:
[0155] An irrelevant base segment determination unit, used for determining sequence irrelevant base segments in each corrected sequence;
[0156] The irrelevant base segment elimination unit is used to trim each corrected sequence from the corrected sequence to eliminate the sequence irrelevant base segments in the corrected sequence to obtain a trimmed sequence.
[0157] Optionally, the assembly unit includes:
[0158] An overlapping segment determination unit, used to determine each overlapping base segment according to each base of each trimmed sequence, and determine the splicing relationship between the trimmed sequences through the overlapping base segments;
[0159] The overlap splicing unit is used to assemble various trimmed sequences according to the splicing relationship between the trimmed sequences and by splicing the trimmed sequences to which each overlapped base segment belongs, so as to obtain the measured genome sequence of the mitochondria of the target species.
[0160] Optionally, the comparison unit includes:
[0161] A first alignment subunit is used for determining, for each read-length sequencing sequence, each base segment aligned and matched on the read-length sequencing sequence by aligning each base on the read-length sequencing sequence with each base of a reference genome sequence of the mitochondria of the target species;
[0162] A second alignment subunit is used to count the total number of bases in each base segment of each read length sequencing sequence;
[0163] The third alignment subunit is used to determine the number of sequence bases in each read length sequencing sequence that match the reference genome sequence, and the number of corresponding sequencing sequence bases;
[0164] A fourth alignment subunit is used to divide the number of sequence bases of the long read sequence that matches each read length sequencing sequence with the reference genome sequence by the number of bases of the read length sequencing sequence to obtain an alignment rate of the read length sequencing sequence;
[0165] The fifth alignment subunit is used to divide the total number of bases of each read-length sequencing sequence that matches the reference genome sequence by the number of bases in the reference sequence to obtain the coverage of the read-length sequencing sequence on the reference genome sequence.
[0166] Optionally, the device further comprises:
[0167] A colinearity analysis unit is used to perform colinearity analysis on the measured genome sequence and the reference genome sequence after the various high-quality sequences are trimmed and assembled in sequence to obtain the measured genome sequence of the target species' mitochondria, so as to determine the similarity between the measured genome sequence and the reference genome sequence, so as to evaluate the homology and consistency between the target species corresponding to the target species' mitochondria and the existing species.
[0168] The apparatus for sequencing mitochondrial genome sequences provided in the embodiments of the present application can be applied to devices for sequencing mitochondrial genome sequences, such as terminals: mobile phones, computers, etc. Optionally, Figure 5 The hardware structure block diagram of the sequencing equipment of the mitochondrial genome sequence is shown, referring to Figure 5 , the hardware structure of the device for sequencing the mitochondrial genome sequence may include: at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4;
[0169] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 communicate with each other through the communication bus 4;
[0170] The processor 1 may be a central processing unit CPU, or an application-specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0171] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;
[0172] The memory stores a program, and the processor can call the program stored in the memory, wherein the program is used to:
[0173] Obtain various read length sequencing sequences in the existing species read length database, and various read length sequencing sequences are all third-generation read length sequences;
[0174] By comparing various read-length sequencing sequences one by one with the reference genome sequence of the mitochondria of the target species, the alignment rate and coverage of each read-length sequencing sequence are determined;
[0175] Among various read length sequencing sequences, a read length sequencing sequence having an alignment rate higher than a first preset threshold and a coverage rate higher than a second preset threshold is determined as a high-quality sequence;
[0176] The various high-quality sequences are trimmed and assembled in sequence to obtain the determined genome sequence of the mitochondria of the target species.
[0177] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0178] The embodiment of the present application further provides a storage medium, which may store a program suitable for execution by a processor, wherein the program is used to:
[0179] Obtain various read length sequencing sequences in the existing species read length database, and various read length sequencing sequences are all third-generation read length sequences;
[0180] By comparing various read-length sequencing sequences one by one with the reference genome sequence of the mitochondria of the target species, the alignment rate and coverage of each read-length sequencing sequence are determined;
[0181] Among various read length sequencing sequences, a read length sequencing sequence having an alignment rate higher than a first preset threshold and a coverage rate higher than a second preset threshold is determined as a high-quality sequence;
[0182] The various high-quality sequences are trimmed and assembled in sequence to obtain the determined genome sequence of the mitochondria of the target species.
[0183] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0184] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0185] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.
[0186] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for sequencing a mitochondrial genome sequence, characterized in that: include: Obtain various read length sequencing sequences in the existing species read length database, and various read length sequencing sequences are all third-generation read length sequences; By comparing various read-length sequencing sequences one by one with the reference genome sequence of the mitochondria of the target species, the alignment rate and coverage of each read-length sequencing sequence are determined; Among various read length sequencing sequences, a read length sequencing sequence having an alignment rate higher than a first preset threshold and a coverage rate higher than a second preset threshold is determined as a high-quality sequence; The various high-quality sequences are trimmed and assembled in sequence to obtain the determined genome sequence of the mitochondria of the target species.
2. The method according to claim 1, characterized in that The method of sequentially trimming and assembling the various high-quality sequences to obtain the determined genome sequence of the mitochondria of the target species comprises: If all the high-quality sequences are obtained by sequencing in high-quality HiFI mode, these sequences are directly used; otherwise, the various sequences are stacked and base-corrected to obtain various corrected sequences, and each corrected sequence is trimmed to obtain a trimmed sequence; The various trimmed sequences are assembled to obtain the determined genomic sequence of the mitochondria of the target species.
3. The method according to claim 2, characterized in that The various high-quality sequences are stacked and base corrected to obtain various corrected sequences, including: By stacking various high-quality sequences, determining sequencing error sites in the stacked various high-quality sequences; According to the sequencing error sites, base correction is performed on the various stacked high-quality sequences to obtain various corrected sequences.
4. The method according to claim 2, characterized in that: The step of trimming each correction sequence to obtain a trimmed sequence comprises: Determine the sequence-irrelevant base segments in each corrected sequence; From each corrected sequence, the corrected sequence is trimmed to remove the sequence-irrelevant base segments in the corrected sequence to obtain a trimmed sequence.
5. The method according to claim 2, characterized in that: The step of assembling various trimmed sequences to obtain the determined genome sequence of the mitochondria of the target species comprises: According to each base of each trimmed sequence, each overlapping base segment is determined, and the splicing relationship between the trimmed sequences is determined through the overlapping base segments; According to the splicing relationship between the trimmed sequences, and by splicing the trimmed sequences to which each overlapping base segment belongs, various trimmed sequences are assembled to obtain the measured genome sequence of the mitochondria of the target species.
6. The method according to claim 1, characterized in that The method determines the alignment rate and coverage of each read length sequencing sequence by comparing each read length sequencing sequence with the reference genome sequence of the target species mitochondria one by one, including: For each sequencing read sequence, each base segment that matches the sequencing read sequence is determined by comparing each base on the sequencing read sequence with each base of the reference genome sequence of the mitochondria of the target species; Counting the total number of bases in each base segment of each read length sequencing sequence; Determine the number of sequence bases in each read length sequencing sequence that match the reference genome sequence, and the number of corresponding sequencing sequence bases; The number of sequence bases of the long read sequence that matches each read length sequencing sequence with the reference genome sequence is divided by the number of bases of the long read sequence sequence to obtain the alignment rate of the long read sequence sequence; The total number of bases that match the reference genome sequence for each sequencing read length is divided by the number of bases in the reference sequence to obtain the coverage of the sequencing read length on the reference genome sequence.
7. The method according to claim 1, characterized in that After the various high-quality sequences are sequentially trimmed and assembled to obtain the determined genome sequence of the mitochondria of the target species, the method further includes: The determined genome sequence and the reference genome sequence are subjected to colinearity analysis to determine the similarity between the determined genome sequence and the reference genome sequence, so as to evaluate the homology and consistency between the target species corresponding to the mitochondria of the target species and the existing species.
8. A mitochondrial genome sequencing device, characterized in that: include: A read length sequencing sequence acquisition unit is used to obtain various read length sequencing sequences in the existing species read length database, and various read length sequencing sequences are all third-generation read length sequences; An alignment unit is used to determine the alignment rate and coverage of each read-length sequencing sequence by aligning each read-length sequencing sequence with the reference genome sequence of the mitochondria of the target species one by one; A high-quality sequence determination unit, used to determine, among various read-length sequencing sequences, a read-length sequencing sequence having an alignment rate higher than a first preset threshold and a coverage rate higher than a second preset threshold as a high-quality sequence; The trimming and assembly unit is used to trim and assemble the various high-quality sequences in sequence to obtain the determined genome sequence of the mitochondria of the target species.
9. A mitochondrial genome sequencing device, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the method for sequencing a mitochondrial genome sequence as claimed in any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method for sequencing a mitochondrial genome sequence according to any one of claims 1 to 7 is implemented.