A method, apparatus, and storage medium for correcting sequencing mutations of mitochondrial genomes

By combining a whitelist of known mutations with multiple filtering indicators, the problem of false positive mutations in mitochondrial genome sequencing was solved, achieving efficient and accurate mutation correction and detection for single samples.

CN116825193BActive Publication Date: 2026-05-19AEGICARE (SHENZHEN) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AEGICARE (SHENZHEN) TECH CO LTD
Filing Date
2023-07-17
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing mitochondrial genome sequencing mutation detection methods cannot effectively remove false-positive mutations caused by sequencing errors and improper experimental methods, especially in single-sample cases, resulting in poor accuracy of analysis results.

Method used

A whitelist filtering method for known mutations was used, combined with 11 filtering indicators (mutation frequency, average read position, average base quality, root mean square of read alignment quality, etc.) to correct mitochondrial genome sequencing mutations, retaining known pathogenic mutations and removing false positives.

Benefits of technology

It improves the accuracy and reliability of mitochondrial genome sequencing mutation detection, reduces the workload of downstream analysis, and is suitable for single-sample calibration, without being limited by the number of samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116825193B_ABST
    Figure CN116825193B_ABST
Patent Text Reader

Abstract

The application discloses a method, device and storage medium for correcting mitochondrial genome sequencing mutations. The mitochondrial genome sequencing mutation correction method of the application comprises the following steps: firstly, retaining possible pathogenic mutations through a known mutation whitelist; and then, removing false positives caused by experimental errors and sequencing errors by using 11 filtering indexes, such as the average value of the position of a read in which the mutation is located, the chain distribution deviation of the mutation and the average base quality value of the mutation, so as to obtain a high-credibility mutation set. The correction method of the application can reduce the workload of mutation interpretation and improve the accuracy of analysis results. Moreover, the correction method of the application is not limited by the number of samples and can be used for correcting single samples or a small number of samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of mitochondrial genome mutation detection technology, and in particular to a method, apparatus and storage medium for correcting mitochondrial genome sequencing mutations. Background Technology

[0002] Mitochondrial genome DNA can be captured and enriched using probe hybridization capture technology from human blood or other tissue samples, followed by high-throughput sequencing to obtain the mitochondrial genome sequence. After quality control of the sequencing data to remove repetitive sequences, sequence alignment software is used for alignment, and mutation detection software is used to detect mutations in the mitochondrial genome, thus obtaining mitochondrial genome sequencing mutation data. However, studies show that a significant portion of these mutations are due to sequencing errors, improper experimental methods, or systematic errors in alignment algorithms; that is, a considerable portion of these mutations are not actually present in the human mitochondrial genome. These mutations increase the burden on downstream analysis and interfere with the results of mitochondrial genome sequencing mutation data detection. Therefore, it is necessary to correct and filter these mutations to reduce the workload of downstream analysis and improve the accuracy of the analytical results.

[0003] Currently, point mutation correction and filtering methods mainly use machine learning to train models or use hard metrics for filtering. Hard metric filtering includes things like mutation quality scores and mutation chain distribution bias. The main method that applies machine learning models is VQSR (Variant Quality Score Recalibration).

[0004] The core principle of VQSR is to construct a classifier that distinguishes between true and false mutations using a Gaussian Mixture Model (GMM). This model is constructed using sites consistent with a known mutation dataset, assigned corresponding confidence weights for training. These known and rigorously validated mutations are treated as true mutations to train the GMM that distinguishes them. Next, the input data is scored, and the mutations with the lowest scores are used as the set of false mutations to construct a GMM that distinguishes them. Finally, the two constructed GMM models—the true mutation GMM and the false mutation GMM—are used simultaneously to score the mutations and determine whether the input mutation belongs to the true or false mutation category.

[0005] To ensure reliable classification results, VQSR has a minimum number of usable sites required for model training. For example, the number of true and false variants available for training must exceed 5000; otherwise, a usable model cannot be trained. While whole-genome sequencing of humans generally meets this minimum requirement, it is insufficient for single-sample human exome sequencing, mitochondrial genome sequencing, or other small genome sequencing. Therefore, single-sample mitochondrial genome sequencing cannot be used for mutation correction and filtering with VQSR. Even with multi-sample mutation correction and filtering, the required sample size is extremely large, severely limiting the efficiency of VQSR mutation correction and filtering.

[0006] In addition, some studies have shown that indicators such as mutation quality value or mutation chain distribution deviation can be used for filtering; however, existing correction methods based on these hard indicators generally have shortcomings such as poor accuracy and many false positives.

[0007] Therefore, how to more effectively and accurately remove false positives caused by experimental and sequencing errors and obtain a more reliable mutation set remains a significant challenge in the analysis of mitochondrial genome sequencing mutation data. Summary of the Invention

[0008] The purpose of this application is to provide a novel method, apparatus, and storage medium for correcting mutations in mitochondrial genome sequencing.

[0009] To achieve the above objectives, this application adopts the following technical solution:

[0010] The first aspect of this application discloses a method for correcting mitochondrial genome sequencing mutations, comprising: comparing the mitochondrial genome sequencing mutations to be corrected with a known mutation whitelist; retaining the mitochondrial genome sequencing mutations to be corrected that appear in the known mutation whitelist as a first mutation set; filtering the mitochondrial genome sequencing mutations to be corrected that do not appear in the known mutation whitelist based on 1) mutation frequency, 2) the average position of the read containing the mutation, 3) whether the mutations are all located at the same position, 4) the average base quality value of the mutation, 5) the root mean square of the read alignment quality value, 6) the quality value per unit depth of the mutation, 7) the average number of mismatches in the read alignment, 8) the signal-to-noise ratio, 9) the rank-sum test of the base quality value of the heterozygous site read, 10) the rank-sum test of the position of the heterozygous site mutation, and 11) the strand distribution bias of the mutation; and using the remaining mutations after filtering as a second mutation set; merging the first mutation set and the second mutation set to obtain the corrected mitochondrial genome sequencing mutations; the known mutation whitelist is a verified set of mutations known to be pathogenic, potentially pathogenic, or of unknown significance.

[0011] The verified set of known mutations can be a set of mutations in a database, such as the mitochondrial-related disease database MSeqDR or the clinical disease and phenotype-related mutation database ClinVar, or other sets of known mutations that have been rigorously verified.

[0012] It should be noted that the correction method in this application first filters some mutations using a known mutation whitelist, and only analyzes and filters mutations not appearing in the known mutation whitelist, reducing the amount of data and cost of subsequent analysis. Furthermore, this application comprehensively considers various factors and employs 11 filtering indicators to remove false positives caused by experimental and sequencing errors, obtaining the final true mutation set. Moreover, the correction method in this application does not require consideration of the number of sites for mutation filtering, nor does it require filtering multiple samples simultaneously; mutation filtering and correction can be performed on a single sample. This application has more mutation filtering indicators, and the constructed known mutation whitelist ensures that no meaningful mutations are missed, thus making the results more reliable. It can be understood that the correction method in this application can be used not only for filtering and correcting mutation data in human mitochondrial genomes, but also for correcting mutation data in mitochondrial genomes of other species such as mice and zebrafish, or small genomes such as fruit fly genomes. In addition, while the correction method in this application is well-suited for filtering and correcting single samples, it can naturally also be used for filtering and correcting mutations in multiple samples, especially for multiple samples with a small number of samples.

[0013] It should also be noted that for human whole-genome or whole-exome sequencing mutation data, the construction of a whitelist is not significant because the number of mutations is too large for the whitelist to encompass. However, for small genomes like the mitochondrial genome, a whitelist of mutations ensures that the vast majority of identified potentially pathogenic mutations are not missed, making the overall results more accurate and reliable. Furthermore, it reduces the workload of downstream analysis and improves the accuracy of the results. Therefore, the correction method proposed in this application is more suitable for sequencing mutation correction in small genomes such as the mitochondrial genome; while other existing correction methods are more suitable for larger genomes.

[0014] In one implementation of this application, filtering mutations based on 1) mutation frequency includes calculating the mutation frequency of all mutations, and for the same mutation site, retaining only the mutation with the highest mutation frequency at the same mutation site.

[0015] In one implementation of this application, filtering mutations based on the average position of the read segment containing the mutation (2) includes: counting the positions of the same mutation site in different read segments, calculating the average position of these positions, and filtering out mutations whose mutation frequency is less than a mutation frequency threshold and whose average position is less than a position threshold.

[0016] In one implementation of this application, filtering mutations based on whether the mutations are all located at the same position includes filtering out mutations where the mutation site is located at the same position in different reads for the same mutation site.

[0017] In one implementation of this application, filtering mutations based on the average base quality value of the mutation (4) includes calculating the average base quality value of the mutation and filtering out mutations whose average base quality value is lower than the base quality value threshold.

[0018] In one implementation of this application, filtering mutations based on the root mean square of the read segment alignment quality value (5) includes calculating the root mean square of the alignment quality values ​​of all read segments containing the mutation, and filtering out mutations whose root mean square of the alignment quality value is less than the root mean square threshold.

[0019] In one implementation of this application, filtering mutations based on the unit depth quality value of the mutation (6) includes calculating the unit depth quality value of the mutation and filtering out mutations whose unit depth quality value is less than a unit depth quality value threshold.

[0020] In one implementation of this application, filtering mutations based on the average number of mismatches in the read segment (7) includes calculating the average number of mismatches in the read segment where the mutation is located, and filtering out mutations whose average number of mismatches is greater than the threshold number of mismatches.

[0021] In one implementation of this application, filtering mutations based on the signal-to-noise ratio (SNR) includes: counting the number of read segments in which the mutation is located where the base quality value is greater than the base quality value threshold and the number of read segments where the base quality value is less than the base quality value threshold; using the quotient of the number of read segments where the base quality value is greater than the base quality value threshold divided by the number of read segments where the base quality value is less than the base quality value threshold as the SNR; and filtering out mutations with an SNR value less than the SNR threshold.

[0022] In one implementation of this application, filtering mutations based on the rank sum test of the base quality values ​​of the heterozygous site read segment (9) includes performing a rank sum test on the base quality values ​​of the read segment containing the mutated base and the base quality values ​​of the read segment containing the reference base, and filtering out mutations whose rank sum is less than the rank sum threshold of the quality value.

[0023] In one implementation of this application, filtering mutations based on the rank sum test of the mutated site position (10) includes performing a rank sum test on the position of the mutated base in the read segment and the position of the reference base in the read segment, and filtering out mutations whose rank sum is less than the position rank sum threshold.

[0024] In one implementation of this application, the formula for calculating the mutation frequency is:

[0025] AF = ALT / total number of reads covering this site.

[0026] or,

[0027] AF = ALT / (REF + all ALT values)

[0028] Where AF is the mutation frequency, ALT is the number of mutated bases at the same site, REF is the number of reference bases at the same site, and all ALT refers to the ALT of all mutations when there are multiple mutations at the same position.

[0029] In one implementation of this application, the mutation frequency threshold is 0.35 and the location threshold is 8.

[0030] In one implementation of this application, the base quality value threshold is 22.5.

[0031] In one implementation of this application, the root mean square threshold is 40.

[0032] In one implementation of this application, the unit depth quality value of the mutation is calculated by dividing the quality value of the mutation by the sum of the depths of all samples containing the mutation, and using the quotient as the unit depth quality value.

[0033] In one implementation of this application, the unit depth quality value threshold is 2.

[0034] In one implementation of this application, the threshold for the number of mismatches is 5.25.

[0035] In one implementation of this application, when filtering mutations based on the signal-to-noise ratio (8), if the number of read segments with base quality values ​​less than the base quality value threshold is 0, then the value is assigned to 0.5.

[0036] In one implementation of this application, the signal-to-noise ratio threshold is 1.5.

[0037] In one implementation of this application, the rank sum threshold of the quality value is -12.5.

[0038] In one implementation of this application, the position rank sum threshold is -8.

[0039] It should be noted that the specific filtering thresholds used above can be adjusted according to experimental conditions and sequencing platforms. For example, the filtering threshold for the average base quality value of mutations can be adjusted upwards or downwards.

[0040] In one implementation of this application, filtering mutations based on the chain distribution bias of mutations (11) includes: using only reads with an average base quality value greater than 22.5; calculating the mutation frequency, denoted as HiAF; calculating the number of reference bases in the positive strand, denoted as SRF; the number of reference bases in the negative strand, denoted as SRR; the number of mutated bases in the positive strand, denoted as SAF; and the number of mutated bases in the negative strand, denoted as SAR; for reference bases, when either SRF or SRR is 0, and SRF + SRR <= 12, the reference base bias value RefBias is 0; calculating SRF / (SRF + SRR), SRR / (SRF + SRR), and SRR / (SRF + SRR), etc. +SRR), when all the above values ​​are greater than 0.05, and both SRF and SRR are greater than 2, the reference base bias value RefBias is 2, otherwise it is 1; for mutant bases, when either SAF or SAR is 0, and SAF+SAR<=12, the mutant base bias value AltBias is equal to 0; calculate SAF / (SAF+SAR) and SAR / (SAF+SAR), when both the above values ​​are greater than 0.05, and both SAF and SAR are greater than 2, the mutant base bias value AltBias is 2, otherwise it is 1; perform Fisher's exact test on SRF, SRR, SAF and SAR to obtain the p value and odd ratio; when HiAF is less than 0.25, RefBias is 2, and AltBias is 1, the p value is less than or equal to 0.01, the odd ratio is greater than 5 or equal to 0, and the mutation length is less than 100bp, the mutation is filtered out.

[0041] In one implementation of this application, mitochondrial mutation data from the mitochondrial-related disease database MSeqDR and the clinical disease and phenotype-related mutation database ClinVar are used to retain pathogenic, potentially pathogenic, or undetermined mutations, and then merge and remove duplicates to form a whitelist of known mutations.

[0042] It should be noted that the known mutation whitelist in this application was obtained by merging and removing duplicates from mitochondrial mutation data from the mitochondrial-related disease database MSeqDR and the clinical disease and phenotype-related mutation database ClinVar. It is understood that, in principle, other rigorously validated mutations can also be added to the known mutation whitelist in this application, not just those from the MSeqDR and ClinVar databases. Therefore, the known mutation whitelist in this application can be continuously adjusted based on existing technology.

[0043] The second aspect of this application discloses an apparatus for correcting mitochondrial genome sequencing mutations, including a whitelist comparison module, a mutation filtering module, and a corrected data acquisition module; wherein, the whitelist comparison module includes a method for comparing the mitochondrial genome sequencing mutation to be corrected with a known mutation whitelist, retaining the mitochondrial genome sequencing mutations to be corrected that appear in the known mutation whitelist as a first mutation set; the mutation filtering module includes a method for filtering mitochondrial genome sequencing mutations to be corrected that do not appear in the known mutation whitelist, based on 1) mutation frequency, 2) the average position of the read containing the mutation, and 3) whether the mutations are all located at the same position. The mutations are filtered using the following methods: 4) average base quality of the mutation, 5) root mean square of read alignment quality, 6) quality per unit depth of the mutation, 7) average number of mismatches in read alignment, 8) signal-to-noise ratio, 9) rank-sum test of base quality of heterozygous site reads, 10) rank-sum test of mutation position at heterozygous site, and 11) chain distribution bias of the mutation. The remaining mutations after filtering are used as the second mutation set. The corrected data acquisition module includes a module for merging the first and second mutation sets to obtain the corrected mitochondrial genome sequencing mutations. The known mutation whitelist is a set of verified mutations known to be pathogenic, potentially pathogenic, or of unknown significance.

[0044] It should be noted that the device for correcting mitochondrial genome sequencing mutations in this application is actually implemented by each module separately to achieve the method for correcting mitochondrial genome sequencing mutations in this application. Therefore, the specific limitations of each module can be referred to the method for correcting mitochondrial genome sequencing mutations in this application. For example, how to filter the 11 filtering indicators and the specific thresholds in each filtering indicator can be referred to the correction method in this application, and will not be elaborated here.

[0045] A third aspect of this application discloses an apparatus for correcting mitochondrial genome sequencing mutations, the apparatus comprising a memory and a processor; the memory including a program for storing a program; and the processor including a method for implementing the method of correcting mitochondrial genome sequencing mutations of this application by executing the program stored in the memory.

[0046] The fourth aspect of this application discloses a computer-readable storage medium storing a program that can be executed by a processor to implement the method of correcting mitochondrial genome sequencing mutations of this application.

[0047] Due to the adoption of the above technical solutions, the beneficial effects of this application are as follows:

[0048] The mitochondrial genome sequencing mutation correction method of this application first retains potentially pathogenic mutations through a known mutation whitelist to avoid missing known potentially pathogenic mutations. Then, it removes false-positive mutations caused by improper experimental methods, sequencing errors, or systematic errors in alignment software through 11 filtering indicators, thereby improving the reliability and authenticity of the final mutation set. This correction method reduces the workload of mutation interpretation while improving the accuracy of the analysis results; furthermore, this correction method is not limited by the number of samples and can perform correction on a single sample. Attached Figure Description

[0049] Figure 1 This is a flowchart of the mitochondrial genome sequencing mutation correction method in the embodiments of this application;

[0050] Figure 2 This is a structural block diagram of the mitochondrial genome sequencing mutation correction device in the embodiments of this application;

[0051] Figure 3 This is a schematic diagram of the VCF file in the embodiments of this application;

[0052] Figure 4 This is a schematic diagram of the BAM file in an embodiment of this application. Detailed Implementation

[0053] The present application will now be described in further detail with reference to specific embodiments and accompanying drawings. In the following embodiments, many details are described to facilitate a better understanding of the present application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other devices, materials, or methods. In some cases, certain operations related to the present application are not shown or described in the specification to avoid obscuring the core parts of the application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; a complete understanding of the related operations can be obtained from the description in the specification and general technical knowledge in the art.

[0054] Existing mitochondrial genome sequencing mutation correction methods, such as VQSR, cannot correct and filter single samples, and are also difficult to apply when the sample size is small. Furthermore, although some studies have used mutation quality values ​​or mutation strand distribution bias for filtering, overall, existing filtering and correction methods suffer from poor accuracy and a high number of false positives.

[0055] Based on the above research and understanding, this application creatively proposes a new method for correcting mitochondrial genome sequencing mutations, such as... Figure 1As shown, it includes whitelist comparison step 11, mutation filtering step 12, and corrected data acquisition step 13.

[0056] Whitelist comparison step 11 includes comparing the mitochondrial genome sequencing mutations to be corrected with the known mutation whitelist, and retaining the mitochondrial genome sequencing mutations to be corrected that appear in the known mutation whitelist as the first mutation set.

[0057] The mutation filtering step 12 includes filtering mitochondrial genome sequencing mutations that do not appear in the known mutation whitelist based on 1) mutation frequency, 2) average position of the read containing the mutation, 3) whether all mutations are located in the same position, 4) average base quality value of the mutation, 5) root mean square of read alignment quality value, 6) quality value per unit depth of the mutation, 7) average number of mismatches in read alignment, 8) signal-to-noise ratio, 9) rank-sum test of base quality value of heterozygous site reads, 10) rank-sum test of the position of heterozygous site mutations, and 11) chain distribution bias of the mutation. The remaining mutations after filtering are used as the second mutation set.

[0058] Step 13, the data acquisition step after correction, includes merging the first mutation set and the second mutation set, i.e., obtaining the corrected mitochondrial genome sequencing mutations. In this application, the known mutation whitelist is a verified set of mutations known to be pathogenic, potentially pathogenic, or of unknown significance.

[0059] The key to this application is:

[0060] 1) Obtain known mutation data of the human mitochondrial genome, perform format conversion and remove duplicates from the data, and construct a whitelist of known mutations.

[0061] 2) Mutations in the input sample data that appear in the whitelist are retained. For mutations that do not appear in the whitelist, 11 filtering indicators, including the average position of the read containing the mutation, the strand distribution deviation of the mutation, and the average base quality value of the mutation, are used to remove false positives caused by experimental errors and sequencing errors, thus obtaining the final true mutation set.

[0062] In one implementation of this application, the method for correcting mitochondrial genome sequencing mutations specifically includes:

[0063] 1. Obtain mitochondrial mutation data from MSeqDR (Mitochondrial Related Diseases Database) and ClinVar (Clinical Disease and Phenotype Related Mutation Database), retain pathogenic, potentially pathogenic, or undetermined mutations, merge and remove duplicates to form a whitelist of known mutations.

[0064] 2. If the sample mutation appears in the known mutation whitelist, the following filtering steps will not be performed; otherwise, for mutations not appearing in the known mutation whitelist, all of the following filtering criteria will be applied:

[0065] 1) Only retain the mutation with the highest mutation frequency at the same site: Let ALT be the number of mutated bases at the same site and REF be the number of reference bases at the same site. Then the mutation frequency AF = ALT / total number of reads covering the site = ALT / (REF + all ALT). Only retain the mutation with the highest mutation frequency at the same site.

[0066] In this context, all ALT values ​​are primers. There may be multiple ALT values ​​or mutations at the same position. In such cases, the ALT values ​​of multiple mutations at the same position need to be added together for calculation.

[0067] 2) Average position of the mutation site in different reads: Statistically count the positions of the same mutation site in different reads, calculate the average position, and filter out mutations whose mutation frequency is less than the mutation frequency threshold and whose average position is less than the position threshold.

[0068] In one implementation of this application, the mutation frequency threshold is 0.35, and the position threshold is 8. That is, mutations with a frequency AF < 0.35 and whose mutation positions occur within the first 7 bp of the read are filtered out. It should be noted that the position threshold of 8 is derived from experience and statistical studies. Generally, read lengths are around 100 bp, and a small value of 8 indicates that the mutation appears early in the read. Studies show that when AF is small, such as AF < 0.35, and the mutation appears early, it suggests that the mutation is likely caused by a sequencing error, and therefore it is filtered out.

[0069] 3) Are all mutations located in the same position? Count the position of mutations in all reads. If the mutation is located in the same position in all reads, the mutation is filtered out.

[0070] It should be noted that this application study believes that if the mutation is in the same position in all reads, it means that the mutation is caused by sequencing error. Therefore, this application creatively filters out mutations in which the mutation site is in the same position in different reads.

[0071] 4) Average base quality of mutation: Calculate the average base quality of mutation. If the average base quality of mutation is lower than the base quality threshold, the mutation is filtered out.

[0072] In one implementation of this application, the base quality threshold is 22.5, meaning that mutations with an average base quality value lower than 22.5 are filtered out. It should be noted that this value of 22.5 was obtained based on experience and statistical research; studies show that if the average base quality value is too low, such as less than 22.5, it indicates poor sequencing quality at those locations, and the mutation is likely a false positive, so it is filtered out.

[0073] 5) Root mean square of read alignment quality values: Calculate the root mean square of the alignment quality values ​​of all reads containing the mutation. If the value is less than the root mean square threshold, the mutation is filtered out.

[0074] In one implementation of this application, the root mean square threshold is 40, meaning that mutations with a root mean square value less than 40 are filtered out. It should be noted that the root mean square threshold in this application is based on the RMSMQ value distribution of GATK's VQSR. Referring to this distribution, this application creatively proposes that a root mean square threshold of 40 can effectively filter out false mutations.

[0075] 6) Quality value per unit depth of mutation: The quality value of mutation divided by the sum of the depths of all samples containing the mutation is the quality value per unit depth. If this value is less than the quality value per unit depth threshold, the mutation is filtered out.

[0076] In one implementation of this application, the quality value threshold per unit depth is 2, meaning that mutations with a quality value per unit depth less than 2 are filtered out. The quality value threshold per unit depth in this application is based on the QD value distribution of GATK's VQSR. Referring to this distribution, this application creatively proposes that a quality value threshold of 2 per unit depth can effectively filter out false mutations.

[0077] 7) Average number of mismatches in read segments: When the average number of mismatches in the read segment containing the mutation is greater than the threshold for the number of mismatches, the mutation is likely due to a comparison error, and therefore the mutation is filtered out.

[0078] In one implementation of this application, the mismatch count threshold is 5.25, meaning that mutations with an average mismatch count greater than 5.25 are filtered out. It should be noted that this mismatch count threshold was obtained based on experience and statistical research. Studies show that an excessive number of mismatches likely indicates that the alignment of these reads is incorrect, meaning these SNVs arise from sequence alignment rather than actual occurrences. Therefore, this application suggests that filtering out mutations with an average mismatch count greater than 5.25 can effectively eliminate false positives caused by this.

[0079] 8) Signal-to-noise ratio (SNR): The number of reads in which the mutation is located, where the base quality value is greater than the base quality value threshold, divided by the number of reads where the base quality value is less than the base quality value threshold (if this value is 0, then a value of 0.5 is assigned), is the SNR. If this ratio is less than the SNR threshold, then the mutation is filtered out.

[0080] In one implementation of this application, the base quality threshold is 22.5, and the signal-to-noise ratio (SNR) threshold is 1.5. The SNR is calculated by dividing the number of reads with base quality values ​​greater than the threshold by the number of reads with base quality values ​​less than the threshold, and then filtering out mutations with an SNR less than 1.5. It should be noted that the SNR threshold in this application is also obtained through empirical and statistical studies. Studies show that a higher SNR indicates a lower probability that the mutation is caused by a sequencing error; conversely, a lower SNR indicates a higher probability that the mutation is caused by a sequencing error. Therefore, this application suggests that filtering out mutations with an SNR less than 1.5 can effectively remove mutations caused by sequencing errors.

[0081] 9) Rank sum test of base quality values ​​of heterozygous site reads: Perform a rank sum test on the base quality values ​​of the reads containing the mutated bases and the reads containing the reference bases. If the rank sum is less than the rank sum threshold, the mutation is filtered out.

[0082] In one implementation of this application, the quality value rank sum threshold is -12.5, that is, mutations with a rank sum less than -12.5 are filtered out. It should be noted that the quality value rank sum threshold of this application is based on the MQRankSum distribution of GATK's VQSR. Referring to this distribution, this application creatively proposes that a quality value rank sum threshold of -12.5 can effectively filter out false mutations.

[0083] 10) Rank sum test of heterozygous mutation positions: Perform a rank sum test on the position of the mutated base in the read segment and the position of the reference base in the read segment. If the rank sum is less than the position rank sum threshold, the mutation is filtered out.

[0084] In one implementation of this application, the positional rank sum threshold is -8, i.e., mutations with a rank sum less than -8 are filtered out. It should be noted that this positional rank sum threshold is based on the ReadPOSRankSum distribution of GATK's VQSR. Referring to this distribution, this application creatively proposes that a positional rank sum threshold of -8 can effectively filter out false mutations.

[0085] 11) Mutation chain distribution bias:

[0086] The mutation frequency was calculated using only reads with an average base mass value greater than 22.5 and denoted as HiAF. The calculation formula for HiAF is the same as that for AF in this application, except that the former uses all reads, while HiAF only uses reads with an average base mass value greater than 22.5.

[0087] Calculate the number of reference bases (SRF) on the positive strand, the number of reference bases (SRR) on the negative strand, the number of mutant bases (SAF) on the positive strand, and the number of mutant bases (SAR) on the negative strand.

[0088] For a reference base, the reference base bias value RefBias is 0 when either SRF or SRR is 0 and SRF + SRR <= 12. Calculate SRF / (SRF + SRR) and SRR / (SRF + SRR). If both values ​​are greater than 0.05 and both SRF and SRR are greater than 2, the reference base bias value RefBias is 2; otherwise, it is 1.

[0089] For mutated bases, when either SAF or SAR is 0 and SAF+SAR<=12, the mutated base bias value AltBias is equal to 0. Calculate SAF / (SAF+SAR) and SAR / (SAF+SAR). If both values ​​are greater than 0.05 and both SAF and SAR are greater than 2, the mutated base bias value AltBias is 2; otherwise, it is 1.

[0090] Fisher's exact test was performed on SRF, SRR, SAF, and SAR to obtain p-values ​​and odd ratios.

[0091] When HiAF is less than 0.25, RefBias is 2 and AltBias is 1, p value is less than or equal to 0.01, odd ratio is greater than 5 (or odd ratio is equal to 0), and mutation length is less than 100bp, it indicates that the extreme distribution of the mutated bases is on the positive or negative strand, and the mutation is filtered out.

[0092] It should be noted that in filtering mutations based on the chain distribution bias, all specific values ​​used were obtained from experience and statistical studies in this application; among them, a p-value < 0.01 is the statistical threshold for significant differences; an odd ratio > 5 or equal to 0 also indicates that the two groups of subjects are significantly different.

[0093] Those skilled in the art will understand that all or part of the functions of the above methods can be implemented in hardware or by computer programs. When all or part of the functions of the above methods are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in storage media such as a server, another computer, disk, optical disk, flash drive, or portable hard drive, and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions of the above methods can be achieved.

[0094] Therefore, based on the method for correcting mitochondrial genome sequencing mutations in this application, this application proposes a device for correcting mitochondrial genome sequencing mutations, such as... Figure 2 As shown, it includes a whitelist comparison module 21, a mutation filtering module 22, and a corrected data acquisition module 23.

[0095] The whitelist comparison module 21 includes a method for comparing the mitochondrial genome sequencing mutations to be corrected with a whitelist of known mutations, and retaining the mitochondrial genome sequencing mutations to be corrected that appear in the whitelist of known mutations as the first mutation set.

[0096] The mutation filtering module 22 includes a method for filtering mutations that do not appear in the known mutation whitelist based on 1) mutation frequency, 2) the average position of the read segment where the mutation is located, 3) whether the mutations are all located in the same position, 4) the average base quality value of the mutation, 5) the root mean square of the read alignment quality value, 6) the quality value per unit depth of the mutation, 7) the average number of mismatches in the read alignment, 8) signal-to-noise ratio, 9) the rank-sum test of the base quality value of the heterozygous site read, 10) the rank-sum test of the position of the heterozygous site mutation, and 11) the chain distribution deviation of the mutation. The remaining mutations after filtering are used as a second mutation set.

[0097] The corrected data acquisition module 23 includes a module for merging the first mutation set and the second mutation set, i.e., obtaining the corrected mitochondrial genome sequencing mutations.

[0098] Another implementation of this application also provides an apparatus for correcting mitochondrial genome sequencing mutations, the apparatus including a memory and a processor; the memory includes a program for storing a program; the processor includes a program for executing the program stored in the memory to implement the following method: comparing the mitochondrial genome sequencing mutation to be corrected with a known mutation whitelist, retaining the mitochondrial genome sequencing mutations to be corrected that appear in the known mutation whitelist as a first mutation set; for the mitochondrial genome sequencing mutations to be corrected that do not appear in the known mutation whitelist, based on 1) mutation frequency, 2) the average value of the position of the read containing the mutation, 3) The mutations are filtered based on the following criteria: whether all mutations are located in the same position, 4) the average base quality of the mutations, 5) the root mean square of the read alignment quality, 6) the quality per unit depth of the mutations, 7) the average number of mismatches in read alignments, 8) the signal-to-noise ratio, 9) the rank-sum test of the base quality of heterozygous site reads, 10) the rank-sum test of the mutation position at heterozygous site, and 11) the chain distribution bias of the mutations. The remaining mutations after filtering are used as the second mutation set. The first and second mutation sets are merged to obtain the corrected mitochondrial genome sequencing mutations. The known mutation whitelist is a set of verified mutations that are known to be pathogenic, potentially pathogenic, or of unknown significance.

[0099] Another implementation of this application also provides a computer-readable storage medium including a program that can be executed by a processor to implement the following method: comparing the mitochondrial genome sequencing mutations to be corrected with a known mutation whitelist, retaining the mitochondrial genome sequencing mutations to be corrected that appear in the known mutation whitelist as a first mutation set; for the mitochondrial genome sequencing mutations to be corrected that do not appear in the known mutation whitelist, based on 1) mutation frequency, 2) the average position of the read segment where the mutation is located, 3) whether the mutations are all located in the same position, 4) mutation The mutations are filtered using the following methods: 5) average base quality value, 6) root mean square of read alignment quality value, 7) quality value per unit depth of mutation, 8) average number of mismatches in read alignment, 9) signal-to-noise ratio, 10) rank sum test of base quality value of heterozygous site reads, 11) rank sum test of mutation position of heterozygous site, and 12) chain distribution bias of mutation. The remaining mutations after filtering are used as the second mutation set. The first and second mutation sets are merged to obtain the corrected mitochondrial genome sequencing mutations. The known mutation whitelist is a set of verified mutations known to be pathogenic, potentially pathogenic, or of unknown significance.

[0100] The following specific experiments will further illustrate this application in detail. These experiments are merely illustrative and should not be construed as limiting the scope of this application.

[0101] Example

[0102] First, mitochondrial mutation data were obtained from MSeqDR (Mitochondrial Related Diseases Database) and ClinVar (Clinical Disease and Phenotype Related Mutation Database). Pathogenic, potentially pathogenic, or undetermined mutations were retained, merged, and duplicates were removed to form a whitelist of known mutations.

[0103] Specifically, a Python script is used to retrieve all data from the MSeqDR database, keeping only rows where the Chr column is MT and the Effect column is + / + (i.e., pathogenic mutations). This data is placed in temporary file 1. Another Python script converts the data in temporary file 1 to obtain information such as the start, end, reference base, and mutated base for each mutation, resulting in temporary file 2. The variant_summary.txt file is downloaded from the ClinVar FTP site, and a Python script converts its format, retaining only rows where the Chromosome column is MT, the ClinicalSignificance column contains either pathogenic or uncertain fields (case-insensitive), and the ReviewStatus column does not contain a no assertion field, resulting in temporary file 3. Temporary files 2 and 3 are merged, deduplicated, and a final whitelist of known mutations is generated.

[0104] The whitelist of known mutations obtained in this example contains a total of 1,473 mutations, all of which are known pathogenic, potentially pathogenic, or undetermined mutations in the human mitochondrial genome.

[0105] The mitochondrial genome sequencing mutations to be corrected are compared with a whitelist of known mutations, and the mitochondrial genome sequencing mutations to be corrected that appear in the whitelist of known mutations are retained as the first mutation set.

[0106] For mutations that do not appear in the known mutation whitelist, the following filtering is performed:

[0107] 1) Only retain the mutation with the highest mutation frequency at the same site: Let ALT be the number of mutated bases at the same site and REF be the number of reference bases at the same site. Then the mutation frequency AF = ALT / total number of reads covering the site = ALT / (REF + all ALT). Only retain the mutation with the highest mutation frequency at the same site.

[0108] In this context, all ALT values ​​are primers. There may be multiple ALT values ​​or mutations at the same position. In such cases, the ALT values ​​of multiple mutations at the same position need to be added together for calculation.

[0109] 2) Average position of the mutation site in different reads: Count the positions of the same mutation site in different reads, calculate the average position, and filter out mutations with a mutation frequency of less than 0.35 and an average position of less than 8.

[0110] 3) Are all mutations located in the same position? Count the position of mutations in all reads. If the mutation is located in the same position in all reads, the mutation is filtered out.

[0111] 4) Average base mass value of the mutation: Calculate the average base mass value of the mutation. If the average base mass value of the mutation is lower than 22.5, the mutation is filtered out.

[0112] 5) Root mean square of read alignment quality values: Calculate the root mean square of the alignment quality values ​​of all reads containing the mutation. If the value is less than 40, the mutation is filtered out.

[0113] 6) Quality value per unit depth of mutation: The quality value of mutation divided by the sum of the depths of all samples containing the mutation is the quality value per unit depth. If this value is less than 2, the mutation is filtered out.

[0114] 7) Average number of mismatches in read segments: When the average number of mismatches in the read segment containing the mutation is greater than 5.25, the mutation is likely due to a comparison error, so the mutation should be filtered out.

[0115] 8) Signal-to-noise ratio: The number of reads with a base quality value greater than 22.5 in the read segment containing the mutation, divided by the number of reads with a base quality value less than 22.5 (if the value is 0, then a value of 0.5 is assigned), is the signal-to-noise ratio. If the ratio is less than 1.5, the mutation is filtered out.

[0116] 9) Rank sum test of base quality values ​​of heterozygous site reads: Perform a rank sum test on the base quality values ​​of the reads containing the mutated bases and the reads containing the reference bases. If the rank sum is less than -12.5, the mutation is filtered out.

[0117] 10) Rank sum test of heterozygous mutation position: Perform a rank sum test on the position of the mutated base in the read segment and the position of the reference base in the read segment. If the rank sum is less than -8, then filter out the mutation.

[0118] 11) Mutation chain distribution bias:

[0119] The mutation frequency was calculated using only reads with an average base mass value greater than 22.5 and denoted as HiAF. The calculation formula for HiAF is the same as that for AF in this application, except that the former uses all reads, while HiAF only uses reads with an average base mass value greater than 22.5.

[0120] Calculate the number of reference bases (SRF) on the positive strand, the number of reference bases (SRR) on the negative strand, the number of mutant bases (SAF) on the positive strand, and the number of mutant bases (SAR) on the negative strand.

[0121] For a reference base, the reference base bias value RefBias is 0 when either SRF or SRR is 0 and SRF + SRR <= 12. Calculate SRF / (SRF + SRR) and SRR / (SRF + SRR). If both values ​​are greater than 0.05 and both SRF and SRR are greater than 2, the reference base bias value RefBias is 2; otherwise, it is 1.

[0122] For mutated bases, when either SAF or SAR is 0 and SAF+SAR<=12, the mutated base bias value AltBias is equal to 0. Calculate SAF / (SAF+SAR) and SAR / (SAF+SAR). If both values ​​are greater than 0.05 and both SAF and SAR are greater than 2, the mutated base bias value AltBias is 2; otherwise, it is 1.

[0123] Fisher's exact test was performed on SRF, SRR, SAF, and SAR to obtain p-values ​​and odd ratios.

[0124] When HiAF is less than 0.25, RefBias is 2 and AltBias is 1, p value is less than or equal to 0.01, odd ratio is greater than 5 (or odd ratio is equal to 0), and mutation length is less than 100bp, it indicates that the extreme distribution of the mutated bases is on the positive or negative strand, and the mutation is filtered out.

[0125] After filtering using the above 11 indicators, the remaining mutations constitute the second mutation set. Combining the first and second mutation sets yields the corrected mitochondrial genome sequencing mutations.

[0126] The above process involves inputting the VCF file and SAM or BAM file of the sample. The VCF file can be generated by mutation detection software such as Frebayes, GATK, and VarDict, while the SAM or BAM file can be generated from sequence alignment files such as BWA and bowtie2. A filtering procedure is then executed, filtering through a whitelist and filtering indicators to obtain the filtered VCF file, which represents the corrected mitochondrial genome sequencing mutations. A typical VCF file looks like this... Figure 3 As shown, SAM files are generally like... Figure 4 As shown.

[0127] This example specifically selected 10 mitochondrial genome next-generation sequencing samples and statistically analyzed the changes in the number of mutations, the number of positive mutations, precision, and false alarm rate before and after filtering. The results are shown in Table 1. Precision refers to the proportion of true positive mutations before and after filtering; false alarm rate refers to the proportion of false positive mutations before and after filtering.

[0128] Table 1. Statistics of results before and after mutation correction in mitochondrial genome next-generation sequencing.

[0129]

[0130] Table 1 shows that filtering reduced the number of mutations per sample by an average of 23.1, improved precision by an average of 27.2%, and reduced the false alarm rate by an average of 26.3%. This significantly reduced the amount of data required for downstream analysis, increased the proportion of true positive mutations, and decreased the proportion of false mutations. Although an average of 1.1 mutations per sample were filtered out, it was verified that the missed mutations were all population polymorphisms, i.e., non-pathogenic mutations. For example, the NC_012920.1:m.16179C>T mutation in sample AS75058. Therefore, the mitochondrial genome sequencing mutation correction method in this example can effectively remove false positive mutations caused by improper experimental methods, sequencing errors, or systematic errors in alignment software, obtaining a highly reliable mutation set. This reduces the workload of mutation interpretation while improving the accuracy of the analysis results. Furthermore, this correction method can be used for single samples or a small number of samples.

[0131] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.

Claims

1. A method for correcting mitochondrial genome sequencing mutations, characterized in that: Includes the following steps, The mitochondrial genome sequencing mutations to be corrected are compared with a whitelist of known mutations, and the mitochondrial genome sequencing mutations to be corrected that appear in the whitelist of known mutations are retained as the first mutation set; For mitochondrial genome sequencing mutations that do not appear in the known mutation whitelist and need to be corrected, the mutations are filtered based on 1) mutation frequency, 2) average position of the read containing the mutation, 3) whether the mutations are all located at the same position, 4) average base quality value of the mutation, 5) root mean square of read alignment quality value, 6) quality value per unit depth of the mutation, 7) average number of mismatches in read alignment, 8) signal-to-noise ratio, 9) rank-sum test of base quality value of heterozygous site reads, 10) rank-sum test of mutation position of heterozygous site, and 11) strand distribution bias of the mutation. The remaining mutations after filtering are used as the second mutation set. The first and second mutation sets are merged to obtain the corrected mitochondrial genome sequencing mutations; The known mutation whitelist is a verified set of mutations that are known to be pathogenic, potentially pathogenic, or of unknown significance. The signal-to-noise ratio is calculated by dividing the number of reads with base quality values ​​greater than the base quality value threshold by the number of reads with base quality values ​​less than the base quality value threshold.

2. The method according to claim 1, characterized in that: Filtering mutations based on mutation frequency includes calculating the mutation frequency of all mutations and, for the same mutation site, retaining only the mutation with the highest mutation frequency at that same mutation site.

3. The method according to claim 2, characterized in that: Filtering mutations based on the average position of the mutation in the read segment (2) includes: counting the positions of the same mutation site in different read segments, calculating the average position, and filtering out mutations whose mutation frequency is less than the mutation frequency threshold and whose average position is less than the position threshold.

4. The method according to claim 1, characterized in that: Based on 3), filtering mutations based on whether they are all located in the same position includes filtering out mutations that are located in the same position in different reads for the same mutation site.

5. The method according to claim 1, characterized in that: Filtering mutations based on the average base quality value of mutations (4) includes calculating the average base quality value of mutations and filtering out mutations with an average base quality value lower than the base quality value threshold.

6. The method according to claim 1, characterized in that: Filtering mutations based on the root mean square of the read alignment quality values ​​(5) includes calculating the root mean square of the alignment quality values ​​of all reads containing the mutation, and filtering out mutations whose root mean square of the alignment quality values ​​is less than the root mean square threshold.

7. The method according to claim 1, characterized in that: Filtering mutations based on the unit depth quality value of mutations (6) includes calculating the unit depth quality value of mutations and filtering out mutations whose unit depth quality value is less than the unit depth quality value threshold.

8. The method according to claim 1, characterized in that: According to 7), filtering mutations by average mismatch count of read segments includes calculating the average mismatch count of the read segment containing the mutation and filtering out mutations whose average mismatch count is greater than the mismatch count threshold.

9. The method according to claim 1, characterized in that: Based on signal-to-noise ratio (8), filtering for mutations includes... Filter out abrupt changes with a signal-to-noise ratio (SNR) value less than the SNR threshold.

10. The method according to claim 1, characterized in that: According to the rank sum test of the base quality values ​​of the heterozygous site read segment in step 9), the mutation filtering includes performing a rank sum test on the base quality values ​​of the read segment containing the mutated base and the read segment containing the reference base, and filtering out mutations whose rank sum is less than the rank sum threshold of the quality value.

11. The method according to claim 1, characterized in that: According to the rank sum test of the position of the heterozygous site mutation in 10), the mutation is filtered out by performing a rank sum test on the position of the mutated base in the read segment and the position of the reference base in the read segment, and filtering out mutations whose rank sum is less than the position rank sum threshold.

12. The method according to claim 3, characterized in that: The formula for calculating the mutation frequency is as follows: AF = ALT / total number of reads covering this site. or, AF = ALT / (REF + all ALT values) Where AF is the mutation frequency, ALT is the number of mutated bases at the same site, REF is the number of reference bases at the same site, and all ALT refers to the ALT of all mutations when there are multiple mutations at the same position.

13. The method according to claim 3, characterized in that: The mutation frequency threshold is 0.35, and the location threshold is 8.

14. The method according to claim 5, characterized in that: The base quality value threshold is 22.

5.

15. The method according to claim 6, characterized in that: The root mean square threshold is 40.

16. The method according to claim 7, characterized in that: The mutation quality value per unit depth is calculated by dividing the mutation quality value by the sum of the depths of all samples containing the mutation.

17. The method according to claim 7, characterized in that: The threshold value for the unit depth quality value is 2.

18. The method according to claim 8, characterized in that: The threshold for the number of mismatches is 5.

25.

19. The method according to claim 9, characterized in that: When filtering mutations based on the signal-to-noise ratio in step 8), if the number of reads with a base quality value less than the base quality value threshold is 0, then assign it a value of 0.

5.

20. The method according to claim 9, characterized in that: The signal-to-noise ratio threshold is 1.

5.

21. The method according to claim 10, characterized in that: The rank sum threshold for the quality value is -12.

5.

22. The method according to claim 11, characterized in that: The position rank sum threshold is -8.

23. The method according to any one of claims 1-22, characterized in that: Filtering mutations based on the chain distribution bias of mutations (11) includes: Only reads with an average base mass value greater than 22.5 were used to calculate the frequency of mutations, which were then labeled as HiAF. The number of reference bases in the positive strand is calculated and labeled as SRF, the number of reference bases in the negative strand is labeled as SRR, the number of mutant bases in the positive strand is labeled as SAF, and the number of mutant bases in the negative strand is labeled as SAR. For a reference base, when either SRF or SRR is 0, and SRF+SRR<=12, the reference base deviation value RefBias is 0; calculate SRF / (SRF+SRR) and SRR / (SRF+SRR). When both of these values ​​are greater than 0.05, and both SRF and SRR are greater than 2, the reference base deviation value RefBias is 2; otherwise, it is 1. For mutated bases, when either SAF or SAR is 0 and SAF+SAR<=12, the mutated base bias value AltBias is equal to 0; calculate SAF / (SAF+SAR) and SAR / (SAF+SAR). When both of these values ​​are greater than 0.05 and both SAF and SAR are greater than 2, the mutated base bias value AltBias is 2; otherwise, it is 1. Fisher's exact test was performed on SRF, SRR, SAF, and SAR to obtain p-values ​​and odd ratios. When HiAF is less than 0.25, RefBias is 2, AltBias is 1, p value is less than or equal to 0.01, odd ratio is greater than 5 or equal to 0, and mutation length is less than 100bp, the mutation is filtered out.

24. The method according to any one of claims 1-22, characterized in that: The whitelist of known mutations is constructed by merging and removing duplicates from mitochondrial mutation data from the mitochondrial-associated disease database MSeqDR and the clinical disease and phenotype-associated mutation database ClinVar, retaining pathogenic, potentially pathogenic, or undetermined mutations.

25. A device for correcting mutations in mitochondrial genome sequencing, characterized in that: This includes a whitelist comparison module, a mutation filtering module, and a corrected data acquisition module; The whitelist comparison module includes a function to compare the mitochondrial genome sequencing mutations to be corrected with a known mutation whitelist, and retain the mitochondrial genome sequencing mutations to be corrected that appear in the known mutation whitelist as the first mutation set. The mutation filtering module includes a method for filtering mitochondrial genome sequencing mutations that are not listed in the known mutation whitelist. The filtering is performed based on the following criteria: 1) mutation frequency, 2) average position of the read containing the mutation, 3) whether all mutations are located at the same position, 4) average base quality value of the mutation, 5) root mean square of read alignment quality value, 6) quality value per unit depth of the mutation, 7) average number of mismatches in read alignment, 8) signal-to-noise ratio, 9) rank-sum test of base quality value of heterozygous site reads, 10) rank-sum test of mutation position at heterozygous site, and 11) strand distribution bias of the mutation. The remaining mutations after filtering are used as a second mutation set. The corrected data acquisition module includes a function to merge the first mutation set and the second mutation set, i.e., to obtain the corrected mitochondrial genome sequencing mutations. The known mutation whitelist is a verified set of mutations that are known to be pathogenic, potentially pathogenic, or of unknown significance. The signal-to-noise ratio is calculated by dividing the number of reads with base quality values ​​greater than the base quality value threshold by the number of reads with base quality values ​​less than the base quality value threshold.

26. The apparatus according to claim 25, characterized in that: Filtering mutations based on mutation frequency includes calculating the mutation frequency of all mutations and, for the same mutation site, retaining only the mutation with the highest mutation frequency at that same mutation site.

27. The apparatus according to claim 26, characterized in that: Filtering mutations based on the average position of the mutation in the read segment (2) includes: counting the positions of the same mutation site in different read segments, calculating the average position, and filtering out mutations whose mutation frequency is less than the mutation frequency threshold and whose average position is less than the position threshold.

28. The apparatus according to claim 25, characterized in that: Based on 3), filtering mutations based on whether they are all located in the same position includes filtering out mutations that are located in the same position in different reads for the same mutation site.

29. The apparatus according to claim 25, characterized in that: Filtering mutations based on the average base quality value of mutations (4) includes calculating the average base quality value of mutations and filtering out mutations with an average base quality value lower than the base quality value threshold.

30. The apparatus according to claim 25, characterized in that: Filtering mutations based on the root mean square of the read alignment quality values ​​(5) includes calculating the root mean square of the alignment quality values ​​of all reads containing the mutation, and filtering out mutations whose root mean square of the alignment quality values ​​is less than the root mean square threshold.

31. The apparatus according to claim 25, characterized in that: Filtering mutations based on the unit depth quality value of mutations (6) includes calculating the unit depth quality value of mutations and filtering out mutations whose unit depth quality value is less than the unit depth quality value threshold.

32. The apparatus according to claim 25, characterized in that: According to 7), filtering mutations by average mismatch count of read segments includes calculating the average mismatch count of the read segment containing the mutation and filtering out mutations whose average mismatch count is greater than the mismatch count threshold.

33. The apparatus according to claim 25, characterized in that: According to 8), filtering mutations by signal-to-noise ratio includes counting the number of reads containing mutations with base quality values ​​greater than a base quality value threshold and the number of reads with base quality values ​​less than a base quality value threshold. The quotient of the number of reads with base quality values ​​greater than a base quality value threshold divided by the number of reads with base quality values ​​less than a base quality value threshold is used as the signal-to-noise ratio. Mutations with a signal-to-noise ratio value less than the signal-to-noise ratio threshold are then filtered out.

34. The apparatus according to claim 25, characterized in that: According to the rank sum test of the base quality values ​​of the heterozygous site read segment in step 9), the mutation filtering includes performing a rank sum test on the base quality values ​​of the read segment containing the mutated base and the read segment containing the reference base, and filtering out mutations whose rank sum is less than the rank sum threshold of the quality value.

35. The apparatus according to claim 25, characterized in that: According to the rank sum test of the position of the heterozygous site mutation in 10), the mutation is filtered out by performing a rank sum test on the position of the mutated base in the read segment and the position of the reference base in the read segment, and filtering out mutations whose rank sum is less than the position rank sum threshold.

36. The apparatus according to claim 27, characterized in that: The formula for calculating the mutation frequency is as follows: AF = ALT / total number of reads covering this site. or, AF = ALT / (REF + all ALT values) Where AF is the mutation frequency, ALT is the number of mutated bases at the same site, REF is the number of reference bases at the same site, and all ALT refers to the ALT of all mutations when there are multiple mutations at the same position.

37. The apparatus according to claim 27, characterized in that: The mutation frequency threshold is 0.35, and the location threshold is 8.

38. The apparatus according to claim 29, characterized in that: The base quality value threshold is 22.

5.

39. The apparatus according to claim 30, characterized in that: The root mean square threshold is 40.

40. The apparatus according to claim 31, characterized in that: The mutation quality value per unit depth is calculated by dividing the mutation quality value by the sum of the depths of all samples containing the mutation.

41. The apparatus according to claim 31, characterized in that: The threshold value for the unit depth quality value is 2.

42. The apparatus according to claim 32, characterized in that: The threshold for the number of mismatches is 5.

25.

43. The apparatus according to claim 33, characterized in that: When filtering mutations based on the signal-to-noise ratio in step 8), if the number of reads with a base quality value less than the base quality value threshold is 0, then assign it a value of 0.

5.

44. The apparatus according to claim 33, characterized in that: The signal-to-noise ratio threshold is 1.

5.

45. The apparatus according to claim 34, characterized in that: The rank sum threshold for the quality value is -12.

5.

46. ​​The apparatus according to claim 35, characterized in that: The position rank sum threshold is -8.

47. The apparatus according to claim 25, characterized in that: Filtering mutations based on the chain distribution bias of mutations (11) includes: Only reads with an average base mass value greater than 22.5 were used to calculate the frequency of mutations, which were then labeled as HiAF. The number of reference bases in the positive strand is calculated and labeled as SRF, the number of reference bases in the negative strand is labeled as SRR, the number of mutant bases in the positive strand is labeled as SAF, and the number of mutant bases in the negative strand is labeled as SAR. For a reference base, when either SRF or SRR is 0, and SRF+SRR<=12, the reference base deviation value RefBias is 0; calculate SRF / (SRF+SRR) and SRR / (SRF+SRR). When both of these values ​​are greater than 0.05, and both SRF and SRR are greater than 2, the reference base deviation value RefBias is 2; otherwise, it is 1. For mutated bases, when either SAF or SAR is 0 and SAF+SAR<=12, the mutated base bias value AltBias is equal to 0; calculate SAF / (SAF+SAR) and SAR / (SAF+SAR). When both of these values ​​are greater than 0.05 and both SAF and SAR are greater than 2, the mutated base bias value AltBias is 2; otherwise, it is 1. Fisher's exact test was performed on SRF, SRR, SAF, and SAR to obtain p-values ​​and odd ratios. When HiAF is less than 0.25, RefBias is 2, AltBias is 1, p value is less than or equal to 0.01, odd ratio is greater than 5 or equal to 0, and mutation length is less than 100bp, the mutation is filtered out.

48. The apparatus according to claim 25, characterized in that: The whitelist of known mutations is constructed by merging and removing duplicates from mitochondrial mutation data from the mitochondrial-associated disease database MSeqDR and the clinical disease and phenotype-associated mutation database ClinVar, retaining pathogenic, potentially pathogenic, or undetermined mutations.

49. A device for correcting mutations in mitochondrial genome sequencing, characterized in that: The device includes a memory and a processor; The memory includes a storage device for storing programs; The processor includes a method for implementing the method of correcting mitochondrial genome sequencing mutations according to any one of claims 1-24 by executing a program stored in the memory.

50. A computer-readable storage medium, characterized in that: The storage medium stores a program that can be executed by a processor to implement the method for correcting mitochondrial genome sequencing mutations as described in any one of claims 1-24.