Method, device and equipment for analyzing mitochondrial copy number and computer readable storage medium

The interval screening and calculation of mitochondrial genomes is solved through whole-genome sequencing data, and the problems of high cost of quantification of mitochondrial genome copy number in the prior art are solved, and efficient and accurate mitochondrial copy number calculation is achieved.

CN120048338APending Publication Date: 2025-05-27AEGICARE (SHENZHEN) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411949702.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has problems of high cost and poor accuracy when quantifying mitochondrial genome copy numbers. In particular, the Real-time PCR method increases detection costs and time costs, and the method based on second-generation sequencing data is poorly accurate.

Method used

Through whole genome sequencing data, the chromosomal genome and mitochondrial genome are divided into multiple intervals, and gene sequence homology screening, genome sequencing complex region screening and GC content screening were performed to exclude inappropriate intervals, and mitochondrial copy number were calculated, using the formula Mi = (ADi * MBi) / (ABi * MDi).

Benefits of technology

Without standards or reference samples, the mitochondrial copy number is calculated directly based on whole genome sequencing data, reducing costs and time, improving accuracy, and the calculation results are close to the true value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048338A_ABST
    Figure CN120048338A_ABST
Patent Text Reader

Abstract

The invention discloses a method, device and equipment for analyzing the mitochondrial copy number and a computer readable storage medium, and the method comprises the following steps: obtaining whole genome sequencing data, and dividing a chromosome genome and a mitochondrial genome into a plurality of intervals according to a preset length; carrying out at least one of the following screening steps on all the intervals: gene sequence homology screening, genome sequencing complex region screening and GC content screening; a screened interval is reserved and obtained; calculating the mitochondrial copy number: calculating the interval number ABi and the sequencing depth sum ADi of the reserved chromosome genome, calculating the interval number MBi and the sequencing depth sum MDi of the reserved mitochondrial genome, and calculating the mitochondrial copy number Mi according to the following formula: # imgabs0 #. According to the invention, the mitochondrial copy number can be accurately calculated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of mitochondrial copy number analysis, and in particular to a method, device, equipment and computer-readable storage medium for analyzing mitochondrial copy number. Background Art

[0002] Mitochondria are one of the most important organelles in eukaryotes. They have genetic material that is semi-independent from the nuclear genome, namely the mitochondrial genome (mtDNA). Unlike the common diploid genome (such as human chromosomes 1 to 22), the mitochondrial genome is a small haploid genome inherited from the mother. There are usually hundreds or thousands of copies of the mitochondrial genome in each cell, depending on the cell type, developmental stage, etc. Too high a copy number of the mitochondrial genome may mean that the function of the mitochondria is impaired, and the energy demand of the cell needs to be compensated by increasing the number of mitochondria. Too low a copy number of the mitochondrial genome may mean that there are mutations in the mitochondrial genome itself or in the nuclear gene DNA that affect the replication of the mitochondrial genome, thereby impairing the function of the mitochondria. By quantifying the copy number of the mitochondrial genome, the pathological patterns of mitochondrial diseases, such as Leigh syndrome and Parkinson's syndrome, can be studied. In addition, studies have shown that human aging is often accompanied by a decrease in the copy number of the mitochondrial genome, so the quantification of the copy number of the mitochondrial genome can also help researchers study the mechanism of human aging.

[0003] At present, the main methods for quantifying the copy number of the mitochondrial genome include Real-time PCR (real-time fluorescence quantitative PCR), methods for calculating the copy number by establishing a baseline using reference samples, and methods for relative quantification of the copy number based on second-generation sequencing data. Real-time PCR technology uses the characteristics of DNA polymerase to synthesize new DNA chains during the PCR process, combined with a fluorescently labeled probe or dye, to measure the extent of the PCR reaction by real-time monitoring of the increase in the fluorescence signal. The intensity of the fluorescence signal is proportional to the amount of target DNA in the PCR reaction. Therefore, by using several standards with known template concentrations as controls, the copy number of the target DNA of the sample to be tested can be obtained. However, the Real-time PCR method adds additional detection costs and time costs, and the accuracy is easily affected by the noise of the fluorescence signal intensity and the accuracy of the standard. Other methods calculate the mitochondrial copy number by sequencing the mitochondrial genome and establishing a baseline with reference samples. However, this method also has its flaws, because the copy number of the mitochondrial genome varies greatly among individuals, and the mitochondrial copy number of the same individual sampled at different times and tissues may also be different. Therefore, it is almost impossible to obtain a baseline with small deviation and applicable to all samples to be tested. Therefore, using this method to calculate the mitochondrial copy number has the problem of poor accuracy. Summary of the invention

[0004] Based on the above problems, the present application provides a method, apparatus, device, and computer-readable storage medium for analyzing mitochondrial copy number.

[0005] The present application adopts the following technical solutions:

[0006] A first aspect of the present application discloses a method for analyzing mitochondrial copy number, including: obtaining whole-genome sequencing data, dividing the chromosomal genome and mitochondrial genome into multiple intervals according to a preset length; performing at least one of the following screening steps on all intervals: gene sequence homology screening, screening of complex regions of genome sequencing, GC content screening; and retaining the intervals obtained after screening; calculating the mitochondrial copy number: calculating the number of intervals ABi of the retained chromosomal genome and the total sequencing depth ADi, calculating the number of intervals MBi of the retained mitochondrial genome and the total sequencing depth MDi, and calculating the mitochondrial copy number Mi according to the following formula:

[0007]

[0008] It should be noted that in the present application, based on the whole-genome sequencing data, by dividing the chromosomal genome and mitochondrial genome into multiple intervals and screening out inappropriate intervals, the mitochondrial copy number can be calculated more accurately.

[0009] In one implementation manner of the present application, the preset length is 50bp to 500bp.

[0010] In one implementation manner of the present application, the preset length is 80bp to 200bp.

[0011] In one implementation manner of the present application, the step of gene sequence homology screening includes: calculating the unique comparison score of each interval, and retaining the intervals with the unique comparison score greater than or equal to a preset value. It should be noted that the unique comparison score can also be called the unique alignment rate, which refers to the proportion of reads that can be uniquely aligned to the genome in the sequencing data, usually between 0 and 1, indicating the proportion of successful alignments. For example, 0.9 means that 90% of the reads can be uniquely aligned to the genome. In the present application, by retaining the intervals with the unique comparison score greater than or equal to the preset value and excluding the intervals with a low unique comparison score, it is possible to exclude the influence of genome sequence homology on the average sequencing depth of each interval of the chromosomal and mitochondrial genomes, calculate the average sequencing depth more accurately, and then calculate the mitochondrial copy number more accurately.

[0012] In one implementation manner of the present application, the preset value is 0.95 to 1.

[0013] In one implementation manner of the present application, the preset value is 1.

[0014] In one implementation of the present application, the steps of screening complex regions for genome sequencing include: comparing each interval with the complex regions for genome sequencing, and retaining the intervals whose positions do not overlap with the complex regions for genome sequencing. It should be noted that the complex regions for genome sequencing refer to the genomic regions that are prone to problems during sequencing, alignment, or subsequent analysis. In the present application, by retaining the intervals that do not overlap with the complex regions for genome sequencing and excluding the intervals that overlap with the complex regions for genome sequencing, it is possible to facilitate the exclusion of the influence of the complex regions for genome sequencing on the average sequencing depth of each interval of the chromosome and mitochondrial genome, more accurately calculate the average sequencing depth, and further more accurately calculate the mitochondrial copy number.

[0015] In one implementation of the present application, the complex regions for genome sequencing include at least one of repetitive sequence regions, GC content abnormal regions, highly conserved regions, chromosome ends and centromere regions, pseudogene regions, and regions prone to deviation in sequencing analysis.

[0016] In one implementation of the present application, the complex regions for genome sequencing are ENCODE blacklist regions.

[0017] In one implementation of the present application, the steps of GC content screening include: retaining the intervals with a GC content greater than or equal to 30% and less than or equal to 60%. It should be noted that regions with too low or too high GC content are usually regions with a relatively high degree of repetition or some relatively complex regions. By excluding regions with too low or too high GC content, it is beneficial to correct the deviation in mitochondrial copy number analysis caused by the GC content.

[0018] In one implementation of the present application, all the following screening steps are performed on all intervals: the gene sequence homology screening, the screening of complex regions for genome sequencing, and the GC content screening. Thus, it is possible to correct the deviation caused by homologous sequences, repetitive sequences, and GC content by excluding the influence of genome sequence homology, complex regions for genome sequencing, and GC content on the sequencing depth, so as to calculate a mitochondrial genome copy number closer to the true value.

[0019] In one implementation of the present application, the calculation of the mitochondrial copy number includes: grouping the retained intervals according to the GC content, calculating the mitochondrial copy number of each group respectively, and the final mitochondrial copy number is the average or median of the mitochondrial copy numbers of each group. It should be noted that by grouping the intervals according to the GC content to calculate the mitochondrial copy number and taking the average or median as the final mitochondrial copy number, it is possible to facilitate the correction of the deviation of the DNA sequence amplification effect caused by different GC contents, thereby further improving the result accuracy.

[0020] In one implementation of the present application, the reserved intervals are divided into 2 to 8 groups according to the GC content.

[0021] In one implementation of the present application, the reserved intervals are divided into 6 groups according to the GC content. The first group is: 30% ≤ GC content < 35%, the second group is: 35% ≤ GC content < 40%, the third group is: 40% ≤ GC content < 45%, the fourth group is: 45% ≤ GC content < 50%, the fifth group is: 50% ≤ GC content < 55%, and the sixth group is: 55% ≤ GC content < 60%.

[0022] In one implementation of the present application, the whole-genome sequencing data is human whole-genome sequencing data.

[0023] In one implementation of the present application, the chromosomal genome is an autosomal genome.

[0024] The second aspect of the present application discloses an apparatus for analyzing mitochondrial copy number, including: an interval division module for obtaining whole-genome sequencing data and dividing the chromosomal genome and the mitochondrial genome into multiple intervals according to a preset length; an interval screening module for performing at least one of the following screening steps on all intervals: gene sequence homology screening, genomic sequencing complex region screening, GC content screening; and retaining the screened intervals; a mitochondrial copy number calculation module for calculating the mitochondrial copy number: calculating the number of intervals ABi of the retained chromosomal genome and the total sequencing depth ADi, calculating the number of intervals MBi of the retained mitochondrial genome and the total sequencing depth MDi, and calculating the mitochondrial copy number Mi according to the following formula:

[0025]

[0026] The third aspect of the present application discloses a device for analyzing mitochondrial copy number, including a memory and a processor. The memory is used to store a program, and the processor realizes the method for analyzing mitochondrial copy number as described in the first aspect of the present application by executing the program stored in the memory.

[0027] The fourth aspect of the present application discloses a computer-readable storage medium, in which a program is stored, and the program can be executed by a processor to realize the method for analyzing mitochondrial copy number as described in the first aspect of the present application.

[0028] The beneficial effects of the present application are as follows:

[0029] In the present application, without a standard product or a reference sample, based on the whole-genome sequencing data, by dividing the chromosomal genome and the mitochondrial genome into multiple intervals and screening out inappropriate intervals, the mitochondrial copy number can be accurately calculated. Brief Description of the Drawings

[0030] Figure 1 It is a schematic flowchart of the method for analyzing mitochondrial copy number involved in the present application.

[0031] Figure 2 It is a schematic flowchart of the method for screening intervals involved in the present application.

[0032] Figure 3 It is a schematic diagram of the device for analyzing mitochondrial copy number involved in the present application.

[0033] Figure 4 It is a schematic diagram of the equipment for analyzing mitochondrial copy number involved in the present application. Detailed Embodiments

[0034] The present invention will be further described in detail below in conjunction with the drawings through specific embodiments. In the following embodiments, many details are described to enable a better understanding of the present application. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other materials and methods. In some cases, some operations related to the present application are not shown or described in the specification, which is to avoid the core part of the present application being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail, and the related operations can be fully understood according to the description in the specification and the general technical knowledge in the art.

[0035] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in an obvious manner by those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment, and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.

[0036] In order to reduce the cost of mitochondrial genome quantification and reduce the influence of standards, experimental techniques, reference samples, etc. on mitochondrial copy number quantification, based on whole-genome next-generation sequencing data, the present application corrects the biases caused by GC content, homologous sequences, and repetitive sequences, so as to calculate a more accurate mitochondrial genome copy number. The advantages of the present application include: 1. There is no need for additional Real-time PCR, and it can be directly calculated from whole-genome sequencing data, saving time and cost; 2. It does not rely on standards and reference samples, saving the cost of purchasing additional standards and screening reference samples; 3. The calculation standard is unified, and the implementation method is simple and fast.

[0037] The present application provides a method, apparatus, device, and computer-readable storage medium for analyzing mitochondrial copy number.

[0038] As described above, the present application relates to a method for analyzing mitochondrial copy number.

[0039] Figure 1 It is a schematic flowchart of the method for analyzing mitochondrial copy number according to the present application.

[0040] As Figure 1 shown, in a specific embodiment, the method for analyzing mitochondrial copy number may include: obtaining whole-genome sequencing data (step S10), dividing the chromosome and mitochondrial genomes into multiple intervals according to a preset length (step S20), screening all intervals and retaining the screened intervals (step S30), and calculating the mitochondrial copy number based on the number of retained intervals of the chromosomal genome and the total sum of their sequencing depths, as well as the number of retained intervals of the mitochondrial genome and the total sum of their sequencing depths (step S40).

[0041] In a specific embodiment, in step S10, the whole-genome sequencing data is human whole-genome sequencing data. The whole-genome sequencing data is the whole-genome sequencing data of normal and healthy humans. The whole-genome sequencing data refers to whole-genome second-generation sequencing data, and the whole-genome sequencing data includes chromosomal genome sequencing data and mitochondrial genome sequencing data. The chromosomal genome sequencing data and the mitochondrial genome sequencing data are obtained in the same experimental batch. In a specific embodiment, the chromosomal genome may refer to the autosomal genome, and when the gender is known, the chromosomal genome may also include the sex chromosome genome. In a specific embodiment, the file format of the obtained whole-genome sequencing data is the BAM format. The BAM (Binary Alignment / Map) file is a standard file format in high-throughput sequencing data analysis, used to store the information of aligned sequencing reads. BAM files are widely used in fields such as genomics, transcriptomics, and epigenetics. Especially in analyses such as alignment, variant detection, and expression quantification, through some common tools, it is possible to easily process and analyze BAM files to extract the target information. The file format of the obtained whole-genome sequencing data may also be the FASTQ format.

[0042] In a specific embodiment, in step S20, the preset length may be 50 bp to 500 bp. For example, the chromosomal genome and mitochondrial genome can be divided into multiple intervals according to sizes of 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, or 500 bp. Considering that the mitochondrial genome is relatively small, preferably, the preset length is 80 bp to 200 bp. More preferably, the preset length is 100 bp. For example, for whole-genome sequencing data, the chromosomes and mitochondria can be divided into multiple intervals according to a size of 100 bp. For example, the first 100 bases of chromosome 1 are the first interval, the 101st to 200th bases of chromosome 1 are the second interval, and so on, dividing all chromosomal genomes and mitochondrial genomes into multiple intervals.

[0043] In a specific embodiment, in step S30, screening all intervals includes performing at least one of the following screening steps on all intervals: gene sequence homology screening, screening of complex regions of genome sequencing, and GC content screening. Preferably, screening all intervals includes performing all of the following screening steps on all intervals: gene sequence homology screening, screening of complex regions of genome sequencing, and GC content screening.

[0044] Reference Figure 2 , Figure 2 is a schematic flowchart of the method for screening intervals involved in the present application. Screening all intervals may include: gene sequence homology screening (step S31), screening of complex regions of genome sequencing (step S32), and GC content screening (step S33). It should be noted that there is no order requirement for step 31, step S32, and step S33. For example, it can be performed in the order of step S31, step S32, and step S33, or in the order of step S31, step S33, and step S32, or in the order of step S32, step S31, and step S33, or in the order of step S32, step S33, and step S31, or in the order of step S33, step S31, and step S32, or in the order of step S33, step S32, and step S31.

[0045] In a specific embodiment, in step S31, the steps of gene sequence homology screening may include: calculating the unique comparison score for each interval and retaining the intervals with a unique comparison score greater than or equal to a preset value. The unique comparison score, also known as the unique alignment rate, refers to the proportion of reads in the sequencing data that can be uniquely aligned to the genome, usually between 0 and 1, representing the proportion of successful alignments. For example, 0.9 means that 90% of the reads can be uniquely aligned to the genome.

[0046] In a specific embodiment, in step S31, a human genome uniquely mappable interval file can be obtained from the UCSC website (address: https: / / hgdownload.cse.ucsc.edu / goldenPath / hg19 / encodeDCC / wgEncodeMapability / wgEncodeCrgMapabilityAlign50mer.bigWig). The uniquely mappable score of all intervals is calculated using the bigWigAverageOverBed tool (address: https: / / hgdownload.cse.ucsc.edu / admin / exe / linux.x86_64 / bigWigAverageOverBed) to obtain the uniquely mappable score of each interval.

[0047] In a specific embodiment, the preset value is 0.95 - 1. In other words, intervals with a uniquely mappable score greater than or equal to 0.95 can be retained. Preferably, the preset value is 1.

[0048] In a specific embodiment, in step S32, the steps for screening genomic sequencing complex regions may include: comparing each interval with the genomic sequencing complex regions and retaining the intervals whose positions do not overlap with the genomic sequencing complex regions. The genomic sequencing complex regions may include at least one of the following: repetitive sequence regions, regions with abnormal GC content, highly conserved regions, chromosomal terminal and centromere regions, pseudogene regions, and regions prone to sequencing analysis deviation. Preferably, the genomic sequencing complex region is the ENCODE blacklist region. A main feature of the ENCODE blacklist region is that it shows strong signal enrichment in multiple experiments and is caused by systematic biases in experimental techniques or data analysis processes. The ENCODE gene sequencing blacklist region file can be downloaded from the github repo (https: / / github.com / Boyle-Lab / Blacklist).

[0049] In a specific embodiment, in step S33, the steps for GC content screening may include: retaining the intervals with a GC content greater than or equal to 30% and less than or equal to 60%. It should be noted that regions with too low or too high GC content are usually regions with a relatively high degree of repetition or some complex regions. By excluding regions with too low or too high GC content, it is beneficial to correct the deviation in mitochondrial copy number analysis caused by GC content.

[0050] In a specific embodiment, step S40 includes: calculating the number of intervals ABi of the retained chromosomal genome and the total sequencing depth ADi, calculating the number of intervals MBi of the retained mitochondrial genome and the total sequencing depth MDi, and calculating the mitochondrial copy number Mi according to the following formula:

[0051]

[0052] It should be noted that the sequencing depth of a certain interval refers to the average number of reads at each position in this region. For example, the number of reads at the first base of the coordinate is 10, and the number of reads at the second base of the coordinate is 15. Then the total number of reads in the interval from coordinate 1 to 2 is (10 + 15) / 2, that is, 12.5. Reads refer to the sequences generated by the sequencer, and one read refers to a short sequence output by the sequencer.

[0053] In a specific embodiment, in step S40, the retained intervals can be grouped according to the GC content, and the mitochondrial copy number of each group is calculated separately. The final mitochondrial copy number is the average or median of the mitochondrial copy numbers of each group. For example, the retained intervals can be divided into 2 to 8 groups according to the GC content. For example, the retained intervals can be divided into 6 groups according to the GC content. The first group is: 30% ≤ GC content < 35%, the second group is: 35% ≤ GC content < 40%, the third group is: 40% ≤ GC content < 45%, the fourth group is: 45% ≤ GC content < 50%, the fifth group is: 50% ≤ GC content < 55%, and the sixth group is: 55% ≤ GC content < 60%. Calculate the mitochondrial copy number of each group, and take the average or median of the mitochondrial copy numbers of the six groups as the final mitochondrial copy number result.

[0054] This application also relates to a device for analyzing mitochondrial copy number.

[0055] Figure 3 It is a schematic diagram of the device 800 for analyzing mitochondrial copy number involved in this application. As Figure 3 shown, the device 800 for analyzing mitochondrial copy number may include an interval division module 801, an interval screening module 802, and a mitochondrial copy number calculation module 803. The functions of each module can be as follows:

[0056] In a specific embodiment, the interval division module 801 can be used to obtain whole-genome sequencing data and divide the chromosomal genome and the mitochondrial genome into multiple intervals according to a preset length. For the specific functions, reference can be made to the corresponding content in the aforementioned method for analyzing mitochondrial copy number, which will not be elaborated here.

[0057] In a specific embodiment, the interval screening module 802 can be used to perform at least one of the following screening steps on all intervals: gene sequence homology screening, screening of complex regions in genome sequencing, GC content screening; and retain the screened intervals. For the specific functions, reference can be made to the corresponding content in the aforementioned method for analyzing mitochondrial copy number, which will not be elaborated here.

[0058] In a specific embodiment, the mitochondrial copy number calculation module 803 can be used to calculate the mitochondrial copy number: calculate the number of intervals ABi of the retained chromosomal genome and the total sequencing depth ADi, calculate the number of intervals MBi of the retained mitochondrial genome and the total sequencing depth MDi, and calculate the mitochondrial copy number Mi according to the formula. For the specific functions, reference can be made to the corresponding content in the aforementioned method for analyzing mitochondrial copy number, which will not be elaborated here.

[0059] This application also relates to a device for analyzing mitochondrial copy number.

[0060] Figure 4 It is a schematic diagram of the device 100 for analyzing mitochondrial copy number involved in this application. As Figure 4 shown, the device 100 for analyzing mitochondrial copy number can include a processor 10, a memory 20, and a computer program 21 (which can also be referred to as a computer-readable storage medium 21) stored in the memory 20. The computer program 21 can run on the processor 10. When the processor 10 executes the computer program 21, it implements the aforementioned method for analyzing mitochondrial copy number, for example, implementing Figure 1 each step shown. Or, when the processor 10 executes the computer program 21, it implements the functions of each module in the aforementioned device 800 for analyzing mitochondrial copy number, for example Figure 3 the functions of the module 801, module 802, and module 803 shown.

[0061] The device 100 for analyzing mitochondrial copy number can be a computing device such as a desktop computer, a notebook, a handheld computer, and a cloud server. The device 100 for analyzing mitochondrial copy number can include, but is not limited to, the processor 10 and the memory 20. Those skilled in the art can understand that Figure 4 this is only an example of the device 100 for analyzing mitochondrial copy number, and does not constitute a limitation on the device 100 for analyzing mitochondrial copy number. For example, it can include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the device 100 for analyzing mitochondrial copy number can also include input / output devices, network access devices, buses, etc.

[0062] The processor 10 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0063] The memory 20 may be an internal storage unit of the device 100 for analyzing mitochondrial copy number, for example, a hard disk or a memory. The memory 20 may also be an external storage device of the device 100 for analyzing mitochondrial copy number, such as a plug-in hard disk equipped on the device 100 for analyzing mitochondrial copy number, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may include both an internal storage unit and an external storage device of the device 100 for analyzing mitochondrial copy number.

[0064] The memory 20 may be used to store the computer program 21 and other programs and data required by the device 100 for analyzing mitochondrial copy number. The memory 20 may also be used to temporarily store the data that has been output or will be output.

[0065] The present application will be further described in detail below through examples and comparative examples. The following examples and comparative examples only further illustrate the present application and should not be construed as a limitation of the present application. In this example, unless otherwise specified, the reagents, instruments, software, or tools used are all obtained from ordinary commercial sources or open-source databases, and the operations are all carried out in accordance with the product manuals and conventional operation specifications.

[0066] Example 1:

[0067] (1) Use the ART tool on the NCBI website (address:

[0068] https: / / www.niehs.nih.gov / research / resources / software / biostatistics / art) to generate whole-genome simulated sample data, which is a bam file of a healthy person with a mitochondrial copy number of 300.

[0069] (2) The chromosome length list of the human reference genome hg19 version was obtained from the UCSC website (address: https: / / hgdownload.cse.ucsc.edu / goldenpath / hg19 / bigZips / hg19.chrom.sizes). Chromosomes 1 to 22 and the mitochondrial genome of the hg19 version of the human reference genome were divided into 100 bp intervals, of which the mitochondrial genome was divided into 166 intervals and chromosomes 1 to 22 were divided into 28,810,343 intervals.

[0070] (3) Obtain the human genome unique alignment interval file from the UCSC website (address: https: / / hgdownload.cse.ucsc.edu / goldenPath / hg19 / encodeDCC / wgEncodeMapability / wgEncodeCrgMapabilityAlign50mer.bigWig). Use the bigWigAverageOverBed tool (https: / / hgdownload.cse.ucsc.edu / admin / exe / linux.x86_64 / bigWigAverageOverBed) to calculate the unique alignment scores of all intervals. Finally, 20 intervals were retained for the mitochondrial genome, and 19,173,604 intervals were retained for chromosomes 1 to 22. Table 1 shows the unique alignment score results for each interval of the mitochondrial genome (due to the large length of the table, only the results of some intervals are shown):

[0071] Table 1

[0072] Chromosome Genomic start position Genomic end position Unique comparison score chrMT 0 100 0.94 chrMT 100 200 1 chrMT 200 300 1 chrMT 300 400 0.945 chrMT 400 500 1 chrMT 500 600 0.961667 chrMT 600 700 0.411167 chrMT 700 800 0.531167 chrMT 800 900 0.27375 chrMT 900 1000 0.774333

[0073] (4) Download the ENCODE gene sequencing blacklist region file from the github repo (https: / / github.com / Boyle-Lab / Blacklist) (as shown in Table 2, due to the large length of the table, only part of the results are shown schematically), use the bedtools intersect tool, use the -v parameter, and obtain the intervals with an overlap length of 0 with the ENCODE hg19 version blacklist region. None of the mitochondrial genome intervals overlap with the blacklist, so there are still 20 intervals retained, and 19,116,562 intervals are retained for chromosomes 1 to 22.

[0074] Table 2

[0075] Chromosome Genomic start position Genomic end position Type chr10 38726200 42489100 HighSignalRegion chr10 42524900 42819200 HighSignalRegion chr10 98560400 98562500 HighSignalRegion chr10 135437600 135534700 HighSignalRegion chr11 0 196300 HighSignalRegion chr11 584400 586500 HighSignalRegion chr11 964000 966100 LowMappability chr11 1015700 1019100 HighSignalRegion chr11 1088800 1094300 HighSignalRegion chr11 1141100 1214300 HighSignalRegion

[0076] (5) Using the bedtools tool and the hg19 reference genome fasta, calculate the GC content of all remaining intervals. Use the mosdetph tool (address: https: / / github.com / brentp / mosdepth) to calculate the average sequencing depth of the remaining intervals in the mitochondrial genome and chromosomes 1 to 22. Table 3 shows the GC content and average sequencing depth of each interval in the mitochondrial genome.

[0077] Table 3

[0078]

[0079] (6) Group the remaining intervals according to the GC content, and divide them into six groups. The first group is: 30% ≤ GC content < 35%, the second group is: 35% ≤ GC content < 40%, the third group is: 40% ≤ GC content < 45%, the fourth group is: 45% ≤ GC content < 50%, the fifth group is: 50% ≤ GC content < 55%, and the sixth group is: 55% ≤ GC content < 60%. Calculate the total mitochondrial genome sequencing depth MDi of each group, the total number of mitochondrial intervals MBi of each group, the total autosomal sequencing depth ADi of each group, and the total number of autosomal intervals ABi of each group. The calculation method of the mitochondrial copy number of each group is Finally, the mitochondrial copy numbers of the 6 groups are calculated, and the median is taken from them to obtain the final mitochondrial copy number. Specifically, the intervals belonging to the first group of the GC content of the mitochondrial genome are the two intervals of 200 - 300 and 10000 - 10100. The total sequencing depth MDi of these two intervals is 9583.47, the number of intervals MBi is 2, the total sequencing depth ADi of the first group intervals of chromosomes 1 - 22 is 686130.99, and the number of intervals ABi is 22000. According to the formula, the calculated mitochondrial copy number of the first group is 307.26. In the same way, the calculated mitochondrial copy numbers of the 2nd to 6th groups are 356.40, 317.12, 244.24, 267.86, and 294.20 respectively. Therefore, the final mitochondrial copy number is the median (294.20 + 307.26) / 2, that is, 300.73. This value is very close to the true copy number value of 300. It can be seen that this embodiment can reliably calculate the copy number of the human mitochondrial genome.

[0080] Example 2:

[0081] Example 2 uses another tool - the bamsurgeon tool (address:

[0082] Generate whole-genome simulated sample data on https: / / github.com / adamewing / bamsurgeon). The sample data is a bam file of a healthy person with a mitochondrial copy number of 300. The remaining steps are the same as those in Example 1.

[0083] In Example 2, the finally calculated mitochondrial copy number is 304.13, which is very close to the true copy number value of 300. It can be seen that the method for calculating the mitochondrial copy number in this application has good stability and accuracy.

[0084] Example 3:

[0085] Compared with Example 1, step (3) is not carried out in Example 3, and the rest is the same as that in Example 1. That is, in Example 3, the intervals with an alignment score not equal to 1 are not removed, and the rest of the steps are the same as those in Example 1. The calculated mitochondrial copy number in Example 3 is 271.24, which is relatively close to the true copy number value of 300, but not as good as Example 1.

[0086] Example 4:

[0087] Compared with Example 1, step (4) is not carried out in Example 4, and the rest is the same as that in Example 1. That is, in Example 4, the intervals overlapping with the ENCODE blacklist region are not removed, and the rest of the steps are the same as those in Example 1. The calculated mitochondrial copy number in Example 4 is 258.31, which is close to the true copy number value of 300, but not as good as Example 1.

[0088] Example 5:

[0089] Compared with Example 1, in Example 5, the intervals are not grouped according to the GC content to calculate the mitochondrial copy number and take the median. Instead, the sum of the sequencing depths of all remaining intervals and the total number of intervals are directly calculated to obtain the mitochondrial copy number. The rest of the steps are the same as those in Example 1. The calculated mitochondrial copy number in Example 5 is 280.46, which is close to the true copy number value of 300, but not as good as Example 1.

[0090] Comparative Example 1:

[0091] Compared with Example 1, in Comparative Example 1, the intervals with an alignment score not equal to 1 are not removed, the intervals overlapping with the ENCODE blacklist region are not removed, and the intervals are not grouped according to the GC content to calculate the mitochondrial copy number and take the median. The calculated mitochondrial copy number in Comparative Example 1 is 412.72, which is extremely different from the true copy number value of 300.

[0092] In summary, in the present application, based on the whole-genome next-generation sequencing data, the deviations caused by homologous sequences, repetitive sequences, and / or GC content are excluded, which can help to calculate a result close to the true mitochondrial genome copy number.

[0093] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0094] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0095] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the module or unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0096] The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0097] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0098] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such understanding, all or part of the processes in the above-described embodiment methods of the present application may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0099] The above content is a further detailed description of the present application in combination with specific implementation manners, and it cannot be determined that the specific implementation of the present application is only limited to these descriptions. For those of ordinary skill in the technical field to which the present application belongs, without departing from the concept of the present application, several simple deductions or substitutions can also be made.

Claims

1. A method for analyzing mitochondrial copy number, characterized in that: include: Obtain whole genome sequencing data, and divide the chromosome genome and mitochondrial genome into multiple intervals according to preset lengths; Perform at least one of the following screening steps on all intervals: gene sequence homology screening, genome sequencing complex region screening, GC content screening; and retain the screened intervals; Calculate the mitochondrial copy number: calculate the number of intervals ABi and the total sequencing depth ADi of the retained chromosome genome, calculate the number of intervals MBi and the total sequencing depth MDi of the retained mitochondrial genome, and calculate the mitochondrial copy number Mi according to the following formula:

2. The method according to claim 1, characterized in that The preset length is 50 bp to 500 bp; Preferably, the preset length is 80 bp to 200 bp.

3. The method according to claim 1, characterized in that: The step of screening for gene sequence homology includes: calculating a unique comparison score for each interval, and retaining intervals with a unique comparison score greater than or equal to a preset value; Preferably, the preset value is 0.95 to 1; Preferably, the preset value is 1.

4. The method according to claim 1, characterized in that: The step of screening the complex region of genome sequencing includes: comparing each interval with the complex region of genome sequencing, and retaining the interval whose position does not overlap with the complex region of genome sequencing; Preferably, the complex genome sequencing region includes at least one of a repetitive sequence region, an abnormal GC content region, a highly conserved region, a chromosome end and centromere region, a pseudogene region, and a region prone to sequencing analysis deviation; Preferably, the complex genome sequencing region is an ENCODE blacklist region.

5. The method according to any one of claims 1 to 4, characterized in that: The step of GC content screening includes: retaining intervals with GC contents greater than or equal to 30% and less than or equal to 60%; Preferably, all the following screening steps are performed on all intervals: the gene sequence homology screening, the genome sequencing complex region screening, and the GC content screening.

6. The method according to claim 5, characterized in that The calculating of the mitochondrial copy number comprises: grouping the reserved intervals according to the GC content, respectively calculating the mitochondrial copy number of each group, and the final mitochondrial copy number is the average or median of the mitochondrial copy number of each group; Preferably, the retained intervals are divided into 2 to 8 groups according to GC content; Preferably, the retained intervals are divided into 6 groups according to the GC content, the first group being: 30% ≦ GC content < 35%, the second group being: 35% ≦ GC content < 40%, the third group being: 40% ≦ GC content < 45%, the fourth group being: 45% ≦ GC content < 50%, the fifth group being: 50% ≦ GC content < 55%, and the sixth group being: 55% ≦ GC content < 60%.

7. The method according to claim 1, characterized in that The whole genome sequencing data is human whole genome sequencing data; Preferably, the chromosomal genome is an autosomal genome.

8. A device for analyzing mitochondrial copy number, characterized in that: include: An interval division module is used to obtain whole genome sequencing data and divide the chromosome genome and mitochondrial genome into multiple intervals according to preset lengths; An interval screening module is used to perform at least one of the following screening steps on all intervals: gene sequence homology screening, genome sequencing complex region screening, and GC content screening; and retain the screened intervals; The mitochondrial copy number calculation module is used to calculate the mitochondrial copy number: calculate the number of intervals ABi and the total sequencing depth ADi of the retained chromosome genome, calculate the number of intervals MBi and the total sequencing depth MDi of the retained mitochondrial genome, and calculate the mitochondrial copy number Mi according to the following formula:

9. A device for analyzing mitochondrial copy number, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store a program, and the processor implements the method for analyzing mitochondrial copy number according to any one of claims 1 to 7 by executing the program stored in the memory.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program can be executed by a processor to implement the method for analyzing mitochondrial copy number as described in any one of claims 1 to 7.