Methods, devices, and storage media for detecting sequences of a mitochondrial-derived nuclear genome
By preprocessing and clustering the sequencing data, and combining the statistics of inconsistent alignment reads, the accuracy and false positive problems of non-reference mitochondrial nuclear genome sequence detection were solved, achieving efficient and accurate detection of non-ref NUMTs and identification of their cumulative mutations.
Patent Information
- Application Number
- CN202310565670.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-05-17
AI Technical Summary
Existing technologies are insufficient for accurately detecting non-ref mitochondrial nuclear genome sequences and their cumulative mutations, and suffer from high false positive rates and insufficient detection accuracy.
Reference NUMTs were removed by sequencing data preprocessing. Clustering and assembly were performed using alignment information of potential linker reads. Inconsistent alignment read pair statistics were combined to determine the homozygous or heterozygous status of non-ref NUMTs. Accumulated mutations were identified through consistency sequence verification and alignment.
It enables accurate detection of mitochondrial-origin nuclear genome sequences from non-reference genomes, reduces false positive rates, improves detection accuracy and efficiency, simplifies the operation process, and reduces time costs.
Smart Images

Figure CN116665775B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of nuclear genome sequence detection technology, and in particular to a method, apparatus and storage medium for detecting nuclear genome sequences of mitochondrial origin. Background Technology
[0002] The genetic material of human cells is located in the nucleus and mitochondria. Nuclear DNA (nDNA) consists of approximately 3.1 billion base pairs and includes 22 pairs of autosomes and 2 sex chromosomes. Mitochondrial DNA (mtDNA) is a double-stranded circular molecule with a length of 16,569 bp. The integration of mtDNA fragments into nDNA is an inevitable result of endosymbiotic events; these integrated mtDNA fragments are called nuclear fragments of mitochondrial origin (NUMTs). Most NUMTs are ancient and neutral, products of long-term cellular evolution, and are documented in the human reference genome. NUMTs documented in the human reference genome are called reference genome mitochondrial origin nuclear genome sequences (ref-NUMTs), while those NUMTs that have recently occurred in non-reference genomes are called non-reference genome mitochondrial origin nuclear genome sequences (non-ref NUMTs). Non-ref NUMTs, especially those in somatic cells, can affect the stability of the nuclear genome and the expression of related genes, and have been reported to be associated with the occurrence and development of various human diseases. Furthermore, after non-ref NUMTs occur, new mutations accumulate in nDNA, which are often mistaken for mtDNA mutations, greatly affecting mtDNA mutation detection and subsequent disease-related investigations. Therefore, detecting NUMTs, especially non-ref NUMTs, and their accumulated mutations is crucial for understanding the development and progression of human diseases.
[0003] Currently, FISH (fluorescence in situ hybridization) technology can effectively detect non-ref NUMTs by performing sequence hybridization between nDNA and mtDNA. Based on this, Koo DH et al. developed "mtFIBER FISH" specifically for detecting non-ref NUMTs, but its resolution is limited, only able to detect non-ref NUMTs with insert lengths >1kb; most non-ref NUMTs have a fragment length <1kb and cannot be detected by mtFIBER FISH. At the same time, mtFIBER FISH also cannot detect accumulated mutations in non-ref NUMTs.
[0004] Whole genome sequencing (WGS) is the technology to date capable of detecting non-ref NUMTs and their accumulated mutations at single-base resolution. With the development of sequencing technology, the cost of WGS has decreased exponentially, and numerous studies both domestically and internationally have accumulated massive amounts of WGS data, which can aid in the study of non-ref NUMTs. However, in conventional 30-50× WGS data, the mtDNA coverage depth can reach tens of thousands of layers, exhibiting extremely high redundancy and heterogeneity, making the detection of non-ref NUMTs very challenging, and currently available tools are very limited. A recently published tool, NUMT-detection, can detect non-ref NUMTs using WGS; however, because its detection principle mainly relies on the alignment information of inconsistent read pairs, it has a high false positive rate, limited accuracy in inferred breakpoints, and cannot detect accumulated mutations on non-ref NUMTs.
[0005] In summary, in the field of mitochondrial origin nuclear genome sequencing technology, how to accurately and effectively detect non-refNUMTs remains a key research focus and challenge; moreover, there are currently no tools or methods that can simultaneously detect non-refNUMTs and their cumulative mutations. Summary of the Invention
[0006] The purpose of this application is to provide a new method, apparatus, and storage medium for detecting mitochondrial-origin nuclear genome sequences.
[0007] To achieve the above objectives, this application adopts the following technical solution:
[0008] The first aspect of this application discloses a method for detecting mitochondrial-origin nuclear genome sequences, comprising the following steps:
[0009] The sequencing data preprocessing steps include: 1) aligning the whole genome sequencing data to the mitochondrial reference genome and retaining only the reads that can be aligned to the mitochondrial reference genome; 2) aligning the reads aligned to the mitochondrial reference genome to a reference sequence containing the 23 pairs of human chromosomes and the mitochondrial reference genome, removing non-unique aligned reads, and obtaining a read set without ref-NUMTs.
[0010] The steps for sequencing the mitochondrial-origin nuclear genome include: 1) Extracting a portion of the reads from the set of reads without ref-NUMTs and aligning them to the nuclear genome reference sequence, while aligning the remaining portion to the mitochondrial reference genome, as potential linker reads; 2) Clustering potential linker reads within a 50bp range based on their alignment positions, identifying the coordinates and orientation of the integrated mtDNA fragments and the location of nuclear genome integration, and classifying the integrated mtDNA fragments as non-ref NUMTs; 3) Searching for inconsistent alignment pairs within 100bp upstream and downstream of the read clusters, where one read aligns to the mitochondrial reference genome and the other to the nuclear genome reference sequence, and counting the number of inconsistent alignment pairs, which are the non-ref read pairs. 4) Calculate whether non-refNUMTs are homozygous or heterozygous, i.e., the ratio of the number of supporting reads to the average nuclear genome coverage depth. The number of supporting reads is the sum of the number of potential linking reads and the number of inconsistent alignment reads. The average nuclear genome coverage depth is the number of reads aligned to the nuclear genome reference sequence multiplied by the read length and then divided by the length of the nuclear genome reference sequence.
[0011] In this application, a non-unique alignment read refers to a read that can be aligned to both the mitochondrial reference genome and the reference sequence of the 23 pairs of human chromosomes. Such reads can be considered potential ref-NUMTs sequences. After removing them, a set of reads without ref-NUMTs is obtained. It can be understood that the mitochondrial origin nuclear genome sequence detection in this application mainly refers to the detection of non-ref NUMTs, so it is necessary to remove ref-NUMTs that have been clearly recorded in the human reference genome.
[0012] In the mitochondrial origin nuclear genome sequence detection step of this application, the presence and location of non-refNUMTs are first determined according to 1) and 2); in 3), the inconsistent read pairs are located "100 bp upstream and downstream of the read cluster position" and "one read is aligned to mtDNA and the paired read is aligned to nDNA". These two pieces of information can determine that these inconsistent read pairs are the result of the presence of non-refNUMTs.
[0013] In the mitochondrial-origin nuclear genome sequence detection steps of this application, in step 4), at the nuclear genome location where non-ref NUMTs are inserted, if the non-ref NUMTs are heterozygous, one chromosome has a normal sequence, and the other has an inserted mtDNA fragment. Approximately half of the reads covering this location are supporting reads determined by steps 1) and 3), and the other half are normally aligned reads. That is, the ratio of the number of supporting reads to the average coverage depth of the nuclear genome is approximately 0.5. If it is homozygous, then it is almost entirely covered by supporting reads, and this ratio is approximately 1. Determining whether it is homozygous or heterozygous can determine whether non-ref NUMTs affect one or two nDNAs, similar to single-base mutation type, genotyping. It also facilitates subsequent determination of the base frequency of accumulated mutations on non-ref NUMTs. Specifically, the frequency of accumulated mutations on heterozygous non-ref NUMTs is 1 / (1 + mtDNA copy number), while the frequency of accumulated mutations on homozygous non-ref NUMTs is 2 / (2 + mtDNA copy number).
[0014] It should be noted that the method for detecting mitochondrial-origin nuclear genome sequences in this application fully utilizes the alignment information of potential linker reads. Through local assembly and clustering, it accurately detects non-ref NUMTs, reducing the false positive rate of current non-ref NUMT detection and obtaining precise breakpoint and fragment information. The detection method in this application is simple and convenient to operate, and can effectively reduce the time cost of non-ref NUMT analysis.
[0015] In one implementation of this application, the method for detecting mitochondrial nuclear genome sequences further includes a mitochondrial nuclear genome sequence verification step, which includes: 1) assembling sequences from reads or read pairs that support the existence of non-ref NUMTs and are aligned to the mitochondrial reference genome to generate a consistent sequence; 2) aligning the generated consistent sequence to the mitochondrial reference genome and verifying non-ref NUMTs based on their alignment positions.
[0016] In one implementation of this application, the method for detecting mitochondrial-origin nuclear genome sequences further includes a step of detecting cumulative mutations in the mitochondrial-origin nuclear genome sequence. This step includes identifying mismatched bases based on the alignment of the assembled homologous sequence to a mitochondrial reference genome, thereby obtaining the cumulative mutations in non-ref NUMTs. For heterozygous non-ref NUMTs, the frequency of cumulative mutations is 1 / (1 + mtDNA copy number), and for homozygous non-ref NUMTs, the frequency of cumulative mutations is 2 / (2 + mtDNA copy number).
[0017] It should be noted that the assembled consensus sequence serves two purposes: first, it can verify the existence of non-ref NUMTs; second, by comparing the consensus sequence with the mitochondrial reference genome, it can accurately identify the accumulated mutations of non-ref NUMTs.
[0018] In one implementation of this application, the method for detecting mitochondrial-origin nuclear genome sequences further includes an annotation step, which includes annotating the breakpoint locations of the nuclear genome reference sequence and the regions and genes containing non-ref NUMTs.
[0019] It should be noted that the annotation step mainly involves annotating the non-ref NUMTs detected in the previous steps, as well as the nDNA breakpoints where the non-ref NUMTs align to the nuclear genome reference sequence.
[0020] The second aspect of this application discloses an apparatus for detecting mitochondrial-origin nuclear genome sequences, which includes a sequencing data preprocessing module and a mitochondrial-origin nuclear genome sequence detection module;
[0021] The sequencing data preprocessing module includes: 1) aligning whole-genome sequencing data to the mitochondrial reference genome, retaining only reads that can be aligned to the mitochondrial reference genome; 2) aligning reads aligned to the mitochondrial reference genome to a reference sequence containing the 23 pairs of human chromosomes and the mitochondrial reference genome, removing non-unique aligned reads, and obtaining a read set without ref-NUMTs.
[0022] The mitochondrial-origin nuclear genome sequence detection module includes the following functions: 1) Extracting a portion of the reads from a set of reads without ref-NUMTs, aligning them to the nuclear genome reference sequence, and aligning the remaining portion to the mitochondrial reference genome, as potential linker reads; 2) Clustering potential linker reads within a 50bp range based on their alignment positions, identifying the coordinates and orientation of the integrated mtDNA fragments, as well as the location of nuclear genome integration, and classifying the integrated mtDNA fragments as non-ref NUMTs; 3) Searching for inconsistent alignment pairs within 100bp upstream and downstream of the read clusters, i.e., one read aligns to the mitochondrial reference genome and the other paired read aligns to the nuclear genome reference sequence, and counting the number of inconsistent alignment pairs, which constitute the supporting information for non-ref NUMTs; 4) Calculating non-ref... Whether NUMTs are homozygous or heterozygous is determined by the ratio of the number of supporting reads to the average nuclear genome coverage depth. The number of supporting reads is the sum of the number of potential linking reads and the number of inconsistent alignment reads. The average nuclear genome coverage depth is the number of reads aligned to the nuclear genome reference sequence multiplied by the read length and then divided by the length of the nuclear genome reference sequence.
[0023] It should be noted that the device for detecting the mitochondrial nuclear genome sequence in this application is actually implemented by each module in the method for detecting the mitochondrial nuclear genome sequence in this application; therefore, the specific limitations of each module can be referred to the method of this application, and will not be repeated here.
[0024] In one implementation of this application, to further ensure the accuracy of non-ref NUMTs, the device for detecting mitochondrial nuclear genome sequences further includes a mitochondrial nuclear genome sequence verification module, which includes: 1) assembling sequences from reads or read pairs that support the existence of non-ref NUMTs and align to the mitochondrial reference genome portion to generate a congruent sequence; 2) aligning the generated congruent sequence to the mitochondrial reference genome and verifying non-ref NUMTs based on their alignment positions.
[0025] In one implementation of this application, in order to further detect the accumulated mutations in the mitochondrial nuclear genome sequence, the device for detecting the accumulated mutations in the mitochondrial nuclear genome sequence further includes a mitochondrial nuclear genome sequence accumulated mutation detection module, which includes a module for identifying mismatched bases based on the result of aligning the assembled consistent sequence to the mitochondrial reference genome, i.e., obtaining the accumulated mutations of non-ref NUMTs.
[0026] In one implementation of this application, the device for detecting mitochondrial-origin nuclear genome sequences further includes an annotation module, which includes regions and genes containing nuclear genome reference sequence breakpoints and non-ref NUMTs.
[0027] A third aspect of this application discloses an apparatus comprising a memory and a processor; the memory including a program for storing a program; and the processor including a method for detecting mitochondrial-origin nuclear genome sequences by executing the program stored in the memory.
[0028] The fourth aspect of this application discloses a computer-readable storage medium storing a program that can be executed by a processor to implement the method of detecting mitochondrial-origin nuclear genome sequences of this application.
[0029] Due to the adoption of the above technical solutions, the beneficial effects of this application are as follows:
[0030] This application presents a method and apparatus for detecting mitochondrial-origin nuclear genome sequences. Utilizing alignment information from potential linker reads, through local assembly and clustering, it accurately detects non-ref NUMTs, reducing the false positive rate and providing precise breakpoint and fragment information. The detection method is simple and convenient to operate, effectively reducing the time cost of non-ref NUMT analysis. Attached Figure Description
[0031] Figure 1 This is a flowchart of the method for detecting non-ref NUMTs in the embodiments of this application;
[0032] Figure 2 This is a structural block diagram of the apparatus for detecting non-ref NUMTs in the embodiments of this application;
[0033] Figure 3 This is a schematic diagram illustrating the principle of detecting non-ref NUMTs and their cumulative mutations in the embodiments of this application;
[0034] Figure 4 This is an IGV map of some non-ref NUMTs detected in the embodiments of this application;
[0035] Figure 5 The non-ref NUMTs detected by the non-ref NUMTs detection method in this application have an insertion sequence of ~530bp at position chr11:49883569 in the third-generation data;
[0036] Figure 6 This refers to the positions and mismatched bases of some non-ref NUMTs detected in the embodiments of this application that are aligned to rCRS;
[0037] Figure 7 This is a distance analysis result between the non-ref NUMTs detection method in this application embodiment and the breakpoints of non-ref NUMTs detected by the prior art NUMT-detection and the breakpoints obtained from three generations of data;
[0038] Figure 8 This is the highest confidence non-ref NUMT IGV map in second-generation data obtained using the existing NUMT-detection technology in the embodiments of this application;
[0039] Figure 9 This is the verification result of the highest confidence non-ref NUMT obtained by using the existing NUMT-detection technology in the embodiments of this application in three generations of data. Detailed Implementation
[0040] The present application will now be described in further detail with reference to specific embodiments and accompanying drawings. In the following embodiments, many details are described to facilitate a better understanding of the present application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other devices, materials, or methods. In some cases, certain operations related to the present application are not shown or described in the specification to avoid obscuring the core parts of the application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; a complete understanding of the related operations can be obtained from the description in the specification and general technical knowledge in the art.
[0041] Existing non-ref NUMTs detection methods generally extract misaligned read pairs directly from WGS sequencing data alignment files (BAM) and cluster read pairs with alignment positions within 500 bp into read pair clusters; then, NUMTs detection is performed on read pair clusters consisting of at least two read pairs; NUMTs with alignment distances within 1000 bp on both nDNA and mtDNA are considered the same NUMTs; then, truncated reads are searched within 1000 bp upstream and downstream of the read pair clusters and aligned back to the mtDNA reference sequence using BLAT, thereby inferring potential breakpoints connecting nDNA and mtDNA.
[0042] The shortcomings of existing detection methods include a high false positive rate, low accuracy in defining breakpoints by searching for join reads upstream and downstream of 1000 bp, and the inability to determine the orientation of the mtDNA integration fragment, which can lead to errors in the integration fragment length information. Furthermore, current technologies can only detect non-ref NUMT; accumulated mutations on it require additional analysis, which is time-consuming for users unfamiliar with WGS data and bioinformatics analysis workflows.
[0043] This application creatively utilizes information from potential linker reads to detect non-ref NUMTs, effectively leveraging their alignment information to cluster based on truncation positions, reducing clustering distance to 50 bp and significantly decreasing false positives. This application fully utilizes the position and sequence information of the truncation reads aligned to the mtDNA and nuclear genome reference sequences to define breakpoints, achieving accurate breakpoint locations and determining the orientation of the mtDNA integration fragment, thus obtaining the true integration fragment length. Furthermore, in one implementation of this application, when using consistent sequences to verify non-ref NUMTs, accumulated mutations on non-ref NUMTs can be detected simultaneously, making the operation simple and convenient. The principle of this application for detecting non-ref NUMTs and their accumulated mutations is as follows: Figure 3 As shown.
[0044] Based on the above research and understanding, this application proposes a method for detecting the nuclear genome sequence of mitochondrial origin, such as... Figure 1 As shown, it includes sequencing data preprocessing step 11 and mitochondrial origin nuclear genome sequence detection step 12.
[0045] Step 11, the sequencing data preprocessing step, mainly involves removing NUMTs present in the reference genome through alignment, i.e., removing potential ref-NUMTs. Specifically:
[0046] 1) Align the whole genome sequencing data to the mitochondrial reference genome and retain only the reads that can be aligned to the mitochondrial reference genome; specifically, align the WGS sequencing data to the mitochondrial reference genome rCRS (Revised Cambridge reference Sequence, NC_012920) using bwa software and retain only the reads aligned to the rCRS.
[0047] 2) Align the reads aligned to the mitochondrial reference genome to a reference sequence containing the human 23 pairs of chromosomes and the mitochondrial reference genome, and remove the reads that are not uniquely aligned to obtain a read set that does not contain ref-NUMTs; In one implementation of this application, the bwa software is used to align the reads aligned to rCRS to a reference sequence containing the human reference genome and rCRS.
[0048] Step 12 of the mitochondrial origin nuclear genome sequence detection is as follows:
[0049] 1) From the read set that does not contain ref-NUMTs, a portion of the sequences are extracted and aligned to the nuclear genome reference sequence, and the remaining portion is aligned to the rCRS reads as potential linker reads;
[0050] 2) Based on the alignment positions of potential linker reads, cluster potential linker reads within 50 bp as read clusters. Locate the coordinates and orientation of the integrated mtDNA fragment, as well as the location of nuclear genome integration. The integrated mtDNA fragment is designated as non-ref NUMTs. For example, if the first part is soft-cleaved and the second part aligns to an rCRS, i.e., the alignment position of the CIGAR "xxSxxM" read cluster is close to the start coordinates of the integrated mtDNA fragment; while the "xxMxxS" read cluster can locate the termination position of the integrated fragment. The nuclear genome integration breakpoint can be located using the alignment information of all linker reads. Figure 3 As shown.
[0051] 3) Search for inconsistent read pairs within 100 bp upstream and downstream of the read cluster position, i.e., one read is aligned to the mitochondrial reference genome and the other paired read is aligned to the nuclear genome reference sequence. Count the number of inconsistent read pairs. These inconsistent read pairs are the supporting information of non-ref NUMTs.
[0052] 4) Calculate whether non-ref NUMTs are homozygous or heterozygous, i.e., the ratio of the number of supporting reads to the average nuclear genome coverage depth. The number of supporting reads is the sum of the number of potential linking reads and the number of inconsistent alignment reads. The average nuclear genome coverage depth is the number of reads aligned to the nuclear genome reference sequence multiplied by the read length and then divided by the length of the nuclear genome reference sequence.
[0053] Furthermore, such as Figure 1 As shown, the method for detecting the mitochondrial nuclear genome sequence in this application also includes a mitochondrial nuclear genome sequence verification step 13, as follows:
[0054] 1) Assemble sequences that align to the mitochondrial reference genome from reads or read pairs that support the presence of non-ref NUMTs, generating a consistent sequence; in one implementation of this application, the CAP3 software is specifically used to assemble the sequence.
[0055] 2) Align the generated consistent sequence to the mitochondrial reference genome and verify non-ref NUMTs based on their alignment positions; in one implementation of this application, blastn is specifically used to align the consistent sequence to rCRS.
[0056] Furthermore, such as Figure 1 As shown, the method for detecting mitochondrial-origin nuclear genome sequences in this application further includes a step 14 for detecting cumulative mutations in mitochondrial-origin nuclear genome sequences. This step involves identifying mismatched bases, i.e., obtaining the cumulative mutations of non-ref NUMTs, based on the alignment of the assembled homologous sequence to the mitochondrial reference genome. Specifically, for example, blastn is used to align the homologous sequence to rCRS, and the presence of non-ref NUMTs is further verified based on its alignment position. At the same time, mismatched bases, i.e., their cumulative mutations, are identified.
[0057] Furthermore, such as Figure 1 As shown, the method for detecting mitochondrial-origin nuclear genome sequences in this application further includes an annotation step 15, which includes annotating nDNA breakpoint locations and the regions and genes containing non-ref NUMTs. In one implementation of this application, ANNOVAR is specifically used to annotate nDNA breakpoint locations and the regions and genes containing mtDNA integration fragments.
[0058] Those skilled in the art will understand that all or part of the functions of the methods described above can be implemented in hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. Alternatively, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or portable hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.
[0059] Therefore, based on the method for detecting the mitochondrial nuclear genome sequence of this application, this application proposes a device for detecting the mitochondrial nuclear genome sequence, such as... Figure 2 As shown, it includes a sequencing data preprocessing module 21 and a mitochondrial origin nuclear genome sequence detection module 22.
[0060] The sequencing data preprocessing module 21 is used to implement the sequencing data preprocessing step 11 in the method for detecting the mitochondrial nuclear genome sequence of this application, and the mitochondrial nuclear genome sequence detection module 22 is used to implement the mitochondrial nuclear genome sequence detection step 12 in the method for detecting the mitochondrial nuclear genome sequence of this application.
[0061] Furthermore, such as Figure 2 As shown, the apparatus of this application also includes a mitochondrial nuclear genome sequence verification module 23, a mitochondrial nuclear genome sequence cumulative mutation detection module 24, and / or annotation module 25. These three modules are used in sequence to implement the mitochondrial nuclear genome sequence verification step 13, the mitochondrial nuclear genome sequence cumulative mutation detection step 14, and the annotation step 15 in the method for detecting the mitochondrial nuclear genome sequence of this application.
[0062] Furthermore, this application also provides an apparatus comprising a memory and a processor; the memory includes a program for storing data; the processor includes a program for executing the program stored in the memory to implement the following method: a sequencing data preprocessing step, comprising 1) aligning whole-genome sequencing data to a mitochondrial reference genome, retaining only reads that can be aligned to the mitochondrial reference genome; 2) aligning the reads aligned to the mitochondrial reference genome to a reference sequence containing the 23 pairs of human chromosomes and the mitochondrial reference genome, removing non-unique aligned reads to obtain a read set free of ref-NUMTs; a mitochondrial-origin nuclear genome sequence detection step, comprising 1) extracting a portion of the reads from the read set free of ref-NUMTs, aligning them to the nuclear genome reference sequence, and the remaining portion aligning them to the mitochondrial reference genome, as potential linker reads; 2) clustering potential linker reads within 50 bp based on their alignment positions, locating the coordinates and orientation of the integrated mtDNA fragments, as well as the location of nuclear genome integration, and classifying the integrated mtDNA fragments as non-ref-NUMTs. NUMTs; 3) Search for inconsistent alignment pairs within 100 bp upstream and downstream of the read cluster position, i.e., one read is aligned to the mitochondrial reference genome and the other paired read is aligned to the nuclear genome reference sequence. Count the number of inconsistent alignment pairs. These inconsistent alignment pairs are the supporting information of non-ref NUMTs; 4) Calculate whether non-ref NUMTs are homozygous or heterozygous, i.e., the ratio of the number of supporting reads to the average coverage depth of the nuclear genome. Furthermore, the method also includes a mitochondrial-origin nuclear genome sequence verification step, which includes 1) assembling sequences from reads or read pairs that support the existence of non-ref NUMTs and align to the mitochondrial reference genome to generate a congruent sequence; 2) aligning the generated congruent sequence to the mitochondrial reference genome and verifying non-ref NUMTs based on its alignment position; a mitochondrial-origin nuclear genome sequence cumulative mutation detection step, which includes identifying mismatched bases based on the alignment of the assembled congruent sequence to the mitochondrial reference genome, i.e., obtaining the cumulative mutations of non-ref NUMTs; and an annotation step, which includes annotating the breakpoint positions of the nuclear genome reference sequence and the regions and genes where non-ref NUMTs are located.
[0063] Another implementation of this application also provides a computer-readable storage medium, which includes a program executable by a processor to implement the following method: a sequencing data preprocessing step, including 1) aligning whole-genome sequencing data to a mitochondrial reference genome, retaining only reads that can be aligned to the mitochondrial reference genome; 2) aligning the reads aligned to the mitochondrial reference genome to a reference sequence containing the 23 pairs of human chromosomes and the mitochondrial reference genome, removing non-unique aligned reads, and obtaining a read set free of ref-NUMTs; a mitochondrial-origin nuclear genome sequence detection step, including 1) extracting a portion of the reads from the read set free of ref-NUMTs and aligning them to the nuclear genome reference sequence, and the remaining portion aligning to the mitochondrial reference genome, as potential linker reads; 2) clustering potential linker reads within a distance of 50 bp according to their alignment positions, locating the coordinates and orientation of the integrated mtDNA fragment, as well as the location of nuclear genome integration, and classifying the integrated mtDNA fragment as non-ref-NUMTs. NUMTs; 3) Search for inconsistent alignment pairs within 100 bp upstream and downstream of the read cluster position, i.e., one read is aligned to the mitochondrial reference genome and the other read is aligned to the nuclear genome reference sequence. Count the number of inconsistent alignment pairs. These inconsistent alignment pairs are the supporting information of non-ref NUMTs; 4) Calculate whether non-ref NUMTs are homozygous or heterozygous, i.e., the ratio of the number of supporting reads to the average coverage depth of the nuclear genome. Furthermore, the method also includes a mitochondrial-origin nuclear genome sequence verification step, which includes 1) assembling sequences from reads or read pairs that support the existence of non-ref NUMTs and align to the mitochondrial reference genome to generate a congruent sequence; 2) aligning the generated congruent sequence to the mitochondrial reference genome and verifying non-ref NUMTs based on its alignment position; a mitochondrial-origin nuclear genome sequence cumulative mutation detection step, which includes identifying mismatched bases based on the alignment of the assembled congruent sequence to the mitochondrial reference genome, i.e., obtaining the cumulative mutations of non-ref NUMTs; and an annotation step, which includes annotating the breakpoint positions of the nuclear genome reference sequence and the regions and genes where non-ref NUMTs are located.
[0064] The following specific experiments will further illustrate this application in detail. These experiments are merely illustrative and should not be construed as limiting the scope of this application.
[0065] Example
[0066] I. Test Subjects
[0067] Two pairs of esophageal squamous cell carcinoma (ESCC) tumor-normal tissues were used, totaling four samples. In this study, second-generation 150bp paired-end WGS sequencing data and third-generation Nanopore long-read WGS sequencing data from these four samples were used to detect and analyze the mitochondrial-origin nuclear genome sequence and its cumulative mutations. The four samples were numbered as follows: tes200124-PN-1, tes200124-TN-1, tes210075-PN-1, and tes210075-TN-1. For details of the four samples and their second-generation and third-generation WGS data, please refer to the reference: Cui, H., et al., Characterization of somatic structural variations in 528 Chinese individuals with Esophageal squamous cell carcinoma. Nat Commun, 2022, 13(1): p.6296.
[0068] II. Overview of the Experimental Process
[0069] 1) The mitochondrial origin nuclear genome sequence detection method of this application and the existing tool NUMT-detection were used to perform non-ref-NUMTs detection on the second-generation WGS data of the four samples respectively;
[0070] 2) The method for detecting mitochondrial-origin nuclear genome sequences in this application was verified using third-generation WGS data from four samples, comparing it with the non-ref-NUMTs detected by the existing NUMT-detection technique;
[0071] 3) Compare the performance of the detection method of this application with that of the prior art NUMT-detection.
[0072] For details on the existing technology NUMT-detection, please refer to the following reference: Wei, W., et al., Nuclear-embedded mitochondrial DNA sequences in 66,083 human genomes. Nature, 2022, 611(7934): p.105-114.
[0073] III. Specific Implementation Process
[0074] 1. Second-generation WGS data preprocessing: alignment, removal of reference genome NUMTs
[0075] 1) First, the second-generation WGS sequencing data of the four samples were aligned to the mitochondrial reference genome rCRS (Revised Cambridge reference Sequence, NC_012920) using bwa software, and only the reads aligned to the rCRS were retained;
[0076] 2) Then, the reads aligned to rCRS are aligned to the reference sequences of 23 pairs of chromosomes containing hg19 and rCRS using bwa software. Non-unique aligned reads, i.e. potential reference genome NUMTs sequences, are removed to obtain a set of reads that do not contain ref-NUMTs.
[0077] 2. Preprocessing of third-generation WGS data: comparison
[0078] The third-generation WGS sequencing data of four samples were aligned to the reference sequences of 23 pairs of chromosomes containing hg19 and rCRS using the minimap2 software, resulting in a BAM file aligned to the whole genome. For details on the minimap2 software, please refer to the reference: Li, H., Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 2018, 34(18): p.3094-3100.
[0079] 3. Detection and verification of non-ref-NUMTs
[0080] The detection method in this example:
[0081] 1) First, extract potential linker reads, that is, align a portion of the sequence to nDNA and the remainder to rCRS linker reads;
[0082] 2) Based on the alignment position of potential linker reads, linker reads within 50 bp are clustered to locate the coordinates and orientation of the integrated mtDNA fragments, as well as the location of nuclear genome integration. The integrated mtDNA fragments are then designated as non-ref NUMTs.
[0083] 3) Then, look for inconsistent read pairs within 100 bp upstream and downstream of the above read clusters, i.e., one read matches rCRS and the other read matches nDNA. Count the number of such read pairs as supporting information for the existence of non-ref NUMTs.
[0084] 4) Calculate whether non-ref NUMTs are homozygous or heterozygous, i.e., the ratio of the number of supporting reads to the average coverage depth of the nuclear genome.
[0085] Further validation of the cumulative mutation detection method in this example is as follows:
[0086] 1) First, use CAP3 software to assemble reads / read pairs that support the presence of non-ref NUMTs and align them to the rCRS region to generate a consistent sequence; for details of CAP3 software, please refer to the following document: Huang, X. and A. Madan, CAP3: A DNA sequence assembly program. Genome Res, 1999.9(9):p.868-77.
[0087] 2) Then, blastn is used to align the consistent sequence to rCRS, and the existence of non-ref NUMTs is further verified based on the alignment position.
[0088] Furthermore, the detection method in this example also includes a annotation step:
[0089] ANNOVAR was used to annotate nDNA breakpoint locations and regions and genes containing mtDNA integration fragments. For details on ANNOVAR, see the reference: Wang, K., M. Li, and H. Hakonarson, ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res, 2010, 38(16): p.e164.
[0090] Non-ref-NUMTs were detected using the method described in this example and existing NUMT-detection techniques, with all parameters set to default. The method described in this example detected one and four germline non-ref-NUMTs in two ESCC patients, respectively. Specific information obtained using this method is shown in Table 1.
[0091] Table 1 shows the non-ref-NUMTs information obtained by the detection method in this example.
[0092]
[0093] The test results showed that the non-ref NUMTs obtained by the detection method in this example were validated in all three generations of sequencing data. That is, based on the IGV map, insertion sequences of similar length were observed at the corresponding nDNA locations. Subsequently, these sequences were extracted and aligned back to rCRS. The alignment positions were close to the breakpoints predicted in this example, with a positional deviation of 0-6 bp, which was consistent with expectations. Figures 4 to 6This is an IGV diagram of one of the non-ref NUMTs in this example: chrM:61-16089-chr11:49883569-49883572 in the second and third generation data, as well as the position of the inserted sequence aligned to the rCRS and the mismatch base information in the third generation data. The results shown in the figure verify the existence of the non-ref NUMTs detected in this example.
[0094] Figure 4 For the non-ref NUMTs detected in this example: chrM:61-16089-chr11:49883569-49883572, here is the IGV plot. Figure 5 For non-ref NUMTs, there is an insertion sequence of ~530bp at position chr11:49883569 in the third-generation data; Figure 6 The non-ref NUMTs sequence was aligned to the position of rCRS and the mismatched bases.
[0095] Table 2 shows the statistical results of non-ref NUMTs obtained by the detection method in this example and non-ref NUMTs obtained by the existing technology NUMT-detection.
[0096] Table 2 shows the number of non-ref-NUMTs obtained by the detection method in this example and existing NUMT-detection techniques.
[0097] sample Existing technology NUMT-detection This example uses the detection method. Total number tes200124-PN-1 42 1 1 tes200124-TN-1 118 1 1 tes210075-PN-1 43 4 3 tes210075-TN-1 174 4 3
[0098] The results in Table 2 show that the number of non-ref NUMTs detected by the detection method in this example is much smaller than that of the existing NUMT-detection technology. Among them, the existing technology did not detect two non-ref NUMTs that were validated in three generations of data. Therefore, it can be inferred that the detection method in this example has higher specificity than the existing NUMT-detection technology.
[0099] Using the breakpoint locations obtained from three generations of data as a standard, the accuracy of the breakpoints detected by the method in this example and existing technologies for non-ref NUMTs is compared. Figure 7 As shown, the breakpoint distance obtained by the detection method in this example is significantly smaller than that of the prior art, meaning that the breakpoints of non-ref NUMTs predicted by the detection method in this example are more accurate.
[0100] Furthermore, based on the information in the existing NUMT-detection output files and the descriptions of result reliability in the articles, this example sorts the results of the existing NUMT-detection technology according to their estimated reliability and analyzes the top ten non-ref NUMTs. It finds that while they do indeed contain reads from mtDNA in second-generation data, they are not non-ref NUMTs, and this was not verified in third-generation data. Figure 8 and Figure 9 As shown in the figure. Therefore, it can be inferred that the detection method in this example has a lower false positive rate compared to the existing NUMT-detection technology.
[0101] Figure 7 The results show the distance analysis between the breakpoints of non-ref NUMTs detected by the detection method in this example and the breakpoints obtained from three generations of data, compared with the existing NUMT-detection technology. Figure 8 The highest confidence non-ref NUMT IGV plot in second-generation data obtained by existing NUMT-detection technology;
[0102] Figure 9 This is the highest confidence level of non-ref NUMT obtained by existing NUMT-detection technology, validated in three generations of data.
[0103] 4. Detect mutations that accumulate in non-ref-NUMTs
[0104] Using blastn, the consistent sequence was aligned to rCRS, and the presence of non-refNUMTs was further verified based on the alignment position, and mismatched bases, i.e., their accumulated mutations, were identified.
[0105] The test results showed that the detection method in this example detected 22, 24, 28, and 27 accumulated mutations on non-ref-NUMTs in four samples, respectively, including two high-frequency indels (m.16258_16259insA and m.16264_16264del), which mainly accumulated on non-ref NUMTs: chrM:61-16089-chr11:49883569-49883572. Existing NUMT-detection technology cannot detect accumulated mutations on non-ref-NUMTs.
[0106] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.
Claims
1. A method for detecting the nuclear genome sequence of mitochondrial origin, characterized in that: Includes the following steps, The sequencing data preprocessing steps include aligning whole-genome sequencing data to the mitochondrial reference genome, retaining only reads that can be aligned to the mitochondrial reference genome; aligning reads aligned to the mitochondrial reference genome to a reference sequence containing the 23 pairs of human chromosomes and the mitochondrial reference genome, removing non-unique aligned reads, and obtaining a read set that does not contain the mitochondrial-origin nuclear genome sequence of the reference genome. The mitochondrial-origin nuclear genome sequence detection step includes extracting a portion of the reads from a set of reads that do not contain a reference mitochondrial-origin nuclear genome sequence, aligning a portion of the sequence to the nuclear genome reference sequence, and aligning the remaining portion to the mitochondrial reference genome as potential linking reads; Based on the alignment positions of potential linker reads, potential linker reads within a 50bp distance are clustered as read clusters. The coordinates and orientation of the integrated mtDNA fragments, as well as the location of nuclear genome integration, are located. The integrated mtDNA fragments are considered as non-reference mitochondrial-origin nuclear genome sequences. Inconsistent alignment pairs are searched within 100bp upstream and downstream of the read clusters, i.e., one read aligns to the mitochondrial reference genome and the other paired read aligns to the nuclear genome reference sequence. The number of inconsistent alignment pairs is counted, and these inconsistent alignment pairs provide supporting information for the existence of non-reference mitochondrial-origin nuclear genome sequences. The homozygous or heterozygous non-reference mitochondrial-origin nuclear genome sequence is calculated, i.e., the ratio of the number of supporting reads to the average nuclear genome coverage depth. The number of supporting reads is the sum of the number of potential linker reads and the number of inconsistent alignment pairs, and the average nuclear genome coverage depth is the number of reads aligned to the nuclear genome reference sequence multiplied by the read length and then divided by the length of the nuclear genome reference sequence.
2. The method according to claim 1, characterized in that: It also includes a step for validating the nuclear genome sequence of mitochondrial origin; The mitochondrial nuclear genome sequence verification step includes assembling reads or read pairs that align to the mitochondrial reference genome portion and generate a consistent sequence; aligning the generated consistent sequence to the mitochondrial reference genome and verifying the non-reference mitochondrial nuclear genome sequence based on its alignment position.
3. The method according to claim 2, characterized in that: It also includes a step for detecting cumulative mutations in the nuclear genome sequence of mitochondrial origin; The step of detecting cumulative mutations in the mitochondrial-origin nuclear genome sequence includes identifying mismatched bases, i.e., cumulative mutations in the non-reference mitochondrial-origin nuclear genome sequence, based on the result of the alignment of the consistent sequence to the mitochondrial reference genome.
4. The method according to any one of claims 1-3, characterized in that: It also includes the annotation steps; The annotation steps include annotating the breakpoint locations of the nuclear genome reference sequence and the regions and genes containing the non-reference genome mitochondrial-origin nuclear genome sequence.
5. An apparatus for detecting the nuclear genome sequence of mitochondrial origin, characterized in that: Includes a sequencing data preprocessing module and a mitochondrial-origin nuclear genome sequence detection module; The sequencing data preprocessing module includes a method for aligning whole-genome sequencing data to a mitochondrial reference genome, retaining only reads that can be aligned to the mitochondrial reference genome; aligning reads aligned to the mitochondrial reference genome to a reference sequence containing the 23 pairs of human chromosomes and the mitochondrial reference genome, removing non-unique aligned reads, and obtaining a read set that does not contain the mitochondrial-origin nuclear genome sequence of the reference genome. The mitochondrial-origin nuclear genome sequence detection module includes reads that are extracted from a set of reads containing mitochondrial-origin nuclear genome sequences without a reference genome, and aligned to the nuclear genome reference sequence, while the remaining reads are aligned to the mitochondrial reference genome, serving as potential linker reads. Based on the alignment positions of potential linker reads, potential linker reads within a 50bp distance are clustered as read clusters. The coordinates and orientation of the integrated mtDNA fragments, as well as the location of nuclear genome integration, are located. The integrated mtDNA fragments are considered as non-reference mitochondrial-origin nuclear genome sequences. Inconsistent alignment pairs are searched within 100bp upstream and downstream of the read clusters, i.e., one read aligns to the mitochondrial reference genome and the other paired read aligns to the nuclear genome reference sequence. The number of inconsistent alignment pairs is counted, and these inconsistent alignment pairs provide supporting information for the existence of non-reference mitochondrial-origin nuclear genome sequences. The homozygous or heterozygous non-reference mitochondrial-origin nuclear genome sequence is calculated, i.e., the ratio of the number of supporting reads to the average nuclear genome coverage depth. The number of supporting reads is the sum of the number of potential linker reads and the number of inconsistent alignment pairs, and the average nuclear genome coverage depth is the number of reads aligned to the nuclear genome reference sequence multiplied by the read length and then divided by the length of the nuclear genome reference sequence.
6. The apparatus according to claim 5, characterized in that: It also includes a mitochondrial origin nuclear genome sequence verification module; The mitochondrial nuclear genome sequence verification module includes assembling reads or read pairs that align to the mitochondrial reference genome portion to support the existence of a non-reference mitochondrial nuclear genome sequence, generating a consistent sequence; aligning the generated consistent sequence to the mitochondrial reference genome, and verifying the non-reference mitochondrial nuclear genome sequence based on its alignment position.
7. The apparatus according to claim 6, characterized in that: It also includes a module for detecting cumulative mutations in the nuclear genome sequence of mitochondrial origin; The mitochondrial-origin nuclear genome sequence cumulative mutation detection module includes a method for identifying mismatched bases, i.e., cumulative mutations in the non-reference mitochondrial-origin nuclear genome sequence, based on the result of the alignment of the consistent sequence to the mitochondrial reference genome.
8. The apparatus according to any one of claims 5-7, characterized in that: It also includes a comment module; The annotation module includes regions and genes containing nuclear genome reference sequence breakpoints and non-reference mitochondrial-origin nuclear genome sequences.
9. An apparatus for detecting the nuclear genome sequence of mitochondrial origin, characterized in that: The device includes a memory and a processor; The memory includes a storage device for storing programs; The processor includes a method for detecting mitochondrial-origin nuclear genome sequences according to any one of claims 1-4 by executing a program stored in the memory.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program that can be executed by a processor to implement the method for detecting mitochondrial-origin nuclear genome sequences as described in any one of claims 1-4.
Citation Information
Patent Citations
Cucumber mitochondrial genome SSR marker development and application of marker in seed purity identification
CN105886602A
Novel rapid DNA separation and sequencing method for mitochondrial genome and kit
CN107881225A