Method, system and equipment for monitoring minimal residual focus of blood tumor fusion gene
By performing high-throughput sequencing and aligning screening of blood tumor samples, candidate fusion gene sequences are screened out and filtered, and the ratio of micro residual lesions is calculated, the problem of insufficient MRD monitoring sensitivity and specificity in the prior art is solved, and the monitoring effect of high sensitivity and specificity is achieved.
Patent Information
- Application Number
- CN202510264124.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-06
AI Technical Summary
Existing micro-residual lesions (MRD) monitoring techniques for blood tumors are less sensitive and specific, making it difficult to accurately detect a small number of residual cancer cells in the body.
By obtaining the fusion gene information and high-throughput sequencing data of the sample, performing PCR alignment screening without deduplication, split reads and discordant pair reads are selected as candidate fusion gene sequences, and filtering is combined with the annotation file of the reference genome, and the copy number of the fusion gene and the housekeeper gene is calculated to obtain the tiny residual lesion ratio.
High sensitivity and specificity monitoring of MRD are achieved, and the sensitivity is 1 to 2 logs higher than conventional techniques, and a small number of residual cancer cells can be accurately detected in the body.
Smart Images

Figure CN120199323A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of biotechnology and bioinformatics, and particularly relates to a method, system and device for monitoring minimal residual disease of blood tumor fusion genes. Background Art
[0002] Minimal Residual Disease (MRD) refers to the state in which, after cancer treatment, tumors cannot be detected by conventional clinical examination means, but in fact, there are still a small number of cancer cells remaining in the body.
[0003] So far, in the domestic and foreign literature that has been published, most authors classify MRD into two states, positive and negative, by determining a threshold; however, due to limitations in sensitivity, MRD negative described in the literature means that cancer cells cannot be detected in the patient's body by existing methods, that is, MRD negative can mean that there are cancer cells in the body but they cannot be detected, or that there are no cancer cells in the body.
[0004] Currently, the molecular biology methods for detecting blood tumor MRD mainly include: multi-parameter flow cytometry (MFC), real-time quantitative polymerase chain reaction (RQPCR) technology, digital PCR (dd-PCR), and next-generation sequencing technology (NGS), etc. However, the existing monitoring technologies have low sensitivity and specificity. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, system and device for monitoring minimal residual disease of blood tumor fusion genes, which can achieve high-sensitivity and high-specificity monitoring of MRD.
[0006] The first aspect of the present invention provides a method for monitoring minimal residual disease of blood tumor fusion genes, including:
[0007] Obtaining the fusion gene information of a sample and the high-throughput sequencing data containing all the genetic information of the sample;
[0008] Without removing duplicates in PCR for the high-throughput sequencing data, comparing the high-throughput sequencing data to a reference genome to obtain a first alignment result, and screening out split reads and discordant pair reads from the first alignment result as candidate fusion gene sequences;
[0009] Aligning the candidate fusion gene sequences to the annotation file of the reference genome to obtain a second alignment result, obtaining a predicted fusion gene sequence according to the overlapping situation of the second alignment result, filtering the predicted fusion gene sequence to obtain a true fusion gene sequence, and calculating the first copy number of the fusion gene according to the true fusion gene sequence;
[0010] Screen out candidate housekeeping gene sequences aligned to the chromosomal region where the housekeeping genes are located from the first alignment results, screen out true housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence level, and calculate the second copy number of the housekeeping genes based on the true housekeeping gene sequences;
[0011] Divide the first copy number by the second copy number to obtain the minimal residual disease ratio of the fusion gene.
[0012] In some embodiments, filtering the predicted fusion gene sequences to obtain true fusion gene sequences includes:
[0013] If the breakpoint of the predicted fusion gene sequence is not annotated, merge the breakpoints with the most sequence support and within a preset base range as the predicted fusion gene;
[0014] If the predicted fusion gene sequence only supports split reads, filter the predicted fusion gene sequences of paired-end sequencing where the number of bases aligned to both ends of the breakpoint does not reach the preset base threshold;
[0015] If the predicted fusion gene sequence has multiple transcripts, filter out the transcripts other than the transcript with the most reads support;
[0016] Filter out the predicted fusion gene sequences with a random matching probability greater than the probability threshold;
[0017] Filter out the predicted fusion gene sequences with paired chaos to obtain true fusion gene sequences.
[0018] In some embodiments, screening out true housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence level includes:
[0019] Screen out candidate housekeeping gene sequences with an alignment distance within the expected range, correct direction, and unique alignment from the candidate housekeeping gene sequences according to the alignment status as predicted housekeeping gene sequences;
[0020] Screen out the predicted housekeeping gene sequences with a MAPQ value greater than or equal to the preset confidence threshold from the predicted housekeeping gene sequences according to the confidence level as true housekeeping gene sequences.
[0021] In some embodiments, calculating the first copy number of the fusion gene based on the true fusion gene sequences includes:
[0022] Calculate the number of the true fusion gene sequences as the first copy number.
[0023] In some embodiments, calculating the second copy number of the housekeeping genes based on the true housekeeping gene sequences includes:
[0024] Calculate the proportion of the number of bases with the tag character M in the CIGAR value of the true housekeeping gene sequence, and calculate the number of true housekeeping gene sequences in which the proportion of the number of bases with the tag character M is above the preset proportion threshold, which is used as the second copy number of the housekeeping gene.
[0025] In some embodiments, obtaining the predicted fusion gene sequence according to the overlapping situation of the second alignment result includes:
[0026] Preset the minimum number of supporting reads or the minimum overlapping length as the overlapping threshold;
[0027] When the minimum number of supporting reads is used as the overlapping threshold, the candidate fusion gene sequence with the number of supporting reads reaching the overlapping threshold is used as the predicted fusion gene sequence;
[0028] When the minimum overlapping length is used as the overlapping threshold, the candidate fusion gene sequence with the overlapping length reaching the overlapping threshold is used as the predicted fusion gene sequence
[0029] In some embodiments, the housekeeping gene is the ABL1 gene.
[0030] The second aspect of the present invention provides a monitoring system for minimal residual disease of blood tumor fusion genes, including:
[0031] An acquisition module for acquiring the fusion gene information of a sample and the high-throughput sequencing data containing all the genetic information of the sample;
[0032] A comparison and screening module for comparing the high-throughput sequencing data to a reference genome without PCR deduplication to obtain a first alignment result, and screening out splitreads and discordantpairreads from the first alignment result as candidate fusion gene sequences;
[0033] A comparison, filtering and calculation module for aligning the candidate fusion gene sequences to the annotation file of the reference genome to obtain a second alignment result, obtaining the predicted fusion gene sequence according to the overlapping situation of the second alignment result, filtering the predicted fusion gene sequence to obtain the true fusion gene sequence, and calculating the first copy number of the fusion gene according to the true fusion gene sequence;
[0034] A screening and calculation module for screening out candidate housekeeping gene sequences aligned to the chromosomal region where the housekeeping gene is located from the first alignment result, screening out true housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence level, and calculating the second copy number of the housekeeping gene according to the true housekeeping gene sequence;
[0035] The MRD ratio calculation module is used to divide the first copy number by the second copy number to obtain the minimal residual disease ratio of the fusion gene.
[0036] The third aspect of the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0037] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0038] The technical solution provided by the present invention has the following advantages and effects: By aligning the high-throughput sequencing data of the fusion gene to the reference genome, then screening and filtering the first alignment result to obtain the true fusion gene sequence and the true housekeeping gene sequence, then calculating the copy number of the fusion gene and the copy number of the housekeeping gene, and calculating the minimal residual disease ratio according to the copy number of the fusion gene and the copy number of the housekeeping gene, that is, obtaining the MRD ratio, realizing the monitoring of high sensitivity and specificity of MRD. Description of the Drawings
[0039] Figure 1 is a schematic flow chart of the method for monitoring minimal residual disease of fusion genes in hematological malignancies provided by the present invention;
[0040] Figure 2 is a structural block diagram of the system for monitoring minimal residual disease of fusion genes in hematological malignancies provided by the present invention;
[0041] Figure 3 is an internal structure diagram of the computer device provided by the embodiment of the present invention. Detailed Embodiments
[0042] To facilitate the understanding of the present invention, the specific embodiments of the present invention will be described in more detail below with reference to the accompanying drawings of the specification.
[0043] Unless otherwise specified or defined, the "first, second..." used herein is only for differentiating names and does not represent a specific quantity or order.
[0044] Unless otherwise specified or defined, the term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0045] It should be noted that "fixed to" and "connected to" herein can be directly fixed or connected to an element, or indirectly fixed or connected to an element.
[0046] Such as Figure 1As shown, in this embodiment, a method for monitoring minimal residual disease of fusion genes in hematological malignancies is provided, including the following steps S1 to S5:
[0047] Step S1: Obtain the fusion gene information of the sample and the high-throughput sequencing data containing all the genetic information of the sample.
[0048] In practical applications, a fusion gene refers to two or more originally independent genes that are joined together at the DNA level due to chromosomal structural variations (such as translocations, inversions, insertions, etc.) to form a new chimeric gene. The fusion gene information includes: exon information of the left breakpoint transcript and exon information of the right breakpoint transcript; the exon information includes: all exon regions, exon numbers, transcript directions, the number of exons, and the position information of the exon information where the breakpoint is located. For the acquisition of high-throughput sequencing data containing all the genetic information of the sample, according to the fusion gene information, primers are designed using primer5 software, primer parameters are designed according to PCR amplification requirements, total RNA of the peripheral blood or bone marrow extract of the subject to which the sample belongs is extracted, the total RNA is specifically reverse-transcribed into cDNA, PCR is used for amplification, the product is purified by magnetic beads, a cDNA library is established, purified and sequenced on a machine, the sequencing data of the peripheral blood or bone marrow extract is automatically monitored for the completion of the signal, and the downloaded data file is split to obtain the original fastq.gz file. The original fastq.gz file is a common format of high-throughput sequencing data, which contains the sequence information directly generated by the sequencer. This file format consists of four lines, and every four lines represent a sequencing read.
[0049] In practical applications, the transcript of the fusion gene is normalized by the transcript of the housekeeping gene for relative quantification. In this application, the housekeeping gene is the ABL1 gene.
[0050] Step S2: Without PCR deduplication of the high-throughput sequencing data, compare the high-throughput sequencing data to the reference genome to obtain a first alignment result, and screen out split reads and discordant pair reads from the first alignment result as candidate fusion gene sequences.
[0051] In practical applications, the reference genome uses the hg19 reference genome, which can be downloaded from the following website: https: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / hg19.fa.gz. After high-throughput sequencing of total RNA, since PCR amplification is performed during the sequencing process, PCR duplicate removal is usually performed after obtaining the high-throughput sequencing data. However, in this application, the high-throughput sequencing data is not subjected to PCR duplicate removal to prevent the duplicate removal operation from accidentally removing true biological signals. Especially when there are a large number of similar but not completely identical sequences in the sample, the non-duplicate removal strategy will improve the detection sensitivity. Therefore, in this application, the high-throughput sequencing data without PCR duplicate removal is aligned with the reference genome to obtain a BAM file. The BAM file is a binary sequence alignment / mapping file, usually used to store the results of the alignment of sequencing reads with the reference genome, that is, to obtain the first alignment result. The sequencing reads with alignment results of split reads and discordant pair reads are selected as candidate fusion gene sequences. The split reads refer to a read that spans a fusion gene and can be split and mapped to the genes on both sides of the fusion gene. Split reads can directly indicate the existence of a fusion gene. Discordant pair reads refer to the situation where the positions of a pair of reads obtained by paired-end sequencing aligned to the reference genome do not match the expectation, that is, they are aligned to different genes respectively, or the alignment distance does not match the insert size during library construction. Discordant pair reads indicate that there may be a fusion event because, under the normal genomic structure, a pair of reads obtained by paired-end sequencing should be aligned to adjacent positions. Therefore, in this application, split reads and discordant pair reads are used as candidate fusion gene sequences.
[0052] Specifically, before aligning the high-throughput sequencing data to the reference genome, data quality control of the high-throughput sequencing data is also required to identify and remove low-quality sequencing reads, such as low-quality bases, short reads, or adapter contamination. Low-quality sequencing reads usually do not provide useful information and may even interfere with subsequent analysis. Therefore, performing data quality control facilitates improving the accuracy and reliability of subsequent analysis.
[0053] Step S3: Align the candidate fusion gene sequences to the annotation file of the reference genome to obtain a second alignment result. According to the overlapping situation of the second alignment result, a predicted fusion gene sequence is obtained. The predicted fusion gene sequence is filtered to obtain a true fusion gene sequence. According to the true fusion gene sequence, the first copy number of the fusion gene is calculated.
[0054] In practical applications, the predicted fusion genes have different splicing patterns. After using split reads and discordant pair reads as candidate fusion gene sequences, these sequences may contain reads that are not fusion genes. Therefore, it is necessary to screen the candidate fusion gene sequences. By using alignment software such as BLAST and Bowtie2, the candidate fusion gene sequences are aligned with the annotation file of the reference genome, which is the annotation file built into the alignment software. During the alignment process, the software will find the best matching positions between the candidate sequences and the reference genome, calculate the overlapping situation, and then obtain the predicted fusion gene sequences based on the overlapping situation.
[0055] Specifically, obtaining the predicted fusion gene sequences according to the overlapping situation of the second alignment result includes:
[0056] Pre-setting the minimum number of supporting reads or the minimum overlapping length as the overlapping threshold;
[0057] In the case where the minimum number of supporting reads is used as the overlapping threshold, the candidate fusion gene sequences with the number of supporting reads reaching the overlapping threshold are used as the predicted fusion gene sequences;
[0058] In the case where the minimum overlapping length is used as the overlapping threshold, the candidate fusion gene sequences with the overlapping length reaching the overlapping threshold are used as the predicted fusion gene sequences.
[0059] In practical applications, by setting the minimum number of supporting reads as the overlapping threshold, the number of reads spanning two genes is counted to judge the credibility of the fusion event. The more supporting reads there are, the lower the false positive rate. In this way, the predicted fusion gene sequences are screened out. By setting the minimum overlapping length as the overlapping threshold, it is ensured that the alignment lengths of the reads on both sides of the fusion breakpoint are long enough to improve the alignment accuracy and avoid misalignment of short sequences. In this way, the predicted fusion gene sequences are screened out, and the predicted fusion gene sequences are filtered to remove false positives to obtain the true fusion gene sequences. In other embodiments, the minimum number of supporting reads and the minimum overlapping length can be set simultaneously, and only the candidate fusion gene sequences with the number of supporting reads reaching the minimum number of reads and the overlapping length reaching the minimum overlapping length are used as the predicted fusion gene sequences to further improve the filtering accuracy.
[0060] Specifically, filtering the predicted fusion gene sequences to obtain the true fusion gene sequences includes:
[0061] Filtering the predicted fusion gene sequences according to the preset filtering rules to obtain the true fusion gene sequences;
[0062] Among them, the preset filtering rules include:
[0063] If the breakpoint of the predicted fusion gene sequence is not annotated, the breakpoints with the most sequence support and within a preset base range are merged as the predicted fusion gene;
[0064] If the predicted fusion gene sequence only supports split reads, the predicted fusion gene sequence of paired-end sequencing with the number of bases aligned to both ends of the breakpoint not reaching the preset base threshold is filtered;
[0065] If there are multiple transcripts in the predicted fusion gene sequence, the transcripts other than the one with the most reads support are filtered;
[0066] Filter the predicted fusion gene sequences with a random matching probability greater than the probability threshold;
[0067] Filter the predicted fusion gene sequences with paired chaos.
[0068] In practical applications, in the case where the breakpoints of the predicted fusion gene sequences are not annotated, merging the breakpoints with the most sequence support and within a preset base range helps reduce breakpoint variations caused by sequencing errors or alignment uncertainties. The preset base range is + / - 5 bases. The preset base threshold is 25 bases. For predicted fusion gene sequences of paired-end sequencing supported only by split-reads, it is required that at least 25 bases can be aligned to both ends of the breakpoint. That is, for predicted fusion gene sequences of paired-end sequencing, predicted fusion gene sequences with the number of bases aligned to both ends of the breakpoint less than 25 bases need to be filtered, which helps ensure the authenticity of the fusion event because shorter split-read support may not be sufficient to exclude accidental alignment errors. For the case where there are multiple transcripts of the predicted fusion gene sequence, only the transcript with the most reads support is retained, and the remaining transcripts are filtered, which helps reduce false positives caused by insufficient sequencing depth or differences in transcript expression levels. The probability threshold can be set to 10^-3. By presetting the probability threshold and using BLAST to filter fusion genes with similar sequences, by setting an appropriate expectation value (representing the probability of random matching), that is, using the preset probability threshold as the expectation value, false positive fusion gene pairs caused by high sequence similarity can be filtered out, which helps distinguish true fusion events from misjudgments caused by sequence repeats or similarities. Fusion events may be false positives caused by sequencing errors or data processing problems. Therefore, it is necessary to filter out fusion gene pairs with chaotic pairings. Chaotic fusion gene pairs refer to the phenomenon that one gene fuses with multiple different genes, usually caused by technical noise or false positives. A mapping software is used to filter out chaotic fusion gene pairs. For example, by setting a filtering threshold, fusion gene pairs with a value less than the filtering threshold are filtered out, and fusion gene pairs with the number of fusion gene partners exceeding 10 are filtered out. Fusion gene partners refer to the two independent genes participating in the formation of a chimeric gene.
[0069] Specifically, calculating the first copy number of the fusion gene based on the true fusion gene sequence includes:
[0070] Calculating the number of the true fusion gene sequences as the first copy number.
[0071] In practical applications, the number of true fusion gene sequences is used as the first copy number, that is, the number of split reads and discordant pair reads of the filtered fusion gene are added to obtain the first copy number, that is, the copy number of the fusion gene is obtained.
[0072] Step S4: Screen out the candidate housekeeping gene sequences aligned to the chromosomal regions where the housekeeping genes are located from the first alignment result, screen out the true housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence level, and calculate the second copy number of the housekeeping gene based on the true housekeeping gene sequences.
[0073] In practical applications, each read in the BAM file will have a corresponding reference name, indicating the chromosome or sequence name on the reference genome to which the read is aligned. Screen out the reads whose coordinates are within the range of the ABL1 gene from the BAM file as candidate housekeeping gene sequences. The coordinate range of the ABL1 gene is chr9:133730369-133730483, chr9:133738149-133738223.
[0074] Specifically, the screening of the true housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence level includes:
[0075] Screen out the candidate housekeeping gene sequences with unique alignments whose alignment distances are within the expected range and in the correct direction from the candidate housekeeping gene sequences according to the alignment status as predicted housekeeping gene sequences;
[0076] Screen out the predicted housekeeping gene sequences with MAPQ values greater than or equal to the preset confidence threshold from the predicted housekeeping gene sequences according to the confidence level as the true housekeeping gene sequences.
[0077] In practical applications, the alignment status (flags) is a set of binary flag bits used to describe the alignment status of reads in the BAM file. Different combinations of flag values can represent different alignment characteristics of reads. Through the flags, candidate housekeeping gene sequences with unique alignments whose alignment distances are within the expected range and in the correct direction can be screened out from the candidate housekeeping gene sequences as predicted housekeeping gene sequences. The confidence level (MAPQ) represents the confidence of the read alignment to the reference genome. The higher the MAPQ value, the more reliable the alignment result. Filter out the predicted housekeeping gene sequences with MAPQ values less than the preset confidence threshold, remove the false housekeeping gene sequences, and obtain the true housekeeping gene sequences.
[0078] Specifically, the calculation of the second copy number of the housekeeping gene based on the true housekeeping gene sequences includes:
[0079] Calculate the proportion of the number of bases with the marker character M in the true housekeeping gene sequences according to the CIGAR values of the true housekeeping gene sequences, and calculate the number of true housekeeping gene sequences with the proportion of the number of bases with the marker character M above the preset proportion threshold as the second copy number of the housekeeping gene.
[0080] In practical applications, the preset proportion threshold is set to 95%. The CIGAR value is used to describe the detailed matching situation when reads are aligned with the reference genome, including operations such as matching (M), insertion (I), deletion (D), etc. and their lengths. In this application, the number of true housekeeping gene sequences with the proportion of the number of bases with the marker character M being more than 95% is calculated, and it is used as the second copy number of the housekeeping gene. The true housekeeping gene sequence with the proportion of the number of bases with the marker character M being more than 95% indicates that the matching degree of this true housekeeping gene sequence with the reference genome is very high, with almost no insertions or deletions.
[0081] Step S5: Divide the first copy number by the second copy number to obtain the minimal residual disease ratio of the fusion gene.
[0082] In practical applications, in order to verify the sensitivity of this application for MRD monitoring, relevant experiments were conducted in this application. Fourteen standard samples (MRD8701, MRD8702, MRD8801, MRD8802, MRD8901, MRD8902, MRD0005-1, MRD0005-2, MRD0002-1, MRD0002-2, MRD0001-1, MRD0001-2, MRD00005-1, MRD00005-2) were selected. By diluting their concentrations, the BCR::ABL1 / ABL1 was adjusted to different ratios. The ratios of BCR::ABL1 / ABL1 for MRD8701 and MRD8702 were 0.05%, the ratios of BCR::ABL1 / ABL1 for MRD8801 and MRD8802 were 0.02%, the ratios of BCR::ABL1 / ABL1 for MRD8901 and MRD8902 were 0.01%, the ratios of BCR::ABL1 / ABL1 for MRD0005-1 and MRD0005-2 were 0.005%, the ratios of BCR::ABL1 / ABL1 for MRD0002-1 and MRD0002-2 were 0.002%, the ratios of BCR::ABL1 / ABL1 for MRD0001-1 and MRD0001-2 were 0.001%, the ratios of BCR::ABL1 / ABL1 for MRD00005-1 and MRD00005-2 were 0.0005%. In addition, two normal negative samples, MRD26E13 and MRD28E13, were selected. Then, whether the detected ratio of BCR::ABL1 / ABL1 was consistent with the theoretical ratio was verified by this method. The last column is the theoretical ratio, and the second-to-last column is the detected ratio by this method. The experimental data are shown in Table 1, and Table 1 is as follows:
[0083] Table 1 Results of different ratio gradients of the NGS method for blood fusion genes
[0084] Verification test number BCR::ABL1 / ABL1 (%) Standard MRD8701 0.0673 0.05% Standard MRD8702 0.2024 0.05% Standard MRD8801 0.0384 0.02% Standard MRD8802 0.1548 0.02% Standard MRD8901 0.0985 0.01% Standard MRD8902 0.0363 0.01% Standard MRD0005-1 0.0013 0.005% Standard MRD0005-2 0.0016 0.005% Standard MRD0002-1 0.0017 0.002% Standard MRD0002-2 0.0007 0.002% Standard MRD0001-1 0.0005 0.001% Standard MRD0001-2 0.0006 0.001% Standard MRD00005-1 0.0004 0.0005% Standard MRD00005-2 0.0006 0.0005% Standard MRD26E13 0 BCR::ABL1P210 Negative MRD28E13 0 BCR::ABL1P210 Negative
[0085] In Table 1, the NGS method (Next-Generation Sequencing) is a high-throughput sequencing method. BCR::ABL1 is a fusion gene. The value under BCR::ABL1 in Table 1 represents the first copy number, the value under ABL1 represents the second copy number, and BCR::ABL1 / ABL1 (%) represents the ratio of the first copy number to the second copy number. According to Table 1, the sensitivity of the method for monitoring minimal residual disease of blood tumor fusion genes in this application can reach 10 -6 , and the sensitivity of the conventional technology RQ-PCR is 10 -4 ~10 -5 , and the sensitivity of MFC is 10 -3 ~10 -5 . Compared with other MRD monitoring methods (RQ-PCR or MFC), the sensitivity of this application is 1-2 logs higher, with high sensitivity and specificity.
[0086] For the method for monitoring minimal residual disease of blood tumor fusion genes of the present invention, by aligning the high-throughput sequencing data of the fusion gene to the reference genome, and then screening and filtering the first alignment result to obtain the true fusion gene sequence and the true housekeeping gene sequence, and then calculating the copy number of the fusion gene and the copy number of the housekeeping gene, and calculating the MRD ratio according to the copy number of the fusion gene and the copy number of the housekeeping gene, the monitoring of high sensitivity and specificity of MRD is realized.
[0087] As Figure 2 shown, the embodiment of the present invention also provides a system for monitoring minimal residual disease of blood tumor fusion genes, including:
[0088] An acquisition module 10 for acquiring the fusion gene information of a sample and the high-throughput sequencing data containing all the genetic information of the sample;
[0089] An alignment and screening module 20 for aligning the high-throughput sequencing data to the reference genome without PCR duplicate removal to obtain a first alignment result, and screening split reads and discordant pair reads from the first alignment result as candidate fusion gene sequences;
[0090] An alignment, filtering and calculation module 30 for aligning the candidate fusion gene sequences to the annotation file of the reference genome to obtain a second alignment result, obtaining a predicted fusion gene sequence according to the overlapping situation of the second alignment result, filtering the predicted fusion gene sequence to obtain a true fusion gene sequence, and calculating the first copy number of the fusion gene according to the true fusion gene sequence;
[0091] A screening calculation module 40, configured to screen out candidate housekeeping gene sequences aligned to the chromosomal regions where housekeeping genes are located from the first alignment result, screen out true housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence level, and calculate the second copy number of the housekeeping gene based on the true housekeeping gene sequences;
[0092] An MRD ratio calculation module 50, configured to divide the first copy number by the second copy number to obtain the minimal residual disease ratio of the fusion gene.
[0093] Each module of the above-mentioned minimal residual disease monitoring system for blood tumor fusion genes can be implemented in whole or in part by software, hardware, and their combinations. The above-mentioned modules and units can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0094] As Figure 3 shown, an embodiment of the present invention discloses a computer device, including a memory and a processor, where the memory stores a computer program;
[0095] Among them, the computer device can be a server, and its internal structure diagram can be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the minimal residual disease monitoring method for blood tumor fusion genes described in the above-mentioned embodiments.
[0096] Those skilled in the art can understand that Figure 3 the structure shown in
[0097] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0098] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0099] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
Claims
1. A method for monitoring minimal residual lesions of blood tumor fusion genes, characterized in that: include: Obtain the fusion gene information of the sample and high-throughput sequencing data containing all the genetic information of the sample; Without performing PCR deduplication on the high-throughput sequencing data, the high-throughput sequencing data is compared to a reference genome to obtain a first comparison result, and split reads and discordant pair reads are screened out from the first comparison result as candidate fusion gene sequences; Aligning the candidate fusion gene sequence to an annotation file of a reference genome to obtain a second alignment result, obtaining a predicted fusion gene sequence according to an overlap of the second alignment result, filtering the predicted fusion gene sequence to obtain a true fusion gene sequence, and calculating a first copy number of the fusion gene according to the true fusion gene sequence; Screening out candidate housekeeping gene sequences aligned to the chromosome region where the housekeeping gene is located from the first alignment result, screening out true housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence, and calculating the second copy number of the housekeeping gene according to the true housekeeping gene sequence; The first copy number is divided by the second copy number to obtain the minimal residual disease ratio of the fusion gene.
2. The method for monitoring minimal residual lesions of blood tumor fusion genes according to claim 1, characterized in that: The filtering the predicted fusion gene sequence to obtain the true fusion gene sequence includes: If the breakpoints of the predicted fusion gene sequence are not annotated, the breakpoints with the most sequence support and a difference distance within a preset base range are merged as the predicted fusion gene; If the predicted fusion gene sequence only supports splitreads, the predicted fusion gene sequence of double-end sequencing whose base numbers aligned to both ends of the breakpoint do not reach the preset base threshold is filtered; If the predicted fusion gene sequence has multiple transcripts, the transcripts except for the transcript with the most support from reads are filtered; Filter the predicted fusion gene sequences whose random matching probability is greater than the probability threshold; The predicted fusion gene sequences with chaotic pairing were filtered out to obtain the true fusion gene sequences.
3. The method for monitoring minimal residual lesions of blood tumor fusion genes according to claim 1, characterized in that: The step of selecting the true housekeeping gene sequence from the candidate housekeeping gene sequences according to the comparison status and the confidence level comprises: According to the alignment status, a candidate housekeeping gene sequence having an alignment distance within an expected range and a correct direction and a unique alignment is selected from the candidate housekeeping gene sequences as a predicted housekeeping gene sequence; According to the confidence, predicted housekeeping gene sequences with MAPQ values greater than or equal to a preset confidence threshold are screened out from the predicted housekeeping gene sequences as true housekeeping gene sequences.
4. The method for monitoring minimal residual lesions of blood tumor fusion genes according to claim 1, characterized in that: The step of calculating the first copy number of the fusion gene according to the true fusion gene sequence comprises: The number of the true fusion gene sequences is calculated as the first copy number.
5. The method for monitoring minimal residual lesions of blood tumor fusion genes according to claim 1, characterized in that: The step of calculating the second copy number of the housekeeping gene according to the real housekeeping gene sequence comprises: The proportion of the number of bases marked with the character M is calculated according to the CIGAR value of the real housekeeping gene sequence, and the number of real housekeeping gene sequences whose proportion of the number of bases marked with the character M is above a preset proportion threshold is calculated as the second copy number of the housekeeping gene.
6. The method for monitoring minimal residual lesions of blood tumor fusion genes according to claim 1, characterized in that: The step of obtaining a predicted fusion gene sequence according to the overlap of the second comparison results includes: The minimum number of supported read segments or the minimum overlap length is pre-set as the overlap threshold; When the minimum number of supported reads is used as the overlap threshold, the candidate fusion gene sequence whose number of supported reads reaches the overlap threshold is used as the predicted fusion gene sequence; When the minimum overlap length is used as the overlap threshold, the candidate fusion gene sequence whose overlap length reaches the overlap threshold is used as the predicted fusion gene sequence.
7. The method for monitoring minimal residual lesions of blood tumor fusion genes according to claim 1, characterized in that: The housekeeping gene is the ABL1 gene.
8. A blood tumor fusion gene minimal residual disease monitoring system, characterized in that: include: An acquisition module is used to obtain the fusion gene information of the sample and high-throughput sequencing data containing all the genetic information of the sample; A comparison and screening module, for comparing the high-throughput sequencing data to a reference genome without performing PCR deduplication on the high-throughput sequencing data, obtaining a first comparison result, and screening splitreads and discordantpairreads from the first comparison result as candidate fusion gene sequences; a comparison filtering calculation module, used for comparing the candidate fusion gene sequence to the annotation file of the reference genome to obtain a second comparison result, obtaining a predicted fusion gene sequence according to the overlap of the second comparison result, filtering the predicted fusion gene sequence to obtain a true fusion gene sequence, and calculating a first copy number of the fusion gene according to the true fusion gene sequence; a screening and calculation module, configured to screen out candidate housekeeping gene sequences that are aligned to the chromosome region where the housekeeping gene is located from the first alignment result, screen out true housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence, and calculate the second copy number of the housekeeping gene according to the true housekeeping gene sequence; The MRD ratio calculation module is used to divide the first copy number by the second copy number to obtain the minimal residual disease ratio of the fusion gene.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Device for detecting blood disease fusion gene
CN111180013A
Method for detecting minimal residual focus of solid tumor, computing device and storage medium
CN114708908A
Detection method, device and equipment for tiny residual focus and storage medium
CN115679000A
Method, system and equipment for detecting internal tandem repetition and storage medium
CN121331232A
Systems and methods for generating and analyzing a customized genomic sequence incorporating gene fusions for therapeutic applications
US20220375545A1
Cited By
Gene fusion detection method based on high-throughput single-ended sequencing
CN121528304A