Method, system and apparatus for monitoring minimal residual disease of hematological tumor fusion genes
By performing PCR-based non-duplicate alignment and filtering on high-throughput sequencing data of hematologic tumor samples, true fusion gene and housekeeping gene sequences were screened out, and the MRD ratio was calculated. This solved the problems of low sensitivity and low specificity in the existing technology for monitoring minimal residual disease in hematologic tumors, and achieved a monitoring effect with high sensitivity and high specificity.
Patent Information
- Application Number
- CN202510264124.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-03-06
AI Technical Summary
Existing technologies for monitoring minimal residual disease in hematological malignancies have low sensitivity and specificity, making it difficult to accurately detect the presence of small amounts of cancer cells in the body.
By acquiring fusion gene information and high-throughput sequencing data from the samples, PCR alignment without deduplication was performed to screen split reads and discordant pair reads as candidate fusion gene sequences. These were then compared and filtered using a reference genome annotation file. The copy number of the fusion gene and the housekeeping gene was calculated to obtain the minimal residual disease ratio.
It achieves highly sensitive and specific monitoring of minimal residual lesions in hematological malignancies, improving detection accuracy and increasing sensitivity by 1-2 logs compared to existing technologies.
Smart Images

Figure CN120199323B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of biotechnology and bioinformatics, and particularly relates to a blood tumor fusion gene minimal residual disease monitoring method, system and device. BACKGROUND
[0002] Minimal Residual Disease (MRD) refers to a state in which a small amount of cancer cells remains in the body after cancer treatment, which cannot be detected by conventional clinical examination methods.
[0003] So far, in the published domestic and foreign literature, most authors divide MRD into positive and negative states by determining the threshold; however, due to the limitation of sensitivity, MRD negative described in the literature means that no cancer cells exist in the patient's body detected by existing methods, that is, MRD negative can be that cancer cells exist in the body but cannot be detected, or that cancer cells no longer exist in the body.
[0004] At present, the molecular biology methods for detecting MRD of blood tumors mainly include multi-parameter flow cytometry (MFC), real-time quantitative polymerase chain reaction (RQPCR) technology, digital PCR (dd-PCR) and next-generation sequencing technology (NGS) and the like, but the existing monitoring technology has low sensitivity and specificity. SUMMARY
[0005] The purpose of the present application is to provide a blood tumor fusion gene minimal residual disease monitoring method, system and device, which can realize high sensitivity and high specificity monitoring of MRD.
[0006] The first aspect of the present application provides a blood tumor fusion gene minimal residual disease monitoring method, comprising:
[0007] Obtaining fusion gene information of the sample and high-throughput sequencing data containing all genetic information of the sample;
[0008] In the case of no PCR deduplication to the high-throughput sequencing data, comparing the high-throughput sequencing data to a reference genome to obtain a first alignment result, and screening split reads and discordant pair reads from the first alignment result as candidate fusion gene sequences;
[0009] Comparing the candidate fusion gene sequences to an annotation file of the reference genome to obtain a second alignment result, obtaining a predicted fusion gene sequence according to the overlap of the second alignment result, filtering the predicted fusion gene sequence to obtain a true fusion gene sequence, and calculating a first copy number of the fusion gene according to the true fusion gene sequence;
[0010] screening a candidate housekeeping gene sequence aligned to a chromosome region where a housekeeping gene locates from the first alignment result, screening a real housekeeping gene sequence from the candidate housekeeping gene sequence according to an alignment state and a confidence, and calculating a second copy number of the housekeeping gene according to the real housekeeping gene sequence;
[0011] dividing the first copy number by the second copy number to obtain a micro residual lesion ratio of the fusion gene.
[0012] In some embodiments, the filtering the predicted fusion gene sequence to obtain a real fusion gene sequence comprises:
[0013] if the breakpoint of the predicted fusion gene sequence is not annotated, merging the breakpoint with the most sequence support and a distance difference within a preset base range as the predicted fusion gene;
[0014] if the predicted fusion gene sequence only supports split reads, filtering the predicted fusion gene sequence of paired-end sequencing with a base number aligned to both ends of the breakpoint not reaching a preset base threshold;
[0015] if the predicted fusion gene sequence has multiple transcripts, filtering the transcripts other than the transcript with the most reads support;
[0016] filtering the predicted fusion gene sequence with a random matching probability greater than a probability threshold;
[0017] filtering the predicted fusion gene sequence with pairing confusion to obtain a real fusion gene sequence.
[0018] In some embodiments, the screening a real housekeeping gene sequence from the candidate housekeeping gene sequence according to an alignment state and a confidence comprises:
[0019] screening a candidate housekeeping gene sequence with a unique alignment with a correct direction and an alignment distance within an expected range from the candidate housekeeping gene sequence as a predicted housekeeping gene sequence according to the alignment state;
[0020] screening a predicted housekeeping gene sequence with a MAPQ value greater than or equal to a preset confidence threshold from the predicted housekeeping gene sequence as a real housekeeping gene sequence according to the confidence.
[0021] In some embodiments, the calculating a first copy number of the fusion gene according to the real fusion gene sequence comprises:
[0022] calculating a number of the real fusion gene sequence as the first copy number.
[0023] In some embodiments, the second copy number of the housekeeping gene is calculated according to the real housekeeping gene sequence, comprising:
[0024] The proportion of the number of bases marked by M is calculated according to the CIGAR value of the real housekeeping gene sequence, and the number of real housekeeping gene sequences with the proportion of the number of bases marked by M above a preset proportion threshold is taken as the second copy number of the housekeeping gene.
[0025] In some embodiments, the predicted fusion gene sequence is obtained according to the overlapping condition of the second alignment result, comprising:
[0026] The minimum support read number or the minimum overlap length is preset as the overlap threshold;
[0027] In the case where the minimum support read number is taken as the overlap threshold, the candidate fusion gene sequence with the support read number reaching the overlap threshold is taken as the predicted fusion gene sequence;
[0028] In the case where the minimum overlap length is taken as the overlap threshold, the candidate fusion gene sequence with the overlap length reaching the overlap threshold is taken as the predicted fusion gene sequence
[0029] In some embodiments, the housekeeping gene is ABL1 gene.
[0030] The second aspect of the present application provides a blood tumor fusion gene minimal residual disease monitoring system, comprising:
[0031] The acquisition module is used for acquiring fusion gene information of a sample and high-throughput sequencing data containing all genetic information of the sample;
[0032] The alignment and screening module is used for comparing the high-throughput sequencing data to a reference genome to obtain a first alignment result under the condition that PCR is not removed from the high-throughput sequencing data, and screening split reads and discordant pair reads from the first alignment result as candidate fusion gene sequences;
[0033] The alignment and filtering calculation module is used for comparing the candidate fusion gene sequences to an annotation file of the reference genome to obtain a second alignment result, obtaining a predicted fusion gene sequence according to the overlapping condition of the second alignment result, filtering the predicted fusion gene sequence to obtain a real fusion gene sequence, and calculating a first copy number of the fusion gene according to the real fusion gene sequence.
[0034] The screening calculation module is configured to screen a candidate housekeeping gene sequence aligned to a chromosome region where a housekeeping gene is located from the first alignment result, screen a real housekeeping gene sequence from the candidate housekeeping gene sequence according to an alignment state and a confidence degree, and calculate a second copy number of the housekeeping gene according to the real housekeeping gene sequence.
[0035] The MRD ratio calculation module is configured to divide the first copy number by the second copy number to obtain a minimal residual disease ratio of the fusion gene.
[0036] The third aspect of the present application provides a computer device including a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0037] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the above method when executed by a processor.
[0038] The technical solution provided by the present application has the following advantages and effects: by aligning high-throughput sequencing data of a fusion gene to a reference genome, then screening and filtering the first alignment result to obtain a real fusion gene sequence and a real housekeeping gene sequence, then calculating a copy number of the fusion gene and a copy number of the housekeeping gene, and according to the copy number of the fusion gene and the copy number of the housekeeping gene, a minimal residual disease ratio is calculated, that is, an MRD ratio is obtained, high sensitivity and specificity of MRD monitoring are realized. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 is a flowchart of the blood tumor fusion gene minimal residual disease monitoring method provided by the present application;
[0040] Figure 2 is a structural block diagram of the blood tumor fusion gene minimal residual disease monitoring system provided by the present application;
[0041] Figure 3 is an internal structure diagram of the computer device provided by the present application. DETAILED DESCRIPTION
[0042] In order to facilitate understanding of the present application, specific embodiments of the present application will be described in more detail below with reference to the accompanying drawings.
[0043] Unless specifically stated or defined otherwise, the terms "first, second,..." used herein are merely used for distinguishing names and do not represent a specific number or order.
[0044] Unless specifically stated or defined otherwise, the term "and / or" used herein includes any and all combinations of one or more related listed items.
[0045] It should be noted that "fixed to", "connected to" herein can be directly fixed or connected to an element, or indirectly fixed or connected to an element.
[0046] As shown in Figure 1 The blood tumor fusion gene microresidual lesion monitoring method provided in the embodiment includes the following steps S1-S5:
[0047] Step S1, obtaining fusion gene information of a sample and high-throughput sequencing data containing all genetic information of the sample.
[0048] In actual application, the fusion gene refers to two or more originally independent genes connected together at DNA level due to chromosomal structural variation (such as translocation, inversion, insertion, etc.), forming a new chimeric gene. The fusion gene information includes: exon information of the left breakpoint transcript and exon information of the right breakpoint transcript; the exon information includes: all exon regions, exon numbers, transcript directions, exon numbers and position information of the breakpoint information. For obtaining the high-throughput sequencing data containing all genetic information of the sample, according to the fusion gene information, primer 5 software is used to design primers, primer parameters are designed according to PCR amplification requirements, total RNA of peripheral blood or bone marrow aspirate of the subject to which the sample belongs is extracted, the total RNA is specifically reverse transcribed into cDNA, PCR is used for amplification, product magnetic bead purification is performed, cDNA library is established, and sequencing is performed after purification and on-machine sequencing, automatic monitoring of peripheral blood or bone marrow aspirate sequencing data is completed, and the data file is split after the machine is turned off, to obtain the original fastq.gz file. The original fastq.gz file is a common format of high-throughput sequencing data, which contains sequence information directly generated by the sequencer. The file format is composed of four lines, and each four lines represent a sequencing read (read).
[0049] In actual application, the transcript of the housekeeping gene is used to normalize the transcript of the fusion gene for relative quantification. In the present application, the housekeeping gene is ABL1 gene.
[0050] Step S2, under the condition that PCR is not removed from the high-throughput sequencing data, the high-throughput sequencing data is compared to the reference genome to obtain a first alignment result, and split reads and discordant pair reads are screened from the first alignment result as candidate fusion gene sequences.
[0051] In practical applications, the reference genome used is the hg19 reference genome, which can be downloaded from this website: https: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / hg19.fa.gz. After high-throughput sequencing of total RNA, PCR amplification is performed during the sequencing process. Therefore, PCR deduplication is usually performed after obtaining the high-throughput sequencing data. However, this application does not perform PCR deduplication on the high-throughput sequencing data to prevent the deduplication operation from unintentionally removing real biological signals, especially when there are a large number of similar but not identical sequences in the sample. The non-deduplication strategy can improve the detection sensitivity. Therefore, this application compares the high-throughput sequencing data without PCR deduplication with the reference genome to obtain a BAM file. The BAM file is a binary sequence alignment / mapping file, which is usually used to store the results of the alignment of sequencing reads with the reference genome, that is, to obtain the first alignment result. Sequencing reads with alignment results of split reads and discordant pair reads are screened out as candidate fusion gene sequences. The split reads refer to reads that span the fusion gene and can be split to match the genes on both sides of the fusion gene. The split reads can directly indicate the presence of the fusion gene. Discordant pair reads refer to a pair of reads obtained from paired-end sequencing that do not align to the expected positions on the reference genome. This means they align to different genes or the alignment distance does not match the size of the insert fragment during library construction. Discordant pair reads indicate the possible existence of a fusion event because, under normal genome structure, a pair of reads obtained from paired-end sequencing should align to adjacent positions. Therefore, in this application, split reads and discordant pair reads are considered as candidate fusion gene sequences.
[0052] Specifically, before comparing the high-throughput sequencing data to the reference genome, data quality control is required to identify and remove low-quality sequencing reads, such as low-quality bases, short reads, or adapter contamination. Low-quality sequencing reads usually do not provide useful information and may even interfere with subsequent analysis. Therefore, data quality control helps improve the accuracy and reliability of subsequent analysis.
[0053] Step S3: Align the candidate fusion gene sequence to the annotation file of the reference genome to obtain a second alignment result. Obtain the predicted fusion gene sequence based on the overlap of the second alignment result. Filter the predicted fusion gene sequence to obtain the real fusion gene sequence. Calculate the first copy number of the fusion gene based on the real fusion gene sequence.
[0054] In practical applications, the predicted fusion gene has different splicing modes. After taking the split reads and discordant pair reads as candidate fusion gene sequences, the candidate fusion gene sequences may contain reads that are not fusion genes. Therefore, the candidate fusion gene sequences need to be screened. The candidate fusion gene sequences can be aligned with an annotation file of a reference genome by using an alignment software such as BLAST or Bowtie2. The annotation file of the reference genome is an annotation file provided by the alignment software. In the alignment process, the software finds the best matching position between the candidate sequence and the reference genome and calculates the overlap. Then, the predicted fusion gene sequence is obtained according to the overlap.
[0055] Specifically, the predicted fusion gene sequence is obtained according to the overlap of the second alignment result, comprising:
[0056] The minimum support read number or the minimum overlap length is set as an overlap threshold value in advance;
[0057] In the case where the minimum support read number is set as the overlap threshold value, the candidate fusion gene sequence whose support read number reaches the overlap threshold value is taken as the predicted fusion gene sequence;
[0058] In the case where the minimum overlap length is set as the overlap threshold value, the candidate fusion gene sequence whose overlap length reaches the overlap threshold value is taken as the predicted fusion gene sequence.
[0059] In practical applications, by setting the minimum support read number as the overlap threshold value, the number of reads crossing two genes is counted to determine the credibility of the fusion event. The more support reads, the lower the false positive rate. In this way, the predicted fusion gene sequence is screened. By setting the minimum overlap length as the overlap threshold value, it is ensured that the alignment length of the reads on both sides of the fusion breakpoint is long enough to improve the alignment accuracy and avoid false alignment of short sequences. In this way, the predicted fusion gene sequence is screened. The predicted fusion gene sequence is filtered to remove false positives to obtain the real fusion gene sequence. In other embodiments, the minimum support read number and the minimum overlap length can be set at the same time. Only the candidate fusion gene sequence whose support read number reaches the minimum read number and whose overlap length reaches the minimum overlap length is taken as the predicted fusion gene sequence to further improve the filtering accuracy.
[0060] Specifically, the predicted fusion gene sequence is filtered to obtain the real fusion gene sequence, comprising:
[0061] The predicted fusion gene sequence is filtered according to a preset filtering rule to obtain the real fusion gene sequence;
[0062] The preset filtering rule comprises:
[0063] If the breakpoint of the predicted fusion gene sequence is not annotated, the breakpoints with the most sequence support and within a preset base range of distance are merged as the predicted fusion gene;
[0064] If the predicted fusion gene sequence only supports split reads, the predicted fusion gene sequence of paired-end sequencing with the number of bases aligned to both ends of the breakpoint not reaching a preset base threshold is filtered;
[0065] If the predicted fusion gene sequence has multiple transcripts, the transcripts other than the transcript with the most read support are filtered;
[0066] The predicted fusion gene sequence with a random matching probability greater than a probability threshold is filtered;
[0067] The predicted fusion gene sequence with pairing confusion is filtered.
[0068] In practical application, in the case that the breakpoint of the predicted fusion gene sequence is not annotated, the breakpoints with the most sequence support and the distance within the preset base range are combined, which helps to reduce the breakpoint variation caused by sequencing errors or alignment uncertainty. The preset base range is + / - 5 bases. The preset base threshold is 25 bases. For the predicted fusion gene sequence of paired-end sequencing supported by split-reads, at least 25 bases need to be aligned to both ends of the breakpoint, that is, for the predicted fusion gene sequence with less than 25 bases aligned to both ends of the breakpoint, the predicted fusion gene sequence is filtered, which helps to ensure the authenticity of the fusion event, because the shorter split-read support may not be sufficient to rule out accidental alignment errors. For the case that the predicted fusion gene sequence has multiple transcripts, only the transcript with the most read support is retained, and the remaining transcripts are filtered, which helps to reduce false positives caused by insufficient sequencing depth or differences in transcript expression levels. The probability threshold can be set to 10^-3. The probability threshold is set in advance, and the fusion genes with similar sequences are filtered using BLAST. By setting a suitable expectation value (indicating the probability of random matching), that is, using the pre-set probability threshold as the expectation value, false positive fusion gene pairs caused by high sequence similarity can be filtered out, which helps to distinguish between true fusion events and false judgments caused by sequence duplication or similarity. Fusion events may be false positives caused by sequencing errors or data processing problems. Therefore, fusion gene pairs with pairing confusion need to be filtered out. Fusion gene pairs with pairing confusion: a gene fused with multiple different genes, which is usually caused by technical noise or false positives. The alignment software is used to filter the fusion gene pairs with pairing confusion, such as setting a filtering threshold, filtering the fusion gene pairs less than the filtering threshold, and filtering out the fusion gene pairs with more than 10 fusion gene partners. Fusion gene partners refer to two independent genes involved in forming a chimeric gene.
[0069] Specifically, the first copy number of the fusion gene is calculated according to the real fusion gene sequence, comprising:
[0070] The number of real fusion gene sequences is calculated as the first copy number.
[0071] In practical application, the number of real fusion gene sequences is taken as the first copy number, that is, the number of split reads of the filtered fusion gene and the number of discordant pair reads are added to obtain the first copy number, that is, the copy number of the fusion gene.
[0072] Step S4, screening a candidate housekeeping gene sequence aligned to the chromosome region where the housekeeping gene is located from the first alignment result, screening a real housekeeping gene sequence from the candidate housekeeping gene sequence according to the alignment state and the confidence, and calculating a second copy number of the housekeeping gene according to the real housekeeping gene sequence.
[0073] In actual application, each read in the BAM file has a corresponding reference name, indicating the chromosome or sequence name on the reference genome to which the read is aligned. Reads with coordinates located in the range of the ABL1 gene are screened from the BAM file as candidate housekeeping gene sequences, and the coordinate range of the ABL1 gene is chr9:133730369-133730483, chr9:133738149-133738223.
[0074] Specifically, the screening of the real housekeeping gene sequence from the candidate housekeeping gene sequence according to the alignment state and the confidence comprises:
[0075] screening a candidate housekeeping gene sequence with correct alignment direction and unique alignment within an expected range of alignment distance from the candidate housekeeping gene sequence according to the alignment state, as a predicted housekeeping gene sequence;
[0076] screening a predicted housekeeping gene sequence with a MAPQ value greater than or equal to a preset confidence threshold from the predicted housekeeping gene sequence according to the confidence, as a real housekeeping gene sequence.
[0077] In actual application, the alignment state (flags) is a set of binary flag bits in the BAM file for describing the alignment state of the read. Different combinations of flag values can represent different alignment characteristics of the read. The candidate housekeeping gene sequence with correct alignment direction and unique alignment within an expected range of alignment distance can be screened from the candidate housekeeping gene sequence according to the flags, as a predicted housekeeping gene sequence. The confidence (MAPQ) represents the confidence of the read alignment to the reference genome. The higher the MAPQ value, the more reliable the alignment result. The predicted housekeeping gene sequence with a MAPQ value less than a preset confidence threshold is filtered to remove false housekeeping gene sequences, and the real housekeeping gene sequence is obtained.
[0078] Specifically, the calculation of the second copy number of the housekeeping gene according to the real housekeeping gene sequence comprises:
[0079] calculating the proportion of the number of bases marked with M according to the CIGAR value of the real housekeeping gene sequence, and calculating the number of real housekeeping gene sequences with a proportion of the number of bases marked with M above a preset proportion threshold, as the second copy number of the housekeeping gene.
[0080] In practical applications, the preset proportion threshold is set to 95%, the CIGAR value is used to describe the detailed matching condition when the read is aligned with the reference genome, including matching (M), insertion (I), deletion (D) and the length of the operation, in this application, the number of real housekeeping gene sequences with the proportion of base number marked as M being more than 95% is calculated, and the number is taken as the second copy number of the housekeeping gene, and the real housekeeping gene sequence with the proportion of base number marked as M being more than 95% indicates that the matching degree of the real housekeeping gene sequence with the reference genome is very high, and there is almost no insertion or deletion.
[0081] Step S5, dividing the first copy number by the second copy number to obtain the micro residual lesion ratio of the fusion gene.
[0082] In practical applications, in order to verify the sensitivity of the present application for MRD monitoring, the present application makes related tests, selects 14 standard samples (MRD8701, MRD8702, MRD8801, MRD8802, MRD8901, MRD8902, MRD0005-1, MRD0005-2, MRD0002-1, MRD0002-2, MRD0001-1, MRD0001-2, MRD00005-1, MRD00005-2), adjusts BCR::ABL1 / ABL1 to different ratios by diluting the concentration, the ratio of BCR::ABL1 / ABL1 of MRD8701 and MRD8702 is 0.05%, the ratio of BCR::ABL1 / ABL1 of MRD8801 and MRD8802 is 0.02%, the ratio of BCR::ABL1 / ABL1 of MRD8901 and MRD8902 is 0.01%, the ratio of BCR::ABL1 / ABL1 of MRD0005-1 and MRD0005-2 is 0.005%, the ratio of BCR::ABL1 / ABL1 of MRD0002-1 and MRD0002-2 is 0.002%, the ratio of BCR::ABL1 / ABL1 of MRD0001-1 and MRD0001-2 is 0.001%, the ratio of BCR::ABL1 / ABL1 of MRD00005-1 and MRD00005-2 is 0.0005%, and selects two normal human negative samples MRD26E13 and MRD28E13, then verifies whether the detection ratio of BCR::ABL1 / ABL1 is consistent with the theoretical ratio by the method, wherein the last column is the theoretical ratio, and the second last column is the detection ratio of the method, and the test data is shown in Table 1, and Table 1 is as follows:
[0083] Table 1 blood fusion gene NGS method different ratio gradient results
[0084] Verification test number BCR::ABL1 / ABL1 (%) Standard MRD8701 0.0673 0.05% standard MRD8702 0.2024 0.05% standard MRD8801 0.0384 0.02% standard MRD8802 0.1548 0.02% standard MRD8901 0.0985 0.01% standard MRD8902 0.0363 0.01% standard MRD0005-1 0.0013 0.005% standard MRD0005-2 0.0016 0.005% standard MRD0002-1 0.0017 0.002% standard MRD0002-2 0.0007 0.002% standard MRD0001-1 0.0005 0.001% standard MRD0001-2 0.0006 0.001% standard MRD00005-1 0.0004 0.0005% standard MRD00005-2 0.0006 0.0005% standard MRD26E13 0 BCR::ABL1 P210 negative MRD28E13 0 BCR::ABL1 P210 negative
[0085] In Table 1, NGS (Next-Generation Sequencing) is a high-throughput sequencing method, BCR::ABL1 is a fusion gene, and the values under BCR::ABL1 in Table 1 represent the first copy number, the values under ABL1 represent the second copy number, and BCR::ABL1 / ABL1 (%) represents the ratio of the first copy number to the second copy number. According to Table 1, the sensitivity of the blood tumor fusion gene minimal residual disease monitoring method of the application can reach 10 -6 , the sensitivity of the conventional technology RQ-PCR is 10 -4 ~ 10 -5 , the sensitivity of MFC is 10 -3 ~ 10 -5 Compared with other MRD monitoring methods (RQ-PCR or MFC), the sensitivity of the application is 1~2 logs higher, and the application has high sensitivity and specificity.
[0086] The blood tumor fusion gene minimal residual disease monitoring method of the application realizes high sensitivity and specificity of MRD monitoring by comparing the high-throughput sequencing data of the fusion gene to the reference genome, then screening and filtering the first alignment result to obtain the real fusion gene sequence and the real housekeeping gene sequence, then calculating the copy number of the fusion gene and the copy number of the housekeeping gene, and calculating the MRD ratio according to the copy number of the fusion gene and the copy number of the housekeeping gene.
[0087] As shown in Figure 2 , the application also provides a blood tumor fusion gene minimal residual disease monitoring system, comprising:
[0088] The acquisition module 10 is configured to acquire fusion gene information of a sample and high-throughput sequencing data containing all genetic information of the sample.
[0089] The comparison and screening module 20 is configured to compare the high-throughput sequencing data to a reference genome without removing PCR duplicates to obtain a first alignment result, and screen split reads and discordant pair reads from the first alignment result as candidate fusion gene sequences.
[0090] The comparison, filtering and calculation module 30 is configured to compare the candidate fusion gene sequences to an annotation file of the reference genome to obtain a second alignment result, obtain a predicted fusion gene sequence according to the overlap of the second alignment result, filter the predicted fusion gene sequence to obtain a real fusion gene sequence, and calculate a first copy number of the fusion gene according to the real fusion gene sequence.
[0091] The screening and calculation module 40 is used to screen out candidate housekeeping gene sequences that are aligned to the chromosomal region where the housekeeping gene is located from the first alignment result, screen out the real housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence level, and calculate the second copy number of the housekeeping gene based on the real housekeeping gene sequence.
[0092] MRD ratio calculation module 50 is used to divide the first copy number by the second copy number to obtain the minimal residual lesion ratio of the fusion gene.
[0093] The various modules of the aforementioned hematologic malignancy fusion gene minimal residual disease monitoring system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules and units can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0094] like Figure 3 As shown, an embodiment of the present invention discloses a computer device, including a memory and a processor, wherein the memory stores a computer program;
[0095] The computer device can be a server, and its internal structure diagram can be as follows: Figure 3 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the method for monitoring minimal residual disease lesions of hematologic malignancies fusion genes described in the above embodiments.
[0096] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0097] This invention also discloses a computer-readable storage medium storing a computer program that causes a computer to execute the method for monitoring minimal residual disease fusion genes in hematologic tumors described in the above embodiments.
[0098] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0099] The technical features of the above embodiments can be combined in any manner. In order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not contradict, they should be considered as the scope of the present disclosure.
Claims
1. A method for monitoring minimal residual disease lesions caused by fusion genes in hematological malignancies, characterized in that, include: Obtain fusion gene information from the sample and high-throughput sequencing data containing all genetic information of the sample; Without deduplication, the high-throughput sequencing data is compared with a reference genome to obtain a first alignment result. Split reads and discordantpair reads are selected as candidate fusion gene sequences from the first alignment result. The candidate fusion gene sequence is aligned to the annotation file of the reference genome to obtain a second alignment result. The predicted fusion gene sequence is obtained based on the overlap of the second alignment result. The predicted fusion gene sequence is filtered to obtain the real fusion gene sequence. The first copy number of the fusion gene is calculated based on the real fusion gene sequence. Candidate housekeeping gene sequences that align to the chromosomal region where the housekeeping gene is located are selected from the first alignment results. The actual housekeeping gene sequences are selected from the candidate housekeeping gene sequences based on the alignment status and confidence level. The second copy number of the housekeeping gene is calculated based on the actual housekeeping gene sequence. Dividing the first copy number by the second copy number yields the ratio of minimal residual lesions of the fusion gene.
2. The method for monitoring minimal residual disease lesions of fusion genes in hematological malignancies as described in claim 1, characterized in that, The step of filtering the predicted fusion gene sequence to obtain the actual fusion gene sequence includes: If the breakpoints of the predicted fusion gene sequence are not annotated, then the breakpoints that support the most sequences and whose distances are within a preset base range are merged as the predicted fusion gene. If the predicted fusion gene sequence only supports split reads, then the predicted fusion gene sequence with paired-end sequencing that does not reach the preset base threshold when the number of bases aligned to both ends of the breakpoint is filtered out. If the predicted fusion gene sequence has multiple transcripts, then the transcripts other than the one supported by the most reads will be filtered out. Filter out predicted fusion gene sequences whose random matching probability is greater than the probability threshold; Filter out the predicted fusion gene sequences with disordered pairings to obtain the true fusion gene sequences.
3. The method for monitoring minimal residual disease lesions of fusion genes in hematological malignancies as described in claim 1, characterized in that, The step of selecting the true housekeeper gene sequence from the candidate housekeeper gene sequences based on alignment status and confidence level includes: Based on the alignment status, candidate housekeeper gene sequences with unique alignments that are within the expected range and have the correct orientation are selected from the candidate housekeeper gene sequences and used as predicted housekeeper gene sequences. Based on the confidence level, predicted housekeeper gene sequences with MAPQ values greater than or equal to a preset confidence threshold are selected from the predicted housekeeper gene sequences and used as true housekeeper gene sequences.
4. The method for monitoring minimal residual disease lesions of fusion genes in hematological malignancies as described in claim 1, characterized in that, The calculation of the first copy number of the fusion gene based on the actual fusion gene sequence includes: The number of the actual fusion gene sequences is calculated as the first copy number.
5. The method for monitoring minimal residual disease lesions of fusion genes in hematological malignancies as described in claim 1, characterized in that, The calculation of the second copy number of the housekeeping gene based on the actual housekeeping gene sequence includes: The percentage of bases marked with the character M is calculated based on the CIGAR value of the real housekeeper gene sequence. The number of real housekeeper gene sequences whose percentage of bases marked with the character M is above a preset percentage threshold is then used as the second copy number of the housekeeper gene.
6. The method for monitoring minimal residual disease lesions of fusion genes in hematological malignancies as described in claim 1, characterized in that, The step of obtaining the predicted fusion gene sequence based on the overlap of the second alignment result includes: Pre-set the minimum number of supported read segments or the minimum overlap length as the overlap threshold; When the minimum number of supported reads is used as the overlap threshold, candidate fusion gene sequences whose number of supported reads reaches the overlap threshold are used as predicted fusion gene sequences. When the minimum overlap length is used as the overlap threshold, candidate fusion gene sequences whose overlap length reaches the overlap threshold are used as predicted fusion gene sequences.
7. The method for monitoring minimal residual disease lesions of fusion genes in hematological malignancies as described in claim 1, characterized in that, The housekeeping gene is the ABL1 gene.
8. A monitoring system for minimal residual disease lesions caused by fusion genes in hematological malignancies, characterized in that, include: The acquisition module is used to acquire the fusion gene information of the sample and the high-throughput sequencing data containing all the genetic information of the sample; The alignment and screening module is used to compare the high-throughput sequencing data with a reference genome without deduplication by performing PCR on the high-throughput sequencing data to obtain a first alignment result, and to screen split reads and discordant pair reads from the first alignment result as candidate fusion gene sequences. The alignment and filtering calculation module is used to align the candidate fusion gene sequence to the annotation file of the reference genome to obtain a second alignment result, obtain a predicted fusion gene sequence based on the overlap of the second alignment result, filter the predicted fusion gene sequence to obtain the real fusion gene sequence, and calculate the first copy number of the fusion gene based on the real fusion gene sequence. The screening and calculation module is used to screen out candidate housekeeping gene sequences that are aligned to the chromosomal region where the housekeeping gene is located from the first alignment results, screen out the real housekeeping gene sequences from the candidate housekeeping gene sequences according to the alignment status and confidence level, and calculate the second copy number of the housekeeping gene based on the real housekeeping gene sequence. The MRD ratio calculation module is used to divide the first copy number by the second copy number to obtain the minimal residual disease ratio of the fusion gene.
9. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Method, system and equipment for detecting internal tandem repetition and storage medium
CN121331232A