High-sensitivity detection method for fusion gene and application thereof
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-03-31
- Publication Date
- 2026-08-07
AI Technical Summary
然而,由于mRNA的不稳定性,所有算法皆对融合基因的转录水平有较严格的限制条件,导致融合基因检测的敏感性受限于测序深度和转录水平过滤阈值,进而丧失了对低丰度融合基因的有效检出能力
[0017]与现有技术相比,本发明具有的有益效果至少包括:
Smart Images

Figure CN122531491A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics analysis technology, specifically relating to a highly sensitive detection method for fusion genes and its application. Background Technology
[0002] Fusion genes are chimeric genes formed by the abnormal joining of the coding sequences of two independent genes, typically driven by genomic structural variations such as chromosomal translocations, inversions, insertions, or fragmentation. When a chromosomal break-rejoining event causes two genes originally located on different chromosomes or at different locations on the same chromosome to join, the transcribed mRNA contains partial exons of both genes, ultimately translating into a fusion protein with a novel function, potentially driving the development of various diseases. For example, the classic... PML - RARA The fusion gene is formed by the t(15;17)(q24;q21) translocation, in which the gene on chromosome 15... PML The gene breakpoint is located in exon 6, while on chromosome 17... RARA The gene breakpoint is located in exon 3, leading to PML The coiled helical structure domain and RARA The DNA-binding domain and ligand-binding domain fuse. This chimeric structure disrupts... RARA This mediates the bone marrow myeloid differentiation function, ultimately leading to the development of acute myeloid leukemia.
[0003] While the importance of fusion genes is widely recognized, their detection still faces multiple technical challenges. Traditional cytogenetic methods (such as karyotype analysis) are the gold standard for fusion gene detection, but their resolution is limited, and some subtle rearrangements may be missed. Other traditional methods, such as fluorescence in situ hybridization and real-time quantitative PCR, are highly sensitive, but they can only target known fusion sites and cannot detect complex or unknown fusion events, such as chromosome fragmentation or occult translocations.
[0004] In recent years, with the development of high-throughput RNA sequencing (RNA-seq) technology, RNA-seq-based strategies for detecting fusion genes have made significant progress. Currently, various fusion gene detection algorithms are available, including the classic STAR-Fusion and Arriba (two default algorithms used in various databases), TrinityFusion and JAFFA-Assembly based on de novo fusion transcript assembly, and FusionCatcher optimized for computational efficiency. A systematic comparison of the performance of these algorithms in the literature revealed that read mapping-based algorithms (such as STAR-Fusion and Arriba) perform best in both accuracy and detection speed. This advantage mainly stems from the fact that both STAR-Fusion and Arriba are based on the STAR (Spliced Transcripts Alignment to a Reference) alignment tool, which is specifically optimized for RNA-seq data. Its high sensitivity and accuracy enable it to effectively identify various mRNA splicing events, including typical and atypical splicing, as well as chimeric (fusion) transcripts. However, due to the instability of mRNA, all algorithms have strict limitations on the transcriptional level of fusion genes, which leads to the sensitivity of fusion gene detection being limited by sequencing depth and transcriptional level filtering thresholds, thus losing the ability to effectively detect low-abundance fusion genes.
[0005] Therefore, developing a highly sensitive method for detecting fusion genes, especially for the accurate identification of low-abundance fusion genes, is of great significance for improving the accuracy of disease diagnosis and research. Summary of the Invention
[0006] In view of the above, the purpose of this invention is to provide a highly sensitive detection method for fusion genes and its application. Through a multi-level quality control strategy including data preprocessing, sequence similarity screening, authenticity screening, specificity screening, and reliability screening, fusion genes can be automatically annotated, effectively identifying potential false positive fusion gene events in the STAR-Fusion output results, and significantly improving the sensitivity and accuracy of fusion gene detection.
[0007] To achieve the above-mentioned objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a highly sensitive detection method for fusion genes, comprising the following steps: We acquired fusion gene data, along with the corresponding wild-type gene expression matrix, reference genome information, and data batch information. After preprocessing the fusion gene data to filter out low-expression fusion genes, we annotated the chaperone sequences at both ends of the fusion genes based on the reference genome information. The preprocessed fusion genes were subjected to sequence similarity screening, which involved comparing the similarity of the chaperone sequences at both ends of a single fusion gene and the similarity between the chaperone sequences at both ends of multiple fusion genes to exclude false positive fusion genes with similarity higher than a preset similarity threshold. After sequence similarity screening, fusion genes are subjected to authenticity screening, specificity screening, and reliability screening. Authenticity screening is used to exclude false positive fusion genes based on at least one of the following: wild-type gene expression matrix, annotation information of chaperone sequences at both ends of the fusion gene, and data batch information. Specificity screening is used to merge scattered fusion gene reads. Reliability screening is used to retain fusion genes that exist in the form of fusion proteins. The data from the retained fusion genes will be compiled for subsequent analysis.
[0008] Preferably, in the preprocessing, filtering out low-expression fusion genes includes: A first threshold is set based on cross-breakpoint reads and cross-breakpoint fragments. The first threshold includes a lower limit for cross-breakpoint reads or a lower limit for the sum of cross-breakpoint reads and cross-breakpoint fragments. Fusion genes below the first threshold are filtered out as low-expression fusion genes.
[0009] Preferably, the sequence similarity screening includes: Similarity detection within fusion genes: Extract the partner sequences of predetermined lengths from both ends of a single fusion gene and perform sequence alignment to obtain the length of similar sequences and similarity scores; Intra-sample fusion gene similarity detection: Construct the mate sequences at both ends of all fusion genes in the sample and perform pairwise sequence alignment to obtain the length of similar sequences and similarity scores; Similarity screening: Based on the similar sequence length and similarity score output by sequence alignment, false positive fusion genes are identified using preset similarity conditions. The similarity conditions include similar sequence length ≥ first length, or second length ≤ similar sequence length < first length and similar score ≥ first score, where the first length is greater than the second length. Fusion genes that meet the similarity conditions are filtered out as false positive fusion genes.
[0010] Preferably, the sequence similarity screening further includes: Based on the number of different fusion genes mapped to the fusion gene reads, the cross-breakpoint reads and cross-breakpoint fragments are recalibrated, and a second threshold is set. The second threshold includes the lower limit of the calibrated cross-breakpoint reads or the lower limit of the sum of the calibrated cross-breakpoint reads and the calibrated cross-breakpoint fragments. Fusion genes below the second threshold are filtered out as false positive fusion genes.
[0011] Preferably, the authenticity screening employs one or more combinations of the following three methods: Based on the wild-type gene expression matrix, fusion genes with retained sequence lengths exceeding the sequencing read length at both ends but corresponding wild-type gene expression levels of zero were removed. Alternatively, based on the annotation information of the mate sequences at both ends of the fusion gene, fusion genes with retained sequences at both ends exceeding the preset support point length threshold and with a detection result ratio lower than the preset proportion that have long support points can be removed. Alternatively, based on data batch information, batch-related fusion genes that appear only in a single batch and whose frequency is higher than a preset batch threshold can be removed.
[0012] Preferably, the specific screening includes: Based on the annotation information of the chaperone sequences at both ends of the fusion gene, fusion gene reads with different breakpoint locations but the same mature mRNA sequence are merged to eliminate scattered fusion gene reads caused by intron non-cleavage.
[0013] Preferably, the reliability screening includes at least one of the following screening criteria: Fusion genes that retain both chaperone sequences at both ends as protein-coding sequences; Alternatively, only the fusion gene that exists as a fusion protein may be retained.
[0014] Preferably, the method further includes: The coexistence of the selected fusion genes was tested to identify twin fusion genes and sister fusion genes that coexist in the sample for subsequent analysis.
[0015] Secondly, embodiments of the present invention also provide a highly sensitive detection device for fusion genes, which is implemented using the above-mentioned highly sensitive detection method for fusion genes, including: a data preprocessing module, a similarity screening module, a multi-dimensional screening module, and a detection data summarization module; The data preprocessing module is used to acquire fusion gene data and the corresponding wild-type gene expression matrix, reference genome information and data batch information. After preprocessing the fusion gene data to filter out low-expression fusion genes, the mate sequences at both ends of the fusion gene are annotated based on the reference genome information. The similarity screening module is used to screen the preprocessed fusion genes for sequence similarity. It excludes false positive fusion genes with similarity higher than a preset similarity threshold by comparing the similarity of the chaperone sequences at both ends of a single fusion gene and the similarity between the chaperone sequences at both ends of multiple fusion genes. The multi-dimensional screening module is used to perform authenticity screening, specificity screening, and reliability screening on the fusion genes after sequence similarity screening. Among them, authenticity screening is used to exclude false positive fusion genes based on at least one of the following: wild-type gene expression matrix, annotation information of chaperone sequences at both ends of the fusion gene, and data batch information; specificity screening is used to merge scattered fusion gene reads; and reliability screening is used to retain fusion genes that exist in the form of fusion proteins. The detection data aggregation module is used to aggregate the data of the finally retained fusion genes for subsequent analysis.
[0016] Thirdly, embodiments of the present invention also provide an application of the highly sensitive detection method for fusion genes as described above, specifically for tumor marker detection, evaluation of tumor treatment efficacy, or auxiliary diagnosis of developmental related diseases.
[0017] Compared with the prior art, the beneficial effects of the present invention include at least the following: This invention filters out low-expression fusion gene data through preprocessing, and then filters out false-positive fusion genes based on multi-level screening. Sequence similarity screening eliminates false positives caused by high sequence homology through two-level sequence comparison within the fusion gene and within the sample. Authenticity screening comprehensively excludes false-positive fusion events from multiple dimensions by considering wild-type expression, breakpoint support, and batch distribution characteristics. Combining specificity screening and reliability screening, scattered gene reads are merged and functional fusion genes with the potential to form fusion proteins are retained, thereby achieving automation, high sensitivity, and accuracy in fusion gene detection. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating a highly sensitive detection method for fusion genes provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the raw data acquisition process required by the SNIFFER algorithm provided in this embodiment of the invention; Figure 3 This is a schematic diagram illustrating the relevant terms related to typical fusion genes in the embodiments of the present invention; Figure 4 This is a schematic diagram illustrating the core assumptions in sequence similarity screening of fusion genes provided in the embodiments of the present invention; Figure 5This is a schematic diagram of similar sequence detection in the sequence similarity screening of fusion genes provided in this embodiment of the invention; Figure 6 This is a schematic diagram illustrating the proportion of data deleted in the four filtering steps of the SNIFFER algorithm provided in this embodiment of the invention; Figure 7 This is a comparison of the correlation between the fusion gene results of the SNIFFER algorithm and the STAR-Fusion algorithm provided in this embodiment of the invention and the chromosomal vulnerable sites; Figure 8 This is a schematic diagram of the structure of a highly sensitive detection device for fusion genes provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.
[0021] The inventive concept of this invention is as follows: In view of the problem that the sensitivity and accuracy of existing fusion gene detection are insufficient, the embodiments of this invention provide a highly sensitive detection method for fusion genes and its application. First, low-expression fusion genes are filtered out through preprocessing. Then, a multi-level detection system combining sequence similarity screening, authenticity screening, specificity screening and reliability screening is constructed. This effectively improves the detection capability of low-abundance fusion genes while significantly reducing the false positive rate, thus achieving high sensitivity and high accuracy in fusion gene detection.
[0022] like Figure 1 As shown in the example, this embodiment provides a highly sensitive detection method for fusion genes (hereinafter referred to as the SNIFFER algorithm), which includes the following steps: S1. Obtain fusion gene data, the corresponding wild-type gene expression matrix, reference genome information, and data batch information. Preprocess the fusion gene data to filter out low-expression fusion genes, and then annotate the partner sequences at both ends of the fusion gene based on the reference genome information.
[0023] In this embodiment, multiple files containing annotation information of fusion gene data are first acquired. The files of fusion gene data are derived from the fusion gene output results of STAR-Fusion with specific parameters, the wild-type gene expression matrix of the sample, the reference genome information used, and the data batch information. The STAR-Fusion parameters include, but are not limited to, alignment strategy and splicing detection parameters (STAR_SJDBoverhangMin: 1; min_novel_junction_support: 1), multiple alignment and mismatch filtering parameters, intron and paired read distance limit parameters (STAR_max_mate_dist: 1000000), alignment quality filtering parameters, splicing database and output parameters, soft splicing parameters, chimeric (fusion) read key parameters (min_junction_reads: 1), fusion gene support threshold parameters (min_sum_frags: 2; min_spanning_frags_only: 3), fusion gene filtering strategy parameters (require_LDAS: 0; no_single_fusion_per_breakpoint; no_annotation_filter), fusion read and expression level parameters (min_FFPM 0.0001), and FusionInspector deep validation parameters. The STAR-Fusion output based on these parameters is used as input data. Further preprocessing of the fusion gene data includes filtering low-quality data in the input file and annotating the fusion genes based on the reference genome information provided by the user, including information on the function of the genes at both ends, gene transcript information, specific breakpoint locations, the sequence segments and lengths retained after fusion, and whether the fusion affects protein coding.
[0024] Specifically, such as Figure 2 As shown, fusion gene data can be obtained from raw RNA sequencing data using STAR-Fusion and recommended custom parameters, or from databases such as TCGA, TARGET, and CCLE, or from experimental data in publicly available literature. For the SNIFFER algorithm, fusion gene data best suited for fusion gene sequencing is recommended to be obtained directly from raw RNA sequencing data.
[0025] After acquiring the fusion gene data, the data needs to be merged, quality controlled, and annotated. Specifically: merging the data involves merging all single-sample data generated by STAR-Fusion and selecting the necessary annotation columns; quality control includes... Figure 3In the definition of fusion genes, a cross-breakpoint read refers to a sequencing read that crosses a fusion breakpoint. Part of this read aligns to one gene and the other part aligns to another gene, which is relatively strong evidence for the detection of fusion genes. On the other hand, a cross-breakpoint fragment refers to a sequencing fragment that crosses two different gene regions. The two ends of these fragments align to different genes, but the middle does not cover the fusion breakpoint, which is auxiliary evidence for the detection of fusion genes. The default first threshold for screening low-expression fusion genes in the SNIFFER algorithm is set as follows: all fusion results that meet the following conditions will be retained for subsequent analysis: (1) cross-breakpoint reads ≥ 1 or (2) cross-breakpoint reads + cross-breakpoint fragments ≥ 2. Those below the above conditions are low-expression fusion genes. The values of breakpoint reads and cross-breakpoint reads + cross-breakpoint fragments are adjustable parameters in SNIFFER and can be customized by users. The fusion gene annotation information is added based on the reference genome annotation file, including the function of each gene at both ends of the fusion, gene transcript information, specific breakpoint location (on which exon or intron), the sequence segment and length retained after fusion, and whether the fusion affects protein coding, etc., for further screening.
[0026] S2, sequence similarity screening is performed on the preprocessed fusion genes. False positive fusion genes with similarity higher than a preset similarity threshold are excluded by comparing the similarity of the partner sequences at both ends of a single fusion gene and the similarity between the partner sequences at both ends of multiple fusion genes.
[0027] In this embodiment, the main reason STAR-Fusion outperforms other alignment tools in fusion gene detection lies in the optimization of its soft clipping algorithm. Soft clipping allows portions of RNA-seq reads to be cut off instead of aligned with the reference genome. Especially in RNA-seq data, soft clipping can flexibly adapt to highly variable or structurally complex gene regions, thus better identifying and handling reads across splice sites, which is crucial for detecting fusion gene transcripts. However, although soft clipping reduces errors in forced alignment, allowing partial read misalignment can introduce low-confidence false positives in some cases, especially short reads that cross fusion breakpoints. Such errors can be amplified dramatically after relaxing the fusion gene expression screening criteria; therefore, this embodiment provides sequence similarity screening based on fusion genes.
[0028] The sequence similarity screening process is based on a core assumption: true fusion genes should exhibit unique sequence characteristics at both ends of their breakpoint regions, while pseudo-fusion genes show high similarity at the sequence level. This mainly occurs in the following two situations: (1) One end of the fusion gene is highly similar to its non-fusion region, leading to mismatch splicing, such as... Figure 4As shown in (a), if the A gene sequence and the B gene sequence to the right of the breakpoint of the fusion gene AB are highly similar, then the AB fusion transcript may be the wild-type transcript of the A gene, and vice versa; (2) If the breakpoint sequences of multiple fusions in a sample are similar to each other, it suggests that it may be a duplicate report of a class of events, such as Figure 4 As shown in (b), if the B gene sequence to the right of the breakpoint of fusion gene AB and the C gene sequence to the right of the breakpoint of another fusion gene AC are highly similar in a single sample, further screening based on the reliability of the fusion genes should be conducted. In addition, there may be cases where fusion genes DE and FG are highly similar in a single sample, or where multiple other fusion genes are paired similarly.
[0029] To detect sequence similarity, the SNIFFER algorithm incorporates BLASTN (Basic Local Alignment Search Tool for Nucleotides), a tool for sequence similarity research. Its principle is to segment the query sequence into short words, quickly search for matching regions of these words in the target database, and then extend this to generate local alignments with high similarity scores. The similarity length and similarity score generated by BLASTN are saved for subsequent filtering. Based on the above sequence similarity hypothesis, the implementation example includes a two-stage BLASTN analysis process, addressing identification and cleanup for these two scenarios respectively.
[0030] The first stage corresponds to Figure 4 In (a): Taking a sequencing length of 150bp as an example, firstly, for each candidate fusion gene, 75bp upstream and downstream sequences are extracted from its left N-terminus and right C-terminus, respectively. BLASTN alignment is then used to determine whether it is a wild-type A gene or a wild-type B gene, and the similarity sequence length and similarity score are obtained. Next, based on the similarity sequence length and similarity score results, such as... Figure 5 As shown, there are two cases. Case 1: If BLASTN detects only one result (hit), its similarity length and similarity score are directly recorded; Case 2: If there are multiple hits, their positional relationship is extracted to determine whether they are translational repetitions of different segments. If the positional spacing of multiple hits is the same, they can be identified as different translational views of a similar region, and the sequence length and similarity score are calculated by merging them. After merging the translational repetition hits, the hits with the highest similarity sequence length and similarity score are selected. In the embodiment, the default screening conditions for sequence similarity are set: if (1) the similarity sequence length is ≥24bp, or (2) 15bp≤similar sequence length<24bp and the similarity score is greater than 20 (recommended 20.6), it is considered to be a pseudo-fusion gene. These two conditions are adjustable parameters in SNIFFER and can be customized by the user.
[0031] The second stage corresponds to Figure 4 In step (b): First, 75bp sequences are extracted from the N-terminus and C-terminus of all candidate fusions in a sample, and combined into a single-sample N-terminal sequence database and a C-terminal sequence database. Next, pairwise BLASTN is performed on all gene sequence combinations within the N-terminal and C-terminal sequence databases. Due to the potentially large data volume, a multi-threaded parallel alignment method is used to increase processing speed. The method for merging Hits is the same as in the first stage. After identifying potentially false positive fusion gene combinations, scores are assigned based on the fusion gene's expression in the sample, its frequency across all samples, its average expression across all samples, the annotation status of the fusion gene, whether the fusion gene has long sequence support, and the Shannon entropy of the fusion gene sequence. The fusion gene combination with the highest score is retained.
[0032] Finally, the embodiments further screened the fusion genes based on formulas (1) and (2), wherein This indicates that a certain cross-breakpoint read segment or cross-breakpoint fragment is mapped to several different fusions. Therefore... and This represents a recalculation of the frequency of mappings for multiple-mapped fusion gene reads. Multiple-mapped fusion gene reads represent high similarity of fusion sequences; the more different the fusion gene reads mapped to, the lower their reliability and importance. Therefore, the SNIFFER algorithm is also based on... and The filtering is performed, and the default second threshold is set as follows: as long as (1) is satisfied. Or (2) All fusion results will be retained for subsequent analysis, among which and The value is an adjustable parameter in SNIFFER, which can be customized by the user.
[0033] (1), (2), in, This represents the total number of junction reads or spanning fragments in a given fusion gene. This represents an index of a junction read or spanning fragment within a fusion gene.
[0034] S3. After sequence similarity screening, the fusion genes are subjected to authenticity screening, specificity screening, and reliability screening. Among them, authenticity screening is used to exclude false positive fusion genes based on at least one of the following: wild-type gene expression matrix, annotation information of mate sequences at both ends of the fusion gene, and data batch information; specificity screening is used to merge scattered fusion gene reads; and reliability screening is used to retain fusion genes that exist in the form of fusion proteins.
[0035] S3.1, Authenticity screening of fusion genes.
[0036] In this embodiment, the authenticity of the fusion gene is judged according to three criteria: (1) Since the fusion gene transcript can only be identified by the cross-breakpoint read and the cross-breakpoint fragment, if the fusion gene is long, the read that does not cross the fusion breakpoint will be identified as the wild-type transcript. Therefore, the transcription status of the wild-type gene can be used as a screening condition. Taking the longest 150bp paired-end sequencing as an example, based on the sequence length result of the fusion after the fusion is obtained by the preprocessing of the fusion gene data, the embodiment sets a screening strategy based on the wild-type gene: that is, if one end of the fusion gene is greater than 150bp, but its wild-type gene transcription result is 0, then the fusion gene is deleted; (2) The anchor length of the fusion gene represents the length of the cross-breakpoint read that covers the two genes after crossing the two ends of the breakpoint, such as Figure 3 As shown, the longer the support point length, the higher the reliability of the fusion gene. Considering that the distribution of mRNA reads should be random, if the sequence length of a fusion gene is long enough and the frequency of fusion gene occurrence is high enough, it should be able to cover a large area near the fusion breakpoint. When STAR-Fusion detects fusion genes, it defines fusion genes with a coverage length of more than 25bp at both ends as having long support points. Therefore, the embodiment sets a screening strategy based on the support points of fusion genes: that is, if the length of both ends of a fusion gene is more than 25bp, but more than 95% of the results of the fusion gene do not have long support points, then the fusion gene is deleted. Among them, 25bp and 95% are adjustable parameters in SNIFFER, which can be customized by the user; (3) The detection of fusion genes is directly related to the length of RNA-seq sequencing and the detection principle. There may be batch-related fusion genes due to different batches between different methodologies. Such fusion genes may not actually exist. Therefore, SNIFFER has a default screening strategy for fusion genes based on batch effects: that is, based on batch characteristics from methodology or other inputs, fusion gene results that only appear in a certain batch and have an occurrence rate higher than 50% are deleted. The batch characteristics of the fusion gene data are user-defined batch matrices; 50% is an adjustable parameter in SNIFFER that can be customized by the user.
[0037] S3.2, Specific screening of fusion genes.
[0038] During preliminary research, it was found that the same fusion gene could appear in a single sample but with different fusion breakpoints, making it impossible to determine which fusion gene was dominant. This situation was particularly likely to occur with fusion genes whose fusion breakpoints were located on intron sequences. Since mRNA splicing is a constantly changing and dynamic process, immature fusion gene RNA with incomplete intron sequence splicing may occur, leading STAR-Fusion to falsely report fusion breakpoint locations and disperse fusion gene reads. Therefore, in this embodiment, based on the annotation information of the preserved fusion gene sequence, fusion genes with different breakpoints but the same final mature mRNA sequence were deleted, and all reads of the same fusion gene were merged.
[0039] S3.3, Reliability screening of fusion genes.
[0040] Differences in the reference genome have a significant impact on the detection of fusion genes. The accuracy and integrity of the reference genome directly determine the alignment effect, especially the identification of fusion reads and cross-splicing sites. This step utilizes the annotation of fusion genes after data preprocessing to identify fusion genes that cannot form fusion proteins, including but not limited to those lacking start or stop codons, premature stop codons, and lncRNA fusion genes. By comprehensively judging whether a fusion gene can form a fusion protein, the reliability of the fusion gene's impact on the phenotype is evaluated. Since most coding genes do not vary significantly in different sources or versions of the reference genome, in order to ensure the reliability, reproducibility, and biological importance of fusion gene data, two optional screening conditions are set in the examples: (1) only fusion genes with protein-coding N-terminus and C-terminus are retained; (2) only fusion genes that will exist in the form of fusion proteins are retained, such as fusion genes with a stop codon at the N-terminus will be screened.
[0041] S4. The data from the final retained fusion genes will be compiled for subsequent analysis.
[0042] S4.1, Coexistence detection of fusion genes.
[0043] When detecting fusion genes, both AB and BA forms often exist simultaneously in the same sample, such as... PML-RARA and RARA-PML However, most research focuses only on PML-RARA Above, prompt RARA-PMLFunctionally unimportant genes or potential pseudo-fusion genes are not included. Therefore, SNIFFER integrates the detection of AB and BA phenomena in the same sample. Twin fusion genes are defined as AB and BA always appearing simultaneously; while if AB always appears, BA will also appear (i.e., AB contains BA), then AB and BA are called a pair of sister fusion genes, with AB being the dominant fusion gene among the sister fusion genes. The default example only outputs a CSV file containing potential twin fusion genes and sister fusion genes without filtering; however, the filtering strategy can be adjusted by customizing parameters.
[0044] S4.2, Data Output and Data Bureau Statistics.
[0045] The example returns all deleted data from each step and indicates the specific reason for the deletion of each fusion gene. Simultaneously, the example provides statistical analysis based on the final selected fusion gene data, according to different fusion gene breakpoints, including information such as fusion frequency and average fusion expression level. Furthermore, it annotates protein information and other information such as potential roles in cancer for fusion partners based on the Uniprot and COSMIC databases. This step, through aggregation and automatic annotation using large databases, facilitates subsequent in-depth research.
[0046] like Figure 6 As shown, the proportions of data deleted in the four screening steps in the embodiment were 13.4%, 4.4%, 0.6%, and 16.6%, respectively. Furthermore, since chromosome breakage caused by chromosome fragility contributes to the generation of fusion genes, there should be a certain positive correlation between the two. Therefore, to indicate the accuracy of fusion gene detection, the embodiment examined the association between chromosome fragile sites identified based on whole-genome sequencing technology and fusion genes. To verify this hypothesis, human chromosome fragile site data obtained from the database were used, and correlation analysis was performed by counting the number of chromosome fragile sites within each chromosome band and the corresponding fusion genes. For example... Figure 7 As shown, compared to the current best STAR-Fusion algorithm, the correlation between the fusion gene results provided in this example and the chromosomal vulnerable sites has improved from 0.46 to 0.63. This result indicates that the fusion gene results provided in this example not only outperform the current best STAR-Fusion algorithm but are also supported by WGS data, further enhancing their reliability.
[0047] In summary, the highly sensitive detection method for fusion genes provided by this invention can effectively improve the detection capability of low-abundance fusion genes and significantly reduce the false positive rate through a multi-level screening strategy, achieving high sensitivity and high accuracy in fusion gene detection. It can be applied to fields such as tumor marker detection, disease diagnosis and prognosis assessment, drug efficacy monitoring, and developmental-related diseases.
[0048] Based on the same inventive concept, such as Figure 8 As shown, this embodiment of the invention also provides a highly sensitive detection device 800 for fusion genes, including: a data preprocessing module 810, a similarity screening module 820, a multi-dimensional screening module 830, and a detection data aggregation module 840.
[0049] The data preprocessing module 810 is used to acquire fusion gene data and the corresponding wild-type gene expression matrix, reference genome information and data batch information. After preprocessing the fusion gene data to filter out low-expression fusion genes, the mate sequences at both ends of the fusion gene are annotated based on the reference genome information.
[0050] The similarity screening module 820 is used to screen the preprocessed fusion genes for sequence similarity. It excludes false positive fusion genes with similarity higher than a preset similarity threshold by comparing the similarity of the partner sequences at both ends of a single fusion gene and the similarity between the partner sequences at both ends of multiple fusion genes.
[0051] The multi-dimensional screening module 830 is used to perform authenticity screening, specificity screening, and reliability screening on fusion genes after sequence similarity screening. Among them, authenticity screening is used to exclude false positive fusion genes based on at least one of the wild-type gene expression matrix, annotation information of the mate sequences at both ends of the fusion gene, and data batch information; specificity screening is used to merge scattered fusion gene reads; and reliability screening is used to retain fusion genes that exist in the form of fusion proteins.
[0052] The detection data aggregation module 840 is used to aggregate the data of the finally retained fusion genes for subsequent analysis.
[0053] It should be noted that the highly sensitive detection device for fusion genes provided in the above embodiments should be described using the above-described functional unit division as an example when performing fusion gene detection. Depending on specific needs, the above functions can be assigned to different functional units. These functional units can be divided into different modules within the internal structure of the terminal device or server to achieve all or part of the functions described above. This invention is not limited to the functional unit division method shown in the above embodiments and can be flexibly changed and adjusted according to actual conditions. This flexible functional unit division method can be adjusted and configured according to actual needs, ensuring that the highly sensitive detection device for fusion genes has adaptability and scalability.
[0054] Based on the same inventive concept, the embodiment also provides a highly sensitive detection device for fusion genes, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the highly sensitive detection method for fusion genes described above.
[0055] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be executed by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium, such as RAM, ROM, FLASH, floppy disk, hard disk, etc., or it can be stored in a remote storage cloud. During execution, the computer program encompasses the processes described in the above method embodiments. In various embodiments of the present invention, the memory may be a local volatile memory, such as RAM, or a non-volatile memory, such as ROM, FLASH, floppy disk, hard disk, etc., or it may be a remote storage cloud. The processor may be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field-programmable gate array (FPGA).
[0056] The above embodiments are merely specific application examples of the present invention, intended to explain the invention and not to limit its scope. Any modifications, variations, and equivalent substitutions made to these embodiments under the guidance of the principles and spirit of the present invention should be considered to fall within the protection scope of the present invention.
Claims
1. A highly sensitive method for detecting fusion genes, characterized in that, Includes the following steps: We acquired fusion gene data, along with the corresponding wild-type gene expression matrix, reference genome information, and data batch information. After preprocessing the fusion gene data to filter out low-expression fusion genes, we annotated the chaperone sequences at both ends of the fusion genes based on the reference genome information. The preprocessed fusion genes were subjected to sequence similarity screening, which involved comparing the similarity of the chaperone sequences at both ends of a single fusion gene and the similarity between the chaperone sequences at both ends of multiple fusion genes to exclude false positive fusion genes with similarity higher than a preset similarity threshold. After sequence similarity screening, fusion genes are subjected to authenticity screening, specificity screening, and reliability screening. Authenticity screening is used to exclude false positive fusion genes based on at least one of the following: wild-type gene expression matrix, annotation information of chaperone sequences at both ends of the fusion gene, and data batch information. Specificity screening is used to merge scattered fusion gene reads. Reliability screening is used to retain fusion genes that exist in the form of fusion proteins. The data from the retained fusion genes will be compiled for subsequent analysis.
2. The highly sensitive detection method for fusion genes according to claim 1, characterized in that, The preprocessing includes filtering out low-expression fusion genes, including: A first threshold is set based on cross-breakpoint reads and cross-breakpoint fragments. The first threshold includes a lower limit for cross-breakpoint reads or a lower limit for the sum of cross-breakpoint reads and cross-breakpoint fragments. Fusion genes below the first threshold are filtered out as low-expression fusion genes.
3. The highly sensitive detection method for fusion genes according to claim 1, characterized in that, The sequence similarity screening includes: Similarity detection within fusion genes: Extract the partner sequences of predetermined lengths from both ends of a single fusion gene and perform sequence alignment to obtain the length of similar sequences and similarity scores; Intra-sample fusion gene similarity detection: Construct the mate sequences at both ends of all fusion genes in the sample and perform pairwise sequence alignment to obtain the length of similar sequences and similarity scores; Similarity screening: Based on the similar sequence length and similarity score output by sequence alignment, false positive fusion genes are identified using preset similarity conditions. The similarity conditions include similar sequence length ≥ first length, or second length ≤ similar sequence length < first length and similar score ≥ first score, where the first length is greater than the second length. Fusion genes that meet the similarity conditions are filtered out as false positive fusion genes.
4. The highly sensitive detection method for fusion genes according to claim 3, characterized in that, The sequence similarity screening also includes: Based on the number of different fusion genes mapped to the fusion gene reads, the cross-breakpoint reads and cross-breakpoint fragments are recalibrated, and a second threshold is set. The second threshold includes the lower limit of the calibrated cross-breakpoint reads or the lower limit of the sum of the calibrated cross-breakpoint reads and the calibrated cross-breakpoint fragments. Fusion genes below the second threshold are filtered out as false positive fusion genes.
5. The highly sensitive detection method for fusion genes according to claim 1, characterized in that, The authenticity screening employs one or more of the following three methods in combination: Based on the wild-type gene expression matrix, fusion genes with retained sequence lengths exceeding the sequencing read length at both ends but corresponding wild-type gene expression levels of zero were removed. Alternatively, based on the annotation information of the mate sequences at both ends of the fusion gene, fusion genes with retained sequences at both ends exceeding the preset support point length threshold and with a detection result ratio lower than the preset proportion that have long support points can be removed. Alternatively, based on data batch information, batch-related fusion genes that appear only in a single batch and whose frequency is higher than a preset batch threshold can be removed.
6. The highly sensitive detection method for fusion genes according to claim 1, characterized in that, The specific screening includes: Based on the annotation information of the chaperone sequences at both ends of the fusion gene, fusion gene reads with different breakpoint locations but the same mature mRNA sequence are merged to eliminate scattered fusion gene reads caused by intron non-cleavage.
7. The highly sensitive detection method for fusion genes according to claim 1, characterized in that, The reliability screening includes at least one of the following screening criteria: Fusion genes that retain both chaperone sequences at both ends as protein-coding sequences; Alternatively, only the fusion gene that exists as a fusion protein may be retained.
8. The highly sensitive detection method for fusion genes according to claim 1, characterized in that, The method further includes: The coexistence of the selected fusion genes was tested to identify twin fusion genes and sister fusion genes that coexist in the sample for subsequent analysis.
9. A highly sensitive detection device for fusion genes, implemented using the highly sensitive detection method for fusion genes according to any one of claims 1 to 8, characterized in that, include: Data preprocessing module, similarity filtering module, multi-dimensional filtering module, and detection data aggregation module; The data preprocessing module is used to acquire fusion gene data and the corresponding wild-type gene expression matrix, reference genome information and data batch information. After preprocessing the fusion gene data to filter out low-expression fusion genes, the mate sequences at both ends of the fusion gene are annotated based on the reference genome information. The similarity screening module is used to screen the preprocessed fusion genes for sequence similarity. It excludes false positive fusion genes with similarity higher than a preset similarity threshold by comparing the similarity of the chaperone sequences at both ends of a single fusion gene and the similarity between the chaperone sequences at both ends of multiple fusion genes. The multi-dimensional screening module is used to perform authenticity screening, specificity screening, and reliability screening on the fusion genes after sequence similarity screening. Among them, authenticity screening is used to exclude false positive fusion genes based on at least one of the following: wild-type gene expression matrix, annotation information of chaperone sequences at both ends of the fusion gene, and data batch information; specificity screening is used to merge scattered fusion gene reads; and reliability screening is used to retain fusion genes that exist in the form of fusion proteins. The detection data aggregation module is used to aggregate the data of the finally retained fusion genes for subsequent analysis.
10. The application of a highly sensitive detection method for fusion genes as described in any one of claims 1 to 8, characterized in that, The method can be used for tumor marker detection, evaluation of tumor treatment efficacy, or auxiliary diagnosis of developmental diseases.