A fusion gene interpretation method based on targeted sequencing customization automation

CN116486912BActive Publication Date: 2026-08-21SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310461322.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2026-08-21
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

[0009]本发明的目的是提供一种基于靶向测序定制化自动化的融合基因判读方法,从而解决现有技术中融合基因预测方法存在假阳性结果数量较高、预测特异性较低的问题

Benefits of technology

[0031]然而,本发明基于靶向测序的方法,增加了引物定制化筛选,相比于现有的融合基因分析软件,在现有的报告基础上进行了充分的结果过滤,可以控制融合基因报告阈值以及二次融合检查;本发明的条件过滤可以应用于高通量的样本检测,并实现了流程自动化。因此,对靶向融合基因进行二次序列比对筛选可以实现高特异性的融合候选整合。本发明基于靶向测序建立针对性筛选现有分析软件结果给出融合候选的过滤机制,能够解决现有分析软件结果假阳率高的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486912B_ABST
    Figure CN116486912B_ABST
Patent Text Reader

Abstract

The application relates to a fusion gene judgment method based on targeted sequencing customization automation, which comprises the following steps: 1) targeted sequencing; 2) obtaining preliminary results by using a fusion gene analysis software; 3) obtaining a read sequence containing a breakpoint; 4) conditional filtering: the read sequence containing the breakpoint is subjected to conditional screening according to filtering conditions set by a filter, if the conditions are met, the result is judged as a credible fusion, otherwise, the result is judged as an incredible fusion, thereby realizing a fusion gene judgment based on targeted sequencing customization automation. The application is based on targeted sequencing customization, can accommodate more targeted gene information, and adjusts filtering conditions for a targeted sequencing library. By controlling different input parameters, the precision of the fusion gene judgment result is adjusted, more customized judgment is realized, the method is suitable for batch automatic detection of multiple samples, the specificity of the fusion gene judgment is improved, and a series of sequencing-related analysis results such as variable splicing can be optimized accordingly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fusion gene detection, and more specifically to a customized and automated method for interpreting fusion genes based on targeted sequencing. Background Technology

[0002] A fusion gene is a rearrangement formed by the connection of two originally independent genes. For information on fusion gene identification, see [link to fusion gene identification documentation]. Figure 1 A fusion gene is a new sequence resulting from the rearrangement of two different genes. In interpreting fusion genes, the gene sequences to be fused must belong to two different genes. If the gene sequences have partial overlap or are separated by several bases, they are considered within the reliable range of fusion genes. If the sequences of the two genes are in an inclusion relationship, this type is not included in the reliable range.

[0003] In cancer research, an increasing number of fusion genes have been found to drive mutations in cancer, with identified fusion genes accounting for up to 20% of human cancer incidence and serving as biomarkers for many cancer subtypes. With the continuous development of Next Generation Sequencing (NGS), the availability of sequencing data and the cost of sequencing are constantly decreasing, making NGS one of the effective means of detecting fusion genes. Although whole genome sequencing (WGS) and RNA sequencing are powerful methods for analyzing gene data and have great potential in detecting potential fusion partners, they still have many limitations. First, due to the massive amounts of data generated across the entire genome or transcriptome, the technical requirements for data mining and analysis are relatively high, such as how to filter sequencing errors and low-quality reads from massive sequencing data and detect fusions of genes of interest. Second, different sequencing depths also affect the capture of fusion genes. Besides the high cost of high depth, genes with low expression levels are prone to insufficient sensitivity in fusion gene detection, thus significantly limiting the detection of fusion genes.

[0004] Currently, analyzing fusion genes from sequencing data presents significant challenges. Several analysis software programs, such as Arriba, STARFIusion, EricScript, and FusionCatcher, have been developed in recent years to detect fusion genes using short-read sequencing. Tools for predicting fusion genes primarily rely on aligning sample sequencing data with reference genome or transcriptome data, reporting candidate fusion genes from discrepancies. However, most of these software programs only validate their performance on simulated or cell line data, lacking verification against clinical sample analysis results. Secondly, the results from different analysis software programs for the same sample show little overlap, possibly because the software itself cannot detect true fusion genes, or because the software contains a large number of spurious fusion genes, resulting in a lack of superior performance between different software programs. Furthermore, the software exhibits poor stability, showing low sensitivity even with large sample sizes.

[0005] The existing technologies mainly employ the following two solutions:

[0006] 1) Filter fusion genes based on sequencing data (e.g., US patent US11473137B2). This method does not target specific primer sequencing libraries and performs indiscriminate analysis.

[0007] 2) Existing analysis software reported in the literature is designed for whole exome sequencing of RNA or whole genome sequencing of DNA. Although it can obtain possible fusion genes, it has a high false positive rate and cannot reflect the customization of targeted sequencing during the analysis process.

[0008] The high false positive rate in existing fusion gene prediction workflows significantly impacts the accuracy of results. Therefore, prediction methods that effectively reduce the number of false positives and improve prediction specificity are crucial. Candidate sequence alignment based on targeted genes can achieve specific screening of fusion genes. However, for targeted sequencing fusion gene screening with small sample sizes, a customizable fusion gene interpretation workflow is lacking. Analysis software for whole-genome DNA sequencing or whole-exome RNA sequencing cannot effectively mine data, resulting in wasted resources and a failure to obtain useful information. Summary of the Invention

[0009] The purpose of this invention is to provide a customized and automated method for interpreting fusion genes based on targeted sequencing, thereby solving the problems of high false positive rates and low prediction specificity in existing fusion gene prediction methods.

[0010] To solve the above problems, the present invention adopts the following technical solution:

[0011] A customized and automated method for interpreting fusion genes based on targeted sequencing is provided, comprising the following steps: 1) Targeted sequencing: Targeting the amplification of fusion sequences containing fusion regions, PCR technology is used to amplify and enrich the target sequences, and the constructed library is sequenced using a sequencing platform; 2) Preliminary detection: The obtained sequencing files are subjected to quality checks, and reads with poor quality or lengths that do not meet the requirements of the sequencing platform are initially filtered out. The results after quality checks are input into existing fusion gene analysis software, and preliminary fusion gene detection is performed using a set human genome reference set as the alignment sequence to obtain preliminary results; 3) Obtaining read sequences containing breakpoints: Using the preliminary results provided by existing fusion gene analysis software as a screening library, reads that meet the conditions are selected from the original sequencing files according to the reported fusion gene breakpoint sequences, and read sequences containing breakpoints are obtained; 4) Conditional filtering: The read sequences containing breakpoints are filtered according to the filtering conditions set by the filter. The filtering conditions include: a) gene information check, b) read information check, c) primer information check, d) secondary sequence alignment, and e) pseudogene and repetitive sequence check; if the conditions are met, it is judged as a reliable fusion; otherwise, it is judged as an unreliable fusion, and finally a customized and automated fusion gene interpretation based on targeted sequencing is realized.

[0012] A flowchart of a customized and automated fusion gene interpretation method based on targeted sequencing provided by the present invention is shown below. Figure 2 As shown.

[0013] It should be understood that, in step 1), the PCR technology includes, but is not limited to, conventional PCR, single-end PCR, nested PCR, and other PCR technologies.

[0014] It should be understood that, in step 2), existing fusion gene analysis software includes, but is not limited to, various short-read or long-read sequencing data mining software such as Arriba, STARFOsion, FusionCatcher, and FusionMap.

[0015] Step 4) also includes: adjusting parameters to achieve precise control of the interpretation of the fusion gene.

[0016] Preferably, the gene information check includes: conditional screening based on the fact that the gene of one of the fusion partners is in the list of target genes.

[0017] Preferably, the reading information check includes: setting a confidence threshold for the reading; if the reading exceeds the threshold, it passes; if it is less than the threshold, it is filtered out.

[0018] Taking a threshold of 5 as an example, if the maximum number of reads containing the fusion breakpoint sequence is greater than 5, the record can be checked through the read information and enter the next screening node; otherwise, the fusion credibility is considered low and it will not be recorded.

[0019] Preferably, the primer information check includes: determining the reliability of the fusion sequence to be determined by setting the amplification primer sequence; when the sequence with the primer sequence and the designed interval as the starting sequence has high reliability, only the sequence with high reliability is retained.

[0020] Preferably, the secondary sequence alignment includes: converting all transcript numbers to obtain a list of aligned gene names and alignment results; screening fusion partners other than the target gene; and filtering out fusion candidates if the sequence alignment results do not contain a fusion partner or the alignment position of the fusion partner sequence is inside the target gene sequence.

[0021] Preferably, the pseudogene and repetitive sequence inspection includes: constructing a pseudogene screening table; when the fusion partner is one of the genes in the pseudogene screening table, the fusion is filtered out to achieve pseudogene inspection; for cases where the same read length corresponds to different fusion partners after secondary sequence alignment, the evalue is used for screening, and records that meet the set evalue conditions are retained, such as less than 2e-2 or taking the minimum value, etc. If the evalues ​​are the same, they are all filtered out to achieve repetitive sequence inspection.

[0022] In step 4), some of the filtering parameters are flexibly adjusted through program input parameters, so as to achieve automated output of fusion gene interpretation results when the input sample and reference sequence are determined.

[0023] It should be understood that in step 1), targeted sequencing includes: designing primer sequences at the exon ends for amplification of fusion-type sequences containing fusion regions, and reserving a spacer sequence for amplification verification, which serves as the starting point for sequence amplification. Target sequences are enriched using techniques such as PCR amplification (polymerase chain reaction, increasing the copy number of nucleic acid molecules), single-end PCR, and nested PCR, and the constructed library is sequenced using a sequencing platform.

[0024] According to a preferred embodiment of the present invention, the entire process of conditional filtering is provided (e.g. Figure 3 As shown in the figure, its filtering process is as follows:

[0025] 1) Genetic testing. Based on the target gene list, the fusion genes reported by existing fusion gene analysis software are initially filtered to ensure that the gene of one of the fusion partners meets the criteria and is in the target gene list. In this automated process, exon-related splicing variations are not considered; therefore, fusions containing the same target gene will be filtered out at this step.

[0026] 2) Reading check. Further check the sequence readings containing fusion breakpoints; if they do not meet the threshold requirements, they are filtered out.

[0027] 3) Primer information check, also known as start sequence check. To ensure the effectiveness of primer amplification to the greatest extent, the start sequence must begin with the designed primer sequence and the reserved spacer sequence. If only the primer sequence is included without the spacer sequence, it may be non-specific amplification, and this reading will be filtered out.

[0028] 4) Secondary sequence alignment. Secondary sequence alignment is performed on fusion candidates that meet the above criteria to reduce false positives caused by gene structure connection errors during processing in existing fusion gene analysis software. In results based on gene alignment methods such as BLASTN, all transcript numbers are converted to obtain a list of aligned gene names and alignment results. This step filters fusion partners other than the target gene. If the sequence alignment results do not contain a fusion partner or the fusion partner sequence alignment position is inside the target gene sequence, the fusion candidate will be filtered out. From the list of candidates that meet the criteria, the alignment result with the lowest evalue is selected as the final reported evalue.

[0029] 5) Pseudogene and Repetitive Sequence Check. Existing fusion gene analysis software reports results where the fusion partner is a pseudogene or the same sequence corresponds to two different fusion partners. To address issue one, a pseudogene screening table is constructed; if the fusion partner is one of the genes in the pseudogene screening table, the fusion is filtered out. To address issue two, only fusions with lower evalues ​​are retained; if the evalues ​​are the same, neither is reported.

[0030] As described in the background section of this invention, existing fusion gene analysis software has a high false positive rate. Due to factors such as sequencing errors, high similarity of short sequences during processing, and splicing errors, fusion gene analysis software is prone to errors in detection, easily mis-splicing the same gene and reporting it as a novel fusion gene. Existing fusion gene analysis workflows lack customization for targeted sequencing. Because there is no input of targeted features during operation, they are all general analyses, using methods applicable to whole-genome sequencing to analyze targeted sequencing, making it impossible to customize the quality control of targeted sequencing. Furthermore, most false positives in the analysis reports are not within the scope of targeted sequencing, increasing the need for useless result screening.

[0031] However, this invention, based on targeted sequencing, incorporates customized primer screening. Compared to existing fusion gene analysis software, it provides thorough result filtering on top of existing reports, allowing control over fusion gene reporting thresholds and secondary fusion checks. The conditional filtering of this invention can be applied to high-throughput sample detection and automates the process. Therefore, secondary sequence alignment screening of targeted fusion genes can achieve highly specific integration of fusion candidates. This invention establishes a targeted filtering mechanism based on targeted sequencing to specifically screen fusion candidates from existing analysis software results, thus solving the problem of high false positive rates in existing analysis software.

[0032] According to this invention, all variable parameters involved in the entire conditional filtering process can be flexibly adjusted to suit targeted sequencing settings and accuracy requirements. Gene information for gene testing can be input via options during program execution, while primer sequences customized for different libraries are input as external files, from which relevant sequence information is read, increasing the operability of the automated process. Threshold screening based on reads can be set with appropriate values ​​according to the accuracy requirements of the results, thereby affecting the number of candidate reads processed for subsequent primer screening and secondary sequence alignment. By adjusting the relevant parameters, flexible and library-customized automatic interpretation of fusion genes can be achieved.

[0033] In summary, the present invention provides a customized and automated method for interpreting fusion genes based on targeted sequencing, which has the following significant advantages over existing technologies:

[0034] 1) This invention is based on targeted sequencing customization, which can accommodate more targeted gene information and adjust the filtering conditions for targeted sequencing libraries;

[0035] 2) This invention can adjust the accuracy of the interpretation results of fusion genes by controlling different input parameters, thereby achieving more customized interpretation;

[0036] 3) In its design, this invention takes into account the contradictory results of a single sequence pointing to two fusions and the meaningless results of pseudogene fusion in the interpretation results, and filters and screens them, thereby optimizing the fusion gene results;

[0037] 4) This invention can be applied to batch automated detection of multiple samples, improving the specificity of fusion gene interpretation;

[0038] 5) This invention can be applied to various sequence-based analyses, such as the analysis of phenomena like fusion genes, alternative splicing, and chromosomal structural variations. Attached Figure Description

[0039] Figure 1 A schematic diagram confirming the fusion gene is shown;

[0040] Figure 2 A schematic diagram of a customized and automated fusion gene interpretation process based on targeted sequencing provided by the present invention is shown.

[0041] Figure 3 A flowchart illustrating the conditional filter is shown;

[0042] Figure 4 The results of fusion gene detection in 5 samples using the method of the present invention are shown in comparison with the results of fusion gene detection using only existing fusion gene analysis software. Detailed Implementation

[0043] The present invention will be further described below with reference to specific embodiments. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0044] According to a preferred embodiment of the present invention, a customized and automated method for interpreting fusion genes based on targeted sequencing is provided, which includes the following steps:

[0045] 1) Targeted sequencing: The goal is to amplify fusion sequences with fusion partners. PCR, single-end PCR, nested PCR and other technologies are used to amplify and enrich the target sequences, and the constructed libraries are sequenced using a sequencing platform.

[0046] 2) Preliminary detection: The obtained sequencing files are subjected to quality detection to initially filter out readings with poor quality and lengths that do not meet the requirements of the sequencing platform. The results after quality detection are then input into the existing fusion gene analysis software. Using the set human genome reference set as the alignment sequence, preliminary fusion gene detection is performed to obtain the preliminary results of the existing analysis software.

[0047] 3) Obtaining readout sequences containing breakpoints: Based on the preliminary analysis results of the existing analysis software, the reported fusion readout sequences are extracted. The standard extraction length is set with the fixed number of bases before the gene fusion breakpoint, such as 5-50bp, to generate a fixed-length fusion breakpoint sequence containing gene 1 and gene 2.

[0048] 4) Conditional Filtering: The reading sequence containing breakpoints is filtered according to the filter conditions set by the filter. The filter conditions include the following sub-steps:

[0049] a) Genetic analysis: The two fused genes are examined. Based on the targeted sequencing library, fusion records where neither gene is listed in the library are deleted. A preliminary screening table is generated by combining the fusion breakpoint sequence and fusion gene information.

[0050] b) Reader Information Check: Based on the preliminary screening table of the existing analysis software results, reads containing fusion breakpoint sequences are selected from the raw sequencing files, and the individual read with the highest number of reads and its quantity are selected for analysis. A threshold is applied to the read count, with a fixed threshold selected as the filtering condition. This condition can be adjusted according to the analysis task. For example, if the threshold is set to 5, reads with a count higher than 5 will proceed to the next step; otherwise, they will be filtered out.

[0051] c) Primer Information Check: Based on targeted sequencing, the ideal amplification state is achieved with amplification primers and reads containing reserved spacer sequences, resulting in higher amplification reliability. Reads that meet the reading threshold in the previous step undergo starter sequence verification. Based on the fusion gene information, the specified target gene primer sequence check process is initiated. A complete match screening is performed based on the sequences in the input primer sequence file, recording the primer orientation as + / - (forward or reverse) and recording the primer number.

[0052] d) Secondary sequence alignment: Sequence alignment was performed using genome alignment methods such as BLASTN. For aligned transcript information, it was converted into gene names corresponding to the transcript numbers in the reference sequence set for easier and more intuitive gene name identification. If the results included gene 2 (excluding the target gene) from the preliminary screening table of existing fusion gene analysis software, its alignment position and evalue were recorded.

[0053] e) Pseudogene and Repetitive Sequence Check: For records in existing fusion gene analysis software results where the fusion partner is a pseudogene, all pseudogenes are extracted using the genome alignment annotation file. A pseudogene screening table is constructed to remove pseudogene fusions from the existing fusion gene analysis software results. For cases where the same read corresponds to different fusion partners after secondary sequence alignment, evalue is used for filtering. Records that meet the set evalue criteria are retained, while records with the same evalue are filtered out. The final results are output as Table 1.

[0054] Table 1. Results diagram a.gene1_gene2_brk: Represents the gene name and fusion breakpoint sequence of the fusion. b. direction: primer direction c.fp: Primer number d.gene1_pos: The position of the target gene in the fusion sequence e.gene2_pos: The position of the fusion partner in the fusion sequence. f.partner: Fusion partner gene name g.evalue: The e-value of the fusion partner alignment to the sequence. h.reads: Readings of the fusion sequence in the sequencing file. i.fastq: Fusion sequence in the sequencing file

[0055] The specific information on the fusion standards used in this embodiment is shown in Table 2 below. This embodiment includes 5 standard experiments with different sample sizes: 100ng, 50ng, 20ng, 10ng, and 5ng. RNA sequencing results for all samples were obtained using targeted sequencing, targeting the genes NTRK1, NTRK2, NTRK3, FGFR2, and FGFR3.

[0056] Table 2. Information on fusion genes of standard samples

[0057] Fusion gene interpretation was performed on all 5 samples using this workflow. The reference sequence was version hg38. The read count threshold was set to 5, and the evalue filter was set to the smallest evalue less than 2e-2. The fusion gene detection results of the 5 samples were compared with those obtained using only existing fusion gene analysis software. The results are as follows: Figure 4 As shown, for a given fusion, all targeted fusions were detected with a high sample size, while only one fusion was not detected when the sample size was reduced to 10 ng or below. No false positives were produced in the detection results across all sample sizes.

[0058] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention. Various variations can be made to the above embodiments of the present invention. All simple and equivalent changes and modifications made in accordance with the claims and description of this application fall within the protection scope of the claims of this patent. All aspects not described in detail in this invention are conventional technical content.

Claims

1. A customized and automated method for interpreting fusion genes based on targeted sequencing, characterized in that, Includes the following steps: 1) Targeted sequencing: With the goal of amplifying fusion-type sequences with fusion regions, primer sequences that can be used for amplification are designed at the exon ends, and a spacer sequence for amplification is reserved. This is used as the starting point for sequence amplification. PCR technology is used to amplify and enrich the target sequence, and a sequencing platform is used to sequence the constructed library. 2) Preliminary detection: The obtained sequencing files are subjected to quality detection to initially filter out reads with poor quality and lengths that do not meet the requirements of the sequencing platform. The results after quality detection are then input into the existing fusion gene analysis software. Using the set human genome reference set as the alignment sequence, preliminary fusion gene detection is performed to obtain preliminary results. 3) Obtaining readout sequences containing breakpoints: Using the preliminary results provided by the existing fusion gene analysis software as a screening library, select readouts that meet the criteria from the original sequencing files according to the reported fusion gene breakpoint sequences to obtain readout sequences containing breakpoints; 4) Conditional Filtering: The read sequences containing breakpoints are filtered according to the filter conditions set by the filter. The filter conditions include: a) Gene Information Check: The condition of filtering is that the gene of one of the fusion partners is in the target gene list; b) Reading Information Check: A confidence threshold is set for the readings. If the reading exceeds the threshold, it passes; if it is less than the threshold, it is filtered out; c) Primer Information Check: The reliability of the fusion sequence to be determined is judged by setting the amplification primer sequence. The sequence with high reliability is when the primer sequence and the designed interval are used as the starting sequence. Only the sequences with high reliability are retained; d) Secondary Sequence Alignment: All transcript numbers are transformed to obtain a list of aligned gene names and alignment results. For fusions other than the target gene... Partner screening: If the sequence alignment result does not contain a fusion partner or the fusion partner sequence alignment position is inside the target gene sequence, the fusion candidate will be filtered out; e) Pseudogene and repetitive sequence check: Construct a pseudogene screening table. When the fusion partner is one of the genes in the pseudogene screening table, the fusion is filtered out, realizing pseudogene detection; For cases where the same read length corresponds to different fusion partners after secondary sequence alignment, the evalue is used for screening. Records that meet the set evalue conditions are retained. If the evalues ​​are the same, they are all filtered out, realizing repetitive sequence detection; If the above conditions are met, it is judged as a reliable fusion; otherwise, it is judged as an unreliable fusion, ultimately realizing a customized and automated fusion gene interpretation based on targeted sequencing.

2. The fusion gene interpretation method according to claim 1, characterized in that, Step 4) also includes: adjusting parameters to achieve precise control of the interpretation of the fusion gene.

3. The fusion gene interpretation method according to claim 1, characterized in that, In step 4), some of the filtering parameters are flexibly adjusted through program input parameters, so as to achieve automated output of fusion gene interpretation results when the input sample and reference sequence are determined.

Citation Information

Patent Citations

  • Alignment free filtering for identifying fusions

    US11473137B2

  • Rapid and ultrahigh-sensitivity DNA fusion gene detection method

    CN113035273A

  • Method and device for detecting genetic mutation and expression

    WO2022089033A1