Fusion gene detection method and device based on second-generation sequencing, equipment, medium
Patent Information
- Application Number
- CN202410859793.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2044-06-28
AI Technical Summary
现有技术中尽管已经存在一些能够检测DNA水平融合基因的工具,但通常只能查看个别基因的基因融合,为了探索多个疾病的潜在风险,可能需要使用多个工具或软件
[0052]The beneficial effects of this invention are as follows: By sequencing the sample to be tested and comparing it with a genomic reference sequence, second-generation sequencing data of the sample is obtained. Multiple structural variations of the sample are detected, annotated, and used as candidate fusion gene variations, which are then merged to obtain a candidate fusion gene set. Event identification is performed on the candidate fusion gene variations in the candidate fusion gene set, and candidate fusion gene variations identified as belonging to the same event are annotated with the same event number. Then, the target variation frequency for each event is calculated. The candidate fusion gene set is then filtered through a list to obtain a first fusion gene set; and the first fusion gene set is further filtered through a list to obtain a second fusion gene set, which is then subjected to quality filtering. A third fusion gene set is obtained. Finally, based on the target mutation frequency of each event, the third fusion gene set is filtered according to fusion properties to obtain a fourth fusion gene set. The target fusion gene mutations in the first and fourth fusion gene sets are then annotated and merged to obtain the target fusion gene set. This allows for gene mutation frequency calculation at the event level, which can more accurately estimate the mutation frequency of fusion events. Furthermore, multi-level filtering based on list, quality, and/or fusion status can improve the accuracy of mutation detection, ensure accurate identification of target mutations from a large amount of data, effectively reduce false alarm rates, and thus improve the accuracy and reliability of detection results.
Smart Images

Figure CN118737278B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gene detection technology, specifically relating to a method, device, equipment, and storage medium for detecting fusion genes based on next-generation sequencing. Background Technology
[0002] Fusion genes are formed by the linking of partial sequences of two or more genes. They are often formed due to abnormal genomic structures such as chromosomal rearrangements or aberrant transcription, and play an important role in the malignant transformation and development of many human cancers. Fusion genes can be used as diagnostic and prognostic markers to confirm cancer diagnoses and to monitor patient responses to molecular therapies.
[0003] Early detection of fusion genes primarily relied on traditional tissue biopsies, including qRT-PCR, fluorescence in situ hybridization (FISH), immunohistochemistry (IHC), and chromosome banding analysis. These methods are widely used in routine clinical practice for detecting chromosomal rearrangement events.
[0004] With the advent of next-generation sequencing, high-throughput sequencing can more comprehensively identify various fusion genes, offering higher accuracy compared to traditional tissue biopsies and pinpointing specific sequences. Many bioinformatics tools have been developed for detecting fusion events, such as TopHat-Fusion, STAR-Fusion, and Arriba based on RNA sequencing data, and Delly and Manta based on DNA sequencing data.
[0005] Common methods for detecting disease-related DNA variants include panel sequencing, whole-exome sequencing (WES), and whole-genome sequencing (WGS). While some tools exist capable of detecting DNA-level gene fusions, they typically only examine individual gene fusions. To explore the potential risks of multiple diseases, multiple tools or software may be required. Furthermore, in addition to routine fusion gene variant detection, detecting fusion genes integrating immune cell Ig / TR family genes is also challenging. General tools cannot detect fusions of other genes that frequently fuse with Ig / TR family genes, which are also crucial for detecting diseases such as lymphoma.
[0006] Furthermore, due to the challenges in sequence homology and identifying accurate sequence breakpoints where structural variations occur, currently available tools generally suffer from inaccurate allele frequency (AF) detection. In addition, the pursuit of high sensitivity may also result in a high false positive rate.
[0007] First, traditional tissue biopsies are costly, have long testing cycles, and are prone to false positives. Furthermore, traditional methods are often ineffective in detecting variations in cancerous tissue caused by genotypes containing low concentrations of fusion genes. Moreover, existing detection technologies are demanding and difficult to standardize.
[0008] Secondly, typical fusion gene detection software can only detect a few fusions, resulting in a narrow detection range. More seriously, existing software suffers from low accuracy in calculating fusion affinity (AF), leading to low accuracy in test results and a high likelihood of false positives. Due to a certain degree of misjudgment and confusion, non-target fusion gene variants may be incorrectly identified as target variants, resulting in low specificity and thus reducing the reliability of the test results. Summary of the Invention
[0009] The purpose of this invention is to provide a method, device, equipment, and storage medium for detecting fusion genes based on next-generation sequencing, which can more accurately estimate the mutation frequency of fusion events, thereby improving the accuracy and reliability of the detection results.
[0010] The first aspect of this invention discloses a method for detecting fusion genes based on next-generation sequencing, comprising:
[0011] The sample to be tested is sequenced and compared with a genomic reference sequence to obtain the next-generation sequencing data of the sample to be tested;
[0012] The second-generation sequencing data were analyzed to obtain multiple structural variations in the sample to be tested;
[0013] Each structural variation of the sample to be tested is annotated and used as a candidate fusion gene variation, which is then merged to obtain a candidate fusion gene set.
[0014] Event identification is performed on the candidate fusion gene variants in the candidate fusion gene set, and candidate fusion gene variants identified as the same event are annotated with the same event number.
[0015] Calculate the target mutation frequency for each event;
[0016] The candidate fusion gene set is filtered according to the first preset list to obtain the first fusion gene set;
[0017] According to the second preset list and / or the third preset list, the first fusion gene set is filtered by list to obtain the second fusion gene set; the second fusion gene set is filtered by quality to obtain the third fusion gene set; wherein the first preset list, the second preset list and the third preset list are different;
[0018] Based on the target mutation frequency of each event, the third fusion gene set is filtered according to fusion properties to obtain the fourth fusion gene set;
[0019] The target fusion gene variants in the first fusion gene set and the fourth fusion gene set are annotated and merged to obtain the target fusion gene set.
[0020] In some embodiments, calculating the target mutation frequency for each event includes:
[0021] Determine whether each event is a complex event;
[0022] If the event is determined to be a non-complex event, calculate the first and second variation frequencies for each event when using different evidence;
[0023] Based on the magnitude of the first mutation frequency, the candidate mutation frequency for the event is selected from the first mutation frequency and the second mutation frequency.
[0024] The potential mutation frequency of this event is determined as the target mutation frequency.
[0025] In some embodiments, calculating the first and second variation frequencies for each event when using different evidence includes:
[0026] For each event, obtain SR and PR evidence for all candidate fusion gene variants, and then perform deduplication to obtain the total number of unique SR and unique PR evidence for each event.
[0027] Calculate the first and second numbers of non-variant reads in the first and second breakpoint neighborhoods of each candidate fusion gene variant in each event, which have a specified proportion of similarity to the reference genome sequence.
[0028] The first mutation frequency for each event is calculated based on the total number of unique SR evidences and the first and second quantities corresponding to each candidate fusion gene mutation.
[0029] The second mutation frequency for each event is calculated based on the total number of unique SR evidence, the total number of unique PR evidence, the first number and the second number corresponding to each candidate fusion gene mutation.
[0030] In some embodiments, the method further includes:
[0031] If the event is determined to be a complex event, the third variation frequency of the event is calculated based on the number of SR evidence for each candidate fusion gene variant, the number of PR evidence for each candidate fusion gene variant, the number of split reads output by each candidate fusion gene variant that have a specified similarity to the reference genome sequence, and the number of cross reads output by each candidate fusion gene variant that have a specified similarity to the reference genome sequence.
[0032] The third mutation frequency of this event is determined as the target mutation frequency.
[0033] In some embodiments, the first preset list is a protection list; filtering the candidate fusion gene set according to the first preset list to obtain the first fusion gene set includes:
[0034] The candidate fusion gene variants present in the protection list are retained in the candidate fusion gene set to obtain the first fusion gene set.
[0035] In some embodiments, the second preset list is a blacklist list, and the third preset list is a specified gene list; filtering the first fusion gene set according to the second preset list and / or the third preset list to obtain the second fusion gene set includes:
[0036] Remove fusion gene variants from the first fusion gene set that exist in the blacklist or have at least one breakpoint in a region of the blacklist, and / or retain at least one candidate fusion gene variant that has at least one breakpoint in a specified gene list to obtain the second fusion gene set.
[0037] In some embodiments, event identification is performed on candidate fusion gene variants in the candidate fusion gene set, including:
[0038] If the first and second breakpoints of a candidate fusion gene mutation are within a specified distance from the corresponding breakpoint of another candidate fusion gene mutation, and these two candidate fusion gene mutations belong to the same mutation type, then they are considered to be the same event.
[0039] A second aspect of this invention discloses a fusion gene detection device based on next-generation sequencing, comprising:
[0040] The alignment unit is used to sequence the sample to be tested and align it with the genome reference sequence to obtain the next-generation sequencing data of the sample to be tested.
[0041] The detection unit is used to detect the next-generation sequencing data and obtain multiple structural variations of the sample to be tested;
[0042] The first annotation unit is used to annotate each structural variation of the sample to be tested as a candidate fusion gene variation, and then merge them to obtain a candidate fusion gene set.
[0043] An event identification unit is used to identify the candidate fusion gene variants in the candidate fusion gene set as events, and to annotate the candidate fusion gene variants identified as the same event with the same event number.
[0044] A computational unit is used to calculate the target mutation frequency for each event;
[0045] The first filtering unit is used to filter the candidate fusion gene set according to a first preset list to obtain the first fusion gene set.
[0046] The second filtering unit is used to perform list filtering on the first fusion gene set according to the second preset list and / or the third preset list to obtain the second fusion gene set;
[0047] The third filtering unit is used to perform quality filtering on the second fusion gene set to obtain a third fusion gene set; wherein the first preset list, the second preset list and the third preset list are all different.
[0048] The fourth filtering unit is used to filter the third fusion gene set based on fusion properties according to the target mutation frequency of each event to obtain the fourth fusion gene set.
[0049] The second annotation unit is used to annotate the target fusion gene variants in the first fusion gene set and the fourth fusion gene set, and merge them to obtain the target fusion gene set.
[0050] A third aspect of the present invention discloses an electronic device, including a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the fusion gene detection method based on next-generation sequencing disclosed in the first aspect.
[0051] The fourth aspect of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program causes a computer to execute the fusion gene detection method based on next-generation sequencing disclosed in the first aspect.
[0052] The beneficial effects of this invention are as follows: By sequencing the sample to be tested and comparing it with a genomic reference sequence, second-generation sequencing data of the sample is obtained. Multiple structural variations of the sample are detected, annotated, and used as candidate fusion gene variations, which are then merged to obtain a candidate fusion gene set. Event identification is performed on the candidate fusion gene variations in the candidate fusion gene set, and candidate fusion gene variations identified as belonging to the same event are annotated with the same event number. Then, the target variation frequency for each event is calculated. The candidate fusion gene set is then filtered through a list to obtain a first fusion gene set; and the first fusion gene set is further filtered through a list to obtain a second fusion gene set, which is then subjected to quality filtering. A third fusion gene set is obtained. Finally, based on the target mutation frequency of each event, the third fusion gene set is filtered according to fusion properties to obtain a fourth fusion gene set. The target fusion gene mutations in the first and fourth fusion gene sets are then annotated and merged to obtain the target fusion gene set. This allows for gene mutation frequency calculation at the event level, which can more accurately estimate the mutation frequency of fusion events. Furthermore, multi-level filtering based on list, quality, and / or fusion status can improve the accuracy of mutation detection, ensure accurate identification of target mutations from a large amount of data, effectively reduce false alarm rates, and thus improve the accuracy and reliability of detection results. Attached Figure Description
[0053] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.
[0054] Unless otherwise specified or defined, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.
[0055] Figure 1 This is a flowchart of a fusion gene detection method based on next-generation sequencing disclosed in an embodiment of the present invention;
[0056] Figure 2 This is a schematic diagram of the structure of a fusion gene detection device based on next-generation sequencing disclosed in an embodiment of the present invention;
[0057] Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention.
[0058] Explanation of reference numerals in the attached figures:
[0059] 201. Comparison unit; 202. Detection unit; 203. First annotation unit; 204. Event identification unit; 205. Calculation unit; 206. First filtering unit; 207. Second filtering unit; 208. Third filtering unit; 209. Fourth filtering unit; 210. Second annotation unit; 301. Memory; 302. Processor. Detailed Implementation
[0060] Unless otherwise specified or defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. When combined with the technical solutions of the invention in a real-world scenario, all technical and scientific terms used herein may also have meanings corresponding to the purpose of achieving the technical solutions of the invention. The terms "first," "second," etc., used herein are merely for distinguishing names and do not represent a specific number or order. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0061] It should be noted that when a component is considered "fixed" to another component, it can be directly fixed to the other component or there can be an intervening component; when a component is considered "connected" to another component, it can be directly connected to the other component or there can be an intervening component; when a component is considered "mounted" on another component, it can be directly mounted on the other component or there can be an intervening component; when a component is considered "placed" on another component, it can be directly placed on the other component or there can be an intervening component.
[0062] Unless otherwise specified or defined, the terms "described" or "the" as used herein refer to the technical features or technical content mentioned or described prior to the relevant section, which may be the same as or similar to the technical features or technical content mentioned herein. Furthermore, the terms "comprising" and "having," and any variations thereof, as used herein, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0063] This invention discloses a method for detecting fusion genes based on next-generation sequencing, which can be implemented through computer programming. The execution subject of this method can be an electronic device such as a computer, laptop, or tablet, or a next-generation sequencing-based fusion gene detection device embedded in an electronic device; this invention does not limit this. To facilitate understanding of this invention, specific embodiments will be described in more detail below with reference to the accompanying drawings.
[0064] like Figure 1 As shown, the method includes the following steps 110-180:
[0065] 110. Sequencing the sample to be tested and comparing it with the genome reference sequence to obtain the next-generation sequencing data of the sample to be tested.
[0066] The sequencing method can be either paired-end sequencing or single-end sequencing. Second-generation sequencing data includes, but is not limited to, sequence files such as BAM, FASTQ, SAM, or VCF. In this embodiment of the invention, BAM sequence files are used for detection.
[0067] 120. Detect the second-generation sequencing data to obtain multiple structural variations in the sample to be tested.
[0068] Structural variations (SVs) include deletions, insertions, duplications, inversions, and translocations. Structural variations are formed by the breaking and reconnection of sequences; the locations of these breaks and reconnections are called breakpoints. In the aligned BAM file, these breakpoints are observed on the two sequences respectively. Therefore, breakpoints are both the result and a marker of structural variations. Two sequences break and connect to form a fusion variation. By identifying the breakpoints on the two sequences and recording the positions of this pair of breakpoints, the fusion variation is identified.
[0069] Specifically, to ensure both specificity and sensitivity, at least two modes are used in parallel to detect second-generation sequencing data. The structural variant detection results from these different modes are then combined to obtain multiple structural variants in the sample. By using at least two different structural variant detection modes, comprehensive and efficient detection of variants at different levels and confidence levels can be ensured. These modes are designed to consider the diversity and complexity of variants, aiming to improve the accuracy and reliability of detection.
[0070] The next-generation sequencing data was analyzed in parallel using at least two modes, typically a master mode and a high-confidence mode. An additional Ig / TR (IGTR) mode was also included for the detection of fusions in the Ig / TR family of genes. Different read processing conditions were applied to the short reads obtained from the next-generation sequencing data for each mode, allowing for the detection of variations at different levels and confidence levels, thus improving detection accuracy.
[0071] Main mode: Only reads with a mapping quality (MAPQ) score higher than 30 are used. There must be at least two split reads (SR reads) supporting the variant and at least three spanning reads (PR reads) supporting the variant. The call region is the entire genome. MAPQ represents the alignment of each read; a higher MAPQ indicates better alignment quality. MAPQ values range from 0 to 60, reflecting the confidence level of a read's alignment to its corresponding position in the reference genome sequence. MAPQ = 60 is considered very reliable, while lower MAPQ values indicate less reliable alignment. SR reads refer to two or more read fragments formed after a single read has broken; PR reads refer to a pair of reads that spans two breakpoints of a structural variant, aligning completely to the reference genome on both sides of the variant.
[0072] High-confidence mode: Only reads with a MAPQ score higher than 30 are used, with at least 20 SR reads and at least 5 PR reads supporting the mutation. The mutation call region is located throughout the entire genome.
[0073] IGTR mode: Use all reads with a MAPQ greater than 0, at least 2 SR reads, and at least 3 PR reads supporting the mutation. Exclude complex regions containing more than 30 other connections. The mutation call region is limited to Ig / TR family genes and partner genes that frequently fuse with these family genes. Due to the wide range of influence of immune cell gene fusions, to maintain the same or higher accuracy as standard experimental results, the following regions containing partner genes are defined based on the FISH probe coverage area (based on the hg19 reference genome):
[0074] MYC and its upstream and downstream regions: chr8:127742218-130573118.
[0075] BCL2 and its upstream and downstream regions: chr18:60375125-61425158.
[0076] CCND1 and its upstream and downstream regions: chr11:69185886-69717397.
[0077] CCND3 and its upstream and downstream regions: chr6:41475652-42432766.
[0078] MALT1 and its upstream and downstream regions: chr18:55708094-56945477.
[0079] BCL6 and its upstream and downstream regions: chr3:186771265-188259733.
[0080] CCND2 and its upstream and downstream regions: chr12:3750019-5244466.
[0081] NSD2 and its upstream and downstream regions: chr4:1465336-2389637.
[0082] BIRC3 and its upstream and downstream regions: chr11:101545997-102851609.
[0083] 130. After annotating the structural variations of the test samples, they are used as candidate fusion gene variations and merged to obtain a candidate fusion gene set.
[0084] In this step, the gene containing the two breakpoints of the structural variant or its neighboring genes, the distance between the two breakpoints and their respective gene or neighboring genes, the transcript, and the intron or exon number are annotated based on the transcript database. Additionally, for structural variants that are not insertion types, the fusion gene direction is annotated based on gene orientation, variant type, and read connection direction. The annotation rules are shown in Table 1 below:
[0085] Table 1. Annotation rules for candidate fusion gene variants.
[0086]
[0087] Translocation type 1 refers to: the downstream segment of breakpoint 1 connects to the upstream segment of breakpoint 2;
[0088] Translocation type 2 refers to: the upstream segment of breakpoint 1 connects to the downstream segment of breakpoint 2;
[0089] Translocation type 3 refers to: the downstream segment of breakpoint 1 connects to the downstream segment of breakpoint 2;
[0090] Translocation type 4 refers to: the upstream segment of breakpoint 1 is connected to the upstream segment of breakpoint 2.
[0091] 140. Perform event identification on candidate fusion gene variants in the candidate fusion gene set, and annotate candidate fusion gene variants identified as the same event with the same event number.
[0092] In a test sample, a candidate fusion gene variant contains two breakpoints: a first breakpoint (breakpoint 1) and a second breakpoint (breakpoint 2). By comparing two candidate fusion gene variants, if the breakpoints 1 and 2 of one candidate fusion gene variant are within a specified length (e.g., 50 bp) of the corresponding breakpoint of the other candidate fusion gene variant, and these two candidate fusion gene variants belong to the same variant type, they are identified as the same event, and the candidate fusion gene variants identified as the same event are annotated with the same event ID.
[0093] 150. Calculate the target mutation frequency for each event.
[0094] To ensure that the detected mutation frequencies of fusion genes more closely reflect reality, the target mutation frequency AF is calculated for each event using the following steps S11–S15:
[0095] S11. Determine whether each event is a complex event. If not, proceed to steps S12-S14; if yes, proceed to step S15.
[0096] An event is considered a complex event when the breakpoints involved in that event contain multiple events, and those events require individual handling:
[0097]
[0098] in The coverage of the region where the breakpoint 1 of the i-th candidate fusion gene mutation is located (e.g., 40 bp before and after breakpoint 1). denoted as the coverage of the region where the breakpoint 2 of the i-th candidate fusion gene mutation is located (e.g., 40 bp before and after breakpoint 2), and n is the total number of candidate fusion gene mutations in the event. Represents the number of SR evidences for the i-th candidate fusion gene variant, Ref 1i Ref represents the first number of non-variant reads in the neighborhood of the first breakpoint in the i-th candidate fusion gene variant that have a specified proportion of similarity to the reference genome sequence. 2i This represents the second number of non-variant reads in the neighborhood of the second breakpoint in the i-th candidate fusion gene variant that have a specified similarity to the reference genome sequence.
[0099] S12. Calculate the first and second variation frequencies for each event when using different evidence.
[0100] Specifically, step S12 may include the following steps S1201 to S1204 (not shown):
[0101] S1201. Obtain SR evidence and PR evidence for all candidate fusion gene variants corresponding to each event, and perform deduplication to obtain the total number of unique SR evidence and the total number of unique PR evidence for each event.
[0102] This involves obtaining all evidence supporting candidate fusion gene variants in each event, including SR evidence and PR evidence. The requirement of "unique" is because each event may contain multiple similar candidate fusion gene variants, and there may be duplicate evidence among these multiple similar candidate fusion gene variants. Therefore, "unique" means removing duplicate evidence to ensure the correctness of the calculation results.
[0103] Total number of unique SR evidence Total number of unique PR evidence These refer to the total number of SR and PR evidences obtained after deduplication of all n candidate fusion gene variants in the event, respectively. This represents the number of SR evidences for the i-th candidate fusion gene variant. This represents the number of PR evidences for the i-th candidate fusion gene variant.
[0104] S1202. Calculate the first and second number of non-variant reads in the first breakpoint neighborhood and the second breakpoint neighborhood of each candidate fusion gene variant in each event, respectively, which have a specified similarity to the reference genome sequence.
[0105] The size of the neighborhood can be specified, such as 40bp or 50bp. The first breakpoint neighborhood refers to the 40bp / 50bp region near breakpoint 1, and the second breakpoint neighborhood refers to the 40bp / 50bp region near breakpoint 2. The specified percentage can be set to 100% or a value close to 100%, such as 96%, 98%, or 99%. When the neighborhood length is short, such as only 40bp, a stricter threshold of 100% can be specified, which has shown good performance in testing.
[0106] In this invention, fusion gene mutations can be understood as connections between the current position and another position. The first breakpoint (breakpoint 1) is the breakpoint at the current position, and the second breakpoint (breakpoint 2) is the breakpoint at another position. The total number of non-mutant reads includes the first number of non-mutant reads (Ref) whose neighborhood of the first breakpoint has a specified similarity to the reference genome sequence. 1i And the second number of non-variant reads whose similarity to the reference genome sequence in the neighborhood of the second breakpoint reaches a specified proportion. 2i sum.
[0107] S1203. Calculate the first mutation frequency for each event based on the total number of unique SR evidences and the first and second quantities corresponding to each candidate fusion gene mutation.
[0108] S1204. Calculate the second mutation frequency for each event based on the total number of unique SR evidence, the total number of unique PR evidence, the first number and the second number corresponding to each candidate fusion gene mutation.
[0109] Among them, based on the total number of unique SR evidence The first number of Refs corresponding to each candidate fusion gene variant 1i Second quantity Ref 2i Calculate the first mutation frequency AF1 for each event:
[0110]
[0111] And, based on the total number of unique SR evidence Total number of unique PR evidence The first number of Refs corresponding to each candidate fusion gene variant 1i Second quantity Ref 2i Calculate the second mutation frequency AF2 for each event:
[0112]
[0113] In the formula, n is the total number of candidate fusion gene variants in each event.
[0114] S13. Based on the magnitude of the first mutation frequency, select the event's candidate mutation frequency from the first mutation frequency and the second mutation frequency.
[0115] Specifically, if the first mutation frequency is greater than the specified frequency, then the first mutation frequency is determined as the candidate mutation frequency; if the first mutation frequency is less than the specified frequency, then the second mutation frequency is determined as the candidate mutation frequency. This is illustrated in equation (4) below:
[0116]
[0117] AF pipe This indicates the frequency of variation to be used, with the specified frequency set to 0.1.
[0118] S14. If the event is determined to be a non-complex event, the potential mutation frequency of the event is determined as the target mutation frequency.
[0119] S15. If the event is determined to be a complex event, calculate the third variation frequency of the event and determine the third variation frequency of the event as the target variation frequency.
[0120] Specifically, the method for calculating the third variation frequency of this event is as follows: the third variation frequency of this event is calculated based on the number of SR evidence for each candidate fusion gene variation, the number of PR evidence for each candidate fusion gene variation, the number of split reads output by each candidate fusion gene variation that have a specified similarity to the reference genome sequence, and the number of cross reads output by each candidate fusion gene variation that have a specified similarity to the reference genome sequence.
[0121] In this embodiment of the invention, the target mutation frequency of each event is determined based on the judgment result of whether the event is a complex event. Specifically, the target mutation frequency AF is determined by the following formula (5):
[0122]
[0123] Where n is the total number of candidate fusion gene variants in each event, complex is the Boolean value that satisfies equation (1), complex = True indicates that the event satisfies the condition shown in equation (1), complex = False indicates that the event does not satisfy the condition shown in equation (1); This refers to the number of split reads (SRs) output in structural variation detection from candidate fusion gene variant i that achieve a specified proportion of similarity to the reference genome sequence. This is the sum of the number of cross-reads (PRs) output by candidate fusion gene variant i in structural variant detection, whose similarity to the reference genome sequence reaches a specified proportion. The number of SR evidences for candidate fusion gene variant i. The number of PR evidence for candidate fusion gene variant i.
[0124] 160. Filter the candidate fusion gene set according to the first preset list to obtain the first fusion gene set.
[0125] To ensure the specificity of the results, the candidate fusion gene set is filtered by a list: specifically, the list filtering excludes candidate fusion gene variants that are concentrated in the blacklist, not in the protection list, or not in the specified gene list. Specifically, assuming the first preset list is the protection list, step 160 may include the following step (a) not shown in the figure:
[0126] (a) Retain the candidate fusion gene variants that exist in the protection list in the candidate fusion gene set to obtain the first fusion gene set.
[0127] The protection list includes the following fusion genes: EML4::ALK, ALK::EML4, ETV6::NTRK3, NTRK3::ETV6, EWSR1::FLI1, FLI1::EWSR1, KIAA1549::BRAF, BRAF::KIAA1549, TMPRSS2::ERG, ERG::TMPRSS2, BCR::ABL1, ABL1::BCR, KMT2A::AFF1, AFF1::KMT2A, ETV6::RUNX1, RUNX1::ETV6.
[0128] 170. Based on the second preset list and / or the third preset list, the first fusion gene set is filtered by list to obtain the second fusion gene set; the second fusion gene set is filtered by quality to obtain the third fusion gene set; wherein the first preset list, the second preset list and the third preset list are different.
[0129] Specifically, assuming the second preset list is a blacklist and the third preset list is a specified gene list, the method for obtaining the second fusion gene set in step 170 may include the following steps (b) and / or (c) not shown:
[0130] (b) Remove fusion gene variants that exist in the blacklist or have at least one breakpoint in a region of the blacklist from the first fusion gene set.
[0131] The blacklist includes the following fusion genes or regions: HLA-C::HLA-B, HLA-B::HLA-C, HLA-DRB1::HLA-DRB5, HLA-DRB5::HLA-DRB1, HLA-C::HLA-A, HLA-A::HLA-C, HLA-DQA1::HLA-DQA2, HLA-DQA2::HLA-DQA1, HLA-DQB2::HLA-DQB1, HLA-DQB1::HLA-DQB2, chr2:33140864-33141692 (hg19 / hs37d5 reference genome).
[0132] (c) Retain at least one breakpoint in a candidate fusion gene variant that exists in a specified gene list.
[0133] Different gene lists are used for different purposes. The following gene lists are set for conventional genes and Ig / TR family genes respectively:
[0134] i. When detecting rearrangements of routine genes, genes that frequently undergo fusions are included: ALK, BRAF, CD74, CLDN18, EGFR, ETV4, RSPO2, TFE3, ETV5, ETV6, EWSR1, EZR, FGFR1, FGFR2, TPM3, TMPRSS2, FGFR3, FOXO1, FUS, KIAA1549, MET, MYB, SDC4, RET, NR4A3, NRG1, NTRK1, NTRK2, NTRK3, NUTM1, USP6, ROS1, PAX3, PAX7, PAX8, PDGFB, PDGFRA, PDGFRB, SLC34A2, SS18, PLAG1, QKI, RAF1, RELA.
[0135] ii. When detecting rearrangements of Ig / TR genes in the immune system, the following genes are included: Ig / TR family genes and partner genes that often fuse with this family of genes: IG, TR, BCL2, BCL6, MYC, CCND1, CCND3, NSD2, MALT1, BIRC3, CCND2.
[0136] In step 170, the second fusion gene set is further subjected to quality filtering, which is mainly used to exclude candidate fusion gene variants with low confidence. Specifically, candidate fusion gene variants in the second fusion gene set that satisfy the following equations (6)-(9) are retained as the third fusion gene set:
[0137] Fusion length ≥1000bp (6)
[0138] Support MAPQ30 ≥3 (7)
[0139] AF ≥ 0.5% (8)
[0140] AF pop <1% (9)
[0141] Among them, Fusion length The length of the structural variation is given. For the variation type of translocation, the default condition (6) is satisfied. The translocation type includes mutual translocation, Robertson translocation, or insertion translocation, etc.
[0142] Support MAPQ30 The number of reads supported by MAPQ greater than or equal to 30.
[0143] AF represents the target mutation frequency of the event corresponding to the candidate fusion gene mutation. It should be noted that a candidate fusion gene mutation will only correspond to one event, while an event may have multiple mutations. Multiple mutations in the same event will collectively correspond to the AF of that event.
[0144] AF pop AF represents the mutation frequency of a candidate fusion gene variant in a population, i.e., the probability of a variant occurring in a population. If a variant occurs in 30 out of 1000 people, then AF is considered a high probability. pop That is, 3%. AF pop Higher AF values indicate a higher likelihood that the variant is frequently occurring in the population, is not a target variant, or is a false positive due to genomic or sequencing reasons. pop The lower the value, the more credible the variant is considered.
[0145] 180. Based on the target mutation frequency of each event, filter the third fusion gene set according to fusion properties to obtain the fourth fusion gene set.
[0146] In step 180, the filtered third fusion gene set is filtered based on fusion properties, primarily to filter candidate fusion gene variants that do not fall within the fusion gene category. Specifically, this may include the following filtering steps (d) to (g):
[0147] (d) When detecting rearrangements in conventional genes, retain candidate fusion gene variants in the third fusion gene set whose two breakpoints are both less than 1000 bp apart; or,
[0148] When detecting rearrangements of Ig / TR genes in the immune system, candidate fusion gene variants in the third fusion gene set whose breakpoint distance is less than 1000 bp are retained.
[0149] (e) Remove candidate fusion gene variants of the insertion type from the third fusion gene set.
[0150] (f) When detecting rearrangements of conventional genes, only candidate fusion gene variants from the normal fusion of upstream and downstream genes in the third fusion gene set are retained; or,
[0151] When detecting rearrangements of Ig / TR genes in the immune system, candidate fusion gene variants with arbitrary fusion directions are retained in the third fusion gene set.
[0152] (g) Remove candidate fusion gene variants in the third fusion gene set where the two breakpoints are the same gene.
[0153] 190. Annotate the target fusion gene variants in the first and fourth fusion gene sets and merge them to obtain the target fusion gene set.
[0154] Specifically, based on a pre-defined fusion variant database, the target fusion gene variants in the first fusion gene set filtered in step 160 and the fourth fusion gene set filtered in step 180 are annotated and merged to obtain the final detection result, i.e., the target fusion gene set. The pre-defined fusion variant database includes databases such as COSMIC, FusionGDB2.0, and Mitelman. This step mainly involves annotating relevant information from the fusion variant database.
[0155] By employing a three-layer filtering system based on list size, quality, and fusion, the system ensures accurate identification of target variants from large datasets. This three-layer filtering system not only improves the accuracy of variant detection but also effectively reduces the false alarm rate.
[0156] Example 1
[0157] In this case, 12 samples were selected for fusion gene detection. These samples were sequenced in 700 or 88 target regions and were experimentally verified to carry fusion gene variants such as EML4::ALK.
[0158] Mode selection: Based on the target of routine gene rearrangement detection, two modes (main mode and high confidence mode) are used for variant detection.
[0159] Sample detection: Structural variation detection is performed on the obtained genome alignment data based on the two selected modes.
[0160] (1.1) The detected structural variations are annotated as candidate fusion gene sets as follows: based on the transcript database, the gene or neighboring gene where the two breakpoints of the structural variation are located, the distance between the two breakpoints and the gene or neighboring gene where they are located, the transcript and the intron or exon number where they are located are annotated; in addition, for structural variations of non-insertion type, the fusion gene direction is annotated based on gene direction, variation type and read connection direction.
[0161] (1.2) The obtained candidate fusion gene set is identified as an event, and the candidate fusion gene variants identified as the same event are annotated with the same event ID.
[0162] (1.3) Calculate the target mutation frequency at the event level for the candidate fusion gene set.
[0163] (1.4) Filter the candidate fusion gene set according to the protection list to obtain the first fusion gene set, and filter the first fusion gene set according to the blacklist list and the specified gene list to obtain the second fusion gene set.
[0164] (1.5) Perform quality filtering on the second fusion gene set to obtain the third fusion gene set.
[0165] (1.6) Filter the third fusion gene set according to the target mutation frequency of each event to obtain the fourth fusion gene set.
[0166] (1.7) Annotate the target fusion gene variants in the first and fourth fusion gene sets and merge them to obtain the final fusion gene variant detection results.
[0167] The final fusion gene variant detection results for each sample are summarized in Table 2 below. The results show that corresponding fusion events were detected in all 12 samples, and the fusion directions were consistent. Furthermore, the low number of detected variants reduces the workload of downstream manual review and improves clinical efficiency.
[0168] Table 2 Summary of Detection Results in Example 1
[0169]
[0170]
[0171] Example 2
[0172] In this case, eight samples that were WES sequenced and experimentally verified to carry Ig-related fusion genes or isolated rearrangement variants were selected for fusion gene detection.
[0173] Mode selection: Based on the target of the detection of Ig / TR gene rearrangements in the immune system, three modes (main mode, high confidence mode, and Ig / TR mode (IGTR)) are used for variant detection.
[0174] Sample detection: Structural variation detection is performed on the obtained genome alignment data based on the three selected modes.
[0175] (2.1) The detected structural variations are annotated as a candidate fusion gene set.
[0176] (2.2) The obtained candidate fusion gene set is identified as an event, and the candidate fusion gene variants identified as the same event are annotated with the same event ID.
[0177] (2.3) Calculate the target mutation frequency at the event level for the candidate fusion gene set.
[0178] (2.4) The candidate fusion gene set is filtered according to the protection list to obtain the first fusion gene set, and the first fusion gene set is filtered according to the blacklist list to obtain the second fusion gene set.
[0179] (2.5) Perform quality filtering on the second fusion gene set to obtain the third fusion gene set.
[0180] (2.6) Based on the target mutation frequency of each event, filter the third fusion gene set to obtain the fourth fusion gene set.
[0181] (2.7) Annotate the target fusion gene variants in the first and fourth fusion gene sets and merge them to obtain the final fusion gene variant detection results.
[0182] The final fusion gene mutation detection results for each sample are summarized in Table 3 below. The detection results show that the corresponding fusion events were detected in all 8 samples.
[0183] IgCaller is a widely used software in Ig gene recombination and oncogenic translocation analysis. In this embodiment, IgCaller's default parameters were used to test each sample. It was found that BCL6 segregation rearrangement and MYC segregation rearrangement variants existed in test sample 2-2, and BCL6 segregation rearrangement existed in sample 2-3. IgCaller failed to detect these variants, while this method accurately detected them.
[0184] Table 3 Summary of Detection Results in Example 2
[0185]
[0186] Example 3
[0187] This example selected seven samples for testing. These samples had been sequenced whole-genome and experimentally verified to carry germline structural variations. It should be noted that fusion genes are also a type of structural variation. In this embodiment, only a single mode was used for testing, and the specified gene list filtering, fusion filtering, and annotations related to fusion genes were disabled to preserve structural variations in non-gene regions.
[0188] Mode selection: Based on the target of routine gene rearrangement detection, the main mode is used for variant detection.
[0189] Sample detection: Structural variation detection is performed on the obtained genome alignment data based on a selected mode.
[0190] (3.1) The detected structural variations are annotated as a candidate fusion gene set.
[0191] (3.2) Perform event identification on the obtained candidate fusion gene set, and annotate the candidate fusion gene variants identified as the same event with the same event ID.
[0192] (3.3) Calculate the target mutation frequency at the event level for the candidate fusion gene set.
[0193] (3.4) Filter the candidate fusion gene set according to the protection list to obtain the first fusion gene set, and filter the first fusion gene set according to the specified gene list to obtain the second fusion gene set.
[0194] (3.5) Perform quality filtering on the second fusion gene set to obtain the third fusion gene set.
[0195] (3.6) Filter the third fusion gene set according to the target mutation frequency of each event to obtain the fourth fusion gene set.
[0196] (3.7) Annotate the target fusion gene variants in the first and fourth fusion gene sets and merge them to obtain the final fusion gene variant detection results.
[0197] The final results of fusion gene variant detection for each sample are summarized in Table 4 below. The results show that corresponding structural variants were detected in all 7 samples. Furthermore, the detected variant frequency values contained only one frequency value representing the variant event, and were closer to the theoretical values of 0.5 or 1 for germline variants. Manta is a method for detecting structural variants from next-generation sequencing, and it offers competitive speed and sensitivity. In Manta's detection results, a single variant time point was detected with multiple variant records and corresponding to multiple variant frequency values, which were further away from the theoretical values of 0.5 or 1.
[0198] Table 4 Summary of Detection Results in Example 3
[0199]
[0200]
[0201] In summary, the next-generation sequencing-based fusion gene detection method provided by this invention, based on three structural variant detection modes and event-level algebraic error (AF) calculation, can detect fusion gene variants at different levels and confidence levels, and has an AF close to that of real-world scenarios. Furthermore, the three-layer filtering module results in a low number of fusion variants detected and high specificity. This invention is applicable to high-sensitivity, high-accuracy, and high-specificity detection of conventional fusion genes and fusion genes targeting Ig / TR family genes in immune cells.
[0202] like Figure 2As shown, this embodiment of the invention discloses a fusion gene detection device based on next-generation sequencing, including an alignment unit 201, a detection unit 202, a first annotation unit 203, an event identification unit 204, a calculation unit 205, a first filtering unit 206, a second filtering unit 207, a third filtering unit 208, a fourth filtering unit 209, and a second annotation unit 210, wherein...
[0203] The alignment unit 201 is used to sequence the sample to be tested and align it with the genome reference sequence to obtain the second-generation sequencing data of the sample to be tested.
[0204] The detection unit 202 is used to detect second-generation sequencing data and obtain multiple structural variations of the sample to be tested;
[0205] The first annotation unit 203 is used to annotate each structural variation of the sample to be tested as a candidate fusion gene variation, and then merge them to obtain a candidate fusion gene set.
[0206] Event identification unit 204 is used to identify event variations of candidate fusion genes in the candidate fusion gene set, and to annotate candidate fusion gene variations identified as the same event with the same event number.
[0207] Calculation unit 205 is used to calculate the target variation frequency for each event;
[0208] The first filtering unit 206 is used to filter the candidate fusion gene set according to the first preset list to obtain the first fusion gene set.
[0209] The second filtering unit 207 is used to perform list filtering on the first fusion gene set according to the second preset list and / or the third preset list to obtain the second fusion gene set;
[0210] The third filtering unit 208 is used to perform quality filtering on the second fusion gene set to obtain the third fusion gene set; wherein the first preset list, the second preset list and the third preset list are different.
[0211] The fourth filtering unit 209 is used to filter the third fusion gene set based on fusion properties according to the target mutation frequency of each event to obtain the fourth fusion gene set.
[0212] The second annotation unit 210 is used to annotate the target fusion gene variants in the first fusion gene set and the fourth fusion gene set, and merge them to obtain the target fusion gene set.
[0213] As an optional implementation, the computing unit 205 specifically includes the following sub-units (not shown):
[0214] The judgment sub-unit is used to determine whether each event is a complex event;
[0215] The first calculation subunit is used to calculate the first and second variation frequencies of each event when different evidence is used, when the event is determined to be a non-complex event.
[0216] The selection subunit is used to select the candidate mutation frequency for the event from the first mutation frequency and the second mutation frequency based on the magnitude of the first mutation frequency.
[0217] The first determining subunit is used to determine the candidate mutation frequency of the event as the target mutation frequency.
[0218] Further optionally, the computing unit 205 may also include the following sub-units (not shown):
[0219] The second calculation subunit is used to calculate the third variation frequency of an event when it is determined to be a complex event.
[0220] The second determining subunit is used to determine the third variation frequency of the event as the target variation frequency.
[0221] Furthermore, the aforementioned first calculation subunit is specifically used to acquire SR evidence and PR evidence for all candidate fusion gene variants corresponding to each event, and to perform deduplication to obtain the total number of unique SR evidence and the total number of unique PR evidence for each event; and to calculate the first and second numbers of non-variant reads in the first breakpoint neighborhood and the second breakpoint neighborhood of each candidate fusion gene variant in each event that have a similarity to the reference genome sequence reaching a specified proportion; and to calculate the first mutation frequency for each event based on the total number of unique SR evidence, the first number and the second number corresponding to each candidate fusion gene variant; and to calculate the second mutation frequency for each event based on the total number of unique SR evidence, the total number of unique PR evidence, the first number and the second number corresponding to each candidate fusion gene variant.
[0222] Furthermore, the aforementioned second calculation subunit is specifically used to calculate the third mutation frequency of the event based on the number of SR evidence for each candidate fusion gene variant, the number of PR evidence for each candidate fusion gene variant, the number of split reads output by each candidate fusion gene variant that have a specified similarity to the reference genome sequence, and the number of cross reads output by each candidate fusion gene variant that have a specified similarity to the reference genome sequence.
[0223] Furthermore, the first preset list is a protection list; the first filtering unit 206 is specifically used to retain the candidate fusion gene variants that exist in the protection list in the candidate fusion gene set, so as to obtain the first fusion gene set.
[0224] Furthermore, the second preset list is a blacklist list, and the third preset list is a specified gene list; the second filtering unit 207 is specifically used to remove fusion gene variants that exist in the blacklist list or fusion gene variants with at least one breakpoint in the region of the blacklist list from the first fusion gene set, and / or retain at least one candidate fusion gene variant with a breakpoint in the specified gene list to obtain the second fusion gene set.
[0225] Furthermore, the aforementioned event identification unit 204 is specifically used to identify two candidate fusion gene variants as the same event if the first and second breakpoints of one candidate fusion gene variant are within a specified distance from the corresponding breakpoint of another candidate fusion gene variant, and the two candidate fusion gene variants belong to the same variant type. The candidate fusion gene variants identified as the same event are annotated with the same event number.
[0226] like Figure 3 As shown, an embodiment of the present invention discloses an electronic device, including a memory 301 storing executable program code and a processor 302 coupled to the memory 301;
[0227] The processor 302 calls the executable program code stored in the memory 301 to execute the fusion gene detection method based on second-generation sequencing described in the above embodiments.
[0228] This invention also discloses a computer-readable storage medium storing a computer program that causes a computer to execute the second-generation sequencing-based fusion gene detection method described in the above embodiments.
[0229] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.
[0230] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.
Claims
1. A method for detecting fusion genes based on next-generation sequencing, characterized in that, include: The sample to be tested is sequenced and compared with a genomic reference sequence to obtain the next-generation sequencing data of the sample to be tested; The second-generation sequencing data were analyzed to obtain multiple structural variations in the sample to be tested; Each structural variation of the sample to be tested is annotated and used as a candidate fusion gene variation, which is then merged to obtain a candidate fusion gene set. Event identification is performed on the candidate fusion gene variants in the candidate fusion gene set, and candidate fusion gene variants identified as the same event are annotated with the same event number. Calculate the target mutation frequency for each event; The candidate fusion gene set is filtered according to the first preset list to obtain the first fusion gene set; Based on the second preset list and / or the third preset list, the first fusion gene set is filtered to obtain the second fusion gene set; The second fusion gene set is subjected to quality filtering to obtain a third fusion gene set; wherein, the first preset list is a protection list, the second preset list is a blacklist list, and the third preset list is a specified gene list; Based on the target mutation frequency of each event, the third fusion gene set is filtered according to fusion properties to obtain the fourth fusion gene set; The target fusion gene variants in the first fusion gene set and the fourth fusion gene set are annotated and merged to obtain the target fusion gene set. Calculate the target mutation frequency for each event, including: Determine whether each event is a complex event; If the event is determined to be a non-complex event, calculate the first and second variation frequencies for each event when using different evidence; Based on the magnitude of the first mutation frequency, the candidate mutation frequency for the event is selected from the first mutation frequency and the second mutation frequency. The potential mutation frequency of this event is determined as the target mutation frequency; Calculate the first and second frequencies of variation for each event when using different evidence, including: For each event, obtain SR and PR evidence for all candidate fusion gene variants, and then perform deduplication to obtain the total number of unique SR and unique PR evidence for each event. Calculate the first and second numbers of non-variant reads in the first and second breakpoint neighborhoods of each candidate fusion gene variant in each event, which have a specified proportion of similarity to the reference genome sequence. The first mutation frequency for each event is calculated based on the total number of unique SR evidences and the first and second quantities corresponding to each candidate fusion gene mutation. Based on the total number of unique SR evidence, the total number of unique PR evidence, and the first and second quantities corresponding to each candidate fusion gene variant, calculate the second variant frequency for each event; The method further includes: If the event is determined to be a complex event, the third variation frequency of the event is calculated based on the number of SR evidence for each candidate fusion gene variant, the number of PR evidence for each candidate fusion gene variant, the number of split reads output by each candidate fusion gene variant that have a specified similarity to the reference genome sequence, and the number of cross reads output by each candidate fusion gene variant that have a specified similarity to the reference genome sequence. The third mutation frequency of this event is determined as the target mutation frequency.
2. The method for detecting fusion genes based on next-generation sequencing as described in claim 1, characterized in that, The first preset list is a protection list; The candidate fusion gene set is filtered according to a first preset list to obtain a first fusion gene set, including: The candidate fusion gene variants present in the protection list are retained in the candidate fusion gene set to obtain the first fusion gene set.
3. The method for detecting fusion genes based on next-generation sequencing as described in claim 1, characterized in that, The second preset list is a blacklist, and the third preset list is a list of specified genes; Based on the second preset list and / or the third preset list, the first fusion gene set is filtered to obtain the second fusion gene set, including: Remove fusion gene variants from the first fusion gene set that exist in the blacklist or have at least one breakpoint in a region of the blacklist, and / or retain at least one candidate fusion gene variant that has at least one breakpoint in a specified gene list to obtain the second fusion gene set.
4. The method for detecting fusion genes based on next-generation sequencing as described in claim 1, characterized in that, Event identification of candidate fusion gene variants in the candidate fusion gene set, including: If the first and second breakpoints of a candidate fusion gene mutation are within a specified distance from the corresponding breakpoint of another candidate fusion gene mutation, and these two candidate fusion gene mutations belong to the same mutation type, they are considered to be the same event.
5. A fusion gene detection device based on next-generation sequencing, characterized in that, include: The alignment unit is used to sequence the sample to be tested and align it with the genome reference sequence to obtain the next-generation sequencing data of the sample to be tested. The detection unit is used to detect the next-generation sequencing data and obtain multiple structural variations of the sample to be tested; The first annotation unit is used to annotate each structural variation of the sample to be tested as a candidate fusion gene variation, and then merge them to obtain a candidate fusion gene set. The event identification unit is used to identify the candidate fusion gene variants in the candidate fusion gene set as events, and to annotate the candidate fusion gene variants identified as the same event with the same event number. A computational unit is used to calculate the target mutation frequency for each event; The first filtering unit is used to filter the candidate fusion gene set according to a first preset list to obtain the first fusion gene set. The second filtering unit is used to perform list filtering on the first fusion gene set according to the second preset list and / or the third preset list to obtain the second fusion gene set; The third filtering unit is used to perform quality filtering on the second fusion gene set to obtain a third fusion gene set; wherein, the first preset list is a protection list, the second preset list is a blacklist list, and the third preset list is a specified gene list; The fourth filtering unit is used to filter the third fusion gene set based on fusion properties according to the target mutation frequency of each event to obtain the fourth fusion gene set. The second annotation unit is used to annotate the target fusion gene variants in the first fusion gene set and the fourth fusion gene set, and merge them to obtain the target fusion gene set. Calculate the target mutation frequency for each event, including: Determine whether each event is a complex event; If the event is determined to be a non-complex event, calculate the first and second variation frequencies for each event when using different evidence; Based on the magnitude of the first mutation frequency, the candidate mutation frequency for the event is selected from the first mutation frequency and the second mutation frequency. The potential mutation frequency of this event is determined as the target mutation frequency; Calculate the first and second frequencies of variation for each event when using different evidence, including: For each event, obtain SR and PR evidence for all candidate fusion gene variants, and then perform deduplication to obtain the total number of unique SR and unique PR evidence for each event. Calculate the first and second numbers of non-variant reads in the first and second breakpoint neighborhoods of each candidate fusion gene variant in each event, which have a specified proportion of similarity to the reference genome sequence. The first mutation frequency for each event is calculated based on the total number of unique SR evidences and the first and second quantities corresponding to each candidate fusion gene mutation. Based on the total number of unique SR evidence, the total number of unique PR evidence, and the first and second quantities corresponding to each candidate fusion gene variant, calculate the second variant frequency for each event; The method further includes: If the event is determined to be a complex event, the third variation frequency of the event is calculated based on the number of SR evidence for each candidate fusion gene variant, the number of PR evidence for each candidate fusion gene variant, the number of split reads output by each candidate fusion gene variant that have a specified similarity to the reference genome sequence, and the number of cross reads output by each candidate fusion gene variant that have a specified similarity to the reference genome sequence. The third mutation frequency of this event is determined as the target mutation frequency.
6. An electronic device, characterized in that, It includes a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the fusion gene detection method based on next-generation sequencing as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program causes a computer to perform the next-generation sequencing-based fusion gene detection method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Structural variation detection method
CN111326212A
Fusion gene identification method and apparatus, device, program, and storage medium
WO2023184065A1