Fusion Gene Identification via Spanning and Split Read Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying gene fusions, particularly through targeted sequencing, are limited in detecting unknown fusion genes and are often designed for specific known fusion genes, making them unsuitable for identifying various types of fusion genes in conventional target intervals, with high sequencing costs and large data volumes leading to computational and storage challenges.
Innovation Solution
A method and apparatus for identifying fusion genes by acquiring and aligning target and reference gene sequences, screening for spanning and split reads, calculating breakpoint positions, and filtering candidate fusion gene pairs based on quality and sequencing depth to determine fusion gene scores, allowing for the identification of fusion gene pairs without prior knowledge of specific fusion genes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If targeted sequencing is used to identify known fusion genes, then detection accuracy for specific fusion genes is improved, but the ability to detect unknown fusion genes deteriorates
Solution Approach 1:
The patent develops a universal identification method that can detect both known and unknown fusion genes using the same targeted sequencing approach. The method employs a reference gene sequence database and alignment algorithms that work for any fusion gene type, making the system multi-functional rather than species-specific.
Solution Approach 2:
The patent performs preliminary alignment of sequencing reads to reference gene sequences before fusion gene identification. This preliminary action creates a foundation that enables subsequent detection of various fusion gene types without requiring separate methodologies for each gene species.
2Adaptability or versatility
If whole genome sequencing is performed to detect all possible fusion genes, then detection coverage is improved, but sequencing cost and data volume increase
Solution Approach 1:
The patent extracts and focuses only on the relevant portions of the genome that contain fusion genes by using targeted sequencing with reference gene sequences. This extraction approach captures necessary information for fusion gene detection while discarding unnecessary genomic data, significantly reducing data volume compared to whole genome sequencing.
Solution Approach 2:
The patent applies local quality enhancement by concentrating sequencing depth and computational resources on specific target regions where fusion genes are likely to occur. This localized approach improves detection sensitivity in relevant areas while minimizing overall data generation.
3Measurement precision
If multiple dedicated methods are developed for different fusion gene species, then detection precision for each species is improved, but method complexity increases
Solution Approach 1:
The patent creates a unified identification system that handles multiple fusion gene species through a single methodological framework. The reference gene sequence database and alignment-based identification process work universally across different gene types, eliminating the need for multiple dedicated methods.
Solution Approach 2:
The patent adjusts identification parameters such as breakpoint position calculation and spanning read thresholds based on the specific characteristics of different fusion gene candidates, rather than creating entirely separate methods. This parameter-based adaptation maintains precision while reducing overall system complexity.
Data Source
AI summary
The present disclosure provides a method and apparatus for identifying a fusion gene, a device, a program and a storage medium, belonging to the technical field of gene detection. The method includes: acquiring a target gene sequencing sequence to be identified and a reference gene sequence; aligning the target gene sequencing sequence to the reference gene sequence, and acquiring distribution and targeted capturing results of spanning reads and split reads of the target sequencing sequence located in a target area; screening a target fusion gene pair from the reads based on the distribution and the targeted capture results; and outputting an identification result regarding the target fusion gene pair.


