Sequencing Data Processing Method for DNA Fragment Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Processing of genetic data is time-consuming due to the large volumes of sequencing reads that need to be processed to determine genetic sequences, necessitating faster and more efficient methods for data processing.
Innovation Solution
A sequencing data processing method that includes multiple adapter trimming passes, stitching, extracting, first matching, deduplication, and second matching steps to efficiently identify DNA fragments from sequencing reads, utilizing adapter trimming to remove adapter data and improve sequence quality scores, and re-labeling insert bases for accurate DNA fragment identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional processing methods are used to handle voluminous sequencing reads, then comprehensive data processing is achieved, but processing time becomes excessively long
Solution Approach 1:
The patent segments the processing of voluminous sequencing reads into multiple passes, where each pass handles a specific subset of reads or a specific processing stage. This segmentation allows the system to process data in manageable chunks rather than attempting to process all reads simultaneously, thereby reducing overall processing time while maintaining comprehensive analysis.
Solution Approach 2:
The patent performs preliminary actions by pre-processing sequencing reads before main analysis, including initial quality filtering, adapter trimming, and basic error correction. This preliminary processing reduces the burden on subsequent analysis stages and accelerates the overall workflow by preparing data in advance for more efficient processing.
2Measurement precision
If adapter trimming is performed to improve sequence quality, then sequence quality scores are improved, but processing complexity increases
Solution Approach 1:
The adapter trimming process is divided into multiple sequential passes, with each pass focusing on removing adapters from specific regions or types of reads. This segmented approach to trimming reduces the computational complexity of any single pass while achieving comprehensive adapter removal across all reads, thereby improving sequence quality without overwhelming processing complexity.
Solution Approach 2:
The patent applies partial trimming actions in early passes, removing only the most obvious or abundant adapter sequences first, then progressively addressing remaining adapters in subsequent passes. This staged approach improves quality incrementally while managing processing complexity by not attempting to remove all possible adapters in a single exhaustive pass.
Data Source
AI summary
Embodiments of the present disclosure are directed to systems, apparatuses, devices and methods for processing sequencing data for determining the identity of DNA fragments from a plurality of reads contained in a sequencing data file.


