Parallel High-Throughput Sequencing Analysis with Differential Outputs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for comparative genomic analysis face challenges in processing large datasets efficiently, leading to slow and difficult management of genomic data, especially when comparing multiple large datasets for genomic changes, and require significant computational resources.
Innovation Solution
A method involving incremental synchronization of genetic sequence strings using known positions of corresponding sub-strings to generate local alignments, reducing the need for processing massive files and generating minimal output files with high information density, using a sequence analysis engine to produce differential genetic sequence objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple large genomic datasets are processed using traditional comparative analysis methods, then comprehensive genomic comparison is achieved, but processing time increases and computational resources are excessively consumed
Solution Approach 1:
The patent divides large genomic datasets into smaller contig segments that can be processed independently and in parallel. This segmentation allows the system to compare genomic regions separately, reducing the computational burden on any single processing unit while maintaining comprehensive coverage of the entire genome through aggregation of results.
Solution Approach 2:
The patent introduces a new dimension of processing by implementing parallel computation across multiple processors simultaneously analyzing different contig segments. This dimensional expansion from sequential to parallel processing dramatically reduces overall processing time while maintaining the accuracy of genomic comparisons through coordinated aggregation of results.
2Reliability
If high genome sequencing coverage is used to obtain statistically relevant data, then data reliability is improved, but data storage requirements and processing complexity increase significantly
Solution Approach 1:
The patent extracts and processes only the essential genomic information from high-coverage sequencing data by analyzing contig segments and their mappings. This extraction approach retrieves statistically relevant signals while discarding redundant information, maintaining data reliability without requiring storage and processing of all raw sequencing data.
Solution Approach 2:
The patent implements partial action by focusing computational resources on analyzing only the portions of genomic data that contain relevant variations and differences. Rather than processing every base pair equally, the system identifies and analyzes only those regions that contribute to statistically significant findings, reducing overall data volume requirements.
3Loss of information
If traditional genomic analysis methods are used to compare multiple large datasets, then complete genomic coverage is achieved, but output file size becomes unmanageably large with low information density
Solution Approach 1:
The patent extracts only the essential differential information from comprehensive genomic comparisons by identifying and reporting only significant variations between samples. This extraction process maintains complete genomic coverage for analysis purposes while producing compact output files that contain only the high-information-density results needed for interpretation.
Solution Approach 2:
The patent applies local quality by providing different levels of detail in output files based on the specific genomic region and its importance. Rather than uniformly detailing every genomic position, the system concentrates detailed information on regions with significant variations while providing summarized information for conserved regions, optimizing output information density.
4Measurement precision
If multiple massive genomic files are processed simultaneously, then comprehensive differential analysis is achieved, but computational resource requirements become prohibitive
Solution Approach 1:
The patent segments the computational task of comparing multiple genomic files into independent contig-level analyses that can be distributed across multiple processors. This segmentation allows comprehensive differential analysis to be achieved through coordinated processing of smaller units, reducing the peak computational resource requirements compared to loading entire genomes into memory simultaneously.
Solution Approach 2:
The patent performs preliminary actions by pre-processing and indexing contig segments before main analysis, and by pre-identifying regions of interest through initial scanning. This preliminary work reduces the computational intensity of the main comparison operations, allowing accurate differential analysis with reduced real-time resource consumption.
Data Source
AI summary
A differential sequence object is constructed on the basis of alignment of sub-strings via incremental synchronization of sequence strings using known positions of the sub-strings relative to a reference genome sequence. An output file is then gencrated that comprises only relevant changes with respect to the reference genome.


