Parallel High-Throughput Sequencing Analysis with Differential Outputs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for comparative genomic analysis face challenges in processing large datasets efficiently, leading to slow and difficult management of genomic data, especially when comparing multiple large datasets for genomic changes, and require significant computational resources.

Innovation Solution

A method involving incremental synchronization of genetic sequence strings using known positions of corresponding sub-strings to generate local alignments, reducing the need for processing massive files and generating minimal output files with high information density, using a sequence analysis engine to produce differential genetic sequence objects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple large genomic datasets are processed using traditional comparative analysis methods, then comprehensive genomic comparison is achieved, but processing time increases and computational resources are excessively consumed

Engineering Contradiction:
Improvegenomic comparison accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides large genomic datasets into smaller contig segments that can be processed independently and in parallel. This segmentation allows the system to compare genomic regions separately, reducing the computational burden on any single processing unit while maintaining comprehensive coverage of the entire genome through aggregation of results.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of processing by implementing parallel computation across multiple processors simultaneously analyzing different contig segments. This dimensional expansion from sequential to parallel processing dramatically reduces overall processing time while maintaining the accuracy of genomic comparisons through coordinated aggregation of results.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If high genome sequencing coverage is used to obtain statistically relevant data, then data reliability is improved, but data storage requirements and processing complexity increase significantly

Engineering Contradiction:
Improvestatistical relevanceVSAvoiddata volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts and processes only the essential genomic information from high-coverage sequencing data by analyzing contig segments and their mappings. This extraction approach retrieves statistically relevant signals while discarding redundant information, maintaining data reliability without requiring storage and processing of all raw sequencing data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements partial action by focusing computational resources on analyzing only the portions of genomic data that contain relevant variations and differences. Rather than processing every base pair equally, the system identifies and analyzes only those regions that contribute to statistically significant findings, reducing overall data volume requirements.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of information

If traditional genomic analysis methods are used to compare multiple large datasets, then complete genomic coverage is achieved, but output file size becomes unmanageably large with low information density

Engineering Contradiction:
Improvegenomic information completenessVSAvoidoutput management complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts only the essential differential information from comprehensive genomic comparisons by identifying and reporting only significant variations between samples. This extraction process maintains complete genomic coverage for analysis purposes while producing compact output files that contain only the high-information-density results needed for interpretation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by providing different levels of detail in output files based on the specific genomic region and its importance. Rather than uniformly detailing every genomic position, the system concentrates detailed information on regions with significant variations while providing summarized information for conserved regions, optimizing output information density.

Inventive Principle:
Principle #3Local quality

4Measurement precision

If multiple massive genomic files are processed simultaneously, then comprehensive differential analysis is achieved, but computational resource requirements become prohibitive

Engineering Contradiction:
Improvedifferential analysis accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the computational task of comparing multiple genomic files into independent contig-level analyses that can be distributed across multiple processors. This segmentation allows comprehensive differential analysis to be achieved through coordinated processing of smaller units, reducing the peak computational resource requirements compared to loading entire genomes into memory simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-processing and indexing contig segments before main analysis, and by pre-identifying regions of interest through initial scanning. This preliminary work reduces the computational intensity of the main comparison operations, allowing accurate differential analysis with reduced real-time resource consumption.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250292872A1Bambam: parallel comparative analysis of high-throughput sequencing data
Publication Date: 2025.09.18 RGT UNIV OF CALIFORNIA
  • US20250292872A1 patent drawing
  • US20250292872A1 patent drawing
  • US20250292872A1 patent drawing

AI summary

A differential sequence object is constructed on the basis of alignment of sub-strings via incremental synchronization of sequence strings using known positions of the sub-strings relative to a reference genome sequence. An output file is then gencrated that comprises only relevant changes with respect to the reference genome.