Parallel Surprisal Data Reduction for Genome Sequencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The large volume of data generated by DNA gene sequencing, requiring significant storage space and hindering data transmission and analysis, particularly due to the time-consuming comparison of genetic sequences using single computer processors.

Innovation Solution

A method utilizing multiple computer processing elements to divide and compare genetic sequences and reference genomes, identifying and storing 'surprisal data' which represents differences, thereby reducing data volume and accelerating analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If all 3 billion nucleotide bases are transmitted and stored, then complete genetic sequence information is preserved, but storage space requirements increase significantly and data transmission becomes hindered

Engineering Contradiction:
Improvegenetic sequence informationVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent extracts only the surprising or novel nucleotide sequences from the complete genetic sequence, separating them from the redundant common sequences. By taking out only the essential informative portions (surprisal data) rather than transmitting the entire sequence, the system reduces data volume while preserving critical genetic information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates a compressed representation of the genetic sequence by copying only the surprising portions and using the reference genome to reconstruct the complete sequence. This copying approach allows reconstruction of the full genome from a much smaller set of surprisal data, reducing storage and transmission requirements.

Inventive Principle:
Principle #26Copying

2Measurement precision

If sequence comparison is performed by a single computer processor, then complete analysis is achieved, but analysis time becomes significantly large

Engineering Contradiction:
Improvesequence comparison accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the sequence comparison task by dividing the reference genome into multiple parts and distributing them across different computer processors. Each processor independently compares its assigned genome segment with the corresponding sequence portion, then the results are combined. This segmentation enables parallel processing that maintains comparison accuracy while significantly reducing total analysis time.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If common nucleotide sequences are transmitted and stored, then complete genome information is preserved, but data storage needs increase significantly

Engineering Contradiction:
Improvegenome informationVSAvoidstorage space
Core Design Contradiction:
Loss of informationVSVolume of stationary object

Solution Approach 1:

The system extracts and stores only the surprising nucleotide sequences that differ from the reference genome, eliminating the need to store common sequences. By taking out only the informative surprisal data and using the reference genome as a template for reconstruction, the system dramatically reduces storage space requirements while maintaining complete genome information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the representation parameter from storing the complete sequence to storing only the differences (surprisal data). This parameter transformation allows the same genome information to be represented in a much more compact form, reducing storage volume from 3 gigabytes to a fraction of that size.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8751166B2Parallelization of surprisal data reduction and genome construction from genetic data for transmission, storage, and analysis
Publication Date: 2014.06.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8751166B2 patent drawing
  • US8751166B2 patent drawing
  • US8751166B2 patent drawing

AI summary

A method, computer product, and computer system of reducing an amount of data representing a genetic sequence of an organism, comprising: a computer dividing a reference genome and a sequence of the organism into parts and assigning the parts to one of a plurality of computer processing elements. Within each computer processing element, comparing nucleotides of the genetic sequence of the organism to nucleotides from a part of the reference genome, to find differences where nucleotides of the genetic sequence of the organism which are different from the nucleotides of the reference genome; and storing the surprisal data in a repository. Combining the parts of the surprisal data from the repository to form a complete set of surprisal data representing the differences between the genetic sequence of the organism and the reference genome; and storing the complete set of surprisal data in the repository.