Parallel Surprisal Data Reduction for Genome Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The large volume of data generated by DNA gene sequencing, requiring significant storage space and hindering data transmission and analysis, particularly due to the time-consuming comparison of genetic sequences using single computer processors.
Innovation Solution
A method utilizing multiple computer processing elements to divide and compare genetic sequences and reference genomes, identifying and storing 'surprisal data' which represents differences, thereby reducing data volume and accelerating analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all 3 billion nucleotide bases are transmitted and stored, then complete genetic sequence information is preserved, but storage space requirements increase significantly and data transmission becomes hindered
Solution Approach 1:
The patent extracts only the surprising or novel nucleotide sequences from the complete genetic sequence, separating them from the redundant common sequences. By taking out only the essential informative portions (surprisal data) rather than transmitting the entire sequence, the system reduces data volume while preserving critical genetic information.
Solution Approach 2:
The system creates a compressed representation of the genetic sequence by copying only the surprising portions and using the reference genome to reconstruct the complete sequence. This copying approach allows reconstruction of the full genome from a much smaller set of surprisal data, reducing storage and transmission requirements.
2Measurement precision
If sequence comparison is performed by a single computer processor, then complete analysis is achieved, but analysis time becomes significantly large
Solution Approach 1:
The patent segments the sequence comparison task by dividing the reference genome into multiple parts and distributing them across different computer processors. Each processor independently compares its assigned genome segment with the corresponding sequence portion, then the results are combined. This segmentation enables parallel processing that maintains comparison accuracy while significantly reducing total analysis time.
3Loss of information
If common nucleotide sequences are transmitted and stored, then complete genome information is preserved, but data storage needs increase significantly
Solution Approach 1:
The system extracts and stores only the surprising nucleotide sequences that differ from the reference genome, eliminating the need to store common sequences. By taking out only the informative surprisal data and using the reference genome as a template for reconstruction, the system dramatically reduces storage space requirements while maintaining complete genome information.
Solution Approach 2:
The patent changes the representation parameter from storing the complete sequence to storing only the differences (surprisal data). This parameter transformation allows the same genome information to be represented in a much more compact form, reducing storage volume from 3 gigabytes to a fraction of that size.
Data Source
AI summary
A method, computer product, and computer system of reducing an amount of data representing a genetic sequence of an organism, comprising: a computer dividing a reference genome and a sequence of the organism into parts and assigning the parts to one of a plurality of computer processing elements. Within each computer processing element, comparing nucleotides of the genetic sequence of the organism to nucleotides from a part of the reference genome, to find differences where nucleotides of the genetic sequence of the organism which are different from the nucleotides of the reference genome; and storing the surprisal data in a repository. Combining the parts of the surprisal data from the repository to form a complete set of surprisal data representing the differences between the genetic sequence of the organism and the reference genome; and storing the complete set of surprisal data in the repository.


