Genomic Surprisal Data Compression via Reference Genome Hierarchy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The large volume of DNA gene sequencing data, typically requiring 3 gigabytes of storage space per 3 billion nucleotide base pairs, poses challenges in data transmission and storage, especially when including annotations, due to the significant resources needed for handling and transferring such large datasets.
Innovation Solution
A method that identifies characteristics of a genetic sequence, generates a hierarchy, compares it to a repository of reference genomes, breaks matched genomes into pieces, and creates surprisal data by highlighting differences, thereby reducing data volume and focusing on unique nucleotide variations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If the entire genome sequence is transmitted and stored, then complete genetic information is preserved, but storage space and transmission resources are significantly increased
Solution Approach 1:
The patent extracts only the unique or variable portions of the genome sequence (surprisal data) rather than transmitting the entire genome. By identifying and isolating only the nucleotide differences from a reference genome, the system removes redundant information while preserving all essential genetic variations.
Solution Approach 2:
The genome sequence is segmented into reference portions and variable portions. The system divides the complete genome into a reference genome (which can be stored separately) and surprisal data (the unique variations), allowing selective transmission of only the necessary segments.
2Adaptability or versatility
If annotations and genetic specifics are included in the data, then comprehensive analysis capability is improved, but storage requirements increase to terabytes
Solution Approach 1:
The patent extracts only the essential variable information needed for analysis rather than storing complete annotated genomes. By pulling out only the surprisal data (unique nucleotide variations), the system maintains analytical capability while dramatically reducing storage requirements.
3Adaptability or versatility
If large genetic datasets are transmitted between institutions, then data sharing and collaboration are enabled, but transmission time and cost are significantly increased
Solution Approach 1:
The patent extracts only the essential variable portions of genomic data for transmission. By sending only the surprisal data (unique variations) rather than complete genomes, the system enables rapid data sharing between institutions while preserving all scientifically relevant information.
Solution Approach 2:
The genomic data is segmented into reference portions (stored locally) and variable portions (transmitted). This segmentation allows institutions to share only the necessary differential data, dramatically reducing transmission time and cost while maintaining collaboration capabilities.
Data Source
AI summary
A method, computer product, and computer system of minimizing surprisal data comprising: at a source, reading and identifying characteristics of a genetic sequence of an organism; receiving an input of rank of at least two identified characteristics of the genetic sequence of the organism; generating a hierarchy of ranked, identified characteristics based on the rank of the at least two identified characteristics of the genetic sequence of the organism; comparing the hierarchy of ranked, identified characteristics to a repository of reference genomes; and if at least one reference genome from the repository matches the hierarchy of ranked, identified characteristics, breaking the matched reference genomes into pieces, combining pieces associated with the identified characteristics from at least one matched reference genome to form a filter pattern to be compared to the nucleotides of the genetic sequence of the organism, to obtain differences and create surprisal data.


