Genomic Surprisal Data Compression via Reference Genome Hierarchy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The large volume of DNA gene sequencing data, typically requiring 3 gigabytes of storage space per 3 billion nucleotide base pairs, poses challenges in data transmission and storage, especially when including annotations, due to the significant resources needed for handling and transferring such large datasets.

Innovation Solution

A method that identifies characteristics of a genetic sequence, generates a hierarchy, compares it to a repository of reference genomes, breaks matched genomes into pieces, and creates surprisal data by highlighting differences, thereby reducing data volume and focusing on unique nucleotide variations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the entire genome sequence is transmitted and stored, then complete genetic information is preserved, but storage space and transmission resources are significantly increased

Engineering Contradiction:
Improvegenetic information completenessVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent extracts only the unique or variable portions of the genome sequence (surprisal data) rather than transmitting the entire genome. By identifying and isolating only the nucleotide differences from a reference genome, the system removes redundant information while preserving all essential genetic variations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The genome sequence is segmented into reference portions and variable portions. The system divides the complete genome into a reference genome (which can be stored separately) and surprisal data (the unique variations), allowing selective transmission of only the necessary segments.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If annotations and genetic specifics are included in the data, then comprehensive analysis capability is improved, but storage requirements increase to terabytes

Engineering Contradiction:
Improveanalysis capabilityVSAvoidstorage space
Core Design Contradiction:
Adaptability or versatilityVSVolume of stationary object

Solution Approach 1:

The patent extracts only the essential variable information needed for analysis rather than storing complete annotated genomes. By pulling out only the surprisal data (unique nucleotide variations), the system maintains analytical capability while dramatically reducing storage requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If large genetic datasets are transmitted between institutions, then data sharing and collaboration are enabled, but transmission time and cost are significantly increased

Engineering Contradiction:
Improvedata sharing capabilityVSAvoidtransmission time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent extracts only the essential variable portions of genomic data for transmission. By sending only the surprisal data (unique variations) rather than complete genomes, the system enables rapid data sharing between institutions while preserving all scientifically relevant information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The genomic data is segmented into reference portions (stored locally) and variable portions (transmitted). This segmentation allows institutions to share only the necessary differential data, dramatically reducing transmission time and cost while maintaining collaboration capabilities.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10353869B2Minimization of surprisal data through application of hierarchy filter pattern
Publication Date: 2019.07.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10353869B2 patent drawing
  • US10353869B2 patent drawing
  • US10353869B2 patent drawing

AI summary

A method, computer product, and computer system of minimizing surprisal data comprising: at a source, reading and identifying characteristics of a genetic sequence of an organism; receiving an input of rank of at least two identified characteristics of the genetic sequence of the organism; generating a hierarchy of ranked, identified characteristics based on the rank of the at least two identified characteristics of the genetic sequence of the organism; comparing the hierarchy of ranked, identified characteristics to a repository of reference genomes; and if at least one reference genome from the repository matches the hierarchy of ranked, identified characteristics, breaking the matched reference genomes into pieces, combining pieces associated with the identified characteristics from at least one matched reference genome to form a filter pattern to be compared to the nucleotides of the genetic sequence of the organism, to obtain differences and create surprisal data.