Genomic Data Compression via Boolean Array Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Whole genome sequencing data from microorganisms is not being widely adopted for clinical diagnostics due to the lack of efficient methods to compress and compare genomic information across large databases in a timely and accurate manner.

Innovation Solution

A diagnostic analysis method that converts nucleotide sequences into a Boolean array format, using nucleotide sequence anchors and finite length nucleotide sequences, allowing for rapid comparison and identification of microorganisms by scoring against a reference database, resulting in a diagnostic score for species identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If whole genome sequencing data is stored and compared using traditional formats, then comprehensive genomic information is preserved, but processing time increases and storage requirements become unmanageable

Engineering Contradiction:
Improveidentification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts specific diagnostic features from complete genome sequences by identifying and comparing only the most variable and informative genomic regions. This selective extraction maintains identification accuracy while dramatically reducing the data volume that must be processed and stored, resolving the contradiction between comprehensive information preservation and processing efficiency

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The genome sequence is segmented into discrete diagnostic features or markers that can be independently analyzed. By dividing the complete genome into these functional segments, the system achieves rapid comparison across large databases without sacrificing the ability to accurately identify microbial species, thus resolving the time-accuracy tradeoff

Inventive Principle:
Principle #1Segmentation

2Reliability

If complete genome sequences are stored in databases, then comprehensive reference data is available, but storage space requirements become prohibitive

Engineering Contradiction:
Improvediagnostic reliabilityVSAvoiddata storage volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential diagnostic information from complete genome sequences, storing merely the extracted features rather than the entire sequences. This extraction approach maintains diagnostic reliability by preserving the most informative elements while reducing storage requirements from gigabytes to kilobytes per genome

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of storing complete genomes and extracting features during analysis, the patent inverts the approach by pre-extracting and storing only the diagnostic features. This inversion fundamentally reduces storage volume while maintaining the ability to perform reliable diagnostic comparisons

Inventive Principle:
Principle #13The other way round (Inversion)

3Measurement precision

If traditional genome comparison methods are used, then comprehensive genomic analysis is performed, but scalability to large databases is limited

Engineering Contradiction:
Improvespecies identification accuracyVSAvoiddatabase processing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The comparison process is segmented into efficient computational steps operating on extracted features rather than complete sequences. This segmentation enables parallel processing and scalable database queries, maintaining high identification accuracy while achieving the productivity needed for clinical diagnostic throughput

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the fundamental parameters of comparison by working with extracted diagnostic features instead of complete sequences. This parameter change transforms the computational complexity from O(n*m) for full sequence comparison to a much more efficient operation on condensed feature sets, enabling scalability to large databases

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10752958B2Identification of microorganisms from genome sequencing data
Publication Date: 2020.08.25 SANTA CLARA UNIVERSITY
  • US10752958B2 patent drawing
  • US10752958B2 patent drawing
  • US10752958B2 patent drawing

AI summary

A diagnostic analysis method and system is provided for identifying a microorganism from a genome sequence. Partially or fully assembled microbial genomes or short reads from whole-genome sequencing of microbial genomes are processed into a 4 MB Boolean array while preserving 1% of the genomic information in a way that allows for rapid comparison of a query genome to a large reference database. This represents a critical savings in storage space and speed by which large reference libraries can be queried.