Sequence Read Storage by Deduplicating Human-Readable Text Files

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for compressing sequencing data generated by next-generation sequencing technologies are inadequate as they create binary, lossy, or specialized files that are not human-readable, posing a bottleneck in storage and transfer costs, which hinders the ability of clinics and labs to perform genetic carrier screening and other research services.

Innovation Solution

A method that identifies and stores only unique sequence information from redundant sequence read files, allowing for lossless compression and storage in human-readable text formats, significantly reducing storage and transfer requirements by exploiting the redundancy in sequencing data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If existing compression methods are used on sequencing data, then file size is reduced, but the files become binary, lossy, or specialized formats that are not human-readable

Engineering Contradiction:
Improvefile sizeVSAvoidhuman-readability
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent extracts only the essential unique sequence information from redundant sequencing reads, storing just the distinct sequences rather than all duplicate reads. This extraction approach reduces file size while maintaining human-readability through text-based formats like FASTA, as the compressed files contain actual sequence data that can be directly read and analyzed by researchers.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the representation parameter from storing complete redundant reads to storing only unique sequences with reduced information. By identifying and retaining only distinct sequences while discarding exact duplicates, the file size is reduced while the data remains in a human-readable text format that preserves the essential biological information.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If redundancy in sequencing data is maintained, then data completeness is preserved, but storage costs and transfer bandwidth increase significantly

Engineering Contradiction:
Improvedata completenessVSAvoidstorage space
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates a simplified copy of the sequencing data that retains only unique sequences. Instead of storing multiple redundant copies of identical reads, the system stores a single representative copy of each unique sequence, significantly reducing storage requirements while maintaining the essential data content needed for analysis.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent discards redundant duplicate sequences that provide no additional information value, while preserving the unique sequences that contain the essential biological data. This selective discarding of duplicates maintains data completeness for analysis purposes while dramatically reducing the total quantity of stored information.

Inventive Principle:
Principle #34Discarding and recovering

3Loss of information

If all sequence reads are stored in original format, then no information is lost, but the data volume creates a bottleneck for storage and transfer

Engineering Contradiction:
Improveinformation retentionVSAvoiddata transfer efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent extracts and stores only the essential unique sequence information, removing redundant data that does not contribute to biological insight. This extraction maintains sufficient information for downstream analysis while improving data transfer efficiency by reducing the total data volume that needs to be moved between systems.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of compressing data by creating binary or specialized formats that sacrifice readability, the patent inverts the approach by using text-based formats that are inherently more compact for sequence data and maintain human-readability. This inversion achieves compression while preserving the ability to directly read and analyze the sequence information.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS9292527B2Methods and systems for storing sequence read data
Publication Date: 2016.03.22 LABORATORY CORPORATION OF AMERICA HOLDINGS INC
  • US9292527B2 patent drawing
  • US9292527B2 patent drawing
  • US9292527B2 patent drawing

AI summary

The present invention generally relates to storing sequence read data. The invention can involve obtaining a plurality of sequence reads from a sample, identifying one or more sets of duplicative sequence reads within the plurality of sequence reads, and storing only one of the sequence reads from each set of duplicative sequence reads in a text file using nucleotide characters.