Sequence Read Storage Using Unique Text-Based Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for compressing sequencing data generated by next-generation sequencing technologies are inadequate as they create binary, lossy, or specialized files that are not human-readable, posing a bottleneck for storage and transfer, which hinders the ability of clinics and labs to perform genetic carrier screening and other research services.
Innovation Solution
The method involves identifying and storing only unique sequence information from redundant sequence reads in human-readable text files, using meta information to correlate with unique sequences, thereby reducing storage and transfer costs and maintaining lossless data retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing compression methods are used for sequencing data, then storage space is reduced, but the files become binary, lossy, or specialized and not human-readable
Solution Approach 1:
The patent changes the fundamental parameter of data representation from binary encoding to text-based encoding (FASTA/FASTQ formats). This allows the data to remain human-readable while implementing compression through selective storage of unique sequences and their metadata, resolving the contradiction between space reduction and readability.
Solution Approach 2:
The patent extracts and stores only the essential unique sequence information and metadata, discarding redundant duplicate sequences. This extraction approach reduces storage requirements while maintaining data integrity and human-readability through standard text formats.
2Reliability
If all sequence reads are stored in full, then data integrity is maintained, but storage costs and transfer bandwidth requirements increase significantly
Solution Approach 1:
The patent creates a compressed representation by copying only the unique sequence information and its metadata, rather than storing every redundant read. This copying strategy maintains data integrity for analysis purposes while dramatically reducing storage space requirements.
Solution Approach 2:
The patent changes the storage parameter from storing complete redundant reads to storing unique sequences with metadata pointers. This parameter change preserves the ability to reconstruct and analyze data while reducing storage requirements by eliminating redundancy.
3Productivity
If compression is applied to sequencing data, then storage and transfer efficiency improves, but data may become lossy or require specialized alignment programs
Solution Approach 1:
The patent implements self-service compression where the unique sequence data and metadata structure inherently supports both compression and future analysis operations. The metadata contains sufficient information to reconstruct and analyze reads without requiring external specialized programs, eliminating information loss while improving efficiency.
Solution Approach 2:
The patent creates a universal file format that serves multiple functions: compression storage, human readability, and direct analysis capability. This multi-functional approach eliminates the need for specialized alignment programs while maintaining data integrity, resolving the contradiction between efficiency and information preservation.
4Reliability
If redundant sequence reads are stored, then complete sequencing coverage is maintained, but file sizes become extremely large
Solution Approach 1:
The patent extracts and retains only the unique sequence information from redundant reads, removing duplicate data while preserving the essential sequencing coverage information. This extraction maintains the scientific value of the data while reducing file size dramatically.
Solution Approach 2:
The patent segments the sequencing data into unique sequences and their associated metadata, storing them in a structured format. This segmentation allows efficient storage of coverage information without storing redundant full read sequences, reducing file size while maintaining coverage reliability.
Data Source
AI summary
The present invention generally relates to storing sequence read data. The invention can involve obtaining a plurality of sequence reads from a sample, identifying one or more sets of duplicative sequence reads within the plurality of sequence reads, and storing only one of the sequence reads from each set of duplicative sequence reads in a text file using nucleotide characters.


