Sequence Read Storage by Deduplicating Human-Readable Text Files
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for compressing sequencing data generated by next-generation sequencing technologies are inadequate as they create binary, lossy, or specialized files that are not human-readable, posing a bottleneck in storage and transfer costs, which hinders the ability of clinics and labs to perform genetic carrier screening and other research services.
Innovation Solution
A method that identifies and stores only unique sequence information from redundant sequence read files, allowing for lossless compression and storage in human-readable text formats, significantly reducing storage and transfer requirements by exploiting the redundancy in sequencing data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing compression methods are used on sequencing data, then file size is reduced, but the files become binary, lossy, or specialized formats that are not human-readable
Solution Approach 1:
The patent extracts only the essential unique sequence information from redundant sequencing reads, storing just the distinct sequences rather than all duplicate reads. This extraction approach reduces file size while maintaining human-readability through text-based formats like FASTA, as the compressed files contain actual sequence data that can be directly read and analyzed by researchers.
Solution Approach 2:
The patent changes the representation parameter from storing complete redundant reads to storing only unique sequences with reduced information. By identifying and retaining only distinct sequences while discarding exact duplicates, the file size is reduced while the data remains in a human-readable text format that preserves the essential biological information.
2Reliability
If redundancy in sequencing data is maintained, then data completeness is preserved, but storage costs and transfer bandwidth increase significantly
Solution Approach 1:
The patent creates a simplified copy of the sequencing data that retains only unique sequences. Instead of storing multiple redundant copies of identical reads, the system stores a single representative copy of each unique sequence, significantly reducing storage requirements while maintaining the essential data content needed for analysis.
Solution Approach 2:
The patent discards redundant duplicate sequences that provide no additional information value, while preserving the unique sequences that contain the essential biological data. This selective discarding of duplicates maintains data completeness for analysis purposes while dramatically reducing the total quantity of stored information.
3Loss of information
If all sequence reads are stored in original format, then no information is lost, but the data volume creates a bottleneck for storage and transfer
Solution Approach 1:
The patent extracts and stores only the essential unique sequence information, removing redundant data that does not contribute to biological insight. This extraction maintains sufficient information for downstream analysis while improving data transfer efficiency by reducing the total data volume that needs to be moved between systems.
Solution Approach 2:
Instead of compressing data by creating binary or specialized formats that sacrifice readability, the patent inverts the approach by using text-based formats that are inherently more compact for sequence data and maintain human-readability. This inversion achieves compression while preserving the ability to directly read and analyze the sequence information.
Data Source
AI summary
The present invention generally relates to storing sequence read data. The invention can involve obtaining a plurality of sequence reads from a sample, identifying one or more sets of duplicative sequence reads within the plurality of sequence reads, and storing only one of the sequence reads from each set of duplicative sequence reads in a text file using nucleotide characters.


