NGS Data Compression via Block Addressing for Selective Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Next-generation sequencing (NGS) technologies generate vast amounts of genomic data, leading to storage and analysis challenges due to large file sizes, with existing compression methods requiring full decompression to access specific sequencing reads, which is inefficient for targeted analyses.
Innovation Solution
A method and apparatus for compressing and decompressing genetic information using an addressing scheme that groups aligned sequencing reads into blocks based on intervals of a reference sequence, allowing for random access and selective decompression of specific regions, thereby reducing storage needs and improving analysis efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing compression methods are used for NGS data, then file size is reduced, but random access to specific sequencing reads becomes impossible requiring full decompression
Solution Approach 1:
The patent divides the NGS data into multiple blocks, where each block contains sequencing reads aligned to specific intervals of a reference sequence. This segmentation allows the system to store compressed data in a structured format that enables selective access to specific blocks without decompressing the entire file, thus resolving the contradiction between compression and random access capability
Solution Approach 2:
The patent introduces an addressing scheme that maps blocks to intervals on the reference sequence, adding a dimensional structure to the compressed data. This addressing layer enables the system to navigate and access specific regions of the genome efficiently, transforming the data organization from a flat compressed structure to a hierarchical one that supports random access
2Ease of operation
If the entire compressed file is decompressed to access specific reads, then random access is enabled, but storage efficiency and processing time are reduced
Solution Approach 1:
The patent extracts and stores only the necessary addressing information and block structures in the compressed file format. This extraction allows the system to retrieve specific blocks by their addresses without processing the entire file, enabling efficient access while minimizing the computational resources required for decompression operations
3Volume of stationary object
If NGS data is stored without compression, then random access is straightforward, but storage requirements become prohibitively large
Solution Approach 1:
By segmenting the data into addressable blocks mapped to reference sequence intervals, the patent creates a manageable structure that balances compression efficiency with access capability. This segmentation reduces storage requirements while maintaining reasonable data management complexity through systematic organization
Solution Approach 2:
The patent introduces an addressing scheme as an intermediary layer between the compressed data blocks and the reference sequence. This intermediary structure simplifies data management by providing a clear mapping relationship, making the compressed data as easy to manage as uncompressed data while achieving significant space savings
Data Source
AI summary
Provided are methods and apparatuses for compressing genetic information, the methods and apparatuses obtaining read information about reads and alignment information about positions of the reads that are aligned to a reference sequence, and generating a compressed file comprising information about an address of a block corresponding to the aligned reads. Also, a method and apparatus for decompressing genetic information obtains a compressed file with respect to the genetic information, determines an address of a block corresponding to input gene search information, from the compressed file, and selectively decompresses genetic information corresponding to the determined address.


