Sequencing Read Depth Compression Using Ordinal RRD Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The analysis of large nucleotide datasets generated by high-throughput sequencing is costly and time-consuming, requiring significant data storage capacity and processing resources.
Innovation Solution
A method involving converting read depth values to Relative Read Depths (RRDs) by ordinal class assignment and mapping these to a reference nucleotide sequence at base-pair resolution, allowing for data compression and efficient processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If read depth values are stored and processed in full resolution, then measurement precision is improved, but data storage capacity and processing resources increase significantly
Solution Approach 1:
The patent segments the continuous read depth values into discrete ordinal classes (e.g., 0-10, 11-20, etc.). This segmentation reduces the data storage requirements while preserving the essential quantitative information needed for biological interpretation, directly resolving the contradiction between measurement precision and storage capacity.
Solution Approach 2:
The patent changes the parameter representation from continuous numerical values to discrete ordinal categories. This parameter transformation maintains the relative ordering and quantitative relationships of read depth values while dramatically reducing the storage space and computational resources required to handle the data.
2Measurement precision
If full-resolution read depth data is processed, then analysis accuracy is improved, but processing time increases
Solution Approach 1:
By segmenting read depth values into ordinal classes, the patent reduces the computational complexity of processing large datasets. The segmented data requires less computational power for operations like sorting, filtering, and statistical analysis, thereby reducing processing time while maintaining sufficient accuracy for biological conclusions.
Solution Approach 2:
The transformation of read depth parameters from continuous to discrete ordinal values simplifies subsequent data processing operations. This parameter change enables more efficient algorithms to be applied, reducing processing time while preserving the essential quantitative relationships needed for accurate biological analysis.
3Loss of information
If high-throughput sequencing is performed to obtain comprehensive nucleotide datasets, then information completeness is improved, but data storage and processing requirements increase
Solution Approach 1:
The patent applies segmentation to compress the comprehensive nucleotide sequencing data by grouping read depth values into ordinal classes. This approach retains the essential information about gene expression levels and genomic variations while significantly reducing the storage space required to hold the complete dataset.
Solution Approach 2:
By changing the parameter representation of sequencing data from detailed continuous measurements to ordinal categories, the patent reduces storage requirements while maintaining the information necessary for biological interpretation, including gene expression quantification and variant detection.
Data Source
AI summary
The present invention relates to the field of genetics, more in particular to quantitative sequencing data. A method of processing and/or compressing quantitative sequencing data is provided, a computer-readable storage medium to comprise such processed data and a computing device comprising at least one processor configured to process such data.

