Sequencing Read Depth Compression Using Ordinal RRD Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The analysis of large nucleotide datasets generated by high-throughput sequencing is costly and time-consuming, requiring significant data storage capacity and processing resources.

Innovation Solution

A method involving converting read depth values to Relative Read Depths (RRDs) by ordinal class assignment and mapping these to a reference nucleotide sequence at base-pair resolution, allowing for data compression and efficient processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If read depth values are stored and processed in full resolution, then measurement precision is improved, but data storage capacity and processing resources increase significantly

Engineering Contradiction:
Improveread depth measurement precisionVSAvoiddata storage capacity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the continuous read depth values into discrete ordinal classes (e.g., 0-10, 11-20, etc.). This segmentation reduces the data storage requirements while preserving the essential quantitative information needed for biological interpretation, directly resolving the contradiction between measurement precision and storage capacity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation from continuous numerical values to discrete ordinal categories. This parameter transformation maintains the relative ordering and quantitative relationships of read depth values while dramatically reducing the storage space and computational resources required to handle the data.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If full-resolution read depth data is processed, then analysis accuracy is improved, but processing time increases

Engineering Contradiction:
Improveanalysis accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

By segmenting read depth values into ordinal classes, the patent reduces the computational complexity of processing large datasets. The segmented data requires less computational power for operations like sorting, filtering, and statistical analysis, thereby reducing processing time while maintaining sufficient accuracy for biological conclusions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The transformation of read depth parameters from continuous to discrete ordinal values simplifies subsequent data processing operations. This parameter change enables more efficient algorithms to be applied, reducing processing time while preserving the essential quantitative relationships needed for accurate biological analysis.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If high-throughput sequencing is performed to obtain comprehensive nucleotide datasets, then information completeness is improved, but data storage and processing requirements increase

Engineering Contradiction:
Improveinformation completenessVSAvoiddata storage requirements
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent applies segmentation to compress the comprehensive nucleotide sequencing data by grouping read depth values into ordinal classes. This approach retains the essential information about gene expression levels and genomic variations while significantly reducing the storage space required to hold the complete dataset.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

By changing the parameter representation of sequencing data from detailed continuous measurements to ordinal categories, the patent reduces storage requirements while maintaining the information necessary for biological interpretation, including gene expression quantification and variant detection.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260105986A1Nucleotide sequencing data compression
Publication Date: 2026.04.16 KEYGENE NV
  • US20260105986A1 patent drawing
  • US20260105986A1 patent drawing

AI summary

The present invention relates to the field of genetics, more in particular to quantitative sequencing data. A method of processing and/or compressing quantitative sequencing data is provided, a computer-readable storage medium to comprise such processed data and a computing device comprising at least one processor configured to process such data.