Molecular Tag Consensus Compression for Efficient Variant Calling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for storing and processing large amounts of molecular tagged nucleic acid sequence data are inefficient in terms of memory usage and computational efficiency, particularly for variant calling operations, without compromising data quality.

Innovation Solution

A method involving the grouping of sequence reads with the same molecular tag to calculate consensus flow space signal measurements and standard deviations, determining a consensus base sequence and alignment, and generating a compressed data structure that includes these elements, thereby reducing data size while maintaining quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If molecular tagged nucleic acid sequence data is stored and processed using conventional methods, then complete data quality is maintained for variant calling, but memory requirements are excessively high and computational efficiency is poor

Engineering Contradiction:
Improvedata qualityVSAvoidmemory requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent groups multiple sequence reads that share the same molecular tag into a single family and computes a consensus sequence for the entire family. This merging approach consolidates redundant data while preserving the essential genetic information needed for accurate variant calling, thereby reducing memory requirements without sacrificing data quality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

Instead of storing and processing every individual sequence read, the patent creates a consensus copy that represents the entire family of reads. This consensus sequence serves as a compressed representation that captures the essential information from multiple reads, enabling efficient storage and processing while maintaining the ability to perform accurate variant calling.

Inventive Principle:
Principle #26Copying

2Reliability

If molecular tagged nucleic acid sequence data is stored and processed using conventional methods, then complete data quality is maintained for variant calling, but computational efficiency is poor

Engineering Contradiction:
Improvedata qualityVSAvoidcomputational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent merges multiple sequence reads into consensus families, reducing the total number of sequences that need to be processed individually. This consolidation maintains data quality for variant calling while significantly improving computational efficiency by reducing the processing burden.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary grouping and consensus sequence computation before the actual variant calling process. By pre-processing the data to create consolidated consensus families, the system reduces the computational workload for subsequent variant calling operations, thereby improving overall computational efficiency without compromising accuracy.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If sequence reads are grouped by molecular tag to form families, then data compression is achieved, but data structure complexity increases

Engineering Contradiction:
Improvedata sizeVSAvoiddata structure
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the large set of sequence reads into smaller, manageable families based on shared molecular tags. Each family is then represented by a single consensus sequence. This segmentation approach compresses the overall data size while organizing the information into a structured format that is easier to process and analyze.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10892037B2Methods for compression of molecular tagged nucleic acid sequence data
Publication Date: 2021.01.12 LIFE TECHNOLOGIES CORP
  • US10892037B2 patent drawing
  • US10892037B2 patent drawing
  • US10892037B2 patent drawing

AI summary

A method for compressing molecular tagged sequence data includes: grouping sequence reads associated with a molecular tag sequence to form a family of sequence reads, corresponding vectors of flow space signal measurements and corresponding sequence alignments, calculating an arithmetic mean of the corresponding vectors of flow space signal measurements to form a vector of consensus flow space signal measurements, calculating a standard deviation of the corresponding vectors of flow space signal measurements to form a vector of standard deviations, determining a consensus base sequence based on the vector of consensus flow space signal measurements, determining a consensus sequence alignment and generating a compressed data structure comprising consensus compressed data, the consensus compressed data including for each family, the consensus base sequence, the consensus sequence alignment, the vector of consensus flow space signal measurements, the vector of standard deviations and the number of members.