Molecular Tag Consensus Compression for Efficient Variant Calling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for storing and processing large amounts of molecular tagged nucleic acid sequence data are inefficient in terms of memory usage and computational efficiency, particularly for variant calling operations, without compromising data quality.
Innovation Solution
A method involving the grouping of sequence reads with the same molecular tag to calculate consensus flow space signal measurements and standard deviations, determining a consensus base sequence and alignment, and generating a compressed data structure that includes these elements, thereby reducing data size while maintaining quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If molecular tagged nucleic acid sequence data is stored and processed using conventional methods, then complete data quality is maintained for variant calling, but memory requirements are excessively high and computational efficiency is poor
Solution Approach 1:
The patent groups multiple sequence reads that share the same molecular tag into a single family and computes a consensus sequence for the entire family. This merging approach consolidates redundant data while preserving the essential genetic information needed for accurate variant calling, thereby reducing memory requirements without sacrificing data quality.
Solution Approach 2:
Instead of storing and processing every individual sequence read, the patent creates a consensus copy that represents the entire family of reads. This consensus sequence serves as a compressed representation that captures the essential information from multiple reads, enabling efficient storage and processing while maintaining the ability to perform accurate variant calling.
2Reliability
If molecular tagged nucleic acid sequence data is stored and processed using conventional methods, then complete data quality is maintained for variant calling, but computational efficiency is poor
Solution Approach 1:
The patent merges multiple sequence reads into consensus families, reducing the total number of sequences that need to be processed individually. This consolidation maintains data quality for variant calling while significantly improving computational efficiency by reducing the processing burden.
Solution Approach 2:
The patent performs preliminary grouping and consensus sequence computation before the actual variant calling process. By pre-processing the data to create consolidated consensus families, the system reduces the computational workload for subsequent variant calling operations, thereby improving overall computational efficiency without compromising accuracy.
3Quantity of substance
If sequence reads are grouped by molecular tag to form families, then data compression is achieved, but data structure complexity increases
Solution Approach 1:
The patent segments the large set of sequence reads into smaller, manageable families based on shared molecular tags. Each family is then represented by a single consensus sequence. This segmentation approach compresses the overall data size while organizing the information into a structured format that is easier to process and analyze.
Data Source
AI summary
A method for compressing molecular tagged sequence data includes: grouping sequence reads associated with a molecular tag sequence to form a family of sequence reads, corresponding vectors of flow space signal measurements and corresponding sequence alignments, calculating an arithmetic mean of the corresponding vectors of flow space signal measurements to form a vector of consensus flow space signal measurements, calculating a standard deviation of the corresponding vectors of flow space signal measurements to form a vector of standard deviations, determining a consensus base sequence based on the vector of consensus flow space signal measurements, determining a consensus sequence alignment and generating a compressed data structure comprising consensus compressed data, the consensus compressed data including for each family, the consensus base sequence, the consensus sequence alignment, the vector of consensus flow space signal measurements, the vector of standard deviations and the number of members.


