Polymer Sequence Trace Reconstruction with Quality-Weighted Voting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for error correction in DNA data storage using oligonucleotides are limited in accuracy and computational efficiency, leading to challenges in reliably recovering digital data from noisy sequencer outputs.

Innovation Solution

Implementing a trace reconstruction system that uses quality scores from sequencers to perform weighted majority voting, generating a consensus output sequence by assigning weights based on error probabilities to improve the accuracy of nucleotide sequence reconstruction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing error correction techniques are used, then data can be recovered from sequencer output, but the accuracy of error correction is limited

Engineering Contradiction:
Improveaccuracy of error correctionVSAvoidreliability of data recovery
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the parameter of error correction by transitioning from traditional majority voting to quality-score-weighted majority voting. Each nucleotide call is weighted by its quality score, allowing the system to prioritize more reliable calls and reduce the impact of erroneous ones, thereby improving both accuracy and reliability of data recovery

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the simple mechanical majority voting mechanism with a more sophisticated statistical approach that incorporates quality scores. This substitution allows the system to differentiate between high-confidence and low-confidence nucleotide calls, improving the overall accuracy of trace reconstruction

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If traditional majority voting is used, then consensus sequence can be generated, but computational cost is high

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoidcomputational cost
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent optimizes the computational parameters by using quality scores to weight votes, which allows for more efficient convergence to the correct consensus sequence. The weighted approach reduces the number of iterations needed compared to traditional unweighted majority voting, thereby improving productivity while managing computational cost

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If quality scores are used for weighted voting, then reconstruction accuracy improves, but system complexity increases

Engineering Contradiction:
Improveaccuracy of sequence reconstructionVSAvoidcomplexity of error correction system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces quality scores as an additional parameter to enhance reconstruction accuracy. While this increases system complexity, the complexity is managed through efficient algorithms that leverage the quality score information without requiring overly complex computational structures

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4052260B1Trace reconstruction of polymer sequences using quality scores
Publication Date: 2025.07.02 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4052260B1 patent drawingFigure 1
  • EP4052260B1 patent drawingFigure 2
  • EP4052260B1 patent drawingFigure 3

AI summary

Polymeric molecules such as deoxyribose nucleic acid (DNA) provide a storage medium for digital data that has advantages over conventional storage media. Accessing digital data stored in polymers includes decoding the output of sequencers which detect the physical order of monomer subunits in the polymers. This output includes errors which are corrected through the process of trace reconstruction. Trace reconstruction identifies a consensus output sequence from a set of noisy output reads provided by a sequencer. The accuracy of trace reconstruction is improved by using weighted majority voting to determine the consensus output sequence. Weights are based on quality labels assigned by the sequencer to its output. Quality labels may be derived from empirical error data. A quality label for a single position in an output read may be determined independently or it may be influenced by quality labels of other nearby positions in the read.