Polymer Sequence Determination via Non-Canonical Unit Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining polymer sequences through nanopore translocation face challenges in accuracy, particularly with long reads, due to high computational expense and errors associated with non-canonical polymer units, which are not recognized by traditional base-calling techniques.
Innovation Solution
A method involving machine learning techniques, such as recurrent neural networks, to attribute measurements of non-canonical polymer units to corresponding canonical units, reducing computational requirements and improving sequencing accuracy by recognizing non-canonical units as canonical during analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional base-calling techniques are used to determine polymer sequences, then the sequencing process is straightforward, but accuracy deteriorates due to unrecognized non-canonical polymer units and high computational expense
Solution Approach 1:
The patent extracts and separately handles non-canonical polymer units from the sequencing analysis. By identifying these units as distinct entities and creating specific mappings for them, the system removes the source of recognition errors that plague traditional base-calling techniques, thereby improving sequencing accuracy without proportionally increasing computational complexity
Solution Approach 2:
The patent introduces an intermediary mapping layer between raw nanopore measurements and final sequence determination. This mapping explicitly connects non-canonical polymer units to their corresponding canonical representations, serving as a mediator that translates complex measurement data into accurate sequence information while managing computational requirements
2Measurement precision
If non-canonical polymer units are recognized and processed, then sequencing accuracy improves, but computational expense increases
Solution Approach 1:
The patent performs preliminary classification and mapping of non-canonical polymer units before the main sequence determination process. By pre-establishing mapping relationships and identifying non-canonical units in advance, the system avoids redundant computational work during sequence assembly, thereby improving accuracy while controlling computational energy consumption
Solution Approach 2:
The patent changes the parameter representation of non-canonical polymer units by mapping them to canonical equivalents. This parameter transformation simplifies the data structure and reduces the computational burden of processing diverse polymer unit types, enabling accurate recognition without proportional increases in computational energy requirements
3Measurement precision
If comprehensive analysis of all polymer units is performed, then sequencing accuracy improves, but systematic errors increase due to complexity
Solution Approach 1:
The patent converts the harmful effect of non-canonical polymer unit complexity into a benefit by creating explicit mapping relationships. Instead of treating non-canonical units as sources of error, the system leverages their distinct characteristics to improve recognition accuracy, transforming what was previously a harmful factor into an advantageous feature for comprehensive analysis
Data Source
AI summary
The invention resides in a method of determining a sequence of a target polymer, or part thereof, comprising polymer units comprising canonical and non-canonical polymer units. The method comprises taking a series of measurements of a signal relating to the target polymer wherein a measurement of the signal is dependent upon a plurality of polymer units, and wherein the polymer units of the target polymer modulate the signal, and wherein a non-canonical polymer unit modulates the signal differently from a corresponding canonical polymer unit. The series of measurements are analysed using a machine learning technique that attributes a measurement of a non-canonical polymer unit to being a measurement of a respective corresponding canonical polymer unit. The sequence of the target polymer, or part thereof, is determined from the analysed series of measurements. A non-canonical polymer unit identified from the analysis can be additionally or alternatively determined. Two or more types of non-canonical polymer units corresponding to the two or more types of canonical polymer unit can be used. The polynucleotide can be DNA.


