Adaptive Nanopore Preprocessing for Large-Array Sequencing Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Nanopore-based DNA sequencing systems generate vast amounts of data, making it inefficient to process all data from large arrays of sequencing sensor cells, as only relevant information for calibration and base determination is needed.
Innovation Solution
A pre-processing circuit extracts relevant information from the data generated by nanopore-based sequencing sensor chips, reducing the data transferred by selecting and sending only the necessary information for further processing, allowing for adaptive extraction based on system requests.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all data from 100,000 or more sensor cells is transferred for processing, then complete information is available for analysis, but data transfer volume and processing time become excessively large
Solution Approach 1:
The patent extracts only the essential features from raw nanopore data - specifically, the minimum and maximum voltage values within each measurement window - and transmits only these extracted features for further processing. This extraction principle reduces data volume dramatically while preserving the critical information needed for base calling and quality assessment.
Solution Approach 2:
The patent performs preliminary processing of the data at the sensor chip level before transmission, calculating statistical features (min, max, mean, standard deviation) and filtering out redundant information in advance. This preliminary action ensures that only processed, condensed data needs to be transmitted and further processed, reducing overall system processing time.
2Productivity
If a large number of sensor cells are used for parallel sequencing, then sequencing throughput increases, but data generation volume increases proportionally
Solution Approach 1:
The patent segments the data processing task by dividing it into two stages: (1) local extraction of essential features at each sensor cell, and (2) centralized processing of only the extracted features. This segmentation allows parallel sequencing of many cells while keeping the data transmission and central processing workload manageable.
Solution Approach 2:
The patent applies partial action by selecting only the most critical data features (min, max, mean, standard deviation) for transmission and processing, rather than processing all raw data. This partial processing approach maintains sequencing throughput while significantly reducing the quantity of data that must be handled.
3Measurement precision
If detailed processing of all raw data is performed, then accurate base determination is achieved, but computational resources and processing time are excessively consumed
Solution Approach 1:
The patent extracts the most discriminative features from the raw nanopore voltage data - specifically the minimum and maximum voltage values that correspond to different base positions - and uses only these extracted features for base determination. This extraction maintains measurement precision while dramatically reducing computational requirements.
Solution Approach 2:
The patent transforms the raw voltage data into a different parameter space by calculating statistical features (min, max, mean, standard deviation) and using these transformed parameters for base calling. This parameter change simplifies the computational task while preserving the information needed for accurate base determination.
Data Source
AI summary
Techniques described herein relate to systems and methods for parallel DNA molecules sequencing. A preprocessor can receive raw data frames from a sensor chip including 100,000 or more cells, where each raw data frame can include detection signals from the 100,000 or more cells at a given time during the formation of the 100,000 or more cells or during the DNA molecules sequencing using the 100,000 or more cells. The preprocessor can then extract relevant information for determining states of the cells from the raw data frames, generate one or more digested frames that includes the extracted information, and send the digested frames to a processor for processing, such as base determination. Because the number of digested frames sent to the processor is less than a number of the raw data frames and the digested frames include preprocessed data, the amount of data being transferred to the processor and the amount of data processing by the processor can be reduced.


