Neural Network Base Calling Using 3D Convolutions for Sequencing Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current base calling methods struggle to accurately account for technical artifacts, biases, and error profiles in sequencing data, requiring substantial technical expertise and explicit programming for feature engineering and kinetic modeling.
Innovation Solution
A neural network-based base caller that uses a combination of 3D convolutions, 1D convolutions, and pointwise convolutions to automatically extract features from assay data and learn to detect and account for biases such as phasing, prephasing, spatial crosstalk, emission overlap, and fading.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If kinetic models with explicit programming for feature engineering are used, then base calling accuracy can be maintained with known technical artifacts, but the system requires substantial technical expertise and biochemistry intuition, increasing device complexity and reducing ease of operation
Solution Approach 1:
The patent replaces manual kinetic modeling and explicit programming with a deep neural network system. The neural network automatically learns features and error profiles from data without requiring manual programming of kinetic models, substituting the mechanical process of explicit feature engineering with an automated learning-based approach that reduces technical expertise requirements while maintaining accuracy
Solution Approach 2:
The neural network system performs self-learning and self-adjustment by automatically extracting features and detecting biases from sequencing data. The system serves itself by learning error profiles and adjusting base calling decisions without external intervention or manual programming, thereby reducing the need for substantial technical expertise while maintaining high accuracy
2Reliability
If manual feature engineering and kinetic modeling are performed, then specific biases can be accounted for, but the process is time-consuming and reduces throughput
Solution Approach 1:
The neural network system performs preliminary learning of error profiles and bias patterns during training on known data. This preliminary action allows the system to automatically account for biases like phasing, prephasing, and spatial crosstalk during actual sequencing without time-consuming manual feature engineering, thereby maintaining reliability while increasing throughput
Solution Approach 2:
The patent substitutes manual kinetic modeling processes with automated neural network inference. The neural network rapidly processes sequencing data and automatically adjusts for biases using learned models, replacing the time-consuming manual feature engineering and kinetic modeling steps, thus maintaining accuracy while significantly improving sequencing throughput
3Extent of automation
If deep neural networks with multiple convolution types are used, then automatic feature extraction and bias detection are achieved, but computational resource requirements increase
Solution Approach 1:
The patent segments the convolutional neural network into distinct types of convolutions (1D, 2D, and 3D) that process different aspects of the data. This segmentation allows the system to automatically extract features through specialized convolution layers while managing computational resources by applying the appropriate type of convolution to the appropriate data dimension, achieving automation without excessive computational burden
Solution Approach 2:
The neural network utilizes multiple dimensional convolutions (1D for temporal patterns, 2D for spatial patterns, 3D for spatio-temporal patterns) to extract features from sequencing data. This dimensional approach enables automatic feature extraction across multiple scales and types of biases while optimizing computational efficiency by selecting the appropriate convolution dimensionality for each specific feature extraction task
Data Source
AI summary
We propose a neural network-implemented method for base calling analytes. The method includes accessing a sequence of per-cycle image patches for a series of sequencing cycles, where pixels in the image patches contain intensity data for associated analytes, and applying three-dimensional (3D) convolutions on the image patches on a sliding convolution window basis such that, in a convolution window, a 3D convolution filter convolves over a plurality of the image patches and produces at least one output feature. The method further includes beginning with output features produced by the 3D convolutions as starting input, applying further convolutions and producing final output features and processing the final output features through an output layer and producing base calls for one or more of the associated analytes to be base called at each of the sequencing cycles.


