Neural Network Base Calling for Overlapping Sequencing Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current nucleic acid cluster-based sequencing technologies face limitations in resolving data from closely proximate or overlapping clusters, leading to reduced throughput and quality of nucleic acid sequencing data, which is a challenge in various genomic and diagnostic applications.
Innovation Solution
The use of neural network-based methods and systems that process image data from sequencing images to improve base calling accuracy and throughput by segregating processing among sequencing cycles, utilizing distance and scaling channels, and refining cluster center positions, enabling more precise classification of bases and quality scoring.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If neural network-based methods are used to process image data from sequencing images, then base calling accuracy is improved, but computational resources required increase
Solution Approach 1:
The patent segments the neural network processing into multiple sequencing cycles, where the network processes image data from different cycles in separate steps. This allows for optimized computational resource allocation at each stage rather than requiring all resources simultaneously, thereby maintaining high base calling accuracy while managing computational demands.
Solution Approach 2:
The patent implements preliminary processing steps including generating distance channels and scaling channels before the main neural network base calling process. Cluster center positions are refined in advance, and image data is preprocessed with additional channels that encode spatial and intensity relationships. This preliminary action reduces the computational burden during the actual base calling by providing pre-computed features to the neural network.
2Productivity
If neural network-based methods process data from closely proximate or overlapping clusters, then throughput is improved, but measurement precision deteriorates
Solution Approach 1:
The patent adds multiple dimensions to the image data by generating distance channels and scaling channels. The distance channel encodes the distance from each pixel to the nearest cluster center, while the scaling channel encodes the relative intensity scaling. These additional dimensions provide the neural network with richer spatial context, enabling it to accurately resolve closely proximate or overlapping clusters that would be indistinguishable in the original image data alone, thus maintaining high measurement precision while improving throughput.
3Use of energy by moving object
If processing is segregated among sequencing cycles, then computational resources required are reduced, but device complexity increases
Solution Approach 1:
The patent segments the base calling process into distinct sequencing cycle steps, where the neural network processes image data from different cycles separately. Each cycle's image data is enhanced with distance and scaling channels specific to that cycle, and the neural network performs base calling for each cycle in sequence. This segmentation reduces the computational resources required at any given moment compared to processing all cycles simultaneously, while the modular structure makes the complexity manageable through systematic organization.
Data Source
AI summary
The technology disclosed processes input data through a neural network and produces an alternative representation of the input data. The input data includes per-cycle image data for each of one or more sequencing cycles of a sequencing run. The per-cycle image data depicts intensity emissions of one or more analytes and their surrounding background captured at a respective sequencing cycle. The technology disclosed processes the alternative representation through an output layer and producing an output and base calls one or more of the analytes at one or more of the sequencing cycles based on the output.


