Cluster detection and filtering based on artificial intelligence predicted basecalls
A neural network-based system enhances sequencing accuracy by detecting and filtering unreliable clusters, addressing the challenges of existing technologies in base calling processes.
Patent Information
- Application Number
- JP2022581614
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-25
- Filing Date
- 2021-08-26
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2041-08-26
AI Technical Summary
Existing sequencing technologies struggle with accurately identifying and filtering unreliable clusters, leading to reduced accuracy and efficiency in base calling processes.
Utilizing a neural network-based system that analyzes sequencing images to detect and filter unreliable clusters by generating probability quartiles and applying specialized convolutional architectures to improve base calling accuracy.
Enhances the accuracy and efficiency of base calling by effectively identifying and filtering unreliable clusters, improving the overall quality of sequencing results.
Smart Images

Figure 0007767335000004 
Figure 0007767335000005 
Figure 0007767335000006
Abstract
Description
[Technical Field]
[0001] Priority application This application claims priority to U.S. Patent No. 17 / 411,980, entitled "DETECTING AND FILTERING CLUSTERS BASED ON ARTIFICIAL INTELLIGENCE-PREDICTED BASE CALLS," filed August 25, 2021 (Attorney Docket No. ILLM1018-2 / IP-1860-US), which claims the benefit of U.S. Provisional Patent Application No. 63 / 072,032, entitled "DETECTING AND FILTERING CLUSTERS BASED ON ARTIFICIAL INTELLIGENCE-PREDICTED BASE CALLS," filed August 28, 2020 (Attorney Docket No. ILLM1018-1 / IP-1860-PRV), which is incorporated herein by reference.
[0002] Built-in The following are incorporated by reference for all purposes as if fully set forth herein: U.S. Provisional Patent Application No. 62 / 821,602, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed March 21, 2019 (Attorney Docket No. ILLM1008-1 / IP-1693-PRV); U.S. Provisional Patent Application No. 62 / 821,618, entitled "Artificial Intelligence-Based Generation of Sequencing Metadata," filed March 21, 2019 (Attorney Docket No. ILLM1008-3 / IP-1741-PRV); U.S. Provisional Patent Application No. 62 / 821,681, entitled "Artificial Intelligence-Based Base Calling," filed March 21, 2019 (Attorney Docket No. ILLM1008-4 / IP-1744-PRV); U.S. Provisional Patent Application No. 62 / 821,724, entitled "Artificial Intelligence-Based Quality Scoring," filed March 21, 2019 (Attorney Docket No. ILLM1008-7 / IP-1747-PRV); U.S. Provisional Patent Application No. 62 / 821,766, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV); Dutch Patent Application No. 2023310 (Attorney Docket No. ILLM1008-11 / IP-1693-NL), entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed on June 14, 2019; Dutch Patent Application No. 2023311 (Attorney Docket No. ILLM1008-12 / IP-1741-NL), entitled "Artificial Intelligence-Based Generation of Sequencing Metadata," filed on June 14, 2019; Dutch Patent Application No. 2023312 entitled "Artificial Intelligence-Based Base Calling" filed on June 14, 2019 (Attorney Docket No. ILLM1008-13 / IP-1744-NL); Dutch Patent Application No. 2023314 (Attorney Docket No. ILLM1008-14 / IP-1747-NL), entitled "Artificial Intelligence-Based Quality Scoring," filed on June 14, 2019; Dutch Patent Application No. 2023316 (Attorney Docket No. ILLM1008-15 / IP-1752-NL), entitled "Artificial Intelligence-Based Sequencing," filed on June 14, 2019; U.S. Provisional Patent Application No. 62 / 849,091, entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing," filed May 16, 2019 (Attorney Docket No. ILLM1011-1 / IP-1750-PRV); U.S. Provisional Patent Application No. 62 / 849,132, entitled "Base Calling Using Convolutions," filed May 16, 2019 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2); U.S. Provisional Patent Application No. 62 / 849,133, entitled "Base Calling Using Compact Convolutions," filed May 16, 2019 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3); U.S. Provisional Patent Application No. 62 / 979,384, entitled "Artificial Intelligence-Based Base Calling of Index Sequences," filed February 20, 2020 (Attorney Docket No. ILLM1015-1 / IP-1857-PRV); U.S. Provisional Patent Application No. 62 / 979,414, entitled "Artificial Intelligence-Based Many-To-Many Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1016-1 / IP-1858-PRV); U.S. Provisional Patent Application No. 62 / 979,385, entitled "Knowledge Distillation-Based Compression of Artificial Intelligence-Based Base Caller," filed February 20, 2020 (Attorney Docket No. ILLM1017-1 / IP-1859-PRV); U.S. Provisional Patent Application No. 62 / 979,412, entitled "Multi-Cycle Cluster Based Real Time Analysis System," filed February 20, 2020 (Attorney Docket No. ILLM1020-1 / IP-1866-PRV); U.S. Provisional Patent Application No. 62 / 979,411, entitled "Data Compression for Artificial Intelligence-Based Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1029-1 / IP-1964-PRV); and U.S. Provisional Patent Application No. 62 / 979,399, entitled "Squeezing Layer for Artificial Intelligence-Based Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1030-1 / IP-1982-PRV).
[0003] The disclosed technology relates to artificial intelligence-based computers and digital data processing systems, and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the disclosed technology relates to using deep neural networks, such as deep convolutional neural networks, to analyze data. [Background technology]
[0004] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or problems associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to embodiments of the claimed technology.
[0005] Base calls assign bases and associated quality values to each position in a read. The quality of sequenced bases is assessed by an Illumina sequencer using a procedure called chastity filtering. Chastity can be determined as the highest intensity value divided by the sum of the highest intensity value and the second-highest intensity value. Quality assessment can include identifying reads in a first subset of base calls whose second-worst chastity is below a threshold and marking those reads as poor-quality data. The first subset of base calls can be any suitable number of base calls. For example, the subset can be the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or more than the first 25 base calls. This is sometimes called read filtering, and clusters that meet this cutoff are said to be "filtered."
[0006]
number
[0007] In some embodiments, the purity of the signal from each cluster is examined over the first 25 cycles and calculated as a chastity value: up to one cycle may be below a chastity threshold (e.g., 0.6) or the read will not pass the chastity filter.
[0008] Illumina calculates a Phred score, which is used to evaluate the error probability of a base call. The Phred score is calculated based on the intensity profile (shifted purity: how much signal is accounted for by the brightest channel?) and the signal-to-noise ratio (signal overlap with background: is the signal from the colony well delineated from the surrounding area of the flow cell?). Illumina attempts to quantify the chastity of the strongest base signal, whether the signal for a given base call is much stronger than the signals for nearby bases, whether the spot representing the colony is suspiciously fuzzy during the sequencing process (intensity decay), and whether the signals in previous and subsequent cycles appear clean.
[0009] Based on artificial intelligence predicted base calling, there is an opportunity to detect and filter unreliable clusters, which can result in improved accuracy and quality of base calling. [Brief explanation of the drawings]
[0010] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Figure 1] 1A-1D are block diagrams illustrating various aspects of the disclosed technology. [Figure 2A] 1 illustrates an exemplary softmax function. [Figure 2B] 10 shows exemplary cluster-by-cluster, cycle-by-cycle probability quartiles generated by the disclosed techniques. [Figure 3] An example of using filter values to identify unreliable clusters is shown. [Figure 4] FIG. 1 is a flow diagram showing one embodiment of a method for identifying unreliable clusters to improve the accuracy and efficiency of base calling. [Figure 5A]1 illustrates one embodiment of a sequencing system, the sequencing system including a configurable processor. [Figure 5B] 1 illustrates one embodiment of a sequencing system, the sequencing system including a configurable processor. [Figure 5C] FIG. 1 is a simplified block diagram of a system for analysis of sensor data from a sequencing system, such as base call sensor output. [Figure 6] FIG. 1 illustrates one embodiment of the disclosed data flow logic that enables a host processor to filter unreliable clusters based on base calls predicted by a neural network running on a configurable processor, and further enables the configurable processor to generate a reliable intermediate representation of the remaining clusters using data that identifies the unreliable clusters. [Figure 7] FIG. 1 illustrates another embodiment of the disclosed data flow logic that enables a host processor to filter unreliable clusters based on base calls predicted by a neural network running on a configurable processor, and further enables the host processor to base call only reliable clusters using data that identifies unreliable clusters. [Figure 8] FIG. 10 illustrates yet another embodiment of the disclosed data flow logic that enables a host processor to filter unreliable clusters based on base calls predicted by a neural network running on a configurable processor, and further uses the data identifying unreliable clusters to generate reliable data for each remaining cluster. [Figure 9] 1 shows the results of a comparative analysis of the detection of empty and non-empty wells using the disclosed technology, referred to herein as "DeepRTA," versus Illumina's conventional base caller, referred to as Real-Time Analysis (RTA) software. [Figure 10]1 shows the results of a comparative analysis of the detection of empty and non-empty wells using the disclosed technology, referred to herein as "DeepRTA," versus Illumina's conventional base caller, referred to as Real-Time Analysis (RTA) software. [Figure 11] 1 shows the results of a comparative analysis of the detection of empty and non-empty wells using the disclosed technology, referred to herein as "DeepRTA," versus Illumina's conventional base caller, referred to as Real-Time Analysis (RTA) software. [Figure 12] 1 shows the results of a comparative analysis of the detection of empty and non-empty wells using the disclosed technology, referred to herein as "DeepRTA," versus Illumina's conventional base caller, referred to as Real-Time Analysis (RTA) software. [Figure 13] 1 shows the results of a comparative analysis of the detection of empty and non-empty wells using the disclosed technology, referred to herein as "DeepRTA," versus Illumina's conventional base caller, referred to as Real-Time Analysis (RTA) software. [Figure 14] A computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE INVENTION
[0011] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0012] The present disclosure provides methods and systems for artificial intelligence-based image analysis that are particularly useful for detecting and filtering unreliable clusters. Figure 1 illustrates an exemplary data analysis and filtering system, along with some of its components. The system includes an image generation system 132, per-cycle cluster data 112, a data provider 102, a neural network-based base caller 104, probability quartiles 106, detection and filtering logic 146, and data identifying unreliable clusters 124. The system may be formed by one or more programmed computers, with programming stored on one or more machine-readable media having code executed to perform one or more steps of the methods described herein. For example, in the illustrated embodiment, the system includes an image generation system 132 configured to output per-cycle cluster data 112 as digital image data, e.g., image data representing individual picture elements or pixels that together form an image of an array or other object.
[0013] (Neural network-based base calling) Base calling is the process of determining the nucleotide composition of a sequence. Base calling involves the analysis of image data, or sequencing images, generated during a sequencing reaction performed by sequencing instruments such as Illumina's iSeq, HiSeqX, HiSeq3000, HiSeq4000, HiSeq2500, NovaSeq6000, NextSeq550, NextSeq1000, NextSeq2000, NextSeqDx, MiSeq, and MiSeqDx. The following description outlines how sequencing images are generated and what they depict, according to one embodiment.
[0014] Base calling decodes the raw signal of the sequencing instrument, i.e., the intensity data extracted from the sequencing image, into a nucleotide sequence. In one embodiment, the Illumina platform employs cyclic reversible termination (CRT) chemistry for base calling. This process relies on extending a nascent strand complementary to the template strand with fluorescently labeled nucleotides while tracking the emission signal of each newly added nucleotide. The fluorescently labeled nucleotides have a 3' removable block that anchors the fluorophore signal of the nucleotide.
[0015] Sequencing is performed in iterative cycles, each of which includes three steps: (a) extending the emerging strand by adding fluorescently labeled nucleotides; (b) exciting fluorophores using one or more lasers in the sequencing instrument's optical system and generating a sequencing image by imaging through different filters in the optical system; and (c) cleaving the fluorophores and removing the 3' block in preparation for the next sequencing cycle. The capture and imaging cycles are repeated for a specified number of sequencing cycles, defining the read length. Using this approach, each cycle locates a new position along the template strand.
[0016] The enormous power of Illumina sequencers stems from their ability to simultaneously perform and detect CRT reactions of millions or even billions of clusters (e.g., clusters). A cluster contains approximately 1,000 identical copies of a template strand, but the size and shape of the clusters vary. Clusters are grown from template strands by bridge amplification or exclusion amplification of the input library prior to a sequencing run. The purpose of amplification and cluster extension is to increase the intensity of the emitted signal, since imaging devices cannot reliably sense the fluorophore signal of a single strand. However, because the physical distance between strands within a cluster is small, imaging devices perceive a cluster of strands as a single spot.
[0017] Sequencing occurs in a flow cell, a small glass slide that holds the input strand. The flow cell is connected to an optical system that includes microscope imaging, an excitation laser, and fluorescence filters. The flow cell contains multiple chambers called lanes. The lanes are physically separated from each other and can contain distinct tagged sequencing libraries without cross-contamination of samples. The sequencing instrument's imaging device (e.g., a solid-state imager such as a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) sensor) takes snapshots at multiple locations along the lane in a series of non-overlapping regions called tiles. For example, Illumina's Genome Analyzer II has 100 tiles per lane, and Illumina's HiSeq2000 has 68 tiles per lane. Tiles hold hundreds of thousands to millions of clusters.
[0018] The output of sequencing is a sequencing image showing the intensity emissions of each cluster and their surrounding background. The sequencing image shows the intensity emissions generated as a result of incorporating nucleotides into a sequence during sequencing. The intensity emissions come from the associated clusters and their surrounding background.
[0019] The description is organized as follows: First, the inputs to the neural network-based base caller 104 are described according to one embodiment. Then, an example of the structure and configuration of the neural network-based base caller 104 is provided. Finally, the output of the neural network-based base caller 104 is described according to one embodiment.
[0020] Further details regarding the neural network-based base caller 104 can be found in U.S. Provisional Patent Application No. 62 / 821,766, entitled "ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV), which is incorporated herein by reference.
[0021] In one embodiment, image patches are extracted from sequencing images. Data provider 102 provides the extracted image patches to neural network-based base caller 104 as "input image data" for base calling. The image patches have dimensions w x h, where w (width) and h (height) are any number ranging from 1 to 10,000 (e.g., 3 x 3, 5 x 5, 7 x 7, 10 x 10, 15 x 15, 25 x 25). In some embodiments, w and h are the same. In other embodiments, w and h are different.
[0022] Sequencing generates m images per sequencing cycle for m corresponding imaging channels. In one embodiment, each image channel corresponds to one of a plurality of filter wavelength bands. In another embodiment, each image channel corresponds to one of a plurality of imaging events in a sequencing cycle. In yet another embodiment, each image channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter.
[0023] To prepare input image data for a particular sequencing cycle, image patches are extracted from each of the m images. In different embodiments, such as for 4-, 2-, and 1-channel chemistries, m is 4 or 2. In other embodiments, m is greater than 1, 3, or 4. The input image data is in the optical pixel domain in some embodiments and in the upsampled sub-pixel domain in other embodiments.
[0024] For example, consider that sequencing uses two different image channels, namely, a red channel and a green channel. Then, in each sequencing cycle, signal determination generates a red image and a green image. In this way, for a series of k sequencing cycles, a sequence having k pairs of red and green images is generated as output.
[0025] The input image data includes an array of per-cycle image patches generated for a series of k sequencing cycles of a sequencing run. The per-cycle image patches include intensity data for associated clusters and their surrounding background in one or more image channels (e.g., red and green channels). In one embodiment, when a single target cluster (e.g., cluster) is base called, the per-cycle image patch is centered on a central pixel that includes intensity data for the target-associated cluster, and pixels other than the central pixel of the per-cycle image patch include intensity data for associated clusters adjacent to the target-associated cluster. The per-cycle image patches for multiple sequencing cycles are stored as per-cycle cluster data 112.
[0026] The input image data includes data from multiple sequencing cycles (e.g., a current sequencing cycle, one or more preceding sequencing cycles, and one or more consecutive sequencing cycles). In one embodiment, the input image data includes data from three sequencing cycles, such that the data from the current (time t) sequencing cycle being base called is accompanied by data from (i) the left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, and (ii) the right adjacent / context / next / consecutive / subsequent (time t+1) sequencing cycle. In another embodiment, the input image data includes data from five sequencing cycles, and the data for the current (time t) sequencing cycle being base called is accompanied by (i) data from the first left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, (ii) data from the second left adjacent / context / previous / preceding / previous (time t-2) sequencing cycle, (iii) data from the first right adjacent / context / next / successive / subsequent (time t+1) sequencing cycle, and (iv) data from the second right adjacent / context / next / successive / subsequent (time t+2) sequencing cycle. In yet another embodiment, the input image data includes data from seven sequencing cycles, and the data for the current (time t) sequencing cycle to be base called includes (i) data from the first left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, (ii) data from the second left adjacent / context / previous / preceding / previous (time t-2) sequencing cycle, (iii) data from the third left adjacent / context / previous / preceding / previous (time t-3) sequencing cycle, (iv) data from the first right adjacent / context / next / contiguous / subsequent (time t+1) sequencing cycle, (v) data from the second right adjacent / context / next / contiguous / subsequent (time t+2) sequencing cycle, and (vi) data from the third right adjacent / context / next / contiguous / subsequent (time t+3) sequencing cycle. In other embodiments, the input image data includes data from a single sequencing cycle.In still other embodiments, the input image data includes 58, 75, 92, 130, 168, 175, 209, 225, 230, 275, 318, 325, 330, 525, or 625 sequencing cycles of data.
[0027] In one embodiment, a sequencing image from a current (time t) sequencing cycle includes sequencing images from the first and second preceding (time t-1, time t-2) sequencing cycles and sequencing images from the first and second subsequent (time t+1, time t+2) sequencing cycles. According to one embodiment, the neural network-based base caller 104 processes the sequencing images through its convolutional layers to generate alternative representations. The alternative representations are then used by an output layer (e.g., a softmax layer) to generate base calls for either the current (time t) sequencing cycle or each of the sequencing cycles (i.e., the current (time t) sequencing cycle, the first and second preceding (time t-1, time t-2) sequencing cycles, and the first and second subsequent (time t+1, time t+2) sequencing cycles). The resulting base calls form sequencing reads.
[0028] In another embodiment, a sequencing image from a current (time t) sequencing cycle is accompanied by a sequencing image from a preceding (time t-1) sequencing cycle and a sequencing image from a subsequent (time t+1) sequencing cycle. According to one embodiment, the neural network-based base caller 104 processes the sequencing images through its convolutional layers to generate alternative representations. The alternative representations are then used by an output layer (e.g., a softmax layer) to generate base calls for the current (time t) sequencing cycle or for each of the sequencing cycles, i.e., the current (time t) sequencing cycle, the preceding (time t-1) sequencing cycle, and the subsequent (time t+1) sequencing cycle. The resulting base calls form sequencing reads.
[0029] In one embodiment, the neural network-based base caller 104 outputs a base call for a single target cluster for a particular sequencing cycle. In another embodiment, the neural network-based base caller 104 outputs a base call for each target cluster within a plurality of target clusters at a particular sequencing cycle. In yet another embodiment, the neural network-based base caller 104 generates a base call sequence for each target cluster by outputting a base call for each target cluster within a plurality of target clusters at each sequencing cycle within the plurality of sequencing cycles.
[0030] In one embodiment, the neural network-based base caller 104 is a multilayer perceptron (MLP). In another embodiment, the neural network-based base caller 104 is a feed-forward neural network. In yet another embodiment, the neural network-based base caller 104 is a fully connected neural network. In a further embodiment, the neural network-based base caller 104 is a fully convolutional neural network. In yet another embodiment, the neural network-based base caller 104 is a semantic segmentation neural network. In yet another further embodiment, the neural network-based base caller 104 is a generative adversarial network (GAN).
[0031] In one embodiment, the neural network-based base caller 104 is a convolutional neural network (CNN) having multiple convolutional layers. In another embodiment, the neural network-based base caller 104 is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), or a gated recurrent unit (GRU). In yet another embodiment, the neural network-based base caller 104 includes both a CNN and an RNN.
[0032] In yet other embodiments, the neural network-based base caller 104 may use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth-wise separable convolution, point-wise convolution, 1×1 convolution, grouped convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The neural network-based base caller 104 may use one or more loss functions, such as logistic regression / logarithmic loss, multiclass cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It can use any parallel, efficient, and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharding, parallel calls to map transforms, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). It can include nonlinear transformation functions, such as upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (e.g., LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, Pejol connections, activation functions (e.g., nonlinear transformation functions such as rectified linear unit (ReLU), leaky ReLU, exponential linear unit (ELU), sigmoid, and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or mean pooling), global mean pooling layers, and attention mechanisms.
[0033] The neural network-based base caller 104 trains using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that the neural network-based base caller 104 may use to train include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that the neural network-based base caller 104 may use to train include Momentum, Nestorv accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.
[0034] The neural network-based base caller 104 uses a dedicated architecture to separate the processing of data for different sequencing cycles. The motivation for using such a dedicated architecture is first explained. As described above, the neural network-based base caller 104 processes intensity-contextualized patches for the current sequencing cycle, one or more preceding sequencing cycles, and one or more following sequencing cycles. Data from additional sequencing cycles provides unique context for each sequence. The neural network-based base caller 104 learns sequence-specific contexts during training and calls them as bases. Furthermore, data from pre- and post-sequencing cycles provide secondary contributions of pre-phasing and phasing signals to the current sequencing cycle.
[0035] However, images captured in different sequencing cycles and in different image channels are misaligned and have residual registration errors with each other. To account for this misalignment, the specialized architecture includes a spatial convolution layer that does not mix information between sequencing cycles, but only mixes information within the same sequencing cycle.
[0036] The spatial convolutional layer uses so-called "decoupled convolutions" that operate on decoupling by processing the data for each of multiple sequencing cycles independently through a "dedicated, non-shared" array of convolutions. Decoupled convolutions convolve on the data and resulting feature maps only within a given sequencing cycle, i.e., the cycle, without convolving on the data and resulting feature maps of any other sequencing cycles.
[0037] For example, suppose the input data includes (i) a current intensity-contextualized patch for the current (time t) sequencing cycle to be base-called, (ii) a previous intensity-contextualized patch for the previous (time t-1) sequencing cycle, and (iii) a next intensity-contextualized patch for the next (time t+1) sequencing cycle. The dedicated architecture then initiates three separate convolution pipelines: a current convolution pipeline, a previous convolution pipeline, and a next convolution pipeline. The current data processing pipeline receives the current intensity-contextualized patch for the current (time t) sequencing cycle as input and processes it independently through multiple spatial convolution layers 784 to generate a so-called "current spatial convolutional representation" as the output of the final spatial convolutional layer. The previous convolution pipeline receives the previous intensity-contextualized patch for the previous (time t-1) sequencing cycle as input and processes it independently through multiple spatial convolutional layers to generate a so-called "previous spatial convolutional representation" as the output of the final spatial convolutional layer. The next convolution pipeline receives as input the next intensity-contextualized patch for the next (time t+1) sequencing cycle and processes it independently through multiple spatial convolution layers to produce the so-called “next spatial convolutional representation” as the output of the final spatial convolutional layer.
[0038] In some implementations, the current, previous, and next convolution pipelines run in parallel. In some implementations, the spatial convolution layer is part of a spatial convolution network (or sub-network) within a dedicated architecture.
[0039] The neural network-based base caller 104 further includes temporal convolutional layers that blend information between sequencing cycles, i.e., between cycles. The temporal convolutional layers receive their inputs from the spatial convolutional networks and operate on the spatial convolutional representations produced by the final spatial convolutional layer for each data processing pipeline.
[0040] The inter-cycle operational freedom of the temporal convolutional layers arises from the fact that misalignment features present in the image data supplied as input to the spatial convolutional network are purged from the spatial convolutional representation by the stack or cascade of separated convolutions performed by the array of spatial convolutional layers.
[0041] The temporal convolutional layer uses so-called "combinatorial convolution," which convolves group-wise on input channels with subsequent inputs on a sliding window basis. In one implementation, the subsequent inputs are subsequent outputs generated by previous spatial or temporal convolutional layers.
[0042] In some embodiments, the temporal convolutional layer is part of a temporal convolutional network (or sub-network) within a dedicated structure. The temporal convolutional network receives its input from a spatial convolutional network. In one embodiment, the first temporal convolutional layer of the temporal convolutional network combines the spatial convolutional representations between sequencing cycles by group. In another embodiment, subsequent temporal convolutional layers of the temporal convolutional network combine successive outputs of previous temporal convolutional layers. The output of the final temporal convolutional layer is fed to an output layer, which generates an output. The output is used to base call one or more clusters in one or more sequencing cycles.
[0043] In one embodiment, bypassing base calls of unreliable clusters refers to processing unreliable clusters only through the spatial convolutional layers of the neural network-based base caller 104, and not processing unreliable clusters through the temporal convolutional layers of the neural network-based base caller 104.
[0044] In the context of the present application, unreliable clusters are also identified by pixels that do not represent any cluster, and such pixels are discarded from processing by the temporal convolutional layers. In some embodiments, this occurs when the well into which the biological sample is deposited is empty.
[0045] Detecting and filtering unreliable clusters The disclosed techniques detect and filter unreliable clusters. The following discussion describes unreliable clusters.
[0046] An unreliable cluster is a low-quality cluster that emits an insignificant amount of the desired signal compared to the background signal. The signal-to-noise ratio of an unreliable cluster is substantially low, e.g., less than 1. In some embodiments, an unreliable cluster may not generate any desired signal at all. In other embodiments, an unreliable cluster may generate only a very small amount of signal compared to the background. In one embodiment, the signal is an optical signal, which is intended to include, for example, fluorescence, luminescence, scattering, or absorption signals. Signal level refers to the amount of detected energy or encoded information having a desired or predetermined characteristic. For example, optical signals can be quantified by one or more of intensity, wavelength, energy, frequency, power, brightness, etc. Other signals can be quantified according to characteristics such as voltage, current, electric field strength, magnetic field strength, frequency, power, temperature, etc. The absence of signal in an unreliable cluster is understood to be a signal level of zero or a signal level that is not significantly distinguishable from noise.
[0047] There are many potential reasons for the insufficient quality of signals in unreliable clusters. If there are polymerase chain reaction (PCR) errors in colony amplification, such that a significant proportion of the approximately 1,000 molecules in an unreliable cluster contain a different base at a specific position, signals for two bases may be observed, which is interpreted as a sign of insufficient quality and is referred to as a phase error. Phase errors occur when individual molecules in an unreliable cluster do not incorporate nucleotides in some cycles, lagging behind other molecules (e.g., due to incomplete removal of the 3' terminator, known as phasing), or when individual molecules incorporate two or more nucleotides in a single cycle (e.g., due to incorporation of nucleotides without an effective 3' block, known as prephasing). This results in a loss of synchronization in the readout of sequence copies. The proportion of sequences in an unreliable cluster affected by phasing and prephasing increases with increasing cycle number, which is the main reason why read quality tends to decrease at higher cycle numbers.
[0048] Unreliable clusters also result from fading, which is the exponential decay in the signal intensity of unreliable clusters as a function of cycle number. As the sequencing operation progresses, strands of unreliable clusters are excessively washed, exposed to laser radiation that creates reactive species, and subjected to harsh environmental conditions. All of this results in the gradual loss of fragments in unreliable clusters, reducing their signal intensity.
[0049] Unreliable clusters can also result from poorly developed colonies, i.e., small cluster sizes of unreliable clusters, which can result in empty or partially filled wells on a patterned flow cell. That is, in some embodiments, unreliable clusters represent empty wells, polyclonal wells, and ambiguous wells on a patterned flow cell. Unreliable clusters can also result from overlapping colonies caused by non-exclusive amplification. Unreliable clusters can also result from insufficient or uneven illumination, for example, due to being located at the edge of the flow cell. Unreliable clusters can also result from impurities on the flow cell that obscure the emitted signal. Unreliable clusters also include polyclonal clusters, which occur when multiple clusters are deposited in the same well.
[0050] This section discusses how unreliable clusters are detected and filtered by the detection and filtering logic 146 to improve base calling accuracy and efficiency. The data provider 102 provides per-cycle cluster data 112 to the neural network-based base caller 104. The per-cycle cluster data 112 is for a plurality of clusters and for a first subset of sequencing cycles of a sequencing operation. For example, consider a sequencing operation having 150 sequencing cycles. The first subset of sequencing cycles can then include any subset of the 150 sequencing cycles, such as the first 5, 10, 15, 25, 35, 40, 50, or 100 sequencing cycles of the 150-cycle sequencing operation. Additionally, each sequencing cycle produces a sequencing image depicting the intensity emissions of clusters within the plurality of clusters. Thus, the cycle-by-cycle cluster data 112 for the plurality of clusters and for the first subset of sequencing cycles of the sequencing operation includes only sequencing images for the first 5, 10, 15, 25, 35, 40, 50, or 100 sequencing cycles of the 150-cycle sequencing operation, and does not include sequencing images for the remaining sequencing cycles of the 150-cycle sequencing operation.
[0051] The neural network-based base caller 104 base calls each cluster among the plurality of clusters in each sequencing cycle within the first subset of sequencing cycles. To do so, the neural network-based base caller 104 processes the per-cycle cluster data 112 and generates an intermediate representation of the per-cycle cluster data 112. The neural network-based base caller 104 then processes the intermediate representation through an output layer to generate per-cluster, per-cycle probability quartiles for each cluster and for each sequencing cycle. Examples of output layers include a softmax function, a log-softmax function, an ensemble output mean function, a multi-layer perceptron uncertainty function, a Bayesian Gaussian distribution function, and a cluster strength function. The per-cluster, per-cycle probability quartiles are stored as probability quartiles 106.
[0052] The following discussion focuses on cluster-wise, cycle-wise probability quartiles, using the softmax function as an example. First, we describe the softmax function, then cluster-wise, cycle-wise probability quartiles.
[0053] The softmax function is a preferred function for multi-class classification. The softmax function calculates the probability of each target class across all possible target classes. The output of the softmax function ranges between zero and one, and the sum of all probabilities equals one. The softmax function calculates the exponent of a given input value and the sum of the exponent values of all input values. The ratio of the exponent of the input value to the sum of the exponent values is the output of the softmax function, referred to herein as "exponential normalization."
[0054] Formally, training a so-called softmax classifier is a regression onto class probabilities rather than a true classifier, since it returns not the classes but rather confidence predictions of the probabilities of each class. The softmax function takes some kind of value and converts them into probabilities that sum to 1. The softmax function compresses any real-valued n-dimensional vector into an n-dimensional vector of real values between 0 and 1. Therefore, using a softmax function guarantees that the output is a valid, exponentially normalized probability mass function (non-negative and sums to 1).
[0055] Intuitively, the softmax function is a "soft" version of the max function. The term "soft" comes from the fact that the softmax function is continuous and differentiable. Instead of selecting a single maximum element, it decomposes the vector into parts of the whole, such that the maximum input element gets a proportionally larger value, while the others get a smaller proportion of the value. This property of outputting a probability distribution makes the softmax function suitable for probabilistic interpretation in classification tasks.
[0056] Consider z as a vector of inputs to a softmax layer. The softmax layer units are the number of nodes in the softmax layer, so the length of the z vector is the number of units in the softmax layer (if you have 10 output units, there will be 10 z elements).
[0057] n-dimensional vector Z=[Z1,Z2,...Z n ], the softmax function uses exponential normalization (exp) to generate another n-dimensional vector p(Z) with normalized values in the range [0,1] whose sum equals 1.
[0058]
number
[0059] Figure 2A shows an example softmax function. The softmax function is
[0060]
number
[0061] A particular cluster-by-cluster, cycle-by-cycle probability quartile identifies the probability of A, C, T, and G being incorporated into a particular cluster in a particular sequencing cycle. If the output layer of the neural network-based base caller 104 uses a softmax function, the probabilities in the cluster-by-cycle probability quartiles are exponentially normalized classification scores that sum to 1. Figure 2B shows exemplary cluster-by-cycle probability quartiles 222 generated by the softmax function for cluster 1 (202, shown in brown) and for sequencing cycles 1 through S (212), respectively. In other words, the first subset of sequencing cycles includes S sequencing cycles.
[0062] The detection and filtering logic 146 identifies unreliable clusters based on generating filter values from the per-cluster, per-cycle probability quartiles. In this application, the per-cluster, per-cycle probability quartiles are also referred to as base call classification scores or normalized base call classification scores or initial base call classification scores or normalized initial base call classification scores or initial base calls.
[0063] The filter calculator 116 determines a filter value for each cluster-by-cluster, cycle-by-cycle probability quartile based on the probability that each cluster-by-cluster, cycle-by-cycle probability quartile identifies, thereby generating an array of filter values 232 for each cluster. The array of filter values 232 is stored as the filter value 126.
[0064] The filter values for the per-cluster, per-cycle probability quartiles are determined based on an arithmetic operation involving one or more of the probabilities. In one implementation, the arithmetic operation used by filter calculator 116 is subtraction. For example, in the implementation shown in FIG. 2B , the filter values for the per-cluster, per-cycle probability quartiles are determined by subtracting the second-highest of the probabilities (shown in blue) from the highest of the probabilities (shown in magenta).
[0065] In another embodiment, the arithmetic operation used by filter calculator 116 is division. For example, the filter value for each cluster, each cycle probability quartile is determined as the ratio of the highest one of the probabilities (shown in magenta) to the second highest one of the probabilities (shown in blue). In yet another embodiment, the arithmetic operation used by filter calculator 116 is addition. In yet a further embodiment, the arithmetic operation used by filter calculator 116 is multiplication.
[0066] In one implementation, filter calculator 116 uses a filtering function to generate filter values 126. In one example, the filtering function is a Chastity filter that defines Chastity as the ratio of the brightest base intensity divided by the sum of the brightest and second brightest base intensities. In another example, the filtering function is at least one of a maximum log-probability function, a least squares error function, a mean signal-to-noise ratio (SNR), and a least absolute error function.
[0067] The unreliable cluster identifier 136 uses the filter value 126 to identify some clusters within the plurality of clusters as unreliable clusters 124. Data identifying the unreliable clusters 124 may be in a computer-readable format or medium. The unreliable clusters may be identified by an instrument ID, a run number on the instrument, a flow cell ID, a lane number, a tile number, an X coordinate of the cluster, a Y coordinate of the cluster, and a unique molecular identifier (UMI). The unreliable cluster identifier 136 identifies as an unreliable cluster 124 a cluster within the plurality of clusters whose sequence of filter values includes "N" filter values below a threshold "M." In one embodiment, "N" ranges from 1 to 5. In another embodiment, "M" ranges from 0.5 to 0.99.
[0068] FIG. 3 illustrates an example of identifying an unreliable cluster 124 using filter values 126. In FIG. 3, the threshold "M" is 0.5, and the number of filter values "N" is 2. FIG. 3 illustrates three filter value arrays 302, 312, and 322 for three clusters 1, 2, and 3, respectively. In the first array 302 for cluster 1, there are two filter values (shown in purple) that are less than M, i.e., N=2; therefore, cluster 1 is identified as an unreliable cluster. In the second array 312 for cluster 2, there are three filter values (shown in pink) that are less than M, i.e., N=3; therefore, cluster 2 is identified as an unreliable cluster. In the third array 322 for cluster 3, there is only one filter value (shown in green) that is less than M, i.e., N=1; therefore, cluster 3 is identified as a reliable cluster.
[0069] Here, we discuss bypass logic 142 implemented by data provider 102. Bypass logic 142 bypasses base calling of unreliable clusters (e.g., clusters 1 and 2) for the remainder of the sequencing cycles of a sequencing operation, thereby making base calls only for clusters among a plurality of clusters that are not identified as unreliable clusters for the remainder of the sequencing cycles. For example, assume that a first subset of sequencing cycles of a sequencing operation includes 25 sequencing cycles, and the sequencing operation has a total of 100 sequencing cycles. Then, after the first 25 sequencing cycles, clusters 1, 2, and 3 each have a respective sequence of 25 filter values based on the above filtering function.
[0070] The remainder of the sequencing cycles then includes the last 75 cycles of the 100-cycle sequencing operation. Then, after the first 25 sequencing cycles and before the 26th sequencing cycle, the unreliable cluster identifier 136 determines which of clusters 1, 2, and 3 is an unreliable cluster based on the sequence of each of the 25 filter values. Then, for the remaining sequencing cycles, i.e., the last 75 cycles of the 100-cycle sequencing operation, the bypass logic 142 does not base call (i.e., stops base calling) clusters identified as unreliable by the unreliable cluster identifier 136 (e.g., clusters 1 and 2), and continues to base call only clusters not identified as unreliable by the unreliable cluster identifier 136 (e.g., cluster 3). In other words, unreliable clusters are base called only for cycles 1 through 25 of the sequencing operation and not for cycles 26 through 100 of the sequencing operation, while reliable clusters are base called for all cycles 1 through 100 of the sequencing operation.
[0071] The term filtering, when used with respect to clusters and base calls, refers to discarding or ignoring clusters as data points. Thus, any clusters of insufficient intensity or quality can be filtered and not included in the output data set. In some embodiments, filtering of low-quality clusters occurs at one or more distinct points during the sequencing operation. In some embodiments, filtering occurs during template generation. Alternatively, or additionally, in some embodiments, filtering occurs after a predefined cycle. In certain embodiments, filtering occurs at or after cycle 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or after cycle 30. In some embodiments, filtering occurs at cycle 25, thereby filtering unreliable clusters based on the sequence of filter values determined for the first 25 cycles.
[0072] Figure 4 is a flow diagram illustrating one embodiment of a method for identifying unreliable clusters to improve base calling accuracy and efficiency. Various processes and steps of the methods described herein can be performed using a computer. The computer can include a processor that is part of a detection device, networked with a detection device used to acquire data processed by the computer, or separate from the detection device. In some embodiments, information (e.g., image data) can be transmitted between components of the systems disclosed herein directly or via a computer network. A local area network (LAN) or wide area network (WAN) can be an enterprise computing network, including access to the Internet, to which computers and computing devices comprising the system are connected. In one embodiment, the LAN conforms to the Transmission Control Protocol / Internet Protocol (TCP / IP) industry standard. In some cases, information (e.g., image data) is entered into the systems disclosed herein via an input device (e.g., a disk drive, compact disc player, USB port, etc.). In some cases, information is received by loading the information from a storage device, such as a disk or flash drive.
[0073] A processor used to execute the algorithms or other processes described herein may include a microprocessor. The microprocessor may be any conventional general-purpose single-chip or multi-chip microprocessor, such as a Pentium™ processor manufactured by Intel Corporation. A particularly useful computer may utilize an Intel Ivybridge dual-12 core processor, an LSI RAID controller, with 128 GB of RAM and a 2 TB solid-state disk drive. Additionally, the processor may include any conventional special-purpose processor, such as a digital signal processor or a graphics processor. Processors typically have conventional address lines, conventional data lines, and one or more conventional control lines.
[0074] The embodiments disclosed herein may be implemented as a method, apparatus, system, or article using standard programming or engineering techniques to generate software, firmware, hardware, or any combination thereof. As used herein, the term "article of manufacture" refers to code or logic implemented in hardware or computer-readable media, such as optical storage devices, as well as volatile or non-volatile memory devices. Such hardware may include, but is not limited to, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microprocessors, or other similar processing devices. In certain embodiments, the information or algorithms described herein reside in non-transitory storage media.
[0075] In certain embodiments, the computer-implemented methods described herein can be performed in real time while multiple images of an object are being acquired. Such real-time analysis is particularly useful for nucleic acid sequencing applications, where nucleic acid sequences are subjected to repeated cycles of fluidization and detection steps. While analysis of sequencing data can often be beneficial to perform the methods described herein in real time or in the background, it can also be beneficial to perform the methods described herein while other data collection or analysis algorithms are in process. Examples of real-time analysis methods that can be used in the present methods are those commercially available from Illumina, Inc. (San Diego, Calif.) and / or used in the MiSeq and HiSeq sequencing instruments described in U.S. Patent Application Publication No. 2012 / 0020537 A1, which is incorporated herein by reference.
[0076] At action 402, the method includes accessing cycle-by-cycle cluster data for a plurality of clusters and for a first subset of sequencing cycles of the sequencing operation.
[0077] In action 412, the method includes base calling each cluster among the plurality of clusters in each sequencing cycle in the first subset of sequencing cycles.
[0078] In action 422, the method includes processing the per-cycle cluster data to generate an intermediate representation of the per-cycle cluster data.
[0079] In action 432, the method includes processing the intermediate representation through an output layer to generate per-cluster, per-cycle probability quartiles for each cluster and for each sequencing cycle, where a particular per-cluster, per-cycle probability quartile identifies the probability of A, C, T, and G being incorporated into a particular cluster in a particular sequencing cycle.
[0080] In action 442, the method includes generating an array of filter values for each cluster by determining filter values for the cluster-by-cluster, cycle-by-cycle probability quartiles based on the probabilities that the cluster-by-cluster, cycle-by-cycle probability quartiles identify.
[0081] At action 452, the method includes identifying, among the plurality of clusters, a cluster whose sequence of filter values includes at least "N" filter values below a threshold "M" as an unreliable cluster.
[0082] In action 462, the method includes base calling only clusters of the plurality of clusters that are not identified as unreliable clusters in the remainder of the sequencing cycle of the sequencing operation by bypassing base calling of unreliable clusters in the remainder of the sequencing cycle.
[0083] Sequencing System 5A and 5B show one embodiment of a sequencing system 500A. The sequencing system 500A includes a configurable processor 546. The configurable processor 546 implements the base calling techniques disclosed herein. A sequencing system is also referred to as a "sequencer."
[0084] The sequencing system 500A can operate to obtain any information or data relating to at least one of biological or chemical substances. In some embodiments, the sequencing system 500A is a workstation, which can be similar to a benchtop device or desktop computer. For example, most (or all) of the systems and components for performing the desired reactions can be within a common housing 502.
[0085] In certain embodiments, the sequencing system 500A is a nucleic acid sequencing system configured for various applications, including, but not limited to, de novo sequencing, whole genome or targeted genomic region resequencing, and metagenomics. The sequencer may also be used for DNA or RNA analysis. In some embodiments, the sequencing system 500A may also be configured to generate reaction sites within a biosensor. For example, the sequencing system 500A may be configured to receive a sample and generate surface-bound clusters of clonally amplified nucleic acids from the sample. Each cluster may constitute or be part of a reaction site within a biosensor.
[0086] The exemplary sequencing system 500A may include a system receptacle or interface 510 configured to interact with a biosensor 512 to effect a desired reaction within the biosensor 512. In the following description with respect to FIG. 5A , the biosensor 512 is loaded into the system receptacle 510. However, it is understood that a cartridge containing the biosensor 512 may be inserted into the system receptacle 510, and that in some conditions the cartridge may be temporarily or permanently removed. As discussed above, the cartridge may include, among other things, fluid control and fluid storage components.
[0087] In certain embodiments, the sequencing system 500A is configured to perform multiple parallel reactions within the biosensor 512. The biosensor 512 includes one or more reaction sites where desired reactions can occur. The reaction sites may be immobilized, for example, on a solid surface of the biosensor or on beads (or other movable substrates) located within corresponding reaction chambers of the biosensor. The reaction sites may include, for example, clusters of clonally amplified nucleic acids. The biosensor 512 may include a solid-state imaging device (e.g., a CCD or CMOS imager) and a flow cell attached thereto. The flow cell may include one or more flow paths that receive solutions from the sequencing system 500A and direct the solutions toward the reaction sites. Optionally, the biosensor 512 may be configured to engage a thermal element for transferring thermal energy into and out of the flow paths.
[0088] Sequencing system 500A may include various components, assemblies, and systems (or subsystems) that interact with each other to perform a predetermined method or assay protocol for biological or chemical analysis. For example, sequencing system 500A includes a system controller 506 that can communicate with the various components, assemblies, and subsystems of sequencing system 500A, as well as a biosensor 512. For example, in addition to system receptacle 510, sequencing system 500A may also include a fluid control system 508 for controlling fluid flow throughout the fluidic network of sequencing system 500A and biosensor 512, a fluid reservoir system 514 configured to hold any fluids (e.g., gases or liquids) that may be used by the bioassay system, a temperature control system 504 that can regulate the temperature of the fluids in the fluidic network, fluid reservoir system 514, and / or biosensor 512, and an illumination system 516 configured to illuminate biosensor 512. As mentioned above, when a cartridge having a biosensor 512 is loaded into the system receptacle 510, the cartridge may also include fluid control and fluid storage components.
[0089] The sequencing system 500A may also include a user interface 518 for interacting with a user. For example, the user interface 518 may include a display 520 for displaying or requesting information from the user and a user input device 522 for receiving user input. In some embodiments, the display 520 and the user input device 522 are the same device. For example, the user interface 518 may include a touch-sensitive display configured to detect the presence of an individual touch and identify the location of the touch on the display. However, other user input devices 522, such as a mouse, touchpad, keyboard, keypad, handheld scanner, voice recognition system, motion recognition system, etc., may also be used. As described in more detail below, the sequencing system 500A may communicate with various components, including a biosensor 512 (e.g., in the form of a cartridge), to perform desired reactions. The sequencing system 500A may also be configured to analyze data obtained from the biosensor to provide desired information to the user.
[0090] The system controller 506 may include a microcontroller, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a coarse-grained reconfigurable architecture (CGRA), a logic circuit, and any other circuit or processor capable of performing the functions described herein. The above examples are merely illustrative and, thus, are not intended to limit the definition and / or meaning of the term system controller. In an exemplary embodiment, the system controller 506 executes a set of instructions stored in one or more storage elements, memories, or modules for at least one of acquiring and analyzing detection data. The detection data may include multiple sequences of pixel signals, such that sequences of pixel signals from each of millions of sensors (or pixels) can be detected over many base call cycles. The storage elements may be in the form of information sources or physical memory elements within the sequencing system 500A.
[0091] The instruction set may include various commands that instruct the sequencing system 500A or biosensor 512 to perform specific operations, such as the methods and processes of various embodiments described herein. The instruction set may be in the form of a software program, which may form part of a tangible, non-transitory computer-readable medium or media. As used herein, the terms "software" and "firmware" are used interchangeably and include any computer program stored in memory executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are exemplary only and thus not limiting of the types of memory that may be used to store a computer program.
[0092] The software may be in various forms, such as system software or application software. Furthermore, the software may be in the form of a collection of separate programs, or a program module or portion of a program module within a larger program. The software may also include modular programming in the form of object-oriented programming. After acquiring the detection data, the detection data may be processed automatically by the sequencing system 500A in response to user input, or in response to a request made by another processing machine (e.g., a remote request via a communications link). In the illustrated embodiment, the system controller 506 includes an analysis module 544. In other embodiments, the system controller 506 does not include the analysis module 544, but instead has access to the analysis module 544 (e.g., the analysis module 544 may be separately hosted on the cloud).
[0093] The system controller 506 may be connected to the biosensor 512 and other components of the sequencing system 500A via a communication link. The system controller 506 may also be communicatively connected to an off-site system or server. The communication link may be a wire, a cord, or wireless. The system controller 506 may receive user input or commands from a user interface 518 and a user input device 522.
[0094] The fluid control system 508 includes a fluid network and is configured to regulate the flow of one or more fluids through the fluid network. The fluid network may be in fluid communication with the biosensor 512 and the fluid reservoir system 514. For example, a selected fluid may be drawn from the fluid reservoir system 514 and directed to the biosensor 512 in a controlled manner, or fluid may be drawn from the biosensor 512 and directed to, for example, a waste reservoir within the fluid reservoir system 514. Although not shown, the fluid control system 508 may include a flow sensor that detects the flow rate or pressure of the fluid within the fluid network. The sensor may be in communication with the system controller 506.
[0095] The temperature control system 504 is configured to regulate the temperature of fluids in different regions of the fluid network, the fluid reservoir system 514, and / or the biosensor 512. For example, the temperature control system 504 may include a thermal circulator that interacts with the biosensor 512 and controls the temperature of fluids flowing along reaction sites within the biosensor 512. The temperature control system 504 may also regulate the temperature of solid elements or components of the sequencing system 500A or the biosensor 512. Although not shown, the temperature control system 504 may include sensors for detecting the temperature of the fluids or other components. The sensors may be in communication with the system controller 506.
[0096] The fluid storage system 514 is in fluid communication with the biosensor 512 and may store various reaction components or reactants used to carry out a desired reaction. The fluid storage system 514 may also store fluids for washing or cleaning the fluidic network and the biosensor 512 and for diluting reactants. For example, the fluid storage system 514 may include various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, etc. Additionally, the fluid storage system 514 may also include a waste reservoir for receiving waste from the biosensor 512. In embodiments that include a cartridge, the cartridge may include one or more of a fluid storage system, a fluid control system, or a temperature control system. Accordingly, one or more of the components described herein for these systems may be contained within the cartridge housing. For example, the cartridge may have various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, waste, etc. Thus, one or more of the fluid reservoir system, fluid control system, or temperature control system may be removably engaged with the bioassay system via a cartridge or other biosensor.
[0097] The illumination system 516 may include a light source (e.g., one or more light-emitting diodes (LEDs)) and multiple optical components for illuminating the biosensor. Examples of light sources may include lasers, arc lamps, LEDs, or laser diodes. The optical components may be, for example, reflectors, polarizers, beam splitters, collimators, lenses, filters, wedges, prisms, mirrors, detectors, etc. In embodiments using an illumination system, the illumination system 516 may be configured to direct excitation light to the reaction sites. As an example, a fluorophore may be excited by a green wavelength of light, so the wavelength of the excitation light may be approximately 532 nm. In one embodiment, the illumination system 516 is configured to generate illumination parallel to a surface normal of the surface of the biosensor 512. In another embodiment, the illumination system 516 is configured to generate illumination that is off-angled relative to the surface normal of the surface of the biosensor 512. In yet another embodiment, the illumination system 516 is configured to generate illumination having multiple angles, including some parallel illumination and some off-angle illumination.
[0098] The system receptacle or interface 510 is configured to engage the biosensor 512 in at least one of mechanical, electrical, and fluidic ways. The system receptacle 510 may hold the biosensor 512 in a desired orientation to facilitate fluid flow through the biosensor 512. The system receptacle 510 may also include electrical contacts configured to engage the biosensor 512, thereby allowing the sequencing system 500A to communicate with and / or provide power to the biosensor 512. Additionally, the system receptacle 510 may include a fluid port (e.g., a nozzle) configured to engage the biosensor 512. In some embodiments, the biosensor 512 is removably coupled to the system receptacle 510 in mechanical, electrical, and / or fluidic ways.
[0099] Additionally, the sequencing system 500A may communicate remotely with other systems or networks, or with other bioassay systems 500A. Detection data obtained by the bioassay system(s) 500A may be stored in a remote database.
[0100] FIG. 5B is a block diagram of a system controller 506 that may be used in the system of FIG. 5A. In one embodiment, the system controller 506 includes one or more processors or modules that may communicate with each other. Each of the processors or modules may include an algorithm (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or sub-algorithm for performing a particular process. While the system controller 506 is conceptually illustrated as a collection of modules, it may also be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 506 may be implemented using an off-the-shelf PC with a single processor or multiple processors, with functional operations distributed among the processors. As a further option, the modules described below may be implemented using a hybrid configuration in which certain modular functions are implemented using dedicated hardware, while remaining modular functions are implemented using an off-the-shelf PC, etc. The modules may also be implemented as software modules within a processing unit.
[0101] During operation, the communication port 550 may send information (e.g., commands) to or receive information (e.g., data) from the biosensor 512 (FIG. 5A) and / or the subsystems 508, 514, 504 (FIG. 5A). In embodiments, the communication port 550 may output multiple arrays of pixel signals. The communication link 534 may receive user input from the user interface 518 (FIG. 5A) and send data or information to the user interface 518. Data from the biosensor 512 or the subsystems 508, 514, 504 may be processed in real time by the system controller 506 during a bioassay session. Additionally or alternatively, the data may be temporarily stored in system memory during a bioassay session and processed slower than in real time or for offline operation.
[0102] As shown in FIG. 5B, the system controller 506 may include multiple modules 526-548 in communication with a main control module 524 along with a central processing unit (CPU) 552. The main control module 524 may be in communication with a user interface 518 (FIG. 5A). While the modules 526-548 are shown in direct communication with the main control module 524, the modules 526-548 may also be in direct communication with each other, the user interface 518, and the biosensor 512. The modules 526-548 may also be in communication with the main control module 524 through other modules.
[0103] The plurality of modules 526-548 includes system modules 528-532, 526 that communicate with subsystems 508, 514, 504, and 516, respectively. The fluid control module 528 may communicate with the fluid control system 508 to control valves and flow sensors in the fluid network to control the flow of one or more fluids through the fluid network. The fluid storage module 530 may notify a user when fluid is low or when a waste reservoir is at or near full capacity. The fluid storage module 530 may also communicate with a temperature control module 532 so that fluids can be stored at a desired temperature. The illumination module 526 may communicate with the illumination system 516 to illuminate reaction sites at specified times during a protocol, such as after a desired reaction (e.g., a binding event) has occurred. In some embodiments, the illumination module 526 may communicate with the illumination system 516 to illuminate the reaction sites at a specified angle.
[0104] The plurality of modules 526-548 may also include a device module 536 that communicates with the biosensor 512 and an identification module 538 that determines identification information associated with the biosensor 512. The device module 536 may, for example, communicate with the system receptacle 510 to confirm that the biosensor has established electrical and fluidic connection with the sequencing system 500A. The identification module 538 may receive a signal that identifies the biosensor 512. The identification module 538 may use the identification information of the biosensor 512 to provide other information to the user. For example, the identification module 538 may determine and subsequently display the lot number, manufacturing date, or recommended protocol for operating the biosensor 512.
[0105] The plurality of modules 526-548 also includes an analysis module 544 (also referred to as a signal processing module or signal processor) that receives and analyzes signal data (e.g., image data) from the biosensor 512. The analysis module 544 includes memory (e.g., RAM or flash) for storing the detection / image data. The detection data can include multiple sequences of pixel signals, such that sequences of pixel signals from each of millions of sensors (or pixels) can be detected over many base call cycles. The signal data can be stored for subsequent analysis or transmitted to the user interface 518 to display desired information to a user. In some embodiments, the signal data can be processed by a solid-state image sensor (e.g., a CMOS image sensor) before the analysis module 544 receives the signal data.
[0106] Analysis module 544 is configured to acquire image data from the photodetector in each of a plurality of sequencing cycles. The image data is derived from the luminescence signals detected by the photodetector and processes the image data for each of the plurality of sequencing cycles via neural network-based base caller 104 to generate base calls for at least some of the analytes in each of the plurality of sequencing cycles. The photodetector may be part of one or more overhead cameras (e.g., a CCD camera in an Illumina GAIIx that takes images of the clusters on biosensor 512 from above) or may be part of biosensor 512 itself (e.g., a CMOS image sensor in an Illumina iSeq that is below the clusters on biosensor 512 and takes images of the clusters from the bottom).
[0107] The output of the photodetectors is a sequencing image showing the intensity emissions of each cluster and their surrounding background. The sequencing image shows the intensity emissions generated as a result of incorporating nucleotides into a sequence during sequencing. The intensity emissions are from the associated analytes and their surrounding background. The sequencing image is stored in memory 548.
[0108] Protocol modules 540 and 542 communicate with main control module 524 to control the operation of subsystems 508, 514, and 504 in implementing a predetermined assay protocol. Protocol modules 540 and 542 may include instruction sets for instructing sequencing system 500A to perform specific operations according to a predetermined protocol. As shown, the protocol module may be a sequencing-by-synthesis (SBS) module 540 configured to issue various commands to execute a sequencing-by-synthesis process. In SBS, the extension of nucleic acid primers along a nucleic acid template is monitored to determine the sequence of nucleotides in the template. The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase enzyme) or ligation (e.g., catalyzed by a ligase enzyme). In certain polymer-based SBS embodiments, fluorescently labeled nucleotides are added to primers (thereby extending the primers) in a template-dependent manner, such that detection of the order and type of nucleotides added to the primers can be used to determine the sequence of the template. For example, to initiate the first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc. can be delivered into / through a flow cell containing an array of nucleic acid templates. The nucleic acid templates may be located at corresponding reaction sites. Primer extension can detect incorporated labeled nucleotides through an imaging event, and these reaction sites can be detected. During the imaging event, an illumination system 516 can provide excitation light to the reaction sites. Optionally, the nucleotides can further include a reversible termination feature that terminates further primer extension once the nucleotide is added to the primer. For example, a nucleotide analog with a reversible terminator moiety can be added to the primer to prevent further extension until a deblocking agent is delivered to remove the moiety. Thus, in another embodiment using reversible termination, a command can be given to deliver a deblocking reagent to the flow cell (either before or after detection).One or more commands can be given to provide wash(s) between the various delivery steps.Then, by repeating the cycle n times to extend the primer by n nucleotides, a sequence of length n can be detected.Exemplary sequencing techniques are described, for example, in Bentley et al., Nature 456:53-59 (2005), WO 04 / 015497, U.S. Patent No. 7,057,026, WO 91 / 06675, U.S. Patent No. 07 / 123744, U.S. Patent Nos. 7,329,492, 7,211,414, 7,315,019, 7,405,251, and 2005 / 014705052, each of which is incorporated herein by reference.
[0109] In the nucleotide delivery step of the SBS cycle, any single type of nucleotide can be delivered at a time, or multiple different nucleotide types (e.g., A, C, T, and G) can be delivered. In nucleotide delivery configurations where only a single type of nucleotide is present at a time, different nucleotides do not need to have distinct labels because they can be distinguished based on the temporal separation inherent in individualized delivery. Thus, a sequencing method or device can use single-color detection. For example, the excitation source only needs to provide excitation at a single wavelength or a single wavelength range. In nucleotide delivery configurations where delivery results in multiple different nucleotides being present in the flow cell at a given time, the sites at which different nucleotide types incorporate can be distinguished based on the different fluorescent labels attached to each nucleotide type in the mixture. For example, four different nucleotides, each bearing one of four different fluorophores, can be used. In one embodiment, four different fluorophores can be distinguished using excitation in four different regions of the spectrum. For example, four different excitation radiation sources can be used. Alternatively, fewer than four different excitation sources can be used, but optical filtering of the excitation radiation from a single source can be used to generate different excitation radiation ranges in the flow cell.
[0110] In some embodiments, fewer than four different colors can be detected in a mixture having four different nucleotides. For example, pairs of nucleotides can be detected at the same wavelength but can be distinguished based on differences in intensity for one member of the pair, or based on a change to one member of the pair (e.g., through chemical modification, photochemical modification, or physical modification) that causes a distinct signal to appear or disappear compared to the signal detected for the other member of the pair. Exemplary devices and methods for distinguishing four different nucleotides using detection of fewer than four colors are described, for example, in U.S. Patent Application Nos. 61 / 535,294 and 61 / 619,575, which are incorporated herein by reference in their entireties. U.S. Patent Application No. 13 / 624,200, filed September 21, 2012, is incorporated by reference in its entirety.
[0111] The multiple protocol modules may also include a sample preparation (or generation) module 542 configured to issue commands to the fluidic control system 508 and the temperature control system 504 to amplify the product in the biosensor 512. For example, the biosensor 512 may be coupled to a sequencing system 500A. The amplification module 542 can issue instructions to the fluidic control system 508 to deliver the necessary amplification components to a reaction chamber in the biosensor 512. In other embodiments, the reaction site may already contain some components for amplification, such as template DNA and / or primers. After delivering the amplification components to the reaction chamber, the amplification module 542 can instruct the temperature control system 504 to cycle through different temperature steps according to a known amplification protocol. In some embodiments, amplification and / or nucleotide incorporation is performed isothermally.
[0112] The SBS module 540 can issue commands to perform bridge PCR, in which clusters of clonal amplicons are formed over localized regions within the flow cell channel. After generating amplicons via bridge PCR, the amplicons may be "linearized" to create single-stranded template DNA, and sstDNA and sequencing primers may be hybridized to universal sequences flanking the region of interest. For example, reversible terminator-based sequencing by synthesis methods can be used, as described above or as follows.
[0113] Each base calling or sequencing cycle can extend the sstDNA by a single base, which can be achieved, for example, by using a modified DNA polymerase and a mixture of four types of nucleotides. Different types of nucleotides can have unique fluorescent labels, and each nucleotide can further have a reversible terminator that allows only a single base to be incorporated in each cycle. After a single base is added to the sstDNA, excitation light can be incident on the reaction site and fluorescent emission can be detected. After detection, the fluorescent label and terminator can be chemically cleaved from the sstDNA. Another similar base calling or sequencing cycle can be as follows: In such a sequencing protocol, the SBS module 540 can instruct the fluid control system 508 to direct the flow of reagents and enzyme solutions through the biosensor 512. Exemplary reversible terminator-based SBS methods that can be utilized with the devices and methods described herein are described in U.S. Patent Application Publication Nos. 2007 / 0166705(A1), 2006 / 0156705(A1), and 2007 / 0156705(A1). *3901(A1), U.S. Patent No. 7,057,026, U.S. Patent Application Publication Nos. 2006 / 0240439(A1), 2006 / 02514714709(A1), WO 05 / 065514, U.S. Patent Application Publication Nos. 2005 / 014700900(A1), WO 06 / 05B199 and WO 07 / 01470251, each of which is incorporated herein by reference in its entirety. Exemplary reagents for reversible terminator-based SBS are described in U.S. Pat. Nos. 7,541,444, 7,057,026, 7,414,14716, 7,427,673, 7,566,537, 7,592,435, and WO 07 / 14535365, each of which is incorporated herein by reference.
[0114] In some embodiments, the amplification and SBS modules may operate in a single assay protocol, eg, template nucleic acids are amplified and subsequently sequenced within the same cartridge.
[0115] The sequencing system 500A may also allow the user to reconfigure the assay protocol. For example, the sequencing system 500A may provide the user with options through the user interface 518 to modify the determined protocol. For example, if it is determined that the biosensor 512 will be used for amplification, the sequencing system 500A may request the temperature of the annealing cycle. Furthermore, the sequencing system 500A may issue a warning to the user if the user provides user input that is not generally accepted for the selected assay protocol.
[0116] In an embodiment, biosensor 512 includes millions of sensors (or pixels), each of which generates a sequence of pixel signals over successive base call cycles. Analysis module 544 detects the sequences of pixel signals and attributes them to corresponding sensors (or pixels) according to the row-wise and / or column-wise location of the sensors on the array of sensors.
[0117] Configurable Processor FIG. 5C is a simplified block diagram of a system for analyzing sensor data, such as base call sensor output, from a sequencing system 500A. In the example of FIG. 5C, the system includes a configurable processor 546. The configurable processor 546 can execute a base caller (e.g., neural network-based base caller 104) in coordination with a runtime program executed by a central processing unit (CPU) 552 (i.e., a host processor). The sequencing system 500A includes a biosensor 512 and a flow cell. The flow cell can include one or more tiles in which clusters of genetic material are exposed to a series of analyte flows that are used to trigger reactions within the clusters to identify bases in the genetic material. A sensor senses the reaction for each cycle of sequencing in each tile of the flow cell to provide tile data. Genetic sequencing is a data-intensive operation that converts base call sensor data into a sequence of base calls for each group of genetic material sensed during the base calling operation.
[0118] The system in this example includes a CPU 552 that executes a runtime program for coordinating base calling operations, and memory 548B for storing the sequences of arrays of tile data, base call reads generated by the base calling operations, and other information used in the base calling operations. In this figure, the system also includes memory 548A that stores configuration file(s), e.g., FPGA bit files, and neural network model parameters used to configure and reconfigure configurable processor 546 and to run the neural network. Sequencing system 500A may include a program for configuring the configurable processor, and in some embodiments, may include a reconfigurable processor that runs the neural network.
[0119] Sequencing system 500A is coupled to configurable processor 546 by bus 589. Bus 589 may be implemented using a high-throughput technology, for example, in one embodiment, bus technology compatible with the PCIe (Peripheral Component Interconnect Express) standard currently maintained and developed by the PCI-SIG (Peripheral Component Interconnect Special Interest Group) standard. Also in this embodiment, memory 548A is coupled to configurable processor 546 by bus 593. Memory 548A may be on-board memory disposed on a circuit board with configurable processor 546. Memory 548A is used for fast access by configurable processor 546 of working data used in base calling operations. Bus 593 may also be implemented using a high-throughput technology, such as bus technology compatible with the PCIe standard.
[0120] Configurable processors, including field programmable gate arrays (FPGAs), coarse-grained configurable reconfigurable arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than can be achieved using general-purpose processors running computer programs. Configuring a configurable processor involves compiling a functional description to generate a configuration file, sometimes referred to as a bitstream or bitfile, and distributing the configuration file to configurable elements on the processor. The configuration file configures the circuit to set dataflow patterns, including the use of distributed memory and other on-chip memory resources, lookup table contents, the operation of configurable logic blocks, and configurable execution units such as configurable interconnects and other elements of the configurable array. A configuration file is reconfigurable if it can be changed in the field by modifying a loaded configuration file. For example, the configuration file may be stored in a volatile SRAM element, a non-volatile read-write memory element, or distributed among an array of configurable elements on a configurable or reconfigurable processor. Various commercially available configurable processors are suitable for use in basecall operations as described herein.Examples include Google's Tensor Processing Unit (TPU)™, GX4 Rackmount Series™, GX9 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ (Snapdragon processors™), NVIDIA Volta™, NVIDIA's Drive PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, Arm DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, Xilinx Alveo™ U200, Xilinx Alveo™ U250, Xilinx Alveo™ U280, Intel / Altera Stratix™ GX2800, Intel / Altera Stratix™ GX2800, and Intel Stratix™ GX10M. In some embodiments, the host CPU may be implemented on the same integrated circuit as the configurable processor.
[0121] The embodiments described herein implement the neural network-based basis caller 104 using a configurable processor 546. The configuration file for the configurable processor 546 may be implemented by specifying the logic functions to be performed using a high-level description language HDL or a register-transfer level RTL language specification. This specification can be compiled using resources designed for a selected configurable processor to generate the configuration file. The same or similar specifications can be compiled to generate a design for an application-specific integrated circuit, which may not be a configurable processor.
[0122] Accordingly, alternatives to the configurable processor 546 in all embodiments described herein include a configured processor including an application specific ASIC or dedicated integrated circuit or set of integrated circuits, or a system-on-chip SOC device, or a system-on-chip SOC device, or a graphics processing unit (GPU) processor or a Coarse-Grained Reconfigurable Architecture (CGRA) processor configured to perform neural network-based base call operations as described herein.
[0123] In general, the configurable and configured processors described herein that are configured to perform neural network operations are referred to herein as neural network processors.
[0124] Configurable processor 546 is configured to perform base call functions by a configuration file or other source loaded using a program executed by CPU 552, which in this example configures an array of configurable elements 591 (e.g., Configuration Logic Blocks (CLBs), such as Look Up Tables (LUTs), flip-flops, arithmetic processing units (PMUs), and compute memory units (CMUs), configurable I / O blocks, programmable interconnect) on the configurable processor. In this example, the configuration includes data flow logic 597 coupled to buses 589 and 593 that performs functions to distribute data and control parameters among elements used in base calling operations.
[0125] Configurable processor 546 is also configured with data flow logic 597 to execute neural network-based base caller 104. Logic 597 includes multi-cycle execution clusters (e.g., 579), which in this example include execution cluster 1 through execution cluster X. The number of multi-cycle execution clusters may be selected according to tradeoffs involving the desired throughput of operation and available resources on configurable processor 546.
[0126] The multi-cycle execution clusters are coupled to data flow logic 597 by data flow paths 599 implemented using configurable interconnect and memory resources on configurable processor 546. The multi-cycle execution clusters are also coupled to data flow logic 597 by control paths 595 implemented using configurable interconnect and memory resources, for example, on configurable processor 546, to provide control signals indicating available execution clusters, provisions for providing input units to the available execution clusters for execution of operations of neural network-based base caller 104, provisions for providing learned parameters of neural network-based base caller 104, provisions for providing output patches of base call classification data, and other control data used in the execution of neural network-based base caller 104.
[0127] The configurable processor 546 is configured to execute the operation of the neural network-based base caller 104 using the learned parameters to generate classification data for the detection cycles of the base calling operation. The operation of the neural network-based base caller 104 is executed to generate classification data for the subject detection cycles of the base calling operation. The operation of the neural network-based base caller 104 operates on an array including a number N of arrays of tile data from each detection cycle of the N detection cycles, where the N detection cycles provide sensor data for different base calling operations for one base position per operation in the time sequence in the examples described herein. Optionally, some of the N detection cycles can be removed from the array as needed according to the particular neural network model being implemented. The number N can be any number greater than 1. In some examples described herein, the detection cycles of the N detection cycles represent a set of detection cycles for at least one detection cycle preceding the subject detection cycle and at least one detection cycle following the subject detection cycle. Examples described herein include an integer number N of 5 or greater.
[0128] Data flow logic 597 is configured to use an input unit for a given operation that includes tile data for an array of N spatially aligned patches to move the tile data and at least some of the learned parameters of the model parameters from memory 548A to configurable processor 546 for operation of neural network-based base caller 104. The input unit can be moved by direct memory access operations in a single DMA operation, or in smaller units that move during available time slots in coordination with the execution of the deployed neural network.
[0129] The tile data of the sensing cycles described herein can include an array of sensor data having one or more features. For example, the sensor data can include two images analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The tile data can also include metadata about the images and sensors. For example, in a base calling implementation, the tile data can include information about the alignment of the images with clusters, such as distance from center information indicating the distance of each pixel in the array of sensor data from the center of the group of genetic material on the tile.
[0130] As described below, during execution of the neural network-based base caller 104, the tile data may also include data generated during execution of the neural network-based base caller 104. This data is referred to as intermediate data that can be reused rather than recalculated during operation of the neural network-based base caller 104. For example, during execution of the neural network-based base caller 104, the data flow logic 597 may write the intermediate data to memory 548A in place of sensor data for a given patch of the array of tile data. Such implementations are described in more detail below.
[0131] As shown, a system for analyzing base calling sensor output is described that includes a runtime program-accessible memory (e.g., 548A) that stores tile data including tile sensor data from a detection cycle of a base calling operation. The system also includes a neural network processor, such as configurable processor 546, that has access to the memory. The neural network processor is configured to perform neural network operations using trained parameters to generate classification data for the detection cycle. As described herein, the neural network operations operate on an arrangement of N arrays of tile data from each of the N detection cycles comprising a subject cycle to generate classification data for the subject cycle. Data flow logic 908 is provided to move the tile data and trained parameters from the memory to the neural network processor for execution of the neural network using input units including data for the N arrays of spatially aligned patches from each of the N detection cycles.
[0132] Also described is a system in which a neural network processor has access to memory and includes a plurality of execution clusters, the execution clusters being configured to execute a neural network. Data flow logic 597 has access to memory and to an execution cluster among the plurality of execution clusters to provide an input unit of tile data to an available execution cluster among the plurality of execution clusters, the input unit including N spatially aligned patches of an array of tile data from each detection cycle including a subject detection cycle, and cause the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patch of the subject detection cycle, where N is greater than 1.
[0133] Data Flow Logic FIG. 6 illustrates one embodiment of the disclosed data flow logic that enables a host processor to filter unreliable clusters based on base calls predicted by a neural network running on a configurable processor, and further enables the configurable processor to generate a reliable intermediate representation of the remainder using data that identifies the unreliable clusters.
[0134] In action 1, data flow logic 597 requests initial cluster data from memory 548B. The initial cluster data, as described above, includes sequencing images showing the intensity emissions of clusters in the initial sequencing cycles of the sequencing operation, i.e., the first subset of sequencing cycles of the sequencing operation. For example, the initial cluster data may include sequencing images of the first 25 sequencing cycles of the sequencing operation (initial sequencing cycles).
[0135] Note that because the clusters are arranged on the flow cell at high spatial density (e.g., low micrometer or sub-micrometer resolution), a sequencing image of the initial cluster data will show intensity emissions from multiple clusters, which may include both reliable and unreliable clusters. That is, if a particular unreliable cluster is adjacent to a particular reliable cluster, the corresponding sequencing image of the initial cluster data will show intensity emissions from both the unreliable and reliable clusters, because the sequencing image of the initial cluster data is captured at an optical resolution that captures light or signals emitted from multiple clusters.
[0136] In action 2, memory 548B sends the initial cluster data to data flow logic 597.
[0137] In action 3, the data flow logic 597 provides the initial cluster data to the configurable processor 546.
[0138] In action 4, the neural network-based base caller 104 running on the configurable processor 546 generates initial intermediate representations (e.g., feature maps) from the initial cluster data (e.g., by processing the initial cluster data through its spatial and temporal convolutional layers) and generates initial base call classification scores for the plurality of clusters and for the initial sequencing cycle based on the initial intermediate representations. In one embodiment, the initial base call classification scores are not normalized, e.g., are not subjected to exponential normalization by a softmax function.
[0139] In action 5, the configurable processor 546 sends the initial unnormalized base call classification scores to the data flow logic 597.
[0140] In action 6, the data flow logic 597 provides the initial unnormalized base call classification scores to the host processor 552.
[0141] In action 7, the host processor 552 normalizes the unnormalized initial base call classification scores (e.g., by applying a softmax function) to generate normalized initial base call classification scores, i.e., initial base calls.
[0142] In action 8, the detection and filtering logic 146 running on the host processor 552 uses the normalized initial base call classification scores / initial base calls to identify unreliable clusters among the multiple clusters based on generating filter values as described above in the section entitled "Detecting and Filtering Unreliable Clusters."
[0143] In action 9, the host processor 552 sends data identifying the unreliable cluster to the data flow logic 597. The unreliable cluster can be identified by the instrument ID, the run number on the instrument, the flow cell ID, the lane number, the tile number, the X coordinate of the cluster, the Y coordinate of the cluster, and the unique molecular identifier (UMI).
[0144] In action 10, data flow logic 597 requests remaining cluster data from memory 548B. The remaining cluster data, as described above, includes sequencing images showing intensity emissions of clusters in the remaining sequencing cycles of the sequencing operation, i.e., sequencing cycles of the sequencing operation that do not include the first subset of sequencing cycles of the sequencing operation. For example, the remaining cluster data may include sequencing images for the 26th through 100th sequencing cycles (the last 75 sequencing cycles) of a 100-cycle sequencing operation.
[0145] Note that because the clusters are arranged on the flow cell at high spatial density (e.g., low micrometer or sub-micrometer resolution), the sequencing image of the remaining cluster data will show intensity emissions from multiple clusters, which may include both reliable and unreliable clusters. That is, if a particular unreliable cluster is adjacent to a particular reliable cluster, the corresponding sequencing image of the remaining cluster data will show intensity emissions from both the unreliable and reliable clusters, because the sequencing image of the remaining cluster data is captured at an optical resolution that captures light or signals emitted from multiple clusters.
[0146] In action 11, the memory 548B sends the remaining cluster data to the data flow logic 597.
[0147] In action 12, the data flow logic 597 sends data identifying the unreliable cluster to the configurable processor 546. The unreliable cluster can be identified by an instrument ID, a run number on the instrument, a flow cell ID, a lane number, a tile number, an X coordinate of the cluster, a Y coordinate of the cluster, and a unique molecular identifier (UMI).
[0148] In action 13, the data flow logic 597 sends the remaining cluster data to the configurable processor 546.
[0149] In action 14, the neural network-based base caller 104, running on the configurable processor 546, generates a residual intermediate representation (e.g., a feature map) from the remaining cluster data (e.g., by processing the remaining cluster data through its spatial convolutional layer). The configurable processor 546 generates a reliable residual intermediate representation by using the data identifying unreliable clusters to remove portions of the remaining cluster data representing unreliable clusters from the remaining intermediate representation. In one embodiment, the data identifying unreliable clusters identifies pixels indicative of intensity emissions of unreliable clusters in the initial cluster data and the remaining cluster data. In some embodiments, the configurable processor 546 is further configured to generate the reliable residual intermediate representation by discarding, from the pixelated feature map generated by the neural network-based base caller 104 from the remaining cluster data, feature map pixels resulting from pixels of the remaining cluster data indicative of intensity emissions of unreliable clusters captured for the remaining sequencing cycles.
[0150] In action 15, the configurable processor 546 is further configured to provide the reliable residual intermediate representation to the neural network-based base caller 104 and to cause the neural network-based base caller 104 to generate residual base call classification scores only for clusters among the plurality of clusters that are not reliable clusters and for the remaining sequencing cycles, thereby bypassing generation of residual base call classification scores for the unreliable clusters. In one embodiment, the residual base call classification scores are not normalized, e.g., are not subjected to exponential normalization by a softmax function.
[0151] In action 16, the configurable processor 546 sends the remaining unnormalized base call classification scores to the data flow logic 597.
[0152] In action 17, the data flow logic 597 provides the remaining unnormalized base call classification scores to the host processor 552.
[0153] In action 18, the host processor 552 normalizes the unnormalized residual base call classification scores (e.g., by applying a softmax function) to generate normalized residual base call classification scores, i.e., residual base calls.
[0154] FIG. 7 illustrates another embodiment of the disclosed data flow logic that enables a host processor to filter unreliable clusters based on base calls predicted by a neural network running on a configurable processor, and further enables the host processor to base call only reliable clusters using data that identifies unreliable clusters.
[0155] In action 1, data flow logic 597 requests initial cluster data from memory 548B. The initial cluster data, as described above, includes sequencing images showing the intensity emissions of clusters in the initial sequencing cycles of the sequencing operation, i.e., the first subset of sequencing cycles of the sequencing operation. For example, the initial cluster data may include sequencing images of the first 25 sequencing cycles of the sequencing operation (initial sequencing cycles).
[0156] Note that because the clusters are arranged on the flow cell at high spatial density (e.g., low micrometer or sub-micrometer resolution), a sequencing image of the initial cluster data will show intensity emissions from multiple clusters, which may include both reliable and unreliable clusters. That is, if a particular unreliable cluster is adjacent to a particular reliable cluster, the corresponding sequencing image of the initial cluster data will show intensity emissions from both the unreliable and reliable clusters, because the sequencing image of the initial cluster data is captured at an optical resolution that captures light or signals emitted from multiple clusters.
[0157] In action 2, memory 548B sends the initial cluster data to data flow logic 597.
[0158] In action 3, the data flow logic 597 provides the initial cluster data to the configurable processor 546.
[0159] In action 4, the neural network-based base caller 104 running on the configurable processor 546 generates initial intermediate representations (e.g., feature maps) from the initial cluster data (e.g., by processing the initial cluster data through its spatial and temporal convolutional layers) and generates initial base call classification scores for the plurality of clusters and for the initial sequencing cycle based on the initial intermediate representations. In one embodiment, the initial base call classification scores are not normalized, e.g., are not subjected to exponential normalization by a softmax function.
[0160] In action 5, the configurable processor 546 sends the initial unnormalized base call classification scores to the data flow logic 597.
[0161] In action 6, the data flow logic 597 provides the initial unnormalized base call classification scores to the host processor 552.
[0162] In action 7, the host processor 552 normalizes the unnormalized initial base call classification scores (e.g., by applying a softmax function) to generate normalized initial base call classification scores, i.e., initial base calls.
[0163] In action 8, the detection and filtering logic 146 running on the host processor 552 uses the normalized initial base call classification scores / initial base calls to identify unreliable clusters among the multiple clusters based on generating filter values as described above in the section entitled "Detecting and Filtering Unreliable Clusters."
[0164] In action 9, the host processor 552 sends data identifying the unreliable cluster to the data flow logic 597.
[0165] In action 10, data flow logic 597 requests remaining cluster data from memory 548B. The remaining cluster data, as described above, includes sequencing images showing intensity emissions of clusters in the remaining sequencing cycles of the sequencing operation, i.e., sequencing cycles of the sequencing operation that do not include the first subset of sequencing cycles of the sequencing operation. For example, the remaining cluster data may include sequencing images for the 26th through 100th sequencing cycles (the last 75 sequencing cycles) of a 100-cycle sequencing operation.
[0166] Note that because the clusters are arranged on the flow cell at high spatial density (e.g., low micrometer or sub-micrometer resolution), the sequencing image of the remaining cluster data will show intensity emissions from multiple clusters, which may include both reliable and unreliable clusters. That is, if a particular unreliable cluster is adjacent to a particular reliable cluster, the corresponding sequencing image of the remaining cluster data will show intensity emissions from both the unreliable and reliable clusters, because the sequencing image of the remaining cluster data is captured at an optical resolution that captures light or signals emitted from multiple clusters.
[0167] In action 11, the memory 548B sends the remaining cluster data to the data flow logic 597.
[0168] In action 12, the data flow logic 597 sends the remaining cluster data to the configurable processor 546.
[0169] In action 13, the neural network-based base caller 104, running on the configurable processor 546, generates residual intermediate representations (e.g., feature maps) from the residual cluster data (e.g., by processing the residual cluster data through its spatial and temporal convolutional layers). The neural network-based base caller 104 further generates residual base call classification scores for the plurality of clusters and for the remaining sequencing cycles based on the residual intermediate representations. In one embodiment, the residual base call classification scores are not normalized, e.g., are not subjected to exponential normalization by a softmax function.
[0170] In action 14, the configurable processor 546 sends the remaining unnormalized base call classification scores to the data flow logic 597.
[0171] In action 15, the data flow logic 597 sends data identifying the unreliable cluster to the host processor 552.
[0172] In action 16, the data flow logic 597 provides the remaining unnormalized base call classification scores to the host processor 552.
[0173] In action 17, the host processor 552 normalizes the unnormalized remaining base call classification scores (e.g., by applying a softmax function) and bypasses base calls of unreliable clusters in the remaining sequencing cycles by generating normalized remaining base call classification scores, i.e., remaining base calls, by using the data identifying unreliable clusters to base call only those clusters among the plurality of clusters that are not unreliable clusters. In one embodiment, the data identifying unreliable clusters identifies location coordinates of the unreliable clusters.
[0174] FIG. 8 illustrates yet another embodiment of the disclosed data flow logic that enables a host processor to filter unreliable clusters based on base calls predicted by a neural network running on a configurable processor, and further uses the data identifying unreliable clusters to generate reliable data for each remaining cluster.
[0175] In action 1, the data flow logic 597 requests initial per-cluster data from memory 548B. The per-cluster data refers to an image patch centered on a target cluster extracted from the sequencing image and base-called. The central pixel of the image patch includes the center of the target cluster. The image patch shows signals from additional clusters adjacent to the target cluster in addition to the target cluster. The initial per-cluster data includes an image patch centered on the target cluster, as described above, and shows the intensity emission of the target cluster in the initial sequencing cycles of the sequencing operation, i.e., the first subset of sequencing cycles of the sequencing operation. For example, the initial per-cluster data may include image patches for the first 25 sequencing cycles (initial sequencing cycles) of the sequencing operation.
[0176] In action 2, memory 548B sends the initial per-cluster data to data flow logic 597.
[0177] In action 3, the data flow logic 597 provides the initial per-cluster data to the configurable processor 546.
[0178] In action 4, the neural network-based base caller 104 running on the configurable processor 546 generates an initial intermediate representation (e.g., a feature map) from the initial cluster-by-cluster data (e.g., by processing the initial cluster-by-cluster data through its spatial and temporal convolutional layers), and generates initial base call classification scores for the multiple clusters and for the initial sequencing cycle based on the initial intermediate representation. In one embodiment, the initial base call classification scores are not normalized, e.g., are not subjected to exponential normalization by a softmax function.
[0179] In action 5, the configurable processor 546 sends the initial unnormalized base call classification scores to the data flow logic 597.
[0180] In action 6, the data flow logic 597 provides the initial unnormalized base call classification scores to the host processor 552.
[0181] In action 7, the host processor 552 normalizes the unnormalized initial base call classification scores (e.g., by applying a softmax function) to generate normalized initial base call classification scores, i.e., initial base calls.
[0182] In action 8, the detection and filtering logic 146 running on the host processor 552 uses the normalized initial base call classification scores / initial base calls to identify unreliable clusters among the multiple clusters based on generating filter values as described above in the section entitled "Detecting and Filtering Unreliable Clusters."
[0183] In action 9, the host processor 552 sends data identifying the unreliable cluster to the data flow logic 597. The unreliable cluster can be identified by the instrument ID, the run number on the instrument, the flow cell ID, the lane number, the tile number, the X coordinate of the cluster, the Y coordinate of the cluster, and the unique molecular identifier (UMI).
[0184] In action 10, data flow logic 597 requests remaining per-cluster data from memory 548B. The remaining per-cluster data, as described above, includes image patches centered on the target clusters and indicates the intensity emission of the target clusters for the remaining sequencing cycles of the sequencing operation, i.e., sequencing cycles of the sequencing operation that do not include the first subset of sequencing cycles of the sequencing operation. For example, the remaining per-cluster data may include image patches for the 26th through 100th sequencing cycles (the last 75 sequencing cycles) of a 100-cycle sequencing operation.
[0185] In action 11, the memory 548B sends the remaining per-cluster data to the data flow logic 597.
[0186] In action 12, the data flow logic 597 uses the data identifying the unreliable clusters to generate reliable remaining per-cluster data by removing per-cluster data representing the unreliable clusters from the remaining per-cluster data.
[0187] In action 13, the data flow logic 597 provides the remaining trusted per-cluster data to the configurable processor 546.
[0188] In action 14, the neural network-based base caller 104 running on the configurable processor 546 is further configured to bypass generating residual base call classification scores for unreliable clusters by generating only residual base call classification scores for clusters among the plurality of clusters that are not unreliable clusters and for the remaining sequencing cycles. In one embodiment, the residual base call classification scores are not normalized, e.g., are not subjected to exponential normalization by a softmax function.
[0189] In action 15, the configurable processor 546 sends the remaining unnormalized base call classification scores to the data flow logic 597.
[0190] In action 16, the data flow logic 597 provides the remaining unnormalized base call classification scores to the host processor 552.
[0191] In action 17, the host processor 552 normalizes the unnormalized residual base call classification scores (e.g., by applying a softmax function) to generate normalized residual base call classification scores, i.e., residual bases.
[0192] technical improvements Figures 9, 10, 11, 12, and 13 show the results of a comparative analysis of the detection of empty and non-empty wells using the data flow logic disclosed herein, referred to as "DeepRTA," versus Illumina's conventional base caller, called Real-Time Analysis (RTA) software.
[0193] In Figure 9, in all three plots, the x-axis is the minimum score difference over the first 25 cycles, and the score difference is the result of subtracting the second-highest likelihood from the highest likelihood. The y-axis is the number of clusters in a tile. The first plot is the result for clusters that passed the RTA chastity filter. The middle plot is for empty wells (no clusters in these nanowells according to RTA). The third plot is the result for clusters that failed the RTA chastity filter. The majority of clusters detected as unreliable using the RTA chastity filter have at least one instance of a low score difference in the first 25 cycles.
[0194] In Figure 10, the alignment metrics for one tile are shown. The last column shows the alignment metrics using the RTA chastity filter and reliable clusters based on RTA base calls. The penultimate column shows the alignment metrics using the RTA chastity filter and reliable clusters based on DeepRTA base calls. The first two columns show the alignment metrics using DeepRTA base calls and reliable clusters based on the disclosed dataflow logic, with a threshold of 0.8 (first column) or 0.9 (second column), where two of the first 25 cycles must not meet the threshold to be considered unreliable.
[0195] In Figure 11, a 0.97 threshold has been added, similar to Figure 10. Using the disclosed data flow logic and a 0.97 threshold, more clusters are detected as reliable compared to using the RTA chastity filter, while maintaining a similar (or better) alignment metric.
[0196] Figure 12 shows alignment metrics based on data from 18 tiles of a sequencing run. The first column shows DeepRTA base calls and reliable clusters using a threshold of 0.97 (highest likelihood minus second-highest likelihood), where two of the first 25 cycles must be below the threshold to be considered unreliable. The last column shows DeepRTA base calls and reliable clusters using the RTA chastity filter. Using the disclosed dataflow logic, more clusters are detected as reliable compared to using the RTA chastity filter, while maintaining similar alignment metrics.
[0197] Figure 13 shows a comparison between the RTA chastity filter and the disclosed dataflow logic using different thresholds. A large percentage of the unreliable clusters detected by the disclosed dataflow logic were also detected as unreliable by the RTA chastity filter.
[0198] Computer Systems 14 illustrates a computer system 1400 that may be used by sequencing system 500A to implement the base calling techniques disclosed herein. Computer system 1400 includes at least one central processing unit (CPU) 1472 that communicates with a number of peripheral devices via a bus subsystem 1455. These peripheral devices may include, for example, storage subsystem 858, including memory devices and file storage subsystem 1436, user interface input devices 1438, user interface output devices 1476, and network interface subsystem 1474. The input and output devices enable user interaction with computer system 1400. Network interface subsystem 1474 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0199] In one embodiment, the system controller 506 is communicatively linked to a storage subsystem 1410 and a user interface input device 1438 .
[0200] The user interface input devices 1438 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, scanners, touchscreens integrated into displays, audio input devices such as voice recognition systems and microphones, and other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and ways of inputting information into the computer system 1400.
[0201] The user interface output devices 1476 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display such as an audio output device. In general, use of the term "output device" is intended to encompass all possible types of devices and ways for outputting information from the computer system 1400 to a user or to another machine or computer system.
[0202] The storage subsystem 858 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the deep learning processor 1478.
[0203] The deep learning processor 1478 may be a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The deep learning processor 1478 may be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of deep learning processors 1478 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX14 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.
[0204] The memory subsystem 1422 used in the storage subsystem 858 may include multiple memories, including a main random access memory (RAM) 1432 for storing instructions and data during program execution, and a read only memory (ROM) 1434 in which fixed instructions are stored. The file storage subsystem 1436 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of particular embodiments may be stored in the storage subsystem 858 by the file storage subsystem 1436, or in other machines accessible by the processor.
[0205] Bus subsystem 1455 provides a mechanism for allowing the various components and subsystems of computer system 1400 to communicate with each other as intended. Although bus subsystem 1455 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0206] The computer system 1400 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a widely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 1400 shown in Figure 14 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 1400 can have more or fewer components than the computer system shown in Figure 14.
[0207] Specific Implementations Various embodiments of cluster filtering based on artificial intelligence predicted basecalls are described. One or more features of the embodiments can be combined with the basic embodiments and implemented as a system, method, or article. Non-mutually exclusive embodiments are taught as combinable. One or more features of the embodiments can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of a repeating list of these options should not be construed as limiting the combinations taught in the preceding sections. These descriptions are incorporated herein by reference into each of the following embodiments.
[0208] In one embodiment, the disclosed technology proposes a computer-implemented method for identifying unreliable clusters to improve the accuracy and efficiency of neural network-based base calling. The disclosed technology accesses cycle-by-cycle cluster data for a plurality of clusters and for a first subset of sequencing cycles of a sequencing operation.
[0209] The disclosed technology uses a neural network-based base caller to base call each cluster among a plurality of clusters in each sequencing cycle within a first subset of sequencing cycles. This includes processing the per-cycle cluster data through the neural network-based base caller to generate an intermediate representation of the per-cycle cluster data. This further includes processing the intermediate representation through an output layer to generate per-cluster, per-cycle probability quartiles for each cluster and for each sequencing cycle. A particular per-cluster, per-cycle probability quartile identifies the probability of A, C, T, and G being incorporated into a particular cluster in a particular sequencing cycle.
[0210] The disclosed technique generates an array of filter values for each cluster by determining a filter value for each cluster-by-cluster, cycle-by-cycle probability quartile based on the probability that each cluster-by-cluster, cycle-by-cycle probability quartile identifies.
[0211] The disclosed technique identifies a cluster among multiple clusters whose filter value sequence contains "N" filter values below a threshold "M" as an unreliable cluster.
[0212] The disclosed technology uses a neural network-based base caller to call bases only in clusters among multiple clusters that are not identified as unreliable clusters in the remainder of the sequencing cycle of a sequencing operation by bypassing base calls in unreliable clusters in the remainder of the sequencing cycle.
[0213] item 1. A computer-implemented method for identifying unreliable clusters to improve base calling accuracy and efficiency, the method comprising: accessing cycle-by-cycle cluster data for a plurality of clusters and for a first subset of sequencing cycles of the sequencing operation; calling a base for each cluster among the plurality of clusters in each sequencing cycle within a first subset of sequencing cycles; processing the per-cycle cluster data to generate an intermediate representation of the per-cycle cluster data; processing the intermediate representation through an output layer to generate per-cluster, per-cycle probability quartiles for each cluster and for each sequencing cycle, wherein a particular per-cluster, per-cycle probability quartile identifies the probabilities of A, C, T, and G being incorporated into a particular cluster in a particular sequencing cycle; generating an array of filter values for each cluster by determining a filter value for each cluster-by-cluster, cycle-by-cycle probability quartile based on the probability that each cluster-by-cluster, cycle-by-cycle probability quartile identifies; identifying a cluster among the plurality of clusters as an unreliable cluster, the cluster having an array of filter values that includes at least "N" filter values below a threshold "M"; and bypassing base calling of unreliable clusters in the remainder of a sequencing cycle of a sequencing operation, thereby base calling only clusters among the plurality of clusters that are not identified as unreliable clusters in the remainder of the sequencing cycle. 2. The computer-implemented method of item 1, wherein the filter values for the probability quartiles per cluster and per cycle are determined based on an arithmetic operation involving one or more of the probabilities. 3. The computer-implemented method according to items 1-2, wherein the arithmetic operation is subtraction. 4. The computer-implemented method of items 1 to 3, wherein the filter value for the probability quartile for each cluster and each cycle is determined by subtracting the second highest of the probabilities from the highest of the probabilities. 5. The computer-implemented method according to items 1 to 4, wherein the arithmetic operation is division. 6. The computer-implemented method according to items 1 to 5, wherein the filter value for the probability quartile per cluster per cycle is determined as the ratio of the highest probability of the probabilities to the second highest probability of the probabilities. 7. The computer-implemented method of items 1 to 6, wherein the arithmetic operation is addition. 8. The computer-implemented method according to items 1 to 7, wherein the arithmetic operation is multiplication. 9. The computer-implemented method of items 1 to 8, wherein "N" is in the range of 1 to 5. 10. The computer-implemented method of any one of items 1 to 9, wherein "M" is in the range of 0.5 to 0.99. 11. The computer-implemented method of items 1 to 10, wherein the first subset comprises 1 to 25 sequencing cycles of sequencing operations. 12. The computer-implemented method of items 1 to 11, wherein the first subset comprises 1 to 50 sequencing cycles of sequencing operations. 13. The computer-implemented method of items 1 to 12, wherein the output layer is a softmax layer, and the probabilities at the probability quartiles per cluster and per cycle are exponentially normalized classification scores that sum to 1. 14. The computer-implemented method according to items 1 to 13, wherein unreliable clusters represent empty wells, polyclonal wells, and ambiguous wells on the patterned flow cell. 15. The computer-implemented method of items 1 to 14, wherein the filter value is generated by a filtering function. 16. The computer-implemented method of items 1 to 15, wherein the filtering function is a chastity filter that defines chastity as the ratio of the brightest base intensity divided by the sum of the brightest base intensity and the second brightest base intensity. 17. The computer-implemented method of any one of items 1 to 16, wherein the filtering function is at least one of a maximum log-probability function, a least square error function, a mean signal-to-noise ratio (SNR), and a least absolute error function. 18. determining an average SNR of the sequencing cycles within the first subset of sequencing cycles for each cluster based on intensity data of the cluster data for each cycle, the intensity data representing intensity emissions of the clusters among the plurality of clusters and intensity emissions of a surrounding background; 18. The computer-implemented method of claim 1, further comprising: identifying clusters among the plurality of clusters whose average SNR is below a threshold as unreliable clusters. 19. determining a mean probability score for each cluster based on the maximum probability scores in the per-cluster, per-cycle probability quartiles generated for the sequencing cycles in the first subset of sequencing cycles; 19. The computer-implemented method of claim 1, further comprising: identifying clusters among the plurality of clusters whose average probability scores are below a threshold as unreliable clusters. 20. A system for improving the accuracy and efficiency of neural network-based base calling, the system comprising: a memory that stores, for a plurality of clusters, initial cluster data for an initial sequencing cycle of the sequencing operation and remaining cluster data for remaining sequencing cycles of the sequencing operation; a host processor having access to memory and configured to execute detection and filtering logic to identify unreliable clusters; a configurable processor having access to a memory and configured to execute a neural network to generate base call classification scores; data flow logic having access to a memory, a host processor, and a configurable processor; providing initial cluster data to a neural network and causing the neural network to generate initial base call classification scores for the plurality of clusters and for the initial sequencing cycles based on generating initial intermediate representations from the initial cluster data; providing the initial base call classification scores to detection and filtering logic, and causing the detection and filtering logic to identify unreliable clusters among the plurality of clusters based on generating filter values from the initial base call classification scores; providing the remaining cluster data to a neural network and causing the neural network to generate a remaining intermediate representation from the remaining cluster data; and data flow logic configured to: provide data identifying unreliable clusters to a configurable processor; and cause the configurable processor to generate a reliable remaining intermediate representation by removing from the remaining intermediate representation portions that represent unreliable clusters that arise from portions of the remaining cluster data. 21. The system of item 20, wherein the configurable processor is further configured to provide the reliable residual intermediate representation to the neural network and have the neural network generate residual base call classification scores only for clusters among the plurality of clusters that are not unreliable clusters and for the remaining sequencing cycles, thereby bypassing generation of residual base call classification scores for unreliable clusters. 22. The system according to items 20 to 21, wherein the initial and remaining base call classification scores are not normalized. 23. The data flow logic is further configured to provide the unnormalized initial and remaining base call classification scores to a host processor and cause the host processor to apply an output function to generate exponentially normalized initial and remaining base call classification scores that sum to one and indicate the probabilities of A, C, T, and G being incorporated into a particular cluster in a particular sequencing cycle; 23. The system of items 20 to 22, wherein the output function is at least one of a softmax function, a log-softmax function, an ensemble output mean function, a multilayer perceptron uncertainty function, a Bayesian Gaussian distribution function, and a cluster strength function. 24. The system of items 20 to 23, wherein the host processor is further configured to generate filter values from the exponentially normalized initial base call classification scores based on an arithmetic operation involving one or more of the following: probabilities. 25. The system according to items 20 to 24, wherein the arithmetic operation is subtraction. 26. The system of items 20-25, wherein the filter value is generated by subtracting the second highest of the probabilities from the highest of the probabilities. 27. The system according to items 20 to 26, wherein the arithmetic operation is division. 28. The system according to items 20-27, wherein the filter value is generated as a ratio of the highest probability of the probabilities to the second highest probability of the probabilities. 29. The system according to items 20 to 28, wherein the arithmetic operation is addition. 30. The system according to items 20 to 29, wherein the arithmetic operation is multiplication. 31. The system of items 20-30, wherein the host processor is further configured to generate filter values based on an average signal-to-noise ratio (SNR) determined for each cluster from intensity data in the initial cluster data, the intensity data being indicative of the intensity radiation of a cluster among the plurality of clusters and the intensity radiation of a surrounding background. 32. The system described in items 20 to 31, wherein the host processor is further configured to generate filter values based on an average probability score determined for each cluster from the maximum classification score among the initial base call classification scores. 33. The system according to items 20 to 32, wherein the data identifying unreliable clusters identifies location coordinates of the unreliable clusters. 34. The system described in items 20 to 33, wherein the host processor is further configured to identify clusters among the plurality of clusters that have filter values of "N" initial sequencing cycles that are below a threshold "M" as unreliable clusters. 35. The system according to items 20 to 34, wherein "N" is in the range of 1 to 5. 36. The system according to items 20 to 35, wherein "M" is in the range of 0.5 to 0.99. 37. The system described in items 20 to 36, wherein the host processor is further configured to bypass base calls of unreliable clusters in the remaining sequencing cycles by base calling only clusters that are not unreliable clusters among the plurality of clusters in the remaining sequencing cycles based on the highest score among the remaining exponentially normalized base call classification scores. 38. The initial cluster data and the remaining cluster data are pixelated data; The intermediate representation is a pixelated feature map, 38. The system of items 20-37, wherein the portions are pixels. 39. The system described in items 20 to 38, wherein the data identifying unreliable clusters identifies pixels that exhibit intensity radiation of unreliable clusters in the initial cluster data and the remaining cluster data. 40. The system according to items 20 to 39, wherein the data identifying unreliable clusters identifies pixels that do not exhibit any intensity emission. 41. The system described in items 20 to 40, wherein the configurable processor is further configured to generate a reliable residual intermediate representation from the pixelated feature map generated from the residual cluster data by the spatial convolutional layer of the neural network by discarding feature map pixels resulting from pixels of the residual cluster data that indicate intensity emissions of unreliable clusters captured for the remaining sequencing cycles. 42. The system described in items 20 to 41, wherein the remaining intermediate representation has 4 to 9 times the total number of pixels of the reliable remaining intermediate representation. 43. The system of items 20 to 42, wherein discarding causes the neural network to generate the remaining base call classification scores by operating on fewer pixels and thereby performing fewer computational operations. 44. The system of items 20-43, wherein discarding reduces the amount of data exchanged with the configurable processor, including cluster strength state information, and the amount of data storage. 45. The system according to items 20 to 44, wherein the unreliable clusters represent empty wells, polyclonal wells, and ambiguous wells on the patterned flow cell. 46. A system for improving the accuracy and efficiency of neural network-based base calling, the system comprising: a memory that stores, for a plurality of clusters, initial cluster data for an initial sequencing cycle of the sequencing operation and remaining cluster data for remaining sequencing cycles of the sequencing operation; a host processor having access to memory and configured to execute detection and filtering logic to identify unreliable clusters; a configurable processor having access to a memory and configured to execute a neural network to generate base call classification scores; data flow logic having access to a memory, a host processor, and a configurable processor; providing initial cluster data to a neural network and causing the neural network to generate initial base call classification scores for the plurality of clusters and for the initial sequencing cycles based on generating initial intermediate representations from the initial cluster data; providing the initial base call classification scores to detection and filtering logic, and causing the detection and filtering logic to identify unreliable clusters among the plurality of clusters based on generating filter values from the initial base call classification scores; providing the remaining cluster data to a neural network and causing the neural network to generate remaining base call classification scores for the plurality of clusters and for the remaining sequencing cycles based on generating the remaining intermediate representations from the remaining cluster data; and data flow logic configured to: provide the remaining base call classification scores to a host processor; and have the host processor use the data identifying the unreliable clusters to base call only those clusters among the plurality of clusters that are not unreliable clusters, thereby bypassing base calling of the unreliable clusters in the remaining sequencing cycles. 47. A system for improving the accuracy and efficiency of neural network-based base calling, the system comprising: a memory that stores, for a plurality of clusters, data for an initial cluster for an initial sequencing cycle of the sequencing operation and data for remaining clusters for remaining sequencing cycles of the sequencing operation; a host processor having access to memory and configured to execute detection and filtering logic to identify unreliable clusters; a configurable processor having access to a memory and configured to execute a neural network to generate base call classification scores; data flow logic having access to a memory, a host processor, and a configurable processor; providing initial per-cluster data to a neural network and causing the neural network to generate initial base call classification scores for the plurality of clusters and for the initial sequencing cycles based on generating initial intermediate representations from the initial per-cluster data; providing the initial base call classification scores to detection and filtering logic, and causing the detection and filtering logic to identify unreliable clusters among the plurality of clusters based on generating filter values from the initial base call classification scores; generating remaining reliable per-cluster data by using the data identifying the unreliable clusters and removing the per-cluster data representing the unreliable clusters from the remaining per-cluster data; and data flow logic configured to: provide data for each reliable remaining cluster to a neural network; and bypass generation of remaining base call classification scores for unreliable clusters by having the neural network generate only remaining base call classification scores for clusters among the plurality of clusters that are not unreliable clusters and for the remaining sequencing cycles. 48. A non-transitory computer-readable storage medium having stored thereon computer program instructions for identifying unreliable clusters and improving base calling accuracy and efficiency, the instructions, when executed on a processor, performing: accessing cycle-by-cycle cluster data for a plurality of clusters and for a first subset of sequencing cycles of the sequencing operation; calling a base for each cluster among the plurality of clusters in each sequencing cycle within a first subset of sequencing cycles; processing the per-cycle cluster data to generate an intermediate representation of the per-cycle cluster data; processing the intermediate representation through an output layer to generate per-cluster, per-cycle probability quartiles for each cluster and for each sequencing cycle, wherein a particular per-cluster, per-cycle probability quartile identifies the probabilities of A, C, T, and G being incorporated into a particular cluster in a particular sequencing cycle; generating an array of filter values for each cluster by determining a filter value for each cluster-by-cluster, cycle-by-cycle probability quartile based on the probability that each cluster-by-cluster, cycle-by-cycle probability quartile identifies; identifying a cluster among the plurality of clusters as an unreliable cluster, the cluster having an array of filter values that includes at least "N" filter values below a threshold "M"; A non-transitory computer-movable storage medium that implements a method including: bypassing base calling of unreliable clusters in the remainder of a sequencing cycle of a sequencing operation, thereby base calling only clusters among a plurality of clusters that are not identified as unreliable clusters in the remainder of the sequencing cycle. 49. A system including one or more processors coupled to a memory, the memory being loaded with computer instructions for performing base calling, the instructions, when executed on the processors, accessing cycle-by-cycle cluster data for a plurality of clusters and for a first subset of sequencing cycles of the sequencing operation; calling a base for each cluster among the plurality of clusters in each sequencing cycle within a first subset of sequencing cycles; processing the per-cycle cluster data to generate an intermediate representation of the per-cycle cluster data; processing the intermediate representation through an output layer to generate per-cluster, per-cycle probability quartiles for each cluster and for each sequencing cycle, wherein a particular per-cluster, per-cycle probability quartile identifies the probabilities of A, C, T, and G being incorporated into a particular cluster in a particular sequencing cycle; generating an array of filter values for each cluster by determining a filter value for each cluster-by-cluster, cycle-by-cycle probability quartile based on the probability that each cluster-by-cluster, cycle-by-cycle probability quartile identifies; identifying a cluster among the plurality of clusters as an unreliable cluster, the cluster having an array of filter values that includes at least "N" filter values below a threshold "M"; The system performs an action including: bypassing base calling of unreliable clusters in the remainder of a sequencing cycle of a sequencing operation, thereby base calling only clusters among a plurality of clusters that are not identified as unreliable clusters in the remainder of the sequencing cycle.
[0214] While the present invention has been disclosed with reference to the above-described preferred embodiments and examples, it should be understood that these examples are intended in an illustrative and not a limiting sense. Modifications and combinations will readily occur to those skilled in the art, and such modifications and combinations are deemed to be within the spirit of the invention and the scope of the following claims. [Explanation of symbols]
[0215] 102 Data Providers 104 Neural Network-Based Base Caller 106 Probability Quartiles Cluster data for each 112 cycles 116 Filter Calculator 124 Untrusted Cluster 126 filter values 132 Image Generation System 136 Unreliable cluster identifier 142 Bypass Logic 146 Detection and Filtering Logic 500 Sequencing System 502 Common Housing 504 Temperature Control System 506 System Controller 508 Fluid Control System 510 System Receptacle 512 Biosensors 514 Fluid Storage System 516 Lighting System 518 User Interface 520 Display 522 User Input Devices 524 Main Control Module 526 Lighting Module 528 Fluid Control Module 530 Fluid Storage Module 532 Temperature Control Module 534 Communication Links 536 Device Module 538 Identification Module 542 Amplification Module 544 Analysis Module 546 Configurable Processor 548 memory 550 communication port 552 host processor 589 Bus 593 Bus 595 Control Path 597 Data Flow Logic 599 Data Channel 1400 Computer Systems 1410 Memory Subsystem 1422 Memory Subsystem Used 1432 RAM 1434ROM 1436 File Storage Subsystem 1438 User Interface Input Devices 1455 Bus Subsystem 1472 Central Processing Unit (CPU) 1474 Network Interface Subsystem 1476 User Interface Output Devices 1478 Deep Learning Processor
Claims
1. 1. A computer-implemented method for identifying unreliable clusters to improve base calling accuracy and efficiency, the method comprising: accessing cycle-by-cycle cluster data for a plurality of clusters and for a first subset of sequencing cycles of the sequencing operation; calling bases on a multi-pixel image with a neural network-based base caller in each cluster among the plurality of clusters at each sequencing cycle in a first subset of sequencing cycles; processing the per-cycle cluster data and generating an intermediate representation of the per-cycle cluster data via the neural network-based base caller; processing the intermediate representation through a normalization function in an output layer of the neural network-based base caller to generate per-cluster, per-cycle probability quartiles for each cluster and for each sequencing cycle, wherein a particular per-cluster, per-cycle probability quartile identifies the probabilities of A, C, T, and G being incorporated into a particular cluster in a particular sequencing cycle; using the identified probabilities to determine a filter value as a function of the highest probability of making a base call for each cluster and for each cycle for each probability quartile, thereby generating an array of filter values for each cluster; identifying those clusters among said plurality of clusters whose sequence of filter values includes at least "N" filter values below a threshold "M" as unreliable clusters; bypassing base calling of the unreliable cluster in the remainder of a sequencing cycle of the sequencing operation, thereby base calling only clusters among the plurality of clusters that are not identified as the unreliable cluster in the remainder of the sequencing cycle.
2. The computer-implemented method of claim 1 , wherein the filter values for probability quartiles per cluster per cycle are determined based on an arithmetic operation involving one or more of the probabilities.
3. The computer-implemented method of claim 2 , wherein the arithmetic operation is subtraction.
4. 4. The computer-implemented method of claim 1, wherein the filter value for the probability quartile for each cluster and cycle is determined by subtracting the second-highest one of the probabilities from the highest one of the probabilities.
5. The computer-implemented method of any one of claims 2 to 4, wherein the arithmetic operation is division.
6. 4. The computer-implemented method of claim 1, wherein the filter value for the probability quartile for each cluster and cycle is determined as a ratio of the highest one of the probabilities to the second highest one of the probabilities.
7. The computer-implemented method of claim 2 , wherein the arithmetic operation is addition.
8. The computer-implemented method of claim 2 , wherein the arithmetic operation is multiplication.
9. The computer-implemented method of any one of claims 1 to 8, wherein the at least "N" filter values range from 1 to 5.
10. The computer-implemented method of any one of claims 1 to 9, wherein "M" is in the range of 0.5 to 0.
99.
11. A computer-implemented method described in any one of claims 1 to 10, wherein the first subset of sequencing cycles includes 1 to 25 sequencing cycles of the sequencing operation.
12. A computer-implemented method described in any one of claims 1 to 11, wherein the first subset of sequencing cycles includes 1 to 50 sequencing cycles of the sequencing operation.
13. 13. The computer-implemented method of any one of claims 1 to 12, wherein the output layer is a softmax layer and the probabilities in the per-cluster, per-cycle probability quartiles are exponentially normalized classification scores that sum to one.
14. The computer-implemented method of any one of claims 1 to 13, wherein the unreliable clusters represent empty wells, polyclonal wells, and equivocal wells on a patterned flow cell.
15. A system for improving the accuracy and efficiency of neural network-based base calling, said system comprising: a memory that stores, for a plurality of clusters, initial cluster data as a multi-pixel image for an initial sequencing cycle of a sequencing operation and remaining cluster data for remaining sequencing cycles of the sequencing operation; a host processor having access to the memory and configured to execute detection and filtering logic to identify unreliable clusters; a configurable processor having access to the memory and configured to run a neural network to generate base call classification scores; data flow logic having access to the memory, the host processor, and the configurable processor; providing the initial cluster data to the neural network, causing the neural network to generate initial intermediate representations from the initial cluster data, and generating initial base call classification scores for the plurality of clusters and for the initial sequencing cycle by processing the initial intermediate representations through a normalization function in an output layer of the neural network; providing the initial base call classification scores to the detection and filtering logic, causing the detection and filtering logic to generate filter values using the initial base call classification scores as a function of the highest classification score determining a base call, and identifying unreliable clusters within the plurality of clusters based on the generated filter values; providing the remaining cluster data to the neural network and causing the neural network to generate a remaining intermediate representation from the remaining cluster data; and data flow logic configured to: provide data identifying the unreliable clusters to the configurable processor; and cause the configurable processor to generate a reliable remaining intermediate representation by removing from the remaining intermediate representation portions representing the unreliable clusters that arise from portions of the remaining cluster data.
16. The system of claim 15, further configured to generate initial base call classification scores for the plurality of clusters and for the initial sequencing cycle by generating cluster-wise, cycle-wise probability quartiles for each cluster and for each sequencing cycle, wherein a particular cluster-wise, cycle-wise probability quartile identifies the probability of A, C, T, and G being bases incorporated into a particular cluster in a particular sequencing cycle.
17. The system described in claim 16, wherein filter values for probability quartiles for each cluster and each cycle are determined based on arithmetic operations involving one or more of the probabilities.
18. The system of claim 17, wherein the arithmetic operation is subtraction, division, addition, or multiplication.
19. The system described in claim 17, wherein the filter value for the probability quartile for each cluster and cycle is determined by subtracting the second highest probability among the probabilities from the highest probability among the probabilities.
20. The system described in claim 17, wherein the filter value for the probability quartile for each cluster and cycle is determined as the ratio of the highest probability among the probabilities to the second highest probability among the probabilities.
Citation Information
Patent Citations
Methods and systems for analyzing image data
US20180274023A1