Many-to-many base interpretation based on artificial intelligence

Through the base calling method based on deep neural network, the problem of difficulty in distinguishing and identifying in high-density nucleic acid cluster sequencing is solved, and efficient and low-cost nucleic acid sequencing data acquisition is achieved, which is suitable for a variety of biological and medical applications.

CN115136244BActive Publication Date: 2025-09-30ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180015480.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-20
Filing Date
2021-02-19
Publication Date
2025-09-30
Estimated Expiration
2041-02-19

AI Technical Summary

Technical Problem

Existing high-density nucleic acid cluster sequencing technologies have difficulty distinguishing and identifying closely adjacent or overlapping nucleic acid clusters, resulting in a compromise between the quantity and quality of nucleic acid sequence information, limiting the efficiency and cost-effectiveness of high-throughput nucleic acid sequencing.

Method used

A base calling method based on deep neural networks is adopted. Sequencing image data is processed through convolutional neural networks and recurrent neural networks. Spatial and temporal convolutional layers are used to separate sequencing cycle information. The softmax function is combined for base calling, which improves the accuracy and efficiency of base calling.

Benefits of technology

It improves the throughput level of nucleic acid sequencing, increases the quality and quantity of nucleic acid sequencing data, is suitable for a variety of biological and medical applications, reduces costs and increases the speed of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115136244B_ABST
    Figure CN115136244B_ABST
Patent Text Reader

Abstract

The technology disclosed in the present invention relates to artificial intelligence-based base calling. The technology disclosed in the present invention relates to accessing the progress of a per-cycle analyte channel set generated for sequencing cycles of a sequencing run; processing a window of the per-cycle analyte channel set in progress for a window of sequencing cycles of the sequencing run by a neural network-based base caller (NNBC), causing the NNBC to process an object window of the per-cycle analyte channel set in progress for an object window of sequencing cycles of the sequencing run, and using the NNBC to generate provisional base call predictions for three or more sequencing cycles in the object window of the sequencing cycle from multiple windows that appear at different positions for a particular sequencing cycle to generate a provisional base call prediction for the particular sequencing cycle; and determining a base call for the particular sequencing cycle based on the multiple base call predictions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to artificial intelligence-type computers and digital data processing systems, as well as corresponding data processing methods and products for simulating intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems). The technology also includes systems for inferring uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the technology disclosed herein relates to the use of deep neural networks, such as deep convolutional neural networks, for analyzing data.

[0002] Priority application

[0003] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 979,414, filed on February 20, 2020, entitled “ARTIFICIAL INTELLIGENCE-BASEDMANY-TO-MANY BASE CALLING” (Attorney Docket No. ILLM1016-1 / IP-1858-PRV), and U.S. Patent Application No. 17 / 180,542, filed on February 19, 2021, entitled “ARTIFICIAL INTELLIGENCE-BASEDMANY-TO-MANY BASE CALLING” (Attorney Docket No. ILLM 1016-2 / IP-1858-US). These priority applications are hereby incorporated by reference as if fully set forth herein for all purposes.

[0004] Literature Incorporation

[0005] The following documents are incorporated by reference as if fully set forth herein:

[0006] U.S. Provisional Patent Application No. 62 / 979,384, filed on February 20, 2020, entitled “ARTIFICIAL INTELLIGENCE-BASED BASE CALLING OF INDEX SEQUENCES” (Attorney Docket No. ILLM 1015-1 / IP-1857-PRV);

[0007] U.S. Provisional Patent Application No. 62 / 979,385, filed on February 20, 2020, entitled “KNOWLEDGE DISTILLATION-BASED COMPRESSION OF ARTIFICIAL INTELLIGENCE-BASED BASE CALLER” (Attorney Docket No. ILLM 1017-1 / IP-1859-PRV);

[0008] U.S. Provisional Patent Application No. 63 / 072,032, filed on August 28, 2020, entitled “DETECTING AND FILTERING CLUSTERS BASED ON ARTIFICIAL INTELLIGENCE-PREDICTED BASE CALLS” (Attorney Docket No. ILLM 1018-1 / IP-1860-PRV);

[0009] U.S. Provisional Patent Application No. 62 / 979,412, filed on February 20, 2020, entitled “MULTI-CYCLE CLUSTER BASED REAL TIMEANALYSIS SYSTEM” (Attorney Docket No. ILLM 1020-1 / IP-1866-PRV);

[0010] U.S. Provisional Patent Application No. 62 / 979,411, entitled “DATA COMPRESSION FOR ARTIFICIALINTELLIGENCE-BASED BASE CALLING,” filed on February 20, 2020 (Attorney Docket No. ILLM 1029-1 / IP-1964-PRV);

[0011] U.S. Provisional Patent Application No. 62 / 979,399, filed on February 20, 2020, entitled “SQUEEZING LAYER FOR ARTIFICIALINTELLIGENCE-BASED BASE CALLING” (Attorney Docket No. ILLM 1030-1 / IP-1982-PRV);

[0012] U.S. non-provisional patent application No. 16 / 825,987, filed on March 20, 2020, entitled “TRAINING DATA GENERATION FOR ARTIFICIALINTELLIGENCE-BASED SEQUENCING” (Attorney Docket No. ILLM 1008-16 / IP-1693-US);

[0013] U.S. non-provisional patent application No. 16 / 825,991, entitled “ARTIFICIAL INTELLIGENCE-BASED GENERATION OF SEQUENCING METADATA,” filed on March 20, 2020 (Attorney Docket No. ILLM 1008-17 / IP-1741-US);

[0014] U.S. non-provisional patent application No. 16 / 826,126, filed on March 20, 2020, and entitled “ARTIFICIAL INTELLIGENCE-BASED BASE CALLING” (Attorney Docket No. ILLM 1008-18 / IP-1744-US);

[0015] U.S. non-provisional patent application No. 16 / 826,134, entitled “ARTIFICIAL INTELLIGENCE-BASED QUALITYSCORING,” filed on March 20, 2020 (Attorney Docket No. ILLM 1008-19 / IP-1747-US); and

[0016] U.S. non-provisional patent application No. 16 / 826,168, filed on March 21, 2020, and entitled “ARTIFICIAL INTELLIGENCE-BASED SEQUENCING” (Attorney Docket No. ILLM 1008-20 / IP-1752-PRV-US). Background Art

[0017] The subject matter discussed in this section should not be considered prior art simply because it is mentioned in this section. Similarly, problems mentioned in this section or associated with the subject matter provided as background technology should not be considered to have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to specific implementations of the claimed technology.

[0018] A deep neural network is an artificial neural network that uses multiple layers of nonlinear and complex transformations to continuously model high-level features. Deep neural networks provide feedback via backpropagation, which uses the difference between observed and predicted outputs to adjust parameters. Deep neural networks have evolved with the availability of large training datasets, the power of parallel and distributed computing, and sophisticated training algorithms. They have driven significant advances in many fields, such as computer vision, speech recognition, and natural language processing.

[0019] Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are components of deep neural networks. Convolutional neural networks are particularly successful in image recognition with an architecture comprising convolutional layers, nonlinear layers, and pooling layers. Recurrent neural networks are designed to exploit the sequential information of input data and have cyclic connections between building blocks such as perceptrons, long short-term memory units, and gated recurrent units. In addition, many other emerging deep neural networks for limited contexts have been proposed, such as deep spatiotemporal neural networks, multidimensional recurrent neural networks, and convolutional autoencoders.

[0020] The goal of training a deep neural network is to optimize the weight parameters in each layer, gradually combining simpler features into complex ones, so that the most appropriate layered representation can be learned from the data. A single cycle of the optimization process proceeds as follows. First, given a training dataset, a forward pass sequentially computes the output of each layer and propagates the function signal forward through the network. In the final output layer, the target loss function measures the error between the inferred output and the given label. To minimize the training error, a backward pass uses the chain rule to backpropagate the error signal and calculate the gradient with respect to all weights in the entire neural network. Finally, an optimization algorithm based on stochastic gradient descent is used to update the weight parameters. While batch gradient descent performs parameter updates for each complete dataset, stochastic gradient descent provides a stochastic approximation by performing updates for each small data example. Several optimization algorithms are derived from stochastic gradient descent. For example, the Adagrad and Adam training algorithms perform stochastic gradient descent while adaptively modifying the learning rate based on the update frequency and momentum of the gradient of each parameter, respectively.

[0021] Another core element in deep neural network training is regularization, which refers to strategies designed to avoid overfitting and thus achieve good generalization performance. For example, weight decay adds a penalty factor to the objective loss function so that the weight parameters converge to a small absolute value. Dropout randomly removes hidden units from a neural network during training and can be thought of as an ensemble of possible subnetworks. To enhance the capabilities of dropout, new activation functions, maximum output, and a dropout variant of recurrent neural networks (called rnnDrop) have been proposed. In addition, batch normalization provides a new regularization method by normalizing the scalar features of each activation within a mini-batch and learning each mean and variance as parameters.

[0022] Given that sequence data is multidimensional and high-dimensional, deep neural networks have great prospects in bioinformatics research due to their wide applicability and enhanced predictive power. Convolutional neural networks have been used to solve sequence-based problems in genomics, such as motif discovery, pathogenic variant identification, and gene expression inference. Convolutional neural networks use a weight-sharing strategy that is particularly useful for studying deoxyribonucleic acid (DNA) because they can capture sequence motifs, which are short, recurring local patterns in DNA that are assumed to have significant biological functions. The hallmark of convolutional neural networks is the use of convolutional filters.

[0023] Unlike traditional classification methods based on carefully designed and handcrafted features, convolutional filters perform adaptive learning of features, similar to the process of mapping raw input data into an information representation of knowledge. In this sense, convolutional filters act as a series of motif scanners, as a set of such filters is able to identify relevant patterns in the input and update themselves during the training process. Recurrent neural networks can capture long-range dependencies in sequence data of varying lengths, such as protein or DNA sequences.

[0024] Therefore, there is an opportunity to use a principled deep learning-based framework for template generation and base calling.

[0025] In the era of high-throughput technologies, accumulating the most interpretable data at the lowest cost per work remains a major challenge. Cluster-based nucleic acid sequencing methods, such as those that use bridge amplification to form clusters, have made important contributions to the goal of increasing nucleic acid sequencing throughput. These cluster-based methods rely on sequencing dense populations of nucleic acids immobilized on a solid support and typically involve using image analysis software to deconvolute the optical signals generated during the simultaneous sequencing of multiple clusters located at different locations on the solid support.

[0026] The present invention relates to the sequencing technology of solid-phase nucleic acid cluster.However, this type of sequencing technology based on solid-phase nucleic acid cluster is still faced with sizable obstacles, and these obstacles limit achievable flux.For example, in the sequencing method based on cluster, determining physically too close to each other and unable to spatially resolve or in fact may have an obstacle aspect the nucleic acid sequence of two or more clusters that physically overlap on solid support.For example, current image analysis software may need valuable time and computing resources to determine which cluster in two overlapping clusters has sent light signal.Therefore, for multiple detection platforms, the compromise about the quantity and / or quality of obtainable nucleic acid sequence information is inevitable.

[0027] Genomic approaches based on high-density nucleic acid clusters have also expanded into other areas of genomic analysis. For example, nucleic acid cluster-based genomics can be used for sequencing applications, diagnostics and screening, gene expression analysis, epigenetic analysis, and genetic analysis of polymorphisms. Each of these nucleic acid cluster-based genomic technologies is also limited when data generated from closely spaced or spatially overlapping nucleic acid clusters cannot be interpreted.

[0028] Clearly, there remains a need to increase the quality and quantity of nucleic acid sequencing data that can be rapidly and cost-effectively obtained for a wide variety of uses, including genomics (e.g., for genomic characterization of any and all animal, plant, microbial, or other biological species or populations), pharmacogenetics, transcriptomics, diagnostics, prognosis, biomedical risk assessment, clinical and research genetics, personalized medicine, drug efficacy and drug interaction assessment, veterinary medicine, agriculture, evolutionary and biodiversity studies, aquaculture, forestry, oceanography, ecological and environmental management, and other purposes.

[0029] The disclosed technology provides neural network-based methods and systems that address these and similar needs, including increasing throughput levels in high-throughput nucleic acid sequencing technologies, and provide other related advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In the drawings, like reference characters generally refer to like parts throughout the different views. Additionally, the drawings are not necessarily drawn to scale, emphasis instead being placed on illustrating the principles of the disclosed technology. In the following description, various implementations of the disclosed technology are described with reference to the following drawings, wherein:

[0031] Figure 1A 、 Figure 1B and Figure 1C The disclosed many-to-many base calls are shown.

[0032] Figure 1D and Figure 1E Different examples of the disclosed many-to-many base calls are shown.

[0033] Figure 2 、 Figure 3 and Figure 4 Different implementations of the base call generator are shown.

[0034] Figure 5 A specific implementation of the disclosed multi-cycle gradient backpropagation is shown.

[0035] Figure 6 is a flowchart of a specific implementation of the disclosed technology.

[0036] Figure 7 The technical effects and advantages of the disclosed technology are shown.

[0037] Figure 8A and Figure 8B A specific implementation of a sequencing system is described. The sequencing system includes a configurable processor.

[0038] Figure 9is a simplified block diagram of a system for analyzing sensor data (such as base call sensor output) from a sequencing system.

[0039] Figure 10 is a simplified diagram illustrating aspects of a base calling operation including the functionality of a runtime program executed by a host processor.

[0040] Figure 11 Is a configurable processor (such as, Figure 9 Simplified diagram of the configuration of a configurable processor).

[0041] Figure 12 is a computer system that can be used by the disclosed sequencing system to implement the base calling technology disclosed herein. DETAILED DESCRIPTION

[0042] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology, and is provided in the context of a specific application and its requirements. Various modifications to the disclosed implementations will be apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the disclosed technology. Therefore, the disclosed technology is not intended to be limited to the specific implementations shown but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0043] Sequencing images

[0044] Base calling is the process of determining the nucleotide composition of a sequence. Base calling involves analyzing image data, i.e., sequencing images, generated during a sequencing run (or sequencing reaction) performed by sequencing instruments such as Illumina's iSeq, HiSeqX, HiSeq 3000, HiSeq 4000, HiSeq 2500, NovaSeq 6000, NextSeq 550, NextSeq 1000, NextSeq 2000, NextSeqDx, MiSeq, and MiSeqDx.

[0045] The following discussion summarizes a method for generating a sequencing image and its depictions according to one specific implementation.

[0046] Base calling decodes the intensity data encoded in the sequencing image into the nucleotide sequence. In one embodiment, the Illumina sequencing platform uses cyclic reversible termination (CRT) chemistry to perform base calling. This process relies on growing a nascent chain complementary to the template chain with fluorescently labeled nucleotides while tracking the emission signal of each newly added nucleotide. The fluorescently labeled nucleotides have a 3' removable block that anchors the fluorophore signal of the nucleotide type.

[0047] Sequencing is performed in repeated cycles, each of which consists of three steps: (a) extending the nascent chain by adding fluorescently labeled nucleotides; (b) exciting the fluorophore using one or more lasers of the sequencing instrument's optical system and imaging it through different filters of the optical system, thereby generating a sequencing image; and (c) cleaving the fluorophore and removing the 3' block in preparation for the next sequencing cycle. The incorporation and imaging cycles are repeated until a specified number of sequencing cycles is reached, thereby defining the read length. Using this method, each cycle interrogates a new position along the template chain.

[0048] The enormous power of Illumina sequencers stems from its ability to simultaneously perform and sense millions or even billions of clusters (also called "analytes") undergoing CRT reactions. A cluster comprises approximately one thousand identical copies of a template strand, but the size and shape of the clusters are different. Before the sequencing run, clusters from the template strand are grown by performing bridge amplification or exclusion amplification on the input library. The purpose of amplification and cluster growth is to increase the intensity of the emission signal because the imaging device cannot reliably sense the fluorophore signal of a single strand. However, the physical distance between the strands within a cluster is small, so the imaging device perceives the cluster of strands as a single point.

[0049] Sequencing occurs in a flow cell (or biosensor), i.e., a small glass slide that accommodates the input chain. The flow cell is connected to an optical system that includes microscope imaging, excitation lasers, and fluorescence filters. The flow cell includes a plurality of chambers referred to as grooves. The grooves are physically separated from each other and can include different labeled sequencing libraries that can be distinguished without sample cross-contamination. In some implementations, the flow cell includes a patterned surface. "Patterned surface" refers to the arrangement of different regions in or on the exposed layer of a solid support.

[0050] The sequencing instrument's imaging device (e.g., a solid-state imaging device such as a charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) sensor) takes snapshots at multiple locations along the channel, in a series of non-overlapping areas called blocks. For example, there may be sixty-four blocks or ninety-six blocks per channel. Blocks hold hundreds of thousands to millions of clusters.

[0051] The output of the sequencing run is a sequencing image. The sequencing image uses a grid (or array) of pixelated units (e.g., pixels, superpixels, subpixels) to depict the intensity emission of the cluster and its surrounding background. The intensity emission is stored as the intensity value of the pixelated unit. The sequencing image has the dimensions w, x, h of the grid of the pixelated unit, where w (width) and h (height) are any numbers in the range of 1 to 100,000 (e.g., 115 × 115, 200 × 200, 1800 × 2000, 2200 × 25000, 2800 × 3600, 4000 × 400). In some specific implementations, w and h are the same. In other specific implementations, w and h are different. The sequencing image depicts the intensity emission generated due to the incorporation of nucleotides into the nucleotide sequence during the sequencing run. The intensity emission comes from the associated cluster and its surrounding background.

[0052] Neural network-based base calling

[0053] The following discussion focuses on the neural network-based base caller 102 described herein. First, the inputs to the neural network-based base caller 102 are described according to one specific implementation. Then, an example of the structure and form of the neural network-based base caller 102 is provided. Finally, the outputs of the neural network-based base caller 102 are described according to one specific implementation.

[0054] The data flow logic provides the sequencing image to the neural network-based base caller 102 for base calling. The neural network-based base caller 102 accesses the sequencing image on a block-by-block basis (or block-by-block basis). Each block in the block is a subgrid (or subarray) of pixelated cells in a grid of pixelated cells that forms a sequencing image. The block has a size of q×r of the subgrid of pixelated cells, where q (width) and r (height) are any numbers in the range of 1 to 10,000 (e.g., 3×3, 5×5, 7×7, 10×10, 15×15, 25×25, 64×64, 78×78, 115×115). In some embodiments, q and r are the same. In other embodiments, q and r are different. In some embodiments, the blocks extracted from the sequencing image have the same size. In other embodiments, the blocks have different sizes. In some implementations, a block can have overlapping pixelated units (eg, on an edge).

[0055] Sequencing produces m sequencing images per sequencing cycle for the corresponding m image channels. That is, each sequencing image in the sequencing image has one or more image (or intensity) channels (similar to the red, green, blue (RGB) channels of a color image). In one specific implementation, each image channel corresponds to one filter wavelength band in a plurality of filter wavelength bands. In another specific implementation, each image channel corresponds to one imaging event in a plurality of imaging events in a sequencing cycle. In yet another specific implementation, each image channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter. Image blocks are tiled (or accessed) from each of the m image channels for a particular sequencing cycle. In different specific implementations such as four-channel chemistry, two-channel chemistry, and single-channel chemistry, m is 4 or 2. In other specific implementations, m is 1, 3, or greater than 4.

[0056] For example, consider a sequencing run performed using two different image channels (a blue channel and a green channel). Then, at each sequencing cycle, the sequencing run generates a blue image and a green image. Thus, for a series of k sequencing cycles of the sequencing run, a sequence of k pairs of blue and green images is generated as output and stored as sequencing images. Thus, a sequence of k pairs of blue and green image blocks is generated by the neural network-based base caller 102 for block-level processing.

[0057] The input image data to the neural network-based base caller 102 for a single iteration of base calling (or a single instance of a forward pass or a single forward traversal) includes data for a sliding window of multiple sequencing cycles. The sliding window may include, for example, the current sequencing cycle, one or more previous sequencing cycles, and one or more subsequent sequencing cycles.

[0058] In one embodiment, the input image data includes data for three sequencing cycles, such that the data to be base called for the current (time t) sequencing cycle is accompanied by: (i) data for the left flank / context / previous / previous / before (time t-1) sequencing cycle and (ii) data for the right flank / context / next / subsequent / after (time t+1) sequencing cycle.

[0059] In another specific implementation, the input image data includes data for five sequencing cycles, such that the data for the current (time t) sequencing cycle to be base called is accompanied by: (i) data for the first left flank / context / previous / previous / previous (time t-1) sequencing cycle; (ii) data for the second left flank / context / previous / previous / previous (time t-2) sequencing cycle; (iii) data for the first right flank / context / next / subsequent / after (time t+1) sequencing cycle; and (iv) data for the second right flank / context / next / subsequent / after (time t+2) sequencing cycle.

[0060] In yet another embodiment, the input image data includes data for seven sequencing cycles, such that the data for the current (time t) sequencing cycle to be base called is accompanied by: (i) data for the first left flank / context / previous / previous / previous (time t-1) sequencing cycle; (ii) data for the second left flank / context / previous / previous / previous (time t-2) sequencing cycle; (iii) data for the third left flank / context / previous / previous / previous (time t-3) sequencing cycle; (iv) data for the first right flank / context / next / subsequent / after (time t+1) sequencing cycle; (v) data for the second right flank / context / next / subsequent / after (time t+2) sequencing cycle; and (vi) data for the third right flank / context / next / subsequent / after (time t+3) sequencing cycle. In other embodiments, the input image data includes data for a single sequencing cycle. In other embodiments, the input image data includes data for 10, 15, 20, 30, 58, 75, 92, 130, 168, 175, 209, 225, 230, 275, 318, 325, 330, 525, or 625 sequencing cycles.

[0061] According to one specific implementation, the neural network-based base caller 102 processes the image patch through its convolutional layers and generates an alternative representation. The alternative representation is then used by the output layer (e.g., a softmax layer) to generate base calls for only the current (time t) sequencing cycle or each sequencing cycle in the sequencing cycle, i.e., the current (time t) sequencing cycle, the first and second previous (time t-1, time t-2) sequencing cycles, and the first and second subsequent (time t+1, time t+2) sequencing cycles. The resulting base calls form sequencing reads.

[0062] In one implementation, the neural network-based base caller 102 outputs base calls for a single target cluster for a particular sequencing cycle. In another implementation, the neural network-based base caller 102 outputs base calls for each target cluster in a plurality of target clusters for a particular sequencing cycle. In yet another implementation, the neural network-based base caller 102 outputs base calls for each target cluster in a plurality of target clusters for each sequencing cycle in a plurality of sequencing cycles, thereby generating a base call sequence for each target cluster.

[0063] In one embodiment, the neural network-based base caller 102 is a multilayer perceptron (MLP). In another embodiment, the neural network-based base caller 102 is a feedforward neural network. In yet another embodiment, the neural network-based base caller 102 is a fully connected neural network. In a further embodiment, the neural network-based base caller 102 is a fully convolutional neural network. In yet another embodiment, the neural network-based base caller 102 is a semantic segmentation neural network. In yet another embodiment, the neural network-based base caller 102 is a generative adversarial network (GAN).

[0064] In one embodiment, the neural network-based base caller 102 is a convolutional neural network (CNN) having multiple convolutional layers. In another embodiment, the neural network-based base caller 102 is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), or a gated recurrent unit (GRU). In yet another embodiment, the neural network-based base caller 102 includes both a CNN and an RNN.

[0065] In other specific implementations, the neural network-based base caller 102 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or dilated convolution, transposed convolution, depthwise separable convolution, pointwise convolution, 1×1 convolution, grouped convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The neural network-based base caller 102 can use one or more loss functions, such as logistic regression / logarithmic loss, multi-class cross entropy / softmax loss, binary cross entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. The neural network-based base caller 102 can use any parallelism, efficiency, and compression scheme, such as TFRecords, compressed encoding (e.g., PNG), sharding, parallel calling of map transformations, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). The neural network-based base caller 102 may include upsampling layers, downsampling layers, recurrent connections, gate and gate memory units (e.g., LSTM or GRU), residual blocks, residual connections, high-speed connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential lining units (ELU), sigmoid, and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout layers, pooling layers (e.g., max or average pooling), global average pooling layers, and attention mechanisms.

[0066] The neural network-based base caller 102 is trained using a back-propagation-based gradient update technique. Exemplary gradient descent techniques that can be used to train the neural network-based base caller 102 include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used to train the neural network-based base caller 102 are Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.

[0067] In one embodiment, the neural network-based base caller 102 uses a specialized architecture to separate the processing of data for different sequencing cycles. The motivation for using a specialized architecture is first described. As described above, the neural network-based base caller 102 processes image blocks for the current sequencing cycle, one or more previous sequencing cycles, and one or more subsequent sequencing cycles. The data from the additional sequencing cycles provides sequence-specific context. The neural network-based base caller 102 learns the sequence-specific context during training and performs base calls based on this sequence-specific context. In addition, the data from the previous and next sequencing cycles provide second-order contributions to the prephasing and phasing signals for the current sequencing cycle.

[0068] However, images captured at different sequencing cycles and in different image channels are misaligned with respect to each other and have residual registration errors. To account for this misalignment, a specialized architecture includes spatial convolutional layers that do not mix information between sequencing cycles and only mix information within a sequencing cycle.

[0069] The spatial convolution layer (or spatial logic) uses so-called "isolated convolutions," which achieve isolation by independently processing the data for each of multiple sequencing cycles via a "dedicated, unshared" sequence of convolutions. This isolated convolution convolves the data and resulting feature maps only for a given sequencing cycle (i.e., within the cycle), and not for any other sequencing cycles.

[0070] For example, consider that the input image data includes: (i) the current image block for the current (time t) sequencing cycle for which base calls are to be made; (ii) the previous image block for the previous (time t-1) sequencing cycle; and (iii) the next image block for the next (time t+1) sequencing cycle. The specialized architecture then initiates three separate convolutional pipelines, namely the current convolutional pipeline, the previous convolutional pipeline, and the next convolutional pipeline. The current data processing pipeline receives the current image block for the current (time t) sequencing cycle as input and independently processes the current image block through multiple spatial convolutional layers to produce a so-called "current spatial convolutional representation" as the output of the final spatial convolutional layer. The previous convolutional pipeline receives the previous image block for the previous (time t-1) sequencing cycle as input and independently processes the previous image block through multiple spatial convolutional layers to produce a so-called "previous spatial convolutional representation" as the output of the final spatial convolutional layer. The latter convolution pipeline receives the latter image patch for the latter (time t+1) sequencing cycle as input and processes the latter image patch independently through multiple spatial convolution layers to produce the so-called "latter spatial convolution representation" as the output of the final spatial convolution layer.

[0071] In some implementations, the current convolution pipeline, the previous convolution pipeline, and the next convolution pipeline are executed in parallel. In some implementations, the spatial convolution layer is part of a spatial convolution network (or subnetwork) within a specialized architecture.

[0072] The neural network-based base caller 102 further includes a temporal convolutional layer (or temporal logic) that mixes information between sequencing cycles (i.e., inter-cycle). The temporal convolutional layer receives its input from the spatial convolutional network and operates on the spatial convolutional representation produced by the final spatial convolutional layer of the corresponding data processing pipeline.

[0073] The freedom of inter-cycle operability of temporal convolutional layers stems from the fact that the misalignment property present in the image data fed as input to the spatial convolutional network is cleaned from the spatial convolutional representation by the stacking or cascading of isolated convolutions performed by a sequence of spatial convolutional layers.

[0074] The temporal convolution layer uses so-called "combined convolutions" that convolve the input channels in subsequent inputs group by group on a sliding window basis. In one specific implementation, these subsequent inputs are the subsequent outputs produced by the previous spatial convolution layer or the previous temporal convolution layer.

[0075] In some implementations, the temporal convolutional layer is part of a temporal convolutional network (or subnetwork) within a specialized architecture. The temporal convolutional network receives its input from the spatial convolutional network. In one implementation, the first temporal convolutional layer of the temporal convolutional network combines the spatial convolutional representations between sequencing cycles in groups. In another implementation, subsequent temporal convolutional layers of the temporal convolutional network combine the subsequent outputs of previous temporal convolutional layers. The output of the final temporal convolutional layer is fed to an output layer that produces an output. The output is used to perform base calling on one or more clusters at one or more sequencing cycles.

[0076] Data flow logic provides per cycle cluster data to the base caller 102 based on neural network.Per cycle cluster data is used for the first subset of the sequencing cycles of multiple clusters and sequencing runs.For example, consider that sequencing runs have 150 sequencing cycles.Then, the first subset of sequencing cycles can include any subset of 150 sequencing cycles, for example, the first 5, 10, 15, 25, 35, 40, 50 or 100 sequencing cycles of 150 cycle sequencing runs.Moreover, each sequencing cycle produces a sequencing image depicting the intensity emission of the clusters in multiple clusters.Like this, the per cycle cluster data of the first subset of the sequencing cycles of multiple clusters and sequencing runs include sequencing images only for the first 5, 10, 15, 25, 35, 40, 50 or 100 sequencing cycles of 150 cycle sequencing runs, and do not include sequencing images for the remaining sequencing cycles of 150 cycle sequencing runs.

[0077] Each sequencing cycle in the first subset of the sequencing cycle based on the neural network base caller 102 carries out base calling to each cluster in a plurality of clusters.For this reason, the base caller 102 based on the neural network processes the per cycle cluster data and generates the intermediate representation of the per cycle cluster data.Then, the base caller 102 based on the neural network processes the intermediate representation through the output layer and produces per cluster, per cycle probability quadruple for each cluster and each sequencing cycle.Examples of the output layer include softmax function, log-softmax function, integrated output average function, multilayer perceptron uncertainty function, Bayesian Gaussian distribution function and cluster strength function.Per cluster, per cycle probability quadruple is stored as probability quadruple and is referred to as "base-by-base possibility" in this article because there are four nucleotide bases A, C, T and G.

[0078] The softmax function is a preferred function for multi-class classification. The softmax function calculates the probability of each target class across all possible target classes. The output of the softmax function ranges from zero to one, and the sum of all probabilities equals one. The softmax function calculates the exponential of a given input value and the sum of the exponentials of all input values. The ratio of the exponential of the input value to the sum of the exponentials is the output of the softmax function and is referred to as "exponential normalization" in this article.

[0079] Formally, training a so-called softmax classifier is regressing to class probabilities, not to a true classifier, because it does not return classes, but rather confidence predictions for each class probability. The softmax function takes a class value and converts them into probabilities that sum to 1. The softmax function compresses an n-dimensional vector of arbitrary real values ​​into an n-dimensional vector of real values ​​in the range 0 to 1. Therefore, using the softmax function ensures that the output is a valid, exponentially normalized probability mass function (non-negative and sums to 1).

[0080] Intuitively, the softmax function is a "soft" version of the max function. The term "soft" comes from the fact that the softmax function is continuous and differentiable. Instead of selecting a single largest element, it breaks the vector into components where the largest input element receives a proportionally larger value and the others receive proportionally smaller values. This property of the output probability distribution makes the softmax function well-suited for probabilistic interpretation in classification tasks.

[0081] Let's think of z as the input vector to the softmax layer. The softmax layer units is the number of nodes in the softmax layer, and therefore, the length of the z vector is the number of units in the softmax layer (if there are ten output units, there are ten z elements).

[0082] For an n-dimensional vector Z = [z1, z2, ...z n ], the softmax function uses exponential normalization (exp) to produce another n-dimensional vector p(Z) with normalized values ​​in the range [0,1] that sum to one.

[0083]

[0084]

[0085] For example, applying the softmax function to three categories as Note that the three outputs always sum to 1. Therefore, they define a discrete probability mass function.

[0086] A particular per-cluster, per-cycle probability quadruplet identifies the probability that the base incorporated into a particular cluster at a particular sequencing cycle is A, C, T, and G. When the output layer of the neural network-based base caller 102 uses a softmax function, the probabilities in the per-cluster, per-cycle probability quadruplet are exponentially normalized classification scores that sum to one.

[0087] In one implementation, the method includes processing the convolutional representation through an output layer to generate a probability of a base that is A, C, T, and G incorporated into the target analyte at the current sequencing cycle, and classifying the base as A, C, T, or G based on the probability. In one implementation, the probability is an exponentially normalized score produced by the softmax layer.

[0088] In one embodiment, the method includes deriving an output pair for the target analyte from the output, the output pair identifying a class label for a base that is A, C, T, or G incorporated into the target analyte at the current sequencing cycle, and performing a base call for the target analyte based on the class label. In one embodiment, class label 1,0 identifies an A base; class label 0,1 identifies a C base; class label 1,1 identifies a T base; and class label 0,0 identifies a G base. In another embodiment, class label 1,1 identifies an A base; class label 0,1 identifies a C base; class label 0.5,0.5 identifies a T base; and class label 0,0 identifies a G base. In yet another embodiment, class label 1,0 identifies an A base; class label 0,1 identifies a C base; class label 0.5,0.5 identifies a T base; and class label 0,0 identifies a G base. In yet a further embodiment, class labels 1,2 identify an A base; class labels 0,1 identify a C base; class labels 1,1 identify a T base; and class labels 0,0 identify a G base. In one embodiment, the method includes deriving a class label for the target analyte from the output, the class label identifying the bases that are A, C, T, or G incorporated into the target analyte at the current sequencing cycle, and performing base calling for the target analyte based on the class label. In one embodiment, class label 0.33 identifies an A base; class label 0.66 identifies a C base; class label 1 identifies a T base; and class label 0 identifies a G base. In another embodiment, class label 0.50 identifies an A base; class label 0.75 identifies a C base; class label 1 identifies a T base; and class label 0.25 identifies a G base. In one embodiment, the method includes: deriving a single output value from the output; comparing the single output value to a range of class values ​​corresponding to bases A, C, T, and G based on the comparison; assigning the single output value to a particular class value range; and performing a base call for the target analyte based on the assignment. In one embodiment, the single output value is derived using a sigmoid function, and the single output value is in the range of 0 to 1. In another embodiment, a class value range of 0 to 0.25 represents an A base, a class value range of 0.25 to 0.50 represents a C base, a class value range of 0.50 to 0.75 represents a T base, and a class value range of 0.75 to 1 represents a G base.

[0089] More details about the neural network-based base caller 102 can be found in U.S. Provisional Patent Application No. 62 / 821,766, entitled “ARTIFICIAL INTELLIGENCE-BASED SEQUENCING,” filed on March 21, 2019 (Attorney Docket No. ILLM 1008-9 / IP-1752-PRV), which is incorporated herein by reference.

[0090] Many-to-many base calls

[0091] According to one implementation, the disclosed techniques enable the neural network-based base caller 102 to generate base calls for not only the central sequencing cycle, but also the flanking sequencing cycles, for a given input window. That is, in one implementation, the disclosed techniques simultaneously generate base calls for cycle N, cycle N+1, cycle N-1, cycle N+2, cycle N-2, and so on, for a given input window. That is, a single forward pass / pass / base calling iteration of the neural network-based base caller 102 generates base calls for multiple sequencing cycles within the input window of sequencing cycles, which base calls are referred to herein as "many-to-many base calls."

[0092] The disclosed techniques then use the disclosed many-to-many base calls to generate multiple base calls for the same target sequencing cycle, where the multiple base calls occur across multiple sliding input windows. For example, the target sequencing cycle may occur at different positions across multiple sliding input windows (e.g., starting at position N+2 in the first sliding window, proceeding to position N+1 in the second sliding window, and ending at position N in the third sliding window).

[0093] Multiple base calls are performed on a target sequencing cycle to generate multiple candidates for the correct base call for the target sequencing cycle. The disclosed technology then evaluates the multiple candidates for the correct base call as an aggregate and determines the final base call for the target sequencing cycle. Aggregate analysis techniques such as averaging, consensus, and weighted consensus can be used to select the final base call for the target sequencing cycle.

[0094] Figure 1A 、 Figure 1B and Figure 1C Shown is a disclosed many-to-many base call 100. According to one specific implementation of the disclosed technology, a neural network-based base caller 102 (i.e., base caller 102) processes at least a right wing input, a center input, and a left wing input, and produces at least a right wing output, a center output, and a left wing output.

[0095] Many-to-many base calls 100 are configured to provide data for n sequencing cycles as input to base caller 102 and generate base calls for any number of the n cycles in one iteration of base calling (i.e., one forward pass instance). A target sequencing cycle 108 can be base called n times and can occur / occur / fall at various positions in the n base calling iterations.

[0096] The target sequencing cycle 108 can be the central sequencing cycle in some base calling iterations ( Figure 1B In other iterations, the target sequencing cycle 108 may be a right wing / context sequencing cycle adjacent to the center sequencing cycle ( Figure 1A ), or it may be a left wing / context sequencing cycle adjacent to the center sequencing cycle ( Figure 1C ). The offset to the right or left from the center sequencing cycle can also be variable. That is, the target sequencing cycle 108 in n base calling iterations can fall at the center, immediately to the right of the center, immediately to the left of the center, any offset to the right of the center, any offset to the left of the center, or at any other position in the n base calling iterations. The base calling iteration for the target sequencing cycle can have inputs of sequencing cycles of varying lengths within a given input window of the sequencing cycle, and can also have multiple base call outputs for sequencing cycles of various lengths.

[0097] In one embodiment, the disclosed technology includes accessing a process for generating a set of per-cycle analyte channels for sequencing cycles of a sequencing run; processing a window of the per-cycle analyte channel set in the process for a window of sequencing cycles of the sequencing run by a neural network-based base caller 102, such that the neural network-based base caller 102 processes an object window of the per-cycle analyte channel set in the process for an object window of sequencing cycles of the sequencing run, and generating, using the neural network-based base caller 102, provisional base call predictions for three or more sequencing cycles in the object window of sequencing cycles from a plurality of windows in which a particular sequencing cycle occurs at different positions to generate a provisional base call prediction for the particular sequencing cycle; and determining a base call for the particular sequencing cycle based on the provisional base call predictions.

[0098] In one embodiment, the disclosed technology includes accessing a series of per-cycle analyte channel sets generated for sequencing cycles of a sequencing run; processing, by the neural network-based base caller 102, a window of the per-cycle analyte channel set in the series for a window of sequencing cycles of the sequencing run, such that the neural network-based base caller 102 processes an object window of the per-cycle analyte channel set in the series for the object window of sequencing cycles of the sequencing run and generates base call predictions for two or more sequencing cycles in the object window of sequencing cycles; and processing, by the neural network-based base caller 102, a plurality of windows of the per-cycle analyte channel set in the series for the plurality of windows of sequencing cycles of the sequencing run and generating an output for each of the plurality of windows.

[0099] Each window in the plurality of windows may include a particular set of per-cycle analyte channels for a particular sequencing cycle of a sequencing run. The output for each window in the plurality of windows includes (i) a base call prediction for the particular sequencing cycle and (ii) one or more additional base call predictions for one or more additional sequencing cycles of the sequencing run, thereby generating a plurality of base call predictions for the particular sequencing cycle across the plurality of windows (e.g., generated in parallel or simultaneously by the output layer). Finally, the disclosed technology includes determining a base call for the particular sequencing cycle based on the plurality of base call predictions.

[0100] The right wing input 132 includes current image data 108 for a current sequencing cycle (e.g., cycle 4) of a sequencing run, supplemented with previous image data 104 and previous image data 106 for one or more previous sequencing cycles (e.g., cycle 2 and cycle 3) preceding the current sequencing cycle. The right wing output 142 includes right wing base call predictions 114 for the current sequencing cycle and base call predictions 110 and 112 for the previous sequencing cycles.

[0101] Central input 134 includes current image data 108, supplemented with previous image data 106 (e.g., cycle 3) and subsequent image data 116 for one or more subsequent sequencing cycles following the current sequencing cycle (e.g., cycle 5). Central output 144 includes central base call prediction 120 for the current sequencing cycle and base call predictions 118 and 122 for the previous and subsequent sequencing cycles.

[0102] Left wing input 136 includes current image data 108, supplemented with subsequent image data 116 and subsequent image data 124. Left wing output 146 includes left wing base call predictions 126 for the current sequencing cycle and base call predictions 128 and 130 for subsequent sequencing cycles (e.g., cycle 5 and cycle 6).

[0103] Figure 1D and Figure 1E Different examples of the disclosed many-to-many base calls are shown. Figure 1D and Figure 1EIn the figure, the blue box represents a specific or target sequencing cycle (or its data). Specific sequencing cycles are also taken into account, and the current sequencing cycle is various specific implementations of the disclosed technology. The orange box represents a sequencing cycle (or its data) that is different from the specific sequencing cycle. The green circle represents one or more base calls generated for a specific sequencing cycle. The base calls can be generated by any base caller (such as Illumina's real-time analysis (RTA) software or the disclosed neural network-based base caller 102). The data used for the sequencing cycle can be an image or some other type of input data, such as current readings, voltage changes, pH scale data, etc.

[0104] Go to Figure 1D , the first multi-pair multi-base call example 180 shows three base call iterations 180a, 180b, and 180c, and three input windows / groups of corresponding sequencing cycles w1, w2, and w3 (or their data). In one specific implementation, the base call iteration generates base calls for each sequencing cycle in the corresponding input window of the sequencing cycle. In another specific implementation, the base call iteration generates base calls only for some sequencing cycles (e.g., only specific sequencing cycles) in the sequencing cycle in the corresponding input window of the sequencing cycle. Moreover, a specific sequencing cycle may appear in different positions in the input windows / groups of sequencing cycles w1, w2, and w3. In other specific implementations (not shown), two or more input windows / groups of sequencing cycles may have a specific sequencing cycle in the same position. In addition, the input windows / groups of sequencing cycles w1, w2, and w3 have a specific sequencing cycle as at least one overlapping cycle and also have one or more non-overlapping cycles. That is, the orange boxes at different positions in the different input windows / groups of sequencing cycles represent different non-overlapping cycles. Finally, the three base calling iterations 180a, 180b, and 180c generate three base calls (e.g., three green circles) for a particular sequencing cycle, which can be considered provisional base calls and are subsequently analyzed as an aggregate to make the final base call for the particular sequencing cycle. Figure 2 、 Figure 3 and Figure 4 Different examples of analysis are described later in .

[0105] The second and third examples of multi-to-multi base calls 181 and 182 illustrate that a particular sequencing cycle can be anywhere in the input window / group of sequencing cycles and have any number of right-side flank cycles and left-side flank cycles, or no flank cycles at all (e.g., the third window (w3) in the third-base multi-to-multi base call example 182). The three base calling iterations 181a, 181b, and 181c generate three base calls (e.g., three green circles) for a particular sequencing cycle, which can be considered provisional base calls and are subsequently analyzed as an aggregate to make the final base call for the particular sequencing cycle. Figure 2 、 Figure 3 and Figure 4 Different examples of analysis are described later in . Three base calling iterations 182a, 182b, and 182c generate three base calls (e.g., three green circles) for a particular sequencing cycle, which can be considered provisional base calls and are subsequently analyzed as an aggregate to make the final base call for the particular sequencing cycle. Figure 2 、 Figure 3 and Figure 4 Different examples of analysis are described later in .

[0106] Figure 1E A many-to-many base calling example 183 is shown with five base calling iterations 183a-183e, each base calling iteration generating base call predictions for a particular sequencing cycle by processing five corresponding windows / sets / groups of input data, where the data for the particular sequencing cycle occurs at different positions. The five base calling iterations 183a-183e generate five base calls (e.g., five green circles) for the particular sequencing cycle, which can be considered provisional base calls and are subsequently analyzed as an aggregate to make a final base call for the particular sequencing cycle. Figure 2 、 Figure 3 and Figure 4 Different examples of analysis are described later in .

[0107] Figure 2 、 Figure 3 and Figure 4 Different implementations of a base call generator are shown. Base call generator 202 (e.g., running on a host processor) is coupled to neural network-based base caller 102 (e.g., running on a chip) (e.g., via a PCIe bus or Ethernet or InfiniBand (IB)) and is configured to generate base calls for a current sequencing cycle (e.g., cycle 4) based on the right wing base call predictions, center base call predictions, and left wing base call predictions for the current sequencing cycle.

[0108] The current image data for the current sequencing cycle depicts the intensity emission of the analyte and its surrounding background captured at the current sequencing cycle. The right wing base call prediction 114, the center base call prediction 120, and the left wing base call prediction 126 for the current sequencing cycle (e.g., cycle 4) identify the likelihood of one or more analyte bases spiked into the analyte being A, C, T, and G at the current sequencing cycle. In one embodiment, the likelihood is an exponentially normalized score generated by a softmax layer used as the output layer of the base caller 102.

[0109] In one implementation, the right wing base call prediction 114 for the current sequencing cycle takes into account the prephasing effect between the current sequencing cycle (e.g., cycle 4) and the previous sequencing cycle. In one implementation, the center base call prediction 120 for the current sequencing cycle (e.g., cycle 4) takes into account the prephasing effect between the current sequencing cycle and the previous sequencing cycle, as well as the phasing effect between the current sequencing cycle and the subsequent sequencing cycle. In one implementation, the left wing base call prediction 126 for the current sequencing cycle (e.g., cycle 4) takes into account the phasing effect between the current sequencing cycle and the subsequent sequencing cycle.

[0110] like Figure 2 As shown, the base call generator is further configured to include an averager 204 that sums the likelihoods across the right wing base call predictions 114, the center base call predictions 120, and the left wing base call predictions 126 for the current sequencing cycle (e.g., cycle 4) on a base-by-base basis; determines a base-by-base average 212 based on the base-by-base summation; and generates a base call 214 for the current sequencing cycle (e.g., cycle 4) based on a highest one of the base-by-base averages (e.g., 0.38).

[0111] like Figure 3 As shown, the base call generator is further configured to include a consensus sensor 304 that determines a preliminary base call for each of the right wing base call predictions 114, the center base call predictions 120, and the left wing base call predictions 126 for the current sequencing cycle (e.g., cycle 4) based on a highest one of the likelihoods, thereby producing a preliminary base call sequence 306, and generates a base call for the current sequencing cycle based on a most common base call 308 in the preliminary base call sequence.

[0112] like Figure 4As shown, the base call generator is further configured to include a weighted consensus sensor 404 that determines a preliminary base call for each of the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle based on a highest likelihood among the likelihoods, thereby generating a preliminary base call sequence 406; applies a base-by-base weight 408 to a corresponding one of the preliminary base calls in the preliminary base call sequence and generates a weighted preliminary base call sequence 410; and generates a base call for the current sequencing cycle (e.g., cycle 4) based on the most weighted base call 412 in the weighted preliminary base call sequence. In some implementations, for example, the base-by-base weight 408 is preset on a cycle-by-cycle basis. In other implementations, for example, the base-by-base weight 408 is learned using a least squares method.

[0113] exist Figure 6 In one illustrated implementation, the disclosed technology includes accessing current image data for a current sequencing cycle of a sequencing run, previous image data for one or more previous sequencing cycles before the current sequencing cycle, and subsequent image data for one or more subsequent sequencing cycles after the current sequencing cycle (act 602); processing different groups of the current image data, previous image data, and subsequent image data through a neural network-based base caller and generating a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle (act 612); and generating a base call for the current sequencing cycle based on the first base call prediction, the second base call prediction, and the third base call prediction (act 622).

[0114] In one embodiment, the different groups include a first group including current image data and previous image data; a second group including current image data, previous image data and subsequent image data; and a third group including current image data and subsequent image data.

[0115] In one specific implementation, the disclosed technology includes processing a first grouping by a neural network-based base caller to generate a first base call prediction; processing a second grouping by the neural network-based base caller to generate a second base call prediction; and processing a third grouping by the neural network-based base caller to generate a third base call prediction.

[0116] In one implementation, the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle identify the likelihood of the bases in one or more analytes spiked into the analytes being A, C, T, and G at the current sequencing cycle.

[0117] In one specific implementation, the disclosed technology includes generating a base call for the current sequencing cycle by performing a base-by-base summation across a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle; determining a base-by-base average based on the base-by-base summation; and generating a base call for the current sequencing cycle based on a highest one of the base-by-base averages.

[0118] In one specific implementation, the disclosed technology includes generating base calls for the current sequencing cycle by determining preliminary base calls for each of a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle based on a highest likelihood among the likelihoods, thereby producing a preliminary base call sequence; and generating base calls for the current sequencing cycle based on a most common base call in the preliminary base call sequence.

[0119] In one specific implementation, the disclosed technology includes generating base calls for the current sequencing cycle by determining a preliminary base call for each of a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle based on a highest one of the likelihoods, thereby producing a sequence of preliminary base calls; applying a base-by-base weight to a corresponding one of the preliminary base calls in the sequence of preliminary base calls and producing a sequence of weighted preliminary base calls; and generating the base calls for the current sequencing cycle based on a most weighted base call in the sequence of weighted preliminary base calls.

[0120] In one implementation, referred to as "multi-cycle training, single-cycle inference," base caller 102 is trained to generate two or more base call predictions for two or more sequencing cycles during training using a base call generator, but only generates base call predictions for a single sequencing cycle during inference.

[0121] In one implementation, referred to as “multi-cycle training, multi-cycle inference,” base caller 102 is trained to generate two or more base call predictions for two or more sequencing cycles during training, and does the same during inference using base call generator 202 .

[0122] Multi-cycle gradient backpropagation

[0123] Figure 5 FIG. 5 shows a specific implementation of the disclosed “Multi-cycle gradient back propagation 500”. Figure 5As shown, multi-to-multi base calling 100 is further configured to include a trainer that calculates errors 512, 532, and 552 between base calls generated by base call generator 202 for a current sequencing cycle (e.g., cycle 3), a previous sequencing cycle (e.g., cycle 2), and a subsequent sequencing cycle (e.g., cycle 4) based on the right wing output 506, the center output 504, and the left wing output 502 of the neural network-based base caller 102 and corresponding ground truth base calls 554, 534, and 514, determines corresponding gradients 542, 522, and 562 for the current sequencing cycle, the previous sequencing cycle, and the subsequent sequencing cycle based on the errors, and updates parameters of the neural network-based base caller by back-propagating the gradients.

[0124] Technical effects / advantages

[0125] Figure 7 The technical effects and advantages of the disclosed technology are shown.

[0126] "Multi-loop training, single-loop inference" is specifically implemented in Figure 7 The DL 3C median value was used in the study and the base calling error rate was improved by 8% using real-time analysis base calling software based on traditional non-neural networks.

[0127] "Multi-loop training and multi-loop inference" are specifically implemented in Figure 7 The DL 3C median value is referred to as the "DL 3C average" in the literature and improves the base call error rate by another 8%.

[0128] Base Call Sequencing cycles multiple times to improve base call accuracy and detect and resolve base call discrepancies and ambiguous base calls.

[0129] Multiple cycles of gradient backpropagation also improve the gradient of the base caller 102 and its base calling accuracy throughout the base calling training task.

[0130] Sequencing system

[0131] Figure 8A and Figure 8B A specific implementation of a sequencing system 800A is depicted. The sequencing system 800A includes a configurable processor 846. The configurable processor 846 implements the base calling technology disclosed herein. A sequencing system is also referred to as a "sequencer."

[0132] Sequencing system 800A can be operated to obtain any information or data related to at least one of a biological substance or a chemical substance. In some implementations, sequencing system 800A is a workstation that can be similar to a desktop device or desktop computer. For example, most (or all) systems and components for performing the desired reactions can be located within a common housing 802.

[0133] In certain embodiments, the sequencing system 800A is a nucleic acid sequencing system configured for various applications, including but not limited to de novo sequencing, resequencing of whole genomes or target genome regions, and metagenomics. The sequencer can also be used for DNA or RNA analysis. In some embodiments, the sequencing system 800A can also be configured to generate reaction sites in a biosensor. For example, the sequencing system 800A can be configured to receive a sample and generate surface-attached clusters of cloned amplified nucleic acids derived from the sample. Each cluster can constitute or be part of a reaction site in a biosensor.

[0134] The exemplary sequencing system 800A may include a system receptacle or interface 810 configured to interact with a biosensor 812 to perform a desired reaction within the biosensor 812. Figure 8A In the description, the biosensor 812 is loaded into the system receptacle 810. However, it should be understood that a cartridge including the biosensor 812 can be inserted into the system receptacle 810, and in some cases, the cartridge can be temporarily or permanently removed. As described above, the cartridge can include, among other things, a fluid control component and a fluid storage component.

[0135] In a specific embodiment, the sequencing system 800A is configured to perform a large number of parallel reactions within the biosensor 812. The biosensor 812 includes one or more reaction sites where the desired reaction can occur. The reaction site can, for example, be fixed to the solid surface of the biosensor or fixed to a bead (or other removable substrate) within the corresponding reaction chamber of the biosensor. The reaction site can include, for example, a cluster of cloned amplified nucleic acids. The biosensor 812 can include a solid-state imaging device (e.g., a CCD or CMOS imaging device) and a flow cell mounted thereon. The flow cell can include one or more flow channels that receive a solution from the sequencing system 800A and direct the solution to the reaction site. Optionally, the biosensor 812 can be configured to engage a thermal element for transferring heat energy into or from the flow channel.

[0136] The sequencing system 800A may include various components, assemblies, and systems (or subsystems) that interact with each other to execute a predetermined method or assay protocol for biological or chemical analysis. For example, the sequencing system 800A includes a system controller 806 that can communicate with the various components, assemblies, and subsystems of the sequencing system 800A and the biosensor 812. For example, in addition to the system receptacle 810, the sequencing system 800A may also include a fluid control system 808 to control the flow of fluids throughout the fluid network of the sequencing system 800A and the biosensor 812; a fluid storage system 814 configured to accommodate all fluids (e.g., gases or liquids) that can be used by the bioassay system; a temperature control system 804 that can regulate the temperature of the fluids in the fluid network, the fluid storage system 814, and / or the biosensor 812; and an illumination system 816 configured to illuminate the biosensor 812. As described above, if a cartridge with the biosensor 812 is loaded into the system receptacle 810, the cartridge may also include a fluid control component and a fluid storage component.

[0137] As also shown, the sequencing system 800A may include a user interface 818 for interacting with a user. For example, the user interface 818 may include a display 820 for displaying or requesting information from a user and a user input device 822 for receiving user input. In some implementations, the display 820 and the user input device 822 are the same device. For example, the user interface 818 may include a touch-sensitive display that is configured to detect the presence of individual touches and also identify the location of the touch on the display. However, other user input devices 822 may be used, such as a mouse, touchpad, keyboard, keypad, handheld scanner, voice recognition system, motion recognition system, etc. As will be discussed in more detail below, the sequencing system 800A may communicate with various components including a biosensor 812 (e.g., in the form of a cartridge) to perform the desired reaction. The sequencing system 800A may also be configured to analyze the data obtained from the biosensor to provide the desired information to the user.

[0138] The system controller 806 may include any processor-based or microprocessor-based system, including the use of a microcontroller, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a coarse-grained reconfigurable architecture (CGRA), a logic circuit, and any other circuit or processor capable of performing the functions described herein. The above examples are merely exemplary and are not intended to limit the definition and / or meaning of the term system controller in any way. In an exemplary embodiment, the system controller 806 executes an instruction set stored in one or more storage elements, memories, or modules to obtain at least one of the detection data and the analysis detection data. The detection data may include multiple pixel signal sequences so that the pixel signal sequence of each sensor (or pixel) from millions of sensors (or pixels) can be detected within many base call cycles. The storage element may be in the form of an information source or physical memory element in the sequencing system 800A.

[0139] The instruction set may include various commands that instruct the sequencing system 800A or the biosensor 812 to perform specific operations (such as the methods and processes of the various specific implementations described herein). The instruction set may be in the form of a software program that may form part of one or more tangible non-transitory computer-readable media. As used herein, the terms "software" and "firmware" are interchangeable and include any computer program stored in a memory for execution by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are exemplary only and do not limit the memory types that can be used to store computer programs.

[0140] The software can be in various forms, such as system software or application software. In addition, the software can be in the form of a collection of independent programs, or in the form of a program module or a part of a program module within a larger program. The software can also include modular programming in the form of object-oriented programming. After obtaining the detection data, the detection data can be automatically processed by the sequencing system 800A, processed in response to user input, or processed in response to a request (e.g., a remote request through a communication link) proposed by another processing machine. In the specific implementation shown, the system controller 806 includes an analysis module 844. In other specific implementations, the system controller 806 does not include the analysis module 844, but can access the analysis module 844 (e.g., the analysis module 844 can be hosted separately on the cloud).

[0141] The system controller 806 can be connected to the biosensor 812 and other components of the sequencing system 800A via a communication link. The system controller 806 can also be communicatively connected to an off-site system or server. The communication link can be hardwired, wired, or wireless. The system controller 806 can receive user input or commands from the user interface 818 and the user input device 822.

[0142] Fluid control system 808 includes a fluid network and is configured to direct and regulate the flow of one or more fluids through the fluid network. The fluid network can be in fluid communication with biosensor 812 and fluid storage system 814. For example, a selected fluid can be drawn from fluid storage system 814 and directed to biosensor 812 in a controlled manner, or a fluid can be drawn from biosensor 812 and directed toward, for example, a waste reservoir in fluid storage system 814. Although not shown, fluid control system 808 can include a flow sensor that detects the flow rate or pressure of the fluid within the fluid network. The sensor can communicate with system controller 806.

[0143] The temperature control system 804 is configured to regulate the temperature of fluids at different areas of the fluid network, fluid storage system 814, and / or biosensor 812. For example, the temperature control system 804 may include a thermal cycler that interfaces with the biosensor 812 and controls the temperature of the fluids flowing along the reaction sites in the biosensor 812. The temperature control system 804 may also regulate the temperature of solid elements or components of the sequencing system 800A or the biosensor 812. Although not shown, the temperature control system 804 may include sensors for detecting the temperature of the fluids or other components. The sensors may communicate with the system controller 806.

[0144] Fluid storage system 814 is in fluid communication with biosensor 812, and can store various reaction components or reactants for carrying out desired reaction therein. Fluid storage system 814 can also store fluid for washing or cleaning fluid network and biosensor 812 and for diluting reactants. For example, fluid storage system 814 can include various reservoirs to store samples, reagents, enzymes, other biomolecules, buffer solutions, aqueous solutions and non-polar solutions, etc. In addition, fluid storage system 814 can also include waste storage for receiving waste from biosensor 812. In the specific implementation including cartridge, cartridge can include one or more of fluid storage system, fluid control system or temperature control system. Therefore, one or more components related to those systems described herein can be contained in cartridge housing. For example, cartridge can have various reservoirs to store samples, reagents, enzymes, other biomolecules, buffer solutions, aqueous solutions and non-polar solutions, waste, etc. Therefore, one or more of fluid storage system, fluid control system or temperature control system can be removably engaged with bioassay system via cartridge or other biosensors.

[0145] The illumination system 816 may include a light source (e.g., one or more LEDs) and a plurality of optical components for illuminating the biosensor. Examples of light sources may include lasers, arc lamps, LEDs, or laser diodes. The optical components may be, for example, reflectors, dichroic mirrors, beam splitters, collimators, lenses, filters, wedge mirrors, prisms, reflectors, detectors, and the like. In an implementation using an illumination system, the illumination system 816 may be configured to direct excitation light to the reaction site. As an example, a fluorophore may be excited by green wavelength light, so the wavelength of the excitation light may be approximately 532 nm. In one implementation, the illumination system 816 is configured to generate illumination parallel to the surface normal of the surface of the biosensor 812. In another implementation, the illumination system 816 is configured to generate illumination at an angle relative to the surface normal of the surface of the biosensor 812. In yet another implementation, the illumination system 816 is configured to generate illumination at multiple angles, including some parallel illumination and some angled illumination.

[0146] The system socket or interface 810 is configured to engage the biosensor 812 in at least one of a mechanical, electrical, and fluidic manner. The system socket 810 can hold the biosensor 812 in a desired orientation to facilitate fluid flow through the biosensor 812. The system socket 810 can also include electrical contacts configured to engage the biosensor 812 so that the sequencing system 800A can communicate with the biosensor 812 and / or provide power to the biosensor 812. In addition, the system socket 810 can include a fluid port (e.g., a nozzle) configured to engage the biosensor 812. In some implementations, the biosensor 812 is removably coupled to the system socket 810 mechanically, electrically, and fluidically.

[0147] In addition, sequencing system 800A can communicate remotely with other systems or networks or with other bioassay systems 800 A. Detection data obtained by bioassay system 800A can be stored in a remote database.

[0148] Figure 8B Yes, you can Figure 8A 806 is a block diagram of a system controller 806 used in a system of FIG. In one embodiment, the system controller 806 includes one or more processors or modules that can communicate with each other. Each of the processors or modules can include an algorithm (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or a sub-algorithm for performing a specific process. The system controller 806 is conceptually shown as a collection of modules, but can be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 806 can be implemented using an off-the-shelf PC with a single processor or multiple processors, where functional operations are distributed between the processors. As a further option, the modules described below can be implemented using a hybrid configuration, where certain modular functions are performed using dedicated hardware, while other modular functions are performed using an off-the-shelf PC, etc. The modules can also be implemented as software modules within a processing unit.

[0149] During operation, the communication port 850 can send data to the biosensor 812 ( Figure 8A ) and / or subsystems 808, 814, 804 ( Figure 8A ) transmits information (e.g., commands) to or receives information (e.g., data) from the user interface 818. In a specific implementation, the communication port 850 can output a plurality of pixel signal sequences. The communication link 834 can be connected to the user interface 818 ( Figure 8A) receives user input and transmits data or information to the user interface 818. Data from the biosensor 812 or subsystems 808, 814, 804 can be processed in real time during a biometric session by the system controller 806. Additionally or alternatively, data can be temporarily stored in system memory during a biometric session and processed at a slower rate than real time or offline operation.

[0150] like Figure 8B As shown, the system controller 806 may include a plurality of modules 824-848 that communicate with a main control module 824 and a central processing unit (CPU) 852. The main control module 824 may communicate with a user interface 818 ( Figure 8A Although modules 824-848 are shown as communicating directly with the main control module 824, modules 824-848 may also communicate directly with each other, the user interface 818, and the biosensor 812. In addition, modules 824-848 may communicate with the main control module 824 through other modules.

[0151] A plurality of modules 824-848 include system modules 828-832, 826 that communicate with subsystems 808, 814, 804, and 816, respectively. Fluid control module 828 can communicate with fluid control system 808 to control valves and flow sensors of fluid network to control the flow of one or more fluids through the fluid network. Fluid storage module 830 can notify the user when the fluid volume is low or when the waste reservoir is at or near capacity. Fluid storage module 830 can also communicate with temperature control module 832 so that fluid can be stored at a desired temperature. Lighting module 826 can communicate with lighting system 816 to illuminate the reaction site at a specified time during the protocol, such as after desired reaction (e.g., in conjunction with an event) has occurred. In some implementations, lighting module 826 can communicate with lighting system 816 to illuminate the reaction site at a specified angle.

[0152] The plurality of modules 824-848 may also include a device module 836 that communicates with the biosensor 812 and an identification module 838 that determines identification information associated with the biosensor 812. The device module 836 may, for example, communicate with the system receptacle 810 to confirm that the biosensor has established electrical and fluidic connection with the sequencing system 800A. The identification module 838 may receive a signal identifying the biosensor 812. The identification module 838 may use the identity of the biosensor 812 to provide additional information to the user. For example, the identification module 838 may determine and subsequently display a batch number, a manufacturing date, or a recommended protocol to be run with the biosensor 812.

[0153] The plurality of modules 824-848 also includes an analysis module 844 (also referred to as a signal processing module or signal processor) that receives and analyzes signal data (e.g., image data) from the biosensor 812. The analysis module 844 includes a memory (e.g., RAM or flash memory) for storing detection / image data. The detection data may include a plurality of pixel signal sequences so that a pixel signal sequence from each sensor (or pixel) in millions of sensors (or pixels) may be detected within many base calling cycles. The signal data may be stored for subsequent analysis or may be transmitted to the user interface 818 to display the desired information to the user. In some implementations, the signal data may be processed by a solid-state imaging device (e.g., a CMOS image sensor) before the analysis module 844 receives the signal data.

[0154] The analysis module 844 is configured to obtain image data from the photodetector at each sequencing cycle of the plurality of sequencing cycles. The image data is derived from an emission signal detected by the photodetector, and the image data for each sequencing cycle of the plurality of sequencing cycles is processed by the base caller 102, and base calls are generated for at least some of the analytes at each sequencing cycle of the plurality of sequencing cycles. The photodetector can be part of one or more overhead cameras (e.g., a CCD camera of Illumina's GAIIx captures an image of the clusters on the biosensor 812 from the top), or can be part of the biosensor 812 itself (e.g., a CMOS image sensor of Illumina's iSeq is located below the clusters on the biosensor 812 and captures an image of the clusters from the bottom).

[0155] The output of the photodetector is a sequencing image, each of which depicts the intensity emission of a cluster and its surrounding background. The sequencing image depicts the intensity emission generated by the incorporation of nucleotides into the sequence during sequencing. The intensity emission comes from the associated analyte and its surrounding background. The sequencing image is stored in memory 848.

[0156] The protocol module 840 and the protocol module 842 communicate with the main control module 824 to control the operation of the subsystems 808, 814 and 804 when performing a predetermined determination protocol. The protocol module 840 and the protocol module 842 can include an instruction set for instructing the sequencing system 800A to perform a specific operation according to a predetermined protocol. As shown in the figure, the protocol module can be a sequencing by synthesis (SBS) module 840, which is configured to issue various commands for executing a sequencing by synthesis process. In SBS, the extension of the nucleic acid primer along the nucleic acid template is monitored to determine the sequence of the nucleotides in the template. The basic chemical process can be polymerization (for example, catalyzed by a polymerase) or connection (for example, catalyzed by a ligase). In a specific polymerase-based SBS implementation, fluorescently labeled nucleotides are added to the primers (so that the primers are extended) in a template-dependent manner so that the detection of the order and type of nucleotides added to the primers can be used to determine the sequence of the template. For example, in order to start the first SBS cycle, a command can be issued to deliver one or more labeled nucleotides, DNA polymerase, etc. to / through a flow cell containing a nucleic acid template array. Nucleic acid template can be located at corresponding reaction site.Those reaction sites that wherein primer extension causes the nucleotide of labelling to be incorporated can be detected by imaging event.During imaging event, illumination system 816 can provide excitation light to reaction site.Optionally, nucleotide can also comprise the reversible termination property that just stops further primer extension once nucleotide is added to primer.For example, nucleotide analog with reversible terminator part can be added to primer so that subsequent extension just occurs until delivery deblocking agent to remove this part.Therefore, for the specific implementation using reversible termination, can issue order so that deblocking agent is delivered to circulation cell (before or after detection occurs).Can issue one or more orders to realize the washing between each delivery step.Then can repeat this cycle n times, so that primer is extended n nucleotides, so that the sequence of detection length is n. Exemplary sequencing technologies are described in, e.g., Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, US 7,057,026, WO 91 / 06678, WO 07 / 123744, US 7,329,492, US 7,211,414, US 7,315,019, US 7,405,281, and US 2008 / 014708082, each of which is incorporated herein by reference.

[0157] For the nucleotide delivery step of SBS circulation, the nucleotide of a single type can be delivered once, or a plurality of different nucleotide types (for example, A, C, T and G together) can be delivered. For the nucleotide delivery configuration in which only a single type of nucleotide is present once, different nucleotides do not need to have different labels, because they can be distinguished based on the time interval inherent in individualized delivery. Therefore, sequencing method or device can use monochromatic detection. For example, the excitation source only needs to provide the excitation in a single wavelength or a single wavelength range. For the nucleotide delivery configuration in which delivery causes a plurality of different nucleotides to be present in the flow cell simultaneously, the site of incorporating different nucleotide types can be distinguished based on the different fluorescent labels attached to the corresponding nucleotide type in the mixture. For example, four different nucleotides can be used, and each nucleotide has one of four different fluorophores. In a specific implementation, the excitation in four different regions of the spectrum can be used to distinguish four different fluorophores. For example, four different excitation radiation sources can be used. Alternatively, less than four different excitation sources can be used, but the optical filtering of the excitation radiation from a single source can be used to produce excitation radiation of different ranges at the flow cell.

[0158] In some specific implementations, can detect in the mixture with four kinds of different nucleotides and be less than four kinds of different colors.For example, nucleotide pair can detect under the same wavelength, but based on the intensity difference of a member in the pair relative to another member, or based on the change (for example, by chemical modification, photochemical modification or physical modification) of a member causing to be compared with the signal of another member of the pair detected that obvious signal appears or disappears and distinguishes.For using the detection that is less than four kinds of colors to distinguish the exemplary device and method of four different nucleotides and be described in for example U.S. patent application serial number 61 / 538,294 and 61 / 619,878, it is incorporated herein by reference in its entirety.The U.S. application 13 / 624,200 that submitted on September 21, 2012 is also incorporated to by reference in its entirety.

[0159] A plurality of protocol modules can also include sample preparation (or generation) module 842, which is configured to issue commands to fluid control system 808 and temperature control system 804 to increase the product in biosensor 812. For example, biosensor 812 can be attached to sequencing system 800A. Amplification module 842 can issue instructions to fluid control system 808 to deliver necessary amplification components to the reaction chamber in biosensor 812. In other specific implementations, the reaction site may have included some components for amplification, such as template DNA and / or primers. After amplification components are delivered to the reaction chamber, amplification module 842 can instruct temperature control system 804 to cycle through different temperature stages according to known amplification protocols. In some specific implementations, amplification and / or nucleotide incorporation isothermal is carried out.

[0160] The SBS module 840 can issue a command to perform bridge PCR, in which clusters of clonal amplicons are formed on localized regions within the channels of the flow cell. After amplicons are generated by bridge PCR, the amplicons can be "linearized" to prepare single-stranded template DNA or sstDNA, and sequencing primers can be hybridized to universal sequences flanking the region of interest. For example, a sequencing-by-synthesis method based on reversible terminators can be used as described above or as follows.

[0161] Each base call or sequencing cycle can extend sstDNA by a single base, which can be completed, for example, by using a mixture of modified DNA polymerase and four types of nucleotides. Different types of nucleotides can have unique fluorescent labels, and each nucleotide can also have a reversible terminator that only allows single-base incorporation to occur in each cycle. After a single base is added to sstDNA, the excitation light can be incident on the reaction site and can detect fluorescent emission. After detection, the fluorescent label and terminator can be chemically cut from the sstDNA. Next, another similar base call or sequencing cycle can be performed. In this type of sequencing protocol, the SBS module 840 can instruct the fluid control system 808 to guide reagents and enzyme solution to flow through the biosensor 812. Exemplary reversible terminator-based SBS methods that can be used with the devices and methods described herein are described in U.S. Patent Application Publication No. 2007 / 0166705 Al, U.S. Patent Application Publication No. 2006 / 0188901 Al, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439 Al, U.S. Patent Application Publication No. 2006 / 02814714709 Al, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 014700900 Al, PCT Publication No. WO 06 / 08B199, and PCT Publication No. WO 07 / 01470251, each of which is incorporated herein by reference in its entirety. Exemplary reagents for reversible terminator-based SBS are described in: US 7,541,444, US 7,057,026, US 7,414,14716, US 7,427,673, US 7,566,537, US 7,592,435, and WO 07 / 14835368, each of which is incorporated herein by reference in its entirety.

[0162] In some implementations, the amplification module and the SBS module can operate in a single assay protocol, where, for example, a template nucleic acid is amplified and then sequenced within the same cartridge.

[0163] Sequencing system 800A can also allow the user to reconfigure the assay protocol. For example, sequencing system 800A can provide the user with an option to modify the determined protocol via user interface 818. For example, if biosensor 812 is determined to be used for amplification, sequencing system 800A can request a temperature for the annealing cycle. Furthermore, if the user has provided user input that is generally unacceptable for the selected assay protocol, sequencing system 800A can issue a warning to the user.

[0164] In a specific implementation, biosensor 812 includes millions of sensors (or pixels), each of which generates multiple pixel signal sequences during subsequent base calling cycles. Analysis module 844 detects the multiple pixel signal sequences based on the row-by-row and / or column-by-column positions of the sensors on the sensor array and attributes them to the corresponding sensors (or pixels).

[0165] Figure 9 is a simplified block diagram of a system for analyzing sensor data (such as base call sensor output) from sequencing system 800A. Figure 9 In an example of , the system includes a configurable processor 846. The configurable processor 846 can execute a base caller (e.g., a neural network-based base caller 102) in coordination with a runtime program executed by a central processing unit (CPU) 852 (i.e., a host processor). The sequencing system 800A includes a biosensor 812 and a flow cell. The flow cell may include one or more blocks in which clusters of genetic material are exposed to a sequence of an analyte flow that is used to cause a reaction in the cluster to identify the bases in the genetic material. The sensor senses the reaction of each cycle of the sequence in each block of the flow cell to provide block data. Genetic sequencing is a data-intensive operation that converts base call sensor data into a base call sequence for each cluster of genetic material sensed during the base calling operation.

[0166] The system in this example includes a CPU 852 that executes a runtime program to coordinate base calling operations, a memory 848B for storing the sequence of the block data array, the base-called reads generated by the base-calling operations, and other information used in the base-calling operations. In addition, in this illustration, the system includes a memory 848A to store a configuration file (or files) such as an FPGA bit file and model parameters for configuring and reconfiguring the neural network of the configurable processor 846 and executing the neural network. The sequencing system 800A can include a program for configuring the configurable processor, and in some embodiments, the reconfigurable processor, to execute the neural network.

[0167] Sequencing system 800A is coupled to configurable processor 846 via bus 902. Bus 902 can be implemented using high-throughput technology, such as, in one example, bus technology compatible with the PCIe standard (Peripheral Component Interconnect Express) currently maintained and developed by PCI-SIG (PCI Special Interest Group). Also in this example, memory 848A is coupled to configurable processor 846 via bus 906. Memory 848A can be on-board memory provided on a circuit board with configurable processor 846. Memory 848A is used by configurable processor 846 to access working data used in base calling operations at high speed. Bus 906 can also be implemented using high-throughput technology, such as bus technology compatible with the PCIe standard.

[0168] Configurable processors, including field programmable gate arrays (FPGAs), coarse-grained reconfigurable arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than would be possible using a general-purpose processor executing a computer program. Configuring a configurable processor involves compiling a functional description to produce a configuration file, sometimes called a bitstream or bitfile, and distributing the configuration file to the configurable elements on the processor. The configuration file defines the logical functions to be performed by the configurable processor by configuring the circuit to set data flow patterns, the use of distributed memory and other on-chip memory resources, the contents of lookup tables, the operation of configurable logic blocks and configurable execution units (such as multiply-accumulate units, configurable interconnects, and other elements of the configurable array). A configurable processor is reconfigurable if the configuration file can be changed in the field by changing a loaded configuration file. For example, the configuration file can be stored in volatile SRAM elements, non-volatile read-write memory elements, and combinations thereof, distributed across an array of configurable elements on a configurable or reconfigurable processor. A variety of commercially available configurable processors are suitable for base calling operations as described herein. Examples include Google's Tensor Processing Unit (TPU). TM , rack solutions (such as GX4 Rackmount Series TM 、GX9 Rackmount Series TM ), NVIDIA DGX-1 TM , Microsoft's Stratix V FPGA TM , Graphcore’s Intelligent Processor Unit (IPU) TM 、Qualcomm's Snapdragon processors TM Zeroth Platform TM、NVIDIA's Volta TM 、NVIDIA's DRIVE PX TM , NVIDIA's JETSON TX1 / TX2MODULE TM 、Intel's Nirvana TM 、Movidius VPU TM ,Fujitsu DPI TM ARM's DynamicIQ TM 、IBM TrueNorth TM , with Testa V100s TM Lambda GPU server, Xilinx Alveo TM U200, Xilinx Alveo TM U250, Xilinx Alveo TM U280, Intel / Altera Stratix TM GX2800、Intel / AlteraStratix TM GX2800 and Intel Stratix TM GX10M. In some examples, the host CPU can be implemented on the same integrated circuit as the configurable processor.

[0169] The embodiments described herein implement the neural network-based base caller 102 using a configurable processor 846. The configuration file for the configurable processor 846 can be implemented by specifying the logic functions to be performed using a high-level description language (HDL) or a register transfer level (RTL) language specification. The specification can be compiled using resources designed for the selected configurable processor to generate the configuration file. To generate a design for an application-specific integrated circuit (ASIC) that may not be a configurable processor, the same or similar specification can be compiled.

[0170] Thus, in all embodiments described herein, alternatives to the configurable processor 846 include a configured processor comprising a dedicated ASIC or application specific integrated circuit or group of integrated circuits, or a system on a chip (SOC) device, or a graphics processing unit (GPU) processor or a coarse-grained reconfigurable architecture (CGRA) processor, configured to perform neural network-based base calling operations as described herein.

[0171] In general, configurable processors and processors configured as described herein that are configured to perform operations of a neural network are referred to herein as neural network processors.

[0172] In this example, the configurable processor 846 is configured by a configuration file loaded by a program executed by the CPU 852, or by other sources that configure an array of configurable elements 916 (e.g., configuration logic blocks (CLBs), such as lookup tables (LUTs), flip-flops, computational processing units (PMUs) and computational memory units (CMUs), configurable I / O blocks, and programmable interconnects) on the configurable processor to perform base calling functions. In this example, the configuration includes data flow logic 908, which is coupled to bus 902 and bus 906 and performs functions for distributing data and control parameters between elements used in base calling operations.

[0173] In addition, the configurable processor 846 is configured with base calling execution data flow logic 908 to execute the neural network-based base caller 102. The data flow logic 908 includes multi-cycle execution clusters (e.g., 914), which in this example include execution cluster 1 through execution cluster X. The number of multi-cycle execution clusters can be selected based on a trade-off between the desired throughput of the operations involved and the available resources on the configurable processor 846.

[0174] The multi-cycle execution clusters are coupled to the dataflow logic 908 via a dataflow path 910 implemented using configurable interconnect and memory resources on the configurable processor 846. In addition, the multi-cycle execution clusters are coupled to the dataflow logic 908 via a control path 912 implemented using, for example, configurable interconnect and memory resources on the configurable processor 846, which provides control data indicating available execution clusters, input cells ready to be provided to the available execution clusters for executing operations of the neural network-based base caller 102, ready to provide trained parameters to the neural network-based base caller 102, ready to provide output patches of base call classification data, and other control data for executing the neural network-based base caller 102.

[0175] The configurable processor 846 is configured to use the trained parameters to execute the operation of the neural network-based base caller 102 to generate classification data for the sensing cycle of the base calling operation. The operation of the neural network-based base caller 102 is executed to generate classification data for the subject sensing cycle of the base calling operation. The operation of the neural network-based base caller 102 operates on a sequence (a digital N array of block data including corresponding sensing cycles from N sensing cycles), wherein the N sensing cycles provide sensor data for different base calling operations for a base position of each operation in the time series in the example described herein. Optionally, if necessary, some of the N sensing cycles may be out of order according to the specific neural network model being executed. The number N can be any number greater than 1. In some examples described herein, the sensing cycles in the N sensing cycles represent a set of sensing cycles of at least one sensing cycle before the subject sensing cycle and at least one sensing cycle after the subject cycle (subject cycle) in the time series. Examples where the number N is an integer equal to or greater than five are described herein.

[0176] The data flow logic 908 is configured to move tile data and at least some trained parameters of the model parameters from the memory 848A to the configurable processor 846 for the operation of the neural network-based base caller 102 using input units for a given operation, the input units including tile data for spatially aligned patches of the N arrays. The input units may be moved via direct memory access operations in one DMA operation or in smaller units that are moved during available time slots in coordination with the execution of the deployed neural network.

[0177] Block data for sensing cycles as described herein may include a sensor data array having one or more features. For example, the sensor data may include two images that are analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The block data may also include metadata about the images and sensors. For example, in an embodiment of a base calling operation, the block data may include information about the alignment of the image with the cluster, such as distance-to-center information indicating the distance of each pixel in the sensor data array from the center of the cluster of genetic material on the block.

[0178] During execution of the neural network-based base caller 102, as described below, the tile data may also include data generated during execution of the neural network-based base caller 102, referred to as intermediate data, that can be reused rather than recalculated during operation of the neural network-based base caller 102. For example, during execution of the neural network-based base caller 102, the data flow logic 908 may write the intermediate data to the memory 848A in place of the sensor data for a given patch of the tile data array. Embodiments similar to this are described in more detail below.

[0179] As shown, a system for analyzing base call sensor output is described. The system includes a memory (e.g., 848A) accessible by a runtime program, the memory storing block data comprising sensor data for blocks of sensing cycles from base calling operations. Additionally, the system includes a neural network processor, such as a configurable processor 846 that can access the memory. The neural network processor is configured to execute a neural network using trained parameters to generate classification data for the sensing cycles. As described herein, the neural network executes a sequence of N arrays of block data from corresponding sensing cycles of N sensing cycles (including a subject cycle) to generate classification data for the subject cycle. Data flow logic 908 is provided to move the block data and trained parameters from the memory to the neural network processor using input units (including data from spatially aligned patches of N arrays from corresponding sensing cycles of N sensing cycles) for execution of the neural network.

[0180] Additionally, a system is described in which a neural network processor has access to a memory and includes a plurality of execution clusters, wherein an execution cluster in the plurality of execution clusters is configured to execute a neural network. Data flow logic 908 can access the memory and an execution cluster in the plurality of execution clusters to provide an input unit of tile data to an available execution cluster in the plurality of execution clusters, the input unit including a number N spatially aligned patches from a tile data array for a corresponding sensing cycle (including a subject sensing cycle), and cause the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patches for the subject sensing cycle, where N is greater than 1.

[0181] like Figure 9 and Figure 10As shown, in one embodiment, the disclosed technology includes an artificial intelligence-based system for base calling. The system includes a host processor, a memory, a configurable processor, and data flow logic. The memory can be accessed by the host processor and stores image data for sequencing cycles of a sequencing run, wherein the current image data for the current sequencing cycle of the sequencing run depicts the intensity emission of the analyte captured in the current sequencing cycle and its surrounding background; the configurable processor can access the memory, including multiple execution clusters, and the execution clusters in the multiple execution clusters are configured to execute a neural network; the data flow logic can access the memory and the execution clusters in the multiple execution clusters and is configured to provide the current image data, the image data for the current sequencing cycle, and the image data before the current sequencing cycle to an available execution cluster in the multiple execution clusters. Previous image data for one or more previous sequencing cycles, and subsequent image data for one or more subsequent sequencing cycles after the current sequencing cycle, cause the execution cluster to apply different groupings of the current image data, the previous image data, and the subsequent image data to the neural network to generate a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle, and feed the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle back to the memory for use in generating a base call for the current sequencing cycle based on the first base call prediction, the second base call prediction, and the third base call prediction.

[0182] In one embodiment, the different groups include a first group including current image data and previous image data; a second group including current image data, previous image data and subsequent image data; and a third group including current image data and subsequent image data.

[0183] In one implementation, the execution cluster applies the first grouping to the neural network to generate a first base call prediction, applies the second grouping to the neural network to generate a second base call prediction, and applies the third grouping to the neural network to generate a third base call prediction.

[0184] In one implementation, the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle identify the likelihood of the bases in one or more analytes spiked into the analytes being A, C, T, and G at the current sequencing cycle.

[0185] In one implementation, the data flow logic is further configured to generate a base call for the current sequencing cycle by performing a base-by-base summing of the likelihoods across a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle, determining a base-by-base average based on the base-by-base summing, and generating the base call for the current sequencing cycle based on a highest one of the base-by-base averages.

[0186] In one implementation, the data flow logic is further configured to generate base calls for the current sequencing cycle by determining preliminary base calls for each of the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle based on a highest likelihood among the likelihoods, thereby producing a preliminary base call sequence; and generating the base calls for the current sequencing cycle based on the most common base call in the preliminary base call sequence.

[0187] In one implementation, the data flow logic is further configured to generate base calls for the current sequencing cycle by determining a preliminary base call for each of the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle based on a highest one of the likelihoods, thereby producing a sequence of preliminary base calls; applying a base-by-base weight to a corresponding one of the preliminary base calls in the sequence of preliminary base calls and producing a sequence of weighted preliminary base calls; and generating the base calls for the current sequencing cycle based on a most weighted base call in the sequence of weighted preliminary base calls.

[0188] Figure 10 is a simplified diagram illustrating aspects of a base calling operation including functionality of a runtime program executed by a host processor. In the diagram, the output of the image sensor from the flow cell is provided on line 1000 to an image processing thread 1001 which may perform processing on the image, such as alignment and arrangement in the sensor data array of the individual tiles and resampling of the image, and may be used by a process of computing a tile cluster mask for each tile in the flow cell which identifies pixels in the sensor data array corresponding to clusters of genetic material on the corresponding tile of the flow cell. Depending on the state of the base calling operation, the output of the image processing thread 1001 is provided on line 1002 to scheduling logic 1010 in the CPU which routes the tile data array on a high speed bus 1003 to a data cache 1004 (e.g., an SSD storage device) or on a high speed bus 1005 to neural network processor hardware 1020, such as Figure 9Configurable processor 846. The processed and transformed image can be stored in data cache 1004 for use in the previously used sensing cycle. Hardware 1020 returns the classified data output by the neural network to scheduling logic 1010, which passes the information to data cache 1004 or to thread 1002 on thread 1011, which performs base calling and quality score calculations using the classified data and arranges the data for base called reads in a standard format. The output of thread 1002 performing base calling and quality score calculations is provided on line 1012 to thread 1003, which aggregates the base called reads, performs other operations such as data compression, and writes the resulting base call output to a designated destination for client utilization.

[0189] In some implementations, the host may include a thread (not shown) that performs final processing of the output of hardware 1020 to support the neural network. For example, hardware 1020 may provide the output of classification data from the final layer of a multi-cluster neural network. The host processor may perform an output activation function, such as a softmax function, on the classification data to configure the data for use by base calling and quality scoring thread 1002. Additionally, the host processor may perform input operations (not shown), such as batch normalization of the block data before input to hardware 1020.

[0190] Figure 11 is a configurable processor 846 (such as Figure 9 A simplified diagram of the configuration of a configurable processor). Figure 11 In the embodiment, the configurable processor 846 includes an FPGA with multiple high-speed PCIe interfaces. The FPGA is configured with a wrapper 1100, which includes a reference Figure 9The data flow logic 908 described above is shown in Figure 1. The wrapper 1100 manages the interface and coordination with the runtime program in the CPU via CPU communication link 1109 and manages communication with onboard DRAM 1102 (e.g., memory 848A) via DRAM communication link 1110. The data flow logic 908 in the wrapper 1100 provides patch data retrieved by traversing the number N loops of tile data arrays on onboard DRAM 1102 to cluster 1101 and retrieves process data 1115 from cluster 1101 for delivery back to onboard DRAM 1102. The wrapper 1100 also manages data transfers between onboard DRAM 1102 and host memory for both input arrays of tile data and output patches of classification data. The wrapper transmits the patch data on line 1113 to the assigned cluster 1101. The wrapper provides trained parameters, such as weights and biases, to cluster 1101 retrieved from onboard DRAM 1102 on line 1112. The wrapper provides configuration and control data to cluster 1101 on line 1111, which the cluster provides from or generates in response to a runtime program on the host via CPU communication link 1109. The cluster may also provide status signals to wrapper 1100 on line 1116, which are used in conjunction with control signals from the host to manage the traversal of the tile data array to provide spatially aligned patch data and to execute multi-recurrent neural networks on the patch data using the resources of cluster 1101.

[0191] As described above, multiple clusters can exist on a single configurable processor managed by the wrapper 1100, the multiple clusters configured to execute on corresponding patches in the multiple patches of block data. Each cluster can be configured to use the block data for the multiple sensing cycles described herein to provide classification data for base calls in the sensing cycles of the subject.

[0192] In an example of a system, model data (including kernel data, such as filter weights and biases) can be sent from the host CPU to a configurable processor so that the model can be updated according to the number of cycles. As a representative example, a base calling operation may include about hundreds of sensing cycles. In some embodiments, the base calling operation may include double-end reads. For example, the model training parameters can be updated every 20 cycles (or other number of cycles), or according to an update mode implemented for a specific system and neural network model. In some embodiments including double-end reads, where the sequence of a given string in a genetic cluster on a block includes a first portion extending downward (or upward) along the string from a first end and a second portion extending upward (or downward) along the string from a second end, the trained parameters can be updated in the transition from the first portion to the second portion.

[0193] In some examples, image data for multiple cycles in the sensory data for a tile can be sent from the CPU to the wrapper 1100. The wrapper 1100 can optionally perform some preprocessing and conversion on the sensory data and write the information to the onboard DRAM 1102. The input tile data for each sensing cycle can include an array of sensor data, including approximately 4000×3000 pixels or more per tile per sensing cycle, with two features representing the colors of the two images of the tile, and each feature having one or two bytes per pixel. For embodiments where the number N is three sensing cycles to be used in each run of the multi-cycle neural network, the tile data array for each run of the multi-cycle neural network can consume on the order of hundreds of megabytes per tile. In some embodiments of the system, the tile data also includes an array of DFC data stored once per tile, or other types of metadata about the sensor data and the tile.

[0194] In operation, when a multi-cycle cluster is available, the encapsulator assigns patches to the cluster. The encapsulator retrieves the next patch of data for the chunk during a traversal of the chunk and sends it to the assigned cluster along with appropriate control and configuration information. The cluster can be configured with sufficient memory on the configurable processor to hold data patches being processed in-place, including patches from multiple cycles in some systems, as well as data patches to be processed when the current patch is finished processing using ping-pong buffering or raster scanning techniques in various embodiments.

[0195] When the assigned cluster completes its execution of the neural network for the current patch and produces an output patch, it signals the wrapper. The wrapper reads the output patch from the assigned cluster, or alternatively, the assigned cluster pushes the data to the wrapper. The wrapper then assembles the output patch for the processed tile in DRAM 1102. When processing of the entire tile is complete and the output patch of data has been transferred to DRAM, the wrapper sends the tile's processed output array back to the host / CPU in the specified format. In some implementations, the onboard DRAM 1102 is managed by memory management logic in the wrapper 1100. The runtime program can control the sequencing operation to complete the analysis of all arrays of tile data for all cycles in the run in a continuous stream, thereby providing real-time analysis.

[0196] Computer system

[0197] Figure 121 is a computer system 1200 that can be used by sequencing system 800A to implement the base calling techniques disclosed herein. Computer system 1200 includes at least one central processing unit (CPU) 1272 that communicates with a number of peripheral devices via a bus subsystem 1255. These peripheral devices may include a storage subsystem 1210, including, for example, memory devices and a file storage subsystem 1236, a user interface input device 1238, a user interface output device 1276, and a network interface subsystem 1274. The input and output devices allow a user to interact with computer system 1200. Network interface subsystem 1274 provides an interface to an external network, including interfaces to corresponding interface devices in other computer systems.

[0198] In one implementation, the system controller 806 may be communicatively linked to the storage subsystem 1210 and the user interface input device 1238 .

[0199] User interface input devices 1238 may include: a keyboard; a pointing device such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touch screen incorporated into a display; audio input devices such as a voice recognition system and a microphone; and other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 1200.

[0200] The user interface output devices 1276 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include an LED display, a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as an audio output device. Generally speaking, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computer system 1200 to a user or to another machine or computer system.

[0201] The storage subsystem 1210 stores programming and data structures that provide some or all of the functionality and methods of the modules described herein. These software modules are typically executed by the deep learning processor 1278.

[0202] The deep learning processor 1278 may be a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The deep learning processor 1278 may be provided by a deep learning cloud platform such as Google Cloud Platform. TM 、Xilinx TM and CirrascaleTM Hosting. Examples of deep learning processors 1278 include Google's Tensor Processing Unit (TPU) TM , rack solutions (such as GX4RackmountSeries TM 、GX12 Rackmount Series TM ), NVIDIA DGX-1 TM , Microsoft's Stratix V FPGA TM , Graphcore’s Intelligent Processor Unit (IPU) TM 、Qualcomm's Snapdragon processors TM Zeroth Platform TM 、NVIDIA's Volta TM 、NVIDIA's DRIVE PX TM , NVIDIA's JETSON TX1 / TX2 MODULE TM 、Intel's Nirvana TM 、Movidius VPU TM ,Fujitsu DPI TM ARM's DynamicIQ TM 、IBM TrueNorth TM , with Testa V100s TM Lambda GPU servers, etc.

[0203] The memory subsystem 1222 used in the storage subsystem 1210 may include multiple memories, including a main random access memory (RAM) 1232 for storing instructions and data during program execution and a read-only memory (ROM) 1234 in which fixed instructions are stored. The file storage subsystem 1236 may provide persistent storage for program files and data files and may include a hard drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media tape. Modules that implement the functionality of certain specific implementations may be stored by the file storage subsystem 1236 in the storage subsystem 1210 or in other machines accessible to the processor.

[0204] The bus subsystem 1255 provides a mechanism for the various components and subsystems of the computer system 1200 to communicate with each other as intended. Although the bus subsystem 1255 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.

[0205] The computer system 1200 itself can be of different types, including a personal computer, a laptop computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a widely distributed group of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the Figure 12 The description of the computer system 1200 depicted in FIG is intended only as a specific example for illustrating a preferred embodiment of the present invention. Many other configurations of the computer system 1200 are possible, with more Figure 12 The computer system depicted in FIG may have more or fewer components.

[0206] Terms

[0207] The present invention discloses the following clauses:

[0208] 1. An artificial intelligence-based system for base calling, comprising:

[0209] a neural network-based base caller that processes at least a right wing input, a center input, and a left wing input and produces at least a right wing output, a center output, and a left wing output;

[0210] wherein the right wing input comprises current image data for a current sequencing cycle of a sequencing run, supplemented with previous image data for one or more previous sequencing cycles preceding the current sequencing cycle, and wherein the right wing output comprises right wing base call predictions for the current sequencing cycle and base call predictions for previous sequencing cycles;

[0211] wherein the central input comprises current image data, supplemented with previous image data and subsequent image data for one or more subsequent sequencing cycles following the current sequencing cycle; and wherein the central output comprises a central base call prediction for the current sequencing cycle and base call predictions for previous and subsequent sequencing cycles;

[0212] wherein the left wing input comprises current image data, supplemented with subsequent image data, and wherein the left wing output comprises left wing base call predictions for the current sequencing cycle and base call predictions for subsequent sequencing cycles; and

[0213] A base call generator is coupled to the neural network-based base caller and is configured to generate a base call for the current sequencing cycle based on the right wing base call predictions, the center base call predictions, and the left wing base call predictions for the current sequencing cycle.

[0214] 2. An artificial intelligence-based system according to claim 1, wherein the current image data for the current sequencing cycle depicts the intensity emission of the analyte captured in the current sequencing cycle and its surrounding background.

[0215] 3. An artificial intelligence-based system according to claim 2, wherein the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle identify the likelihood of the bases in one or more analytes incorporated into the analyte being A, C, T, and G at the current sequencing cycle.

[0216] 4. The artificial intelligence-based system of clause 3, wherein the base call generator is further configured to include an averager,

[0217] summing the likelihood across the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle, base by base;

[0218] determining a base-by-base average based on the base-by-base summation; and

[0219] Based on the highest of the base-by-base averages, a base call is generated for the current sequencing cycle.

[0220] 5. The artificial intelligence-based system of clause 3, wherein the base call generator is further configured to include a consensus sensor

[0221] determining a preliminary base call for each of the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle based on a highest one of the likelihoods, thereby generating a preliminary base call sequence; and

[0222] Base calls for the current sequencing cycle are generated based on the most common base call in the preliminary base call sequence.

[0223] 6. The artificial intelligence-based system of clause 3, wherein the base call generator is further configured to include a weighted consensus sensor, the weighted consensus sensor

[0224] determining a preliminary base call for each of the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle based on a highest one of the likelihoods, thereby generating a preliminary base call sequence;

[0225] applying the base-by-base weight to a corresponding one of the preliminary base calls in the sequence of preliminary base calls and producing a weighted preliminary base call sequence; and generating a base call for the current sequencing cycle based on the most weighted base call in the sequence of weighted preliminary base calls.

[0226] 7. An artificial intelligence-based system according to claim 3, wherein the likelihood is an exponentially normalized score produced by a softmax layer.

[0227] 8. The artificial intelligence-based system according to clause 1, further configured to include a trainer, which during training,

[0228] calculating errors between base calls generated by a base call generator for a current sequencing cycle, a previous sequencing cycle, and a subsequent sequencing cycle based on a right wing output, a center output, and a left wing output of the neural network-based base caller and a reference true base call;

[0229] determining a gradient for a current sequencing cycle, a previous sequencing cycle, and a subsequent sequencing cycle based on the error; and

[0230] The parameters of the neural network-based base caller are updated by back-propagating gradients.

[0231] 9. An artificial intelligence-based system according to claim 1, wherein the right wing base call prediction for the current sequencing cycle takes into account the pre-phasing effect between the current sequencing cycle and the previous sequencing cycle.

[0232] 10. The artificial intelligence-based system of claim 9, wherein the prediction of the central base call for the current sequencing cycle takes into account a prephasing effect between the current sequencing cycle and the previous sequencing cycle, as well as a phasing effect between the current sequencing cycle and the subsequent sequencing cycle.

[0233] 11. An artificial intelligence-based system according to claim 10, wherein the left-wing base call prediction for the current sequencing cycle takes into account the phasing effect between the current sequencing cycle and the subsequent sequencing cycle.

[0234] 12. An artificial intelligence-based system for base calling, the system comprising:

[0235] Host processor;

[0236] a memory accessible to a host processor that stores image data for sequencing cycles of a sequencing run, wherein current image data for a current sequencing cycle of the sequencing run depicts intensity emissions of an analyte captured at the current sequencing cycle and its surrounding background; and

[0237] A configurable processor, the configurable processor being capable of accessing a memory, the configurable processor comprising:

[0238] a plurality of execution clusters, an execution cluster in the plurality of execution clusters configured to execute a neural network; and

[0239] Data flow logic, the data flow logic having access to a memory and an execution cluster from among the plurality of execution clusters, is configured to provide current image data, previous image data for one or more previous sequencing cycles prior to the current sequencing cycle, and subsequent image data for one or more subsequent sequencing cycles subsequent to the current sequencing cycle to an available execution cluster from among the plurality of execution clusters, causing the execution cluster to apply different groupings of the current image data, the previous image data, and the subsequent image data to a neural network to generate a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle, and to feed the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle back to the memory for use in generating a base call for the current sequencing cycle based on the first base call prediction, the second base call prediction, and the third base call prediction.

[0240] 13. An artificial intelligence-based system according to claim 12, wherein the different groups include a first group, a second group, and a third group, the first group including current image data and previous image data; the second group including current image data, previous image data, and subsequent image data; and the third group including current image data and subsequent image data.

[0241] 14. The artificial intelligence-based system of clause 13, wherein the execution cluster applies the first grouping to the neural network to generate a first base call prediction, applies the second grouping to the neural network to generate a second base call prediction, and applies the third grouping to the neural network to generate a third base call prediction.

[0242] 15. An artificial intelligence-based system according to claim 12, wherein the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle identify the likelihood of the base in one or more analytes incorporated into the analyte being A, C, T, and G at the current sequencing cycle.

[0243] 16. The artificial intelligence-based system of clause 15, wherein the data flow logic is further configured to generate base calls for the current sequencing cycle,

[0244] by summing the likelihood across the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle, base by base;

[0245] determining a base-by-base average based on the base-by-base summation; and

[0246] Based on the highest of the base-by-base averages, a base call is generated for the current sequencing cycle.

[0247] 17. The artificial intelligence-based system of clause 15, wherein the data flow logic is further configured to generate base calls for the current sequencing cycle,

[0248] generating a preliminary sequence of base calls by determining a preliminary base call for each of the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle based on a highest one of the likelihoods; and

[0249] Base calls for the current sequencing cycle are generated based on the most common base call in the preliminary base call sequence.

[0250] 18. The artificial intelligence-based system of clause 15, wherein the data flow logic is further configured to generate base calls for the current sequencing cycle,

[0251] generating a preliminary base call sequence by determining a preliminary base call for each of the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle based on a highest one of the likelihoods;

[0252] applying the base-by-base weight to a corresponding one of the preliminary base calls in the sequence of preliminary base calls and producing a weighted preliminary base call sequence; and generating a base call for the current sequencing cycle based on the most weighted base call in the sequence of weighted preliminary base calls.

[0253] 19. An artificial intelligence-based method for base calling, comprising:

[0254] accessing current image data for a current sequencing cycle of a sequencing run, previous image data for one or more previous sequencing cycles before the current sequencing cycle, and subsequent image data for one or more subsequent sequencing cycles after the current sequencing cycle;

[0255] processing different groups of current image data, previous image data, and subsequent image data through a neural network-based base caller and generating a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle; and

[0256] A base call for the current sequencing cycle is generated based on the first base call prediction, the second base call prediction, and the third base call prediction.

[0257] 20. The artificial intelligence-based method of clause 19, wherein the different groupings include:

[0258] a first grouping, the first grouping including current image data and previous image data;

[0259] a second grouping including current image data, previous image data, and subsequent image data; and

[0260] The third group includes current image data and subsequent image data.

[0261] 21. The artificial intelligence-based method according to clause 20, further comprising:

[0262] processing the first grouping through a neural network-based base caller to generate a first base call prediction;

[0263] processing the second grouping through a neural network-based base caller to generate second base call predictions; and

[0264] The third grouping is processed by the neural network-based base caller to produce a third base call prediction.

[0265] 22. An artificial intelligence-based method according to claim 19, wherein the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle identify the likelihood of the bases in one or more analytes incorporated into the analyte being A, C, T, and G at the current sequencing cycle.

[0266] 23. The artificial intelligence-based method of clause 22, further comprising generating a base call for the current sequencing cycle,

[0267] by summing the likelihood across the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle, base by base;

[0268] determining a base-by-base average based on the base-by-base summation; and

[0269] Based on the highest of the base-by-base averages, a base call is generated for the current sequencing cycle.

[0270] 24. The artificial intelligence-based method of clause 22, further comprising generating a base call for the current sequencing cycle,

[0271] generating a preliminary sequence of base calls by determining a preliminary base call for each of the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle based on a highest one of the likelihoods; and

[0272] Base calls for the current sequencing cycle are generated based on the most common base call in the preliminary base call sequence.

[0273] 25. The artificial intelligence-based method of clause 22, further comprising generating a base call for the current sequencing cycle,

[0274] generating a preliminary base call sequence by determining a preliminary base call for each of the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle based on a highest one of the likelihoods;

[0275] applying the base-by-base weight to a corresponding one of the preliminary base calls in the sequence of preliminary base calls and producing a weighted preliminary base call sequence; and generating a base call for the current sequencing cycle based on the most weighted base call in the sequence of weighted preliminary base calls.

[0276] 26. An artificial intelligence-based method for base calling, comprising:

[0277] processing at least a right wing input, a center input, and a left wing input by a neural network-based base caller and generating at least a right wing output, a center output, and a left wing output;

[0278] wherein the right wing input comprises current image data for a current sequencing cycle of a sequencing run, supplemented with previous image data for one or more previous sequencing cycles preceding the current sequencing cycle, and wherein the right wing output comprises right wing base call predictions for the current sequencing cycle and base call predictions for previous sequencing cycles;

[0279] wherein the central input comprises current image data, supplemented with previous image data and subsequent image data for one or more subsequent sequencing cycles following the current sequencing cycle; and wherein the central output comprises a central base call prediction for the current sequencing cycle and base call predictions for previous and subsequent sequencing cycles;

[0280] wherein the left wing input comprises current image data, supplemented with subsequent image data, and wherein the left wing output comprises left wing base call predictions for a current sequencing cycle and base call predictions for subsequent sequencing cycles; and

[0281] A base call for the current sequencing cycle is generated based on the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle.

[0282] 27. The artificial intelligence-based method of clause 26, wherein the current image data for the current sequencing cycle depicts the intensity emission of the analyte captured in the current sequencing cycle and its surrounding background.

[0283] 28. An artificial intelligence-based method according to claim 26, wherein the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle identify the likelihood of the bases in one or more analytes incorporated into the analyte being A, C, T, and G at the current sequencing cycle.

[0284] 29. The artificial intelligence-based method of clause 28, further comprising generating a base call for the current sequencing cycle,

[0285] by summing the likelihood across the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle, base by base;

[0286] determining a base-by-base average based on the base-by-base summation; and

[0287] Based on the highest of the base-by-base averages, a base call is generated for the current sequencing cycle.

[0288] 30. The artificial intelligence-based method of clause 28, further comprising generating a base call for the current sequencing cycle,

[0289] generating a preliminary base call sequence by determining a preliminary base call for each of the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle based on a highest one of the likelihoods; and

[0290] Base calls for the current sequencing cycle are generated based on the most common base call in the preliminary base call sequence.

[0291] 31. The artificial intelligence-based method of clause 28, further comprising generating a base call for the current sequencing cycle,

[0292] generating a preliminary base call sequence by determining a preliminary base call for each of the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle based on a highest one of the likelihoods;

[0293] applying the base-by-base weight to a corresponding one of the preliminary base calls in the sequence of preliminary base calls and producing a weighted preliminary base call sequence; and generating a base call for the current sequencing cycle based on the most weighted base call in the sequence of weighted preliminary base calls.

[0294] 32. An artificial intelligence-based method according to clause 28, wherein the likelihood is an exponentially normalized score produced by a softmax layer.

[0295] 33. The artificial intelligence-based method according to clause 26, further comprising: during training,

[0296] calculating errors between base calls generated by a base call generator for a current sequencing cycle, a previous sequencing cycle, and a subsequent sequencing cycle based on a right wing output, a center output, and a left wing output of the neural network-based base caller and a reference true base call;

[0297] determining a gradient for a current sequencing cycle, a previous sequencing cycle, and a subsequent sequencing cycle based on the error; and

[0298] The parameters of the neural network-based base caller are updated by back-propagating gradients.

[0299] 34. An artificial intelligence-based method according to clause 26, wherein the right wing base call prediction for the current sequencing cycle takes into account a prephasing effect between the current sequencing cycle and the previous sequencing cycle.

[0300] 35. An artificial intelligence-based method according to claim 34, wherein the prediction of the central base call for the current sequencing cycle takes into account the pre-phasing effect between the current sequencing cycle and the previous sequencing cycle, and the phasing effect between the current sequencing cycle and the subsequent sequencing cycle.

[0301] 36. An artificial intelligence-based method according to claim 35, wherein the left-wing base call prediction for the current sequencing cycle takes into account the phasing effect between the current sequencing cycle and the subsequent sequencing cycle.

[0302] 37. An artificial intelligence-based method for base calling, comprising:

[0303] processing at least a first input, a second input, and a third input by a neural network-based base caller and generating at least a first output, a second output, and a third output;

[0304] wherein the first input comprises particular image data for a particular sequencing cycle of a sequencing run, supplemented with previous image data for one or more previous sequencing cycles preceding the particular sequencing cycle, and wherein the first output comprises a first base call prediction for the particular sequencing cycle and base call predictions for previous sequencing cycles;

[0305] wherein the second input comprises particular image data supplemented with previous image data and subsequent image data for one or more subsequent sequencing cycles following the particular sequencing cycle; and wherein the second output comprises a second base call prediction for the particular sequencing cycle and base call predictions for the previous sequencing cycle and the subsequent sequencing cycle;

[0306] wherein the third input comprises particular image data, supplemented with subsequent image data, and wherein the third output comprises a third base call prediction for a particular sequencing cycle and base call predictions for subsequent sequencing cycles; and

[0307] A base call for the particular sequencing cycle is generated based on the first base call prediction, the second base call prediction, and the third base call prediction for the particular sequencing cycle.

[0308] 38. An artificial intelligence-based method according to clause 37, which implements each of the clauses ultimately dependent on clause 1.

[0309] 39. A non-transitory computer-readable storage medium having embodied thereon computer program instructions for performing artificial intelligence-based base calling, the instructions when executed on a processor implementing a method comprising:

[0310] accessing current image data for a current sequencing cycle of a sequencing run, previous image data for one or more previous sequencing cycles before the current sequencing cycle, and subsequent image data for one or more subsequent sequencing cycles after the current sequencing cycle;

[0311] processing different groups of current image data, previous image data, and subsequent image data through a neural network-based base caller and generating a first base call prediction, a second base call prediction, and a third base call prediction for the current sequencing cycle; and

[0312] A base call for the current sequencing cycle is generated based on the first base call prediction, the second base call prediction, and the third base call prediction.

[0313] 40. The non-transitory computer-readable storage medium of clause 39, which implements each of the clauses ultimately dependent upon clause 1.

[0314] 41. A non-transitory computer-readable storage medium having embodied thereon computer program instructions for performing artificial intelligence-based base calling, the instructions when executed on a processor implementing a method comprising:

[0315] processing at least a first input, a second input, and a left input by a neural network-based base caller and generating at least a first output, a second output, and a left output;

[0316] wherein the first input comprises particular image data for a particular sequencing cycle of a sequencing run, supplemented with previous image data for one or more previous sequencing cycles preceding the particular sequencing cycle, and wherein the first output comprises a first base call prediction for the particular sequencing cycle and base call predictions for previous sequencing cycles;

[0317] wherein the second input comprises particular image data supplemented with previous image data and subsequent image data for one or more subsequent sequencing cycles following the particular sequencing cycle; and wherein the second output comprises a second base call prediction for the particular sequencing cycle and base call predictions for the previous sequencing cycle and the subsequent sequencing cycle;

[0318] wherein the left input comprises particular image data, supplemented with subsequent image data, and wherein the left output comprises a left base call prediction for a particular sequencing cycle and base call predictions for subsequent sequencing cycles; and

[0319] Base calls for the particular sequencing cycle are generated based on the first base call prediction, the second base call prediction, and the left base call prediction for the particular sequencing cycle.

[0320] 44. The non-transitory computer-readable storage medium of clause 43, which implements each of the clauses ultimately dependent upon clause 1.

[0321] 45. An artificial intelligence-based method for base calling, the method comprising:

[0322] Access the progress of the per-cycle analyte channel sets generated for the sequencing cycles of a sequencing run;

[0323] The neural network-based base caller processes the windows of the sequencing cycles of the sequencing run and the windows of the analyte channels per cycle in progress, so that

[0324] Neural network-based base caller

[0325] an object window for sequencing cycles of a sequencing run, an object window for the analyte channel set per cycle in progress; and

[0326] generating provisional base call predictions for three or more sequencing cycles in an object window of sequencing cycles;

[0327] using a neural network-based base caller to generate a tentative base call prediction for a particular sequencing cycle from multiple windows where the base call occurs at different positions in the particular sequencing cycle; and

[0328] The base call for a particular sequencing cycle is determined based on multiple base call predictions.

[0329] 46. ​​An artificial intelligence-based method according to clause 45, which implements each of the clauses ultimately dependent on clause 1.

[0330] 47. A system comprising one or more processors coupled to a memory loaded with computer instructions for performing artificial intelligence-based base calling, the instructions, when executed on the processors, performing the following actions, comprising:

[0331] Access the progress of the per-cycle analyte channel sets generated for the sequencing cycles of a sequencing run;

[0332] The neural network-based base caller processes the windows of the sequencing cycles of the sequencing run and the windows of the analyte channels per cycle in progress, so that

[0333] Neural network-based base caller

[0334] an object window for sequencing cycles of a sequencing run, an object window for the analyte channel set per cycle in progress; and

[0335] generating provisional base call predictions for three or more sequencing cycles in an object window of sequencing cycles;

[0336] using a neural network-based base caller to generate a tentative base call prediction for a particular sequencing cycle from multiple windows where the base call occurs at different positions in the particular sequencing cycle; and

[0337] The base call for a particular sequencing cycle is determined based on multiple base call predictions.

[0338] 48. A system according to clause 47, which implements each of the clauses ultimately dependent on clause 1.

[0339] 49. A non-transitory computer-readable storage medium having embodied thereon computer program instructions for performing artificial intelligence-based base calling, the instructions when executed on a processor implementing a method comprising:

[0340] Access the progress of the per-cycle analyte channel sets generated for the sequencing cycles of a sequencing run;

[0341] The neural network-based base caller processes the windows of the sequencing cycles of the sequencing run and the windows of the analyte channels per cycle in progress, so that

[0342] Neural network-based base caller

[0343] an object window for sequencing cycles of a sequencing run, an object window for the analyte channel set per cycle in progress; and

[0344] generating provisional base call predictions for three or more sequencing cycles in an object window of sequencing cycles;

[0345] using a neural network-based base caller to generate a tentative base call prediction for a particular sequencing cycle from multiple windows where the base call occurs at different positions in the particular sequencing cycle; and

[0346] The base call for a particular sequencing cycle is determined based on multiple base call predictions.

[0347] 50. The non-transitory computer-readable storage medium of clause 49, which implements each of the clauses ultimately dependent upon clause 1.

[0348] 51. An artificial intelligence-based method for base calling, the method comprising:

[0349] Access a series of per-cycle analyte channel sets generated for sequencing cycles of a sequencing run;

[0350] The windows of analyte channels per cycle in the sequence are processed by a neural network-based base caller for a window of sequencing cycles of the sequencing run, such that

[0351] Neural network-based base caller

[0352] Object windows for sequencing cycles for a sequencing run, object windows for analyte channel sets per cycle in a processing series,

[0353] and generating base call predictions for two or more sequencing cycles in an object window of sequencing cycles;

[0354] Through a neural network-based base caller,

[0355] Multiple windows for sequencing cycles of a sequencing run, multiple windows for analyte channel sets per cycle in a processing series,

[0356] and generating an output for each window in the plurality of apertures,

[0357] wherein each window in the plurality of windows comprises a particular set of analyte per cycle channels for a particular sequencing cycle of the sequencing run, and

[0358] The output for each of the multiple windows includes:

[0359] (i) Base call prediction and

[0360] (ii) one or more additional base call predictions for one or more additional sequencing cycles of the sequencing run, thereby generating a plurality of base call predictions for a particular sequencing cycle across the plurality of windows; and

[0361] The base call for a particular sequencing cycle is determined based on multiple base call predictions.

[0362] 52. A system comprising one or more processors coupled to a memory loaded with computer instructions for performing artificial intelligence-based base calling, the instructions, when executed on the processors, performing the following actions, comprising:

[0363] Access a series of per-cycle analyte channel sets generated for sequencing cycles of a sequencing run;

[0364] The windows of analyte channels per cycle in the sequence are processed by a neural network-based base caller for a window of sequencing cycles of the sequencing run, such that

[0365] Neural network-based base caller

[0366] Object windows for sequencing cycles for a sequencing run, object windows for analyte channel sets per cycle in a processing series,

[0367] and generating base call predictions for two or more sequencing cycles in an object window of sequencing cycles;

[0368] Through a neural network-based base caller,

[0369] Multiple windows for sequencing cycles of a sequencing run, multiple windows for analyte channel sets per cycle in a processing series,

[0370] and generating an output for each window in the plurality of apertures,

[0371] wherein each window in the plurality of windows comprises a particular set of analyte per cycle channels for a particular sequencing cycle of the sequencing run, and

[0372] The output for each of the multiple windows includes:

[0373] (i) Base call prediction and

[0374] (ii) one or more additional base call predictions for one or more additional sequencing cycles of the sequencing run, thereby generating a plurality of base call predictions for a particular sequencing cycle across the plurality of windows; and

[0375] The base call for a particular sequencing cycle is determined based on multiple base call predictions.

[0376] 53. A system according to clause 52, which implements each of the clauses ultimately dependent on clause 1.

[0377] 54. A non-transitory computer-readable storage medium having embodied thereon computer program instructions for performing artificial intelligence-based base calling, the instructions when executed on a processor implementing a method comprising:

[0378] Access a series of per-cycle analyte channel sets generated for sequencing cycles of a sequencing run;

[0379] The windows of analyte channels per cycle in the sequence are processed by a neural network-based base caller for a window of sequencing cycles of the sequencing run, such that

[0380] Neural network-based base caller

[0381] Object windows for sequencing cycles for a sequencing run, object windows for analyte channel sets per cycle in a processing series,

[0382] and generating base call predictions for two or more sequencing cycles in an object window of sequencing cycles;

[0383] Through a neural network-based base caller,

[0384] Multiple windows for sequencing cycles of a sequencing run, multiple windows for analyte channel sets per cycle in a processing series,

[0385] and generating an output for each window in the plurality of apertures,

[0386] wherein each window in the plurality of windows comprises a particular set of analyte per cycle channels for a particular sequencing cycle of the sequencing run, and

[0387] The output for each of the multiple windows includes:

[0388] (i) Base call prediction and

[0389] (ii) one or more additional base call predictions for one or more additional sequencing cycles of the sequencing run, thereby generating a plurality of base call predictions for a particular sequencing cycle across the plurality of windows; and

[0390] The base call for a particular sequencing cycle is determined based on multiple base call predictions.

[0391] 55. The non-transitory computer-readable storage medium of clause 54, which implements each of the clauses ultimately dependent upon clause 1.

[0392] Other implementations of the above methods may include a non-transitory computer-readable storage medium storing instructions, the instructions being executable by a processor to perform any of the above methods. Yet another implementation of the methods described in this section may include a system comprising a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above methods.

Claims

1. An artificial intelligence-based system for base calling, comprising: a neural network-based base caller comprising a plurality of parallel convolutional neural network pipelines, each trained to process at least a right wing input, a center input, and a left wing input of an image depicting a cluster and its surrounding background, and to produce at least a right wing output, a center output, and a left wing output for the cluster in a current sequencing cycle; wherein the right wing input comprises current image data for the current sequencing cycle of a sequencing run, supplemented with previous image data for one or more previous sequencing cycles preceding the current sequencing cycle, and wherein the right wing output comprises right wing base call predictions for the current sequencing cycle and base call predictions for the previous sequencing cycles; wherein the central input comprises the current image data, supplemented with the previous image data and subsequent image data for one or more subsequent sequencing cycles following the current sequencing cycle; and wherein the central output comprises a central base call prediction for the current sequencing cycle and base call predictions for the previous sequencing cycle and the subsequent sequencing cycle; wherein the left wing input comprises the current image data, supplemented with the subsequent image data, and wherein the left wing output comprises a left wing base call prediction for the current sequencing cycle and a base call prediction for the subsequent sequencing cycle; wherein training of the trained parallel convolutional neural network pipeline comprises providing a training dataset of right wing inputs, center inputs, and left wing inputs of an image depicting a cluster and its surrounding background, and ground truth base calls for the cluster in the current sequencing cycle, and applying back-propagation-based gradient updates; and a base call generator coupled to the neural network-based base caller and configured to generate a base call for the cluster in the current sequencing cycle based on the right wing base call predictions, the center base call predictions, and the left wing base call predictions for the current sequencing cycle.

2. The artificial intelligence-based system of claim 1 , wherein the current image data for the current sequencing cycle depicts the intensity emission of the analyte captured in the current sequencing cycle and its surrounding background.

3. The artificial intelligence-based system of claim 2, wherein the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle identify the likelihood of bases incorporated into one or more of the analytes being A, C, T, and G at the current sequencing cycle.

4. The artificial intelligence-based system of claim 3, wherein the base call generator is further configured to include an averager, performing a base-by-base summation of the likelihoods across the right wing base call predictions, the center base call predictions, and the left wing base call predictions for the cluster applied to the current sequencing cycle; determining a base-by-base average based on the base-by-base summation; and The base call for the current sequencing cycle is generated based on a highest one of the base-by-base averages.

5. The artificial intelligence-based system of claim 3, wherein the base call generator is further configured to include a consensus sensor, the consensus sensor determining a preliminary base call for each of the right wing base call predictions, the center base call predictions, and the left wing base call predictions for the cluster applicable to the current sequencing cycle based on a highest one of the likelihoods, thereby generating a preliminary base call sequence; and The base calls for the cluster applicable to the current sequencing cycle are generated based on the most common base calls in the preliminary base call sequence.

6. The artificial intelligence-based system of claim 3, wherein the base call generator is further configured to include a weighted consensus sensor, the weighted consensus sensor determining a preliminary base call for each of the right wing base call prediction, the center base call prediction, and the left wing base call prediction for the current sequencing cycle based on a highest one of the likelihoods, thereby generating a preliminary base call sequence; applying a base-by-base weight to a corresponding one of the preliminary base calls in the preliminary base call sequence and generating a weighted preliminary base call sequence; and The base call for the current sequencing cycle is generated based on the most weighted base call in the sequence of weighted preliminary base calls.

7. The artificial intelligence-based system of any one of claims 3-6, wherein the likelihood is an exponentially normalized score produced by a softmax layer.

8. The artificial intelligence-based system according to any one of claims 1 to 6, further configured to include a trainer, wherein during training, calculating errors between base calls generated by the base call generator for the current sequencing cycle, the previous sequencing cycle, and the subsequent sequencing cycle based on the right wing output, the center output, and the left wing output of the neural network-based base caller and a reference true base call; determining a gradient for the current sequencing cycle, the previous sequencing cycle, and the subsequent sequencing cycle based on the error; and The parameters of the neural network-based base caller are updated by back-propagating the gradients.

9. The artificial intelligence-based system of any one of claims 1-6, wherein the right wing base call prediction for the current sequencing cycle takes into account a pre-phasing effect between the current sequencing cycle and the previous sequencing cycle.

10. The artificial intelligence-based system of claim 9, wherein the central base call prediction for the current sequencing cycle takes into account the prephasing effect between the current sequencing cycle and the previous sequencing cycle, and the phasing effect between the current sequencing cycle and the subsequent sequencing cycle.

11. The artificial intelligence-based system of claim 10, wherein the left-wing base call prediction for the current sequencing cycle takes into account the phasing effect between the current sequencing cycle and the subsequent sequencing cycle.

12. An artificial intelligence-based system for base calling, the system comprising: Host processor; a memory accessible to the host processor and storing image data depicting clusters and their surrounding background for sequencing cycles of a sequencing run, wherein current image data for a current sequencing cycle of the sequencing run depicts intensity emissions of an analyte captured at the current sequencing cycle and its surrounding background; and a configurable processor capable of accessing the memory, the configurable processor comprising: a plurality of execution clusters, the execution cluster of the plurality of execution clusters being configured to execute a neural network; and Data flow logic, the data flow logic being capable of accessing the memory and an execution cluster of the plurality of execution clusters, configured to providing the current image data, previous image data for one or more previous sequencing cycles before the current sequencing cycle, and subsequent image data for one or more subsequent sequencing cycles after the current sequencing cycle to an available execution cluster of the plurality of execution clusters, and Make the execution cluster applying different groupings of the current image data, the previous image data, and the subsequent image data to a parallel convolutional neural network pipeline in the neural network, the groupings comprising recurring windows, wherein the current sequencing cycle occurs at different positions in the windows, the parallel convolutional neural network pipeline being trained using a training dataset of recurring windows and ground truth base calls for clusters in the current sequencing cycle and applying backpropagation-based gradient updates, wherein the current sequencing cycle occurs at different positions in the windows, generating, from the grouping, a first base call prediction, a second base call prediction, and a third base call prediction for the cluster that can be applied to the current sequencing cycle, and Feeding back the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle to the memory for use in generating base calls for the cluster applicable to the current sequencing cycle based on the first base call prediction, the second base call prediction, and the third base call prediction.

13. The artificial intelligence-based system according to claim 12, wherein the different groups include a first group including the current image data and the previous image data; a second group including the current image data, the previous image data, and the subsequent image data; and a third group including the current image data and the subsequent image data.

14. The artificial intelligence-based system of claim 13, wherein the execution cluster applies the first grouping to the neural network to generate the first base call prediction, applies the second grouping to the neural network to generate the second base call prediction, and applies the third grouping to the neural network to generate the third base call prediction.

15. The artificial intelligence-based system of any one of claims 12-14, wherein the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle identify the likelihood of bases incorporated into one or more of the analytes being A, C, T, and G at the current sequencing cycle.

16. The artificial intelligence-based system of claim 15, wherein the data flow logic is further configured to generate the base call for the current sequencing cycle by, by performing a base-by-base summation of the likelihood across the first base call prediction, the second base call prediction, and the third base call prediction for the current sequencing cycle; determining a base-by-base average based on the base-by-base summation; and The base call for the current sequencing cycle is generated based on a highest one of the base-by-base averages.

17. An artificial intelligence-based method for base calling, comprising: accessing current image data depicting the cluster and its surrounding background for a current sequencing cycle of a sequencing run, previous image data for one or more previous sequencing cycles preceding the current sequencing cycle, and subsequent image data for one or more subsequent sequencing cycles following the current sequencing cycle; processing different groups of the current image data, the previous image data, and the subsequent image data with a neural network-based base caller comprising a plurality of respectively trained parallel convolutional neural network pipelines and generating a plurality of call predictions applicable to clusters of the current sequencing cycle, the groups comprising windows of cycles wherein the current sequencing cycle occurs at different positions within the windows; wherein training of the trained parallel convolutional neural network pipeline comprises providing a training dataset of right wing input, center input, and left wing input of an image depicting a cluster and its surrounding background, and ground truth base calls of the cluster in the current sequencing cycle, and applying back-propagation-based gradient updates; as well as A base call for the cluster applicable to the current sequencing cycle is generated based on the plurality of base call predictions for the cluster.

18. An artificial intelligence-based method for base calling, comprising: generating corresponding base call candidates for the image of the cluster and its surrounding background in a particular sequencing cycle in response to executing corresponding iterations of the base caller; wherein the respective iterative processing is for a respective input set of a respective plurality of windows of sequencing cycles, the particular sequencing cycle occurring at a different position in the plurality of windows; and wherein the corresponding window of sequencing cycles has a particular sequencing cycle as at least one overlapping cycle and one or more non-overlapping cycles; and Based on the corresponding base call candidates, a final base call is generated for the cluster in the particular sequencing cycle.