Hardware Execution and Acceleration of an Artificial Intelligence-Based Base Caller
By using deep neural networks and configurable processors in nucleic acid sequencing, the problem of throughput limitation in high-throughput nucleic acid sequencing is solved, and resource-efficient rapid data analysis and real-time result output are achieved.
Patent Information
- Application Number
- CN202180015454.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-15
- Filing Date
- 2021-02-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-02-17
AI Technical Summary
Existing high-throughput nucleic acid sequencing technologies face problems of flux limitation and high resource consumption, especially when dealing with closely close or overlapping nucleic acid clusters, resulting in a trade-off in quantity and quality of nucleic acid sequence information.
Using a deep neural network-based method, configurable or reconfigurable processors such as FPGA and CGRA are used to perform multi-cycle neural networks, and by classifying base detection sensor data, the throughput of nucleic acid sequencing and optimize resource use.
It realizes efficient and rapid increase in the quality and quantity of nucleic acid sequencing data on hardware, supports real-time analysis, and reduces computing time and resource requirements.
Smart Images

Figure CN115136243B_ABST
Abstract
Description
Technical Field
[0001] The disclosed technology relates to artificial intelligence type computers and digital data processing systems, as well as corresponding data processing methods and products for simulating intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems); and includes systems for uncertain inference (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the disclosed technology relates to using deep neural networks such as deep convolutional neural networks for data analysis.
[0002] Priority Applications
[0003] This PCT application claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 979,412, filed on February 20, 2020, entitled "MULTI-CYCLE CLUSTER BASED REALTIME ANALYSIS SYSTEM" (Attorney Docket No. ILLM 1020-1 / IP-1866-PRV) and U.S. Patent Application No. 17 / 176,147, filed on February 15, 2021, entitled "MULTI-CYCLE CLUSTER BASED REAL TIME ANALYSIS SYSTEM" (Attorney Docket No. ILLM 1020-2 / IP-1866-US). These priority applications are hereby incorporated by reference in their entirety, as if fully set forth herein, for all purposes.
[0004] This PCT application claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 979,385, filed on February 20, 2020, entitled "KNOWLEDGE DISTILLATION-BASED COMPRESSION OF ARTIFICIAL INTELLIGENCE-BASED BASE CALLER" (Attorney Docket No. ILLM 1017-1 / IP-1859-PRV) and U.S. Patent Application No. 17 / 176,151, filed on February 15, 2021, entitled "KNOWLEDGE DISTILLATION-BASED COMPRESSION OF ARTIFICIAL INTELLIGENCE-BASED BASE CALLER" (Attorney Docket No. ILLM 1017-2 / IP-1859-US). These priority applications are hereby incorporated by reference in their entirety, as if fully set forth herein, for all purposes.
[0005] This PCT application claims the priority and benefit of U.S. Provisional Patent Application No. 63 / 072,032, filed on August 28, 2020, titled "DETECTING AND FILTERING CLUSTERS BASED ON ARTIFICIAL INTELLIGENCE-PREDICTED BASE CALLS" (Attorney Docket No. ILLM 1018-1 / IP-1860-PRV). These priority applications are hereby incorporated by reference in their entirety, as if fully set forth herein, for all purposes.
[0006] This PCT application claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 979,411, filed on February 20, 2020, titled "DATA COMPRESSION FOR ARTIFICIAL INTELLIGENCE-BASED BASE CALLING" (Attorney Docket No. ILLM 1029-1 / IP-1964-PRV). These priority applications are hereby incorporated by reference in their entirety, as if fully set forth herein, for all purposes.
[0007] This PCT application claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 979,399, filed on February 20, 2020, titled "SQUEEZING LAYER FOR ARTIFICIAL INTELLIGENCE-BASED BASE CALLING" (Attorney Docket No. ILLM 1030-1 / IP-1982-PRV). These priority applications are hereby incorporated by reference in their entirety, as if fully set forth herein, for all purposes.
[0008] Incorporation of Literature
[0009] The following documents are hereby incorporated by reference in their entirety, as if fully set forth herein:
[0010] U.S. Provisional Patent Application No. 62 / 979,384, filed on February 20, 2020, titled "ARTIFICIAL INTELLIGENCE-BASED BASE CALLING OF INDEX SEQUENCES" (Attorney Docket No. ILLM 1015-1 / IP-1857-PRV);
[0011] U.S. Provisional Patent Application No. 62 / 979,414, filed on February 20, 2020, with the title "ARTIFICIAL INTELLIGENCE-BASED MANY-TO-MANYBASE CALLING" (Attorney Docket No. ILLM 1016-1 / IP-1858-PRV);
[0012] U.S. Non-Provisional Patent Application No. 16 / 825,987, filed on March 20, 2020, with the title "TRAINING DATA GENERATION FOR ARTIFICIALINTELLIGENCE-BASED SEQUENCING" (Attorney Docket No. ILLM1008-16 / IP-1693-US);
[0013] U.S. Non-Provisional Patent Application No. 16 / 825,991, filed on March 20, 2020, with the title "ARTIFICIAL INTELLIGENCE-BASED GENERATION OFSEQUENCING METADATA" (Attorney Docket No. ILLM 1008-17 / IP-1741-US);
[0014] U.S. Non-Provisional Patent Application No. 16 / 826,126, filed on March 20, 2020, with the title "ARTIFICIAL INTELLIGENCE-BASED BASE CALLING" (Attorney Docket No. ILLM 1008-18 / IP-1744-US);
[0015] U.S. Non-Provisional Patent Application No. 16 / 826,134, filed on March 20, 2020, with the title "ARTIFICIAL INTELLIGENCE-BASED QUALITYSCORING" (Attorney Docket No. ILLM 1008-19 / IP-1747-US); and
[0016] U.S. Non-Provisional Patent Application No. 16 / 826,168, filed on March 21, 2020, with the title "ARTIFICIAL INTELLIGENCE-BASED SEQUENCING" (Attorney Docket No. ILLM 1008-20 / IP-1752-PRV-US). BACKGROUND OF THE INVENTION
[0017] Subjects discussed in this section should not be regarded as prior art merely because they are mentioned in this section. Similarly, problems mentioned in this section or associated with the subjects provided as background art should not be regarded as having been previously recognized in the prior art. The subjects in this section merely represent different methods, which themselves may also correspond to specific implementations of the technologies protected by the claims.
[0018] A deep neural network is an artificial neural network that uses multiple non-linear and complex transformation layers to continuously model high-level features. The deep neural network provides feedback via backpropagation, which carries the difference between the observed output and the predicted output to adjust the parameters. The deep neural network has evolved with the availability of large training datasets, the ability of parallel and distributed computing, and complex training algorithms. The deep neural network has promoted significant progress in many fields such as computer vision, speech recognition, and natural language processing.
[0019] Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are components of deep neural networks. Convolutional neural networks are particularly successful in image recognition with architectures including convolutional layers, non-linear layers, and pooling layers. Recurrent neural networks are designed to utilize the sequential information of the input data and have recurrent connections between building blocks such as perceptrons, long short-term memory units, and gated recurrent units. In addition, many other emerging deep neural networks for limited scenarios have been proposed, such as deep spatio-temporal neural networks, multi-dimensional recurrent neural networks, and convolutional autoencoders.
[0020] The goal of training a deep neural network is to optimize the weight parameters in each layer, which gradually combines simpler features into complex features so that the most appropriate hierarchical representation can be learned from the data. A single cycle of the optimization process proceeds as follows. First, given a training dataset, the forward pass sequentially calculates the outputs in each layer and propagates the function signals forward through the network. In the final output layer, the objective loss function measures the error between the inferred output and the given label. To minimize the training error, the backward pass uses the chain rule to backpropagate the error signal and calculate the gradients with respect to all the weights in the entire neural network. Finally, an optimization algorithm is used to update the weight parameters based on stochastic gradient descent. While batch gradient descent performs parameter updates for each complete dataset, stochastic gradient descent provides a stochastic approximation by performing updates for each small set of data examples. Several optimization algorithms are derived from stochastic gradient descent. For example, the Adagrad and Adam training algorithms perform stochastic gradient descent while adaptively modifying the learning rate based on the update frequency and momentum of the gradient of each parameter, respectively.
[0021] Another core element in deep neural network training is regularization, which refers to strategies aimed at avoiding overfitting and thus achieving good generalization performance. For example, weight decay adds a penalty factor to the objective loss function, causing the weight parameters to converge to smaller absolute values. Dropout randomly removes hidden units from the neural network during training and can be considered an ensemble of possible sub-networks. To enhance the capabilities of dropout, new activation functions, maxout, and dropout variants for recurrent neural networks (referred to as rnnDrop) have been proposed. Additionally, batch normalization provides a new regularization method by normalizing the scalar features of each activation within a mini-batch and learning each mean and variance as parameters.
[0022] Given that sequence data is multi-dimensional and high-dimensional, deep neural networks have great promise in bioinformatics research due to their broad applicability and enhanced predictive power. Convolutional neural networks have been used to address sequence-based problems in genomics, such as motif discovery, pathogenic variant identification, and gene expression inference. Convolutional neural networks use a weight-sharing strategy, which is particularly useful for studying DNA as it can capture sequence motifs, which are short and recurring local patterns in DNA that are assumed to have significant biological functions. The hallmark of convolutional neural networks is the use of convolutional filters.
[0023] Unlike traditional classification methods based on precisely designed features and handcrafted features, convolutional filters perform adaptive learning of features, similar to the process of mapping raw input data to an information representation of knowledge. In this sense, convolutional filters act as a series of motif scanners, as a set of such filters is able to identify relevant patterns in the input and update themselves during the training process. Recurrent neural networks can capture long-range dependencies in sequence data of different lengths, such as protein or DNA sequences.
[0024] Therefore, there is an opportunity to use a deep learning-based principle framework for template generation and base calling.
[0025] In the era of high-throughput technologies, it remains a major challenge to accumulate the most interpretable data at the lowest cost per job. Cluster-based nucleic acid sequencing methods, such as those that use bridge amplification to form clusters, have made important contributions to the goal of increasing nucleic acid sequencing throughput. These cluster-based methods rely on sequencing dense populations of nucleic acids immobilized on solid supports and typically involve using image analysis software to deconvolve the optical signals generated during the simultaneous sequencing of multiple clusters located at different positions on the solid support.
[0026] However, such solid-phase nucleic acid cluster-based sequencing technologies still face considerable obstacles that limit the achievable throughput. For example, in cluster-based sequencing methods, there may be obstacles in determining the nucleic acid sequences of two or more clusters that are physically too close to each other to be spatially resolved or that physically overlap on a solid support. For example, current image analysis software may require valuable time and computational resources to determine from which of two overlapping clusters a light signal is emitted. Thus, for a variety of detection platforms, a trade-off in the quantity and / or quality of the nucleic acid sequence information obtainable is inevitable.
[0027] Genomics methods based on high-density nucleic acid clusters have also been extended to other areas of genomic analysis. For example, nucleic acid cluster-based genomics can be used for sequencing applications, diagnostics and screening, gene expression analysis, epigenetic analysis, polymorphic genetic analysis, and the like. Each of these nucleic acid cluster-based genomic technologies is also limited when data generated by closely proximate or spatially overlapping nucleic acid clusters cannot be resolved.
[0028] Clearly, there remains a need to increase the quality and quantity of nucleic acid sequencing data that can be obtained quickly and cost-effectively for a wide variety of uses, including genomics (e.g., for genomic characterization of any and all animal, plant, microbial, or other biological species or groups), pharmacogenetics, transcriptomics, diagnostics, prognosis, biomedical risk assessment, clinical and research genetics, personalized medicine, drug efficacy and drug interaction assessment, veterinary medicine, agriculture, evolutionary and biodiversity research, aquaculture, forestry, oceanography, ecological and environmental management, and other purposes.
[0029] The disclosed technology provides neural network-based methods and systems that address these and similar needs, including increasing the throughput level in high-throughput nucleic acid sequencing technologies, and providing other related advantages.
[0030] The use of deep neural networks and other complex machine learning algorithms can require substantial resources in terms of hardware computing and storage capacity. Additionally, it is desirable to minimize the time required to perform sensing and analysis operations such that the computation. It is desirable to achieve a computation time that makes the results effectively available to the customer in real time. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In the drawings, like reference symbols generally refer to like components throughout the various different views. Additionally, the drawings are not necessarily to scale, but rather emphasize the principles of the disclosed technology. In the following description, various specific implementations of the disclosed technology are described with reference to the following drawings, in which:
[0032] Figure 1 is a simplified diagram of a base calling computing system that includes a configurable processor.
[0033] Figure 2 is a simplified data flow diagram executable by a system such as Figure 1 The same system executes a simplified data flow diagram.
[0034] Figure 3 Illustrates the configuration architecture of components of a configurable or reconfigurable array that supports base calling operations.
[0035] Figure 4 is a diagram of a neural network architecture executable using a configurable or reconfigurable array configured as described herein.
[0036] Figure 5 is a simplified illustration of the organization of blocks of sensor data used by a neural network architecture such as Figure 4 The same neural network architecture uses a simplified illustration of the organization of blocks of sensor data.
[0037] Figure 6 is a simplified illustration of a patch of blocks of sensor data used by a neural network architecture such as Figure 4 The same neural network architecture uses a simplified illustration of a patch of blocks of sensor data.
[0038] Figure 7 Illustrates the configuration of a patch of input blocks used by a neural network architecture such as Figure 4 The same neural network architecture uses a simplified illustration of the configuration of a patch of input blocks.
[0039] Figure 8 Illustrates a portion of the configuration of a neural network such as Figure 4 The same neural network on a configurable or reconfigurable array (such as a field programmable gate array (FPGA)).
[0040] Figure 9 Illustrates the configuration of a multi-cycle machine learning cluster that can be used to execute a neural network such as Figure 4 The same neural network uses a simplified illustration of the configuration of a multi-cycle machine learning cluster.
[0041] Figure 10 is a diagram of an alternative neural network architecture executable using a configurable or reconfigurable array configured as described herein.
[0042] Figure 11 is a diagram of another alternative neural network architecture executable using a configurable or reconfigurable array configured as described herein.
[0043] Figure 12 is a diagram of yet another alternative neural network architecture executable using a configurable or reconfigurable array as described herein, the yet another alternative neural network architecture utilizing a mask that can reduce the memory and processing requirements of all neural network implementations described herein.
[0044] Figure 13Illustrates an implementation of a deep neural network using residual blocks, which can be implemented using a configurable or reconfigurable array as described herein.
[0045] Figure 14 Is an illustration of Figure 1 A host runtime flowchart of a base calling operation that utilizes resources as described herein.
[0046] Figure 15 Is a logical cluster execution flowchart that illustrates the data flow configuration of a logical cluster as described herein.
[0047] Figure 16 Is an alternative logical cluster execution flowchart that illustrates the data flow configuration of a logical cluster as described herein.
[0048] Figure 17 Illustrates a particular implementation of a specialized architecture of a neural network-based base caller for isolating the processing of data for different sequencing cycles.
[0049] Figure 18 Depicts a particular implementation of an isolation layer, each of which may include a convolution.
[0050] Figure 19A Depicts a particular implementation of a combination layer, each of which may include a convolution.
[0051] Figure 19B Depicts another particular implementation of a combination layer, each of which may include a convolution.
[0052] Figure 20 Is a block diagram of a base calling system according to a particular implementation.
[0053] Figure 21 Is Figure 20 A block diagram of a system controller that can be used in the system of
[0054] Figure 22 Is a simplified block diagram of a computer system that can be used to implement the disclosed technology. Detailed Description
[0055] The method of the disclosed technology is specifically implemented as follows: storing block data in a memory, the block data including sensor data of blocks from the sensing cycles of a base calling operation; running a neural network using trained parameters to generate classification data for the sensing cycles, the running of the neural network operating on sequences of N arrays of block data of corresponding sensing cycles from among N sensing cycles including the subject's cycles to generate the classification data for the subject's cycles; and input units moving the block data and the trained parameters from the memory to the neural network for the running of the neural network, the input units including data of N spatially aligned patches of the N arrays from the corresponding sensing cycles among the N sensing cycles.
[0056] Another method disclosed by the present technology is specifically implemented as follows: storing block data in a memory, the block data including an array of sensor data of blocks from the sensing cycles of a base calling operation; and performing a neural network on the block data using a plurality of execution clusters. In this specific implementation, performing the neural network includes: providing input units of the block data to available execution clusters among the plurality of execution clusters, the input units including digital N spatially aligned patches of an array of block data of corresponding sensing cycles including the subject's sensing cycles, and causing the execution clusters to apply the N spatially aligned patches to the neural network to generate output patches of classification data for the spatially aligned patches of the subject's sensing cycles, where N is greater than 1. The output patches may include only data from pixels corresponding to the clusters that generate the base calling data and may thus have a different size and dimension from the input patches.
[0057] Additionally, the methods described herein may include image resampling in a system where the input image orientation jitters around due to camera movement / vibration. To compensate, the image data is resampled and provided as an input to the network to ensure a common basis across images to simplify the operation of the neural network. Resampling involves, for example, applying an affine transformation (translation, shear, rotation) to the image data. The resampling may be performed in software executed by a host processor. In other specific implementations, the resampling may be performed in a configured gate array or other hardware.
[0058] The bit files, model parameters, and runtime programs described herein may be implemented individually or in any combination using one or more non-transitory computer-readable storage media (CRMs) storing instructions including configuration data and model parameters, the instructions being executable by a processor to perform the methods described herein. Each of the features discussed in the specific implementation part of the specific implementation of the method applies equally to the CRM implementation. As shown above, all method features are not repeated here and should be considered repeated by reference.
[0059] Another specific implementation may include a system that includes a memory and one or more processors operable to execute instructions stored in the memory to perform the above-described method.
[0060] A system specific implementation of the disclosed technology includes: a memory that is accessible by a runtime program storing block data, the block data including sensor data from blocks of a sensing cycle of a base calling operation; a neural network processor that can access the memory, the neural network processor being configured to perform an operation of a neural network using trained parameters to generate classification data for the sensing cycle, the operation of the neural network operating on sequences of N arrays of block data from corresponding blocks of N sensing cycles including a subject cycle to generate the classification data for the subject cycle; and data flow logic that uses an input unit to move the block data and the trained parameters from the memory to the neural network processor for the operation of the neural network, the input unit including data of N spatially aligned patches from the N arrays of the corresponding sensing cycle of the N sensing cycles. The neural network processor and the data flow logic can be implemented using a configurable or reconfigurable processor such as an FPGA or a CGRA.
[0061] Another system specific implementation of the disclosed technology includes: a host processor; a memory that can be accessed by the host processor, the memory storing block data, the block data including an array of sensor data from blocks of a sensing cycle of a base calling operation; and a neural network processor that can access the memory, the neural network processor may include a plurality of execution clusters, and the execution logic cluster among the plurality of execution clusters is configured to execute a neural network; and data flow logic that can access the memory and the execution clusters among the plurality of execution clusters to provide an input unit of the block data to an available execution cluster among the plurality of execution clusters, the input units including digital N spatially aligned patches of an array of block data from corresponding sensing cycles (including a subject sensing cycle), and cause the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patches of the subject sensing cycle, where N is greater than 1. The neural network processor and the data flow logic can be implemented using a configurable or reconfigurable processor such as an FPGA or a CGRA.
[0062] Each of the features discussed in the specific implementation part of the method specific implementation equally applies to the system specific implementation. As shown above, all method features are not repeated here and should be considered repeated by reference.
[0063] The disclosed technology (e.g., the disclosed base caller (e.g., Figure 4 and Figure 10)) It can be implemented on processors such as central processing unit (CPU), graphics processing unit (GPU), field programmable gate array (FPGA), coarse grained reconfigurable architecture (CGRA), application specific integrated circuit (ASIC), application specific instruction set processor (ASIP), and digital signal processor (DSP).
[0064] Figure 1 is a simplified block diagram of a system for analyzing sensor data (such as base calling sensor output) from a sequencing system. (See also Figure 21 ). In Figure 1 's example, the system includes a sequencing machine 100 and a configurable processor 150. The configurable processor 150 can execute a neural network-based base caller (e.g., Figure 21 's 2158) in coordination with a runtime program executed by a central processing unit CPU 102. The sequencing machine 100 includes a base calling sensor and a flow cell 101. The flow cell may include one or more blocks where clusters of genetic material are exposed to a sequence of an analysis fluid that is used to cause a reaction in the clusters to identify bases in the genetic material. The sensor senses the reaction for each cycle of the sequence in each block of the flow cell to provide block data. Examples of this technology are described in more detail below. Genetic sequencing is a data-intensive operation that converts base calling sensor data into a base calling sequence for each cluster of genetic material sensed during a base calling operation.
[0065] The system in this example includes a central processing unit 102 that executes a runtime program to coordinate base calling operations, a memory 103 for storing an array of sequences of block data, base calling reads generated by the base calling operations, and other information used in the base calling operations. Additionally, in this illustration, the system includes a memory 104 for storing one configuration file (or files) such as an FPGA bit file and model parameters of a neural network for configuring and reconfiguring the configurable processor 150 and executing the neural network. The machine 100 may include a program for configuring the configurable processor and, in some embodiments, a reconfigurable processor to execute the neural network.
[0066] The sequencing machine 100 is coupled to a configurable processor 150 via a bus 105. The bus 105 can be implemented using high-throughput technologies. For example, in one example, the bus technology is compatible with the PCIe standard (Peripheral Component Interconnect Express) currently maintained and developed by the PCI-SIG (PCI Special Interest Group). Additionally, in this example, a memory 160 is coupled to the configurable processor 150 via a bus 161. The memory 160 can be an on-board memory disposed on a circuit board having the configurable processor 150. The memory 160 is used for the configurable processor 150 to access the working data used in the base calling operation at high speed. The bus 161 can also be implemented using high-throughput technologies such as bus technologies compatible with the PCIe standard.
[0067] Configurable processors, including field-programmable gate arrays (FPGAs), coarse-grained reconfigurable arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than might be achievable using a general-purpose processor that executes a computer program. Configuring a configurable processor involves compiling a functional description to produce a configuration file, sometimes referred to as a bitstream or bitfile, and distributing the configuration file to the configurable elements on the processor.
[0068] The configuration file defines the logical functions to be executed by the configurable processor by configuring the circuit to set the data flow pattern, the use of distributed memory and other on-chip memory resources, the contents of lookup tables, the operation of configurable logic blocks, and configurable execution units (such as multiply-accumulate units, configurable interconnects, and other elements of the configurable array). If the configuration file can be changed in the field by changing the loaded configuration file, the configurable processor is reconfigurable. For example, the configuration file can be stored in volatile SRAM elements, non-volatile read-write memory elements, and combinations thereof, distributed in an array of configurable elements on the configurable or reconfigurable processor. A variety of commercially available configurable processors are suitable for base calling operations as described herein. Examples include commercially available products such as Xilinx Alveo TM U200, Xilinx Alveo TM U250, Xilinx Alveo TM U280, Intel / Altera Stratix TM GX2800, Intel / Altera Stratix TM GX2800, and Intel Stratix TM GX10M. In some examples, the host CPU can be implemented on the same integrated circuit as the configurable processor.
[0069] The embodiments described herein implement a multi-recurrent neural network using a configurable processor 150. The configuration file of the configurable processor can be implemented by specifying the logic functions to be executed using a high-level description language HDL or a register transfer level RTL language specification. The specification can be compiled using the resources designed for the selected configurable processor to generate a configuration file. To generate a design for an application specific integrated circuit that may not be a configurable processor, the same or a similar specification can be compiled.
[0070] Thus, in all embodiments described herein, alternatives to the configurable processor include a configured processor that includes an application specific ASIC or an application specific integrated circuit or a set of integrated circuits, or a system on chip (SOC) device, the configured processor being configured to perform neural network-based base calling operations as described herein.
[0071] Generally speaking, the configurable processor and the configured processor as described herein that are configured to execute the operation of a neural network are referred to herein as neural network processors.
[0072] In this example, the configurable processor 150 is configured by using a configuration file loaded by a program executed by the CPU 102 or other source, the configuration file configuring an array of configurable elements on the configurable processor 150 to perform the base calling function. In this example, the configuration includes data flow logic 151 that is coupled to buses 105 and <161> and performs the function of distributing data and control parameters among the elements used in the base calling operation.
[0073] In addition, the configurable processor 150 is configured with base calling execution logic 152 to execute a multi-recurrent neural network. The logic 152 includes a plurality of multi-recurrent execution clusters (e.g., 153), which in this example includes multi-recurrent cluster 1 to multi-recurrent cluster X. The number of multi-recurrent clusters can be selected based on a trade-off involving the required throughput of the operation and the available resources on the configurable processor.
[0074] The multi-recurrent clusters are coupled to the data flow logic 151 via data flow paths 154 implemented using configurable interconnect and memory resources on the configurable processor. In addition, the multi-recurrent clusters are coupled to the data flow logic 151 via control paths 155 implemented using, for example, configurable interconnect and memory resources on the configurable processor, the control paths providing control signals indicating available clusters, ready to provide input units for executing the operation of the neural network to the available clusters, ready to provide trained parameters for the neural network, ready to provide output patches of base calling classification data, and other control data for executing the neural network.
[0075] The configurable processor is configured to perform a run of a multi-loop neural network using trained parameters to generate classification data for a sensing cycle of a base flow operation. A run of the neural network is performed to generate classification data for a subject sensing cycle for a base calling operation. The run of the neural network operates on a sequence (including N digital arrays of block data from respective sensing cycles out of N sensing cycles), where the N sensing cycles provide sensor data for different base calling operations for one base position for each operation in a time series in the examples described herein. Optionally, if desired, some of the N sensing cycles may be out of order depending on the particular neural network model being executed. The number N can be any number greater than 1. In some examples described herein, the sensing cycles out of the N sensing cycles represent a set of sensing cycles that includes at least one sensing cycle before a subject sensing cycle and at least one sensing cycle after the subject cycle in a time series. Examples are described herein where the number N is an integer equal to or greater than five.
[0076] The data flow logic is configured to move at least some of the trained parameters of the block data and the model parameters from the memory 160 to the configurable processor for the run of the neural network using input units (including block data of spatially aligned patches of N arrays) for a given run. The input units may be moved by a direct memory access operation in one DMA operation or in smaller units moved in coordination with the execution of the deployed neural network during available time slots.
[0077] The block data for a sensing cycle as described herein may include an array of sensor data having one or more features. For example, the sensor data may include two images that are analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The block data may also include metadata about the images and the sensors. For example, in an embodiment of a base calling operation, the block data may include information about the alignment of an image with a cluster, such as information about the distance from the center, which indicates the distance of each pixel in the sensor data array from the center of the cluster of genetic material on the block.
[0078] During the execution of the multi-loop neural network as described below, the block data may also include data generated during the execution of the multi-loop neural network, called intermediate data, which may be reused during the run of the multi-loop neural network instead of being recomputed. For example, during the execution of the multi-loop neural network, the data flow logic may write the intermediate data to the memory 160 in place of the sensor data for a given patch of the block data array. Embodiments similar to this are described in more detail below.
[0079] As shown, a system for analyzing base detection sensor outputs is described. The system includes a memory (e.g., 160) accessible by a runtime program, which stores block data including sensor data from blocks of sensing cycles of a base detection operation. Additionally, the system includes a neural network processor, such as a configurable processor 150 capable of accessing the memory. The neural network processor is configured to execute a run of a neural network using trained parameters to generate classification data for a sensing cycle. As described herein, the run of the neural network operates on sequences of N arrays of block data from corresponding sensing cycles out of N sensing cycles (including a subject cycle) to generate classification data for the subject cycle. Data flow logic 151 is provided to move block data and trained parameters from the memory to the neural network processor for the run of the neural network using input units (including data of N spatially aligned patches from N arrays of corresponding sensing cycles out of N sensing cycles).
[0080] Additionally, a system is described in which the neural network processor is capable of accessing a memory and includes a plurality of execution clusters, and an execution logic cluster among the plurality of execution clusters is configured to execute a neural network. The data flow logic is capable of accessing the memory and an execution cluster among the plurality of execution clusters to provide input units of block data to an available execution cluster among the plurality of execution clusters, the input units including digital N spatially aligned patches from arrays of block data from corresponding sensing cycles (including a subject sensing cycle), and to cause the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patches of the subject sensing cycle, where N is greater than 1.
[0081] Figure 2is a simplified diagram showing aspects of a base calling operation that includes functions of a runtime program executed by a host processor. In this diagram, the output from the image sensor of the flow cell is provided on line 200 to an image processing thread 201 that can perform processing on the image, such as resampling, alignment, and arrangement in the sensor data array of each tile, and can be used by a process that calculates a tile cluster mask for each tile in the flow cell that identifies pixels in the sensor data array corresponding to clusters of genetic material on the corresponding tile of the flow cell. To calculate the cluster mask, an exemplary algorithm is based on a process for using a metric derived from the softmax output to detect clusters that are unreliable in early sequencing cycles, then discarding data from those wells / clusters and not producing output data for those clusters. For example, the process can identify clusters with high reliability during the first N (e.g., 25) base calls and reject other clusters. The rejected clusters may be polyclonal or very weak in intensity or have ambiguous fiducials. The program can be executed on the host CPU. In an alternative embodiment, this information will potentially be used to identify the necessary clusters of interest to pass back to the CPU, thereby limiting the storage required for intermediate data (i.e., the "dehydration" step described below can view all pixels with wells, or only processing pixels with wells / clusters that pass a filter test can be achieved more efficiently).
[0082] Depending on the state of the base calling operation, the output of the image processing thread 201 is provided on line 213 to scheduling logic 210 in the CPU that routes the tile data array on a high-speed bus 214 to a data cache 204, or on a high-speed bus 205 to a multi-cluster neural network processor hardware 220, such as Figure 1 a configurable processor. The hardware 220 returns the classification data output by the neural network to the scheduling logic 210 that passes the information to the data cache 204, or on line 211 to a thread 202 that performs base calling and quality score calculations using the classification data and can arrange the data for base call reads in a standard format. The output of the thread 202 that performs base calling and quality score calculations is provided on line 212 to a thread 203 that aggregates the base call reads, performs other operations such as data compression, and writes the resulting base call output to a specified destination for customer use.
[0083] In some embodiments, the host may include a thread (not shown) that performs final processing of the output of the execution hardware 220 to support a neural network. For example, the hardware 220 may provide the output of classification data from the final layer of a multi-cluster neural network. The host processor may perform an output activation function, such as a softmax function, on the classification data to configure the data for use by the base calling and quality scoring threads 202. Additionally, the host processor may perform input operations (not shown), such as resampling, batch normalization, or other adjustments to the block data before input to the hardware 220.
[0084] Figure 3 is a simplified diagram of the configuration of a configurable processor such as Figure 1 the configurable processor. In Figure 3 , the configurable processor includes an FPGA having a plurality of high-speed PCIe interfaces. The FPGA is configured with a wrapper 300 that includes data flow logic described with reference to Figure 1 . The wrapper 300 manages the interface and coordination with the runtime program in the CPU via the CPU communication link 309 and manages communication with the on-board DRAM 302 (e.g., memory 160) via the DRAM communication link 310. The data flow logic in the wrapper 300 provides patch data retrieved by traversing a digital N-cycle block data array on the on-board DRAM 302 to the cluster 301 and retrieves process data 315 from the cluster 301 for delivery back to the on-board DRAM 302. The wrapper 300 also manages data transfer between the on-board DRAM 302 and the host memory for both the input array of block data and the output patch of classification data. The wrapper assigns patch data to the cluster 301 on line 313. The wrapper provides trained parameters, such as weights and biases, to the cluster 301 retrieved from the on-board DRAM 302 on line 312. The wrapper provides configuration and control data to the cluster 301 on line 311, which is provided by or generated in response to the runtime program on the host via the CPU communication link 309. The cluster may also provide a status signal to the wrapper 300 on line 316, which is used in cooperation with control signals from the host to manage the traversal of the block data array to provide spatially aligned patch data and perform a multi-cycle neural network on the patch data using the resources of the cluster 301.
[0085] As described above, there may be multiple clusters on a single configurable processor managed by the wrapper 300, which are configured to perform on corresponding patches of multiple patches of block data. Each cluster may be configured to provide classification data for base calling in a subject sensing cycle using the multiple sensing cycles of block data described herein.
[0086] In an example of the system, model data (including kernel data such as filter weights and biases) can be sent from the host CPU to the configurable processor so that the model can be updated according to the number of cycles. As a representative example, a base calling operation can include approximately hundreds of sensing cycles. In some embodiments, the base calling operation can include paired-end reads. For example, the model training parameters can be updated every 20 cycles (or other number of cycles), or according to an update pattern implemented for a particular system and neural network model. In some embodiments that include paired-end reads, where the sequence of a given string in a genetic cluster on a tile includes a first portion extending down (or up) the string from a first end and a second portion extending up (or down) the string from a second end, the trained parameters can be updated at the transition from the first portion to the second portion.
[0087] In some examples, image data for multiple cycles of sensing data for a tile can be sent from the CPU to wrapper 300. Wrapper 300 can optionally perform some preprocessing and transformation on the sensing data and write the information to on-board DRAM 302. The input tile data for each sensing cycle can include a sensor data array, including approximately 4000×3000 pixels or more per tile per sensing cycle, where two features represent the colors of two images of the tile, and each feature is one or two bytes per pixel. For embodiments where the number N is three sensing cycles to be used in each run of the multi-cycle neural network, the tile data array for each run of the multi-cycle neural network can consume approximately hundreds of megabytes per tile. In some embodiments of the system, the tile data also includes an array of DFC data stored once per tile, or other types of metadata regarding the sensor data and the tile.
[0088] In operation, when a multi-cycle cluster is available, the wrapper assigns patches to the cluster. The wrapper fetches the next patch of tile data in the traversal of the tile and sends it, along with appropriate control and configuration information, to the assigned cluster. The cluster can be configured to have sufficient memory on the configurable processor to hold data patches that include patches from multiple cycles in some systems and are being processed in-place, as well as data patches that will be processed when the processing of the current patch is completed using ping-pong buffering techniques or raster scanning techniques in various embodiments.
[0089] When an allocated cluster finishes its run of the neural network for the current patch and produces an output patch, it signals the wrapper. The wrapper reads the output patch from the allocated cluster, or alternatively, the allocated cluster pushes the data to the wrapper. The wrapper then assembles the output patches for the processed blocks in DRAM 302. When the processing of the entire block is complete and the output patches of the data have been transferred to DRAM, the wrapper sends the processed output array of the block back to the host / CPU in a specified format. In some embodiments, the on-board DRAM 302 is managed by memory management logic in the wrapper 300. The runtime program can control the sequencing operations to complete the analysis of all arrays of block data for all cycles in the run in a continuous stream manner, providing real-time analysis.
[0090] Figure 4 is a diagram of a multi-cycle neural network model that can be executed using the system described herein. Figure 4 The example shown can be referred to as a five-cycle input, one-cycle output neural network. The input to the multi-cycle neural network model includes five spatially aligned patches (e.g., 400) of block data arrays from five sensing cycles of a given block. The spatially aligned patches have the same aligned row and column dimensions (x,y) as the other patches in the set, such that the information pertains to the same clusters of genetic material on the block in the sequence of cycles. In this example, the subject patch is a patch of the block data array from cycle N. A set of five spatially aligned patches includes a patch from cycle N-2, which is two cycles before the subject patch, a patch from cycle N-1, which is one cycle before the subject patch, a patch from cycle N+1, which is one cycle after the patch from the subject cycle, and a patch from cycle N+2, which is two cycles after the patch from the subject cycle.
[0091] The model includes an isolated stack 401 of layers of a neural network for each input patch in the input patches. Thus, stack 401 receives as input the block data of the patches from cycle N+2 and is isolated from stacks 402, 403, 404, and 405 such that they do not share input data or intermediate data. In some embodiments, all of the stacks 410-405 may have the same model and the same trained parameters. In other embodiments, the model and the trained parameters may be different in different stacks. Stack 402 receives as input the block data of the patches from cycle N+1. Stack 403 receives as input the block data of the patches from cycle N. Stack 404 receives as input the block data of the patches from cycle N-1. Stack 405 receives as input the block data of the patches from cycle N-2. The layers of the isolated stack each perform a convolutional operation of a kernel that includes a plurality of filters on the input data of the layer. As in the above example, patch 400 may include three features. The output of layer 410 may include more features, such as 10 to 20 features. Similarly, the output of each of layers 411 to 416 may include any number of features suitable for a particular implementation. The parameters of the filters are the trained parameters of the neural network, such as weights and biases. The set of output features (intermediate data) from each of stacks 401-405 is provided as input to an inverse hierarchy 420 of a temporal combination layer, where the intermediate data from multiple cycles is combined. In the illustrated example, inverse hierarchy 420 includes: a first layer that includes three combination layers 421, 422, 423, each combination layer receiving the intermediate data from three of the isolated stacks in the isolated stack; and a final layer that includes a combination layer 430 that receives the intermediate data from the three temporal layers 421, 422, 423.
[0092] The output of the final combination layer 430 is an output patch of classification data of the clusters located in the corresponding patches of the blocks from cycle N. The output patches can be assembled into an output array of classification data of the blocks of cycle N. In some embodiments, the output patches may have a different size and dimension than the input patches. In some embodiments, the output patches may include per-pixel data that can be filtered by a host to select the cluster data.
[0093] According to a particular implementation, the output classification data may then be applied to a softmax function 440 (or other output activation function) optionally executed by a host or on a configurable processor. An output function different from softmax may be used (e.g., generating a base call output parameter based on the maximum output and then giving a base quality using a learned non-linear mapping of the context / network output).
[0094] Finally, the output of the softmax function 440 can be provided as the base calling probability (450) for cycle N and stored in the host memory for use in subsequent processing. Other systems may use another function for output probability calculation, e.g., another non-linear model.
[0095] A configurable processor with multiple execution clusters can be used to implement the neural network so as to complete the evaluation of a block cycle within a duration equal to or close to the duration of one sensing cycle, thereby effectively providing output data in real time. The data flow logic can be configured to distribute the input units of the block data and the trained parameters to the execution clusters and distribute the output patches for aggregation in the memory.
[0096] Reference Figure 5 and Figure 6 describes the input units of data for a five-cycle input, one-cycle output neural network for base calling operations using dual-channel sensor data as Figure 4 For example, for a given base in a gene sequence, the base calling operation can perform two analytical streams and two reactions that generate two signal (e.g., image) channels, and these images can be processed to identify which of the four bases is at the current position of the genetic sequence in each cluster of the genetic material. In other systems, a different number of channels of sensing data can be utilized.
[0097] Figure 5 Shows a block data array for five cycles for a given block (block M), which is used for the purpose of implementing a five-cycle input, one-cycle output neural network. The five-cycle input block data in this example can be written to the on-board DRAM or other memory in the system accessible by the data flow logic, and includes an array 501 for channel 1 and an array 511 for channel 2 for cycle N-2, an array 502 for channel 1 and an array 512 for channel 2 for cycle N-1, an array 503 for channel 1 and an array 513 for channel 2 for cycle N, an array 504 for channel 1 and an array 514 for channel 2 for cycle N+1, and an array 505 for channel 1 and an array 515 for channel 2 for cycle N+2. Additionally, an array 520 of the metadata of the block can be written to the memory once, in which case, it includes a DFC file to be used as an input to the neural network along with each cycle.
[0098] The data flow logic constitutes the input units of the block data, which can be referenced Figure 6It is understood that the block data includes spatial alignment patches of the block data arrays of each execution cluster, and each execution cluster is configured to execute the operation of a neural network on the input patches. The input units of the execution cluster for allocation are constituted by the data flow logic in the following manner: reading spatial alignment patches (e.g., 601, 602, 611, 612, 620) from each of the block data arrays 501-505, 511, 515, 520 of five input loops, and delivering them via a data path (schematically, 600) to a memory on a configurable processor configured to be used by the allocated execution cluster. The allocated execution cluster executes the operation of a five-loop input / one-loop output neural network, and delivers an output patch of classification data for the same patch of blocks in subject loop N for subject loop N.
[0099] Figure 7 Illustrates the mapping of patches on the block data array of a given block. In this example, the input array 700 of the block data has a width of X pixels and a height of Y pixels. After convolving a kernel (such as a 3×3 kernel with a step size of one pixel) in multiple layers of the neural network, the output block 701 can reduce by two rows and two columns for each layer of the neural network. In this example, the reduction of two rows / columns is caused by the 3×3 kernel size and the type of (edge) padding in use, and can vary with the configuration. Thus, for example, for a neural network with L / 2 layers including this type of convolution, the output block 701 of the classification data will have a width of X - L pixels. Similarly, for a neural network with L / 2 layers, the output block of the classification data will have a height of Y - L pixels. For example, taking a neural network with six layers as an example, L can be 12 pixels. In Figure 7 the example shown, the patch regions are not drawn to scale.
[0100] The input patches are formed in an overlapping manner to account for the lost pixels generated by convolution beyond the patch size. The size of the input patches can be selected according to a specific implementation. In one example, the input patches can have a size of 76×76 pixels, where each pixel has three channels with one or more bytes. The output patches can have a size of 64×64 pixels. In an implementation, the base detection operation for A / C / T / G base detection outputs classification and the output patches can include four channels with one or more bytes for each pixel, thereby representing the confidence score of the classification. In Figure 4 the example, the output on line 435 is the unnormalized confidence scores of four base detections.
[0101] The data stream logic can address a block data array to patches in a raster scan manner or other scan manner to provide input patches (e.g., 705). For example, for the first available cluster, patch P0,0 can be provided. For the next available cluster, patch P0,1 can be provided. This sequence can continue in a raster pattern until all patches of the block are delivered to the available clusters for processing.
[0102] In some embodiments, the output patches (e.g., 706) can be rewritten to the same address space aligned with their subject input patches, thereby accounting for any differences in the number of bytes per pixel used to encode the data. The area (number of pixels) of the output patches decreases relative to the input patches based on the number of convolutional layers and the nature of the convolutions performed.
[0103] Figure 8 is a simplified representation of a stack of neural networks that can be used in a system such as Figure 4 (e.g., 401 and 420). In this example, some functions of the neural network are executed on a host (e.g., 800, 802), and other parts of the neural network are executed on a configurable processor (801).
[0104] The first function can be batch normalization (layer 810) formed on the CPU. Batch normalization is a training technique that improves overall performance results by normalizing data on a per-batch basis, but other techniques can be used. Several parameters of each layer are calculated and updated during training.
[0105] During inference, the batch normalization parameters are not adjusted and are fixed at long-term averages. Multiplication operations can be fused into adjacent layers to reduce the total operation count. This fusion is performed during fixed-point model retraining, so the inference neural network can include a single addition within each BN layer of the BN layers that can be implemented in the configurable processor. The first batch normalization layer 810 is executed on the CPU during training. This is the only layer executed on the CPU. The output of the quantized batch normalization calculation is transmitted to the configurable processor for further processing. After training, batch normalization is replaced with fixed scaling and bias addition. In the first layer, this scaling and bias addition occurs on the CPU. In a pure inference implementation, the batch normalization term is not actually needed.
[0106] As discussed above with respect to the configurable processor, multiple spatially isolated convolutional layers are performed as the first set of convolutional layers of the neural network. In this example, the first set of convolutional layers applies 2D convolutions spatially.
[0107] These convolutions can be efficiently implemented using stacks of 2D Winograd convolutions. The operations are applied independently to each patch in each cycle. The multi-cycle structure is retained through these layers. There are different ways to implement the convolution, which are effective ways for digital logic and programmable processors.
[0108] As Figure 8 shown, for each of the L / 2 (where L is as referenced Figure 7 described) spatially separated neural network layers in each stack, a first spatial convolution 821 is performed, followed by a second spatial convolution 822, followed by a third spatial convolution 823, and so on. As indicated at 823A, the number of spatial layers can be any actual number, which for the context can range from a few to more than 20 in different embodiments.
[0109] For SP_CONV_0, the kernel weights are stored, for example, in a (1, 6, 6, 3, L) structure because there are 3 input channels for this layer. In this example, the "6" in this structure is attributed to storing the coefficients in the transformed Winograd domain (the kernel size is 3×3 in the spatial domain but is extended in the transformed domain).
[0110] For this example, for the other SP_CONV layers, the kernel weights are stored in a (1, 6, 6L) structure because for each of these layers, there are K (= L) inputs and outputs.
[0111] The output of the stack of spatial layers is provided to the temporal layers, including convolution layers 824, 825 executed on the FPGA. Layers 824 and 825 can be convolution layers that apply 1D convolutions across cycles. As indicated at 824A, the number of temporal layers can be any actual number, which for the context can range from a few to more than 20 in different embodiments.
[0112] The first temporal layer, TEMP_CONV_0 layer 824, reduces the number of cyclic channels from 5 to 3, as Figure 4 shown. The second temporal layer (layer 825) reduces the number of cyclic channels from 3 to 1, as Figure 4 shown, and reduces the number of feature map maps to four outputs for each pixel, thereby representing the confidence in each base call.
[0113] The output of the temporal layers is accumulated in the output patch and delivered to the host CPU to apply, for example, a softmax function 830 or other function to normalize the base call probabilities.
[0114] Figure 9It is a block diagram of a configuration of an execution cluster suitable for executing a multi - loop neural network as described herein. In this example, the execution cluster includes a plurality of execution engines ENG.0 to ENG.E - 1 (e.g., 900, 901, 902). The number N can be any value selected based on design trade - offs. For a practical example, when the engines are implemented on a single FPGA, for each cluster, the number N can be in the range of 6 to 10, but more or fewer engines can be configured. Thus, the execution cluster includes a set of computing engines having a plurality of members, configured to perform convolution on the input data of multiple layers of a neural network using trained parameters, where the input data of the first layer is from an input unit, and the data of subsequent layers is from the activation data output from the previous layer.
[0115] The engines include: a front - end that provides current data, filter parameters, and control to the assigned engines that execute loops of a multiply - accumulate function that supports convolution; and a back - end that includes circuitry for applying biases and other functions that support the neural network. In this example, the front - end includes a plurality of selectors 920, 921, 922 that are used to select a data source for the assigned engines from input patches (PATCH) delivered by data - flow logic from on - board DRAM or other memory sources or from activation data fed back from the previous layer on line 950 from the back - end 910. Additionally, the front - end includes a filter storage area 925 that is connected to a source that stores filter parameters for performing convolution. The filter storage area 925 is coupled to each of the engines, and provides appropriate parameters according to the portion of the convolution being performed in a particular engine. The output of the engines includes a first path 930, 931, 932 that provides the result to the back - end 910. Additionally, the output of the engines includes a second path 940, 941, 942 that is connected to the data - flow logic in the wrapper for routing the data back to on - board memory or other memory utilized by the system. Additionally, the cluster includes a bias storage area 926 loaded with various bias values BN_BIAS and BIAS utilized when executing the neural network. The bias storage area 926 is operated to provide a particular bias value to the back - end (910) process according to the particular state of the neural network being executed.
[0116] A cluster among the plurality of clusters in a neural network processor can be configured to include a kernel memory to store trained parameters, and the data - flow logic can be configured to provide an instance of the trained parameters to the kernel memory of an execution cluster among the plurality of execution clusters for executing the neural network. The weights for a layer can be, for example, trained parameters that are changed according to a sensing cycle (such as every X cycles), where X can be a constant selected empirically such as 15 or 20. Additionally, the number X may be variable according to the characteristics of the sequencing system that provides the input block data.
[0117] Figure 10A diagram of a multi-recurrent neural network model that can be equivalent to Figure 4 the model shown, but reuses intermediate data for execution to save computing resources. As Figure 10 shown, the model is a five-recurrent input and one-recurrent output neural network. The input to the neural network includes the current patch Px,y(1000) from the block data array, which includes sensor data from the current sensing cycle of a given block. Additionally, the input includes intermediate data Int(Px,y) from previous spatially aligned patches from cycles N-2, N-1, N, and N+1.
[0118] The input patch 1000 from cycle N+2 is applied to an isolation stack including layers 1010, 1011, 1012, 1013, 1014, 1015, 1016. The output of the final layer 1016 of the isolation stack is applied on line 1001 to be used as intermediate data for subsequent runs of the neural network and passed down for repeated use multiple times as shown for lines 1053, 1054, 1055. This intermediate data can be rewritten to the on-board DRAM as block data for a specific block and specific cycle, or rewritten to other memory resources. In some embodiments, the intermediate data can be written at the same location as the original sensor data array of the block data of the spatially aligned patch.
[0119] In Figure 10 it, the input patch 1000 has multiple features F1 applied to the first layer 1010 of the neural network. Layer 1010 outputs multiple features F2 to the next layer 1011. The next layer outputs multiple features F3, and so on, such that subsequent layers respectively output F4 features, F5 features, F6 features, F7 features, and F8 features. In some embodiments, the number of features F1 can be 3, as discussed above with respect to the sensor data of the input patch. The number of features F2 to F8 can be 12 or greater and can be the same or different. In this case, the 12 (or more) features of the final layer 1016 may consume more storage than the input patch 1000 because it includes more features. In other embodiments, the number of features F2 to F8 may vary.
[0120] In some storage-saving embodiments, the multiple features F8 can be reduced when used as intermediate data to conserve memory resources. For example, the number of the multiple features F8 can be 2 features or 3 features, such that it consumes the same or less memory space than the original input patch when stored as block data for the subject cycle.
[0121] In Figure 10In the model shown, the intermediate data of loops N-2, N-1, N, and N+1 are used as inputs to time layers 1021, 1022, 1023 and are applied on lines 1002, 1003, 1004, 1005 in the manner discussed. Additionally, the output of layer 1016 of the isolation stack is applied as an input to layer 1021 in the time layer. The final time layer 1030 receives the outputs of time layers 1021-1023 as inputs. The output of layer 1030 is applied on line 1035 to the softmax function 1040 or other activation function, and the output of this function is provided as the base calling probability 1050 of the subject loop (loop N). Figure 4 For the neural network model of , the block data of the sensing loop in the memory includes the sensor data of the current sensing loop (N+2) and the intermediate data fed back from the neural network of earlier sensing loops (N-2, N-1, N, and N+1).
[0122] Thus, for Figure 10 the neural network model, the block data of the sensing loop in the memory includes the sensor data of the current sensing loop (N+2) and the intermediate data fed back from the neural network of earlier sensing loops (N-2, N-1, N, and N+1).
[0123] Figure 11 Illustrates an alternative embodiment of a 10-input, six-output neural network that can be performed for base calling operations. In a system such as Figure 11 as, a method of saving and reusing as Figure 10 can be used to substantially improve efficiency. In this example, the block data of the spatially aligned input patches from loops 0 to 9 are applied to the isolation stack of the spatial layer, such as stack 1101 of loop 9. The output of the isolation stack is applied to the inverse hierarchical arrangement of the time stack 1120 with outputs 1135(2) to 1135(7), thereby providing the base calling classification data for subject loops 2 to 7.
[0124] Figure 12 Illustrates Figure 10 an improvement to the model shown and includes similar features with the same reference numerals. Figure 12 Different from Figure 10 in that it includes a "dehydration" filter layer 1251 and a block cluster mask 1252 at the output of layer 1016. The block cluster mask 1252 can be generated by preprocessing the sensor data to identify the pixels in the image corresponding to the clusters of the genetic material being sequenced. The block cluster mask 1252 can be applied in a configurable processor that includes mask logic to apply the block cluster mask to the intermediate data in the neural network. The mask logic can include a "dehydration" filter layer to select only the pixels with data relevant to the base calling operation (e.g., reliable cluster data) and rearrange the pixels in a smaller intermediate data array, and thus can substantially reduce the size of the intermediate data, as well as the size of the data applied to the time layer.
[0125] The block cluster mask 1252 can be generated, for example, as described in U.S. Patent Application Publication No. US 2012 / 0020537 to Garcia et al., which is incorporated by reference in its entirety as if fully set forth herein, where the feature identification locations can be the locations of clusters of genetic material.
[0126] The block cluster mask 1252 can identify those pixels corresponding to unreliable clusters and can be used by the dehydration filter layer to discard / filter out such pixels, and thereby apply the temporal layer only to those pixels corresponding to reliable clusters. In a particular implementation, the base calling classification score is generated by the output layer. Examples of output layers include the softmax function, log-softmax function, integrated output averaging function, multi-layer perceptron uncertainty function, Bayesian Gaussian distribution function, and cluster intensity function. In a particular implementation, the output layer produces a probability quadruple per cluster per cycle for each cluster and for each sequencing cycle.
[0127] The following discussion focuses on the probability quadruple per cluster per cycle using the softmax function as an example of the output layer. The softmax function is explained first, and then the probability quadruple per cluster per cycle, which are both used to identify unreliable clusters.
[0128] The softmax function is a preferred function for multi-class classification. The softmax function calculates the probability of each target class relative to all possible target classes. The output of the softmax function ranges between zero and one, and the sum of all probabilities equals one. The softmax function calculates the sum of the exponential values of the given input value and all input values. The ratio of the exponential value of the input value to the sum of the exponential values is the output of the softmax function, referred to herein as "exponential normalization".
[0129] Formally, training a so-called softmax classifier is a regression to class probabilities, rather than a regression to a true classifier, because it does not return the class, but rather a confidence prediction of the probability of each class. The softmax function takes a set of class values and converts them to probabilities that sum to 1. The softmax function compresses an n-dimensional vector of arbitrary real values into an n-dimensional vector of real values in the range 0 to 1. Thus, using the softmax function ensures that the output is a valid, exponentially normalized probability mass function (non-negative and summing to 1).
[0130] Intuitively, the softmax function is a "softened" version of the maximum function. The term "soft" comes from the fact that the softmax function is continuous and differentiable. Instead of choosing a single maximum element, it breaks the vector into fractional parts, where the largest input element gets a proportionally larger value and the other elements get proportionally smaller values. The property of outputting a probability distribution makes the softmax function suitable for probabilistic interpretation in classification tasks.
[0131] Let z be the vector of inputs to the softmax layer. The softmax layer units are the number of nodes in the softmax layer, and thus the length of the z vector is the number of units in the softmax layer (if there are ten output units, there are ten z elements).
[0132] For an n-dimensional vector Z = [z1, z2,... z n , the softmax function uses exponential normalization (exp) to produce another n-dimensional vector p(Z), where the normalized values are in the range [0, 1] and sum to one:
[0133]
[0134] The Softmax function is applied to three classes as follows: Note that the three outputs always sum to 1. Thus, they define a discrete probability mass function.
[0135] The specific per-cluster per-cycle probability quadruplets identify the probabilities of base incorporation into a specific cluster when the bases are A, C, T, and G at a specific sequencing cycle. When the output layer of a neural network-based base caller uses the softmax function, the probabilities in the per-cluster per-cycle probability quadruplets are exponentially normalized classification scores that sum to one. The unreliable cluster identifier identifies unreliable clusters based on generating filter values from the per-cluster per-cycle probability quadruplets. In this application, the per-cluster per-cycle probability quadruplets are also referred to as base calling classification scores or normalized base calling classification scores or initial base calling classification scores or normalized initial base calling classification scores or initial base calls.
[0136] The filter calculator determines the filter value for each per-cluster per-cycle probability quadruplet based on the probabilities it identifies, thereby generating a sequence of filter values for each cluster. The sequence of filter values is stored as the filter values.
[0137] The filter value for the per-cluster per-cycle probability quadruplets is determined based on arithmetic operations involving one or more of the probabilities. In one specific implementation, the arithmetic operation used by the filter calculator is subtraction. In one specific implementation, the filter value for the per-cluster per-cycle probability quadruplets is determined by subtracting the second-highest probability from the highest probability among the probabilities.
[0138] In another specific implementation, the arithmetic operation used by the filter calculator is division. For example, the filter value for each of the four elements of the probability per cluster per cycle is determined as the ratio of the highest probability in the probabilities to the second highest probability in the probabilities. In yet another specific implementation, the arithmetic operation used by the filter calculator is addition. In still another specific implementation, the arithmetic operation used by the filter calculator is multiplication.
[0139] In one specific implementation, the filter calculator uses a filtering function to generate filter values. In one example, the filtering function is a chastity filter that defines chastity as the ratio of the brightest base intensity divided by the sum of the brightest base intensity and the second brightest base intensity. In another example, the filtering function is at least one of a maximum log probability function, a least squares error function, an average signal-to-noise ratio (SNR), and a least absolute error function.
[0140] The unreliable cluster identifier uses the filter values to identify some of the multiple clusters as unreliable clusters. The data for identifying the unreliable clusters can be in a computer-readable format or located on a computer-readable medium. The unreliable clusters can be identified by an instrument ID, a run number on the instrument, a flow cell ID, a lane number, a tile number, an X coordinate of the cluster, a Y coordinate of the cluster, and a unique molecular identifier (UMI). The unreliable cluster identifier identifies those clusters among the multiple clusters as unreliable clusters whose sequence of filter values contains "N" filter values below a threshold "M". In one specific implementation, the range of "N" is from 1 to 5. In another specific implementation, the range of "M" is from 0.5 to 0.99. In one specific implementation, the unreliable cluster identification corresponds to those pixels that correspond to the unreliable clusters (i.e., depict the intensity emissions of these unreliable clusters).
[0141] Unreliable clusters are low-quality clusters that emit a negligible amount of the desired signal compared to the background signal. The signal-to-noise ratio of the unreliable clusters is substantially lower, e.g., less than one. In some specific implementations, the unreliable clusters may not produce any amount of the desired signal. In other specific implementations, the unreliable clusters may produce a very low amount of signal relative to the background. In one specific implementation, the signal is an optical signal and is intended to include, for example, a fluorescence signal, a luminescence signal, a scattering signal, or an absorption signal. The signal level refers to the amount or quantity of detected energy or encoded information having a desired or predefined characteristic. For example, an optical signal can be quantified by one or more of intensity, wavelength, energy, frequency, power, brightness, etc. Other signals can be quantified according to characteristics such as voltage, current, electric field strength, magnetic field strength, frequency, power, temperature, etc. The absence of a signal in the unreliable clusters is understood as a signal level of zero or a signal level that is not significantly distinguishable from the noise.
[0142] There are many potential reasons for the poor signal quality of unreliable clusters. If polymerase chain reaction (PCR) errors already exist in colony amplification such that a substantial proportion of the approximately 1000 molecules in an unreliable cluster contain different bases at a certain position, signals of two bases can be observed - this is interpreted as a sign of poor quality and is called a phase error. Phase errors occur when individual molecules in an unreliable cluster do not incorporate nucleotides in a certain cycle (e.g., due to incomplete removal of 3' terminators, called phasing) and then lag behind other molecules, or when individual molecules incorporate more than one nucleotide in a single cycle (e.g., due to ineffective 3'-blocking of nucleotide incorporation, called pre-phasing). This results in a loss of synchronization in the reads of sequence copies. The proportion of sequences in unreliable clusters affected by phasing and pre-phasing increases with the number of cycles, which is the main reason why the quality of reads tends to decline at high cycle numbers.
[0143] Unreliable clusters are also caused by fading. Fading is the exponential decay of the signal intensity of unreliable clusters according to the number of cycles. As the sequencing run progresses, the strands in unreliable clusters are overwashed, exposed to laser emissions that generate reactive substances, and subjected to harsh environmental conditions. All of these lead to a gradual loss of fragments in unreliable clusters, thereby reducing their signal intensity.
[0144] Unreliable clusters are also caused by underdeveloped colonies, i.e., unreliable clusters of small cluster sizes that produce empty or partially filled wells on the patterned flow cell. That is, in some specific embodiments, unreliable clusters indicate empty, polyclonal, and dim wells on the patterned flow cell. Unreliable clusters are also caused by overlapping colonies due to unrestricted amplification. Unreliable clusters are also caused by insufficient illumination or uneven illumination, for example, due to being located on the edge of the flow cell. Unreliable clusters are also caused by impurities on the flow cell that obscure the emitted signal. When multiple clusters are deposited in the same well, unreliable clusters also include polyclonal clusters.
[0145] Figure 13 An alternative specific embodiment of an executable deep neural network is illustrated. Figure 13 In the form of Figure 8 and includes the same reference numerals for the same components. In Figure 13 , one or more spatial convolutional layers in the spatial convolutional layer in the isolation stack are replaced by a residual block structure 1301 (ResBLOCK). The residual block structure in an exemplary configuration may include a first convolutional layer and a second convolutional layer that receive as input the sum of the input to the first convolutional layer and the output of the first convolutional layer. The second convolutional layer in the residual block 1301 may not have an activation function after addition.
[0146] Figure 14 is for using as Figure 1A simplified flowchart of base calling operations performed by the same system. In this simplified process, the runtime program executed by the CPU 102 initiates the base calling operation (1401). This can include instantiating one or more processing threads that can run on multiple cores of the host CPU. In the base calling operation, the sensor and the flow cell execute in a cyclic sequence in which a block data array is generated. The processing thread waits for and receives data (1402) from the sensed cyclic sequence (e.g., block image) of clusters across the genetic material in the block of the flow cell. The output from the flow cell is processed, for example, between the sensing operation and the neural network, by resampling such that each image has a common structure and a common reference frame, e.g., fiducial points are aligned, clusters / wells are aligned, and DFC data is constant; and arranging the data in a sequence of block data arrays, where each sequence can include a block data array that includes F digital features for each sensing cycle (1403). This resampling and arrangement of the block image is performed because the neural network performance is improved when the same pixels in each training input in the training input of the neural network convey the same information about the underlying signal from the flow cell. The processing thread transfers the block data array and the trained neural network parameters to the memory 1060 via a wrapper (data flow logic 151) on the configurable processor. When sufficient data (such as N complete arrays of block data including subject blocks and multiple adjacent blocks of N cycles) has been loaded into the memory 140 (e.g., in response to control signals such as control triggers and / or control tokens (feedforward single pulses) formed or issued on the bus), the processing thread can submit a job to the inference engine by signaling an event to the wrapper, or the wrapper can detect the event to start the operation of the neural network on a specific subject block (1404). The wrapper loads the trained parameters of the neural network into available logic clusters, allocates the available logic clusters, and loads spatially aligned patches from the digital N cycles (1405). The wrapper can update the neural network parameters, for example, using ping-pong buffering, according to the number of cycles in the sequence (1406). For example, in a particular implementation, the neural network model parameters (weights and biases) can be updated every 20 cycles. Thus, the wrapper is capable of accessing the memory as well as an execution cluster among multiple execution clusters, and includes logic for providing an input unit of block data to an available execution cluster among the multiple execution clusters and causing the execution cluster to apply the block data of the spatially aligned patches of each of the N cycles from the input unit to the neural network to produce an output patch of classification data for the subject cycle. The DRAM can be configured with sufficient allocated space for storage to support a five-cycle inference engine, where for each block on the flow cell, there are images for four cycles.Using pointer juggling, it is inferred that the required fifth buffer may exist in the host processor memory allocated to a particular processing thread (the number of blocks is typically much larger than the number of instantiated processor threads).
[0147] After running the neural network on the allocated patches, the cluster returns the output patches of the N-cycle subject cycle to the memory 140 via the wrapper and signals availability (1407). The wrapper tracks the sequence of patches to traverse the block data array, arranges the output patches, and signals the waiting processor thread in the host when the output block or other unit of the block data has been formed in the memory 140 (1408). Thus, the wrapper may include assembly logic for assembling the output patches from multiple execution clusters to provide the basecall classification data array of the subject cycle and storing the basecall classification data array in the memory. The formed output block is retrieved by the host or pushed by the wrapper to the host for further processing (1409). Optionally, as described above, the output block may include classification data in an unnormalized form, such as four logits for each pixel or corresponding to each pixel of the cluster. In this case, the host may perform a softmax or other type of normalization function on the block data to prepare for forming the basecall file and quality score (1410).
[0148] Figure 15 is for a neural network as Figure 4 and is a simplified flowchart of one embodiment of the logical cluster flow in a system as Figure 1 The logical cluster flow is initiated by the wrapper as discussed in Figure 16 and moves the input unit including N spatially aligned patches to the memory of the cluster (1501). The logical cluster allocates one or more available engines and routes the spatially aligned sub-patches of the N spatially aligned patches to the allocated engines (1502). The block data of the current input cycle of the spatially aligned patches includes N patches from the current cycle (cycle N+2) and N-1 previous cycles (cycles N-2, N-1, N, and N+1) of block data (1503). Each engine of the cluster can be configured with sufficient memory to hold a set of spatially aligned sub-patches of the data currently being processed in-place and sub-patches of the data to be processed when the current batch processing is completed using, for example, a ping-pong input buffer structure or other memory management techniques.
[0149] The allocated engine cycle passes through filters in the layers applied to the network, including applying a filter bank for the current network layer and returning activation data (1504) for the next layer. In this example, the neural network includes a first stage that includes a stack of N isolated spatial stacks, and this first stage feeds a second stage that includes a set of inverse hierarchical time layers. The last layer of the network generates an output of four features, each feature representing the classification of each of the four bases A / C / G / T (adenine (A), cytosine (C), guanine (G), and thymine (T)) when executed to classify DNA, or the four bases A / C / G / U (adenine (A), cytosine (C), guanine (G), and uracil (U)) when executed to classify bases (1505).
[0150] In this example, the neural network parameters include a block cluster mask for identifying the positions of gene clusters in the blocks (1506). In a specific implementation, the template generation step identifies the xy position coordinates of reliable clusters (e.g., the reliable clusters disclosed in U.S. Patent Application Publication US2012 / 0020537 by Garcia et al.). Intermediate data generated by the final spatial layer in the stack (dehydrating it) is reduced by an engine (in this exemplary flow, by mask logic in a configurable processor), and the engine uses the mask to remove pixels that do not correspond to the positions of the clusters in the blocks, thereby generating a smaller amount of activation data for subsequent layers (1507).
[0151] The engine returns the features of the output sub-patches of the subject cycle (mid-cycle) and signals the logical cluster about their availability (1508). The logical cluster continues to execute on the current patch until all the features of the current layer of the network for the patch are completed, and then traverses to the next allocated patch that assigns sub-patches to the engine and assembles the output sub-patches until all the allocated patches are completed (1509). Finally, the encapsulator transfers the output patch to the memory to assemble it into an output block (1510).
[0152] Figure 16 is a simplified flowchart of an embodiment of the logical cluster flow in a system such as Figure 10 for a neural network such as Figure 1 as shown in reference Figure 16The wrapper under discussion initiates a logical cluster stream and moves an input unit that includes N spatially aligned patches to the cluster's memory (1601). The logical cluster allocates one or more available engines and routes the spatial alignment sub-patches of the N spatially aligned patches to the allocated engines (1602). The spatially aligned patches include: a patch of the block data array for the current cycle (e.g., cycle N+2), which includes sensor data having F features (or channels); and N-1 patches of intermediate data calculated using the previous cycles (cycles N-2, N-1, N, and N+1) (1603). Each engine of the cluster can be configured with sufficient memory to hold a set of spatially aligned sub-patches of the data currently being processed in-place, as well as the sub-patches to be processed when the current sub-patch processing is completed using, for example, a ping-pong input buffer structure or other memory management techniques.
[0153] The allocated engines loop through the filters in the spatial layer and the network layer for the isolation stack applied to the current cycle (cycle N+2), including applying the filter bank for the current network layer for a given layer and returning the activation data for the next layer (1604). In this example, the neural network includes: a first stage that includes one isolation spatial stack for the current cycle; and a second stage that includes a set of inverse hierarchical time layers, as Figure 10 shown.
[0154] The engine returns the activation data of the current input cycle sub-patch to the logical cluster at the end of the first stage to be used as intermediate data in subsequent cycles (1606). This intermediate data can be stored back in the on-board memory at the location of the block data of the response cycle, thus replacing the sensor data, or stored in other locations. For example, the final layer of the spatial layer has a digital F or fewer features, such that the amount of memory for the block data of the cycle is not expanded due to the intermediate data (1607). Additionally, in this example, the neural network parameters include a block cluster mask that identifies the location of the gene clusters in the block (1608). The intermediate data from the final spatial layer is "dehydrated" in the engine using the mask or in another engine allocated by the cluster (1609). This reduces the size of the output sub-patch that is used as intermediate data and applied to the second stage.
[0155] After processing the isolation spatial stack for the current cycle, the engine loops through the time layers of the network using the activation data from the current input cycle (cycle N+2) and the intermediate data from the N-1 previous cycles (1610).
[0156] The engine returns the features of the output sub-patch of the subject cycle (cycle N) and signals the logical cluster about the availability (1611). The logical cluster accumulates the output sub-patches to form an output patch (1612). Finally, the encapsulator transmits the output patch to the memory for assembly into an output block (1613).
[0157] Figure 17 Illustrates a particular implementation of a specialized architecture of a neural network-based base caller (e.g., Figure 4 and Figure 10 ) for isolating the processing of data for different sequencing cycles. First, the motivation for using the specialized architecture is described.
[0158] The neural network-based base caller processes data from the current sequencing cycle, one or more previous sequencing cycles, and one or more subsequent sequencing cycles. The data from additional sequencing cycles provides sequence-specific context. The neural network-based base caller learns the sequence-specific context during training and performs base calling on this sequence-specific context. In addition, the data from the pre-sequencing cycle and the post-sequencing cycle provides a second-order contribution of the predetermined phase and the phase reference signal for the current sequencing cycle.
[0159] The images captured at different sequencing cycles and in different image channels are misaligned with respect to each other and have a residual registration error. Taking this misalignment into account, the specialized architecture includes a spatial convolutional layer that does not mix information between sequencing cycles and only mixes information within a sequencing cycle.
[0160] The spatial convolutional layer uses a so-called "isolation convolution" that achieves isolation by independently processing the data of each sequencing cycle in multiple sequencing cycles via a "dedicated non-shared" convolution sequence. The isolation convolution convolves only the data and the resulting feature map of a given sequencing cycle (i.e., within the cycle), and does not convolve the data and the resulting feature map of any other sequencing cycle.
[0161] For example, consider that the input data includes (i) current data of the current (time t) sequencing cycle for which base calling is to be performed, (ii) previous data of the previous (time t-1) sequencing cycle, and (iii) subsequent data of the subsequent (time t+1) sequencing cycle. Then, the specialized architecture initiates three separate data processing pipelines (or convolutional pipelines), namely, the current data processing pipeline, the previous data processing pipeline, and the subsequent data processing pipeline. The current data processing pipeline receives the current data of the current (time t) sequencing cycle as input and independently processes the current data through a plurality of spatial convolutional layers to produce a so-called "current spatial convolutional representation" as the output of the final spatial convolutional layer. The previous data processing pipeline receives the previous data of the previous (time t-1) sequencing cycle as input and independently processes the previous data through a plurality of spatial convolutional layers to produce a so-called "previous spatial convolutional representation" as the output of the final spatial convolutional layer. The subsequent data processing pipeline receives the subsequent data of the subsequent (time t+1) sequencing cycle as input and independently processes the subsequent data through a plurality of spatial convolutional layers to produce a so-called "subsequent spatial convolutional representation" as the output of the final spatial convolutional layer.
[0162] In some specific embodiments, the current processing pipeline, the previous processing pipeline, and the subsequent processing pipeline are executed synchronously.
[0163] In some specific embodiments, the spatial convolutional layer is part of a spatial convolutional network (or sub-network) within the specialized architecture.
[0164] The neural network-based base caller further includes a temporal convolutional layer that mixes information between sequencing cycles (i.e., inter-cycle). The temporal convolutional layer receives its input from the spatial convolutional network and operates on the spatial convolutional representations produced by the final spatial convolutional layers of the respective data processing pipelines.
[0165] The inter-cycle operability freedom of the temporal convolutional layer stems from the fact that the misalignment attribute is cleared from the spatial convolutional representations by a stack or cascade of isolation convolutions performed by the sequence of spatial convolutional layers, and this misalignment attribute exists in the image data fed as input to the spatial convolutional network.
[0166] The temporal convolutional layer uses a so-called "combinatorial convolution" that convolves the input channels in subsequent inputs group by group on a sliding window basis. In one specific embodiment, these subsequent inputs are the subsequent outputs produced by the previous spatial convolutional layer or the previous temporal convolutional layer.
[0167] In some specific implementations, the temporal convolutional layer is part of a temporal convolutional network (or sub-network) within a specialized architecture. The temporal convolutional network receives its input from a spatial convolutional network. In one specific implementation, the first temporal convolutional layer of the temporal convolutional network combines the spatial convolutional representations between sequencing cycles in groups. In another specific implementation, subsequent temporal convolutional layers of the temporal convolutional network combine the subsequent outputs of the previous temporal convolutional layers.
[0168] The output of the final temporal convolutional layer is fed to an output layer that produces an output. The output is used for base calling of one or more clusters at one or more sequencing cycles.
[0169] During forward propagation, the specialized architecture processes information from multiple inputs in two stages. In the first stage, isolation convolutions are used to prevent information mixing between the inputs. In the second stage, combination convolutions are used to mix the information between the inputs. The result from the second stage is used to make a single inference on the multiple inputs.
[0170] This is different from batch mode techniques where convolutional layers process multiple inputs in a batch simultaneously and make corresponding inferences for each input in the batch. In contrast, the specialized architecture maps the multiple inputs to the single inference. The single inference can include more than one prediction, such as classification scores (e.g., softmax or pre-softmax base-by-base classification scores or base-by-base regression scores) for each of the four bases (A, C, T, and G).
[0171] In one specific implementation, these inputs have a temporal order such that each input is generated at a different time step and has multiple input channels. For example, the multiple inputs can include the following three inputs: a current input generated by the current sequencing cycle at time step (t), a previous input generated by the previous sequencing cycle at time step (t - 1), and a subsequent input generated by the subsequent sequencing cycle at time step (t + 1). In another specific implementation, each input respectively originates from the current output, previous output, and subsequent output produced by one or more previous convolutional layers and includes k feature maps.
[0172] In one specific implementation, each input can include the following five input channels: a red image channel (red), a red distance channel (yellow), a green image channel (green), a green distance channel (purple), and a scaling channel (blue). In another specific implementation, each input can include k feature maps produced by a previous convolutional layer, and each feature map is treated as an input channel.
[0173] Figure 18Depicts a specific implementation of an isolation layer, where each isolation layer may include a convolution. The isolated convolution processes the multiple inputs by synchronously applying a convolution filter to each input once. With isolated convolution, the convolution filter combines input channels within the same input and does not combine input channels from different inputs. In one specific implementation, the same convolution filter is synchronously applied to each input. In another specific implementation, different convolution filters are synchronously applied to each input. In some specific implementations, each spatial convolution layer includes a set of k convolution filters, where each convolution filter is synchronously applied to each input.
[0174] Figure 19A Depicts a specific implementation of a combination layer, where each combination layer may include a convolution. Figure 19B Depicts another specific implementation of a combination layer, where each combination layer may include a convolution. The combined convolution mixes information between different inputs by grouping corresponding input channels of different inputs and applying a convolution filter to each group. The grouping of these corresponding input channels and the application of the convolution filter occur on a sliding window basis. In this context, the window spans two or more subsequent input channels, which represent, for example, the outputs of two subsequent sequencing cycles. Since this is a sliding window, most input channels are used in two or more windows.
[0175] In some specific implementations, the different inputs are derived from the output sequences produced by a previous spatial convolution layer or a previous temporal convolution layer. In this output sequence, these different inputs are arranged as subsequent outputs and are thus treated as subsequent inputs by a subsequent temporal convolution layer. Then, in this subsequent temporal convolution layer, these combined convolutions apply a convolution filter to the corresponding groups of input channels in these subsequent inputs.
[0176] In one specific implementation, these subsequent inputs have a temporal order such that the current input is generated by the current sequencing cycle at time step (t), the previous input is generated by the previous sequencing cycle at time step (t - 1), and the subsequent input is generated by the subsequent sequencing cycle at time step (t + 1). In another specific implementation, each subsequent input is respectively derived from the current output, previous output, and subsequent output produced by one or more previous convolution layers and includes k feature maps.
[0177] In one specific implementation, each input may include the following five input channels: a red image channel (red), a red distance channel (yellow), a green image channel (green), a green distance channel (purple), and a scaling channel (blue). In another specific implementation, each input may include k feature maps produced by a previous convolution layer, and each feature map is treated as an input channel.
[0178] The depth B of the convolutional filter depends on the number of subsequent inputs, and the corresponding input channels of these subsequent inputs are convolved by the convolutional filter group by group on the basis of a sliding window. In other words, the depth B is equal to the number of subsequent inputs in each sliding window and the group size.
[0179] In Figure 19A , the corresponding input channels from two subsequent inputs are combined in each sliding window, and thus B = 2. In Figure 19B , the corresponding input channels from three subsequent inputs are combined in each sliding window, and thus B = 3.
[0180] In a specific implementation, the sliding windows share the same convolutional filter. In another specific implementation, different convolutional filters are used for each sliding window. In some specific implementations, each temporal convolutional layer includes a set of k convolutional filters, where each convolutional filter is applied to subsequent inputs on the basis of a sliding window.
[0181] Figure 20 is a block diagram of a base calling system 2000 according to a specific implementation. The base calling system 2000 is operable to obtain any information or data related to at least one of a biological substance or a chemical substance. In some specific implementations, the base calling system 2000 is a workstation similar to a benchtop device or a desktop computer. For example, most (or all) of the systems and components for performing the required reactions may be located within a common housing 2016.
[0182] In a particular specific implementation, the base calling system 2000 is a nucleic acid sequencing system (or sequencer) configured for various applications, including but not limited to de novo sequencing, resequencing of whole genomes or target genomic regions, and metagenomics. The sequencer can also be used for DNA or RNA analysis. In some specific implementations, the base calling system 2000 can also be configured to generate reaction sites in a biosensor. For example, the base calling system 2000 can be configured to receive a sample and generate surface-attached clusters of clonally amplified nucleic acids derived from the sample. Each cluster can constitute or be part of a reaction site in the biosensor.
[0183] The exemplary base calling system 2000 may include a system socket or interface 2012 configured to interact with a biosensor 2002 to perform required reactions within the biosensor 2002. In the following description with respect to Figure 20 , the biosensor 2002 is loaded into the system socket 2012. However, it should be understood that a cartridge including the biosensor 2002 can be inserted into the system socket 2012, and in some states, the cartridge can be removed temporarily or permanently. As described above, among other things, the cartridge can also include fluid control components and fluid storage components.
[0184] In certain specific embodiments, the base detection system 2000 is configured to perform a large number of parallel reactions within the biosensor 2002. The biosensor 2002 includes one or more reaction sites where the desired reactions can occur. The reaction sites can be immobilized, for example, on the solid surface of the biosensor or on beads (or other movable substrates) located within the corresponding reaction chambers of the biosensor. The reaction sites can include, for example, clusters of clonally amplified nucleic acids. The biosensor 2002 can include a solid-state imaging device (e.g., a CCD or CMOS imager) and a flow cell mounted thereon. The flow cell can include one or more flow channels that receive solutions from the base detection system 2000 and direct the solutions to the reaction sites. Optionally, the biosensor 2002 can be configured to engage a thermal element for transferring thermal energy into or out of the flow channels.
[0185] The base detection system 2000 can include various components, assemblies, and systems (or subsystems) that interact with each other to perform a predetermined method or assay protocol for biological or chemical analysis. For example, the base detection system 2000 includes a system controller 2004 that can communicate with various components, assemblies, and subsystems of the base detection system 2000 and the biosensor 2002. For example, in addition to the system socket 2012, the base detection system 2000 can also include a fluid control system 2006 to control the flow of fluids throughout the fluid network of the base detection system 2000 and the biosensor 2002; a fluid storage system 2008 that is configured to store all fluids (e.g., gases or liquids) that can be used by the bioassay system; a temperature control system 2010 that can regulate the temperature of the fluids in the fluid network, the fluid storage system 2008, and / or the biosensor 2002; and an illumination system 2009 that is configured to illuminate the biosensor 2002. As described above, if a cartridge having the biosensor 2002 is loaded into the system socket 2012, the cartridge can also include fluid control components and fluid storage components.
[0186] As also shown, the base detection system 2000 may include a user interface 2014 for interacting with a user. For example, the user interface 2014 may include a display 2013 for displaying or requesting information from the user and a user input device 2015 for receiving user input. In some embodiments, the display 2013 and the user input device 2015 are the same device. For example, the user interface 2014 may include a touch-sensitive display configured to detect the presence of an individual touch and also identify the location of the touch on the display. However, other user input devices 2015 may be used, such as a mouse, a touchpad, a keyboard, a keypad, a handheld scanner, a voice recognition system, a motion recognition system, etc. As will be discussed in more detail below, the base detection system 2000 may communicate with various components including a biosensor 2002 (e.g., in the form of a cartridge) to perform the required reactions. The base detection system 2000 may also be configured to analyze data obtained from the biosensor to provide the required information to the user.
[0187] The system controller 2004 may include any processor-based or microprocessor-based system, including those using a microcontroller, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), logic circuits, and any other circuit or processor capable of performing the functions described herein. The above examples are merely exemplary and are not intended to limit in any way the definition and / or meaning of the term system controller. In an exemplary embodiment, the system controller 2004 executes an instruction set stored in one or more storage elements, memories, or modules to perform at least one of obtaining detection data and analyzing the detection data. The detection data may include a plurality of pixel signal sequences such that pixel signal sequences from each of millions of sensors (or pixels) may be detected over a number of base detection cycles. The storage element may be in the form of an information source or a physical memory element within the base detection system 2000.
[0188] The instruction set may include various commands that direct the base detection system 2000 or the biosensor 2002 to perform specific operations such as the methods and processes of the various embodiments described herein. The instruction set may be in the form of a software program that may form part of one or more tangible non-transitory computer-readable media. As used herein, the terms "software" and "firmware" are interchangeable and include any computer program stored in memory for execution by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely exemplary and thus do not limit the types of memory that may be used to store a computer program.
[0189] The software can be in various forms, such as system software or application software. In addition, the software can be in the form of a collection of independent programs, or in the form of program modules or parts of program modules within a larger program. The software can also include modular programming in the form of object-oriented programming. After the detection data is obtained, the detection data can be automatically processed by the base detection system 2000, processed in response to user input, or processed in response to a request made by another processing machine (e.g., a remote request via a communication link). In the illustrated specific implementation, the system controller 2004 includes an analysis module 2138. In other specific implementations, the system controller 2004 does not include the analysis module 2138, but is capable of accessing the analysis module 2138 (e.g., the analysis module 2138 can be hosted separately on the cloud).
[0190] The system controller 2004 can be connected via a communication link to the biosensor 2002 and other components of the base detection system 2000. The system controller 2004 can also be communicatively connected to an off-site system or server. The communication link can be hardwired, wired, or wireless. The system controller 2004 can receive user input or commands from the user interface 2014 and the user input device 2015.
[0191] The fluid control system 2006 includes a fluid network and is configured to direct and regulate the flow of one or more fluids through the fluid network. The fluid network can be in fluid communication with the biosensor 2002 and the fluid storage system 2008. For example, a selected fluid can be aspirated from the fluid storage system 2008 and directed to the biosensor 2002 in a controlled manner, or the fluid can be aspirated from the biosensor 2002 and directed towards, for example, a waste reservoir in the fluid storage system 2008. Although not shown, the fluid control system 2006 can include a flow sensor for detecting the flow rate or pressure of the fluid within the fluid network. The sensor can communicate with the system controller 2004.
[0192] The temperature control system 2010 is configured to regulate the temperature of the fluid at different regions of the fluid network, the fluid storage system 2008, and / or the biosensor 2002. For example, the temperature control system 2010 can include a thermal cycler that docks with the biosensor 2002 and controls the temperature of the fluid flowing along the reaction sites in the biosensor 2002. The temperature control system 2010 can also regulate the temperature of the solid elements or components of the base detection system 2000 or the biosensor 2002. Although not shown, the temperature control system 2010 can include sensors for detecting the temperature of the fluid or other components. The sensors can communicate with the system controller 2004.
[0193] The fluid storage system 2008 is in fluid communication with the biosensor 2002 and can store various reaction components or reactants for performing desired reactions therein. The fluid storage system 2008 can also store fluids for washing or cleaning the fluid network and the biosensor 2002 and for diluting reactants. For example, the fluid storage system 2008 can include various reservoirs to store samples, reagents, enzymes, other biomolecules, buffer solutions, aqueous solutions, and non-polar solutions, etc. In addition, the fluid storage system 2008 can also include a waste reservoir for receiving waste from the biosensor 2002. In a specific implementation including a cartridge, the cartridge can include one or more of a fluid storage system, a fluid control system, or a temperature control system. Thus, one or more components related to those systems described herein can be housed within the cartridge housing. For example, the cartridge can have various reservoirs to store samples, reagents, enzymes, other biomolecules, buffer solutions, aqueous solutions, non-polar solutions, waste, etc. Thus, one or more of the fluid storage system, the fluid control system, or the temperature control system can be removably engaged with the bioassay system via the cartridge or other biosensors.
[0194] The illumination system 2009 can include a light source (e.g., one or more LEDs) and a plurality of optical components for illuminating the biosensor. Examples of the light source can include lasers, arc lamps, LEDs, or laser diodes. The optical components can be, for example, reflectors, dichroic mirrors, beam splitters, collimators, lenses, filters, wedge prisms, prisms, mirrors, detectors, etc. In a specific implementation using the illumination system, the illumination system 2009 can be configured to direct excitation light to the reaction site. As an example, a fluorophore can be excited by light of a green wavelength, so the wavelength of the excitation light can be approximately 532 nm. In a specific implementation, the illumination system 2009 is configured to produce illumination parallel to the surface normal of the surface of the biosensor 2002. In another specific implementation, the illumination system 2009 is configured to produce illumination at an angle offset with respect to the surface normal of the surface of the biosensor 2002. In yet another specific implementation, the illumination system 2009 is configured to produce illumination having multiple angles, including some parallel illumination and some offset illumination.
[0195] The system socket or interface 2012 is configured to engage the biosensor 2002 in at least one of a mechanical, electrical, and fluidic manner. The system socket 2012 can hold the biosensor 2002 in a desired orientation to facilitate fluid flow through the biosensor 2002. The system socket 2012 can also include electrical contacts configured to engage the biosensor 2002 such that the base detection system 2000 can communicate with and / or provide power to the biosensor 2002. Additionally, the system socket 2012 can include fluid ports (e.g., nozzles) configured to engage the biosensor 2002. In some embodiments, the biosensor 2002 is removably coupled to the system socket 2012 in a mechanical, electrical, and fluidic manner.
[0196] In addition, the base detection system 2000 can communicate remotely with other systems or networks or with other biometric systems 2000. Detection data obtained by the biometric system 2000 can be stored in a remote database.
[0197] Figure 21 is a block diagram of the system controller 2004 that can be used in a Figure 20 system. In one embodiment, the system controller 2004 includes one or more processors or modules that can communicate with each other. Each of the processors or modules can include algorithms (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or sub-algorithms for performing specific processes. The system controller 2004 is conceptually illustrated as a collection of modules but can be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 2004 can be implemented using an off-the-shelf PC with a single processor or multiple processors, where the functional operations are distributed among the processors. As a further option, the modules described below can be implemented using a hybrid configuration, where certain modular functions are performed using dedicated hardware and the remaining modular functions are performed using an off-the-shelf PC, etc. The modules can also be implemented as software modules within a processing unit.
[0198] During operation, the communication port 2120 can transmit information (e.g., commands) to and / or receive information (e.g., data) from the biosensor 2002 ( Figure 20 ) and / or the subsystems 2006, 2008, 2010 ( Figure 20 ). In an embodiment, the communication port 2120 can output multiple pixel signal sequences. The communication link 2120 can receive from the user interface 2014 ( Figure 20)Receives user input and transfers data or information to the user interface 2014. Data from the biosensor 2002 or subsystems 2006, 2008, 2010 can be processed in real time by the system controller 2004 during a biometric session. Additionally or alternatively, the data can be temporarily stored in the system memory during the biometric session and processed at a slower rate than real-time or offline operation.
[0199] As Figure 21 shown, the system controller 2004 can include a plurality of modules 2131 - 2139 that communicate with the main control module 2130. The main control module 2130 can communicate with the user interface 2014( Figure 20 ). Although the modules 2131 - 2139 are shown as communicating directly with the main control module 2130, the modules 2131 - 2139 can also communicate directly with each other, directly with the user interface 2014 and the biosensor 2002. Additionally, the modules 2131 - 2139 can communicate with the main control module 2130 through other modules.
[0200] The plurality of modules 2131 - 2139 include system modules 2131 - 2133, 2139 that communicate with subsystems 2006, 2008, 2010, and 2009 respectively. The fluid control module 2131 can communicate with the fluid control system 2006 to control the valves and flow sensors of the fluid network, thereby controlling the flow of one or more fluids through the fluid network. The fluid storage module 2132 can notify the user when the fluid level is low or when the waste reservoir is at or near capacity. The fluid storage module 2132 can also communicate with the temperature control module 2133 so that the fluid can be stored at the desired temperature. The illumination module 2139 can communicate with the illumination system 2009 to illuminate the reaction site at a specified time during the protocol, such as after the desired reaction (e.g., binding event) has occurred. In some embodiments, the illumination module 2139 can communicate with the illumination system 2009 to illuminate the reaction site at a specified angle.
[0201] The plurality of modules 2131 - 2139 can also include a device module 2134 that communicates with the biosensor 2002 and an identification module 2135 that determines identification information related to the biosensor 2002. The device module 2134 can communicate with, for example, the system socket 2012 to confirm that the biosensor has established an electrical connection and a fluid connection with the base detection system 2000. The identification module 2135 can receive a signal that identifies the biosensor 2002. The identification module 2135 can use the identity of the biosensor 2002 to provide other information to the user. For example, the identification module 2135 can determine and then display the lot number, manufacturing date, or protocol recommended to run with the biosensor 2002.
[0202] The plurality of modules 2131 - 2139 further includes an analysis module 2138 (also referred to as a signal processing module or signal processor) that receives and analyzes signal data (e.g., image data) from the biosensor 2002. The analysis module 2138 includes a memory (e.g., RAM or flash memory) for storing detection data. The detection data may include a plurality of pixel signal sequences such that pixel signal sequences from each of millions of sensors (or pixels) can be detected over many base calling cycles. The signal data may be stored for subsequent analysis or may be transmitted to the user interface 2014 to display the desired information to the user. In some embodiments, the signal data may be processed by a solid-state imager (e.g., a CMOS image sensor) before being received by the analysis module 2138.
[0203] The analysis module 2138 is configured to obtain image data from the light detector at each of a plurality of sequencing cycles. The image data is derived from the emission signals detected by the light detector and is processed by a neural network (e.g., a neural network-based template generator 2148, a neural network-based base caller 2158 (e.g., Figure 4 and Figure 10 ), and / or a neural network-based quality score generator 2168) for each of the plurality of sequencing cycles, and base calling is generated for at least some of the analytes at each of the plurality of sequencing cycles.
[0204] The protocol modules 2136 and 2137 communicate with the main control module 2130 to control the operation of subsystems 2006, 2008, and 2010 when performing a pre-determined assay protocol. The protocol modules 2136 and 2137 may include instruction sets for instructing the base calling system 2000 to perform specific operations according to a pre-determined protocol. As shown, the protocol module may be a sequencing by synthesis (SBS) module 2136, which is configured to issue various commands for performing a sequencing by synthesis process. In SBS, the extension of a nucleic acid primer along a nucleic acid template is monitored to determine the sequence of nucleotides in the template. The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase) or ligation (e.g., catalyzed by a ligase). In a particular polymerase-based SBS implementation, fluorescently labeled nucleotides are added to the primer in a template-dependent manner (thereby extending the primer) such that detection of the order and type of nucleotides added to the primer can be used to determine the sequence of the template. For example, to initiate the first SBS cycle, commands may be issued to deliver one or more labeled nucleotides, DNA polymerase, etc. to / through the flow cell containing the nucleic acid template array. The nucleic acid templates may be located at corresponding reaction sites. Those reaction sites where primer extension results in incorporation of labeled nucleotides can be detected by an imaging event. During the imaging event, the illumination system 2009 may provide excitation light to the reaction sites. Optionally, the nucleotides may also include a reversible termination property that terminates further primer extension once the nucleotide is added to the primer. For example, nucleotide analogs having reversible terminator moieties may be added to the primer such that subsequent extension does not occur until a deblocking agent is delivered to remove the moiety. Thus, for implementations using reversible termination, commands may be issued to deliver the deblocking agent to the flow cell (before or after detection). One or more commands may be issued to effect washing between the respective delivery steps. The cycle may then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary sequencing techniques are described in, for example, Bentley et al., Nature 456:53-59 (2008); WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,329,492; US 7,211,414; US 7,315,019; US 7,405,281 and US 2008 / 014708082, each of which is incorporated herein by reference.
[0205] For the nucleotide delivery step of the SBS cycle, a single type of nucleotide can be delivered at a time, or multiple different nucleotide types can be delivered (e.g., A, C, T, and G together). For nucleotide delivery configurations where only a single type of nucleotide is present at a time, the different nucleotides do not need to have different labels because they can be distinguished based on the time intervals inherent in the individualized delivery. Thus, a sequencing method or apparatus can use single-color detection. For example, the excitation source only needs to provide excitation at a single wavelength or within a single wavelength range. For nucleotide delivery configurations where the delivery results in multiple different nucleotides being present in the flow cell simultaneously, the sites of incorporation of different nucleotide types can be distinguished based on different fluorescent labels attached to the corresponding nucleotide types in the mixture. For example, four different nucleotides can be used, each with one of four different fluorophores. In one specific implementation, excitation in four different regions of the spectrum can be used to distinguish the four different fluorophores. For example, four different excitation radiation sources can be used. Alternatively, fewer than four different excitation sources can be used, but optical filtering of the excitation radiation from a single source can be used to produce different ranges of excitation radiation at the flow cell.
[0206] In some specific implementations, fewer than four different colors can be detected in a mixture of four different nucleotides. For example, nucleotide pairs can be detected at the same wavelength, but distinguished based on the intensity difference of one member of the pair relative to the other member, or based on a change in one member of the pair that results in a distinct signal appearance or disappearance compared to the signal of the other member of the pair being detected (e.g., through chemical modification, photochemical modification, or physical modification). Exemplary devices and methods for distinguishing four different nucleotides using detection of fewer than four colors are described, for example, in U.S. Patent Application Serial Numbers 61 / 538,294 and 61 / 619,878, which are incorporated herein by reference in their entirety. U.S. Application 13 / 624,200, filed September 21, 2012, is also incorporated herein by reference in its entirety.
[0207] The plurality of scenario modules can also include a sample preparation (or generation) module 2137 that is configured to issue commands to the fluid control system 2006 and the temperature control system 2010 for amplifying the product within the biosensor 2002. For example, the biosensor 2002 can be coupled to the base calling system 2000. The amplification module 2137 can issue instructions to the fluid control system 2006 to deliver the necessary amplification components into the reaction chamber within the biosensor 2002. In other specific implementations, the reaction site may already contain some components for amplification, such as template DNA and / or primers. After delivering the amplification components into the reaction chamber, the amplification module 2137 can instruct the temperature control system 2010 to cycle through different temperature stages according to a known amplification protocol. In some specific implementations, amplification and / or nucleotide incorporation occur isothermally.
[0208] The SBS module 2136 can issue commands to perform bridge PCR, where clusters of cloned amplicons are formed on local regions within the channels of the flow cell. After generating amplicons by bridge PCR, the amplicons can be "linearized" to prepare single-stranded template DNA or sstDNA, and sequencing primers can be hybridized to the universal sequences flanking the regions of interest. For example, reversible terminator-based sequencing-by-synthesis methods can be used as described above or below.
[0209] Each base detection or sequencing cycle can extend the sstDNA by a single base, which can be accomplished, for example, by using a modified DNA polymerase and a mixture of four types of nucleotides. The different types of nucleotides can have unique fluorescent labels, and each nucleotide can also have a reversible terminator that allows only single-base incorporation in each cycle. After adding a single base to the sstDNA, excitation light can be incident on the reaction site and the fluorescence emission can be detected. After detection, the fluorescent label and terminator can be chemically cleaved from the sstDNA. Next, another similar base detection or sequencing cycle can occur. In such a sequencing scheme, the SBS module 2136 can direct the fluid control system 2006 to flow reagent and enzyme solutions over the biosensor 2002. Exemplary reversible terminator-based SBS methods that can be used with the devices and methods described herein are described in U.S. Patent Application Publication 2007 / 0166705A1, U.S. Patent Application Publication 2006 / 0188901A1, U.S. Patent 7,057,026, U.S. Patent Application Publication 2006 / 0240439A1, U.S. Patent Application Publication 2006 / 02814714709A1, PCT Publication WO 05 / 065814, U.S. Patent Application Publication 2005 / 014700900A1, PCT Publication WO06 / 064199, and PCT Publication WO 07 / 01470251, each of which is incorporated herein by reference in its entirety. Exemplary reagents for reversible terminator-based SBS are described in: US 7,541,444; US 7,057,026; US7,414,14716; US 7,427,673; US 7,566,537; US 7,592,435, and WO 07 / 14835368, each of which is incorporated herein by reference in its entirety.
[0210] In some specific embodiments, the amplification module and the SBS module can be operated in a single assay protocol, where, for example, the template nucleic acid is amplified and then sequenced in the same cartridge.
[0211] The base detection system 2000 may also allow a user to reconfigure the assay protocol. For example, the base detection system 2000 may provide options for modifying the determined protocol to the user via the user interface 2014. For example, if it is determined that the biosensor 2002 will be used for amplification, the base detection system 2000 may request the temperature of the annealing cycle. Additionally, if the user has provided user input that is not normally acceptable for the selected assay protocol, the base detection system 2000 may issue a warning to the user.
[0212] In a particular implementation, the biosensor 2002 includes millions of sensors (or pixels), and each sensor (or pixel) generates multiple pixel signal sequences within subsequent base detection cycles. The analysis module 2138 detects the multiple pixel signal sequences and attributes them to corresponding sensors (or pixels) based on the row-by-row and / or column-by-column positions of the sensors on the sensor array.
[0213] Each sensor in the sensor array may generate sensor data for a block of the flow cell, where the block is located in the region of the flow cell where clusters of genetic material are set during base detection operations. The sensor data may include image data in a pixel array. For a given cycle, the sensor data may include more than one image, resulting in multi-feature per pixel as block data.
[0214] As used herein, "logic" (e.g., data flow logic) may be implemented in the form of a computer product that includes a non-transitory computer-readable storage medium having computer-usable program code for performing the method steps described herein. "Logic" may be implemented in the form of an apparatus that includes a memory and at least one processor coupled to the memory and operable to execute exemplary method steps. "Logic" may be implemented in the form of an apparatus for performing one or more of the method steps described herein; the apparatus may include (i) a hardware module, (ii) a software module executed on one or more hardware processors, or (iii) a combination of a hardware module and a software module; any of (i)-(iii) implements the specific techniques set forth herein, and the software module is stored in a computer-readable storage medium (or multiple such media). In a particular implementation, the logic implements data processing functions. The logic may be a general single-core or multi-core processor with a computer program, a digital signal processor with a computer program, a configurable logic (such as an FPGA) with a configuration file, a dedicated circuit (such as a state machine), or any combination thereof. Additionally, a computer program product may embody the computer program and the configuration file portion of the logic.
[0215] Figure 22FIG. 2200 is a simplified block diagram of a computer system 2200 that can be used to implement the disclosed technology. The computer system 2200 includes at least one central processing unit (CPU) 2272 that communicates with a plurality of peripheral devices via a bus subsystem 2255. These peripheral devices can include a storage subsystem 2210, which includes, for example, memory devices and a file storage subsystem 2236, a user interface input device 2238, a user interface output device 2276, and a network interface subsystem 2274. The input and output devices allow a user to interact with the computer system 2200. The network interface subsystem 2274 provides an interface to an external network, including providing an interface to corresponding interface devices in other computer systems.
[0216] The user interface input device 2238 can include: a keyboard; pointing devices such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touchscreen incorporated into a display; audio input devices such as a speech recognition system and a microphone; and other types of input devices. In general, the term "input device" is intended to include all possible types of devices and ways of inputting information into the computer system 2200.
[0217] The user interface output device 2276 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a light-emitting diode (LED) display, a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for generating a visible image. The display subsystem can also provide a non-visual display, such as an audio output device. In general, the term "output device" is intended to include all possible types of devices and ways of outputting information from the computer system 2200 to a user or to another machine or computer system.
[0218] The storage subsystem 2210 stores programming and data structures that provide the functionality and methods of some or all of the modules described herein. These software modules are typically executed by a deep learning processor 2278.
[0219] In one particular implementation, a neural network is implemented using deep learning processors 2278, which can be configurable and reconfigurable processors, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and / or coarse grained reconfigurable architectures (CGRAs) and graphics processing units (GPUs) or other configured devices. The deep learning processors 2278 can be provided by a deep learning cloud platform such as Google Cloud Platform TM 、Xilinx TM and Cirrascale TM)Hosting. Examples of deep learning processors 14978 include Google's Tensor Processing Unit (TPU) TM , rack solutions such as the GX4 Rackmount Series TM , GX149 Rackmount Series TM ), NVIDIA DGX-1 TM , Microsoft's Stratix V FPGA TM , Graphcore's Intelligent Processor Unit (IPU) TM , Qualcomm's Zeroth Platform with Snapdragon processors TM TM , NVIDIA's Volta TM , NVIDIA's DRIVE PX TM , NVIDIA's JETSON TX1 / TX2 MODULE TM , Intel's Nirvana TM , Movidius VPU TM , Fujitsu DPI TM , ARM's DynamicIQ TM , IBM TrueNorth TM and so on.
[0220] The memory subsystem 2222 used in the storage subsystem 2210 may include multiple memories, including a main random access memory (RAM) 2232 for storing instructions and data during program execution and a read-only memory (ROM) 2234 in which fixed instructions are stored. The file storage subsystem 2236 may provide persistent storage for program files and data files and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical disk drive or a removable media tape drive. Modules implementing the functions of certain specific embodiments may be stored in the storage subsystem 2210 by the file storage subsystem 2236, or in other machines accessible to the processor.
[0221] The bus subsystem 2255 provides the mechanism for enabling the various components and subsystems of the computer system 2200 to communicate with each other as expected. Although the bus subsystem 2255 is schematically shown as a single bus, alternative specific embodiments of the bus subsystem may use multiple buses.
[0222] The computer system 2200 itself can have different types, including personal computers, portable computers, workstations, computer terminals, network computers, televisions, mainframes, server farms, a group of widely distributed and loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the Figure 22 description of the computer system 2200 depicted in Figure 22 is only intended as a specific example for illustrating a preferred specific implementation of the present invention. Many other configurations of the computer system 2200 are possible, having more or fewer components than the
[0223] Clause
[0224] 1. A system for analyzing the output of a base detection sensor, the system comprising:
[0225] A host processor;
[0226] A memory accessible by the host processor, the memory storing block data, the block data including an array of sensor data from a block of a sensing cycle of a base detection operation; and
[0227] A neural network processor accessible by the memory, the neural network processor comprising:
[0228] A plurality of execution clusters, the execution clusters in the plurality of execution clusters being configured to execute a neural network; and
[0229] Data flow logic capable of accessing the memory and the execution clusters in the plurality of execution clusters to provide input units of block data to available execution clusters in the plurality of execution clusters, the input units including digital N spatially aligned patches from an array of block data of a corresponding sensing cycle including a subject sensing cycle, and causing the execution clusters to apply the N spatially aligned patches to the neural network to produce an output patch of classification data of the spatially aligned patches of the subject sensing cycle, where N is greater than 1.
[0230] 2. The system according to clause 1, the system comprising: assembly logic for assembling these output patches from the plurality of execution clusters to provide base detection classification data of the subject cycle, and storing the base detection classification data in a memory.
[0231] 3. The system according to clause 1, wherein an execution cluster among the plurality of execution clusters includes: a set of computing engines having a plurality of members, configured to perform convolution on input data of multiple layers of the neural network using trained parameters, wherein the input data of the first layer is from these input units, and the data of subsequent layers is from activation data output from the previous layer.
[0232] 4. The system according to clause 3, the system includes: a memory that stores multiple versions of the trained parameters of the neural network, and wherein a cluster among the multiple clusters in the neural network processor is configured to include a kernel memory to store these trained parameters, and the data flow logic is configured to provide an instance of the trained parameters to the kernel memory of an execution cluster among the plurality of execution clusters for executing the neural network.
[0233] 5. The system according to clause 4, wherein instances of these trained parameters are applied according to the number of cycles in these sensing cycles of the base detection operation.
[0234] 6. The system according to clause 1, wherein an execution cluster among the plurality of execution clusters includes a set of computing engines having a plurality of members, the set of computing engines is configured to apply multiple configurable filters for a corresponding layer of the neural network to sub-patches from the input units and sub-patches from activation data output from layers of the neural network.
[0235] 7. The system according to clause 1, wherein the neural network performs an isolated stack of spatial layers for each spatially aligned patch of input units, and provides data from the N spatially aligned patches output from these isolated stacks to one or more combination layers.
[0236] 8. The system according to clause 7, wherein an execution cluster among the plurality of execution clusters is configured to use the sensor data of the input N spatially aligned patches to perform N isolated stacks of spatial layers to generate an output patch of classification data of the spatially aligned patches of the subject sensing cycle.
[0237] 9. The system according to clause 7, wherein an execution cluster among the plurality of execution clusters is configured to perform an isolation stack of a spatial layer using sensor data of a current spatially aligned patch, feed intermediate data of the spatially aligned patch back to the memory to be used as block data in an input unit of other loops, and provide the intermediate data from a previous loop and an output of the isolation stack of the current loop to the one or more combination layers to generate the output patch of the subject sensing loop. In a specific implementation, the memory is an on-chip or off-chip DRAM or an on-chip memory such as SRAM or BRAM, and a processing unit of the processor executes the neural network. In such specific implementations, the intermediate data from a previous loop is sent from the memory to these on-chip processing elements (e.g., via a DMA engine) to be processed by, for example, combination layers, which are executed by the chip as part of the execution of the neural network on the chip / processor.
[0238] 10. The system according to clause 1, wherein the digital N is an integer equal to 5 or greater.
[0239] 11. The system according to clause 1, wherein the logic circuit is configured to load an input unit in a sequence of these block data arrays traversing multiple sensing loops, and the loading includes writing the next input unit in the sequence to the neural network processor for the execution cluster during execution of the neural network by the execution cluster for a previous input unit.
[0240] 12. The system according to clause 1, the system including: logic in the host processor for performing an activation function on these output patches.
[0241] 13. The system according to clause 1, the system including: logic in the host processor for performing a softmax function on these output patches.
[0242] 14. The system according to clause 1, wherein these block data arrays include M features, where M is greater than one.
[0243] 15. The system according to clause 7, the system including: a block cluster mask for these blocks for a base calling operation, the block cluster mask being stored in the memory, and these execution clusters being configured to use the block cluster mask to remove data from intermediate data from at least one of these layers.
[0244] 16. A computer-implemented method for analyzing base calling sensor outputs, the method including:
[0245] Storing block data in a memory, the block data including a sensor data array of a block from a sensing loop of a base calling operation; and
[0246] Performing a neural network on the block data using multiple execution clusters, the performing including:
[0247] Providing input units of the block data to available execution clusters among the multiple execution clusters, the input units including a digital N spatial alignment patches of a block data array from a respective sensing cycle including a subject sensing cycle, and causing the execution clusters to apply the N spatial alignment patches to the neural network to produce output patches of classification data for the spatial alignment patches of the subject sensing cycle, where N is greater than 1.
[0248] 17. The computer-implemented method according to clause 16, the method including: assembling these output patches from the multiple execution clusters to provide base calling classification data for the subject cycle, and storing the base calling classification data in a memory.
[0249] 18. The computer-implemented method according to clause 16, where an execution cluster among the multiple execution clusters includes: a set of computing engines having a plurality of members configured to perform convolution on input data of multiple layers of the neural network using trained parameters, where the input data of the first layer is from the input unit and the data of subsequent layers is from activation data output from a previous layer.
[0250] 19. The computer-implemented method according to clause 18, the method including: storing multiple versions of the trained parameters of the neural network, and where a cluster among the multiple execution clusters in a neural network processor is configured to include a kernel memory to store the trained parameters; and providing an instance of the trained parameters to the kernel memory of an execution cluster among the multiple execution clusters for performing the neural network.
[0251] 20. The computer-implemented method according to clause 19, the method including: applying an instance of the trained parameters according to the number of cycles among the sensing cycles of the base calling operation.
[0252] 21. The computer-implemented method according to clause 16, where an execution cluster among the multiple execution clusters includes a set of computing engines having a plurality of members, the set of computing engines being configured to apply multiple configurable filters for a corresponding layer of the neural network to sub-patches from the input unit and sub-patches of activation data output from a layer of the neural network.
[0253] 22. The computer-implemented method according to clause 16, where the neural network includes an isolated stack of spatial layers for each spatial alignment patch of the input unit, and providing data from the N spatial alignment patches output from the isolated stacks to one or more combination layers.
[0254] 23. The computer-implemented method according to clause 22, wherein an execution cluster among the plurality of execution clusters is configured to use the sensor data of the N spatially aligned patches input thereto to perform N isolation stacks of a spatial layer to generate an output patch of classification data of the spatially aligned patches of the subject sensing cycle.
[0255] 24. The computer-implemented method according to clause 22, wherein an execution cluster among the plurality of execution clusters is configured to use the sensor data of a current spatially aligned patch to perform an isolation stack of a spatial layer, feedback intermediate data of the spatially aligned patch to the memory to be used as block data in an input unit of other cycles, and provide the intermediate data from a previous cycle and the output of the isolation stack of the current cycle to the one or more combination layers to generate the output patch of the subject sensing cycle.
[0256] 25. The computer-implemented method according to clause 16, wherein the number N is an integer equal to 5 or greater.
[0257] 26. The computer-implemented method according to clause 16, the method comprising: loading an input unit in a sequence of these block data arrays traversing a plurality of sensing cycles, the loading including writing the next input unit in the sequence to the execution cluster during execution of the neural network by the execution cluster for a previous input unit.
[0258] 27. The computer-implemented method according to clause 16, the method comprising: performing an activation function on these output patches.
[0259] 28. The computer-implemented method according to clause 16, the method comprising: performing a softmax function on these output patches.
[0260] 29. The computer-implemented method according to clause 16, wherein these block data arrays include M features, where M is greater than one.
[0261] 30. The computer-implemented method according to clause 16, the method comprising: storing a block cluster mask of these blocks for a base calling operation stored in the memory, and using the block cluster mask to remove data of the intermediate data from at least one of these layers.
[0262] 31. A system for analyzing base calling sensor output, the system comprising:
[0263] a memory accessible by the runtime program, the memory storing block data, the block data including sensor data of blocks of a sensing cycle from a base calling operation;
[0264] A neural network processor that can access the memory, the neural network processor being configured to execute the operation of the neural network using the trained parameters to generate classification data for the sensing cycles, the operation of the neural network operating on a sequence of N arrays of block data for the respective sensing cycles from among N sensing cycles including the subject cycle to generate the classification data for the subject cycle; and
[0265] Data flow logic that uses an input unit to move the block data and the trained parameters from the memory to the neural network processor for the operation of the neural network, the input unit including data of spatially aligned patches from the N arrays for the respective sensing cycles from among the N sensing cycles.
[0266] 32. The system according to clause 31, wherein the block data of the sensing cycles in the memory includes one or both of the sensor data of the sensing cycle and the intermediate data fed back from the neural network for the sensing cycle.
[0267] 33. The system according to clause 31, wherein the memory stores a block cluster mask that identifies elements in the sensor data arrays representing the positions of the flow cell clusters in the block, and the neural network processor includes mask logic for applying the block cluster mask to the intermediate data in the neural network.
[0268] 34. The system according to clause 31, the data flow logic including elements of the neural network processor configured using configuration data.
[0269] 35. The system according to clause 31, the system including: a host processing system that includes a runtime program, and the data flow logic including logic for coordinating these operations with the runtime program on the host.
[0270] 36. The system according to clause 31, the system including: a host processing system that includes a runtime program, the logic of the runtime program providing the block data from these sensing cycles and the trained parameters of the neural network to the memory.
[0271] 37. The system according to clause 31, wherein the block data of the sensing cycles in the memory includes one or both of the sensor data of the sensing cycle and the intermediate data fed back from the neural network for the sensing cycle.
[0272] 38. The system according to clause 31, wherein the sensor data of the sensing cycle includes data representing signals for detecting one base for each flow cell cluster.
[0273] 39. The system according to clause 31, wherein the neural network processor is configured to form a plurality of execution clusters, and these execution logic clusters in the plurality of execution clusters run the neural network on spatially aligned patches of the block data; and
[0274] wherein the data flow logic is capable of accessing the memory and the execution clusters in the plurality of execution clusters to provide input units of the block data to the available execution clusters in the plurality of execution clusters and cause these execution clusters to apply the block data of the N spatially aligned patches of these input units to the neural network to generate output patches of classification data for the subject cycle.
[0275] 40. The system according to clause 39, the system comprising: assembly logic for assembling these output patches from the plurality of execution clusters to provide base calling classification data for the subject cycle and storing the base calling classification data in a memory.
[0276] 41. The system according to clause 39, wherein the execution clusters in the plurality of execution clusters include: a set of computing engines having a plurality of members configured to perform convolution on input data of multiple layers of the neural network using trained parameters, wherein the input data of the first layer is from these input units and the data of subsequent layers is from activation data output from the previous layer.
[0277] 42. The system according to clause 41, the system comprising: a memory storing multiple versions of the trained parameters of the neural network, and wherein the clusters in the plurality of clusters in the neural network processor are configured to include kernel memories to store these trained parameters, and the data flow logic is configured to provide instances of the trained parameters to the kernel memories of the execution clusters in the plurality of execution clusters for executing the neural network.
[0278] 43. The system according to clause 42, wherein the instances of these trained parameters are applied according to the number of cycles in these sensing cycles of the base calling operation.
[0279] 44. The system according to clause 39, wherein the execution clusters in the plurality of execution clusters include a set of computing engines having a plurality of members, and the set of computing engines is configured to apply a plurality of configurable filters for corresponding layers of the neural network to sub-patches from the input units and sub-patches from activation data output from the layers of the neural network.
[0280] 45. The system according to clause 39, wherein an execution cluster among the multiple execution clusters is configured to perform an isolation stack of a spatial layer using sensor data of a current spatially aligned patch, feed intermediate data of the spatially aligned patch back to the memory to be used as block data in other input units, and provide the intermediate data from a previous cycle and the output of the isolation stack of the current cycle to the one or more combination layers to generate the output patch of the subject sensing cycle.
[0281] 46. The system according to clause 31, wherein the number N is an integer equal to 5 or greater.
[0282] 47. The system according to clause 31, wherein the logic circuits are configured to load input units in a sequence of block data arrays traversing multiple sensing cycles, the loading including writing the next input unit in the sequence to the neural network processor for the execution cluster during execution of the neural network by the execution cluster for a previous input unit.
[0283] 48. The system according to clause 31, the system comprising: logic in the host processor for performing an activation function on the output patches.
[0284] 49. The system according to clause 31, the system comprising: logic in the host processor for performing a softmax function on the output patches.
[0285] 50. The system according to clause 31, wherein the block data arrays include M features, where M is greater than one.
[0286] 51. The system according to any one of the preceding clauses, wherein the neural network processor is a reconfigurable processor.
[0287] 52. The system according to any one of the preceding clauses, wherein the neural network processor is a configurable processor.
[0288] 53. A computer-implemented method for analyzing base calling sensor output, the method comprising:
[0289] Storing block data in a memory, the block data including sensor data of blocks of a sensing cycle from a base calling operation;
[0290] Performing a run of a neural network using trained parameters to generate classification data of a sensing cycle, the run of the neural network operating on a sequence of N arrays of block data from corresponding sensing cycles among N sensing cycles including a subject cycle to generate the classification data of the subject cycle; and
[0291] Using input units to move the block data and these trained parameters from a memory to the neural network for operation of the neural network, the input units including data of spatial alignment patches of the N arrays from respective sensing cycles out of N sensing cycles.
[0292] 54. The computer-implemented method according to clause 53, wherein the block data of the sensing cycle in the memory includes one or both of sensor data of the sensing cycle and intermediate data fed back from the neural network for the sensing cycle.
[0293] 55. The computer-implemented method according to clause 53, the method including: storing a block cluster mask that identifies elements in the sensor data arrays representing positions of flow cell clusters in the block, and applying the block cluster mask to intermediate data in the neural network.
[0294] 56. The computer-implemented method according to clause 53, wherein the data stream logic includes elements of the neural network processor configured using configuration data.
[0295] 57. The computer-implemented method according to clause 53, wherein the block data of the sensing cycle in the memory includes one or both of sensor data of the sensing cycle and intermediate data fed back from the neural network for the sensing cycle.
[0296] 58. The computer-implemented method according to clause 53, wherein the sensor data of the sensing cycle includes data representing signals for detecting one base for each flow cell cluster.
[0297] 59. The computer-implemented method according to clause 53, the method including: running a plurality of execution clusters of the neural network using spatial alignment patches of the block data; and
[0298] providing input units of the block data to available execution clusters among the plurality of execution clusters and causing the execution clusters to apply the block data of the N spatial alignment patches of the input units to the neural network to produce output patches of classification data for the subject cycle.
[0299] 60. The computer-implemented method according to clause 59, the method including: assembling the output patches from the plurality of execution clusters to provide base call classification data for the subject cycle, and storing the base call classification data in a memory.
[0300] 61. The computer-implemented method according to clause 59, wherein an execution cluster among the plurality of execution clusters includes: a set of computing engines having a plurality of members, configured to perform convolution on input data of multiple layers of the neural network using trained parameters, wherein the input data of the first layer is from the input unit, and the data of subsequent layers is from the activation data output from the previous layer.
[0301] 62. The computer-implemented method according to clause 61, the method includes: storing multiple versions of the trained parameters of the neural network, and wherein a cluster among the plurality of clusters in the neural network processor is configured to include a kernel memory to store these trained parameters; and providing an instance of the trained parameters to the kernel memory of an execution cluster among the plurality of execution clusters for executing the neural network.
[0302] 63. The computer-implemented method according to clause 62, wherein instances of these trained parameters are applied according to the number of cycles in these sensing cycles of the base detection operation.
[0303] 64. The computer-implemented method according to clause 59, wherein an execution cluster among the plurality of execution clusters includes a set of computing engines having a plurality of members, the set of computing engines being configured to apply a plurality of configurable filters for a corresponding layer of the neural network to sub-patches from the input unit and sub-patches of activation data output from a layer of the neural network.
[0304] 65. The computer-implemented method according to clause 59, wherein an execution cluster among the plurality of execution clusters is configured to perform an isolation stack of a spatial layer using sensor data of a current spatially aligned patch, feed intermediate data of the spatially aligned patch back to the memory to be used as block data in other input units, and provide the intermediate data from a previous cycle and the output of the isolation stack of the current cycle to the one or more combination layers to generate the output patch of the subject sensing cycle.
[0305] 66. The computer-implemented method according to clause 53, wherein the number N is an integer equal to 5 or greater.
[0306] 67. The computer-implemented method according to clause 53, the method includes: loading an input unit in a sequence of these block data arrays traversing multiple sensing cycles, the loading including writing the next input unit in the sequence to the neural network during execution of the neural network for a previous input unit.
[0307] 68. The computer-implemented method according to clause 53, wherein these block data arrays include M features, where M is greater than one.
[0308] 69. A computer-implemented method according to any of the preceding computer-implemented method clauses, the method comprising: configuring a configurable processor to execute the neural network.
[0309] 70. A computer-implemented method according to any of the preceding computer-implemented method clauses, the method comprising: configuring a reconfigurable processor to execute the neural network.
Claims
1. A system for analyzing the output of a base detection sensor, the system comprising: A host processor; A memory accessible by the host processor, the memory storing block data, the block data including a pixel array of a block from a sensing cycle of cluster image data in a base detection operation, wherein the cluster image depicts intensity emissions resulting from nucleotide incorporation in clusters of associated analytes on a substrate during a sequencing cycle of a sequencing run, and the block corresponds to a region of a set of clusters of a flow cell; And A neural network processor accessible by the memory, the neural network processor comprising: A plurality of execution clusters, an execution cluster among the plurality of execution clusters being configured to execute a neural network; Data flow logic accessible by the memory and the execution cluster among the plurality of execution clusters to provide an input unit including the block data including the pixel array to an available execution cluster among the plurality of execution clusters, the input unit including a digital N spatially aligned patches of the pixel array of the block data from a corresponding sensing cycle including a subject sensing cycle, and causing the execution cluster to apply the digital N spatially aligned patches to the neural network to generate an output patch of classification data of the spatially aligned patches of the cluster image data in the subject sensing cycle, wherein N is greater than 1; and An output layer of the neural network for generating a base detection probability of the subject sensing cycle based on the output patch from the plurality of execution clusters.
2. The system according to claim 1, the system comprising: Assembly logic for assembling the output patches from the plurality of execution clusters to provide base detection classification data of the subject sensing cycle, and storing the base detection classification data in the memory.
3. The system according to claim 1 or 2, wherein one of the plurality of execution clusters includes: A set of computing engines having a plurality of members, configured to convolve input data of a plurality of layers of the neural network using trained parameters, wherein the input data of the first layer is from the input unit, and the data of subsequent layers is from activation data output from a previous layer.
4. The system according to claim 3, the system comprising: A memory storing a plurality of versions of the trained parameters of the neural network, and wherein a cluster among the plurality of execution clusters in the neural network processor is configured to include a kernel memory for storing the trained parameters, and the data flow logic is configured to provide an instance of the trained parameters to the kernel memory of the execution cluster among the plurality of execution clusters for executing the neural network.
5. The system according to claim 4, wherein an instance of the trained parameters is applied according to the number of cycles in the sensing cycle of the base detection operation.
6. The system according to claim 3, wherein the execution cluster among the plurality of execution clusters includes a set of computing engines having a plurality of members, the set of computing engines being configured to apply a plurality of configurable filters for a corresponding layer of the neural network to a sub-patch from the input unit and a sub-patch of the activation data output from a layer of the neural network.
7. The system according to claim 1, wherein the neural network performs an isolated stack of spatial layers for each spatially aligned patch of the input unit and provides data from the N spatially aligned patches output from the isolated stack of the spatial layer to one or more combination layers.
8. The system according to claim 7, wherein an execution cluster among the plurality of execution clusters is configured to use the sensor data of the N spatially aligned patches of the input to perform N isolated stacks of spatial layers to generate an output patch of the classification data of the spatially aligned patches of the cluster image data in the subject sensing cycle.
9. The system according to claim 7, wherein an execution cluster among the plurality of execution clusters is configured to use the sensor data of the current spatially aligned patch of the cluster image data in the subject sensing cycle to perform the isolated stack of the spatial layer, feedback the intermediate data of the spatially aligned patch of the cluster image data in the subject sensing cycle to the memory to be used as block data in the input unit of other cycles, and provide the intermediate data from the previous cycle and the output of the isolated stack of the spatial layer of the current cycle to the one or more combination layers to generate the output patch of the classification data of the spatially aligned patch of the cluster image data in the subject sensing cycle.
10. The system according to claim 7, wherein the number N is an integer equal to 5 or greater.
11. The system according to claim 1, wherein the logic circuit is configured to load an input unit in a sequence of pixel arrays from the block data traversing a plurality of sensing cycles, the loading including writing the next input unit in the sequence to the neural network processor for the execution cluster during the execution of the neural network by the execution cluster for the previous input unit.
12. The system according to claim 1, wherein the system comprises: Logic in the host processor for performing an activation function on the output patch.
13. The system according to claim 1, the system comprising: Logic in the host processor for performing a softmax function on the output patch.
14. The system according to claim 1, wherein the pixel array from the block data includes M features, where M is greater than one.
15. The system according to claim 7, the system comprising: A block cluster mask for the block for base calling operation, the block cluster mask being stored in the memory, and the execution cluster being configured to use the block cluster mask to remove data from the intermediate data from at least one layer in the spatial layer.
16. A computer-implemented method for analyzing base calling sensor output, the method comprising: Storing block data in a memory, the block data including a pixel array of a block of a sensing cycle of cluster image data in a base calling operation, wherein the cluster image depicts intensity emissions resulting from nucleotide incorporation in an associated analyte cluster on a substrate during a sequencing cycle of a sequencing run, and the block corresponds to a region of a set of clusters of a flow cell; Performing a neural network on the block data including the pixel array using a plurality of execution clusters, the performing including: Provide input units of block data including units of the pixel array to available execution clusters among the multiple execution clusters of a neural network, the input units including digital N spatially-aligned patches of the pixel array of the block data from a respective sensing cycle including a subject sensing cycle, and output patches of classification data of the spatially-aligned patches causing the available execution clusters to apply the digital N spatially-aligned patches to the neural network to generate cluster image data in the subject sensing cycle, where N is greater than 1; and Using an output layer of the neural network, generate base calling probabilities for the subject sensing cycle based on the output patches from the multiple execution clusters.
17. The computer-implemented method according to claim 16, the method comprising: Assemble the output patches from the multiple execution clusters to provide base calling classification data of the spatially-aligned patches of the subject sensing cycle, and store the base calling classification data of the spatially-aligned patches of the subject sensing cycle in the memory.
18. The computer-implemented method according to claim 16 or 17, wherein the execution clusters among the plurality of execution clusters include: A set of computing engines having multiple members, configured to perform convolution on input data of multiple layers of the neural network using trained parameters, where the input data of the first layer is from an input unit, and the input data of subsequent layers is from activation data output from a previous layer.
19. The computer-implemented method according to claim 18, the method comprising: Store multiple versions of the trained parameters of the neural network, and where clusters among the multiple execution clusters in a neural network processor are configured to include kernel memories to store the trained parameters; And provide an instance of the trained parameters to the kernel memories of the execution clusters among the multiple execution clusters for executing the neural network.
20. A computer-implemented method for analyzing base calling sensor output, the method comprising: Store block data in a memory, the block data including a pixel array of a block depicting cluster image data of multiple clusters from a sensing cycle of a base calling operation, where the cluster image depicts intensity emissions resulting from nucleotide incorporation in associated analytes in clusters on a substrate during a sequencing cycle of a sequencing run, and the block corresponds to a region of a set of clusters of a flow cell; Perform a run of a neural network using trained parameters to generate classification data of cluster image data in a sensing cycle, the run of the neural network operating on sequences of N pixel arrays of the block data from respective sensing cycles among N sensing cycles including a subject sensing cycle to generate output patches of the classification data of the subject sensing cycle; Use an input unit to move the block data and the trained parameters from the memory to the neural network for the run of the neural network, the input unit including data of spatially-aligned patches of the N pixel arrays from the respective sensing cycles among the N sensing cycles to generate output patches of classification data of the spatially-aligned patches of the N pixel arrays from the respective sensing cycles among the N sensing cycles; and Provide the output patch to an output layer to generate base calling probabilities for the spatially aligned patches of the N pixel arrays from the respective sensing cycles of the N sensing cycles.
Citation Information
Patent Citations
Artificial intelligence-based generation of sequencing metadata
US11210554B2
Training data generation for artificial intelligence-based sequencing
US11347965B2
Artificial intelligence-based quality scoring
US11676685B2
Labelled nucleotides
US20060188901A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1