State-based base calling

JP2024536665A5Active Publication Date: 2025-10-01ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023579829
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-14
Filing Date
2022-09-21
Publication Date
2025-10-01
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

Existing base calling methods in next-generation sequencing face challenges with accuracy due to stochastic and circumstantial sources of variation, such as k-mer bias and PCR bias, which affect the precision of base calls, especially in GC-rich regions, and require large amounts of computer memory for training neural networks, leading to increased complexity and memory strain.

Method used

The method incorporates state information from previous sequencing cycles into the analysis of current sequencing cycles using state generators that generate summary statistics and channel-specific state values, which are then processed by a base caller, including neural networks, to improve base call accuracy and reduce memory requirements.

Benefits of technology

This approach enhances base call accuracy by reducing error rates and memory demands, allowing for more efficient and precise base calling, especially in regions prone to k-mer bias, while minimizing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The disclosed technology relates to state-based base calling. In particular, the disclosed technology relates to incorporating state information about data from previous sequencing cycles into the analysis of data from a current sequencing cycle when generating base calls for the current sequencing cycle. For example, when generating base calls for an Nth sequencing cycle, the disclosed technology can incorporate into the base call logic state information about data from sequencing cycles 1 through N-1.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority application) This application claims priority to and the benefit of U.S. Nonprovisional Patent Application No. 17 / 944,809 (Attorney Docket No. ILLM 1043-3 / IP-2073-US), entitled "STATE-BASED BASE CALLING," filed September 14, 2022, which claims the benefit of U.S. Provisional Patent Application No. 63 / 247,296 (Attorney Docket No. ILLM 1043-1 / IP-2073-PRV), filed September 22, 2021.

[0002] This application claims priority to and the benefit of U.S. Non-provisional Patent Application No. 17 / 944,948 (Attorney Docket No. ILLM 1043-4 / IP-2208-US), entitled "COMPRESSED STATE-BASED BASE CALLING," filed September 14, 2022, which claims the benefit of U.S. Provisional Patent Application No. 63 / 247,301 (Attorney Docket No. ILLM 1043-2 / IP-2208-PRV).

[0003] The priority application is incorporated herein by reference for all purposes as if fully set forth herein.

[0004] FIELD OF THE INVENTION The disclosed technology relates to artificial intelligence-based computers and digital data processing systems, and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the disclosed technology relates to using deep neural networks, such as deep convolutional neural networks, to analyze data.

[0005] (built-in) The following are incorporated by reference for all purposes as if fully set forth herein:

[0006] U.S. Nonprovisional Patent Application No. 17 / 308,035 (Attorney Docket No. ILLM 1032-2 / IP-1991-US), entitled "EQUALIZATION-BASED IMAGE PROCESSING AND SPATIAL CROSSTALK ATTENUATOR," filed May 4, 2021; U.S. Provisional Patent Application No. 63 / 106,256, entitled "SYSTEMS AND METHODS FOR PER-CLUSTER INTENSITY CORRECTION AND BASE CALLING," filed October 27, 2020 (Attorney Docket No. ILLM 1034-1 / IP-2026-PRV); U.S. Non-Provisional Patent Application No. 15 / 909,437, entitled "OPTICAL DISTORTION CORRECTION FOR IMAGED SAMPLES," filed March 1, 2018; U.S. Nonprovisional Patent Application No. 16 / 825,987, entitled "TRAINING DATA GENERATION FOR ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 20, 2020 (Attorney Docket No. ILLM 1008-16 / IP-1693-US); U.S. Nonprovisional Patent Application No. 16 / 825,991, entitled "ARTIFICIAL INTELLIGENCE-BASED GENERATION OF SEQUENCING METADATA," filed March 20, 2020 (Attorney Docket No. ILLM 1008-17 / IP-1741-US); U.S. Nonprovisional Patent Application No. 16 / 826,126, entitled "ARTIFICIAL INTELLIGENCE-BASED BASE CALLING," filed March 20, 2020 (Attorney Docket No. ILLM 1008-18 / IP-1744-US); U.S. Nonprovisional Patent Application No. 16 / 826,134, entitled "ARTIFICIAL INTELLIGENCE-BASED QUALITY SCORING," filed March 20, 2020 (Attorney Docket No. ILLM 1008-19 / IP-1747-US); U.S. Nonprovisional Patent Application No. 16 / 826,168, entitled "ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 21, 2020 (Attorney Docket No. ILLM 1008-20 / IP-1752-US); U.S. Nonprovisional Patent Application No. 17 / 175,546 (Attorney Docket No. ILLM 1015-2 / IP-1857-US), entitled "ARTIFICIAL INTELLIGENCE-BASED BASE CALLING OF INDEX SEQUENCES," filed February 12, 2021; U.S. Nonprovisional Patent Application No. 17 / 180,542, entitled "ARTIFICIAL INTELLIGENCE-BASED MANY-TO-MANY BASE CALLING," filed February 19, 2021 (Attorney Docket No. ILLM 1016-2 / IP-1858-US); U.S. Nonprovisional Patent Application No. 17 / 176,151 (Attorney Docket No. ILLM 1017-2 / IP-1859-US), entitled "KNOWLEDGE DISTILLATION-BASED COMPRESSION OF ARTIFICIAL INTELLIGENCE-BASED BASE CALLER," filed February 15, 2021; U.S. Provisional Patent Application No. 63 / 072,032, entitled "DETECTING AND FILTERING CLUSTERS BASED ON ARTIFICIAL INTELLIGENCE-PREDICTED BASE CALLS," filed August 28, 2020 (Attorney Docket No. ILLM1018-1 / IP-1860-PRV); U.S. Provisional Patent Application No. 63 / 161,880, entitled "TILE LOCATION AND / OR CYCLE BASED WEIGHT SET SELECTION FOR BASE CALLING," filed March 16, 2021 (Attorney Docket No. ILLM 1019-1 / IP-1861-PRV); U.S. Provisional Patent Application No. 63 / 161,896, entitled "NEURAL NETWORK PARAMETER QUANTIZATION FOR BASE CALLING," filed March 16, 2021 (Attorney Docket No. ILLM 1019-2 / IP-2049-PRV); U.S. Nonprovisional Patent Application No. 17 / 176,147, entitled "HARDWARE EXECUTION AND ACCELERATION OF ARTIFICIAL INTELLIGENCE-BASED BASE CALLER," filed February 15, 2021 (Attorney Docket No. ILLM 1020-2 / IP-1866-US). U.S. Provisional Patent Application No. 63 / 228,954, entitled "BASE CALLING USING MULTIPLE BASE CALLER MODELS," filed August 3, 2021 (Attorney Docket No. ILLM 1021-1 / IP-1856-PRV); U.S. Nonprovisional Patent Application No. 17 / 179,395 (Attorney Docket No. ILLM 1029-2 / IP-1964-US), entitled "DATA COMPRESSION FOR ARTIFICIAL INTELLIGENCE-BASED BASE CALLING," filed February 18, 2021; U.S. Nonprovisional Patent Application No. 17 / 180,480 (Attorney Docket No. ILLM 1030-2 / IP-1982-US), entitled "SPLIT ARCHITECTURE FOR ARTIFICIAL INTELLIGENCE-BASED BASE CALLER," filed February 19, 2021; U.S. Nonprovisional Patent Application No. 17 / 180,513 (Attorney Docket No. ILLM 1031-2 / IP-1965-US), entitled "BUS NETWORK FOR ARTIFICIAL INTELLIGENCE-BASED BASE CALLER," filed February 19, 2021; U.S. Provisional Patent Application No. 63 / 169,163, entitled "ARTIFICIAL INTELLIGENCE-BASED BASE CALLER WITH CONTEXTUAL AWARENESS," filed March 31, 2021 (Attorney Docket No. ILLM 1033-1 / IP-2007-PRV); U.S. Provisional Patent Application No. 63 / 216,419, entitled "SELF-LEARNED BASE CALLER, TRAINED USING OLIGO SEQUENCES," filed June 29, 2021 (Attorney Docket No. ILLM 1038-1 / IP-2050-PRV); U.S. Provisional Patent Application No. 63 / 216,404, entitled "SELF-LEARNED BASE CALLER, TRAINED USING ORGANISM SEQUENCE," filed June 29, 2021 (Attorney Docket No. ILLM 1038-2 / IP-2094-PRV); U.S. Provisional Patent Application No. 63 / 223,408, entitled "SPECIALIST SIGNAL PROFILERS FOR BASE CALLING," filed July 19, 2021 (Attorney Docket No. ILLM 1041-1 / IP-2063-PRV); U.S. Provisional Patent Application No. 63 / 226,707, entitled "QUALITY SCORE CALIBRATION OF BASECALLING SYSTEMS," filed July 28, 2021 (Attorney Docket No. ILLM 1045-1 / IP-2093-PRV); U.S. Provisional Patent Application No. 63 / 217,644, entitled "EFFICIENT ARTIFICIAL INTELLIGENCE-BASED BASE CALLING OF INDEX SEQUENCES," filed July 1, 2021 (Attorney Docket No. ILLM 1046-1 / IP-2135-PRV); U.S. Non-Provisional Patent Application No. 14 / 530,299, entitled "IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS," filed October 31, 2014; U.S. Non-Provisional Patent Application No. 15 / 153,953, entitled "METHODS AND SYSTEMS FOR ANALYZING IMAGE DATA," filed December 3, 2014; U.S. Nonprovisional Patent Application No. 15 / 863,241, entitled "PHASING CORRECTION," filed January 5, 2018; U.S. Nonprovisional Patent Application No. 14 / 020,570, filed September 6, 2013, entitled "CENTROID MARKERS FOR IMAGE ANALYSIS OF HIGH DENSITY CLUSTERS IN COMPLEX POLYNUCLEOTIDE SEQUENCING"; U.S. Nonprovisional Patent Application No. 12 / 565,341, entitled "METHOD AND SYSTEM FOR DETERMINING THE ACCURACY OF DNA BASE IDENTIFICATIONS," filed September 23, 2009; U.S. Non-Provisional Patent Application No. 12 / 295,337, entitled "SYSTEMS AND DEVICES FOR SEQUENCE BY SYNTHESIS ANALYSIS," filed March 30, 2007; U.S. Non-Provisional Patent Application No. 12 / 020,739, entitled "IMAGE DATA EFFICIENT GENETIC SEQUENCING METHOD AND SYSTEM," filed January 28, 2008; U.S. Nonprovisional Patent Application No. 13 / 833,619, entitled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR SAME," filed March 15, 2013 (Attorney Docket No. IP-0626-US); U.S. Nonprovisional Patent Application No. 15 / 175,489 (Attorney Docket No. IP-0689-US), entitled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND METHODS OF MANUFACTURING THE SAME," filed June 7, 2016; U.S. Non-Provisional Patent Application No. 13 / 882,088 (Attorney Docket No. IP-0462-US), entitled "MICRODEVICES AND BIOSENSOR CARTRIDGES FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR THE SAME," filed April 26, 2013; U.S. Nonprovisional Patent Application No. 13 / 624,200, entitled "METHODS AND COMPOSITIONS FOR NUCLEIC ACID SEQUENCING," filed September 21, 2012 (Attorney Docket No. IP-0538-US); U.S. Non-Provisional Patent Application No. 13 / 006,206, entitled "DATA PROCESSING SYSTEM AND METHODS," filed January 13, 2011; U.S. Nonprovisional Patent Application No. 15 / 936,365, filed March 26, 2018, entitled "DETECTION APPARATUS HAVING A MICROFLUOROMETER, A FLUIDIC SYSTEM, AND A FLOW CELL LATCH CLAMP MODULE"; U.S. Nonprovisional Patent Application No. 16 / 567,224, entitled "FLOW CELLS AND METHODS RELATED TO SAME," filed September 11, 2019; U.S. Nonprovisional Patent Application No. 16 / 439,635, entitled "DEVICE FOR LUMINESCENT IMAGING," filed June 12, 2019; U.S. Nonprovisional Patent Application No. 15 / 594,413, filed May 12, 2017, entitled "INTEGRATED OPTOELECTRONIC READ HEAD AND FLUIDIC CARTRIDGE USEFUL FOR NUCLEIC ACID SEQUENCING"; U.S. Nonprovisional Patent Application No. 16 / 351,193, entitled "ILLUMINATION FOR FLUORESCENCE IMAGING USING OBJECTIVE LENS," filed March 12, 2019; U.S. Non-provisional Patent Application No. 12 / 638,770, entitled "DYNAMIC AUTOFOCUS METHOD AND SYSTEM FOR ASSAY IMAGER," filed December 15, 2009; and U.S. Non-Provisional Patent Application No. 13 / 783,043, filed March 1, 2013, entitled "KINETIC EXCLUSION AMPLIFICATION OF NUCLEIC ACID LIBRARIES." [Background technology]

[0007] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section or problems associated with the subject matter provided as background have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which, as such, may also correspond to implementations of the claimed technology.

[0008] Rapid improvements in computing power have enabled deep convolutional neural networks (CNNs) to achieve great success in many computer vision tasks in recent years, with significantly improved accuracy. During the inference phase, many applications require low-latency processing of a single image with strict power consumption requirements, which reduces the efficiency of graphics processing units (GPUs) and other general-purpose platforms. This creates opportunities for specific acceleration hardware, such as field programmable gate arrays (FPGAs), by customizing digital circuits to be particularly effective for inferencing deep learning algorithms. However, deploying CNNs in portable and embedded systems remains challenging due to large data volumes, intensive computations, diverse algorithm structures, and frequent memory accesses.

[0009] Because convolution provides most of the operations in CNNs, the convolution acceleration scheme significantly impacts the efficiency and performance of hardware CNN accelerators. Convolution involves a multiply-and-accumulate (MAC) operation with four levels of loops that slide along kernels and feature maps. The first loop level calculates the MAC for pixels within a kernel window. The second loop level accumulates the sum of MAC products across various different input feature maps. After completing the first and second loop levels, the final output pixel is obtained by adding a bias. The third loop level slides the kernel window within the input feature map. The fourth loop level generates various different output feature maps.

[0010] FPGAs, particularly for accelerating inference tasks, have attracted increasing interest and become more widely used. This is due to their (1) high reconfigurability, (2) superiority over application-specific integrated circuits (ASICs) in terms of the development time required to keep up with the rapid evolution of CNNs, (3) good performance, and (4) superior energy efficiency compared to GPUs. The high performance and efficiency of FPGAs can be achieved by synthesizing circuits customized for specific computations and directly processing billions of operations with a customized memory system. For example, hundreds to thousands of digital signal processing (DSP) blocks in modern FPGAs support core convolution operations, such as multiply-and-accumulate operations with high parallelism. Dedicated data buffers between external on-chip memory and on-chip processing engines (PEs) can be designed to achieve prioritized data flow by configuring tens of megabytes of on-chip block random access memory (BRAM) on field-programmable gate array (FPGA) chips.

[0011] Efficient data flow and hardware architecture for CNN acceleration is desired to minimize data communication while maximizing resource utilization to achieve high performance. This creates an opportunity to design methodologies and frameworks to accelerate the inference process of various CNN algorithms on acceleration hardware and achieve high performance, high efficiency, and high flexibility.

[0012] A key feature of next-generation sequencing (NGS) technologies is parallelization, and the primary mechanism underlying several sequencing platforms is sequencing-by-synthesis (SBS). Briefly, tens to hundreds of millions of random DNA fragments are simultaneously sequenced by sequentially building complementary bases on a single-stranded DNA template and capturing the synthesis information in a series of raw fluorescent images.

[0013] Extracting the actual sequence information (i.e., strings of characters in {A,C,G,T}) from image data involves two computational tasks: image analysis and base calling. The main function of image analysis is to translate image data into fluorescence intensity data for each DNA fragment, while the goal of base calling is to infer sequence information from the resulting intensity data.

[0014] There are several stochastic and situational sources of variation that can reduce base calling accuracy. For example, k-mer bias in base calling is influenced by the GC content of the sequenced genome. Base callers, when applied to GC-rich regions of DNA, primarily result from reduced sequence complexity, but can also exhibit bias as a result of polymerase chain reaction (PCR) bias during the amplification step.

[0015] Base calling accuracy is essential for various downstream applications, including sequence assembly, SNP calling, and genotype calling. Improving base calling accuracy can enable achieving the desired performance of downstream applications with lower sequencing coverage, which leads to reduced sequencing costs.

[0016] Training neural networks for base calling requires large amounts of computer memory, which increases exponentially with increasing image size and number. Computer memory becomes a limiting factor because the backpropagation algorithm for optimizing deep neural networks requires storage of intermediate activations. The size and number of these intermediate activations increase proportionally to the input size and number, so memory fills up quickly with larger and more numerous images.

[0017] Base callers using neural networks, e.g., those disclosed in commonly owned patent applications 16 / 826,126, 16 / 826,134, 16 / 826,168, 17 / 175,546, 17 / 180,542, 17 / 176,151, 63 / 072,032, 63 / 161,880, 63 / 161,896, In accordance with one embodiment, those disclosed in US Pat. Nos. 17 / 176,147, 63 / 228,954, 17 / 179,395, 17 / 180,480, 17 / 180,513, 63 / 169,163, and 63 / 217,644 use image data for a sliding window of sequencing cycles to make base call predictions. Increasing the size of the sliding window to include image data from more sequencing cycles increases the complexity of the neural network and also places additional strain on available computation and memory.

[0018] Opportunities arise to configure base calling operations to incorporate contextual information from multiple past sequencing cycles, which may result in more accurate base calls with reduced error rates, especially for attenuation of k-mer bias. [Brief explanation of the drawings]

[0019] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Figure 1] FIG. 1 is a high-level diagram of the disclosed state-based base calling. [Figure 2] 1 illustrates one embodiment of performing base calling operations using state information generated based on past intensity values ​​of pixels in a sequencing image. [Figure 3] 1 illustrates one embodiment for generating pixel-by-pixel states for pixels in a sequencing image that are processed as input for base calling. [Figure 4] 1 illustrates one embodiment of generating base calls based on pixel-by-pixel states. [Figure 5] 1 illustrates one embodiment for generating per-channel states for pixels in a sequencing image that are processed as input for base calling. [Figure 6] 1 illustrates one embodiment for generating per-channel MIN (minimum) states for pixels in a sequencing image that are processed as input for base calling. [Figure 7] 1 illustrates one embodiment for generating pan-channel MIN (minimum) states for pixels in a sequencing image that are processed as input for base calling. [Figure 8] 1 illustrates one embodiment for generating per-channel MAX states for pixels in a sequencing image that are processed as input for base calling. [Figure 9] 1 illustrates one embodiment for generating pan channel MAX states for pixels in a sequencing image that are processed as input for base calling. [Figure 10] 1 illustrates one embodiment for generating per-channel AVG (average) states of pixels in a sequencing image that are processed as input for base calling. [Figure 11]1 illustrates one embodiment for generating pan-channel AVG (average) states for pixels in a sequencing image that are processed as input for base calling. [Figure 12] 12 shows one embodiment for generating per-channel MIN / MAX (minimum and maximum) states 1200 for pixels in a sequencing image that are processed as input for base calling. [Figure 13] 13 shows one embodiment for generating per-channel MIN AVG (minimum and average) states 1300 for pixels in a sequencing image that are processed as input for base calling. [Figure 14] 14 shows one embodiment that generates a per-channel MAX AVG (maximum and average) state 1400 for pixels in a sequencing image that are processed as input for base calling. [Figure 15] 1 illustrates one embodiment for generating per-channel MIN MAX AVG (minimum, maximum, and average) states for pixels in a sequencing image processed as input for base calling. [Figure 16] 1 shows one embodiment that relies on previously called bases to generate state information for use in future base calling. [Figure 17A] Base call-directed ON and OFF state generation for the current sequencing cycle 11 (cycle 11) is shown. [Figure 17B] Base call-directed ON and OFF state generation for the current sequencing cycle 11 (cycle 11) is shown. [Figure 17C] Base call-directed ON and OFF state generation for the current sequencing cycle 11 (cycle 11) is shown. [Figure 18A] Figures 17A, 17B, and 17C are expanded to show how the OFF state is updated in the next sequencing cycle 12 (cycle 12). [Figure 18B] Figures 17A, 17B, and 17C are expanded to show how the OFF state is updated in the next sequencing cycle 12 (cycle 12). [Figure 18C]Figures 17A, 17B, and 17C are expanded to show how the OFF state is updated in the next sequencing cycle 12 (cycle 12). [Figure 19A] Figures 18A, 18B, and 18C are further expanded to show how the ON state is updated in the next sequencing cycle thirteen (cycle 13). [Figure 19B] Figures 18A, 18B, and 18C are further expanded to show how the ON state is updated in the next sequencing cycle thirteen (cycle 13). [Figure 19C] Figures 18A, 18B, and 18C are further expanded to show how the ON state is updated in the next sequencing cycle thirteen (cycle 13). [Figure 20A] 19A, 19B, and 19C are further expanded to show how both the ON and OFF states are updated in the next sequencing cycle, 14 (Cycle 14). [Figure 20B] 19A, 19B, and 19C are further expanded to show how both the ON and OFF states are updated in the next sequencing cycle, 14 (Cycle 14). [Figure 20C] 19A, 19B, and 19C are further expanded to show how both the ON and OFF states are updated in the next sequencing cycle, 14 (Cycle 14). [Figure 21] 1 illustrates one embodiment of generating state information for base calls using exponentially weighted averaging. [Figure 22] 10 illustrates one embodiment of using exponentially weighted averaging to generate per-channel state information for base calling. [Figure 23] 10 shows an example of using exponentially weighted averaging to generate per-channel state information for base calling. [Figure 24] 10 illustrates one embodiment that uses exponentially weighted averaging directed by previously called bases to generate per-channel state information for base calling. [Figure 25]An example is shown of generating per-channel state information for base calling using exponentially weighted averaging directed by previously called bases. [Figure 26] 10 illustrates one embodiment of generating state information for non-cluster pixels using previously called bases. [Figure 27] 1 illustrates different embodiments for generating state data for base calls. [Figure 28] 10A-10C illustrate different implementations of supply state data in various aspects of the processing of a base calling operation. [Figure 29] 1 illustrates one embodiment of state-based and neural network-based base calling. [Figure 30] 1 illustrates one embodiment that uses per-cluster state data for base call clusters. [Figure 31] 10 illustrates one implementation of generating a state for each cluster using past intensity values ​​of corresponding cluster pixels. [Figure 32] 10 illustrates one embodiment of generating a state for each cluster using past feature values ​​of spatial convolution features corresponding to cluster pixels. [Figure 33] 1 shows one embodiment in which pixel intensities are interpolated to generate cluster intensities, and the interpolated cluster intensities are used to generate per-cluster states for base calls. [Figure 34] 1 illustrates one embodiment of generating states for each cluster using a real-time analysis (RTA) base caller that is separate from the neural network-based base caller. [Figure 35] 1 illustrates one embodiment of compression logic for generating compressed features. [Figure 36] 10 illustrates one embodiment of generating a state for each cluster using past feature values ​​of compressed spatial convolution features corresponding to cluster pixels. [Figure 37]For the incorporation of cluster-wise states with compressed features, we present one embodiment that generates cluster-wise states using an RTA base caller that is separate from the neural network-based base caller. [Figure 38] To incorporate the state for each cluster into the compressed features, one embodiment is shown in which the state for each cluster is generated using past intensity values ​​of the corresponding cluster pixels. [Figure 39] 1 illustrates one embodiment for generating dense per-pixel states from sparse per-well states. [Figure 40] 10 illustrates one embodiment that provides sparse per-well states, dense per-pixel states, and per-pixel intensity values ​​as inputs to a base caller to perform base calling operations. [Figure 41] 10 illustrates one embodiment that uses sparse per-well states, dense per-pixel states, and per-pixel intensity values ​​as inputs to a neural network-based base caller to perform base calling operations. [Figure 42A] 1 illustrates one embodiment of a sequencing system that includes a configurable processor. [Figure 42B] 1 illustrates one embodiment of a sequencing system that includes a configurable processor. [Figure 42C] FIG. 1 is a simplified block diagram of a system for analysis of sensor data from a sequencing system, such as base call sensor output. [Figure 43A] FIG. 1 is a simplified diagram illustrating aspects of base calling operations, including the functionality of a runtime program executed by a host processor. [Figure 43B] 1 is a simplified diagram of the configuration of a configurable processor. [Figure 44] 10 illustrates one embodiment of determining state data on the CPU and loading the state data from the CPU into the FPGA for base calls. [Figure 45]We compare the base calling performance of a non-neural network-based base caller RTA, a neural network-based base caller without state information DeepRTA(5ci_k14), and the disclosed state-based and neural network-based base caller DeepRTA with State(5ci_k14_DC), which uses pixel-by-pixel MIN state information as additional / supplementary input. [Figure 46] We compare the base calling performance of the disclosed state-based and neural network-based base callers DeepRTA with State(5ci_k14_DC) across different state channels. [Figure 47] Figure 1 shows the base calling performance of the disclosed state-based and neural network-based base caller DeepRTA with State (DC) for k-mers, specifically 5-mers. A 5-mer refers to a repeating base pattern for five base positions (e.g., ACGCG, GGGGG, TCGCG). [Figure 48] We compare the base calling performance of a non-neural network-based base caller RTA, a neural network-based base caller without state information DeepRTA(5ci_k14), a disclosed state-based and neural network-based base caller DeepRTA with State(5ci_k14_DC_MIN) that uses per-pixel MIN state information as additional / supplementary input, and a disclosed state-based and neural network-based base caller DeepRTA with State(5ci_k14_DC_MIN_MAX) that uses per-pixel MIN and MAX state information as additional / supplementary input. [Figure 49]We compare the base calling performance of a non-neural network-based base caller, RTA; a neural network-based base caller, DeepRTA(5ci_k14), without state information and a 5-cycle neighboring input window; a disclosed state-based and neural network-based base caller, DeepRTA with State(5ci_k14_DC_MIN), which uses per-pixel MIN state information as additional / supplementary input with a 5-cycle neighboring input window; a disclosed state-based and neural network-based base caller, DeepRTA with State(3ci_k14_DC_MIN_MAX), which uses per-pixel MIN and MAX state information as additional / supplementary input with a 3-cycle neighboring input window; and a disclosed state-based and neural network-based base caller, DeepRTA with State(5ci_k14_DC_MIN_MAX), which uses per-pixel MIN and MAX state information as additional / supplementary input with a 5-cycle neighboring input window. [Figure 50] The base calling performance of the disclosed state-based and neural network-based base caller DeepRTA with State(5ci_k14_DC_MIN_MAX), which uses per-pixel MIN and MAX state information as additional / supplemental input along with a 5-cycle neighboring input window on k-mers (especially 5-mers), is compared with the base calling performance of the neural network-based base caller DeepRTA(5ci_k14) without state information and a 5-cycle neighboring input window. 5-mers refer to repeating base patterns for five base positions (e.g., ACGCG, GGGGG, TCGCG). [Figure 51]We compare the base calling performance of a non-neural network-based base caller RTA with an equalizer implementation, a disclosed state-based and neural network-based base caller DeepRTA with State(5ci_k14_DC_MIN_MAX) that uses per-pixel MIN and MAX state information as additional / supplementary inputs along with a 5-cycle neighboring input window, and a disclosed state-based and neural network-based base caller DeepRTA with State(5ci_k14_DC_AVG) that uses per-pixel AVG state information as additional / supplementary inputs along with a 5-cycle neighboring input window. [Figure 52] We compare the base calling performance of a non-neural network based base caller RTA with an equalizer implementation with the disclosed state-based and neural network based base caller DeepRTA with State(5ci_k14_DC_MIN_MAX_AVG), which uses per-pixel MIN, MAX, and AVG state information as additional / supplementary input along with a 5-cycle neighboring input window. [Figure 53] We compare the base calling performance of a non-neural network-based base caller RTA with an equalizer implementation, a disclosed state-based and neural network-based base caller DeepRTA with State(5ci_k14_DC_MIN_MAX) that uses per-pixel MIN and MAX state information as additional / supplemental inputs with a 5-cycle neighboring input window and a filter bank of size 14 (K=14), and a disclosed state-based and neural network-based base caller DeepRTA with State(5ci_k14_DC_MIN_MAX) that uses per-pixel MIN and MAX state information as additional / supplemental inputs with a 5-cycle neighboring input window and a filter bank of size 32 (K=32). [Figure 54]As discussed above, we compare the base calling performance of the disclosed state-based and neural network-based base callers DeepRTA with State for different alpha parameter configurations (e.g., 0.05, 0.07, 0.10, 0.12) of the exponentially weighted averaging implementation. [Figure 55] 1 is a graph tracking state values ​​determined by exponentially weighted averaging, according to one embodiment of the disclosed technology. [Figure 56] 1 is a graph tracking state values ​​determined by exponentially weighted averaging, according to one embodiment of the disclosed technology. [Figure 57] We compare the base calling performance of RTA with an equalizer implementation (RTA+Eq), DeepRTA without state information (DeepRTA(k14_1m)), and the disclosed DeepRTA with State (DeepRTA_extrachannels(k14_1,_onoff_expavg)) in terms of secondary analysis tasks and metrics such as the number of called reads, read mismatch error rates for Read 1 and Read 2 and their averages, single nucleotide polymorphism (SNP) reproducibility, SNP accuracy, SNP calling accuracy (F1 score), insertion / deletion (Indel) reproducibility, Indel accuracy, and Indel calling accuracy (F1 score). DETAILED DESCRIPTION OF THE INVENTION

[0020] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0021] The detailed description of various embodiments can be better understood when read in conjunction with the accompanying drawings. To the extent that the figures illustrate diagrams of functional blocks of various embodiments, the functional blocks are not necessarily indicative of a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., modules, processors, or memories) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or block of random access memory, hard disk, etc.) or in multiple pieces of hardware. Similarly, a program may be a stand-alone program, may be incorporated as a subroutine within an operating system, may be a function within an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentality shown in the figures.

[0022] The processing engines and databases in the figures designated as modules can be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some modules may be implemented on different processors, computers, or servers, or may be spread across multiple different processors, computers, or servers. In addition, it will be understood that some of the modules may operate in parallel or in a different order than shown in the figures without affecting the functionality achieved. The modules in the figures may also be considered flowchart steps in a method. Also, a module need not necessarily have all its code located contiguously in memory. Some portions of code may be separated from other portions of code, with code from other modules or other functions located between them.

[0023] Introduction The disclosed technology relates to state-based base calling. In particular, the disclosed technology relates to incorporating state information about data from previous sequencing cycles into the analysis of data from a current sequencing cycle when generating base calls for the current sequencing cycle. For example, when generating base calls for an Nth sequencing cycle, the disclosed technology can incorporate into base call logic state information about data from sequencing cycles 1 through N-1.

[0024] The following description describes various implementations of the disclosed state-based base calling. Implementations vary based on different data and processing aspects. For example, different "methods" or "logic" for generating state information will result in different types of states. Also, the "things" whose past values ​​are tracked for state generation may vary from implementation to implementation. Furthermore, once generated, the "when," "where," and "how" the state information is processed for base calling will result in various implementations of the disclosed technology.

[0025] State-based base calling 1 is a high-level diagram of the disclosed state-based base calling 100. Base calling is the process of determining the nucleotide composition of a sequence. In one embodiment, base calling involves analysis of image data, i.e., sequencing images, generated during a sequencing run (or sequencing reaction) performed by a sequencing system such as Illumina's iSeq, HiSeqX, HiSeq3000, HiSeq4000, HiSeq2500, NovaSeq6000, NextSeq550, NextSeq1000, NextSeq2000, NextSeqDx, MiSeq, and MiSeqDx. In other embodiments, base calling can involve inferring sequence reads from non-image sequencing data.

[0026] The sequencing system 104 can be used to sequence nucleic acids. Applicable techniques include those in which nucleic acids are attached to fixed locations in an array (e.g., wells of a flow cell) and the array is repeatedly imaged. In such an embodiment, the sequencing system 104 can acquire images in two different color channels that can be used to distinguish a particular nucleotide base type from another nucleotide base. More specifically, the sequencing system 104 can perform a process called "base calling," which generally refers to the process of determining a base call (e.g., adenine (A), cytosine (C), guanine (G), or thymine (T)) for a given spot location in an image during an imaging cycle. During two-channel base calling, for example, image data extracted from two images can be used to determine the presence of one of four base types by encoding the base identity as a combination of the intensities of the two images. For a given spot or location in each of the two images, the base identity can be determined based on whether the signal identity combination is [on, on], [on, off], [off, on], or [off, off].

[0027] Output data from sequencing system 104 can be communicated to a real-time analysis module (not shown). The real-time analysis module, in various embodiments, executes computer-readable instructions for analyzing image data (e.g., image quality scoring, base calling, etc.), reporting or displaying beam characteristics (e.g., focus, shape, intensity, power, brightness, position) in a graphical user interface (GUI), etc. These operations can be performed in real time during an imaging cycle to minimize downstream analysis time and provide real-time feedback and troubleshooting during an imaging run. In an embodiment, the real-time analysis module can be a computing device that is communicatively coupled to and controls the imaging subsystem of sequencing system 104.

[0028] The following description outlines how sequencing images are generated and what they depict, according to one embodiment.

[0029] In some embodiments, base calling decodes the intensity data encoded in the sequencing image into a nucleotide sequence. In one embodiment, Illumina's sequencing platform employs cyclic reversible termination (CRT) chemistry for base calling. This process relies on extending a nascent strand complementary to the template strand with fluorescently labeled nucleotides while tracking the emission signal of each newly added nucleotide. Fluorescently labeled nucleotides have a 3'-removable block that anchors the fluorophore signal of the nucleotide.

[0030] Sequencing is performed in iterative cycles, each of which includes three steps: (a) extending the emerging strand by adding fluorescently labeled nucleotides; (b) exciting fluorophores using one or more lasers in the optical subsystem of the sequencing system 104 and generating a sequencing image by imaging through different filters in the optical subsystem; and (c) cleaving the fluorophores and removing the 3' block in preparation for the next sequencing cycle. The capture and imaging cycles are repeated for a specified number of sequencing cycles, which defines the read length. Using this approach, each cycle locates a new position along the template strand.

[0031] The tremendous power of Illumina sequencers comes from their ability to simultaneously run and sense millions or even billions of clusters (also called "clusters") undergoing CRT reactions. A cluster contains approximately 1,000 identical copies of a template strand, but the size and shape of the cluster vary. Clusters are grown from the template strands by bridge amplification or exclusion amplification of the input library prior to a sequencing run. The purpose of amplification and cluster extension is to increase the intensity of the emitted signal, since imaging devices cannot reliably sense the fluorophore signal of a single strand. However, because the physical distance between strands within a cluster is small, imaging devices perceive a cluster of strands as a single spot.

[0032] Sequencing is performed in a flow cell (or biosensor), a small glass slide that holds the input strands. The flow cell is connected to an optical system that includes microscope imaging, an excitation laser, and a fluorescence filter. The flow cell contains multiple chambers called lanes. The lanes are physically separated from each other and can contain different tagged sequencing libraries that are distinguishable without cross-contamination of samples. In some embodiments, the flow cell contains a patterned surface. "Patterned surface" refers to the arrangement of different regions within or on an exposed layer of a solid support.

[0033] The imaging device of the sequencing system 104 (e.g., a solid-state imaging device such as a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) sensor) takes snapshots at multiple locations along the lane in a series of non-overlapping regions called tiles. For example, there may be 64 or 96 tiles per lane. A tile holds hundreds of thousands to millions of clusters.

[0034] The output of a sequencing run is a sequencing image. A sequencing image shows the intensity radiation of clusters and their surrounding background using a grid (or array) of pixelated units (e.g., pixels, superpixels, subpixels). The intensity radiation is stored as intensity values ​​of the pixelated units. A sequencing image has dimensions w by h of the grid of pixelated units, where w (width) and h (height) are any number ranging from 1 to 100,000 (e.g., 115 x 115, 200 x 200, 1800 x 2000, 2200 x 25000, 2800 x 3600, 4000 x 400). In some embodiments, w and h are the same. In other embodiments, w and h are different. A sequencing image shows the intensity radiation generated as a result of incorporating nucleotides into a nucleotide sequence during a sequencing run. The intensity radiation is from the associated clusters and their surrounding background.

[0035] In one embodiment, data flow logic (not shown) provides the sequencing image to base caller 144 for base calling. In one embodiment, base caller 144 accesses the sequencing image patch by patch (or tile by tile). Each patch is a subgrid (or subarray) of pixelated units within the grid of pixelated units that forms the sequencing image. A patch has dimensions q by r of the subgrid of pixelated units, where q (width) and r (height) are any number ranging from 1 to 10,000 (e.g., 3x3, 5x5, 7x7, 10x10, 15x15, 25x25, 64x64, 78x78, 115x115). In some embodiments, q and r are the same. In other embodiments, q and r are different from one another. In some embodiments, patches extracted from a single sequencing image are the same size. In other embodiments, patches are of different sizes. In some implementations, patches can have overlapping pixelated units (eg, on edges).

[0036] According to some embodiments, sequencing generates m sequencing images per sequencing cycle for the corresponding m image channels. That is, each sequencing image has one or more image (or intensity) channels (similar to the red, green, and blue (RGB) channels of a color image). In one embodiment, each image channel corresponds to one of a plurality of filter wavelength bands. In another embodiment, each image channel corresponds to one of a plurality of imaging events in a sequencing cycle. In yet another embodiment, each image channel corresponds to a combination of illumination with a particular laser and imaging through a particular optical filter. Image patches are tiled (or accessed) from each of the m image channels for a particular sequencing cycle. In different embodiments, such as 4-, 2-, and 1-channel chemistries, m is 4 or 2. In other embodiments, m is greater than 1, 3, or 4. In other embodiments, images can be in blue and violet channels instead of or in addition to red and green channels.

[0037] For example, consider a sequencing run performed using two different imaging channels, i.e., a blue channel and a green channel. Then, in each sequencing cycle, the sequencing run generates a blue image and a green image. In this manner, for a series of k sequencing cycles of the sequencing run, a sequence of k pairs of blue and green images is generated as output and stored as the sequencing image. Thus, a sequence of k pairs of blue and green image patches is generated for patch-level processing by the base caller 144.

[0038] The sequencing system 104 can also generate non-image sequencing data according to other embodiments. In one embodiment, sequencing data can be based on pH changes induced by the release of hydrogen ions during molecular elongation. The pH change is detected and converted to a voltage change proportional to the number of incorporated bases. In yet another embodiment, sequencing data can be constructed from nanopore sensing, for example, using a biosensor to determine the identity of the bases while simultaneously measuring the current disruption as the cluster passes through or near the opening of the nanopore. In one embodiment, nanopore-based sequencing can be based on the following concept: a single strand of DNA (or RNA) is passed through a membrane via a nanopore, and a potential difference is applied across the membrane. Nucleotides present within the pore can affect the electrical resistance of the pore, so that current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal (due to its appearance of "squishing" when plotted) is the raw data collected by the sequencer. These measurements can be stored as 16-bit integer Data Acquisition (DAC) values ​​taken at a 4 kHz frequency (for example). With a DNA strand velocity of ∼450 base pairs per second, this can give, on average, approximately 9 raw observations per base. This signal can then be processed to identify breaks in the aperture signal corresponding to individual reads. Extensions of these raw signals can be base-called, a process that converts DAC values ​​into sequences of DNA bases. In some embodiments, sequencing data can include normalized or scaled DAC values.

[0039] Current sequencing data 112 includes sequencing data generated by sequencing data 104 for the current sequencing cycle of the sequencing run. In Figure 1, the current sequencing cycle is identified as the "N" sequencing cycle.

[0040] Previous sequencing data 116 includes sequencing data generated by sequencing data 104 for one or more previous sequencing cycles of a sequencing run. The previous sequencing cycles precede the current sequencing cycle. In FIG. 1, the previous sequencing cycles are identified as "1 to N-1" sequencing cycles. In other embodiments, previous sequencing data 116 includes sequencing data for a subset of the 1 to N-1 sequencing cycles.

[0041] The state generator 126 uses the current sequencing data 112 and the previous sequencing data 116 to generate current state data 136 for the current sequencing cycle. The state generator 126 may be a value or a function that is applied to the sequencing data to produce a desired result. The state generator 126 may be applied to the sequencing data by any of a variety of mathematical operations, including, but not limited to, addition, subtraction, division, multiplication, or combinations thereof. The state generator 126 may be a mathematical formula, a logical function, a computer-implemented algorithm, etc. The sequencing data may be image data, electrical data, or a combination thereof.

[0042] In one embodiment, the state generator 126 generates the current state data 136 by accumulating summary statistics of the current sequencing data 112 and the previous sequencing data 116. Examples of summary statistics include maximum, minimum, average (mean), exponentially weighted average, running average, exponential moving average, mode, standard deviation, variance, skewness, kurtosis, percentile, and entropy. In another embodiment, the state generator 126 determines secondary statistics based on the summary statistics. Examples of secondary statistics include delta, sum, a set of maxima, a set of minima, the minimum of a set of maxima, and the maximum of a set of minima.

[0043] Base caller 144 generates current base call data 154 for the current sequencing cycle in response to processing current sequencing data 112 and current state data 136. Current base call data 154 may include base calls for one or more clusters. In some embodiments, current sequencing data 112 and current state data 136 are combined prior to processing by base caller 144. Combinations can be effected, for example, by summation operations, element-wise multiplication operations, element-wise multiplication and sum (convolution) operations, and concatenation operations.

[0044] Examples of base callers 144 include different base calling procedures available on the Illumina platform, such as Real Time Analysis (RTA), BlindCall, freeIbis, Softy, AYB, OnlineCall, BM-BC, ParticleCall, TotalReCaller, naiveBayesCall, Srfim, BayesCall, Ibis, Rolexa, Alta-Cyclic, and Bustard. Examples of base callers 144 include commonly owned patent applications 16 / 825,987, 16 / 825,991, 16 / 826,126, 16 / 826,134, 16 / 826,134, 16 / 826,168, 17 / 175,546, 17 / 180,542, 17 / 176,151, 63 / 072,032, 63 / 072,032, and 63 / 072,032. Also included are Illumina's neural network-based offerings, such as those disclosed in US Patent Nos. 161,880, 63 / 161,896, 17 / 176,147, 63 / 228,954, 17 / 179,395, 17 / 180,480, 17 / 180,513, 63 / 169,163, and 63 / 217,644, which are collectively referred to herein as "DeepRTA" or "Deep Learning Primary Analysis." Further examples of base callers 144 include different base calling procedures available from Oxford Nanopore Technologies (ONT), such as Metrichor, Nanocall, DeepNano, Nanonet, Scrappie, Albacore, Guppy, Basecrawler, Chiron, Halcyon, MinCall, SACall, Causalcall, and WaveNano.

[0045] Pixelated State Data FIG. 2 illustrates one embodiment for performing base calling operations using state information 200 generated based on past intensity values ​​of pixels in a sequencing image.

[0046] In action 202, the memory stores (e.g., in data store 112) a respective current intensity value for each pixel in the plurality of pixels for a current sequencing cycle of the sequencing run. In action 212, the memory stores (e.g., in data store 116) a sequence of respective previous intensity values ​​for each pixel for one or more previous sequencing cycles of the sequencing run preceding the current sequencing cycle. In one embodiment, each intensity value for each pixel is characterized by channel-specific intensity values ​​for a plurality of channels. In one embodiment, a channel corresponds to a combination of illumination by a particular laser and imaging through a particular optical filter. In another embodiment, a channel corresponds to a filter wavelength band. In yet another embodiment, a channel corresponds to an imaging event in a sequencing cycle.

[0047] In operation 222, state generator 126, which has access to memory, generates a respective current state value for each pixel depending on (i) the respective current intensity value and (ii) each previous intensity value in the sequence of respective previous intensity values. State generator 126 also stores the respective current state values ​​in memory. In some implementations, each state value for each pixel is characterized by channel-specific state values ​​for a plurality of channels. In one implementation, each current state value is encoded on a pixel-by-pixel basis with the respective current intensity values. In some implementations, channel-specific state values ​​for a subset of channels in the plurality of channels are encoded on a pixel-by-pixel basis with the respective current intensity values. In some implementations, the channel-specific state values ​​are averaged across channels in the plurality of channels to generate a respective pan channel current state value. In one implementation, each pan channel current state value is encoded on a pixel-by-pixel basis with the respective current intensity values.

[0048] In some implementations, the pixel-wise encoding comprises pixel-wise concatenation. In other implementations, the pixel-wise encoding comprises pixel-wise summation. In still other implementations, the pixel-wise encoding comprises element-wise multiplication. In still other implementations, the pixel-wise encoding comprises element-wise multiplication and summation (convolution).

[0049] In one embodiment, each state value is configured to characterize a past intensity pattern of the respective pixel. In another embodiment, the past intensity pattern is configured to compensate for loss of base calling accuracy of the base caller 144. In yet another embodiment, the past intensity pattern is configured to compensate for loss of base calling accuracy of the base caller 144 when base calling bases of the k-mer. In yet another embodiment, each state value is configured to identify a respective signal profile of the cluster.

[0050] In one embodiment, the respective current state values ​​are respective current average intensities determined for the respective pixels in the current sequencing cycle from (i) the respective previous intensity values, and (ii) the respective current intensity values, hi some embodiments, the respective current average intensities are each characterized by a channel-specific current average intensity.

[0051] In another embodiment, the respective current state values ​​are respective current maximum intensities determined for the respective pixels in the current sequencing cycle from (i) the respective previous intensity values, and (ii) the respective current intensity values, hi some embodiments, the respective current maximum intensities are each characterized by a channel-specific current maximum intensity.

[0052] In yet another embodiment, the respective current state values ​​are respective current maximum intensities determined for the respective pixels in the current sequencing cycle from (i) the respective previous intensity values, and (ii) the respective current intensity values, hi some embodiments, the respective current maximum intensities are each characterized by a channel-specific current maximum intensity.

[0053] In yet another embodiment, the respective current state values ​​are respective current minimum intensities determined for the respective pixels in the current sequencing cycle from (i) the respective previous intensity values, and (ii) the respective current intensity values, hi some embodiments, the respective current minimum intensities are each characterized by a channel-specific current minimum intensity.

[0054] In yet another embodiment, the respective current state values ​​are respective current exponentially weighted average intensities determined for the respective pixels in the current sequencing cycle from (i) the respective previous intensity values ​​and (ii) the respective current intensity values. In some embodiments, the respective current exponentially weighted average intensities are determined based on weighting more recent sequencing cycles than previous sequencing cycles. In other embodiments, the respective current exponentially weighted average intensities are each characterized by a channel-specific current exponentially weighted average intensity.

[0055] In yet another embodiment, the respective current state values ​​are (i) the respective current intensity values, and (ii) the respective current running average intensities determined for the respective pixels in the current sequencing cycle from a rolling subset of the respective previous intensity values. In some embodiments, the respective current running average intensities are each characterized by a channel-specific current running average intensity.

[0056] In some embodiments, each respective current intensity value is assigned to an active-state bucket or an inactive-state bucket based on a comparison of the respective current intensity value with the respective current active-state intensity and the respective current inactive-state intensity. In one embodiment, the respective current active-state intensity is (i) a respective current global maximum intensity determined for the respective pixel in the current sequencing cycle from the respective previous intensity value. In some embodiments, each respective current global maximum intensity is characterized by a channel-specific current global maximum intensity.

[0057] In one embodiment, each current inactive state intensity is (i) a respective current global minimum intensity determined for each pixel in the current sequencing cycle from each previous intensity value, hi some embodiments, each current global minimum intensity is characterized by a channel-specific current global minimum intensity.

[0058] In one embodiment, the respective current state values ​​further include a respective current active state value and a respective current inactive state value generated for the respective pixel in the current sequencing cycle. In some embodiments, the respective current active state values ​​are each characterized by a channel-specific current active state value. In other embodiments, the respective current inactive state values ​​are each characterized by a channel-specific current inactive state value.

[0059] In one embodiment, the current active state value of a pixel in the target channel is determined in the current sequencing cycle from (i) the current intensity value for the pixel in the target channel detected in the current sequencing cycle and belonging to the active state bucket, and (ii) the previous intensity value for the pixel in the target channel detected in the previous sequencing cycle.

[0060] In some embodiments, the current intensity value is assigned to an active state bucket based on a comparison of the current intensity value to a global maximum and a global minimum determined from previous intensity values ​​in the target channel. In other embodiments, the current intensity value is assigned to an active state bucket based on a comparison of the current intensity value to previous active state values ​​and previous inactive state values ​​determined in a previous sequencing cycle in the target channel. In one embodiment, the current active state value is an exponentially weighted average determined from the current intensity value and the previous intensity value. In another embodiment, the current active state value is an average determined from the current intensity value and the previous intensity value. In yet another embodiment, the current active state value is a moving average determined from a rolling subset of the current intensity value and the previous intensity value. In yet another embodiment, the current active state value is a minimum value determined from the current intensity value and the previous intensity value. In yet another embodiment, the current active state value is a maximum value determined from the current intensity value and the previous intensity value.

[0061] In one embodiment, the current active state value is carried from a previous sequencing cycle and is not redetermined in the current sequencing cycle if the current intensity value belongs to the inactive state bucket in the current sequencing cycle.

[0062] In one embodiment, the current inactivity value of a pixel in the target channel is determined in the current sequencing cycle from (i) the current intensity value for the pixel in the target channel detected in the current sequencing cycle and attributed to the inactivity bucket, and (ii) the previous intensity value for the pixel in the target channel detected in the previous sequencing cycle.

[0063] In some embodiments, the current intensity value is assigned to the inactivity bucket based on a comparison of the current intensity value to a global maximum and a global minimum determined from previous intensity values ​​in the target channel, while in other embodiments, the current intensity value is assigned to the inactivity bucket based on a comparison of the current intensity value to a previous inactivity value and a previous inactivity value determined in a previous sequencing cycle in the target channel.

[0064] In some embodiments, when the current intensity value is closer to the global minimum than to the global maximum, the current intensity value is assigned to the inactive state bucket. In some embodiments, when the current intensity value is closer to the global maximum than to the global minimum, the current intensity value is assigned to the active state bucket.

[0065] In one embodiment, the current inactivity value is an exponentially weighted average determined from the current intensity value and the previous intensity value. In another embodiment, the current inactivity value is an average determined from the current intensity value and the previous intensity value. In yet another embodiment, the current inactivity value is a moving average determined from a rolling subset of the current intensity value and the previous intensity value. In yet another embodiment, the current inactivity value is a minimum value determined from the current intensity value and the previous intensity value. In yet another embodiment, the current inactivity value is a maximum value determined from the current intensity value and the previous intensity value.

[0066] In some embodiments, the current inactive state value is carried from a previous sequencing cycle and is not redetermined in the current sequencing cycle if the current intensity value belongs to an active state bucket in the current sequencing cycle.

[0067] In action 232, the base caller 144, which has access to memory, generates a base call for the current sequencing cycle in response to processing (i) each current intensity value and (ii) each current state value. In some embodiments, the base call for the current sequencing cycle includes a base call for one or more clusters in which each current signal value and each previous signal value is detected.

[0068] In one embodiment, base caller 144 is a neural network. In some embodiments, the neural network is a convolutional neural network. In one embodiment, the convolutional neural network includes multiple spatial convolutional layers and multiple temporal convolutional layers.

[0069] In some embodiments, base caller 144 is trained using sequencing images generated offline for previously performed sequencing runs. In one embodiment, each state value for each pixel is calculated offline for each sequencing cycle of a previously performed sequencing run prior to base calling by base caller 144. In one embodiment, base caller 144 is trained to use each state value to compensate for loss of base calling accuracy.

[0070] In some embodiments, each current state value is generated iteratively in each sequencing cycle of the sequencing run. In some embodiments, the memory is further configured to store a sequence of respective next intensity values ​​for each pixel for a next sequencing cycle of the sequencing run. In some embodiments, base caller 144 is further configured to generate a base call for the current sequencing cycle in response to processing (i) the respective current intensity value, (ii) each previous intensity value in the sequence of respective previous intensity values ​​for one or more of the previous sequencing cycles, (iii) each next intensity value in the sequence of respective next intensity values ​​for one or more of the next sequencing cycles, and (iv) the respective state value. In some embodiments, each state value is encoded on a pixel-by-pixel basis using the respective previous intensity value and the respective next intensity value.

[0071] Per-pixel state FIG. 3 illustrates one embodiment for generating pixel-by-pixel states 300 for pixels in a sequencing image that are processed as input for base calling. FIG. 3 illustrates exemplary pixel grids 302N, 302N-1, ... 302 of size 3 x 3. FIG. 3 also illustrates a pixel-by-pixel state grid 304. Pixel grids 302N, 302N-1, ... 302 represent the sequencing image that is provided as input to base caller 144 for base calling. Pixels in pixel grids 302N, 302N-1, ... 302 can have intensity values ​​for multiple image channels, e.g., red and blue channels. Pixel grid 302N is generated for a current sequencing cycle N of a sequencing run and shows intensity emissions generated as a result of nucleotide incorporation in a set of clusters in the current sequencing cycle N. Similarly, pixel grid 302N-1 is generated for a previous sequencing cycle N-1 of a sequencing run and shows intensity emissions generated as a result of nucleotide incorporation in a set of clusters in the previous sequencing cycle N-1. Similarly, pixel grid 302 is generated for the first sequencing cycle 1 of the sequencing run and shows the intensity emissions generated as a result of nucleotide incorporation into the set of clusters in the first sequencing cycle 1.

[0072] From a spatial perspective, pixel grids 302N, 302N-1, ... 302 can be thought of as sharing nine pixels 1 through 9. In FIG. 3, the nine pixels 1 through 9 are indexed by the same pixel-by-pixel superscript across pixel grids 302N, 302N-1, ... 302. From a temporal perspective, the nine pixels 1 through 9 can be thought of as representing a first set of intensity values ​​at the current time step N, a second set of intensity values ​​at the previous time step N-1, and a third set of intensity values ​​at the first time step 1. The intensity values ​​within the first, second, and third sets of intensity values ​​can be thought of as being arranged in a temporal order. In FIG. 3, the time-varying intensity values ​​are indexed by different pixel-by-pixel superscripts across pixel grids 302N, 302N-1, ... 302.

[0073] A current pixel state 318 for a given pixel may be calculated in the current sequencing cycle N based on the current intensity value 312 for the given pixel in the current sequencing cycle N and the previous intensity values ​​314, ... 316 for the given pixel in each previous sequencing cycle N-1 to 1. The previous intensity values ​​314, ... 316 for the given pixel may be thought of as being arranged in a previous intensity sequence starting from previous intensity position N-1 and ending with previous intensity position 1.

[0074] In one embodiment, the current pixel state 318 is the output of state generation logic (or a function) that takes as input the current intensity value 312 and the previous intensity values ​​314, ... 316. In one embodiment, the state generation logic selects the maximum value from among the current intensity value 312 and the previous intensity values ​​314, ... 316 and uses that maximum value as the current pixel state 318. In another embodiment, the state generation logic selects the minimum value from among the current intensity value 312 and the previous intensity values ​​314, ... 316 and uses that minimum value as the current pixel state 318. In yet another embodiment, the state generation logic determines the average (mean) of the current intensity value 312 and the previous intensity values ​​314, ... 316 and uses that average as the current pixel state 318. In yet another embodiment, the state generation logic determines an exponentially weighted average of the current intensity value 312 and the previous intensity values ​​314, ... 316 and uses the exponentially weighted average as the current pixel state 318. In yet another embodiment, the state generation logic may determine a running average of the current intensity value 312 and the previous intensity values ​​314, ... 316 and use the running average as the current pixel state 318. In yet another embodiment, the state generation logic may determine an exponential moving average of the current intensity value 312 and the previous intensity values ​​314, ... 316 and use the exponential moving average as the current pixel state 318. In yet another embodiment, the state generation logic may determine a standard deviation of the current intensity value 312 and the previous intensity values ​​314, ... 316 and use the standard deviation as the current pixel state 318. This process is performed for each of the nine pixels 1-9, such that a respective current pixel state is determined for each of the nine pixels 1-9. According to the example shown in FIG. 3 , the current pixel states generated by the state generation logic for the nine pixels 1-9 form the per-pixel state grid 304. In one embodiment, the state generation logic is performed by the state generator 126.

[0075] Pixel-wise state encoding 4 illustrates one embodiment for generating base calls based on per-pixel states 304. In one embodiment, encoding logic 402 combines 400 the pixel grid 302N with the per-pixel state grid 304, for each pixel, for example, by summation, concatenation, element-wise multiplication, or element-wise multiplication and summation (convolution). The combination of the pixel grid 302N and the per-pixel state grid 304 is processed by the base caller 144 to generate current base call data 154.

[0076] Status for each channel FIG. 5 illustrates one embodiment for generating per-channel states 500 for pixels in a sequencing image that are processed as input for base calling. FIG. 5 relates to an exemplary pixel P. In FIG. 5, pixel P has intensity values ​​across two exemplary channels (or intensity channels): channel 1 502 and channel 2 506. In other embodiments, pixel P may have fewer or more channels. In the example shown in FIG. 5, a first channel 502 has a first set of intensity values ​​512, 513, ... 516 for five sequencing cycles 1-5. Similarly, a second channel 506 has a second set of intensity values ​​522, 523, ... 526 for five sequencing cycles 1-5.

[0077] A first channel state 532 is then determined for the first channel 502 based on the first set of intensity values. Similarly, a second channel state 536 is determined for the second channel 506 based on the second set of intensity values. The first and second channel states 532 and 536 may be determined by the state generation logic by implementing, for example, a minimum selection function, a maximum selection function, an average function, an exponentially weighted average function, a moving average function, an exponential moving average function, or a standard deviation function.

[0078] At least one base call is then generated for the fifth sequencing cycle (cycle 5) based on the first channel state 532 and the second channel state 536. This, in one embodiment, includes combining the channel 1 intensity values ​​516 for cycle 5 with the first channel state 532 to generate a first combination for channel 1 502, and combining the channel 2 intensity values ​​526 for cycle 5 with the second channel state 536 to generate a second combination for channel 2 506. The base caller 144 then processes the first combination for channel 1 502 and the second combination for channel 2 506 to generate one or more base calls (e.g., for one or more clusters) for cycle 5.

[0079] In another embodiment, base caller 144 processes cycle 5 channel 1 intensity values ​​516, cycle 5 channel 2 intensity values ​​526, cycle 5 first channel state 532, and cycle 5 second channel state 536 as four separate input channels (e.g., RGB-style image channels) to generate base calls for cycle 5. Those skilled in the art will understand that other contemporary or future methods of combining data / channels and processing them as a combination are similarly applicable and may be used equally within the scope of this disclosure.

[0080] In yet another embodiment, the first and second channel states 532 and 536 are combined to generate a so-called "pan channel state," for example, by applying an averaging function to the first and second channel states 532 and 536. This pan channel state for cycle 5 is then processed along with the cycle 5 channel 1 intensity values ​​516 and the cycle 5 channel 2 intensity values ​​526 to generate the base call for cycle 5, i.e., a total of three separate input channels as opposed to four in the per-channel embodiment.

[0081] While Figure 5 relates to only one pixel P for purposes of illustration and simplicity, it will be understood that the state generation and state encoding-responsive base calling steps are performed on a pixel-by-pixel and channel-by-channel basis for multiple pixels in the input sequencing image, in part because cluster intensity profiles and their states are defined by groups of adjacent pixels and are therefore analyzed as a group to generate base calls for the cluster.

[0082] MIN state for each channel Figure 6 shows one embodiment for generating per-channel MIN (minimum) states 600 for pixels in a sequencing image that are processed as input for base calling. In Figure 6, column 601 indexes pixels of the input sequencing image that are processed by base caller 144 for base calling. The input sequencing image represents an intensity profile captured for a set of clusters in the current sequencing cycle N of a sequencing run. In Figure 6, column 602 represents the pixel-by-pixel intensity value in the first channel (image channel) for the pixel indexed in column 601 and the current sequencing cycle N. In Figure 6, column 603 represents the pixel-by-pixel intensity value in the second channel (image channel) for the pixel indexed in column 601 and the current sequencing cycle N.

[0083] In Figure 6, column 604 represents the minimum intensity value, pixel by pixel, in a first channel selected from the intensity values ​​observed in the first channel between the first sequencing cycle 1 and the current sequencing cycle N for the pixel indexed in column 601. In Figure 6, column 605 represents the minimum intensity value, pixel by pixel, in a second channel selected from the intensity values ​​observed in the second channel between the first sequencing cycle 1 and the current sequencing cycle N for the pixel indexed in column 601.

[0084] In one embodiment, base caller 144 processes columns 602, 603, 604, and 605 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0085] Pan channel MIN state Figure 7 shows one embodiment for generating pan channel MIN (minimum) states 700 for pixels in a sequencing image that are processed as input for base calling. In Figure 7, column 708 averages columns 604 and 605 from Figure 6 on a pixel-by-pixel basis. In one embodiment, base caller 144 processes columns 602, 603, and 708 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0086] MAX status for each channel Figure 8 shows one embodiment for generating per-channel MAX states 800 for pixels in a sequencing image that are processed as input for base calling. In Figure 8, column 804 represents the maximum intensity value per pixel in a first channel selected from the intensity values ​​observed in the first channel between the first sequencing cycle 1 and the current sequencing cycle N for the pixel indexed in column 601. In Figure 8, column 805 represents the maximum intensity value per pixel in a second channel selected from the intensity values ​​observed in the second channel between the first sequencing cycle 1 and the current sequencing cycle N for the pixel indexed in column 601.

[0087] In one embodiment, base caller 144 processes columns 602, 603, 804, and 805 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0088] Pan channel MAX state Figure 9 shows one embodiment for generating a pan channel MAX state 900 for pixels in a sequencing image that are processed as input for base calling. In Figure 9, column 908 averages columns 804 and 805 from Figure 8 on a pixel-by-pixel basis. In one embodiment, base caller 144 processes columns 602, 603, and 908 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0089] AVG status for each channel Figure 10 shows one embodiment for generating per-channel AVG (average) states 1000 for pixels in a sequencing image that are processed as input for base calling. In Figure 10, column 1004 represents a pixel-by-pixel average intensity value in a first channel calculated as the pixel-by-pixel average of each of the intensity values ​​observed in the first channel between the first sequencing cycle 1 and the current sequencing cycle N for the pixel indexed in column 601. In Figure 10, column 1005 represents a pixel-by-pixel average intensity value in a second channel calculated as the pixel-by-pixel average of each of the intensity values ​​observed in the second channel between the first sequencing cycle 1 and the current sequencing cycle N for the pixel indexed in column 601.

[0090] In one embodiment, base caller 144 processes columns 602, 603, 1004, and 1005 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0091] Pan Channel AVG Status Figure 11 shows one embodiment for generating a pan-channel AVG (average) state 1100 for pixels in a sequencing image that are processed as input for base calling. In Figure 11, column 1108 averages columns 1004 and 1005 from Figure 10 on a pixel-by-pixel basis. In one embodiment, base caller 144 processes columns 602, 603, and 1108 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0092] MIN MAX status for each channel 12 shows one embodiment for generating per-channel MIN / MAX (minimum and maximum) states 1200 for pixels in a sequencing image that are processed as input for base calling. In one embodiment, base caller 144 processes columns 602, 603, 604, 605, 804, and 805 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0093] MIN AVG status for each channel 13 shows one embodiment for generating per-channel MIN AVG (minimum and average) states 1300 for pixels in a sequencing image that are processed as input for base calling. In one embodiment, base caller 144 processes columns 602, 603, 604, 605, 1004, and 1005 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0094] MAX AVG status for each channel 14 shows one embodiment for generating per-channel MAX AVG (maximum and average) states 1400 for pixels in a sequencing image that are processed as input for base calling. In one embodiment, base caller 144 processes columns 602, 603, 804, 805, 1004, and 1005 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0095] MIN MAX AVG status for each channel 15 shows one embodiment for generating per-channel MIN MAX AVG (minimum, maximum, and average) states 1500 for pixels in a sequencing image that are processed as input for base calling. In one embodiment, base caller 144 processes columns 602, 603, 604, 605, 804, 805, 1004, and 1005 as separate input channels to generate one or more base calls for the current sequencing cycle N.

[0096] Those skilled in the art will appreciate that any combination or sequencing or arrangement of channels and states discussed above can be provided as input to base caller 144 for base calling. Other current and future methods of combining, sequencing, and arranging the channels and states described above are within the scope of this disclosure. For example, pan-channel states can be generated by combining per-channel states using other aggregation / accumulation functions, such as exponentially weighted averages, moving averages, standard deviations, variances, etc. In another example, per-channel MIN states can be concatenated with pan-channel AVG MAX states, etc.

[0097] Base-call-directed state generation Figure 16 shows one embodiment for generating 1600 state information dependent on previously called bases for use in future base calls. Figures 17A, 17B, and 17C show base-call-directed ON and OFF state generation 1700 for the current sequencing cycle 11 (cycle 11). Figures 18A, 18B, and 18C expand on Figures 17A, 17B, and 17C to show how 1800 the OFF state is updated in the next sequencing cycle 12 (cycle 12). Figures 19A, 19B, and 19C further expand on Figures 18A, 18B, and 18C to show how 1900 the ON state is updated in the next sequencing cycle 13 (cycle 13). Figures 20A, 20B, and 20C further expand on Figures 19A, 19B, and 19C and show how both the ON and OFF states are updated (2000) in the next sequencing cycle 14 (cycle 14).

[0098] One means of distinguishing between different strategies for detecting nucleotide incorporation in sequencing reactions using a single fluorescent dye (or two or more dyes with the same or similar excitation / emission spectra) is by characterizing the incorporation in terms of the presence or relative absence of, or levels between, fluorescence transitions that occur during a sequencing cycle. Thus, sequencing strategies can be illustrated by their fluorescence profile over a sequencing cycle. For the strategies disclosed herein, a "1" or a "0" indicates a fluorescence state (1 / ON) in which the nucleotide is in a signal state (e.g., detectable by fluorescence) or a fluorescence state (0 / OFF) in which the nucleotide is in a dark state (e.g., not detected or minimally detected in an imaging step). A "0," "OFF," or "dark" state does not necessarily refer to a complete lack or absence of signal. In some embodiments, there may be a total lack or absence of signal (e.g., fluorescence), while in other embodiments, there may be some detectable signal even in the OFF state. Minimal or reduced fluorescence signals (e.g., background signals) are also considered to be included within the "0," "OFF," or "dark" state range, as long as the change in fluorescence from the first image to the second image (or vice versa) can be reliably distinguished.

[0099] As used herein, the terms "dark" or "OFF" are intended to refer to an amount of desired signal detected by a detector that is insignificant compared to the background signal detected by the detector. For example, an object feature may be considered dark or OFF when the signal-to-noise ratio of that feature is substantially low, e.g., less than 1. In some embodiments, a dark or OFF feature may not produce any amount of desired signal (i.e., no signal is produced or detected). In other embodiments, a very low amount of signal relative to the background may be considered dark or OFF.

[0100] In one embodiment, an exemplary strategy for detecting and determining nucleotide incorporation in a sequencing reaction using one fluorescent dye (or two dyes of the same or similar excitation / emission spectra) and two imaging events is illustrated by the detection table below.

[0101] [Table 1]

[0102] Other strategies for detecting nucleotide incorporation (eg, four-channel chemistry and one-channel chemistry) are within the scope of this disclosure and will not be discussed separately.

[0103] In action 1602, the disclosed technology accesses, for a given pixel, the underlying channel-wise active (ON) and inactive (OFF) classifications of bases called in previous sequencing cycles, i.e., previously called bases. This is illustrated by the example in FIG. 17A. In FIG. 17A, pixel P shows the intensity emission of the corresponding cluster to which the base is called for cycles 1 through 10. In some embodiments, pixel P includes the center of the corresponding cluster, as determined by the location coordinates of the corresponding cluster center. The called base results from channel-wise active (ON) and inactive (OFF) classifications 1754 and 1756 for the intensity values ​​in the first channel (channel 1) and the intensity values ​​in the second channel (channel 2).

[0104] In action 1612, the disclosed technique accumulates channel-wise summary statistics for the active (ON) and inactive (OFF) states for each channel based on the per-channel active (ON) and inactive (OFF) classifications 1754 and 1756 for the given pixel and current sequencing cycle. This is shown by example in FIG. 17B, where the inactive (OFF) state of the first channel 1702 is determined for sequencing cycle 11 (cycle 11) by applying an accumulation function to the intensity values ​​in column 1754 classified as inactive (OFF). Examples of accumulation functions include maximum selection, minimum selection, average (mean), exponentially weighted average, running average, exponential moving average, mode, standard deviation, variance, skewness, kurtosis, percentile, and entropy. In Figure 17B, the active (ON) state of the first channel 1712 is determined for cycle 11 by applying a cumulative function to the intensity values ​​in column 1754 that are categorized as an active (ON) state. In Figure 17B, the inactive (OFF) state of the second channel 1722 is determined for cycle 11 by applying a cumulative function to the intensity values ​​in column 1756 that are categorized as an inactive (OFF) state. In Figure 17B, the active (ON) state of the second channel 1732 is determined for cycle 11 by applying a cumulative function to the intensity values ​​in column 1756 that are categorized as an active (ON) state.

[0105] In action 1622, the disclosed technology combines the accumulated channel-wise summary statistics for the active (ON) and inactive (OFF) states per channel for a given pixel and current sequencing cycle with the current intensity channel for the given pixel. This is shown by example in FIG. 17C, where 1764 and 1766 are the intensity values ​​for the first and second channels, respectively, registered for pixel P in cycle 11. In FIG. 17C, base caller 144 processes 1764, 1766, 1702, 1712, 1722, and 1732 as separate input channels to generate a base call for the corresponding cluster of pixel P in sequencing cycle 11.

[0106] In action 1632, the disclosed technology generates at least one base call based on a combination of the active (ON) and inactive (OFF) states per channel of a given pixel and accumulated channel-wise summary statistics for the intensity channel for the current sequencing cycle. This is shown by example in Figure 18A. In Figure 18A, base call "G" is generated for the corresponding cluster of pixel P in cycle 11 based on "OFF" and "OFF" channel-wise classifications 1854 and 1856, which are made by base caller 144 in response to operations 1764, 1766, 1702, 1712, 1722, and 1732 in Figure 18C.

[0107] In FIG. 18B , in response to the “OFF” and “OFF” channel-wise classifications 1854 and 1856 (and the absence of an “ON” classification) in cycle 11, only the OFF channel states of the first and second channels 1802 and 1822 are updated in cycle 12, so that the accumulation function is reapplied to also consider the channel-wise intensity values ​​in cycle 11 that both have “OFF” classifications 1854 and 1856. In some embodiments, the channel-wise ON and OFF states are maintained in memory as a single variable and are back-calculated for each update instance, for example, by multiplying by the divisor used to calculate the average when the accumulation function was an averaging function. The ON channel states for the first and second channels 1812 and 1832 remain unchanged in cycle 12 and inherit the same values ​​from cycle 11, i.e., 1712 and 1732, respectively.

[0108] In Figure 18C, 1864 and 1866 are the intensity values ​​for the first and second channels, respectively, registered for pixel P in cycle 12. In Figure 18C, base caller 144 processes 1864, 1866, 1802, 1812, 1822, and 1832 as separate input channels to generate base calls for the corresponding cluster of pixel P in sequencing cycle 12.

[0109] In Figure 19A, a base call "T" is generated for the corresponding cluster of pixel P in cycle 12 based on "ON" and "ON" channel-wise classifications 1954 and 1956, which are made by base caller 144 in response to processes 1864, 1866, 1802, 1812, 1822, and 1832 in Figure 18C.

[0110] 19B, in response to the "ON" and "ON" channel-wise classifications 1954 and 1956 in cycle 12 (and the absence of an "OFF" classification), only the ON channel states for the first and second channels 1912 and 1932 are updated in cycle 13, so that the accumulation function is reapplied to also take into account the channel-wise intensity values ​​in cycle 12, both of which have "ON" classifications 1954 and 1956. The OFF channel states of the first and second channels 1902 and 1922 remain unchanged in cycle 13, inheriting the same values ​​from cycle 12, i.e., 1802 and 1822, respectively.

[0111] In Figure 19C, 1964 and 1966 are the intensity values ​​for the first and second channels, respectively, registered for pixel P in cycle 13. In Figure 19C, base caller 144 processes 1964, 1966, 1902, 1912, 1922, and 1932 as separate input channels to generate base calls for the corresponding cluster of pixel P in sequencing cycle 13.

[0112] In Figure 20A, base call "A" is generated for the corresponding cluster of pixel P in cycle 13 based on "ON" and "OF" channel-wise classifications 2054 and 2056, which are made by base caller 144 in response to operations 1964, 1966, 1902, 1912, 1922, and 1932 in Figure 19C.

[0113] In FIG. 20B, in response to the "ON" and "OFF" channel-by-channel classifications 2054 and 2056 in cycle 13, both the ON and OFF channel states of the first and second channels 2002, 2012, 2022, and 2032, respectively, are updated in cycle 14, so that the accumulation function is reapplied to also take into account the channel-by-channel intensity values ​​of cycle 13 having the "ON" and "OFF" classifications 2054 and 2056.

[0114] In Figure 20C, 2064 and 2066 are the intensity values ​​for the first and second channels, respectively, registered for pixel P in cycle 14. In Figure 20C, base caller 144 processes 2064, 2066, 2002, 2012, 2022, and 2032 as separate input channels to generate base calls for the corresponding cluster of pixel P in sequencing cycle 14.

[0115] Note that in base-call-directed state generation, state information is delayed by one sequencing cycle. That is, state information 1802, 1812, 1822, and 1832 used to base call cycle 12 is based on cycles 1 through 11 and does not include cycle 12. Similarly, state information 1902, 1912, 1922, and 1932 used to base call cycle 13 is based on cycles 1 through 12 and does not include cycle 13.

[0116] Exponentially weighted averaging Figure 21 shows one embodiment of generating 2100 state information for a base call using exponentially weighted averaging. In one embodiment, exponentially weighted averaging is used to give greater weight to input signals (e.g., intensity values) from more recent sequencing cycles when calculating the state information. As a result, state information estimates generated by exponentially weighted averaging track signal changes and are more robust to outliers.

[0117] At action 2102, the disclosed technique accumulates summary statistics for past intensity values ​​for a given pixel from previous sequencing cycles (e.g., the first 10 or 20 sequencing cycles). Examples of accumulation functions for accumulating summary statistics include max-pick, min-pick, average (mean), exponentially weighted average, running average, exponential moving average, mode, standard deviation, variance, skewness, kurtosis, percentile, and entropy.

[0118] In action 2112, the disclosed technique initializes a starting exponential weighted average for a given pixel based on the accumulated summary statistics. In one implementation, the starting exponential weighted average can be an average of past intensity values ​​(e.g., an average of past intensity values ​​of all other pixels or all other cluster pixels), a global maximum (global MAX) selected from past intensity values ​​(e.g., a global MAX of past intensity values ​​of all other pixels or all other cluster pixels), or a global minimum (global MIN) selected from past intensity values ​​(e.g., a global MIN of past intensity values ​​of all other pixels or all other cluster pixels).

[0119] In action 2122, the disclosed technique determines a current exponentially weighted average for the given pixel and the current sequencing cycle based on a weighted combination of the starting exponentially weighted average and the current pixel intensity value.

[0120] In action 2132, the disclosed techniques base call the cluster corresponding to the given pixel, e.g., use the current exponentially weighted average as the current state value for base calling in the current sequencing cycle. In one embodiment, this involves base caller 144 processing the current state value and the current pixel intensity value to generate at least one base call for the current sequencing cycle.

[0121] In action 2142, the disclosed technique determines a next exponentially weighted average for the given pixel and the next sequencing cycle based on a weighted combination of the current exponentially weighted average and the next pixel intensity value.

[0122] In action 2152, the disclosed technology uses the next exponentially weighted average as the next state value for a base call in the next sequencing cycle (e.g., a base call for a cluster corresponding to a given pixel). In one embodiment, this involves base caller 144 processing the next state value and the next pixel intensity value to generate at least one base call for the next sequencing cycle.

[0123] In one embodiment, the exponential weighted average logic is expressed as follows: y[k]=(1-alpha)*y[k-1]+alpha*x[k] During the ceremony, y[k] is the exponentially weighted average of the current sequencing cycle, y[k-1] is the exponentially weighted average of the previous sequencing cycles, x[k] is the input signal (e.g., intensity value) for the current sequencing cycle; Alpha is a weighting parameter,

[0124] The weighting parameter "alpha" can be a value between 0 and 1 {0≦alpha≦1}. When alpha is equal to 0 {alpha=0} the output is y[k-1] (no averaging). When alpha is equal to 1 {alpha=1} the output is x[k]. For other alpha values ​​the output is an exponentially weighted average of the intensities. Alpha represents how quickly the filter reacts to updates. Smaller values ​​react slower and do more averaging, while larger values ​​place more weight on more recent input values.

[0125] Note that y[.] is not used except when updating the above formula, and can therefore be stored as a single variable and updated in place. y[k-1] can be initialized, for example, to the expected intensity (e.g., the average initial intensity across all clusters / wells). In one embodiment, the first 20 sequencing cycles (cycles 1-20) can be used to estimate the initial average intensity without online estimation. The initial average intensity can be plugged in as an initial estimate for y[k-1], and then online exponentially weighted averaging can be used after sequencing cycle 20 (cycle 20).

[0126] When implemented on a CPU, the above formula can be efficiently implemented with a single multiplication by rewriting it as follows: y[k]=y[k-1]*alpha(x[k]-y[k-1])

[0127] In some implementations, alpha can be a small value close to 0, thus performing more averaging. In other implementations, alpha can be varied to find a value that trades off averaging v / s response time that best suits a particular application.

[0128] Exponentially weighted averaging per channel Figure 22 shows one embodiment of using exponentially weighted averaging to generate per-channel state information for base calls 2200. Figure 23 shows an example of using exponentially weighted averaging to generate per-channel state information for base calls 2300.

[0129] In action 2202, the disclosed technique accumulates, for a given pixel, channel-wise summary statistics for past intensity values ​​per channel from previous sequencing cycles, e.g., the first 10 or 20 sequencing cycles. Examples of channel-wise accumulation functions for accumulating channel-wise summary statistics include maximum / global maximum selection, minimum / global minimum selection, average (mean), exponentially weighted average, running average, exponential moving average, mode, standard deviation, variance, skewness, kurtosis, percentile, and entropy. This is illustrated by example in FIG. 23.

[0130] In Figure 23, a first channel (channel 1) has past intensity values ​​2374 for the first five sequencing cycles (cycles 1-5). In Figure 23, a second channel (channel 2) has past intensity values ​​2376 for cycles 1-5. Additionally, a global maximum (GLOBAL MAX) value 2312 is selected for channel 1 from the past intensity values ​​2374, and a global minimum (GLOBAL MIN) value 2314 is selected for channel 1 from the past intensity values ​​2374. Similarly, a global MAX value 2316 is selected for channel 2 from the past intensity values ​​2376, and a global MIN value 2318 is selected for channel 2 from the past intensity values ​​2376.

[0131] In operation 2212, the disclosed technique initializes, for each channel, a pair of starting active (ON) and inactive (OFF) states based on the accumulated per-channel summary statistics. This is shown by example in FIG. 23. In FIG. 23, after channel 1 and cycle 5, global MAX 2312 is initialized as the active (ON) state of channel 1, and global MIN 2314 is initialized as the inactive (OFF) state of channel 1. In FIG. 23, after channel 2 and cycle 5, global MAX 2316 is initialized as the active (ON) state of channel 2, and global MIN 2318 is initialized as the inactive (OFF) state of channel 2.

[0132] In operation 2222, the disclosed technique, for a given pixel and current sequencing cycle, on a channel-by-channel basis, attributes a current channel intensity value to a starting active (ON) state or a starting inactive (OFF) state based on a comparison to a pair of starting active (ON) and inactive (OFF) states. This is illustrated by the example of FIG. 23. In FIG. 23, for channel 1 and cycle 6, intensity value 2322 is registered. Because intensity value 2322 is quantitatively closer to global MAX 2312, global MAX 2312 serves as a proxy for the active (ON) state of channel 1 and is used as the starting exponentially weighted average of channel 1 in the active (ON) state, so intensity value 2322 is attributed to the active (ON) state. In FIG. 23, for channel 2 and cycle 6, intensity value 2326 is registered.

[0133] Since intensity value 2326 is quantitatively closer to global MIN 2318, global MIN 2318 serves as a proxy for the inactive (OFF) state of channel 2 and is used as the starting exponentially weighted average of channel 2 in the inactive (OFF) state, so intensity value 2326 is attributed to the inactive (OFF) state.

[0134] In action 2232, the disclosed technique updates the attribute state for a given pixel, within a given channel, based on a weighted combination of the current channel intensity value and the previous state value of the attribute state using an exponentially weighted average.

[0135] In action 2242, the disclosed technique maintains, for a given pixel, in a given channel, an unattributed state from a previous state value of the unattributed state.

[0136] Actions 2232 and 2242 are shown by way of example in Figure 23. In Figure 23, for cycle 6, state information is determined for channels 1 and 2 as follows: The active (ON) state 2342 for cycle 6 and channel 1 is determined using an exponentially weighted average 2332 because the intensity value 2322 for cycle 6 and channel 1 is attributed to the active (ON) state of channel 1. The exponentially weighted average 2332 uses an alpha value of "0.12" as described with respect to the equation above. The inactive (OFF) state 2344 for channel 1 is kept unchanged and is inherited from cycle 5 as the global MIN 2314 for channel 1.

[0137] The inactive (OFF) state 2348 of cycle 6 and channel 2 is determined using the exponentially weighted average 2338 because the intensity value 2326 of cycle 6 and channel 2 is attributed to the inactive (OFF) state of channel 2. The exponentially weighted average 2338 also uses an alpha value of "0.12" as described with respect to the equation above. The active (ON) state 2346 of channel 2 is left unchanged and inherited from cycle 5 as the global MAX 2316 for channel 2.

[0138] Next, the channel-by-channel active (ON) and inactive (OFF) states 2342, 2344, 2346, and 2348 for cycle 6 are used along with the intensity values ​​2322 and 2326 registered for cycle 6 to generate at least one base call for cycle 6.

[0139] 23, for cycle 7, state information is determined for channels 1 and 2 as follows: The inactive (OFF) state 2374 for cycle 7 and channel 1 is determined using the exponentially weighted average 2364 because the intensity value 2352 for cycle 7 and channel 1 is attributed to the inactive (OFF) state of channel 1. The exponentially weighted average 2364 also uses an alpha value of "0.12" as described with respect to the equation above. The active (ON) state 2372 for channel 1 is kept unchanged and is inherited from cycle 6 as the exponentially weighted average 2342 for channel 1.

[0140] The active (ON) state 2376 of cycle 7 and channel 2 is determined using the exponentially weighted average 2366 because the intensity value 2356 of cycle 7 and channel 2 is attributed to the active (ON) state of channel 2. The exponentially weighted average 2366 also uses an alpha value of "0.12" as described with respect to the formula above. The inactive (OFF) state 2378 of channel 2 is left unchanged and inherited from cycle 6 as the exponentially weighted average 2348 of channel 2.

[0141] Next, the channel-by-channel active (ON) and inactive (OFF) states 2372, 2374, 2376, and 2378 for cycle 7 are used along with the intensity values ​​2352 and 2356 registered for cycle 7 to generate at least one base call for cycle 7.

[0142] Note that in Figure 23, the channel-wise attribution to active (ON) and inactive (OFF) states in a given sequencing cycle is based on the intensity values ​​per channel registered for the given sequencing cycle (e.g., cycles 6 and 7 in Figure 23). We now describe base-call-directed channel-wise exponentially weighted averaging, in which the channel-wise attribution to active (ON) and inactive (OFF) states in a given sequencing cycle is based on the base call made in the previous sequencing cycle.

[0143] Base-call-directed exponentially weighted averaging per channel Figure 24 shows one embodiment of generating 2400 per-channel state information for base calling using exponentially weighted averaging directed by previously called bases. Figure 25 depicts an example of generating 2500 per-channel state information for base calling using exponentially weighted averaging directed by previously called bases.

[0144] In Figure 24, actions 2402 and 2412 are similar to actions 2202 and 2212 in Figure 22. In action 2422, per-channel attribution to active (ON) and inactive (OFF) states is made based on bases called in the previous sequencing cycle, delayed by one sequencing cycle. This is shown by example in Figure 25, where in cycle 6, base call "A" 2502 is made based on the "ON" classification 2504 and "OFF" classification 2506 of channels 1 and 2, respectively.

[0145] In action 2432, the "ON" and "OFF" classifications 2504 and 2506 and base call 2502 in cycle 6 are used to determine state information 2342, 2344, 2346, and 2348 for cycle 7 by applying exponentially weighted averaging. In action 2442, the unassigned state for cycle 7 is also identified based on the base call 2502 in cycle 6 and is kept unchanged.

[0146] Note that in cycle 7, based on the state information 2342, 2344, 2346, and 2348 determined in cycle 6 and the intensity values ​​per channel 2372 and 2376 registered in cycle 7, "OFF" and "ON" classifications 2514 and 2516 and then the base call "C" 2512 are generated.

[0147] The "OFF" and "ON" classifications 2514 and 2516 and base call 2512 in cycle 7 are then used to determine state information 2372, 2374, 2376, and 2378 for cycle 8 by applying exponentially weighted averaging. The unassigned state in cycle 8 is also identified based on the base call 2512 in cycle 7 and remains unchanged.

[0148] Although not shown, a base call for cycle 8 is generated based on the channel-by-channel intensity values ​​registered for cycle 8 and the state information 2372, 2374, 2376, and 2378 determined in cycle 7 for cycle 8.

[0149] In different embodiments, state information determined using exponentially weighted averaging (EWA) may be used instead of, in addition to, or in combination with state information determined using some other logic, such as minimum selection logic, maximum selection logic, averaging logic, etc.

[0150] Non-clustered pixels Next, we distinguish between cluster pixels and non-cluster pixels. Cluster pixels are pixels that contain the center of a cluster, as determined by the location coordinates of the cluster center. Non-cluster pixels do not contain the cluster center. Note that non-cluster pixels exhibit cluster intensity (or background intensity). They do not coincide with a cluster center.

[0151] A brief description of the cluster-pixel-base call relationship is also useful here. A cluster is represented by a plurality of pixels, e.g., a pixel grid of 3x3 pixels. In some embodiments, base caller 144 processes a sequencing image having different intensity values ​​for pixels within a given pixel grid to characterize the respective intensity profiles of a given cluster in different sequencing cycles. In response to the processing, in one embodiment, base caller 144 generates an output specifying the respective base calls for a given cluster in different sequencing cycles by referencing only the central pixel within the given pixel grid that contains the center of the given cluster. That is, even though the entire given pixel grid characterizes the intensity profile of a given cluster, each base call is made with respect to only the central pixel. Non-central pixels of a given pixel grid are analyzed by base caller 144 to generate base calls. However, only the central pixel is used to represent the base calls.

[0152] In implementations where only the central pixel is used to represent the base call, base call-directed state generation can be difficult because the base call is not available to determine the state of non-cluster pixels. This limitation can be compensated for by the techniques described below.

[0153] Figure 26 shows one embodiment for generating state information for non-cluster pixels using previously called bases 2600. For pixels that contain cluster centers, i.e., cluster pixels 2602, state data is generated 2612 using the previous base calls of the cluster, e.g., as described above with respect to Figures 16-20C and 24-25.

[0154] For pixels that do not contain a cluster center, i.e., non-cluster pixel 2608, according to one embodiment, the state data for the nearest cluster pixel is used to generate state data 2618. In another embodiment, the average of the state data for all other cluster pixels is used to generate state data for the non-cluster pixel 2608 2628.

[0155] Status Input 27 illustrates different embodiments for generating state data 2700 for base calling. The disclosed technology generates state information 2722 for an "entity" 2702 based on its current configuration i and past configurations i-1, i-2, ... 1. Examples of entity 2702 include any type of signal measurement, pixel measurement, voltage measurement, current measurement, pH scale measurement, intermediate output of processing, convolution features (e.g., feature maps), compressed features (e.g., compressed feature maps), different types of output (e.g., softmax scores, sigmoid scores, regression scores), base calls, and reads. A current base call is then made depending on the state information 2722 and the current configuration i of entity 2702.

[0156] Note that in order to combine state information 2722 with entity 2702, in some implementations, the dimensionality of state information 2722 may need to be adjusted with the dimensionality of entity 2702 (e.g., made compatible by matching), or vice versa. In different implementations, this may be achieved by dimensionality-changing operations such as cloning, padding (e.g., zero-padding), concatenation, convolution, sum, transpose convolution, etc. For example, if state information 2722 has a dimensionality of 1×1 and entity 2702 has a dimensionality of 3×3, then nine clones of state information 2722 may be concatenated with entity 2702.

[0157] Supply status input FIG. 28 illustrates different implementations of provided state data 2800 in various aspects of the processing of a base calling operation. In one implementation, base caller 144 can have various processing modules 1-n (e.g., pre-processing layer, neural network layer, post-processing layer, output layer). State data 2802 can be provided to any of processing modules 1-n of base caller 144 to generate base calls 2808. For example, state data 2802 can be combined with input data 2802 (e.g., input image) for processing by a first processing module (e.g., a first convolutional layer) of base caller 144. In another example, state data 2802 can be combined with intermediate output 2812 of base caller 144 (e.g., intermediate feature maps generated by a preceding convolutional layer). In yet another example (not shown), state data 2802 can be combined with the final output of base caller 144 (e.g., a final feature map generated by a final convolutional layer). In such an implementation, the combination of state data 2802 and final output may be processed by an output / base caller layer (e.g., a softmax layer, a regression layer, a sigmoid layer) to generate base calls 2808.

[0158] State-based and neural network-based base callers The following discussion focuses on the neural network-based base caller 2900 described herein. The neural network-based base caller 2900 is one implementation of the base caller 144 and is collectively referred to herein as "DeepRTA." First, the inputs to the neural network-based base caller 2900 are described according to one implementation. Then, an example of the structure and form of the neural network-based base caller 2900 is provided. Finally, the output of the neural network-based base caller 2900 is described according to one implementation.

[0159] The dataflow logic provides the sequencing image to a neural network-based base caller 2900 for base calling. The neural network-based base caller 2900 accesses the sequencing image patch by patch (or tile by tile). Each patch is a subgrid (or subarray) of pixelated units within the grid of pixelated units that forms the sequencing image. A patch has dimensions q by r of the subgrid of pixelated units, where q (width) and r (height) are any number ranging from 1 to 10,000 (e.g., 3x3, 5x5, 7x7, 10x10, 15x15, 25x25, 64x64, 78x78, 115x115). In some embodiments, q and r are the same. In other embodiments, q and r are different from each other. In some embodiments, patches extracted from a single sequencing image are the same size. In other embodiments, patches are of different sizes. In some implementations, patches can have overlapping pixelated units (eg, on edges).

[0160] Sequencing generates m sequencing images per sequencing cycle for the corresponding m image channels. That is, each sequencing image has one or more image (or intensity) channels (similar to the red, green, and blue (RGB) channels of a color image). In one embodiment, each image channel corresponds to one of multiple filter wavelength bands. In another embodiment, each image channel corresponds to one of multiple imaging events in a sequencing cycle. In yet another embodiment, each image channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter. Image patches are tiled (or accessed) from each of the m image channels for a particular sequencing cycle. In different embodiments, such as 4-, 2-, and 1-channel chemistries, m is 4 or 2. In other embodiments, m is greater than 1, 3, or 4. In other embodiments, images can be in blue and violet channels instead of or in addition to red and green channels.

[0161] For example, consider a sequencing run performed using two different imaging channels, i.e., a blue channel and a green channel. Then, in each sequencing cycle, the sequencing run generates a blue image and a green image. In this manner, for a series of k sequencing cycles of the sequencing run, a sequence of k pairs of blue and green images is generated as output and stored as the sequencing image. Thus, a sequence of k pairs of blue and green image patches is generated for patch-level processing by the neural network-based base caller 2900.

[0162] The input image data to the neural network-based base caller 2900 for one base calling iteration (or one instance of a forward pass or single forward traversal) includes data for a sliding window that includes multiple sequencing cycles. The sliding window can include, for example, the current sequencing cycle, one or more preceding sequencing cycles, and one or more subsequent sequencing cycles.

[0163] In one embodiment, the input image data includes data from three sequencing cycles, such that data from the current (time t) sequencing cycle being base called is accompanied by data from (i) the left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, and (ii) the right adjacent / context / next / successive / subsequent (time t+1) sequencing cycle.

[0164] In another embodiment, the input image data includes data from five sequencing cycles, and the data for the current (time t) sequencing cycle being base called is accompanied by (i) data from the first left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, (ii) data from the second left adjacent / context / previous / preceding / previous (time t-2) sequencing cycle, (iii) data from the first right adjacent / context / next / successive / subsequent (time t+1) sequencing cycle, and (iv) data from the second right adjacent / context / next / successive / subsequent (time t+2) sequencing cycle.

[0165] In yet another embodiment, the input image data includes data from seven sequencing cycles, and the data for the current (time t) sequencing cycle to be base called includes (i) data from the first left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, (ii) data from the second left adjacent / context / previous / preceding / previous (time t-2) sequencing cycle, (iii) data from the third left adjacent / context / previous / preceding / previous (time t-3) sequencing cycle, (iv) data from the first right adjacent / context / next / contiguous / subsequent (time t+1) sequencing cycle, (v) data from the second right adjacent / context / next / contiguous / subsequent (time t+2) sequencing cycle, and (vi) data from the third right adjacent / context / next / contiguous / subsequent (time t+3) sequencing cycle. In other embodiments, the input image data includes data from a single sequencing cycle. In still other embodiments, the input image data includes data from 10, 15, 20, 30, 58, 75, 92, 130, 168, 175, 209, 225, 230, 275, 318, 325, 330, 525, or 625 sequencing cycles.

[0166] According to one embodiment, the neural network-based base caller 2900 processes image patches through its convolutional layers to generate alternative representations. The alternative representations are then used by an output layer (e.g., a softmax layer) to generate base calls for either the current sequencing cycle (time t) or each of the sequencing cycles (i.e., the current sequencing cycle (time t), the first and second preceding sequencing cycles (time t-1, time t-2), and the first and second subsequent sequencing cycles (time t+1, time t+2)). The resulting base calls form sequencing reads.

[0167] In one embodiment, neural network-based base caller 2900 outputs a base call for a single target cluster for a particular sequencing cycle. In another embodiment, neural network-based base caller 2900 outputs a base call for each target cluster within a plurality of target clusters at a particular sequencing cycle. In yet another embodiment, neural network-based base caller 2900 outputs a base call for each target cluster within a plurality of target clusters at each sequencing cycle within the plurality of sequencing cycles, thereby generating a base call sequence for each target cluster.

[0168] In one embodiment, neural network-based base caller 2900 is a multilayer perceptron (MLP). In another embodiment, neural network-based base caller 2900 is a feed-forward neural network. In yet another embodiment, neural network-based base caller 2900 is a fully connected neural network. In a further embodiment, neural network-based base caller 2900 is a fully convolutional neural network. In yet another embodiment, neural network-based base caller 2900 is a semantic segmentation neural network. In yet another further embodiment, neural network-based base caller 2900 is a generative adversarial network (GAN). In yet another embodiment, neural network-based base caller 2900 includes a multi-head attention mechanism such as Transformer, BERT, and DETR.

[0169] In one embodiment, neural network-based base caller 2900 is a convolutional neural network (CNN) having multiple convolutional layers. In another embodiment, neural network-based base caller 2900 is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bi-directional LSTM (Bi-LSTM), or a gated recurrent unit (GRU). In yet another embodiment, neural network-based base caller 2900 includes both a CNN and an RNN.

[0170] In yet other embodiments, the neural network-based base caller 2900 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or dilated convolution, transposed convolution, depth-wise separable convolution, point-wise convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, Transformer, BERT, spatially separable convolution, and deconvolution. The neural network-based base caller 2900 can use one or more loss functions, such as logistic regression / logarithmic loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. The neural network-based base caller 2900 can use any parallel, efficient, and compression scheme, such as TFRecord, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). The neural network-based base caller 2900 can include nonlinear transformation functions, such as upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (e.g., LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions (rectifying linear unit (ReLU), leaky ReLU, exponential liner unit (ELU), sigmoid, and hyperbolic tangent (tanh)), etc.), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or mean pooling), global mean pooling layers, and attention mechanisms.

[0171] The neural network-based base caller 2900 trains using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that the neural network-based base caller 2900 may use to train include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that the neural network-based base caller 2900 may use to train include Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.

[0172] In one embodiment, the neural network-based base caller 2900 uses a dedicated architecture to separate the processing of data for different sequencing cycles. The motivation for using such a dedicated architecture is first explained. As described above, the neural network-based base caller 2900 processes image patches for the current sequencing cycle, one or more previous sequencing cycles, and one or more subsequent sequencing cycles. Data from additional sequencing cycles provides sequence-specific context. The neural network-based base caller 2900 learns a unique context for each sequence during training and calls them as bases. Furthermore, data from pre- and post-sequencing cycles provide secondary contributions of pre-phasing and phasing signals to the current sequencing cycle.

[0173] However, images captured in different sequencing cycles and in different image channels are misaligned and have residual registration errors with each other. To account for this misalignment, the specialized architecture includes a spatial convolution layer that does not mix information between sequencing cycles, but only mixes information within the same sequencing cycle.

[0174] The spatial convolutional layer (or spatial logic) uses so-called "decoupled convolutions" that operate on decoupling by processing the data for each of multiple sequencing cycles independently through a "dedicated, unshared" sequence of convolutions. Decoupled convolutions convolve on the data and resulting feature maps only within a given sequencing cycle, i.e., the cycle, without convolving on the data and resulting feature maps of any other sequencing cycles.

[0175] For example, consider input image data including (i) a current image patch for the current (time t) sequencing cycle to be base-called, (ii) a previous image patch for the previous (time t-1) sequencing cycle, and (iii) a next image patch for the next (time t+1) sequencing cycle. The dedicated architecture then initiates three separate convolution pipelines: a current convolution pipeline, a previous convolution pipeline, and a next convolution pipeline. The current data processing pipeline receives the current image patch for the current (time t) sequencing cycle as input and processes it independently through multiple spatial convolution layers to generate a so-called "current spatial convolution representation" as the output of the final spatial convolution layer. The previous convolution pipeline receives the previous image patch for the previous (time t-1) sequencing cycle as input and processes it independently through multiple spatial convolution layers to generate a so-called "previous spatial convolution representation" as the output of the final spatial convolution layer. The next convolution pipeline receives the next data for the next (time t+1) sequencing cycle as input and processes it independently through multiple spatial convolution layers to produce the so-called “next spatial convolution representation” as the output of the final spatial convolution layer.

[0176] In some implementations, the current, previous, and next convolution pipelines run in parallel. In some implementations, the spatial convolution layer is part of a spatial convolution network (or sub-network) within a dedicated architecture.

[0177] The neural network-based base caller 2900 further includes temporal convolutional layers (or temporal logic) that blend information between sequencing cycles, i.e., between cycles. The temporal convolutional layers receive their inputs from the spatial convolutional networks and operate on the spatially convolved representations produced by the final spatial convolutional layer for each data processing pipeline.

[0178] The inter-cycle operational freedom of the temporal convolutional layers arises from the fact that misalignment features present in the image data fed as input to the spatial convolutional network are purged from the spatial convolutional representation by the stack or cascade of separated convolutions performed by the sequence of spatial convolutional layers.

[0179] The temporal convolutional layer uses so-called "combinatorial convolution," which convolves group-wise on input channels with subsequent inputs on a sliding window basis. In one embodiment, the subsequent inputs are subsequent outputs generated by previous spatial or temporal convolutional layers.

[0180] In some embodiments, the temporal convolutional layer is part of a temporal convolutional network (or sub-network) within a dedicated architecture. The temporal convolutional network receives its input from a spatial convolutional network. In one embodiment, the first temporal convolutional layer of the temporal convolutional network combines the spatial convolutional representations between sequencing cycles, group by group. In another embodiment, subsequent temporal convolutional layers of the temporal convolutional network combine subsequent outputs of previous temporal convolutional layers. The output of the final temporal convolutional layer is fed to an output layer, which generates an output. The output is used to base call one or more clusters in one or more sequencing cycles.

[0181] Further details regarding the neural network-based base caller 2900 can be found in U.S. Provisional Patent Application No. 62 / 821,766, entitled "ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV), which is incorporated herein by reference.

[0182] Per-pixel state input 29 illustrates one embodiment of state-based and neural network-based base calling. In the example shown in FIG. 29, the first window of sequencing cycles includes sequencing cycles N-2, N-1, N, N+1, and N+2. In one embodiment, each image patch (sequencing image / cluster image) 2914, 2924, 2934, 2944, and 2954 for each sequencing cycle N-2, N-1, N, N+1, and N+2 is combined (e.g., concatenated) pixel-by-pixel with state data 2912 for the current sequencing cycle N. State data 2912 can be determined using any of the techniques described above.

[0183] Image patches 2914, 2924, 2934, 2944, and 2954 are collectively referred to as cluster image (sequencing image) 2950. Per-pixel, per-channel state data 2940 includes five instances of state data 2912 because there are five sequencing cycles: N-2, N-1, N, N+1, and N+2.

[0184] 3 and 4, state data 2912 includes pixel-by-pixel state values ​​for corresponding pixels in cluster image 2950. Also, in one embodiment, state data 2912 includes two state channels because cluster image 2950 includes two channels (intensity channels). In other embodiments, state data 2912 can include only one channel or more than two channels, as described above in different embodiments.

[0185] At a high level, per-pixel, per-channel state data 2940 is combined with cluster images 2950 for processing by spatial logic 106 (or a spatial network or spatial sub-network or spatial convolutional neural network). The combination of per-pixel, per-channel state data 2940 and cluster images 2950 is separately processed via spatial logic 106 to generate respective spatial maps 2916, 2926, 2936, 2946, and 2956 (or intermediate results or spatial output sets or spatial feature map sets) for each sequencing cycle N-2, N-1, N, N+1, and N+2. The spatial convolutional network 106 can use 1D, 2D, or 3D convolution.

[0186] The spatial logic 106 includes a sequence (or cascade) of spatial convolutional layers. Each spatial convolutional layer has a filter bank with multiple spatial convolutional filters that implement decoupled convolutions. Therefore, each spatial convolutional layer generates multiple spatial feature maps as output. The number of spatial feature maps generated by a given subject spatial convolutional layer is a function of the number of spatial convolutional filters configured in that subject spatial convolutional layer. For example, if a subject spatial convolutional layer has 14 spatial convolutional filters, the subject spatial convolutional layer generates 14 spatial feature maps. From a collective perspective, the 14 spatial feature maps can be viewed as a volume (or tensor) of spatial feature maps with 14 channels (or depth dimension = 14).

[0187] Furthermore, the next spatial convolutional layer following the target spatial convolutional layer may also be composed of 14 spatial convolutional filters. In such a case, the next spatial convolutional layer processes the 14 spatial feature maps as input to generate the target spatial convolutional layer, which itself generates 14 new spatial feature maps as output. Figure 29 shows five spatial feature map sets 2916, 2926, 2936, 2946, and 2956 generated by the final spatial convolutional layer of the spatial network 106 for sequencing cycles N-2, N-1, N, N+1, and N+2, respectively. For example, each of the five spatial feature map sets 2916, 2926, 2936, 2946, and 2956 has 14 feature maps.

[0188] In another example, a sequence of seven spatial feature map sets can be generated by a cascade of seven spatial convolutional layers of the spatial network 106. The combination of input patch data and state data for each cycle for target sequencing cycle i can have, for example, a spatial dimensionality of 115x115 and a depth dimensionality of 2 (due to the two image channels of the original sequencing image). In one implementation, each of the seven spatial convolutional layers uses a 3x3 convolution, which reduces the spatial dimensionality of the subsequent spatial feature map volume by two, e.g., from 10x10 to 8x8.

[0189] The first spatial feature map volume may have a spatial dimensionality of 113x113 (i.e., reduced from 115x115 by 3x3 convolutions in the first spatial convolution layer) and a depth dimensionality of 14 (i.e., 14 feature maps or 14 channels due to 14 spatial convolution filters in the first spatial convolution layer). The second spatial feature map volume may have a spatial dimensionality of 111x111 (i.e., reduced from 113x113 by 3x3 convolutions in the second spatial convolution layer) and a depth dimensionality of 14 (i.e., 14 feature maps or 14 channels due to 14 spatial convolution filters in the second spatial convolution layer). The third spatial feature map volume may have a spatial dimensionality of 10x10 (i.e., reduced from 111x111 by the 3x3 convolutions of the third spatial convolutional layer) and a depth dimensionality of 14 (i.e., 14 feature maps or 14 channels due to the 14 spatial convolution filters in the third spatial convolutional layer). The fourth spatial feature map volume may have a spatial dimensionality of 10x10 (i.e., reduced from 10x10 by the 3x3 convolutions of the fourth spatial convolutional layer) and a depth dimensionality of 14 (i.e., 14 feature maps or 14 channels due to the 14 spatial convolution filters in the fourth spatial convolutional layer). The fifth spatial feature map volume may have a spatial dimensionality of 10×10 (i.e., reduced from 10×10 by the 3×3 convolutions of the fifth spatial convolutional layer) and a depth dimensionality of 14 (i.e., 14 feature maps or 14 channels due to the 14 spatial convolution filters in the fifth spatial convolutional layer). The sixth spatial feature map volume may have a spatial dimensionality of 10×10 (i.e., reduced from 10×10 by the 3×3 convolutions of the sixth spatial convolutional layer) and a depth dimensionality of 14 (i.e., 14 feature maps or 14 channels due to the 14 spatial convolution filters in the sixth spatial convolutional layer).The seventh spatial feature map volume may have a spatial dimensionality of 101x101 (i.e., reduced from 103x103 by the 3x3 convolutions of the seventh spatial convolution layer) and a depth dimensionality of 14 (i.e., 14 feature maps or 14 channels due to the 14 spatial convolution filters in the seventh spatial convolution layer).

[0190] Similar to the multi-cycle example shown in FIG. 29, for five sequencing cycles N-2, N-1, N, N+1, and N+2, and five per-cycle image patches 2914, 2924, 2934, 2944, and 2954, and five instances of state data 2912, the spatial logic 106 generates five separate sequences of seven spatial feature map volumes, and spatial maps 2916, 2926, 2936, 2946, and 2956 in FIG. 29 are equivalent to five separate instances of the final spatial feature map volume.

[0191] Spatial maps 2916, 2926, 2936, 2946, and 2956 are processed by temporal logic 2928 to generate base calls 2930 for the current sequencing cycle N.

[0192] Cluster / well status Figure 30 illustrates one embodiment that uses per-cluster state data for base call clusters. Unlike Figure 29, in Figure 30, the state data is not combined with the input cluster images 2914, 2924, 2934, 2944, and 2954. Instead, the state data is combined with the output of spatial logic 106. Note that cluster images 2914, 2924, 2934, 2944, and 2954 include both cluster and non-cluster pixels. As a result, spatial maps 3016, 3026, 3036, 3046, and 3056 include spatially convolved features emanating from both cluster and non-cluster pixels.

[0193] In some implementations, the cluster feature filtering logic 3030 uses the cluster center locations 3002 to filter out those features from the spatial maps 3016, 3026, 3036, 3046, and 3056 that correspond to non-cluster pixels. The resulting filtered per-cluster spatial maps 3054, 3064, 3074, 3084, and 3094 contain only features that correspond to cluster pixels, collectively referred to as cluster features 3004. In one implementation, the spatial maps 3016, 3026, 3036, 3046, and 3056 have a dimensionality of 101 x 101, and the filtered per-cluster spatial maps 3054, 3064, 3074, 3084, and 3094 have a dimensionality of 25 x 25. In such an embodiment, the 25×25×k tensor contains spatial features generated as a result of applying successive convolution operations to data corresponding to the cluster pixels, including the original cluster pixels and successively generated convolution features of the cluster pixels.

[0194] Unlike FIG. 29, where state data 2940 is combined with cluster images 2950 at the starting input to neural network-based base caller 2900, in FIG. 30, per-cluster states 3042 for the current sequencing cycle N are combined with intermediate outputs of neural network-based base caller 2900, namely, filtered per-cluster spatial maps 3054, 3064, 3074, 3084, and 3094.

[0195] In some implementations, the dimensionality of the per-cluster state 3042 is modified (e.g., by trimming dimensionality, adding dimensionality, padding (e.g., zero-padding), cloning, etc.) to match the dimensionality of the filtered per-cluster spatial maps 3054, 3064, 3074, 3084, and 3094. In different implementations, the five instances of each of the per-cluster state 3042 and the filtered per-cluster spatial maps 3054, 3064, 3074, 3084, and 3094 may each be combined using the techniques described above, such as concatenation, summation, element-wise multiplication, and element-wise multiplication and summation (convolution).

[0196] 30 , five combinations of each of the filtered spatial maps 3054, 3064, 3074, 3084, and 3094 for each cluster, and five instances of each of the states 3042 for each cluster, are then processed by the temporal logic 2928 to generate a base call 3060 for the current sequencing cycle N. The base call 3060 is made for the cluster corresponding to the cluster pixel, and therefore the cluster feature 3004.

[0197] Next, different variations of the per-cluster state 3042 and how these variations are determined according to different embodiments of the disclosed technology will be described.

[0198] Figure 31 illustrates one implementation of generating 3100 a per-cluster state using past intensity values ​​of corresponding cluster pixels. In Figure 31, the per-cluster state 3042 is determined by accumulating summary statistics of past intensity values ​​of cluster pixels from previous and current sequencing cycles, using, for example, minimum-selection functions, maximum-selection functions, averaging functions, exponentially weighted averaging functions, etc., as described above. The per-cluster state 3042 is then combined with spatially convolved features of the cluster pixels and the current base calling iteration / operation to generate a base call for the corresponding cluster. These steps are indicated by actions 3102, 3112, 3122, and 3132 in Figure 31.

[0199] Figure 32 illustrates one implementation of generating a per-cluster state (3200) using past feature values ​​of spatially convolved features corresponding to cluster pixels. In Figure 31, the per-cluster state 3042 is determined by accumulating summary statistics of past feature values ​​of spatially convolved features corresponding to cluster pixels from previous and current sequencing cycles, using, for example, minimum selection functions, maximum selection functions, averaging functions, exponentially weighted averaging functions, etc., as described above. The per-cluster state 3042 is then combined with the cluster pixels and the spatially convolved features of the current base calling iteration / operation to generate a base call for the corresponding cluster. These steps are indicated by actions 3202, 3212, 3222, and 3232 in Figure 32.

[0200] Figure 33 shows one embodiment in which cluster intensities are generated by interpolating pixel intensities (3300) and the interpolated cluster intensities are used to generate per-cluster states for base calling. In one embodiment, intensity values ​​of groups of pixels in a sequencing image can be interpolated (e.g., by using linear, bilinear, or cubic interpolation) to estimate cluster intensities. In some embodiments, cluster center location and cluster shape data can be used to select a given group of pixels and interpolate cluster intensities for the corresponding cluster. In other embodiments, cluster shape data can be used to determine cluster intensities using weighted intensities of partial pixels.

[0201] For example, minimum selection functions, maximum selection functions, averaging functions, exponentially weighted averaging functions, etc. may be used to accumulate the interpolated cluster intensity summary statistics from previous and current sequencing cycles, as described above, to generate per-cluster state 3042 for the current sequencing cycle. Per-cluster state 3042 is then combined with the spatially convolved features of the cluster pixels and the current base calling iteration / operation to generate a base call for the corresponding cluster. These steps are indicated by actions 3302, 3312, 3322, and 3332 in Figure 32.

[0202] RTA coefficients as states per cluster FIG. 34 shows one embodiment in which a real-time analysis (RTA) base caller separate from the neural network-based base caller 2900 is used to generate states for each cluster 3400 .

[0203] Regarding the RTA base caller, RTA is a base caller that can extract features from sequencing images for base calling using a linear intensity extractor. The following discussion describes one embodiment of intensity extraction and base calling by RTA. In this embodiment, RTA can perform a template generation step to generate a template image that identifies the location of clusters on a tile using sequencing images from a number of initial sequencing cycles, called template cycles. The template image can be used as a reference for the subsequent alignment and intensity extraction steps. The template image can be generated by detecting and merging bright spots in each sequencing image of the template cycle, which involves sharpening the sequencing image (e.g., using Laplacian convolution), determining an "ON" threshold using the spatially separated Otsu method, and then performing 5-pixel local maximum detection using subpixel location interpolation. In another example, the location of clusters on a tile can be identified using fiducial markers. The solid support on which the biological sample is imaged can include such fiducial markers to facilitate determining the orientation of the sample or its image relative to the probes attached to the solid support. Exemplary criteria may include, but are not limited to, beads (with or without moieties such as fluorescent moieties or nucleic acids to which labeled probes can be attached), fluorescent molecules bound to known or determinable features, or structures that combine morphological shapes with fluorescent moieties. Exemplary criteria are described in U.S. Patent Application Publication No. 2002 / 0150909, which is incorporated herein by reference.

[0204] RTA can then register the current sequencing image to the template image, which can be achieved by aligning the current sequencing image to the template image over a subregion using image correlation, or by using a nonlinear transformation (e.g., a full six-parameter linear affine transformation).

[0205] The RTA can generate a color matrix to correct for crosstalk between color channels in the sequencing image. The RTA can perform empirical phase correction to compensate for noise in the sequencing image caused by phase errors.

[0206] After different corrections are applied to the sequencing image, the RTA can extract signal intensities for each spot location in the sequencing image. For example, for a given spot location, the signal intensity can be extracted by determining a weighted average of the intensities of the pixels within the spot location. For example, the weighted average of the central pixel and neighboring pixels can be performed using bilinear or bicubic interpolation. In some embodiments, each spot location in the image can comprise several pixels (e.g., 1 to 5 pixels).

[0207] RTA can then spatially normalize the extracted signal intensities to account for variations in illumination across the sampled image. For example, intensity values ​​may be normalized so that the 5th and 95th percentiles have values ​​of 0 and 1, respectively. The normalized signal intensities of the image (e.g., the normalized intensities of each channel) can be used to calculate the average sharpness of multiple spots in the image.

[0208] In some embodiments, the RTA can use an equalizer to maximize the signal-to-noise ratio of the extracted signal intensities. The equalizer can be trained (e.g., using least-squares estimation, an adaptive equalization algorithm) to maximize the signal-to-noise ratio of the cluster intensity data in the sequencing image. In one embodiment, the equalizer can include trained coefficients for correcting for spatial crosstalk and other forms of crosstalk and noise. In some embodiments, the equalizer can be a single look-up table (LUT). In other embodiments, the equalizer can be a LUT bank with multiple LUTs with subpixel resolution, also referred to as "equalizer filters" or "convolution kernels." In one embodiment, the number of LUTs in the equalizer can depend on the number of subpixels into which a pixel of the sequencing image can be divided. For example, if a pixel can be divided into n x n subpixels (e.g., 5 x 5 subpixels), the equalizer can generate n LUTs (e.g., 25 LUTs). In a further embodiment, the equalizer may be based on a Fast Fourier Transform (FFT). In a further embodiment, the equalizer may be based on Winograd convolution.

[0209] In one training embodiment, data from sequencing images are binned by well subpixel location. For example, for a 5x5 LUT, the center of 1 / 25th of the wells is in bin (1,1) (e.g., the upper left corner of the sensor pixel), the center of 1 / 25th of the wells is in bin (1,2), and so on. In one embodiment, the equalizer coefficients for each bin are determined using least-squares estimation on the subset of data from the wells corresponding to the respective bin. In this way, the resulting estimated equalizer coefficients will vary from bin to bin.

[0210] Each LUT / equalizer filter / convolution kernel has multiple coefficients learned from training. In one embodiment, the number of coefficients in the LUT corresponds to the number of pixels used to base call clusters. For example, if the local grid of pixels (image or pixel patch) used to base call clusters is of size p×p (e.g., a 9×9 pixel patch), then each LUT has p coefficients (e.g., 81 coefficients).

[0211] In one embodiment, the training generates equalizer coefficients configured to mix / combine pixel intensity values ​​representing intensity radiation from a base-called target cluster and intensity radiation from one or more neighboring clusters to maximize the signal-to-noise ratio. The signal that maximizes the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that minimizes the signal-to-noise ratio is the intensity radiation from neighboring clusters, i.e., spatial crosstalk, plus some random noise (e.g., to account for background intensity radiation). The equalizer coefficients are used as weights, and the mixing / combining involves performing element-wise multiplication between the equalizer coefficients and the pixel intensity values ​​to calculate a weighted sum of the pixel intensity values, i.e., a convolution operation.

[0212] RTA can perform base calling by fitting a mathematical model to the optimized intensity data. Suitable mathematical models that can be used include, for example, k-means clustering algorithms, k-means-like clustering algorithms, expectation maximization clustering algorithms, histogram-based methods, etc. Four Gaussian distributions can be fitted to a set of two-channel intensity data, such that one distribution is applied to each of the four nucleotides represented in the dataset. In one specific embodiment, an expectation maximization (EM) algorithm can be applied. As a result of the EM algorithm, for each X, Y value (referring to each of the two channel intensities), a value can be generated that represents the likelihood that a given X, Y intensity value belongs to one of the four Gaussian distributions to which the data is fitted. If four bases give four distinct distributions, each X, Y intensity value also has four associated likelihood values, one for each of the four bases. The maximum of the four likelihood values ​​indicates the base call. For example, if a cluster is "OFF" in both channels, the base call is G. If the cluster is "OFF" in one channel and "ON" in another channel, the base call is either C or T (depending on which channel is ON); if the cluster is "ON" in both channels, the base call is A.

[0213] Further details regarding RTA can be found in U.S. Non-Provisional Patent Application No. 15 / 909,437, filed March 1, 2018, entitled "OPTICAL DISTORTION CORRECTION FOR IMAGED SAMPLES." U.S. Non-provisional Patent Application No. 14 / 530,299, entitled "IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS," filed October 31, 2014; U.S. Non-provisional Patent Application No. 15 / 153,953, entitled "METHODS AND SYSTEMS FOR ANALYZING IMAGE DATA," filed December 3, 2014; U.S. Non-provisional Patent Application No. 13 / 006,206, entitled "DATA PROCESSING SYSTEM AND METHODS," filed January 13, 2011; and U.S. Non-provisional Patent Application No. 17 / 308,035, entitled "EQUALIZATION-BASED IMAGE PROCESSING AND SPATIAL CROSSTALK ATTENUATOR," filed May 4, 2021 (Attorney Docket No. ILLM 1032-2 / IP-1991-US), all of which are incorporated by reference as if fully set forth herein.

[0214] Inter-cluster intensity profile variation in the intensity profiles of a large number of clusters (e.g., thousands, millions, hundreds of millions, etc.) within a cluster population causes a decrease in base calling throughput and an increase in base calling error rate. To correct for inter-cluster intensity profile variation, RTA generates a cluster-specific variation correction factor for each cluster.

[0215] In a two-channel implementation, the per-cluster variation correction factor includes an amplification factor that accounts for scale variations in inter-cluster intensity profile variations and two-channel-specific offset factors that account for shift variations along the first and second intensity channels, respectively, in the inter-cluster intensity profile variations. In another implementation, the shift variations are accounted for by using a common offset factor for different intensity channels (e.g., the first and second intensity channels).

[0216] It will be apparent to one skilled in the art that the variation correction logic of the RTA can be similarly applied to sequencing images generated using a one-channel implementation, a four-channel implementation, etc. For example, in the case of a four-channel implementation, four channel-specific offset coefficients are determined to correct for shift variations in each of the four intensity channels.

[0217] A cluster-specific variation correction factor for the target cluster is generated in a current sequencing cycle of the sequencing run based on a combination of an analysis of past intensity statistics determined for the target cluster in a previous sequencing cycle of the sequencing run and an analysis of current intensity statistics determined for the target cluster in the current sequencing cycle. The cluster-specific variation correction factor is used to correct the next intensity read registered for the target cluster in the next sequencing cycle of the sequencing run. The corrected next intensity read is used to base call the target cluster in the next sequencing cycle. Repeated application of each cluster-specific variation correction factor to each intensity profile of each cluster in successive sequencing cycles of the sequencing run results in the intensity profiles becoming consistent and anchored to the origin (e.g., at the lower corner of a trapezoid).

[0218] Current sequencing cycle In a current sequencing cycle i, the sequencer generates a sequencing image, which includes current intensity data registered to multiple clusters in the cluster population and current intensity data registered to a target cluster in the current sequencing cycle i.

[0219] The current intensity data is provided to the RTA, which processes the current intensity data and generates current base calls for the target cluster in the current sequencing cycle i.

[0220] In the current sequencing cycle i, the intensity profile of the target cluster includes the current intensity data and the current past intensity data registered for the target cluster in the sequencing cycles of the sequencing run preceding the current sequencing cycle i, i.e., the preceding sequencing cycles 1 to i-1. The current intensity data and the current past intensity data are collectively referred to as the currently available intensity data.

[0221] In the intensity profile, the four intensity distributions correspond to the four bases A, C, T, and G. In one embodiment, the current base call is made by determining which of the four intensity distributions the current intensity data belongs to. In some embodiments, this is accomplished by using an expectation-maximization algorithm, which iteratively maximizes the likelihood of observing a mean (centroid) and distribution (covariance) that best fits the currently available intensity data.

[0222] By using the expectation maximization algorithm, once the four intensity distributions are determined in the current sequencing cycle i, the likelihood of the current intensity data belonging to each of the four intensity distributions is calculated. The largest likelihood gives the current base call. As an example, let "m, n" be the intensity values ​​of the current intensity data in the first and second intensity channels, respectively. The expectation maximization algorithm generates four values ​​representing the likelihood of the intensity value of "m, n" belonging to each of the four intensity distributions. The maximum of the four values ​​identifies the called base.

[0223] In other embodiments, k-means clustering algorithms, k-means-like clustering algorithms, histogram-based methods, etc. may be used for base calling.

[0224] Next sequencing cycle At the next sequencing cycle i+1, the intensity correction parameter determiner determines intensity correction parameters for the target cluster based on the current base call. In a two-channel embodiment, the intensity correction parameters include a distribution intensity in the first intensity channel, a distribution intensity in the second intensity channel, an intensity error in the first intensity channel, an intensity error in the second intensity channel, a distance from the distribution centroid to the origin, and a similarity measure of the distribution intensity versus intensity error.

[0225] Each of the intensity correction parameters is defined as follows: 1) The distribution intensity in the first intensity channel is the intensity value in the first intensity channel at the centroid of the base-specific intensity distribution to which the target cluster belongs in the current sequencing cycle i. Note that the base-specific intensity distribution is the basis for making the current base call. 2) The distribution intensity in the second intensity channel is the intensity value in the second intensity channel at the centroid of the base-specific intensity distribution. 3) The intensity error in the first intensity channel is the difference between the measured intensity value of the current intensity data in the first intensity channel and the distribution intensity in the first intensity channel. 4) The intensity error in the second intensity channel is the difference between the measured intensity value of the current intensity data in the second intensity channel and the distribution intensity in the second intensity channel. 5) The distance of the distribution centroid to the origin is the Euclidean distance between the centroid of the base-specific intensity distribution and the origin of the multidimensional space to which the base-specific intensity distribution was fitted (e.g., by using an expectation-maximization algorithm). In other embodiments, distance metrics such as Mahalanobis distance and minimum covariance determinant (MCD) distance, and their associated centroid estimates, can be used. 6) The similarity measure of distribution intensity versus intensity error is the sum of channel-wise dot products between the distribution intensity and intensity error in the first and second intensity channels.

[0226] The cumulative intensity correction parameter determiner accumulates the intensity correction parameter with a previous cumulative intensity correction parameter from the preceding sequencing cycle i-1 to determine the cumulative intensity correction parameter. Examples of accumulation include summing and averaging.

[0227] The fluctuation correction factor determiner determines a fluctuation correction factor based on the determined cumulative intensity correction parameter.

[0228] At the next sequencing cycle i+1, the sequencer generates a sequencing image that includes the next intensity data registered to the clusters in the cluster population and the next intensity data registered to the target cluster at the next sequencing cycle i+1.

[0229] The intensity corrector applies the variation correction factor to the next intensity data to generate the next corrected intensity data.

[0230] In the next sequencing cycle i+1, the intensity profile of the target cluster includes the corrected next intensity data and the next previous intensity data registered with the target cluster in the sequencing cycle of the sequencing run preceding the next sequencing cycle i+1, i.e., the preceding sequencing cycles 1 through i. The corrected next intensity data and the next previous intensity data are collectively referred to as the next available intensity data.

[0231] The corrected next intensity data is provided to the RTA, which can process the corrected next intensity data and generate a next base call for the target cluster in the next sequencing cycle i+1. To generate the next base call, an expectation maximization algorithm can observe the mean (centroid) and distribution (covariates) based on the corrected next intensity data to best fit the next available intensity data.

[0232] By using the expectation maximization algorithm, once the four intensity distributions are determined in the next sequencing cycle i+1, the likelihood of the corrected next intensity data belonging to each of the four intensity distributions is calculated, and the highest likelihood gives the next base call.

[0233] Note that the base calling pipeline is run per cluster, in parallel for multiple clusters within a cluster population, and is run iteratively for successive sequencing cycles of a sequencing run (e.g., for 150 consecutive sequencing cycles of read 1 and another 150 consecutive sequencing cycles of read 2 in a paired-end sequencing run).

[0234] least squares method The least squares method determines closed-form expressions for the cumulative intensity correction parameters and the variation correction coefficients. The least squares method determiner includes an intensity modeler and a minimizer.

[0235] The intensity model models the relationship between the measured intensity of the target cluster and the variation correction factor according to the following equation: y C,i =ax C,i +d i +n C,i Equation (1) During the ceremony, a is the amplification factor of the target cluster. d i is the channel-specific offset coefficient for intensity channel i. x C,i is the distribution intensity in intensity channel i for the target cluster at the current sequencing cycle C. y C,i is the measured intensity in intensity channel i for the target cluster in the current sequencing cycle C. n C,i is the additive noise of intensity channel i for the target cluster at the current sequencing cycle C.

[0236] The minimizer uses the least squares method to minimize the following equation:

[0237]

number

[0238]

number

[0239]

number

[0240] Using the chain rule, the minimizer calculates the amplification factor

[0241]

number

[0242]

number

[0243]

number

[0244] Channel-specific intensity error e c,i is defined as follows: e C,i =y C,i -x C,i Equation (5)

[0245] Closed-form equation The first partial derivative is the amplification factor

[0246]

number

[0247]

number

[0248] Closed-form expressions for cumulative intensity correction parameters

[0249]

number

[0250]

number

[0251]

number

[0252] Each of the cumulative intensity correction parameters is defined as follows: 1) First cumulative intensity correction parameter

[0253]

number

[0254]

number

[0255]

number

[0256]

number

[0257]

number

[0258]

number

[0259] The second partial derivative is the offset coefficient

[0260]

number

[0261]

number

[0262] Then, for each intensity channel:

[0263]

number

[0264] For the first intensity channel, i.e., i=1:

[0265]

number

[0266]

number

[0267] For the second intensity channel, i.e., i=2:

[0268]

number

[0269]

number

[0270] Substituting equations 17 and 18 into equation 11:

[0271]

number

[0272]

number

[0273] In another embodiment, to reduce memory requirements per cluster, a common offset factor for different intensity channels (e.g., the first and second intensity channels) is determined by the constraint

[0274]

number

[0275]

number

[0276] It will be apparent to one skilled in the art that a least squares method is performed prior to a sequencing run to determine a closed-form equation, which, once determined, is applied to the intensity values ​​generated during the sequencing run for each cluster iteratively in each sequencing cycle of the sequencing run.

[0277] Further details regarding per-cluster variation correction factors and how they are determined can be found in U.S. Provisional Patent Application No. 63 / 106,256, entitled "SYSTEMS AND METHODS FOR PER-CLUSTER INTENSITY CORRECTION AND BASE CALLING," filed October 27, 2021 (Attorney Docket No. ILLM 1034-1 / IP-2026-PRV), which is incorporated herein by reference in its entirety.

[0278] In Figure 34, per-cluster state 3042 is the per-cluster variation correction factor determined by RTA for the current sequencing cycle, according to one embodiment. Per-cluster state 3042 is then combined with the spatial convolution features of the cluster pixels and the current base call iteration / operation to generate a base call for the corresponding cluster. These steps are indicated by actions 3402, 3412, 3422, and 3432 in Figure 34.

[0279] Compression Network As described above, the specialized architecture of the neural network-based base caller 2900 processes a sliding window of image patches for a corresponding sequencing cycle. There is overlap between the sequencing cycles of subsequent sliding windows. This causes the neural network-based base caller 2900 to redundantly process image patches for overlapping sequencing cycles. This results in wasted computational resources. For example, in one embodiment, each spatial convolutional layer of the neural network-based base caller 2900 has approximately 100 million multiplication operations. Then, for a window of five sequencing cycles and a cascade (or sequence) of seven spatial convolutional layers, the spatial convolutional neural network performs approximately 620 million multiplication operations. Furthermore, the temporal convolutional neural network performs approximately 10 million multiplication operations.

[0280] Because image data for cycle N-1 in the current sliding window (or current iteration of base calling) is processed as cycle N in the previous sliding window (or previous iteration of base calling), an opportunity is provided to store intermediate results of processing performed in the current sliding window and intermediate results of processing performed in subsequent windows, thereby avoiding (or eliminating) redundant processing (or reprocessing) of input image data for overlapping sequencing cycles between subsequent sliding windows.

[0281] However, the intermediate results are several terabytes of data, requiring impractical amounts of storage. To overcome this technical problem, the disclosed technique proposes to compress the intermediate results when they are first generated by the neural network-based base caller 2900, and then repurpose and utilize the compressed intermediate results in a subsequent sliding window to avoid redundant computations, thereby not regenerating the intermediate results (or generating them only once).

[0282] 35, according to one embodiment, compression logic 3530 (or compression network or compression sub-network or compression layer or squeeze layer) processes the output of the cluster feature filtering logic 3030 and generates a compressed representation of the output. In one embodiment, the compression network 3530 comprises a compressed convolutional layer that reduces the depth dimensionality of the feature maps generated by the cluster feature filtering logic 3030.

[0283] For example, consider that the depth dimensionality of the filtered per-cluster spatial maps 3054, 3064, 3074, 3084, and 3094 is 14 (i.e., 14 feature maps or 14 channels per spatial output). The compression network 3530 attenuates the filtered per-cluster spatial maps 3054, 3064, 3074, 3084, and 3094 into respective compressed filtered per-cluster spatial maps 3554, 3564, 3574, 3584, and 3594 for respective sequencing cycles N-2, N-1, N, N+1, and N+2, collectively referred to as compressed cluster features 3504. In one implementation, each of the compressed filtered per-cluster spatial maps 3554, 3564, 3574, 3584, and 3594 has a depth dimensionality of 2 (i.e., two feature maps or two channels per compressed spatial output). In other embodiments, the compressed filtered spatial maps per cluster 3554, 3564, 3574, 3584, and 3594 can have a depth dimensionality of three or four (i.e., three or four feature maps or three or four channels per compressed spatial output). In yet other embodiments, the compressed filtered spatial maps per cluster 3554, 3564, 3574, 3584, and 3594 can have a depth dimensionality of one (i.e., one feature map or one channel per compressed spatial output). In one embodiment, the compression layer 3530 does not include an activation function such as ReLU. In other embodiments, the compression layer 3530 can include an activation function. In other embodiments, the compression logic 3530 can configure a corresponding set of compressed spatial maps to each have more than four feature maps.

[0284] We now discuss how the compression logic 3530 generates the compressed output.

[0285] In one implementation, the compression logic 3530 uses 1x1 convolutions to reduce the number of feature maps (i.e., the depth dimension, or number of channels) while introducing nonlinearity. A 1x1 convolution has a kernel size of 1. A 1x1 convolution can convert volumetric depth to another squeezed or dilated representation without changing the spatial dimension. A 1x1 convolution acts like a fully connected linear layer over the input channels. This is useful for mapping from feature maps with many channels to a smaller number of feature maps. A single 1x1 convolution can be applied to an input tensor with two feature maps. A 1x1 convolution compresses a two-channel input to a single-channel output.

[0286] The number of compressed outputs (or compressed feature maps or compressed spatial maps or compressed temporal maps) generated by the compression layer 108 is a function of the number of 1×1 convolution filters (or compressed convolution filters or compression filters) configured in the compression layer 3530. In one embodiment, the compression layer 3530 may have two 1×1 convolution filters. The first 1×1 convolution filter may process a spatial feature volume having 14 feature maps and generate a first feature map while preserving the 101×101 spatial dimensionality. The second 1×1 convolution filter may also process a spatial feature volume having 14 feature maps and generate a second feature map while preserving the 101×101 spatial dimensionality. Thus, the compression layer 3530 reduces a spatial feature volume having 14 feature maps to a compressed output having two spatial feature maps (i.e., a compression ratio of 7).

[0287] In some embodiments, the disclosed technology saves approximately 80% of convolutions in the spatial network of the neural network-based base caller 2900. In one embodiment, 80% savings are observed in spatial convolutions when compression logic and repurposing of compressed feature maps are used in subsequent sequencing cycles for an input window of 5 sequencing cycles (e.g., cycle N, cycle N+1, cycle N-1, cycle N+2, and cycle N-2). In another embodiment, 90% savings are observed in spatial convolutions when compression logic and repurposing of compressed feature maps in subsequent sequencing cycles are used for an input window of 10 sequencing cycles (e.g., cycle N, cycle N+1, cycle N-1, cycle N+2, cycle N-2, cycle N+3, and cycle N-3). That is, the larger the window size, the greater the savings from using compression logic and repurposing of compressed feature maps, and the larger the window size, the better the base calling performance due to the incorporation of greater context from additional neighboring cycles. Greater savings for larger windows improve overall performance for a given computational power.

[0288] The computational efficiency and compact computational footprint afforded by the compression logic facilitates hardware implementation of the neural network-based basis caller 2900 on resource-constrained processors such as central processing units (CPUs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), coarse-grained reconfigurable architectures (CGRAs), application-specific integrated circuits (ASICs), application-specific instruction-set processors (ASIPs), and digital signal processors (DSPs).

[0289] The computations saved by the compression logic allow for the incorporation of more convolution operators into the neural network-based base caller 2900. Examples include adding more convolution filters to the spatial and temporal convolution layers, increasing the size of the convolution filters, and increasing the number of spatial and temporal convolution layers. The additional convolution operations improve the intensity pattern detection and overall base calling accuracy of the neural network-based base caller 2900.

[0290] Further details regarding the compression logic and its compressed output can be found in U.S. Non-Provisional Patent Application No. 17 / 179,395, entitled "DATA COMPRESSION FOR ARTIFICIAL INTELLIGENCE-BASED BASE CALLING," filed February 18, 2021 (Attorney Docket No. ILLM 1029-2 / IP-1964-US), which is incorporated herein by reference in its entirety.

[0291] In Figure 35, the per-cluster state 3042 for the current sequencing cycle N is combined with the intermediate outputs of the neural network-based base caller 2900, namely, the compressed and filtered per-cluster spatial maps 3554, 3564, 3574, 3584, and 3594.

[0292] In some implementations, the dimensionality of the per-cluster state 3042 is modified (e.g., by trimming dimensions, adding dimensions, padding (e.g., zero-padding), cloning, etc.) to match the dimensionality of the compressed and filtered per-cluster spatial maps 3554, 3564, 3574, 3584, and 3594. In a different implementation, the five instances of each of the per-cluster state 3042 and the compressed and filtered per-cluster spatial maps 3554, 3564, 3574, 3584, and 3594 may each be combined using the techniques described above, such as concatenation, summation, element-wise multiplication, and element-wise multiplication and summation (convolution).

[0293] 35 , each of the five combinations of each of the compressed and filtered per-cluster spatial maps 3554, 3564, 3574, 3584, and 3594 with each of the five instances of per-cluster state 3042 is then processed by temporal logic 2928 to generate base calls 3080 for the current sequencing cycle N. Base calls 3080 are made for clusters corresponding to cluster pixels, and therefore cluster features 3004.

[0294] Next, different variations of the per-cluster state 3042 and how these variations are determined according to different embodiments of the disclosed technology will be described.

[0295] Figure 36 illustrates one implementation of generating a per-cluster state (3600) using past feature values ​​of compressed-space convolved features corresponding to cluster pixels. In Figure 36, the per-cluster state 3042 is determined by accumulating summary statistics of past feature values ​​of compressed-space convolved features corresponding to cluster pixels from previous and current sequencing cycles, as described above, using, for example, a minimum-selection function, a maximum-selection function, an averaging function, an exponentially weighted averaging function, etc. The per-cluster state 3042 is then combined with the cluster pixels and the compressed-space convolved features of the current base calling iteration / operation to generate a base call for the corresponding cluster. These steps are indicated by actions 3602, 3612, 3622, 3632, and 3642 in Figure 36.

[0296] In Figure 37, per-cluster state 3042 is the per-cluster variation correction factor determined by RTA for the current sequencing cycle, according to one embodiment. Per-cluster state 3042 is then combined with the compressed spatial convolution features of the cluster pixels and the current base call iteration / operation to generate a base call for the corresponding cluster. These steps are indicated by actions 3702, 3712, 3722, 3732, and 3742 in Figure 37.

[0297] Figure 38 illustrates one implementation of generating 3800 a per-cluster state 3042 using past intensity values ​​of corresponding cluster pixels to incorporate the per-cluster state 3042 into compressed features. In Figure 38, the per-cluster state 3042 is determined by accumulating summary statistics of past intensity values ​​of cluster pixels from previous and current sequencing cycles, as described above, using, for example, a minimum selection function, a maximum selection function, an averaging function, an exponentially weighted averaging function, etc. The per-cluster state 3042 is then combined with the compressed spatial convolution features of the cluster pixels and the current base calling iteration / operation to generate a base call for the corresponding cluster. These steps are indicated by actions 3802, 3812, 3822, 3832, and 3842 in Figure 38.

[0298] Conversion of well state data to pixel state data Figure 39 shows one embodiment of generating dense per-pixel states from sparse per-well states. In Figure 39, per-well states 3902 are sparse and can be converted to dense per-pixel states 3906 using transpose convolution or interpolation (e.g., linear, bilinear, cubic interpolation) 3904.

[0299] "Sparse" and "dense" refer to the number of zero elements versus non-zero elements in an array (e.g., a vector or matrix). A sparse array is an array that contains mostly zeros and a few non-zero entries. A dense array contains mostly non-zeros. In a dense array, in one embodiment, each per-pixel state may be surrounded by at least two, four, or eight neighboring per-pixel states. In a sparse array, in one embodiment, each per-well state is not surrounded by even two, four, or eight neighboring per-well states. A set of states is called dense if every row between the first and last row is defined with a non-zero element and given a value. A set of states is called sparse if there are gaps or zero elements in the rows. Sparsity can mean that many elements (e.g., every other, every third, or every fourth) are zero or very close to zero.

[0300] Figure 40 illustrates one embodiment of providing sparse per-well state, dense per-pixel state, and per-pixel intensity values ​​as inputs to a base caller to perform base calling operations. In Figure 40, sparse per-well state 3902 can be provided as input to base caller 144, e.g., neural network-based base caller 2900, both as a direct input and as a redundant input using, e.g., residual connections and / or skip connections. In some embodiments, sparse per-well state 3902 is provided as a supplemental input along with dense per-pixel state 3906 and dense per-pixel intensity values ​​4012 (i.e., a cluster image), which are processed by base caller 144 to generate base calls.

[0301] 41 shows an embodiment that uses sparse per-well states, dense per-pixel states, and per-pixel intensity values ​​as inputs to a neural network-based base caller 2900 to perform base calling operations. The inputs to the neural network-based base caller 2900 are dense per-pixel states 3906 and dense per-pixel intensity values ​​4012 (i.e., the cluster image).

[0302] Sequencing System 42A and 42B show one embodiment of a sequencing system 4200A. The sequencing system 4200A includes a configurable processor 4246. The configurable processor 4246 implements the base calling techniques disclosed herein. A sequencing system is also referred to as a "sequencer."

[0303] The sequencing system 4200A can operate to obtain any information or data relating to at least one of biological or chemical substances. In some embodiments, the sequencing system 4200A is a workstation, which can be similar to a benchtop device or desktop computer. For example, most (or all) of the systems and components for performing the desired reactions can be within a common housing 4202.

[0304] In certain embodiments, the sequencing system 4200A is a nucleic acid sequencing system configured for various applications, including, but not limited to, de novo sequencing, whole genome or targeted genomic region resequencing, and metagenomics. The sequencer may also be used for DNA or RNA analysis. In some embodiments, the sequencing system 4200A may also be configured to generate reaction sites within a biosensor. For example, the sequencing system 4200A may be configured to receive a sample and generate surface-bound clusters of clonally amplified nucleic acids from the sample. Each cluster may constitute or be part of a reaction site within a biosensor.

[0305] The exemplary sequencing system 4200A may include a system receptacle or interface 4210 configured to interact with a biosensor 4212 to effect a desired reaction within the biosensor 4212. In the description that follows with respect to FIG. 42A , the biosensor 4212 is loaded into the system receptacle 4210. However, it is understood that a cartridge containing the biosensor 4212 may be inserted into the system receptacle 4210, and that in some conditions the cartridge may be temporarily or permanently removed. As noted above, the cartridge may include, among other things, fluid control and fluid storage components.

[0306] In certain embodiments, the sequencing system 4200A is configured to perform multiple parallel reactions within the biosensor 4212. The biosensor 4212 includes one or more reaction sites where desired reactions can occur. The reaction sites may be immobilized, for example, on a solid surface of the biosensor or on beads (or other movable substrates) located within corresponding reaction chambers of the biosensor. The reaction sites may include, for example, clusters of clonally amplified nucleic acids. The biosensor 4212 may include a solid-state imaging device (e.g., a CCD or CMOS imager) and a flow cell attached thereto. The flow cell may include one or more flow paths that receive solutions from the sequencing system 4200A and direct the solutions toward the reaction sites. Optionally, the biosensor 4212 may be configured to engage a thermal element for transferring thermal energy into and out of the flow channel.

[0307] The sequencing system 4200A may include various components, assemblies, and systems (or subsystems) that interact with each other to perform a predetermined method or assay protocol for biological or chemical analysis. For example, the sequencing system 4200A includes a system controller 4206 that can communicate with the various components, assemblies, and subsystems of the sequencing system 4200A, and also a biosensor 4212. For example, in addition to the system receptacle 4210, the sequencing system 4200A may also include a fluid control system 4208 for controlling fluid flow throughout the fluidic network of the sequencing system 4200A and the biosensor 4212, a fluid reservoir system 4214 configured to hold any fluids (e.g., fluids, gases, or liquids) that may be used by the bioassay system, a temperature control system 4204 that can regulate the temperature of the fluids in the fluidic network, the fluid reservoir system 4214, and / or the biosensor 4212, and an illumination system 4216 configured to illuminate the biosensor 4212. As described above, when a cartridge having a biosensor 4212 is loaded into the system receptacle 4210, the cartridge may also include fluid control and fluid storage components.

[0308] The sequencing system 4200A may also include a user interface 4218 for interacting with a user. For example, the user interface 4218 may include a display 4220 for displaying or requesting information from the user and a user input device 4222 for receiving user input. In some embodiments, the display 4220 and the user input device 4222 are the same device. For example, the user interface 4218 may include a touch-sensitive display configured to detect the presence of an individual touch and identify the location of the touch on the display. However, other user input devices 4222, such as a mouse, touchpad, keyboard, keypad, handheld scanner, voice recognition system, motion recognition system, etc., may also be used. As described in more detail below, the sequencing system 4200A may communicate with various components, including a biosensor 4212 (e.g., in the form of a cartridge), to perform desired reactions. The sequencing system 4200A may also be configured to analyze data obtained from the biosensor to provide desired information to the user.

[0309] The system controller 4206 may include a microcontroller, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a coarse-grained reconfigurable architecture (CGRA), a logic circuit, and any other circuit or processor capable of performing the functions described herein. The above examples are merely illustrative and thus are not intended to limit the definition and / or meaning of the term system controller. In an exemplary implementation, the system controller 4206 executes a set of instructions stored in one or more storage elements, memories, or modules to at least one of acquire and analyze detection data. The detection data may include multiple sequences of pixel signals, thereby allowing sequences of pixel signals from each of millions of sensors (or pixels) to be detected over many base call cycles. The storage elements may be in the form of an information source or physical memory element within the sequencing system 4200A.

[0310] The instruction set may include various commands that instruct the sequencing system 4200A or biosensor 4212 to perform specific operations, such as the methods and processes of various embodiments described herein. The instruction set may be in the form of a software program, which may form part of a tangible, non-transitory computer-readable medium or media. As used herein, the terms "software" and "firmware" are used interchangeably and include any computer program stored in memory executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are exemplary only and thus not limiting of the types of memory that may be used to store a computer program.

[0311] The software may be in various forms, such as system software or application software. Furthermore, the software may be in the form of a collection of separate programs, or a program module or portion of a program module within a larger program. The software may also include modular programming in the form of object-oriented programming. After acquiring the detection data, the detection data may be processed automatically by the sequencing system 4200A in response to user input, or in response to a request made by another processing machine (e.g., a remote request via a communications link). In the illustrated embodiment, the system controller 4206 includes an analysis module 4244. In other alternative embodiments, the system controller 4206 does not include the analysis module 4244, but instead has access to the analysis module 4244 (e.g., the analysis module 4244 may be separately hosted on the cloud).

[0312] The system controller 4206 may be connected to the biosensor 4212 and other components of the sequencing system 4200A via a communication link. The system controller 4206 may also be communicatively connected to an off-site system or server. The communication link may be a wire, a cord, or wireless. The system controller 4206 may receive user inputs or commands from a user interface 4218 and a user input device 4222.

[0313] The fluid control system 4208 includes a fluid network and is configured to direct the flow of one or more fluids through the fluid network. The fluid network may be in fluid communication with the biosensor 4212 and the fluid storage system 4214. For example, a selected fluid may be drawn from the fluid storage system 4214 and directed to the biosensor 4212 in a controlled manner, or fluid may be drawn from the biosensor 4212 and directed to, for example, a waste reservoir within the fluid storage system 4214. Although not shown, the fluid control system 4208 may include a flow sensor that detects the flow rate or pressure of the fluid within the fluid network. The sensor may be in communication with the system controller 4206.

[0314] The temperature control system 4204 is configured to regulate the temperature of fluids in different regions of the fluid network, the fluid reservoir system 4214, and / or the biosensor 4212. For example, the temperature control system 4204 may include a thermal circulator that interacts with the biosensor 4212 and controls the temperature of fluids flowing along reaction sites within the biosensor 4212. The temperature control system 4204 may also regulate the temperature of solid elements or components of the sequencing system 4200A or the biosensor 4212. Although not shown, the temperature control system 4204 may include sensors for detecting the temperature of fluids or other components. The sensors may be in communication with the system controller 4206.

[0315] The fluid storage system 4214 is in fluid communication with the biosensor 4212 and may store various reaction components or reactants used to carry out a desired reaction. The fluid storage system 4214 may also store fluids for washing or cleaning the fluidic network and biosensor 4212 and for diluting reactants. For example, the fluid storage system 4214 may include various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, etc. Additionally, the fluid storage system 4214 may also include a waste reservoir for receiving waste from the biosensor 4212. In implementations that include a cartridge, the cartridge may include one or more of a fluid storage system, a fluid control system, or a temperature control system. Accordingly, one or more of the components described herein for these systems may be contained within the cartridge housing. For example, the cartridge may have various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, waste, etc. Thus, one or more of the fluid reservoir system, fluid control system, or temperature control system may be removably engaged with the bioassay system via a cartridge or other biosensor.

[0316] The illumination system 4216 may include a light source (e.g., one or more LEDs) and multiple optical components for illuminating the biosensor. Examples of light sources include lasers, arc lamps, LEDs, or laser diodes. The optical components may be, for example, reflectors, polarizers, beam splitters, collimators, lenses, filters, wedges, prisms, mirrors, detectors, etc. In embodiments using an illumination system, the illumination system 4216 may be configured to direct excitation light to the reaction sites. As an example, a fluorophore may be excited by a green wavelength of light, and therefore the wavelength of the excitation light may be approximately 4232 nm. In one embodiment, the illumination system 4216 is configured to generate illumination parallel to the surface normal of the surface of the biosensor 4212. In another embodiment, the illumination system 4216 is configured to generate illumination that is off-angled relative to the surface normal of the surface of the biosensor 4212. In yet another embodiment, the illumination system 4216 is configured to generate illumination having multiple angles, including some parallel illumination and some off-angle illumination.

[0317] The system receptacle or interface 4210 is configured to engage the biosensor 4212 in at least one of mechanical, electrical, and fluidic manners. The system receptacle 4210 can hold the biosensor 4212 in a desired orientation to facilitate fluid flow through the biosensor 4212. The system receptacle 4210 may also include electrical contacts configured to engage the biosensor 4212, thereby allowing the sequencing system 4200A to communicate with and / or provide power to the biosensor 4212. Additionally, the system receptacle 4210 may include a fluid port (e.g., a nozzle) configured to engage the biosensor 4212. In some embodiments, the biosensor 4212 is removably coupled to the system receptacle 4210 both electrically and fluidically.

[0318] Additionally, the sequencing system 4200A may communicate remotely with other systems or networks, or with other bioassay systems 4200A. Detection data obtained by the bioassay system 4200A may be stored in a remote database.

[0319] FIG. 42B is a block diagram of a system controller 4206 that may be used in the system of FIG. 42A. In one embodiment, the system controller 4206 includes one or more processors or modules that can communicate with each other. Each of the processors or modules may include algorithms (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or sub-algorithms for performing a particular process. The system controller 4206 is conceptually illustrated as a collection of modules, but may be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 4206 may be implemented using a single processor or an off-the-shelf PC with multiple processors, with functional operations distributed among the processors. As a further option, the modules described below may be implemented using a hybrid configuration in which certain modular functions are performed using dedicated hardware, while remaining modular functions are performed using an off-the-shelf PC, etc. The modules may also be implemented as software modules within a processing unit.

[0320] During operation, the communication port 4250 may transmit information (e.g., commands) to the biosensor 4212 ( FIG. 42A ) and / or the subsystems 4208, 4214, 4204 ( FIG. 42A ). In an embodiment, the communication port 4250 may output multiple sequences of pixel signals. The communication link 4234 may receive user input from the user interface 4218 ( FIG. 42A ) and transmit data or information to the user interface 4218. Data from the biosensor 4212 or the subsystems 4208, 4214, 4204 may be processed in real time by the system controller 4206 during a bioassay session. Additionally or alternatively, the data may be temporarily stored in system memory during a bioassay session and processed slower than real time or for offline operation.

[0321] As shown in FIG. 42B, the system controller 4206 may include multiple modules 4226-4248 in communication with a main control module 4224 along with a central processing unit (CPU) 4252. The main control module 4224 may be in communication with a user interface 4218 (FIG. 42A). While the modules 4226-4248 are shown in direct communication with the main control module 4224, the modules 4226-4248 may also be in direct communication with each other, the user interface 4218, and the biosensor 4212. The modules 4226-4248 may also be in communication with the main control module 4224 through other modules.

[0322] The plurality of modules 4226-4248 includes system modules 4228-4232, 4226 that communicate with subsystems 4208, 4214, 4204, and 4216, respectively. The fluid control module 4228 may communicate with the fluid control system 4208 to control valves and flow sensors in the fluid network to control the flow of one or more fluids through the fluid network. The fluid storage module 4230 may notify a user when fluid is low or when a waste reservoir is at or near full capacity. The fluid storage module 4230 may also communicate with a temperature control module 4232 so that fluids can be stored at a desired temperature. The illumination module 4226 may communicate with the illumination system 4216 to illuminate the reaction sites at specified times during a protocol, such as after a desired reaction (e.g., a binding event) has occurred. In some embodiments, the illumination module 4226 can communicate with the illumination system 4216 to illuminate the reaction sites at a specified angle.

[0323] The plurality of modules 4226-4248 may also include a device module 4236 that communicates with the biosensor 4212 and an identification module 4238 that determines identification information associated with the biosensor 4212. The device module 4236 may, for example, communicate with the system receptacle 4210 to confirm that the biosensor has established electrical and fluidic connection with the sequencing system 4200A. The identification module 4238 may receive a signal that identifies the biosensor 4212. The identification module 4238 may use the identification information of the biosensor 4212 to provide other information to the user. For example, the identification module 4238 may determine and subsequently display the lot number, manufacturing date, or recommended protocol for operating the biosensor 4212.

[0324] The plurality of modules 4226-4248 also includes an analysis module 4244 (also referred to as a signal processing module or signal processor) that receives and analyzes signal data (e.g., image data) from the biosensor 4212. The analysis module 4244 includes memory (e.g., RAM or flash) for storing the detection / image data. The detection data can include multiple sequences of pixel signals, such that sequences of pixel signals from each of millions of sensors (or pixels) can be detected over many base call cycles. The signal data can be stored for subsequent analysis or transmitted to the user interface 4218 to display desired information to the user. In some embodiments, the signal data can be processed by a solid-state image sensor (e.g., a CMOS image sensor) before the analysis module 4244 receives the signal data.

[0325] Analysis module 4244 is configured to acquire image data from the photodetector in each of a plurality of sequencing cycles. The image data is derived from the luminescence signals detected by the photodetector and processes the image data for each of the plurality of sequencing cycles via neural network-based base caller 2900 to generate base calls for at least some of the clusters in each of the plurality of sequencing cycles. The photodetector may be part of one or more overhead cameras (e.g., a CCD camera in an Illumina GAIIx that takes images of the clusters on biosensor 4212 from above) or may be part of biosensor 4212 itself (e.g., a CMOS image sensor in an Illumina iSeq that is below the clusters on biosensor 4212 and takes images of the clusters from the bottom).

[0326] The output of the photodetectors is a sequencing image showing the intensity emissions of each cluster and their surrounding background. The sequencing image shows the intensity emissions generated as a result of incorporating nucleotides into a sequence during sequencing. The intensity emissions come from the associated clusters and their surrounding background. The sequencing image is stored in memory 4248.

[0327] Protocol modules 4240 and 4242 communicate with main control module 4224 to control the operation of subsystems 4208, 4214, and 4204 in implementing a predetermined assay protocol. Protocol modules 4240 and 4242 may include instruction sets for instructing sequencing system 4200A to perform specific operations according to a predetermined protocol. As shown, the protocol module may be a sequencing-by-synthesis (SBS) module 4240 configured to issue various commands to execute a sequencing-by-synthesis process. In SBS, the extension of nucleic acid primers along a nucleic acid template is monitored to determine the nucleotide sequence in the template. The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase enzyme) or ligation (e.g., catalyzed by a ligase enzyme). In certain polymer-based SBS implementations, fluorescently labeled nucleotides are added to primers (thereby extending the primers) in a template-dependent manner, such that detection of the order and type of nucleotides added to the primers can be used to determine the sequence of the template. For example, to initiate the first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc. can be delivered into / through a flow cell containing an array of nucleic acid templates. The nucleic acid templates may be located at corresponding reaction sites. Primer extension can detect incorporated labeled nucleotides through an imaging event, and these reaction sites can be detected. During the imaging event, an illumination system 4216 can provide excitation light to the reaction sites. Optionally, the nucleotides can further include a reversible termination feature that terminates further primer extension once the nucleotide is added to the primer. For example, a nucleotide analog with a reversible terminator moiety can be added to the primer so that further extension does not occur until a deblocking agent is delivered to remove the moiety.Thus, in another implementation using reversible termination, a command can be given to deliver a deblocking reagent to the flow cell (before or after detection occurs). One or more commands can be given to provide washing between various delivery steps. The cycle is then repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary sequencing techniques are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, U.S. Patent No. 7,057,026, WO 91 / 06678, WO 07 / 123744, U.S. Patent Nos. 7,329,492, 7,211,414, 7,315,019, 7,405,281, and U.S. Patent Application Publication No. 2008 / 014708082, each of which is incorporated herein by reference.

[0328] In the nucleotide delivery step of the SBS cycle, any single type of nucleotide can be delivered at a time, or multiple different nucleotide types (e.g., A, C, T, and G together) can be delivered. In nucleotide delivery configurations where only a single type of nucleotide is present at a time, different nucleotides do not need to have distinct labels because they can be distinguished based on the temporal separation inherent in individualized delivery. Thus, a sequencing method or apparatus can use single-color detection. For example, the excitation source need only provide excitation at a single wavelength or a single wavelength range. In nucleotide delivery configurations where delivery results in multiple different nucleotides being present in the flow cell at a given time, the sites at which different nucleotide types incorporate can be distinguished based on the different fluorescent labels attached to each nucleotide type in the mixture. For example, four different nucleotides, each bearing one of four different fluorophores, can be used. In one embodiment, the four different fluorophores can be distinguished using excitation in four different regions of the spectrum. For example, four different excitation radiation sources can be used. Alternatively, fewer than four different excitation sources can be used, but optical filtering of the excitation radiation from a single source can be used to generate different excitation radiation ranges in the flow cell.

[0329] In some implementations, fewer than four different colors can be detected in a mixture having four different nucleotides. For example, pairs of nucleotides can be detected at the same wavelength but can be distinguished based on differences in intensity for one member of the pair, or based on a change to one member of the pair (e.g., through chemical modification, photochemical modification, or physical modification) that causes a distinct signal to appear or disappear compared to the signal detected for the other member of the pair. Exemplary devices and methods for distinguishing four different nucleotides using detection of fewer than four colors are described, for example, in U.S. Patent Application Nos. 61 / 538,294 and 61 / 619,878, which are incorporated herein by reference in their entireties. U.S. Patent Application No. 13 / 624,200, filed September 21, 2012, is incorporated herein by reference in its entirety.

[0330] The multiple protocol modules may also include a sample preparation (or generation) module 4242 configured to issue commands to the fluidic control system 4208 and the temperature control system 4204 to amplify the product in the biosensor 4212. For example, the biosensor 4212 may be coupled to a sequencing system 4200A. The amplification module 4242 can issue instructions to the fluidic control system 4208 to deliver the necessary amplification components to a reaction chamber in the biosensor 4212. In other implementations, the reaction site may already contain some components for amplification, such as template DNA and / or primers. After delivering the amplification components to the reaction chamber, the amplification module 4242 can instruct the temperature control system 4204 to cycle through different temperature steps according to a known amplification protocol. In some implementations, amplification and / or nucleotide incorporation are performed isothermally.

[0331] The SBS module 4240 can issue commands to perform bridge PCR, in which clusters of clonal amplicons are formed over localized regions within the flow cell channel. After generating amplicons via bridge PCR, the amplicons may be "linearized" to create single-stranded template DNA, and sstDNA and sequencing primers may be hybridized to universal sequences flanking the region of interest. For example, reversible terminator-based sequencing by synthesis methods can be used, as described above or as follows.

[0332] Each base calling or sequencing cycle can extend the sstDNA by a single base, which can be accomplished, for example, by using a modified DNA polymerase and a mixture of four types of nucleotides. Different types of nucleotides can have unique fluorescent labels, and each nucleotide can further have a reversible terminator that allows only a single base to be incorporated in each cycle. After a single base is added to the sstDNA, excitation light can be incident on the reaction site and fluorescent emission can be detected. After detection, the fluorescent label and terminator can be chemically cleaved from the sstDNA. Another similar base calling or sequencing cycle can be as follows: In such a sequencing protocol, the SBS module 4240 can instruct the fluid control system 4208 to direct the flow of reagents and enzyme solutions through the biosensor 4212. Exemplary reversible terminator-based SBS methods that can be utilized with the devices and methods described herein are described in U.S. Patent Application Publication No. 2007 / 0166705(A1), U.S. Patent Application Publication No. 2006 / 0188901(A1), U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439(A1), U.S. Patent Application Publication No. 2006 / 02814714709(A1), WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 014700900(A1), WO 06 / 08B199, and WO 07 / 01470251, each of which is incorporated by reference in its entirety. Exemplary reagents for reversible terminator-based SBS are described in U.S. Pat. Nos. 7,541,444, 7,057,026, 7,414,14716, 7,427,673, 7,566,537, 7,592,435, and WO 07 / 14835368, each of which is incorporated herein by reference.

[0333] In some implementations, the amplification and SBS modules may operate in a single assay protocol, for example, template nucleic acids are amplified and subsequently sequenced within the same cartridge.

[0334] The sequencing system 4200A may also allow the user to reconfigure the assay protocol. For example, the sequencing system 4200A may provide the user with options through the user interface 4218 to modify the determined protocol. For example, if it is determined that the biosensor 4212 will be used for amplification, the sequencing system 4200A may request the temperature of the annealing cycle. Furthermore, the sequencing system 4200A may issue a warning to the user if the user provides user input that is not generally accepted for the selected assay protocol.

[0335] In an embodiment, biosensor 4212 includes millions of sensors (or pixels), each of which generates multiple sequences of pixel signals over successive base call cycles. Analysis module 4244 detects the multiple sequences of pixel signals and attributes them to corresponding sensors (or pixels) according to the row-wise and / or column-wise location of the sensors on the array of sensors.

[0336] FIG. 42C is a simplified block diagram of a system for analyzing sensor data, such as base call sensor output, from a sequencing system 4200A. In the example of FIG. 42C, the system includes a configurable processor 4246. The configurable processor 4246 can execute a base caller (e.g., neural network-based base caller 2900) in coordination with a runtime program / logic 4280 executed by a central processing unit (CPU) 4252 (i.e., a host processor). The sequencing system 4200A includes a biosensor 4212 and a flow cell. The flow cell may include one or more tiles in which clusters of genetic material are exposed to a series of cluster flows that are used to trigger reactions within the clusters to identify bases in the genetic material. A sensor senses the reaction for each cycle of sequencing in each tile of the flow cell to provide tile data. Genetic sequencing is a data-intensive operation that converts base call sensor data into a sequence of base calls for each group of genetic material sensed during the base calling operation.

[0337] The system in this example includes a CPU 4252 that executes runtime program / logic 4280 for coordinating base calling operations, and memory 4248B that stores sequences of arrays of tile data, base call reads generated by the base calling operations, and other information used in the base calling operations. In this figure, the system also includes memory 4248A that stores configuration file(s), e.g., FPGA bit files, and neural network model parameters used to configure and reconfigure configurable processor 4246 and to run the neural network. Sequencing system 4200A may include programs for configuring the configurable processor, and in some embodiments, may include a reconfigurable processor that runs the neural network.

[0338] Sequencing system 4200A is coupled to configurable processor 4246 by bus 4289. Bus 4289 can be implemented using high-throughput technology, such as bus technology compatible with the PCIe (Peripheral Component Interconnect Express) standard currently maintained and developed by the PCI-SIG (PCI Special Interest Group) standard. Also, in this embodiment, memory 4248A is coupled to configurable processor 4246 by bus 4293. Memory 4248A can be on-board memory disposed on a circuit board having configurable processor 4246. Memory 4248A is used for high-speed access by configurable processor 4246 of working data used in base calling operations. Bus 4293 can also be implemented using high-throughput technology, such as bus technology compatible with the PCIe standard.

[0339] Configurable processors, including field programmable gate arrays (FPGAs), coarse-grained configurable reconfigurable arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than can be achieved using general-purpose processors running computer programs. Configuring a configurable processor involves compiling a functional description to generate a configuration file, sometimes called a bitstream or bitfile, and distributing the configuration file to configurable elements on the processor. The configuration file configures the circuit to set dataflow patterns, including the use of distributed memory and other on-chip memory resources, lookup table contents, the operation of configurable logic blocks, and configurable execution units such as configurable interconnects and other elements of the configurable array. A configuration file is reconfigurable if it can be changed in the field by modifying a loaded configuration file. For example, the configuration file may be stored in a volatile SRAM element, a non-volatile read-write memory element, or distributed among an array of configurable elements on a configurable or reconfigurable processor. Various commercially available configurable processors are suitable for use in basecall operations as described herein.Examples include Google's Tensor Processing Unit (TPU)™, GX4 Rackmount Series™, GX9 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ (Snapdragon processors™), NVIDIA Volta™, NVIDIA DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, Arm DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, Xilinx Alveo™ U200, Xilinx Alveo™ U2190, Xilinx These include the Alveo™ U280, the Intel / Altera Stratix™ GX2800, the Intel / Altera Stratix™ GX2800, and the Intel Stratix™ GX10M. In some embodiments, the host CPU may be implemented on the same integrated circuit as the configurable processor.

[0340] The embodiments described herein implement the neural network-based basis caller 2900 using a configurable processor 4246. The configuration file for the configurable processor 4246 may be implemented by specifying the logic functions to be performed using a high-level description language HDL or a register-transfer level RTL language specification. This specification can be compiled using resources designed for a selected configurable processor to generate the configuration file. The same or similar specifications can be compiled to generate the design of an application-specific integrated circuit, which may not be a configurable processor.

[0341] Accordingly, alternative examples of the configurable processor 4246 in all embodiments described herein include a configured processor including an application specific ASIC or dedicated integrated circuit or set of integrated circuits, or a system on a chip SOC device, or a system on a chip SOC device, or a graphics processing unit (GPU) processor or a Coarse-Grained Reconfigurable Architecture (CGRA) processor configured to perform neural network-based base call operations as described herein.

[0342] In general, the configurable processors and configured processors described herein that are configured to perform neural network operations are referred to herein as neural network processors.

[0343] Configurable processor 4246, in this example, is configured by a configuration file loaded using a program executed by CPU 4252 or by other source to configure an array of configurable elements 4291 (e.g., configuration logic blocks (CLBs), such as look up tables (LUTs), flip-flops, arithmetic processing units (PMUs), and compute memory units (CMUs), configurable I / O blocks, programmable interconnects) to perform base calling functions. In this example, the configuration includes data flow logic 4297 coupled to buses 4289 and 4293, which performs the function of distributing data and control parameters among elements used in base calling operations.

[0344] Configurable processor 4246 is also configured with dataflow logic 4297 to execute neural network-based base caller 2900. Logic 4297 includes multi-cycle execution clusters (e.g., 4279), which in this example include execution cluster 1 through execution cluster X. The number of multi-cycle execution clusters can be selected according to tradeoffs involving the desired throughput of operation and the available resources on configurable processor 4246.

[0345] The multi-cycle execution clusters are coupled to the data flow logic 4297 by data flow paths 4299 implemented using configurable interconnect and memory resources on the configurable processor 4246. The multi-cycle execution clusters are also coupled to the data flow logic 4297 by control paths 4295 implemented using configurable interconnect and memory resources, for example on the configurable processor 4246, to provide control signals indicating available execution clusters, provisions for providing input units to the available execution clusters for execution of operations of the neural network based base caller 2900, provisions for providing learned parameters of the neural network based base caller 2900, provisions for providing output patches of base call classification data, and other control data used in the execution of the neural network based base caller 2900.

[0346] The configurable processor 4246 is configured to execute the operations of the neural network-based base caller 2900 using the learned parameters to generate classification data for the detection cycles of the base calling operation. The neural network-based base caller 2900 executes the operations to generate classification data for the subject detection cycles of the base calling operation. The neural network-based base caller 2900 operates in a sequence including a number N of arrays of tile data from each of the N detection cycles, where the N detection cycles provide sensor data for different base calling operations for one base position per operation in the time sequence described herein. Optionally, some of the N detection cycles can be removed from the sequence as needed according to the particular neural network model being implemented. The number N can be any number greater than 1. In some examples described herein, the detection cycles of the N detection cycles represent a set of detection cycles for at least one detection cycle preceding the subject detection cycle and at least one detection cycle following the subject cycle in time sequence. Examples described herein include an integer number N of 5 or greater.

[0347] The data flow logic 4297 is configured to use an input unit for a given operation that includes tile data for an array of N spatially aligned patches to move the tile data and at least some of the learned parameters of the model parameters from memory 4248A to the configurable processor 4246 for operation of the neural network-based bass caller 2900. The input unit can be moved by a direct memory access operation in a single DMA operation, or in smaller units that move during available time slots in coordination with the execution of the deployed neural network.

[0348] The tile data of the sensing cycles described herein can include an array of sensor data having one or more features. For example, the sensor data can include two images analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The tile data can also include metadata about the images and sensors. For example, in a base calling implementation, the tile data can include information about the alignment of the images with clusters, such as distance from center information indicating the distance of each pixel in the array of sensor data from the center of the group of genetic material on the tile.

[0349] As described below, during execution of the neural network-based base caller 2900, the tile data may also include data generated during execution of the neural network-based base caller 2900. This data is referred to as intermediate data that can be reused rather than recalculated during operation of the neural network-based base caller 2900. For example, during execution of the neural network-based base caller 2900, the data flow logic 4297 may write the intermediate data to memory 4248A in place of sensor data for a given patch of the array of tile data. Such implementations are described in more detail below.

[0350] As shown, a system for analyzing base calling sensor output is described that includes a memory (e.g., 4248A) accessible by runtime program / logic 4280 that stores tile data including sensor data for tiles from a sensing cycle of a base calling operation. The system also includes a neural network processor, such as configurable processor 4246, with access to the memory. The neural network processor is configured to perform neural network operations using trained parameters to generate classification data for the sensing cycle. As described herein, the neural network operations operate on a sequence of N arrays of tile data from each sensing cycle of the N sensing cycles comprising a subject cycle to generate classification data for the subject cycle. Data flow logic 4297 is provided to move the tile data and trained parameters from the memory to the neural network processor for execution of the neural network using input units including data for the N arrays of spatially aligned patches from each sensing cycle of the N sensing cycles.

[0351] Also described is a system in which a neural network processor has access to a memory and includes a plurality of execution clusters, the execution clusters being configured to execute a neural network. Data flow logic 1997 has access to the memory and to an execution cluster among the plurality of execution clusters to provide an input unit of tile data to an available execution cluster among the plurality of execution clusters, the input unit including N spatially aligned patches of an array of tile data from each detection cycle including a subject detection cycle, and cause the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patch of the subject detection cycle, where N is greater than 1.

[0352] Figure 43A is a simplified diagram illustrating aspects of a base calling operation, including the functionality of a runtime program (e.g., runtime logic 4280) executed by a host processor. In this diagram, the output of an image sensor from a flow cell is provided on line 4300 to an image processing thread 4301, which can perform processes on the image, such as aligning and positioning individual tiles within an array of sensor data and resampling the image, which can be used by a process to calculate a tile cluster mask for each tile within the flow cell, which can be used by a process to identify pixels within the array of sensor data that correspond to clusters of genetic material on the corresponding tile of the flow cell. The output of image processing thread 4301 is provided on line 4302 to dispatch logic 4303 within the CPU, which transfers the array of tile data to neural network processor hardware 4307, such as configurable processor 1946 of Figure 19C, on high-speed bus 4304 or on high-speed bus 4306 to data cache 4305 (e.g., SSD storage), depending on the state of the base calling operation. The processed and transformed images may be stored on a data cache 4305 to track previously used cycles. Hardware 4307 returns the classification data output by the neural network to dispatch logic 4303, which passes the information to data cache 4305 or on line 4308 to thread 4309, which can perform base calling and quality score calculations using the classification data and arrange the data in a standard format for base called reads. The output of thread 4309, which performs base calling and quality score calculations, is provided on line 4310 to thread 4311, which aggregates the base called reads, performs other operations such as data compression, and writes the resulting base calling output to a specified destination for consumption by the customer.

[0353] In some implementations, the host may include a thread (not shown) that performs final processing of the output of the hardware 4307 supporting the neural network. For example, the hardware 4307 may provide classification data output from the final layer of a multi-cluster neural network. The host processor may perform output activation functions, such as a softmax function, over the classification data to populate the data used by the base calling and quality score thread 4302. The host processor may also perform input operations (not shown), such as batch normalization of the tile data before input to the hardware 4307.

[0354] FIG. 43B is a simplified diagram of a configurable processor 1946 configuration such as that of FIG. 19C. In FIG. 43B, the configurable processor 1946 includes an FPGA with multiple high-speed PCIe interfaces. The FPGA is configured with a wrapper 4390 including dataflow logic 1997 as described with reference to FIG. 19C. The wrapper 4390 manages interfacing and coordination with runtime programs in the CPU via CPU communication link 4377 and communication with onboard DRAM 4399 (e.g., memory 1448A) via DRAM communication link 4397. The dataflow logic 1997 in the wrapper 4390 provides patch data obtained by traversing an array of tile data on onboard DRAM 4399 to clusters 4385, and retrieves process data 4387 from clusters 4385 and delivers it to onboard DRAM 4399 for a number N of cycles. The wrapper 4390 also manages the transfer of data between onboard DRAM 4399 and host memory for both the input array of tile data and the output patches of classification data. The wrapper transfers the patch data to the assigned cluster 4385 on line 4383. The wrapper provides the cluster 4385 with trained parameters, such as weights and biases, obtained from the onboard DRAM 4399 on line 4381. The wrapper provides the cluster 4385 with configuration and control data provided by or generated in response to a runtime program on the host on line 4379 via CPU communication link 4377. The cluster can also provide state signals to the wrapper 4390 on line 4389 that are used in conjunction with control signals from the host to manage the traversal of the array of tile data to provide spatially aligned patch data and to run a multi-cycle neural network on the patch data using the resources of the cluster 4385.

[0355] As described above, multiple clusters may reside on a single configurable processor managed by a wrapper 4390 configured to run on corresponding ones of the multiple patches of tile data. Each cluster may be configured to provide classification data for base calls in a subject sensing cycle using the tile data of multiple sensing cycles described herein.

[0356] In an example system, model data, including kernel data such as filter weights and biases, can be sent from the host CPU to the configurable processor, so that the model can be updated as a function of cycle number. Base calling operations can typically involve hundreds of sensing cycles. In some implementations, base calling operations can include paired end reads. For example, model-trained parameters can be updated every 20 cycles (or other number of cycles) or according to an update pattern implemented in a particular system and neural network model. In some implementations, where a sequence for a given string within a genetic cluster on a tile includes paired end reads that include a first portion extending downward (or upward) from a first end of the string and a second portion extending upward (or downward) from a second end of the string, the learned parameters can be updated at the transition from the first portion to the second portion.

[0357] In some embodiments, image data for multiple cycles of sensor data for a tile can be sent from the CPU to the wrapper 4390. The wrapper 4390 can optionally perform some preprocessing and transformation of the sensor data and write the information to onboard DRAM 4399. The input tile data for each sensing cycle can include an array of sensor data including 4000 x 3000 pixels or more per tile per sensing cycle, with two features representing the colors of two images of the tile and including one or two bytes per pixel. For embodiments where the number N is three sensing cycles used in each operation of the multi-cycle neural network, the array of tile data for each operation of the multi-cycle neural network can consume several hundred megabytes per number. In some implementations of the system, the tile data also includes an array of distance-from-cluster center (DFC) data stored once per tile, or other types of metadata about the sensor data and tile.

[0358] In operation, if a multi-cycle cluster is available, the wrapper assigns the patch to the cluster. The wrapper fetches the next patch of tile data for the cross section of the tile and sends it to the assigned cluster along with the appropriate control and configuration information. The cluster can be configured with enough memory on the configurable processor to have enough memory to hold the patch of data, including the patch, from multiple cycles in some systems being processed in place, and in various implementations are processed using ping-pong buffer or raster scan techniques.

[0359] When an assigned cluster completes its operation of the neural network for the current patch and generates an output patch, it signals the wrapper. The wrapper either reads the output patch from the assigned cluster, or the assigned cluster pushes the data to the wrapper. The wrapper will then assemble the output patch for the processed tile in DRAM 4399. Once processing of the entire tile is complete and the output patch of data has been transferred to DRAM, the wrapper sends the processed output array back to the host / CPU in a specific format. In some embodiments, the on-board DRAM 4399 is managed by memory management logic within the wrapper 4390. The runtime program can control the sequencing operations to complete analysis of all tile data arrays for every cycle, operating in a continuous flow to provide real-time analysis.

[0360] FIG. 44 illustrates one embodiment of determining state data on the CPU and loading the state data from the CPU into the FPGA for base calling.

[0361] Terminology and Additional Embodiments Base calling involves incorporating or attaching a fluorescently labeled tag with the analyte. The analyte may be a nucleotide or oligonucleotide, and the tag may be a specific nucleotide type (A, C, T, or G). Excitation light is directed at the tagged analyte, causing the tag to emit a detectable fluorescent signal or intensity emission. The intensity emission indicates the photons emitted by the excitation tag chemically bound to the analyte.

[0362] Throughout this application, including the claims, when "images, image data, or image regions showing the intensity radiation of analytes and their surrounding background are used, they refer to the intensity radiation of tags attached to the analytes. Those skilled in the art will understand that the intensity radiation of an attached tag represents or corresponds to the intensity radiation of the analyte to which the tag is attached, and are therefore used interchangeably. Similarly, a characteristic of an analyte refers to a characteristic of the tag attached to the analyte, or the intensity radiation from the attached tag. For example, the center of the analyte refers to the center of the intensity radiation emitted by the tag attached to the analyte. In another example, the background around the analyte refers to the background around the intensity radiation emitted by the tag attached to the analyte.

[0363] All literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and web pages, regardless of the form such literature and similar materials may be in, are expressly incorporated by reference in their entirety. In the event that one or more of the incorporated literature and similar materials differs from or contradicts this application in terms of, but not limited to, defined terms, term usage, described techniques, etc., this application controls.

[0364] The disclosed technology uses neural networks to improve the quality and quantity of nucleic acid sequence information that can be obtained from a nucleic acid template or its complement, e.g., a nucleic acid sample, such as a DNA or RNA polynucleotide or other nucleic acid sample. Accordingly, certain implementations of the disclosed technology provide higher throughput polynucleotide sequencing, e.g., higher rates of collection of DNA or RNA sequence data, greater efficiency in sequence data collection, and / or lower costs of obtaining such sequence data, compared to previously available methods.

[0365] The disclosed technology uses neural networks to identify the centers of solid-phase nucleic acid clusters and analyze optical signals generated during sequencing of such clusters, unambiguously distinguishing between adjacent, neighboring, or overlapping clusters and assigning sequencing signals to single, discrete source clusters. These and related embodiments thus enable the recovery of meaningful information, such as sequence data, from regions of high-density cluster arrays, where useful information may not have been previously obtained from such regions due to confounding effects of overlapping or closely spaced neighboring clusters, including the effects of overlapping signals (e.g., as used in nucleic acid sequencing).

[0366] As described in more detail below, in certain embodiments, compositions are provided that include a solid support immobilized with one or more nucleic acid clusters, as provided herein. Each cluster contains multiple immobilized nucleic acids of the same sequence and has a distinct center with a detectable central label as provided herein, which is distinguishable from the nucleic acids immobilized in the surrounding area within the cluster. Also described herein are methods for making and using such clusters with distinct centers.

[0367] Embodiments of the present disclosure will find use in many situations where benefits derive from the ability to identify, determine, annotate, record, or otherwise assign the location of a substantially central location within a cluster, including high-throughput nucleic acid sequencing, development of image analysis algorithms for assigning optical or other signals to individual source clusters, and other applications where recognition of the center of immobilized nucleic acid clusters is desirable and beneficial.

[0368] In certain embodiments, the present invention contemplates methods related to high-throughput nucleic acid analysis, such as nucleic acid sequencing (e.g., "sequencing"). Exemplary high-throughput nucleic acid analyses include, but are not limited to, de novo sequencing, resequencing, whole genome sequencing, gene expression analysis, gene expression monitoring, epigenetics analysis, genome methylation analysis, allele-specific primer extension (APSE), genetic diversity profiling, whole genome polymorphism discovery and analysis, single nucleotide polymorphism analysis, hybridization-based sequencing, and the like. Those skilled in the art will understand that a variety of different nucleic acids can be analyzed using the methods and compositions of the present invention.

[0369] Although implementations of the present invention are described in the context of nucleic acid sequencing, they are applicable in any field in which image data acquired at different times, spatial locations, or other temporal or physical aspects are analyzed. For example, the methods and systems described herein are useful in the fields of molecular biology and cell biology, where image data from microarrays, biological specimens, cells, organisms, etc., are acquired and analyzed at different times or perspectives. Images can be obtained using any number of techniques known in the art, including, but not limited to, fluorescence microscopy, optical microscopy, confocal microscopy, optical imaging, magnetic resonance imaging, tomographic scanning, etc. As another example, the methods and systems described herein can be applied when image data acquired by surveillance, aerial, or satellite imaging techniques, etc., are acquired and analyzed at different times or perspectives. The methods and systems are particularly useful for analyzing images acquired within a field of view, in which observed specimens remain in the same location relative to each other within the field of view. However, specimens may have different properties in separate images; for example, specimens may appear different in separate images of the field of view. For example, the analyte may appear to be different in color for a given analyte detected in different images, may show changes in the intensity of the signal detected for a given analyte in different images, or even the appearance of a signal for a given analyte in one image and the disappearance of the signal for that analyte in another image.

[0370] The examples described herein may be used in a variety of biological or chemical processes and systems for academic or commercial analysis. More specifically, the examples described herein may be used in a variety of processes and systems in which it is desirable to detect an event, characteristic, quality, or property indicative of a specified reaction. For example, the examples described herein include optical detection devices, biosensors, and components thereof, as well as bioassay systems operating with biosensors. In some embodiments, the devices, biosensors, and systems may include a flow cell and one or more optical sensors coupled (removably or fixedly) together in a substantially monolithic structure.

[0371] The devices, biosensors, and bioassay systems may be configured to perform multiple designated reactions that can be detected individually or collectively. The devices, biosensors, and bioassay systems may be configured to perform multiple cycles in which multiple designated reactions occur in parallel. For example, the devices, biosensors, and bioassay systems may be used to sequence high-density sequences of DNA features through repeated cycles of enzymatic manipulation and light or image detection / capture. Thus, the devices, biosensors, and bioassay systems (e.g., via one or more cartridges) may include one or more microfluidic channels that deliver reagents or other reaction components into the reaction solution, biosensors, and bioassay systems. In some examples, the reaction solution may be substantially acidic, such as comprising a pH of about 5 or less, or about 4 or less, or about 3 or less. In some other examples, the reaction solution may be substantially alkaline / basic, such as comprising a pH of about 8 or more, or about 9 or more, or about 10 or more. As used herein, the term "acidic" and grammatical variants thereof refer to pH values ​​less than about 7, and the terms "basic," "alkaline," and grammatical variants thereof refer to pH values ​​greater than about 7.

[0372] In some embodiments, the reaction sites are provided or spaced in a predetermined manner, such as a uniform or repeating pattern. In some other embodiments, the reaction sites are randomly distributed. Each of the reaction sites can be associated with one or more light guides and one or more light sensors that detect light from the associated reaction site. In some embodiments, the reaction sites are located within a reaction recess or chamber that can at least partially compartmentalize a designated reaction.

[0373] As used herein, a "designated reaction" includes a change in at least one of the chemical, electrical, physical, or optical properties (or qualities) of a chemical or biological substance of interest, such as an analyte of interest. In certain examples, the designated reaction is a positive binding event, such as the incorporation of a fluorescently labeled biomolecule with a fluorescently labeled biomolecule. More generally, the designated reaction may be a chemical conversion, chemical change, or chemical interaction. The designated reaction may also be a change in an electrical property. In certain examples, the designated reaction includes the incorporation of an analyte with a fluorescently labeled molecule. The analyte may be an oligonucleotide, and the fluorescently labeled molecule may be a nucleotide. The designated reaction may be detected when excitation light is directed at the oligonucleotide bearing the labeled nucleotide and the fluorophore emits a detectable fluorescent signal. In alternative examples, the detected fluorescence is the result of chemiluminescence or bioluminescence. The specified reaction can also, for example, increase fluorescence (or Förster) resonance energy transfer (FRET) by bringing a donor fluorophore into close proximity with an acceptor fluorophore, decrease FRET by separating the donor and acceptor fluorophores, increase fluorescence by separating a quencher from a fluorophore, or decrease fluorescence by bringing a quencher and a fluorophore into close proximity.

[0374] As used herein, "reaction solution," "reaction component," or "reactant" includes any substance that can be used to obtain at least one specified reaction. For example, potential reaction components include, for example, reagents, enzymes, samples, other biomolecules, and buffers. A reaction component may be delivered to a reaction site in solution and / or immobilized at the reaction site. A reaction component may interact directly or indirectly with another substance, such as an analyte of interest, immobilized at the reaction site. As noted above, the reaction solution may be substantially acidic (i.e., comprising a relatively high degree of acidity) (e.g., including a pH of about 5 or less, a pH of about 4 or less), or a pH of about 3 or less, or substantially alkaline / basic (i.e., comprising a relatively high degree of alkaline / basicity) (e.g., including a pH of about 8 or more, a pH of about 9 or more, or a pH of about 10 or more).

[0375] As used herein, the term "reaction site" refers to a localized region where at least one specified reaction can occur. A reaction site may include a support surface of a reaction structure or substrate onto which a substance can be immobilized. For example, a reaction site may include the surface of a reaction structure (which may be positioned within a channel of a flow cell) having reaction components thereon, e.g., colonies of nucleic acids thereon. In some such examples, the nucleic acids in the colonies have the same sequence, e.g., are clonal copies of a single-stranded or double-stranded template. However, in some examples, a reaction site may contain only a single nucleic acid molecule, e.g., in single-stranded or double-stranded form.

[0376] The multiple reaction sites may be randomly distributed along the reaction structure or may be arranged in a predetermined manner (e.g., parallel in a matrix such as a microarray). A reaction site may also include a reaction chamber or recess that at least partially defines a spatial region or volume configured to compartmentalize a specified reaction. As used herein, the term "reaction chamber" or "reaction recess" includes a defined spatial region of a support structure (often in fluid communication with a flow path). A reaction recess may be at least partially isolated from the surrounding environment or spatial region. For example, multiple reaction recesses may be separated from each other by a shared wall, such as a detection surface. As a more specific example, a reaction recess may be a nanocell that includes a depression, well, groove, cavity, or depression defined by the inner surface of the detection surface, and may have an opening or aperture (i.e., an open side) to allow the nanocell to be in fluid communication with the flow path.

[0377] In some embodiments, the reaction recess of a reaction structure is sized and shaped relative to a solid (including a semi-solid) so that the solid can be fully or partially inserted therein. For example, the reaction recess may be sized and shaped to accommodate a capture bead. The capture bead may have clonally amplified DNA or other material thereon. Alternatively, the reaction recess may be sized and shaped to receive an approximate number of beads or solid substrates. As another example, the reaction recess may be filled with a porous gel or material configured to control diffusion or filter fluids or solutions that may flow into the reaction recess.

[0378] In some embodiments, an optical sensor (e.g., a photodiode) is associated with a corresponding reaction site. The optical sensor associated with a reaction site is configured to detect light emission from the associated reaction site via at least one light guide when a designated reaction occurs at the associated reaction site. In some cases, multiple optical sensors (e.g., several pixels of a light detection or camera device) may be associated with a single reaction site. In other cases, a single optical sensor (e.g., a single pixel) may be associated with a single reaction site or with a group of reaction sites. The optical sensors, reaction sites, and other features of the biosensor may be configured such that at least a portion of the light is directly detected by the optical sensor without being reflected.

[0379] As used herein, "biological or chemical substances" include biomolecules, subject samples, subject analytes, and other chemical compounds. Biological or chemical substances may be used to detect, identify, or analyze other chemical compounds, or to act as intermediaries for studying or analyzing other chemical compounds. In certain examples, biological or chemical substances include biomolecules. As used herein, "biomolecules" include at least one of biopolymers, nucleotides, nucleic acids, polynucleotides, oligonucleotides, proteins, enzymes, polypeptides, antibodies, antigens, ligands, receptors, polysaccharides, carbohydrates, polyphosphates, cells, tissues, organisms, or fragments thereof, or any other biologically active chemical compounds, such as analogs or mimetics of the foregoing species. In a further example, a biological or chemical substance or biomolecule detects the product of another reaction, such as an enzyme or reagent, e.g., the product of an enzyme or reagent, such as an enzyme or reagent used to detect pyrophosphate in a pyrosequencing reaction. Enzymes and reagents useful for pyrophosphate detection are described, for example, in U.S. Patent Publication No. 2005 / 0244870 A1, which is incorporated by reference in its entirety.

[0380] The biomolecules, samples, and biological substances or chemicals may be naturally occurring or synthetic and may be suspended in a solution or mixture within the reaction wells or regions. The biomolecules, samples, and biological substances or chemicals may also be bound to a solid phase or gel material. The biomolecules, samples, and biological substances or chemicals may also include pharmaceutical compositions. In some cases, the biomolecules, samples, and biological substances or chemicals of interest may be referred to as targets, probes, or analytes.

[0381] As used herein, "biosensor" includes devices including a reaction structure with multiple reaction sites configured to detect a specified reaction occurring at or near the reaction site. A biosensor may include a solid-state photodetector or "imaging" device (e.g., a CCD or CMOS photodetector) and, optionally, a flow cell attached thereto. The flow cell may include at least one flow path in fluid communication with the reaction site. As one specific example, the biosensor is configured to be fluidically and electrically coupled to a bioassay system. The bioassay system may deliver reaction solutions to the reaction sites according to a predetermined protocol (e.g., sequence number synthesis) and perform multiple imaging events. For example, the bioassay system may flow reaction solutions along the reaction sites. At least one of the reaction solutions may contain four types of nucleotides with the same or different fluorescent labels. The nucleotides may bind to corresponding oligonucleotides, etc., in the reaction sites. The bioassay system can then illuminate the reaction sites using an excitation light source (e.g., a solid-state light source such as a light-emitting diode (LED)). The excitation light may have a predetermined wavelength or multiple wavelengths, including a range of wavelengths. Fluorescent labels excited by incident excitation light can provide an emission signal (e.g., of a different wavelength or wavelengths of light than the excitation light, and potentially different from each other) that can be detected by a light sensor.

[0382] As used herein, the term "immobilized," when used with reference to a biomolecule or biological substance or chemical, includes substantially attaching the biomolecule or biological substance or chemical to a surface, such as the detection surface of an optical detection device or a reaction structure. For example, a biomolecule or biological substance or chemical may be immobilized on the surface of a reaction structure using adsorption techniques, including non-covalent bonding (e.g., electrostatic forces, van der Waals, and hydrophobic interfacial dehydration), as well as covalent bonding techniques in which functional groups or linkers facilitate binding of the biomolecule to the surface. Immobilizing a biomolecule or biological substance or chemical on a surface may be based on the properties of the surface, the liquid medium carrying the biomolecule or biological substance or chemical, and the properties of the biomolecule or biological substance or chemical itself. In some cases, the surface may be functionalized (e.g., chemically or physically modified) to facilitate immobilization of a biomolecule (or biological substance or chemical) on the surface.

[0383] In some embodiments, nucleic acids can be immobilized on a reaction structure, such as the surface of a reaction well. In certain embodiments, the devices, biosensors, bioassay systems, and methods described herein can include the use of naturally occurring nucleotides and enzymes configured to interact with naturally occurring nucleotides. Naturally occurring nucleotides include, for example, ribonucleotides or deoxyribonucleotides. Naturally occurring nucleotides can be in monophosphate, diphosphate, or triphosphate form and can have a base selected from adenine (A), thymine (T), uracil (U), guanine (G), or cytosine (C). However, it will be understood that non-naturally occurring nucleotides, modified nucleotides, or analogs of the above nucleotides can be used.

[0384] As described above, biomolecules, biological substances, or chemicals may be immobilized at reaction sites within the reaction recesses of the reaction structure. Such biomolecules or biological substances may be physically held or immobilized within the reaction recesses by interference fitting, adhesion, covalent bonding, or entrapment. Examples of articles or solids that can be placed within the reaction recesses include polymer beads, pellets, agarose gel, powders, quantum dots, or other solids that can be compressed and / or held within the reaction chamber. In certain embodiments, the reaction recesses may be coated or filled with a hydrogel layer that can covalently bond to DNA oligonucleotides. In certain examples, nucleic acid superstructures such as DNA balls can be placed within the reaction recesses by, for example, attaching them to the inner surface of the reaction recesses or by immersing them in a liquid within the reaction recesses. DNA balls or other nucleic acid superstructures can be formed and then placed within the reaction recesses. Alternatively, DNA balls can be synthesized in situ in the reaction recesses. The substances immobilized within the reaction recesses can be in a solid, liquid, or gaseous state.

[0385] As used herein, the term "analyte" is intended to mean a point or region of a pattern that can be distinguished from other points or regions according to their relative position. An individual analyte can contain one or more molecules of a particular type. For example, an analyte can contain a single target nucleic acid molecule having a particular sequence, or an analyte can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). Different molecules that are different analytes of a pattern can be differentiated from one another according to the location of the analyte within the pattern. Exemplary analytes include wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, pads of gel material on a substrate, or channels in a substrate.

[0386] Any of a variety of target analytes to be detected, characterized, or identified can be used in the devices, systems, or methods described herein. Exemplary analytes include, but are not limited to, nucleic acids (e.g., DNA, RNA, or analogs thereof), proteins, polysaccharides, cells, antibodies, epitopes, receptors, ligands, enzymes (e.g., kinases, phosphatases, or polymerases), small molecule drug candidates, cells, viruses, organisms, etc.

[0387] The terms "analyte," "nucleic acid," "nucleic acid molecule," and "polynucleotide" are used interchangeably herein. In various embodiments, a nucleic acid may be used as a template (e.g., a nucleic acid template or a nucleic acid complement complementary to a nucleic acid template) as provided herein for a particular type of nucleic acid analysis, including, but not limited to, nucleic acid amplification, nucleic acid expression analysis, and / or nucleic acid sequencing, or a suitable combination thereof. Nucleic acids in certain embodiments include, for example, linear polymers of deoxyribonucleotides in 3'-5' phosphodiester chains, or deoxyribonucleic acid (DNA), e.g., single- and double-stranded DNA, genomic DNA, copy DNA or complementary DNA (cDNA), recombinant DNA, or any form of synthetic or modified DNA. In other embodiments, nucleic acids include, for example, linear polymers of ribonucleotides in 3'-5' phosphodiester or other linkages, such as ribonucleic acid (RNA), including single- and double-stranded RNA, messenger (mRNA), copy or complementary RNA (cRNA), alternatively spliced ​​mRNA, ribosomal RNA, small nuclear RNA (snoRNA), microRNA (miRNA), small interfering RNA (sRNA), piwi RNA (piRNA), or any form of synthetic or modified RNA. Nucleic acids used in the compositions and methods of the invention may vary in length and may be intact or full-length molecules or fragments, or smaller portions of larger nucleic acid molecules. In certain embodiments, nucleic acids may carry one or more detectable labels, as described elsewhere herein.

[0388] The terms "specimen," "cluster," "nucleic acid cluster," "nucleic acid colony," and "DNA cluster" are used interchangeably and refer to multiple copies of a nucleic acid template and / or its complement attached to a solid support. Typically, in certain preferred embodiments, a nucleic acid cluster comprises multiple copies of a template nucleic acid and / or its complement attached to a solid support via their 5' ends. The copies of the nucleic acid strands that make up a nucleic acid cluster may be in single-stranded or double-stranded form. The copies of the nucleic acid template present within a cluster may have nucleotides at corresponding positions that differ from each other due to, for example, the presence of a labeled moiety. Corresponding positions may also include analog structures with different chemical structures but similar Watson-Crick base pairing properties, such as uracil and thymine.

[0389] Colonies of nucleic acids may also be referred to as "nucleic acid clusters." Nucleic acid colonies can optionally be generated by cluster amplification or bridge amplification techniques, as described in more detail elsewhere herein. Multiple repeats of a target sequence can be present in a single nucleic acid molecule, such as a disruptor generated using a rolling circle amplification procedure.

[0390] The nucleic acid clusters of the present invention can have different shapes, sizes, and densities depending on the conditions used. For example, the clusters can be substantially circular, polyhedral, donut-shaped, or ring-shaped. The diameter of the nucleic acid cluster can be designed to be about 0.2 μm to about 6 μm, about 0.3 μm to about 4 μm, about 0.4 μm to about 3 μm, about 0.5 μm to about 2 μm, about 0.75 μm to about 1.5 μm, or any intermediate diameter. In certain embodiments, the diameter of the nucleic acid cluster is about 0.5 μm, about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The diameter of the nucleic acid cluster can be affected by numerous parameters, including, but not limited to, the number of amplification cycles performed in producing the cluster, the length of the nucleic acid template, or the density of primers attached to the surface on which the cluster is formed. The density of the nucleic acid cluster is typically less than 0.1 μm / mm 2 , 1 / mm 2 , 10 / mm 2 , 100 / mm 2 , 1,000 / mm 2 , 10,000 / mm 2 ~100,000 / mm 2 The present invention, in part, is directed to higher density nucleic acid clusters, e.g., 100,000 / mm 2 ~1,000,000 / mm 2 , and 1,000,000 / mm 2 ~10,000,000 / mm 2 Further plans are underway.

[0391] As used herein, an "analyte" is a specimen or region of interest within a field of view. When used in connection with a microarray device or other molecular analysis device, an analyte refers to a region occupied by similar or identical molecules. For example, an analyte can be an amplification oligonucleotide or any other group of polynucleotides or polypeptides having the same or similar sequence. In other embodiments, an analyte can be any element or group of elements that occupy a physical region on a sample. For example, an analyte can be a parcel of land, a body of water, etc. When analytes are imaged, each analyte has some area. Thus, in many embodiments, an analyte is not simply a single pixel.

[0392] The distance between specimens can be described in any number of ways. In some embodiments, the distance between specimens can be described from the center of one specimen to the center of another specimen. In other embodiments, the distance can be described from the edge of one specimen to the edge of another specimen, or between the outermost identifiable points of each specimen. The edge of the specimen can be described as a theoretical or actual physical boundary on the chip, or some point within the boundary of the specimen. In other embodiments, the distance can be described with respect to a fixed point on the specimen, or an image of the specimen.

[0393] Generally, some embodiments are described herein with respect to analytical methods. It will be understood that systems for performing the methods in an automated or semi-automated manner are also provided. Thus, the present disclosure provides a neural network-based template generation and base calling system, the system including a processor, a storage device, and a program for image analysis, the program including instructions for performing one or more of the methods described herein. Thus, the methods described herein can be performed, for example, on a computer having components described herein or known in the art.

[0394] The methods and systems described herein are useful for analyzing any of a variety of objects. Particularly useful objects are solid supports or solid surfaces with analytes attached. The methods and systems described herein offer advantages when used with objects having repeating patterns of analytes in the xy plane. One example is a microarray having a collection of cells, viruses, nucleic acids, proteins, antibodies, carbohydrates, small molecules (such as drug candidates), biologically active molecules, or other analytes of interest.

[0395] The number of applications of arrays containing analytes containing biological molecules such as nucleic acids and polypeptides has increased. Such microarrays typically contain deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) probes, which are specific for nucleotide sequences present in humans and other organisms. In certain applications, for example, individual DNA or RNA probes can be attached to individual analytes on the array. Test samples, such as those from known humans or organisms, can be exposed to the array so that target nucleic acids (e.g., gene fragments, mRNA, or amplicons) hybridize to complementary probes for each analyte in the array. The probes can be labeled through target-specific processes (e.g., due to labels present on the target nucleic acids or due to enzyme labels on the probes or targets present in hybridized form in the analyte). The analytes can then be examined by scanning specific light frequencies over them to identify which target nucleic acids are present in the sample.

[0396] Biological microarrays can be used for gene sequencing and similar applications. Generally, gene sequencing involves determining the order of nucleotides in a length of target nucleic acid, such as a fragment of DNA or RNA. Relatively short sequences are typically sequenced for each specimen, and the resulting sequence information can be used in various bioinformatics methods to reliably determine the sequence of many diverse lengths of genetic material from which the fragments originate. Automated computer-based algorithms for identifying signature fragments have been developed and have more recently been used in genome mapping, gene identification, and their functions. Microarrays are particularly useful for characterizing genome content because of the large number of variants present, which is an alternative to conducting numerous experiments for individual probes and targets. Microarrays are an ideal format for conducting such studies in a practical manner.

[0397] Any of a variety of analyte arrays (also referred to as "microarrays") known in the art can be used in the methods or systems described herein. A typical array contains analytes, each having an individual probe or a population of probes. In the latter case, the population of probes in each analyte is typically homogeneous, having a single type of probe. For example, in the case of nucleic acid sequences, each analyte can have multiple nucleic acid molecules, each having a common sequence. However, in some embodiments, the population in each analyte in the array can be heterogeneous. Similarly, protein sequences can have analytes having a single protein or a population of proteins, typically, but not necessarily, having the same amino acid sequence. Probes can be attached to the surface of the array, for example, by covalently linking the probe to the surface or through non-covalent interactions between the probe and the surface. In some embodiments, probes, such as nucleic acid molecules, can be attached to the surface via a gel layer, as described, for example, in U.S. Patent Application No. 13 / 784,368 and U.S. Patent Application Publication No. 2011 / 0059865(A1), each of which is incorporated herein by reference.

[0398] Exemplary arrays include, but are not limited to, BeadChip arrays available from Illumina, Inc. (San Diego, Calif.), or others, such as those described in U.S. Pat. Nos. 6,266,459, 6,355,431, 6,770,441, 6,859,570, or 7,622,294, or International Publication No. WO 00 / 63437, each of which is incorporated herein by reference, in which probes are attached to beads present on a surface (e.g., beads in wells on a surface). Further examples of commercially available microarrays that can be used include, for example, Affymetrix® GeneChip® microarrays or other microarrays synthesized according to a technique sometimes referred to as VLSIPS™ (Very Large Scale Immobilized Polymer Synthesis) technology. Spotted microarrays can also be used in methods or systems according to some embodiments of the present disclosure. An exemplary spotted microarray is the CodeLink™ Array available from Amersham Biosciences. Another useful microarray is one manufactured using inkjet printing methods, such as SurePrint™ Technology available from Agilent Technologies.

[0399] Other useful arrays include those used in nucleic acid sequencing applications.For example, the arrays with amplicons of genome fragments (often referred to as clusters) are described in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, International Publication No. 91 / 06678, International Publication No. 07 / 123744, US Patent No. 7,329,492, US Patent No. 7,211,414, US Patent No. 7,315,019, US Patent No. 7,405,281 or US Patent No. 7,057,026, or US Patent Application Publication No. 2008 / 0108082 (A1), each of which is incorporated herein by reference.Another type of array that is useful for nucleic acid sequencing is the array of particles that are generated from emulsion PCR technology. Examples are described in Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), WO 05 / 010145, U.S. Patent Application Publication No. 2005 / 0130173, or U.S. Patent Application Publication No. 2005 / 0064460, each of which is incorporated herein by reference in its entirety.

[0400] Arrays used for nucleic acid sequencing often have random spatial patterns of nucleic acid analytes. For example, the HiSeq or MiSeq sequencing platforms available from Illumina Inc. (San Diego, Calif.) utilize flow cells in which nucleic acid sequences are formed by random seeding followed by bridge amplification. However, patterned arrays can also be used for nucleic acid sequencing or other analytical applications. Examples of patterned arrays, their fabrication methods, and their uses are described in U.S. Patent Application Nos. 13 / 787,396, 13 / 783,043, 13 / 784,368, U.S. Patent Application Publication Nos. 2013 / 0116153 (A1), and 2012 / 0316086 (A1), each of which is incorporated herein by reference. Such patterned array analytes can be used to capture single nucleic acid template molecules for subsequent formation of homogeneous colonies, e.g., via bridge amplification. Such patterned arrays are particularly useful for nucleic acid sequencing applications.

[0401] The size of the specimens on an array (or other object used in the methods or systems herein) can be selected to suit a particular application. For example, in some embodiments, the specimens on the array can have a size that accommodates only a single nucleic acid molecule. A surface with multiple specimens in this size range is useful for constructing an array of molecules for detection with single molecule resolution. Specimens in this size range are also useful for use in arrays with specimens each comprising a colony of nucleic acid molecules. Thus, the specimens on the array can each be approximately 1 mm 2 Below, approximately 500μm 2 Below, approximately 100μm 2 Below, approximately 10μm 2 Below, approximately 1μm 2 Below, about 500nm 2 Less than or equal to about 100 nm 2 Below, approximately 10nm 2 Below, approximately 5nm 2 Less than or equal to 1 nm 2Alternatively or additionally, the specimens in the array may have an area of ​​about 1 mm 2 More than approximately 500μm 2 More than approximately 100μm 2 or more, about 10μm 2 or more, approximately 1μm 2 or more, about 500nm 2 Over 100nm 2 or more, about 10nm 2 or more, about 5nm 2 or more, or about 1 nm 2 That's it. In practice, analytes can have a size within a range between upper and lower limits selected from those exemplified above. While several size ranges for surface analytes have been exemplified with respect to nucleic acids and nucleic acid scales, it will be understood that analytes in these size ranges can be used in applications that do not involve nucleic acids. It will further be understood that analyte sizes need not necessarily be limited to the scales used in nucleic acid applications.

[0402] In embodiments involving an object having multiple analytes, such as an array of analytes, the analytes can be distinct, separated by a space between them. Arrays useful in the invention can have analytes separated by an edge-to-edge distance of at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, or less. Alternatively or additionally, arrays can have analytes separated by an edge-to-edge distance of at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can apply to the average edge-to-edge spacing and edge-to-edge spacing of the analytes, as well as to the minimum or maximum spacing.

[0403] In some embodiments, the analytes in the array need not be distinct; instead, adjacent analytes can abut one another. Whether the analytes are distinct or not, the size of the analytes and / or the pitch of the analytes can be varied to allow the array to have a desired density. For example, the average analyte pitch in a regular pattern can be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, or less. Alternatively or additionally, the average analyte pitch in a regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can also apply to the maximum or minimum pitch of a regular pattern. For example, the maximum specimen pitch in the regular pattern can be 100 μm or less, 50 μm or less, 10 μm or less, 5 μm or less, 1 μm or less, 0.5 μm or less, and / or the minimum specimen pitch in the regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more.

[0404] The density of analytes in an array can also be understood in terms of the number of analytes present per unit area. For example, the average density of analytes for an array is at least about 1×10 3 specimens / mm 2 , 1×10 4 specimens / mm 2 , 1×10 5 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 7 specimens / mm 2 , 1×10 8 specimens / mm 2 , or 1×10 9 specimens / mm 2 Alternatively or additionally, the average density of analytes on the array can be at most about 1 x 10 9 specimens / mm 2 , 1×10 8 specimens / mm 2 , 1×10 7 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 5specimens / mm 2 , 1×10 4 specimens / mm 2 , or 1×10 3 specimens / mm 2 It can be the following:

[0405] The above ranges may apply, for example, to all or part of a regular pattern that includes all or part of an array of analytes.

[0406] The analytes within the pattern can have any of a variety of shapes. For example, when viewed in a two-dimensional plane, such as on the surface of an array, the analytes may appear rounded, circular, oval, rectangular, square, symmetrical, asymmetrical, triangular, polygonal, etc. The analytes can be arranged in a regular repeating pattern, including, for example, a hexagonal or rectilinear pattern. The pattern can be selected to achieve a desired level of packing. For example, circular analytes are optimally packed in a hexagonal arrangement. Of course, other packing configurations can also be used for circular analytes, and vice versa.

[0407] A pattern can be characterized in terms of the number of analytes present in a subset that forms the smallest geometric unit of the pattern. A subset can include, for example, at least about 2, 3, 4, 5, 6, 10, or more analytes. Depending on the size and density of the analytes, a geometric unit can be as small as 1 mm. 2 , 500 μm 2 , 100 μm 2 , 50 μm 2 , 10 μm 2 , 1 μm 2 , 500nm 2 , 100 nm 2 , 50nm 2 , 10nm 2 Alternatively or additionally, the geometric unit may be 10 nm 2 , 50nm 2 , 100 nm 2 , 500nm 2 , 1 μm 2 , 10 μm 2 , 50 μm 2, 100 μm 2 , 500 μm 2 , 1mm 2 The properties of the specimens in the geometric unit, such as shape, size, pitch, etc., can be selected from those described herein more generally for specimens in an array or pattern.

[0408] An array with a regular pattern of analytes may be ordered with respect to the relative location of the analytes, but random with respect to one or more other characteristics of each analyte. For example, in the case of nucleic acid sequences, the nucleic acid analytes may be regular with respect to their relative location, but random with respect to the sequence knowledge of the nucleic acid species present in any particular analyte. As a more specific example, a nucleic acid sequence formed by seeding a repeating pattern of analytes with template nucleic acids and amplifying the template in each analyte to form copies of the template in the analyte (e.g., via cluster amplification or bridge amplification) will have a regular pattern of nucleic acid analytes, but will be random with respect to the distribution of the sequences of the nucleic acids across the array. Thus, detection of the presence of nucleic acid material on an array can result in a repeating pattern of analytes, whereas sequence-specific detection can result in a non-repeating distribution of signal across the array.

[0409] It will be understood that descriptions of pattern, order, randomness, etc. herein relate not only to analytes on an object, such as analytes on an array, but also to analytes in an image. Thus, the pattern, order, randomness, etc. can exist in any of a variety of formats used to store, manipulate, or communicate image data, including, but not limited to, computer-readable media or computer components such as a graphical user interface or other output device.

[0410] As used herein, the term "image" is intended to mean a representation of all or a portion of an object. The representation may be an optically detected reproduction. For example, an image may be obtained from fluorescence, luminescence, scattering, or absorption signals. The portion of an object present in an image may be the surface or other xy plane of the object. Typically, an image is a two-dimensional representation, but in some cases, information in an image can be derived from three or more dimensions. An image need not include optically detected signals. Non-optical signals may instead be present. An image may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.

[0411] As used herein, "image" refers to a reproduction or representation of at least a portion of a sample or other object. In some embodiments, the reproduction is an optical reproduction, e.g., produced by a camera or other optical detector. The reproduction may be a non-optical reproduction, e.g., a representation of electrical signals obtained from an array of nanopore analytes or a representation of electrical signals obtained from an ion-sensitive CMOS detector. In certain embodiments, non-optical reproductions may be excluded from the methods or apparatus described herein. The image may have a resolution capable of distinguishing between analytes present at any of a variety of intervals, including, for example, those spaced less than 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm apart.

[0412] As used herein, "acquisition," "capture," and like terms refer to any part of the process of acquiring an image file. In some embodiments, data acquisition can include generating an image of the specimen, looking for a signal in the specimen, directing a detection device to look for or generate an image of the signal, and providing instructions for further analysis or transformation of the image file, and instructions for any number of transformations or manipulations of the image file.

[0413] As used herein, the term "template" refers to a representation of the location or relationship between signals or analytes. Thus, in some embodiments, the template is a physical grid having a representation of signals corresponding to analytes in a sample. In some embodiments, the template can be a chart, table, text file, or other computer file that indicates locations corresponding to analytes. In the embodiments presented herein, a template is generated to track the location of analytes across a set of images of the sample captured at different reference points. For example, the template can be a set of x, y coordinates or a set of values ​​that describe the orientation and / or distance of one analyte relative to another analyte.

[0414] As used herein, the term "specimen" can refer to an object or region of an object from which an image is captured. For example, in an embodiment in which an image is taken from the surface of soil, a parcel of land may be the specimen. In other embodiments in which biomolecular analysis is performed within a flow cell, the flow cell may be divided into any number of subdivisions, each of which may be a specimen. For example, the flow cell may be divided into various channels or lanes, and each lane may be further divided into 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 140, 160, 180, 200, 400, 600, 800, 1000, or more distinct regions to be imaged. One example flow cell has eight lanes, each divided into 120 specimens or tiles. In other embodiments, samples may be generated in multiple tiles, or even the entire flow cell. Thus, each specimen image can represent a larger surface area than is imaged.

[0415] References to ranges and sequential lists of numbers set forth herein will be understood to include not only the numbers recited but all real numbers between the recited numbers.

[0416] As used herein, a "reference point" refers to any temporal or physical distinction between images. In another preferred embodiment, the reference point is a time point. In a more preferred embodiment, the reference point is a time point or cycle during the sequencing reaction. However, the term "reference point" can also include other aspects that distinguish or separate images, such as angle, rotation, time, or other aspects that can distinguish or separate the images.

[0417] As used herein, "subset of images" refers to a group of images in a set. For example, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or any number of images selected from a set of images. In certain other embodiments, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or less or any number of images selected from a set of images. In another preferred embodiment, images are obtained from one or more sequencing cycles, with four images associated with each cycle. Thus, for example, a subset may be a group of 16 images acquired over four cycles.

[0418] Base refers to a nucleotide base or nucleotide, (adenine), C (cytosine), T (thymine), or G (guanine). This application uses "base" and "nucleotide" interchangeably.

[0419] The term "chromosome" refers to a gene carrier of the present invention in a living cell, derived from a chromatin strand containing DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is used herein.

[0420] The term "site" refers to a unique location (e.g., chromosome ID, chromosomal location and orientation) on a reference genome. In some embodiments, a site may be the location of a residue, sequence tag, or segment on a sequence. The term "locus" may be used to refer to a specific location of a nucleic acid sequence or polymorphism on a reference chromosome.

[0421] The term "sample" as used herein typically refers to a sample derived from a biological fluid, cell, tissue, organ, or organism containing nucleic acids to be sequenced and / or phased, or a sample derived from a mixture of nucleic acids containing at least one nucleic acid sequence to be sequenced and / or phased. Such samples include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, fine needle biopsy samples (e.g., surgical biopsies, needle biopsies, etc.), urine, peritoneal fluid, pleural fluid, tissue explants, organ cultures, and any other tissue or cell preparations thereof, or fractions or derivatives thereof. While samples are often collected from human subjects (e.g., patients), samples can be collected from any organism that has chromosomes, including, but not limited to, dogs, cats, horses, goats, sheep, cows, pigs, etc. Samples can be used directly as obtained from a biological source or after pretreatment to modify the characteristics of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, and the like.

[0422] The term "sequence" includes or refers to a chain of nucleotides linked together. The nucleotides can be based on DNA or RNA. It should be understood that a sequence may contain multiple subsequences. For example, a single sequence (e.g., a PCR amplicon) may have 350 nucleotides. A sample read may contain multiple subsequences within these 350 nucleotides. For example, a sample read may include first and second flanking subsequences, each having, for example, 20-50 nucleotides. The first and second flanking subsequences may be located on either side of a repeat segment with a corresponding subsequence (e.g., 40-100 nucleotides). Each of the flanking subsequences may include (or a portion of) a primer subsequence (e.g., 10-30 nucleotides). For ease of interpretation, the term "subsequence" is referred to as "sequence," but it is understood that two sequences need not be distinct from each other on a common strand. To distinguish between the various sequences described herein, the sequences may be given different labels (e.g., target sequence, primer sequence, flanking sequence, reference sequence, etc.). Other terms, such as "allele," may be given different labels to distinguish between similar entities. Applications use "read" and "sequence read" interchangeably.

[0423] The term "paired-end sequencing" refers to a sequencing method in which both ends of a target fragment are sequenced. Paired-end sequencing can facilitate the detection of genome rearrangements and repeated segments, as well as the detection of gene fusions and novel transcripts. Methods for paired-end sequencing are described in International Publication No. 07010252, International Application No. GB2007 / 003798, and U.S. Patent Application Publication No. 2009 / 0088327, each of which is incorporated herein by reference. In one example, a series of operations can be performed as follows: (a) generating clusters of nucleic acids, (b) linearizing the nucleic acids, (c) hybridizing a first sequencing primer and repeatedly performing cycles of extension, scanning, and deblocking as described above, (d) "flipping" the target nucleic acid on the flow cell surface by synthesizing a complementary copy, (e) linearizing the resynthesized strand, and (f) hybridizing a second sequencing primer and repeatedly performing cycles of extension, scanning, and deblocking as described above. The flipping operation can deliver the reagents described above for a single cycle of bridge amplification.

[0424] The term "reference genome" or "reference sequence" refers to a specific known genome sequence, either partial or complete, of any organism that can be used to reference an identified sequence from a subject. For example, reference genomes used for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information (ncbi.nlm.nih.gov). "Genome" refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences. A genome includes both genetic and non-coding sequences of DNA. A reference sequence may be larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 10 times larger, or at least about 10 times larger, or at least about 10 times larger. In one example, the reference genome sequence is of the full-length human genome. In another example, the reference genome sequence is limited to a specific human chromosome, such as chromosome 13. In some embodiments, the reference chromosome is a chromosome sequence from the human genome version hg19. Such sequences are sometimes referred to as chromosomal reference sequences, although the term reference genome is intended to encompass such sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, partial chromosomal regions (such as strands) of any species. In various embodiments, a reference genome is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, a reference sequence may be taken from a specific individual. In other embodiments, "genome" also covers so-called "graph genomes," which use specific storage formats and representations of genome sequences. In one embodiment, a graph genome stores data in a linear file. In another embodiment, a graph genome refers to a representation in which alternative sequencing (e.g., different copies of a chromosome with small differences) are stored as different paths in a graph.Additional information regarding the implementation of graph genomes can be found at https: / / www.biorxiv.org / content / biorxiv / early / 2018 / 03 / 20 / 194530.full.pdf, the contents of which are incorporated herein by reference in their entirety.

[0425] The term "read" refers to a collection of sequence data describing a fragment of a nucleotide sample or reference. The term "read" can refer to a sample read and / or a reference read. Typically, but not necessarily, a read represents a short sequence of consecutive base pairs in a sample or reference. A read may be symbolically represented by the base pair sequence (ACTG) of the sample or reference fragment. A read may be stored in a memory device and appropriately processed to determine whether the read matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing device or indirectly from stored sequence information about the sample. In some cases, the read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify larger sequences or regions that can be aligned and specifically assigned to, for example, a chromosome or genomic region or gene.

[0426] Next-generation sequencing methods include, for example, sequencing by synthesis (Illumina), pyrosequencing (454), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing, and sequencing by ligation (SOLiD sequencing). Depending on the sequencing method, the length of each read can vary from about 30 bp to over 10,000 bp. For example, DNA sequencing using a SOLiD sequencer generates nucleic acid reads of about 50 bp. In another example, Ion Torrent sequencing generates nucleic acid reads of up to 400 bp, while 454 pyrosequencing generates nucleic acid reads of about 700 bp. In yet another example, single-molecule real-time sequencing can generate reads of 10,000 bp to 15,000 bp. Thus, in certain embodiments, nucleic acid sequence reads have lengths of 30 to 100 bp, 50 to 200 bp, or 50 to 400 bp.

[0427] The terms "sample read," "sample sequence," or "sample fragment" refer to sequence data relating to a genomic sequence of interest from a sample. For example, a sample read includes sequence data from a PCR amplicon having forward and reverse primer sequences. The sequence data can be obtained from any selected sequencing methodology. A sample read can be, for example, a sequencing-by-synthesis (SBS) reaction, a sequencing-ligation reaction, or any other suitable sequencing methodology in which it is desirable to determine the length and / or identity of repetitive elements. A sample read can be a consensus (e.g., average or weighted) sequence derived from multiple sample reads. In certain embodiments, providing a reference sequence includes identifying a locus of interest based on primer sequences of a PCR amplicon.

[0428] The term "raw fragment" refers to a portion of sequence data of a genome sequence of interest that at least partially overlaps a specified or secondary location of interest within a sample read or sample fragment. Non-limiting examples of raw fragments include double-stitched fragments, simple stitched fragments, and simple unstitched fragments. The term "raw" is used to indicate that a raw fragment contains sequence data that has some relationship to the sequence data in a sample read, regardless of whether the raw fragment exhibits supporting variants that correspond to and authenticate or confirm potential variants in the sample read. The term "raw fragment" does not necessarily indicate that the fragment contains supporting variants that verify the variant call in the sample read. For example, when a sample read is determined by a variant calling application to exhibit a first variant, the variant calling application may determine that one or more raw fragments lack a corresponding type of "supporting" variant that would otherwise be expected to occur given the variants in the sample read.

[0429] The terms "mapping," "aligned," "aligning," or "aligning" refer to the process of comparing a read or tag to a reference sequence, thereby determining whether the reference sequence contains the read sequence. If the reference sequence is read, the read may be mapped to the reference sequence, or in certain other embodiments, may be mapped to a specific location within the reference sequence. In some cases, alignment simply conveys whether a read is a member of a specific reference sequence (i.e., whether the read is present or absent in the reference sequence). For example, alignment of a read to a reference sequence for human chromosome 13 conveys whether the read is present in the reference sequence for chromosome 13. Tools that provide this information are sometimes called set membership testers. In some cases, alignment also indicates the location within the reference sequence where the read or tag map is. For example, if the reference sequence is the entire human genome sequence, alignment may indicate that the read is present on chromosome 13, and may further indicate that the read is on a specific strand and / or site of chromosome 13.

[0430] The term "indel" refers to the insertion and / or deletion of bases in an organism's DNA. Microindels refer to indels that result in a net change of 1 to 50 nucleotides. Unless the length of the indel is a multiple of three, a frameshift mutation occurs in coding regions of the genome. Indels can be contrasted with point mutations. An indel insertion deletes nucleotides from the sequence, whereas a point mutation is a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels can also be contrasted with tandem base mutations (TBMs), which can be defined as substitutions at adjacent nucleotides (mostly two adjacent nucleotides, although substitutions at three adjacent nucleotides have been observed).

[0431] The term "variant" refers to a nucleic acid sequence that differs from a reference nucleic acid. Typical nucleic acid sequence variants include, but are not limited to, single nucleotide polymorphisms (SNPs), short deletion and insertion polymorphisms (indels), copy number variations (CNVs), microsatellite markers, or short tandem repeats and structural variants. Somatic variant calling is an effort to identify variants present at low frequencies in a DNA sample. Calling somatic variants is of interest in the context of cancer treatment. Cancer is caused by the accumulation of mutations in DNA. DNA samples derived from tumors are generally heterogeneous, containing some normal cells, early-stage cancer progression (with fewer mutations), and some late-stage cells (with more mutations). Due to this heterogeneity, somatic mutations often appear at low frequencies when sequencing tumors (e.g., from FFPE samples). For example, an SNV may be found in only 10% of reads covering a given base. A variant that is classified as somatic or germline by a variant classifier is also referred to herein as a "variant under test."

[0432] The term "noise" refers to erroneous variant calls that result from one or more errors in the sequencing process and / or variant calling application.

[0433] The term "variant frequency" refers to the relative frequency of an allele (variant of a gene) at a particular locus within a population, expressed as a fraction or percentage. For example, the fraction or percentage may be the proportion of all chromosomes in the population that carry that allele. As an example, a sample variant frequency refers to the relative frequency of an allele / variant at a particular locus / position along a genome sequence of interest across a "population" corresponding to the number of reads and / or samples obtained for the genome sequence of interest from individuals. As another example, a baseline variant frequency refers to the relative frequency of an allele / variant at a particular locus / position along one or more baseline genome sequences, where a "population" corresponds to the number of reads and / or samples obtained for one or more baseline genome sequences from a population of normal individuals.

[0434] The term "Variant Allele Frequency (VAF)" refers to the proportion of sequenced reads that carry a variant divided by the overall coverage at the target position. VAF is a measure of the proportion of sequenced reads that carry a variant.

[0435] The terms "position," "designated position," and "locus" refer to the position or coordinates of one or more nucleotides within a nucleotide sequence. The terms "position," "designated position," and "locus" also refer to the position or coordinates of one or more base pairs in a sequence of nucleotides.

[0436] The term "haplotype" refers to a combination of alleles at adjacent sites on a chromosome that are inherited from one another. A haplotype, if present, may be one locus, several loci, or an entire chromosome, depending on the number of recombination events that have occurred between a given set of loci.

[0437] The term "threshold" herein refers to a number or values ​​used as a cutoff for characterizing a sample, a nucleic acid, or a portion thereof (e.g., a readout). The threshold may be varied based on empirical analysis. The threshold can be compared to a measured or calculated value to determine whether the source giving rise to such a value should be classified in a particular manner. The threshold can be identified empirically or analytically. The choice of threshold depends on the confidence level with which the user desires the classification to be made. The threshold may be selected for a particular purpose (e.g., to balance sensitivity and selectivity). As used herein, the term "threshold" refers to a point at which the course of an analysis may change and / or an action may be triggered. A threshold need not be a predetermined number. Instead, the threshold may be, for example, a function based on multiple factors. The threshold may be adaptive to the situation. Furthermore, a threshold may indicate an upper limit, a lower limit, or a range between limits.

[0438] In some embodiments, an index or score based on sequencing data may be compared to a threshold. As used herein, the terms "metric" or "score" may include a value or result determined from sequencing data or a function based on a value or result determined from sequencing data. Like a threshold, an index or score may be adaptive to the context. For example, an index or score may be a normalized value. As an example of a score or metric, one or more embodiments may use a count score when analyzing data. The count score may be based on the number of sample reads. The sample reads may have undergone one or more filtering stages so that the sample reads have at least one common characteristic or quality. For example, each of the sample reads used to determine the count score may be aligned with a reference sequence or assigned as a potential allele. The number of sample reads with a common characteristic may be counted to determine a read count. The count score may be based on the read count. In some embodiments, the count score may be a value equal to the read count. In other examples, the count score may be based on the read count and other information. For example, the counting score may be based on the read count of a particular allele at a locus and the total number of reads at the locus. In some embodiments, the counting score may be based on the read count of a locus and previously obtained data. In some embodiments, the counting score may be a normalized score between predetermined values. The counting score may also be a function of read counts from other loci in the sample, or read counts from other samples run simultaneously with the sample of interest. For example, the counting score may be a function of the read count of a particular allele and the read counts of other loci in the sample, and / or read counts from other samples. As an example, the read counts from other loci and / or read counts from other samples may be used to normalize the counting score for a particular allele.

[0439] The term "coverage" or "fragment coverage" refers to a count or other measure of multiple sample reads for the same fragment of a sequence. A read count may represent a count of the number of reads that cover the corresponding fragment. Alternatively, coverage may be determined by multiplying the read count by a specified factor based on historical knowledge, knowledge of the sample, knowledge of the locus, etc.

[0440] The term "read depth" (conventionally a number followed by "x") refers to the number of sequenced reads with overlapping alignment at the target position. This is often expressed as an average or percentage above a cutoff for a set of intervals (such as an exon, gene, or panel). For example, a clinical report may state that the panel average coverage is 1,105x, with 98% of the targeted base coverage >100x.

[0441] The term "base call quality score" or "Q score" refers to a PHRED-scaled probability ranging from 0-50 that is inversely proportional to the probability that a single sequenced base is correct. For example, a T base call with a Q of 20 is considered correct 99.99% of the time. Any base call with a Q < 20 should be considered low quality, and any variant identified when a significant proportion of sequenced reads supporting the variant is low should be considered a potential false positive.

[0442] The term "variant read" or "variant read number" refers to the number of sequenced reads that support the presence of a variant.

[0443] Regarding "strandedness" (or DNA twist), the genetic message in DNA can be represented as the letters A, G, C, and T, e.g., 5'-AGGACA-3'. Often, sequences are written in the orientation shown herein, i.e., 5'-left and 3'-right. While DNA can occur as a single-stranded molecule (as in certain viruses), we typically find DNA as a double-stranded unit, which has a double helix structure with two antiparallel strands. In this case, the term "antiparallel" means that the two strands run parallel but have opposite polarity. Double-stranded DNA is held together by base pairing, and pairing is always maintained, such that adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). This pairing is called complementarity, and one DNA strand is said to be the complement of another. Thus, double-stranded DNA can be represented as two strings, such as 5'-AGGACA-3' and 3'-TCCTGT-5'. Note that the two strands have opposite polarities. Thus, the strandedness of the two DNA strands can be referred to as a reference strand and its complement, a forward and reverse strand, a top and bottom strand, a sense and antisense strand, or a Watson and Crick strand.

[0444] Read alignment (also called read mapping) is the process of referring to where sequences in a genome originate. Once aligned, the "mapping quality" or "mapping quality score (MAPQ)" of a given read quantifies the probability that its location on the genome is correct. Mapping quality is encoded on a phase scale, where P is the probability that the alignment is incorrect. The probability is P=10 (-MAQ / 10)MAPQ is calculated as follows: where MAPQ is the mapping quality. For example, a mapping quality of 40 = 10 with a power of -4 means that there is a 0.01% chance that the ...

Claims

1. 1. A system comprising: storing a respective current intensity value for each pixel within the plurality of pixels for a current sequencing cycle of the sequencing run; a memory that stores a sequence of respective previous intensity values ​​for each pixel for one or more previous sequencing cycles of the sequencing run that precede the current sequencing cycle; and having access to the memory, and determining a respective current state value for each of the pixels. (i) the respective current intensity values; and (ii) each previous intensity value in said sequence of said each previous intensity value and storing said respective current state values ​​in said memory; having access to the memory; (i) the respective current intensity values; and (ii) each of said current state values; a base caller configured to generate base calls for the current sequencing cycle in response to processing the

2. The system of claim 1 , wherein each respective intensity value for each pixel is characterized by a channel-specific intensity value for a plurality of channels.

3. The system described in claim 2, wherein each state value for each pixel is characterized by a channel-specific state value for the multiple channels.

4. The system described in claim 3, wherein each of the current state values ​​is encoded pixel by pixel using each of the current intensity values.

5. The system of claim 3 , wherein channel eigenstate values ​​for a subset of channels in the plurality of channels are encoded pixel-by-pixel with the respective current intensity values.

6. 4. The system of claim 3, wherein the channel-specific state values ​​are averaged across channels in the plurality of channels to generate respective pan channel current state values, and the respective pan channel current state values ​​are encoded on a pixel-by-pixel basis with the respective current intensity values.

7. The system of claim 4, wherein the pixel-by-pixel encoding includes pixel-by-pixel concatenation.

8. The system of claim 4, wherein the pixel-by-pixel encoding includes a pixel-by-pixel sum.

9. Each of said current state values ​​is (i) each of said previous intensity values; and (ii) each of said current intensity values; 2. The system of claim 1, wherein the respective current average intensities are determined for the respective pixels in the current sequencing cycle from

10. The system of claim 9 , wherein the respective current average powers are each characterized by a channel-specific current average power.

11. The respective current state values: (i) each of said previous intensity values; and (ii) each of said current intensity values; 11. The system of claim 10, wherein the respective current maximum intensities determined for the respective pixels in the current sequencing cycle are from 12. The system of claim 11, wherein each of the current maximum intensities is characterized by a channel-specific current maximum intensity.

13. Each of said current state values ​​is (i) each of said previous intensity values; and (ii) each of said current intensity values; 13. The system of claim 12, wherein the respective current minimum intensities determined for the respective pixels in the current sequencing cycle are:

14. The system of claim 13 , wherein each of the respective current minimum intensities is characterized by a channel-specific current minimum intensities.

15. Each of said current state values ​​is (i) each of said previous intensity values; and (ii) each of said current intensity values; 15. The system of claim 14, wherein each of the current exponentially weighted average intensities determined for the respective pixel in the current sequencing cycle is:

16. 16. The system of claim 15, wherein the respective current exponentially weighted average intensities are determined based on weighting more recent sequencing cycles than previous sequencing cycles.

17. 17. The system of claim 16, wherein the respective current exponentially weighted average powers are each characterized by a channel-specific current exponentially weighted average power.

18. Each of said current state values ​​is (i) the respective current intensity values; and (ii) a rolling subset of each of said previous intensity values; 18. The system of claim 17, wherein the current running average intensity is determined for the respective pixel in the current sequencing cycle from

19. 20. The system of claim 18, wherein each of the respective current moving average strengths is characterized by a channel-specific current moving average strength.

20. 20. The system of claim 19, wherein each of the respective current intensity values ​​is assigned to an active state bucket or an inactive state bucket based on a comparison of the respective current intensity value with a respective current active state intensity and a respective current inactive state intensity.

21. each of said current active state intensities; (i) each of said previous intensity values; 21. The system of claim 20, wherein each current global maximum intensity is determined for the respective pixel in the current sequencing cycle from 22. The system of claim 21, wherein each of the current global maximum intensities is characterized by a channel-specific current global maximum intensity.

23. The respective current inactive state intensities: (i) each of said previous intensity values; 23. The system of claim 22, wherein each current global minimum intensity is determined for the respective pixel in the current sequencing cycle from

24. The system described in claim 23, wherein each of the respective current global minimum intensities is characterized by a channel-specific current global minimum intensities.

25. The system described in claim 24, wherein the respective current state values ​​further include respective current active state values ​​and respective current inactive state values ​​generated for the respective pixels in the current sequencing cycle.

26. The system of claim 25, wherein each of the respective current active state values ​​is characterized by a channel-specific current active state value.

27. The system of claim 26, wherein each of the respective current inactive state values ​​is characterized by a channel-specific current inactive state value.

28. The current active state value for the pixel in the target channel is (i) current intensity values ​​for the pixels in the channel of interest detected in the current sequencing cycle and attributed to active state buckets; and (ii) a previous intensity value for the pixel in the channel of interest detected in the one or more previous sequencing cycles; 28. The system of claim 27, wherein the current sequencing cycle is determined from 29. The method of claim 29, further comprising: accessing current window sequencing signal data for a first plurality of sequencing cycles within a current sliding window of sequencing cycles; accessing previous sequencing signal data for a second plurality of sequencing cycles that precedes a sequencing cycle in the first plurality of sequencing cycles; generating status signal data based on the current window sequencing signal data and the previous sequencing signal data; and calling a base for at least one sequencing cycle in the first plurality of sequencing cycles in response to processing the current window sequencing signal data and the state signal data.

30. A system comprising: a memory that stores a current sequencing image for a cluster for a current sequencing cycle of a sequencing run and a previous sequencing image for the cluster for one or more previous sequencing cycles of the sequencing run that precede the current sequencing cycle; a host processor having access to the memory and configured to generate, for the current sequencing cycle, current per-pixel states for at least pixels in the current sequencing image based on the current sequencing image and the previous sequencing image; a configurable processor having access to the memory and configured to execute a neural network to generate base call classification scores; a neural network having access to the memory, the host processor, and the configurable processor, providing the current pixel-by-pixel state and the current sequencing image to the neural network, and causing the neural network to generate a current base call classification score for the cluster for the current sequencing cycle in response to processing the current pixel-by-pixel state and the current sequencing image; data flow logic configured to provide the current base call classification score to the host processor; A system comprising: