Specialist signal profiler for base calling

JP2024528511A5Pending Publication Date: 2025-07-24ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023579788
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-13
Filing Date
2022-07-14
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The variation in intensity profiles of clusters during sequencing runs, caused by factors such as spatial crosstalk, photobleaching, and optical distortions, leads to reduced data throughput and increased error rates in high-throughput sequencing.

Method used

The use of specialist signal profilers, trained to maximize the signal-to-noise ratio for specific categories of data, such as surface-specific, lane-specific, and subtile-specific, to correct for spatial crosstalk and optical distortions in image data, optimizing base calling accuracy.

Benefits of technology

This approach significantly improves base call accuracy and reduces sequencing errors by fine-tuning the signal processing for each subpopulation of clusters, enhancing the overall quality of sequencing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system is disclosed, the system comprising a memory and runtime logic, the memory storing a plurality of specialist signal profilers, each specialist signal profiler in the plurality of specialist signal profilers being trained to maximize a signal-to-noise ratio of sequence signals in a particular signal profile detected for analytes in a particular analyte class and characterized in a particular training data set, the runtime logic having access to the memory is configured to perform a base calling operation by applying each specialist signal profiler in the plurality of specialist signal profilers to sequence signals in the respective signal profiles detected for analytes in the respective analyte class during a base calling operation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Priority Application This application claims priority to and the benefit of U.S. Nonprovisional Patent Application No. 17 / 839,353, entitled "Specialist Signal Profiler for Base Calling," filed on June 13, 2022 (Attorney Docket No. ILLM 1041-2 / IP-2063-US), which claims priority to U.S. Provisional Patent Application No. 63 / 223,408, entitled "Specialist Signal Profiler for Base Calling," filed on July 19, 2021 (Attorney Docket No. ILLM 1041-1 / IP-2063-PRV), which is incorporated herein by reference in its entirety for all purposes.

[0002] The disclosed technology relates to an apparatus and corresponding methods for automated analysis or pattern recognition of an image. This specification includes systems that transform an image to (a) improve its visual quality before recognition, (b) reduce the amount of image data by registering and aligning the image to a sensor or a stored prototype, or by discarding irrelevant data, and (c) measure significant characteristics of the image. In particular, the disclosed technology relates to removing spatial crosstalk from sensor pixels using equalization-based image processing techniques.

[0003] Reference The following are incorporated for all purposes as if fully set forth herein:

[0004] U.S. Nonprovisional Patent Application No. 17 / 308,035, entitled “EQUALIZATION-BASED IMAGE PROCESSING AND SPATIAL CROSSTALK ATTENUATOR,” filed on May 4, 2021 (Attorney Docket No. ILLM 1032-2 / IP-1991-US); U.S. Provisional Patent Application No. 63 / 106,256, entitled “Systems and Methods for Per-Cluster Intensity Correction and Base Calling,” filed on October 27, 2020; U.S. Nonprovisional Patent Application No. 15 / 909,437, entitled “Optical Distortion Correction for Imaged Samples,” filed March 1, 2018; U.S. Nonprovisional Patent Application No. 14 / 530,299, entitled “IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS,” filed on October 31, 2014; U.S. Nonprovisional Patent Application No. 15 / 153,953, entitled “METHODS AND SYSTEMS FOR ANALYZING IMAGE DATA,” filed December 3, 2014; U.S. Nonprovisional Patent Application No. 15 / 863,241, entitled “Phasing Correction,” filed January 5, 2018; U.S. Nonprovisional Patent Application No. 14 / 020,570, entitled “CENTROID MARKERS FOR IMAGE ANALYSIS OF HIGH DENSITY CLUSTERS IN COMPLEX POLYNUCLEOTIDE SEQUENCING,” filed on September 6, 2013; U.S. Nonprovisional Patent Application No. 12 / 565,341, entitled "METHOD AND SYSTEM FOR DETERMINING THE ACCURACY OF DNA BASE IDENTIFICATIONS," filed September 23, 2009; U.S. Nonprovisional Patent Application No. 12 / 295,337, entitled "SYSTEMS AND DEVICES FOR SEQUENCE BY SYNTHESIS ANALYSIS," filed March 30, 2007; U.S. Non-provisional Patent Application No. 12 / 020,739, entitled "IMAGE DATA EFFICIENT GENETIC SEQUENCING METHOD AND SYSTEM," filed on January 28, 2008; U.S. Nonprovisional Patent Application No. 13 / 833,619, entitled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR SAME," filed March 15, 2013 (Attorney Docket No. IP-0626-US); U.S. Nonprovisional Patent Application No. 15 / 175,489, entitled “BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND METHODS OF MANUFACTURING THE SAME,” filed on June 7, 2016 (Attorney Docket No. IP-0689-US); U.S. Non-Provisional Patent Application No. 13 / 882,088, entitled “MICRODEVICES AND BIOSENSOR CARTRIDGES FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR THE SAME,” filed April 26, 2013 (Attorney Docket No. IP-0462-US); U.S. Nonprovisional Patent Application No. 13 / 624,200, entitled "METHODS AND COMPOSITIONS FOR NUCLEIC ACID SEQUENCING," filed on September 21, 2012 (Attorney Docket No. IP-0538-US); U.S. Nonprovisional Patent Application No. 13 / 006,206, entitled "DATA PROCESSING SYSTEM AND METHODS," filed on January 13, 2011; U.S. Nonprovisional Patent Application No. 15 / 936,365, entitled “DETECTION APPARATUS HAVING A MICROFLUOROMETER, A FLUIDIC SYSTEM, AND A FLOW CELL LATCH CLAMP MODULE,” filed March 26, 2018; U.S. Nonprovisional Patent Application No. 16 / 567,224, entitled “FLOW CELLS AND METHODS RELATED TO SAME,” filed on September 11, 2019; U.S. Nonprovisional Patent Application No. 16 / 439,635, entitled “DEVICE FOR LUMINESCENT IMAGING,” filed June 12, 2019; U.S. Nonprovisional Patent Application No. 15 / 594,413, entitled “INTEGRATED OPTOELECTRONIC READ HEAD AND FLUIDIC CARTRIDGE USEFUL FOR NUCLEIC ACID SEQUENCING,” filed May 12, 2017; U.S. Nonprovisional Patent Application No. 16 / 351,193, entitled “ILLUMINATION FOR FLUORESCENCE IMAGING USING OBJECTIVE LENS,” filed March 12, 2019; U.S. Nonprovisional Patent Application No. 12 / 638,770, entitled "DYNAMIC AUTOFOCUS METHOD AND SYSTEM FOR ASSAY IMAGER," filed December 15, 2009; No. 13 / 783,043, entitled "KINETIC EXCLUSION AMPLIFICATION OF NUCLEIC ACID LIBRARIES," filed March 1, 2013; and U.S. patent application Ser. No. 16 / 826,168, entitled “ARTIFICIAL INTELLIGENCE-BASED SEQUENCING,” filed March 21, 2020 (Attorney Docket No. ILLM 1008-20 / IP-1752-PRV). [Background technology]

[0005] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which as such may also correspond to implementations of the claimed technology. Summary of the Invention [Means for solving the problem]

[0006] Base calling accuracy is crucial for high-throughput sequencing and downstream analysis such as read mapping and genome assembly. The present disclosure relates to optimizing image data to accurately base call clusters during sequencing runs. One challenge with optimizing image data is the variation of the intensity profile (or intensity distribution) of clusters in the cluster population to be base called. This is particularly detrimental to multi-cycle imaging of substrates (e.g., flow cells) with a large number of clusters (e.g., thousands, millions, billions, etc.). This makes the scale of variation unmanageable, thereby causing a decrease in data throughput and an increase in error rate.

[0007] The intensity profile of millions of clusters on a flow cell may vary between each cluster or between subpopulations of clusters. There are many potential reasons for this variation. The variation may be due to differences in cluster brightness caused by the fragment length distribution of the cluster population, or unwanted light emission from neighboring clusters (spatial crosstalk). The variation may be due to phase errors that occur when molecules in a cluster do not incorporate nucleotides in some sequencing cycles and lag behind other molecules, or when molecules incorporate more than one nucleotide in a single sequencing cycle. The variation may be due to bleaching, i.e., the exponential decay of the signal intensity of a cluster as a function of the number of sequencing cycles due to excessive washing and laser exposure as the sequencing run progresses. The variation may be due to small cluster sizes that produce poorly developed cluster colonies, i.e., empty or only partially filled wells on the patterned flow cell. The variation may be due to overlapping cluster colonies caused by non-exclusive amplification. The variation may be due to poor or uneven illumination, for example, due to clusters being located at the edge of the flow cell. The variation may be due to impurities (e.g., air bubbles) on the flow cell that obscure the emitted signal. The variation may be due to polyclonal clusters, i.e., when multiple clusters are deposited in the same well. The variation may result from different types of distortions in the image caused by the geometry of the optical lens. Such distortions may include, for example, magnification distortion, skew distortion, translation distortion, and nonlinear distortions such as barrel distortion and pincushion distortion.

[0008] This variation can be corrected in a coarse way by training an intensity corrector on the entire cluster population. This can be different from training each intensity corrector on each subpopulation of the cluster. Here, the subpopulations are segmented in a way that minimizes sequencing errors and maximizes base calling accuracy within the available computation. The present disclosure relates to the latter, more granular approach. Further details are provided below. [Brief description of the drawings]

[0009] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings, in which: [Figure 1] 1 illustrates an exemplary sequencing environment having an imaging system. [Diagram 2] FIG. 1 is a block diagram illustrating an exemplary two-channel line-scan modular optical imaging system that may be implemented in certain embodiments. [Diagram 3] Shown is one implementation of respective signal profilers 1 to N trained to maximize the signal-to-noise ratio of respective image data subsets 1 to N generated for respective classes 1 to N of clusters located on a flow cell in respective spatial configurations 1 to N. [Figure 4] 1 illustrates an exemplary configuration of a flow cell that can be imaged according to embodiments disclosed herein. [Figure 5A] The lanes on the top surface of the flow cell are shown. [Figure 5B] Shown is a representation of a tile within a lane on the top surface of a flow cell. [Figure 5C] Shown are tiles within a representation of a lane on the top surface of a flow cell. [Figure 5D] Shown are subtiles within tiles within a representation of a lane on the top surface of a flow cell. [Figure 6A] One implementation is shown that trains each surface-specific specialist signal profiler for each cluster class during a sequencing run 600. [Figure 6B] 1 illustrates one implementation that applies a trained, surface-specific specialist signal profiler to image data subsets corresponding to each cluster class. [Figure 7A]13 illustrates one implementation that trains a respective lane group specific specialist signal profiler for each cluster class. [Figure 7B] 1 illustrates one implementation that trains a respective lane-specific specialist signal profiler for each cluster class. [Figure 7C] 1 shows one implementation that trains a respective exemplar-specific specialist signal profiler for each cluster class. [Figure 7D] 13 illustrates one implementation that trains each tile-specific specialist signal profiler for each cluster class. [Figure 7E] 13 illustrates one implementation that trains each subtile-specific specialist signal profiler for each cluster class. [Figure 8] 1 illustrates one implementation that applies trained lane group-specific specialist signal profilers to image data subsets corresponding to respective cluster classes during a sequencing run. [Figure 9] An implementation is shown that applies trained, lane-specific specialist signal profilers to image data subsets corresponding to each cluster class during a sequencing run. [Figure 10] One implementation is shown that applies trained, exemplar-specific specialist signal profilers to image data subsets corresponding to each cluster class during a sequencing run. [Figure 11] We show one implementation that applies trained tile-specific specialist signal profilers to image data subsets corresponding to respective cluster classes during a sequencing run. [Figure 12] 1 illustrates one implementation that applies trained subtile-specific specialist signal profilers to image data subsets corresponding to respective cluster classes during a sequencing run. [Figure 13]1 shows one implementation of a respective / separate / different / independent specialist signal profiler for each sub-series of sequencing cycles of a sequencing run with a total of N sequencing cycles. [Figure 14] 1 shows one implementation of respective / separate / different / independent specialist signal profilers for combinations of different spatial configurations (e.g., different subtiles) and different temporal configurations (e.g., different subseries of a sequencing cycle). [Figure 15] 1 shows one implementation of a respective / separate / different / independent specialist signal profiler for each cluster / well sequenced during a sequencing run. [Figure 16] 1 illustrates one implementation of offline training of a specialist signal profiler on sequenced data from one or more completed / already performed sequencing runs, and application of the trained specialist signal profiler to sequenced data from an ongoing sequencing run. [Figure 17] 3 illustrates an example tile image divided into subtile images. [Figure 18] 1 illustrates one implementation of online training of a specialist signal profiler on sequenced data from early sequencing cycles of an ongoing sequencing run, and application of the trained specialist signal profiler to sequenced data from later sequencing cycles of an ongoing sequencing run. [Figure 19] 1 shows one implementation that trains a respective / separate / different / independent specialist signal profiler for each signal distribution observed in the sequenced data. [Figure 20] 1 shows an example of a signal distribution / signal profile / cluster intensity profile. [Figure 21] 1 illustrates one implementation of a processing pipeline that implements the disclosed techniques. [Figure 22A]4 illustrates an example spatial equalizer coefficient set. [Figure 22B] 4 illustrates an example spatial equalizer coefficient set. [Figure 22C] 4 illustrates an example spatial equalizer coefficient set. [Figure 22D] 4 illustrates an example spatial equalizer coefficient set. [Figure 22E] 4 illustrates an example spatial equalizer coefficient set. [Figure 23] 1 illustrates one implementation for training an equalizer. [Figure 24] 1 illustrates one implementation for training an equalizer. [Diagram 25] 1 illustrates one implementation for training an equalizer. [Figure 26] 1 illustrates one implementation for training an equalizer. [Figure 27] 1 illustrates one implementation for training an equalizer. [Figure 28] 1 illustrates one implementation for training an equalizer. [Figure 29] 1 illustrates one implementation for training an equalizer. [Diagram 30] 1 illustrates one implementation for training an equalizer. [Diagram 31] 1 illustrates one implementation for training an equalizer. [Figure 32A] 1 shows the signal distribution per base for a cluster population with no equalizer and a signal-to-noise ratio of 11.96 decibels (dB). [Figure 32B] 1 shows the base-by-base signal distribution for the same cluster population using an equalizer, improving the signal-to-noise ratio to 13.13 dB. [Diagram 33] We show how the cost function of the specialist signal profiler improves with each iteration of gradient descent. [Diagram 34] 34 is a plot showing the initial and final values ​​of the cost function of FIG. 33 as the specialist signal profiler is adapted / trained / configured / updated at each sequencing cycle. [Figure 35A]We show improvements in primary analysis metrics for sequencing runs when fitting / training / configuring / updating a specialist signal profiler. [Figure 35B] We show improvements in primary analysis metrics for sequencing runs when fitting / training / configuring / updating a specialist signal profiler. [Figure 36A] 13 shows two plots evaluating the number of subtiles into which a sequencing tile can be divided for adaptive equalization of each specialist signal profiler. [Figure 36B] 13 shows two plots evaluating the number of subtiles into which a sequencing tile can be divided for adaptive equalization of each specialist signal profiler. [Figure 37] FIG. 1 illustrates an exemplary computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] The following discussion is presented to enable those skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0011] First, the signal profiler is described, followed by a description of the disclosed specialist signal profiler.

[0012] Signal Profiler As used herein, a "signal profiler" maximizes the signal-to-noise ratio of a signal that is corrupted by noise. A signal profiler may be a value or function applied to data to modify the data in a desired manner. For example, data may be modified to increase its accuracy, relevance, or applicability with respect to a particular situation. A signal profiler may be applied to data by any of a variety of mathematical operations, including, but not limited to, addition, subtraction, division, multiplication, or combinations thereof. A signal profiler may be a mathematical formula, a logical function, a computer-implemented algorithm, or the like. Data may be image data, electrical data, or combinations thereof.

[0013] In one implementation, the signal profiler is an equalizer (e.g., a spatial equalizer). The equalizer can be trained (e.g., using least squares estimation, adaptive equalization algorithms) to maximize the signal-to-noise ratio of the cluster intensity data in the sequence image. In some implementations, the equalizer is a LUT bank with multiple look-up tables (LUTs), also called "equalizer filters" or "convolution kernels," with sub-pixel resolution. In one implementation, the number of LUTs in the equalizer depends on the number of sub-pixels into which a pixel in the sequencing image can be divided. For example, if a pixel can each be divided into n×n sub-pixels (e.g., 5×5 sub-pixels), the equalizer can be trained to maximize the signal-to-noise ratio of the cluster intensity data in the sequence image (e.g., using least squares estimation, adaptive equalization algorithms). In some implementations, the equalizer is a LUT bank with multiple LUTs with sub-pixel resolution, also called "equalizer filters" or "convolution kernels." In one implementation, the number of LUTs in the equalizer depends on the number of sub-pixels into which a pixel in the sequencing image can be divided. For example, if a pixel can each be divided into n×n sub-pixels (e.g., 5×5 sub-pixels), the equalizer can be trained to maximize the signal-to-noise ratio of the cluster intensity data in the sequence image (e.g., using least squares estimation, adaptive equalization algorithms). 2 LUTs (e.g., 25 LUTs) are generated.

[0014] In one embodiment of equalizer training, data from the sequencing image is binned by well subpixel location. For example, for a 5×5 LUT, approximately 1 / 25th of the wells are centered in bin (1,1) (e.g., the top left corner of the sensor pixel), 1 / 25th of the wells are in bin (1,2), and so on. In one embodiment, the equalizer coefficients for each bin are determined using least squares estimation on the subset of data from the wells that correspond within each bin. In this way, the resulting estimated equalizer coefficients will vary from bin to bin.

[0015] Each LUT / equalizer filter / convolution kernel has multiple coefficients learned from training. In one embodiment, the number of coefficients in the LUT corresponds to the number of pixels used to base call the clusters. For example, if the local grid of pixels (image or pixel patch) used to base call the clusters is of size p×p (e.g., a 9×9 pixel patch), then each LUT has p 2 (e.g., a coefficient of 81).

[0016] In one embodiment, the training generates equalizer coefficients that are configured to mix / combine the intensity values ​​of pixels representing intensity radiation from the base-called target cluster and intensity radiation from one or more neighboring clusters to maximize the signal-to-noise ratio. The signal that is maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that is minimized in the signal-to-noise ratio is the intensity radiation from the neighboring clusters, i.e., spatial crosstalk, plus some random noise (e.g., to account for background intensity radiation). The equalizer coefficients are used as weights, and the mixing / combining involves performing element-wise multiplications between the equalizer coefficients and the intensity values ​​of the pixels to calculate a weighted sum (i.e., a convolution operation) of the intensity values ​​of the pixels. Furthermore, if the image data spans multiple color channels, a set of equalizer coefficients is generated for each color channel (e.g., one channel, three channels, four channels, etc.).

[0017] 22A-22E show equalizer coefficient sets for an exemplary spatial equalizer. As shown by the heatmaps, different equalizer coefficient sets are configured to attenuate and boost a pixel's signal differently depending on the pixel's location.

[0018] Figure 23 illustrates one implementation of training an equalizer. In a first sequencing cycle (cycle 1), the equalizer of Figure 23 has a first set of equalizer coefficients 2302 for the green channel and a second set of equalizer coefficients 2304 for the blue channel. Additionally, for the first sequencing cycle (cycle 1), a first cluster (cluster 1) has input image pixel 2306 for the green channel and input image pixel 2308 for the blue channel.

[0019] In Figure 24, a first set of equalizer coefficients 2302 are multiplied 2402 element-wise with input image pixels 2306 to generate a weighted sum 2316 for the green channel. In Figure 25, a second set of equalizer coefficients 2304 are multiplied 2502 element-wise with input image pixels 2308 to generate a weighted sum 2318 for the blue channel.

[0020] Next, the base calling logic 2322 uses the expectation-maximization (EM) algorithm described above to predict a base call 2324 based on the weighted sums 2316, 2318. In FIG. 26, based on the predicted base calls 2324, the weighted sum for the green channel 2316 is compared to the centroid value 2612 of the called bases for the green channel using the base calling logic 2322. This comparison results in a base call error 2336 for the green channel. Further, in FIG. 26, based on the predicted base calls 2324, the weighted sum for the blue channel 2318 is compared to the centroid value 2712 of the called bases for the blue channel using the base calling logic 2322. This comparison results in a base call error 2338 for the blue channel.

[0021] 28 and 29, the base call errors 2336, 2338 are used by update logic 2342 to generate an updated first set of equalizer coefficients 2356 for the green channel and an updated second set of equalizer coefficients 2358 for the blue channel.

[0022] In some implementations, the above steps are performed for multiple clusters. For example, for the three clusters shown in Figures 30 and 31, three updated versions of the first set of equalizer coefficients for the green channel and three updated versions of the second set of equalizer coefficients for the blue channel are generated. The three updated versions are used to calculate the first set of equalizer coefficients 2362 for the green channel and the second set of equalizer coefficients 2354 for the blue channel for the second sequencing cycle (cycle 2).

[0023] Figure 32A shows the signal distribution per base for a cluster population without an equalizer and with a signal-to-noise ratio of 11.96 decibels (dB). Figure 32B shows the signal distribution per base for the same cluster population with an equalizer and with an improved signal-to-noise ratio of 13.13 dB. The improvement in signal-to-noise ratio is further visually observable by the tighter / more discrete per base cloud / distribution in Figure 32B compared to the per base cloud in Figure 32A.

[0024] Further details about the equalizer can be found in U.S. Provisional Patent Application No. 17 / 308,035, entitled "Equalization-Based Image Processing and Spatial Crosstalk Attenuator," filed May 4, 2021 (Attorney Docket No. ILLM 1032-2 / IP-1991-US), which is incorporated by reference as if fully set forth herein.

[0025] Next, the specialist signal profiler will be described.

[0026] Specialist Signal Profiler Global intensity correction applied at the pan flow cell level or pan sequencing run level cannot account for various noises in the image data. For example, nonlinear distortions and noises may be caused by the shape of the optical lens capturing the image data. In addition, the imaged flow cell may also introduce distortions of the well pattern due to the manufacturing process (e.g., 3D bathtub effect introduced by binding or movement of wells due to non-rigidity of the substrate). Finally, tilting of the flow cell in the holder is not accounted for by the global intensity correction.

[0027] As used herein, a "specialist signal profiler" is a signal profiler configured / trained to maximize the signal-to-noise ratio of a particular category / type / configuration / characteristic / class / bin of data. Various specialist signal profilers are disclosed. For example, a "surface-specific specialist signal profiler" is configured / trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular surface or a particular surface type / category / class (e.g., top or bottom surface of a flow cell or surfaces 1-N). Similarly, a "lane-specific specialist signal profiler" is configured / trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular lane or a particular lane type / category / class (e.g., center or peripheral lanes of a flow cell or lanes 1-N). Furthermore, a "tile-specific specialist signal profiler" is configured / trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular tile or a particular tile type / category / class (e.g., center or peripheral tiles of a flow cell or tiles 1-N). Furthermore, "subtile specific specialist signal profilers" are configured / trained to maximize the signal to noise ratio of sequencing data for clusters located on a particular subtile or a particular subtile type / category / class (e.g., central or peripheral subtiles or subtiles 1-N of a flow cell). Further examples and details of the disclosed specialist signal profilers are described below.

[0028] In some implementations, a single signal profiler may comprise multiple specialist coefficient sets, such that each specialist coefficient set is configured / trained to maximize the signal-to-noise ratio of a particular category / type / configuration / characteristic / class / bin of data. In some implementations, a single signal profiler may comprise various specialist coefficient sets. For example, a "surface-specific specialist coefficient set" is configured / trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular surface or a particular surface type / category / class (e.g., top or bottom surface or surfaces 1-N of a flow cell). Similarly, a "lane-specific specialist coefficient set" is configured / trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular lane or a particular lane type / category / class (e.g., center or peripheral lane or lanes 1-N of a flow cell). Further, the "tile-specific specialist coefficient sets" are configured / trained to maximize the signal-to-noise ratio of sequencing data for clusters located on a particular tile or a particular tile type / category / class (e.g., central or peripheral tiles of a flow cell or tiles 1-N). Further, the "subtile-specific specialist coefficient sets" are configured / trained to maximize the signal-to-noise ratio of sequencing data for clusters located on a particular subtile or a particular subtile type / category / class (e.g., central or peripheral subtiles of a flow cell or subtiles 1-N). Further examples and details of the disclosed specialist coefficient sets are described below.

[0029] The disclosed specialist signal profiler is applicable to clusters located on both patterned and non-patterned surfaces of a flow cell. On a non-patterned surface, the clusters are randomly distributed on the flow cell. The randomly distributed clusters and the data therefor (e.g., images) can be binned spatially, temporally, by signal, or by any combination of the above. Thus, the specialist signal profiler can be configured and trained for different configurations of differently binned and randomly distributed clusters. On a patterned surface, the clusters are placed on patterned wells with fixed positions. The patterned wells and component clusters can be binned spatially, temporally, by signal, or by any combination of the above. Thus, the specialist signal profiler can be configured and trained for different configurations of differently binned patterned clusters.

[0030] The disclosed specialist signal profiler is a configuration-specific signal profiler trained to maximize the signal-to-noise ratio of image data generated for different configurations of a sequencing run. These configurations may be spatial configurations associated with different regions on a flow cell, temporal configurations associated with different sequencing / imaging cycles of a sequencing run, signal distribution configurations associated with different distributions / patterns of signal profiles observed / encoded in the imaging data, or combinations of the above. Before describing in more detail the various implementations of the systems and methods disclosed herein, it is useful to describe an example environment in which the techniques disclosed herein may be implemented.

[0031] Sequencing Environment 1 illustrates an exemplary sequencing environment having an imaging system 100. The exemplary imaging system 100 may include a device for acquiring or generating an image of a sample. The example outlined in FIG. 1 illustrates an exemplary imaging configuration of an embodiment of a backlight design. It should be noted that although systems and methods may sometimes be described herein in the context of the exemplary imaging system 100, these are merely examples in which embodiments of the specialist signal profiler disclosed herein may be implemented.

[0032] As can be seen in the example of FIG. 1, a sample of interest is located on a sample container 110 (e.g., a flow cell as described herein) and is positioned on a sample stage 170 under an objective lens 142. A light source 160 and associated optics direct a light beam, such as a laser light, to a selected sample location on the sample container 110. The sample fluorescence and resulting light are collected by the objective lens 142 and directed to an image sensor of the camera system 140 to detect the fluorescence. The sample stage 170 is moved relative to the objective lens 142 to position the next sample location on the sample container 110 at the focal point of the objective lens 142. The movement of the sample stage 110 relative to the objective lens 142 can be accomplished by moving the sample stage itself, the objective lens, some other component of the imaging system 100, or any combination of the above. Further embodiments may also include moving the entire imaging system 100 over a stationary sample.

[0033] The fluid delivery module or device 100 directs the flow of reagents (e.g., fluorescently labeled nucleotides, buffers, enzymes, cleavage reagents, etc.) to (and through) the sample container 110 and waste valve 120. The sample container 110 may include one or more substrates on which the sample is provided. For example, in the case of a system that analyzes a large number of different nucleic acid sequences, the sample container 110 may include one or more substrates on which the nucleic acids to be sequenced are bound, attached, or associated. In various embodiments, the substrate may include any inert substrate or substrate to which nucleic acids may be attached, such as, for example, glass surfaces, plastic surfaces, latex, dextran, polystyrene surfaces, polypropylene surfaces, polyacrylamide gels, gold surfaces, and silicon wafers. In some applications, the substrates are within channels or other areas in multiple locations formed in a matrix or array across the sample container 110.

[0034] In some embodiments, the sample container 110 may contain a biological sample to be imaged using one or more fluorescent dyes. For example, in certain embodiments, the sample container 110 may be implemented as a patterned flow cell that includes a transparent cover plate, a substrate, and a liquid sandwiched between the transparent cover plate and the substrate, and the biological sample may be disposed on the inner surface of the transparent cover plate or the inner surface of the substrate. The flow cell may include a large number (e.g., thousands, millions, or billions) of wells or regions within the substrate that are patterned into a defined array (e.g., a hexagonal array, a rectangular array, etc.). Each region may form a cluster (e.g., a monoclonal cluster) of biological samples, such as DNA, RNA, or another genomic material, that may be sequenced using sequencing by synthesis. The flow cell may be further divided into a large number of spaced lanes (e.g., eight lanes), each lane containing a hexagonal array of clusters. An example of a flow cell that may be used in the embodiments disclosed herein is described in U.S. Pat. No. 8,778,848.

[0035] The system further comprises a temperature station actuator 130 and a heater / cooler 135 that can optionally adjust the temperature of the fluid state in the sample vessel 110. A camera system 140 can be included to monitor and track the sequencing of the sample vessel 110. The camera system 140 can be implemented, for example, as a charge-coupled device (CCD) camera (e.g., a time difference integration (TDI) CCD camera) and can interact with various filters in the filter switching assembly 145, the objective lens 142, and the focused laser / focused laser assembly 150. The camera system 140 is not limited to a CCD camera, and other camera and image sensor technologies can be used. In certain implementations, the camera sensor can have a pixel size between about 5 pm and about 15 pm.

[0036] Output data from the sensor of camera system 140 may be communicated to a real-time analysis module (not shown), which may be implemented as a software application that analyzes the image data (e.g., image quality scoring), reports or displays characteristics of the laser beam (e.g., focus, shape, intensity, power, brightness, position) in a graphical user interface (GUI), and dynamically corrects intensity noise in the image data as further described below.

[0037] A light source 160 (e.g., an excitation laser in an assembly optionally including multiple lasers) or other light source may be included to illuminate the fluorescent sequencing reaction in the sample by illumination through a fiber optic interface (which may optionally include one or more reimaging lenses, fiber optic attachments, etc.). In the illustrated example, a low wattage lamp 165, a focused laser 150, and an inverse dichroic 185 are also presented. In some implementations, the focused laser 150 may be turned off during imaging. In other implementations, an alternative focus configuration may include a second focused camera (not shown), which may be a quadrant detector, a position sensitive detector (PSD), or similar detector for measuring the location of the scattered beam reflected from the surface simultaneously with data collection.

[0038] Although illustrated as a backlit device, other examples may include light from a laser or other light source directed through the objective lens 142 to the sample on the sample container 110. The sample container 110 may ultimately be mounted to a sample stage 170 to provide movement and alignment of the sample container 110 relative to the objective lens 142. The sample stage may have one or more actuators that allow it to move in any of three dimensions. For example, actuators may be provided that allow the stage to move in the X, Y, and Z directions relative to the objective lens in terms of a Cartesian coordinate system. This may allow one or more sample locations on the sample container 110 to be positioned in optical alignment with the objective lens 142.

[0039] In this example, a focus (z-axis) component 175 is shown as being included to control the positioning of the optical components relative to the sample container 110 in a focus direction (typically referred to as the z-axis or z-direction). The focus component 175 may include one or more actuators physically coupled to the optical stage or the sample stage, or both, to move the sample container 110 on the sample stage 170 relative to the optical components (e.g., the objective lens 142) to provide proper focusing for the imaging operation. For example, the actuators may be physically coupled to the respective stages, e.g., by direct or indirect mechanical, magnetic, fluidic, or other connections or contacts with the stages. The one or more actuators may be configured to move the stage in the z-direction while keeping the sample stage in the same plane (e.g., while maintaining a level or horizontal orientation perpendicular to the optical axis). The one or more actuators may also be configured to tilt the stage. This is done, for example, so that the sample container 110 may be dynamically flattened to compensate for any tilt in its surface.

[0040] Focusing the system generally refers to aligning the focal plane of the objective lens with the sample being imaged at a selected sample location. However, focusing can also refer to adjusting the system to obtain desired characteristics for the representation of the sample, such as a desired level of sharpness or contrast for the image of the test sample. Because the usable depth of field of the focal plane of the objective lens may be small (sometimes on the order of 1 pm or less), the focusing component 175 closely follows the surface being imaged. Because the sample container, as fixed to the instrument, is not perfectly flat, the focusing component 175 can be set to follow this profile while moving along the scan direction (referred to herein as the y-axis).

[0041] Light emanating from the test sample at the sample location may be directed to one or more detectors of the camera system 140. An aperture may be included and positioned to allow only light emanating from the focal region to pass to the detector. The aperture may be included to improve image quality by filtering components of the light emanating from regions outside the focal region. An emission filter may be included in the filter switching assembly 145 and selected to record the determined emission wavelength and to cut out any stray laser light.

[0042] Although not shown, a controller may be provided to control the operation of the scanning system. The controller may be implemented to control aspects of the system operation, such as, for example, focusing, stage movement, and imaging operations. In various embodiments, the controller may be implemented using hardware, algorithms (e.g., machine executable instructions), or a combination of the above. For example, in some embodiments, the controller may include one or more CPUs or processors having associated memory. As another example, the controller may include hardware or other circuitry for controlling operations, such as a computer processor and a non-transitory computer readable medium having machine readable instructions stored thereon. For example, the circuitry may include one or more of the following: a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logic device (CPLD), a programmable logic array (PLA), a programmable array logic (PAL), or other similar processing devices or circuitry. As yet another example, a controller may include a combination of this circuitry with one or more processors.

[0043] 2 is a block diagram illustrating an exemplary two-channel line-scan modular optical imaging system 200 as may be implemented in certain embodiments. Although systems and methods may sometimes be described herein in the context of the exemplary imaging system 200, it should be noted that these are merely examples in which embodiments of the techniques disclosed herein may be implemented.

[0044] In some embodiments, the system 200 can be used for sequencing nucleic acids. Applicable techniques include those in which nucleic acids are attached to fixed locations in an array (e.g., wells of a flow cell) and the array is repeatedly imaged. In such embodiments, the system 200 can acquire images of two different color channels that can be used to distinguish a particular nucleotide base type from another nucleotide base. More specifically, the system 200 can perform a process called "base calling," which generally refers to the process of determining a base call (e.g., adenine (A), cytosine (C), guanine (G), or thymine (T)) for a given spot location of an image in an imaging cycle. During two-channel base calling, image data extracted from the two images can be used to determine the presence of one of the four base types by encoding the base identity as a combination of the intensities of the two images. For a given spot or location in each of the two images, the base identity can be determined based on whether the signal identity combination is [on, on], [on, off], [off, on], or [off, off].

[0045] Referring again to the imaging system 200, the system includes a line generation module (LGM) 210 having two light sources 211 and 212 disposed therein. The light sources 211 and 212 may be coherent light sources such as laser diodes that output laser beams. The light source 211 may emit light at a first wavelength (e.g., a red wavelength) and the light source 212 may emit light at a second wavelength (e.g., a green wavelength). The light beams output from the laser sources 211 and 212 may be directed through a beam shaping lens(es) 213. In some implementations, a single light shaping lens may be used to shape the light beams output from both light sources. In other implementations, separate beam shaping lenses may be used for each light beam. In some examples, the beam shaping lens is a Powell lens, so that the light beams are shaped into a line pattern. Beam-Shaping Lenses or Other Optical Components of LGM 210 Imaging system 200 may be configured to shape the light emitted by light sources 211 and 212 into a line pattern (e.g., by using one or more Powell lenses, or other beam-shaping lenses, diffractive or scattering components).

[0046] The LGM 210 may further include a mirror 214 and a semi-reflective mirror 215 for configuring to direct the light beam through a single interface port to an emission optics module (EOM) 230. The light beam may pass through a shutter element 216. The EOM 230 may include an objective lens 235 and a z-stage 236 for moving the objective lens 235 longitudinally toward or away from a target 250. For example, the target 250 may include a liquid layer 252 and a semi-transparent cover plate 251, and the biological sample may be located on an inner surface of the semi-transparent cover plate as well as an inner surface of a substrate layer located below the liquid layer. The z-stage may then move the objective lens to focus the light beam on any inner surface of the flow cell (e.g., to focus on the biological sample). The biological sample may be DNA, RNA, protein, or other biological material responsive to optical sequencing as known in the art.

[0047] The EOM 230 may include a semi-reflective mirror 233 for reflecting a focus tracking light beam emitted from a focus tracking module (FTM) 240 onto the target 250 and then reflecting light returned from the target 250 back to the FTM 240. The FTM 240 may include a focus tracking optical sensor for detecting characteristics of the returned focus tracking light beam and generating a feedback signal for optimizing the focus of the objective lens 235 on the target 250.

[0048] EOM 230 may further include a semi-reflective mirror 234 for directing the light through objective lens 235 while allowing light returned from target 250 to pass through. In some embodiments, EOM 230 may include a tube lens 232. Light transmitted through tube lens 232 may pass through filter element 231 and into camera module (CAM) 220. CAM 220 may include one or more optical sensors 221 for detecting light emitted from the biological sample in response to the incident light beam (e.g., fluorescence in response to red and green light received from light sources 211 and 212).

[0049] Output data from the sensors of the CAM 220 may be communicated to a real-time analysis module 225. The real-time analysis module, in various embodiments, executes computer-readable instructions for analyzing image data (e.g., image quality scoring, base calling, etc.), reporting or displaying beam characteristics (e.g., focus, shape, intensity, power, brightness, position) in a graphical user interface (GUI), etc. These operations may be performed in real-time during an imaging cycle to minimize downstream analysis time and provide real-time feedback and troubleshooting during an imaging run. In an embodiment, the real-time analysis module may be a computing device (e.g., computing device 1000) communicatively connected to and controlling the imaging system 200. In implementations described further below, the real-time analysis module 225 may further execute computer-readable instructions for maximizing the signal-to-noise ratio of output image data received from the CAM 220.

[0050] Sequencing produces m sequence images for each sequencing cycle for the corresponding m image channels. That is, each sequence image has one or more image (or intensity) channels (similar to the red, green, and blue (RGB) channels of a color image). In one embodiment, each image channel corresponds to one of a number of filter wavelength bands. In another embodiment, each image channel corresponds to one of a number of imaging events in a sequencing cycle. In yet another embodiment, each image channel corresponds to a combination of illumination by a particular laser and imaging through a particular optical filter. Image patches are tiled (or accessed) from each of the m image channels for a particular sequencing cycle. In different embodiments, such as 4-channel chemistry, 2-channel chemistry, and 1-channel chemistry, m is 4 or 2. In other embodiments, m is greater than 1, 3, or 4. In other implementations, the images can be blue and violet channels instead of or in addition to red and green channels.

[0051] Spatial configuration specific specialist signal profiler 3 illustrates one implementation of respective / separate / different / independent specialist signal profilers 1-N trained to maximize the signal to noise ratio of respective image data subsets 1-N generated for respective classes 1-N of clusters located on a flow cell in respective spatial configurations 1-N. A sequencing run 300 sequences a population of clusters over multiple sequencing / imaging cycles 1-K and generates image data indicative of the intensity emission of the population of clusters.

[0052] With respect to image data generated during the sequencing run 300, imaging cycles of the flow cell are performed such that image data can be collected for the entire flow cell by scanning the flow cell area with one or more coherent light sources (e.g., using a line scanner). As an example, the imaging system 200 can use the LGM 210 in coordination with the system's optics to line scan the flow cell with light having wavelengths in the red spectrum and line scan the sample with light having wavelengths in the green spectrum. In response to the line scan, fluorescent dyes located in different clusters of the flow cell fluoresce, and the resulting light can be collected by the objective lens 235 and directed to the image sensor of the CAM 220 to detect the fluorescence. For example, the fluorescence of each cluster can be detected by several pixels of the CAM 220. The image data output from the CAM 220 can then be communicated to the real-time analysis module 225 for image noise correction and base calling.

[0053] The population of clusters is spatially distributed across the flow cell. The flow cell, and therefore the population of clusters and image data, can be divided into different spatial configurations defined by different regions of the flow cell. Thus, for example, if the flow cell can be divided into three rectangular regions, this results in three spatial configurations of the flow cell, three subpopulations or classes of clusters, three subsets of image data, and three specialist signal profilers. At the next granularity scale, each of the three rectangular regions of the flow cell can be further divided into three squares, resulting in a total of nine squares. This then results in nine spatial configurations of the flow cell, nine subpopulations or classes of clusters, nine subsets of image data, and nine specialist signal profilers.

[0054] With respect to segmenting the image data, the image data is divided into a number of image data subsets corresponding to respective regions of the flow cell. In various embodiments, the size of the image data subsets may be determined using the placement of fiducial markers or fiducials within the field of view of the imaging system 200, within the flow cell, or on the flow cell. The image data subsets may be divided such that the pixels of each image data subset have a predetermined number of fiducial points (e.g., at least three fiducials, four fiducials, six fiducials, eight fiducials, etc.). For example, the total number of pixels of the image data subset may be predetermined based on a predetermined pixel distance between the boundary of the image data subset and the fiducials.

[0055] Additionally, the intensity data in each subset of image data may span multiple color channels (e.g., red, blue, green, and / or purple image channels), and thus each specialist signal profiler has a respective set of coefficients trained to maximize the signal-to-noise ratio of the intensity data in the corresponding color channel.

[0056] During inference, new / unseen / wild image data is segmented with the same criteria as the spatial configurations defined and used to generate each / separate / different / independent specialist signal profiler during training. Each / separate / different / independent specialist signal profiler is applied to a segmented subset of new image data such that application of a particular specialist signal profiler is limited to only the corresponding subset of new image data.

[0057] Application of the Specialist Signal Profilers 1-N to the image data subsets 1-N produces signal-to-noise ratio maximized image data subsets 1-N, additional details of which may be found in U.S. Nonprovisional Patent Application No. 17 / 308,035, entitled "Equalization-Based Image Processing and Spatial Crosstalk Attenuator," filed May 4, 2021 (Attorney Docket No. ILLM 1032-2 / IP-1991-US). Other corrections for channel noise, spatial noise, or phase noise may be applied a priori to the unprocessed image data or a posteriori to the signal-to-noise ratio maximized image data subsets 1-N.

[0058] The output of the specialist signal profilers 1-N, i.e., the signal-to-noise ratio maximized image data subsets 1-N, are provided as input to the base caller 332 to generate base calls 1-N for the population of clusters. The base calls may be made by fitting a mathematical model to the intensity data. Suitable mathematical models that may be used include, for example, k-means clustering algorithms, k-means-like clustering algorithms, expectation maximization clustering algorithms, histogram-based methods, and the like. Four Gaussian distributions may be fitted to the set of two-channel intensity data such that one distribution is applied to each of the four nucleotides represented in the data set. In one particular implementation, an expectation maximization (EM) algorithm may be applied. As a result of the EM algorithm, for each X,Y value (each referring to each of the two channel intensities), a value may be generated that represents the likelihood that the X,Y intensity value belongs to one of the four Gaussian distributions to which the data is fitted. If the four bases give four separate distributions, then each X,Y intensity value further has four associated likelihood values, one for each of the four bases. The maximum of the four likelihood values ​​indicates the base call. For example, if the cluster is "off" in both channels, the base call is G. If the cluster is "off" in one channel and "on" in another, the base call is either C or T (depending on which channel is on), and if the cluster is "on" in both channels, the base call is A.

[0059] FIG. 4 shows an exemplary configuration of a flow cell 400 that can be imaged according to embodiments disclosed herein. In some implementations, the flow cell 400 is patterned with regular clusters or hexagonal arrays of spots or features that can be imaged simultaneously during an imaging run. In other implementations, the flow cell 400 can be patterned using a linear array, a circular array, an octagonal array, or some other array pattern. The flow cell 400 can have tens, hundreds, thousands, millions, or billions of clusters that are imaged. In certain embodiments, the flow cell 400 can be patterned with millions or billions of wells divided into lanes. In this particular implementation, each well of the flow cell 400 can include at least one cluster.

[0060] Surface-specific specialist signal profiler In some implementations, the flow cell 400 can be a multi-sided sample that includes multiple sides of a cluster that are sampled during an imaging run. Examples of multiple sides of a cluster include a top side 402 and a bottom side 412. Thus, in one implementation, the flow cell 400 can have two spatial configurations corresponding to the top side 402 and the bottom side 412, resulting in two subpopulations or classes of clusters, two subsets of image data, and two specialist signal profilers.

[0061] 6A illustrates one implementation of training a respective surface-specific specialist signal profiler 604 for each cluster class 602. Each cluster class 602 includes a cluster group located on the top surface 402 and a cluster group located on the bottom surface 412. The result is a first specialist signal profiler 604a configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the top surface 402 and a second specialist signal profiler 604b configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the bottom surface 412.

[0062] 6B illustrates an implementation in which a trained surface-specific specialist signal profiler 604 is applied to image data subsets 632, 642 corresponding to each cluster class 602 during a sequencing run 600. In one implementation, the flow cell 400 is imaged at the tile level. Thus, assuming the top surface 402 has a first set of 1600 tiles and the bottom surface 412 has a second set of 1600 tiles, the first image data subset 632 includes 1600 tile images of the first set of 1600 tiles, and the second image data subset 642 includes 1600 tile images of the second set of 1600 tiles.

[0063] The first specialist signal profiler 604a is configured to maximize the signal-to-noise ratio of the intensity data in the first image data subset 632 to generate a signal-to-noise ratio maximized version 634 of the first image data subset 632. The second specialist signal profiler 604b is configured to maximize the signal-to-noise ratio of the intensity data in the second image data subset 642 to generate a signal-to-noise ratio maximized version 644 of the second image data subset 642. The base caller 332 processes the signal-to-noise ratio maximized versions 634, 644 and generates base calls 638, 648.

[0064] As used herein, the term "maximized version" refers to data generated as an output by a specialist signal profiler. A maximized version of an input is an output generated by a specialist signal profiler in response to processing of a corresponding input. The maximized version of an input has a greater signal-to-noise ratio (SNR, S / R) than the corresponding input processed by the specialist signal profiler. For example, the maximized version of an input (i.e., output) generated by a specialist signal profiler is corrected by the specialist signal profiler to have less spatial crosstalk and background noise compared to the corresponding input. Similarly, the maximized version of an input (i.e., output) generated by a specialist signal profiler is corrected by the specialist signal profiler to have less fading and pre-fading noise compared to the corresponding input.

[0065] Lane-specific specialist signal profiler In one implementation, the top surface 402 can be divided or partitioned into multiple lanes 508a, 508b, ..., 5081. In the example shown in Figure 5A, the top surface 402 has eight lanes, although the number of lanes is implementation specific. Thus, in one implementation, the flow cell 400 can have 16 spatial configurations corresponding to the eight lanes on the top surface 402 and the eight lanes on the bottom surface 412, resulting in 16 subpopulations or classes of clusters, 16 subsets of image data, and 16 specialist signal profilers.

[0066] 7B illustrates one implementation of training a respective lane-specific specialist signal profiler 714 for each cluster class 712. Each cluster class 712 includes a group of clusters located on lanes 508a, 508b, ..., 5081, respectively. The result is a first specialist signal profiler 714a configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the first lane 508a, a second specialist signal profiler 714b configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the second lane 508b, etc. (following the first specialist signal profiler 7141 configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the first lane 5081).

[0067] 9 illustrates an implementation in which a trained lane-specific specialist signal profiler 714 is applied to image data subsets 902, 912, ..., 922 corresponding to each cluster class 712 during a sequencing run 900. In one implementation, the flow cell 400 is imaged at the tile level. Thus, assuming each lane has 200 tiles, the first image data subset 902 includes a first set of 200 tile images for the 200 tiles on the first lane 508a, the second image data subset 912 includes a second set of 200 tile images for the 200 tiles on the second lane 508b, and so on (following the first image data subset 922).

[0068] The first specialist signal profiler 714a is configured to maximize the signal-to-noise ratio of the intensity data in the first image data subset 902 to generate a signal-to-noise ratio maximized version 904 of the first image data subset 902. The second specialist signal profiler 714b is configured to maximize the signal-to-noise ratio of the intensity data in the second image data subset 912 to generate a signal-to-noise ratio maximized version 914 of the second image data subset 912, and so on (following a signal-to-noise ratio maximized version 924 of the first image data subset 922). The base caller 332 processes the signal-to-noise ratio maximized versions 904, 914, ..., 924 and generates base calls 908, 918, ..., 928.

[0069] Lane-group specific specialist signal profilers In some implementations, the lanes may be grouped into lane groups 502a, 502b, and 502c. Examples of lane groups include upper peripheral lanes, central lanes, and lower peripheral lanes. Thus, in one implementation, the flow cell 400 may have six spatial configurations corresponding to three lane groups on the top surface 402 and three lane groups on the bottom surface 412, resulting in six subpopulations or classes of clusters, six subsets of image data, and six specialist signal profilers.

[0070] 7A illustrates one implementation of training a respective lane group specific specialist signal profiler 704 for each cluster class 702. Each cluster class 702 includes a group of clusters located on lane groups 502a, 502b, 502c, respectively. The result is a first specialist signal profiler 704a configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the first lane group 502a, a second specialist signal profiler 704b configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the second lane group 502b, and a third specialist signal profiler 704c configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the third lane group 502c.

[0071] 8 illustrates an implementation in which trained lane group specific specialist signal profilers 704 are applied to image data subsets 802, 812, 822 corresponding to respective cluster classes 702 during a sequencing run 800. In one implementation, the flow cell 400 is imaged at the tile level. Thus, assuming that the first lane group 502a has 600 tiles, the second lane group 502b has 600 tiles, and the third lane group 502c has 400 tiles, the first image data subset 802 includes a first set of 600 tile images for the 600 tiles on the first lane group 502a, the second image data subset 812 includes a second set of 600 tile images for the 600 tiles on the second lane group 502b, and the third image data subset 822 includes a third set of 400 tile images for the 400 tiles on the third lane group 502c.

[0072] The first specialist signal profiler 704a is configured to maximize the signal-to-noise ratio of the intensity data in the first image data subset 802 to generate a signal-to-noise ratio maximized version 804 of the first image data subset 802. The second specialist signal profiler 704b is configured to maximize the signal-to-noise ratio of the intensity data in the second image data subset 812 to generate a signal-to-noise ratio maximized version 814 of the second image data subset 812. The third specialist signal profiler 704c is configured to maximize the signal-to-noise ratio of the intensity data in the third image data subset 822 to generate a signal-to-noise ratio maximized version 824 of the third image data subset 822. The base caller 332 processes the signal-to-noise ratio maximized versions 804, 814, 824 and generates base calls 808, 818, 828.

[0073] Specimen-specific specialist signal profiler In some implementations, each lane includes one or more rows / samples of tiles 506a, 506b, as shown in Figure 5B. Thus, in one implementation, the flow cell 400 can have 32 spatial configurations corresponding to 32 samples of tiles on the top surface 402 and 32 samples of tiles on the bottom surface 412, resulting in 32 subpopulations or classes of clusters, 32 subsets of image data, and 32 specialist signal profilers.

[0074] 7C illustrates one implementation of training a respective exemplar-specific specialist signal profiler 724 for each cluster class 722. Each cluster class 722 includes a group of clusters located on the exemplars 506a, 506b, respectively. The result is a first specialist signal profiler 724a configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the first exemplar 506a, and a second specialist signal profiler 724b configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the second exemplar 506b.

[0075] 10 illustrates an implementation in which a trained exemplar-specific specialist signal profiler 724 is applied to image data subsets 1002, 1012 corresponding to respective cluster classes 722 during a sequencing run 1000. In one implementation, the flow cell 400 is imaged at the tile level. Thus, assuming a first exemplar 506a has 100 tiles and a second exemplar 506b has 100 tiles, the first image data subset 1002 includes a first set of 100 tile images for the 100 tiles on the first exemplar 506a, and the second image data subset 1012 includes a second set of 100 tile images for the 100 tiles on the second exemplar 506b.

[0076] The first specialist signal profiler 724a is configured to maximize the signal-to-noise ratio of the intensity data in the first image data subset 1002 to generate a signal-to-noise ratio maximized version 1004 of the first image data subset 1002. The second specialist signal profiler 724b is configured to maximize the signal-to-noise ratio of the intensity data in the second image data subset 1012 to generate a signal-to-noise ratio maximized version 1014 of the second image data subset 1012. The base caller 332 processes the signal-to-noise ratio maximized versions 1004, 1014, 1024 and generates base calls 1008, 1018.

[0077] Tile-specific specialist signal profiler Each swatch includes multiple tiles 512a, 512b, ..., 512t, as shown in Figure 5B. The number of tiles in each swatch is implementation specific and may be 50 tiles, 60 tiles, 80 tiles, etc. in different embodiments. For example, consider each swatch to have 100 tiles. The flow cell 400 then has 200 tiles per lane, resulting in 1600 tiles for the top surface 402 and another 1600 tiles for the bottom surface 412 (i.e., a total of 3200 tiles). Thus, in one implementation, the flow cell 400 may have 3200 spatial configurations corresponding to the 3200 tiles, resulting in 3200 subpopulations or classes of clusters, 3200 subsets of image data, and 3200 specialist signal profilers.

[0078] 7D illustrates one implementation that trains a respective tile-specific specialist signal profiler 734 for each cluster class 732. Each cluster class 732 includes a group of clusters located on tiles 512a, 512b, ..., 512t, respectively. The result is a first specialist signal profiler 734a configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the first tile 512a, a second specialist signal profiler 734b configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the second tile 512b, etc. (followed by a first specialist signal profiler 734t configured to maximize the signal-to-noise ratio of the intensity data of the clusters located on the first tile 512t).

[0079] FIG. 11 illustrates one implementation that applies a trained, tile-specific specialist signal profiler to image data subsets 1102, 1112, ..., 1122 corresponding to respective cluster classes 732 during a sequencing run 1100. In one implementation, the flow cell 400 is imaged at the tile level. Thus, for tiles 512a, 512b, ..., 512t, a first image data subset 1102 includes a first tile image of the first tile 512a, a second image data subset 1112 includes a second tile image of the second tile 512b, and so on (the tth tile image of the tth tile 512t including the tth tile image of the tth tile 512t). th Continue with image data subset 1122 (as shown in FIG. 5C).

[0080] The first specialist signal profiler 734a is configured to maximize the signal-to-noise ratio of the intensity data in the first image data subset 1102 to generate a signal-to-noise ratio maximized version 1104 of the first image data subset 1102. The second specialist signal profiler 734b is configured to maximize the signal-to-noise ratio of the intensity data in the second image data subset 1112 to generate a signal-to-noise ratio maximized version 1114 of the second image data subset 1112, and so on (following a signal-to-noise ratio maximized version 1124 of the tth image data subset 1122). The base caller 332 processes the signal-to-noise ratio maximized versions 1104, 1114, ..., 1124 and generates base calls 1108, 1118, ..., 1128.

[0081] Sub-tile specific specialist signal profiler Each tile can be divided into multiple subtiles 518a, 518b, ..., 518s, as shown in Figure 5D. The number of subtiles into which a tile is divided is implementation specific, and in different embodiments there may be 10 subtiles, 30 subtiles, 50 subtiles, etc. For example, consider that each tile is divided into 9 subtiles. The flow cell 400 then has a total of 28,800 subtiles for 3200 tiles. Thus, in one implementation, the flow cell 400 can have 28,800 spatial configurations corresponding to the 28,800 tiles, resulting in 28,800 subpopulations or classes of clusters, 28,800 subsets of image data, and 28,800 specialist signal profilers.

[0082] 7E illustrates one implementation of training each subtile-specific specialist signal profiler 744 for each cluster class 742. Each cluster class 742 includes a group of clusters that are located on subtiles 518a, 518b, 518c, 518d, ..., 518s, respectively. The result is a first specialist signal profiler 744a configured to maximize the signal-to-noise ratio of intensity data of clusters located on the first subtile 518a, a second specialist signal profiler 744b configured to maximize the signal-to-noise ratio of intensity data of clusters located on the second subtile 518b, a third specialist signal profiler 744c configured to maximize the signal-to-noise ratio of intensity data of clusters located on the third subtile 518c, a fourth specialist signal profiler 744d configured to maximize the signal-to-noise ratio of intensity data of clusters located on the fourth subtile 518d, and so on (following the sth specialist signal profiler 744s configured to maximize the signal-to-noise ratio of intensity data of clusters located on the sth subtile 518s).

[0083] 12 illustrates an implementation in which trained subtile-specific specialist signal profilers 744 are applied to image data subsets 1202, 1212, 1222, 1232, ..., 1242 corresponding to respective cluster classes 742 during a sequencing run 1200. In one implementation, the flow cell 400 is imaged at the subtile level. Thus, for 518a, 518b, 518c, 518d, ..., 518s, the first image data subset 1202 includes the first subtile image patch of the first subtile 518a, the second image data subset 1212 includes the second subtile image patch of the second subtile 518b, the third image data subset 1222 includes the third subtile image patch of the third subtile 518c, the fourth image data subset 1232 includes the fourth subtile image patch of the fourth subtile 518d, and so on (following the sth image data subset 1242 which includes the sth subtile image patch of the sth subtile 518s).

[0084] The first specialist signal profiler 744a is configured to maximize the signal-to-noise ratio of the intensity data in the first image data subset 1202 to generate a signal-to-noise ratio maximized version 1204 of the first image data subset 1202. The second specialist signal profiler 744b is configured to maximize the signal-to-noise ratio of the intensity data in the second image data subset 1212 to generate a signal-to-noise ratio maximized version 1214 of the second image data subset 1212. The third specialist signal profiler 744c is configured to maximize the signal-to-noise ratio of the intensity data in the third image data subset 1222 to generate a signal-to-noise ratio maximized version 1224 of the third image data subset 1222. A fourth specialist signal profiler 744d is configured to maximize the signal-to-noise ratio of the intensity data in the fourth image data subset 1232 to generate a signal-to-noise ratio maximized version 1234 of the fourth image data subset 1232, and so on (following a signal-to-noise ratio maximized version 1244 of the sth image data subset 1242). The base caller 332 processes the signal-to-noise ratio maximized versions 1204, 1214, 1224, 1234, ..., 1244 and generates base calls 1208, 1218, 1228, 1238, ..., 1248.

[0085] Temporal configuration-specific specialist signal profiler FIG. 13 illustrates one implementation of respective / separate / different / independent specialist signal profilers for each subseries of sequencing cycles of a sequencing run 1300 with a total of N sequencing cycles. A first specialist signal profiler 1312 is located on subtile M and configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles 1 to N1. A second specialist signal profiler 1314 is located on subtile M and configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles N1+1 to N2. A third specialist signal profiler 1318 is located on subtile M and configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles N2+1 to N. Other examples of temporal configurations include a sequencing cycle of a first read of a sequencing run (read 1) and a sequencing cycle of a second read of a sequencing run (read 2).

[0086] 14 illustrates one implementation of respective / separate / different / independent specialist signal profilers for combinations of different spatial configurations (e.g., different subtiles) and different temporal configurations (e.g., different subseries of a sequencing cycle). In one implementation, cluster class 1410 is defined by different spatial configurations (e.g., different subtiles). In one implementation, cluster subclasses 1412, 1414, and 1416 are defined by different temporal configurations (e.g., different subseries of a sequencing cycle).

[0087] A first specialist signal profiler 1422 is located on subtile 518a and configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles 1 to N1. A second specialist signal profiler 1424 is located on subtile 518a and configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles N1+1 to N2. A third specialist signal profiler 1428 is located on subtile 518a and configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles N2+1 to N.

[0088] A fourth specialist signal profiler 1432 is located on subtile 518b and is configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles 1 to N1. A fifth specialist signal profiler 1434 is located on subtile 518b and is configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles N1+1 to N2. A sixth specialist signal profiler 1438 is located on subtile 518b and is configured to maximize the signal-to-noise ratio of the intensity data of clusters generated during sequencing cycles N2+1 to N.

[0089] Cluster / well specific specialist signal profiler FIG. 15 shows one implementation of a respective / separate / different / independent specialist signal profiler for each cluster / well sequenced during a sequencing run. The clusters / wells on the flow cell can be pre-identified by their position coordinates. These position coordinates of the clusters / wells can be used to train a specialist signal profiler per cluster / well during training on the intensity data per training cluster / well, and during inference on the intensity data per inference cluster / well, the trained specialist signal profiler per cluster can be applied based on the position coordinates of the cluster / well. In FIG. 15, the cluster / well population 1502 has N clusters / wells, and thus the specialist signal profiler 1508 comprises a specialist signal profiler per N clusters / wells, respectively. In other implementations, different per cluster / well signal profilers can also be trained for different temporal configurations, as described above.

[0090] Offline Training 16 illustrates one implementation of offline training of a specialist signal profiler on sequenced data from one or more completed / already performed sequencing runs, and application of the trained specialist signal profiler on sequenced data from an ongoing sequencing run. Training data 1612 is generated during the training stage 1602. The training data 1612 includes sequenced data from one or more completed / already performed sequencing runs.

[0091] The segmentation logic 1622 segments the training data 1612 based on one or more configurations selected from different spatial configurations, temporal configurations, signal distribution configurations, or any combination thereof. The result is segmented training data 1632 having configuration-specific training data subsets 1 through N. For example, the training data 1612 may include K images of a tile from K imaging cycles of a completed / already performed sequencing run, with each tile image having multiple color channels. FIG. 17 shows an example of a tile image having C color channels. In this case, the segmentation logic 1622 logically partitions each tile image in the training data 1612 into subtile images by specifying pixel ranges. For example, the first subtile image of a tile may range from 1 to 500 pixels, the second subtile image of a tile may range from 501 to 1000 pixels, and so on. The pixel ranges may be defined using fiducial markers, as described above.

[0092] The offline training logic 1642 trains a respective / separate / different / independent specialist signal profiler 1-N for each configuration-specific training data subset 1-N. The result is trained specialist signal profilers 1-N. Returning to the example of FIG. 17, the corresponding specialist signal profiler is trained to maximize the signal-to-noise ratio of each subtile image in the training data 1612.

[0093] Inference data 1618 is generated during the inference stage 1608. The inference data 1618 includes sequenced data from an ongoing sequencing run (e.g., the first i cycles of the ongoing sequencing run).

[0094] The segmentation logic 1622 segments the inferred data 1618 based on the same configuration or configurations used to segment the training data 1612 during the training stage 1602. The result is segmented inferred data 1638 having configuration-specific inferred data subsets 1 through N. For example, the inferred data 1618 may include K images of tiles from K imaging cycles of an ongoing sequencing run, with each tile image having multiple color channels. Returning to the example of FIG. 17, the segmentation logic 1622 logically partitions each tile image in the inferred data 1618 into subtile images by specifying the same pixel ranges used to partition the training data 1612.

[0095] The runtime logic 1648 applies a respective trained specialist signal profiler 1-N 1658 to each configuration-specific inferred data subset 1-N. Returning to the example of FIG. 17 , a corresponding trained specialist signal profiler is applied to maximize the signal-to-noise ratio of each subtile image in the segmented inferred data 1638.

[0096] Online Training 18 illustrates one implementation of online training of a specialist signal profiler on sequenced data from early sequencing cycles of an ongoing sequencing run, and application of the trained specialist signal profiler on sequenced data from later sequencing cycles of the ongoing sequencing run. Inference data 1802 is generated during the inference stage 1802. Inference data 1812 includes sequenced data of early sequencing cycles (e.g., from cycles 1 through N1) of the ongoing sequencing run.

[0097] The segmentation logic 1622 segments the inferred data 1812 based on one or more configurations selected from different spatial configurations, temporal configurations, signal distribution configurations, or any combination thereof. The result is segmented inferred data 1832 having configuration-specific training data subsets 1 through N. For example, the inferred data 1812 may include N1 images of a tile from N1 initial sequencing cycles of an ongoing sequencing run, with each tile image having multiple color channels. FIG. 17 illustrates an example of a tile image having C color channels. In this case, the segmentation logic 1622 logically partitions each tile image in the inferred data 1812 into subtile images by specifying pixel ranges. For example, a first subtile image of a tile may range from pixels 1 through 500, a second subtile image of a tile may range from pixels 501 through 1000, and so on. The pixel ranges may be defined using fiducial markers, as described above.

[0098] The offline training logic 1842 trains a respective / separate / different / independent specialist signal profiler 1-N for each configuration-specific training data subset 1-N. The result is trained specialist signal profilers 1-N. Returning to the example of FIG. 17, the corresponding specialist signal profiler is trained to maximize the signal-to-noise ratio of each subtile image in the inferred data 1812.

[0099] Inference data 1818 is also generated during the inference stage 1802. The inference data 1818 includes sequenced data from later sequencing cycles (e.g., from cycles N1+1 to N2) of an ongoing sequencing run.

[0100] The segmentation logic 1622 segments the inference data 1818 based on the same one or more configurations used to segment the inference data 1812 during the initial sequencing cycles (e.g., cycles 1-N1) of the ongoing sequencing run. The result is segmented inference data 1838 having configuration-specific inference data subsets 1-N. For example, the inference data 1818 may include N2 images of tiles from N2 later sequencing cycles of the ongoing sequencing run, with each tile image having multiple color channels. Returning to the example of FIG. 17, the segmentation logic 1622 logically partitions each tile image in the inference data 1818 into subtile images by specifying the same pixel ranges used to partition the inference data 1812.

[0101] The runtime logic 1648 applies a respective trained specialist signal profiler 1-N 1858 to each configuration-specific inferred data subset 1-N. Returning to the example of FIG. 17 , a corresponding trained specialist signal profiler is applied to maximize the signal-to-noise ratio of each subtile image in the segmented inferred data 1838.

[0102] In some implementations, the training process is repeated iteratively, with each trained specialist signal profiler 1-N 1858 being retrained / further trained on segmented inference data 1838 from a later sequencing cycle of the ongoing sequencing run (e.g., cycles N1+1-N2) and applied to segmented inference data from an even later sequencing cycle of the ongoing sequencing run (e.g., cycles N2+1-N3).

[0103] The control logic (not shown) may repeat, in each successive sequencing cycle of an ongoing sequencing run, or in a successive subseries of sequencing cycles of an ongoing sequencing run (e.g., every 10 to 20 sequencing cycles), (i) segmenting the current batch of image data based on one or more configurations selected from different spatial configurations, temporal configurations, signal distribution configurations, or any combination thereof, (ii) retraining each of the trained specialist signal profilers 1 to N on the segmented current batch of image data, (iii) segmenting the next batch of image data with the same criteria as the current batch of image data, and (iv) applying each of the retrained specialist signal profilers 1 to N to the segmented next batch of image data.

[0104] Specialist signal profiler specific to your signal distribution configuration Figure 19 shows an implementation of training each / separate / different / independent specialist signal profiler for each signal distribution observed in sequenced data. In some implementations, each signal distribution can be observed in sequenced data from offline / already performed sequencing run. In other embodiments, each signal distribution can also be observed in sequenced data from online / ongoing sequencing run (e.g., observed in the first 10 sequencing cycles of an ongoing sequencing run).

[0105] FIG. 20 shows an example of a signal distribution / signal profile / cluster intensity profile. The cluster intensity profile depicted in FIG. 20 follows an attenuation pattern. In this case, the cluster signal is strongest at the cluster center and attenuates as it propagates away from the cluster center. Subpopulations / groups / sets of clusters 1904 in a cluster population may have similar signal distributions. Clusters that share the same or similar signal distributions may be bucketed together (e.g., by grouping / addressing the clusters by their location coordinates) so that a specialist signal profiler may be trained on the corresponding cluster group / set that exhibits the corresponding signal distribution. Unlike spatial grouping, where spatially adjacent clusters are grouped, signal distribution-based grouping may group non-adjacent clusters. For example, edge clusters on opposing edges of a tile may have similar signal distributions and may be grouped together (e.g., by grouping / addressing the clusters by their location coordinates) whose intensity data are corrected by the same specialist signal profiler.

[0106] As used herein, the phrase "similar signal distributions" refers to signal distributions that share substantially overlapping signal patterns. For example, two signal patterns of similar shape (e.g., trapezoidal) but different shape sizes (e.g., one larger trapezoid and one smaller trapezoid) may be considered to have similar signal distributions. Similarly, two signal patterns that have a common center of gravity within, for example, 1 to 5 units in each dimension may be considered to have similar signal distributions.

[0107] 19 , each / separate / different / independent specialist signal profiler 1-N 1908 is trained to maximize the signal to noise ratio of a respective signal distribution 1-N 1902 corresponding to a respective cluster set 1-N 1904. Of course, different signal distributions may be observed in different sequencing cycles and therefore different specialist signal profilers may be trained and configured for use at different time stages of an ongoing sequencing run.

[0108] As used herein, the phrase "different time stages of an ongoing sequencing run" refers to different sequencing cycles or different sequencing cycle groups of an ongoing sequencing run. For example, if a sequencing run has 150 sequencing cycles, each successive sequencing cycle can be considered as a different time stage, or a group of sequencing cycles such as cycles 1-20, cycles 20-40, cycles 40-70, etc. can be considered as a different time stage.

[0109] Processing Pipeline 21 shows one implementation of a processing pipeline that implements the disclosed technology. The processing pipeline may be implemented by a real-time analysis module 225. According to one implementation, the processing pipeline is executed 2100 every cycle and repeated 2102 every new cycle. In one implementation, the input to the processing pipeline is a tiled image having a first (green) channel and a second (blue) channel.

[0110] In operation 2113, a template image is generated that identifies the location of the clusters on the tile using sequence images from some initial sequence cycles, called template cycles. The template image is used as a reference for the subsequent alignment and intensity extraction steps. The template image is generated by detecting and merging bright points in each sequence image of the template cycle, which includes sharpening the sequence image 2114 (e.g., using Laplacian convolution), determining the "on" threshold by the spatially separated Otsu method, and then detecting 5-pixel local maxima using sub-pixel position interpolation. The phrase "above threshold" can refer to intensity values ​​that exceed a preset value, e.g., 200 or 320, such that the intensity value is detected as being greater than the background intensity value or noise intensity value.

[0111] The processing pipeline then aligns the current sequencing image to the template image, which is achieved by using image correlation to align the current sequencing image to the template image on the sub-region, or by using a nonlinear transformation (e.g., a full six-parameter linear affine transformation).

[0112] In operation 2115, the processing pipeline applies a nonlinear distortion to each spot, for example to account for optical distortion caused by the geometry of the optical lens. The nonlinear distortion may be applied as channel-dependent coefficients of a third order polynomial.

[0113] In operation 2116, the processing pipeline segments the tile image into subtile images based on one or more configurations selected from different spatial configurations, temporal configurations, signal distribution configurations, or any combination thereof.

[0114] In operation 2118, intensities are extracted from the segmented subtile images using corresponding specialist signal profilers 1-N.

[0115] In operation 2123, the subtile intensities are spatially normalized, for example, by making the 90th percentile of their extracted intensities equal.

[0116] In operation 2124, the subtile intensities are compressed.

[0117] In operation 2125, the processing pipeline applies an empirical phase correction to compensate for noise in the image data caused by the phase error and the a priori phase error.

[0118] In operation 2125, the processing pipeline spatially normalizes the extracted signal intensities to account for variations in illumination across the sampled image. For example, the intensity values ​​may be normalized such that the 5th and 95th percentiles have values ​​of 0 and 1, respectively. The normalized signal intensities for the image (e.g., the normalized intensities for each channel) may be used to calculate an average purity for multiple spots in the image.

[0119] In operation 2133, the processing pipeline scales the intensity for each cluster to account for variations in brightness of the cluster.

[0120] In operation 2134, the processing pipeline generates base calls using the expectation-maximization (EM) algorithm, as described above.

[0121] In operation 2135, the processing pipeline assigns quality scores to the called bases using a quality table (Q-table) 2152.

[0122] In operation 2136, the processing pipeline aligns the called bases to a reference genome (eg, the PhiX bacterial reference genome) and calculates the mismatch rate.

[0123] The processing pipeline generates specific outputs such as base calls and quality scores 2128, InterOp files 2138 (binary report files for the sequencing analysis view), and logs 2148 (e.g., error log, general event log, processing event log, warning event log).

[0124] Other configurations Other examples of configurations encompassed by the present disclosure include segmenting sequencing data and training corresponding specialist signal profilers by library type, sample type, indexing type (first index read vs. second index read), read type (forward read vs. reverse read), physical characteristics of the sample, noise type (e.g., bubbles), and reagent type.

[0125] Performance Results: Technical Effect and Advantage as Objective Evidence of Nonobviousness and Inventive Step FIG. 33 shows how the cost function of the specialist signal profiler improves with each iteration of gradient descent. The cost function (or loss function) measures the performance of the model for the given data. In FIG. 33, the lines of different colors correspond to several subtiles into which the tile is partitioned, so that for each subtile, the specialist signal profiler is adapted / trained / configured / updated independently. As the number of subtiles increases from 1 to 16 (4×4 subtiles), the number of fitting parameters in the specialist signal profiler also increases. As a result, the cost function becomes lower. The cost function is the sum of squared Euclidean distances between the well intensities and the base call centroids for the samples of a well (cluster). Improving this cost function also improves the base call accuracy.

[0126] Figure 34 is a plot showing the initial and final values ​​of the cost function of Figure 33 as the specialist signal profiler is adapted / trained / configured / updated at each sequencing cycle. At each successive sequencing cycle, we started with the same initial specialist signal profiler and adapted the specialist signal profiler using gradient descent. The plot shows that the specialist signal profiler can be adapted / trained / configured / updated at any sequencing cycle.

[0127] Figures 35A and 35B show the improvement in primary analysis metrics for a sequencing run when fitting / training / configuring / updating a specialist signal profiler. In this case, the average PhiX error rate improved from 0.3520% to 0.3316%.

[0128] Figures 36A and 36B show two plots evaluating the number of subtiles into which a sequencing tile can be divided for adaptive equalization of the respective specialist signal profiler. For each subtile, a separate specialist signal profiler is adapted / trained / configured / updated using wells from the corresponding subtile. Using more subtiles allows for more accurate modeling of spatially varying phenomena within the tile. This is why the error rate and Q30 improve significantly as the number of subtiles increases from 1 to 9. However, the number of wells available to adapt the model decreases as the subtiles become smaller. This is why there is a trade-off in choosing the right number of subtiles. In this particular case, increasing the number of subtiles from 9 to 16 slightly degrades the error rate and Q30.

[0129] Computer Systems 37 illustrates an exemplary computer system 3700 that can be used to implement the disclosed techniques. The computer system 3700 includes at least one central processing unit (CPU) 3772 that communicates with a number of peripheral devices via a bus subsystem 3755. These peripheral devices can include, for example, a storage subsystem 3710 including memory devices and a file storage subsystem 3736, a user interface input device 3738, a user interface output device 3776, and a network interface subsystem 3774. The input and output devices enable user interaction with the computer system 3700. The network interface subsystem 3774 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.

[0130] In one embodiment, the specialist signal profiler 3718 is communicatively linked to the storage subsystem 3710 and the user interface input device 3738 .

[0131] The user interface input devices 3738 can include pointing devices such as a keyboard, a mouse, a trackball, a touch pad, or a graphics tablet, a scanner, a touch screen integrated into a display, audio input devices such as a voice recognition system and a microphone, as well as other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and manners for inputting information into the computer system 3700.

[0132] The user interface output devices 3776 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual displays such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and manners for outputting information from the computer system 3700 to a user or to another machine or computer system.

[0133] The storage subsystem 3710 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 3778.

[0134] The processor 3778 can be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 3778 can be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processor 3778 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX37 Rackmount Series™, NVIDIA DGX-1™, Microsoft' Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa VI 00s™, and others.

[0135] The memory subsystem 3722 used in the storage subsystem 3710 may include multiple memories including a main random access memory (RAM) 3732 for storing instructions and data during program execution, and a read only memory (ROM) 3734 in which fixed instructions are stored. The file storage subsystem 3736 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular embodiment may be stored by the file storage subsystem 3736 in the storage subsystem 3710 or in another machine accessible by the processor.

[0136] The bus subsystem 3755 provides a mechanism for allowing the various components and subsystems of the computer system 3700 to communicate with each other as intended. Although the bus subsystem 3755 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0137] The computer system 3700 itself can be of a variety of types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 3700 shown in Figure 37 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 3700 can have more or fewer components than the computer system shown in Figure 37.

[0138] Neural network based base cola The following discussion focuses on the neural network-based base caller described herein, which may be used in conjunction with a specialist signal profiler. First, the input to the neural network-based base caller is described according to one embodiment. Then, an example of the structure and form of the neural network-based base caller is provided. Finally, the output of the neural network-based base caller is described according to one embodiment.

[0139] The data flow logic provides the sequence image to a neural network based base caller for base calling. The neural network based base caller accesses the sequence image on a patch-by-patch (or tile-by-tile) basis. Each patch is a subgrid (or subarray) of pixelated units within the grid of pixelated units that forms the sequence image. A patch has dimensions q by r of the subgrid of pixelated units, where q (width) and r (height) are any number in the range of 1 to 10,000 (e.g., 3×3, 5×5, 7×7, 10×10, 15×15, 25×25, 64×64, 78×78, 115×115). In some implementations, q and r are the same. In other implementations, q and r are different from each other. In some implementations, patches extracted from one sequence image are the same size. In other implementations, the patches are of different sizes. In some implementations, patches can have overlapping pixelated units (e.g., on edges).

[0140] Sequencing produces m sequence images for each sequencing cycle for the corresponding m image channels. That is, each sequence image has one or more image (or intensity) channels (similar to the red, green, and blue (RGB) channels of a color image). In one embodiment, each image channel corresponds to one of a number of filter wavelength bands. In another embodiment, each image channel corresponds to one of a number of imaging events in a sequencing cycle. In yet another embodiment, each image channel corresponds to a combination of illumination by a particular laser and imaging through a particular optical filter. Image patches are tiled (or accessed) from each of the m image channels for a particular sequencing cycle. In different embodiments, such as 4-channel chemistry, 2-channel chemistry, and 1-channel chemistry, m is 4 or 2. In other embodiments, m is greater than 1, 3, or 4. In other implementations, the images can be blue and violet channels instead of or in addition to red and green channels.

[0141] For example, consider that a sequencing run is performed using two different imaging channels, namely, a blue channel and a green channel. Then, in each sequencing cycle, the sequencing run generates a blue image and a green image. In this way, for a series of k sequencing cycles of the sequencing run, a sequence of k pairs of blue and green images is generated as output and stored as sequence image. Thus, a sequence of k pairs of blue and green image patches is generated for patch-level processing by a neural network-based base caller.

[0142] The input image data to the neural network-based base caller for one base calling iteration (or one instance of a forward pass or single forward traversal) includes data for one sliding window that includes multiple sequencing cycles. The sliding window can include, for example, the current sequencing cycle, one or more preceding sequencing cycles, and one or more subsequent sequencing cycles.

[0143] In one embodiment, the input image data includes data from three sequencing cycles such that data from a current (time t) sequencing cycle being base called is accompanied by data from (i) the left adjacent / context / previous / preceding / prior (time t-1) sequencing cycle, and (ii) the right adjacent / context / next / successive / successive (time t+1) sequencing cycle.

[0144] In another embodiment, the input image data includes data from five sequencing cycles, and the data of the current (time t) sequencing cycle being base called is accompanied by (i) data from the first left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, (ii) data from the second left adjacent / context / previous / preceding / previous (time t-2) sequencing cycle, (iii) data from the first right adjacent / context / next / contiguous / subsequent (time t+1) sequencing cycle, and (iv) data from the second right adjacent / context / next / contiguous / subsequent (time t+2) sequencing cycle.

[0145] In yet another embodiment, the input image data includes data from seven sequencing cycles, and the data for the current (time t) sequencing cycle being base called includes (i) data from the first left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, (ii) data from the second left adjacent / context / previous / preceding / previous (time t-2) sequencing cycle, (iii) data from the third left adjacent / context / previous / preceding / previous (time t-3) sequencing cycle, (iv) data from the first right adjacent / context / next / successive / subsequent (time t+1) sequencing cycle, (v) data from the second right adjacent / context / next / successive / subsequent (time t+2) sequencing cycle, and (vi) data from the third right adjacent / context / next / successive / subsequent (time t+3) sequencing cycle. In other embodiments, the input image data includes data from a single sequencing cycle. In yet other embodiments, the input image data includes data for 10, 15, 20, 30, 58, 75, 92, 130, 168, 175, 209, 225, 230, 275, 318, 325, 330, 525, or 625 sequencing cycles.

[0146] According to one embodiment, the neural network-based base caller processes image patches through its convolutional layers to generate alternative representations. The alternative representations are then used by an output layer (e.g., a softmax layer) to generate base calls for either the current sequencing cycle (time t) or each of the sequencing cycles (i.e., the current sequencing cycle (time t), the first and second preceding sequencing cycles (time t-1, time t-2), and the first and second subsequent sequencing cycles (time t+1, time t+2)). The resulting base calls form a sequencing read.

[0147] In one embodiment, the neural network-based base caller outputs a base call for a single target cluster for a particular sequencing cycle. In another embodiment, the neural network-based base caller outputs a base call for each target cluster in the plurality of target clusters at a particular sequencing cycle. In yet another embodiment, the neural network-based base caller outputs a base call for each target cluster in the plurality of target clusters at each sequencing cycle in the plurality of sequencing cycles, thereby generating a base call sequence for each target cluster.

[0148] In one embodiment, the neural network-based base caller is a Multilayer Perceptron (MLP). In another embodiment, the neural network-based base caller is a feed-forward neural network. In yet another embodiment, the neural network-based base caller is a fully connected neural network. In a further embodiment, the neural network-based base caller is a fully convolutional neural network. In yet another embodiment, the neural network-based base caller is a semantic segmentation neural network. In yet another further embodiment, the neural network-based base caller is a generative adversarial network (GAN).

[0149] In one embodiment, the neural network-based base caller is a convolutional neural network (CNN) having multiple convolutional layers. In another embodiment, the neural network-based base caller is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bi-directional LSTM (Bi-LSTM), or a gated recurrent unit (GRU). In yet another embodiment, the neural network-based base caller includes both CNNs and RNNs.

[0150] In yet other implementations, the neural network-based base caller can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, deep-separable convolution, point-wise convolution, 1×1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The neural network-based base caller can use one or more loss functions, such as logistic regression / logarithmic loss, multiclass cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. Neural network-based base callers can use any parallel, efficient, and compression schemes, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallel, data parallel, and synchronous / asynchronous stochastic gradient descent (SDG). Neural network-based base callers can include nonlinear transformation functions, such as upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions (such as rectifying linear unit (ReLU), leaky ReLU, exponential liner unit (ELU), sigmoid, and hyperbolic tangent (tanh))), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or mean pooling), global average pooling layers, and attention mechanisms.

[0151] The neural network-based base caller trains using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that the neural network-based base caller may use to train include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that the neural network-based base caller may use to train include Momentum, Nestorv accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.

[0152] In one embodiment, the neural network-based base caller uses a dedicated architecture to separate the processing of data for different sequencing cycles. The motivation for using such a dedicated architecture is first described. As described above, the neural network-based base caller processes image patches for the current sequencing cycle, one or more previous sequencing cycles, and one or more subsequent sequencing cycles. Data from additional sequencing cycles provide sequence-specific context. The neural network-based base caller learns sequence-specific contexts during training and base calls them. In addition, data from pre- and post-sequencing cycles provide secondary contributions of pre-phasing and phasing signals to the current sequencing cycle.

[0153] However, images captured in different sequencing cycles and in different image channels are misaligned and have residual registration errors with each other. To account for this misalignment, the dedicated architecture includes a spatial convolution layer that does not mix information between sequencing cycles, but only mixes information within the same sequencing cycle.

[0154] The spatial convolutional layer (or spatial logic) uses so-called "separate convolutions" that operate on separation by processing the data for each of multiple sequencing cycles independently through a "dedicated, non-shared" sequence of convolutions. Separate convolutions convolve on the data and resulting feature maps only within a given sequencing cycle, i.e., the cycle, without convolving on the data and resulting feature maps of any other sequencing cycles.

[0155] For example, consider that the input image data includes (i) a current image patch for the current (time t) sequencing cycle to be base called, (ii) a previous image patch for the previous (time t-1) sequencing cycle, and (iii) a next image patch for the next (time t+1) sequencing cycle. The dedicated architecture then starts three separate convolution pipelines, namely, a current convolution pipeline, a previous convolution pipeline, and a next convolution pipeline. The current data processing pipeline receives the current image patch for the current (time t) sequencing cycle as input and processes it independently through multiple spatial convolution layers to generate a so-called "current spatial convolution representation" as the output of the final spatial convolution layer. The previous convolution pipeline receives the previous image patch for the previous (time t-1) sequencing cycle as input and processes it independently through multiple spatial convolution layers to generate a so-called "previous spatial convolution representation" as the output of the final spatial convolution layer. The next convolution pipeline receives the next data for the next (time t+1) sequencing cycle as input and processes it independently through multiple spatial convolution layers to produce a so-called “next spatial convolutional representation” as the output of the final spatial convolution layer.

[0156] In some implementations, the current, previous, and next convolution pipelines run in parallel. In some implementations, the spatial convolution layer is part of a spatial convolution network (or sub-network) in a dedicated architecture.

[0157] The neural network-based base caller further includes a temporal convolutional layer (or temporal logic) that blends information between sequencing cycles, i.e., between cycles. The temporal convolutional layers receive their input from the spatial convolutional network and operate on the spatial convolutional representations produced by the final spatial convolutional layer for each data processing pipeline.

[0158] The inter-cycle steerability of the temporal convolutional layers arises from the fact that misalignment features present in the image data provided as input to the spatial convolutional network are purged from the spatial convolutional representation by the stack or cascade of separated convolutions performed by the sequence of spatial convolutional layers.

[0159] The temporal convolutional layer uses so-called "combinatorial convolution" that convolves group-wise on the input channels with the subsequent input on a sliding window basis. In one implementation, the subsequent input is the subsequent output generated by the previous spatial convolutional layer or the previous temporal convolutional layer.

[0160] In some embodiments, the temporal convolutional layer is part of a temporal convolutional network (or sub-network) in the dedicated architecture. The temporal convolutional network receives its input from a spatial convolutional network. In one embodiment, the first temporal convolutional layer of the temporal convolutional network combines the spatial convolutional representations between sequencing cycles by group. In another embodiment, subsequent temporal convolutional layers of the temporal convolutional network combine subsequent outputs of previous temporal convolutional layers. The output of the final temporal convolutional layer is fed to an output layer that generates an output. The output is used to base call one or more clusters in one or more sequencing cycles.

[0161] Further details regarding neural network-based base collation can be found in U.S. Provisional Patent Application No. 62 / 821,766, entitled "ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV), which is incorporated herein by reference.

[0162] Terms The following items are part of this disclosure.

[0163] Clause 1. A system comprising: a memory storing a plurality of specialist signal profilers, each specialist signal profiler in the plurality of specialist signal profilers being trained to maximize a signal to noise ratio of sequence signals in a particular signal profile detected for analytes in a particular analyte class and characterized in a particular training data set; and runtime logic having access to a memory and configured to perform a base calling operation by applying a respective specialist signal profiler in a plurality of specialist signal profilers to sequence signals in a respective signal profile detected for analytes in a respective analyte class during a base calling operation. Clause 2. The system of clause 1, wherein each analyte class represents a distinct spatial arrangement of analytes that contribute to the generation of a respective signal profile during a base calling operation. Clause 3. The system of clause 2, wherein the different spatial configurations include analytes that are disposed on different surfaces of a biosensor on which base calling operations are performed. Clause 4. The system of clause 3, wherein the distinct surfaces include a top surface and a bottom surface. Clause 5. The system of any of clauses 2-4, wherein the different spatial configurations include analytes disposed on different lanes of the biosensor. Clause 6. The system of any of clauses 2-5, wherein the different spatial configurations include analytes disposed on different lane groups of the biosensor. Clause 7. The system of clause 6, wherein the distinct lane groups include top perimeter lanes, center lanes, and bottom perimeter lanes. Clause 8. The system of clause 6 or 7, wherein the different lane groups include edge lanes and non-edge lanes. Clause 9. The system of any of clauses 2-8, wherein the different spatial configurations include the analytes being disposed on different swatches in different lanes of the biosensor. Clause 10. The system of clause 9, wherein the different swatches include a top peripheral swatch, a central swatch, and a bottom peripheral swatch. Clause 11. The system of clause 9 or 10, wherein the different swatches include an edge swatch and a center swatch. Clause 12. The system of any of clauses 9-11, wherein the different spatial configurations include the analytes being disposed on different tiles of different swatches in different lanes of the biosensor. Clause 13. The system of any of clauses 2-12, wherein the different spatial configurations include analytes disposed on different tile groups of the biosensor. Clause 14. The system of clause 13, wherein the different tile groups include edge tiles, central tiles, and near-edge tiles. Clause 15. The system of any of clauses 12-14, wherein the different spatial configurations include analytes being disposed on different subtiles of different tiles of different swatches of different lanes of the biosensor. Clause 16. The system of any of clauses 2-15, wherein the different spatial configurations include analytes disposed on different sections of the biosensor. Clause 17. The system of clause 16, wherein the different sections include an upper right section, an upper center section, an upper left section, a center right section, a center section, a center left section, a lower left section, a lower center section, and a lower left section. Clause 18: Each specialist signal profiler is further trained to maximize the signal-to-noise ratio of sequence signals in a particular signal profile detected for analytes of a particular analyte subclass and characterized in a particular training data subset; The system of any of clauses 1-17, wherein the runtime logic is further configured to perform a base calling operation by applying a respective specialist signal profiler to sequence signals in the respective signal profiles detected for analytes of the respective analyte subclasses during the base calling operation. Clause 19. The system of clause 18, wherein each sample subclass represents a different spatial arrangement of samples that generated sequence signals at different periods of the base calling operation, and different combinations of the different spatial arrangements and different periods contribute to the creation of each signal profile detected during the base calling operation. Clause 20. The system of clause 19, wherein the different time periods correspond to different sensing cycles in a series of sensing cycles of the base calling operation. Clause 21. The system of clause 19 or 20, wherein the different periods correspond to different subseries of sensing cycles in a series of sensing cycles of the base calling operation. Clause 22. The system of any one of clauses 1 to 21, wherein each specialist signal profiler is configured with a channel-specific equalizer, each channel-specific equalizer having a plurality of convolution kernels. Clause 23. The system of any of clauses 1-22, wherein the runtime logic is further configured to iteratively train each specialist signal profiler during the base calling operation. Clause 24. The system of clause 23, wherein for the current training iteration, the runtime logic is further configured to: perform expectation maximization to iteratively maximize the likelihood of observing, for each channel, a signal centroid and signal distribution per base that best fits the sequence signal detected so far during the base calling operation; determine a signal-to-noise ratio maximizing sequence signal for each channel in response to applying a respective specialist signal profiler to the sequence signal; make base calls based on the signal-to-noise ratio maximizing sequence signal; determine a base calling error for each channel based on comparing the signal-to-noise ratio maximizing sequence signal to the signal centroid of the called base; and update coefficients of a convolution kernel of each specialist signal profiler for each channel based on the base calling error. Clause 25. The system of any one of clauses 1-24, wherein when the biosensor is a patterned biosensor, the analyte corresponds to a well. Clause 26. A system according to any one of clauses 1 to 25, wherein the sequence signal is an intensity signal. Clause 27. A system according to any one of clauses 1 to 25, wherein the sequence signal is a voltage signal. Clause 28. A system according to any one of clauses 1 to 25, wherein the sequence signal is a current signal. Clause 29 A system comprising: a memory for storing an initial sequence signal detected during an initial sequencing cycle of a sequencing run; fitting logic having access to the memory and configured to initially fit a plurality of signal distributions to the sequence signal and store the plurality of signal distributions in the memory; online training logic having access to the memory and configured to train each specialist signal profiler in the plurality of specialist signal profilers to maximize a signal-to-noise ratio of each signal distribution in the plurality of signal distributions and to store the trained each specialist signal profiler in the memory; and runtime logic having access to the memory and configured to uniquely map subsequent sequence signals detected during subsequent sequencing cycles of the sequencing run to respective signal distributions, and apply respective specialist signal profilers trained based on the unique mappings to the respective signal distributions to the subsequent sequence signals to generate base calls for the subsequent sequencing cycles. Clause 30. The system of clause 29, wherein at least some of the signal distributions in the plurality of signal distributions represent different underlying sequencing events that contribute to the generation of some of the signal distributions. Clause 31. The system of clause 30, wherein the underlying sequencing event comprises the formation of an air bubble on a biosensor on which the sequencing run is performed. Clause 32. The system of any of clauses 29-31, wherein at least some of the signal distributions in the plurality of signal distributions represent different analyte locations on the biosensor that contribute to generating some of the signal distributions. Clause 33. A system according to any of clauses 29 to 32, wherein at least some of the signal distributions in the plurality of signal distributions represent different sequencing cycles of a sequencing run that contribute to the generation of some of the signal distributions. Clause 34. A system according to any of clauses 29 to 33, wherein each trained specialist signal profiler has a respective set of coefficients of variation corresponding to a respective signal distribution. Clause 35. The system of clause 34, wherein each set of variation coefficients includes a channel-specific amplification factor that compensates for scale variations in a corresponding signal distribution. Clause 36. The system of clause 34 or 35, wherein each set of variation coefficients includes a channel-specific offset coefficient that compensates for shift variations in the corresponding signal distribution. Clause 37. The system of any of clauses 29-36, further configured to comprise control logic configured to repeat execution of the adaptation logic, the online training logic, and the runtime logic in each sequencing cycle of a sequencing run. Clause 38. The system of any of clauses 29-37, further configured to comprise control logic configured to repeat execution of the adaptation logic, the online training logic, and the runtime logic after a certain number of sequencing cycles of a sequencing run. Clause 39. The system of any of clauses 29-38, further configured to comprise control logic configured to repeat execution of the adaptation logic, the online training logic, and the runtime logic in a particular predetermined sequencing cycle of a sequencing run. Clause 40 A system comprising: a memory storing a plurality of specialist signal profilers configured for use in a base calling operation, each specialist signal profiler in the plurality of specialist signal profilers being trained to maximize a signal to noise ratio of sensor data within a respective signal distribution observed in a respective sequencing event of the base calling operation and characterized in a respective training data set; and runtime logic having access to a memory and configured to select a specialist signal profiler from a plurality of specialist signal profilers based on a subject sequencing event that generated the subject sensor data, and apply the selected specialist signal profiler to the subject sensor data to generate base call classification data for the subject sequencing event. Clause 41. The system of clause 40, wherein the signal-to-noise ratio of the sensor data generated by each sequencing event deteriorates with the temporal progression of the base calling operation, and each specialist signal profiler selected based on each sequencing event and applied throughout the temporal progression of the base calling operation reverses the deterioration of the signal-to-noise ratio of the sensor data. Clause 42. The system of clause 41, wherein each sequencing event is a temporal progression of a base calling operation through each sensing cycle in a series of sensing cycles of a base calling operation. Clause 43. The system of clause 42, wherein each sequencing event is a temporal progression of a base calling operation through a respective sub-series of sensing cycles in a series of sensing cycles. Clause 44. The system of clause 41, wherein each sequencing event is a spatial progression of a base calling operation through each sample location on the biosensor at which the base calling operation is performed. Clause 45. The system of clause 40, wherein the runtime logic is further configured to iteratively train each specialist signal profiler during the sequencing run at each sequencing event. Clause 46. The system of clause 45, wherein for the current training iteration, the runtime logic is further configured to: perform expectation maximization to iteratively maximize the likelihood of observing per-base signal centroids and signal distributions for each channel that best fit the sensor data; determine signal-to-noise ratio maximized sensor data for each channel in response to applying a respective specialist signal profiler to the sensor data; make base calls based on the signal-to-noise ratio maximized sensor data; determine base calling errors for each channel based on comparing the signal-to-noise ratio maximized sensor data to the signal centroids of the called bases; and update coefficients of a convolution kernel of each specialist signal profiler for each channel based on the base calling errors. Clause 47 A system comprising: a memory that stores an initial sequence signal detected during an initial sequencing cycle of a sequencing run on the population of analytes; fitting logic having access to the memory and configured to process the initial sequence signals for each analyte, fit a respective signal profile for each analyte in the population of analytes, and store the respective signal profile in the memory; online training logic having access to a memory and configured to train each specialist signal profiler of the plurality of specialist signal profilers to maximize a signal-to-noise ratio of a respective signal profile matched to a respective analyte and store the respective trained specialist signal profiler in the memory; The system comprises: runtime logic having access to the memory and configured to uniquely map subsequent sequence signals detected during subsequent sequencing cycles of the sequencing run to respective signal profiles for each analyte, and apply respective trained specialist signal profilers to the subsequent sequence signals based on the unique mapping to the respective signal profiles to generate base calls for the subsequent sequencing cycles for each analyte. Clause 48. The system of clause 47, wherein each trained specialist signal profiler has a respective set of coefficients of variation corresponding to a respective signal profile. Clause 49. The system of clause 48, wherein each set of variation coefficients includes a channel-specific amplification factor that compensates for scale variations in a corresponding signal profile. Clause 50. The system of clause 48, wherein each set of variation coefficients includes a channel-specific offset coefficient that compensates for shift variations in a corresponding signal profile. Clause 51. The system of clause 47, further configured to comprise control logic configured to repeat execution of the adaptation logic, the online training logic, and the runtime logic in each sequencing cycle of a sequencing run. Clause 52. The system of clause 51, further configured to comprise control logic configured to repeat execution of the adaptation logic, the online training logic, and the runtime logic after a particular number of sequencing cycles of a sequencing run. Clause 53. The system of clause 51, further configured to comprise control logic configured to repeat execution of the adaptation logic, the online training logic, and the runtime logic in a particular predetermined sequencing cycle of a sequencing run. Clause 54 A system comprising: a memory storing a plurality of specialist signal profilers, each specialist signal profiler in the plurality of specialist signal profilers being trained to maximize a signal to noise ratio of sequence signals in a particular signal profile detected for a particular analyte and characterized in a particular training data set; and runtime logic having access to a memory and configured to perform a base calling operation by applying each specialist signal profiler in a plurality of specialist signal profilers to sequence signals in each signal profile detected for each analyte during a base calling operation. Clause 55 A system comprising: spatial classification logic configured to segment the population of specimens into spatial classes using specimen locations, each spatial class comprising a non-overlapping set of specimens selected from the population of specimens; signal profiling logic configured to estimate one or more signal profiles for each spatial class using the detected sequence signals for the population of analytes, each signal profile including a non-overlapping subset of analytes selected from the non-overlapping set of analytes; and online training logic configured to train at least one specialist signal profiler to maximize a signal-to-noise ratio of each estimated signal profile for each spatial class. and runtime logic configured to base call analytes in each non-overlapping subset of analytes using each trained specialist signal profiler. Clause 56. The system of clause 55, wherein the online training logic is further configured to limit the training data used to train a particular specialist signal profiler to sequence signals detected for a particular non-overlapping subset of analytes. Clause 57. The system of clause 55 or 56, further configured to merge adjacent non-overlapping subsets of analytes if a particular non-overlapping subset of analytes lacks suitable analytes for which suitable training data is available. Clause 58. The system of any of clauses 55-57, wherein the number of exemplars in each non-overlapping subset of exemplars is configurable to optimize training. Clause 59. The system of any of clauses 55-58, wherein the online training logic is further configured to estimate a set of coefficients of variation for each trained specialist signal profiler. Clause 60. The system of clause 59, wherein the set of variation coefficients includes channel-specific amplification factors that compensate for scale variations in the corresponding signal profiles. Clause 61. The system of clause 59 or 60, wherein the set of variation coefficients includes a channel-specific offset coefficient that compensates for shift variations in the corresponding signal profile. Clause 62. The system of any of clauses 55-61, further configured to comprise control logic configured to repeat execution of the signal profiling logic, the online training logic, and the runtime logic in each sequencing cycle of a sequencing run. Clause 63. The system of any of clauses 55-62, further configured to comprise control logic configured to repeat execution of the signal profiling logic, the online training logic, and the runtime logic after a particular number of sequencing cycles of a sequencing run. Clause 64. The system of any of clauses 55-63, further configured to comprise control logic configured to repeat execution of the signal profiling logic, the online training logic, and the runtime logic in a particular predetermined sequencing cycle of a sequencing run. [Explanation of symbols]

[0164] 100 Imaging System 110 Sample container 120 Waste Valve 130 Temperature Station Actuator 135 Cooler 140 Camera System 142 Objective Lens 145 Filter Switching Assembly 150 Focused Laser 160 light source 165 Low Watt Lamp 170 Sample Stage 175 Focus (z-axis) component 175 Focus component 185 Inverse Dichroic 200 Imaging System 200 Optical Component Imaging System 200 Channel Line Scan Modular Optical Imaging System 200 Systems 210 Line generation module (LGM) 211 Light source 212 Light source 213 Beam shaping lens(es) 214 Mirror 215 Semi-reflective mirror 216 Shutter Element 220 Camera Module (CAM) 221 Optical Sensor 225 Real-time Analysis Module 230 emission optics module (EOM) 231 Filter Elements 232 Tube Lens 233 Semi-reflective mirror 234 Semi-reflective mirror 235 Objective Lens 236 z-stage 240 Focus tracking module (FTM) 250 targets 251 Translucent cover plate 252 Liquid layer 300 Sequencing Runs 332 Base Cola 400 Flow Cell 402 Top surface 412 Bottom 500 pixels 501 pixels 502 Lane Group 502a 1st Lane Group 502b Second Lane Group 502c 3rd Lane Group 506 Sample 506a First sample 506b Second sample 508 Lane 508a 1st lane 508b 2nd lane 512 tiles 512a 1st tile 512b 2nd tile 518 subtiles 518a 1st subtile 518b 2nd subtile 518c 3rd subtile 518d 4th subtile 600 Sequencing Runs 602 Cluster Class 604 Specialist Signal Profiler 604a #1 Specialist Signal Profiler 604b 2nd Specialist Signal Profiler 632 Image Data Subset 634 Signal to Noise Ratio Maximized Version 638 Base Call 642 Image Data Subset 644 Signal to Noise Ratio Maximized Version 648 Base Call 702 Cluster Class 704 Specialist Signal Profiler 704a #1 Specialist Signal Profiler 704b 2nd Specialist Signal Profiler 704c 3rd Specialist Signal Profiler 712 Cluster Class 714 Specialist Signal Profiler 714a #1 Specialist Signal Profiler 714b 2nd Specialist Signal Profiler 722 Cluster Class 724 Specialist Signal Profiler 724a #1 Specialist Signal Profiler 724b 2nd Specialist Signal Profiler 732 Cluster Class 734 Specialist Signal Profiler 734a #1 Specialist Signal Profiler 734b 1st Specialist Signal Profiler 742 Cluster Class 744 Specialist Signal Profiler 744a #1 Specialist Signal Profiler 744b 2nd Specialist Signal Profiler 744c 3rd Specialist Signal Profiler 744d 4th Specialist Signal Profiler 800 Sequencing Runs 802 First image data subset 804 signal-to-noise ratio maximized version 808 base call 812 Second Image Data Subset 814 Signal to Noise Ratio Maximized Version 818 Base Call 822 Third Image Data Subset 824 signal-to-noise ratio maximized version 828 Base Call 900 Sequencing Runs 902 First image data subset 904 Maximized Signal to Noise Ratio Version 908 Base Call 912 Second Image Data Subset 914 Maximized Signal to Noise Ratio Version 918 Base Call 922 First Image Data Subset 924 signal-to-noise ratio maximized version 1000 Sequencing Runs 1002 First image data subset 1004 Maximized signal-to-noise ratio version 1008 Base Calls 1012 Second image data subset 1014 Maximized Signal to Noise Ratio Version 1018 Base Call 1024 Maximized Signal to Noise Ratio Version 1100 Sequencing Runs 1102 First image data subset 1104 Maximized signal-to-noise ratio version 1108 Base Call 1112 Second image data subset 1114 Maximized signal-to-noise ratio version 1118 Base Call 1122 Image Data Subset 1124 Maximized Signal to Noise Ratio Version 1200 Sequencing Runs 1202 First image data subset 1204 Maximized Signal to Noise Ratio Version 1208 Base Call 1212 Second image data subset 1214 Maximized Signal to Noise Ratio Version 1218 Base Call 1222 Third image data subset 1224 Maximized Signal to Noise Ratio Version 1228 Base Call 1232 Fourth Image Data Subset 1234 Maximized Signal to Noise Ratio Version 1238 Base Call 1242 Image Data Subset 1244 signal-to-noise ratio maximized version 1300 Sequencing Runs 1312 No. 1 Specialist Signal Profiler 1314 2nd Specialist Signal Profiler 1318 3rd Specialist Signal Profiler 1410 Cluster Class 1412 Cluster Subclass 1414 Cluster Subclass 1422 No. 1 Specialist Signal Profiler 1424 2nd Specialist Signal Profiler 1428 3rd Specialist Signal Profiler 1432 4th Specialist Signal Profiler 1434 5th Specialist Signal Profiler 1438 6th Specialist Signal Profiler 1502 cluster / well population 1508 Specialist Signal Profiler 1602 Training Stage 1608 Inference Stage 1612 Training Data 1618 Inference Data 1622 Segmentation Logic 1632 training data 1638 Inference Data 1642 Offline Training Logic 1648 Runtime Logic 1658 Specialist Signal Profiler 1802 Inference Data 1812 Inference Data 1818 Inference Data 1832 Inference Data 1838 Inference Data 1842 Offline Training Logic 1858 Specialist Signal Profiler 1902 Signal distribution 1904 Cluster Set 1908 Specialist Signal Profiler 2113 operations 2114 Sequencing Images 2115 Operation 2116 operations 2118 operations 2123 operations 2124 operations 2125 Operation 2128 Quality Score 2133 operations 2134 operations 2135 Operation 2136 operations 2138 InterOp File 2148 Log 2152 Quality Table (Q Table) 2302 Equalizer coefficients 2302 First set of equalizer coefficients 2304 Second set of equalizer coefficients 2306 input image pixels 2308 input image pixels 2322 Base Call Logic 2324 Base Call 2336 Base call error 2338 Base call error 2342 Update Logic 2354 Second set 2356 Equalizer Coefficients 2358 Equalizer Coefficients 2362 1st set 2612 Center of gravity 2712 Center of gravity 3700 Exemplary Computer System 3700 Computer Systems 3710 Storage Subsystem 3718 Specialist Signal Profiler 3722 Memory Subsystem 3732 Main random access memory (RAM) 3734 Dedicated memory (read only memory, ROM) 3736 File Storage Subsystem 3738 User Interface Input Devices 3755 Bus Subsystem 3772 Central Processing Unit (CPU) 3774 Network Interface Subsystem 3776 User Interface Output Device 3778 processor 5081 1st Lane 7141 #1 Specialist Signal Profiler

Claims

1. A system comprising: at least one processor; and a non-transitory computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to: store a plurality of specialist signal profilers for sequence signals from a sample, each specialist signal profiler within the plurality of specialist signal profilers being trained to maximize the signal-to-noise ratio of a particular sequence signal within a particular signal profile using intensity data classified according to different signal distribution configurations of signal profiles encoded in different spatial configurations or intensity data of the position of the sample on the biosensor, perform a basecall operation by applying, to each sequence signal within each signal profile detected for a particular sample from the particular intensity data classified according to the different spatial configurations or different signal distribution configurations, a respective specialist signal profiler from among the plurality of specialist signal profilers; a non-transitory computer-readable medium.

2. The system according to claim 1, wherein the particular intensity data classified according to the different spatial configurations of the particular sample contributes to the creation of the respective signal profiles during the basecall operation.

3. The system according to claim 2, wherein the different spatial configurations include the particular sample disposed on different surfaces of the biosensor on which the basecall operation is performed.

4. The system according to claim 2, wherein the different spatial configurations include samples disposed in different tile groups of the biosensor.

5. The system according to claim 3, wherein the different spatial configurations include the particular sample disposed on different lanes of the biosensor.

6. The system according to any one of claims 3 to 5, wherein the different spatial configurations include samples disposed on different lane groups of the biosensor.

7. The system according to claim 6, wherein the different lane groups include upper peripheral lanes, central lanes, and lower peripheral lanes.

8. The system according to claim 6, wherein the different lane groups include edge lanes and non-edge lanes.

9. The system according to claim 5, wherein the different spatial configurations include the specific sample disposed on different samples of different lanes of the biosensor.

10. The system according to claim 9, wherein the different samples include an upper peripheral sample, a central sample, and a bottom peripheral sample.

11. The system according to claim 9, wherein the different samples include an edge sample and a central sample.

12. The system according to any one of claims 9 to 11, wherein the different spatial configurations include the specific sample disposed on different tiles of the different samples of different lanes of the biosensor.

13. A non-transitory computer-readable medium including instructions that, when executed by at least one processor, cause a computer device to store a plurality of specialist signal profilers of sequence signals from a sample, wherein each specialist signal profiler in the plurality of specialist signal profilers is trained to maximize the signal-to-noise ratio of a specific sequence signal in a specific signal profile using intensity data classified according to different signal distribution configurations of signal profiles encoded in different spatial configurations or intensity data of the position of the sample on the biosensor, perform a basecall operation by applying each specialist signal profiler from among the plurality of specialist signal profilers to each sequence signal in each signal profile detected for the sample for the specific sample from the specific intensity data classified according to the different spatial configurations or different signal distribution configurations. Non-transitory computer-readable storage medium.

14. The different spatial configurations include that the specific sample is disposed on different surfaces of the biosensor on which the basecall operation is performed, The non-transitory computer-readable storage medium according to claim 13, wherein the different spatial configurations include the specific sample disposed on different tile groups of the biosensor.

15. The non-transitory computer-readable storage medium according to claim 13, wherein the different spatial configurations include the specific sample disposed on different samples of different lanes of the biosensor.

16. The non - transitory computer - readable storage medium according to claim 14 or 15, wherein the different spatial configurations include the specific sample disposed on different sections of the biosensor.

17. The non - transitory computer - readable storage medium according to claim 14, wherein the different sections of the biosensor include a top - right section, a top - center section, a top - left section, a center - right section, a center section, a center - left section, a bottom - left section, a bottom - center section, and a bottom - right section.

18. A computer - implemented method, storing a plurality of specialist signal profilers for sequence signals from a sample, wherein each specialist signal profiler within the plurality of specialist signal profilers is trained to (ii) use intensity data classified according to different signal distribution configurations of signal profiles encoded in different spatial configurations or intensity data of the position of a sample on a biosensor to maximize the signal - to - noise ratio of a specific sequence signal within a specific signal profile; performing a basecall operation by applying each specialist signal profiler from among the plurality of specialist signal profilers to each sequence signal within each signal profile detected for a specific sample from the specific intensity data classified according to the different spatial configurations or different signal distribution configurations; A computer - implemented method comprising the above.

19. The computer - implemented method according to claim 18, wherein the specific intensity data classified according to the different spatial configurations of the specific sample contributes to the creation of the respective signal profiles during the basecall operation.

20. The computer - implemented method according to claim 19, wherein the different spatial configurations include the specific sample disposed on different surfaces of the biosensor on which the basecall operation is performed.