Dependence of base calling on flow cell tilt
By measuring flow cell height and applying adaptive defocus correction, the technique addresses focus variations in sequencing-by-synthesis, enhancing base calling accuracy and reducing errors in high-density DNA sequencing.
Patent Information
- Application Number
- JP2024572199
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-09
- Filing Date
- 2023-06-09
- Publication Date
- 2025-08-15
AI Technical Summary
Sequencing-by-synthesis techniques face challenges due to focus variations caused by flow cell tilt and non-planarity, leading to reduced base-calling accuracy and data loss, particularly in high-density DNA sequencing processes.
The technique involves measuring the flow cell surface height across the entire flow cell and adaptively setting the focal height of the imaging device, segmenting images based on focal height differences, and applying defocus correction filters to improve base calling accuracy.
This approach enhances base calling quality by reducing the effects of defocus and spatial crosstalk, resulting in improved sequencing accuracy and reduced errors.
Smart Images

Figure 2025526537000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 350,776, filed June 9, 2022, the entire disclosure of which is incorporated herein by reference in its entirety.
[0002] The disclosed technology relates to sequencing by synthesis, which determines gene sequences in parallel by base calling many nucleotides of a gene sequence in parallel. Base calling is enhanced by relying on the focus / tilt of a flow cell that holds portions of the genetic material. Base calling enhancement is related to the imaging process.
[0003] Incorporation by Reference The following are incorporated by reference for all purposes as if fully set forth herein: U.S. Non-Provisional Patent Application No. 8,422,031(B2), filed April 16, 2013, entitled "Focusing Methods and Optical Systems and Assemblies Using the Same."
[0004] U.S. Nonprovisional Patent Application No. 15 / 936,365, filed March 26, 2018, entitled "DETECTION APPARATUSHAVING A MICROFLUOROMETER, A FLUIDIC SYSTEM, AND A FLOW CELL LATCH CLAMP MODULE"; U.S. Nonprovisional Patent Application No. 16 / 567,224, filed September 11, 2019, entitled "FLOW CELLS AND METHODS RELATED TO SAME"; U.S. Nonprovisional Patent Application No. 16 / 439,635, filed June 12, 2019, entitled "DEVICE FOR LUMINESCENT IMAGING"; U.S. Nonprovisional Patent Application No. 15 / 594,413, filed May 12, 2017, entitled "INTEGRATED OPTOELECTRONIC READ HEAD AND FLUIDIC CARTRIDGE USEFUL FOR NUCLEIC ACID SEQUENCING"; U.S. Nonprovisional Patent Application No. 16 / 351,193, filed March 12, 2019, entitled "ILLUMINATION FOR FLUORESCENCE IMAGING USING OBJECTIVE LENS"; U.S. Non-Provisional Patent Application No. 12 / 638,770, filed December 15, 2009, entitled "DYNAMIC AUTOFOCUS METHOD AND SYSTEM FOR ASSAY IMAGER"; U.S. Nonprovisional Patent Application No. 13 / 783,43, filed March 1, 2013, entitled "KINETIC EXCLUSION AMPLIFICATION OF NUCLEIC ACID LIBRARIES"; U.S. Non-Provisional Patent Application No. 13 / 006,206, filed January 13, 2011, entitled "DATA PROCESSING SYSTEM AND METHODS"; U.S. Nonprovisional Patent Application No. 14 / 530,299, filed October 31, 2014, entitled "IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS"; U.S. Nonprovisional Patent Application No. 15 / 153,953, filed December 3, 2014, entitled "METHODS AND SYSTEMS FOR ANALYZING IMAGE DATA"; U.S. Nonprovisional Patent Application No. 14 / 20,570, filed September 6, 2013, entitled "Centroid Markers for Image Analysis of High Density Clusters in Complex Polynucleotide Sequencing"; U.S. Nonprovisional Patent Application No. 14 / 530,299, filed October 31, 2014, entitled "IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS"; U.S. Nonprovisional Patent Application No. 12 / 565,341, filed September 23, 2009, entitled "METHOD AND SYSTEM FOR DETERMINING THE ACCURACY OF DNA BASE IDENTIFICATIONS"; U.S. Nonprovisional Patent Application No. 12 / 295,337, filed March 30, 2007, entitled "SYSTEMS AND DEVICES FOR SEQUENCE BY SYNTHESIS ANALYSIS"; U.S. Non-Provisional Patent Application No. 12 / 20,739, filed January 28, 2008, entitled "IMAGE DATA EFFICIENT GENETIC SEQUENCING METHOD AND SYSTEM"; U.S. Nonprovisional Patent Application No. 13 / 833,619 (Attorney Docket No. IP-0626-US), filed March 15, 2013, entitled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR SAME"; U.S. Nonprovisional Patent Application No. 15 / 175,489 (Attorney Docket No. IP-0689-US), filed June 7, 2016, entitled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND METHODS OF MANUFACTURING THE SAME"; U.S. Nonprovisional Patent Application No. 13 / 882,088 (Attorney Docket No. IP-0462-US), filed April 26, 2013, entitled "MICRODEVICES AND BIOSENSOR CARTRIDGES FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR THE SAME"; U.S. Nonprovisional Patent Application No. 13 / 624,200 (Attorney Docket No. IP-0538-US), filed September 21, 2012, entitled "METHODS AND COMPOSITIONS FOR NUCLEIC ACID SEQUENCING"; U.S. Non-Provisional Patent Application No. 17 / 308,35 (Attorney Docket No. ILLM 1032-2 / 1991-US), filed May 4, 2021, entitled "EQUALIZATION-BASED IMAGE PROCESSING AND SPATIAL CROSSTALK ATTENUATOR."
[0005] U.S. Provisional Patent Application No. 62 / 821,602 (Attorney Docket No. ILLM1008-1 / IP-1693-PRV), filed March 21, 2019, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing"; U.S. Provisional Patent Application No. 62 / 821,618 (Attorney Docket No. ILLM1008-3 / IP-1741-PRV), filed March 21, 2019, entitled "Artificial Intelligence-Based Generation of Sequencing Metadata"; U.S. Provisional Patent Application No. 62 / 821,681 (Attorney Docket No. ILLM1008-4 / IP-1744-PRV), filed March 21, 2019, entitled "Artificial Intelligence-Based Base Calling"; U.S. Provisional Patent Application No. 62 / 821,724 (Attorney Docket No. ILLM1008-7 / IP-1747-PRV), filed March 21, 2019, entitled "Artificial Intelligence-Based Quality Scoring"; U.S. Provisional Patent Application No. 62 / 821,766 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV), filed March 21, 2019, entitled "Artificial Intelligence-Based Sequencing"; Dutch Patent Application No. 2023310 (Attorney Reference Number ILLM1008-11 / IP-1693-NL), filed on June 14, 2019, invention title: "Training Data Generation for Artificial Intelligence-Based Sequencing" Dutch Patent Application No. 2023311 (Attorney Reference No. ILLM1008-12 / IP-1741-NL), filed on June 14, 2019, Title of invention: "Artificial Intelligence-Based Generation of Sequencing Metadata" Dutch Patent Application No. 2023312 (Attorney Reference Number ILLM1008-13 / IP-1744-NL), filed on June 14, 2019, Title of invention: "Artificial Intelligence-Based Base Calling" Dutch Patent Application No. 2023314 (Attorney Reference No. ILLM1008-14 / IP-1747-NL), filed on June 14, 2019, entitled "Artificial Intelligence-Based Quality Scoring"; Dutch Patent Application No. 2023316 (Attorney Reference Number ILLM1008-15 / IP-1752-NL), filed on June 14, 2019, title of invention: "Artificial Intelligence-Based Sequencing."
[0006] U.S. Nonprovisional Patent Application No. 16 / 825,987 (Attorney Docket No. ILLM1008-16 / IP-1693-US), filed March 20, 2020, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing"; U.S. Nonprovisional Patent Application No. 16 / 825,991 (Attorney Docket No. ILLM1008-17 / IP-1741-US), filed March 20, 2020, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing"; U.S. Nonprovisional Patent Application No. 16 / 826,126 (Attorney Docket No. ILLM1008-18 / IP-1744-US), filed March 20, 2020, entitled "Artificial Intelligence-Based Base Calling"; U.S. Nonprovisional Patent Application No. 16 / 826,134 (Attorney Docket No. ILLM1008-19 / IP-1747-US), filed March 20, 2020, entitled "Artificial Intelligence-Based Quality Scoring"; U.S. Nonprovisional Patent Application No. 16 / 826,168 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV), filed March 21, 2020, entitled "Artificial Intelligence-Based Sequencing"; U.S. Nonprovisional Patent Application No. 17 / 511,483 (Attorney Docket No. ILLM 1053-1 / IP-2214-US), filed October 26, 2021, entitled "Intensity Extraction with Interpolation and Adaptation for Base Calling"; U.S. Provisional Patent Application No. 17 / 687,586 (Attorney Docket No. ILLM 1033-2 / IP-2007-US), filed March 4, 2022, entitled "Artificial Intelligence-Based Base Caller with Contextual Awareness"; U.S. Nonprovisional Patent Application No. 16 / 826,126 (Attorney Docket No. ILLM1008-18 / IP-1744-US), filed March 30, 2020, entitled "Artificial Intelligence-Based Base Calling"; U.S. Nonprovisional Patent No. 10,830,700(B2), filed March 1, 2019, entitled "Solid Inspection Apparatus and Method of Use"; U.S. Nonprovisional Patent Application No. 17 / 179,395 (Attorney Docket No. ILLM 1029-2 / IP-1964-US), filed February 18, 2021, entitled "Data Compression for Artificial Intelligence-Based Base Calling"; U.S. Nonprovisional Patent Application No. 17 / 180,480 (Attorney Docket No. ILLM 1030-2 / IP-1982-US), filed February 19, 2021, entitled "Split Architecture for Artificial Intelligence-Based Base Caller"; U.S. Nonprovisional Patent Application No. 17 / 180,513 (Attorney Docket No. ILLM 1031-2 / IP-1965-US), filed February 19, 2021, entitled "Bus Network for Artificial Intelligence-Based Base Caller"; U.S. Provisional Patent Application No. 62 / 849,091 (Attorney Docket No. ILLM1011-1 / IP-1750-PRV), filed May 16, 2019, entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing"; U.S. Provisional Patent Application No. 62 / 849,132 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2), filed May 16, 2019, entitled "Base Calling Using Convolutions"; U.S. Provisional Patent Application No. 62 / 849,133, filed May 16, 2019 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3, entitled "Base Calling Using Compact Convolutions"); U.S. Provisional Patent Application No. 62 / 979,384 (Attorney Docket No. ILLM1015-1 / IP-1857-PRV), filed February 20, 2020, entitled "Artificial Intelligence-Based Base Calling of Index Sequences"; U.S. Provisional Patent Application No. 62 / 979,414 (Attorney Docket No. ILLM1016-1 / IP-1858-PRV), filed February 20, 2020, entitled "Artificial Intelligence-Based Many-To-Many Base Calling"; U.S. Provisional Patent Application No. 62 / 979,385 (Attorney Docket No. ILLM1017-1 / IP-1859-PRV), filed February 20, 2020, entitled "Knowledge Distillation-Based Compression of Artificial Intelligence-Based Base Caller"; U.S. Provisional Patent Application No. 62 / 979,412 (Attorney Docket No. ILLM1020-1 / IP-1866-PRV), filed February 20, 2020, entitled "Multi-Cycle Cluster Based Real Time Analysis System"; U.S. Provisional Patent Application No. 62 / 979,411 (Attorney Docket No. ILLM1029-1 / IP-1964-PRV), filed February 20, 2020, entitled "Data Compression for Artificial Intelligence-Based Base Calling"; U.S. Provisional Patent Application No. 62 / 979,399 (Attorney Docket No. ILLM1030-1 / IP-1982-PRV), filed February 20, 2020, entitled "Squeezing Layer for Artificial Intelligence-Based Base Calling"; U.S. Provisional Patent Application No. 63 / 228,954 (Attorney Docket No. ILLM 1021-1 / IP-1856-PRV), filed August 3, 2021, entitled "Self-Learned Base Caller"; U.S. Provisional Patent Application No. 63 / 300,531 (Attorney Docket No. IP-2205-PRV), filed January 18, 2022, entitled "Dynamic Detilt Focus Tracking," and U.S. Provisional Patent Application No. 63 / 072,032 (Attorney Docket No. ILLM1018-1 / IP-1860-PRV), filed August 28, 2020, entitled "Detecting and Filtering Clusters Based on Artificial Intelligence-Predicted Base Calls."
[0007] Entire image capture The following documents, filed with this provisional patent application, are incorporated in their entirety into and should be considered part of this provisional patent application:
[0008] Appendix, page 37. [Background technology]
[0009] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or problems associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to implementations of the claimed technology.
[0010] Various protocols in biological or chemical research involve conducting multiple controlled reactions on a localized support surface or within a predetermined reaction chamber. The desired reactions may then be observed or detected, and subsequent analysis can help identify or characterize the chemicals involved in the reactions. For example, in some multiplex assays, an unknown analyte bearing a distinguishable label (e.g., a fluorescent label) may be exposed to thousands of known probes under controlled conditions. Each known probe may be deposited in a corresponding well of a microplate. Observing any chemical reactions that occur between the known probes and the unknown analyte in the well can help identify or characterize the analyte. Other examples of such protocols include well-known DNA sequencing processes, such as sequencing-by-synthesis or circular array sequencing. In circular array sequencing, a high-density array of DNA features (e.g., template nucleic acids) is sequenced through repeated cycles of enzymatic manipulation. After each cycle, an image may be captured and subsequently analyzed with other images to determine the sequence of the DNA features.
[0011] As a first specific example, one well-known DNA sequencing system uses the pyrosequencing process and includes a chip with a fused fiber optic faceplate containing millions of wells. A single capture bead containing clonally amplified sstDNA from a genome of interest is deposited in each well. After the capture bead is deposited in the well, nucleotides are sequentially added to the well by flowing a solution containing specific nucleotides along the faceplate. The environment within the well is such that if the nucleotide flowing through a particular well complements the DNA strand on the corresponding capture bead, the nucleotide is added to the DNA strand. The colonies of DNA strands are called clusters. The incorporation of nucleotides into the clusters initiates a process that ultimately generates a chemiluminescent light signal. The system includes a CCD camera positioned directly adjacent to the faceplate and configured to detect the light signal from the DNA clusters in the wells. Subsequent analysis of images obtained throughout the pyrosequencing process allows the sequence of the genome of interest to be determined.
[0012] However, the pyrosequencing system described above, in addition to other systems, may have certain limitations. For example, the fiber optic faceplate is acid-etched to form millions of tiny wells. While the wells may be roughly spaced apart, it is difficult to know the exact location of each well relative to other adjacent wells. If a CCD camera were positioned directly adjacent to the faceplate, the wells would not be evenly distributed along the CCD camera's pixels, and thus would not be aligned with the pixels in a known manner. Spatial crosstalk, or well-to-well crosstalk between adjacent wells, makes it difficult to distinguish true optical signals from the well of interest from other unwanted optical signals in subsequent analysis. Furthermore, fluorescence emission is substantially isotropic. As analyte density increases, it becomes increasingly difficult to manage or account for unwanted emissions (e.g., crosstalk) from neighboring analytes. As a result, data recorded during a sequencing cycle must be carefully analyzed.
[0013] In a second specific example of sequencing by synthesis, the genetic sequence associated with a sample of DNA, RNA, protein, and / or other genetic material having a base sequence is determined. Gene sequences are useful for many purposes, including the diagnosis and treatment of disease.
[0014] As a third specific example related to sequencing-by-synthesis, tilt and / or non-planarity of the flow cell introduce focus variations across the flow cell. Focus and / or tilt adjustment techniques allow some sequencing imaging to proceed through establishing a best-fit plane across the sample so that the entire sample remains within the DoF of the optical imaging system. However, increasing the numerical aperture (NA), for example, reduces the available DoF. Thus, global and / or local variations in tilt and / or height result in displacement of portions of the sample outside the DoF, resulting in out-of-focus image portions and, therefore, reduced data quality and / or data loss. This results in reduced base-calling accuracy.
[0015] Sequencing-by-synthesis is a parallel technique for determining gene sequences that operates on a large number of oligonucleotides (sometimes referred to as oligos) of a sample at a time, one base position at a time for each oligo. Some implementations of sequencing-by-synthesis operate by cloning oligos onto a substrate, such as a slide and / or flow cell, arranged in multiple lanes and imaged as individual tiles within each lane. In some implementations, cloning is configured to preferentially clone each of multiple starting oligos into a respective cluster of oligos, such as a respective nanowell of a patterned flow cell.
[0016] Sequencing-by-synthesis proceeds in a series of sequencing cycles (sometimes simply referred to as cycles). Each sequencing cycle involves chemical, image capture, and base-calling processes. The result is a base (e.g., one of the four amino acids, adenine (A), guanine (G), thymine (T), and cytosine (C)) determined for each oligo in parallel. The chemistry is designed to add one dye-tagged complementary nucleotide (sometimes referred to as a fluorophore) to each clone (e.g., oligo) in each cluster in each cycle. The image capture process generally involves focusing and aligning an imager (e.g., a camera) with the tiles in the flow cell lane, illuminating the tiles (e.g., with one or more lasers) to stimulate fluorophore fluorescence, and capturing multiple images of the fluorescence (e.g., one to four images, each corresponding to a tile and each corresponding to a different wavelength). The base calling operation results in the identification of the determined base (e.g., one of A, G, T, and C) for each oligo in parallel. In some implementations, the image capture operation corresponds to a discrete point-and-shoot operation, e.g., the imager and flow cell are moved relative to each other, and then an image capture operation is performed on the tile. In some implementations, the image capture operation corresponds to a continuous scanning operation, e.g., the imager and flow cell are continuously moved relative to each other, and image capture is performed during the movement. In various continuous scanning implementations, a tile corresponds to any contiguous region of the sample.
[0017] Some implementations of sequencing-by-synthesis use fluorescently labeled nucleotides, such as fluorescently labeled deoxyribonucleoside triphosphates (dNTPs), as fluorophores. During each sequencing cycle, a single fluorophore is added to each of the oligos in parallel. An excitation source, such as a laser, stimulates the fluorescence of many fluorophores in parallel, and the fluorescing fluorophores are imaged in parallel via one or more imaging runs. Once imaging of the fluorophores added in a sequencing cycle is complete, the fluorophores added in that cycle are removed and / or inactivated, and sequencing proceeds to the next sequencing cycle. During the next sequencing cycle, a next single fluorophore is added to each of the oligos in parallel, and an excitation source stimulates the fluorescence of many fluorophores added in the next sequencing cycle in parallel, and the fluorescing fluorophores are imaged in parallel via one or more imaging runs. The sequencing cycle is repeated as necessary based on how many bases are in the oligo and / or other termination conditions.
[0018] Base calling accuracy is crucial for high-throughput DNA sequencing and downstream analyses such as read mapping and genome assembly. In various scenarios, tilt and / or height resulting in suboptimal focus is caused by the flow cell holder or flow cell or its elements (e.g., the glass / substrate of the flow cell and / or the patterned nanowells of the flow cell). In some implementations, spatial crosstalk between adjacent clusters and / or focus variations due to variations in tilt and / or flatness of the flow cell, etc., are responsible for the majority of sequencing errors. Therefore, by accounting for and / or correcting spatial crosstalk in cluster intensity data and / or by accounting for and / or correcting focus variations due to tilt and / or non-planarity relative to the image, etc., opportunities arise to reduce DNA sequencing errors and improve base calling accuracy. Summary of the Invention
[0019] The disclosed technology relates to sequencing-by-synthesis, which determines gene sequences by base calling many nucleotides of a gene sequence in parallel. A flow cell holds a portion of genetic material. Defocus is introduced during sequencing by tilting the flow cell and by variations in the flatness of the flow cell. Using techniques that exploit the dependency of base calling on flow cell tilt, the effects of defocus are reduced and base calling quality is improved. For example, the flow cell surface height is measured across the entire flow cell. The focal height of an imaging device having a sequencing sensor is optionally adaptively set one or more times during sequencing. Each image captured by the sensor is segmented based, for example, on the difference between the focal height and the measured flow cell surface height across the area of the sensor. For example, a filter associated with defocus correction is selected based at least in part on the difference between the focal height and the measured flow cell surface height in a particular region of the image being corrected for defocus. Defocus correction is performed using the selected filter, and base calling is performed using the resulting image information. [Brief explanation of the drawings]
[0020] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings may also be available in PAIR via the Supplemental Content tab.
[0021] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the disclosed technology. In the following description, various implementations of the disclosed technology are described with reference to the following drawings: [Figure 1AA] 1 illustrates an example of the dependence of base calling on flow cell tilt. [Figure 1AB] 1A illustrates the operation of an example of the dependence of base calling on flow cell tilt as depicted in FIG. 1A. [Figure 1AC] 1A and 1B illustrate schematically elements for imaging a flow cell, including selected details regarding the tilt of the flow cell. [Figure 1AD] 1 illustrates selected details regarding the flow cell slope. [Figure 1AE] 1 illustrates selected details regarding the non-planarity of the flow cell. [Figure 1A] 1 illustrates one implementation of generating a lookup table (LUT) / equalizer filter by training an equalizer. [Figure 1B] 1A depicts one implementation that uses the LUT / equalizer filter of FIG. 1A to attenuate spatial crosstalk from sensor pixels and base call clusters using the crosstalk-corrected sensor pixels. [Figure 2] Visualize one example of a sequencing image containing at least five clusters / well centers / point sources on a flow cell. [Figure 3] 3 is a visualization of an example of extracting a pixel patch (yellow) from the sequencing image of FIG. 2 such that the center of target cluster 1 (blue) is contained in the central pixel of the pixel patch. [Figure 4] 10 visualizes an example of cluster-to-pixel signaling. [Figure 5] 10 visualizes an example of signal overlap from cluster to pixel. [Figure 6] 10 visualizes an example of a cluster signal pattern. [Figure 7] 4 visualizes one example of a sub-pixel LUT grid used to attenuate spatial crosstalk from the pixel patches of FIG. 3. [Figure 8]1B illustrates the selection of a LUT / equalizer filter from the LUT bank of FIG. 1B based on the sub-pixel location of the cluster / well center within the pixel. [Figure 9] 1 illustrates an implementation in which the center of target cluster 1 (blue) is not substantially concentric with the center of the pixel. [Figure 10] 1 depicts one implementation that interpolates between a set of selected LUTs to generate respective LUT weights. [Figure 11] 10 shows a weight kernel generator that uses the calculated weights of LUTs 12, 7, 8, and 13 to generate a weight kernel. [Figure 12] 10 shows an element-wise multiplier that multiplies the interpolated pixel coefficients of a weight kernel by the luminance values of the pixels in a pixel patch element-wise and sums the intermediate products of the multiplication to produce an output. [Figure 13A] Examples of coefficients for LUTs 12, 7, 8, and 13 are shown. [Figure 13B] Examples of coefficients for LUTs 12, 7, 8, and 13 are shown. [Figure 13C] Examples of coefficients for LUTs 12, 7, 8, and 13 are shown. [Figure 13D] Examples of coefficients for LUTs 12, 7, 8, and 13 are shown. [Figure 13E] Examples of coefficients for LUTs 12, 7, 8, and 13 are shown. [Figure 13F] Examples of coefficients for LUTs 12, 7, 8, and 13 are shown. [Figure 14A] 1 illustrates an example of a weight kernel. [Figure 14B] 1 illustrates one embodiment of weight kernel generation logic used by the weight kernel generator to generate weight kernels from the calculated weights of LUTs 12, 7, 8, and 13. [Figure 14C] 1 illustrates one embodiment of weight kernel generation logic used by the weight kernel generator to generate weight kernels from the calculated weights of LUTs 12, 7, 8, and 13. [Figure 15A]We show and explain how the interpolated pixel coefficients of the weight kernel maximize the signal-to-noise ratio and restore the underlying signal of target cluster 1 from the signal corrupted by crosstalk from clusters 2, 3, 4, and 5. [Figure 15B] We show and explain how the interpolated pixel coefficients of the weight kernel maximize the signal-to-noise ratio and restore the underlying signal of target cluster 1 from the signal corrupted by crosstalk from clusters 2, 3, 4, and 5. [Figure 16] 10 shows one implementation of a base-by-base Gaussian fit centered around the base-by-base intensity target, which is used as the ground truth value for error calculation during training. [Figure 17A] FIG. 1 is a block diagram of an exemplary computer system. [Figure 17B] 1 illustrates training and generation elements that implement aspects of base calling that depend on the tilt of the flow cell. [Figure 18] 1 illustrates one implementation of an adaptive equalization technique that may be used to train an equalizer. [Figure 19A] 1 illustrates various performance metrics of the disclosed technology. [Figure 19B] 1 illustrates various performance metrics of the disclosed technology. [Figure 19C] 1 illustrates various performance metrics of the disclosed technology. [Figure 19D] 1 illustrates various performance metrics of the disclosed technology. [Figure 20A] 1 illustrates an example of a reference point. [Figure 20B] Illustrates exemplary criteria in the context of various focuses. [Figure 20C] 1 illustrates an example cross-correlation equation for a discrete function. [Figure 20D] 1 illustrates an exemplary scoring formula. DETAILED DESCRIPTION OF THE INVENTION
[0022] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology, and is provided in the context of one or more particular applications and related requirements. Various modifications to the disclosed implementations will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the disclosed implementations, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0023] The following detailed description is made with reference to the figures. Example implementations are described to illustrate the disclosed technology, not to limit its scope, which is defined by the claims. Those skilled in the art will recognize various equivalent variations to the following description.
[0024] Examples of selected terms Depending on the implementation, elements of an equalizer (e.g., a spatial equalizer), such as elements enabled to perform convolution, execute convolution, and / or manage layers, loss functions, and / or objective functions as well as lookup table information, may variously correspond to one or more hardware elements, one or more software elements, and / or various combinations of hardware and software elements. In a first example, a convolution element, such as an N×M×D convolution element, is implemented as hardware logic included in an Application Specific Integrated Circuit (ASIC). In a second example, multiple convolution layers are implemented in the TensorFlow machine learning framework on a collection of internet-connected servers. In a third example, a first one or more portions of a spatial equalizer, such as one or more convolution layers, are each implemented in hardware logic according to the first example, and a second one or more portions of the spatial equalizer, such as one or more convolution layers, are implemented on a collection of internet-connected servers according to the second example. Various implementations are contemplated using different combinations of hardware and software elements to provide corresponding price and performance points.
[0025] Exemplary implementations of a Real Time Analysis (RTA) architecture (e.g., an equalizer such as a spatial equalizer) include various collections of software and / or hardware elements that collectively perform operations according to the RTA architecture. Various RTA implementations differ according to machine learning frameworks, programming languages, runtime systems, operating systems, and underlying hardware resources. The underlying hardware resources variously include one or more computer systems having any combination of central processing units (CPUs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), coarse-grained reconfigurable architectures (CGRAs), application-specific integrated circuits (ASICs), application-specific instruction-set processors (ASIPs), and digital signal processors (DSPs), as well as computing systems in general, e.g., elements capable of executing programmed instructions specified via a programming language. Various RTA implementations are capable of storing programming information (such as code and data) on non-transitory computer-readable media and are further capable of executing the code and referencing the data in accordance with a program that implements the RTA architecture.
[0026] Examples of programming languages, code and / or data libraries, and / or operating environments that can be used for techniques implementing the dependence of base calling on flow cell tilt, as related to the expression of signal processing functions (e.g., equating and / or expectation maximization), include Python, Numpy, R, Java, Javascript, C#, C++, Julia, Shell, Go, TypeScript, and Scala.
[0027] An example of image collection is to use an imager to simultaneously capture light emitted by multiple fluorescently tagged nucleotides as a collected image when the nucleotides fluoresce in response to excitation energy (such as laser excitation energy). The image has one or more dimensions, such as a line of pixels or a two-dimensional array of pixels. The pixels are represented according to one or more values. In a first example, each pixel is represented by a single integer (such as an 8-bit integer) that represents the pixel's brightness (such as grayscale). In a second example, each pixel is represented by multiple integers (such as three 24-bit integers), each representing the pixel's brightness according to a respective band of wavelengths (such as a respective color).
[0028] An example of focus is when the element being imaged (e.g., a tile of a flow cell or a portion thereof) is nominally coincident with the focal plane of the imager, such that the element is within the depth of field (DoF) of the imager. Focus corresponds to nominal maximum clarity or maximum definition of the element. An example of upper focus is when the element is above the focal plane, such that the element is above the DoF (e.g., the element is too close to the imager to be in focus). An example of lower focus is when the element is below the focal plane, such that the element is below the DoF (e.g., the element is too far from the imager to be in focus).
[0029] Dependence of base calling on flow cell tilt In this disclosure, a training context and a production context are described. In some implementations, a laboratory instrument (sometimes referred to as a biological sequencing instrument) is used in the training context, and a production instrument (sometimes referred to as a biological sequencing instrument) is used in the production context. In some implementations, the laboratory instrument and the production instrument are used in the training context. The training context and the production context implement various RTA-related processes (such as one or more equalizer functions directed to implementing the dependency of base calling on flow cell tilt). In various implementations, all or any portion of the RTA-related processes in the training context are variously implemented on any one or more of the laboratory instruments, any one or more of the production instruments, and / or any one or more computer systems separate from the laboratory instruments and the production instruments. In various implementations, all or any portion of the RTA-related processing in the generation context is variously implemented on one or more of the lab equipment, on one or more of the generating equipment, and / or on one or more computer systems (e.g., one or more servers) separate from the lab equipment and the generating equipment. In various implementations, all or any portion of the RTA-related processing on the lab equipment is performed by one or more computer systems of the lab equipment. Similarly, in various implementations, all or any portion of the RTA-related processing on the generating equipment is performed by one or more computer systems of the generating equipment. In various implementations, all or any portion of the lab equipment is used primarily for image acquisition, and the RTA-related processing in the associated training context is performed on one or more computer systems separate from the lab equipment used primarily for image acquisition.
[0030] The dependence of base calling on flow cell tilt allows for enhanced sequencing-by-synthesis, which determines the sequence of bases in genetic material with improved accuracy compared to ignoring flow cell tilt. The improved accuracy, in turn, allows for improved performance and / or reduced cost. Sequencing-by-synthesis proceeds in parallel, one base at a time, for each of multiple oligos attached to all or any portion of the flow cell. Each base-by-base process for multiple oligos involves imaging a tile of the flow cell and using flow cell tilt-dependent base calling to enhance base-calling accuracy.
[0031] Recall that sequencing-by-synthesis proceeds, in part, by capturing and processing images, e.g., tiles of a flow cell. Consider an imager capturing an image of a portion of a flow cell having multiple portions. In some scenarios, the focusing technique used during image capture brings one of the portions into sharp focus, but one or more other portions are not in sharp focus due to a limited DoF and the portions being different distances from the imager. In some scenarios, the portions are at different distances from the imager because the flow cell is tilted relative to the optical surface of the imager and / or the flow cell is not uniformly flat (e.g., not uniformly planar) and therefore has different heights. The tilt and / or lack of uniform flatness results in non-uniformity of focus between portions of a single image, between different images, and between images of different flow cells. This non-uniformity of focus, in some scenarios, results in reduced base-calling accuracy.
[0032] Variously, the slope may be aligned with the scanning direction of the imager, perpendicular to the scanning direction, or oblique at any angle relative to the scanning direction. Variously, the slope may be essentially uniform across the flow cell, substantially variable across the flow cell, or variable between relatively uniform and relatively variable across the flow cell. In various scenarios, the non-planarity of the flow cell may vary similarly to the slope variation discussed above. In some implementations, slope is considered a vector having a magnitude (e.g., how much slope is present) and a direction (e.g., which direction the slope is pointing). In contrast, height is a scalar having only a magnitude (e.g., how far a point on the flow cell surface is from the image plane of the imager). In some implementations, two height measurements at two respective points on the flow cell surface can be used to determine slope as a vector having a magnitude and a direction.
[0033] Recall that the tilt of the flow cell relative to the optical surface affects focus due to different distances between different locations on the flow cell and the optical surface. In a first example, the center portion is sharply focused and the edge portions are not sharply focused. In a second example, the center portion is sharply focused, a first edge portion is at an upper focus, and a second edge portion, orthogonally opposite the first edge portion, is at a lower focus. Other examples are characterized by various portions that are in focus, at an upper focus, and at a lower focus. Flow cells are manufactured with various flatness tolerances. Thus, in some scenarios, the aforementioned focus variability is due, in part, to imperfect flatness of the flow cell. Furthermore, the flatness of a flow cell varies between various flow cells and within a single flow cell. Thus, in some scenarios, the tilt varies from flow cell to flow cell and from flow cell tile to flow cell tile.
[0034] Reducing the flatness tolerance of a flow cell can potentially reduce the cost of the flow cell. Reduced flow cell flatness increases the maximum slope and / or slope variability. Increasing product base calling throughput can potentially be accompanied by an increase in cluster density on the flow cell. Increased cluster density can potentially result in a reduced DoF (such as due to increased NA), thus increasing the impact of slope and / or flow cell non-planarity. Increased cluster density can potentially increase optical crosstalk, thus increasing the difficulty in making accurate base calls.
[0035] Some imagers and / or imaging systems are capable of measuring and / or determining the tilt and / or height of a flow cell. In a first example, a multi-spot focus tracker measures defocus at multiple locations in the image plane. The defocus measurements are processed to determine tilt at the multiple locations. In a second example, a grid of resolution features, such as isolated nanowells, is included in the flow cell to enable defocus monitoring. The monitored defocus is processed to determine tilt at the grid locations. In a third example, optical aberrations are introduced into the imager's optical train (e.g., using a phase mask) so that the point spread function is asymmetric between the upper and lower foci, allowing for easy discrimination between defocus as upper versus lower foci. The discrimination is processed to determine tilt information. In a fourth example, the height of the flow cell is measured at multiple locations and used to create a surface map. The surface map is processed to determine tilt at the multiple locations. In various usage scenarios, the height of the flow cell remains stable during sequencing of the entire flow cell, lane, and / or column, allowing the surface map to be used for processing of the entire flow cell, lane, and / or column, respectively. In some implementations, the non-planarity of the flow cell is assessed via height measurement and / or height determination. In various implementations, the height of the flow cell is measured and / or determined according to various combinations of the aforementioned techniques for measuring and / or determining the slope of the flow cell.
[0036] In addition to the ability to measure the tilt and / or height of the flow cell, the base calling technique can be adapted according to various imaging conditions. Various imaging conditions include the sub-pixel location of the cluster within a pixel, the ratio of signal light to background light, and the size and / or shape of the point spread function. The inventors have recognized that various imaging conditions further include various degrees of defocus, e.g., the base calling technique can be further adapted according to various degrees of defocus, such as across different portions of the field of view of the imager. More specifically, the inventors have recognized that measuring and / or determining flow cell tilt (e.g., by processing focus / defocus information) and providing a flow cell tilt measurement to inform base calling allows for improved accuracy of base calling.
[0037] According to various implementations, measurements of the slope (and / or height) of the flow cell and / or measurements of information used to determine the slope (and / or height) of the flow cell are collected at various time points. The determination of the slope (and / or height) of the flow cell is determined at various time points according to various implementations. According to various implementations, base calling is informed of the measurements of the slope (and / or height) of the flow cell and / or the determination of the slope (and / or height) of the flow cell at various time points. Focus adjustment is optionally performed at various time points according to various implementations. Tilt adjustment is optionally performed at various time points. The various time points include one or more times over the lifetime of the instrument, one or more times per sequencing-by-synthesis run, one or more times per sequencing-by-synthesis cycle, one or more times per flow cell, one or more times per lane of the flow cell, one or more times per column of the lane, one or more times per tile, and / or one or more times per portion of the flow cell, and vary according to various implementations.
[0038] In various implementations, measurements of the flow cell slope (and / or height) and / or measurements of information used to determine the flow cell slope (and / or height) are collected at a first time set, a determination of the flow cell slope (and / or height) is determined at a second time set, and base calling is notified of the flow cell slope (and / or height) measurements and / or flow cell slope (and / or height) determinations at a third time set. Some implementations configure the first, second, and third time sets to have some predetermined relationship to one another. For example, the entire surface of the flow cell is first mapped out, e.g., before any images are captured. Then, once per tile, the base caller is notified of the tile slope measurements based on the map. Alternatively, four times per tile, corresponding to each of the four quarters of the tile, the base caller is notified of the slope measurements for each quarter of the tile. In another embodiment, the tilt is determined for each tile as it is imaged, and the base caller is notified of the measured tilt as each tile is processed by the base caller.
[0039] According to various implementations, the slope (and / or height) of the flow cell is variously determined at the aforementioned time, in coordination with the measurement of the slope (and / or height) of the flow cell, or alternatively at a time different from the time of the measurement of the slope (and / or height) of the flow cell.
[0040] According to various implementations, base calling is variously informed of the flow cell slope (and / or height) measurement at said time, either in coordination with the flow cell slope (and / or height) determination or at a time different from the time of the flow cell slope (and / or height) determination.
[0041] For clarity of explanation, the dependence of base calling on flow cell tilt is described in the context of an assumed base calling implementation based on a spatial equalizer, generally referred to herein as RTA-based base calling. Spatial equalizer implementations are now described in the context of a single base caller. Other implementations of base calling dependence on flow cell tilt use other than spatial equalizer techniques, depending on the implementation.
[0042] Multiple Base Caller FIG. 1AA illustrates an example of the dependence of base calling on flow cell tilt. The top of the diagram illustrates a training context (such as using laboratory sequencing by synthesis instruments), and the bottom illustrates a production context (such as using one or more sequencing by synthesis production instruments). From left to right, the diagram shows the flow cell, imaging, and RTA sections. As shown, the RTA section is implemented using multiple base callers, each implemented with its own equalizer and LUT (lookup table) elements.
[0043] Conceptually, base calling is performed using knowledge of the slope of the flow cell. The flow cell slope is measured (slope measurement) and evaluated (slope evaluation). The evaluation determines whether the flow cell (or any portion thereof, e.g., a lane, column, tile, or part thereof) is in super-focus, in-focus, or in-focus. Alternatively, the evaluation determines whether all or any region of the image (e.g., one or more clustered patches within the image, a contiguous region of pixels in the image, or one or more regular sections of the image, etc.) is in super-focus, in-focus, or in-focus. Still alternatively, the evaluation is used to determine how to divide the image into regions according to super-focus, in-focus, or in-focus. A base caller is selected from among multiple base callers based on the slope evaluation.
[0044] In a first example, if an image region is determined to be in focus, a base caller that has been trained or previously trained for use with the in-focus region is selected, and the image region is processed with the in-focus base caller. The "=base caller" element in the figure is an example of an in-focus base caller.
[0045] In a second example, if an image region is determined to be in super-focus, a base caller that has been trained for use with the super-focus region or that has previously been trained for use with the super-focus region is selected, and the image region is processed with the super-focus base caller. The "+ base caller" element in the figure is an example of a super-focus base caller.
[0046] In a third example, if an image region is determined to be hypofocused, a base caller that has been trained for use with a hypofocused region or that has previously been trained for use with a hypofocused region is selected, and the image region is processed with the hypofocused base caller. The "-base caller" element in the figure is an example of a hypofocused base caller.
[0047] During training, images (e.g., one image for each of multiple tiles of one or more flow cells) are collected and used in conjunction with ground truth (GT) to learn the parameters (sometimes referred to as weights) of the base callers in the training context, such as determining the information stored as coefficients in the LUT. Each image is processed as a single element or divided into multiple elements, depending on the implementation. Processing involves evaluating a gradient associated with each respective element of the single element or multiple elements. The evaluated gradient determines which of the multiple base callers will be trained for each respective element. Each base caller is associated with a respective GT set, and each base caller includes a respective LUT set. During training, each base caller is trained independently according to the gradient evaluation. After training is complete, the information in the LUT is provided to the RTA base caller in the production context for use in improving base calling compared to base calling without the benefit of training.
[0048] During generation, images, such as one image for each of multiple tiles of a flow cell, are collected and then processed for base calling. Each image is processed as a single element or divided into multiple elements, depending on the implementation. Similar to training, processing involves evaluating a gradient associated with each respective element of the single element or multiple elements. The evaluated gradient determines which of multiple base callers is selected to perform base calling for the element. Because each base caller includes its own set of LUTs, the set of LUTs used to determine a base call depends on the evaluated gradient.
[0049] In some implementations, initial training is performed in the context of dedicated training (sometimes referred to as pre-training), and additional training is performed in the context of generation specific to each generation device.
[0050] Further details of training base callers and using them to perform base calling during generation are disclosed elsewhere herein in the context of a single base caller as described with respect to Figures 1A-19D. Further details of evaluating gradients are disclosed elsewhere herein.
[0051] Figure 1AB illustrates operations for an example of the dependence of base calling on flow cell tilt, as shown in Figure 1AA. This operation is repeated for all tiles of the flow cell. The operation begins by capturing an image of the tile and the tilt information associated with the image. Optionally, the image is divided into multiple portions. The entire image (or each image portion in turn) is then processed as an image region as follows:
[0052] These portions may be, for example, various geometrically regular portions (such as any of a 2x2, 3x3, or 4x4 grid of substantially equal area portions), edge versus interior (non-edge) portions, and / or one or more contiguous regions, each of which collectively form the entire image, with each of these contiguous regions determined to be within a respective tilt range (and / or focus section, such as upper focus, in-focus, and lower focus) depending on the implementation.
[0053] The tilt and / or focus of the image region is estimated based on, for example, the tilt information of the tile or the tilt information of the image region determined from the tilt information of the tile.
[0054] In response to an image determined to be in focus, a focused base calling technique (=base caller) is selected. The focused base calling technique is appropriately used for training or generation, depending on the operational context. For training, a set of GTs corresponding to the focused context (=GT) is used to train the selected base caller, resulting in zero or more updates to the coefficients stored in the LUT of the selected base caller (=LUT).
[0055] In response to determining that the image is in superfocus, a superfocus base calling technique (+ base caller) is selected. The superfocus base calling technique is used for training or generation, as appropriate, depending on the context of operation. For training, a set of GTs corresponding to the superfocus context (+GT) is used to train the selected base caller, resulting in zero or more updates to the coefficients stored in the LUT of the selected base caller (+LUT).
[0056] In response to determining that the image is in down-focus, a down-focus base calling technique (-Base Caller) is selected. The down-focus base calling technique is used for training or generation, as appropriate, depending on the operational context. For training, a set of GTs corresponding to the down-focus context (-GT) is used to train the selected base caller, resulting in zero or more updates to the coefficients stored in the LUT of the selected base caller (-LUT).
[0057] The implementations described above to which Figures 1AA and 1AB relate are specific to flow cell tilts that result in various super-focused, in-focus, and sub-focused imaging. Other implementations are specific to flow cell heights that result in various super-focused, in-focus, and sub-focused imaging. Conceptually, the tilt assessment in Figure 1AA is instead a height assessment: height above the DoF-using (+ base caller) technique, height within the DoF-using (= base caller) technique, and height below the DoF-using (- base caller) technique. Further explanation is provided with respect to Figure 1AE.
[0058] FIG. 1AC schematically illustrates elements for imaging a flow cell, including selected details regarding the tilt of the flow cell. Depending on the implementation, a tilt measurement element is included that conceptually represents one or more dedicated elements, one or more capabilities present in non-dedicated elements, or a combination of both. Depending on the implementation, tilt measurement is implemented by one or more direct and / or indirect measurements and / or determinations based on one or more factors. Factors include tilt, focus, and / or distance. The section "Determining Tilt, Focus, and / or Distance" (located elsewhere herein) describes various techniques for measuring and / or determining tilt. In some implementations, the tilt measurement element includes capabilities for tilt measurement as well as height measurement. Further description is provided with respect to FIG. 1AE.
[0059] The flow cell is generally planar and includes multiple, generally parallel lanes that are imaged sequentially (fully automated) as a series of tiles organized, for example, in one or more rows, or alternatively, imaged and processed continuously (in a continuous scan) as a series of one or more tiles. The imager comprises a sensor, a semi-reflective mirror, and an objective lens. In some implementations, the laser and imager, as well as a mirror positioned to direct the laser radiation onto the semi-reflective mirror, are disposed within a module.
[0060] In some implementations, the imager and flow cell are moved relative to one another (such as by the flow cell traveling on a movable platform along a predetermined path, or by repositioning the imager and laser relative to the flow cell as the image is taken). In a continuous scanning implementation, some continuous area of the lane of the flow cell is imaged, corresponding to elements of a series of tiles.
[0061] In some implementations, the movable platform (sometimes referred to as a stage) comprises a flow cell receiving surface capable of supporting a flow cell. In some implementations, a controller is coupled to the stage and the optical assembly. Some implementations of the controller are configured to move the stage and the optical assembly relative to one another in a step-and-shoot manner, sometimes referred to as a step and settle technique. In various implementations, all or various portions of the tilt measurement and / or height measurement are implemented in the controller. In various implementations, a biological sequencing instrument (such as a laboratory instrument or a production instrument) includes all or some of the elements depicted in the figures. In various implementations, the biological sequencing instrument includes a stage, an optical assembly, and / or a controller.
[0062] During operation, the imager and flow cell are moved relative to one another, thus repositioning the imager from alignment with the (previous) tile to alignment with the (current) tile. Imaging proceeds by operating a laser. The laser radiation is reflected from the mirror onto a semi-reflective mirror, and then reflected from the semi-reflective mirror to illuminate the tiles of the flow cell, as indicated by the dashed arrow (power) pointing towards . In response to the illumination, fluorophores in the tiles fluoresce. Light from the fluorescence passes through an objective lens for focusing and subsequently passes through a semi-reflective mirror to form an image. The image is captured by a sensor.
[0063] For example, tilt and / or out-of-plane blur (conceptually depicted in the figure by the curved double-headed arrow "Tilt") is introduced by differences in distance between the imager and various regions of the tile being imaged. For example, a nominally planar flow cell may be out of optical alignment (e.g., tilted) with respect to the imager such that different portions (e.g., one or more edges) of the same tile are at different distances from the imager. Thus, depending on the DoF of the imager, one of the portions will be improperly focused and therefore degraded. In another example, an otherwise nominally planar flow cell has imperfections that cause one portion of the tile to be closer to the imager than another portion of the tile.
[0064] The diagram depicts an exemplary tilt of the flow cell. The flow cell is tilted with the left side up and the right side down. The direction of motion of the imager relative to the flow cell is left to right. Therefore, the tilt is aligned with the direction of motion. Other scenarios arise where the tilt is in any direction relative to the direction of motion. Returning to the diagram and the direction of tilt therein, the band on the left side of the flow cell is in upper focus, the middle band of the flow cell is in sharp focus, and the band on the right side of the flow cell is in lower focus. The upper focus band, the sharp focus band, and the lower focus band are illustrated in the image formed on the flow cell tiles and the sensor. In response to measuring, determining, and / or evaluating the tilt and / or focus of various regions of the image, base calling depends on the tilt of the flow cell. In particular, in response to an image region being in upper focus, a "+ Base Caller" is used for the image region (in training to determine parameters and in generation to call bases). In response to an image region being in sharp focus, a "= Base Caller" is used for the image region. In response to the image region being in down focus, the "-Base Caller" is used for the image region.
[0065] Some implementations of imagers use a point imaging technique that collects a relatively small collection of one or more pixels. Some implementations of imagers use an area imaging technique that collects a relatively large collection of pixels, such as a rectangular (e.g., square) shape. Some implementations of imagers use a line imaging technique that collects a relatively large collection of pixels, such as in a rectangular area with a relatively high aspect ratio. Some implementations of imagers, such as some variations of area imaging, use an area sensor that is coplanar with the collection area, with minimal optical components between the fluorescent fluorophores and the area sensor. Exemplary area sensors are based on semiconductor technology, such as Complementary Metal-Oxide Semiconductor (CMOS) chips.
[0066] Figure 1AD illustrates selected details regarding the flow cell tilt. Like-named elements in Figures 1AC and 1AD correspond to each other. The upper part of the figure (top view) is a view looking up into the sensor and depicts various focal zones of the image: above, sharp, and below. The lower part of the figure (side view) is a view looking from the side of the imager objective lens, with a portion of the flow cell being imaged. The tilt is such that the flow cell surface is above the image plane on the left and below the image plane on the right. Note that the figure is not to scale and that the tilt has been exaggerated for ease of understanding. Note also that the flow cell is shown as uniformly flat for ease of understanding. The bands of the image that are in sharp focus correspond to the depth of focus (DoF) of the imager. Bands of the image that are beyond the DoF relative to the image plane (either above or below the imaging plane) are blurred. As in Figure 1AC, a "+ base caller" is used in response to image regions that are in upper focus. In response to image regions that are in sharp focus, a "= base caller" is used. In response to image regions that are in under-focus, a "- base caller" is used. In some implementations, the point spread function (PSF) is asymmetric for blurred images that are in under-focus versus over-focus, allowing the blurred images to be classified as under-focus or over-focus based on the difference in the PSF.
[0067] FIG. 1AE illustrates selected details regarding the non-planarity of the flow cell. Likely named elements in FIGS. 1AC and 1AE correspond to one another. As with FIG. 1AD, the upper part (top view) of FIG. 1AE is a view looking up into the sensor and depicts various focal zones of imaging: above, sharp, and below. The lower part (side view) of the figure is a view looking from the side of the imager's objective lens, with a portion of the flow cell being imaged. The non-planarity of the flow cell is such that (as in FIG. 1AD) the flow cell surface is above the image plane on the left and below the image plane on the right. Note that the figure is not to scale and that the non-planarity has been exaggerated for ease of understanding. Furthermore, note that the depicted flow cell surface is a two-dimensional cross-section of a three-dimensional object (the flow cell), and the focal zones assume uniformity in the third dimension for ease of understanding. The sharply focused image zones correspond to the depth of focus (DoF) of the imager. Bands of the image beyond the DoF (above or below the imaging plane) relative to the image plane are blurred. Similar to FIG. 1AC, in response to image regions that are in super-focus, a "+ base caller" is used. In response to image regions that are in sharp focus, a "= base caller" is used. In response to image regions that are in sub-focus, a "- base caller" is used. In some implementations, the point spread function (PSF) is asymmetric for blurred images that are in super-focus versus sub-focus, allowing the blurred images to be classified as super-focus or sub-focus based on the difference in the PSF.
[0068] Similar to the inclusion of a tilt measurement element in FIG. 1AC, FIG. 1AD includes a height measurement element, which conceptually represents one or more dedicated elements, one or more capabilities present in non-dedicated elements, or a combination of both, depending on the implementation. Height measurement is implemented by one or more direct and / or indirect measurements and / or determinations based on one or more factors, depending on the implementation. Factors include tilt, focus, and / or distance. The section "Determining Tilt, Focus, and / or Distance" (located elsewhere in this specification) describes various techniques for measuring and / or determining height.
[0069] FIG. 1AD shows that tilt by itself is insufficient to determine whether an image is super-focused, in-focus, or sub-focused. In the figure, tilt is uniform across the entire image. However, a first portion of the image is super-focused, a second portion of the image is in-focus, and a third portion of the image is sub-focused. In contrast, FIG. 1AE shows that height alone is sufficient to determine whether an image is in-focus, sub-focused, or super-focused. The first portion of the image is super-focused, the second portion of the image is in-focus, and the third portion of the image is sub-focused.
[0070] Single Base Caller The preceding description assumes a spatial equalizer-based base calling implementation, collectively referred to herein as RTA-based base calling. Specific techniques for implementing RTA-based base calling using a spatial equalizer are described below. In various implementations, the equalizer base caller 104 (sometimes referred to as equalizer 104) of FIG. 1A is an exemplary implementation of the "+", "=", and "-" base caller elements of FIGS. 1AA-1AD. The ground truth base call 112 of FIG. 1A is an example of the "+", "=", and "-" GT elements of FIGS. 1AA-1AD. The lookup table 106 (sometimes referred to as LUT 106 or LUT bank 106) of FIG. 1A is an exemplary implementation of the "+", "=", and "-" LUT elements of FIGS. 1AA-1AD. Correspondingly, the sequencing image 102 of FIG. 1A corresponds to the image elements of FIGS. 1AA-1AD, and the trainer 114 of FIG. 1A corresponds to the trainer element of FIG. 1AA and the training element of FIG. 1AB.
[0071] Lookup table generation 1A illustrates one implementation for generating a look-up table (LUT) (or LUT bank) 106 by training an equalizer 104. The equalizer 104 is also referred to herein as an equalizer-based base caller 104. The system 100A includes a trainer 114 that trains the equalizer 104 using least-squares estimation. Additional details regarding equalizers and least-squares estimation can be found in appendices included with this application.
[0072] The sequencing image 102 is generated during a sequencing run performed by a sequencing instrument such as Illumina's iSeq, HiSeqX, HiSeq3000, HiSeq4000, HiSeq2500, NovaSeq6000, NextSeq550, NextSeq1000, NextSeq2000, NextSeqDx, MiSeq, and MiSeqDx. In one implementation, Illumina sequencers employ cyclic reversible termination (CRT) chemistry for base calling. This process relies on extending a nascent strand complementary to a template strand with fluorescently labeled nucleotides while tracking the emission signal of each newly added nucleotide. The fluorescently labeled nucleotides have a 3' removable block that anchors the fluorophore signal of the nucleotide type.
[0073] Sequencing occurs in repetitive cycles, each of which involves three steps: (a) extending the nascent strand by adding fluorescently labeled nucleotides; (b) exciting fluorophores using one or more lasers in the sequencing instrument's optical system and generating a sequencing image by imaging through different filters in the optical system; and (c) cleaving the fluorophores and removing the 3' block in preparation for the next sequencing cycle. The incorporation and imaging cycles are repeated for a specified number of sequencing cycles, defining the read length. Using this approach, each cycle locates a new position along the template strand.
[0074] The enormous power of Illumina sequencers stems from their ability to simultaneously perform and sense CRT reactions of millions or even billions of samples (e.g., clusters). Clusters contain approximately 1,000 identical copies of a template strand, but the size and shape of the clusters vary. Clusters are grown from template strands by bridge amplification or exclusion amplification of the input library prior to a sequencing run. The purpose of amplification and cluster growth is to increase the brightness of the emitted signal, since imaging devices cannot reliably sense the fluorophore signal of a single strand. However, because the physical distance between strands within a cluster is small, imaging devices perceive a cluster of strands as a single bright spot.
[0075] Sequencing occurs in a flow cell, a small glass slide that holds the input strand. The flow cell is connected to an optical system that includes microscope imaging, an excitation laser, and a fluorescence filter. The flow cell contains multiple chambers, called lanes. The lanes are physically separated from one another and can contain different tagged sequencing libraries that can be distinguished without sample cross-contamination. In some implementations, the flow cell includes a patterned surface. A "patterned surface" refers to the arrangement of different regions within or on an exposed layer of a solid support. For example, one or more of the regions can be features in which one or more amplification primers are present. The features can be separated by interstitial regions in which amplification primers are absent. In some implementations, the pattern can be an xy format of features in rows and columns. In some implementations, the pattern can be a repeating arrangement of features and / or interstitial regions. In some implementations, the pattern can be a random arrangement of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Pat. No. 8,778,849, U.S. Pat. No. 9,079,148, U.S. Pat. No. 8,778,848, and U.S. Patent Application Publication No. 2014 / 0243224, each of which is incorporated herein by reference.
[0076] In some implementations, the flow cell comprises an array of wells or depressions in a surface, which may be fabricated as generally known in the art using a variety of techniques, including, but not limited to, photolithography, stamping techniques, molding techniques, and microetching techniques. As will be appreciated by those skilled in the art, the technique used will depend on the composition and shape of the array substrate.
[0077] The features within the patterned surface can be wells in an array of wells (e.g., microwells or nanowells) on glass, silicon, plastic, or other suitable solid support with a patterned covalently attached gel, such as poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication No. 2013 / 184796, WO 2016 / 066586, and WO 2015-002813, each of which is incorporated by reference in its entirety). This process creates gel pads used for sequencing, which can be stable over many cycles and across sequencing runs. Covalently attaching a polymer to the wells is useful for maintaining the gel in the structured features throughout the life of the structured substrate during various uses. However, in many implementations, the gel need not be covalently attached to the wells. For example, in some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Pat. No. 8,563,477, the entirety of which is incorporated herein by reference) that is not covalently bonded to any part of the structured substrate can be used as the gel material.
[0078] In certain other implementations, a structured substrate can be created by patterning a solid support material with wells (e.g., microwells or nanowells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof, such as an azido-SFA version of SFA (azido-SFA)), and polishing the gel-coated support, e.g., by chemical or mechanical polishing, thereby retaining the gel within the wells but removing or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the wells. Primer nucleic acids can be attached to the gel material. A solution of target nucleic acids (e.g., a fragmented human genome) can then be contacted with the polished substrate such that individual target nucleic acids seed individual wells through interaction with primers attached to the gel material. However, due to the absence or inactivity of the gel material, the target nucleic acids do not occupy the interstitial regions. Amplification of the target nucleic acids is confined to the wells because the absence or inactivity of gel within the interstitial regions prevents outward migration of growing nucleic acid colonies. The process is manufacturable, scalable, and utilizes conventional micro- or nano-fabrication methods.
[0079] The sequencing instrument's imaging device (e.g., a solid-state imaging device such as a charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) sensor) takes snapshots at multiple locations along the lane in a series of non-overlapping regions called tiles. For example, there may be 64 or 96 tiles per lane. A tile holds hundreds of thousands to millions of clusters.
[0080] The output of a sequencing run is a set of sequencing images, each showing the intensity radiation of a cluster and its surrounding background. Sequencing images show the intensity radiation generated as a result of incorporating nucleotides into a sequence during sequencing. The intensity radiation arises from the associated specimens / clusters and their surrounding background.
[0081] Sequencing images 102 are provided from multiple sequencing instruments, sequencing runs, cycles, flow cells, tiles, wells, and clusters. In one implementation, the sequencing images are processed by the equalizer 104 on an imaging channel basis. A sequencing run generates m images per sequencing cycle corresponding to m imaging channels. In one implementation, each imaging channel corresponds to one of multiple filter wavelength bands. In another implementation, each imaging channel corresponds to one of multiple imaging events in a sequencing cycle. In yet another implementation, each imaging channel corresponds to a combination of illumination by a specific laser and imaging through a specific optical filter. In different implementations, such as 4-channel chemistry, 2-channel chemistry, and 1-channel chemistry, m is 4 or 2. In other implementations, m is 1, 3, or greater than 4.
[0082] In another implementation, the input data is based on pH changes induced by the release of hydrogen ions during molecular elongation. The pH change is detected and converted to a voltage change proportional to the number of incorporated bases (e.g., in the case of Ion Torrent). In yet another implementation, the input data is constructed from nanopore sensing, which uses a biosensor to measure the disruption of current as an analyte passes through or near the opening of the nanopore, determining the identity of the bases. For example, Oxford Nanopore Technologies (ONT) sequencing is based on the following concept: a single strand of DNA (or RNA) is passed through a membrane via a nanopore, and a voltage difference is applied across the membrane. Nucleotides present within the pore affect the electrical resistance of the pore, so that current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal (hence the "squiggle" that appears when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer Data Acquisition (DAC) values taken at (for example) a 4 kHz frequency. At a DNA strand speed of approximately 450 base pairs per second, this gives, on average, approximately 9 raw observations per base. This signal is then processed to identify breaks in the open pore signal that correspond to individual reads. Extension of these raw signals is base-called, a process that converts DAC values into sequences of DNA bases. In some implementations, the input data includes normalized or scaled DAC values.Additional information regarding non-image-based sequencing data can be found in U.S. Provisional Patent Application No. 62 / 849,132 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2), filed May 16, 2019, entitled "Base Calling Using Convolutions," U.S. Provisional Patent Application No. 62 / 849,133 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3), filed May 16, 2019, entitled "Base Calling Using Compact Convolutions," and U.S. Non-Provisional Patent Application No. 16 / 826,168 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV), filed March 21, 2020, entitled "Artificial Intelligence-Based Sequencing."
[0083] training The equalizer 104 generates a LUT bank having multiple LUTs (equalizer filters) 106 with sub-pixel resolution. In one implementation, the number of LUTs 106 generated by the equalizer 104 for the LUT bank depends on the number of sub-pixels into which the sensor pixels of the sequencing image 102 are divided or can be divided. For example, if the sensor pixels of the sequencing image 102 can each be divided into n×n sub-pixels (e.g., 5×5 sub-pixels), the equalizer 104 generates n 2 LUTs 106 (for example, 25 LUTs) are generated.
[0084] In one training implementation, data from the sequencing image is binned by well subpixel location. For example, for a 5x5 LUT, the center of 1 / 25th of the well is in bin (1,1) (e.g., the upper left corner of the sensor pixel), the center of 1 / 25th of the well is in bin (1,2), and so on. The equalizer coefficients for each well-center-bin are determined using least-squares estimation for the subset of data from the wells within each bin. The input to the equalizer 104 is the raw sensor pixels of the sequencing image for those bins. The resulting estimated equalizer coefficients vary from bin to bin.
[0085] Each LUT has multiple coefficients learned from training. In one implementation, the number of coefficients in the LUT corresponds to the number of sensor pixels used to base-call the cluster. For example, if the local grid of sensor pixels (image or pixel patch) used to base-call the cluster is of size p×p (e.g., a 9×9 pixel patch), then each LUT has p 2 has a coefficient of (e.g., a coefficient of 81).
[0086] The training generates equalizer coefficients configured to mix / combine pixel intensity values representing intensity radiation from the target cluster being base-called and intensity radiation from one or more neighboring clusters to maximize the signal-to-noise ratio. The signal that maximizes the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that minimizes the signal-to-noise ratio is the intensity radiation from neighboring clusters, i.e., spatial crosstalk, plus some random noise (e.g., to account for background intensity radiation). The equalizer coefficients are used as weights, and the mixing / combining involves performing element-wise multiplications between the equalizer coefficients and the pixel intensity values to calculate a weighted sum of the pixel intensity values.
[0087] During training, the equalizer 104, according to one implementation, learns to maximize the signal-to-noise ratio through least-squares estimation. Using least-squares estimation, the equalizer 104 is trained to estimate shared equalizer coefficients from the pixel intensities around the target well and the desired output. Least-squares estimation is suitable for this purpose because it minimizes the squared error and outputs coefficients that take into account the effects of noise amplification.
[0088] The desired output is an impulse at the well location (point source) when the luminance channel is on, and a background level when the luminance channel is off. In some implementations, ground truth base calls 112 are used to generate the desired output. In some implementations, the ground truth base calls 112 are modified to account for DC offsets per well, amplification factors, the degree of polyclonality, and gain offset parameters included in the least squares estimation. In one implementation, during training, a DC offset, i.e., a fixed offset, is calculated as part of the least squares estimation. During inference, the DC offset is added to each equalizer calculation as a bias.
[0089] In one implementation, the desired output is estimated using Illumina's Real-time Analysis (RTA) base caller, which does not use an equalizer. Details regarding RTA can be found in U.S. Patent Application No. 13 / 006,206, which is incorporated by reference as if fully set forth herein. The RTA base caller is used to generate ground truth base calls 112. This is because RTA has a low base calling error rate. Base calling errors are averaged over many training examples. In another implementation, ground truth base calls 112 are generated using aligned genomic data, which has better quality because it can use reference genomes and truth information that incorporate knowledge gained from multiple sequencing platforms and sequencing runs to average out noise.
[0090] The ground truth base calls 112 are base-specific intensity values that reliably represent the intensity profiles of bases A, C, G, and T, respectively. A base caller, such as an RTA, base calls clusters by processing the sequencing image 102 and generating per-color intensity values / output for each base call. The per-color intensity values can be considered per-base intensity values because a color is mapped to each of the bases A, C, G, and T depending on the type of chemistry (e.g., two-color chemistry or four-color chemistry). The base that matches the closest intensity profile is called.
[0091] FIG. 16 shows one implementation of a per-base Gaussian fit centered around a per-base intensity target, which is used as a ground truth value for error calculations during training. The per-base intensity outputs generated by the base caller for a large number of base calls in the training data (e.g., tens, hundreds, thousands, or millions of base calls) are used to generate a per-base intensity distribution. FIG. 16 shows a chart of four Gaussian clouds, which are probability distributions of the per-base intensity outputs for bases A, C, G, and T, respectively. The intensity values at the centers of the four Gaussian clouds are used as ground truth intensity targets given the ground truth base calls 112 for bases A, C, G, and T, respectively, and are referred to herein as intensity targets.
[0092] During training, consider that the input image data provided to the equalizer 104 is annotated with the base "A" as the ground truth base call. Then, the target / desired output of the equalizer 104 is the intensity value at the center of the green cloud in FIG. 16, i.e., the intensity target for base A. Similarly, for a ground truth base call of base "C," the desired output of the equalizer 104 is the intensity value at the center of the blue cloud in FIG. 16, i.e., the intensity target for base C. Thus, the target or desired output during training of the equalizer 104 is the average intensity for each base A, C, G, and T after averaging over the training data. In one implementation, the trainer 114 uses least-squares estimation to fit the coefficients of the equalizer 104 to minimize the equalizer output error to these intensity targets.
[0093] In one implementation, during training, the equalizer 104 applies coefficients in a given look-up table (LUT) to pixels of a sequencing image labeled with a given base. This involves element-wise multiplying the coefficients by the pixel's intensity value to generate a weighted sum of the intensity values, with the coefficients acting as weights. The weighted sum becomes the predicted output of the equalizer 104. Then, based on a cost / error function (e.g., sum of squared errors (SSE)), the error (e.g., least square error, minimum mean square error) between the weighted sum and an intensity target determined for a given base (e.g., as the average intensity observed at the given base from the center of the corresponding intensity Gaussian fit) is calculated. Cost functions such as SSE are differentiable functions used to estimate equalizer coefficients using an adaptive approach; therefore, the derivative of the error with respect to the coefficients can be evaluated, and these derivatives are then used to update the coefficients with values that minimize the error. This process is repeated until the updated coefficients no longer reduce the error. In other implementations, batch least squares is used to train the equalizer 104.
[0094] In other implementations, the per-base intensity distribution / Gaussian cloud shown in Figure 16 may be generated for each well and noise corrected by adding a DC offset, amplification factor, and / or phase parameter. In this way, depending on the well location of a particular well, the corresponding per-base Gaussian cloud may be used to generate a target intensity value for that particular well.
[0095] In one implementation, a bias term is added to the dot product that generates the output of the equalizer 104. During training, the bias parameter can be estimated using a similar approach used to learn the equalizer coefficients, i.e., least squares or least mean squares (LMS). In some implementations, the value of the bias parameter is a constant value equal to 1, i.e., a value that does not change with input pixel luminance. There is one bias per equalizer coefficient set. The bias is learned during training and then fixed for use during inference. The learned bias, along with the learned coefficients of each LUT, represents a DC offset used in all equalizer calculations during inference. This bias accounts for random noise caused by different cluster sizes, different background luminance, varying stimulus response, varying focus, varying sensor sensitivity, and varying lens aberrations.
[0096] In yet other decision-directed implementations, the output of the equalizer 104 is assumed to be correct for training purposes.
[0097] In another implementation of training, the equalizer 104 generates only a single LUT (equalizer filter) for a bin, and then uses multiple per-bin interpolation filters 108 to generate the remaining equalizer filters for the remaining bins. In this implementation, the sensor pixels around every well for every training example are resampled / interpolated to a well-aligned space (i.e., the wells are centered in their respective pixel patches / local grids). The resampled pixels for all examples are then consistently aligned across all wells.
[0098] However, to apply the single equalizer filter generated by the equalizer 104 in a practical online system for base calling, the raw sensor pixels of the sequencing image must be preprocessed to return them to a well-aligned space. That is, interpolation must be performed on the raw pixels surrounding each well, with the interpolation parameters varying depending on the subpixel location of a given well. To avoid this interpolation process, an overall response for a given well subpixel location is precalculated. By interpolating the raw pixel intensities into the well-aligned pixel space, well-aligned equalizer input values are calculated. The interpolated and equalizer responses are convolved together to reduce computation. Because the interpolation filter varies depending on the subpixel well location, this results in a different set of equalizer coefficients / filters for each subpixel well location, which in turn generates the remaining LUTs for the remaining bins. Therefore, in this implementation of training, while only the coefficients of the single equalizer filter are trained during training, the precalculation process generates a bank of LUT-based equalizers by applying the bin-specific interpolation filters 108 along with the single equalizer filter. where the LUT index is the subpixel well location.
[0099] The trainer 114 may train the equalizer 104 and generate trained coefficients for the LUT 106 using a number of training techniques. Examples of training techniques include least squares estimation, least squares, least mean squares, and recursive least squares. In least squares techniques, the parameters of a function are adjusted to best fit a data set so that the sum of squares of residuals is minimized. Additional details on least squares estimation algorithms may be found in Least Squares, https: / / en.wikipedia.org / w / index.php?title=Least_squares&oldid=951737821 (last visited April 28, 2020), which is incorporated by reference as if fully set forth herein. Ordinary least squares is a type of least squares method for estimation in linear regression models. Additional details about the least squares algorithm can be found in Ordinary least squares, https: / / en.wikipedia.org / w / index.php?title=Ordinary_least_squares&oldid=951770366 (last visited April 28, 2020), which is incorporated by reference as if fully set forth herein. In other implementations, other estimation and adaptive equalization algorithms can be used to train the equalizer 104.
[0100] The equalizer 104 may be trained in an offline mode, in which, according to one implementation, the trained coefficients of the LUT 106 are generated using the following batch least squares equalization logic:
[0101]
number
[0102] In the above equation, the LUT coefficients are the beta hats, the pixel intensities are X, and the target is y. A DC term is also added to the pixel intensities and a coefficient (e.g., an additional intensity term that is fixed to 1 in all cases). Next, as an example, consider X to be a matrix of size 82 (= 9 × 9 input intensities + constant DC term) × the number of training examples in the batch, and Y is the target output for each training example, i.e., each value is the intensity center of the ON / OFF cloud according to the training example truth. The beta hat is the set of coefficients that minimizes the sum of squares of the residuals, and is also of size 82 (= 9 × 9 coefficients + 1 DC term).
[0103] The equalizer 104 can also be trained in online mode, adapting the coefficients of the LUT 106 to track changes in temperature (e.g., optical distortion), focus, chemical, machine-specific variations, etc., on a tile-by-tile or subtile-by-subtile basis while the sequencer is operating and sequencing runs are cyclically progressing. In online mode, the trained coefficients of the LUT 106 are generated using adaptive equalization. In online mode, the least mean squares method, a form of stochastic gradient descent, is used as the training algorithm. Additional details about the least mean squares algorithm can be found in Least mean squares filter, https: / / en.wikipedia.org / w / index.php?title=Least_mean_squares_filter&oldid=941899198 (last visited April 28, 2020), which is incorporated by reference as if fully set forth herein.
[0104] The least mean squares technique uses the gradient of the squared error for each coefficient to move the coefficient in the direction that minimizes a cost function, which is the expected value of the squared error. This has a very low computational cost, and only multiplication and accumulation operations are performed per coefficient. No long-term storage is required except for the coefficients. The least mean squares technique is suitable for processing large amounts of data (e.g., processing data from billions of clusters in parallel). Extensions of the least mean squares technique include normalized least mean squares and frequency-domain least mean squares, which can also be used herein. In some implementations, the least mean squares technique can be applied in a decision-directed manner, assuming that our decisions are correct, i.e., our error rate is very low and a small mu value filters out impeded updates due to inaccurate base calls.
[0105] Figure 18 shows one implementation of an adaptive equalization technique that can be used to train the equalizer 104. Here, the equalization logic is y = x.h + d, where x is the input pixel luminance, h is the equalizer coefficient, and d is the DC offset. In one implementation, x and h are row and column vectors with length 81, respectively. This vector model corresponds to the dot product of a 9x9 matrix representing the input pixel and the coefficients. The cost is the expected squared error. The gradient update moves each coefficient in the direction that reduces the expected squared error, which leads to the next update.
[0106]
number
[0107] where h is a vector of equalizer coefficients (e.g., 9x9 equalizer coefficients), x is a vector of equalizer input intensities (e.g., 9x9 pixels in a pixel patch), and e is the error of the equalizer calculation performed using 81 values of x, i.e., only one error term per equalizer output.
[0108] Applying this update produces new estimates of the 9x9 equalizer coefficients. This estimate moves the equalizer coefficients (on average) in the direction that reduces the Mean Squared Error (MSE). 81 updates are made, one for each equalizer coefficient. In some implementations, Mu is a small constant used to change the adaptation rate / convergence speed. The DC term update may be calculated in a similar manner. The gain term update may be calculated in a similar manner.
[0109] Coefficient sets can be shared, for example, between tiles, regions of tiles, or flow cell surfaces by saving and restoring coefficient sets when the input data changes.
[0110] In some implementations, linear interpolation is applied to the coefficient set, so the updates are applied slightly differently as follows: h(q,n+1)=h(q,n)+lambda_q.mu.x(n).e(n)
[0111] where h(q,n) is the weight q at cycle n, and lambda_q is the linear interpolation weight for a particular set of coefficients, which may involve four updates per equalizer output by linear interpolation in two dimensions.
[0112] Recursive least squares techniques extend least squares techniques to recursive algorithms. Additional details of recursive least squares algorithms can be found in Recursive least squares filter, https: / / en.wikipedia.org / w / index.php?title=Recursive_least_squares_filter&oldid=916406502 (last visited April 28, 2020), which is incorporated by reference as if fully set forth herein.
[0113] In a multi-domain implementation, the LUTs 106 and their trained coefficients may be generated along multiple domains. Examples of domains include sequencer or sequencing equipment / machine (e.g., Illumina's NextSeq, MiSeq, HiSeq and their respective models), sequencing protocol and chemistry (e.g., blind amplification, exclusion amplification), sequencing run (e.g., forward and reverse), sequencing illumination (e.g., structured, unstructured, angled), sequencing device (e.g., overhead CCD camera, underlying CMOS sensor, one laser, multiple lasers), imaging technology (one channel, two channels, four channels), flow cell (e.g., patterned, unpatterned, embedded in a CMOS chip, underlying CCD camera), and spatial resolution on the flow cell (e.g., differentiating between different regions within the flow cell). These domains include different regions or quadrants (e.g., different tiles on a flow cell (e.g., in the case of edge wells on a tile closer to a laser or camera or fluidic system)) and different regions within a tile (e.g., different lanes on a tile (e.g., in the case of edge wells on a lane closer to a laser or camera or fluidic system)). Those skilled in the art will understand that other selectable domains and parameters typically associated with sequencing are similarly included (e.g., image processing algorithms, image alignment algorithms, ground truth annotation schemes (e.g., continuous labels such as intensity values, hard labels such as one-hot coding, soft labels such as softmax scores), temperature, focus, lenses, sequencing reagents, sequencing buffers).
[0114] A separate and distinct training set may be created for each domain using the sequencing images generated using each domain. The separate training sets may be used to train the equalizer 104, generating a LUT with trained coefficients for the corresponding regions. Trained coefficients specifically trained for each domain in the multiple domains may be stored and accessed during online mode depending on which domain or combination of domains is being used in a current or ongoing sequencing operation. For example, a first coefficient set more suitable for edge wells of a flow cell may be used for a sequencing operation, along with a second coefficient set more suitable for center wells of the same flow cell.
[0115] In one implementation, a configuration file may specify different combinations of domains and may be analyzed during online mode to select different coefficient sets specific to the domains identified by the configuration file.
[0116] In a multiple-training implementation, the equalizer 104 undergoes not only training but also pre-training. That is, the LUT 106 and its coefficients are first trained during a pre-training phase using a first training technique, and then re-trained or further trained during a further training phase using a second training technique. The first and second training techniques can be any of the training techniques described above. The first and second training techniques can be the same or different. For example, the pre-training phase can be an offline mode using a batch least-squares training technique, and the training phase can be an online mode using an iterative probability least-mean-squares technique.
[0117] In some implementations, multi-domain and multi-training implementations can be combined such that domain-specific coefficients are pre-trained and then further trained in a domain-specific manner. That is, further training (e.g., online mode) retrains the coefficients for that particular domain using only data that represents that particular domain and that is similar to the data used in the pre-training stage. In other knowledge transfer implementations, pre-training and training can use training data from the entire domain; for example, a coefficient set is generated during pre-training using images from a patterned flow cell, but is retrained during a subsequent training stage using images from an unpatterned flow cell.
[0118] Spatial Crosstalk Attenuator 2 depicts one implementation that uses the trained LUT / equalizer filter 106 of FIG. 1A to attenuate spatial crosstalk from sensor pixels and base call call clusters using the crosstalk-corrected sensor pixels. The trained equalizer base caller 104 operates during the inference stage in which base calling is performed. In some implementations, the operations shown in FIG. 2 are performed in a preprocessing stage before the base calling stage to generate crosstalk-corrected image data used by the base caller for base calling.
[0119] In one implementation, the equalizer coefficients are applied to pixel patches 120 (image patches or local grids of sensor pixels) extracted from the sequencing image 116 on an imaging channel basis and a target cluster basis. Regarding the imaging channel basis, in some implementations, each sequencing image has image data from multiple imaging channels. Consider the optical system of an Illumina sequencer that uses two different imaging channels, namely, a red channel and a green channel. Then, in each sequencing cycle, the optical system generates a red image with red channel intensity and a green image with green channel intensity, which together form a single sequencing image (like the RGB channels of a typical color image).
[0120] During training, the coefficients are trained / configured to maximize the signal-to-noise ratio (SNR) by minimizing the error between the predicted / estimated output and the desired / actual output. An example of an error is the mean squared error (MSE) or mean squared deviation (MSD). The signal maximized in the signal-to-noise ratio is the intensity radiation from the base-called target cluster (e.g., the cluster at the center of the image patch), and the noise minimized in the signal-to-noise ratio is the intensity radiation from one or more neighboring clusters, i.e., spatial crosstalk, as well as other noise sources (e.g., considering background intensity radiation). The trained coefficients are element-wise multiplied to the pixels of the image patch to calculate a weighted sum of the pixel intensity values. The weighted sum is then used to base call the target cluster.
[0121] In one implementation, the patch extractor 118 extracts red pixel patches from the red channel and green pixel patches for the green channel from a single sequencing image. In other implementations, red pixel patches are extracted from the red sequencing image of the target sequencing cycle, and green pixel patches are extracted from the green sequencing image of the target sequencing cycle. The coefficients of the LUT 106 are used to generate a red weighted sum for the red pixel patches and a green weighted sum for the green pixel patches. Both the red weighted sum and the green weighted sum are then used to base call the target cluster. The pixel patch 120 has dimensions w×h, where w (width) and h (height) are any number ranging from 1 to 10,000 (e.g., 3×3, 5×5, 7×7, 9×9, 15×15, 25×25). In some implementations, w and h are the same. In other implementations, w and h are different. Those skilled in the art will understand that one, two, three, four, or more channels or images of data can be generated per sequencing cycle for a target cluster, with one, two, three, four, or more patches extracted, respectively, to generate one, two, three, four, or more weighted sums, respectively, for base-calling the target cluster.
[0122] For target cluster-based extraction of pixel patches 120 from the sequencing image 116, the pixel extractor 118 extracts pixel patches 120 based on the locations of the cluster / well centers on the sequencing image 116, such that the central pixel of each extracted pixel patch contains the center of the target cluster / well. In some implementations, the patch extractor 118 locates the cluster / well centers on the sequencing image, identifies a pixel in the sequencing image that contains the cluster / well center (i.e., the central pixel), and extracts pixel patches of contiguous neighboring pixels around the central pixel.
[0123] 2 visualizes an example of a sequencing image 200 containing central / point sources of at least five clusters / wells on a flow cell. The pixels of sequencing image 200 show intensity emissions from target cluster 1 (blue) and additional neighboring clusters 2 (purple), 3 (orange), 4 (brown), and 5 (green).
[0124] Figure 3 visualizes one example of extracting pixel patch 300 (yellow) from sequencing image 200 such that the center of target cluster 1 (blue) is contained within central pixel 206 of pixel patch 300. Figure 3 also shows other pixels 202, 204, 214, and 216 that contain the centers of neighboring clusters 2 (purple), 3 (orange), 4 (brown), and 5 (green), respectively.
[0125] FIG. 4 visualizes one example of a cluster-to-pixel signal 400. In one implementation, the sensor pixels (yellow) are in the pixel plane. Spatial crosstalk is caused by clusters 412 periodically distributed in the sample plane (e.g., a flow cell). In one implementation, the target cluster and additional neighboring clusters are periodically distributed in a diamond shape on the flow cell and fixed onto the wells of the flow cell. In another implementation, the target cluster and additional neighboring clusters are periodically distributed on a hexagonal flow cell and fixed onto the wells of the flow cell. Signal cones 402 from the clusters are optically coupled to a local grid of sensor pixels (e.g., pixel patch 300) via at least one lens (e.g., one or more lenses of an overhead or adjacent CCD camera).
[0126] In addition to diamonds and hexagons, the clusters can be arranged in other regular shapes, such as squares, diamonds, triangles, etc. In still other implementations, the clusters are arranged on the sample plane in a random, non-periodic arrangement. Those skilled in the art will understand that the clusters can be arranged on the sample plane in any arrangement, as required by a particular sequencing implementation.
[0127] 5 visualizes one example of cluster-to-pixel signal overlap 500. Signal cones 402 overlap and impinge on sensor pixels, creating spatial crosstalk 502.
[0128] 6 visualizes one example of a cluster signal pattern 600. In one implementation, the cluster signal pattern 600 follows a decay pattern 602, where the cluster signal is strongest at the cluster center and decays as it propagates away from the cluster center.
[0129] 6 also shows one example of equalizer coefficients 604 trained / configured to maximize the signal-to-noise ratio by calculating a weighted sum of luminance radiation from target cluster 1 and luminance radiation from neighboring clusters 2, 3, 4, and 5. The equalizer coefficients 604 act as weights. The weighted sum is calculated by element-wise multiplying a first matrix containing the equalizer coefficients 604 by a second matrix containing pixel luminance values, where each pixel luminance value is the sum of radiation from one or more of clusters 1, 2, 3, 4, and 5, and other noise sources in the system as measured by the pixel sensor.
[0130] FIG. 7 visualizes one example of a subpixel LUT grid 700 used to attenuate spatial crosstalk from a pixel patch 300. Each pixel in the pixel patch 300 can be divided into multiple subpixels. In FIG. 7, pixel 206, which contains the center of target cluster 1 (blue), is divided into as many subpixels as the number of trained LUTs 106. That is, pixel 206 is divided into as many subpixels as the number of bins for which equalizer 104 generated LUTs 106 during training. As a result, each subpixel of pixel 206 corresponds to a respective LUT in the LUT bank generated by equalizer 104 using decision-directed feedback and least-squares estimation.
[0131] In the example shown in FIG. 7 , pixel 206 (center pixel) is divided into a 5×5 subpixel LUT grid 700 to generate 25 subpixels corresponding to the 25 LUTs (equalizer filters) generated by adaptive filter 104 as a result of training. Each of the 25 LUTs includes coefficients configured to mix / combine the luminance values of pixels in pixel patch 300 representing luminance radiation from target cluster 1 with luminance radiation from neighboring clusters 2, 3, 4, and 5, so as to maximize the signal-to-noise ratio. The signal that maximizes the signal-to-noise ratio is the luminance radiation from the target cluster, and the noise that minimizes the signal-to-noise ratio is the luminance radiation from neighboring clusters 2, 3, 4, and 5, i.e., spatial crosstalk plus some random noise (e.g., to account for background luminance radiation). The LUT coefficients are used as weights, and mixing / combining involves performing element-wise multiplication between the LUT coefficients and the luminance values of the pixels in pixel patch 300 to calculate a weighted sum of the pixel luminance values.
[0132] The number of coefficients in each of the 25 LUTs is the same as the number of pixels in pixel patch 300, i.e., a 9×9 coefficient grid in each LUT for the 9×9 pixels in pixel patch 300. This is because the coefficients are multiplied element-wise with the pixels in pixel patch 300.
[0133] In one implementation, a pixel-to-subpixel converter (not shown in FIG. 1B ) divides the pixels 206 into the subpixel LUT grid 700 based on a preset pixel divisor parameter (e.g., 1 / 5 pixel per subpixel to generate a 5×5 subpixel LUT grid 700). For example, the pixels may be divided into five subpixel bins with the following boundaries: −0.5, −0.3, −0.1, 0.1, 0.3, 0.5.
[0134] Note that in FIG. 7 , the center of target cluster 1 (blue) is substantially concentric with the center of transformed pixel 702. This is because the sequencing image 200, and therefore pixel patch 300, is resampled so that the center of target cluster 1 (blue) is substantially concentric with the center of transformed pixel 702 by (i) aligning the sequencing image 200 with respect to the template image and determining affine and nonlinear transformation parameters, (ii) using the parameters to transform the position coordinates of target cluster 1 (blue) into the image coordinates of sequencing image 200, and (iii) applying interpolation using the transformed position coordinates of target cluster 1 (blue) to make its center substantially concentric with the center of transformed pixel 702. The locations of the wells in the sample plane are known and can be used to calculate where the equalizer input for a particular well is in raw pixel space. Interpolation can then be used to recover the intensity at those locations from the raw image.
[0135] 8 illustrates the selection of a LUT / equalizer filter from LUT bank 106 based on the subpixel location of a cluster / well center within a pixel. Because the center of the target cluster (blue) is at a particular subpixel 12 in subpixel LUT grid 700, and because that particular subpixel 12 of pixel 206 corresponds to a LUT 12 in LUT bank 106, LUT selector 122 selects a LUT 12 and its coefficients from LUT bank 106 to apply to the pixels of pixel patch 300. Next, element-wise multiplier 134 multiplies the luminance values of the pixels in pixel patch 300 by the coefficients of LUT 12 element-wise and sums the products of the multiplications to generate an output (e.g., weighted sum 136). This output is used to base call target cluster 1 (e.g., provide this output as an input to base caller 138).
[0136] The equalizer 104 implements the following equalization logic when the target cluster is substantially concentric with the center of the pixel, as described above with respect to FIGS.
[0137]
number
[0138] where the well center coordinates (m,n) are integers to ensure that the wells are substantially aligned with the pixels; p(i,j) is the pixel intensity at location i,j; w(i,j) is the equalizer weight for the pixel at location i,j. i,j are summation limits operating over the range of pixels surrounding the well centered on p(m,n), e.g., -4<=i<=4, -4<=j<=4, and the output is a weighted average of the input pixels.
[0139] 9 illustrates an implementation in which the center of target cluster 1 (blue) is not substantially concentric with the center of pixel 206 because no resampling is performed as discussed with respect to FIG. 8. In such an implementation, interpolation is performed between a set of selected LUTs 124 to generate an interpolation LUT with interpolation coefficients. The interpolation LUT with interpolation coefficients is also referred to herein as a weight kernel 132.
[0140] First, as in Figure 8, a first LUT, i.e., LUT12, is selected that corresponds to the particular subpixel that contains the center of target cluster 1 (blue). Next, LUT selector 122 selects additional subpixel lookup tables from bank of subpixel lookup tables 106 that correspond to the subpixels that are closest in contiguous proximity to the particular subpixel. In Figure 9, the closest contiguous proximity subpixels that contact particular subpixel 12 are subpixels 7, 8, and 13, and therefore LUTs 7, 8, and 13 are selected from LUT bank 106, respectively.
[0141] 10 depicts one implementation for interpolating between a set of selected LUTs to generate respective LUT weights. Interpolator 126 is configured with interpolation logic (e.g., linear, bilinear, or bicubic interpolation) that uses coefficients from selected LUTs 12, 7, 8, and 13 to generate weights 128 for each of LUTs 12, 7, 8, and 13.
[0142] 13A, 13B, 13C, 13D, 13E, and 13F show example coefficients for LUTs 12, 7, 8, and 13. These figures also show example interpolation logic 1312, 1322, and 1332 used by interpolator 126 to calculate weights 128 for LUTs 12, 7, 8, and 13. These figures also show example weights 128 calculated for LUTs 12, 7, 8, and 13. These figures are snapshots of Excel sheets. The blue arrows and color coding in these figures are generated by Excel's Track Precedence function to indicate the interpolation logic.
[0143] FIG. 11 illustrates weight kernel generator 130 generating weight kernel 132 using weights 128 calculated for LUTs 12, 7, 8, and 13. FIG. 14A illustrates an example of weight kernel 132. FIGS. 14B and 14C illustrate example weight kernel generation logic 1402 used by weight kernel generator 130 to generate weight kernel 132 from weights 128 calculated for LUTs 12, 7, 8, and 13. Weight kernel 132 includes interpolated pixel coefficients 1412 configured to blend / combine luminance values of pixels in pixel patch 300 representing luminance radiation from target cluster 1 and neighboring clusters 2, 3, 4, and 5 to maximize the signal-to-noise ratio. The signal that is maximized in the signal-to-noise ratio is the luminance radiation from the target cluster, and the noise that is minimized in the signal-to-noise ratio is the luminance radiation from neighboring clusters 2, 3, 4, and 5, i.e., spatial crosstalk, plus some random noise (e.g., to account for background luminance radiation). The interpolated pixel coefficients 1412 are used as weights, and the mixing / combining involves performing element-wise multiplications between the LUT coefficients and the luminance values of the pixels in the pixel patch 300 to calculate a weighted sum of the luminance values of the pixels.
[0144] FIG. 12 shows the element-wise multiplier 134, which element-wise multiplies the interpolated pixel coefficients 1412 of the weight kernel 132 by the intensity values of the pixels in the pixel patch 300 and sums the intermediate products 1202 of the multiplication to generate the weighted sum 136. For each well, the optical system operates on a point source (cluster intensity at the well) with a point spread function (the response of the optical system). In some implementations, a bias is added to the operation to account for noise caused by different cluster sizes, different background intensities, varying stimulus responses, varying focus, varying sensor sensitivity, and varying lens aberrations. The captured image is a superposition of responses from all wells. The selected LUT equalizes the system response around each well to estimate the intensity of the point source from that well. That is, it processes the PSF intensity over a local neighborhood / grid of sensor pixels to estimate the intensity of the point source that generated the local grid of sensor pixels. This equalizer operation is a dot product on the sensor pixels in the local grid with the equalizer coefficients.
[0145] The equalizer 104 implements the following equalization logic when the target cluster is not substantially concentric with the center of the central pixel, as described above with respect to Figures 9, 10, 11, and 12: When the well is not at the center of the pixel, the output of the equalizer 104 is calculated as a function of a virtual pixel intensity p'(i,j) derived from the actual pixel intensity of the pixel in the sequencing image. (1)y m,n =Σ i,j p'(m+i,n+j).w(i,j)
[0146] In the above equation, the well center coordinates (m, n) may have a fractional part. Each "virtual" equalizer input p'(i, j) is generated by applying an interpolation filter to a pixel neighborhood. In one implementation, a windowed sinc low-pass filter h(x, y) is used for the interpolation. In other implementations, other filters, such as a bilinear interpolation filter, may be used.
[0147] The virtual pixel at location (i,j) is calculated using an interpolation filter as follows: (2) p'(i,j) = Σ u,v p(u,v).h(iu,jv)
[0148] By combining equations (1) and (2), the equalizer 104 uses only the raw pixel intensities as follows:
[0149]
number
[0150] where h is fixed given the sub-pixel offsets frac(m), frac(n), u, v specify the range of pixels used for interpolation to generate the equalizer input, and i, j specify the range of virtual pixels used as input to the equalizer 104.
[0151] For a given sub-pixel offset, only the input pixel changes, not the filter or weights. Therefore, for the center of each binned sub-pixel offset, we compute a fixed set of interpolated equalizer coefficients. The output is:
[0152]
number
[0153] In the above equation, h fm,fn represents the LUT equalizer coefficients for a well with binned fractional subpixel offsets fm, fn, where (fm,fn) are the LUT indices.
[0154] 15A and 15B show how the weight kernel interpolated pixel coefficients 1412 maximize the signal-to-noise ratio and restore the underlying signal of target cluster 1 from the signal corrupted by crosstalk from clusters 2, 3, 4, and 5.
[0155] The weighted sum 136 is provided as an input to a base caller 138 to generate a base call 140. The base caller 138 can be a non-neural network based base caller or a neural network based base caller, examples of both of which are described in applications incorporated herein by reference, such as U.S. Patent Application Nos. 62 / 821,766 and 16 / 826,168.
[0156] In yet other implementations, the need for interpolation is eliminated by having large LUTs, each with a large number of subpixel bins (eg, 50, 75, 100, 150, 200, 300, etc. subpixel bins per LUT).
[0157] Figure 19A shows a graph depicting base calling error rates using images from a NovaSeq sequencer. Error rates are shown in cycles on the x-axis. 0.004 on the y-axis represents a base calling error rate of 0.4%. The error rates here are calculated after mapping and aligning the reads to the Phi-X reference, which is a highly reliable ground truth set. The blue line is the legacy base caller. The red line is the improved equalizer-based base caller 104 disclosed herein. The overall error rate is reduced by 57% at the expense of limited extra computation. The base error rate in later cycles is higher due to extra noise in the system (e.g., pre-phasing / phase alignment, cluster attenuation). The improved performance in later cycles is significant because it indicates the ability to support longer reads. Cycle-to-cycle performance variability is also significantly reduced.
[0158] 19B shows another example of performance results of the disclosed equalizer-based base caller 104 on sequencing data from a NovaSeq sequencer and a Vega sequencer. For the NovaSeq sequencer, the disclosed equalizer-based base caller 104 reduces the base calling error rate by more than 50%. For the Vega sequencer, the disclosed equalizer-based base caller 104 reduces the base calling error rate by more than 35%.
[0159] 19C shows another example of performance results of the disclosed equalizer-based base caller 104 on sequencing data from a NextSeq2000 sequencer. For the NextSeq2000 sequencer, the disclosed equalizer-based base caller 104 reduces the base calling error rate by an average of 10%, not including throughput.
[0160] 19D shows one implementation of the computational resources required by the disclosed equalizer-based base caller 104. As shown, the disclosed equalizer-based base caller 104 can be executed using a small number of CPU threads, ranging from 2 to 7 threads. Thus, the disclosed equalizer-based base caller 104 is a computationally efficient base caller that significantly reduces the base error rate and, therefore, can be integrated into most existing sequencers without requiring any additional computation or specialized processors such as GPUs, FPGAs, ASICs, etc.
[0161] Dependence of base calling on flow cell tilt - additional implementation details In some implementations, the imager is capable of determining the orientation of the surface plane of the sample being imaged, such as in the X-axis (sometimes referred to as tip), Y-axis (sometimes referred to as tilt), and / or Z-axis (sometimes referred to as twist). In some implementations, the imager is enabled to use the determined orientation, such as in combination with elements associated with holding and / or moving the flow cell relative to the imager, to reduce portions of the image that would otherwise be out of focus. Reducing out-of-focus portions is achieved, for example, by controlling the X-axis orientation, the Y-axis orientation, and / or the Z-axis orientation to increase portions of the image that are within the DoF of the imager. Orientation control is achieved, for example, by one or more actuators and / or motor drivers, according to various implementations. Further description is found in U.S. Provisional Patent Application No. 63 / 300,531, entitled "Dynamic Detilt Focus Tracking," filed January 18, 2022 (Attorney Docket No. IP-2205-PRV). Despite the aforementioned reduction of out-of-focus image portions, some implementations provide measurements and / or determinations of tilt, focus, and / or distance that can be used to implement techniques described elsewhere herein to improve base-calling accuracy via flow cell tilt-dependent base-calling.
[0162] Determining tilt, focus, and / or distance Measuring and / or determining tilt, focus, and / or distance (e.g., with respect to a portion of an image being captured) can occur via various techniques, depending on the implementation. Tilt of the flow cell and / or image region can be determined by measuring the distance between projected bright spots, using a grid of resolution features, and / or by intentionally introduced optical aberrations. Multiple tilt values across the flow cell and / or one or more image regions can be determined through the creation and processing of a surface height map. Image region focus can be determined by conjugate lens beam separation. Image region defocus can be measured using a multi-bright spot focus tracker.
[0163] The aforementioned techniques are described in more detail below.
[0164] Tilt determination from bright spot separation measurements In some implementations, tilt is determined by measuring the separation between a pair of bright spots projected onto the sample area being imaged. In some implementations, tilt is determined by measuring a first separation between a first pair of bright spots projected onto the sample image area (e.g., the area being imaged and / or to be imaged) and by measuring a second separation between a second pair of bright spots projected onto the sample area. One or more pairs of bright spots are projected by a light source. In some implementations, the tilt determination is along one dimension, while in some other implementations, the tilt determination is along multiple dimensions, e.g., two dimensions that are substantially perpendicular to each other. In some implementations, the first separation is used to determine a first sample height, the second separation is used to determine a second sample height, and the first and second sample heights are used to determine a corresponding sample tilt. In some implementations, the tilt map is determined from multiple separation measurements of pairs of bright spots from multiple imaging at multiple sample locations. Further description is provided in U.S. Provisional Patent Application No. 63 / 300,531, entitled "Dynamic Detilt Focus Tracking," filed January 18, 2022 (Attorney Docket No. IP-2205-PRV).
[0165] Focus determination from conjugate lens beam separation. In some implementations, the degree of focus is determined by providing a pair of incident light beams to a conjugate lens. The conjugate lens directs the incident light beams toward a focal region. The incident light beams are reflected from the sample image region (e.g., the region being and / or to be imaged). The reflected light beams return to and propagate through the conjugate lens. The relative separation between the reflected light beams is measured and used to determine the degree of focus, working distance, and / or surface profile for the sample based on the relative separation. Further description is found in U.S. Non-Provisional Patent No. 8,422,031 (B2), entitled "Focusing Methods and Optical Systems and Assemblies Using the Same," filed April 16, 2013.
[0166] Grid of resolution features In some implementations, the tilt of a sample is determined by collecting through-focus stacks of images of a grid of resolution features (such as a pinhole array and / or multiple isolated nanowells contained in a flow cell) and analyzing the images to determine tilt. For example, tilt can be measured as an angle by performing multiple through-focus stacks at different X coordinates and comparing the best-focus Z position at each X coordinate. Additionally or alternatively, tilt can be measured as an angle by using an autofocus system to detect the Z position of an element observable by an imaging device, such as a cluster and / or fiducial, at multiple X positions. Further description is found in U.S. Nonprovisional Patent No. 10,830,700(B2), entitled "Solid Inspection Apparatus and Method of Use," filed March 1, 2019.
[0167] Multi-spot focus tracker A multi-spot focus tracker measures defocus at multiple locations in the image plane. The defocus measurements are processed to determine the tilt at the multiple locations.
[0168] Optical aberration Optical aberrations are introduced into the imager optical train (e.g., using a phase mask) so that the point spread function is asymmetric between the upper and lower foci, allowing for easy discrimination between out-of-focus states as upper versus lower foci. The discrimination is processed to determine tilt information.
[0169] Surface Map The height of the flow cell is measured at multiple locations and used to create a surface map, which is processed to determine the slope at multiple locations.
[0170] Criteria and Targets One example of a fiducial is an identifiable reference point in or on an object. For example, the reference point may be present in an image of the object, in a data set derived from detecting the object, or in any other representation of the object suitable for representing information about the reference point relative to the object. The reference point may be specified by an x and / or y coordinate in the plane of the object. Alternatively or additionally, the reference point may be specified by a z coordinate orthogonal to the xy plane, e.g., defined by the relative position of the object and the detector. One or more coordinates for the reference point may be specified relative to one or more other features of the object or an image or other data set derived from the object.
[0171] FIG. 20A shows an example of a fiducial. The top of the figure is a close-up of a single fiducial with four concentric bull's-eye rings. The bottom of the figure is an image of a tile with six exemplary bull's-eye ring fiducials in the image. In various implementations, each dot represents a respective oligocluster, a respective nanowell of a patterned flow cell, or a respective nanowell having one or more oligoclusters therein. In some implementations, the bull's-eye ring fiducial comprises a light ring surrounded by a dark border, such as to enhance contrast. The fiducial can be used as a reference point for aligning an imaged tile with other images of the same tile (e.g., at various wavelengths). For example, the location of the fiducial in the image is determined via cross-correlation with the location of a reference virtual fiducial, determining the location as where the cross-correlation score is maximized. In some implementations, the cross-correlation is performed using a cross-correlation equation for a discrete function (see, e.g., FIG. 20C).
[0172] FIG. 20B shows an exemplary fiducial in various focus contexts. The exemplary fiducial is constructed in the form of a plus sign using a selective chrome layer so that areas with chrome appear dark and areas without chrome appear white. In some contexts, a fiducial implemented using chrome is called an "Uber Target," a chrome target, or simply a target. As shown from top to bottom, a camera (e.g., of an imager) is focused above the chrome layer, on the chrome layer, and below the chrome layer. When focused on the chrome layer, the edges of the chrome appear sharp. When focused above or below the chrome layer, the edges of the chrome appear blurry. In some implementations, the chrome target can be used to perform focus characterization.
[0173] The fiducials (as shown in FIG. 20A and / or FIG. 20B ) can be used as reference image data (e.g., ground truth image data) according to various implementations as described elsewhere herein. In some implementations, a measure of fit between the fiducial in the image and the virtual fiducial is calculated using a scoring equation (see, e.g., FIG. 20D ). In various implementations, various image registration operations use information based on evaluating one or more cross-correlation equations (e.g., as shown in FIG. 20C ) and / or one or more scoring equations (e.g., as shown in FIG. 20D ). In various implementations, various criterion loss functions use information based on evaluating one or more cross-correlation equations (e.g., as shown in FIG. 20C ) and / or one or more scoring equations (e.g., as shown in FIG. 20D ). In various implementations, various criterion quality assessments use information based on evaluating one or more cross-correlation equations (e.g., as shown in FIG. 20C ) and / or one or more scoring equations (e.g., as shown in FIG. 20D ).
[0174] For the example of the dependency of base calling on flow cell tilt in FIG. 1AA, consider various use examples of the criteria in FIG. 20B. As a first example, the camera is focused on the chrome layer, as in the top of FIG. 20B. The image is blurred. In the training context of FIG. 1AA, in response to determining that the image (or a portion thereof) is in superfocus, the "+base caller" is selected for training. The "+GT" element is used to train the "+LUT" element in the superfocus context. Similarly, in the generation context of FIG. 1AA, in response to determining that the image (or a portion thereof) is in superfocus, the "+base caller" is selected for generation. The "+LUTs" element is used to perform base calling in the superfocus context.
[0175] As a second example, the camera is focused on the chrome layer, as in the center portion of FIG. 20B. The image is sharp. In the training context of FIG. 1AA, in response to determining that the image (or portion thereof) is in focus, a "=base caller" is selected for training. The "=GT" element is used to train the "=LUTs" element in the focus context. Similarly, in the generation context of FIG. 1AA, in response to determining that the image (or portion thereof) is in focus, a "=base caller" is selected for generation. The "=LUT" element is used to perform base calling in the focus context.
[0176] As a third example, the camera is focused below the chrome layer, as in the bottom of FIG. 20B. The image is blurred. In the training context of FIG. 1AA, in response to determining that the image (or a portion thereof) is in down focus, the "-base caller" is selected for training. The "-GT" element is used to train the "-LUT" element in the down focus context. Similarly, in the generation context of FIG. 1AA, in response to determining that the image (or a portion thereof) is in down focus, the "-base caller" is selected for generation. The "-LUT" element is used to perform base calling in the down focus context.
[0177] 20C shows an example cross-correlation equation for a discrete function, which can be used, for example, to determine the location of a fiducial (see, e.g., FIG. 20A) using an example scoring equation (see, e.g., FIG. 20D).
[0178] 20D shows an example scoring equation, where Minimum_CC is the maximum value of the cross-correlation, Maximum_CC is the maximum value of the cross-correlation, and RunnerUp_CC is the maximum cross-correlation value outside a radius of, for example, 4 pixels from the location of Maximum_CC.
[0179] Additional Technology In various implementations, one or more techniques for base calling dependency on flow cell tilt are used in a system directed to a self-learning base caller. For various examples, neural network-based and / or non-neural network-based base callers are trained and used for base calling using one or more techniques for base calling dependency on flow cell tilt. Further description is found in U.S. Provisional Patent Application No. 63 / 228,954, entitled "Self-Learned Base Caller," filed August 3, 2021.
[0180] In various implementations, one or more sharpening masks are trained according to flow cell tilt dependency. For example, training (and subsequent use) of sharpening masks, as described in U.S. Nonprovisional Patent Application No. 17 / 511,483 (Attorney Docket No. ILLM 1053-1 / IP-2214-US), filed October 26, 2021, entitled "Intensity Extraction with Interpolation and Adaptation for Base Calling," is adapted to train and use a set of sharpening masks for each of multiple focus contexts. A first set of sharpening masks is trained using image information and / or ground truth information associated with super-focus imaging. A second set of sharpening masks is trained using image information and / or ground truth information associated with in-focus imaging. A third set of sharpening masks is trained using image information and / or ground truth information associated with sub-focus imaging. Following training of the three sets of sharpening masks, the masks are used during base calling. During base calling, a first set of sharpening masks is used to sharpen images determined to be in focus, a second set of sharpening masks is used to sharpen images determined to be in focus, and a third set of sharpening masks is used to sharpen images determined to be in focus.
[0181] Figure 21 shows an overview of an RTA pipeline implementation. Images are collected in two channels (e.g., corresponding to a first wavelength for image 1 and a second wavelength for image 2). Processing of the images begins with alignment as shown and progresses through fitting one or more Gaussians to determine the most likely base call.
[0182] In some implementations, processing associated with image sharpening (e.g., using Laplacian masks) is adapted to use various techniques as described herein with respect to flow cell tilt dependency. For example, one or more Laplacian masks are associated with a first tilt state (e.g., substantially in focus), and one or more other Laplacian masks are associated with a second tilt state (e.g., substantially out of focus). Alternatively, the first one or more Laplacian masks are associated with an up-focus tilt condition, the second one or more Laplacian masks are associated with a focus tilt condition, and the third one or more Laplacian masks are associated with a down-focus tilt condition.
[0183] In some implementations, processing associated with spatially normalizing subtile intensities is adapted to use various techniques as described herein with respect to flow cell tilt dependency. For example, a first spatial normalization of the subtiles is selected for use with image regions determined to be in focus, a second spatial normalization of the subtiles is selected for use with image regions determined to be in focus, and a third spatial normalization of the subtiles is selected for use with image regions determined to be in focus. In some implementations, the spatial normalization is trained according to the flow cell tilt, such as training in the context of images in focus, focus, and down-focus, respectively. In some implementations, the spatial normalization uses techniques such as an equalizer, as shown in FIG. 1AA, used in the base caller in focus ("+ base caller"), the base caller in focus ("= base caller"), and the base caller in down-focus ("- base caller").
[0184] In some implementations that use expectation maximization during base calling, flow cell tilt dependency is introduced by maintaining separate and independent statistical models for each focus context. For example, an up-focus statistical model is used for processing associated with image regions that are at the upper focus, an in-focus statistical model is used for processing associated with image regions that are in focus, and a down-focus statistical model is used for processing associated with image regions that are at the lower focus. Each statistical model has the same architecture, but is EM-optimized separately and independently with respect to the other statistical models.
[0185] Some implementations are based on one, four, or more focus divisions rather than two (e.g., in focus / sharp vs. out of focus / blur) or three (e.g., upper focus, in focus, and lower focus) focus divisions. For example, an implementation is based on five focus divisions: heavily upper focus, slightly upper focus, in focus, slightly lower focus, and heavily lower focus. Other implementations are based on other numbers of focus divisions.
[0186] Some implementations are not based on focus segments, but on gradients, such as gradient magnitude, gradient direction, or gradient magnitude and direction. For example, consider an implementation similar to that shown in FIGS. 1AA and 1AB. Rather than evaluating gradients to determine up-focus, focus, and down-focus, the gradient evaluation determines "uphill," "even," and "downhill," which correspond to gradient vectors that are upward, horizontal, and downward, respectively, relative to the scan direction. The corresponding GTs are referenced to train the respective LUTs in the respective base callers. The corresponding base callers (including the respective LUTs) are referenced to perform base calling according to the gradient classification. As another example, consider an implementation similar to that described above but with multiple upward segments (e.g., significantly upward and slightly upward) and multiple downward segments (e.g., significantly downward and slightly downward). As yet another example, consider an implementation similar to that described above, but with multiple segments associated with the direction of tilt (e.g., with equal angular ranges centered around central angles of 0, 90, 180, and 270 degrees relative to the scan direction). As yet another example, consider an implementation that combines multiple upward and downward segments with segments in the direction of tilt.
[0187] Some implementations measure the slope and / or height before starting a sequencing-by-synthesis run. Some implementations selectively measure the slope and / or height multiple times during a sequencing-by-synthesis run. Some implementations measure the slope and / or height once at the beginning of a sequencing-by-synthesis run. Some implementations (e.g., some implementations using an area sensor combined with a flow cell) do not measure the slope or height during a sequencing-by-synthesis run. Some implementations that do not measure the slope or height during a sequencing-by-synthesis run use the techniques described with respect to FIG. 1AA and / or FIG. 1AB.
[0188] Some implementations adjust the focus and / or tilt before starting a sequencing-by-synthesis run. Some implementations selectively adjust the focus and / or tilt multiple times during a sequencing-by-synthesis run. Some implementations adjust the focus and / or tilt once at the beginning of a sequencing-by-synthesis run. Some implementations (e.g., some implementations using an area sensor combined with a flow cell) do not adjust the focus or tilt during a sequencing-by-synthesis run. Some implementations that do not adjust the focus or tilt during a sequencing-by-synthesis run use techniques such as those described with respect to FIG. 1AA and / or FIG. 1AB.
[0189] Some implementations that do not perform focus and tilt adjustments during sequencing-by-synthesis enable improved throughput in some usage scenarios compared to implementations that perform focus and / or tilt adjustments during sequencing-by-synthesis. Consider a first specific example of an implementation that does not perform focus and tilt adjustments during sequencing-by-synthesis. During training, various base callers are selectively trained using images and associated tilt information (e.g., magnitude, direction, or both). The images and associated tilt information are collected without the benefit of focus and tilt adjustments. Training is selective by using the tilt information associated with an image to select which base callers to train on that image. During generation, various (trained) base callers are selectively used to perform base calling from the image and associated tilt information. Base calling is selective by using the tilt information associated with an image to select which base callers will perform base calling on an image.
[0190] Consider a second specific example of an implementation in which neither focus nor tilt adjustment is performed during sequencing-by-synthesis. The second specific example is similar to the first specific example. However, rather than selecting a base caller based on tilt information, the tilt information is used directly as a parameter for training and generating one or more base callers. The tilt information used can vary depending on the implementation, including the magnitude of the tilt, the direction of the tilt, or both.
[0191] In some implementations according to certain examples above, as well as in some implementations according to the techniques described with respect to Figures 1AA and / or 1AB, the base caller uses one or more artificial intelligence (AI) techniques, and the training is within the context of at least some of the AI techniques. Some implementations that use one or more AI techniques are referred to as "deepRTA" implementations. Further information about DeepRTA can be found in U.S. Patent Application Nos. 16 / 825,987, 16 / 825,991, 16 / 826,126, 16 / 826,134, 16 / 826,168, 62 / 979,412, 62 / 979,411, 17 / 179,395, 62 / 979,399, 17 / 180,480, 17 / 180,513, 62 / 979,414, 62 / 979,385, and 63 / 072,032.
[0192] In some implementations according to certain examples above, as well as in some implementations according to the techniques described with respect to Figures 1AA and / or 1AB, the base caller uses techniques other than AI techniques. Some implementations that use one or more techniques other than AI techniques are referred to as "RTA" implementations.
[0193] In some implementations, the slope is measured in situ during imaging of the sample's fluorescence or for one or more sequencing-by-synthesis cycles of a sequencing-by-synthesis run, before and / or after imaging of the fluorescence. In some situations (e.g., when the slope of one or more regions of the flow cell is relatively constant across the tile and / or relatively constant over time), the slope is measured once per sequencing-by-synthesis cycle or over one sequencing-by-synthesis run. In some situations (e.g., when the slope of one or more regions of the flow cell changes over time due to thermal fluctuations, etc.), the slope is measured more than once per sequencing-by-synthesis cycle or over two or more sequencing-by-synthesis cycles.
[0194] Technical Improvements and Terminology In this application, the terms "cluster," "well," "sample," and "fluorescence sample" are used interchangeably, as a well contains a corresponding cluster / sample / fluorescence sample. As defined herein, "sample" and its derivatives are used in the broadest sense and include any specimen, culture, etc. suspected of containing a target. In some implementations, a sample contains DNA, RNA, PNA, LNA, chimeric, or hybrid forms of nucleic acid. A sample can include any biological, clinical, surgical, agricultural, air, or water sample containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, fresh-frozen, or formalin-fixed, paraffin-embedded nucleic acid sample. It is also contemplated that the sample may be derived from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, (matched) nucleic acid samples from a single individual such as a tumor sample and a normal tissue sample, or a sample from a single source containing two different forms of genetic material such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample containing plant or animal DNA. In some implementations, the source of nucleic acid material may include, for example, nucleic acids obtained from a newborn as typically used for newborn screening.
[0195] A nucleic acid sample may include high molecular weight material such as genomic DNA (gDNA). A sample may include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, the low molecular weight material includes enzymatically or mechanically fragmented DNA. A sample may include cell-free circulating DNA. In some implementations, a sample may include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resection, and other clinical or laboratory samples. In some implementations, a sample may be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, a sample may include nucleic acid molecules obtained from animals, such as humans or mammalian sources. In other implementations, a sample may include nucleic acid molecules obtained from non-mammalian sources, such as plants, bacteria, viruses, or fungi. In some implementations, the source of the nucleic acid molecules may be a preserved or extinct sample or species.
[0196] Additionally, the methods and compositions disclosed herein may be useful for amplifying nucleic acid samples with low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one implementation, a forensic sample may include nucleic acids obtained from a crime scene, from a missing persons DNA database, from a laboratory associated with a forensic investigation, or obtained by a law enforcement agency, one or more military forces, or similar personnel. A nucleic acid sample may be crude DNA, including purified samples or lysates, derived from, for example, a buccal swab, paper, cloth, or other substrate that may be impregnated with saliva, blood, or other bodily fluids. Thus, in some implementations, a nucleic acid sample may contain small or fragmented portions of DNA, such as genomic DNA. In some implementations, target sequences may be present in one or more bodily fluids, including, but not limited to, blood, sputum, plasma, semen, urine, and serum. In some implementations, target sequences may be obtained from hair, skin, tissue samples, autopsies, or the remains of a victim. In some implementations, nucleic acids containing one or more target sequences may be obtained from a deceased animal or human. In some implementations, the target sequence may include nucleic acids obtained from non-human DNA, such as microbial, plant, or entomological DNA. In some implementations, the target sequence or amplified target sequence is intended for human identification purposes. In some implementations, the present disclosure generally relates to methods for identifying features of forensic samples. In some implementations, the present disclosure generally relates to human identification methods using one or more target-specific primers disclosed herein or one or more target-specific primers designed using the primer design criteria outlined herein. In one implementation, a forensic sample or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.
[0197] As used herein, the term "adjacent," when used in reference to two reaction sites, means that there are no other reaction sites between the two reaction sites. The term "adjacent" can have a similar meaning when used in reference to adjacent detection paths and adjacent photodetectors (e.g., adjacent photodetectors have no other photodetectors between them). In some cases, a reaction site may not be adjacent to another reaction site but may still be in close proximity to the other reaction site. A first reaction site may be in close proximity to a second reaction site if a fluorescent emission signal from the first reaction site is detected by a photodetector associated with the second reaction site. More specifically, a first reaction site may be in close proximity to a second reaction site if the photodetector associated with the second reaction site detects, for example, crosstalk from the first reaction site. Adjacent reaction sites may be contiguous so as to abut each other, or adjacent sites may be non-contiguous with an intervening space between them.
[0198] All literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and web pages, regardless of the form such literature and similar materials are in, are expressly incorporated by reference in their entirety. In the event that one or more of the incorporated literature and similar materials, including but not limited to defined terms, term usage, described technology, etc., differs from or contradicts this application, this application controls. Further information regarding the term can be found in U.S. Non-Provisional Patent Application No. 16 / 826,168, filed March 21, 2020 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV), entitled "Artificial Intelligence-Based Sequencing," and U.S. Provisional Patent Application No. 62 / 821,766, filed March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV), entitled "Artificial Intelligence-Based Sequencing."
[0199] The disclosed technology uses neural networks to improve the quality and quantity of nucleic acid sequence information that can be obtained from a nucleic acid template or its complement, e.g., a nucleic acid sample, such as a DNA or RNA polynucleotide or other nucleic acid sample. Accordingly, certain implementations of the disclosed technology provide higher throughput polynucleotide sequencing, e.g., higher rates of collection of DNA or RNA sequence data, greater efficiency in sequence data collection, and / or obtaining such sequence data at lower cost, compared to previously available methods.
[0200] The disclosed technology uses neural networks to identify the centers of solid-phase nucleic acid clusters and analyze optical signals generated during sequencing of such clusters to unambiguously distinguish between adjacent, abutting, or overlapping clusters and assign sequencing signals to single, discrete source clusters. These and related implementations thus enable the acquisition of meaningful information, such as sequence data, from regions of high-density cluster arrays where useful information was previously unavailable due to the confusing effects of overlapping or closely spaced adjacent clusters, including the effects of overlapping signals (e.g., as used in nucleic acid sequencing) emanating from such regions.
[0201] As described in more detail below, in certain implementations, compositions are provided that include a solid support immobilized with one or more nucleic acid clusters, as provided herein. Each cluster contains multiple immobilized nucleic acids of the same sequence and has a distinct center with a detectable central label as provided herein, which is distinguishable from the nucleic acids immobilized in the surrounding area within the cluster. Also described herein are methods for producing and using such clusters with distinct centers.
[0202] Implementations of the present disclosure find use in numerous situations in which benefits derive from the ability to identify, determine, annotate, record, or otherwise assign the location of a substantially central location within a cluster, such as high-throughput nucleic acid sequencing, development of image analysis algorithms for assigning optical or other signals to individual source clusters, and other applications in which recognition of the center of a fixed nucleic acid cluster is desirable and beneficial.
[0203] In certain implementations, the present invention contemplates methods related to high-throughput nucleic acid analysis, such as nucleic acid sequencing (e.g., "sequencing"). Exemplary high-throughput nucleic acid analyses include, but are not limited to, de novo sequencing, resequencing, whole-genome sequencing, gene expression analysis, gene expression monitoring, epigenetics analysis, genome methylation analysis, allele-specific primer extension (APSE), genetic diversity profiling, whole-genome polymorphism discovery and analysis, single-nucleotide polymorphism analysis, hybridization-based sequencing, and the like. Those skilled in the art will understand that a variety of different nucleic acids can be analyzed using the methods and compositions of the present invention.
[0204] Although implementations of the present invention are described in the context of nucleic acid sequencing, they are applicable in any field in which image data acquired at different times, spatial locations, or other temporal or physical perspectives are analyzed. For example, the methods and systems described herein are useful in the fields of molecular biology and cell biology, where image data from microarrays, biological specimens, cells, organisms, etc., are acquired and analyzed at different times or perspectives. Images can be obtained using any number of techniques known in the art, including, but not limited to, fluorescence microscopy, optical microscopy, confocal microscopy, optical imaging, magnetic resonance imaging, tomography scanning, etc. As another example, the methods and systems described herein can be applied when image data acquired by surveillance, aerial, or satellite imaging techniques, etc., are acquired and analyzed at different times or perspectives. The methods and systems are particularly useful for analyzing images acquired within a field of view, where specimens seen within the field of view remain in the same location relative to each other within the field of view. However, specimens may have different characteristics in separate images; for example, specimens may appear differently in separate images of the field of view. For example, analytes may appear different in terms of the color of a given analyte detected in different images, changes in the brightness of the signal detected for a given analyte in different images, or even the appearance of a signal for a given analyte in one image and the disappearance of the signal for that analyte in another image.
[0205] As used herein, the term "analyte" is intended to mean a point or region of a pattern that can be distinguished from other points or regions according to their relative position. An individual analyte can include one or more molecules of a particular type. For example, an analyte can include a single target nucleic acid molecule having a particular sequence, or an analyte can include several nucleic acid molecules having the same sequence (and / or its complementary sequence). Different molecules that are different analytes of a pattern can be distinguished from one another according to the location of the analyte within the pattern. Exemplary analytes include wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, pads of gel material on a substrate, or channels in a substrate.
[0206] Any of a variety of target analytes to be detected, characterized, or identified may be used in the devices, systems, or methods described herein. Exemplary analytes include, but are not limited to, nucleic acids (e.g., DNA, RNA, or analogs thereof), proteins, polysaccharides, cells, antibodies, epitopes, receptors, ligands, enzymes (e.g., kinases, phosphatases, or polymerases), small molecule drug candidates, cells, viruses, organisms, etc.
[0207] The terms "analyte," "nucleic acid," "nucleic acid molecule," and "polynucleotide" are used interchangeably herein. In various implementations, a nucleic acid may be used as a template (e.g., a nucleic acid template or a nucleic acid complement complementary to a nucleic acid template) as provided herein for a particular type of nucleic acid analysis, including, but not limited to, nucleic acid amplification, nucleic acid expression analysis, and / or nucleic acid sequencing, or a suitable combination thereof. In certain implementations, nucleic acids include, for example, linear polymers of deoxyribonucleotides in 3'-5' phosphodiester or other linkages such as deoxyribonucleic acid (DNA), e.g., single- and double-stranded DNA, genomic DNA, copy or complementary DNA (cDNA), recombinant DNA, or any form of synthetic or modified DNA. In other implementations, nucleic acids include, for example, linear polymers of ribonucleotides in 3'-5' phosphodiester or other linkages, such as ribonucleic acid (RNA), including single- and double-stranded RNA, messenger (mRNA), copy RNA, or complementary RNA (cRNA), alternatively spliced mRNA, ribosomal RNA, small nuclear RNA (snoRNA), microRNA (miRNA), small interfering RNA (sRNA), piwi RNA (piRNA), or any form of synthetic or modified RNA. Nucleic acids used in the compositions and methods of the invention may vary in length and may be intact or full-length molecules or fragments, or smaller portions of larger nucleic acid molecules. In certain implementations, nucleic acids may carry one or more detectable labels, as described elsewhere herein.
[0208] The terms "analyte," "cluster," "nucleic acid cluster," "nucleic acid colony," and "DNA cluster" are used interchangeably and refer to multiple copies of a nucleic acid template and / or its complement attached to a solid support. In some implementations, a nucleic acid cluster comprises multiple copies of a template nucleic acid and / or its complement attached to a solid support via their 5' ends. The copies of the nucleic acid strands that make up a nucleic acid cluster may be in single-stranded or double-stranded form. The copies of the nucleic acid template present within a cluster may have nucleotides at corresponding positions that differ from each other due to, for example, the presence of a labeling moiety. The corresponding positions may also contain analog structures that have different chemical structures but similar Watson-Crick base pairing properties, such as uracil and thymine.
[0209] Colonies of nucleic acids may also be referred to as "nucleic acid clusters." Nucleic acid colonies may optionally be generated by cluster amplification or bridge amplification techniques, as described in more detail elsewhere herein. Multiple repeats of a target sequence may be present in a single nucleic acid molecule, such as concatemers generated using rolling circle amplification procedures.
[0210] The nucleic acid clusters of the present invention can have different shapes, sizes, and densities depending on the conditions used. For example, the clusters can be substantially circular, multifaceted, donut-shaped, or ring-shaped. The diameter of the nucleic acid clusters can be designed to be about 0.2 μm to about 6 μm, about 0.3 μm to about 4 μm, about 0.4 μm to about 3 μm, about 0.5 μm to about 2 μm, about 0.75 μm to about 1.5 μm, or any intermediate diameter. In certain implementations, the diameter of the nucleic acid clusters is about 0.5 μm, about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The diameter of the nucleic acid clusters can be affected by numerous parameters, including, but not limited to, the number of amplification cycles performed in producing the clusters, the length of the nucleic acid template, or the density of primers attached to the surface on which the clusters are formed. The density of the nucleic acid clusters is typically less than 0.1 μm / mm 2 , 1 / mm2 , 10 / mm 2 , 100 / mm 2 , 1,000 / mm 2 , 10,000 / mm 2 ~100,000 / mm 2 The present invention, in part, is directed to higher density nucleic acid clusters, e.g., 100,000 / mm 2 ~1,000,000 / mm 2 , and 1,000,000 / mm 2 ~10,000,000 / mm 2 Further plans are underway.
[0211] As used herein, an "analyte" is a sample or region of interest within a field of view. When used in connection with a microarray device or other molecular analysis device, an analyte refers to a region occupied by similar or identical molecules. For example, an analyte can be an amplification oligonucleotide or any other group of polynucleotides or polypeptides having the same or similar sequence. In other implementations, an analyte can be any element or group of elements that occupy a physical area on a sample. For example, an analyte can be a piece of land, a body of water, etc. When analytes are imaged, each analyte has some area. Thus, in many implementations, an analyte is not simply a pixel.
[0212] The distance between specimens may be described in any number of ways. In some implementations, the distance between specimens may be described from the center of one specimen to the center of another specimen. In other implementations, the distance may be described from the edge of one specimen to the edge of another specimen, or between the outermost identifiable points of each specimen. The edge of the specimen may be described as a theoretical or actual physical boundary on the chip, or some point within the boundary of the specimen. In other implementations, the distance may be described with respect to a fixed point on the specimen, or an image of the specimen.
[0213] Generally, several implementations are described herein with respect to analytical methods. It will be understood that systems for performing the methods in an automated or semi-automated manner are also provided. Thus, the present disclosure provides a neural network-based template generation and base calling system, which can include a processor, a storage device, and a program for image analysis, the program including instructions for performing one or more of the methods described herein. Thus, the methods described herein can be performed, for example, on a computer having components described herein or known in the art.
[0214] The methods and systems described herein are useful for analyzing any of a variety of objects. Particularly useful objects are solid supports or solid surfaces with analytes attached. The methods and systems described herein offer advantages when used with objects having repeating patterns of analytes in the xy plane. One example is a microarray with attached collections of cells, viruses, nucleic acids, proteins, antibodies, carbohydrates, small molecules (such as drug candidates), biologically active molecules, or other analytes of interest.
[0215] There has been an increase in the number of applications of arrays containing analytes containing biological molecules such as nucleic acids and polypeptides. Such microarrays typically contain deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) probes, which are specific for nucleotide sequences present in humans and other organisms. In certain applications, for example, individual DNA or RNA probes can be attached to individual analytes on the array. A test sample, such as one from a known human or organism, can be exposed to the array so that target nucleic acids (e.g., gene fragments, mRNA, or amplicons thereof) hybridize to complementary probes in each analyte in the array. The probes can be labeled for target-specific processes (e.g., due to labels present on the target nucleic acids or due to enzyme labels of the probes or targets present in hybridized form in the analyte). The array can then be examined by scanning specific light frequencies over the analytes to identify which target nucleic acids are present in the sample.
[0216] Biological microarrays can be used for gene sequencing and similar applications. Generally, gene sequencing involves determining the order of nucleotides in a length of target nucleic acid, such as a fragment of DNA or RNA. Relatively short sequences are typically sequenced for each specimen, and the resulting sequence information can be used in various bioinformatics methods to logically match sequence fragments together to reliably determine the sequence of a much larger and wider range of genetic material from which the fragments are derived. Automated computer-based algorithms for identifying characteristic fragments have been developed and have been more recently used in genome mapping, gene identification, and their functions. Microarrays are particularly useful for characterizing genome content due to the large number of variants present, which replaces the need to perform numerous experiments on individual probes and targets. Microarrays are an ideal format for practically conducting such studies.
[0217] Any of a variety of analyte arrays (also referred to as "microarrays") known in the art can be used in the methods or systems described herein. A typical array includes analytes, each having an individual probe or a population of probes. In the latter case, the population of probes in each analyte is typically homogeneous, with a single type of probe. For example, in the case of a nucleic acid array, each analyte may have multiple nucleic acid molecules, each having a common sequence. However, in some implementations, the population in each analyte of the array may be heterogeneous. Similarly, a protein array may have analytes having a single protein or a population of proteins, typically, but not necessarily, having the same amino acid sequence. Probes can be attached to the surface of the array, for example, by covalently binding the probe to the surface or through non-covalent interactions between the probe and the surface. In some implementations, probes, such as nucleic acid molecules, can be attached to the surface via a gel layer, as described, for example, in U.S. Patent Application No. 13 / 784,368 and U.S. Patent Application Publication No. 2011 / 0059865(A1), each of which is incorporated herein by reference.
[0218] Exemplary arrays include, but are not limited to, BeadChip Arrays available from Illumina, Inc. (San Diego, Calif.), or others, such as those in which probes are attached to beads present on a surface (e.g., beads in wells on a surface), such as those described in U.S. Patent Nos. 6,266,459, 6,355,431, 6,770,441, 6,859,570, or 7,622,294, or International Publication No. PCT / WO00 / 63437, each of which is incorporated herein by reference. Further examples of commercially available microarrays that can be used include, for example, Affymetrix® GeneChip® microarrays or other microarrays synthesized according to a technology sometimes referred to as VLSIPS™ (Very Large Scale Immobilized Polymer Synthesis) technology. Spotted microarrays can also be used in methods or systems according to some implementations of the present disclosure. An exemplary spotted microarray is the CodeLink™ Array available from Amersham Biosciences. Another useful microarray is one manufactured using inkjet printing methods, such as SurePrint™ Technology available from Agilent Technologies.
[0219] Other useful arrays include those used in nucleic acid sequencing applications.For example, the arrays with amplicons of genome fragments (often referred to as clusters) are particularly useful, such as those described in Bentley et al., Nature 456:53-59 (2008), WO 4 / 018497, WO 91 / 06678, WO 7 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,19, U.S. Patent No. 7,405,281 or U.S. Patent No. 7,057,26, U.S. Patent Application Publication No. 2008 / 0108082 (A1), each of which is incorporated herein by reference.Another type of array that is useful for nucleic acid sequencing is the array of particles that are produced by emulsion PCR technology. Examples are described in Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), WO 5 / 010145, U.S. Patent Application Publication No. 2005 / 0130173, or U.S. Patent Application Publication No. 2005 / 0064460, each of which is incorporated herein by reference in its entirety.
[0220] Arrays used for nucleic acid sequencing often have a random spatial pattern of nucleic acid analytes. For example, the HiSeq or MiSeq sequencing platforms available from Illumina Inc. (San Diego, Calif.) utilize flow cells in which nucleic acid arrays are formed by random seeding followed by bridge amplification. However, patterned arrays can also be used for nucleic acid sequencing or other analytical applications. Examples of patterned arrays, their fabrication methods, and their use are described in U.S. Patent Nos. 13 / 787,396, 13 / 783,43, 13 / 784,368, U.S. Patent Application Publication Nos. 2013 / 0116153 (A1), and 2012 / 0316086 (A1), each of which is incorporated herein by reference. Such patterned array analytes can be used to capture single nucleic acid template molecules and seed the subsequent formation of homogeneous colonies, for example, via bridge amplification. Such patterned arrays are particularly useful for nucleic acid sequencing applications.
[0221] The size of the specimens on an array (or other object used in a method or system herein) can be selected to suit a particular application. For example, in some implementations, the specimens on the array can have a size that accommodates only a single nucleic acid molecule. A surface with multiple specimens in this size range is useful for constructing an array of molecules for detection with single molecule resolution. Specimens in this size range are also useful for use in arrays with specimens each comprising a colony of nucleic acid molecules. Thus, the specimens on the array can each be approximately 1 mm 2 Below, approximately 500μm 2 Below, approximately 100μm 2 Below, approximately 10μm 2 Below, approximately 1μm 2 Below, about 500nm 2 Less than or equal to about 100 nm 2 Below, approximately 10nm 2 Below, approximately 5nm 2 Less than or equal to 1 nm 2Alternatively or additionally, the specimens in the array may have an area of about 1 mm 2 or more, about 500μm 2 More than approximately 100μm 2 or more, about 10μm 2 or more, approximately 1μm 2 or more, about 500nm 2 Over 100nm 2 or more, about 10nm 2 or more, about 5nm 2 More than or about 1 nm 2 That's it. In practice, analytes can have sizes within a range between upper and lower limits selected from those exemplified above. While some size ranges for surface analytes have been exemplified with respect to nucleic acids and nucleic acid scales, it will be understood that analytes in these size ranges can be used in applications that do not involve nucleic acids. It will further be understood that analyte sizes need not necessarily be limited to the scales used in nucleic acid applications.
[0222] In implementations involving an object having multiple analytes, such as an array of analytes, the analytes may be distinct, separated by a space between them. Arrays useful in the present invention may have analytes separated by an edge-to-edge distance of at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, or less. Alternatively or additionally, arrays may have analytes separated by an edge-to-edge distance of at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges may apply to the average edge-to-edge spacing of the analytes, as well as to minimum or maximum spacing.
[0223] In some implementations, the analytes in the array need not be distinct; instead, adjacent analytes can abut one another. Whether the analytes are distinct or not, the size of the analytes and / or the pitch of the analytes can be varied to allow the array to have a desired density. For example, the average analyte pitch in a regular pattern can be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, or less. Alternatively or additionally, the average analyte pitch in a regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can also apply to the maximum or minimum pitch of a regular pattern. For example, the maximum specimen pitch in the regular pattern can be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm or less, and / or the minimum specimen pitch in the regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm or more.
[0224] The density of analytes in an array can also be understood in terms of the number of analytes present per unit area. For example, the average density of analytes in an array can be at least about 1×10 3 specimens / mm 2 , 1×10 4 specimens / mm 2 , 1×10 5 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 7 specimens / mm 2 , 1×10 8 specimens / mm 2 , or 1×10 9 specimens / mm 2 Alternatively or additionally, the average density of analytes in the array can be at most about 1 x 10 9 specimens / mm 2 , 1×10 8 specimens / mm 2 , 1×10 7 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 5 specimens / mm 2, 1×10 4 specimens / mm 2 , or 1×10 3 specimens / mm 2 It can be the following:
[0225] The above ranges may apply, for example, to all or part of a regular pattern that includes all or part of an array of analytes.
[0226] The analytes within the pattern can have any of a variety of shapes. For example, when viewed in a two-dimensional plane, such as on the surface of an array, the analytes may appear rounded, circular, oval, rectangular, square, symmetrical, asymmetrical, triangular, polygonal, etc. The analytes can be arranged in a regular repeating pattern, including, for example, a hexagonal or rectilinear pattern. The pattern can be selected to achieve a desired level of packing. For example, circular analytes are optimally packed in a hexagonal arrangement. Of course, other packing configurations can also be used for circular analytes, and vice versa.
[0227] A pattern can be characterized in terms of the number of analytes present in a subset that forms the smallest geometric unit of the pattern. A subset can include, for example, at least about 2, 3, 4, 5, 6, 10, or more analytes. Depending on the size and density of the analytes, a geometric unit can be as small as 1 mm. 2 , 500 μm 2 , 100 μm 2 , 50 μm 2 , 10 μm 2 , 1 μm 2 , 500nm 2 , 100 nm 2 , 50nm 2 , 10nm 2 Alternatively or additionally, the geometric unit may be 10 nm 2 , 50nm 2 , 100 nm 2 , 500nm 2 , 1 μm 2 , 10 μm 2 , 50 μm 2 , 100 μm 2 , 500 μm 2 , 1mm2 The properties of the specimen in the geometric unit, such as shape, size, pitch, etc., may be selected from those described herein more generally with respect to specimens in an array or pattern.
[0228] An array with a regular pattern of analytes may be regular in terms of the relative location of the analytes, but random in terms of one or more other characteristics of each analyte.For example, in the case of a nucleic acid array, the nucleic acid analytes may be regular in terms of their relative location, but random in terms of the knowledge of the sequence of the nucleic acid species present in any particular analyte.As a more specific example, a nucleic acid array formed by seeding a repeating pattern of analytes with template nucleic acids and amplifying the template in each analyte to form copies of the template in the analyte (e.g., via cluster amplification or bridge amplification) has a regular pattern of nucleic acid analytes, but is random in terms of the distribution of the sequences of the nucleic acids across the array.Thus, generally, detecting the presence of nucleic acid material on an array can result in a repeating pattern of analytes, while sequence-specific detection can result in a non-repeated distribution of signals across the array.
[0229] It will be understood that descriptions of pattern, regularity, randomness, etc. herein relate not only to analytes on an object, such as analytes on an array, but also to analytes in an image. Thus, the pattern, regularity, randomness, etc. may exist in any of a variety of formats used to store, manipulate, or communicate image data, including, but not limited to, computer-readable media or computer components such as a graphical user interface or other output device.
[0230] As used herein, the term "imaging" is intended to mean a representation of all or a portion of an object. The representation may be an optically detected reproduction. For example, an image may be obtained from fluorescence, luminescence, scattering, or absorption signals. The portion of an object present in an image may be the surface or other xy plane of the object. Typically, an image is a two-dimensional representation, but in some cases, information in an image may be derived from three or more dimensions. An image need not include optically detected signals. Non-optical signals may instead be present. An image may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.
[0231] As used herein, "image" refers to a reproduction or representation of at least a portion of a sample or other object. In some implementations, the reproduction is an optical reproduction, for example, produced by a camera or other optical detector. The reproduction may be a non-optical reproduction, for example, a representation of electrical signals obtained from an array of nanopore analytes or a representation of electrical signals obtained from an ion-sensitive CMOS detector. In certain implementations, non-optical reproductions may be excluded from the methods or apparatus described herein. The image may have a resolution capable of distinguishing sample analytes present at any of a variety of intervals, including, for example, those spaced less than 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm apart.
[0232] As used herein, "acquiring," "acquisition," and like terms refer to any part of the process of acquiring an image file. In some implementations, data acquisition can include generating an image of the sample, looking for a signal in the sample, instructing a detection device to look for or generate an image of the signal, and providing instructions for further analysis or transformation of the image file, and instructions for any number of transformations or manipulations of the image file.
[0233] As used herein, the term "template" refers to a representation of the location or relationship between signals or analytes. Thus, in some implementations, the template is a physical grid with a representation of signals corresponding to analytes in a sample. In some implementations, the template can be a chart, table, text file, or other computer file that indicates locations corresponding to analytes. In the implementations presented herein, the template is generated to track the location of analytes on a sample across a set of images of the sample captured at different reference points. For example, the template can be a set of x, y coordinates or a set of values that describe the orientation and / or distance of one analyte relative to another analyte.
[0234] As used herein, the term "specimen" may refer to an object or region of an object from which an image is captured. For example, in an implementation in which an image is taken from the surface of soil, a patch of land may be a sample. In other implementations in which biomolecular analysis is performed within a flow cell, the flow cell may be divided into any number of subdivisions, each of which may be a sample. For example, a flow cell may be divided into various channels or lanes, and each lane may be further divided into 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 140, 160, 180, 200, 400, 600, 800, 1000, or more separate regions to be imaged. One example flow cell has eight lanes, each divided into 120 samples or tiles. In another implementation, the sample may be made up of multiple tiles, or even the entire flow cell, so that each sample image may represent a larger area of the surface being imaged.
[0235] References to ranges and sequential lists of numbers set forth herein will be understood to include not only the numbers recited but all real numbers between the recited numbers.
[0236] As used herein, a "reference point" refers to any temporal or physical distinction between images. In some implementations, the reference point is a time point. In other implementations, the reference point is a time point or cycle during a sequencing reaction. However, the term "reference point" can include other aspects that distinguish or separate images, such as angle, rotation, time, or other aspects that may distinguish or separate images.
[0237] As used herein, a "subset of images" refers to a group of images within a set. For example, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60, or any number of images selected from the set of images. In certain implementations, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60, or fewer images, or any number selected from the set of images. In some implementations, the images are obtained from one or more sequencing cycles, with four images correlating to each cycle. Thus, for example, a subset may be a group of 16 images acquired over four cycles.
[0238] Base refers to a nucleotide base or nucleotide, A (adenine), C (cytosine), T (thymine), or G (guanine). This application uses "base" and "nucleotide" interchangeably.
[0239] The term "chromosome" refers to the genetic carrier of a living cell, derived from a chromatin strand containing DNA and protein components (especially histones). The conventional internationally recognized system of numbering individual human genome chromosomes is employed herein.
[0240] The term "site" refers to a unique location (e.g., chromosome ID, chromosomal location and orientation) on a reference genome. In some implementations, a site may be a residue, a sequence tag, or the location of a segment on a sequence. The term "locus" may be used to refer to a specific location of a nucleic acid sequence or polymorphism on a reference chromosome.
[0241] The term "sample" as used herein typically refers to a sample derived from a biological fluid, cell, tissue, organ, or organism containing the nucleic acid to be sequenced and / or phased, or a sample derived from a mixture of nucleic acids containing at least one nucleic acid sequence to be sequenced and / or phased. Such samples include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, fine-needle biopsy samples (e.g., surgical biopsies, fine-needle biopsies, etc.), urine, peritoneal fluid, pleural fluid, tissue explants, organ cultures, and any other tissue or cell preparations, fractions, or derivatives thereof or isolated therefrom. While samples are often obtained from human subjects (e.g., patients), samples may be obtained from any organism that has chromosomes, including, but not limited to, dogs, cats, horses, goats, sheep, cows, pigs, etc. Samples may be used directly as obtained from the biological source or after pre-processing to modify the characteristics of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, etc.
[0242] The term "sequence" includes or refers to a chain of nucleotides linked together. The nucleotides may be based on DNA or RNA. It should be understood that a sequence may contain multiple subsequences. For example, a single sequence (e.g., a PCR amplicon) may have 350 nucleotides. A sample read may contain multiple subsequences within these 350 nucleotides. For example, a sample read may include first and second adjacent subsequences, e.g., 20-50 nucleotides. The first and second adjacent subsequences may be located on either side of a repeat segment with corresponding subsequences (e.g., 40-100 nucleotides). Each adjacent subsequence may include a primer subsequence (e.g., 10-30 nucleotides) (or a portion thereof). For ease of interpretation, the term "subsequence" is referred to as "sequence," but it should be understood that two sequences need not be distinct from each other on a common strand. To distinguish between the various sequences described herein, sequences may be labeled differently (e.g., target sequence, primer sequence, flanking sequence, reference sequence, etc.). Other terms, such as "allele," may be given different labels to distinguish between similar entities. Applications use "read" and "sequence read" interchangeably.
[0243] The term "paired-end sequencing" refers to a sequencing method in which both ends of a target fragment are sequenced. Paired-end sequencing can facilitate the detection of genome rearrangements and repeated segments, as well as the detection of gene fusions and novel transcripts. Methods for paired-end sequencing are described in International Publication No. PCT / WO07010252, International Application No. PCT / GB2007 / 003798, and U.S. Patent Application Publication No. 2009 / 0088327, each of which is incorporated herein by reference. In one example, the series of operations can be performed as follows: (a) generating clusters of nucleic acids, (b) linearizing the nucleic acids, (c) hybridizing a first sequencing primer and repeatedly performing cycles of extension, scanning, and deblocking as described above, (d) "flipping" the target nucleic acid on the flow cell surface by synthesizing a complementary copy, (e) linearizing the resynthesized strand, and (f) hybridizing a second sequencing primer and repeatedly performing cycles of extension, scanning, and deblocking as described above. The flipping operation can be performed to deliver the reagents described above for a single cycle of bridge amplification.
[0244] The term "reference genome" or "reference sequence" refers to a specific, known genomic sequence, whether partial or complete, of any organism that can be used to reference an identified sequence from a subject. For example, reference genomes used for human subjects as well as many other organisms can be found at the National Center for Biotechnology Information (ncbi.nlm.nih.gov). "Genome" refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences. A genome includes both genetic and non-coding sequences of DNA. A reference sequence may be larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 10 times larger, or at least about 10 times larger, or at least about 10 times larger. In one example, the reference genome sequence is of the full-length human genome. In another example, the reference genome sequence is limited to a specific human chromosome, such as chromosome 13. In some implementations, the reference chromosome is a chromosome sequence from human genome version hg19. Such sequences are sometimes referred to as chromosomal reference sequences, although the term "reference genome" is intended to encompass such sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes of any species, partial chromosomal regions (such as strands), and the like. In various implementations, a reference genome is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, a reference sequence may be taken from a specific individual. In other implementations, "genome" also covers so-called "graph genomes," which use specific storage formats and representations of genome sequences. In one implementation, a graph genome stores data in a linear file. In another implementation, a graph genome refers to a representation in which alternative sequences (e.g., different copies of a chromosome with small differences) are stored as different paths in a graph.Additional information regarding graph genome implementations can be found at https: / / www.biorxiv.org / content / biorxiv / early / 2018 / 3 / 20 / 194530.full.pdf, the contents of which are incorporated herein by reference in their entirety.
[0245] The term "read" refers to a collection of sequence data describing a fragment of a nucleotide sample or reference. The term "read" can refer to a sample read and / or a reference read. Typically, but not necessarily, a read represents a short sequence of consecutive base pairs in a sample or reference. A read may be symbolically represented by the base pair sequence (ATCG) of the sample or reference fragment. Optionally, the read may be stored in a memory device and processed to determine whether the read matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing device or indirectly from stored sequence information about the sample. In some cases, the read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region that can be aligned and specifically assigned to, for example, a chromosome or genomic region or gene.
[0246] Next-generation sequencing methods include, for example, sequencing-by-synthesis (Illumina), pyrosequencing (454), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), and sequencing by ligation (SOLiD sequencing). Depending on the sequencing method, the length of each read can vary from approximately 30 bp to over 10,000 bp. For example, DNA sequencing using a SOLiD sequencer generates nucleic acid reads of approximately 50 bp. In another example, Ion Torrent sequencing generates nucleic acid reads of up to 400 bp, while 454 pyrosequencing generates nucleic acid reads of approximately 700 bp. In yet another example, single-molecule real-time sequencing methods can generate reads of 10,000 bp to 15,000 bp. Thus, in certain implementations, the nucleic acid sequence reads have a length of 30-100 bp, 50-200 bp, or 50-400 bp.
[0247] The terms "sample read," "sample sequence," or "sample fragment" refer to sequence data relating to a genomic sequence of a subject from a sample. For example, a sample read includes sequence data from a PCR amplicon having forward and reverse primer sequences. The sequence data can be obtained from any selected sequence methodology. The sample read can be, for example, a sequencing-by-synthesis (SBS) reaction, a sequencing-by-ligation reaction, or any other suitable sequencing methodology in which it is desirable to determine the length and / or identity of repetitive elements. The sample read can be a consensus (e.g., average or weighted) sequence derived from multiple sample reads. In certain implementations, providing a reference sequence includes identifying a locus of interest based on primer sequences of a PCR amplicon.
[0248] The term "raw fragment" refers to sequence data of a portion of a subject's genome sequence that at least partially overlaps a specified or secondary position of interest within a sample read or sample fragment. Non-limiting examples of raw fragments include double-stitched fragments, single-stitched fragments, double-unstitched fragments, and single-unstitched fragments. The term "raw" is used to indicate that a raw fragment contains sequence data that has some relationship to sequence data in a sample read, regardless of whether the raw fragment exhibits supporting variants that correspond to and authenticate or confirm potential variants in the sample read. The term "raw fragment" does not indicate that the fragment necessarily contains supporting variants that verify the variant call in the sample read. For example, if a sample read is determined by a variant calling application to exhibit a first variant, the variant calling application may determine that one or more raw fragments lack a corresponding type of "supporting" variant that would otherwise be expected to occur given the variants in the sample read.
[0249] The terms "mapping," "aligned," "alignment," or "aligning" refer to the process of comparing a read or tag to a reference sequence, thereby determining whether the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence, or in certain implementations, may be mapped to a specific location within the reference sequence. In some cases, alignment simply tells whether a read is a member of a particular reference sequence (i.e., whether the read is present or absent in the reference sequence). For example, alignment of a read to the reference sequence of human chromosome 13 tells whether the read is present in the reference sequence of chromosome 13. Tools that provide this information are sometimes called set membership testers. In some cases, alignment also indicates the location within the reference sequence to which the read or tag maps. For example, if the reference sequence is the entire human genome sequence, alignment may indicate that the read is present on chromosome 13, and may further indicate that the read is on a specific strand and / or site of chromosome 13.
[0250] The term "indel" refers to the insertion and / or deletion of bases in an organism's DNA. Microindels refer to indels that result in a net change of 1 to 50 nucleotides. Unless the length of the indel is a multiple of three, a frameshift mutation occurs in coding regions of the genome. Indels can be contrasted with point mutations. Indels insert and delete nucleotides from a sequence, whereas point mutations are a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels can also be contrasted with tandem base mutations (TBMs), which can be defined as substitutions at adjacent nucleotides (mostly substitutions at two adjacent nucleotides are observed, although substitutions at three adjacent nucleotides have also been observed).
[0251] The term "variant" refers to a nucleic acid sequence that differs from a reference nucleic acid. Typical nucleic acid sequence variants include, but are not limited to, single nucleotide polymorphisms (SNPs), short deletion and insertion polymorphisms (indels), copy number variations (CNVs), microsatellite markers, or short tandem repeats and structural variants. Somatic variant calling is an effort to identify variants present at low frequencies in a DNA sample. Somatic variant calling is of interest in the context of cancer treatment. Cancer is caused by the accumulation of mutations in DNA. DNA samples derived from tumors are generally heterogeneous, containing some normal cells, some cells at an early stage of cancer progression (with fewer mutations), and some late-stage cells (with more mutations). Due to this heterogeneity, somatic mutations often appear at low frequencies when sequencing tumors (e.g., from FFPE samples). For example, an SNV may be found in only 10% of reads covering a given base. A variant that is classified as somatic or germline by a variant classifier is also referred to herein as a "variant under test."
[0252] The term "noise" refers to erroneous variant calls that result from one or more errors in the sequencing process and / or variant calling application.
[0253] The term "variant frequency" refers to the relative frequency of an allele (a genetic variant) at a particular locus within a population, expressed as a fraction or percentage. For example, the fraction or percentage may be the fraction of all chromosomes in a population that carry that allele. As one example, a sample variant frequency refers to the relative frequency of an allele / variant at a particular locus / position along a subject's genome sequence across a "population" that corresponds to the number of reads and / or samples obtained for the subject's genome sequence from individuals. As another example, a baseline variant frequency refers to the relative frequency of an allele / variant at a particular locus / position along one or more baseline genome sequences, where a "population" corresponds to the number of reads and / or samples obtained for one or more baseline genome sequences from a population of normal individuals.
[0254] The term "Variant Allele Frequency (VAF)" refers to the proportion of sequenced reads observed to match a variant divided by the overall coverage at the target position. VAF is a measure of the proportion of sequenced reads that carry a variant.
[0255] The terms "position," "designated position," and "locus" refer to the position or coordinates of one or more nucleotides within a nucleotide sequence. The terms "position," "designated position," and "locus" also refer to the position or coordinates of one or more base pairs in a sequence of nucleotides.
[0256] The term "haplotype" refers to a combination of alleles at adjacent sites on a chromosome that are inherited together. A haplotype, if present, may be one locus, several loci, or an entire chromosome, depending on the number of recombination events that have occurred between a given set of loci.
[0257] The term "threshold" herein refers to a numeric or non-numeric value used as a cutoff for characterizing a sample, a nucleic acid, or a portion thereof (e.g., a read). The threshold may be varied based on empirical analysis. The threshold may be compared to a measured or calculated value to determine whether a source giving rise to such a value should be classified in a particular manner. The threshold may be identified empirically or analytically. The choice of threshold depends on the confidence with which a user desires the classification to be made. The threshold may be selected for a specific purpose (e.g., to balance sensitivity and selectivity). As used herein, the term "threshold" refers to a point at which the course of an analysis may change and / or an action may be triggered. A threshold need not be a predetermined number. Instead, the threshold may be, for example, a function based on multiple factors. The threshold may be adaptive to the situation. Furthermore, a threshold may indicate an upper limit, a lower limit, or a range between limits.
[0258] In some implementations, an index or score based on the sequencing data may be compared to a threshold. As used herein, the terms "metric" or "score" may include a value or result determined from sequencing data or a function based on a value or result determined from sequencing data. Like a threshold, an index or score may be adaptive to the context. For example, an index or score may be a normalized value. As an example of a score or index, one or more implementations may use a count score when analyzing data. The count score may be based on the number of sample reads. The sample reads may have undergone one or more filtering stages so that the sample reads have at least one common characteristic or quality. For example, each of the sample reads used to determine the count score may be aligned with a reference sequence or assigned as a potential allele. The number of sample reads with a common characteristic may be counted to determine a read count. The count score may be based on the read count. In some implementations, the count score may be a value equal to the read count. In other implementations, the count score may be based on the read counts and other information. For example, the count score may be based on the read counts of a particular allele at a locus and the total number of reads at the locus. In some implementations, the count score may be based on the read counts of a locus and previously obtained data. In some implementations, the count score may be a normalized score between predetermined values. The count score may also be a function of read counts from other loci in the sample or read counts from other samples run simultaneously with the sample of interest. For example, the count score may be a function of the read counts of a particular allele and read counts of other loci in the sample and / or read counts from other samples. As an example, the read counts from other loci and / or read counts from other samples may be used to normalize the count score for a particular allele.
[0259] The term "coverage" or "fragment coverage" refers to a count or other measure of multiple sample reads for the same fragment of a sequence. The read count may represent a count of the number of reads that cover the corresponding fragment. Alternatively, coverage may be determined by multiplying the read count by a specified factor based on historical knowledge, knowledge of the sample, knowledge of the locus, etc.
[0260] The term "read depth" (conventionally a number followed by "x") refers to the number of sequenced reads with overlapping alignment at a target position. It is often expressed as an average or percentage above a cutoff for a set of intervals (such as an exon, gene, or panel). For example, a clinical report may state that the panel average coverage is 1.105x, with 98% of the target bases covered at >100x.
[0261] The term "base call quality score" or "Q score" refers to a PHRED-scaled probability ranging from 0 to 50 that is inversely proportional to the probability that a single sequenced base is correct. For example, a T base call with a Q of 20 is considered 99.99% likely to be correct. Any base call with a Q < 20 should be considered low quality, and any variant in which a substantial proportion of sequenced reads supporting the variant are identified as being of low quality should be considered a potential false positive.
[0262] The term "variant read" or "variant read number" refers to the number of sequenced reads that support the presence of a variant.
[0263] With respect to "strandedness" (or DNA strandedness), the genetic message in DNA can be represented as the letters A, G, C, and T, e.g., 5'-AGGACA-3'. Often, sequences are written in the orientation shown herein, i.e., with the 5' end to the left and the 3' end to the right. While DNA can occur as a single-stranded molecule (such as certain viruses), we typically view DNA as a double-stranded unit. It has a double helix structure with two antiparallel strands. In this case, the term "antiparallel" means that the two strands run parallel but have opposite polarity. Double-stranded DNA is held together by base pairing; adenine (A) always pairs with thymine (T) and cytosine (C) pairs with guanine (G). This pairing is called complementarity, and one DNA strand is said to be the complement of another. Thus, double-stranded DNA can be represented as two strings, such as 5'-AGGACA-3' and 3'-TCCTGT-5'. Note that the two strands have opposite polarities. Thus, the strands of two DNA can be referred to as a reference strand and its complement, a forward and reverse strand, a top and bottom strand, a sense and an antisense strand, or a Watson and Crick strand.
[0264] Read alignment (also called read mapping) is the process of finding where a sequence is located in the genome. Once aligned, the "mapping quality" or "mapping quality score (MAPQ)" of a given read quantifies the probability that its location on the genome is correct. Mapping quality is encoded on a phase scale, where P is the probability that the alignment is incorrect. The probability is P=10 (-MAQ / 10)MAPQ is calculated as follows: where MAPQ is the mapping quality. For example, a mapping quality of 40 = 10 -4 means that there is a 0.01% chance that the reads are incorrectly aligned. Therefore, mapping quality is related to several alignment factors, such as the base quality of the reads, the complexity of the reference genome, and paired-end information. First, if the base quality of the reads is low, it means that the observed sequence may be incorrect, and therefore the alignment is incorrect. Second, mapping ability refers to the complexity of the genome. Repetitive regions make mapping more difficult, and reads contained in these regions usually have low mapping quality. In this context, MAPQ reflects the fact that reads are not uniquely aligned and their actual origin cannot be determined. Third, in the case of paired-end sequencing data, matched pairs are more likely to be aligned. The higher the mapping quality, the better the alignment. Reads aligned with good mapping quality usually mean that the reads are well aligned and aligned with few mismatches within high mapping ability regions. MAPQ values can be used as a quality control for alignment results. A percentage of aligned reads with a MAPQ higher than 20 is usually suitable for downstream analysis.
[0265] As used herein, "signal" refers to a detectable event, such as a light emission, e.g., a light emission in an image. Thus, in some implementations, a signal can represent any detectable light emission (i.e., a "bright spot") captured in an image. Thus, as used herein, "signal" can refer to both actual emissions from analytes in a sample and spurious emissions that do not correlate with actual analytes. Thus, a signal may result from noise and may be subsequently discarded because it does not represent actual analytes in the test strip.
[0266] As used herein, the term "clamp" refers to a group of signals. In certain implementations, the signals are derived from different analytes. In some implementations, a signal clump is a group of signals that cluster together. In other implementations, a signal clump represents a physical area covered by one amplification oligonucleotide. Ideally, each signal clump should be observed as several signals (one per template cycle, possibly more due to crosstalk). Thus, overlapping signals are detected, where two (or more) signals are included in the template from the same signal clump.
[0267] As used herein, terms such as "minimum," "maximum," "minimize," "maximize," and grammatical variations thereof may include values that are not absolute maximums or minimums. In some implementations, values include near maxima and near minimums. In other examples, values may include local maxima and / or local minima. In some implementations, values include only absolute maximums or minimums.
[0268] As used herein, "crosstalk" refers to the detection of a signal in one image that is also detected in a separate image. In some implementations, crosstalk can occur when an emitted signal is detected in two separate detection channels. For example, if an emitted signal occurs in one color, the emission spectrum of that signal may overlap with another emitted signal in another color. In some implementations, fluorescent molecules used to indicate the presence of nucleotide bases A, C, G, and T are detected in separate channels. However, because the emission spectra of A and C overlap, a portion of the C color signal may be detected during detection using the A color channel. Thus, crosstalk between the A and C signals allows a signal from one color image to appear in the other color image. In some implementations, there is crosstalk between G and T. In some implementations, the amount of crosstalk between channels is asymmetric. It will be understood that the amount of crosstalk between channels can be controlled by, among other things, selecting signal molecules with appropriate emission spectra and selecting the size and wavelength range of the detection channels.
[0269] As used herein, "register," "registering," "registration," and similar terms refer to any process for correlating signals in an image or dataset from a first time point or perspective with signals in an image or dataset from another time point or perspective. For example, registration may be used to align signals from a set of images to form a template. In another example, registration may be used to align signals from other images to the template. One signal may be directly or indirectly registered to another signal. For example, a signal from image "S" may be directly registered to image "G." As another example, a signal from image "N" may be directly registered to image "G," or alternatively, a signal from image "N" may be registered to image "S," which was previously registered to image "G." Thus, the signal from image "N" is indirectly registered to image "G."
[0270] As used herein, the term "fiducial" is intended to mean a distinguishable reference point in or on an object. A fiducial point may be, for example, a mark, a second object, a shape, an edge, a region, an irregularity, a channel, a pit, a post, etc. A fiducial point may be present in an image of the object or in another data set derived from detecting the object. A fiducial point may be specified by an x and / or y coordinate in the plane of the object. Alternatively or additionally, a fiducial point may be specified by a z coordinate orthogonal to the xy plane, for example, defined by the relative positions of the object and the detector. One or more coordinates of a fiducial point may be specified relative to one or more other specimens of the object or an image or other data set derived from the object.
[0271] As used herein, the term "optical signal" is intended to include, for example, fluorescence, luminescence, scattering, or absorption signals. Optical signals may be detected in the ultraviolet (UV) range (approximately 200-390 nm), the visible (VIS) range (approximately 391-770 nm), the infrared (IR) range (approximately 0.771-25 micrometers), or other ranges of the electromagnetic spectrum. Optical signals may be detected in a manner that excludes all or part of one or more of these ranges.
[0272] As used herein, the term "optical signal" is intended to mean the amount or quantity of detected energy or encoded information having a desired or predetermined characteristic. For example, optical signals may be quantified by one or more of intensity, wavelength, energy, frequency, power, brightness, etc. Other signals may be quantified according to characteristics such as voltage, current, electric field strength, magnetic field strength, frequency, power, temperature, etc. The absence of a signal is understood to be a signal level of zero or a signal level that is not significantly distinguishable from noise.
[0273] As used herein, the term "simulate" is intended to mean creating a representation or model of a physical object or action that predicts characteristics of the object or action. The representation or model may often be distinguishable from the object or action. For example, the representation or model may be distinguishable from the object with respect to one or more characteristics, such as color, strength of a signal detected from all or a portion of the object, size, or shape. In particular implementations, the representation or model may be idealized, exaggerated, muted, or incomplete compared to the object or action. Thus, in some implementations, a representation of a model may be distinguishable from the object or action it represents with respect to at least one of the above characteristics, for example. The representation or model may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.
[0274] As used herein, the term "specific signal" is intended to mean a detected energy or encoded information that is selectively observed over other energies or information, such as background energy or information. For example, a specific signal can be an optical signal detected at a particular intensity, wavelength, or color, an electrical signal detected at a particular frequency, power, or field strength, or other signal known in the art related to spectroscopy and analytical detection.
[0275] As used herein, the term "swath" is intended to mean a rectangular portion of an object. A swath may be an elongated strip that is scanned by relative movement between the object and the detector in a direction parallel to the longest dimension of the strip. Generally, the width of the rectangular portion or strip is constant along its entire length. Multiple swaths of an object may be parallel to one another. Multiple swaths of an object may be adjacent to one another, overlap one another, abut one another, or separated from one another by interstitial regions.
[0276] As used herein, the term "variance" is intended to mean the difference between an expected and an observed difference, or the difference between two or more observations. For example, variance can be the discrepancy between an expected value and a measured value. Statistical functions such as standard deviation, the square of the standard deviation, the coefficient of variation, etc. can be used to express variance.
[0277] As used herein, the term "x-y coordinates" is intended to mean information that specifies a position, size, shape, and / or orientation within an x-y plane. The information may be, for example, numerical coordinates in a Cartesian coordinate system. The coordinates may be provided relative to one or both of the x-axis and y-axis, or may be provided relative to another location within the x-y plane. For example, the coordinates of an analyte on an object may specify the location of the analyte relative to a reference or other analyte locations on the object.
[0278] As used herein, the term "xy-plane" is intended to mean a two-dimensional region defined by linear axes x and y. When used in reference to a detector and an object observed by the detector, the region may be further specified as being orthogonal to the direction of observation between the detector and the object being detected.
[0279] As used herein, the term "z-coordinate" is intended to mean information specifying the location of a point, line, or region along an axis orthogonal to the x-y plane. In certain implementations, the z-axis is orthogonal to the region of the object observed by the detector. For example, the direction of the focus of an optical system may be specified along the z-axis.
[0280] In some implementations, the acquired signal data is transformed using an affine transformation. In some such implementations, template generation utilizes the fact that affine transformations between color channels are consistent across runs. Because of this consistency, a set of default offsets may be used when determining the coordinates of analytes in a sample. For example, a default offset file may contain the relative transformations (shifts, scales, skews) of different channels with respect to one channel, such as the A channel. However, in other implementations, the offsets between color channels vary during and / or between runs, making offset-driven template generation challenging. In such implementations, the methods and systems provided herein may utilize offset-reduced template generation, which is further described below.
[0281] In some aspects of the above implementations, the system may include a flow cell. In some aspects, the flow cell includes lanes or other configurations of tiles, at least some of which include an array of one or more analytes. In some aspects, the analytes include a plurality of molecules, such as nucleic acids. In certain aspects, the flow cell is configured to deliver labeled nucleotide bases to the array of nucleic acids, thereby extending primers hybridized to nucleic acids within the analytes to generate signals corresponding to the analytes, including the nucleic acids. In some implementations, the nucleic acids within the analytes are identical or substantially identical to one another.
[0282] In some of the image analysis systems described herein, each image in an image set includes a color signal, with different colors corresponding to different nucleotide bases. In some embodiments, each image in an image set includes a signal having a single color selected from at least four different colors. In some embodiments, each image in an image set includes a signal having a single color selected from four different colors. In some of the systems described herein, nucleic acids can be sequenced by providing four different labeled nucleotide bases to an array of molecules to generate four different images, each image including a signal having a single color, with the color of the signal being different for each of the four different images, thereby generating a cycle of four color images corresponding to the four possible nucleotides present at a particular position in the nucleic acid. In certain embodiments, the system includes a flow cell configured to deliver additional labeled nucleotide bases to the array of molecules, thereby generating a cycle of multiple color images.
[0283] In some implementations, the methods provided herein may include determining whether a processor is actively acquiring data or whether the processor is in a low-activity state. Acquiring and storing a large number of high-quality images typically requires a large amount of storage capacity. Furthermore, once acquired and stored, analyzing the image data can be resource-intensive and may impede the processing power of other functions, such as the ongoing acquisition and storage of additional image data. Thus, as used herein, the term low-activity state refers to the processing power of a processor at a given time. In some implementations, a low-activity state occurs when the processor is not acquiring and / or storing data. In some implementations, a low-activity state occurs when some data acquisition and / or storage is occurring, but additional processing power remains so that image analysis can occur simultaneously without interfering with other functions.
[0284] As used herein, "identifying a conflict" refers to identifying a situation in which multiple processes compete for a resource. In some such implementations, one process is prioritized over another. In some implementations, the conflict may relate to the need to prioritize allocation of time, processing power, storage power, or any other resource that is prioritized. Thus, in some implementations, when processing time or capacity is distributed between two processes, such as analyzing a data set and acquiring and / or storing a data set, a conflict between the two processes exists and may be resolved by prioritizing one of the processes.
[0285] Also provided herein are systems for performing image analysis. The system may include a processor, a storage capacity, and a program for image analysis, the program including instructions for processing a first dataset for storage and a second dataset for analysis, the processing including acquiring and / or storing the first dataset on the storage device and analyzing the second dataset when the processor is not acquiring the first dataset. In certain aspects, the program includes instructions for identifying at least one instance of a conflict between acquiring and / or storing the first dataset and analyzing the second dataset, and instructions for resolving the conflict in which acquiring and / or storing the image data takes priority, such that acquiring and / or storing the first dataset takes priority. In certain aspects, the first dataset includes an image file obtained from an optical imaging device. In certain aspects, the system further includes an optical imaging device. In some aspects, the optical imaging device comprises a light source and a detection device.
[0286] As used herein, the term "program" refers to instructions or commands for performing a task or process. The term "program" may be used interchangeably with the term module. In certain implementations, a program may be a compilation of various instructions executed under the same command set. In other implementations, a program may refer to a separate batch or file.
[0287] The following are some of the surprising benefits of using the methods and systems for performing image analysis described herein. In some sequencing implementations, an important measure of the usefulness of a sequencing system is its overall efficiency. For example, the amount of mappable data generated per day and the total cost of installing and running the equipment are important aspects of an economical sequencing solution. To shorten the time required to generate mappable data and increase the efficiency of the system, real-time base calling can be enabled on the equipment computer and can be performed in parallel with sequencing chemistry and imaging. This allows most of the data processing and analysis to be completed before finalizing the sequencing chemistry. Furthermore, the storage required for intermediate data can be reduced, limiting the amount of data that needs to be transferred across the network.
[0288] Although sequence output is increasing, the data per run transferred from the systems provided herein to the network and secondary analytical processing hardware is substantially reduced. By converting the data on the instrument computer (acquisition computer), the network load is dramatically reduced. Without these on-instrument, off-network data reduction techniques, the image output of a fleet of DNA sequencing instruments would overwhelm most networks.
[0289] The widespread adoption of high-throughput DNA sequencing instruments has been driven in part by their ease of use, support for a wide range of applications, and adaptability to virtually any laboratory environment. The highly efficient algorithms presented herein allow for the addition of significant analytical capabilities to simple workstations capable of controlling sequencing instruments. This reduction in computational hardware requirements has several practical advantages that become even more important as sequencing output levels continue to increase. For example, by performing image analysis and base calling in simple towers, heat generation, laboratory footprint, and power consumption are kept to a minimum. In contrast, other commercial sequencing technologies have recently ramped up their computing infrastructure by up to five times the processing power for primary analysis, resulting in increased heat output and power consumption. Thus, in some implementations, the computational efficiency of the methods and systems provided herein allows customers to increase their sequencing throughput while minimizing server hardware expenditures.
[0290] Thus, in some implementations, the methods and / or systems presented herein function as a state machine, keeping track of the individual state of each sample and, upon detecting that a sample is ready to progress to the next state, taking appropriate action to advance the sample to that state. A more detailed example of how a state machine monitors the file system to determine if a sample is ready to progress to the next state according to a particular implementation is provided in Example 1 below.
[0291] In some implementations, the methods and systems provided herein are multi-threaded and can work with a configurable number of threads. Thus, for example, in the context of nucleic acid sequencing, the methods and systems provided herein can work behind the scenes during live sequencing runs for real-time analysis, or can be run using existing image datasets for offline analysis. In certain implementations, the methods and systems handle multi-threading by giving each thread its own subset of the samples it is involved in. This minimizes the possibility of thread hoarding.
[0292] The disclosed methods can include using a detection device to acquire a target image of an object, the image including a repeating pattern of analytes on the object. Detection devices capable of high-resolution imaging of surfaces are particularly useful. In certain implementations, the detection device has sufficient resolution to distinguish analytes at the densities, pitches, and / or analyte sizes described herein. Detection devices capable of acquiring images or image data from a surface are particularly useful. Exemplary detectors are those configured to acquire area images while maintaining a static relationship between the object and the detector. Scanning devices can also be used. For example, devices that acquire continuous area images (e.g., so-called "step and shoot" detectors) can be used. Devices that continuously scan points or lines on the surface of an object and accumulate data to construct an image of the surface are also useful. Point-scan detectors can be configured to scan points (i.e., small detection areas) on the surface of an object via a raster motion in the x-y plane of the surface. Line-scan detectors can be configured to scan a line along the y-dimension of the surface of the object, with the longest dimension of the line occurring along the x-dimension. It should be understood that the detection device, the object, or both can be moved to achieve scanning detection.For example, the detection device that is particularly useful in nucleic acid sequencing applications is described in US Patent Application Publication No. 2012 / 0270305 (A1), No. 2013 / 0023422 (A1) and No. 2013 / 0260372 (A1), and US Patent No. 5,528,050, No. 5,719,391, No. 8,158,926 and No. 8,241,573, each of which is incorporated herein by reference.
[0293] Implementations disclosed herein may be implemented as a method, apparatus, system, or article of manufacture using programming or engineering techniques to generate software, firmware, hardware, or any combination thereof. As used herein, the term "article of manufacture" refers to code or logic implemented in hardware or computer-readable media, such as optical storage devices, and volatile or non-volatile memory devices. Such hardware may include, but is not limited to, field programmable gate arrays (FPGAs), coarse-grained reconfigurable architectures (CGRAs), application-specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microprocessors, or other similar processing devices. In certain implementations, the information or algorithms described herein reside in non-transitory storage media.
[0294] In certain implementations, the computer-implemented methods described herein can be performed in real time while multiple images of an object are being acquired. Such real-time analysis is particularly useful for nucleic acid sequencing applications, in which nucleic acid arrays are subjected to repeated cycles of fluidic and detection steps. Analysis of sequencing data can often be computationally intensive, so it may be beneficial to perform the methods described herein in real time or behind the scenes while other data acquisition or analysis algorithms are ongoing. Examples of real-time analysis methods that can be used in the present methods are those used in the MiSeq and HiSeq sequencing devices commercially available from Illumina, Inc. (San Diego, Calif.) and / or described in U.S. Patent Application Publication No. 2012 / 0020537(A1), which is incorporated herein by reference.
[0295] An exemplary data analysis system is formed by one or more programmed computers, and has programming stored on one or more machine-readable media having code executed to perform one or more steps of the methods described herein. In one implementation, for example, the system includes an interface designed to enable networking of the system to one or more detection systems (e.g., optical imaging systems) configured to acquire data from target objects. The interface may receive and condition the data, as appropriate. In certain implementations, the detection systems output digital image data, e.g., image data representing individual image elements or pixels that together form an image of an array or other object. A processor processes the received detection data according to one or more routines defined by the processing code. The processing code may be stored in various types of memory circuits.
[0296] According to currently contemplated implementations, processing code executed on the detection data includes data analysis routines designed to analyze the detection data to determine the location and metadata of individual analytes visualized or encoded within the data, as well as locations where no analyte is detected (i.e., locations where no analyte is present or where no significant signal is detected from an existing analyte). In certain implementations, analyte locations within the array typically appear brighter than non-analyte locations due to the presence of a fluorescent dye attached to the imaged analyte. It will be understood that analytes need not appear brighter than their surrounding areas if, for example, a probe's target in the analyte is not present within the array being detected. The color in which individual analytes appear can be a function of the dye used as well as the wavelength of light used by the imaging system for imaging purposes. Analytes without bound targets or specific labels can be identified according to other characteristics, such as their expected location within the microarray.
[0297] Once the data analysis routine has located individual analytes in the data, value assignment can be performed. Generally, value assignment assigns a digital value to each analyte based on the characteristics of the data represented by the detector elements (e.g., pixels) at the corresponding location. That is, for example, when the imaging data is processed, the value assignment routine may be designed to recognize that a particular color or wavelength of light has been detected at a particular location, as indicated by a group or cluster of pixels at that location. In a typical DNA imaging application, for example, the four common nucleotides are represented by four separate, distinguishable colors. Each color may then be assigned a value corresponding to that nucleotide.
[0298] As used herein, the terms “module,” “system,” or “system controller” may include hardware and / or software systems and circuits that operate to perform one or more functions. For example, a module, system, or system controller may include a computer processor, controller, or other log-based device that performs operations based on instructions stored on a tangible and non-transitory computer-readable storage medium, such as a computer memory. Alternatively, a module, system, or system controller may include a wired device that performs operations based on hardwired logic and circuitry. A module, system, or system controller shown in the accompanying figures may represent hardware and circuitry that operates based on software or hardware instructions, software that directs hardware to perform operations, or a combination thereof. A module, system, or system controller may include or represent hardware circuitry or circuitry that includes and / or is connected to one or more processors, such as a computer microprocessor.
[0299] As used herein, the terms "software" and "firmware" are used interchangeably and include any computer program stored in memory that is executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely examples and are therefore not limiting of the types of memory that can be used to store computer programs.
[0300] In the field of molecular biology, one of the processes for nucleic acid sequencing in use is sequencing-by-synthesis. This technology can be applied to large-scale parallel sequencing projects. For example, by using an automated platform, it is possible to perform millions of sequencing reactions simultaneously. Therefore, one implementation of the present invention relates to an apparatus and method for acquiring, storing, and analyzing image data generated during nucleic acid sequencing.
[0301] The enormous gain in the amount of data that can be acquired and stored makes streamlined image analysis methods even more beneficial. For example, the image analysis methods described herein allow both designers and end users to efficiently utilize existing computer hardware. Thus, methods and systems are presented herein that reduce the computational complexity of processing data in the face of rapidly increasing data output. For example, in the field of DNA sequencing, yields have expanded 15-fold in recent years, potentially reaching hundreds of gigabytes in a single run of a DNA sequencing device. Large-scale genome-scale experiments would remain out of reach for most researchers if computational infrastructure requirements increased proportionally. Therefore, the generation of more raw sequence data increases the need for secondary analysis and data storage, making optimization of data transport and storage highly beneficial. Some implementations of the methods and systems presented herein can reduce the time, hardware, networking, and laboratory infrastructure requirements required to generate usable sequence data.
[0302] This disclosure describes various methods and systems for performing the methods. Some example methods are described as a series of steps. However, it should be understood that implementations are not limited to the specific steps and / or order of steps described herein. Steps may be omitted, steps may be modified, and / or other steps may be added. Furthermore, steps described herein may be combined, steps may be performed simultaneously, steps may be performed in parallel, steps may be divided into multiple substeps, steps may be performed in a different order, or steps (or series of steps) may be re-performed iteratively. It should be understood that although different methods are described herein, other implementations may combine different methods (or steps of different methods).
[0303] In some implementations, a processing unit, processor, module, or computing system that is "configured" to perform a task or operation may be understood to be specifically structured to perform the task or operation (e.g., having stored thereon or used in conjunction with one or more programs or instructions that are adapted or intended to perform the task or operation, and / or having an arrangement of processing circuitry that is adapted or intended to perform the task or operation). For clarity and avoidance of doubt, a general-purpose computer (which may be "configured" to perform a task or operation when suitably programmed) is not "configured" to perform a task or operation unless or until it is specifically programmed or structurally modified to perform the task or operation.
[0304] Furthermore, the operations of the methods described herein may be sufficiently complex such that the operations cannot be intelligently performed by the average person or skilled in the art within a commercially reasonable period of time. For example, the methods may rely on relatively complex calculations such that such a person cannot complete the methods within a commercially reasonable period of time.
[0305] Throughout this application, various publications, patents, or patent applications are referenced. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which this invention pertains.
[0306] The term "comprising," as used herein, is intended to be open-ended, including not only the recited elements, but also including any additional elements.
[0307] As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set. Exceptions may occur where explicit disclosure or context clearly dictates otherwise.
[0308] Although the invention has been described with reference to the above embodiments, it will be understood that various modifications can be made without departing from the invention.
[0309] The modules of the present application can be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some may be implemented on different processors or computers, or may be spread across multiple different processors or computers. In addition, it will be understood that some of the modules may be combined and operated in parallel or in a different order than shown in the figures without affecting the functionality achieved. Also, as used herein, a "module" may include "sub-modules," which may be considered herein to constitute a module. Blocks in the figures designated as modules may also be considered flowchart steps in a method.
[0310] As used herein, "identification" of an item of information does not necessarily require direct specification of that item of information. Information may be "identified" within a field by simply referencing the actual information through one or more layers in one direction, or by identifying one or more different items of information that are sufficient to determine the actual item of information. Note that the term "specify" is used herein to mean the same thing as "identify."
[0311] As used herein, a given signal, event, or value is "in dependence upon" a preceding signal, a preceding signal event or value, or an event or value that is influenced by the given signal, event, or value. If there are intervening processing elements, steps, or periods, a given signal, event, or value may still "depend" on the preceding signal, event, or value. If an intervening processing element or step combines two or more signals, events, or values, the signal output of the processing element or step is considered to be "dependent" on each of the signal, event, or value inputs. If a given signal, event, or value is the same as a preceding signal, event, or value, this is simply considered to mean that the given signal, event, or value is still "in dependence upon" or "dependent on" or "based on" the preceding signal, event, or value. The "responsiveness" of a given signal, event, or value to another signal, event, or value is defined similarly.
[0312] As used herein, "concurrently" or "in parallel" does not require exact simultaneity. It is sufficient if the assessment of one of the individuals begins before the assessment of another of the individuals is completed.
[0313] Computer Systems FIG. 17A is a block diagram of an exemplary computer system. The computer system includes a storage subsystem, a user interface input device, a CPU, a network interface, a user interface output device, and an optional deep learning processor (illustrated as a GPU, FPGA, and CGRA for simplicity), interconnected by a bus subsystem. The storage system includes a memory subsystem and a file storage subsystem. The memory subsystem includes randomly accessible read / write memory (RAM) and read-only memory (ROM). The ROM and file storage subsystem elements provide non-transitory computer-readable media capabilities for storing and executing instructions programmed to implement, for example, all or any portion of the RTA functionality described elsewhere herein. According to various implementations, the deep learning processor is enabled to implement all or any portion of the RTA functionality described elsewhere herein. In various implementations, the deep learning processor elements include various combinations of CPUs, GPUs, FPGAs, CGRAs, ASICs, ASIPs, and DSPs.
[0314] Generally, computer system 1700 can be used to implement the disclosed techniques. More specifically, computer system 1700 includes at least one central processing unit (CPU) 1772 capable of communicating with a number of peripheral devices via a bus subsystem 1755. The peripheral devices variously include a storage subsystem 1710, including, for example, a memory device and file storage subsystem 1736, a user interface input device 1738, a user interface output device 1776, and a network interface subsystem 1774. The input and output devices enable user interaction with computer system 1700. Network interface subsystem 1774 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0315] User interface input devices 1738 may include various types of input devices, such as keyboards, pointing devices such as mice, trackballs, touchpads, and / or graphics tablets, scanners, touchscreens integrated into displays, voice input devices such as voice recognition systems and microphones, and other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and methods for inputting information into computer system 1700.
[0316] The user interface output devices 1776 variously include a display subsystem, a printer, fax capabilities, and / or a non-visual display such as an audio output device. The display subsystem variously includes a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem optionally provides a non-visual display such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from computer system 1700 to a user or to another machine or computer system.
[0317] The storage subsystem 1710 is capable of storing software modules containing programming and data components that provide the functionality of some or all of the techniques described herein. The software modules are generally executed by the processor 1778.
[0318] The processor 1778 may include any combination of a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 1778 may be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processors 1778 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX17 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu's DPI™, ARM's DynamicIQ™, IBM's TrueNorth™, and Lambda's GPU Server with Testa V100s™.
[0319] The memory subsystem 1722 of the storage subsystem 1710 variously includes multiple memories, including a main random access memory (RAM) 1732 for storing instructions and data during program execution, and a read-only memory (ROM) 1734 for storing fixed information such as instructions and constants. The file storage subsystem 1736 can provide persistent storage for program files and data files, and variously includes a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, and / or a removable media cartridge. Modules that implement the functionality of particular implementations are variously stored by the file storage subsystem 1736 within the storage subsystem 1710 or in other machines accessible by the processor.
[0320] Bus subsystem 1755 allows communication between various components and subsystems of computer system 1700. Although bus subsystem 1755 is illustrated schematically as a single bus, alternative implementations of the bus subsystem use multiple buses.
[0321] The computer system 1700 itself may be of various types according to this implementation, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a widely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of the computer system 1700 shown in Figure 17A is intended only as a specific example for purposes of illustrating various implementations of the present invention. Many other configurations of the computer system 1700 may have more or fewer components than the computer system shown in Figure 17A.
[0322] In various implementations, the equalizer base caller 104 is communicatively linked to the storage subsystem 1710 and / or the user interface input device(s) 1738 .
[0323] In various implementations, one or more of the laboratory equipment and / or generation equipment described elsewhere herein comprises one or more computer systems the same as or similar to the exemplary computer system of the figures. In various implementations, any one or more of the training context and / or generation context uses any one or more computer systems the same as or similar to the exemplary computer system of the figures to perform RTA-related processing, such as acting as one or more servers associated with training data collection and / or processing and production data collection and / or processing.
[0324] In various implementations, the memory subsystem and / or file storage subsystem may store information associated with the RTA, such as all or any portion of information associated with or related to the GT and / or LUT elements of the various equalizers and / or base callers described elsewhere herein. For example, all or any portion of the stored information may variously correspond to any combination of initialization information for an equalizer used in a training context, trained information for an equalizer used in a training context, and / or trained information for an equalizer used in a production context. In another example, all or any portion of the stored information may correspond to one or more intermediate representations, such as those related to information provided by a training context to a production context, as shown and described elsewhere herein.
[0325] FIG. 17B illustrates training and production elements implementing aspects of base calling that depend on flow cell tilt. The upper portion of the figure illustrates one or more training contexts, and the lower portion illustrates one or more production contexts. Each training context includes one or more training data collection / processing capabilities, each having one or more respective training servers. Each training server is capable of storing respective training data, such as information obtained from training via one or more RTA-related activities. In some implementations, all or any portion of one of the training contexts corresponds to laboratory equipment. Each production context includes one or more production devices. Each production device is capable of storing production data.
[0326] In various implementations, the memory subsystem and / or file storage subsystem can store image and tilt data and representations thereof, such as information enabling pixel intensity and / or tilt determination of one or more regions of the image. In various implementations, the computer system can process the image in real time, including extracting the intensity of specific pixels in real time. In some implementations based on real-time pixel intensity extraction, all or any portion of the image data corresponding to the extracted region is not specifically saved to the file storage subsystem. In various implementations, the computer system can process tilt measurements and / or information related to the determination of tilt measurements in real time. In various implementations, the computer system can process image and / or tilt information in real time, e.g., enable base calling in real time.
[0327] The training context in the diagram represents various training contexts illustrated and described elsewhere herein. The production context in the diagram represents various production contexts illustrated and described elsewhere herein. The training context uses collected and / or synthesized training data to train one or more RTA-related elements, such as an equalizer and / or a LUT associated with or included in the equalizer. The results of the training are then provided to a production context for us, as indicated by the dashed arrow "Deploy Trained Information," to provide, for example, gradient-dependent base calling.
[0328] As a first specific example, one of the training contexts in Figure 17B corresponds to the training context in Figure 1AA, and the corresponding one or more of the generation contexts in Figure 17B correspond to one or more instances of the generation context in Figure 1AA. Deployment of the trained information in Figure 17B corresponds to providing information from any one or more of the LUTs of the training contexts in Figure 17B after training is completed to the corresponding LUTs of the generation contexts in Figure 17B in preparation for generation of gradient-dependent base calling.
[0329] As a second specific example, one of the training contexts in FIG. 17B corresponds to the system 100A of FIG. 1A used for training, and one of the generation contexts in FIG. 17B corresponds to the system 100A of FIG. 1A used for generation (e.g., to perform base calling using information determined by training and stored in LUT 106).
[0330] In some implementations, the same servers are used in the training context and the production context, e.g., one or more servers used to implement the training context of Figure 1AA are also used to implement the production context of Figure 1AA.
[0331] Specific Implementations The disclosed technology attenuates spatial crosstalk from sensor pixels using equalization-based image processing techniques. The disclosed technology may be implemented as a system, method, or product. One or more features of the implementations may be combined with a base implementation. Implementations that are not mutually exclusive are taught as combinable. One or more features of the implementations may be combined with other implementations. The present disclosure periodically notifies the user of these options. The omission from some implementations of a repeating list of these options should not be construed as limiting the combinations taught in the preceding sections. These lists are incorporated by reference into each of the following implementations.
[0332] In one implementation, the disclosed technology proposes a computer-implemented method for attenuating spatial crosstalk from sensor pixels.
[0333] The disclosed technique addresses spatial crosstalk on sensor pixels in the pixel plane caused by fluorescent samples periodically distributed within the sample plane. Signal cones from the fluorescent samples are optically coupled to a local grid of sensor pixels via at least one lens. The signal cones overlap and impinge on the sensor pixels, thereby generating spatial crosstalk.
[0334] The disclosed technique captures, in at least one sub-pixel lookup table, the characteristic extent of a characteristic signal cone projected through a lens and the resulting contribution of the characteristic signal cone to the fluorescence detected by sensor pixels within a local grid of sensor pixels, the local grid of sensor pixels being substantially concentric with the center of the characteristic signal cone.
[0335] The disclosed technique interpolates between a set of sub-pixel lookup tables that represent the characteristic spread at sub-pixel resolution to generate an interpolated lookup table based on the target fluorescent sample center.
[0336] The disclosed technique isolates signal from a target fluorescent sample that projects the center of its signal cone onto substantially the center of the target local grid of sensor pixels by convolving an interpolation lookup table with the sensor pixels in the target local grid.
[0337] The disclosed technique uses the sum of the convoluted contributions of the isolated signals as the fluorescence intensity from the target fluorescent sample.
[0338] The disclosed technique then base calls the first target fluorescent sample using the fluorescence intensities. The fluorescence intensities are determined for the first target fluorescent sample for each imaging channel in the multiple imaging channels. Consider a four-channel chemistry, which uses four imaging channels to generate four images per sequencing cycle. Next, for the first target fluorescent sample, four fluorescence intensities are determined using the disclosed technique as described above. The four fluorescence intensities are then processed by a base caller to base call the first target fluorescent sample. Similarly, in a two-channel chemistry, two fluorescence intensities are used to base call the first target fluorescent sample.
[0339] The methods described in this section and other sections of the disclosed technology may include one or more of the following features and / or features described in connection with the additional methods disclosed. For purposes of brevity, combinations of features disclosed in this application are not individually listed and are not repeated for each base set of features. The reader will understand how features identified in this method can be readily combined with the sets of basic features identified as implementations in other sections of this application.
[0340] In some implementations, the periodically distributed fluorescent samples are arranged in a diamond shape, while in other implementations, the periodically distributed fluorescent samples are arranged in a hexagonal shape.
[0341] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.
[0342] In another implementation, the disclosed technology proposes a computer-implemented method for base calling.
[0343] The disclosed technique accesses an image whose pixels represent luminance radiation from a target cluster and luminance radiation from additional adjacent clusters, the pixels including a central pixel that includes a center of the target cluster, each pixel within the central pixel being divisible into multiple sub-pixels.
[0344] In response to a particular subpixel among a plurality of subpixels of a central pixel that includes a center of a target cluster, the disclosed technique selects a subpixel lookup table corresponding to the particular subpixel from a bank of subpixel lookup tables, the selected subpixel lookup table including pixel coefficients configured to accept luminance radiation from the target cluster and reject luminance radiation from adjacent clusters.
[0345] The disclosed technique multiplies the luminance values of pixels in an image element-wise by pixel coefficients and sums the products of the multiplications to generate an output.
[0346] The disclosed technique uses the output to base call target clusters.
[0347] Each of the features discussed in this specific implementation section for other implementations applies equally to this method implementation, and as noted above, all method features are not repeated here and should be considered repeated by reference.
[0348] In some implementations, the disclosed techniques further include (i) selecting, from the bank of subpixel lookup tables, an additional subpixel lookup table corresponding to the subpixel most contiguously adjacent to the particular subpixel; (ii) interpolating between pixel coefficients of the selected subpixel lookup table and the selected additional subpixel lookup table to generate interpolated pixel coefficients configured to accept luminance radiation from the target cluster and reject luminance radiation from adjacent clusters; (iii) element-wise multiplying the interpolated pixel coefficients by luminance values of pixels in the image and summing the products of the multiplications to generate an output; and (iv) base calling the target cluster using the output.
[0349] In some implementations, the target cluster and additional adjacent clusters are periodically distributed in a diamond pattern on the flow cell and immobilized on the wells of the flow cell, while in other implementations, the target cluster and additional adjacent clusters are periodically distributed in a hexagonal pattern on the flow cell and immobilized on the wells of the flow cell.
[0350] In some implementations, the interpolation is based on at least one of linear interpolation, bilinear interpolation, and bicubic interpolation.
[0351] In some implementations, the pixel coefficients of the subpixel lookup tables in the bank of subpixel lookup tables are learned as a result of training an equalizer using decision-directed equalization. In one implementation, the decision-directed equalization uses least-squares estimation as a loss function. In one implementation, the least-squares estimation minimizes the squared error using ground truth base calls. In one implementation, the ground truth base calls are modified to account for DC offsets, amplification factors, and the degree of polyclonality.
[0352] In some implementations, pixel coefficients of the subpixel lookup tables in the bank of subpixel lookup tables are derived from a combination of (i) a single subpixel lookup table whose pixel coefficients are learned as a result of training an equalizer using decision-directed equalization, and (ii) a set of pre-computed interpolation filters, each interpolation filter corresponding to a respective subpixel in the plurality of subpixels.
[0353] The disclosed technique further includes making the center of the target cluster substantially concentric with the center of the central pixel by (i) aligning the image with respect to the template image and determining affine transformation parameters and nonlinear transformation parameters; (ii) using the parameters to transform position coordinates of the target cluster and additional adjacent clusters into image coordinates of the image to generate a transformed image having transformed pixels; and (iii) applying interpolation using the transformed position coordinates of the target cluster and additional adjacent clusters to make each cluster center substantially concentric with the center of each transformed pixel that comprises the cluster center.
[0354] The disclosed techniques further include generating an output for each of the multiple images captured using the respective imaging channels in a particular sequencing cycle, and base calling the target cluster using the output generated for each image.
[0355] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.
[0356] While the present invention has been disclosed with reference to the implementations and examples detailed above, it should be understood that these examples are intended in an illustrative and not a limiting sense. Modifications and combinations may be readily made by those skilled in the art, and such modifications and combinations are deemed to be within the spirit of the invention and the scope of the following claims.
Claims
1. 1. A method for selectively performing base calling on in-focus and out-of-focus elements of images collected during sequencing, comprising: capturing an image of a portion of the flow cell using a sensor with a depth of focus (DoF); classifying each element of the image as either in focus or out of focus in spatial relationship to the DoF based at least in part on one or more tilt measurements of the flow cell portion; selecting one or more respective base callers for each of the in-focus and out-of-focus segments based at least in part on a scalar gradient measurement; performing base calling of the image using each of the one or more selected base callers.
2. The method of claim 1 , wherein the out-of-focus segments include an upper focal segment for out-of-focus elements above the DoF and a lower focal segment for out-of-focus elements below the DoF.
3. The method of claim 2 , wherein selecting one or more base callers for the in-focus section comprises selecting a base caller adapted to process an in-focus image.
4. 3. The method of claim 2, wherein selecting one or more base callers for the out-of-focus section comprises selecting a base caller adapted to process a super-focused image of at least a portion of each element in the super-focused section.
5. 3. The method of claim 2, wherein selecting one or more base callers for the out-of-focus section comprises selecting a base caller adapted to process a down-focused image of at least a portion of each element in the down-focused section.
6. The method of claim 1 , wherein the images are acquired during sequencing by synthesis.
7. 2. The method of claim 1, wherein each of the one or more base callers for the out-of-focus section comprises an equalizer adapted to apply defocus correction based on a set of trained coefficients in each of a plurality of look-up tables (LUTs).
8. 8. The method of claim 7, wherein at least one equalizer of the one or more base callers is at least partially trained using ground truth (GT) corresponding to a down-focused context.
9. 8. The method of claim 7, wherein at least one equalizer of the one or more base callers is at least partially trained using ground truth (GT) corresponding to an up-focused context.
10. 1. A method for selectively performing equalizer-based base calling on in-focus and out-of-focus elements of images collected during sequencing, comprising: capturing an image of a portion of the flow cell using a sensor with a depth of focus (DoF); classifying each element of the image as either in focus or out of focus in spatial relationship to the DoF based at least in part on one or more tilt measurements of the flow cell portion; selecting one or more respective equalizer-based base callers for each of the in-focus and out-of-focus segments based at least in part on a scalar gradient measurement; performing base calling of the image using the one or more selected equalizer-based base callers, wherein performing base calling comprises: applying a set of coefficients in each LUT to the luminance values of a corresponding set of image pixels of the target base; determining a weighted sum of luminance values of the image pixels based on application of the set of coefficients; and performing base calling for the image, comprising outputting the weighted sum as a base call prediction.
11. The method of claim 10 , wherein the out-of-focus section comprises an upper focus section and a lower focus section.
12. 12. The method of claim 11, wherein selecting one or more base callers for the in-focus section comprises selecting a base caller adapted to process an in-focus image.
13. 12. The method of claim 11, wherein selecting one or more base callers for the top-focused section comprises selecting a base caller adapted to process top-focused images.
14. 12. The method of claim 11, wherein selecting one or more base callers for the lower-focus section comprises selecting an equalizer-based base caller adapted to process lower-focused images.
15. The method of claim 10 , wherein the images are acquired during sequencing by synthesis.
16. 1. A method of training an equalizer to perform equalizer-based base calling on in-focus and out-of-focus elements of images collected during sequencing, comprising: acquiring a training dataset of images of respective portions of one or more flow cells, each image including a known out-of-focus region based on a tilt measurement of the respective flow cell portion; inputting the training data set into the equalizer; causing the equalizer to perform base calling on known out-of-focus regions in the training data set, wherein performing base calling comprises: applying a set of coefficients in each LUT of the equalizer to the luminance values of a set of corresponding image pixels of a target base; determining a weighted sum of luminance values of the image pixels based on application of the set of coefficients; calculating an error for the weighted sum based on a brightness target determined for the target base; updating the set of coefficients with values that reduce the error using a derivative of the error; and repeatedly causing the equalizer to perform base calling on known out-of-focus regions in the training dataset until an updated set of coefficients no longer reduces the error.
17. 17. The method of claim 16, wherein the equalizer is trained to perform equalizer-based base calling in a training context, the method further comprising exporting the updated set of coefficients for use in an equalizer in a production context.
18. The method of claim 16 , wherein the equalizer is trained in a generative context.
19. 1. A system for selectively performing base calling on in-focus and out-of-focus elements of images collected during sequencing, comprising: at least one processor; at least one memory system having computer-executable instructions stored thereon, the computer-executable instructions, when executed by the processor, causing the processor to: receiving an image of a portion of the flow cell captured using a sensor having a depth of focus (DoF); determining a segmentation of each element of the image as being either in focus or out of focus in spatial relationship to the DoF based at least in part on one or more tilt measurements of the flow cell portion; selecting one or more respective base callers for each of the in-focus and out-of-focus segments based at least in part on a scalar gradient measurement; at least one memory system that selectively causes each of the one or more selected base callers to perform base calling on respective in-focus and out-of-focus elements of the image.
20. 20. The method of claim 19, wherein the out-of-focus segments include an upper focal segment for out-of-focus elements above the DoF and a lower focal segment for out-of-focus elements below the DoF.
21. 21. The method of claim 20, wherein the computer-executable instructions, when executed by the processor, cause the processor to select one or more base callers for the in-focus section, the base callers adapted to process the in-focus image.
22. 21. The method of claim 20, wherein the computer-executable instructions, when executed by the processor, cause the processor to select one or more base callers for the out-of-focus sections, the base callers adapted to process super-focused images.
23. 21. The method of claim 20, wherein the computer-executable instructions, when executed by the processor, cause the processor to select one or more base callers for the out-of-focus sections, the base callers adapted to process under-focused images.
24. 20. The method of claim 19, wherein the image of the portion of the flow cell is collected during sequencing-by-synthesis.
25. 20. The method of claim 19, wherein each of the one or more base callers for the out-of-focus section comprises an equalizer adapted to apply defocus correction based on a set of trained coefficients in each of a plurality of look-up tables (LUTs).
26. 26. The method of claim 25, wherein at least one equalizer of the one or more base callers is at least partially trained using ground truth (GT) corresponding to a down-focused context.
27. 26. The method of claim 25, wherein at least one equalizer of the one or more base callers is at least partially trained using ground truth (GT) corresponding to an up-focused context.
Citation Information
Patent Citations
Optical distortion correction for imaging sample
JP2018173943A
Analyzer for three-dimensional analysis of medical specimens using a light field camera
JP2021525866A
Equalization-based image processing and spatial crosstalk attenuator
JP2023525993A
Focusing methods and optical systems and assemblies using the same
US20110188053A1
Equalization-based image processing and spatial crosstalk attenuator
WO2021226285A1