Equalizer-based intensity correction for base calling

The equalizer-based method addresses spatial crosstalk in DNA sequencing by generating a lookup table to enhance signal-to-noise ratio, thereby improving base calling accuracy and reducing sequencing errors.

JP7767672B2Active Publication Date: 2025-11-11ILLUMINA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025115949
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-04
Filing Date
2025-07-09
Publication Date
2025-11-11
Estimated Expiration
2041-05-05

AI Technical Summary

Technical Problem

Existing DNA sequencing systems face challenges in accurately distinguishing true optical signals from neighboring wells due to spatial crosstalk, leading to sequencing errors and reduced base calling accuracy.

Method used

A computer-implemented method using an equalizer-based approach to generate a lookup table that corrects for spatial crosstalk by convolving pixel coefficients with intensity values to maximize the signal-to-noise ratio for base calling.

Benefits of technology

Improves base calling accuracy by effectively reducing sequencing errors caused by spatial crosstalk, enhancing the reliability of DNA sequencing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767672000008
    Figure 0007767672000008
  • Figure 0007767672000009
    Figure 0007767672000009
  • Figure 0007767672000010
    Figure 0007767672000010
Patent Text Reader

Abstract

To attenuate spatial crosstalk from sequencing images for base calling.SOLUTION: The disclosed technology accesses an image whose pixels depict intensity emissions from a target cluster and intensity emissions from additional adjacent clusters. The pixels include a center pixel that contains a center of the target cluster. Each pixel in the pixels is divisible into a plurality of subpixels. Depending upon a particular subpixel, in a plurality of subpixels of the center pixel which contains the center of the target cluster, the disclosed technology selects, from a bank of subpixel lookup tables, a subpixel lookup table that corresponds to the particular subpixel. The selected subpixel lookup table contains pixel coefficients that are configured to maximize a signal-to-noise ratio. The disclosed technology element-wise multiplies the pixel coefficients with the pixels and determines a weighted sum.SELECTED DRAWING: Figure 1B
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority application) This PCT application claims the benefit of U.S. Provisional Patent Application No. 63 / 020,449, entitled "EQUALIZATION-BASED IMAGE PROCESSING AND SPATIAL CROSSTALK ATTENUATOR," filed May 5, 2021 (Attorney Docket No. ILLM1032-1 / IP-1991-PRV), and U.S. Provisional Patent Application No. 17 / 308,035, entitled "EQUALIZATION-BASED IMAGE PROCESSING AND SPATIAL CROSSTALK ATTENUATOR," filed May 4, 2020 (Attorney Docket No. ILLM1032-2 / IP-1991-US), the priority applications of which are incorporated by reference for all purposes.

[0002] FIELD OF THE INVENTION The disclosed technology relates to an apparatus and corresponding method for automated analysis or pattern recognition of an image. This specification includes systems that transform an image to (a) improve its visual quality before recognition, (b) reduce the amount of image data by registering and aligning the image with a sensor or stored prototype or discarding irrelevant data, and (c) measure significant characteristics of the image. Specifically, the disclosed technology relates to removing spatial crosstalk from sensor pixels using equalization-based image processing techniques.

[0003] (CROSS-REFERENCE TO RELATED APPLICATIONS) Built-in The following are incorporated by reference for all purposes as if fully set forth herein: U.S. Nonprovisional Patent Application No. 15 / 936,365, filed March 26, 2018, entitled "DETECTION APPARATUS HAVING A MICROFLUOROMETER, A FLUIDIC SYSTEM, AND A FLOW CELL LATCH CLAMP MODULE"; U.S. Nonprovisional Patent Application No. 16 / 567,224, entitled "FLOW CELLS AND METHODS RELATED TO SAME," filed September 11, 2019; U.S. Nonprovisional Patent Application No. 16 / 439,635, entitled "DEVICE FOR LUMINESCENT IMAGING," filed June 12, 2019; U.S. Nonprovisional Patent Application No. 15 / 594,413, filed May 12, 2017, entitled "INTEGRATED OPTOELECTRONIC READ HEAD AND FLUIDIC CARTRIDGE USEFUL FOR NUCLEIC ACID SEQUENCING"; U.S. Nonprovisional Patent Application No. 16 / 351,193, entitled "ILLUMINATION FOR FLUORESCENCE IMAGING USING OBJECTIVE LENS," filed March 12, 2019; U.S. Non-Provisional Patent Application No. 12 / 638,770, entitled "DYNAMIC AUTOFOCUS METHOD AND SYSTEM FOR ASSAY IMAGER," filed December 15, 2009; U.S. Nonprovisional Patent Application No. 13 / 783,043, entitled "KINETIC EXCLUSION AMPLIFICATION OF NUCLEIC ACID LIBRARIES," filed March 1, 2013; U.S. Non-Provisional Patent Application No. 13 / 006,206, entitled "DATA PROCESSING SYSTEM AND METHODS," filed January 13, 2011; U.S. Non-Provisional Patent Application No. 14 / 530,299, entitled "IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS," filed October 31, 2014; U.S. Non-Provisional Patent Application No. 15 / 153,953, entitled "METHODS AND SYSTEMS FOR ANALYZING IMAGE DATA," filed December 3, 2014; U.S. Nonprovisional Patent Application No. 14 / 020,570, filed September 6, 2013, entitled "CENTROID MARKERS FOR IMAGE ANALYSIS OF HIGH DENSITY CLUSTERS IN COMPLEX POLYNUCLEOTIDE SEQUENCING"; U.S. Non-Provisional Patent Application No. 14 / 530,299, entitled "IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS," filed October 31, 2014; U.S. Nonprovisional Patent Application No. 12 / 565,341, entitled "METHOD AND SYSTEM FOR DETERMINING THE ACCURACY OF DNA BASE IDENTIFICATIONS," filed September 23, 2009; U.S. Non-Provisional Patent Application No. 12 / 295,337, entitled "SYSTEMS AND DEVICES FOR SEQUENCE BY SYNTHESIS ANALYSIS," filed March 30, 2007; U.S. Non-Provisional Patent Application No. 12 / 020,739, entitled "IMAGE DATA EFFICIENT GENETIC SEQUENCING METHOD AND SYSTEM," filed January 28, 2008; U.S. Nonprovisional Patent Application No. 13 / 833,619, entitled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR SAME," filed March 15, 2013 (Attorney Docket No. IP-0626-US); U.S. Nonprovisional Patent Application No. 15 / 175,489 (Attorney Docket No. IP-0689-US), entitled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND METHODS OF MANUFACTURING THE SAME," filed June 7, 2016; U.S. Non-Provisional Patent Application No. 13 / 882,088 (Attorney Docket No. IP-0462-US), entitled "MICRODEVICES AND BIOSENSOR CARTRIDGES FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR THE SAME," filed April 26, 2013; U.S. Nonprovisional Patent Application No. 13 / 624,200, entitled "METHODS AND COMPOSITIONS FOR NUCLEIC ACID SEQUENCING," filed September 21, 2012 (Attorney Docket No. IP-0538-US); U.S. Provisional Patent Application No. 62 / 821,602, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed March 21, 2019 (Attorney Docket No. ILLM1008-1 / IP-1693-PRV); U.S. Provisional Patent Application No. 62 / 821,618, entitled "Artificial Intelligence-Based Generation of Sequencing Metadata," filed March 21, 2019 (Attorney Docket No. ILLM1008-3 / IP-1741-PRV); U.S. Provisional Patent Application No. 62 / 821,681, entitled "Artificial Intelligence-Based Base Calling," filed March 21, 2019 (Attorney Docket No. ILLM1008-4 / IP-1744-PRV); U.S. Provisional Patent Application No. 62 / 821,724, entitled "Artificial Intelligence-Based Quality Scoring," filed March 21, 2019 (Attorney Docket No. ILLM1008-7 / IP-1747-PRV); U.S. Provisional Patent Application No. 62 / 821,766, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV); Dutch Patent Application No. 2023310 (Attorney Docket No. ILLM1008-11 / IP-1693-NL), entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed on June 14, 2019; Dutch Patent Application No. 2023311 (Attorney Docket No. ILLM1008-12 / IP-1741-NL), entitled "Artificial Intelligence-Based Generation of Sequencing Metadata," filed on June 14, 2019; Dutch Patent Application No. 2023312 entitled "Artificial Intelligence-Based Base Calling" filed on June 14, 2019 (Attorney Docket No. ILLM1008-13 / IP-1744-NL); Dutch Patent Application No. 2023314 entitled "Artificial Intelligence-Based Quality Scoring", filed on June 14, 2019 (Attorney Docket No. ILLM1008-14 / IP-1747-NL), and Dutch Patent Application No. 2023316 (Attorney Docket No. ILLM1008-15 / IP-1752-NL), entitled "Artificial Intelligence-Based Sequencing," filed on June 14, 2019. U.S. Nonprovisional Patent Application No. 16 / 825,987, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed March 20, 2020 (Attorney Docket No. ILLM1008-16 / IP-1693-US); U.S. Nonprovisional Patent Application No. 16 / 825,991, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed March 20, 2020 (Attorney Docket No. ILLM1008-17 / IP-1741-US); U.S. Nonprovisional Patent Application No. 16 / 826,126, entitled "Artificial Intelligence-Based Base Calling," filed March 20, 2020 (Attorney Docket No. ILLM1008-18 / IP-1744-US); U.S. Nonprovisional Patent Application No. 16 / 826,134, entitled "Artificial Intelligence-Based Quality Scoring," filed March 20, 2020 (Attorney Docket No. ILLM1008-19 / IP-1747-US); U.S. Nonprovisional Patent Application No. 16 / 826,168, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2020 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV); U.S. Provisional Patent Application No. 62 / 849,091, entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing," filed May 16, 2019 (Attorney Docket No. ILLM1011-1 / IP-1750-PRV); U.S. Provisional Patent Application No. 62 / 849,132, entitled "Base Calling Using Convolutions," filed May 16, 2019 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2); U.S. Provisional Patent Application No. 62 / 849,133, entitled "Base Calling Using Compact Convolutions," filed May 16, 2019 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3); U.S. Provisional Patent Application No. 62 / 979,384, entitled "Artificial Intelligence-Based Base Calling of Index Sequences," filed February 20, 2020 (Attorney Docket No. ILLM1015-1 / IP-1857-PRV); U.S. Provisional Patent Application No. 62 / 979,414, entitled "Artificial Intelligence-Based Many-To-Many Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1016-1 / IP-1858-PRV); U.S. Provisional Patent Application No. 62 / 979,385, entitled "Knowledge Distillation-Based Compression of Artificial Intelligence-Based Base Caller," filed February 20, 2020 (Attorney Docket No. ILLM1017-1 / IP-1859-PRV); U.S. Provisional Patent Application No. 62 / 979,412, entitled "Multi-Cycle Cluster Based Real Time Analysis System," filed February 20, 2020 (Attorney Docket No. ILLM1020-1 / IP-1866-PRV); U.S. Provisional Patent Application No. 62 / 979,411, entitled "Data Compression for Artificial Intelligence-Based Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1029-1 / IP-1964-PRV); and U.S. Provisional Patent Application No. 62 / 979,399, entitled "Squeezing Layer for Artificial Intelligence-Based Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1030-1 / IP-1982-PRV). [Background technology]

[0004] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or problems associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to embodiments of the claimed technology.

[0005] Various protocols in biological or chemical research involve conducting multiple controlled reactions on a localized support surface or within a predetermined reaction chamber. The desired reactions can then be observed or detected, and subsequent analysis can help identify or clarify the properties of the chemicals involved in the reactions. For example, in some multiplex assays, an unknown analyte bearing a distinguishable label (e.g., a fluorescent label) can be exposed to thousands of known probes under controlled conditions. Each known probe can be deposited into a corresponding well of a microplate. Observing any chemical reactions that occur between the known probe and the unknown analyte in the well can help identify or clarify the properties of the analyte. Other examples of such protocols include known DNA sequencing processes, such as sequencing-by-synthesis or circular array sequencing. In circular array sequencing, a high-density array of DNA features (e.g., template nucleic acids) is sequenced through repeated cycles of enzymatic manipulation. After each cycle, an image is captured and can subsequently be analyzed with other images to determine the sequence of the DNA features.

[0006] As a more specific example, one known DNA sequencing system uses a pyrosequencing process and includes a chip with a fused fiber optic faceplate containing millions of wells. A single capture bead containing clonally amplified sstDNA from a genome of interest is deposited in each well. After the capture bead is deposited in the well, nucleotides are added sequentially to the well by flowing a solution containing specific nucleotides along the faceplate. The environment within the well is such that if a nucleotide flowing through a particular well complements the DNA strand on the corresponding capture bead, the nucleotide is added to the DNA strand. The colonies of DNA strands are called clusters. The incorporation of nucleotides into the clusters initiates a process that ultimately generates a chemiluminescent signal. The system includes a CCD camera positioned directly adjacent to the faceplate and configured to detect the optical signal from the DNA clusters in the wells. Subsequent analysis of images obtained throughout the pyrosequencing process allows the sequence of the genome of interest to be determined.

[0007] However, the pyro-sequencing system described above, in addition to other systems, may have certain limitations. For example, the fiber optic faceplate is acid-etched, creating millions of tiny wells. While the wells may be approximately spaced apart, it is difficult to know the exact location of a well relative to other adjacent wells. If a CCD camera were positioned directly adjacent to the faceplate, the wells would not be evenly distributed along the CCD camera's pixels, and thus would not be aligned with the pixels in a known manner. Spatial crosstalk, or well-to-well crosstalk between adjacent wells, makes it difficult to distinguish true optical signals from the well of interest from other unwanted optical signals in subsequent analysis. Furthermore, fluorescence emission is substantially isotropic. As analyte density increases, it becomes increasingly difficult to manage or account for unwanted emissions (e.g., crosstalk) from neighboring analytes. As a result, data recorded during a sequencing cycle must be carefully analyzed.

[0008] Base calling accuracy is crucial for high-throughput DNA sequencing and downstream analyses such as read mapping and genome assembly. Spatial crosstalk between neighboring clusters explains a large portion of sequencing error. Therefore, correcting for spatial crosstalk in cluster intensity data presents an opportunity to reduce DNA sequencing error and improve base calling accuracy. Summary of the Invention [Means for solving the problem]

[0009] One aspect of the invention provides a computer-implemented method for base calling, the method including accessing an image, wherein pixels of the image indicate intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, selecting a lookup table including pixel coefficients configured to maximize signal-to-noise ratio, convolving the pixel coefficients with intensity values ​​of the pixels in the image to generate an output, and base calling the target cluster based on the output.

[0010] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings may also be available in PAIR (patent application information retrieval) via the Supplemental Content tab.

[0011] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Brief explanation of the drawings]

[0012] [Figure 1A] 1 illustrates one embodiment of generating a lookup table (LUT) / equalizer filter by training an equalizer. [Figure 1B] 10 shows one embodiment using the LUT / equalizer filter of FIG. 1 to attenuate spatial crosstalk from sensor pixels and basecall clusters using the crosstalk-corrected sensor pixels. [Figure 2] Visualize an example of a sequencing image containing the centers / point sources of at least five clusters / wells on the flow cell. [Figure 3] This is a visualization of an example of extracting a pixel patch (yellow) from the sequencing image in Figure 2, such that the center of target cluster 1 (blue) is contained in the central pixel of the pixel patch. [Figure 4] Visualize an example of cluster-to-pixel signal. [Figure 5] Visualize an example of signal overlap from cluster to pixel. [Figure 6] An example of a cluster signal pattern is visualized. [Figure 7] 4 visualizes an example of a sub-pixel LUT grid used to attenuate spatial crosstalk from the pixel patches of FIG. 3. [Figure 8] 1C illustrates the selection of a LUT / equalizer filter from the LUT bank of FIG. 1B based on the sub-pixel location of the cluster / well center within the pixel. [Figure 9] 1 shows an embodiment in which the center of target cluster 1 (blue) is not substantially concentric with the center of the pixel. [Figure 10] 1 illustrates one embodiment for interpolating between a set of selected LUTs and generating respective LUT weights. [Figure 11] 10 shows a weight kernel generator that uses the calculated weights of LUTs 12, 7, 8, and 13 to generate a weight kernel. [Figure 12]10 shows an element-wise multiplier that multiplies the interpolated pixel coefficients of a weight kernel by the intensity values ​​of the pixels in a pixel patch element-wise and sums the intermediate products of the multiplication to produce an output. [Figure 13A] Examples of coefficients of LUT12, 7, 8, and 13 are shown below. [Figure 13B] Examples of coefficients of LUT12, 7, 8, and 13 are shown below. [Figure 13C] Examples of coefficients of LUT12, 7, 8, and 13 are shown below. [Figure 13D] Examples of coefficients of LUT12, 7, 8, and 13 are shown below. [Figure 13E] Examples of coefficients of LUT12, 7, 8, and 13 are shown below. [Figure 13F] Examples of coefficients of LUT12, 7, 8, and 13 are shown below. [Figure 14A] An example of a weight kernel is shown below. [Figure 14B] 10 shows an example of weight kernel generation logic used by the weight kernel generator to generate weight kernels from the calculated weights of LUTs 12, 7, 8, and 13. [Figure 14C] 10 shows an example of weight kernel generation logic used by the weight kernel generator to generate weight kernels from the calculated weights of LUTs 12, 7, 8, and 13. [Figure 15A] We show how the interpolated pixel coefficients of the weight kernel maximize the signal-to-noise ratio and recover the underlying signal of target cluster 1 from the signal corrupted by crosstalk from clusters 2, 3, 4, and 5. [Figure 15B] We show how the interpolated pixel coefficients of the weight kernel maximize the signal-to-noise ratio and recover the underlying signal of target cluster 1 from the signal corrupted by crosstalk from clusters 2, 3, 4, and 5. [Figure 16] 10 shows one implementation of a base-by-base Gaussian fit centered around a base-by-base intensity target that is used as a ground truth value for error calculation during training. [Figure 17]A computer system that can be used to implement the disclosed techniques. [Figure 18] 1 illustrates one embodiment of an adaptive equalization technique that can be used to train an equalizer. [Figure 19A] 1 illustrates various performance metrics of the disclosed technology. [Figure 19B-1] 1 illustrates various performance metrics of the disclosed technology. [Figure 19B-2] 1 illustrates various performance metrics of the disclosed technology. [Figure 19C] 1 illustrates various performance metrics of the disclosed technology. [Figure 19D] 1 illustrates various performance metrics of the disclosed technology. DETAILED DESCRIPTION OF THE INVENTION

[0013] The following description is typically made with reference to specific structural embodiments and methods. It is not intended to limit the present technology to the specifically disclosed embodiments and methods, but it should be understood that the present technology can be implemented using other features, elements, methods, and embodiments. Preferred embodiments are described to illustrate the present technology, not to limit the scope defined by the claims. Those skilled in the art will recognize various equivalent variations to the following description.

[0014] Lookup table generation 1 illustrates one embodiment for generating a look-up table (LUT) (or LUT bank) 106 by training an equalizer 104. The equalizer 104 is also referred to herein as an equalizer-based basecaller 104. The system 100A includes a trainer 114 that trains the equalizer 104 using least squares estimation. Additional details regarding equalizers and least squares estimation are provided in appendices included with this application.

[0015] The sequencing image 102 is generated during a sequencing run performed by a sequencing instrument such as Illumina's iSeq, HiSeqX, HiSeq3000, HiSeq4000, HiSeq2500, NovaSeq6000, NextSeq550, NextSeq1000, NextSeq2000, NextSeqDx, MiSeq, and MiSeqDx. In one embodiment, Illumina sequencers use cyclic reversible termination (CRT) chemistry for base calling. This process relies on extending a nascent strand complementary to a template strand with fluorescently labeled nucleotides while tracking the emission signal of each newly added nucleotide. The fluorescently labeled nucleotides have a 3' removable block that anchors the fluorophore signal of the nucleotide.

[0016] Sequencing is performed in iterative cycles, each of which includes three steps: (a) extending the emerging strand by adding fluorescently labeled nucleotides; (b) exciting fluorophores using one or more lasers in the sequencing instrument's optical system and generating a sequencing image by imaging through different filters in the optical system; and (c) cleaving the fluorophores and removing the 3' block in preparation for the next sequencing cycle. The capture and imaging cycles are repeated for a specified number of sequencing cycles, defining the read length. Using this approach, each cycle locates a new position along the template strand.

[0017] The enormous power of Illumina sequencers stems from their ability to simultaneously perform and sense CRT reactions of millions or even billions of samples (e.g., clusters). Clusters contain approximately 1,000 identical copies of a template strand, but clusters vary in size and shape. Clusters are grown from template strands by bridge amplification or exclusion amplification of the input library prior to a sequencing run. The purpose of amplification and cluster extension is to increase the intensity of the emitted signal, since imaging devices cannot reliably sense the fluorophore signal of a single strand. However, because the physical distance between strands within a cluster is small, imaging devices perceive a cluster of strands as a single spot.

[0018] Sequencing occurs in a flow cell, a small glass slide that holds the input strand. The flow cell is connected to an optical system that includes microscope imaging, an excitation laser, and a fluorescence filter. The flow cell contains multiple chambers, called lanes. The lanes are physically separated from one another and can contain distinct tagged sequencing libraries without cross-contamination of samples. In some embodiments, the flow cell comprises a patterned surface. A "patterned surface" refers to the arrangement of distinct regions within or on an exposed layer of a solid support. For example, one or more regions can be features in which one or more amplification primers are present. The features can be separated by interstitial regions in which amplification primers are absent. In some embodiments, the pattern can be an xy format of features in rows and columns. In some embodiments, the pattern can be a repetitive sequence of features and / or interstitial regions. In some embodiments, the pattern can be a random sequence of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Pat. No. 8,778,849, U.S. Pat. No. 9,079,148, U.S. Pat. No. 8,778,848, and U.S. Patent Application Publication No. 2014 / 0243224, each of which is incorporated herein by reference.

[0019] In some embodiments, the flow cell comprises an array of wells or depressions in a surface, which can be fabricated as commonly known in the art using a variety of techniques, including but not limited to photolithography, stamping techniques, molding techniques, and microetching techniques. As understood in the art, the technique used will depend on the composition and shape of the array substrate.

[0020] The features within the patterned surface may be wells in an array of wells (e.g., microwells or nanowells) on glass, silicon, plastic, or other suitable solid support with a patterned covalently attached gel, such as poly(N-(5-azidoacetamylpentyl)acrylamide-co-acrylamide) (PAZAM; see, e.g., U.S. Patent Application Publication Nos. 2013 / 184796, WO 2016 / 066586, and 2015-002813, each of which is incorporated by reference in its entirety). This process creates a gel pad used for sequencing, which can be stable over many cycles of sequencing operations. Covalently attaching a polymer to the wells is useful for maintaining the gel in the structured features throughout the life of the structured substrate during various applications. However, in many embodiments, the gel need not be covalently attached to the wells. For example, in some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Pat. No. 8,563,477, incorporated herein by reference in its entirety) that is not covalently bonded to any part of the structured matrix can be used as the gel material.

[0021] In certain other embodiments, structured substrates can be fabricated by patterning a solid support material with wells (e.g., microwells or nanocells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof), such as an azido-SFA version, and polishing the gel-coated support, e.g., by chemical or mechanical polishing, thereby retaining the gel within the wells but removing or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the wells. Primer nucleic acids can be attached to the gel material. A solution of target nucleic acids (e.g., a fragmented human genome) can then be contacted with the polished substrate such that individual target nucleic acids are seeded into individual wells through interaction with primers attached to the gel material. Due to the absence or inactivity of the gel material, the target nucleic acids do not occupy the interstitial regions. Amplification of the target nucleic acids will be confined to the wells because the absence or inactivity of gel in the interstitial regions prevents outward migration of growing nucleic acid colonies. The process is manufacturable, scalable, and utilizes conventional micro- or nano-fabrication methods.

[0022] The sequencing instrument's imaging device (e.g., a solid-state imager such as a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) sensor) takes snapshots at multiple locations along the lane in a series of non-overlapping regions called tiles. For example, there may be 64 or 96 tiles per lane. A tile holds hundreds of thousands to millions of clusters.

[0023] The output of a sequencing run are sequencing images, each showing the intensity radiation of a cluster and its surrounding background. Sequencing images show the intensity radiation generated as a result of incorporating nucleotides into a sequence during sequencing. The intensity radiation comes from the associated specimens / clusters and their surrounding background.

[0024] Sequencing images 102 are provided from multiple sequencing instruments, sequencing runs, cycles, flow cells, tiles, wells, and clusters. In one embodiment, sequencing images are processed by equalizer 104 on an imaging channel basis. A sequencing run generates m images per sequencing cycle corresponding to m imaging channels. In one embodiment, each imaging channel corresponds to one of multiple filter wavelength bands. In another embodiment, each imaging channel corresponds to one of multiple imaging events in a sequencing cycle. In yet another embodiment, each imaging channel corresponds to a combination of illumination by a specific laser and imaging through a specific optical filter. In different embodiments, such as 4-channel chemistry, 2-channel chemistry, and 1-channel chemistry, m is 4 or 2. In other embodiments, m is 1, 3, or greater than 4.

[0025] In another embodiment, the input data is based on pH changes induced by the release of hydrogen ions during molecular elongation. The pH change is detected and converted to a voltage change proportional to the number of incorporated bases (e.g., in the case of Ion Torrent). In yet another embodiment, the input data is constructed from nanopore sensing, which uses a biosensor to measure the disruption of current as an analyte passes through or near the opening of the nanopore. For example, Oxford Nanopore Technologies (ONT) sequencing is based on the following concept: a single strand of DNA (or RNA) is passed through a membrane via a nanopore, and a potential difference is applied across the membrane. Nucleotides present within the pore affect the electrical resistance of the pore, so that current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal (due to its "squish" appearance when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer Data Acquisition (DAC) values ​​taken at a frequency of 4 kHz (e.g., 1 kHz). With a DNA strand velocity of ~450 base pairs per second, this gives, on average, about 9 raw observations per base. This signal is then processed to identify breaks in the aperture signal that correspond to individual reads. Extension of these raw signals is base called, a process that converts DAC values ​​into sequences of DNA bases. In some embodiments, the input data includes normalized or scaled DAC values.Additional information regarding non-image-based sequencing data can be found in U.S. Provisional Patent Application No. 62 / 849,132, entitled "Base Calling Using Convolutions," filed May 16, 2020 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2), U.S. Provisional Patent Application No. 62 / 849,133, entitled "Base Calling Using Compact Convolutions," filed May 16, 2019 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3), and U.S. Non-Provisional Patent Application No. 16 / 826,168, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2019 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV).

[0026] training The equalizer 104 generates a LUT bank having multiple LUTs (equalizer filters) 106 with sub-pixel resolution. In one embodiment, the number of LUTs 106 generated by the equalizer 104 for the LUT bank depends on the number of sub-pixels into which the sensor pixels of the sequenced image 102 are divided or can be divided. For example, if the sensor pixels of the sequenced image 102 are each divisible into n×n sub-pixels (e.g., 5×5 sub-pixels), the equalizer 104 generates n 2 LUTs 106 (for example, 25 LUTs) are generated.

[0027] In one training implementation, data from the sequencing image is binned by well subpixel location. For example, for a 5x5 LUT, the center of 1 / 25th of the well is in bin (1,1) (e.g., the upper left corner of the sensor pixel), the 1 / 25th of the well is in bin (1,2), and so on. The equalizer coefficients for each well-center-bin are determined using least-squares estimation on the subset of data from wells that fall within each bin. The input to the equalizer 104 is the raw sensor pixels of the sequencing image for those bins. The resulting estimated equalizer coefficients vary from bin to bin.

[0028] Each LUT has multiple coefficients learned from training. In one embodiment, the number of coefficients in the LUT corresponds to the number of sensor pixels used to base call the clusters. For example, if the local grid of sensor pixels (image or pixel patches) used to base call the clusters is of size p×p (e.g., a 9×9 pixel patch), then each LUT has p 2 has a coefficient of (e.g., a coefficient of 81).

[0029] Training generates equalizer coefficients configured to mix / combine pixel intensity values ​​representing intensity radiation from a base-called target cluster and intensity radiation from one or more neighboring clusters to maximize the signal-to-noise ratio. The signal that maximizes the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that minimizes the signal-to-noise ratio is the intensity radiation from neighboring clusters, i.e., spatial crosstalk, plus some random noise (e.g., to account for background intensity radiation). The equalizer coefficients are used as weights, and mixing / combining involves performing element-wise multiplications between the equalizer coefficients and the pixel intensity values ​​to calculate a weighted sum of the pixel intensity values.

[0030] During training, the equalizer 104, according to one embodiment, learns to maximize the signal-to-noise ratio through least-squares estimation. Using least-squares estimation, the equalizer 104 is trained to estimate shared equalizer coefficients from the pixel intensities around the well of interest and the desired output. Least-squares estimation is suitable for this purpose because it minimizes the squared error and outputs coefficients that account for the effects of noise amplification.

[0031] The desired output is an impulse at the well location (point source) when the intensity channel is on, and a background level when the intensity channel is off. In some embodiments, ground truth base calls 112 are used to generate the desired output. In some embodiments, the ground truth base calls 112 are modified to account for DC offset per well, amplification factor, degree of polyclonality, and gain offset parameters included in the least squares estimate. In one embodiment, during training, a DC offset, i.e., a fixed offset, is calculated as part of the least squares estimate. During inference, the DC offset is added as a bias to each equalizer calculation.

[0032] In one embodiment, the desired output is estimated using Illumina's Real-time Analysis (RTA) base caller, which does not use an equalizer. Details regarding RTA can be found in U.S. Patent Application No. 13 / 006,206, which is incorporated by reference as if fully set forth herein. The RTA base caller is used to generate ground truth base calls 112. This is because RTA has a low base calling error rate. Base calling errors are averaged over many training examples. In another embodiment, the ground truth base calls 112 are sourced using aligned genomic data, which has better quality because it can use reference genomes and truth information that incorporate knowledge gained from multiple sequencing platforms and sequencing runs to average out noise.

[0033] The ground truth base calls 112 are base-specific intensity values ​​that reliably represent the intensity profiles of bases A, C, G, and T, respectively. A base caller, such as RTA, base calls clusters by processing the sequencing image 102 and generating per-color intensity values / output for each base call. The per-color intensity values ​​can be considered per-base intensity values ​​because, depending on the type of chemistry (e.g., two-color chemistry or four-color chemistry), a color is mapped to each of the bases A, C, G, and T. The base with the closest intensity profile match is called.

[0034] Figure 16 illustrates one implementation of a per-base Gaussian fit centered on a per-base target that is used as a ground truth value for error calculations during training. The per-base intensity outputs generated by the base caller for a large number of base calls in the training data (e.g., tens, hundreds, thousands, or millions of base calls) are used to generate a per-base intensity distribution. Figure 16 illustrates a chart of four Gaussian clouds, which are probability distributions of the per-base intensity outputs for bases A, C, G, and T, respectively. The intensity values ​​at the centers of the four Gaussian clouds are used as ground truth intensity targets given the ground truth base calls 112 for bases A, C, G, and T, respectively, and are referred to herein as intensity targets.

[0035] Consider that during training, the input image data fed to equalizer 104 is annotated with the base "A" as the ground truth base call. Then, the target / desired output of equalizer 104 is the intensity value at the center of the green cloud in FIG. 16, i.e., the intensity target for base A. Similarly, for a ground truth base call of base "C," the desired output of equalizer 104 is the intensity value at the center of the blue cloud in FIG. 16, i.e., the intensity target for base C. Thus, the target or desired output during training of equalizer 104 is the average intensity for each base A, C, G, and T after averaging over the training data. In one embodiment, trainer 114 uses least-squares estimation to fit the coefficients of equalizer 104 to minimize the equalizer output error to these intensity targets.

[0036] In one implementation, during training, the equalizer 104 applies coefficients in a given look-up table (LUT) to pixels of a sequenced image labeled with a given base. This involves element-wise multiplying the coefficients with the pixel's intensity values ​​to generate a weighted sum of the intensity values, with the coefficients acting as weights. The weighted sum becomes the predicted output of the equalizer 104. Then, based on a cost / error function (e.g., sum of squared errors (SSE)), the error (e.g., least square error, minimum mean square error) between the weighted sum and an intensity target determined for the given base (e.g., as the average intensity observed at the given base from the center of the corresponding intensity Gaussian fit) is calculated. A cost function such as the SSE is a differentiable function used to estimate the equalizer coefficients using an adaptive approach, where the derivative of the error with respect to the coefficients can be evaluated and these derivatives are used to update the coefficients with values ​​that minimize the error. This process is repeated until the updated coefficients no longer reduce the error. In another embodiment, batch least squares is used to train the equalizer 104 .

[0037] In other embodiments, the per-base intensity distribution / Gaussian cloud shown in Figure 16 can be generated for each well and noise corrected by adding a DC offset, amplification factor, and / or phase parameter. In this way, depending on the well location of a particular well, the corresponding per-base Gaussian cloud can be used to generate a target intensity value for that particular well.

[0038] In one implementation, a bias term is added to the dot product that generates the output of the equalizer 104. During training, the bias parameter can be estimated using a similar approach used to learn the equalizer coefficients, i.e., least squares or least mean squares (LMS). In some implementations, the value of the bias parameter is a constant value equal to 1, i.e., a value that does not change with input pixel intensity. There is one bias per equalizer coefficient set. The bias is learned during training and then fixed for use during inference. The learned bias, along with the learning coefficient for each LUT, represents a DC offset used in all equalizer calculations during inference. This bias accounts for random noise caused by different cluster sizes, different background intensities, varying stimulus responses, varying focus, varying sensor sensitivity, and varying lens aberrations.

[0039] In yet another decision-directed implementation, the output of the equalizer 104 is assumed to be correct for training purposes.

[0040] In another embodiment of training, the equalizer 104 generates only a single LUT (equalizer filter) for a bin, and then uses multiple per-bin interpolation filters 108 to generate the remaining equalizer filters for the remaining bins. In this embodiment, the sensor pixels around every well for every training example are resampled / interpolated to a well-aligned space (i.e., the wells are centered in their respective pixel patches / local grids). The resampled pixels for all examples are then consistently aligned across all wells.

[0041] However, to apply the single equalizer filter generated by the equalizer 104 in a practical online system for base calling, the raw sensor pixels of the sequencing image must be preprocessed to return them to a well-aligned space. That is, interpolation must be performed on the raw pixels surrounding each well, with the interpolation parameters varying depending on the subpixel location of a given well. To avoid this interpolation process, the overall response for a given well subpixel location is precalculated. By interpolating the raw pixel intensities into the well-aligned pixel space, well-aligned equalizer input values ​​are calculated. The interpolated response and the equalizer response are convolved together to reduce computation. Because the interpolation filter varies with subpixel well location, this results in a different set of equalizer coefficients / filters for each subpixel well location, thereby generating the remaining LUTs for the remaining bins. Therefore, in this training implementation, while only the coefficients of a single equalizer filter are trained during training, the precalculation process generates a bank of LUT-based equalizers by applying the bin-specific interpolation filters 108 together with the single equalizer filter, where the LUT index is the subpixel well location.

[0042] The trainer 114 can train the equalizer 104, train multiple trainers, and generate training coefficients for the LUT 106. Examples of training techniques include least squares estimation, least squares, least mean squares, and recursive least squares. In least squares, the parameters of a function are adjusted to best fit a data set so that the sum of squared residuals is minimized. For more information about least squares estimation algorithms, see "Ordinary Least Squares," https: / / en.wikipedia.org / w / index.php?title=Least_squares&oldid=951737821 (last visited April 28, 2020), which is incorporated by reference as if fully set forth herein. Ordinary least squares is a type of least squares method for estimation in linear regression models. For more information about least squares algorithms, see "Ordinary least squares," https: / / en.wikipedia.org / w / index.php?title=Ordinary_least_squares&oldid=951770366 (last visited April 28, 2020), which is incorporated by reference as if fully set forth herein. In other embodiments, other estimation and adaptive equalization algorithms can be used to train equalizer 104.

[0043] The equalizer 104 can be trained in an offline mode, in which, according to one embodiment, the trained coefficients of the LUT 106 are generated using the following batch least squares equalization logic:

[0044]

number

[0045] In the above equation, the LUT coefficients are the beta hats, the pixel intensities are X, and the target is y. A DC term is also added to the pixel intensities and a coefficient (e.g., an additional intensity term that is fixed to 1 in all cases). Next, as an example, consider X to be a matrix of size 82 (= 9 × 9 input intensities + constant DC term) × the number of training examples in the batch, and Y is the target output for each training example. That is, each value is the intensity center of an ON / OFF cloud that depends on the training example truth. The beta hat is the set of coefficients that minimizes the sum of squared residuals, and is also of size 82 (= 9 × 9 coefficients + 1 DC term).

[0046] The equalizer 104 can also be trained in online mode, adapting the coefficients of the LUT 106 to track changes in temperature (e.g., optical distortion), focus, chemistry, machine-specific variations, etc., on a tile-by-tile or subtile-by-subtile basis while the sequencer is operating and sequencing runs are cyclically progressing. In online mode, the trained coefficients of the LUT 106 are generated using adaptive equalization. In online mode, the least mean squares method, a form of stochastic gradient descent, is used as the training algorithm. For more information about the least mean squares algorithm, see "Least Mean Squares Filter," https: / / en.wikipedia.org / w / index.php?title=Least_mean_squares_filter&oldid=941899198 (last visited April 28, 2020), which is incorporated by reference as if fully set forth herein.

[0047] The least mean squares method uses the gradient of the squared error for each coefficient to move the coefficient in the direction that minimizes a cost function, which is the expected value of the squared error. This has a very low computational cost, with only multiplication and accumulation operations per coefficient being performed. No long-term storage is required, except for the coefficients. The least mean squares method is suitable for processing large amounts of data (e.g., processing data from billions of clusters in parallel). Extensions of the least mean squares method include normalized least mean squares and frequency-domain least mean squares, which can also be used here. In some embodiments, the least mean squares method can be applied in a decision-directed manner, where our decisions are assumed to be correct, i.e., our error rate is very low and small μ values ​​filter out impeded updates due to inaccurate base calls.

[0048] Figure 18 shows one embodiment of an adaptive equalization technique that can be used to train the equalizer 104. Here, the equalization logic is y = x.h + d, where x is the input pixel intensity, h is the equalizer coefficient, and d is the DC offset. In one embodiment, x and h are row and column vectors, respectively, with length 81. This vector model corresponds to the dot product of a 9x9 matrix representing the input pixel and the coefficients. The cost is the expected squared error. The gradient update moves each coefficient in the direction that reduces the expected squared error, which leads to the next update.

[0049]

number

[0050] where h is a vector of equalizer coefficients (e.g., 9x9 equalizer coefficients), x is a vector of equalizer input intensities (e.g., 9x9 pixels in a pixel patch), and e is the error of the equalizer calculation performed using 81 values ​​of x, i.e., only one error term per equalizer output.

[0051] Applying this update produces new estimates of the 9x9 equalizer coefficients. This estimate moves the equalizer coefficients (on average) in the direction that reduces the mean squared error (MSE). 81 updates are made, one for each equalizer coefficient. In some implementations, Mu is a small constant used to change the adaptation rate / convergence speed. The DC term update can be calculated in a similar manner. The gain term update can be calculated in a similar manner.

[0052] Coefficient sets can be shared, for example, between tiles, regions of tiles, or flow cell surfaces by saving and restoring coefficient sets when the input data changes.

[0053] In some implementations, since linear interpolation is applied to the coefficient set, the updates are applied slightly differently in the following manner. h(q,n+1)=h(q,n)+lambda_q.mu.x(n).e(n)

[0054] where h(q,n) is the weight q at cycle n and lambda_q is the linear interpolation weight for a particular set of coefficients, which may include four updates per equalizer output by linear interpolation in two dimensions.

[0055] Recursive least squares is an extension of least squares to a recursive algorithm. For more information about the recursive least squares algorithm, see "Recursive least squares filter", https: / / en.wikipedia.org / w / index.php?title=Recursive_least_squares_filter&oldid=916406502 (last visited April 28, 2020), which is incorporated by reference as if fully set forth herein.

[0056] In multi-domain embodiments, the LUTs 106 and their trained coefficients can be generated along multiple domains. Examples of domains include sequencers or sequencing instruments / machines (e.g., Illumina's NextSeq, MiSeq, HiSeq and their respective models), sequencing protocols and chemistries (e.g., blind amplification, exclusion amplification), sequencing runs (e.g., forward and reverse), sequencing illumination (e.g., structured, unstructured, angled), sequencing equipment (e.g., overhead CCD camera, underlying CMOS sensor, one laser, multiple lasers), imaging technology (one channel, two channels, four channels), flow cells (e.g., patterned, unpatterned, embedded in a CMOS chip, underlying CCD camera), and spatial resolution on the flow cell (e.g., different regions within the flow cell). or quadrants (e.g., different tiles on a flow cell (e.g., in the case of edge wells on tiles closer to a laser or camera or fluidic system)) and different regions within a tile (e.g., different lanes on a tile (e.g., in the case of edge wells on lanes closer to a laser or camera or fluidic system)). Those skilled in the art will understand that other selectable domains and parameters typically associated with sequencing (e.g., image processing algorithms, image registration algorithms, ground truth annotation schemes (e.g., continuous labels like intensity values, hard labels like one-hot coding, soft labels like softmax scores), temperature, focus, lenses, sequencing reagents, sequencing buffers) are similarly included.

[0057] A separate and distinct training set can be created for each domain using the sequencing images generated using each domain. The discrete training sets can be used to train the equalizer 104 and generate a LUT with trained coefficients for the corresponding regions. The training coefficients specifically trained for each domain in the multiple domains can be stored and accessed during online mode depending on which domain or combination of domains is being used in a current or ongoing sequencing operation. For example, a first coefficient set more suitable for the edge wells of a flow cell can be used for a sequencing operation, along with a second coefficient set more suitable for the center wells of the same flow cell.

[0058] In one implementation, a configuration file can specify different combinations of domains and can be analyzed during online mode to select different sets of coefficients specific to the domains identified by the configuration file.

[0059] In a multiple-training embodiment, the equalizer 104 undergoes not only training but also pre-training. That is, the LUT 106 and its coefficients are first trained in a pre-training stage using a first training technique, and then re-trained or further trained in a further training stage using a second training technique. The first and second training techniques can be any of the training techniques described above. The pre-training and second training techniques can be the same or different. For example, the pre-training stage can be an offline mode using a batch least-squares training technique, and the training stage can be an online mode using an iterative least-mean-squares technique.

[0060] In some embodiments, multi-domain and multi-training implementations can be combined such that domain-specific coefficients are pre-trained and then further trained in a domain-specific manner. That is, further training (e.g., in an online mode) retrains the coefficients for that particular domain using only data that represents that particular domain and that is similar to the data used in the pre-training stage. In other knowledge transfer embodiments, pre-training and training can use training data from the entire domain; for example, a coefficient set is generated during pre-training using images from a patterned flow cell, but is retrained during a subsequent training stage using images from an unpatterned flow cell.

[0061] Spatial Crosstalk Attenuator 2 illustrates one embodiment that uses the trained LUT / equalizer filter 106 of FIG. 1 to attenuate spatial crosstalk from sensor pixels and reduce call clusters to base calls using the crosstalk-corrected sensor pixels. The trained equalizer base caller 104 operates during the inference stage when base calls are made. In some embodiments, the actions illustrated in FIG. 2 are performed in a preprocessing stage before the base calling stage to generate crosstalk-corrected image data used by the base caller for base calling.

[0062] In one embodiment, equalizer coefficients are applied to pixel patches 120 (image patches or local grids of sensor pixels) extracted from the sequencing image 116 on an imaging channel basis and a target cluster basis. Regarding the imaging channel basis, in some embodiments, each sequencing image has image data from multiple imaging channels. Consider the optical system of an Illumina sequencer that uses two different imaging channels, namely a red channel and a green channel. Then, in each sequencing cycle, the optical system generates a red image with red channel intensities and a green image with green channel intensities, which together form a single sequencing image (like the RGB channels of a typical color image).

[0063] During training, the coefficients are trained / configured to maximize the signal-to-noise ratio (SNR) by minimizing the error between the predicted / estimated output and the desired / actual output. An example of the error is the mean squared error (MSE) or mean squared deviation (MSD). The signal maximized in the signal-to-noise ratio is the intensity emission from the base-called target cluster (e.g., the cluster at the center of the image patch), and the noise minimized in the signal-to-noise ratio is the intensity emission from one or more neighboring clusters, i.e., spatial crosstalk, plus other noise sources (e.g., to account for background intensity emission). The trained coefficients are element-wise multiplied to the pixels of the image patch to calculate a weighted sum of the pixel intensity values. The weighted sum is then used to base-call the target cluster.

[0064] In one embodiment, patch extractor 118 extracts red pixel patches from the red channel and green pixel patches for the green channel from a single sequencing image. In another embodiment, red pixel patches are extracted from the red sequencing image of a target sequencing cycle, and green pixel patches are extracted from the green sequencing image of a target sequencing cycle. The coefficients of LUT 106 are used to generate red weighted sums for the red pixel patches and green weighted sums for the green pixel patches. Both the red and green weighted sums are then used to base call the target cluster. Pixel patch 120 has dimensions w×h, where w (width) and h (height) are any number between 1 and 10,000 (e.g., 3×3, 5×5, 7×7, 9×9, 15×15, 25×25). In some embodiments, w and h are the same. In other embodiments, w and h are different. Those skilled in the art will understand that one, two, three, four, or more channels or images of data can be generated per sequencing cycle for a target cluster, and one, two, three, four, or more patches can be extracted, respectively, to generate one, two, three, four, or more weighted sums, respectively, for base calling the target cluster.

[0065] For target cluster-based extraction of pixel patches 120 from sequencing image 116, pixel extractor 118 extracts pixel patches 120 based on the location of the cluster / well centers on sequencing image 116 such that the center pixel of each extracted pixel patch contains the center of the target cluster / well. In some embodiments, patch extractor 118 locates the cluster / well centers on the sequencing image, identifies the pixel in the sequencing image that contains the cluster / well center (i.e., the center pixel), and extracts pixel patches of contiguous neighboring pixels around the center pixel.

[0066] Figure 2 visualizes an example of a sequencing image 200 containing central / point sources of at least five clusters / wells on a flow cell. The pixels of sequencing image 200 show intensity emissions from target cluster 1 (blue) and additional neighboring clusters 2 (purple), 3 (orange), 4 (brown), and 5 (green).

[0067] Figure 3 visualizes an example of extracting pixel patch 300 (yellow) from sequencing image 200 such that the center of target cluster 1 (blue) is contained within central pixel 206 of pixel patch 300. Figure 3 also shows other pixels 202, 204, 214, and 216 that contain the centers of neighboring clusters 2 (purple), 3 (orange), 4 (brown), and 5 (green), respectively.

[0068] Figure 4 visualizes an example of a cluster-to-pixel signal 400. In one embodiment, the sensor pixels (yellow) are in the pixel plane. Spatial crosstalk is caused by clusters 412 periodically distributed in the sample plane (e.g., a flow cell). In one embodiment, the target cluster and additional neighboring clusters are periodically distributed in a diamond shape on the flow cell and immobilized on the wells of the flow cell. In another embodiment, the target cluster and additional neighboring clusters are periodically distributed on a hexagonal flow cell and immobilized on the wells of the flow cell. Signal cones 402 from the clusters are optically coupled to a local grid of sensor pixels (e.g., pixel patch 300) via at least one lens (e.g., one or more lenses of an overhead or adjacent CCD camera).

[0069] In addition to diamonds and hexagons, the clusters can be arranged in other regular shapes, such as squares, diamonds, triangles, etc. In still other embodiments, the clusters are arranged on the sample plane in a random, non-periodic arrangement. One skilled in the art will understand that the clusters can be arranged on the sample plane in any arrangement, as required by the particular sequencing implementation.

[0070] 5 visualizes an example of cluster-to-pixel signal overlap 500. Signal cones 402 overlap and impinge on sensor pixels, creating spatial crosstalk 502.

[0071] 6 visualizes an example of a cluster signal pattern 600. In one implementation, the cluster signal pattern 600 follows a decay pattern 602, where the cluster signal is strongest at the cluster center and decays as it propagates away from the cluster center.

[0072] 6 also shows an example of equalizer coefficients 604 trained / configured to maximize the signal-to-noise ratio by calculating a weighted sum of the intensity radiation from target cluster 1 and the intensity radiation from neighboring clusters 2, 3, 4, and 5. The equalizer coefficients 604 act as weights. The weighted sum is calculated by element-wise multiplying a first matrix containing the equalizer coefficients 604 by a second matrix containing pixel intensity values, where each pixel intensity value is the sum of radiation from one or more of clusters 1, 2, 3, 4, and 5, as well as other noise sources in the system as measured by the pixel sensor.

[0073] FIG. 7 visualizes an example of a subpixel LUT grid 700 used to attenuate spatial crosstalk from a pixel patch 300. Each pixel in the pixel patch 300 can be divided into multiple subpixels. In FIG. 7, pixel 206, which contains the center of target cluster 1 (blue), is divided into as many subpixels as the number of trained LUTs 106. That is, pixel 206 is divided into as many subpixels as the number of bins for which equalizer 104 generated LUTs 106 during training. As a result, each subpixel of pixel 206 corresponds to a respective LUT in the LUT bank generated by equalizer 104 using decision-directed feedback and least-squares estimation.

[0074] In the example shown in FIG. 7 , pixel 206 (center pixel) is divided into a 5×5 subpixel LUT grid 700 to generate 25 subpixels corresponding to the 25 LUTs (equalizer filters) generated by adaptive filter 104 as a result of training. Each of the 25 LUTs includes coefficients configured to mix / combine intensity values ​​of pixels within pixel patch 300 representing intensity radiation from target cluster 1 and neighboring clusters 2, 3, 4, and 5, so as to maximize the signal-to-noise ratio. The signal that maximizes the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that minimizes the signal-to-noise ratio is the intensity radiation from neighboring clusters 2, 3, 4, and 5, i.e., spatial crosstalk plus some random noise (e.g., to account for background intensity radiation). The LUT coefficients are used as weights, and mixing / combining involves performing element-wise multiplication between the LUT coefficients and the intensity values ​​of the pixels within pixel patch 300 to calculate a weighted sum of the pixel intensity values.

[0075] The number of coefficients in each of the 25 LUTs is the same as the number of pixels in pixel patch 300, i.e., a 9×9 coefficient grid in each LUT for the 9×9 pixels in pixel patch 300. This is because the coefficients are multiplied element-wise with the pixels in pixel patch 300.

[0076] In one implementation, a pixel-to-subpixel converter (not shown in FIG. 1B ) divides the pixels 206 into the subpixel LUT grid 700 based on a preset pixel divisor parameter (e.g., 1 / 5 pixel per subpixel to generate a 5×5 subpixel LUT grid 700). For example, the pixels may be divided into five subpixel bins with the following boundaries: −0.5, −0.3, −0.1, 0.1, 0.3, 0.5.

[0077] Note that in Figure 7, the center of target cluster 1 (blue) is substantially concentric with the center of transformed pixel 702. This is because sequencing image 200, and therefore pixel patch 300, is resampled so that the center of target cluster 1 (blue) is substantially concentric with the center of transformed pixel 702 by (i) aligning sequencing image 200 with respect to the template image and determining affine and nonlinear transformation parameters, (ii) using the parameters to transform the position coordinates of target cluster 1 (blue) into image coordinates of sequencing image 702, and (iii) applying interpolation using the transformed position coordinates of target cluster 1 (blue) to make its center substantially concentric with the center of transformed pixel 200. The locations of the wells in the sample plane are known and can be used to calculate where the equalizer input for a particular well is in raw pixel space. Interpolation can then be used to reconstruct the intensity at those locations from the raw image.

[0078] 8 illustrates the selection of a LUT / equalizer filter from LUT bank 106 based on the subpixel location of a cluster / well center within a pixel. Because the center of the target cluster (blue) is at a particular subpixel 12 in subpixel LUT grid 700, and because that particular subpixel 12 of pixel 206 corresponds to a LUT 12 in LUT bank 106, LUT selector 122 selects a LUT 12 and its coefficients from LUT bank 106 to apply to the pixels of pixel patch 300. Next, element-wise multiplier 134 multiplies the intensity values ​​of the pixels in pixel patch 300 by the coefficients of LUT 12 element-wise and sums the products of the multiplications to generate an output (e.g., weighted sum 136). This output is used to base call target cluster 1 (e.g., provide this output as an input to base caller 138).

[0079] Equalizer 104 implements the following equalization logic when the target cluster is substantially concentric with the center of the pixel, as discussed above with respect to FIGS.

[0080]

number

[0081] In the above equation, the well center coordinates (m,n) are integers to ensure that the wells are substantially aligned with the pixels; p(i,j) is the pixel intensity at location i,j; w(i,j) is the equalizer weight for the pixel at location i,j. i,j are summation limits operating over the range of pixels surrounding the well centered on p(m,n), e.g., -4<=i<=4, -4<=j<=4, and the output is a weighted average of the input pixels.

[0082] 9 illustrates an embodiment in which the center of target cluster 1 (blue) is not substantially concentric with the center of pixel 206 because no resampling is performed as discussed with respect to FIG. 8. In such an embodiment, interpolation is performed between a set of selected LUTs 124 to generate an interpolation LUT with interpolation coefficients. The interpolation LUT with interpolation coefficients is also referred to herein as a weight kernel 132.

[0083] First, as in Figure 8, a first LUT, i.e., LUT12, is selected that corresponds to the particular subpixel that contains the center of target cluster 1 (blue). Next, LUT selector 122 selects additional subpixel lookup tables from bank of subpixel lookup tables 106 that correspond to the subpixels that are most contiguously adjacent to the particular subpixel. In Figure 9, the most closely adjacent subpixels that contact particular subpixel 12 are subpixels 7, 8, and 13, and therefore LUTs 7, 8, and 13 are selected from LUT bank 106, respectively.

[0084] 10 illustrates one implementation for interpolating between a set of selected LUTs to generate respective LUT weights. Interpolator 126 comprises interpolation logic (e.g., linear, bilinear, or bicubic interpolation) that uses the coefficients of selected LUTs 12, 7, 8, and 13 to generate weights 128 for each of LUTs 12, 7, 8, and 13.

[0085] 13A, 13B, 13C, 13D, 13E, and 13F show example coefficients for LUTs 12, 7, 8, and 13. These figures also show examples of interpolation logic 1312, 1322, and 1332 used by interpolator 126 to calculate weights 128 for LUTs 12, 7, 8, and 13. These figures also show example weights 128 calculated for LUTs 12, 7, 8, and 13. These figures are snapshots of Excel sheets. The blue arrows and color coding in these figures are generated by Excel's Track Precedence function to indicate the interpolation logic.

[0086] FIG. 11 illustrates weight kernel generator 130 generating weight kernel 132 using weights 128 calculated for LUTs 12, 7, 8, and 13. FIG. 14A illustrates an example of weight kernel 132. FIGS. 14B and 14C illustrate example weight kernel generation logic 1402 used by weight kernel generator 130 to generate weight kernel 132 from weights 128 calculated for LUTs 12, 7, 8, and 13. Weight kernel 132 includes interpolated pixel coefficients 1412 configured to mix / combine intensity values ​​of pixels in pixel patch 300 representing intensity radiation from target cluster 1 and neighboring clusters 2, 3, 4, and 5 to maximize the signal-to-noise ratio. The signal that is maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that is minimized in the signal-to-noise ratio is the intensity radiation from neighboring clusters 2, 3, 4, and 5, i.e., spatial crosstalk, plus some random noise (e.g., to account for background intensity radiation). The interpolated pixel coefficients 1412 are used as weights, and the mixing / combining involves performing element-wise multiplications between the LUT coefficients and the intensity values ​​of the pixels in the pixel patch 300 to calculate a weighted sum of the pixel intensity values.

[0087] FIG. 12 shows the element-wise multiplier 134, which element-wise multiplies the interpolated pixel coefficients 1412 of the weight kernel 132 by the intensity values ​​of the pixels in the pixel patch 300 and sums the intermediate products 1202 of the multiplication to generate the weighted sum 136. For each well, the optical system operates on a point source (cluster intensity in the well) with a point spread function (the response of the optical system). In some embodiments, a bias is added to the operation to account for noise caused by different cluster sizes, different background intensities, varying stimulus responses, varying focus, varying sensor sensitivity, and varying lens aberrations. The captured image is a superposition of responses from all wells. The selected LUT equalizes the system response around each well to estimate the intensity of the point source from that well. That is, it processes the PSF intensity over a local neighborhood / grid of sensor pixels to estimate the intensity of the point source that generated the local grid of sensor pixels. This equalizer operation is a dot product on the sensor pixels in the local grid with the equalizer coefficients.

[0088] The equalizer 104 implements the following equalization logic when the target cluster is not substantially concentric with the center of the central pixel, as discussed above with respect to Figures 9, 10, 11, and 12. When the well is not at the center of the pixel, the output of the equalizer 104 is calculated as a function of the virtual pixel intensity p'(i,j) derived from the actual pixel intensity of the pixel in the sequencing image.

[0089]

number

[0090] In the above equation, the well center coordinates (m,n) may have a fractional part. Each "virtual" equalizer input p'(i,j) is generated by applying an interpolation filter to a pixel neighborhood. In one implementation, a windowed sinc low-pass filter h(x,y) is used for interpolation. In other implementations, other filters, such as a bilinear interpolation filter, can be used.

[0091] The virtual pixel at location (i,j) is calculated using an interpolation filter as follows:

[0092]

number

[0093] Combining equations (1) and (2), the equalizer 104 uses only the raw pixel intensities as follows:

[0094]

number

[0095] where h is fixed given the sub-pixel offsets frac(m), frac(n), u, v specify the range of pixels used for interpolation to generate the equalizer input, and i, j specify the range of virtual pixels used as input to the equalizer 104.

[0096] For a given sub-pixel offset, only the input pixel changes, not the filter or weights. Therefore, for the center of each binned sub-pixel offset, we compute a fixed set of interpolated equalizer coefficients. The output is:

[0097]

number

[0098] In the above formula, h fm,fn represents the LUT equalizer coefficients for a well with binned fractional sub-pixel offsets fm, fn, where (fm,fn) are the LUT indices.

[0099] 15A and 15B show how the weight kernel interpolated pixel coefficients 1412 maximize the signal-to-noise ratio and restore the underlying signal of target cluster 1 from the signal corrupted by crosstalk from clusters 2, 3, 4, and 5.

[0100] The weighted sum 136 is provided as an input to a base caller 138 to generate a base call 140. The base caller 138 can be a non-neural network based base caller or a neural network based base caller, examples of both of which are described in applications incorporated herein by reference, such as U.S. Patent Application Nos. 62 / 821,766 and 16 / 826,168.

[0101] In yet another embodiment, the need for interpolation is eliminated by having large LUTs, each with a large number of subpixel bins (eg, 50, 75, 100, 150, 200, 300, etc. subpixel bins per LUT).

[0102] Figure 19A shows a graph depicting base calling error rates using images from a NovaSeq sequencer. The error rate is indicated by cycles on the x-axis. 0.004 on the y-axis represents a base calling error rate of 0.4%. The error rate here is calculated after mapping and aligning the reads to the Phi-X reference, which is a high-confidence ground truth set. The blue line is the legacy base caller. The red line is the improved equalizer-based base caller 104 disclosed herein. The overall error rate is reduced by 57% at the expense of limited extra computation. The base error rate in later cycles is higher due to extra noise in the system (e.g., prephasing / phasing, cluster dimming). The improved performance in later cycles is valuable because it indicates that longer reads can be supported. Cycle-to-cycle performance variability is also significantly reduced.

[0103] 19B-1 and 19B-2 show another example of performance results of the disclosed equalizer-based base caller 104 on sequence data from a NovaSeq sequencer and a Vega sequencer. For the NovaSeq sequencer, the disclosed equalizer-based base caller 104 reduces the base calling error rate by more than 50%. For the Vega sequencer, the disclosed equalizer-based base caller 104 reduces the base calling error rate by more than 35%.

[0104] 19C shows another example of performance results of the disclosed equalizer-based base caller 104 on sequence data from a NextSeq2000 sequencer. For the NextSeq2000 sequencer, the disclosed equalizer-based base caller 104 reduces the base calling error rate by an average of 10%, not including throughput.

[0105] 19D illustrates one embodiment of the computational resources required by the disclosed equalizer-based base caller 104. As shown, the disclosed equalizer-based base caller 104 can be executed using a small number of CPU threads, ranging from 2 to 7 threads. Thus, the disclosed equalizer-based base caller 104 is a computationally efficient base caller that significantly reduces the base error rate and, therefore, can be integrated into most existing sequencers without requiring any additional computation or specialized processors such as GPUs, FPGAs, ASICs, etc.

[0106] In this application, the terms "cluster," "well," "sample," and "fluorescent sample" are used interchangeably, as a well contains a corresponding cluster / sample / fluorescent sample. As defined herein, "sample" and its derivatives are used in the broadest sense and include any sample, culture, etc. suspected of containing a target. In some embodiments, a sample contains DNA, RNA, PNA, LNA, chimeric, or hybrid forms of nucleic acid. A sample can include any biological, clinical, surgical, agricultural, air, or water sample containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, fresh-frozen, or formalin-fixed, paraffin-embedded nucleic acid sample. It is also contemplated that a sample can be derived from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples from a single individual (matched), such as a tumor sample and a normal tissue sample, or a sample from a single source containing two different forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample containing plant or animal DNA. In some embodiments, the source of nucleic acid material can include nucleic acid obtained from a newborn, such as is typically used for newborn screening.

[0107] A nucleic acid sample can contain high molecular weight material, such as genomic DNA (gDNA). A sample can contain low molecular weight material, such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another embodiment, the low molecular weight material includes enzymatically or mechanically fragmented DNA. A sample can contain cell-free circulating DNA. In some embodiments, a sample can contain nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resection, and other clinical or laboratory samples. In some embodiments, a sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some embodiments, a sample can contain nucleic acid molecules obtained from animals, such as humans or mammalian sources. In other embodiments, a sample can contain nucleic acid molecules obtained from non-mammalian sources, such as plants, bacteria, viruses, or fungi. In some embodiments, the source of the nucleic acid molecules can be an archived or extinct sample or species.

[0108] Additionally, the methods and compositions disclosed herein may be useful for amplifying nucleic acid samples with low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, a forensic sample may include nucleic acids obtained from a crime scene, from a missing persons DNA database, from a laboratory associated with a forensic investigation, or may include forensic samples obtained by law enforcement agencies, one or more military forces, or similar personnel. A nucleic acid sample may be crude DNA, including purified samples or lysates, derived from, for example, oral swabs, paper, cloth, or other substrates that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, a nucleic acid sample may contain small or fragmented portions of DNA, such as genomic DNA. In some embodiments, target sequences may be present in one or more bodily fluids, including, but not limited to, blood, sputum, plasma, semen, urine, and serum. In some embodiments, target sequences may be obtained from hair, skin, tissue samples, autopsies, or the remains of a victim. In some embodiments, nucleic acids containing one or more target sequences may be obtained from a deceased animal or human. In some embodiments, the target sequence can comprise nucleic acid obtained from a non-human source, such as a microorganism, a plant cell, or an entomological source. In some embodiments, the target sequence or the amplified target sequence is intended for human identification. In some embodiments, the present disclosure generally relates to a method for identifying characteristics of a forensic sample. In some embodiments, the present disclosure generally relates to a human identification method using one or more target-specific primers disclosed herein or one or more target-specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic sample or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.

[0109] As used herein, the term "adjacent," when used in reference to two reaction sites, means that there are no other reaction sites between the two reaction sites. The term "adjacent" can have a similar meaning when used in reference to adjacent detection paths and adjacent photodetectors (e.g., adjacent photodetectors have no other photodetectors between them). In some cases, a reaction site may not be adjacent to another reaction site, but may still be in close proximity to another reaction site. A first reaction site may be in close proximity to a second reaction site if a fluorescent emission signal from the first reaction site is detected by a photodetector associated with the second reaction site. More specifically, a first reaction site may be in close proximity to a second reaction site if the photodetector associated with the second reaction site detects, for example, crosstalk from the first reaction site. Adjacent reaction sites may be contiguous, so as to be adjacent to each other, or adjacent sites may be non-contiguous, with an intervening space between them.

[0110] Technical Improvements and Terminology All literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and web pages, regardless of the format of such literature and similar materials, are expressly incorporated by reference in their entirety. In the event that one or more of the incorporated literature and similar materials differs from or contradicts this application, including, but not limited to, in terms of defined terms, terminology usage, described techniques, etc., this application controls. Further information regarding terminology can be found in U.S. Nonprovisional Patent Application No. 16 / 826,168, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2019 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV), and U.S. Provisional Patent Application No. 62 / 821,766, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2020 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV).

[0111] The disclosed technology uses neural networks to improve the quality and quantity of nucleic acid sequence information that can be obtained from a nucleic acid template or its complement, e.g., a nucleic acid sample, such as a DNA or RNA polynucleotide or other nucleic acid sample. Accordingly, certain implementations of the disclosed technology provide higher throughput polynucleotide sequencing, e.g., higher rates of collection of DNA or RNA sequence data, greater efficiency in collecting sequence data, and / or lower costs of obtaining such sequence data, compared to previously available methods.

[0112] The disclosed technology uses neural networks to identify the centers of solid-phase nucleic acid clusters and analyze optical signals generated during sequencing of such clusters, unambiguously distinguishing between adjacent, neighboring, or overlapping clusters and assigning sequencing signals to single, discrete source clusters. These and related embodiments thus enable the recovery of meaningful information, such as sequence data, from regions of high-density cluster arrays, where useful information may not have been previously obtained from such regions due to confounding effects of overlapping or closely spaced neighboring clusters, including the effects of overlapping signals (e.g., as used in nucleic acid sequencing).

[0113] As described in more detail below, in certain embodiments, compositions are provided that include a solid support immobilized with one or more nucleic acid clusters, as provided herein. Each cluster contains multiple immobilized nucleic acids of the same sequence and has a distinct center with a detectable central label as provided herein, which is distinguishable from the nucleic acids immobilized in the surrounding area within the cluster. Also described herein are methods for making and using such clusters with distinct centers.

[0114] Embodiments of the present disclosure will find use in many situations where benefits derive from the ability to identify, determine, annotate, record, or otherwise assign the location of a substantially central location within a cluster, including high-throughput nucleic acid sequencing, development of image analysis algorithms for assigning optical or other signals to individual source clusters, and other applications where recognition of the center of immobilized nucleic acid clusters is desirable and beneficial.

[0115] In certain embodiments, the present invention contemplates methods related to high-throughput nucleic acid analysis, such as nucleic acid sequencing (e.g., "sequencing"). Exemplary high-throughput nucleic acid analyses include, but are not limited to, de novo sequencing, resequencing, whole genome sequencing, gene expression analysis, gene expression monitoring, epigenetics analysis, genome methylation analysis, allele-specific primer extension (APSE), genetic diversity profiling, whole genome polymorphism discovery and analysis, single nucleotide polymorphism analysis, hybridization-based sequencing, and the like. Those skilled in the art will understand that a variety of different nucleic acids can be analyzed using the methods and compositions of the present invention.

[0116] Although implementations of the present invention are described in the context of nucleic acid sequencing, they are applicable in any field in which image data acquired at different times, spatial locations, or other temporal or physical aspects are analyzed. For example, the methods and systems described herein are useful in the fields of molecular biology and cell biology, where image data from microarrays, biological specimens, cells, organisms, etc., are acquired and analyzed at different times or perspectives. Images can be obtained using any number of techniques known in the art, including, but not limited to, fluorescence microscopy, optical microscopy, confocal microscopy, optical imaging, magnetic resonance imaging, tomographic scanning, etc. As another example, the methods and systems described herein can be applied when image data acquired by surveillance, aerial, or satellite imaging techniques, etc., are acquired and analyzed at different times or perspectives. The methods and systems are particularly useful for analyzing images acquired within a field of view, in which observed specimens remain in the same location relative to each other within the field of view. However, specimens may have different properties in separate images; for example, specimens may appear different in separate images of the field of view. For example, the analyte may appear to be different in color for a given analyte detected in different images, may show changes in the intensity of the signal detected for a given analyte in different images, or even the appearance of a signal for a given analyte in one image and the disappearance of the signal for that analyte in another image.

[0117] As used herein, the term "analyte" is intended to mean a point or region of a pattern that can be distinguished from other points or regions according to their relative position. An individual analyte can contain one or more molecules of a particular type. For example, an analyte can contain a single target nucleic acid molecule having a particular sequence, or an analyte can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). Different molecules that are different analytes of a pattern can be differentiated from one another according to the location of the analyte within the pattern. Exemplary analytes include wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, pads of gel material on a substrate, or channels in a substrate.

[0118] Any of a variety of target analytes to be detected, characterized, or identified can be used in the devices, systems, or methods described herein. Exemplary analytes include, but are not limited to, nucleic acids (e.g., DNA, RNA, or analogs thereof), proteins, polysaccharides, cells, antibodies, epitopes, receptors, ligands, enzymes (e.g., kinases, phosphatases, or polymerases), small molecule drug candidates, cells, viruses, organisms, etc.

[0119] The terms "analyte," "nucleic acid," "nucleic acid molecule," and "polynucleotide" are used interchangeably herein. In various embodiments, a nucleic acid may be used as a template (e.g., a nucleic acid template or a nucleic acid complement complementary to a nucleic acid template) as provided herein for a particular type of nucleic acid analysis, including, but not limited to, nucleic acid amplification, nucleic acid expression analysis, and / or nucleic acid sequencing, or a suitable combination thereof. Nucleic acids in certain implementations include, for example, linear polymers of deoxyribonucleotides in 3'-5' phosphodiester chains, or deoxyribonucleic acid (DNA), such as single- and double-stranded DNA, genomic DNA, copy DNA or complementary DNA (cDNA), recombinant DNA, or any form of synthetic or modified DNA. In other embodiments, nucleic acids include, for example, linear polymers of ribonucleotides in 3'-5' phosphodiester or other linkages, such as ribonucleic acid (RNA), including single- and double-stranded RNA, messenger (mRNA), copy or complementary RNA (cRNA), alternatively spliced ​​mRNA, ribosomal RNA, small nuclear RNA (snoRNA), microRNA (miRNA), small interfering RNA (sRNA), piwi RNA (piRNA), or any form of synthetic or modified RNA. Nucleic acids used in the compositions and methods of the invention may vary in length and may be intact or full-length molecules or fragments, or smaller portions of larger nucleic acid molecules. In certain embodiments, nucleic acids may carry one or more detectable labels, as described elsewhere herein.

[0120] The terms "specimen," "cluster," "nucleic acid cluster," "nucleic acid colony," and "DNA cluster" are used interchangeably and refer to multiple copies of a nucleic acid template and / or its complement attached to a solid support. Typically, in certain preferred embodiments, a nucleic acid cluster comprises multiple copies of a template nucleic acid and / or its complement attached to a solid support via their 5' ends. The copies of the nucleic acid strands that make up a nucleic acid cluster may be in single-stranded or double-stranded form. The copies of the nucleic acid template present within a cluster may have nucleotides at corresponding positions that differ from each other due to, for example, the presence of a labeled moiety. Corresponding positions may also include analog structures with different chemical structures but similar Watson-Crick base pairing properties, such as uracil and thymine.

[0121] Colonies of nucleic acids may also be referred to as "nucleic acid clusters." Nucleic acid colonies can optionally be generated by cluster amplification or bridge amplification techniques, as described in more detail elsewhere herein. Multiple repeats of a target sequence can be present in a single nucleic acid molecule, such as a disruptor generated using a rolling circle amplification procedure.

[0122] The nucleic acid clusters of the present invention can have different shapes, sizes, and densities depending on the conditions used. For example, the clusters can be substantially circular, polyhedral, donut-shaped, or ring-shaped. The diameter of the nucleic acid cluster can be designed to be about 0.2 μm to about 6 μm, about 0.3 μm to about 4 μm, about 0.4 μm to about 3 μm, about 0.5 μm to about 2 μm, about 0.75 μm to about 1.5 μm, or any intermediate diameter. In certain embodiments, the diameter of the nucleic acid cluster is about 0.5 μm, about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The diameter of the nucleic acid cluster can be affected by numerous parameters, including, but not limited to, the number of amplification cycles performed in producing the cluster, the length of the nucleic acid template, or the density of primers attached to the surface on which the cluster is formed. The density of the nucleic acid cluster is typically less than 0.1 μm / mm 2 , 1 / mm 2 , 10 / mm 2 100 / mm 2 1,000 / mm 2 10,000 / mm 2 ~100,000 / mm 2 The present invention, in part, is directed to higher density nucleic acid clusters, e.g., 100,000 / mm 2 ~1,000,000 / mm 2 , and 1,000,000 / mm 2 ~10,000,000 / mm 2 Further plans are underway.

[0123] As used herein, an "analyte" is a specimen or region of interest within a field of view. When used in connection with a microarray device or other molecular analysis device, an analyte refers to a region occupied by similar or identical molecules. For example, an analyte can be an amplification oligonucleotide or any other group of polynucleotides or polypeptides having the same or similar sequence. In other embodiments, an analyte can be any element or group of elements that occupy a physical region on a sample. For example, an analyte can be a parcel of land, a body of water, etc. When analytes are imaged, each analyte has some area. Thus, in many embodiments, an analyte is not simply a single pixel.

[0124] The distance between specimens can be described in any number of ways. In some embodiments, the distance between specimens can be described from the center of one specimen to the center of another specimen. In other embodiments, the distance can be described from the edge of one specimen to the edge of another specimen, or between the outermost identifiable points of each specimen. The edge of the specimen can be described as a theoretical or actual physical boundary on the chip, or some point within the boundary of the specimen. In other embodiments, the distance can be described with respect to a fixed point on the specimen, or an image of the specimen.

[0125] Generally, some embodiments are described herein with respect to analytical methods. It will be understood that systems for performing the methods in an automated or semi-automated manner are also provided. Thus, the present disclosure provides a neural network-based template generation and base calling system, the system including a processor, a storage device, and a program for image analysis, the program including instructions for performing one or more of the methods described herein. Thus, the methods described herein can be performed, for example, on a computer having components described herein or known in the art.

[0126] The methods and systems described herein are useful for analyzing any of a variety of objects. Particularly useful objects are solid supports or solid surfaces with analytes attached. The methods and systems described herein offer advantages when used with objects having repeating patterns of analytes in the xy plane. One example is a microarray having a collection of cells, viruses, nucleic acids, proteins, antibodies, carbohydrates, small molecules (such as drug candidates), biologically active molecules, or other analytes of interest.

[0127] The number of applications of arrays containing analytes containing biological molecules such as nucleic acids and polypeptides has increased. Such microarrays typically contain deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) probes, which are specific for nucleotide sequences present in humans and other organisms. In certain applications, for example, individual DNA or RNA probes can be attached to individual analytes on the array. Test samples, such as those from known humans or organisms, can be exposed to the array so that target nucleic acids (e.g., gene fragments, mRNA, or amplicons) hybridize to complementary probes in each analyte in the array. The probes can be labeled through target-specific processes (e.g., due to labels present on the target nucleic acids or due to enzyme labels on the probes or targets present in hybridized form in the analyte). The analytes can then be examined by scanning specific light frequencies over them to identify which target nucleic acids are present in the sample.

[0128] Biological microarrays can be used for gene sequencing and similar applications. Generally, gene sequencing involves determining the order of nucleotides in a length of target nucleic acid, such as a fragment of DNA or RNA. Relatively short sequences are typically sequenced for each specimen, and the resulting sequence information can be used in various bioinformatics methods to reliably determine the sequence of many widely varying lengths of genetic material from which the fragments are derived. Automated computer-based algorithms for identifying signature fragments have been developed and have more recently been used in genome mapping, gene identification, and their functions. Microarrays are particularly useful for characterizing genome content because of the large number of variants present, which is an alternative to conducting numerous experiments for individual probes and targets. Microarrays are an ideal format for conducting such studies in a practical manner.

[0129] Any of a variety of analyte arrays (also referred to as "microarrays") known in the art can be used in the methods or systems described herein. A typical array contains analytes, each having an individual probe or a population of probes. In the latter case, the population of probes in each analyte is typically homogeneous, with a single type of probe. For example, in the case of nucleic acid sequences, each analyte can have multiple nucleic acid molecules, each having a common sequence. However, in some embodiments, the population in each analyte in the array can be heterogeneous. Similarly, protein sequences can have analytes with a single protein or a population of proteins, typically, but not necessarily, having the same amino acid sequence. Probes can be attached to the surface of the array, for example, by covalently linking the probe to the surface or through non-covalent interactions between the probe and the surface. In some embodiments, probes, such as nucleic acid molecules, can be attached to the surface via a gel layer, as described, for example, in U.S. Patent Application No. 13 / 784,368 and U.S. Patent Application Publication No. 2011 / 0059865(A1), each of which is incorporated herein by reference.

[0130] Exemplary arrays include, but are not limited to, BeadChip arrays available from Illumina, Inc. (San Diego, Calif.), or others, such as those described in U.S. Pat. Nos. 6,266,459, 6,355,431, 6,770,441, 6,859,570, or 7,622,294, or International Publication No. WO 00 / 63437, each of which is incorporated herein by reference, in which probes are attached to beads present on a surface (e.g., beads within wells on a surface). Further examples of commercially available microarrays that can be used include, for example, Affymetrix® GeneChip® microarrays or other microarrays synthesized according to a technique sometimes referred to as VLSIPS™ (Very Large Scale Immobilized Polymer Synthesis) technology. Spotted microarrays can also be used in methods or systems according to some embodiments of the present disclosure. An exemplary spotted microarray is the CodeLink™ Array available from Amersham Biosciences. Another useful microarray is one manufactured using inkjet printing methods, such as SurePrint™ Technology available from Agilent Technologies.

[0131] Other useful arrays include those used in nucleic acid sequencing applications.For example, the arrays with amplicons of genome fragments (often referred to as clusters) are described in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, International Publication No. 91 / 06678, International Publication No. 07 / 123744, US Patent No. 7,329,492, US Patent No. 7,211,414, US Patent No. 7,315,019, US Patent No. 7,405,281 or US Patent No. 7,057,026, or US Patent Application Publication No. 2008 / 0108082 (A1), each of which is incorporated herein by reference.Another type of array that is useful for nucleic acid sequencing is the array of particles that are generated from emulsion PCR technology. Examples are described in Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), WO 05 / 010145, U.S. Patent Application Publication No. 2005 / 0130173, or U.S. Patent Application Publication No. 2005 / 0064460, each of which is incorporated herein by reference in its entirety.

[0132] Arrays used for nucleic acid sequencing often have random spatial patterns of nucleic acid analytes. For example, the HiSeq or MiSeq sequencing platforms available from Illumina Inc. (San Diego, Calif.) utilize flow cells in which nucleic acid sequences are formed by random seeding followed by bridge amplification. However, patterned arrays can also be used for nucleic acid sequencing or other analytical applications. Examples of patterned arrays, their fabrication methods, and their uses are described in U.S. Patent Application Nos. 13 / 787,396, 13 / 783,043, 13 / 784,368, U.S. Patent Application Publication Nos. 2013 / 0116153 (A1), and 2012 / 0316086 (A1), each of which is incorporated herein by reference. Such patterned array analytes can be used to capture single nucleic acid template molecules for subsequent formation of homogeneous colonies, e.g., via bridge amplification. Such patterned arrays are particularly useful for nucleic acid sequencing applications.

[0133] The size of the specimens on an array (or other object used in the methods or systems herein) can be selected to suit a particular application. For example, in some embodiments, the specimens on the array can have a size that accommodates only a single nucleic acid molecule. A surface with multiple specimens in this size range is useful for constructing an array of molecules for detection with single molecule resolution. Specimens in this size range are also useful for use in arrays with specimens each comprising a colony of nucleic acid molecules. Thus, the specimens on the array can each be approximately 1 mm 2 Below, approximately 500μm 2 Below, approximately 100μm 2 Below, approximately 10μm 2 Below, approximately 1μm 2 Below, about 500nm 2 Less than or equal to about 100 nm 2 Below, approximately 10nm 2 Below, approximately 5nm 2 Less than or equal to 1 nm 2Alternatively or additionally, the specimens in the array may have an area of ​​about 1 mm 2 More than approximately 500μm 2 More than approximately 100μm 2 or more, about 10μm 2 or more, approximately 1μm 2 or more, about 500nm 2 Over 100nm 2 or more, about 10nm 2 or more, about 5nm 2 More than or about 1 nm 2 That's it. In practice, analytes can have a size within a range between upper and lower limits selected from those exemplified above. While several size ranges for surface analytes have been exemplified with respect to nucleic acids and nucleic acid scales, it will be understood that analytes in these size ranges can be used in applications that do not involve nucleic acids. It will further be understood that analyte sizes need not necessarily be limited to the scales used in nucleic acid applications.

[0134] In embodiments involving an object having multiple analytes, such as an array of analytes, the analytes can be distinct, separated by a space between them. Arrays useful in the invention can have analytes separated by an edge-to-edge distance of at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, or less. Alternatively or additionally, arrays can have analytes separated by an edge-to-edge distance of at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can apply to the average edge-to-edge spacing and edge-to-edge spacing of the analytes, as well as to the minimum or maximum spacing.

[0135] In some embodiments, the analytes in the array need not be distinct; instead, adjacent analytes can abut one another. Whether the analytes are distinct or not, the size of the analytes and / or the pitch of the analytes can be varied to allow the array to have a desired density. For example, the average analyte pitch in a regular pattern can be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, or less. Alternatively or additionally, the average analyte pitch in a regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can also apply to the maximum or minimum pitch of a regular pattern. For example, the maximum specimen pitch in the regular pattern can be 100 μm or less, 50 μm or less, 10 μm or less, 5 μm or less, 1 μm or less, 0.5 μm or less, and / or the minimum specimen pitch in the regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more.

[0136] The density of analytes in an array can also be understood in terms of the number of analytes present per unit area. For example, the average density of analytes for an array is at least about 1×10 3 specimens / mm 2 , 1×10 4 specimens / mm 2 , 1×10 5 specimens / mm 2 , 1×10 6 Inspection / mm 2 , 1×10 7 specimens / mm 2 , 1×10 8 specimens / mm 2 , or 1 × 10 9 specimens / mm 2 Alternatively or additionally, the average density of analytes on the array can be at most about 1 x 10 9 specimens / mm 2 , 1×10 8 specimens / mm 2 , 1×10 7 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 5specimens / mm 2 , 1×10 4 specimens / mm 2 , or 1 × 10 3 specimens / mm 2 It can be the following:

[0137] The above ranges may apply, for example, to all or part of a regular pattern that includes all or part of an array of analytes.

[0138] The analytes within the pattern can have any of a variety of shapes. For example, when viewed in a two-dimensional plane, such as on the surface of an array, the analytes may appear rounded, circular, oval, rectangular, square, symmetrical, asymmetrical, triangular, polygonal, etc. The analytes can be arranged in a regular repeating pattern, including, for example, a hexagonal or rectilinear pattern. The pattern can be selected to achieve a desired level of packing. For example, circular analytes are optimally packed in a hexagonal arrangement. Of course, other packing configurations can also be used for circular analytes, and vice versa.

[0139] A pattern can be characterized in terms of the number of analytes present in a subset that forms the smallest geometric unit of the pattern. A subset can include, for example, at least about 2, 3, 4, 5, 6, 10, or more analytes. Depending on the size and density of the analytes, a geometric unit can be as small as 1 mm. 2 , 500 μm 2 , 100 μm 2 , 50 μm 2 , 10 μm 2 , 1 μm 2 , 500nm 2 , 100 nm 2 , 50nm 2 , 10nm 2 Alternatively or additionally, the geometric unit may be 10 nm 2 , 50nm 2 , 100 nm 2 , 500nm 2 , 1 μm 2 , 10 μm 2 , 50 μm 2, 100 μm 2 , 500 μm 2 , 1mm 2 The properties of the specimens in the geometric unit, such as shape, size, pitch, etc., can be selected from those described herein more generally for specimens in an array or pattern.

[0140] An array with a regular pattern of analytes may be ordered with respect to the relative location of the analytes, but random with respect to one or more other characteristics of each analyte. For example, in the case of nucleic acid sequences, the nucleic acid analytes may be regular with respect to their relative location, but random with respect to the sequence knowledge of the nucleic acid species present in any particular analyte. As a more specific example, a nucleic acid sequence formed by seeding a repeating pattern of analytes with template nucleic acids and amplifying the template in each analyte to form copies of the template in the analyte (e.g., via cluster amplification or bridge amplification) will have a regular pattern of nucleic acid analytes, but will be random with respect to the distribution of the sequences of the nucleic acids across the array. Thus, detection of the presence of nucleic acid material on an array can result in a repeating pattern of analytes, whereas sequence-specific detection can result in a non-repeating distribution of signal across the array.

[0141] It will be understood that descriptions of pattern, order, randomness, etc. herein relate not only to analytes on an object, such as analytes on an array, but also to analytes in an image. Thus, the pattern, order, randomness, etc. can exist in any of a variety of formats used to store, manipulate, or communicate image data, including, but not limited to, computer-readable media or computer components such as a graphical user interface or other output device.

[0142] As used herein, the term "image" is intended to mean a representation of all or a portion of an object. The representation may be an optically detected reproduction. For example, an image may be obtained from fluorescence, luminescence, scattering, or absorption signals. The portion of an object present in an image may be the surface or other xy plane of the object. Typically, an image is a two-dimensional representation, but in some cases, information in an image can be derived from three or more dimensions. An image need not include optically detected signals. Non-optical signals may instead be present. An image may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.

[0143] As used herein, "image" refers to a reproduction or representation of at least a portion of a sample or other object. In some embodiments, the reproduction is an optical reproduction, e.g., produced by a camera or other optical detector. The reproduction may be a non-optical reproduction, e.g., a representation of electrical signals obtained from an array of nanopore analytes or a representation of electrical signals obtained from an ion-sensitive CMOS detector. In certain embodiments, non-optical reproductions may be excluded from the methods or apparatus described herein. The image may have a resolution capable of distinguishing between analytes present at any of a variety of intervals, including, for example, those spaced less than 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm apart.

[0144] As used herein, "acquisition," "capture," and like terms refer to any part of the process of acquiring an image file. In some embodiments, data acquisition can include generating an image of the specimen, looking for a signal in the specimen, directing a detection device to look for or generate an image of the signal, and providing instructions for further analysis or transformation of the image file, and instructions for any number of transformations or manipulations of the image file.

[0145] As used herein, the term "template" refers to a representation of the location or relationship between signals or analytes. Thus, in some embodiments, the template is a physical grid having a representation of signals corresponding to analytes in a sample. In some embodiments, the template can be a chart, table, text file, or other computer file that indicates locations corresponding to analytes. In the embodiments presented herein, a template is generated to track the location of analytes across a set of images of the sample captured at different reference points. For example, the template can be a set of x, y coordinates or a set of values ​​that describe the orientation and / or distance of one analyte relative to another analyte.

[0146] As used herein, the term "specimen" can refer to an object or region of an object from which an image is captured. For example, in an embodiment in which an image is taken from the surface of soil, a parcel of land may be the specimen. In other embodiments in which biomolecular analysis is performed within a flow cell, the flow cell may be divided into any number of subdivisions, each of which may be a specimen. For example, the flow cell may be divided into various channels or lanes, and each lane may be further divided into 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 140, 160, 180, 200, 400, 600, 800, 1000, or more distinct regions to be imaged. One example flow cell has eight lanes, each divided into 120 specimens or tiles. In other embodiments, samples may be generated in multiple tiles, or even the entire flow cell. Thus, each specimen image can represent a larger surface area than is imaged.

[0147] References to ranges and sequential lists of numbers set forth herein will be understood to include not only the numbers recited but all real numbers between the recited numbers.

[0148] As used herein, a "reference point" refers to any temporal or physical distinction between images. In another preferred embodiment, the reference point is a time point. In a more preferred embodiment, the reference point is a time point or cycle during the sequencing reaction. However, the term "reference point" can also include other aspects that distinguish or separate images, such as angle, rotation, time, or other aspects that can distinguish or separate the images.

[0149] As used herein, "subset of images" refers to a group of images in a set. For example, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or any number of images selected from a set of images. In certain other embodiments, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or less or any number of images selected from a set of images. In another preferred embodiment, images are obtained from one or more sequencing cycles, with four images associated with each cycle. Thus, for example, a subset may be a group of 16 images acquired over four cycles.

[0150] Base refers to a nucleotide base or nucleotide, (adenine), C (cytosine), T (thymine), or G (guanine). This application uses "base" and "nucleotide" interchangeably.

[0151] The term "chromosome" refers to a gene carrier of the present invention in a living cell, derived from a chromatin strand containing DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is used herein.

[0152] The term "site" refers to a unique location (e.g., chromosome ID, chromosomal location and orientation) on a reference genome. In some embodiments, a site may be a residue, sequence tag, or the location of a segment on a sequence. The term "locus" may be used to refer to a specific location of a nucleic acid sequence or polymorphism on a reference chromosome.

[0153] The term "sample" as used herein typically refers to a sample derived from a biological fluid, cell, tissue, organ, or organism containing the nucleic acid to be sequenced and / or phased, or a sample derived from a mixture of nucleic acids containing at least one nucleic acid sequence to be sequenced and / or phased. Such samples include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, fine needle biopsy samples (e.g., surgical biopsy, needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, tissue explants, organ cultures, and any other tissue or cell preparations thereof, or fractions or derivatives thereof. While samples are often collected from human subjects (e.g., patients), samples can be collected from any organism that has chromosomes, including, but not limited to, dogs, cats, horses, goats, sheep, cattle, pigs, etc. Samples can be used directly as obtained from a biological source or after pretreatment to modify the characteristics of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, and the like.

[0154] The term "sequence" includes or refers to a chain of nucleotides linked together. The nucleotides can be based on DNA or RNA. It should be understood that a single sequence may contain multiple subsequences. For example, a single sequence (e.g., a PCR amplicon) may have 350 nucleotides. A sample read may contain multiple subsequences within these 350 nucleotides. For example, a sample read may include first and second flanking subsequences, e.g., 20-50 nucleotides. The first and second flanking subsequences may be located on either side of a repeat segment with corresponding subsequences (e.g., 40-100 nucleotides). Each of the flanking subsequences may include (or a portion of) a primer subsequence (e.g., 10-30 nucleotides). For ease of reading, the term "subsequence" is referred to as "sequence," but it is understood that two sequences need not be distinct from each other on a common strand. To distinguish between the various sequences described herein, the sequences may be given different labels (e.g., target sequence, primer sequence, flanking sequence, reference sequence, etc.). Other terms, such as "allele," may be given different labels to distinguish between similar entities. Applications use "read" and "sequence read" interchangeably.

[0155] The term "paired end sequencing" refers to a sequencing method in which both ends of a target fragment are sequenced. Paired end sequencing can facilitate the detection of genome rearrangements and repeated segments, as well as the detection of gene fusions and novel transcripts. Methods for paired end sequencing are described in International Publication No. WO 07010252, International Application No. GB2007 / 003798, and U.S. Patent Application Publication No. 2009 / 0088327, each of which is incorporated herein by reference. In one example, the series of operations can be performed as follows: (a) generating clusters of nucleic acids, (b) linearizing the nucleic acids, (c) hybridizing a first sequencing primer and repeatedly performing cycles of extension, scanning, and deblocking as described above, (d) "flipping" the target nucleic acid on the flow cell surface by synthesizing a complementary copy, (e) linearizing the resynthesized strand, and (f) hybridizing a second sequencing primer and repeatedly performing cycles of extension, scanning, and deblocking as described above. The flipping operation can deliver the reagents described above for a single cycle of bridge amplification.

[0156] The term "reference genome" or "reference sequence" refers to a specific known genome sequence, either partial or complete, of any organism that can be used to reference an identified sequence from a subject. For example, reference genomes used for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences. A genome includes both genetic and non-coding sequences of DNA. A reference sequence may be larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 10 times larger, or at least about 10 times larger, or at least about 10 times larger. In one example, the reference genome sequence is of the full-length human genome. In another example, the reference genome sequence is limited to a specific human chromosome, such as chromosome 13. In some embodiments, the reference chromosome is a chromosome sequence from the human genome version hg19. Such sequences are sometimes referred to as chromosomal reference sequences, although the term reference genome is intended to encompass such sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes of any species, sub-chromosomal regions (such as strands), and the like. In various embodiments, a reference genome is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, a reference sequence may be taken from a specific individual. In other embodiments, "genome" also covers so-called "graph genomes," which use specific storage formats and representations of genome sequences. In one embodiment, a graph genome stores data in a linear file. In another embodiment, a graph genome refers to a representation in which alternative sequencing (e.g., different copies of a chromosome with small differences) are stored as different paths in a graph.Additional information regarding the implementation of graph genomes can be found at https: / / www.biorxiv.org / content / biorxiv / early / 2018 / 03 / 20 / 194530.full.pdf, the contents of which are incorporated herein by reference in their entirety.

[0157] The term "read" refers to a collection of sequence data describing a fragment of a nucleotide sample or reference. The term "read" can refer to a sample read and / or a reference read. Typically, but not necessarily, a read represents a short sequence of consecutive base pairs in a sample or reference. A read may be symbolically represented by the base pair sequence (ATCG) of the sample or reference fragment. A read may be stored in a memory device and appropriately processed to determine whether the read matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing device or indirectly from stored sequence information about the sample. In some cases, the read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region that can be aligned and specifically assigned to, for example, a chromosome or genomic region or gene.

[0158] Next-generation sequencing methods include, for example, sequencing by synthesis (Illumina), pyrosequencing (454), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), and sequencing by ligation (SOLiD sequencing). Depending on the sequencing method, the length of each read can vary from about 30 bp to over 10,000 bp. For example, DNA sequencing using a SOLiD sequencer generates nucleic acid reads of about 50 bp. In another example, Ion Torrent sequencing generates nucleic acid reads of up to 400 bp, while 454 pyrosequencing generates nucleic acid reads of about 700 bp. In yet another example, single-molecule real-time sequencing can generate reads of 10,000 bp to 15,000 bp. Thus, in certain embodiments, nucleic acid sequence reads have lengths of 30 to 100 bp, 50 to 200 bp, or 50 to 400 bp.

[0159] The terms "sample read," "sample sequence," or "sample fragment" refer to sequence data relating to a genomic sequence of interest from a sample. For example, a sample read includes sequence data from a PCR amplicon having forward and reverse primer sequences. The sequence data can be obtained from any selected sequence methodology. A sample read can be, for example, a sequencing-by-synthesis (SBS) reaction, a sequencing-ligation reaction, or any other suitable sequencing methodology in which it is desirable to determine the length and / or identity of repetitive elements. A sample read can be a consensus (e.g., average or weighted) sequence derived from multiple sample reads. In certain embodiments, providing a reference sequence includes identifying a locus of interest based on primer sequences of a PCR amplicon.

[0160] The term "raw fragment" refers to sequence data of a portion of a genome sequence of interest that at least partially overlaps a specified or secondary location of interest within a sample read or sample fragment. Non-limiting examples of raw fragments include double-stitched fragments, simple stitched fragments, and simple unstitched fragments. The term "raw" is used to indicate that a raw fragment contains sequence data that has some relationship to the sequence data in a sample read, regardless of whether the raw fragment exhibits supporting variants that correspond to and authenticate or confirm potential variants in the sample read. The term "raw fragment" does not necessarily indicate that the fragment contains supporting variants that verify the variant call in the sample read. For example, when a sample read is determined by a variant calling application to exhibit a first variant, the variant calling application may determine that one or more raw fragments lack a corresponding type of "supporting" variant that would otherwise be expected to occur given the variants in the sample read.

[0161] The terms "mapping," "aligned," "aligning," or "aligning" refer to the process of comparing a read or tag to a reference sequence, thereby determining whether the reference sequence contains the read sequence. If the reference sequence is read, the read may be mapped to the reference sequence, or in certain alternative embodiments, may be mapped to a specific location within the reference sequence. In some cases, alignment simply tells whether a read is a member of a particular reference sequence (i.e., whether the read is present or absent in the reference sequence). For example, alignment of a read to a reference sequence for human chromosome 13 tells whether the read is present in the reference sequence for chromosome 13. Tools that provide this information are sometimes called set membership testers. In some cases, alignment also indicates the location within the reference sequence where the read or tag maps. For example, if the reference sequence is the entire human genome sequence, alignment may indicate that the read is present on chromosome 13, and may further indicate that the read is on a specific strand and / or site of chromosome 13.

[0162] The term "indel" refers to the insertion and / or deletion of bases in an organism's DNA. Microindels refer to indels that result in a net change of 1 to 50 nucleotides. Unless the length of the indel is a multiple of three, a frameshift mutation occurs in coding regions of the genome. Indels can be contrasted with point mutations. An indel insertion deletes nucleotides from the sequence, whereas a point mutation is a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels can also be contrasted with tandem base mutations (TBMs), which can be defined as substitutions at adjacent nucleotides (mostly two adjacent nucleotides, although substitutions at three adjacent nucleotides have been observed).

[0163] The term "variant" refers to a nucleic acid sequence that differs from a reference nucleic acid. Typical nucleic acid sequence variants include, but are not limited to, single nucleotide polymorphisms (SNPs), short deletion and insertion polymorphisms (indels), copy number variations (CNVs), microsatellite markers, or short tandem repeats and structural variants. Somatic variant calling is an effort to identify variants present at low frequencies in a DNA sample. Calling somatic variants is of interest in the context of cancer treatment. Cancer is caused by the accumulation of mutations in DNA. DNA samples derived from tumors are generally heterogeneous, containing some normal cells, early-stage cancer progression (with fewer mutations), and some late-stage cells (with more mutations). Due to this heterogeneity, somatic mutations often appear at low frequencies when sequencing tumors (e.g., from FFPE samples). For example, an SNV may be found in only 10% of reads covering a given base. A variant that is classified as somatic or germline by a variant classifier is also referred to herein as a "variant under test."

[0164] The term "noise" refers to erroneous variant calls that result from one or more errors in the sequencing process and / or variant calling application.

[0165] The term "variant frequency" refers to the relative frequency of an allele (a genetic variant) at a particular locus within a population, expressed as a fraction or percentage. For example, the fraction or percentage may be the proportion of all chromosomes in the population that carry that allele. As an example, a sample variant frequency refers to the relative frequency of an allele / variant at a particular locus / position along a genome sequence of interest across a "population" corresponding to the number of reads and / or samples obtained for the genome sequence of interest from an individual. As another example, a baseline variant frequency refers to the relative frequency of an allele / variant at a particular locus / position along one or more baseline genome sequences, where the relative frequency of an allele / variant at a particular locus / position along one or more baseline genome sequences is obtained for one or more baseline genome sequences.

[0166] The term "Variant Allele Frequency (VAF)" refers to the proportion of sequenced reads that carry a variant divided by the overall coverage at the target position. VAF is a measure of the proportion of sequenced reads that carry a variant.

[0167] The terms "position," "designated position," and "locus" refer to the position or coordinates of one or more nucleotides within a nucleotide sequence. The terms "position," "designated position," and "locus" also refer to the position or coordinates of one or more base pairs in a sequence of nucleotides.

[0168] The term "haplotype" refers to a combination of alleles at adjacent sites on a chromosome that are inherited from one another. A haplotype, if present, may be one locus, several loci, or an entire chromosome, depending on the number of recombination events that have occurred between a given set of loci.

[0169] The term "threshold" herein refers to a number or values ​​used as a cutoff for characterizing a sample, a nucleic acid, or a portion thereof (e.g., a readout). The threshold may be varied based on empirical analysis. The threshold can be compared to a measured or calculated value to determine whether the source giving rise to such a value should be classified in a particular manner. The threshold can be identified empirically or analytically. The choice of threshold depends on the confidence level with which the user desires the classification to be made. The threshold may be selected for a particular purpose (e.g., to balance sensitivity and selectivity). As used herein, the term "threshold" refers to a point at which the course of an analysis may change and / or an action may be triggered. A threshold need not be a predetermined number. Instead, the threshold may be, for example, a function based on multiple factors. The threshold may be adaptive to the situation. Furthermore, a threshold may indicate an upper limit, a lower limit, or a range between limits.

[0170] In some embodiments, an index or score based on sequencing data may be compared to a threshold. As used herein, the term "metric" or "score" may include a value or result determined from sequencing data, or may include a function based on a value or result determined from sequencing data. Like a threshold, an index or score may be adaptive to the context. For example, an index or score may be a normalized value. As an example of a score or metric, one or more embodiments may use a count score when analyzing data. The count score may be based on the number of sample reads. The sample reads may have undergone one or more filtering stages so that the sample reads have at least one common characteristic or quality. For example, each of the sample reads used to determine the count score may be aligned with a reference sequence or assigned as a potential allele. The number of sample reads with a common characteristic may be counted to determine a read count. The count score may be based on the read count. In some embodiments, the count score may be a value equal to the read count. In other examples, the count score may be based on the read count and other information. For example, the counting score may be based on the read count of a particular allele at a locus and the total number of reads at the locus. In some embodiments, the counting score may be based on the read count of a locus and previously obtained data. In some embodiments, the counting score may be a normalized score between predetermined values. The counting score may also be a function of read counts from other loci in the sample, or read counts from other samples run simultaneously with the sample of interest. For example, the counting score may be a function of the read count of a particular allele and the read counts of other loci in the sample, and / or read counts from other samples. As an example, the read counts from other loci and / or read counts from other samples may be used to normalize the counting score for a particular allele.

[0171] The term "coverage" or "fragment coverage" refers to a count or other measure of multiple sample reads for the same fragment of a sequence. A read count may represent a count of the number of reads that cover the corresponding fragment. Alternatively, coverage may be determined by multiplying the read count by a specified factor based on historical knowledge, knowledge of the sample, knowledge of the locus, etc.

[0172] The term "read depth" (conventionally a number followed by "x") refers to the number of sequenced reads with overlapping alignment at the target position. This is often expressed as an average or percentage above a cutoff for a set of intervals (such as an exon, gene, or panel). For example, a clinical report may state that the panel average coverage is 1,105x, with 98% of the targeted base coverage >100x.

[0173] The term "base call quality score" or "Q score" refers to a PHRED-scaled probability ranging from 0-50 that is inversely proportional to the probability that a single sequenced base is correct. For example, a T base call with a Q of 20 is considered correct 99.99% of the time. Any base call with a Q < 20 should be considered low quality, and any variant identified when a significant proportion of sequenced reads supporting the variant is low should be considered a potential false positive.

[0174] The term "variant read" or "variant read number" refers to the number of sequenced reads that support the presence of a variant.

[0175] Regarding "strandedness" (or DNA twisting), the genetic message in DNA can be represented as the letters A, G, C, and T, e.g., 5'-AGGACA-3'. Often, sequences are written in the orientation shown herein, i.e., 5'-left and 3'-right. While DNA can occur as a single-stranded molecule (such as certain viruses), we typically find DNA as a double-stranded unit, which has a double helix structure with two antiparallel strands. In this case, the term "antiparallel" means that the two strands run parallel but have opposite polarity. Double-stranded DNA is held together by base pairing, and pairing is always maintained such that adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). This pairing is called complementarity, and one DNA strand is said to be the complement of the other. Thus, double-stranded DNA can be represented as two strings, such as 5'-AGGACA-3' and 3'-TCCTGT-5'. Note that the two strands have opposite polarities. Thus, the strandedness of the two DNA strands can be referred to as a reference strand and its complement, a forward and reverse strand, a top and bottom strand, a sense and antisense strand, or a Watson and Crick strand.

[0176] Read alignment (also called read mapping) is the process of referring to where sequences in a genome originate. Once aligned, the "mapping quality" or "mapping quality score (MAPQ)" of a given read quantifies the probability that its location on the genome is correct. Mapping quality is encoded on a phase scale, where P is the probability that the alignment is incorrect. The probability is P=10 (-MAQ / 10)MAPQ is calculated as follows: where MAPQ is the mapping quality. For example, a mapping quality of 40 = 10 with a power of -4 means that there is a 0.01% chance that the read is incorrectly aligned. Therefore, mapping quality is related to several alignment factors, such as the base quality of the read, the complexity of the reference genome, and paired-end information. First, a low base quality read means that the observed sequence is likely to be erroneous and therefore the alignment is incorrect. Second, mapping ability refers to the complexity of the genome. Repetitive regions are more difficult to map and the reads contained in these regions usually have low mapping quality. In this context, MAPQ reflects the fact that reads are not uniquely aligned and their actual origin cannot be determined. Third, in the case of paired-end sequencing data, concordant pairs are more likely to align. The higher the mapping quality, the better the alignment. Reads aligned with good mapping quality usually mean that the read sequence is well aligned, with only minor mismatches within highly mappable regions. The MAPQ value can be used as a quality control for the alignment results. A percentage of aligned reads with a MAPQ higher than 20 is usually suitable for downstream analysis.

[0177] As used herein, a "signal" refers to a detectable event, such as, for example, a light emission, preferably a light emission, in an image. Thus, in another preferred embodiment, a signal can represent any detectable light emission (i.e., a "spot") captured in an image. Thus, as used herein, a "signal" can refer to both actual emission from an analyte, and spurious emission that does not correlate with the actual analyte. Thus, a signal may result from noise and can be subsequently discarded as not representative of the actual analyte on the test strip.

[0178] As used herein, the term "clamp" refers to a group of signals. In certain embodiments, the signals are derived from different specimens. In another preferred embodiment, the signal clump is a group of signals that cluster together. In a more preferred embodiment, the signal clump represents the physical area covered by one amplification oligonucleotide. Each signal clump should ideally be observed as several signals (one per template cycle, possibly more due to crosstalk). Thus, overlapping signals are detected, where two (or more) signals are included in the template from the same signal clump.

[0179] As used herein, terms such as "minimum," "maximum," "minimize," "maximize," and grammatical variations thereof may include values ​​that are not absolute maximums or minimums. In some embodiments, values ​​include near maximums and minimums. In other examples, values ​​may include local maximums and / or local minimums. In some embodiments, values ​​may include only absolute maximums or minimums.

[0180] As used herein, "crosstalk" refers to the detection of a signal in one image that is also detected in a separate image. In another preferred embodiment, crosstalk can occur when an emitted signal is detected in two separate detection channels. For example, if an emitted signal occurs in one color, the emission spectrum of that signal may overlap with another emitted signal in another color. In a preferred embodiment, fluorescent molecules used to indicate the presence of nucleotide bases A, C, G, and T are detected in separate channels. However, because the emission spectra of A and C overlap, a portion of the C color signal may be detected during detection using a color channel. Thus, crosstalk between the A and C signals allows a signal from one color image to appear in another color image. In some embodiments, there is G and T crosstalk. In some embodiments, the amount of crosstalk between channels is asymmetric. It will be understood that the amount of crosstalk between channels can be controlled by, among other things, selecting signal molecules with appropriate emission spectra and selecting the size and wavelength range of the detection channels.

[0181] As used herein, "register," "registering," "registration," and similar terms refer to any process for correlating signals in an image or dataset with signals in an image or dataset from another time or perspective. For example, registration can be used to align signals from a set of images to form a template. In another example, registration can be used to register signals from other images to a template. One signal may be directly or indirectly registered to another signal. For example, a signal from image "S" may be directly registered to image "G." As another example, a signal from image "N" may be directly registered to image "G," or alternatively, a signal from image "N" may be registered to image "S," which was previously registered to image "G." Thus, the signal from image "N" is indirectly registered to image "G."

[0182] As used herein, the term "fiducial" is intended to mean a distinguishable reference point in or on an object. A fiducial point can be, for example, a mark, a second object, a shape, an edge, a region, an irregularity, a channel, a pit, a post, etc. A fiducial point can be present in an image of the object or in another data set derived from detecting the object. A fiducial point can be specified by an x ​​and / or y coordinate in the plane of the object. Alternatively or additionally, a fiducial point can be specified by a z coordinate orthogonal to the xy plane, for example, defined by the relative positions of the object and the detector. One or more coordinates for a fiducial point can be specified relative to one or more other specimens of the object, or an image or other data set derived from the object.

[0183] As used herein, the term "optical signal" is intended to include, for example, fluorescence, luminescence, scattering, or absorption signals. Optical signals can be detected in the ultraviolet (UV) range (approximately 200-390 nm), visible (VIS) range (approximately 391-770 nm), infrared (IR) range (approximately 0.771-25 micrometers), or other ranges of the electromagnetic spectrum. Optical signals can be detected in a manner that excludes all or part of one or more of these ranges.

[0184] As used herein, the term "signal level" is intended to mean the amount or quantity of detected energy or encoded information having a desired or predetermined characteristic. For example, optical signals can be quantified by one or more of intensity, wavelength, energy, frequency, power, brightness, etc. Other signals can be quantified according to characteristics such as voltage, current, electric field strength, magnetic field strength, frequency, power, temperature, etc. The absence of a signal is understood to be a signal level of zero or a signal level that is not significantly distinguishable from noise.

[0185] As used herein, the term "simulate" is intended to mean creating a representation or model of an object or action that predicts characteristics of the object or action. The representation or model may often be distinguishable from the object or action. For example, the representation or model may be distinguishable with respect to one or more characteristics, such as color, texture, size, or the intensity of a signal detected from all or a portion of the object or action. In certain implementations, the representation or model may be idealized, exaggerated, muted, or incomplete compared to the object or action. Thus, in some implementations, the representation of the model may be representative, for example, of at least one of the above characteristics. The representation or model may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.

[0186] As used herein, the term "specific signal" is intended to mean a detected energy or encoded information that is selectively observed over other energies or information, such as background energy or information. For example, a specific signal can be an optical signal detected at a particular intensity, wavelength, or color; an electrical signal detected at a particular frequency, power, or field strength; or other signal known in the art related to spectroscopy and analytical detection.

[0187] As used herein, the term "swath" is intended to mean a rectangular portion of an object. A swath may be an elongated strip that is scanned by relative movement between the object and the detector in a direction parallel to the longest dimension of the strip. Generally, the width of the rectangular portion or strip is constant along its entire length. Multiple swaths of an object may be parallel to one another. Multiple swaths of an object may overlap one another, be adjacent to one another, or be separated from one another by interstitial regions.

[0188] As used herein, the term "variance" is intended to mean the expected difference and the observed difference, or the difference between two or more observations. For example, variance can be the discrepancy between expected and measured values. Statistical functions such as standard deviation, the square of the standard deviation, and the coefficient of variation can be used to express variance.

[0189] As used herein, the term "x-y coordinates" is intended to mean information that specifies a position, size, shape, and / or orientation within an x-y plane. The information may be, for example, numerical coordinates in a Cartesian coordinate system. The coordinates may be provided relative to one or both of the x-axis and y-axis, or may be provided relative to another location within the x-y plane. For example, the coordinates of an analyte in an object may specify the location of the analyte relative to a fiducial or the location of another analyte in the object.

[0190] As used herein, the term "xy-plane" is intended to mean a two-dimensional region defined by linear axes x and y. When used with reference to a detector and an object observed by the detector, the region may be further specified as being orthogonal to the direction of observation between the detector and the object being detected.

[0191] As used herein, the term "z-coordinate" is intended to mean information specifying the location of a point, line, or region along an axis orthogonal to the x-y plane. In certain embodiments, the z-axis is orthogonal to the region of the object observed by the detector. For example, the direction of the focus of an optical system may be specified along the z-axis.

[0192] In some embodiments, the acquired signal data is transformed using an affine transformation. In some such embodiments, template generation utilizes the fact that affine transformations between color channels are consistent between runs. Because of this consistency, a set of default offsets can be used when determining the coordinates of analytes in a specimen. For example, a default offset file can include relative transformations (shifts, scales, skews) for one channel, such as the A channel, relative to a different channel. However, in other embodiments, offsets between color channels drift during and / or between runs, making offset-driven template generation difficult. In such examples, the methods and systems provided herein can utilize offset template generation, which is further described below.

[0193] In some aspects of the above embodiments, the system may include a flow cell. In some aspects, the flow cell includes lanes or other configurations of tiles, at least some of which contain one or more analytes. In some aspects, the analytes include a plurality of molecules, such as nucleic acids. In certain aspects, the flow cell is configured to extend primers that hybridize to nucleic acids within the analyte to deliver labeled nucleotide bases to a sequence of nucleic acids, thereby generating a signal corresponding to the analyte, which includes the nucleic acid. In preferred embodiments, the nucleic acids within the analyte are identical or substantially identical to one another.

[0194] In some of the image analysis systems described herein, each image in the set of images includes a color signal, with different colors corresponding to different nucleotide bases. In some embodiments, each image in the set of images includes a signal having a single color selected from at least four different colors. In some embodiments, each image in the set of images includes a signal having a single color selected from four different colors. In some of the systems described herein, nucleic acids can be sequenced by providing four different labeled nucleotide bases to the sequence of molecules to generate four different images, each image including a signal having a single color, with the signal color being different for each of the four different images, thereby generating a cycle of four color images corresponding to the four possible nucleotides present at a particular position in the nucleic acid. In certain embodiments, the system includes a flow cell configured to deliver additional labeled nucleotide bases to the sequence of molecules, thereby generating a cycle of multiple color images.

[0195] In preferred embodiments, the methods provided herein may include determining whether a processor is actively collecting data or whether the processor is in a low activity state. Collecting and storing a large number of high-quality images typically requires a large amount of storage capacity. Furthermore, once collected and stored, analyzing the image data can be resource-intensive and may impede the processing power of other functions, such as collecting and storing additional image data. Thus, as used herein, the term "low activity state" refers to the processing power of a processor at a given time. In some embodiments, a low activity state occurs when the processor is not collecting and / or storing data. In some embodiments, a low activity state occurs when some data collection and / or storage occurs, but additional processing power remains so that image analysis can occur simultaneously without interfering with other functions.

[0196] As used herein, "identifying a conflict" refers to identifying a situation in which multiple processes are competing for a resource. In some such embodiments, one process is given priority over another process. In some embodiments, the conflict may relate to the need to give priority to the allocation of time, processing power, storage power, or any other resource that is given priority. Thus, in some embodiments, when processing time or capacity is distributed between two processes, such as one that analyzes a data set and one that acquires and / or stores a data set, a discrepancy between the two processes exists that can be resolved by giving priority to one of the processes.

[0197] Also provided herein are systems for performing image analysis. The system can include a processor, a storage capacity, and a program for image analysis, the program including instructions for processing a first dataset for storage and a second dataset for analysis, the processing including acquiring and / or storing the first dataset on the storage device and analyzing the second dataset when the processor is not acquiring the first dataset. In certain aspects, the program includes instructions for identifying at least one instance of conflict between collecting and / or storing the first dataset and analyzing the second dataset, and prioritizing acquiring and / or storing image data such that collecting and / or storing the first dataset is given priority. In certain aspects, the first dataset includes an image file collected from an optical imaging device. In certain aspects, the system further includes an optical imaging device. In some aspects, the optical imaging device includes a light source and a detection device.

[0198] As used herein, the term "program" refers to instructions or commands for performing a task or process. The term "program" may be used interchangeably with the term "module." In certain embodiments, a program may be a compilation of various instructions executed under the same command set. In other embodiments, a program may refer to a separate batch or file.

[0199] The following are some of the surprising benefits of utilizing the methods and systems for performing image analysis described herein. In some sequencing implementations, an important measure of a sequencing system's usefulness is its overall efficiency. For example, the amount of mappable data generated per day and the total cost of installing and operating the equipment are important aspects of an economical sequencing solution. To reduce the time required to generate mappable data and increase system efficiency, real-time base calling can be enabled on the equipment computer and can run in parallel with sequencing chemistry and imaging. This allows data processing and analysis to be completed before finalizing sequencing chemistry. Furthermore, it can reduce the storage required for intermediate data and limit the amount of data that needs to be transferred across the network.

[0200] While sequence output has increased, the data transferred per operation from the system provided herein to the network and secondary analysis processing hardware has been substantially reduced. By converting data on the instrument computer (acquisition computer), the network load is dramatically reduced. Without these on-instrument, off-network data reduction techniques, the image output of a DNA sequencing instrument's FRET would cripple most networks.

[0201] The widespread adoption of high-throughput DNA sequencing instruments has been driven in part by their ease of use, support for a range of applications, and adaptability to virtually any laboratory environment. The highly efficient algorithms presented herein allow for the addition of significant analytical capabilities to simple workstations capable of controlling sequencing instruments. This reduction in computational hardware requirements has several practical advantages that will become even more important as sequencing output levels continue to increase. For example, image analysis and base calling are kept to a minimum by using simple towers, minimizing heat generation, laboratory footprint, and power consumption. In contrast, other commercial sequencing technologies have recently ramped up their computing infrastructure, with up to five times the processing power for primary analysis, before beginning to increase heat output and power consumption. Thus, in some embodiments, the computational efficiency of the methods and systems provided herein allows for increased sequencing throughput while minimizing server hardware.

[0202] Thus, in some embodiments, the methods and / or systems presented herein function as a state machine, keeping track of the individual state of each sample, and when it detects that a sample is ready to progress to the next state, it takes appropriate action to advance the sample to that state. A more detailed example of how a state machine monitors the file system to determine if a sample is ready to progress to the next state according to a preferred embodiment is provided in Example 1 below.

[0203] In preferred embodiments, the methods and systems provided herein are multi-threaded and can work with a configurable number of threads. Thus, for example, in the context of nucleic acid sequencing, the methods and systems provided herein can operate in the background during live sequencing operations for real-time analysis, or can operate using existing image datasets for offline analysis. In certain preferred embodiments, the methods and systems handle multi-threading by giving each thread its own subset of the analytes it is responsible for. This minimizes the possibility of thread stranding.

[0204] The disclosed methods can include using a detection device to acquire a target image of an object, the image including a repeating pattern of analytes on the object. Detection devices capable of high-resolution imaging of surfaces are particularly useful. In certain embodiments, the detection device will have sufficient resolution to distinguish analytes at the densities, pitches, and / or analyte sizes described herein. Detection devices capable of acquiring images or image data from a surface are particularly useful. Exemplary detectors are those configured to acquire area images while maintaining a static relationship between the object and the detector. Scanning devices can also be used. For example, devices that acquire continuous area images (e.g., referred to as "step and shot" detectors) can be used. Also useful are devices that continuously scan points or lines on the surface of an object and accumulate data to construct an image of the surface. Point-scan detectors can be configured to scan points (i.e., small detection areas) on the surface of an object via a raster motion in the x-y plane of the surface. Line-scan detectors can be configured to scan a line along the y-dimension of the surface of the object, with the longest dimension of the line occurring along the x-dimension. It will be understood that scanning detection can be achieved by moving the detection device, the object, or both. For example, detection devices that are particularly useful in nucleic acid sequencing applications are described in U.S. Patent Application Publication Nos. 2012 / 0270305(A1), 2013 / 0023422(A1), and 2013 / 0260372(A1), as well as U.S. Patent Nos. 5,528,050, 5,719,391, 8,158,926, and 8,241,573, each of which is incorporated herein by reference.

[0205] The embodiments disclosed herein may be implemented as a method, apparatus, system, or article of manufacture using programming or engineering techniques to generate software, firmware, hardware, or any combination thereof. As used herein, the term "article of manufacture" refers to code or logic implemented in hardware or computer-readable media, such as optical storage devices, as well as volatile or non-volatile memory devices. Such hardware includes, but is not limited to, field programmable gate arrays (FPGAs), coarse-grained reconfigurable architectures (CGRAs), application-specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microprocessors, or other similar processing devices. In certain embodiments, the information or algorithms described herein reside in non-transitory storage media.

[0206] In certain embodiments, the computer-implemented methods described herein can occur in real time while multiple images of an object are being acquired. Such real-time analysis is particularly useful for nucleic acid sequencing applications, where nucleic acid sequences are subjected to repeated cycles of fluidization and detection steps. While analysis of sequencing data can often be beneficial to perform the methods described herein in real time or in the background, it can also be beneficial to perform the methods described herein while other data collection or analysis algorithms are in process. Examples of real-time analysis methods that can be used in the present methods are those commercially available from Illumina, Inc. (San Diego, Calif.) and / or used in the MiSeq and HiSeq sequencing instruments described in U.S. Patent Application Publication No. 2012 / 0020537 A1, which is incorporated herein by reference.

[0207] An exemplary data analysis system is formed by one or more programmed computers, and has programming stored on one or more machine-readable media having code executed to perform one or more steps of the methods described herein. In one embodiment, for example, the system includes an interface designed to enable networking of the system to one or more detection systems (e.g., optical imaging systems) configured to acquire data from target objects. The interface can receive and condition the data, as appropriate. In certain embodiments, the detection systems output image data representing, for example, individual image elements or pixels that together form an image of an array or other object. A processor processes the received detection data according to one or more routines defined by the processing code. The processing code may be stored in various types of memory circuits.

[0208] According to presently contemplated embodiments, processing code executed on the detection data includes data analysis routines designed to analyze the detection data to determine the locations of individual analytes visible or encoded within the data, as well as locations where no analyte is detected (i.e., where no analyte is present or where no significant signal is detected from an existing analyte) and metadata. In certain embodiments, analyte locations within the array typically appear brighter than non-analyte locations due to the presence of a fluorescent dye attached to the imaged analyte. It will be understood that analytes need not appear brighter than their surrounding areas if, for example, a probe's target in the analyte is not present within the array being detected. The color in which individual analytes appear can be a function of the dye used as well as the wavelength of light used by the imaging system for imaging purposes. Analytes without bound targets or specific labels can be identified according to other characteristics, such as their expected location within the microarray.

[0209] Once the data analysis routine has located the individual analytes in the data, value assignment can be performed. Generally, value assignment assigns a digital value to each analyte based on the characteristics of the data represented by the detector element (e.g., pixel) at the corresponding location. That is, for example, when imaging data is processed, the value assignment routine may be designed to recognize that a particular color or wavelength of light has been detected at a particular location. In a typical DNA imaging application, for example, the four common nucleotides are represented by four separate, distinguishable colors. Each color may then be assigned a value corresponding to that nucleotide.

[0210] As used herein, the terms “module,” “system,” or “system controller” may include hardware and / or software systems and circuits that operate to perform one or more functions. For example, a module, system, or system controller may include a computer processor, controller, or other log-based device that performs operations based on instructions stored on a tangible and non-transitory computer-readable storage medium, such as a computer memory. Alternatively, a module, system, or system controller may include a hardwired device that performs operations based on hardwired logic and circuitry. A module, system, or system controller shown in the accompanying drawings may represent hardware and circuitry that operates based on software or hardwired instructions, software that instructs hardware to operate, or a combination thereof. A module, system, or system controller may include or represent hardware circuitry or circuitry that includes and / or is connected to one or more processors, such as a computer microprocessor.

[0211] As used herein, the terms "software" and "firmware" are used interchangeably and include any computer program stored in memory that is executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely examples and are not intended to be limiting of the types of memory that can be used to store computer programs.

[0212] In the field of molecular biology, one of the processes for nucleic acid sequencing in use is sequence synthesis. This technology can be applied to highly parallel sequencing projects. For example, by using an automated platform, millions of sequencing reactions can be performed simultaneously. Therefore, one embodiment of the present invention relates to an apparatus and method for collecting, storing, and analyzing image data generated during nucleic acid sequencing.

[0213] The enormous gain in the amount of data that can be collected and stored makes streamlined image analysis methods even more beneficial. For example, the image analysis methods described herein allow both designers and end users to make efficient use of existing computer hardware. Accordingly, methods and systems are presented herein that reduce the computational complexity of processing data in the face of rapidly increasing data output. For example, in the field of DNA sequencing, yields have expanded 15-fold in recent years, potentially reaching hundreds of gigases in a single run of a DNA sequencing device. Large-scale genome-scale experiments are beyond the reach of most researchers, given the proportional increase in computational infrastructure requirements. Therefore, the generation of ever-increasing amounts of raw sequence data increases the need for secondary analysis and data storage, making optimization of data transport and storage highly beneficial. Some embodiments of the methods and systems presented herein can reduce the time, hardware, networking, and laboratory infrastructure requirements required to generate usable sequence data.

[0214] This disclosure describes various methods and systems for performing the methods. Some example methods are described as a series of steps. However, it should be understood that implementations are not limited to the specific steps and / or order of steps described herein. Steps may be omitted, steps may be modified, and / or other steps may be added. Furthermore, steps described herein may be combined, steps may be performed simultaneously, steps may be divided into multiple substeps, steps may be performed in a different order, or steps (or series of steps) may be re-performed iteratively. Additionally, while different methods are described herein, it should be understood that other implementations may combine different methods (or steps of different methods).

[0215] In some embodiments, a processing unit, processor, module, or computing system that is "configured" to perform a task or operation may be understood to be specifically structured to perform the task or operation (e.g., arranged or intended to perform the task or operation, and / or having one or more programs or instructions arranged or intended to perform the task or operation, and / or having an arrangement of processing circuitry arranged or intended to perform the task or operation). For clarity and avoidance of doubt, a general purpose computer (which may be "configured to perform a task or operation when suitably programmed) is not configured as being "configured" to perform a task or operation unless it is specifically programmed or structurally modified to perform the task or operation).

[0216] Furthermore, the operations of the methods described herein may be sufficiently complex such that the operations cannot be performed by the average person or by a person skilled in the art within a commercially reasonable period of time. For example, the methods may rely on relatively complex calculations such that such a person cannot complete the methods within a commercially reasonable period of time.

[0217] Throughout this application, various publications, patents, or patent applications are referenced. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which this invention pertains.

[0218] The term "comprising," as used herein, is intended to be open-ended, encompassing not only the recited elements, but any additional elements as well.

[0219] As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set. Exceptions may occur where explicit disclosure or context clearly dictates otherwise.

[0220] Although the invention has been described with reference to the above embodiments, it will be understood that various modifications can be made without departing from the invention.

[0221] The modules of the present application can be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some may be implemented on different processors or computers, or may be spread across multiple different processors or computers. In addition, it will be understood that some of the modules may operate in parallel or in a different order than shown in the figures without affecting the functionality achieved. Also, as used herein, the term "module" may include "sub-modules," which may be considered herein to constitute a module. Blocks in the figures designated as modules may also be considered flowchart steps in a method.

[0222] As used herein, "identification" of an item of information does not necessarily require direct specification of that item of information. Information may be "identified" within a field by simply referencing the actual information through one or more layers in one direction, or by identifying one or more items of different information that are sufficient to determine the actual item of information. Additionally, the term "designate" is used herein to mean the same thing as "identify."

[0223] As used herein, a given signal, event, or value "depends on a pre-decessor signal, event, or value, an event or value that is affected by the given signal, event, or value. A given signal, event, or value can "exist" depending on a "pre-decessor signal, event, or value" if an intervening processing element, step, or time period is present. If an intervening processing element or step combines two or more signals, events, or values, the signal output of the processing element or step is considered "dependent" on each of the signal, event, or value inputs. If a given signal, event, or value is the same as a pre-decessor signal, event, or value, this is simply considered "dependent" or "depending on" the given signal, event, or value "depending on" or "based on" a "pre-decessor signal, event, or value." The "responsiveness" of a given signal, event, or value to another signal, event, or value is defined similarly.

[0224] As used herein, "concurrently" or "parallel" does not require exact simultaneity. It is sufficient if the assessment of one of the individuals begins before the assessment of another of the individuals is completed.

[0225] Computer Systems 17 illustrates a computer system 1700 that can be used to implement the disclosed techniques. The computer system 1700 includes at least one central processing unit (CPU) 1772 that communicates with a number of peripheral devices via a bus subsystem 1755. These peripheral devices may include, for example, a storage subsystem 1710, including memory devices and a file storage subsystem 1736, a user interface input device 1738, a user interface output device 1776, and a network interface subsystem 1774. The input and output devices enable user interaction with the computer system 1700. The network interface subsystem 1774 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.

[0226] In one embodiment, the equalizer basecaller 104 is communicatively linked to a storage subsystem 1710 and a user interface input device 1738 .

[0227] User interface input devices 1738 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, scanners, touchscreens integrated into displays, audio input devices such as voice recognition systems and microphones, and other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and ways of inputting information into computer system 1700.

[0228] The user interface output devices 1776 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display such as an audio output device. In general, use of the term "output device" is intended to encompass all possible types of devices and ways for outputting information from the computer system 1700 to a user or to another machine or computer system.

[0229] The storage subsystem 1710 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 1778.

[0230] The processor 1778 may be a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 1778 may be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processor 1778 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX17 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.

[0231] The memory subsystem 1722 used in the storage subsystem 1710 may include multiple memories, including a main random access memory (RAM) 1732 for storing instructions and data during program execution, and a read only memory (ROM) 1734 in which fixed instructions are stored. The file storage subsystem 1736 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of particular embodiments may be stored by the file storage subsystem 1736 within the storage subsystem 1710 or in another machine accessible by the processor.

[0232] Bus subsystem 1755 provides a mechanism for allowing the various components and subsystems of computer system 1700 to communicate with each other as intended. Although bus subsystem 1755 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0233] The computer system 1700 itself can be of a variety of types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the varying nature of computers and networks, the description of computer system 1700 shown in Figure 17 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 1700 can have more or fewer components than the computer system shown in Figure 17.

[0234] Specific Implementations The disclosed technology attenuates spatial crosstalk from sensor pixels using equalization-based image processing techniques. The disclosed technology can be implemented as a system, method, or product. One or more features of an embodiment can be combined with a base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of an embodiment can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of a repeating list of these options should not be construed as limiting the combinations taught in the preceding sections. These descriptions are incorporated herein by reference into each of the following embodiments.

[0235] In one embodiment, the disclosed technology proposes a computer-implemented method for attenuating spatial crosstalk from sensor pixels.

[0236] The disclosed technique addresses spatial crosstalk on sensor pixels in the pixel plane caused by fluorescent samples periodically distributed within the sample plane. Signal cones from the fluorescent samples are optically coupled to a local grid of sensor pixels via at least one lens. The signal cones overlap and impinge on the sensor pixels, thereby generating spatial crosstalk.

[0237] The disclosed technique captures, in at least one sub-pixel lookup table, the characteristic extent of a characteristic signal cone projected through a lens and the resulting contribution of the characteristic signal cone to the fluorescence detected by sensor pixels within a local grid of sensor pixels, the local grid of sensor pixels being substantially concentric with the center of the characteristic signal cone.

[0238] The disclosed technique interpolates between a set of sub-pixel lookup tables that represent the characteristic spread at sub-pixel resolution to generate an interpolated lookup table based on the target fluorescent sample center.

[0239] The disclosed technique isolates signals from target fluorescent samples that project the center of their signal cones onto substantially the center of a target local grid of sensor pixels by convolving an interpolation lookup table with the sensor pixels in the target local grid.

[0240] The disclosed technique uses the sum of the convoluted contributions of the isolated signals as the intensity of fluorescence from the target fluorescent sample.

[0241] The disclosed technique then base calls the first target fluorescent sample using the fluorescence intensity. Fluorescence intensities are determined for the first target fluorescent sample for each imaging channel in the multiple imaging channels. Consider a four-channel chemistry that uses four imaging channels to generate four images per sequencing cycle. Then, using the disclosed technique as described above, four fluorescence intensities are determined for the first target fluorescent sample. The four fluorescence intensities are then processed by a base caller to base call the first target fluorescent sample. Similarly, in a two-channel chemistry, two intensities of fluorescence are used to base call the first target fluorescent sample.

[0242] The methods described in this section and other sections of the disclosed technology may include one or more of the following features and / or characteristics described in connection with additional disclosed methods. For purposes of brevity, combinations of features disclosed in this application are not individually listed and are not repeated for each base set of features. The reader will understand how features identified in this method can be readily combined with sets of basic features identified as embodiments in other sections of this application.

[0243] In some embodiments, the periodically distributed fluorescent samples are arranged in a diamond shape, while in other embodiments, the periodically distributed fluorescent samples are arranged in a hexagonal shape.

[0244] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.

[0245] In another embodiment, the disclosed technology proposes a computer-implemented method for base calling.

[0246] The disclosed technique accesses an image whose pixels represent intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, the pixels including a central pixel that includes a center of the target cluster, each pixel within the central pixel being divisible into multiple sub-pixels.

[0247] In response to a particular subpixel in a plurality of subpixels of a central pixel including a center of a target cluster, the disclosed technique selects a subpixel lookup table corresponding to the particular subpixel from a bank of subpixel lookup tables, the selected subpixel lookup table including pixel coefficients configured to accept intensity radiation from the target cluster and exclude intensity radiation from adjacent clusters.

[0248] The disclosed technique multiplies the intensity values ​​of pixels in an image element-wise by pixel coefficients and sums the products of the multiplications to generate an output.

[0249] The disclosed techniques use the output to base call target clusters.

[0250] Each of the features discussed in this specific embodiment section for other embodiments applies equally to this embodiment. As noted above, all methods not repeated here should be repeated by reference.

[0251] In some embodiments, the disclosed techniques further include (i) selecting, from the bank of subpixel lookup tables, an additional subpixel lookup table corresponding to a subpixel that most closely neighbors the particular subpixel; (ii) interpolating between pixel coefficients of the selected subpixel lookup table and the selected additional subpixel lookup table to generate interpolated pixel coefficients configured to accept intensity radiation from the target cluster and reject intensity radiation from adjacent clusters; (iii) element-wise multiplying the interpolated pixel coefficients against intensity values ​​of pixels in the image and summing the products of the multiplications to generate an output; and (iv) base calling the target cluster using the output.

[0252] In some embodiments, the target cluster and additional adjacent clusters are periodically distributed on the flow cell in a diamond pattern and immobilized on the wells of the flow cell, while in other embodiments, the target cluster and additional adjacent clusters are periodically distributed on the flow cell in a hexagonal pattern and immobilized on the wells of the flow cell.

[0253] In some implementations, the interpolation is based on at least one of linear interpolation, bilinear interpolation, and bicubic interpolation.

[0254] In some implementations, pixel coefficients of the subpixel lookup tables in the bank of subpixel lookup tables are learned as a result of training an equalizer using decision-directed equalization. In one implementation, the decision-directed equalization uses least-squares estimation as a loss function. In one implementation, the least-squares estimation minimizes squared error using ground truth base calls. In one implementation, the ground truth base calls are modified to account for DC offsets, amplification factors, and the degree of polyclonality.

[0255] In some implementations, pixel coefficients of the subpixel lookup tables in the bank of subpixel lookup tables are derived from a combination of (i) a single subpixel lookup table whose pixel coefficients are learned as a result of training an equalizer using decision-directed equalization, and (ii) a set of pre-computed interpolation filters, each interpolation filter corresponding to a respective subpixel in the plurality of subpixels.

[0256] The disclosed technique further includes making the center of the target cluster substantially concentric with the center of the central pixel by (i) aligning the image with respect to the template image and determining affine transformation parameters and nonlinear transformation parameters; (ii) using the parameters to transform position coordinates of the target cluster and additional adjacent clusters into image coordinates of the image to generate a transformed image having transformed pixels; and (iii) applying interpolation using the transformed position coordinates of the target cluster and additional adjacent clusters to make each cluster center substantially concentric with the center of each transformed pixel that comprises the cluster center.

[0257] The disclosed techniques further include generating an output for each image in a plurality of images captured using each imaging channel in a particular sequencing cycle, and base calling a target cluster using the output generated for each image.

[0258] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.

[0259] The present inventors disclose the following items. 1. A computer-implemented method for base calling, the method comprising: accessing an image, wherein pixels of the image represent intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, the pixels including a central pixel that includes a center of the target cluster, each pixel within the image being divisible into a plurality of sub-pixels; Selecting, in response to the particular subpixel, a subpixel lookup table corresponding to the particular subpixel from a bank of subpixel lookup tables for a plurality of subpixels of a central pixel including a center of the target cluster, the selected subpixel lookup table including pixel coefficients configured to maximize a signal-to-noise ratio; element-wise multiplying the intensity values ​​of pixels in the image by pixel coefficients and summing the products of the multiplications to generate an output, where the pixel coefficients act as weights and the output is a weighted sum of the intensity values; and using the output to base call the target cluster. 2. The computer-implemented method of item 1, wherein the signal that is maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that is minimized in the signal-to-noise ratio is the intensity radiation from the neighboring cluster. 3. The computer-implemented method of item 1, wherein the element-wise multiplication applies a bias to a given set of equalizer coefficients. 4. The computer-implemented method of item 3, wherein the bias is a DC offset that averages the background noise intensity. 5. selecting an additional subpixel lookup table from the bank of subpixel lookup tables that corresponds to a subpixel that is most closely adjacent to the particular subpixel; interpolating between pixel coefficients of the selected sub-pixel lookup table and the selected additional sub-pixel lookup table to generate interpolated pixel coefficients configured to maximize a signal-to-noise ratio; element-wise multiplying the intensity values ​​of pixels in the interpolated image by pixel coefficients and summing the products of the multiplications to generate an output, wherein the interpolated pixel coefficients act as weights and the output is a weighted sum of the intensity values; 2. The computer-implemented method of claim 1, further comprising: using the output to base call target clusters. 6. The computer-implemented method of item 1, wherein the target cluster and the additional adjacent clusters are periodically distributed in a diamond shape on the flow cell and immobilized on the wells of the flow cell. 7. The computer-implemented method of item 6, wherein the target cluster and additional adjacent clusters are periodically distributed on a hexagonal flow cell and immobilized on the wells of the flow cell. 8. The computer-implemented method of item 1, wherein the interpolation is based on at least one of linear interpolation, bilinear interpolation, and bicubic interpolation. 9. The computer-implemented method of claim 1, wherein the pixel coefficients of the subpixel lookup tables in the bank of subpixel lookup tables are learned as a result of training the equalizer using at least one of least squares estimation, least squares, least mean squares, and recursive least squares. In other embodiments, other estimation and adaptation algorithms can be used to train the equalizer. 10. The computer-implemented method of item 9, further comprising training the equalizer in an offline mode, in which pixel coefficients of the sub-pixel lookup table are fixed after being trained on a batch of training data from a previously executed sequencing run. 11. The computer-implemented method of item 10, further comprising training the equalizer in an online mode, in which pixel coefficients of the sub-pixel lookup table are iteratively updated as training data from an ongoing sequencing run becomes available. 12. The computer-implemented method of item 11, further comprising: accessing per-base intensity distributions for each of the four bases A, C, G, and T generated during previous base calling of images in the training data; selecting the centers of each of the per-base intensity distributions as per-base ground truth target intensities; and training an equalizer using the per-base ground truth target intensities. 13. The computer-implemented method of claim 12, further comprising pre-training the equalizer in an offline mode and re-training the equalizer in an online mode. 14. The computer-implemented method of claim 9, further comprising generating a lookup table in a bank of sub-pixel lookup tables by applying a single set of equalizer coefficients and a set of pre-calculated interpolation filters together, and interpolating pixel intensities to generate inputs to the equalizer. This includes calculating pixel weights for clusters that have a substantially different alignment with respect to the pixel compared to the trained equalizer coefficients by using the interpolated pixel intensity values ​​to generate the equalizer inputs. The interpolation and equalizer filter responses can be convolved together for efficient implementation using a single shared LUT. In other implementations, the interpolation filter calculation can be performed directly without binning into sub-pixels. 15. aligning the image with respect to a template image and determining affine and non-linear transformation parameters; transforming location coordinates of the target cluster and the additional neighboring clusters into image coordinates of the image using the parameters to generate a transformed image having transformed pixels; Item 10. The computer-implemented method of item 1, further comprising: applying interpolation using the transformed position coordinates of the target cluster and additional neighboring clusters to make each cluster center substantially concentric with the center of each transformed pixel that contains the cluster center, thereby making the center of the target cluster concentric with the center of the central pixel. 16. The computer-implemented method of item 4, further comprising generating an output for each image in a plurality of images captured using each imaging channel and / or color channel in a particular sequencing cycle, and base calling a target cluster using the output generated for each image. 17. A computer-implemented method for recovering an underlying signal from a fluorescent sample disposed within a sample plane from signals corrupted by ambient fluorescent sources also within the sample plane, the method comprising: capturing in at least one sub-pixel lookup table a characteristic set of illuminations at the image plane by the sensor pixel array based on sampling that takes into account depletion from surrounding fluorescent sources, and then generating a set of lookup tables for the characteristic set of illuminations by the sensor pixel array when the center coordinates of fluorescent samples are at positions distributed across the center pixel of the sensor array and the positions are distributed relative to the center of the coordinates of the center pixel; receiving an image having a center coordinate of a fluorescent sample anywhere in a center pixel of a sensor pixel array, the image being corrupted by ambient fluorescent sources; and receiving a center coordinate of the fluorescent sample within the center pixel; calculating an interpolated table of a characteristic set of illuminations by the sensor pixel array customized for reception center coordinates of the fluorescent sample based on interpolation between the lookup tables in the set of lookup tables; recovering a signal from the target fluorescent sample that projects the center of the signal cone onto substantially the center of the target local grid of sensor pixels by element-wise multiplying the interpolation lookup table with the sensor pixels in the target local grid; using the sum of the products of the element-wise multiplications as the fluorescence intensity from the target fluorescent sample; and base calling the first target fluorescent sample using the fluorescence intensity. 1. A computer-implemented method for base calling, the method comprising: accessing an image, wherein pixels of the image represent intensity radiation from the target cluster and intensity radiation from additional adjacent clusters; selecting a lookup table containing pixel coefficients configured to maximize the signal-to-noise ratio; convolving the pixel coefficients with intensity values ​​of pixels in the image to generate an output; and base calling the target clusters based on the output. 2. The computer-implemented method of claim 1, wherein the signal that is maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that is minimized in the signal-to-noise ratio is the intensity radiation from adjacent clusters plus additional noise sources. 3. The computer-implemented method of claim 1, wherein the pixels include a central pixel that includes a center of the target cluster, and each pixel within the central pixel is divisible into multiple sub-pixels. 4. The computer-implemented method of claim 3, wherein the lookup table is a sub-pixel lookup table. 5. Selecting, in response to the particular subpixel, a subpixel lookup table corresponding to the particular subpixel from a bank of subpixel lookup tables for a plurality of subpixels of a central pixel including a center of the target cluster, the selected subpixel lookup table including pixel coefficients; element-wise multiplying the intensity values ​​of pixels in the image by pixel coefficients and summing the multiplications to generate an output, where the pixel coefficients act as weights and the output is a weighted sum of the intensity values; 5. The computer-implemented method of claim 4, further comprising base calling a target cluster using the output, comprising generating an output for each imaging channel in the plurality of imaging channels and base calling the target cluster using the output of each imaging channel. 6. The computer-implemented method of claim 5, wherein the element-wise multiplication adds a bias for a given set of equalizer coefficients, the bias being a DC offset that averages background noise intensity. 7. selecting an additional subpixel lookup table from the bank of subpixel lookup tables that corresponds to a subpixel that is consecutively adjacent to the particular subpixel; generating interpolated pixel coefficients based on pixel coefficients of the selected sub-pixel lookup table and the selected additional sub-pixel lookup table, the interpolated pixel coefficients being configured to maximize a signal-to-noise ratio; convolving the interpolated pixel coefficients with intensity values ​​of pixels in the image to generate an output; and base calling a target cluster based on the output. 8. 8. The computer-implemented method of claim 7, further comprising: element-wise multiplying intensity values ​​of pixels in the image by pixel coefficients and summing the products of the multiplications to generate an output, wherein the interpolated pixel coefficients act as weights and the output is a weighted sum of the intensity values. 9. The computer-implemented method of claim 1, further comprising training an equalizer to generate pixel coefficients using at least one of least squares estimation, least squares, least mean squares, and recursive least squares. 10. The computer-implemented method of claim 9, further comprising training the equalizer in an offline mode, in which pixel coefficients of the sub-pixel lookup table are fixed after being trained on a batch of training data from a previously performed sequencing run. 11. The computer-implemented method of claim 10, further comprising training the equalizer in an online mode, in which pixel coefficients of the sub-pixel lookup table are iteratively updated during an ongoing sequencing run. 12. The computer-implemented method of claim 11, further comprising: accessing per-base intensity distributions for each of the four bases A, C, G, and T generated during previous base calling of images in the training data; selecting the center of each of the per-base intensity distributions as a per-base ground truth target intensity for the corresponding color channel; and training an equalizer using the per-base ground truth target intensities. 13. The computer-implemented method of claim 12, further comprising pre-training the equalizer in an offline mode and re-training the equalizer in an online mode. 14. The computer-implemented method of claim 9, further comprising generating a lookup table in a bank of sub-pixel lookup tables by applying a single set of equalizer coefficients together with a set of pre-calculated interpolation filters to interpolate pixel intensities to generate inputs to the equalizer. 15. aligning the image with respect to a template image and determining affine and non-linear transformation parameters; transforming location coordinates of the target cluster and the additional neighboring clusters into image coordinates of the image using the parameters to generate a transformed image having transformed pixels; 2. The computer-implemented method of claim 1, further comprising: applying interpolation using the transformed position coordinates of the target cluster and additional neighboring clusters to make each cluster center concentric with a center of each transformed pixel that contains the cluster center, thereby making the center of the target cluster concentric with the center of the central pixel. 16. A non-transitory computer-readable storage medium storing computer program instructions for performing base calling, the instructions, when executed on a processor, accessing an image, wherein pixels of the image represent intensity radiation from the target cluster and intensity radiation from additional adjacent clusters; selecting a lookup table containing pixel coefficients configured to maximize the signal-to-noise ratio; convolving the pixel coefficients with intensity values ​​of pixels in the image to generate an output; and base calling target clusters based on the output. 17. The non-transitory computer-readable storage medium of claim 16, wherein the signal that is maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that is minimized in the signal-to-noise ratio is the intensity radiation from adjacent clusters plus additional noise sources. 18. The non-transitory computer-readable storage medium of claim 16, implementing a method further comprising training an equalizer using at least one of least squares estimation, least squares, least mean squares, and recursive least squares to generate pixel coefficients. 19. A system including one or more processors coupled to a memory, the memory being loaded with computer instructions for performing base calling, the instructions, when executed on the processor, performing: accessing an image, wherein pixels of the image represent intensity radiation from the target cluster and intensity radiation from additional adjacent clusters; selecting a lookup table containing pixel coefficients configured to maximize the signal-to-noise ratio; convolving the pixel coefficients with intensity values ​​of pixels in the image to generate an output; and base calling a target cluster based on the output. 20. The system of claim 19, further implementing the action of training an equalizer using at least one of least squares estimation, least squares, least mean squares, and recursive least squares to generate pixel coefficients.

[0260] While the present invention has been disclosed with reference to the above-described preferred embodiments and examples, it should be understood that these examples are intended in an illustrative and not a limiting sense. Modifications and combinations will readily occur to those skilled in the art, and such modifications and combinations are deemed to be within the spirit of the invention and the scope of the following claims. [Explanation of symbols]

[0261] 100A System 102 Sequencing Images 104 Equalizer Bass Caller 106 LUTs / LUT banks 108 Interpolation Filter 112 Ground Truth Base Calls 114 Trainer

Claims

1. A system comprising: at least one processor; a non-transitory computer-readable medium comprising instructions; wherein the instructions, when executed by the at least one processor, cause the system to: receiving an image of pixels indicative of intensity emissions from the target cluster and intensity emissions from adjacent clusters for a sequencing cycle; selecting a set of coefficients corresponding to a target pixel, the target pixel indicating the intensity emission from the target cluster and corresponding to a location of the target cluster; adjusting the set of coefficients for the target pixel to generate pixel-specific coefficients specific to the target pixel for the target cluster and pixel-specific coefficient sets specific to pixel sets for the neighboring clusters; determining, for the target cluster, a correction signal that reduces spatial crosstalk of the adjacent clusters at the target cluster based on the pixel-specific coefficients specific to the target pixel, the pixel-specific coefficient set specific to the pixel set, and intensity values ​​of the intensity radiation from the target cluster and the adjacent clusters; determining a base call for the target cluster for the sequencing cycle based on the corrected signal for the target cluster; A system that allows the following to be performed.

2. When executed by the at least one processor, the system: receiving the image of pixels indicative of the intensity radiation from the target cluster overlapping with one or more of the intensity radiation from the adjacent clusters; The system of claim 1 , further comprising instructions to:

3. The method of claim 2, wherein when executed by the at least one processor, the method comprises: receiving an image of the pixels indicative of the intensity radiation from the target cluster and the intensity radiation from the neighboring clusters; receiving an image patch of pixels indicative of a region of a sample plane that includes the intensity radiation from the target cluster and the intensity radiation from the adjacent clusters; The system of claim 1 further comprising instructions to:

4. The method of claim 1, wherein when executed by the at least one processor, the method comprises: adjusting the set of coefficients for the target pixel to generate the pixel-specific coefficient and the pixel-specific coefficient set; generating sub-pixel specific coefficients specific to the target pixel and representing a characteristic signal for the target cluster, and a set of sub-pixel specific coefficients specific to the set of pixels for the neighboring clusters; The system of claim 1 further comprising instructions to:

5. The method of claim 1, wherein when executed by the at least one processor, the method comprises: determining the correction signal for the target cluster based on the pixel-specific coefficients, the set of pixel-specific coefficients, and intensity values ​​of the intensity radiation from the target cluster and the neighboring clusters; determining interpolated pixel-specific coefficients for an array of pixels from the image of pixels from a table of predetermined pixel-specific coefficients; multiplying the interpolated pixel-specific coefficients by intensity values ​​corresponding to the pixel array; The system of claim 1 further comprising instructions to:

6. The method of claim 1, wherein when executed by the at least one processor, the method comprises: determining the interpolated pixel-specific coefficients for the pixel array by interpolating between the table of predetermined pixel-specific coefficients for the pixel array customized to the center coordinates of the target cluster; multiplying the interpolated pixel-specific coefficients with the intensity values ​​of the pixel array by element-wise multiplying the interpolated pixel-specific coefficients with the intensity values ​​corresponding to the pixel array; summing the products of the element-wise multiplications to determine one or more adjusted intensity values ​​for the target cluster; and The system of claim 5 further comprising instructions to:

7. The method of claim 1, wherein when executed by the at least one processor, the method comprises: determining the base call for the target cluster based on the one or more adjusted intensity values ​​for the target cluster; The system of claim 6 , further comprising instructions to:

8. The method of claim 7, wherein when executed by the at least one processor, the method comprises: determining the correction signal for the target cluster; determining a first adjusted intensity value for the target cluster in a first imaging channel based on the pixel-specific coefficients specific to the target pixel, the set of pixel-specific coefficients specific to the set of pixels, and one or more intensity values ​​of the intensity radiation from the target cluster and the neighboring clusters; determining a second adjusted intensity value for the target cluster in a second imaging channel based on the pixel-specific coefficients specific to the target pixel, the set of pixel-specific coefficients specific to the set of pixels, and one or more intensity values ​​of the intensity radiation from the target cluster and the adjacent clusters; determining the base call for the target cluster based on the first adjusted intensity value and the second adjusted intensity value; The system of claim 1 further comprising instructions to:

9. The method of claim 8, wherein when executed by the at least one processor, the method comprises: selecting the coefficient set, selecting a lookup table containing pixel coefficients corresponding to the location of the target cluster; The system of claim 1 further comprising instructions to:

10. A non-transitory computer-readable recording medium for recording instructions, comprising: The instructions, when executed by at least one processor, cause a computing system to: receiving an image of pixels indicative of intensity emissions from the target cluster and intensity emissions from adjacent clusters for a sequencing cycle; selecting a set of coefficients corresponding to a target pixel, the target pixel indicating the intensity emission from the target cluster and corresponding to a location of the target cluster; adjusting the set of coefficients for the target pixel to generate pixel-specific coefficients specific to the target pixel for the target cluster and pixel-specific coefficient sets specific to pixel sets for the neighboring clusters; determining, for the target cluster, a correction signal that reduces spatial crosstalk of the adjacent clusters at the target cluster based on the pixel-specific coefficients specific to the target pixel, the pixel-specific coefficient set specific to the pixel set, and intensity values ​​of the intensity radiation from the target cluster and the adjacent clusters; determining a base call for the target cluster for the sequencing cycle based on the corrected signal for the target cluster; A non-transitory computer-readable recording medium that causes the 11. The computing system according to claim 1, wherein when executed by said at least one processor, said computing system comprises: receiving the image of pixels indicative of the intensity radiation from the target cluster overlapping with one or more of the intensity radiation from the adjacent clusters; 11. The non-transitory computer-readable medium of claim 10, further comprising instructions to:

12. The computing system according to claim 1, wherein when executed by said at least one processor, said computing system comprises: receiving an image of the pixels indicative of the intensity radiation from the target cluster and the intensity radiation from the neighboring clusters; receiving an image patch of pixels indicative of a region of a sample plane that includes the intensity radiation from the target cluster and the intensity radiation from the adjacent clusters; 11. The non-transitory computer-readable storage medium of claim 10, further comprising instructions to:

13. The computing system according to claim 1, wherein when executed by said at least one processor: adjusting the set of coefficients for the target pixel to generate the pixel-specific coefficient and the pixel-specific coefficient set; generating sub-pixel specific coefficients specific to the target pixel and representing a characteristic signal for the target cluster, and a set of sub-pixel specific coefficients specific to the set of pixels for the neighboring clusters; 11. The non-transitory computer-readable storage medium of claim 10, further comprising instructions to:

14. The method of claim 13, wherein when executed by the at least one processor, the method comprises: determining the correction signal for the target cluster based on the pixel-specific coefficients, the set of pixel-specific coefficients, and intensity values ​​of the intensity radiation from the target cluster and the neighboring clusters; determining interpolated pixel-specific coefficients for an array of pixels from the image of pixels from a table of predetermined pixel-specific coefficients; multiplying the interpolated pixel-specific coefficients by intensity values ​​corresponding to the pixel array; 11. The non-transitory computer-readable storage medium of claim 10, further comprising instructions to:

15. The method of claim 1, wherein when executed by the at least one processor, the method comprises: determining the interpolated pixel-specific coefficients for the pixel array by interpolating between the table of predetermined pixel-specific coefficients for the pixel array customized to the center coordinates of the target cluster; multiplying the interpolated pixel-specific coefficients with the intensity values ​​of the pixel array by element-wise multiplying the interpolated pixel-specific coefficients with the intensity values ​​corresponding to the pixel array; summing the products of the element-wise multiplications to determine one or more adjusted intensity values ​​for the target cluster; and 15. The non-transitory computer-readable medium of claim 14, further comprising instructions to:

16. The computing system according to claim 1, wherein when executed by said at least one processor: determining the base call for the target cluster based on the one or more adjusted intensity values ​​for the target cluster; 16. The non-transitory computer-readable medium of claim 15, further comprising instructions to:

17. A computer-implemented method comprising: receiving an image of pixels indicative of intensity emissions from a target cluster and intensity emissions from adjacent clusters for a sequencing cycle; selecting a set of coefficients corresponding to a target pixel, the target pixel indicating the intensity radiation from the target cluster and corresponding to a location of the target cluster; adjusting the set of coefficients for the target pixel to generate pixel-specific coefficients specific to the target pixel for the target cluster and pixel-specific coefficient sets specific to pixel sets for the neighboring clusters; determining, for the target cluster, a correction signal that reduces spatial crosstalk of the adjacent clusters at the target cluster based on the pixel-specific coefficients specific to the target pixel, the pixel-specific coefficient set specific to the pixel set, and intensity values ​​of the intensity radiation from the target cluster and the adjacent clusters; determining a base call for the target cluster for the sequencing cycle based on the corrected signal for the target cluster; 11. A computer-implemented method comprising:

18. A computer-implemented method as described in claim 17, wherein the step of receiving an image of pixels indicating the intensity radiation from the target cluster and the intensity radiation from the adjacent cluster includes the step of receiving an image patch of pixels indicating an area of ​​a sample plane including the intensity radiation from the target cluster and the intensity radiation from the adjacent cluster.

19. The step of determining the correction signal for the target cluster, comprising: determining a first adjusted intensity value for the target cluster in a first imaging channel based on the pixel-specific coefficients specific to the target pixel, the set of pixel-specific coefficients specific to the set of pixels, and one or more intensity values ​​of the intensity radiation from the target cluster and the neighboring clusters; determining a second adjusted intensity value for the target cluster in a second imaging channel based on the pixel-specific coefficients specific to the target pixel, the set of pixel-specific coefficients specific to the set of pixels, and one or more intensity values ​​of the intensity radiation from the target cluster and the adjacent clusters; 20. The computer-implemented method of claim 17, comprising:

20. The computer-implemented method of claim 19, wherein the step of determining the base call includes a step of determining the base call for the target cluster based on the first adjusted intensity value and the second adjusted intensity value.

Citation Information

Patent Citations

  • Analytical systems and methods with software mask

    US20120015825A1

  • Cross-Talk Compensation

    US20190259137A1