Equalizer-based strength correction for base calls
An equalizer-based method corrects spatial crosstalk in DNA sequencing systems by maximizing the signal-to-noise ratio, improving basecall accuracy and reducing sequencing errors.
Patent Information
- Application Number
- JP2022567386
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-04
- Filing Date
- 2021-05-05
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2041-05-05
AI Technical Summary
Existing DNA sequencing systems face challenges in accurately distinguishing optical signals from target wells due to spatial crosstalk between adjacent wells, leading to sequencing errors and reduced basecall accuracy.
A computer-implemented method using an equalizer-based approach to generate a lookup table that convolves pixel coefficients with intensity values to maximize the signal-to-noise ratio, correcting for spatial crosstalk and improving basecalling accuracy.
The method effectively reduces sequencing errors by attenuating spatial crosstalk, thereby enhancing basecall accuracy in DNA sequencing processes.
Smart Images

Figure 0007712297000008 
Figure 0007712297000009 
Figure 0007712297000010
Abstract
Description
Technical Field
[0001] (Priority Application) This PCT application claims the benefit of U.S. Provisional Patent Application No. 63 / 020,449, entitled "EQUALIZATION-BASED IMAGE PROCESSING AND SPATIAL CROSSTALK ATTENUATOR" (Attorney Docket No. ILLM1032-1 / IP-1991-PRV), filed on May 5, 2021, and U.S. Provisional Patent Application No. 17 / 308,035, entitled "EQUALIZATION-BASED IMAGE PROCESSING AND SPATIAL CROSSTALK ATTENUATOR" (Attorney Docket No. ILLM1032-2 / IP-1991-US), filed on May 4, 2020. The priority applications are hereby incorporated by reference for all purposes.
[0002] (Field of the Invention) The disclosed technology relates to apparatuses and corresponding methods for the automatic analysis of images or the recognition of patterns. This specification includes systems for transforming an image for the purposes of (a) improving its visual quality prior to recognition, (b) positioning and aligning the image with respect to a sensor or a stored prototype, or reducing the amount of image data by discarding irrelevant data, and (c) measuring significant characteristics of the image. Specifically, the disclosed technology relates to using equalization-based image processing techniques to remove spatial crosstalk from sensor pixels.
[0003] (Cross-Reference to Related Applications) Incorporation The following are hereby incorporated by reference for all purposes as if fully set forth herein. U.S. Non-Provisional Application No. 15 / 936,365, entitled "DETECTION APPARATUS HAVING A MICROFLUOROMETER, A FLUIDIC SYSTEM, AND A FLOW CELL LATCH CLAMP MODULE", filed on March 26, 2018, U.S. Non-Provisional Application No. 16 / 567,224, entitled "FLOW CELLS AND METHODS RELATED TO SAME", filed on September 11, 2019, U.S. Non-Provisional Application No. 16 / 439,635, entitled "DEVICE FOR LUMINESCENT IMAGING", filed on June 12, 2019, U.S. Non-Provisional Application No. 15 / 594,413, entitled "INTEGRATED OPTOELECTRONIC READ HEAD AND FLUIDIC CARTRIDGE USEFUL FOR NUCLEIC ACID SEQUENCING", filed on May 12, 2017, U.S. Non-Provisional Application No. 16 / 351,193, entitled "ILLUMINATION FOR FLUORESCENCE IMAGING USING OBJECTIVE LENS", filed on March 12, 2019, U.S. Non-Provisional Application No. 12 / 638,770, entitled "DYNAMIC AUTOFOCUS METHOD AND SYSTEM FOR ASSAY IMAGER", filed on December 15, 2009, U.S. Non-Provisional Application No. 13 / 783,043, entitled "KINETIC EXCLUSION AMPLIFICATION OF NUCLEIC ACID LIBRARIES", filed on March 1, 2013, U.S. Non-Provisional Application No. 13 / 006,206, entitled "DATA PROCESSING SYSTEM AND METHODS", filed on January 13, 2011, U.S. Non-Provisional Application No. 14 / 530,299, entitled "IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS", filed on October 31, 2014, U.S. Non-Provisional Patent Application No. 15 / 153,953, titled "METHODS AND SYSTEMS FOR ANALYZING IMAGE DATA", filed on December 3, 2014, U.S. Non-Provisional Patent Application No. 14 / 020,570, titled "CENTROID MARKERS FOR IMAGE ANALYSIS OF HIGH DENSITY CLUSTERS IN COMPLEX POLYNUCLEOTIDE SEQUENCING", filed on September 6, 2013, U.S. Non-Provisional Patent Application No. 14 / 530,299, titled "IMAGE ANALYSIS USEFUL FOR PATTERNED OBJECTS", filed on October 31, 2014, U.S. Non-Provisional Patent Application No. 12 / 565,341, titled "METHOD AND SYSTEM FOR DETERMINING THE ACCURACY OF DNA BASE IDENTIFICATIONS", filed on September 23, 2009, U.S. Non-Provisional Patent Application No. 12 / 295,337, titled "SYSTEMS AND DEVICES FOR SEQUENCE BY SYNTHESIS ANALYSIS", filed on March 30, 2007, U.S. Non-Provisional Patent Application No. 12 / 020,739, titled "IMAGE DATA EFFICIENT GENETIC SEQUENCING METHOD AND SYSTEM", filed on January 28, 2008, U.S. Non-Provisional Patent Application No. 13 / 833,619, titled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR SAME" (Attorney Docket No. IP-0626-US), filed on March 15, 2013, U.S. Non-Provisional Application No. 15 / 175,489 (Attorney Docket No. IP-0689-US), entitled "BIOSENSORS FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND METHODS OF MANUFACTURING THE SAME", filed on June 7, 2016, U.S. Non-Provisional Application No. 13 / 882,088 (Attorney Docket No. IP-0462-US), entitled "MICRODEVICES AND BIOSENSOR CARTRIDGES FOR BIOLOGICAL OR CHEMICAL ANALYSIS AND SYSTEMS AND METHODS FOR THE SAME", filed on April 26, 2013, U.S. Non-Provisional Application No. 13 / 624,200 (Attorney Docket No. IP-0538-US), entitled "METHODS AND COMPOSITIONS FOR NUCLEIC ACID SEQUENCING", filed on September 21, 2012, U.S. Provisional Application No. 62 / 821,602 (Attorney Docket No. ILLM1008-1 / IP-1693-PRV), entitled "Training Data Generation for Artificial Intelligence-Based Sequencing", filed on March 21, 2019, U.S. Provisional Application No. 62 / 821,618 (Attorney Docket No. ILLM1008-3 / IP-1741-PRV), entitled "Artificial Intelligence-Based Generation of Sequencing Metadata", filed on March 21, 2019, U.S. Provisional Application No. 62 / 821,681 (Attorney Docket No. ILLM1008-4 / IP-1744-PRV), entitled "Artificial Intelligence-Based Base Calling", filed on March 21, 2019, U.S. Provisional Patent Application No. 62 / 821,724, entitled "Artificial Intelligence-Based Quality Scoring", filed on March 21, 2019 (Attorney Docket No. ILLM1008-7 / IP-1747-PRV), U.S. Provisional Patent Application No. 62 / 821,766, entitled "Artificial Intelligence-Based Sequencing", filed on March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV), Netherlands Patent Application No. 2023310, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing", filed on June 14, 2019 (Attorney Docket No. ILLM1008-11 / IP-1693-NL), Netherlands Patent Application No. 2023311, entitled "Artificial Intelligence-Based Generation of Sequencing Metadata", filed on June 14, 2019 (Attorney Docket No. ILLM1008-12 / IP-1741-NL), Netherlands Patent Application No. 2023312, entitled "Artificial Intelligence-Based Base Calling", filed on June 14, 2019 (Attorney Docket No. ILLM1008-13 / IP-1744-NL), Netherlands Patent Application No. 2023314, entitled "Artificial Intelligence-Based Quality Scoring", filed on June 14, 2019 (Attorney Docket No. ILLM1008-14 / IP-1747-NL), and Netherlands Patent Application No. 2023316, entitled "Artificial Intelligence-Based Sequencing", filed on June 14, 2019 (Attorney Docket No. ILLM1008-15 / IP-1752-NL). U.S. Non-Provisional Patent Application No. 16 / 825,987 (Attorney Docket No. ILLM1008-16 / IP-1693-US), entitled "Training Data Generation for Artificial Intelligence-Based Sequencing", filed on March 20, 2020, U.S. Non-Provisional Patent Application No. 16 / 825,991 (Attorney Docket No. ILLM1008-17 / IP-1741-US), entitled "Training Data Generation for Artificial Intelligence-Based Sequencing", filed on March 20, 2020, U.S. Non-Provisional Patent Application No. 16 / 826,126 (Attorney Docket No. ILLM1008-18 / IP-1744-US), entitled "Artificial Intelligence-Based Base Calling", filed on March 20, 2020, U.S. Non-Provisional Patent Application No. 16 / 826,134 (Attorney Docket No. ILLM1008-19 / IP-1747-US), entitled "Artificial Intelligence-Based Quality Scoring", filed on March 20, 2020, U.S. Non-Provisional Patent Application No. 16 / 826,168 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV), entitled "Artificial Intelligence-Based Sequencing", filed on March 21, 2020, U.S. Provisional Patent Application No. 62 / 849,091 (Attorney Docket No. ILLM1011-1 / IP-1750-PRV), entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing", filed on May 16, 2019, U.S. Provisional Patent Application No. 62 / 849,132 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2), entitled "Base Calling Using Convolutions", filed on May 16, 2019, U.S. Provisional Patent Application No. 62 / 849,133, entitled "Base Calling Using Compact Convolutions", filed on May 16, 2019 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3), U.S. Provisional Patent Application No. 62 / 979,384, entitled "Artificial Intelligence-Based Base Calling of Index Sequences", filed on February 20, 2020 (Attorney Docket No. ILLM1015-1 / IP-1857-PRV), U.S. Provisional Patent Application No. 62 / 979,414, entitled "Artificial Intelligence-Based Many-To-Many Base Calling", filed on February 20, 2020 (Attorney Docket No. ILLM1016-1 / IP-1858-PRV), U.S. Provisional Patent Application No. 62 / 979,385, entitled "Knowledge Distillation-Based Compression of Artificial Intelligence-Based Base Caller", filed on February 20, 2020 (Attorney Docket No. ILLM1017-1 / IP-1859-PRV), U.S. Provisional Patent Application No. 62 / 979,412, entitled "Multi-Cycle Cluster Based Real Time Analysis System", filed on February 20, 2020 (Attorney Docket No. ILLM1020-1 / IP-1866-PRV), U.S. Provisional Patent Application No. 62 / 979,411, entitled "Data Compression for Artificial Intelligence-Based Base Calling", filed on February 20, 2020 (Attorney Docket No. ILLM1029-1 / IP-1964-PRV), and U.S. Provisional Patent Application No. 62 / 979,399 (Attorney Docket No. ILLM1030-1 / IP-1982-PRV) entitled "Squeezing Layer for Artificial Intelligence-Based Base Calling", filed on February 20, 2020. BACKGROUND OF THE INVENTION
[0004] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, problems mentioned in this section, or problems associated with the subject matter provided as background, should not be assumed to have been previously recognized in the prior art. The subject matter of this section merely represents different approaches, which in themselves may also correspond to embodiments of the claimed technology.
[0005] Various protocols in biological or chemical research involve performing a number of controlled reactions on a local support surface or within a defined reaction chamber. The desired reaction can then be observed or detected, and subsequent analysis can help identify or elucidate the properties of the chemical substances involved in the reaction. For example, in some multiplex assays, an unknown analyte having a distinguishable label (e.g., a fluorescent label) can be exposed to thousands of known probes under controlled conditions. Each known probe can be deposited in a corresponding well of a microplate. Observing any chemical reaction that occurs between the known probe and the unknown analyte in the well can assist in identifying or elucidating the properties of the analyte. Other examples of such protocols include known DNA sequencing processes such as synthesis-based sequencing or circular array sequencing. In circular array sequencing, a high-density array of DNA features (e.g., template nucleic acids) is sequenced through repeated cycles of enzymatic manipulation. After each cycle, an image is captured and subsequently analyzed using other images to determine the sequence of the DNA features.
[0006] As a more specific example, one known DNA sequencing system uses a pyrosequencing process and includes a chip having a fused fiber faceplate with millions of wells. Single capture beads having sstDNA amplified clonally from the target genome are deposited into each well. After the capture beads are deposited in the wells, nucleotides are added continuously to the wells by flowing a solution containing the specific nucleotides along the faceplate. The environment within the well is such that when the nucleotides flowing through a particular well complement the DNA strand on the corresponding capture bead, the nucleotides are added to the DNA strand. A colony of DNA strands is called a cluster. The incorporation of nucleotides into the cluster initiates a process that ultimately generates a chemiluminescent signal. The system includes a CCD camera positioned directly adjacent to the faceplate and configured to detect the optical signals from the DNA clusters in the wells. Subsequent analysis of the images obtained throughout the pyrosequencing process can determine the sequence of the target genome.
[0007] However, the above pyrosequencing system may have certain limitations in addition to other systems. For example, the faceplate of the optical fiber is acid-etched to form millions of small wells. The wells may be arranged approximately spaced apart from each other, but it is difficult to know the exact position of the wells with respect to other adjacent wells. When the CCD camera is positioned directly adjacent to the faceplate, the wells are not evenly distributed along the pixels of the CCD camera, and thus the wells are not aligned with the pixels in a known pattern. Spatial crosstalk is the crosstalk between adjacent wells and makes it difficult to distinguish the true optical signal from the target well from other unwanted optical signals in subsequent analysis. Also, fluorescence emission is substantially isotropic. As the density of the sample increases, it becomes increasingly difficult to manage or account for the unwanted emission (e.g., crosstalk) from adjacent samples. As a result, the data recorded during the sequencing cycle needs to be carefully analyzed.
[0008] Basecall accuracy is extremely important for downstream analysis such as high-throughput DNA sequencing, read mapping, and genome assembly. Spatial crosstalk between adjacent clusters accounts for most of the sequencing errors. Therefore, by correcting the spatial crosstalk in the cluster intensity data, there is an opportunity to reduce DNA sequencing errors and improve basecall accuracy.
SUMMARY OF THE INVENTION
MEANS FOR SOLVING THE PROBLEM
[0009] One aspect of the present invention provides a computer-implemented method for basecalling. The computer-implemented method includes accessing an image, wherein pixels of the image represent intensity emissions from a target cluster and intensity emissions from additional adjacent clusters, selecting a lookup table including pixel coefficients configured to maximize a signal-to-noise ratio, convolving the pixel coefficients with intensity values of the pixels in the image to generate an output, and basecalling the target cluster based on the output.
[0010] The patent or application file includes at least one drawing created in color. Copies of this patent or patent application publication with color drawings (s) will be provided by the Office upon request and payment of the necessary fee. The color drawings may also be available in PAIR (Patent Application Information Retrieval) via the Supplementary Content tab.
[0011] In the drawings, like reference numerals generally refer to like parts throughout different views. Also, the drawings are not necessarily to scale; instead, emphasis is being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings.
BRIEF DESCRIPTION OF THE DRAWINGS
[0012]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13A
Figure 13B
Figure 13C
Figure 13D
Figure 13E
Figure 13F
Figure 14A
Figure 14B
Figure 14C
Figure 15A
Figure 15B
Figure 16
Figure 17
Figure 18
Figure 19A
Figure 19B-1
Figure 19B-2
Figure 19C
Figure 19D
Mode for Carrying Out the Invention
[0013] The following description typically refers to specific structural embodiments and methods. Although there is no intention to limit the present technology to the specifically disclosed embodiments and methods, it should be understood that the present technology can be implemented using other features, elements, methods, and embodiments. The preferred embodiments are described to explain the present technology, not to limit the scope defined by the claims. Those skilled in the art will recognize various equivalent modifications to the following description.
[0014] Generation of Lookup Table FIG. 1 shows an embodiment of generating a lookup table (LUT) (or LUT bank) 106 by training an equalizer 104. The equalizer 104 is also referred to here as an equalizer-based base corrector 104. The system 100A includes a trainer 114 that trains the equalizer 104 using least squares estimation. Additional details regarding the equalizer and least squares estimation are described in the appendix included in this application.
[0015] The array determination image 102 is generated during an array determination run performed by an array determination device such as Illumina's iSeq, HiSeqX, HiSeq3000, HiSeq4000, HiSeq2500, NovaSeq6000, NextSeq550, NextSeq1000, NextSeq2000, NextSeqDx, MiSeq, and MiSeqDx. In one embodiment, the Illumina sequencer uses cyclic reversible termination (CRT) chemistry for base calling. This process relies on extending a nascent strand complementary to a template strand with fluorescently labeled nucleotides while tracking the emission signal of each newly added nucleotide. The fluorescently labeled nucleotides have a 3‘ removable block that anchors a nucleotide-type fluorophore signal.
[0016] Array determination is performed in iterative cycles, each of which includes three steps: (a) extending the emerging strand by adding fluorescently labeled nucleotides, (b) exciting the fluorophore using one or more lasers of the optical system of the array determination device and generating an array determination image by imaging through different filters of the optical system, and (c) cleaving the fluorophore and removing the 3’ block in preparation for the next array determination cycle. The incorporation and imaging cycles are repeated up to a specified number of array determination cycles, defining the read length. Using this approach, each cycle collates a new position along the template strand.
[0017] The vast capabilities of Illumina sequencers stem from the ability to simultaneously perform and detect the CRT reactions of millions or even billions of specimens (e.g., clusters). A cluster contains approximately 1000 identical copies of a template strand, although the size and shape of the clusters vary. Clusters are grown from template strands by bridge amplification or exclusion amplification of the input library prior to the sequencing run. The purpose of amplification and cluster extension is to increase the intensity of the emitted signal because the imaging device cannot reliably detect single-stranded fluorophore signals. However, because the physical distance between the strands within a cluster is small, the imaging device perceives the cluster of strands as a single spot.
[0018] Sequencing is performed within a flow cell, which is a small glass slide that holds the input strands. The flow cell is connected to an optical system that includes a microscope for imaging, an excitation laser, and fluorescence filters. The flow cell contains multiple chambers called lanes. The lanes are physically separated from each other and can contain different tagged sequencing libraries that can be distinguished without cross-contamination of the samples. In some embodiments, the flow cell includes a patterned surface. A "patterned surface" refers to the arrangement of different regions within or on the exposed layer of a solid support. For example, one or more regions can be features where one or more amplification primers are present. This feature can be separated by interstitial regions where no amplification primers are present. In some embodiments, the pattern can be an x-y format of features in rows and columns. In some embodiments, the pattern can be a repeating sequence of features and / or interstitial regions. In some embodiments, the pattern can be a random arrangement of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Patent No. 8,778,849, U.S. Patent No. 9,079,148, U.S. Patent No. 8,778,848, and U.S. Patent Application Publication No. 2014 / 0243224, each of which is incorporated herein by reference.
[0019] In some embodiments, the flow cell includes an array of wells or depressions on the surface. This can be manufactured in a manner generally known in the art using a variety of techniques including, but not limited to, photolithography, stamping techniques, molding techniques, and microetching techniques. As understood in the art, the techniques used depend on the composition and shape of the array substrate.
[0020] Features within the patterned surface may be wells (e.g., microwells or nanowells) in an array on another suitable solid support having a patterned covalent gel such as glass, silicon, plastic, or poly(N-(5-azidoacetamylpentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication No. 2013 / 184796, International Publication No. 2016 / 066586, and No. 2015-002813, each of which is incorporated herein by reference in its entirety). This process creates gel pads for use in sequencing, which can be stable over multiple cycles of the sequencing operation. Covalently bonding the polymer to the wells is useful for maintaining the gel in the structured features over the lifetime of the structured substrate between various applications. However, in many embodiments, the gel need not be covalently bonded to the wells. For example, under some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Patent No. 8,563,477, which is incorporated herein by reference in its entirety) that is not covalently bonded to any part of the structured substrate can be used as the gel material.
[0021] In certain other embodiments, the structured substrate can be made by patterning a solid support material using wells (e.g., microwells or nanocells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof), such as an azidated version of SFA (azido-SFA), and polishing the gel-coated support, e.g., by chemical or mechanical polishing, thereby retaining the gel within the wells but removing or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the wells. A primer nucleic acid can be attached to the gel material. Then, a solution of target nucleic acids (e.g., fragmented human genome) can be contacted with the polished substrate such that individual target nucleic acids are seeded into individual wells via interaction with the primer to which the gel material is attached. Since the gel material is absent or inactive, the target nucleic acids do not occupy the interstitial regions. Amplification of the target nucleic acids will be limited to the wells because the absence or inactivation of the gel in the intervening regions prevents the outward movement of the growing nucleic acid colonies. This process is manufacturable, scalable, and utilizes conventional micro- or nano-fabrication methods.
[0022] The imaging device of the sequencing instrument (e.g., a solid-state imaging device such as a Charge-Coupled Device (CCD) or Complementary Metal-Oxide-Semiconductor (CMOS) sensor) takes snapshots at multiple locations along the lane in a series of non-overlapping regions called tiles. For example, there can be 64 or 96 tiles per lane. Tiles hold hundreds of thousands to millions of clusters.
[0023] The output of a sequencing run is a sequencing image, each showing the intensity emission of clusters and the background around them. The sequencing image shows the intensity emission generated as a result of incorporating nucleotides into the sequence during sequencing. The intensity emission results from the associated specimen / clusters and the background around them.
[0024] The array determination image 102 is supplied from a plurality of array determination devices, array determination runs, cycles, flow cells, tiles, wells, and clusters. In one embodiment, the array determination image is processed by an equalizer 104 on an imaging channel basis. The array determination run generates m images per array determination cycle corresponding to m imaging channels. In one embodiment, each imaging channel corresponds to one of a plurality of filter wavelength bands. In another embodiment, each imaging channel corresponds to one of a plurality of imaging events in an array determination cycle. In yet another embodiment, each imaging channel corresponds to a combination of illumination by a specific laser and imaging through a specific optical filter. In different embodiments such as 4-channel chemistry, 2-channel chemistry, and 1-channel chemistry, m is 4 or 2. In other embodiments, m is 1, 3, or greater than 4.
[0025] In another embodiment, the input data is based on a pH change induced by the release of hydrogen ions during molecular extension. The pH change is detected and converted into a voltage change proportional to the number of incorporated bases (e.g., in the case of Ion Torrent). In yet another embodiment, the input data is constructed from nanopore sensing using a biosensor to measure disruption of the current as an analyte passes through or near the opening of a nanopore. For example, Oxford Nanopore Technologies (ONT) sequencing is based on the following concept: passing a single strand of DNA (or RNA) through a membrane via a nanopore and applying a potential difference across the membrane. Nucleotides present within the pore affect the electrical resistance of the pore, and thus, current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal (the "squiggle" due to its appearance when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer data acquisition (DAC) values taken at a frequency of 4 kHz (for example). Using a DNA strand speed of ~450 base pairs per second, this gives, on average, about nine raw observations per base. This signal is then processed to identify breaks in the pore-opening signal corresponding to individual reads. The extension of these raw signals is base-called, the process of converting the DAC values into the sequence of DNA bases. In some embodiments, the input data includes normalized or scaled DAC values.Additional information regarding non-image-based sequence data can be found in U.S. Provisional Patent Application No. 62 / 849,132, entitled "Base Calling Using Convolutions," filed on May 16, 2020 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2), U.S. Provisional Patent Application No. 62 / 849,133, entitled "Base Calling Using Compact Convolutions," filed on May 16, 2019 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3), and U.S. Non-Provisional Patent Application No. 16 / 826,168, entitled "Artificial Intelligence-Based Sequencing," filed on March 21, 2019 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV).
[0026] Training Equalizer 104 generates a LUT bank having a plurality of LUTs (equalizer filters) 106 with sub-pixel resolution. In one embodiment, the number of LUTs 106 generated by equalizer 104 for the LUT bank depends on the number of sub-pixels into which the sensor pixels of the sequencing image 102 are or can be divided. For example, if the sensor pixels of the sequencing image 102 can each be divided into n×n sub-pixels (e.g., 5×5 sub-pixels), equalizer 104 generates n 2 LUTs 106 (e.g., 25 LUTs).
[0027] In one embodiment of training, the data from the sequencing image is binned by well sub-pixel position. For example, in the case of a 5×5 LUT, the center of the 1 / 25th of the well is in bin (1,1) (e.g., the upper left corner of the sensor pixel), the 1 / 25th of the well is in bin (1,2), and so on. The equalizer coefficients for each well-center-bin are determined using least squares estimation for a subset of the data from the wells within each bin. The input to equalizer 104 is the raw sense pixels of the sequencing image for those bins. The resulting estimated equalizer coefficients are different for each bin.
[0028] Each LUT has a plurality of coefficients learned from training. In one embodiment, the number of coefficients within the LUT corresponds to the number of sensor pixels used to base call the cluster. For example, if the local grid of sensor pixels (image or pixel patch) used to base call the cluster is sized p×p (e.g., 9×9 pixel patch), each LUT has p 2 coefficients (e.g., 81 coefficients).
[0029] Training is configured to generate equalizer coefficients that mix / combine the intensity values of pixels representing the intensity radiation from the target cluster being base called and the intensity radiation from one or more adjacent clusters so as to maximize the signal-to-noise ratio. The signal maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise minimized in the signal-to-noise ratio is the intensity radiation from adjacent clusters, i.e., the spatial crosstalk with some random noise added (e.g., to account for background intensity radiation). The equalizer coefficients are used as weights, and the mixing / combining includes performing an element-wise multiplication between the equalizer coefficients and the intensity values of the pixels to calculate a weighted sum of the intensity values of the pixels.
[0030] During training, the equalizer 104 learns, according to one embodiment, to maximize the signal-to-noise ratio by least squares estimation. Using least squares estimation, the equalizer 104 is trained to estimate the shared equalizer coefficients from the pixel intensities around the target well and the desired output. Least squares estimation is suitable for this purpose because it minimizes the squared error and outputs coefficients that take into account the effect of noise amplification.
[0031] The desired output is an impulse at the well position (point source) when the intensity channel is on and the background level when the intensity channel is off. In some embodiments, a ground truth based call 112 is used to generate the desired output. In some embodiments, the ground truth based call 112 is modified to account for the per-well DC offset, amplification factor, degree of polyclonality, and gain offset parameters included in the least squares estimate. In one embodiment, during training, the DC offset, i.e., the fixed offset, is calculated as part of the least squares estimate. During inference, the DC offset is added as a bias to each equalizer calculation.
[0032] In one embodiment, the desired output is estimated using an Illumina Real-time Analysis (RTA) based caller that does not use an equalizer. Details regarding RTA can be found in U.S. Patent Application No. 13 / 006,206, which is incorporated by reference as if fully set forth herein. The RTA based caller is used to emit the ground truth based call 112. This is because the base call error rate of RTA is low. The base call error is averaged over many training examples. In another embodiment, the ground truth based call 112 is supplied using aligned genomic data, which can use a reference genome and ground truth information incorporating knowledge from multiple sequencing platforms and sequencing runs for averaging noise and thus has better quality.
[0033] The ground truth base calls 112 are base-specific intensity values that reliably represent the intensity profiles of bases A, C, G, and T, respectively. A base caller such as RTA processes the sequencing image 102 and base calls the clusters by generating color-by-color intensity values / outputs for each base call. The color-by-color intensity values can be considered as base-by-base intensity values. This is because depending on the type of chemistry (e.g., two-color chemistry or four-color chemistry), the colors are mapped to each of the bases A, C, G, and T. The base with the closest matching intensity profile is called.
[0034] FIG. 16 shows one embodiment of a Gaussian fit for each base centered on the target for each base used as the ground truth value for error calculation during training. The base-by-base intensity outputs generated by the base caller for a large number of base calls in the training data (e.g., dozens, hundreds, thousands, or millions of base calls) are used to generate the base-by-base intensity distribution. FIG. 16 shows a chart of four Gaussian clouds that are the probability distributions of the base-by-base intensity outputs of bases A, C, G, and T, respectively. The intensity values at the centers of the four Gaussian clouds are used as the ground truth intensity targets given by the ground truth base calls 112 for bases A, C, G, and T, respectively, and are referred to here as intensity targets.
[0035] During training, it should be considered that the input image data supplied to the equalizer 104 is annotated with the base "A" as the ground truth base call. Next, the target / desired output of the equalizer 104 is the intensity value at the center of the green cloud in FIG. 16, i.e., the intensity target for base A. Similarly, for the ground truth base call of base "C", the desired output of the equalizer 104 is the intensity value at the center of the blue cloud in FIG. 16, i.e., the intensity target for base C. Therefore, the target or desired output during training of the equalizer 104 is the average intensity for each of the bases A, C, G, and T after being averaged in the training data. In one embodiment, the trainer 114 uses least squares estimation to adapt the coefficients of the equalizer 104 and minimizes the equalizer output error to these intensity targets.
[0036] In one embodiment, during training, the equalizer 104 applies the coefficients within a given look-up table (LUT) to the pixels of the array determination image labeled with a given base. This includes multiplying the coefficients element-wise using the intensity values of the pixels to generate a weighted sum of the intensity values, where the coefficients function / act / be used as weights. The weighted sum becomes the predicted output of the equalizer 104. Next, based on a cost / error function (e.g., sum of squared errors (SSE)), an error (e.g., least squares error, least mean square error) between the weighted sum and the intensity target determined for a given base (e.g., as the average intensity observed for the given base from the center of the corresponding intensity Gaussian fit) is calculated. A cost function such as SSE is a differentiable function used to estimate the equalizer coefficients using an adaptive approach, capable of evaluating the derivative of the error with respect to the coefficients, and using these derivatives to update the coefficients with values that minimize the error. This process is repeated until the updated coefficients no longer reduce the error. In other embodiments, batch least squares method is used to train the equalizer 104.
[0037] In other embodiments, the intensity distribution / gaussian cloud for each base shown in FIG. 16 is generated for each well, and noise can be corrected by adding a DC offset, an amplification factor, and / or a phase parameter. In this way, according to the well position of a specific well, the corresponding gaussian cloud for each base can be used to generate a target intensity value for that specific well.
[0038] In one embodiment, a bias term is added to the dot product that generates the output of equalizer 104. During training, the bias parameter can be estimated using a similar approach as that used to learn the equalizer coefficients, i.e., least squares or least mean squares (LMS). In some embodiments, the value of the bias parameter is a constant value equal to 1, i.e., a value that does not change with the input pixel intensity. There is one bias for each set of equalizer coefficients. The bias is learned during training and then fixed for use during inference. The learned bias, along with the learned coefficients of each LUT, represents the DC offset used in all equalizer calculations during inference. This bias accounts for random noise caused by different cluster sizes, different background intensities, changing stimulus responses, changing foci, changing sensor sensitivities, and changing lens aberrations.
[0039] In yet other decision-directed embodiments, the output of equalizer 104 is assumed to be correct for training purposes.
[0040] In another embodiment of training, equalizer 104 generates only a single LUT (equalizer filter) for a bin, and then uses interpolation filter 108 for each of the remaining bins to generate the remaining equalizer filters for the remaining bins. In this embodiment, the sensor pixels around all wells for all training examples are resampled / interpolated into a sufficiently aligned space (i.e., the wells are placed at the center of their respective pixel patches / local grids). Then, the resampled pixels for all examples are consistently aligned across all wells.
[0041] However, in order to apply the single equalizer filter generated by equalizer 104 in an actual online system for base calls, it is necessary to preprocess the raw sensor pixels of the array determination image and return them to a well-aligned space. That is, interpolation needs to be performed on the raw pixels around each well, and the interpolation parameters need to vary depending on the sub-pixel position of a given well. To avoid this interpolation process, the overall response for a given well sub-pixel position is pre-computed. By interpolating the raw pixel intensities into a well-aligned pixel space, well-aligned equalizer input values are calculated. The interpolation response and the equalizer response are convolved together, reducing the calculations. Since the interpolation filter varies by sub-pixel well position, this results in a different set of equalizer coefficients / equalizer filter for each sub-pixel well position, thereby generating the remaining LUTs for the remaining bins. Thus, in this embodiment of the training, only the coefficients of the single equalizer filter are trained during training, but the pre-computation process generates a bank of LUT-based equalizers by applying a bin-specific interpolation filter 108 together with the single equalizer filter. Here, the LUT index is the sub-pixel well position.
[0042] The trainer 114 can train the equalizer 104, train a plurality of trainers, and generate the training coefficients of the LUT 106. Examples of training techniques include least squares estimation, least squares method, least mean squares, and recursive least squares. In the least squares method, the parameters of the function are adjusted to best fit the data set such that the sum of the squares of the residuals is minimized. For details of the least squares estimation algorithm, see "Least squares method", https: / / en.wikipedia.org / w / index.php?title=Least_squares&oldid=951737821 (last visited on April 28, 2020). This is incorporated by reference as if fully set forth herein. The ordinary least squares is a kind of least squares method for estimation in a linear regression model. For details of the least squares algorithm, see "Ordinary least squares", https: / / en.wikipedia.org / w / index.php?title=Ordinary_least_squares&oldid=951770366 (last visited on April 28, 2020). This is incorporated by reference as if fully set forth herein. In other embodiments, other estimation algorithms and adaptive equalization algorithms can be used to train the equalizer 104.
[0043] The equalizer 104 can be trained in an offline mode. In the offline mode, according to one embodiment, the trained coefficients of the LUT 106 are generated using the following batch least squares equalization logic.
[0044]
Equation
[0045] In the above equation, the LUT coefficient is beta-hat, the pixel intensity is X, and the target is y. The DC term is also added to the pixel intensity and the coefficient (for example, an additional intensity term fixed to 1 in all cases). Next, as an example, consider that X is a matrix of size 82 (= 9×9 input intensity + a constant DC term) × the number of training examples in the batch, and Y is the target output for each training example. That is, each value is the intensity center of the ON / OFF cloud depending on the training example truth. Beta-hat is a set of coefficients that minimizes the sum of the squared residuals, and its size is also 82 (= 9×9 coefficients + 1 DC term).
[0046] Equalizer 104 can also adapt the coefficients of LUT 106 in an online mode to track changes such as temperature (e.g., optical distortion), focus, chemical, and mechanical inherent variations for each tile or sub-tile while the sequencer is operating and the array determination run is progressing periodically. In the online mode, the trained coefficients of LUT 106 are generated using adaptive equalization. In the online mode, the least mean squares method, which is a form of stochastic gradient descent, is used as the training algorithm. For details of the least mean squares algorithm, see "Least mean squares filter", https: / / en.wikipedia.org / w / index.php?title=Least_mean_squares_filter&oldid=941899198 (last visited on April 28, 2020). This is incorporated by reference as if it were fully described herein.
[0047] In the least mean squares method, the coefficients are moved in the direction of minimizing the cost function, which is the expected value of the mean square error, using the gradient of the mean square error with respect to each coefficient. This has a very low computational cost, with only multiplication and accumulation operations per coefficient being performed. No long-term storage is required except for the coefficients. The least mean squares method is suitable for processing large amounts of data (e.g., processing data from billions of clusters in parallel). Extensions of the least mean squares method include the normalized least mean squares method and the frequency domain least mean squares method, which can also be used here. In some embodiments, the least mean squares method can be applied in a decision-directed manner that assumes the inventors' decisions are correct, i.e., a method in which the inventors' error rate is very low and small μ values filter out updates disturbed by inaccurate base calls.
[0048] FIG. 18 shows one embodiment of an adaptive equalization technique that can be used to train the equalizer 104. Here, the equalization logic is y = x.h + d, where x is the input pixel intensity, h is the equalizer coefficient, and d is the DC offset. In one embodiment, x and h are row and column vectors each having a length of 81, respectively. This vector model corresponds to the inner product of a 9×9 matrix representing the input pixels and the coefficients. The cost is the expected value of the mean square error. With the update of the gradient, each coefficient moves in the direction of decreasing the expected value of the mean square error. This results in the following update.
[0049]
Number
[0050] In the above equation, h is a vector of equalizer coefficients (e.g., 9×9 equalizer coefficients), x is a vector of equalizer input intensities (e.g., 9×9 pixels within a pixel patch), and e is the error of the equalizer calculation performed using 81 values of x, i.e., there is only one error term per equalizer output.
[0051] Applying this update generates new estimated values for the 9×9 equalizer coefficients. These estimated values move the equalizer coefficients (on average) in a direction that reduces the mean squared error (MSE). For each equalizer coefficient, 81 updates are performed, one at a time. In some embodiments, Mu is a small constant used to change the adaptation rate / convergence speed. The update for the DC term can be calculated in a similar manner. The update for the gain term can also be calculated in a similar manner.
[0052] The coefficient set can be shared, for example, between tiles, regions of tiles, or flow cell surfaces. This is done by saving and restoring the coefficient set when the input data changes.
[0053] In some embodiments, since linear interpolation is applied to the coefficient set, the updates are applied slightly differently in the following manner. h(q,n+1)=h(q,n)+lambda_q.mu.x(n).e(n)
[0054] In the above equation, h(q,n) is the weight q at cycle n, lambda_q is the linear interpolation weight for a particular set of coefficients, and can include four updates per equalizer output by linear interpolation in two dimensions.
[0055] The recursive least squares method is an extension of the least squares method to a recursive algorithm. For details of the recursive least squares algorithm, see "Recursive least squares filter", https: / / en.wikipedia.org / w / index.php?title=Recursive_least_squares_filter&oldid=916406502 (last visited on April 28, 2020). This is incorporated by reference as if fully set forth herein.
[0056] In a multi-domain embodiment, the LUTs 106 and their trained coefficients can be generated along multiple domains. Examples of domains include sequencers or sequencing instruments / machines (e.g., Illumina's NextSeq, MiSeq, HiSeq, and their respective models), sequencing protocols and chemistries (e.g., bridge amplification, exonuclease amplification), sequencing runs (e.g., forward and reverse), sequencing illumination (e.g., structured, unstructured, angled), sequencing devices (e.g., overhead CCD cameras, underlying CMOS sensors, one laser, multiple lasers), imaging techniques (1 channel, 2 channels, 4 channels), flow cells (e.g., patterned, unpatterned, embedded in a CMOS chip, underlying CCD camera), and spatial resolution on the flow cell (e.g., different regions or quadrants within the flow cell (e.g., different tiles on the flow cell (e.g., edge wells on tiles close to the laser or camera or fluid system)) and different regions within the tile (e.g., different lanes on the tile (e.g., edge wells on lanes close to the laser or camera or fluid system))). One of ordinary skill in the art will understand that other selectable domains and parameters typically associated with sequencing (e.g., image processing algorithms, image alignment algorithms, ground truth annotation schemes (e.g., continuous labels such as intensity values, hard labels such as one-hot encoding, soft labels such as softmax scores), temperature, focus, lenses, sequencing reagents, sequencing buffers) are similarly included.
[0057] Using the array determination images generated using each domain, different separate training sets can be created for each domain. The equalizer 104 can be trained using the discrete training sets to generate a LUT having coefficients trained for the corresponding regions. The training coefficients specially trained and generated for each domain in a plurality of domains can be stored and accessed during the online mode according to which domain or combination of domains is being used in the current or ongoing array determination operation. For example, for an array determination operation, a first set of coefficients suitable for the edge wells of the flow cell can be used together with a second set of coefficients suitable for the center wells of the same flow cell.
[0058] In one embodiment, the configuration file can specify different combinations of domains and can be analyzed during the online mode to select different sets of coefficients specific to the domains identified by the configuration file.
[0059] In embodiments of multiple trainings, the equalizer 104 undergoes not only training but also pre-training. That is, the LUTs 106 and their coefficients are first trained in a pre-training stage using a first training technique and then re-trained or further trained in a further training stage using a second training technique. The first and second training techniques can be any of the training techniques described above. The first and the second training techniques may be the same or different. For example, the pre-training stage may be an offline mode using a batch least squares training technique, and the training stage may be an online mode using an iterative probabilistic least mean squares technique.
[0060] In some embodiments, the multi - domain and multi - training implementation can be combined such that domain - specific coefficients are pre - trained and then further trained in a domain - specific manner. That is, further training (e.g., in online mode) represents that particular domain and retrains the coefficients of that particular domain using only data similar to the data used in the pre - training stage. In other knowledge transfer embodiments, pre - training and training can use training data from across the domain. For example, a set of coefficients is generated during pre - training using images from patterned flow cells, but is retrained during a subsequent training stage using images from unpatterned flow cells.
[0061] Spatial crosstalk attenuator FIG. 2 shows one embodiment of using the trained LUT / equalizer filter 106 of FIG. 1 to attenuate spatial crosstalk from sensor pixels and use the crosstalk - corrected sensor pixels to base - call cluster calls to a base call. The trained equalizer - based caller 104 operates during the inference stage when the base call is made. In some embodiments, the actions shown in FIG. 2 are performed in a pre - processing stage before the base - call stage, generating crosstalk - corrected image data used by the base - caller for the base call.
[0062] In one embodiment, the equalizer coefficients are applied to pixel patches 120 (local grids of image patches or sensor pixels) extracted from the array determination image 116 based on the imaging channel and the target cluster. With respect to the imaging - channel - based, in some embodiments, each array determination image has image data of a plurality of imaging channels. Consider an optical system of an Illumina sequencer that uses two different imaging channels, namely a red channel and a green channel. Then, in each array determination cycle, the optical system generates a red image with red - channel intensity and a green image with green - channel intensity, which together form a single array determination image (like the RGB channels of a typical color image).
[0063] During training, the coefficients are trained / configured to maximize the signal-to-noise ratio (SNR) by minimizing the error between the predicted / estimated output and the desired / actual output. An example of the error is the mean squared error (MSE) or the mean squared deviation (MSD). The signal maximized in the signal-to-noise ratio is the intensity radiation from the base-called target cluster (e.g., the cluster at the center of the image patch), and the noise minimized in the signal-to-noise ratio is the intensity radiation from one or more adjacent clusters, i.e., spatial crosstalk, in addition to other noise sources (e.g., for explaining background intensity radiation). The trained coefficients are multiplied element-wise to the pixels of the image patch to calculate the weighted sum of the intensity values of the pixels. Next, the weighted sum is used to base-call the target cluster.
[0064] In one embodiment, patch extractor 118 extracts a red pixel patch from the red channel and a green pixel patch for the green channel from a single sequencing image. In other embodiments, the red pixel patch is extracted from the red sequencing image of the target sequencing cycle and the green pixel patch is extracted from the green sequencing image of the target sequencing cycle. The coefficients of LUT 106 are used to generate a weighted sum for the red pixel patch and a weighted sum for the green pixel patch. Next, both the weighted sum for the red pixel patch and the weighted sum for the green pixel patch are used to base call the target cluster. Pixel patch 120 has dimensions w×h, where w (width) and h (height) are any numbers in the range of 1 and 10,000 (e.g., 3×3, 5×5, 7×7, 9×9, 15×15, 25×25). In some embodiments, w and h are the same. In other embodiments, w and h are different. One of ordinary skill in the art will understand that one, two, three, four, or more channels or image data can be generated per sequencing cycle for the target cluster, and one, two, three, four, or more patches are each extracted and one, two, three, four or more weighted sums are each generated to base call the target cluster.
[0065] With respect to the target cluster base for extracting pixel patch 120 from sequencing image 116, pixel extractor 118 extracts pixel patch 120 such that the center pixel of each extracted pixel patch includes the center of the target cluster / well based on the position of the center of the cluster / well on sequencing image 116. In some embodiments, patch extractor 118 positions the cluster / well center on the sequencing image, identifies the pixels of the sequencing image that include the cluster / well center (i.e., the center pixel), and extracts a pixel patch of the pixels in the immediately adjacent neighborhood surrounding the center pixel.
[0066] Figure 2 visualizes an example of a sequencing image 200 that includes at least five cluster / well centers / point sources on a flow cell. The pixels of the sequencing image 200 show the intensity radiation from the target cluster 1 (blue), and the intensity radiation from additional adjacent clusters 2 (purple), cluster 3 (orange), cluster 4 (brown), and cluster 5 (green).
[0067] Figure 3 visualizes an example of extracting a pixel patch 300 (yellow) from the sequencing image 200, such that the center of the target cluster 1 (blue) is included in the central pixel 206 of the pixel patch 300. Figure 3 also shows other pixels 202, 204, 214, and 216 that respectively include the centers of the adjacent clusters 2 (purple), cluster 3 (orange), cluster 4 (brown), and cluster 5 (green).
[0068] Figure 4 visualizes an example of a cluster-to-pixel signal 400. In one embodiment, the sensor pixel (yellow) is in the pixel plane. Spatial crosstalk is caused by clusters 412 that are periodically distributed in the sample plane (e.g., flow cell). In one embodiment, the target cluster and additional adjacent clusters are periodically distributed in a rhombus shape on the flow cell and immobilized on the wells of the flow cell. In another embodiment, the target cluster and additional adjacent clusters are periodically distributed on a hexagonal flow cell and immobilized on the wells of the flow cell. The signal cone 402 from the cluster is optically coupled to the local grid of the sensor pixel (e.g., pixel patch 300) via at least one lens (e.g., one or more lenses of an overhead or adjacent CCD camera).
[0069] In addition to rhombus and hexagon, the clusters can be arranged in other regular shapes such as square, rhomboid, triangle, etc. In still other embodiments, the clusters are arranged on the sample plane in a random and aperiodic arrangement. Those skilled in the art will understand that the clusters can be arranged on the sample plane in any arrangement as required by a particular sequencing embodiment.
[0070] Figure 5 visualizes an example of the cluster pair pixel signal overlap 500. The signal cones 402 overlap and impinge on the sensor pixels, generating spatial crosstalk 502.
[0071] Figure 6 visualizes an example of the cluster signal pattern 600. In one embodiment, the cluster signal pattern 600 follows an attenuation pattern 602. In this case, the cluster signal is strongest at the cluster center and attenuates as it propagates away from the cluster center.
[0072] Figure 6 also shows an example of an equalizer coefficient 604 that is trained / configured to maximize the signal-to-noise ratio by calculating a weighted sum of the intensity emissions from target cluster 1 and the intensity emissions from adjacent clusters 2, 3, 4, and 5. The equalizer coefficient 604 functions as a weight. The weighted sum is calculated by element-wise multiplying a first matrix that includes the equalizer coefficient 604 and a second matrix that includes pixel intensity values, where each pixel intensity value is the sum of the emissions from one or more of clusters 1, 2, 3, 4, and 5 and other noise sources within the system measured by the pixel sensor.
[0073] Figure 7 visualizes an example of a sub-pixel LUT grid 700 used to attenuate spatial crosstalk from pixel patch 300. Each pixel within pixel patch 300 is divisible into a plurality of sub-pixels. In Figure 7, pixel 206 that includes the center of target cluster 1 (blue) is divided into the same number of sub-pixels as the number of trained LUTs 106. That is, pixel 206 is divided into the same number of sub-pixels as the number of bins in which equalizer 104 generated LUT 106 during training. As a result, each sub-pixel of pixel 206 corresponds to a respective LUT within the LUT bank generated by equalizer 104 using decision-directed feedback and least squares estimation.
[0074] In the example shown in FIG. 7, pixel 206 (central pixel) is divided into a 5×5 sub-pixel LUT grid 700, generating 25 sub-pixels each corresponding to one of the 25 LUTs (equalizer filters) generated by the adaptation filter 104 as a result of training. Each of the 25 LUTs includes coefficients configured to mix / combine the intensity radiation from target cluster 1 and the intensity values of the pixels within pixel patch 300 indicative of the intensity radiation from adjacent clusters 2, 3, 4, and 5 so as to maximize the signal-to-noise ratio. The signal maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise minimized in the signal-to-noise ratio is the intensity radiation from adjacent clusters 2, 3, 4, and 5, i.e., the spatial crosstalk with some random noise added (e.g., to account for background intensity radiation). The LUT coefficients are used as weights, and the mixing / combining includes performing an element-wise multiplication between the LUT coefficients and the intensity values of the pixels within pixel patch 300 to calculate a weighted sum of the intensity values of the pixels.
[0075] The number of coefficients in each of the 25 LUTs is the same as the number of pixels in pixel patch 300, i.e., a 9×9 coefficient grid in each LUT for the 9×9 pixels in pixel patch 300. This is because the coefficients are multiplied element-wise using the pixels within pixel patch 300.
[0076] In one embodiment, a pixel-subpixel converter (not shown in FIG. 1B) divides pixel 206 into sub-pixel LUT grid 700 based on a preset pixel divisor parameter (e.g., 1 / 5 pixel per sub-pixel to generate a 5×5 sub-pixel LUT grid 700). For example, the pixel can be divided into 5 sub-pixel bins having the following boundaries, i.e., -0.5, -0.3, -0.1, 0.1, 0.3, 0.5.
[0077] Note that in FIG. 7, the center of the target cluster 1 (blue) is substantially concentric with the center of the conversion pixel 702. This is because the alignment determination image 200 is aligned with the template image to determine the affine transformation and non-linear transformation parameters, (ii) the position coordinates of the target cluster 1 (blue) are converted into the image coordinates of the alignment determination image 702 using the parameters, and (iii) interpolation is applied using the converted position coordinates of the target cluster 1 (blue) to make its center substantially concentric with the center of the converted pixel 200. Thus, the alignment determination image 200, and thus the pixel patch 300, is resampled, and the center of the target cluster 1 (blue) becomes substantially concentric with the center of the converted pixel 702. The positions of the wells in the sample plane are known and can be used to calculate where in the raw pixel space the equalizer input for a particular well is. Next, interpolation can be used to restore the intensities at those positions from the raw image.
[0078] FIG. 8 shows the selection of the LUT / equalizer filter from the LUT bank 106 based on the sub-pixel position of the cluster / well center within the pixel. The center of the target cluster (blue) is at a particular sub-pixel 12 of the sub-pixel LUT grid 700, and since the particular sub-pixel 12 of the pixel 206 corresponds to the LUT12 within the LUT bank 106, the LUT selector 122 selects the LUT12 and its coefficients from the LUT bank 106 for application to the pixels of the pixel patch 300. Next, the element-wise multiplier 134 multiplies the intensity values of the pixels within the pixel patch 300 by the coefficients of the LUT12 element-wise and sums the products of the multiplications to generate an output (e.g., the weighted sum 136). This output is used to base-call the target cluster 1 (e.g., supply this output as an input to the base-caller 138).
[0079] As discussed above with respect to FIGS. 7 and 8, the equalizer 104 implements the following equalization logic when the target cluster is substantially concentric with the center of the pixel.
[0080]
Number
[0081] In the above formula, the well center coordinates (m, n) are integers in order to ensure that the well is substantially aligned with the pixel. p(i, j) is the pixel intensity at positions i and j. w(i, j) is the equalization weight for the pixel at positions i and j. i and j are the summation limits acting over the pixel range surrounding the well centered on p(m, n), for example, -4 <= i <= 4, -4 <= j <= 4, and the output is the weighted average of the input pixels.
[0082] FIG. 9 shows an embodiment in which the center (blue) of target cluster 1 is not substantially concentric with the center of pixel 206 because resampling as considered with respect to FIG. 8 is not performed. In such an embodiment, interpolation is performed between the selected set of LUTs 124 to generate an interpolation LUT with interpolation coefficients. The interpolation LUT with interpolation coefficients is also referred to herein as weight kernel 132.
[0083] First, as in FIG. 8, a first LUT corresponding to a specific sub-pixel containing the center of target cluster 1 (blue), i.e., LUT12, is selected. Next, the LUT selector 122 selects an additional sub-pixel look-up table corresponding to the sub-pixel most continuously adjacent to the specific sub-pixel from the bank 106 of the sub-pixel look-up tables. In FIG. 9, the sub-pixels most closely adjacent to the specific sub-pixel 12 are sub-pixels 7, 8, and 13, and thus, LUTs 7, 8, and 13 are selected from the LUT bank 106, respectively.
[0084] FIG. 10 shows one embodiment of interpolating between a selected set of LUTs and generating respective LUT weights. Interpolator 126 is composed of interpolation logic (e.g., linear, bilinear, or bicubic interpolation) that uses the coefficients of selected LUTs 12, 7, 8, and 13 to generate weights 128 for each of LUTs 12, 7, 8, and 13.
[0085] FIGS. 13A, 13B, 13C, 13D, 13E, 13F show examples of the coefficients of LUTs 12, 7, 8, 13. These figures also show examples 1312, 1322, and 1332 of the interpolation logic used by interpolator 126 to calculate weights 128 for LUTs 12, 7, 8, and 13. These figures also show examples of the weights 128 calculated for LUTs 12, 7, 8, and 13. These figures are snapshots of an Excel sheet. The blue arrows and color-coding in these figures are generated by Excel's Track Precedence feature to show the interpolation logic.
[0086] FIG. 11 shows a weight kernel generator 130 that generates a weight kernel 132 using weights 128 calculated for LUTs 12, 7, 8, and 13. FIG. 14A shows an example of the weight kernel 132. FIGS. 14B and 14C show an example 1402 of weight kernel generation logic used by the weight kernel generator 130 to generate the weight kernel 132 from the weights 128 calculated for LUTs 12, 7, 8, and 13. The weight kernel 132 includes interpolation pixel coefficients 1412 configured to mix / combine the intensity values of pixels within a pixel patch 300 representing the intensity radiation from the target cluster 1 and the intensity radiation from the adjacent clusters 2, 3, 4, and 5 so as to maximize the signal-to-noise ratio. The signal maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise minimized in the signal-to-noise ratio is the intensity radiation from the adjacent clusters 2, 3, 4, and 5, i.e., some random noise added to the spatial crosstalk (e.g., to account for background intensity radiation). The interpolation pixel coefficients 1412 are used as weights, and the mixing / combining includes performing element-wise multiplication between the LUT coefficients and the intensity values of the pixels within the pixel patch 300 to calculate a weighted sum of the intensity values of the pixels.
[0087] FIG. 12 shows an element-by-element multiplier 134 that multiplies the interpolation pixel coefficients 1412 of the weight kernel 132 by the intensity values of the pixels within the pixel patch 300 on a per-element basis and sums the intermediate products 1202 of the multiplication to produce a weighted sum 136. For each well, the optical system operates on a point source (cluster intensity in the well) having a point spread function (response of the optical system). In some embodiments, a bias is added to the operation to account for noise caused by different cluster sizes, different background intensities, varying stimulus responses, varying focus, varying sensor sensitivities, and varying lens aberrations. The captured image is a superposition of responses from all wells. The selected LUT equalizes the system response around each well to estimate the intensity of the point source from that well. That is, it processes the PSF intensity across the local neighborhood / grid of sensor pixels to estimate the intensity of the point source that generated the local grid of sensor pixels. This equalizer operation is a dot product on the sensor pixels within the local grid having equalizer coefficients.
[0088] As discussed above with respect to FIGS. 9, 10, 11, and 12, equalizer 104 implements the following equalization logic when the target cluster is not substantially concentric with the center of the central pixel. When the well is not at the center of the pixel, the output of equalizer 104 is calculated as a function of the virtual pixel intensity p'(i,j) derived from the actual pixel intensity of the pixels of the array determination image.
[0089] [Number]
[0090] In the above formula, the well center coordinates (m,n) can have a fractional part. Each "virtual" equalizer input p'(i,j) is generated by applying an interpolation filter to the pixel neighborhood. In one embodiment, a windowed sinc low-pass filter h(x,y) is used for interpolation. In other embodiments, other filters such as a bilinear interpolation filter can be used.
[0091] The virtual pixel at position (i, j) is calculated as follows using an interpolation filter.
[0092]
Equation
[0093] By combining equations (1) and (2), equalizer 104 uses only the raw pixel intensity as follows.
[0094]
Equation
[0095] In the above equation, h is fixed given the sub-pixel offsets frac(m), frac(n). u, v specify the range of pixels used for interpolation to generate the equalizer input, and i, j specify the range of virtual pixels used as input to equalizer 104.
[0096] For a given sub-pixel offset, only the input pixels change and the filter or weights do not. Therefore, a fixed set of interpolated equalizer coefficients is calculated for the center of each binned sub-pixel offset. The output is as follows.
[0097]
Equation
[0098] In the above equation, h fm,fn represents the LUT equalizer coefficient for a well having binned fractional sub-pixel offsets fm, fn, where (fm, fn) are LUT indices.
[0099] Figures 15A and 15B show how the interpolation pixel coefficient 1412 of the weight kernel maximizes the signal-to-noise ratio and restores the signal underlying target cluster 1 from signals corrupted by crosstalk from clusters 2, 3, 4, and 5.
[0100] The weighted sum 136 is supplied as an input to the base call 138 to generate the base call 140. The base caller 138 can be a non-neural-network-based base caller or a neural-network-based base caller, examples of both of which are described in applications incorporated herein by reference such as U.S. Patent Application Nos. 62 / 821,766 and 16 / 826,168.
[0101] In yet other embodiments, the need for interpolation is eliminated by having large LUTs each having a number of sub-pixel bins (e.g., 50, 75, 100, 150, 200, 300, etc. sub-pixel bins per LUT).
[0102] Figure 19A shows a graph representing the base call error rate using an image from a NovaSeq sequencer. The error rate is shown on the cycles on the x-axis. 0.004 on the y-axis represents a base call error rate of 0.4%. The error rate here is calculated after mapping and aligning the reads to the Phi-X reference. The Phi-X reference is a high-confidence ground truth set. The blue line is the legacy base caller. The red line is the improved equalizer-based base caller 104 disclosed herein. The overall error rate is reduced by 57% at the expense of limited extra calculations. The base error rate in later cycles is higher due to extra noise in the system (e.g., pre-phasing / phasing, cluster dimming). The improved performance in later cycles is valuable as it shows that longer reads can be supported. The cycle-to-cycle performance variation is also significantly reduced.
[0103] Figures 19B-1 and 19B-2 show another example of the performance results of the disclosed equalizer-based base caller 104 for sequence data from the NovaSeq sequencer and the Vega sequencer. For the NovaSeq sequencer, the disclosed equalizer-based base caller 104 reduces the base call error rate by more than 50%. For the Vega sequencer, the disclosed equalizer-based base caller 104 reduces the base call error rate by more than 35%.
[0104] Figure 19C shows another example of the performance results of the disclosed equalizer-based base caller 104 for sequence data from the NextSeq2000 sequencer. For the NextSeq2000 sequencer, the disclosed equalizer-based base caller 104 reduces the base call error rate by an average of 10% without including throughput.
[0105] Figure 19D shows an embodiment of the computing resources required by the disclosed equalizer-based base caller 104. As shown, the disclosed equalizer-based base caller 104 can be executed using a small number of CPU threads in the range of 2 to 7 threads. Therefore, the disclosed equalizer-based base caller 104 is a computationally efficient base caller, which significantly reduces the base error rate and can thus be integrated into most existing sequencers without any additional computation or special processors such as GPUs, FPGAs, ASICs, etc.
[0106] In the present application, the terms "cluster", "well", "sample" and "fluorescent sample" are used interchangeably since a well contains the corresponding cluster / sample / fluorescent sample. As defined herein, "sample" and its derivatives are used in the broadest sense and include any sample, culture, etc. that is suspected of containing a target. In some embodiments, the sample comprises nucleic acids in the form of DNA, RNA, PNA, LNA, chimeras or hybrids. The sample can include any biological sample, clinical sample, surgical sample, agricultural sample, atmospheric sample or water sample that contains one or more nucleic acids. The term also includes any isolated nucleic acid sample, for example, genomic DNA, freshly frozen or formalin-fixed paraffin-embedded nucleic acid samples. The sample can be a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched) from a single individual such as tumor samples and normal tissue samples, or a sample from a single source containing two different forms of genetic material such as maternal and fetal DNA obtained from a maternal subject, or a sample containing plant or animal DNA and which may be derived from the presence of contaminating bacterial DNA. In some embodiments, the source of the nucleic acid material can include nucleic acids obtained from a neonate, such as is typically used in neonatal screening.
[0107] The nucleic acid sample can contain high molecular weight substances such as genomic DNA (gDNA). The sample can contain low molecular weight substances such as nucleic acid molecules obtained from FFPE or stored DNA samples. In another embodiment, the low molecular weight substances contain enzymatically or mechanically fragmented DNA. The sample can contain cell-free circulating DNA. In some embodiments, the sample can contain nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resection, and other clinical or laboratory-obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic or pathogenic sample. In some embodiments, the sample can contain nucleic acid molecules obtained from animals such as human or mammalian sources. In another embodiment, the sample can contain nucleic acid molecules obtained from non-mammalian sources such as plants, bacteria, viruses or fungi. In some embodiments, the source of the nucleic acid molecules can be a preserved or extinct sample or species.
[0108] Furthermore, the methods and compositions disclosed herein can be useful for amplifying nucleic acid samples having low-quality nucleic acid molecules such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation, or forensic samples obtained by a law enforcement agency, one or more military forces or such personnel. The nucleic acid sample can be, for example, a purified sample or a crude DNA containing a lysate derived from a substrate impregnated with a buccal swab, paper, cloth, or saliva, blood, or other body fluid. In itself, in some embodiments, the nucleic acid sample can contain a small amount of DNA such as genomic DNA or a fragmented portion. In some embodiments, the target sequence can be present in one or more body fluids including, but not limited to, blood, sputum, plasma, semen, urine and serum. In some embodiments, the target sequence can be obtained from hair, skin, tissue samples, autopsy or the body of a victim. In some embodiments, the nucleic acid containing one or more target sequences can be obtained from a dead animal or human. In some embodiments, the target sequence can include nucleic acids obtained from non-humans such as microorganisms, plant cells or entomology. In some embodiments, the target sequence or the amplified target sequence is for human identification. In some embodiments, the present disclosure generally relates to methods for identifying characteristics of forensic samples. In some embodiments, the present disclosure generally relates to methods of human identification using one or more of the target-specific primers disclosed herein, or one or more target-specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic sample or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein, or using the primer criteria outlined herein.
[0109] As used herein, the term "adjacent," when used with respect to two reaction sites, means that no other reaction site is present between the two reaction sites. The term "adjacent" may have a similar meaning when used with respect to adjacent detection paths and adjacent photodetectors (e.g., adjacent photodetectors do not have other photodetectors therebetween). In some cases, a reaction site may not be adjacent to other reaction sites, but may still be in the immediate vicinity of other reaction sites. If a fluorescence emission signal from a first reaction site is detected by a photodetector associated with a second reaction site, the first reaction site may be in the immediate vicinity of the second reaction site. More specifically, the first reaction site may be in the immediate vicinity of the second reaction site when the photodetector associated with the second reaction site detects, for example, crosstalk from the first reaction site. Adjacent reaction sites may be continuous so as to be adjacent to each other, or the adjacent sites may be discontinuous with intervening spaces therebetween.
[0110] Technical improvements and terms All documents and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and web pages, are hereby expressly incorporated by reference in their entirety, regardless of the form of such documents and similar materials. If one or more of the incorporated documents and similar materials differ from or conflict with this application, for example, in defined terms, term usage, described techniques, etc., this application prevails. Further information regarding terms can be found in U.S. Patent Application No. 16 / 826,168, filed on March 21, 2019, entitled "Artificial Intelligence-Based Sequencing" (Attorney Docket No. ILLM1008-20 / IP-1752-PRV) and U.S. Provisional Patent Application No. 62 / 821,766, filed on March 21, 2020, entitled "Artificial Intelligence-Based Sequencing" (Attorney Docket No. ILLM1008-9 / IP-1752-PRV).
[0111] The disclosed technology uses neural networks to improve the quality and quantity of nucleic acid sequence information obtainable from a nucleic acid template or its complement, such as a nucleic acid sample like a DNA or RNA polynucleotide or other nucleic acid sample. Thus, certain implementations of the disclosed technology provide higher throughput polynucleotide sequencing, e.g., a higher rate of DNA or RNA sequence data collection, higher efficiency in sequence data collection, and / or lower cost of obtaining such sequence data, compared to previously available methods.
[0112] The disclosed technology uses neural networks to identify the centers of solid-phase nucleic acid clusters and analyze the optical signals generated during the sequencing of such clusters to unambiguously distinguish between adjacent, neighboring, or overlapping clusters and assign sequencing signals to a single discrete source cluster. Thus, these and related embodiments enable the recovery of meaningful information, such as sequence data, from regions of high-density cluster arrays, where useful information may not have been previously obtainable due to the confounding effects of overlapping or very closely spaced neighboring clusters. This includes the effects of overlapping signals (such as those used in nucleic acid sequencing).
[0113] As described in more detail below, in certain embodiments, a composition is provided that includes a solid support to which one or more nucleic acid clusters are immobilized. Each cluster includes a plurality of immobilized nucleic acids of the same sequence and has a distinguishable center with a detectable central label as provided herein, and the distinguishable center is distinguishable from the immobilized nucleic acids in the surrounding region within the cluster. Also described herein are methods for making and using such clusters having a distinguishable center.
[0114] Embodiments of the present disclosure find use in a number of situations, and the advantages are obtained from the ability to identify, determine, annotate, record, or otherwise assign a position that is substantially centered within a cluster, and will find use in many situations. High-throughput nucleic acid sequencing, the development of image analysis algorithms for assigning optical or other signals to individual source clusters, and other applications where the recognition of the center of an immobilized nucleic acid cluster is desirable and beneficial are desired.
[0115] In certain embodiments, the invention contemplates methods related to high-throughput nucleic acid analysis such as nucleic acid sequencing (e.g., “sequencing”). Exemplary high-throughput nucleic acid analyses include, but are not limited to, de novo sequencing, resequencing, whole genome sequencing, gene expression analysis, gene expression monitoring, epigenetics analysis, genome methylation analysis, allele specific primer extension (APSE), genetic diversity profiling, whole genome polymorphism discovery and analysis, single nucleotide polymorphism analysis, hybridization-based sequencing methods, and the like. Those skilled in the art will understand that a variety of different nucleic acids can be analyzed using the methods and compositions of the present invention.
[0116] The practice of the present invention is described in relation to nucleic acid sequencing, but they are applicable in any field where image data acquired at different times, spatial positions, or image data acquired from other temporal or physical perspectives are analyzed. For example, the methods and systems described herein are useful in the fields of molecular biology and cell biology where image data from microarrays, biological specimens, cells, organisms, etc. are acquired, acquired at different times or perspectives, and analyzed. Images can be obtained using any number of techniques known in the art, including but not limited to fluorescence microscopy, optical microscopy, confocal microscopy, optical imaging, magnetic resonance imaging, tomographic scanning, etc. As another example, the methods and systems described herein can be applied when image data acquired by surveillance, aerial, or satellite imaging techniques, etc. are acquired, acquired at different times or perspectives, and analyzed. The methods and systems are particularly useful for analyzing images acquired within a field of view, within which the specimens being observed remain in the same location relative to each other within the field of view. However, the specimens may have different characteristics in separate images, for example, the specimens may appear different in separate images of the field of view. For example, a specimen may appear different in color from a given specimen detected in a different image, show a change in the intensity of the signal detected for a given specimen in different images, or even show the appearance of a signal for a given specimen in one image and the disappearance of the signal for the specimen in another image.
[0117] As used herein, the term "analyte" is intended to mean a point or region of a pattern that can be distinguished from other points or regions according to relative position. An individual analyte can contain one or more molecules of a particular type. For example, an analyte can contain a single target nucleic acid molecule having a particular sequence, or an analyte can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). Different molecules that are different analytes of the pattern can be differentiated from one another according to the location of the analyte within the pattern. Exemplary analytes include wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, raised portions on a substrate, pads of gel material on a substrate, or channels within a substrate.
[0118] Any of the various target analytes to be detected, characterized, or identified can be used with the devices, systems, or methods described herein. Exemplary analytes include, but are not limited to, nucleic acids (e.g., DNA, RNA or their analogs), proteins, polysaccharides, cells, antibodies, epitopes, receptors, ligands, enzymes (e.g., kinases, phosphatases or polymerases), small molecule drug candidates, cells, viruses, organisms, etc.
[0119] The terms "sample", "nucleic acid", "nucleic acid molecule", and "polynucleotide" are used interchangeably herein. In various embodiments, the nucleic acid may be used as a template as provided herein (e.g., a nucleic acid template or a nucleic acid complement complementary to a nucleic acid template) for a particular type of nucleic acid analysis, including, but not limited to, nucleic acid amplification, nucleic acid expression analysis, and / or nucleic acid sequencing, or suitable combinations thereof. Nucleic acids in certain embodiments include, for example, linear polymers of deoxyribonucleotides in 3'-5' phosphodiester linkages, or deoxyribonucleic acid (DNA), such as single-stranded and double-stranded DNA, genomic DNA, copy DNA or complementary DNA (cDNA), recombinant DNA, or any form of synthetic or modified DNA. In other embodiments, nucleic acids include, for example, linear polymers of ribonucleotides in 3'-5' phosphodiester linkages, or other linkages such as ribonucleic acid (RNA), such as single-stranded and double-stranded RNA, messenger (mRNA), copy RNA or complementary RNA (cRNA), alternatively spliced mRNA, ribosomal RNA, small nucleolar RNA (snoRNA), microRNA (miRNA), small interfering RNA (sRNA), piwi RNA (piRNA), or any form of synthetic or modified RNA. The nucleic acids used in the compositions and methods of the present invention may vary in length and may be intact or full-length molecules or fragments, or smaller portions of larger nucleic acid molecules. In certain embodiments, the nucleic acid may have one or more detectable labels, as described elsewhere herein.
[0120] The terms "sample", "cluster", "nucleic acid cluster", "nucleic acid colony", and "DNA cluster" are used interchangeably and refer to multiple copies of a nucleic acid template and / or its complement that are bound to a solid support. Typically, in certain preferred embodiments, a nucleic acid cluster comprises multiple copies of a template nucleic acid and / or its complement that are bound to the solid support via their 5' ends. The copies of the nucleic acid strands that make up the nucleic acid cluster may be in single-stranded or double-stranded form. The copies of the nucleic acid templates present within the cluster can have nucleotides at corresponding positions that differ from one another, for example, due to the presence of a labeled moiety. The corresponding positions can also include analog structures that have different chemical structures but similar Watson-Crick base pairing properties, such as in the case of uracil and thymine.
[0121] A colony of nucleic acids may also be referred to as a "nucleic acid cluster". Nucleic acid colonies can be optionally created by cluster amplification or bridge amplification techniques, as described in more detail elsewhere in this specification. Multiple repeats of a target sequence can be present within a single nucleic acid molecule, such as a scrambling agent created using a rolling circle amplification procedure.
[0122] The nucleic acid clusters of the present invention can have different shapes, sizes, and densities depending on the conditions used. For example, the clusters can have a shape that is substantially circular, polyhedral, doughnut-shaped, or ring-shaped. The diameter of the nucleic acid clusters can be designed to be about 0.2 μm to about 6 μm, about 0.3 μm to about 4 μm, about 0.4 μm to about 3 μm, about 0.5 μm to about 2 μm, about 0.75 μm to about 1.5 μm, or any intervening diameter. In certain embodiments, the diameter of the nucleic acid clusters is about 0.5 μm, about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The diameter of the nucleic acid clusters can be affected by a number of parameters including, but not limited to, the number of amplification cycles performed in the production of the clusters, the length of the nucleic acid template, or the density of the primers attached to the surface on which the clusters are formed. The density of the nucleic acid clusters can typically be designed to be in the range of 0.1 / mm 2 to 1 / mm 2 to 10 / mm 2 to 100 / mm 2 to 1,000 / mm 2 to 10,000 / mm 2 to 100,000 / mm 2 The present invention further contemplates, in part, higher density nucleic acid clusters, for example, 100,000 / mm 2 to 1,000,000 / mm 2 and 1,000,000 / mm 2 to 10,000,000 / mm 2 .
[0123] As used herein, a "specimen" is a specimen or a target region within the field of view. When used in connection with a microarray device or other molecular analysis device, the specimen refers to a region occupied by similar or identical molecules. For example, the specimen can be an amplified oligonucleotide, or any other group of polynucleotides or polypeptides having the same or similar sequences. In other embodiments, the specimen can be any element or group of elements that occupy a physical region on a sample. For example, the specimen can be a parcel of land, a body of water, etc. When a specimen is imaged, each specimen has some regions. Thus, in many embodiments, the specimen is not simply one pixel.
[0124] The distance between specimens can be described in any number of ways. In some embodiments, the distance between specimens can be described from the center of one specimen to the center of another specimen. In other embodiments, the distance can be described from the edge of one specimen to the edge of another specimen, or between the outermost distinguishable points of each specimen. The edge of a specimen can be described as a theoretical or actual physical boundary on the chip, or as some points within the boundary of the specimen. In other embodiments, the distance can be described with respect to a fixed point on the sample, or an image of the sample.
[0125] Generally, with respect to the analysis methods, several embodiments are described herein. It will be understood that systems for performing the methods in an automated or semi-automated manner are also provided. Thus, the present disclosure provides a neural network-based template generation and base call system, the system including a processor, a storage device, and a program for image analysis, the program including instructions for performing one or more of the methods described herein. Thus, the methods described herein can be executed, for example, on a computer having components described herein or known in the art.
[0126] The methods and systems described herein are useful for analyzing any of a variety of objects. Particularly useful objects are solid supports or solid-phase surfaces having an attached analyte. The methods and systems described herein provide advantages when used with objects having a repetitive pattern of analytes in the xy plane. One example is a microarray having a collection of cells, viruses, nucleic acids, proteins, antibodies, carbohydrates, small molecules (such as drug candidates), bioactive molecules, or other analytes of interest.
[0127] The number of uses of arrays having analytes with biological molecules such as nucleic acids and polypeptides has increased. Such microarrays typically include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) probes. These are specific for nucleotide sequences present in humans and other organisms. In certain applications, for example, individual DNA or RNA probes can be attached to individual analytes of the array. Test samples, such as those from known humans or organisms, can be exposed to the array such that target nucleic acids (e.g., gene fragments, mRNA, or amplicons) hybridize to complementary probes at each analyte in the sequence. The probes can be labeled by a target-specific process (e.g., due to a label present on the target nucleic acid or due to an enzyme label of the probe or target present in the hybridized form in the analyte). The array can then be examined by scanning specific frequencies of light over the analyte to identify which target nucleic acids are present in the sample.
[0128] Biological microarrays can be used for gene sequencing and similar applications. Generally, gene sequencing involves determining the order of nucleotides of the length of a target nucleic acid, such as a fragment of DNA or RNA. Relatively short sequences are typically sequenced in each specimen, and the resulting sequence information may be used in various bioinformatics methods to reliably determine the sequences of much broader lengths of genetic material from which the fragments are derived. Automated computer-based algorithms for characteristic fragments have been developed and have been more recently used in genome mapping, gene identification, and their functions, etc. Microarrays are particularly useful for characterizing genomic content because there are a large number of variants, which is an alternative to performing many experiments on individual probes and targets, and thus are particularly useful for characterizing genomic content. Microarrays are an ideal format for performing such investigations in a practical manner.
[0129] Any of a variety of sample arrays (also referred to as "microarrays") known in the art can be used in the methods or systems described herein. Typical arrays contain samples each having an individual probe or a population of probes. In the latter case, the population of probes in each sample is typically homogeneous, having a single type of probe. For example, in the case of nucleic acid sequences, each sample can have a plurality of nucleic acid molecules each having a common sequence. However, in some embodiments, the population in each sample of the array can be heterogeneous. Similarly, protein sequences can have samples having a single protein or a population of proteins, typically having the same amino acid sequence, but not necessarily so. Probes can be attached to the surface of the array, for example, by covalently bonding the probe to the surface or through non-covalent interactions between the probe and the surface. In some embodiments, probes such as nucleic acid molecules can be attached to the surface via a gel layer, as described, for example, in U.S. Patent Application No. 13 / 784,368 and U.S. Patent Application Publication No. 2011 / 0059865(A1), each of which is incorporated herein by reference.
[0130] Exemplary arrays include, but are not limited to, BeadChip arrays available from Illumina, Inc (San Diego, Calif.) or others, for example, where the probes are attached to beads present on a surface (e.g., beads within wells on a surface), as described in U.S. Patent No. 6,266,459, U.S. Patent No. 6,355,431, U.S. Patent No. 6,770,441, U.S. Patent No. 6,859,570, or U.S. Patent No. 7,622,294, or International Publication No. 00 / 63437, or others such as those described therein, each of which is incorporated herein by reference. Further examples of commercially available microarrays that can be used include, for example, Affymetrix® GeneChip® microarrays synthesized according to a technique sometimes referred to as VLSIPS™ (Very Large Scale Immobilized Polymer Synthesis) technology or other microarrays. Spotted microarrays can also be used in the methods or systems according to some embodiments of the present disclosure. An exemplary spotted microarray is the CodeLink™ Array available from Amersham Biosciences. Another useful microarray is one manufactured using an inkjet printing method such as the SurePrint™ Technology available from Agilent Technologies.
[0131] Other useful arrays include those used for nucleic acid sequencing applications. For example, arrays having amplicons of genomic fragments (often referred to as clusters) are described in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, WO 91 / 06678, WO 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent 7,405,281 or U.S. Patent 7,057,026, or U.S. Patent Application Publication No. 2008 / 0108082 (A1), each of which is incorporated herein by reference. Another type of array useful for nucleic acid sequencing is the sequence of particles generated from emulsion PCR technology. Examples are described in Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), WO 05 / 010145, U.S. Patent Application Publication No. 2005 / 0130173 or U.S. Patent Application Publication No. 2005 / 0064460, each of which is incorporated herein in its entirety by reference.
[0132] The sequences used for nucleic acid arrays often have a random spatial pattern of nucleic acid samples. For example, the HiSeq or MiSeq sequencing platforms available from Illumina Inc (San Diego, Calif.) utilize flow cells in which the nucleic acid sequences are formed by random seeding followed by bridge amplification. However, patterned sequences can also be used for nucleic acid arrays or other analytical applications. Examples of patterned arrays, methods for their manufacture, and methods for their use are described in U.S. Patent Application No. 13 / 787,396, U.S. Patent No. 13 / 783,043, U.S. Patent No. 13 / 784,368, U.S. Patent Application Publication No. 2013 / 0116153 (A1), and U.S. Patent Application Publication No. 2012 / 0316086 (A1), each of which is incorporated herein by reference. Samples with such patterned sequences can be used to capture a single nucleic acid template molecule and, for example, through bridge amplification, subsequent formation of homogeneous colonies can be performed. Such patterned sequences are particularly useful for nucleic acid sequencing applications.
[0133] The size of the sample on the array (or other object used in the methods or systems herein) can be selected to be suitable for a particular application. For example, in some embodiments, the sample on the array can have a size that accommodates only a single nucleic acid molecule. A surface having a plurality of samples in this size range is useful for constructing the sequences of molecules for detection at single molecule resolution. Samples in this size range are also useful for use in arrays having samples each containing a colony of nucleic acid molecules. Thus, each sample on the array can be about 1 mm 2 or less, about 500 μm 2 or less, about 100 μm 2 or less, about 10 μm 2 or less, about 1 μm 2 or less, about 500 nm 2 or less, or about 100 nm 2 or less, about 10 nm 2 or less, about 5 nm 2 or less, or about 1 nm 2It can have the following area. Alternatively or additionally, the specimens of the array can be about 1 mm 2 or more, about 500 μm 2 or more, about 100 μm 2 or more, about 10 μm 2 or more, about 1 μm 2 or more, about 500 nm 2 or more, about 100 nm 2 or more, about 10 nm 2 or more, about 5 nm 2 or more, or about 1 nm 2 or more. In fact, the specimens can have a size within the range between the upper and lower limits selected from those exemplified above. Although several size ranges of surface specimens have been exemplified with respect to nucleic acids and the scale of nucleic acids, it will be understood that specimens in these size ranges can be used for applications that do not contain nucleic acids. It will be further understood that the size of the specimens does not necessarily have to be limited to the scale used for nucleic acid applications.
[0134] In embodiments that include an object having a plurality of specimens such as an array of specimens, the specimens can be separate and separated from each other in the space therebetween. An array useful in the present invention can have specimens separated by a distance from edge to edge of up to 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm or less. Alternatively or additionally, the array can have specimens separated by a distance from edge to edge of at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm or more. These ranges can apply to the average edge-to-edge spacing and edge spacing of the specimens, as well as the minimum or maximum spacing.
[0135] In some embodiments, the specimens of the array need not be separate, and instead, adjacent specimens can abut against each other. Whether the specimens are separate or not, the size of the specimens and / or the pitch of the specimens can vary so that the array can have a desired density. For example, the average specimen pitch in a regular pattern can be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm or less. Alternatively or additionally, the average specimen pitch in a regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm or more. These ranges can also apply to the maximum pitch or minimum pitch of the regular pattern. For example, the maximum specimen pitch of the regular pattern can be 100 μm or less, 50 μm or less, 10 μm or less, 5 μm or less, 1 μm or less, 0.5 μm or less, and / or the minimum specimen pitch in the regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm or more.
[0136] The density of the specimens in the array can also be understood in terms of the number of specimens present per unit area. For example, the average density of the specimens for the array can be at least about 1×10 3 specimens / mm 2 , 1×10 4 specimens / mm 2 , 1×10 5 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 7 specimens / mm 2 , 1×10 8 specimens / mm 2 , or 1×10 9 specimens / mm 2 or more. Alternatively or additionally, the average density of the specimens for the array can be at most about 1×10 9 specimens / mm 2 , 1×10 8 specimens / mm 2 , 1×10 7 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 5Specimen / mm 2 、 1×10 4 Specimen / mm 2 、 or 1×10 3 Specimen / mm 2 or less.
[0137] The above range can be applied, for example, to all or part of a regular pattern including all or part of an array of specimens.
[0138] Specimens within the pattern can have any of various shapes. For example, when observed in a two-dimensional plane such as on the surface of an array, the specimens may appear rounded, circular, elliptical, rectangular, square, symmetric, asymmetric, triangular, polygonal, etc. The specimens can be arranged, for example, in a regular repeating pattern including a hexagonal or linear pattern. The pattern can be selected to achieve a desired level of packing. For example, circular specimens are optimally filled in a hexagonal arrangement. Of course, other packing configurations can also be used for circular specimens, and vice versa.
[0139] The pattern can be characterized in terms of the number of specimens present within a subset that forms the minimum geometric unit of the pattern. The subset can include, for example, at least about 2, 3, 4, 5, 6, 10 or more specimens. Depending on the size and density of the specimens, the geometric unit can occupy an area of 1 mm 2 、 500 μm 2 、 100 μm 2 、 50 μm 2 、 10 μm 2 、 1 μm 2 、 500 nm 2 、 100 nm 2 、 50 nm 2 、 10 nm 2 or less. Alternatively or additionally, the geometric unit can occupy an area of 10 nm 2 、 50 nm 2 、 100 nm 2 、 500 nm 2 、 1 μm 2 、 10 μm 2 、 50 μm 2, 100 μm 2 , 500 μm 2 , 1 mm 2 can occupy the area above. The characteristics of the specimen in geometric units such as shape, size, pitch, etc. can be selected from those more generally described herein for the specimens of the array or pattern.
[0140] An array having a regular pattern of specimens is ordered with respect to the relative locations of the specimens, but may be random with respect to one or more other characteristics of each specimen. For example, in the case of nucleic acid sequences, the nucleic acid specimens are regular with respect to their relative positions, but may be random with respect to knowledge of the sequences of the nucleic acid species present in any particular specimen. As a more specific example, a repetitive pattern of specimens having a template nucleic acid is seeded, and the template is amplified in each specimen to form a nucleic acid sequence formed by forming a copy of the template in the specimen (e.g., via cluster amplification or bridge amplification, having a regular pattern of nucleic acid specimens, but being random with respect to the distribution of the sequences of the nucleic acids across the sequences. Thus, detection of the presence of nucleic acid material on the array can result in a repetitive pattern of specimens, whereas sequence-specific detection can result in a non-repetitive distribution of signals across the array.
[0141] It will be understood that the descriptions of pattern, order, randomness, etc. herein relate not only to specimens on objects such as specimens on an array, but also to specimens in an image. Thus, pattern, order, randomness, etc. can be present in any of a variety of formats used to store, manipulate, or communicate image data, including but not limited to computer-readable media or computer components such as a graphical user interface or other output device.
[0142] As used herein, the term "image" is intended to mean a representation of all or a portion of an object. The representation can be an optically detected reproduction. For example, an image can be obtained from fluorescence, luminescence, scattering, or absorption signals. The portion of the object present within the image can be the surface of the object or another xy plane. Typically, an image is a two-dimensional representation, although in some cases, the information within the image can be derived from three or more dimensions. An image need not contain an optically detected signal. Non-optical signals can be present instead. An image can be provided in a computer-readable format or medium, such as one or more of those described elsewhere in this specification.
[0143] As used herein, "image" refers to a reproduction or representation of at least a portion of a sample or other object. In some embodiments, the reproduction is an optical reproduction, for example, generated by a camera or other optical detector. The reproduction can be a non-optical reproduction, such as a representation of an electrical signal obtained from an array of nanopore specimens, or a representation of an electrical signal obtained from an ion-sensitive CMOS detector. In certain embodiments, non-optical reproducibility can be excluded from the methods or apparatuses described herein. An image can have a resolution capable of distinguishing specimens present at any of a variety of intervals, for example, separated by 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or less than 0.5 μm.
[0144] As used herein, "acquire", "acquiring", and similar terms refer to any part of the process of obtaining an image file. In some embodiments, data acquisition can include generating an image of a specimen, searching for signals within the specimen, instructing a detection device to search for or generate an image of the signal, providing instructions for further analysis or conversion of the image file, and instructions for any number of conversions or operations of the image file.
[0145] As used herein, the term "template" refers to a representation of a location or relationship between signals or specimens. Thus, in some embodiments, a template is a physical grid having a representation of a signal corresponding to a specimen in a specimen. In some embodiments, a template can be a chart, table, text file, or other computer file indicating a location corresponding to a specimen. In the embodiments presented herein, a template is generated to track the location of a specimen across a set of images of a sample captured at different reference points. For example, a template can be a set of x, y coordinates, or a set of values describing the orientation and / or distance of one specimen relative to another specimen.
[0146] As used herein, the term "specimen" can refer to an object or region of an object from which an image is captured. For example, in an embodiment where an image is taken from the surface of soil, a parcel of land can be a specimen. In other embodiments where the analysis of biomolecules is performed within a flow cell, the flow cell may be divided into any number of subdivisions, each of which may be a specimen. For example, a flow cell may be divided into various flow channels or lanes, and each lane may be further divided into 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 140, 160, 180, 200, 400, 600, 800, 1000 or more distinct regions to be imaged. An example of a flow cell has eight lanes, and each lane is divided into 120 specimens or tiles. In another embodiment, a sample may be made up of multiple tiles, or even the entire flow cell. Thus, an image of each specimen can represent a region of a larger surface being imaged.
[0147] It will be understood that references to ranges and lists of consecutive numbers herein include not only the recited numbers but also all real numbers between the recited numbers.
[0148] As used herein, "reference point" refers to any temporal or physical distinction between images. In another preferred embodiment, the reference point is a point in time. In a more preferred embodiment, the reference point is a point in time or cycle during a sequencing reaction. However, the term "reference point" can include other ways of distinguishing or separating images, such as angles, rotations, times, or other aspects that can distinguish or separate images.
[0149] As used herein, "subset of images" refers to a group of images within a set. For example, the subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60, or any number of images selected from a set of images. In certain other embodiments, the subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or less, or any number of images selected from a set of images. In another preferred embodiment, the images are obtained from one or more sequencing cycles having four images correlated with each cycle. Thus, for example, the subset can be a group of 16 images obtained over 4 cycles.
[0150] Base refers to a nucleotide base or nucleotide, (adenine), C (cytosine), T (thymine), or G (guanine). This application uses "base" and "nucleotide" interchangeably.
[0151] The term "chromosome" refers to a genetic carrier of the present invention in living cells, derived from chromatin strands containing DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is used herein.
[0152] The term "locus" refers to a unique position on a reference genome (e.g., chromosome ID, chromosomal position, and orientation). In some embodiments, the locus may be the position of a residue, sequence tag, or segment on a sequence. The term "locus" may be used to refer to a specific position of a nucleic acid sequence or polymorphism on a reference chromosome.
[0153] As used herein, the term "sample" typically refers to a biological fluid, cell, tissue, organ, or sample derived from an organism containing nucleic acids to be sequenced and / or phased, or a sample derived from a mixture of nucleic acids containing at least one nucleic acid sequence to be sequenced and / or phased. Such samples include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, fine needle biopsy samples (e.g., surgical biopsy, needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, tissue explants, organ cultures, and any other tissue or cell preparation thereof, or fractions or derivatives thereof. Samples are often taken from human subjects (e.g., patients), but samples can be taken from any organism with chromosomes, including but not limited to dogs, cats, horses, goats, sheep, cows, pigs, etc. Samples can be used directly as obtained from a biological source or after pretreatment to modify the properties of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, etc.
[0154] The term "array" includes or represents a chain of nucleotides joined to each other. Nucleotides can be based on DNA or RNA. It should be understood that one array may contain multiple sub-arrays. For example, a single array (e.g., a PCR amplicon) may have 350 nucleotides. A sample read may contain multiple sub-arrays within these 350 nucleotides. For example, a sample read may contain first and second flanking sub-arrays having, for example, 20 to 50 nucleotides. The first and second adjacent sub-arrays may be located on both sides of a repetitive segment having corresponding sub-arrays (e.g., 40 to 100 nucleotides). Each of the adjacent sub-arrays may contain (or may include a portion of) a primer sub-array (e.g., 10 to 30 nucleotides). For ease of reading, the term "sub-array" is referred to as an "array", but it is understood that two arrays need not be distinct from each other on a common strand. To distinguish the various arrays described herein, arrays may be given different labels (e.g., target array, primer array, adjacent array, reference array, etc.). Other terms such as "allele" may be given different labels to distinguish similar objects. Applications use "read" and "array read" interchangeably.
[0155] The term "paired end sequencing" refers to a sequencing method that sequences both ends of a target fragment. Paired end sequencing can facilitate genome reconstruction and the detection of repetitive segments, as well as the detection of gene fusions and novel transcripts. Methods of paired end sequencing are described in International Publication No. WO 07 / 010252, International Application No. GB 2007 / 003798, and U.S. Patent Application Publication No. 2009 / 0088327, each of which is incorporated herein by reference. In one embodiment, a series of operations may be performed as follows. (a) Generate clusters of nucleic acids, (b) linearize the nucleic acids, (c) hybridize the first sequencing primer as described above and repeatedly perform cycles of extension, scanning, and deblocking, (d) "invert" the target nucleic acid on the flow cell surface by synthesizing a complementary copy, (e) linearize the resynthesized strand, and (f) hybridize the second sequencing primer as described above and repeatedly perform cycles of extension, scanning, and deblocking. The inversion operation can deliver the reagents described above for a single cycle of bridge amplification.
[0156] The term "reference genome" or "reference sequence" refers to any particular known genomic sequence, either partial or complete, of an organism that can be used to refer to an identified sequence from a subject. For example, reference genomes used for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. The term "genome" means the complete genetic information of an organism or virus, expressed as a nucleic acid sequence. The genome includes both the genes and non-coding sequences of DNA. The reference sequence may be larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105 times larger, or at least about 106 times larger, or at least about 107 times larger. In one embodiment, the reference genome sequence is that of the full-length human genome. In another example, the reference genome sequence is limited to a particular human chromosome, such as chromosome 13. In some embodiments, the reference chromosome is a chromosomal sequence from human genome version hg19. Such a sequence may sometimes be referred to as a chromosomal reference sequence, but the term reference genome is intended to encompass such sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, partial chromosomal regions (such as strands), etc. of any species. In various embodiments, the reference genome is a consensus sequence or other combination derived from multiple individuals. However, for certain applications, the reference sequence may be taken from a particular individual. In other embodiments, the term "genome" also encompasses so-called "graph genomes" that use particular storage formats and representations of genomic sequences. In one embodiment, the graph genome stores data in a linear file. In another embodiment, the graph genome refers to a representation in which alternative sequencing (e.g., different copies of a chromosome with small differences) is stored as different paths within the graph.Additional information regarding the implementation of graph genomes can be found at https: / / www.biorxiv.org / content / biorxiv / early / 2018 / 03 / 20 / 194530.full.pdf, the contents of which are hereby incorporated by reference in their entirety.
[0157] The term "read" refers to a set of array data that describes a nucleotide sample or a fragment of a reference. The term "read" can refer to a sample read and / or a reference read. Typically, but not necessarily, a read represents a short sequence of contiguous base pairs in a sample or reference. A read may be symbolically represented by the base pair sequence (ATCG) of a sample or reference fragment. To determine whether a read matches a reference sequence or meets other criteria, it may be stored in a memory device and appropriately processed. A read may be obtained directly from a sequencing device or indirectly from stored sequence information regarding a sample. In some cases, for example, it can be a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region that can be mapped and specifically assigned to a chromosome or genomic region or gene.
[0158] Next-generation sequencing methods include, for example, synthesis technology (Illumina), pyrosequencing (454), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), and sequencing by ligation (SOLiD sequencing). Depending on the sequencing method, the length of each read can vary from about 30 bp to over 10,000 bp. For example, DNA sequencing using a SOLiD sequencer generates nucleic acid reads of about 50 bp. In another example, Ion Torrent Sequencing generates nucleic acid reads up to 400 bp, and 454 pyrosequencing generates nucleic acid reads of about 700 bp. In yet another example, single-molecule real-time sequencing can generate reads of 10,000 bp to 15,000 bp. Thus, in certain embodiments, nucleic acid sequence reads have lengths of 30 - 100 bp, 50 - 200 bp, or 50 - 400 bp.
[0159] The terms "sample read", "sample sequence", or "sample fragment" refer to sequence data regarding the genomic sequence of interest from a sample. For example, a sample read can include sequence data from a PCR amplicon having forward and reverse primer sequences. The sequence data can be obtained from any selected sequencing methodology. A sample read can be, for example, a sequencing-by-synthesis (SBS) reaction, a sequencing-ligation reaction, or any other suitable sequencing methodology where it is desirable to determine the length and / or identity of repetitive elements. A sample read can be a consensus (e.g., average or weighted) sequence derived from multiple sample reads. In certain embodiments, providing a reference sequence includes identifying the locus of interest based on the primer sequences of the PCR amplicon.
[0160] The term "raw fragment" refers to sequence data of a portion of a target genomic sequence that at least partially overlaps a designated or secondary position of interest within a sample read or sample fragment. Non-limiting examples of raw fragments include double-stiched fragments, simple stiched fragments, and simple non-stitched fragments. The term "raw" is used to indicate that a raw fragment includes sequence data having some relationship to the sequence data in a sample read, and is used regardless of whether the raw fragment corresponds to and indicates a supporting variant that authenticates or confirms a potential variant in a sample read. The term "raw fragment" does not necessarily indicate that the fragment includes a supporting variant that validates a variant call in a sample read. For example, when a sample read is determined by a variant calling application to exhibit a first variant, this variant calling application can determine that one or more raw fragments lack the corresponding type of "supporting" variant that would otherwise be expected to occur considering the variant in the sample read.
[0161] The terms "mapping", "aligned", "aligning", or "alignment" refer to the process of reading or comparing a tag to a reference array, thereby determining whether the reference array contains the read array. When the reference array is read, the read may be mapped to the reference array, or in certain other embodiments, may be mapped to a specific position within the reference array. In some cases, the alignment simply conveys whether a read is a member of a particular reference array (i.e., whether the read is present or absent in the reference array). For example, the alignment of a read to a reference array for human chromosome 13 conveys whether the read is present in the reference array of chromosome 13. A tool that provides this information may be referred to as a set membership tester. In some cases, the alignment further indicates the position within the reference array where the read or tag map is located. For example, if the reference array is the entire human genome sequence, the alignment may indicate that the read is present on chromosome 13, and may further indicate that the read is at a particular strand and / or site of chromosome 13.
[0162] The term "indel" refers to an insertion and / or deletion of bases in the DNA of an organism. A microindel represents an indel that results in a net change of 1 to 50 nucleotides. Unless the length of the indel is a multiple of 3, a frameshift mutation occurs when encoding a region of the genome. Indels can be contrasted with point mutations. Indel insertions delete nucleotides from the sequence, while point mutations are in the form of substitutions that replace one of the nucleotides without changing the overall number in the DNA. Indels can also be contrasted with Tandem Base Mutations (TBMs), which can be defined as substitutions in adjacent nucleotides (primarily substitutions in two adjacent nucleotides, although substitutions in three adjacent nucleotides have been observed).
[0163] The term "variant" refers to a nucleic acid sequence that is different from a nucleic acid reference. Typical nucleic acid sequence variants include, but are not limited to, single nucleotide polymorphisms (SNPs), short deletion and insertion polymorphisms (Indels), copy number variations (CNVs), microsatellite markers, or short tandem repeats and structural variations. Somatic variant calling is an effort to identify variants that are present at low frequency in a DNA sample. The calling of somatic variants is of interest in the context of cancer treatment. Cancer is caused by the accumulation of mutations in DNA. Tumor-derived DNA samples are generally heterogeneous and contain some normal cells, early stages of cancer progression (with fewer mutations), and some later-stage cells (with more mutations). Due to this heterogeneity, when sequencing a tumor (e.g., from an FFPE sample), somatic mutations often appear at low frequency. For example, an SNV may be seen in only 10% of the reads covering a given base. Variants classified as somatic or germline by a variant classifier are also referred to herein as "variants under test".
[0164] The term "noise" refers to an erroneous variant call resulting from one or more errors in the sequencing process and / or variant calling application.
[0165] The term "variant frequency" refers to the relative frequency of alleles (variants of a gene) at a specific locus within a population, and is expressed as a fraction or percentage. For example, the fraction or percentage may be the proportion of all chromosomes in the population that carry that allele. As an example, a sample variant frequency represents the relative frequency of an allele / variant at a specific locus / position along a target genomic sequence across a "population" corresponding to the number of reads and / or samples obtained for the target genomic sequence from an individual. As another example, a baseline variant frequency represents the relative frequency of an allele / variant at a specific locus / position along one or more baseline genomic sequences, where the relative frequency of an allele / variant at a specific locus / position along one or more baseline genomic sequences is obtained for one or more baseline genomic sequences.
[0166] The term "Variant Allele Frequency (VAF)" refers to the ratio of the sequenced reads divided by the overall coverage at the target position for the variant. VAF is a measure of the proportion of sequenced reads that carry the variant.
[0167] The terms "position", "designated position", and "locus" refer to the position or coordinates of one or more nucleotides within a nucleotide sequence. The terms "position", "designated position", and "locus" also refer to the position or coordinates of one or more base pairs in a nucleotide sequence.
[0168] The term "haplotype" refers to a combination of alleles at adjacent sites on a chromosome that are inherited together. A haplotype, if present, may be for one locus, several loci, or the entire chromosome, depending on the number of recombination events that have occurred between a given set of loci.
[0169] As used herein, the term "threshold value" refers to a numerical value or values used as a cutoff for characterizing a sample, nucleic acid, or portion thereof (e.g., a reading). The threshold value may vary based on empirical analysis. The threshold value can be compared to a measured or calculated value to determine whether the source giving rise to such a value should be classified in a particular manner. The threshold value can be identified empirically or analytically. The selection of the threshold value depends on the confidence level that the user desires to have in making a classification. The threshold value may be selected for a particular purpose (e.g., for a balance of sensitivity and selectivity). As used herein, the term "threshold value" indicates a point at which the course of an analysis can change and / or an action can be triggered. The threshold value need not be a fixed number. Instead, the threshold value may be, for example, a function based on multiple factors. The threshold value can be adapted to the situation. Further, the threshold value may indicate a range between an upper limit, a lower limit, or bounds.
[0170] In some embodiments, an indicator or score based on sequencing data can be compared to a threshold. As used herein, the term "metric" or "score" may include a value or result determined from sequencing data, or may include a function based on a value or result determined from sequencing data. Similar to a threshold, an indicator or score can be adapted to the situation. For example, an indicator or score may be a normalized value. As an example of a score or metric, in one or more embodiments, a count score can be used when analyzing data. The count score may be based on the number of sample reads. The sample reads may have gone through one or more filtering steps such that the sample reads have at least one common characteristic or quality. For example, each of the sample reads used to determine the count score may be aligned to a reference sequence or may be assigned as a potential allele. The number of sample reads having a common characteristic can be counted to determine a read count. The count score may be based on the read count. In some embodiments, the count score may be a value equal to the read count. In other examples, the count score may be based on the read count and other information. For example, the count score may be based on the read count of a particular allele at a locus and the total number of reads at the locus. In some embodiments, the count score may be based on the read count at a locus and previously obtained data. In some embodiments, the count score may be a normalized score between predetermined values. The count score may also be a function of the read counts from other loci of the sample, or a function of the read counts from other samples run concurrently with the sample of interest. For example, the count score may be a function of the read count of a particular allele and the read counts of other loci in the sample, and / or the read counts from other samples. As an example, the read counts from other loci and / or the read counts from other samples can be used to normalize the count score for a particular allele.
[0171] The term "coverage" or "fragment coverage" refers to the counting or other measure of multiple sample reads for the same fragment of an array. The read count may represent the count of the number of reads that cover the corresponding fragment. Alternatively, the coverage may be determined by multiplying the read count by a specified factor based on historical knowledge, knowledge of the sample, knowledge of the locus, etc.
[0172] The term "read depth" (formerly the number following "×") refers to the number of sequenced reads with overlapping alignments at the target position. This is often expressed as an average or percentage exceeding a cutoff for a set of intervals (such as exons, genes, or panels). For example, a clinical report could state that the panel average coverage is 1,105× with a targeted base coverage > 100× having 98%.
[0173] The term "base call quality score" or "Q-score" refers to a PHRED-scale probability in the range from 0 - 50 that is inversely proportional to the probability that a single sequenced base is correct. For example, a T base call with a Q of 20 is considered to be correct with a 99.99% probability. Any base call with Q < 20 should be considered of low quality, and any variant identified when a significant proportion of the sequenced reads supporting the variant are of low quality should be considered potentially false positive.
[0174] The term "variant read" or "variant read number" refers to the number of sequenced reads that support the presence of a variant.
[0175] Regarding "strandedness" (or DNA strandedness), the genetic message in DNA can be represented as the letters A, G, C, and T, for example, 5'-AGGACA-3'. In many cases, the sequence is written in the direction shown in this specification, that is, the 5' end is written on the left and the 3' end is written on the right. DNA can occur as a single-stranded molecule (such as in certain viruses), but usually, DNA is found as a double-stranded unit. This has a double-helical structure with two anti-parallel strands. In this case, the term "anti-parallel" means that the two strands operate in parallel but have opposite polarities. Double-stranded DNA is held together by base pairing, and the pairing is always maintained such that adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). This pairing is called complementarity, and one DNA strand is said to be the complement of the other. Thus, double-stranded DNA can be represented as two strings, such as 5'-AGGACA-3' and 3'-TCCTGT-5'. Note that the two strands have opposite polarities. Thus, the strandedness of the two DNA strands can be referred to as the reference strand and its complement, the forward and reverse strands, the top and bottom strands, the sense and antisense strands, or the Watson and Crick strands.
[0176] Read alignment (also called read mapping) is the process of referring to where a sequence in the genome is derived from. When the alignment is performed, the "mapping quality" or "mapping quality score (MAPQ)" of a given read quantifies the probability that its position on the genome is correct. The mapping quality is encoded on a phased scale, where P is the probability that the alignment is incorrect. The probability is P = 10 (-MAQ / 10)Calculated as such, where MAPQ in the formula is the mapping quality. For example, a mapping quality of 40 = 10 for a power of -4 means that there is a 0.01% chance that the read is misaligned. Therefore, mapping quality is associated with several alignment factors such as the basic quality of the read, the complexity of the reference genome, and paired-end information. First, if the basic quality of the read is low, the observed sequence may be incorrect, thus meaning that its alignment is incorrect. Second, mapping ability refers to the complexity of the genome. Repetitive regions are more difficult to map the maps and reads contained in these regions, and usually the mapping quality is low. In this context, MAPQ reflects the fact that the reads are not uniquely aligned and their actual origin cannot be determined. Third, in the case of paired-end sequencing data, concordant pairs are more likely to be better aligned. The higher the mapping quality, the better the alignment. Reads aligned with good mapping quality usually mean that the read sequence is good and is aligned with few mismatches within a highly mappable region. The MAPQ value can be used as quality control for the alignment results. The proportion of reads aligned with a MAPQ higher than 20 is usually for downstream analysis.
[0177] As used herein, "signal" refers to a detectable event such as, for example, luminescence within an image, preferably luminescence. Thus, in another preferred embodiment, the signal can represent any detectable luminescence (i.e., "spot") captured within the image. Thus, as used herein, "signal" can refer to both the actual emission from the analyte of the sample and can refer to spurious luminescence that does not correlate with the actual sample. Thus, the signal can arise from noise and can be discarded later so as not to represent the actual analyte of the specimen.
[0178] As used herein, the term "clamp" refers to a group of signals. In certain embodiments, the signals are from different specimens. In another preferred embodiment, a signal clamp is a group of signals that cluster together. In a more preferred embodiment, a signal clamp represents a physical region covered by one amplification oligonucleotide. Each signal clamp should ideally be observed as several signals (one per template cycle, perhaps more due to crosstalk). Thus, duplicate signals are detected where two (or more) signals are included in the template from the same signal clamp.
[0179] As used herein, terms such as "minimum", "maximum", "minimize", "maximize", and their grammatical variants can include values that are not absolute maximum or minimum values. In some embodiments, the values include near the maximum and minimum values. In other examples, the values can include local maximum and / or local minimum values. In some embodiments, the values include only the absolute maximum or minimum value.
[0180] As used herein, "crosstalk" refers to the detection of signals within one image that are also detected in a separate image. In another preferred embodiment, crosstalk can occur when a radiated signal is detected in two separate detection channels. For example, if a radiated signal occurs in one color, the emission spectrum of that signal may overlap with another radiated signal in a different color. In a preferred embodiment, the fluorescent molecules used to indicate the presence of nucleotide bases A, C, G, and T are detected in separate channels. However, since the emission spectra of A and C overlap, a portion of the C-color signal may be detected during detection using color channels. Thus, crosstalk between the A and C signals enables signals from one color image to appear in other color images. In some embodiments, there is G and T crosstalk. In some embodiments, the amount of crosstalk between channels is asymmetric. It will be understood that the amount of crosstalk between channels can be controlled, inter alia, by the selection of signal molecules having appropriate emission spectra, as well as by the selection of the size and wavelength range of the detection channels.
[0181] As used herein, the terms "register", "registering", "registration", and similar terms refer to any process for correlating signals within an image or dataset with signals within an image or dataset from another point in time or perspective. For example, registration can be used to align signals from a set of images to form a template. In another example, registration can be used to align signals from other images to a template. One signal may be registered directly or indirectly to another signal. For example, a signal from image "S" may be registered directly to image "G". As another example, a signal from image "N" may be registered directly to image "G", or alternatively, a signal from image "N" may be registered to image "S" which has previously been registered to image "G". Thus, a signal from image "N" is registered indirectly to image "G".
[0182] As used herein, the term "reference" is intended to mean a distinguishable reference point within or on an object. The reference point can be, for example, a mark, a second object, a shape, an edge, a region, an irregularity, a channel, a pit, a post, etc. The reference point can be present within an image of the object or within another dataset resulting from detecting the object. The reference point can be specified by x and / or y coordinates within the plane of the object. Alternatively or additionally, the reference point can be specified by a z coordinate orthogonal to the xy plane, defined, for example, by the relative position of the object and the detector. One or more coordinates relative to the reference point can be specified relative to one or more other specimens of the object or to an image or other dataset derived from the object.
[0183] As used herein, the term "optical signal" is intended to include, for example, fluorescence, luminescence, scattering, or absorption signals. The optical signal can be detected in the ultraviolet (UV) range (about 200 - 390 nm), the visible (VIS) range (about 391 - 770 nm), the infrared (IR) range (about 0.771 - 25 micrometers), or other ranges of the electromagnetic spectrum. The optical signal can be detected in a manner that excludes all or part of one or more of these ranges.
[0184] As used herein, the term "signal level" is intended to mean the amount or quantity of detected energy or encoded information having a desired or predetermined characteristic. For example, an optical signal can be quantified by one or more of intensity, wavelength, energy, frequency, power, luminance, etc. Other signals can be quantified according to characteristics such as voltage, current, electric field strength, magnetic field strength, frequency, power, temperature, etc. The absence of a signal is understood to be a signal level of zero or a signal level that is not significantly distinguishable from noise.
[0185] As used herein, the term "simulate" is intended to mean creating a representation or model of an object or action that predicts the characteristics of the object or action. The representation or model can often be distinguishable from the object or action. For example, the representation or model can be distinguished from one or more characteristics such as the intensity of a signal detected from all or part of a color, a work piece, a size, or a shape. In certain embodiments, the representation or model can be idealized, exaggerated, muted, or incomplete as compared to the object or action. Thus, in some embodiments, the representation of the model can be such that it represents, for example, with respect to at least one of the above characteristics. The representation or model can be provided in a computer-readable format or medium such as one or more of those described elsewhere herein.
[0186] As used herein, the term "specific signal" is intended to mean a detected energy or encoded information that is selectively observed over other energy or information such as background energy or information. For example, a specific signal can be an optical signal detected at a specific intensity, wavelength, or color; an electrical signal detected at a specific frequency, power, or electric field strength; or other signals known in the art related to spectroscopy and analytical detection.
[0187] As used herein, the term "swath" is intended to mean a rectangular portion of an object. The swath can be an elongated strip that is scanned by relative movement between the object and a detector in a direction parallel to the longest dimension of the strip. Generally, the width of the rectangular portion or strip is constant along its entire length. Multiple swaths of an object can be parallel to each other. Multiple swaths of an object can overlap each other, be adjacent to each other, or be separated from each other by interstitial regions.
[0188] As used herein, the term "dispersion" is intended to mean the expected difference, and the observed difference, or the difference between two or more observations. For example, the dispersion can be the discrepancy between the expected value and the measured value. Statistical functions such as the standard deviation, the square of the standard deviation, the coefficient of variation, etc. can be used to represent the dispersion.
[0189] As used herein, the term "xy coordinate" is intended to mean information that specifies a position, size, shape, and / or orientation within the xy plane. The information can be, for example, numerical coordinates in a Cartesian coordinate system. The coordinates can be provided with respect to one or both of the x-axis and the y-axis, or with respect to another location within the xy plane. For example, the coordinates of a specimen of an object can specify the location of the specimen with respect to a reference of the object or the position of another specimen.
[0190] As used herein, the term "xy plane" is intended to mean a two-dimensional region defined by the straight axes x and y. When used with reference to a detector and an object observed by the detector, the region can be further specified to be orthogonal to the observation direction between the detector and the object being detected.
[0191] As used herein, the term "z coordinate" is intended to mean information that specifies the position of a point, line, or region along an axis orthogonal to the xy plane. In a particular embodiment, the z-axis is orthogonal to the region of the object observed by the detector. For example, the direction of the focus of an optical system can be specified along the z-axis.
[0192] In some embodiments, the acquired signal data is transformed using an affine transformation. In some such embodiments, template generation uses the fact that the affine transformation between color channels is consistent across operations. Due to this consistency, a set of default offsets can be used when determining the coordinates of the specimens in the specimen. For example, the default offset file can include relative transformations (shift, scale, skew) for different channels relative to one channel such as the A channel. However, in other embodiments, offsets between color channel drifts during and / or between operations make offset-driven template generation difficult. In such examples, the methods and systems provided herein can utilize offset template generation, which is further described below.
[0193] In some aspects of the above embodiments, the system can include a flow cell. In some aspects, the flow cell includes lanes, or tiles of other configurations, and at least a portion of the tiles includes one or more groups of specimens. In some aspects, the specimens include a plurality of molecules such as nucleic acids. In certain aspects, the flow cell is configured to deliver labeled nucleotide bases to the sequence of the nucleic acid, thereby extending a primer that hybridizes to the nucleic acid in the specimen to generate a signal corresponding to the specimen containing the nucleic acid. In a preferred embodiment, the nucleic acids in the specimen are identical or substantially identical to each other.
[0194] In some of the image analysis systems described herein, each image in a set of images includes a color signal, and different colors correspond to different nucleotide bases. In some embodiments, each image of the set of images includes a signal having a single color selected from at least four different colors. In some embodiments, each image within the set of images includes a signal having a single color selected from four different colors. In some of the systems described herein, a nucleic acid can be sequenced by providing four different labeled nucleotide bases to the sequence of the molecule so as to generate four different images, each image including a signal having a single color, and the signal colors being different for each of the four different images, thereby generating a cycle of four color images corresponding to the four possible nucleotides present at a particular position within the nucleic acid. In certain embodiments, the system includes a flow cell configured to deliver additional labeled nucleotide bases to the sequence of the molecule, thereby generating multiple cycles of color images.
[0195] In a preferred embodiment, the method provided herein may include determining whether the processor is actively collecting data or whether the processor is in a low activity state. Collecting and storing a large number of high-quality images typically requires a large amount of storage capacity. Further, once collected and stored, the analysis of the image data can be resource-intensive and may interfere with the processing capabilities of other functions such as the collection and storage of additional image data. Thus, as used herein, the term "low activity state" refers to the processing capabilities of the processor at a given time. In some embodiments, the low activity state occurs when the processor is not collecting and / or storing data. In some embodiments, a low activity state occurs when some data collection and / or storage is being performed, but there is sufficient additional processing capacity remaining such that image analysis can occur simultaneously without interfering with other functions.
[0196] As used herein, "identifying a conflict" refers to identifying a situation where multiple processes compete for a resource. In some such embodiments, one process is given priority over another process. In some embodiments, the conflict may relate to the need to assign priorities to the allocation of time, processing power, memory capacity, or any other resource to which priorities are given. Thus, in some embodiments, when processing time or capacity is distributed between two processes, such as whether to analyze a dataset or acquire and / or store the dataset, there is a discrepancy between the two processes that can be resolved by giving priority to one of the processes.
[0197] Also provided herein is a system for performing image analysis. The system can include a processor, a memory capacity, and a program for image analysis. The program can include instructions for processing a first dataset for storage and a second dataset for analysis. The processing can include acquiring and / or storing the first dataset on a storage device and analyzing the second dataset when the processor has not acquired the first dataset. In certain embodiments, the program includes instructions for identifying at least one instance of a conflict between acquiring and / or storing the first dataset and analyzing the second dataset, and acquiring and / or storing the image data is prioritized such that acquiring and / or storing the first dataset is given priority. In certain embodiments, the first dataset includes image files collected from an optical imaging device. In certain embodiments, the system further includes an optical imaging device. In some embodiments, the optical imaging device includes a light source and a detection device.
[0198] As used herein, the term "program" refers to instructions or commands for performing a task or process. The term "program" may be used interchangeably with the term "module". In certain embodiments, a program may be a compilation of various instructions that are executed under the same set of commands. In other embodiments, a program may refer to separate batches or files.
[0199] Described below are some of the surprising effects of utilizing the methods and systems for performing the image analysis described herein. In some examples of sequencing, an important measure of the utility of a sequencing system is its overall efficiency. For example, the amount of mappable data generated per day, as well as the total cost of equipment installation and operation, are important aspects of an economical sequencing solution. To generate mappable data and reduce the time required to enhance the efficiency of the system, real-time basecalling can be enabled on the instrument computer and can operate in parallel with sequencing chemistry and imaging. This allows data processing and analysis to be completed prior to the completion of sequencing chemistry finishing. Further, the storage required for intermediate data can be reduced and the amount of data that needs to be moved across the network can be limited.
[0200] While the sequence output is increasing, the data per operation transferred from the systems provided herein to the network and the secondary analysis processing hardware are substantially decreasing. By converting data on the instrument computer (acquisition computer), the network load is dramatically reduced. Without these on-instrument, off-network data reduction techniques, the image output of a DNA sequencing instrument would cripple most networks.
[0201] The widespread adoption of high-throughput DNA sequencing instruments has been driven in part by ease of use, support for a wide range of applications, and compatibility with virtually any laboratory environment. The highly efficient algorithms presented herein enable adding significant analytical capabilities to a simple workstation that can control the sequencing instrument. This reduction in the computational hardware requirements has several practical advantages that become even more important as the sequencing output level continues to increase. For example, image analysis and basecalling are performed to minimize simple towers, heat generation, laboratory footprint, and power consumption. In contrast, other commercial sequencing technologies have recently ramped up their computing infrastructure with up to five times the processing power for primary analysis, initiating an increase in heat output and power consumption. Thus, in some embodiments, the computational efficiency of the methods and systems provided herein enables increasing their sequencing throughput while minimizing server hardware.
[0202] Thus, in some embodiments, the methods and / or systems presented herein function as a state machine, maintaining track of the individual states of each sample and performing appropriate processing and advancing the sample to its next state when it detects that the sample is ready to proceed to the next state. A more detailed example of a method by which a state machine monitors a file system to determine if a sample is ready to proceed to the next state according to a preferred embodiment is described in Example 1 below.
[0203] In a preferred embodiment, the methods and systems provided herein are multi-threaded and can cooperate with a configurable number of threads. Thus, for example, in the context of nucleic acid sequencing, the methods and systems provided herein can operate in the background during live sequencing operations for real-time analysis or can operate using existing image data sets for offline analysis. In certain preferred embodiments, the methods and systems handle the multi-threading by giving each thread its own subset of the specimens it is involved with. This minimizes the potential for thread holding.
[0204] The method of the present disclosure can include the step of obtaining a target image of an object using a detection device, which image includes a repetitive pattern of an analyte on the object. A detection device capable of high-resolution imaging of the surface is particularly useful. In certain embodiments, the detection device will have a resolution sufficient to distinguish analytes in the density, pitch, and / or analyte size described herein. A detection device capable of obtaining an image or image data from the surface is particularly useful. Exemplary detectors are configured to obtain an area image while maintaining a static relationship between the object and the detector. A scanning device can also be used. For example, a device for obtaining a continuous area image (e.g., referred to as a “step and shot” detector) can be used. Also useful are devices that continuously scan points or lines on the surface of the object and accumulate data to construct an image of the surface. A point-scanning detector can be configured to scan points (i.e., small detection areas) on the surface of the object via a raster movement in the x-y plane of the surface. A line-scanning detector can be configured to scan a line along the y-dimension of the surface of the object, the longest dimension of which line occurs along the x-dimension. It will be understood that scanning detection can be achieved by moving the detection device, the object, or both. Exemplary detection devices particularly useful for nucleic acid sequencing applications are described in U.S. Patent Application Publication Nos. 2012 / 0270305 (A1), 2013 / 0023422 (A1), and 2013 / 0260372 (A1), and U.S. Pat. Nos. 5,528,050, 5,719,391, 8,158,926, and 8,241,573, each of which is incorporated herein by reference.
[0205] Embodiments disclosed herein may be implemented as a manufacturing method, apparatus, system, or article using programming or engineering techniques for generating software, firmware, hardware, or any combination thereof. As used herein, the term "manufactured article" refers to hardware such as an optical storage device or a computer-readable medium, as well as code or logic implemented within a volatile or non-volatile memory device. Such hardware includes, but is not limited to, field programmable gate arrays (FPGAs), coarse grained reconfigurable architectures (CGRAs), application-specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microprocessors, or other similar processing devices. In certain embodiments, the information or algorithms described herein are present in a non-transitory storage medium.
[0206] In certain embodiments, the computer-implemented methods described herein can occur in real time while multiple images of an object are being acquired. Such real-time analysis is particularly useful for nucleic acid sequencing applications where nucleic acid sequences are subjected to repeated cycles of fluidics and detection steps. Analysis of the sequencing data can often be beneficially performed in real time or in the background using the methods described herein, while it can also be beneficial to perform the methods described herein while other data collection or analysis algorithms are in progress. Examples of real-time analysis methods that can be used in the present methods are commercially available from Illumina, Inc (San Diego, Calif), and / or are those used in the MiSeq and HiSeq sequencing instruments described in US Patent Application Publication No. 2012 / 0020537 (A1), which is incorporated herein by reference.
[0207] An exemplary data analysis system having programming formed by one or more programmed computers and having code stored on one or more machine-readable media for performing one or more steps of the methods described herein. In one embodiment, for example, the system includes an interface designed to enable networking of the system to one or more detection systems (e.g., optical imaging systems) configured to acquire data from a target object. The interface can receive and condition data, where appropriate. In certain embodiments, the detection system outputs image data representing, for example, individual image elements or pixels that together form an image of an array or other object. The processor processes the received detection data according to one or more routines defined by the processing code. The processing code may be stored in various types of memory circuits.
[0208] According to the presently contemplated embodiments, the processing code executed on the detection data includes a data analysis routine designed to analyze the detection data to determine the locations of the individual specimens visible or encoded in the data, and the locations where the specimens are not detected (i.e., where there is no specimen or where no significant signal is detected from an existing specimen) and metadata. In certain embodiments, the specimen locations in the array typically appear brighter than the non-specimen locations due to the presence of fluorescent dyes attached to the imaged specimens. It will be understood that a specimen need not appear brighter than its surrounding area if, for example, the target of the probe in the specimen is not detected within the array. The color in which an individual specimen appears can be a function of the dye used and the wavelength of the light used by the imaging system for imaging purposes. Specimens that are not bound to a target or do not have a particular label can be identified according to other characteristics such as their expected location within the microarray.
[0209] When the data analysis routine places the individual specimens in the data, value assignment can be performed. Generally, value assignment assigns a digital value to each specimen based on the characteristics of the data represented by the detector components (e.g., pixels) at the corresponding locations. That is, for example, when imaging data is processed, the value assignment routine may be designed to recognize that light of a particular color or wavelength was detected at a particular location. In a typical DNA imaging application, for example, the four common nucleotides are represented by four distinct distinguishable colors. Each color may then be assigned a value corresponding to that nucleotide.
[0210] As used herein, the terms "module", "system", or "system controller" may include hardware and / or software systems and circuits that operate to perform one or more functions. For example, a module, system, or system controller may include a computer processor, controller, or other log-based device that executes operations based on instructions stored on a tangible and non-transitory computer-readable storage medium such as a computer memory. Alternatively, a module, system, or system controller may include a wired device that executes operations based on wired logic and circuits. A module, system, or system controller shown in the accompanying drawings may represent hardware and circuits that operate based on software or wiring instructions, software that instructs the hardware to operate, or a combination thereof. A module, system, or system controller can include and / or be connected to and / or represent one or more processors, such as a computer microprocessor, and / or include a hardware circuit or circuits.
[0211] As used herein, the terms "software" and "firmware" are interchangeable and include any computer program stored in a memory executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely examples and are not limited to the types of memory that can be used to store computer programs.
[0212] In the field of molecular biology, one of the processes for nucleic acid sequencing in use is sequence number synthesis. This technique can be applied to highly parallel sequencing projects. For example, by using an automated platform, it is possible to perform millions of sequencing reactions simultaneously. Accordingly, one embodiment of the present invention relates to devices and methods for collecting, storing, and analyzing image data generated during nucleic acid sequencing.
[0213] The huge gain in the amount of data that can be collected and stored makes the rationalized image analysis method even more beneficial. For example, the image analysis methods described herein enable both designers and end users to make efficient use of existing computer hardware. Accordingly, methods and systems for reducing the computational amount of processing data in terms of rapidly increasing data output are presented herein. For example, in the field of DNA sequencing, the yield has been expanded 15-fold in recent processes and can reach hundreds of gigabases in a single operation of a DNA sequencing device. When the requirements of the computing infrastructure increase proportionally, large-scale genome-scale experiments are out of reach for most researchers. Thus, the generation of more raw sequence data increases the need for secondary analysis and data storage, making the optimization of data transport and storage highly beneficial. Some embodiments of the methods and systems presented herein can reduce the time, hardware, networking, and laboratory infrastructure requirements necessary to generate useable sequence data.
[0214] This disclosure describes various methods and systems for performing the methods. Some examples of the methods are described as a series of steps. However, it should be understood that the embodiments are not limited to the specific steps and / or order of steps described herein. Steps may be omitted, steps may be modified, and / or other steps may be added. Further, the steps described herein can be combined, steps may be executed simultaneously, steps may be executed simultaneously, steps may be divided into multiple sub-steps, steps may be executed in a different order, or steps (or a series of steps) may be repeatedly re-executed. In addition, although different methods are described herein, it should be understood that in other embodiments, different methods (or steps of different methods) may be combined.
[0215] In some embodiments, a processing unit, processor, module, or computing system “configured” to perform a task or operation can be understood to be specifically structured to perform the task or operation (e.g., having one or more programs or instructions adjusted or intended to perform the task or operation and / or having an arrangement of processing circuitry adjusted or intended to perform the task or operation). For clarity and to avoid doubt, a general-purpose computer (which can be “configured” to perform a task or operation when appropriately programmed) is not configured to be “configured” to perform a task or operation unless specifically programmed or structurally changed to perform the task or operation.
[0216] Furthermore, the operations of the methods described herein can be sufficiently complex such that the operations cannot be performed by an average human or practitioner within a commercially reasonable time period. For example, the method can rely on relatively complex calculations such that such a person cannot complete the method within a commercially reasonable time.
[0217] Throughout this application, various publications, patents, or patent applications are referenced. The entire disclosure of these publications is hereby incorporated by reference into this specification to more fully describe the state of the art to which this invention pertains.
[0218] As used herein, the term “comprising” is intended to be open-ended and to further encompass any additional elements in addition to the recited elements.
[0219] As used herein, the term “each,” when used in reference to a set of items, is intended to identify individual items within the set, but not necessarily to refer to all items within the set. Exceptions can occur where the explicit disclosure or context clearly indicates otherwise.
[0220] The present invention has been described with reference to the above embodiments, but it should be understood that various modifications can be made without departing from the present invention.
[0221] The modules of the present application can be implemented in hardware or software and, as shown in the figures, need not be divided into exactly the same blocks. Some may be implemented on different processors or computers, or may span multiple different processors or computers. Additionally, it will be understood that some of the modules can be operated in parallel with or in a different order than that shown in the figures without affecting the functions achieved. Also, as used herein, the term "module" can include "sub-modules" that can be contemplated herein to make up a module. The blocks in the figures designated as modules can also be considered as flowchart steps in a method.
[0222] As used herein, "identifying" an information item does not necessarily require a direct specification of that information item. The information can be "identified" within a field simply by referring to the actual information through one or more layers in one direction or by identifying one or more items of different information that are sufficient to determine the actual item of information. Additionally, the term "specify" is meant to be the same as "identify" herein.
[0223] As used herein, a given signal, event, or value is "dependent on" a pre-decesser signal, event, or value of the pre-decesser signal, event, or value, a given signal, event, or an event or value affected by the value. If there are intervening processing elements, steps, or periods, a given signal, event, or value can "exist" depending on the "pre-decesser signal, event, or value". When an intervening processing element or step combines two or more signals, events, or values, the signal output of the processing element or step is considered to be "dependent on" each of the signal, event, or value inputs. If a given signal, event, or value is the same as the pre-decesser signal, event, or value, this simply means that the given signal, event, or value is "dependent on" the "pre-decesser signal, event, or value" and is "dependent on" or "dependent on" or "based on the base-decesser signal, event, or value" and is "dependent on" or "depends on". The "responsiveness" of a given signal, event, or value to another signal, event, or value is defined similarly.
[0224] As used herein, "in parallel" or "in parallel" does not require exact simultaneity. It is sufficient if one evaluation of an individual begins before another evaluation of the individual is completed.
[0225] Computer system Figure 17 is a computer system 1700 that can be used to implement the disclosed technology. The computer system 1700 includes at least one central processing unit (CPU) 1772 that communicates with a number of peripheral devices via a bus subsystem 1755. These peripheral devices can include, for example, a memory subsystem 1710 that includes a memory device and a file storage subsystem 1736, a user interface input device 1738, a user interface output device 1776, and a network interface subsystem 1774. The input device and the output device enable user interaction with the computer system 1700. The network interface subsystem 1774 provides an interface to an external network that includes an interface to a corresponding interface device within another computer system.
[0226] In one embodiment, the equalizer-based cooler 104 is communicatively linked to the memory subsystem 1710 and the user interface input device 1738.
[0227] The user interface input device 1738 may include a keyboard, a mouse, a trackball, a touchpad, or a pointing device such as a graphics tablet, a scanner, a touch screen incorporated into a display, an audio input device such as a voice recognition system and a microphone, and other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 1700.
[0228] The user interface output device 1776 can include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem can include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visible image. The display subsystem can also provide a non-visual display such as an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and means for outputting information from the computer system 1700 to a user or another machine or computer system.
[0229] The memory subsystem 1710 stores programming and data constructs that provide some or all of the functionality of some of the modules and methods described herein. These software modules are generally executed by the processor 1778.
[0230] Processor 1778 can be a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or a coarse grained reconfigurable architecture (CGRAs). Processor 1778 can be hosted by deep learning cloud platforms such as Google Cloud Platform (trademark), Xilinx (trademark), and Cirrascale (trademark). Examples of Processor 1778 include Google's Tensor Processing Unit (TPU) (trademark), rackmount solutions such as GX4 Rackmount Series (trademark), GX17 Rackmount Series (trademark), NVIDIA DGX-1 (trademark), Microsoft' Stratix V FPGA (trademark), Graphcore's Intelligent Processor Unit (IPU) (trademark), Qualcomm's Zeroth Platform (trademark) with Snapdragon processors (trademark), NVIDIA's Volta (trademark), NVIDIA's DRIVE PX (trademark), NVIDIA's JETSON TX1 / TX2 MODULE (trademark), Intel's Nirvana (trademark), Movidius VPU (trademark), Fujitsu DPI (trademark), ARM's DynamicIQ (trademark), IBM TrueNorth (trademark), Lambda GPU Server with Testa V100s (trademark), and others.
[0231] The memory subsystem 1722 used in the memory subsystem 1710 can include a number of memories, such as a main random access memory (RAM) 1732 for storing instructions and data during program execution, and a read only memory (ROM) 1734 in which fixed instructions are stored. The file storage subsystem 1736 can provide a persistent storage device for programs and data files, and can include a hard disk drive, a related removable medium, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functions of a particular embodiment can be stored by the file storage subsystem 1736 within the memory subsystem 1710, or in other machines accessible by the processor.
[0232] The bus subsystem 1755 provides a mechanism for enabling the various components and subsystems of the computer system 1700 to communicate with each other as intended. Although the bus subsystem 1755 is shown schematically as a single bus, alternative embodiments of the bus subsystem can use multiple buses.
[0233] The computer system 1700 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely networked set of loosely distributed computers, or any other data processing system or user device. Due to the changing nature of computers and networks, the description of the computer system 1700 shown in FIG. 17 is intended only as a specific example for the purpose of illustrating a preferred embodiment of the present invention. Many other configurations of the computer system 1700 can have more or fewer components than the computer system shown in FIG. 17.
[0234] Particular embodiments The disclosed technology uses equalization-based image processing technology to attenuate spatial crosstalk from sensor pixels. The disclosed technology can be implemented as a system, method, or product. One or more features of an embodiment can be combined with a base embodiment. Embodiments that are not mutually exclusive are taught to be combinable. One or more features of an embodiment can be combined with other embodiments. The present disclosure periodically notifies users of these options. Omissions from some embodiments of the repeated listing of these options should not be construed as limiting the combinations taught in the foregoing sections. These descriptions are incorporated herein by reference to each of the following embodiments.
[0235] In one embodiment, the disclosed technology proposes a computer-implemented method for attenuating spatial crosstalk from sensor pixels.
[0236] The disclosed technology resolves spatial crosstalk on sensor pixels in the pixel plane caused by fluorescent samples periodically distributed in the sample plane. The signal cone from the fluorescent sample is optically coupled to the local grid of sensor pixels via at least one lens. The signal cones overlap and impinge on the sensor pixels, thereby generating spatial crosstalk.
[0237] The disclosed technology captures, in at least one sub-pixel look-up table, the characteristic spread of the characteristic signal cone projected through the lens and the resulting contribution of the characteristic signal cone to the fluorescence detected by the sensor pixels within the local grid of sensor pixels. The local grid of sensor pixels is substantially concentric with the center of the characteristic signal cone.
[0238] The disclosed technology interpolates between sets of sub-pixel look-up tables representing the characteristic spread at sub-pixel resolution to generate an interpolated look-up table based on the target fluorescent sample center.
[0239] The disclosed technique separates signals from a target fluorescent sample that projects the center of the signal cone substantially to the center of the target local grid of the sensor pixel by convolving the sensor pixel within the target local grid with an interpolation look-up table.
[0240] The disclosed technique uses the sum of the convolved contributions of the isolated signals as the intensity of the fluorescence from the target fluorescent sample.
[0241] The then-disclosed technique then basecalls a first target fluorescent sample using the fluorescence intensity. The fluorescence intensity is determined for the first target fluorescent sample for each imaging channel among a plurality of imaging channels. Consider a four-channel chemistry that uses four imaging channels to generate four images per array determination cycle. Then, for the first target fluorescent sample, four fluorescence intensities are determined using the technique disclosed as above. The four fluorescence intensities are then processed by a basecaller to basecall the first target fluorescent sample. Similarly, in a two-channel chemistry, two intensities of fluorescence are used to basecall the first target fluorescent sample.
[0242] The methods described in this section of the disclosure and other sections of the technique can include one or more of the following features and / or characteristics described in relation to the additional methods disclosed. For purposes of brevity, the combinations of features disclosed in this application are not individually enumerated and are not repeated in each base set of features. The reader will understand how the features identified in this way can be readily combined with the set of basic features identified as embodiments in other sections of this application.
[0243] In some embodiments, the periodically distributed fluorescent samples are arranged in a rhombus. In other embodiments, the periodically distributed fluorescent samples are arranged in a hexagonal shape.
[0244] Other embodiments of the methods described in this section can include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another embodiment of the methods described in this section can include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the methods described above.
[0245] In another embodiment, the disclosed technology proposes a computer-implemented method of a base call.
[0246] The disclosed technology accesses an image whose pixels exhibit intensity radiation from a target cluster and intensity radiation from additional adjacent clusters. The pixels include a central pixel that includes the center of the target cluster. Each pixel within a pixel can be divided into a plurality of sub-pixels.
[0247] Depending on a particular sub-pixel, in a plurality of sub-pixels of a central pixel that includes the center of the target cluster, the disclosed technology selects a sub-pixel look-up table corresponding to the particular sub-pixel from a bank of sub-pixel look-up tables. The selected sub-pixel look-up table includes pixel coefficients configured to receive intensity radiation from the target cluster and exclude intensity radiation from adjacent clusters.
[0248] The disclosed technology multiplies pixel coefficients element-wise with the intensity values of the pixels in the image, sums the products of the multiplications to generate an output.
[0249] The disclosed technology uses the output to base call a target cluster.
[0250] Each of the features considered in this particular embodiment section for other embodiments applies equally to this embodiment. As noted above, all methods are not repeated here and should be repeated by reference.
[0251] In some embodiments, the disclosed technology further includes: (i) selecting, from a bank of sub-pixel look-up tables, an additional sub-pixel look-up table corresponding to a sub-pixel that is adjacent and closest to a particular sub-pixel; (ii) interpolating between the pixel coefficients of the selected sub-pixel look-up table and the selected additional sub-pixel look-up table to generate an interpolated pixel coefficient configured to receive intensity radiation from a target cluster and reject intensity radiation from adjacent clusters; (iii) multiplying the interpolated pixel coefficient element-wise with the intensity value of a pixel in the image, summing the products of the multiplications to generate an output; and (iv) using the output to base-call a target cluster.
[0252] In some embodiments, the target cluster and additional adjacent clusters are periodically distributed in a diamond shape on a flow cell and immobilized on the wells of the flow cell. In other embodiments, the target cluster and additional adjacent clusters are periodically distributed on a hexagonal flow cell and immobilized on the wells of the flow cell.
[0253] In some embodiments, the interpolation is based on at least one of linear interpolation, bilinear interpolation, and bicubic interpolation.
[0254] In some embodiments, the pixel coefficients of the sub-pixel look-up tables in the bank of sub-pixel look-up tables are learned as a result of training an equalizer using decision-directed equalization. In one embodiment, decision-directed equalization uses least squares estimation as a loss function. In one embodiment, least squares estimation minimizes the mean squared error using a ground truth base-call. In one embodiment, the ground truth base-call is modified to account for a DC offset, an amplification factor, and a degree of polyclonality.
[0255] In some embodiments, the pixel coefficients of the sub-pixel look-up tables within a bank of sub-pixel look-up tables are derived from a combination of (i) a single sub-pixel look-up table in which the pixel coefficients are learned as a result of training an equalizer using decision-directed equalization, and (ii) a set of pre-computed interpolation filters. Each interpolation filter in the set of interpolation filters corresponds to each of the sub-pixels in a plurality of sub-pixels, respectively.
[0256] The disclosed technique further includes (i) aligning an image with respect to a template image and determining affine transformation parameters and non-linear transformation parameters, and (ii) using the parameters to transform the position coordinates of a target cluster and additional adjacent clusters into image coordinates of the image to generate a transformed image having the transformed pixels, and (iii) applying interpolation using the transformed position coordinates of the target cluster and the additional adjacent clusters to make each cluster center substantially concentric with the center of each transformed pixel including the cluster center, thereby making the center of the target cluster substantially concentric with the center of the center pixel.
[0257] The disclosed technique further includes generating an output for each image in a plurality of images captured using respective imaging channels in a particular sequencing cycle, and base-calling a target cluster using the output generated for each image.
[0258] Other embodiments of the methods described in this section can include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Still other embodiments of the methods described in this section can include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the methods described above.
[0259] The inventors disclose the following items. 1. A computer implementation method for a base call, the method comprising: accessing an image, wherein the pixels of the image exhibit intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, the pixels include a central pixel including the center of the target cluster, and each pixel within the pixel is divisible into a plurality of sub-pixels; selecting, from a bank of sub-pixel look-up tables, a sub-pixel look-up table corresponding to a particular sub-pixel among a plurality of sub-pixels of a central pixel including the center of the target cluster according to the particular sub-pixel, wherein the selected sub-pixel look-up table includes pixel coefficients configured to maximize a signal-to-noise ratio; multiplying the pixel coefficients element by element with respect to the intensity values of the pixels in the image, and summing the products of the multiplications to generate an output, wherein the pixel coefficients function as weights and the output is a weighted sum of the intensity values; using the output to base call the target cluster. 2. The computer implementation method according to item 1, wherein the signal maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise minimized in the signal-to-noise ratio is the intensity radiation from the adjacent clusters. 3. The computer implementation method according to item 1, wherein the element-by-element multiplication adds a bias to a given set of equalizer coefficients. 4. The computer implementation method according to item 3, wherein the bias is a DC offset that averages the background noise intensity. 5. selecting, from a bank of sub-pixel look-up tables, an additional sub-pixel look-up table corresponding to a sub-pixel that is adjacent to and closest to the particular sub-pixel; Interpolating between the pixel coefficients of the selected sub-pixel look-up table and the selected additional sub-pixel look-up table to generate interpolated pixel coefficients configured to maximize the signal-to-noise ratio; For each pixel intensity value in the interpolated image, multiplying the pixel coefficients element by element and summing the products of the multiplications to generate an output, where the interpolated pixel coefficients function as weights and the output is the weighted sum of the intensity values; Using the output to basecall a target cluster, further comprising the computer-implemented method of item 1. 6. The computer-implemented method of item 1, wherein the target cluster and the additional adjacent clusters are periodically distributed in a rhombus pattern on the flow cell and immobilized on the wells of the flow cell. 7. The computer-implemented method of item 6, wherein the target cluster and the additional adjacent clusters are periodically distributed on a hexagonal flow cell and immobilized on the wells of the flow cell. 8. The computer-implemented method of item 1, wherein the interpolation is based on at least one of linear interpolation, bilinear interpolation, and bicubic interpolation. 9. The pixel coefficients of the sub-pixel look-up tables within a bank of sub-pixel look-up tables are learned as a result of training an equalizer using at least one of least squares estimation, least squares method, least mean squares, and recursive least squares. In other embodiments, other estimation algorithms and adaptation algorithms can be used to train the equalizer. The computer-implemented method of item 1. 10. The computer-implemented method of item 9, further comprising training the equalizer in an offline mode, in which the pixel coefficients of the sub-pixel look-up tables are fixed after being trained with a batch of training data from a previously executed array determination run. 11. Further including training the equalizer in an online mode, in which, as training data from an ongoing sequencing run becomes available, the pixel coefficients of the sub-pixel look-up table are iteratively updated, the computer-implemented method according to item 10. 12. Further including accessing the intensity distribution for each base of the four bases A, C, G, and T generated during a previous base call of an image in the training data, selecting the center of each intensity distribution for each base as the ground truth target intensity for each base, and training the equalizer using the ground truth target intensity for each base, the computer-implemented method according to item 11. 13. Further including pre-training the equalizer in an offline mode and re-training the equalizer in an online mode, the computer-implemented method according to item 12. 14. Further including generating a look-up table within a bank of sub-pixel look-up tables by applying a single set of equalizer coefficients and a pre-computed set of interpolation filters together, including interpolating pixel intensities to generate an input to the equalizer. This includes calculating pixel weights for clusters having a substantially different alignment with respect to the pixels as compared to the trained equalizer coefficients by using the interpolated pixel intensity values for generating the equalizer input. The interpolation and equalizer filter responses can be convolved together for an efficient implementation using a single shared LUT. In other embodiments, the interpolation filter calculations can be performed directly without binning to sub-pixels. 15. Aligning an image with respect to a template image and determining affine transformation parameters and non-linear transformation parameters, Using the parameters to transform the position coordinates of a target cluster and additional adjacent clusters to the image coordinates of the image and generating a transformed image having the transformed pixels, Applying interpolation using the transformed position coordinates of the target cluster and additional adjacent clusters, and making each cluster center substantially concentric with the center of each transformed pixel including the cluster center, thereby further including making the center of the target cluster concentric with the center of the central pixel, the computer-implemented method according to item 1. 16. The computer-implemented method according to item 4, further comprising generating an output for each image in a plurality of images captured using respective imaging channels and / or color channels in a specific sequencing cycle, and baselining a target cluster using the output respectively generated for each image. 17. A computer-implemented method for restoring a signal underlying a signal destroyed by surrounding fluorescent light sources in a sample plane from a fluorescent sample disposed in the sample plane, the method comprising: In at least one sub-pixel look-up table, capturing a characteristic set of illuminations in the image plane by a sensor pixel array based on sampling considering the destruction from surrounding fluorescent light sources, and then generating a set of look-up tables for the characteristic set of illuminations by the sensor pixel array when the central coordinates of the fluorescent sample are distributed over the central pixel of the sensor array and the position is distributed with respect to the center of the coordinates of the central pixel; Receiving an image having the central coordinates of the fluorescent sample at any location of the central pixel of the sensor pixel array, the image being destroyed by surrounding fluorescent light sources, receiving the central coordinates of the fluorescent sample within the central pixel; Calculating an interpolation table of the characteristic set of illuminations by the sensor pixel array customized for the received central coordinates of the fluorescent sample based on interpolation between the look-up tables in the set of look-up tables; Restoring the signal from the target fluorescent sample that projects the center of the signal cone substantially to the center of the target local grid of the sensor pixels by multiplying the interpolation look-up table element by element using the sensor pixels in the target local grid. using the sum of the products of each element as the fluorescence intensity from the target fluorescent sample, and using the fluorescence intensity to baseline the first target fluorescent sample, a method comprising. 1. A computer-implemented method for baselining, the method comprising: accessing an image, wherein the pixels of the image indicate intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, selecting a look-up table including pixel coefficients configured to maximize a signal-to-noise ratio, convolving the pixel coefficients with the intensity values of the pixels in the image to generate an output, and baselining the target cluster based on the output. 2. The computer-implemented method according to claim 1, wherein the signal maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise minimized in the signal-to-noise ratio is the intensity radiation from the adjacent clusters plus an additional noise source. 3. The computer-implemented method according to claim 1, wherein the pixels include a central pixel including the center of the target cluster, and each pixel within the pixel is divisible into a plurality of sub-pixels. 4. The computer-implemented method according to claim 3, wherein the look-up table is a sub-pixel look-up table. 5. selecting, in a plurality of sub-pixels of a central pixel including the center of the target cluster, a sub-pixel look-up table corresponding to a specific sub-pixel from a bank of sub-pixel look-up tables according to the specific sub-pixel, wherein the selected sub-pixel look-up table includes pixel coefficients, multiplying the pixel coefficients element-wise by the intensity values of the pixels in the image and summing the products to generate an output, wherein the pixel coefficients function as weights and the output is a weighted sum of the intensity values, Basecalling, further comprising: using the output to basecall a target cluster, generating the output of each imaging channel in a plurality of imaging channels, and using the output of each imaging channel to basecall the target cluster, the computer-implemented method according to claim 4. 6. The computer-implemented method according to claim 5, wherein the multiplication for each element adds a bias of a given equalizer coefficient set, and the bias is a DC offset that averages the background noise intensity. 7. Selecting, from a bank of sub-pixel look-up tables, an additional sub-pixel look-up table corresponding to sub-pixels that are successively adjacent to a particular sub-pixel; Generating interpolation pixel coefficients configured to maximize a signal-to-noise ratio based on the pixel coefficients of the selected sub-pixel look-up table and the selected additional sub-pixel look-up table; Convolving the interpolation pixel coefficients with the intensity values of the pixels in the image to generate an output; Basecalling the target cluster based on the output, the computer-implemented method according to claim 5. 8. Generating, by multiplying pixel coefficients element-wise with the intensity values of the pixels in the image and summing the products of the multiplications to generate an output, wherein the interpolation pixel coefficients function as weights and the output is a weighted sum of the intensity values, the computer-implemented method according to claim 7. 9. The computer-implemented method according to claim 1, further comprising training an equalizer using at least one of least squares estimation, least squares method, least mean squares, and recursive least squares to generate pixel coefficients. 10. The computer-implemented method according to claim 9, further comprising training the equalizer in an offline mode, wherein in the offline mode, the pixel coefficients of the sub-pixel look-up tables are fixed after being trained with a batch of training data from a previously executed array determination run. 11. The computer-implemented method according to claim 10, further comprising training the equalizer in an online mode, wherein in the online mode, the pixel coefficients of the sub-pixel look-up table are repeatedly updated during an ongoing alignment run. 12. The computer-implemented method according to claim 11, further comprising accessing the intensity distribution for each base of the four bases A, C, G, and T generated during a previous base call of an image in training data, selecting the center of each intensity distribution for each base as the ground truth target intensity for each base of the corresponding color channel, and training the equalizer using the ground truth target intensity for each base. 13. The computer-implemented method according to claim 12, further comprising pre-training the equalizer in an offline mode and re-training the equalizer in an online mode. 14. The computer-implemented method according to claim 9, further comprising generating a look-up table in a bank of sub-pixel look-up tables by applying both a single set of equalizer coefficients and a pre-computed set of interpolation filters, including interpolating pixel intensities to generate an input to the equalizer. 15. Aligning an image with respect to a template image and determining affine transformation parameters and non-linear transformation parameters; Using the parameters to transform the position coordinates of a target cluster and additional adjacent clusters to image coordinates of the image and generating a transformed image having the transformed pixels; The computer-implemented method according to claim 1, further comprising making the center of the target cluster concentric with the center of the central pixel by applying interpolation using the transformed position coordinates of the target cluster and the additional adjacent clusters and making the center of each cluster concentric with the center of the respective transformed pixels including the center of the cluster. 16. A non-transitory computer-readable storage medium storing computer program instructions for performing a base call, the instructions, when executed on a processor, Accessing an image, wherein pixels of the image indicate intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, and selecting a look-up table including pixel coefficients configured to maximize a signal-to-noise ratio, and convolving the pixel coefficients with intensity values of pixels in the image to generate an output, and base-calling a target cluster based on the output, the non-transitory computer-readable storage medium implementing instructions including these steps. 17. The non-transitory computer-readable storage medium according to claim 16, wherein the signal maximized in signal-to-noise ratio is intensity radiation from a target cluster, and the noise minimized in signal-to-noise ratio is the intensity radiation from adjacent clusters plus an additional noise source. 18. The non-transitory computer-readable storage medium according to claim 16, further implementing a method that further includes training an equalizer using at least one of least squares estimation, least squares method, least mean squares, and recursive least squares to generate pixel coefficients. 19. A system including one or more processors coupled to a memory, the memory having computer instructions for performing a base call loaded therein, the instructions, when executed on the processor, accessing an image, wherein pixels of the image indicate intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, and selecting a look-up table including pixel coefficients configured to maximize a signal-to-noise ratio, and convolving the pixel coefficients with intensity values of pixels in the image to generate an output, and base-calling a target cluster based on the output, the system implementing actions including these steps. 20. The system according to claim 19, further implementing an action including training an equalizer using at least one of least squares estimation, least squares method, least mean squares, and recursive least squares to generate pixel coefficients.
[0260] The present invention has been disclosed with reference to the above-described preferred embodiments and examples, which should be understood to be intended in an illustrative rather than a limiting sense. Those skilled in the art will readily make changes and combinations, and such changes and combinations are considered to be within the spirit of the present invention and the scope of the following claims.
Explanation of Signs
[0261] 100A system 102 array determination image 104 equalizer-based cooler 106 LUT / LUT bank 108 interpolation filter 112 ground-through base call 114 trainer
Claims
1. A computer-implemented method for a base call, the computer-implemented method comprising: accessing an image, wherein the pixels of the image exhibit intensity radiation from a target cluster and intensity radiation from additional adjacent clusters; selecting, from a bank of lookup tables, a lookup table that includes pixel coefficients configured to increase a signal-to-noise ratio; applying the pixel coefficients to intensity values of the pixels in the image to generate an output; and base calling the target cluster based on the output.
2. The computer-implemented method of claim 1, wherein the pixel coefficients are applied to the intensity values using a convolution operation.
3. The computer-implemented method of claim 1, wherein the pixel coefficients are applied to the intensity values using an interpolation operation.
4. The computer-implemented method of claim 1, wherein the signal increasing in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise decreasing in the signal-to-noise ratio is the intensity radiation from the additional adjacent clusters plus an additional noise source.
5. The computer-implemented method of claim 1, wherein the pixels include a central pixel that includes the center of the target cluster, and each pixel within the pixels is divisible into a plurality of sub-pixels.
6. The computer-implemented method of claim 1, wherein the selected lookup table is a sub-pixel lookup table, and the bank of lookup tables is a plurality of sub-pixel lookup tables.
7. selecting, from the bank of sub-pixel lookup tables, a sub-pixel lookup table corresponding to a particular sub-pixel among a plurality of sub-pixels of a central pixel that includes the center of the target cluster, the selected sub-pixel lookup table including the pixel coefficients. Multiplying each element of the pixel coefficients by the intensity value of the pixel in the image, and summing the products of the multiplication for each element to generate the output, wherein the pixel coefficients function as weights and the output is a weighted sum of the intensity values; generating; Base-calling the target cluster using the output, including generating the output for each imaging channel in a plurality of imaging channels and base-calling the target cluster using the output for each imaging channel; base-calling; further comprising the computer-implemented method according to claim 6. Claim 8 The multiplication for each element adds a bias of a given equalizer coefficient set, and the bias is a DC offset that averages the background noise intensity; the computer-implemented method according to claim 7. Claim 9 Selecting an additional sub-pixel look-up table corresponding to sub-pixels that are continuously adjacent to the specific sub-pixel from the bank of sub-pixel look-up tables; Generating interpolation pixel coefficients configured to increase the signal-to-noise ratio based on the pixel coefficients of the selected sub-pixel look-up table and the selected additional sub-pixel look-up table; Convolving the interpolation pixel coefficients with the intensity values of the pixels in the image to generate an output; Base-calling the target cluster based on the output; further comprising the computer-implemented method according to claim 7. Claim 10 Multiplying each element of the interpolation pixel coefficients by the intensity value of the pixel in the image, and summing the products of the multiplication to generate the output, wherein the interpolation pixel coefficients function as weights and the output is a weighted sum of the intensity values; generating; further comprising the computer-implemented method according to claim 9. Claim 11 Further comprising training an equalizer using at least one of least squares estimation, least squares method, least mean squares, and recursive least squares to generate the pixel coefficients; the computer-implemented method according to claim 1. Claim 12 The computer-implemented method of claim 11, further comprising training the equalizer in an offline mode where the pixel coefficients of the sub-pixel look-up table are fixed after being trained with a batch of training data from a previously executed alignment run.
13. The computer-implemented method of claim 12, further comprising training the equalizer in an online mode where the pixel coefficients of the sub-pixel look-up table are repeatedly updated during an ongoing alignment run.
14. The computer-implemented method of claim 13, further comprising accessing an intensity distribution for each base of four bases A, C, G, and T generated during a previous base call of an image in the training data, selecting a center of each of the intensity distributions for each base as a ground truth target intensity for each base of a corresponding color channel, and training the equalizer using the ground truth target intensities for each base.
15. The computer-implemented method of claim 14, further comprising pre-training the equalizer in the offline mode and re-training the equalizer in the online mode.
16. The computer-implemented method of claim 11, further comprising generating the look-up table within the bank of the sub-pixel look-up table by applying both a single set of equalizer coefficients and a pre-computed set of interpolation filters, including interpolating pixel intensities to generate an input to the equalizer.
17. Aligning the image with respect to a template image and determining affine transformation parameters and non-linear transformation parameters. Using the affine transformation parameters and the non-linear transformation parameters to transform position coordinates of the target cluster and the additional adjacent clusters into image coordinates of the image and generating a transformed image having the transformed pixels. The computer-implemented method of claim 1, further comprising making the center of the target cluster concentric with the center of the central pixel by applying interpolation using the transformed position coordinates of the target cluster and the additional adjacent clusters and making each cluster center concentric with the center of the respective transformed pixels including the cluster center.
18. A non-transitory computer-readable storage medium storing computer program instructions for performing a base call, wherein the computer program instructions, when executed on a processor, access an image, wherein pixels of the image indicate intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, and access the image; select a look-up table containing pixel coefficients configured to increase a signal-to-noise ratio from a bank of look-up tables; apply the pixel coefficients to intensity values of the pixels in the image to generate an output; perform a base call on the target cluster based on the output, and implement a method including the steps above. **Claim 19** The non-transitory computer-readable storage medium according to claim 18, wherein the pixel coefficients are applied to the intensity values using a convolution operation. **Claim 20** The non-transitory computer-readable storage medium according to claim 18, wherein the pixel coefficients are applied to the intensity values using an interpolation operation. **Claim 21** A system including one or more processors coupled to a memory, wherein the memory has loaded therein computer instructions for performing a base call, and the computer instructions, when executed on the one or more processors, access an image, wherein pixels of the image indicate intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, and access the image; select a look-up table containing pixel coefficients configured to increase a signal-to-noise ratio from a bank of look-up tables; apply the pixel coefficients to intensity values of the pixels in the image to generate an output; perform a base call on the target cluster based on the output, and implement actions including the steps above. **Claim 22** The system according to claim 21, wherein the pixel coefficients are applied to the intensity values using a convolution operation. **Claim 23** The system according to claim 21, wherein the pixel coefficients are applied to the intensity values using an interpolation operation. **Claim 24** A system including one or more processors coupled to a memory, wherein the memory has computer instructions for performing a base call loaded, and when the computer instructions are executed on the one or more processors, accessing an image, wherein pixels of the image exhibit intensity radiation from a target cluster and intensity radiation from additional adjacent clusters, selecting a look-up table including pixel coefficients configured to increase a signal-to-noise ratio from a bank of look-up tables, convolving the pixel coefficients with intensity values of the pixels in the image to generate convolved features, interpolating the convolved features to generate an output, base-calling the target cluster based on the output, A system implementing actions including.
25. The system according to claim 24, wherein the signal increasing in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise decreasing in the signal-to-noise ratio is the intensity radiation from the additional adjacent clusters plus an additional noise source.
26. The system according to claim 24, wherein the selected look-up table is a sub-pixel look-up table and the bank of look-up tables is a plurality of sub-pixel look-up tables.
Citation Information
Patent Citations
Fading correction method
JP2020506677A
Method and computer program product for genotype classification
US20140153801A1
Basecaller for DNA sequencing using machine learning
US20150169824A1