Intensity extraction by interpolation and fitting for base calling
Patent Information
- Application Number
- JP2023579791
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-26
- Filing Date
- 2022-07-14
- Publication Date
- 2025-07-23
AI Technical Summary
Existing DNA sequencing systems face challenges in accurately distinguishing between true light signals from wells of interest and unwanted light signals due to spatial crosstalk between adjacent wells, leading to sequencing errors and reduced base call accuracy.
The use of trained sharpening masks to attenuate spatial crosstalk in cluster intensity data by applying convolution operations and interpolation techniques, enhancing the signal-to-noise ratio in DNA sequencing images.
Improves base call accuracy by reducing sequencing errors caused by spatial crosstalk, thereby enhancing the reliability of DNA sequencing results.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Priority Application This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 223,408, entitled "Specialist Signal Profilers for Base Calling," filed July 19, 2021 (Attorney Docket No. ILLM 1041-1 / IP-2063-PRV), which is incorporated herein by reference for all purposes.
[0002] This application also claims the benefit of and priority to U.S. Nonprovisional Patent Application No. 17 / 511,483, entitled “Intensity Extraction with Crosstalk Attenuation Using Interpolation and Adaptation Calling,” filed on October 26, 2021 (Attorney Docket No. ILLM 1053-1 / IP-2214-US).
[0003] Joint ownership statement In accordance with 35 U.S.C. § 102(b)(2)(C) and MPEP § 2146.02(I), Applicant hereby represents that this application, U.S. Provisional Patent Application No. 63 / 020,449, and U.S. Nonprovisional Patent Application No. 17 / 308,035, were owned by or under obligation of assignment to the same person (Illumina, Inc.) prior to the effective filing date of this application, and that Illumina Software, Inc., the applicant and assignee named in this application, is a wholly owned subsidiary of Illumina, Inc.
[0004] The disclosed technology relates to artificial intelligence-based computers and digital data processing systems and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using deep neural networks, such as deep convolutional neural networks, to analyze data.
[0005] Built-in The following are incorporated by reference for all purposes as if fully set forth herein:
[0006] U.S. Nonprovisional Patent Application No. 15 / 936,365, filed March 26, 2018, entitled “Detection Apparatus Having a Microfluorometer, a Fluidic System, and a Flow Cell Latch Clamp Module”; U.S. Nonprovisional Patent Application No. 16 / 567,224, entitled “Flow Cells and Methods Related to Same,” filed on September 11, 2019; U.S. Nonprovisional Patent Application No. 16 / 439,635, entitled “Device for Luminescent Imaging,” filed June 12, 2019; U.S. Nonprovisional Patent Application No. 15 / 594,413, entitled “Integrated Optoelectronic Read Head and Fluidic Cartridge Useful for Nucleic Acid Sequencing,” filed May 12, 2017; U.S. Nonprovisional Patent Application No. 16 / 351,193, entitled “Illumination for Fluorescence Imaging Using Objective Lens,” filed March 12, 2019; U.S. Nonprovisional Patent Application No. 12 / 638,770, entitled "Dynamic Autofocus Method and System for Assay Imager," filed December 15, 2009; U.S. Nonprovisional Patent Application No. 13 / 783,043, entitled "Kinetic Exclusion Amplification of Nucleic Acid Libraries," filed March 1, 2013; U.S. Non-Provisional Patent Application No. 13 / 006,206, entitled "Data Processing System and Method," filed January 13, 2011; U.S. Nonprovisional Patent Application No. 14 / 530,299, entitled “Image Analysis Useful for Patterned Objects,” filed October 31, 2014; U.S. Nonprovisional Patent Application No. 15 / 153,953, entitled “Methods and Systems for Analyzing Image Data,” filed December 3, 2014; U.S. Non-provisional Patent Application No. 14 / 020,570, filed September 6, 2013, entitled "Centroid Markers for Image Analysis of High Density Clusters In Complex Polynucleotide Sequencing"; U.S. Nonprovisional Patent Application No. 14 / 530,299, entitled “Image Analysis Useful for Patterned Objects,” filed October 31, 2014; U.S. Nonprovisional Patent Application No. 12 / 565,341, filed September 23, 2009, entitled "Method and System for Determining the Accuracy of DNA Base Identifications"; U.S. Nonprovisional Patent Application No. 12 / 295,337, entitled "Systems and Devices for Sequence by Synthesis Analysis," filed March 30, 2007; U.S. Nonprovisional Patent Application No. 12 / 020,739, entitled "Image Data Efficient Genetic Sequencing Method and System," filed January 28, 2008; U.S. Nonprovisional Patent Application No. 13 / 833,619, entitled “Biosensors for Biological or Chemical Analysis and Systems and Methods for Same,” filed March 15, 2013 (Attorney Docket No. IP-0626-US); U.S. Nonprovisional Patent Application No. 15 / 175,489, entitled “Biosensors for Biological or Chemical Analysis and Methods of Manufacturing the Same,” filed on June 7, 2016 (Attorney Docket No. IP-0689-US); U.S. Nonprovisional Patent Application No. 13 / 882,088, entitled “Microdevices and Biosensor Cartridges for Biological or Chemical Analysis and Systems and Methods for the Same,” filed April 26, 2013 (Attorney Docket No. IP-0462-US); U.S. Nonprovisional Patent Application No. 13 / 624,200, entitled “Methods and Compositions for Nucleic Acid Sequencing,” filed September 21, 2012 (Attorney Docket No. IP-0538-US); U.S. Provisional Patent Application No. 62 / 821,602, entitled “Training Data Generation for Artificial Intelligence-Based Sequencing,” filed March 21, 2019 (Attorney Docket No. ILLM1008-1 / IP-1693-PRV); U.S. Provisional Patent Application No. 62 / 821,618, entitled “Artificial Intelligence-Based Generation of Sequencing Metadata,” filed March 21, 2019 (Attorney Docket No. ILLM1008-3 / IP-1741-PRV); U.S. Provisional Patent Application No. 62 / 821,681, entitled “Artificial Intelligence-Based Base Calling,” filed March 21, 2019 (Attorney Docket No. ILLM1008-4 / IP-1744-PRV); U.S. Provisional Patent Application No. 62 / 821,724, entitled “Artificial Intelligence-Based Quality Scoring,” filed March 21, 2019 (Attorney Docket No. ILLM1008-7 / IP-1747-PRV); U.S. Provisional Patent Application No. 62 / 821,766, entitled “Artificial Intelligence-Based Sequencing,” filed March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV); Dutch Patent Application No. 2023310, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing", filed on June 14, 2019 (Attorney Reference No. ILLM1008-11 / IP-1693-NL); Dutch patent application No. 2023311, entitled "Artificial Intelligence-Based Generation of Sequencing Metadata", filed on June 14, 2019 (Attorney Docket No. ILLM1008-12 / IP-1741-NL); Dutch patent application No. 2023312, entitled "Artificial Intelligence-Based Base Calling", filed on June 14, 2019 (Attorney Docket No. ILLM1008-13 / IP-1744-NL); Dutch patent application No. 2023314, entitled "Artificial Intelligence-Based Quality Scoring", filed on June 14, 2019 (Attorney Docket No. ILLM1008-14 / IP-1747-NL); Dutch Patent Application No. 2023316, entitled "Artificial Intelligence-Based Sequencing", filed on June 14, 2019 (Attorney Docket No. ILLM1008-15 / IP-1752-NL); U.S. Nonprovisional Patent Application No. 16 / 825,987, entitled “Training Data Generation for Artificial Intelligence-Based Sequencing,” filed on March 20, 2020 (Attorney Docket No. ILLM1008-16 / IP-1693-US); U.S. Nonprovisional Patent Application No. 16 / 825,991, entitled “Training Data Generation for Artificial Intelligence-Based Sequencing,” filed on March 20, 2020 (Attorney Docket No. ILLM1008-17 / IP-1741-US); U.S. Nonprovisional Patent Application No. 16 / 826,126, entitled “Artificial Intelligence-Based Base Calling,” filed on March 20, 2020 (Attorney Docket No. ILLM1008-18 / IP-1744-US); U.S. Nonprovisional Patent Application No. 16 / 826,134, entitled “Artificial Intelligence-Based Quality Scoring,” filed on March 20, 2020 (Attorney Docket No. ILLM1008-19 / IP-1747-US); U.S. Nonprovisional Patent Application No. 16 / 826,168, entitled “Artificial Intelligence-Based Sequencing,” filed on March 21, 2020 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV); U.S. Provisional Patent Application No. 62 / 849,091, entitled “Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing,” filed May 16, 2019 (Attorney Docket No. ILLM1011-1 / IP-1750-PRV); U.S. Provisional Patent Application No. 62 / 849,132, entitled “Base Calling Using Convolutions,” filed May 16, 2019 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2); U.S. Provisional Patent Application No. 62 / 849,133, entitled “Base Calling Using Compact Convolutions,” filed May 16, 2019 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3); U.S. Provisional Patent Application No. 62 / 979,384, entitled “Artificial Intelligence-Based Base Calling of Index Sequences,” filed on February 20, 2020 (Attorney Docket No. ILLM1015-1 / IP-1857-PRV); U.S. Provisional Patent Application No. 62 / 979,414, entitled “Artificial Intelligence-Based Many-To-Many Base Calling,” filed on February 20, 2020 (Attorney Docket No. ILLM1016-1 / IP-1858-PRV); U.S. Provisional Patent Application No. 62 / 979,385, entitled “Knowledge Distillation-Based Compression of Artificial Intelligence-Based Base Caller,” filed on February 20, 2020 (Attorney Docket No. ILLM1017-1 / IP-1859-PRV); U.S. Provisional Patent Application No. 62 / 979,412, entitled “Multi-Cycle Cluster Based Real Time Analysis System,” filed on February 20, 2020 (Attorney Docket No. ILLM1020-1 / IP-1866-PRV); U.S. Provisional Patent Application No. 62 / 979,411, entitled “Data Compression for Artificial Intelligence-Based Base Calling,” filed on February 20, 2020 (Attorney Docket No. ILLM1029-1 / IP-1964-PRV); U.S. Nonprovisional Patent Application No. 63 / 020,449, entitled “Equalization-Based Image Processing and Spatial Crosstalk Attenuator,” filed on May 5, 2020 (Attorney Docket No. ILLM 1032-1 / IP-1991-PRV); U.S. Nonprovisional Patent Application No. 17 / 308,035, entitled “Equalization-Based Image Processing and Spatial Crosstalk Attenuator,” filed on May 4, 2021 (Attorney Docket No. ILLM 1032-2 / IP-1991-US); U.S. Provisional Patent Application No. 62 / 979,399, entitled “Squeezing Layer for Artificial Intelligence-Based Base Calling,” filed February 20, 2020 (Attorney Docket No. ILLM1030-1 / IP-1982-PRV). [Background technology]
[0007] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which as such may also correspond to embodiments of the claimed technology.
[0008] Rapid improvements in computing power have enabled deep Convolutional Neural Networks (CNNs) to achieve great success in many computer vision tasks in recent years, with significantly improved accuracy. During the inference stage, many applications require low-latency processing of a single image with strict power consumption requirements, which reduces the efficiency of Graphics Processing Units (GPUs) and other general-purpose platforms, which creates an opportunity for specific acceleration hardware, such as Field Programmable Gate Arrays (FPGAs), by customizing digital circuits to be particularly effective for inference of deep learning algorithms. However, deploying CNNs in portable and embedded systems remains challenging due to large data volumes, intensive computations, various algorithm structures, and frequent memory accesses.
[0009] Since convolution provides most of the operations in CNN, the convolution acceleration scheme will greatly affect the efficiency and performance of hardware CNN accelerators. Convolution involves multiply and accumulate (MAC) operations with four levels of loops that slide along the kernel and feature maps. The first loop level calculates the MAC of pixels in one kernel window. The second loop level accumulates the sum of MAC products over various different input feature maps. After completing the first and second loop levels, the final output element in the output feature map is obtained by adding a bias. The third loop level slides the kernel window in the input feature map. The fourth loop level generates various different output feature maps.
[0010] FPGAs have attracted more interest and become more widespread, especially for accelerating inference tasks. This is because FPGAs (1) are highly reconfigurable, (2) are superior to application specific integrated circuits (ASICs) in terms of the development time required to catch up with the rapid evolution of CNNs, (3) have good performance, and (4) are more energy efficient than GPUs. The high performance and efficiency of FPGAs can be achieved by synthesizing circuits customized for specific calculations and directly processing billions of operations with customized memory systems. For example, hundreds to thousands of digital signal processing (DSP) blocks in modern FPGAs support core convolution operations, such as multiply-add operations with high parallelism. Dedicated data buffers between external on-chip memories and on-chip processing engines (PEs) can be designed to realize prioritized data flow by configuring tens of megabytes of on-chip block random access memory (BRAM) on the field programmable gate array (FPGA) chip.
[0011] Efficient data flow and hardware architecture for CNN acceleration is desired to minimize data communication while maximizing resource utilization to achieve high performance. This creates an opportunity to design methodologies and frameworks to accelerate the inference process of various CNN algorithms on acceleration hardware and achieve high performance, high efficiency, and high flexibility. CNN algorithms and other machine learning algorithms can be applied to various application domains, including base calling of unknown nucleotides (e.g., A, C, T, or G) using biological sequencing machines.
[0012] Various protocols in biological or chemical research involve carrying out a large number of controlled reactions on a localized support surface or in a predetermined reaction chamber. The desired reactions can then be observed or detected, and subsequent analysis can help identify or reveal the properties of the chemicals involved in the reactions. For example, in some multiplex assays, an unknown analyte with a distinguishable label (e.g., a fluorescent label) can be exposed to thousands of known probes under controlled conditions. Each known probe can be deposited in a corresponding well of a microplate. Observing any chemical reactions that occur between the known probe and the unknown analyte in the well can help identify or reveal the properties of the analyte. Other examples of such protocols include known DNA sequencing processes, such as sequencing by synthesis or circular array sequencing. In circular array sequencing, a high-density array of DNA features (e.g., template nucleic acid) is sequenced through repeated cycles of enzymatic manipulation. After each cycle, an image is captured and can be subsequently analyzed with other images to determine the sequence of the DNA features.
[0013] As a more specific example, one known DNA sequencing system uses a pyrosequencing process and includes a chip with a fused fiber optic faceplate with millions of wells. A single capture bead carrying clonally amplified sstDNA from a genome of interest is deposited in each well. After the capture bead is deposited in the well, nucleotides are added to the wells sequentially by flowing a solution containing specific nucleotides along the faceplate. The environment within the wells is such that if a nucleotide flowing through a particular well complements a DNA strand on the corresponding capture bead, the nucleotide is added to the DNA strand. The colonies of DNA strands are called clusters. Incorporation of the nucleotide into the clusters initiates a process that ultimately produces a chemiluminescent signal. The system includes a CCD camera positioned directly adjacent to the faceplate and configured to detect light signals from the DNA clusters in the wells. Subsequent analysis of images obtained throughout the pyrosequencing process allows the sequence of the genome of interest to be determined.
[0014] However, the pyrosequencing system described above, in addition to other systems, may have certain limitations. For example, the fiber optic faceplate is acid etched to form millions of tiny wells. The wells may be approximately spaced apart from one another, but it is difficult to know the exact location of the wells in relation to other adjacent wells. If a CCD camera were positioned directly adjacent to the faceplate, the wells would not be evenly distributed along the pixels of the CCD camera, and thus the wells would not be aligned with the pixels in a known manner. Spatial crosstalk is inter-well crosstalk between adjacent wells, making it difficult to distinguish true light signals from the wells of interest from other unwanted light signals in subsequent analysis. Also, the fluorescent emission is substantially isotropic. As the density of analytes increases, it becomes increasingly difficult to manage or account for unwanted emissions (e.g., crosstalk) from adjacent analytes. As a result, the data recorded during a sequencing cycle must be carefully analyzed.
[0015] Base calling accuracy is crucial for high-throughput DNA sequencing and downstream analyses such as read mapping and genome assembly. Spatial crosstalk between adjacent clusters explains a large portion of sequencing errors. Therefore, correcting for spatial crosstalk in cluster intensity data presents an opportunity to reduce DNA sequencing errors and improve base calling accuracy. [Brief description of the drawings]
[0016] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings, in which: [Figure 1] 1 shows a cross-sectional view of a biosensor that can be used in various embodiments. [Diagram 2] 1 shows an embodiment of a flow cell that includes clusters within its tiles. [Diagram 3] An exemplary flow cell with eight lanes is shown, along with a zoom-in of one tile and its cluster and their surrounding background. [Figure 4] FIG. 1 is a simplified block diagram of a system for analysis of sensor data from a sequencing system, such as base call sensor output. [Diagram 5] FIG. 2 is a simplified diagram illustrating aspects of a base call operation, including functions of a runtime program executed by a host processor. [Figure 6] 5 is a simplified diagram of a configuration of a configurable processor, such as the configurable processor of FIG. 4. [Figure 7] 1 illustrates a system for generating and / or updating a sharpening mask. [Figure 8A] FIG. 1 shows multiple sharpening masks used for corresponding sections of a sequencing image generated for corresponding regions of a flow cell, where each tile of the flow cell is divided into 3×3 subtile regions, and each subtile region is assigned one or more corresponding sharpening masks. [Figure 8B] FIG. 1 shows multiple sharpening masks used for corresponding sections of a sequencing image generated for corresponding regions of a flow cell, where each tile of the flow cell is divided into 1 x 9 subtile regions, and each subtile region is assigned one or more corresponding sharpening masks. [Figure 8C] Shown are multiple sharpening masks used for corresponding sections of a sequencing image generated for corresponding regions of a flow cell, where each tile of the flow cell is divided into multiple periodically occurring subtile regions, and similar periodically occurring subregions within a tile are assigned one or more corresponding sharpening masks. [Figure 9A] 13 shows one embodiment of a base-by-base Gaussian fit centered on the base-by-base target that is used as a ground truth value for error calculation during training. [Figure 9B] 1 illustrates one embodiment of a fitting technique that can be used to train a basecaller. [Figure 10A] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10B] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10C] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10D] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10E]We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10F] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10G] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10H] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10I] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10J] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 10K] We jointly present various implementations that use trained sharpening masks to attenuate spatial crosstalk from sensor pixels and base call clusters using crosstalk-corrected sensor data. [Figure 11A] 1 shows a method of base calling based on convolution and subsequent interpolation of at least one section of a sequencing image to assign one or more weighted feature values to a cluster, and base calling of the cluster based on the assigned one or more weighted feature values. [Figure 11B]13 shows a comparison of performance results of the disclosed intensity extraction technique using a sharpening mask with various other intensity extraction techniques associated with base calling. [Figure 11C] 13 shows a comparison of other performance results between the disclosed technique using a sharpening mask and various other techniques for base calling. [Figure 12] The present invention shows a method of base calling based on convolution and subsequent interpolation of at least one section of a sequencing image to assign one or more weighted feature values to clusters, and base calling of clusters based on feature values attached to the assigned one or more weighted feature values, where coefficients of a sharpening mask are adaptively updated during sequencing operations. [Figure 13] 13 shows the adaptation of the coefficients of the sharpening mask used for intensity extraction. [Figure 14] 1 shows a comparison of performance results of the disclosed intensity extraction technique using a sharpening mask and adaptation to another intensity extraction technique that does not use adaptation. [Figure 15] 1 shows a comparison of performance results of the disclosed intensity extraction technique using a sharpening mask and adaptation to another intensity extraction technique that does not use adaptation. [Figure 16] A computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] The following description is typically made with reference to specific structural embodiments and methods. It is not intended to limit the present technology to the specifically disclosed embodiments and methods, but it should be understood that the present technology can be implemented using other features, elements, methods and embodiments. Preferred embodiments are described to illustrate the present technology, not to limit the scope defined by the claims. Those skilled in the art will recognize various equivalent variations to the following description.
[0018] As used herein, the term "polynucleotide" or "nucleic acid" refers to deoxyribonucleic acid (DNA); however, where appropriate, one of skill in the art will recognize that the systems and devices herein can also be utilized with ribonucleic acid (RNA). These terms should be understood to include, as equivalents, analogs of either DNA or RNA made from nucleotide analogs. The terms as used herein also encompass cDNA, which is the complementary or copy DNA created from an RNA template, for example, by the action of reverse transcriptase.
[0019] The single-stranded polynucleotide molecules sequenced by the systems and devices herein can originate in single-stranded form as DNA or RNA, or in double-stranded DNA (dsDNA) form (e.g., genomic DNA fragments, PCR and amplification products, etc.). Thus, the single-stranded polynucleotide can be the sense or antisense strand of a polynucleotide duplex. Methods for preparing single-stranded polynucleotide molecules suitable for use in the methods of the present disclosure using standard techniques are known in the art. The exact sequence of the primary polynucleotide molecule is generally not critical to the present disclosure and can be known or unknown. The single-stranded polynucleotide molecule can represent a genomic DNA molecule (e.g., human genomic DNA), including both intron and exon sequences (coding sequences), as well as non-coding regulatory sequences such as promoter and enhancer sequences.
[0020] In certain embodiments, the nucleic acid to be sequenced by use of the present disclosure is immobilized on a substrate (e.g., a substrate in a flow cell or one or more beads on a substrate such as a flow cell). As used herein, the term "immobilized" is intended to encompass direct or indirect, covalent or non-covalent attachment, unless otherwise indicated explicitly or by context. In certain embodiments, covalent attachment may be preferred, but generally, what is required is that the molecule (e.g., nucleic acid) remain immobilized or associated with the support under conditions under which the support is intended to be used, e.g., in applications requiring nucleic acid sequencing.
[0021] The term "solid support" (or "substrate" in certain usages) as used herein refers to any inert substrate or matrix to which nucleic acids may be attached, such as, for example, glass surfaces, plastic surfaces, latex, dextran, polystyrene surfaces, polypropylene surfaces, polyacrylamide gels, gold surfaces, and silicon wafers. In many embodiments, the solid support is a glass surface (e.g., the flat surface of a flow cell channel). In certain embodiments, the solid support may comprise an inert substrate or matrix that has been "functionalized," for example, by applying a layer or coating of an intermediate material that contains reactive groups that allow for covalent attachment to molecules such as polynucleotides. As a non-limiting example, such a substrate may comprise a polyacrylamide hydrogel supported on an inert substrate such as glass. In such embodiments, the molecule (polynucleotide) may be covalently attached directly to the intermediate material (e.g., hydrogel), but the intermediate material may itself be non-covalently attached to the substrate or matrix (e.g., glass substrate). Covalent attachment to a solid support should be interpreted accordingly to encompass this type of configuration.
[0022] As mentioned above, the present disclosure includes novel systems and devices for sequencing nucleic acids. As will be clear to those skilled in the art, reference herein to a particular nucleic acid sequence may also refer to a nucleic acid molecule that includes such a nucleic acid sequence, depending on the context. Sequencing a target fragment means that a chronological reading of bases is established. The bases read do not have to be consecutive, which is preferred, but it is not necessary that all bases on the entire fragment are sequenced during sequencing. Sequencing can be performed using any suitable sequencing technique, where nucleotides or oligonucleotides are added consecutively to a free 3' hydroxyl group, resulting in the synthesis of a polynucleotide chain in the 5' to 3' direction. The nature of the added nucleotide is preferably determined after each nucleotide addition. Sequencing techniques that use sequencing by ligation, where not all consecutive bases are sequenced, and techniques such as massively parallel signature sequencing (MPSS), where bases are removed from a strand rather than added to a strand on a surface, are also amendable for use with the systems and devices of the present disclosure.
[0023] In certain embodiments, the present disclosure discloses sequencing-by-synthesis (SBS), in which four fluorescently labeled modified nucleotides are used to sequence high-density clusters (potentially millions of clusters) of amplified DNA present on the surface of a substrate (e.g., a flow cell). Various additional aspects of SBS procedures and methods that may be utilized with the systems and devices herein are disclosed, for example, in WO 04018497, WO 04018493, and U.S. Pat. No. 7,057,026 (nucleotides), WO 05024010 and WO 06120433 (polymerases), WO 05065814 (surface attachment techniques), and WO 9844151, WO 06064199, and WO 07010251, the contents of each of which are incorporated herein by reference in their entirety.
[0024] In a particular use of the system / device herein, a flow cell containing a nucleic acid sample for sequencing is placed in a suitable flow cell holder. The sample for sequencing can take the form of a single molecule, an amplified single molecule in the form of a cluster, or a bead containing molecules of nucleic acid. The nucleic acid is prepared to contain oligonucleotide primers flanking an unknown target sequence. To initiate the first SBS sequencing cycle, one or more differently labeled nucleotides, and a DNA polymerase, etc., are flowed into / through the flow cell by a fluid flow subsystem (various embodiments of which are described herein). A single nucleotide can be added at a time, or the nucleotides used in the sequencing procedure can be specifically designed to have reversible termination properties, thus allowing each cycle of the sequencing reaction to occur simultaneously in the presence of all four labeled nucleotides (e.g., A, C, T, G). When the four nucleotides are mixed together, the polymerase can select and incorporate the correct base, and each sequence is extended by a single base. In such methods of using the system, the natural competition between all four options results in greater accuracy than if only one nucleotide were present in the reaction mixture (thus, most of the sequence would not be exposed to the correct nucleotide). Sequences in which a particular base is repeated one after the other (e.g., homopolymers) are addressed with high accuracy, as are any other sequences.
[0025] The fluid flow subsystem also flows appropriate reagents to remove blocked 3' ends (if appropriate) and fluorophores from each incorporated base. The substrate can be exposed to either a second round of the four blocked nucleotides, or, optionally, a second round with a different individual nucleotide. Such cycles are then repeated and the sequence of each cluster is read over multiple chemical cycles. Computer aspects of the present disclosure can optionally align sequence data collected from each single molecule, cluster or bead to determine the sequence of longer polymers, etc. Alternatively, image processing and alignment can be performed on separate computers.
[0026] The heating / cooling components of the system regulate the reaction conditions within the flow cell channel and the reagent storage areas / containers (and optionally the camera, optics, and / or other components), while the fluid flow components allow the substrate surface to be exposed to suitable reagents for incorporation (e.g., suitable fluorescently labeled nucleotides to be incorporated) while non-incorporated reagents are washed away. The optionally movable stage on which the flow cell sits allows the flow cell to be properly oriented for laser (or other light) excitation of the substrate, and optionally moved relative to the objective lens to allow reading of different regions of the substrate. Additionally, other components of the system are also optionally movable / adjustable (e.g., camera, objective lens, heaters / coolers, etc.). During laser excitation, images / locations of the fluorescence emitted from the nucleic acids on the substrate are captured by the camera component, thereby recording the identity of the first base for each single molecule, cluster, or bead in the computer component.
[0027] The embodiments described herein may be used in a variety of biological or chemical processes and systems for academic or commercial analysis. More specifically, the embodiments described herein may be used in a variety of processes and systems in which it is desirable to detect an event, characteristic, quality, or property indicative of a desired response. For example, the embodiments described herein include cartridges, biosensors, and components thereof, as well as bioassay systems that operate with the cartridges and biosensors. In certain embodiments, the cartridges and biosensors include a flow cell and one or more sensors, pixels, photodetectors, or photodiodes that are bonded together in a substantially single structure.
[0028] The following detailed description of certain embodiments may be better understood when read in conjunction with the accompanying drawings. To the extent that the figures illustrate diagrams of functional blocks of various embodiments, the functional blocks do not necessarily indicative of a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., a processor or memory) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or random access memory, hard disk, etc.). Similarly, a program may be a stand-alone program, may be incorporated as a subroutine in an operating system, may be a function in an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentalities shown in the figures.
[0029] As used herein, elements or steps described in the singular and followed by the word "a" or "an" should be understood as not excluding a plurality of those elements or steps, unless such exclusion is expressly stated. Furthermore, references to "one embodiment" are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features. Furthermore, unless expressly stated to the contrary, embodiments that "comprise" or "have" or "include" an element or elements having a particular characteristic may include the additional elements, whether or not they have that characteristic.
[0030] As used herein, a "desired reaction" includes a change in at least one of the chemical, electrical, physical, or optical properties (or qualities) of an analyte of interest. In certain embodiments, the desired reaction is a positive binding event (e.g., incorporation of a fluorescently labeled biomolecule into the analyte of interest). More generally, the desired reaction may be a chemical conversion, chemical change, or chemical interaction. The desired reaction may also be a change in an electrical property. For example, the desired reaction may be a change in the concentration of an ion in a solution. Exemplary reactions include, but are not limited to, chemical reactions such as reduction, oxidation, addition, elimination, rearrangement, esterification, amidation, etherification, cyclization, or substitution; binding interactions in which a first chemical binds to a second chemical; dissociation reactions in which two or more chemicals separate from each other; fluorescence, luminescence, bioluminescence, chemiluminescence, and biological reactions such as nucleic acid replication, nucleic acid amplification, nucleic acid hybridization, nucleic acid ligation, phosphorylation, enzyme catalysis, receptor binding, or ligand binding. The desired reaction may also be the addition or removal of a proton, which is detectable, for example, as a change in the pH of the surrounding solution or environment. An additional desired response can be the detection of ion flow across a membrane (e.g., a natural or synthetic bilayer membrane); for example, when ions flow through the membrane, the current is perturbed and this perturbation can be detected.
[0031] In certain embodiments, the desired reaction includes incorporation of a fluorescently labeled molecule into the analyte. The analyte may be an oligonucleotide and the fluorescently labeled molecule may be a nucleotide. The desired reaction may be detected when excitation light is directed to the oligonucleotide with the labeled nucleotide and the fluorophore emits a detectable fluorescent signal. In alternative embodiments, the detected fluorescence is the result of chemiluminescence or bioluminescence. The desired reaction may also increase fluorophore (or Forster) resonance energy transfer (FRET), for example, by bringing a donor fluorophore into close proximity with an acceptor fluorophore, decrease FRET by separating the donor and acceptor fluorophores, increase fluorescence by separating the quencher from the fluorophore, or decrease fluorophore by colocalizing the quencher and fluorophore.
[0032] As used herein, "reaction component" or "reactant" includes any substance that can be used to obtain a desired reaction. For example, reaction components include reagents, enzymes, samples, other biomolecules, and buffers. Reaction components are typically delivered to a reaction site in solution and / or immobilized at the reaction site. A reaction component may directly or indirectly interact with another substance, such as an analyte of interest.
[0033] As used herein, the term "reaction site" is a localized area where a desired reaction can occur. A reaction site may include a support surface of a substrate on which a substance may be immobilized. For example, a reaction site may include a substantially planar surface within a channel of a flow cell having a colony of nucleic acids thereon. Typically, but not always, the nucleic acids in the colonies have the same sequence, e.g., are clonal copies of a single-stranded or double-stranded template. However, in some embodiments, a reaction site may contain only a single nucleic acid molecule, e.g., in single-stranded or double-stranded form. Furthermore, multiple reaction sites may be distributed non-uniformly along the support surface or may be arranged in a predetermined manner (e.g., parallel in a matrix such as a microarray). A reaction site may also include a reaction chamber (or well) that at least partially defines a spatial region or volume configured to compartmentalize a desired reaction.
[0034] This application uses the terms "reaction chamber" and "well" interchangeably. As used herein, the term "reaction chamber" or "well" includes a spatial region in fluid communication with a flow channel. A reaction chamber may be at least partially isolated from the surrounding environment or other spatial regions. For example, multiple reaction chambers may be separated from each other by a shared wall. As a more specific example, a reaction chamber may include a cavity defined by an inner surface of the well and have an opening or aperture such that the cavity is in fluid communication with the flow channel. A biosensor including such a reaction chamber is described in more detail in International Application PCT / US2011 / 057111, filed October 20, 2011, the entirety of which is incorporated herein by reference.
[0035] In some embodiments, the reaction chamber is sized and shaped relative to a solid (including a semi-solid) so that the solid can be fully or partially inserted therein. For example, the reaction chamber can be sized and shaped to accommodate only one capture bead. The capture bead may have clonally amplified DNA or other material thereon. Alternatively, the reaction chamber can be sized and shaped to receive an approximate number of beads or solid substrates. As another example, the reaction chamber can also be filled with a porous gel or material configured to control diffusion or filter fluids that may flow into the reaction chamber.
[0036] In some embodiments, a sensor (e.g., a photodetector, photodiode) is associated with a corresponding pixel area of the sample surface of the biosensor. Thus, a pixel area is a geometric construct that represents an area on the sample surface of the biosensor of one sensor (or pixel). The sensor associated with a pixel area detects luminescence collected from the associated pixel area when a desired reaction occurs at a reaction site or reaction chamber above the associated pixel area. In flat surface embodiments, the pixel areas can overlap. In some cases, multiple sensors can be associated with a single reaction site or a single reaction chamber. In other cases, a single sensor can be associated with a group of reaction sites or a group of reaction chambers.
[0037] As used herein, a "biosensor" includes a structure having multiple reaction sites and / or reaction chambers (or wells). The biosensor may include a solid-state imaging device (e.g., a CCD or CMOS imager) and, optionally, a flow cell attached thereto. The flow cell may include at least one flow channel in fluid communication with the reaction sites and / or reaction chambers. As one particular example, the biosensor is configured to fluidly and electrically couple to a bioassay system. The bioassay system may deliver reactants to the reaction sites and / or reaction chambers according to a predetermined protocol (e.g., sequencing by synthesis) and perform multiple imaging events. For example, the bioassay system may direct solutions to flow along the reaction sites and / or reaction chambers. At least one of the solutions may include four types of nucleotides with the same or different fluorescent labels. The nucleotides may bind to corresponding oligonucleotides located in the reaction sites and / or reaction chambers. The bioassay system may then illuminate the reaction sites and / or reaction chambers using an excitation light source (e.g., a solid-state light source such as a light emitting diode or LED). The excitation light may have a predetermined wavelength or multiple wavelengths including a range of wavelengths. The excited fluorescent label provides a luminescent signal that can be captured by a sensor.
[0038] In alternative embodiments, the biosensor may include electrodes or other types of sensors configured to detect other distinguishable characteristics. For example, the sensor may be configured to detect changes in ion concentration. In another example, the sensor may be configured to detect the flow of ionic current across a membrane.
[0039] As used herein, a "cluster" is a colony of similar or identical molecules or nucleotide sequences or DNA strands. For example, a cluster can be an amplification oligonucleotide or any other group of polynucleotides or polypeptides with the same or similar sequences. In other embodiments, a cluster can be any element or group of elements that occupy a physical area on a sample surface. In embodiments, the cluster is immobilized in a reaction site and / or reaction chamber during the base calling cycle.
[0040] As used herein, the term "immobilized" when used in reference to a biomolecule or biological substance or chemical includes substantially attaching the biomolecule or biological substance or chemical to a surface at a molecular level. For example, the biomolecule or biological substance or chemical may be immobilized to the surface of a substrate material using adsorption techniques including non-covalent bonding (e.g., electrostatic forces, van der Waals, and hydrophobic interfacial dehydration), as well as covalent bonding techniques in which a functional group or linker facilitates attachment of the biomolecule to the surface. Immobilizing the biomolecule or biological substance or chemical to the surface of a substrate material may be based on the properties of the substrate surface, the liquid medium carrying the biomolecule or biological substance or chemical, and the properties of the biomolecule or biological substance or chemical itself. In some cases, the substrate surface may be functionalized (e.g., chemically or physically modified) to facilitate immobilization of the biomolecule (or biological substance or chemical) to the surface. The substrate surface may first be modified to have functional groups attached to the surface. The functional groups may then bind to the biomolecule or biological substance or chemical to immobilize them thereon. Substances can be immobilized on a surface via a gel, for example, as described in US Patent Application Publication No. 2011 / 0059865(A1), which is incorporated herein by reference.
[0041] In some embodiments, nucleic acids can be attached to a surface and amplified using bridge amplification. Useful bridge amplification methods are described, for example, in U.S. Patent No. 5,641,658, International Publication No. WO 2007 / 010251, U.S. Patent No. 6,090,592, U.S. Patent Application Publication No. 2002 / 0055100 (A1), U.S. Patent No. 7,115,400, U.S. Patent Application Publication No. 2004 / 0096853 (A1), U.S. Patent Application Publication No. 2004 / 0002090 (A1), U.S. Patent Application Publication No. 2007 / 0128624 (A1), and U.S. Patent Application Publication No. 2008 / 0009420 (A1), each of which is incorporated herein in its entirety. Another useful method for amplifying nucleic acids on a surface is Rolling Circle Amplification (RCA), for example, using the methods described in more detail below. In some embodiments, the nucleic acid can be attached to a surface and amplified using one or more primer pairs. For example, one of the primers can be in solution and the other primer can be immobilized (e.g., 5'-attached) on the surface. As an example, a nucleic acid molecule can hybridize to one of the primers on the surface, followed by extension of the immobilized primer to create a first copy of the nucleic acid. The primer in solution then hybridizes to the first copy of the nucleic acid, which can be extended using the first copy of the nucleic acid as a template. Optionally, after the first copy of the nucleic acid is created, the original nucleic acid molecule can hybridize to a second immobilized primer on the surface and be extended simultaneously or after the primer in solution is extended. In any embodiment, repeated rounds of extension (e.g., amplification) using the immobilized primer and the primer in solution provide multiple copies of the nucleic acid.
[0042] In certain embodiments, the assay protocols performed by the systems and methods described herein include the use of naturally occurring nucleotides and enzymes configured to interact with the naturally occurring nucleotides. Naturally occurring nucleotides include, for example, ribonucleotides (RNA) or deoxyribonucleotides (DNA). Naturally occurring nucleotides may be in monophosphate, diphosphate, or triphosphate form and may have a base selected from adenine (A), thymine (T), uracil (U), guanine (G), or cytosine (C). However, it will be understood that non-naturally occurring nucleotides, modified nucleotides, or analogs of the above nucleotides may be used. Some examples of useful non-naturally occurring nucleotides are described below with respect to reversible terminator-based sequencing by synthetic methods.
[0043] In embodiments that include a reaction chamber, an article or solid material (including a semi-solid material) may be disposed within the reaction chamber. When disposed, the article or solid may be physically held or immobilized within the reaction chamber via interference fit, adhesion, or confinement. Exemplary articles or solids that may be disposed within the reaction chamber include polymer beads, pellets, agarose gels, powders, quantum dots, or other solids that may be compressed and / or held within the reaction chamber. In certain embodiments, nucleic acid superstructures such as DNA balls may be disposed within or at the reaction chamber, for example, by attaching to the inner surface of the reaction chamber or by residing in a liquid within the reaction chamber. DNA balls or other nucleic acid superstructures may be preformed and then disposed within or at the reaction chamber. Alternatively, DNA balls may be synthesized in situ in the reaction chamber. DNA balls may be synthesized by rolling circle amplification to create concatemers of specific nucleic acid sequences, and the concatemers may be treated under conditions to form relatively compact balls. DNA balls and methods for their synthesis are described, for example, in U.S. Patent Application Publication No. 2008 / 0242560(A1) or U.S. Patent Application Publication No. 2008 / 0234136(A1), each of which is incorporated herein in its entirety. The material held or disposed within the reaction chamber may be in a solid, liquid, or gaseous state.
[0044] As used herein, a "base call" identifies a nucleotide base in a nucleic acid sequence. Base calling refers to the process of determining the base call (A, C, G, T) of every cluster in a particular cycle. As an example, base calling can be performed utilizing the four-channel, two-channel, or one-channel methods and systems described in the incorporated materials of US Patent Application Publication No. 2013 / 0079232. In certain embodiments, a base call cycle is referred to as a "sampling event." In a one-dye and two-channel sequencing protocol, a sampling event includes two illumination stages in chronological order such that a pixel signal occurs at each stage. The first illumination stage induces illumination from a given cluster that indicates nucleotide bases A and T in an AT pixel signal, and the second illumination stage induces illumination from a given cluster that indicates nucleotide bases C and T in a CT pixel signal.
[0045] The disclosed technology, for example the disclosed base callers, can be implemented on processors such as Central Processing Units (CPUs), Graphics Processing Units (GPUs), Field Programmable Gate Arrays (FPGAs), Coarse-Grained Reconfigurable Architectures (CGRAs), Application Specific Integrated Circuits (ASICs), Application Specific Instruction-set Processors (ASIPs), and Digital Signal Processors (DSPs).
[0046] Biosensors FIG. 1 shows a cross-sectional view of a biosensor 100 that can be used in various embodiments. The biosensor 100 has pixel areas 106', 108', 110', 112', and 114', each of which can retain two or more clusters (e.g., two clusters per pixel area) during a base call cycle. As shown, the biosensor 100 can include a flow cell 102 mounted on a sampling device 104. In the illustrated embodiment, the flow cell 102 is fixed directly to the sampling device 104. However, in alternative embodiments, the flow cell 102 can be removably coupled to the sampling device 104. The sampling device 104 has a sample surface 134 that can be functionalized (e.g., chemically or physically modified in a manner suitable for causing a desired reaction). For example, the sample surface 134 may be functionalized and may include multiple pixel regions 106', 108', 110', 112', and 114' each capable of holding two or more clusters during a base calling cycle (e.g., having corresponding cluster pairs 106A, 106B, cluster pairs 108A, 108B, cluster pairs 110A, 110B, cluster pairs 112A, 112B, and cluster pairs 114A, 114B immobilized thereon). Each pixel region is associated with a corresponding sensor (or pixel or photodiode) 106, 108, 110, 112, and 114, such that light received by the pixel region is captured by the corresponding sensor. The pixel region 106' may also be associated with a corresponding reaction site 106" on the reaction surface 134 that holds the cluster pairs, such that light emitted from the reaction site 106" is received by the pixel region 106' and captured by the corresponding sensor 106. As a result of this sensing structure, if two or more clusters are present in a particular sensor pixel area during a base call cycle (e.g., each having a corresponding cluster pair), the pixel signal in that base call cycle carries information based on all of the two or more clusters.As a result, the signal processing described herein is used to distinguish between clusters where there are more clusters than pixel signals at a given sampling event of a particular base call cycle.
[0047] In the illustrated embodiment, the flow cell 102 includes sidewalls 138, 125 and a flow cover 136 supported by the sidewalls 138, 125. The sidewalls 138, 125 are coupled to the sample surface 134 and extend between the flow cover 136 and the sidewalls 138, 125. In some embodiments, the sidewalls 138, 125 are formed from a curable adhesive layer that bonds the flow cover 136 to the sampling device 104.
[0048] The side walls 138, 125 are sized and shaped such that a flow channel 144 exists between the flow cover 136 and the sampling device 104. The flow cover 136 may include a material that is transparent to excitation light 101 propagating from outside the biosensor 100 to the flow channel 144. In one example, the excitation light 101 approaches the flow cover 136 at a non-orthogonal angle.
[0049] As also shown, the flow cover 136 may include inlet and outlet ports 142, 146 configured to fluidly engage other ports (not shown). For example, these other ports may be from the cartridge or the workstation. The flow channel 144 is sized and shaped to direct fluid along the sample surface 134. The height H of the flow channel 144 is 1 / 2 the height of the flow channel 144. 1 and other dimensions may be configured to maintain a substantially uniform flow of fluid along the sample surface 134. The dimensions of the flow channel 144 may also be configured to control bubble formation.
[0050] By way of example, the flow cover 136 (or flow cell 102) may comprise a transparent material such as glass or plastic. The flow cover 136 may comprise a substantially rectangular block having a planar outer surface and a planar inner surface that defines the flow channel 144. The block may be attached onto the side walls 138, 125. Alternatively, the flow cell 102 may be etched to define the flow cover 136 and the side walls 138, 125. For example, a recess may be etched into the transparent material. When the etched material is attached to the sampling device 104, the recess may become the flow channel 144.
[0051] The sampling device 104 may be similar to an integrated circuit comprising, for example, multiple stacked substrate layers 120-126. The substrate layers 120-126 may include a base substrate 120, a solid-state imager 122 (e.g., a CMOS image sensor), a filter or light management layer 124, and a passivation layer 126. Note that the above is merely exemplary and other embodiments may include fewer or additional layers. Additionally, each of the substrate layers 120-126 may include multiple sublayers. The sampling device 104 may be fabricated using processes similar to those used in fabricating integrated circuits such as CMOS image sensors and CCDs. For example, the substrate layers 120-126 or portions thereof may be grown, deposited, etched, etc. to form the sampling device 104.
[0052] The passivation layer 126 is configured to shield the filter layer 124 from the fluid environment of the flow channel 144. In some cases, the passivation layer 126 is also configured to provide a solid surface (i.e., the sample surface 134) onto which biomolecules or other analytes of interest can be immobilized. For example, each of the reaction sites can include a cluster of biomolecules immobilized on the sample surface 134. Thus, the passivation layer 126 can be formed from a material that allows the reaction sites to be immobilized thereon. The passivation layer 126 can also include a material that is at least transparent to the desired fluorescence. By way of example, the passivation layer 126 can be formed from a material that is at least transparent to the desired fluorescence. 2 N 4 ) and / or silica (SiO 2 ). However, other suitable materials may be used. In the illustrated embodiment, the passivation layer 126 may be substantially planar. However, in alternative embodiments, the passivation layer 126 may include recesses, such as pits, wells, trenches, etc. In the illustrated embodiment, the passivation layer 126 has a thickness of about 150-200 nm, and more specifically, about 170 nm.
[0053] The filter layer 124 may include various features that affect the transmission of light. In some embodiments, the filter layer 124 may perform multiple functions. For example, the filter layer 124 may be configured to (a) filter unwanted light signals, such as light signals from an excitation light source, (b) direct luminescence signals from the reaction sites toward corresponding sensors 106, 108, 110, 112, and 114 configured to detect luminescence signals from the reaction sites, or (c) block or prevent detection of unwanted luminescence signals from adjacent reaction sites. Thus, the filter layer 124 may also be referred to as a light management layer. In the illustrated embodiment, the filter layer 124 has a thickness of about 1-5 μm, more specifically about 2-4 μm. In alternative embodiments, the filter layer 124 may include an array of microlenses or other optical components. Each of the microlenses may be configured to direct the luminescence signal from an associated reaction site to a sensor.
[0054] In some embodiments, the solid-state imager 122 and the base substrate 120 may be provided together as a previously constructed solid-state imaging device (e.g., a CMOS chip). For example, the base substrate 120 may be a wafer of silicon, and the solid-state imager 122 may be mounted thereon. The solid-state imager 122 includes a layer of semiconductor material (e.g., silicon) and the sensors 106, 108, 110, 112, and 114. In the illustrated embodiment, the sensors are photodiodes configured to detect light. In other embodiments, the sensors include photodetectors. The solid-state imager 122 may be fabricated as a single chip via a CMOS-based manufacturing process.
[0055] The solid-state imager 122 may include a high density array of sensors 106, 108, 110, 112, and 114 configured to detect activity indicative of a desired response from within or along the flow channel 144. In some embodiments, each sensor has a size of about 1-2 micrometers squared (μm 2 ) The array may include 500,000 sensors, 5 million sensors, 10 million sensors, or even 120 million sensors. Sensors 106, 108, 110, 112, and 114 may be configured to detect predetermined wavelengths of light that are indicative of a desired response.
[0056] In some embodiments, the sampling device 104 includes a microcircuit device, such as the microcircuit device described in U.S. Patent No. 7,595,882, which is incorporated herein by reference in its entirety. More specifically, the sampling device 104 may include an integrated circuit having a planar array of sensors 106, 108, 110, 112, and 114. The circuitry formed within the sampling device 104 may be configured for at least one of signal amplification, digitization, storage, and processing. The circuitry may collect and analyze the detected fluorescence and generate a pixel signal (or detection signal) for communicating the detection data to a signal processor. The circuitry may also perform additional analog and / or digital signal processing in the sampling device 104. The sampling device 104 may include conductive vias 130 that perform signal routing (e.g., transmit the pixel signal to a signal processor). The pixel signal may also be transmitted through electrical contacts 132 of the sampling device 104.
[0057] The sampling device 104 is discussed in more detail with respect to U.S. Non-Provisional Patent Application No. 16 / 874,599, entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing," filed May 14, 2020, which is incorporated by reference as if fully set forth herein. The sampling device 104 is not limited to the above configurations or uses as described above. In alternative embodiments, the sampling device 104 may take other forms. For example, the sampling device 104 may comprise a CCD device, such as a CCD camera, coupled to a flow cell or moved to interface with a flow cell having reaction sites therein.
[0058] Figure 2 shows one embodiment of a flow cell 200 that includes clusters within its tiles. Flow cell 200 corresponds to flow cell 102 of Figure 1, e.g., without flow cover 136. Additionally, the depiction of flow cell 200 is symbolic in nature, and flow cell 200 symbolically shows the various lanes and tiles therein without showing the various other components therein. Figure 2 shows a top view of flow cell 200.
[0059] In one embodiment, flow cell 200 is divided or split into multiple lanes, such as lanes 202a, 202b, ..., 202P, i.e., P lanes. In the example of Figure 2, flow cell 200 is shown as including eight lanes, i.e., P=8 in this example, although the number of lanes in a flow cell is implementation specific.
[0060] In one embodiment, each lane 202 is further divided into non-overlapping regions called "tiles" 212. For example, Figure 2 shows an expanded view of a section 208 of an exemplary lane. Section 208 is shown to include multiple tiles 212.
[0061] In one example, each lane 202 includes one or more tile columns. For example, in FIG. 2, each lane 202 includes two corresponding tile columns 212, as shown in enlarged section 208. The number of tiles in each tile column in each lane is implementation specific, and in one example, there may be 50 tiles, 60 tiles, 100 tiles, or another suitable number of tiles in each tile column in each lane.
[0062] Each tile contains a corresponding number of clusters. During the sequencing procedure, the clusters on the tile and their surrounding background are imaged. For example, Figure 2 shows an example cluster 216 in an example tile.
[0063] FIG. 3 shows an exemplary Illumina GA-IIx™ flow cell with eight lanes, and also shows a zoom-in of one tile and its clusters and their surrounding background. For example, there are 100 tiles per lane in an Illumina Genome Analyzer II, and 68 tiles per lane in an Illumina HiSeq2000. A tile 212 holds hundreds of thousands to millions of clusters. In FIG. 3, an image generated from a tile with clusters shown as bright spots is shown at 308 (e.g., 308 is a magnified image view of a tile), and an exemplary cluster 304 is labeled. Cluster 304 contains about a thousand identical copies of a template molecule, but the clusters differ in size and shape. Clusters are grown from template molecules by bridge amplification of the input library prior to sequencing operations. The purpose of amplification and cluster growth is to increase the intensity of the emitted signal, since imaging devices cannot reliably sense single fluorophores. However, the physical distance between the DNA fragments within the cluster 304 is small, so the imaging device perceives the cluster of fragments as a single spot 304 .
[0064] Clusters and tiles are discussed in further detail in U.S. Non-provisional Patent Application No. 16 / 825,987, filed March 20, 2020, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing."
[0065] FIG. 4 is a simplified block diagram of a system for analysis of sensor data from a sequencing system, such as base calling sensor output (see, e.g., FIG. 1). In the example of FIG. 4, the system includes a sequencing machine 400 and a configurable processor 450. The configurable processor 450 can execute a neural network based base caller and / or a non-neural network based base caller (discussed in more detail herein) in coordination with a runtime program executed by a host processor, such as a central processing unit (CPU) 402. The sequencing machine 400 includes a base calling sensor (e.g., as discussed with respect to FIGS. 1-3) and a flow cell 401. The flow cell can include one or more tiles in which clusters of genetic material are exposed to an array of analyte flows that are used to induce reactions in the clusters to identify bases in the genetic material, as discussed with respect to FIGS. 1-3. A sensor senses the reaction of each cycle of the array in each tile of the flow cell to provide tile data. An example of this technique is described in more detail below. Genetic sequencing is a data-intensive operation that converts base call sensor data into a sequence of base calls for each group of genetic material sensed during the base calling operation.
[0066] The system in this example includes a CPU 402 that executes a runtime program that coordinates the base calling operation, and a memory 403 that stores the sequence of the array of tile data, the base call reads produced by the base calling operation, and other information used in the base calling operation. In this figure, the system also includes a configuration file (or files), e.g., an FPGA bit file, and a memory 404 that stores model parameters of a neural network used to configure and reconfigure the configurable processor 450 and to run the neural network. Sequencing machine 400 can include programs for configuring the configurable processor, and in some embodiments can include a reconfigurable processor that runs the neural network.
[0067] The sequencing machine 400 is coupled to the configurable processor 450 by a bus 405. The bus 405 can be implemented using a high-throughput technology, such as a bus technology compatible with the PCIe standard (Peripheral Component Interconnect Express) currently maintained and developed by the PCI Special Interest Group (PCI-SIG). Also, in this embodiment, the memory 460 is coupled to the configurable processor 450 by a bus 461. The memory 460 can be an on-board memory disposed on a circuit board having the configurable processor 450. The memory 460 is used for fast access by the configurable processor 450 of working data used in base calling operations. The bus 461 can also be implemented using a high-throughput technology, such as a bus technology compatible with the PCIe standard. The memory 460 can store genomic data, for example, variant call format (VCF) files.
[0068] Configurable processors, including Field Programmable Gate Arrays (FPGAs), Coarse Grained Reconfigurable Arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than can be achieved using a general-purpose processor executing a computer program. Configuring a configurable processor involves compiling a functional description to produce a configuration file, sometimes called a bitstream or bitfile, and distributing the configuration file to the configurable elements on the processor.
[0069] The configuration file configures the circuit to set data flow patterns, including the use of distributed memory and other on-chip memory resources, look-up table contents, operation of configurable logic blocks, and configurable execution units such as configurable interconnects and other elements of the configurable array. A configuration file is reconfigurable if it can be changed in the field by modifying a loaded configuration file. For example, the configuration file may be stored in a volatile SRAM element, in a non-volatile read-write memory element, or distributed among an array of configurable elements on a configurable or reconfigurable processor. A variety of commercially available configurable processors are suitable for use in base call operations as described herein. In some embodiments, the host CPU may be implemented on the same integrated circuit as the configurable processor.
[0070] The embodiments described herein use a configurable processor 450 to implement a multi-cycle neural network. The configuration file of the configurable processor may be implemented by specifying the logic functions to be performed using a high-level description language (HDL) or register transfer level (RTL) language specification. The specification may be compiled using resources designed by a selected configurable processor to generate a configuration file. The same or similar specifications may be compiled to generate a design for an application specific integrated circuit, which may not be a configurable processor.
[0071] Thus, alternatives to the configurable processor in all of the embodiments described herein include a configured processor including an application specific ASIC or dedicated integrated circuit or set of integrated circuits, or a system-on-chip SOC device, configured to perform the neural network based base calling operations described herein.
[0072] In general, the configurable and configured processors described herein that are configured to perform neural network operations are referred to herein as neural network processors. In another example, the configurable and configured processors described herein that are configured to perform non-neural network based base caller operations are referred to herein as non-neural network processors. In general, the configurable and configured processors can be used to implement one or both of a neural network based base caller and a non-neural network based base caller, as discussed later in this specification.
[0073] Configurable processor 450 is configured, in this embodiment, by a configuration file loaded using a program executed by CPU 402, or by other sources that configures an array of configurable elements on configurable processor 454 to perform the base calling function. In this embodiment, the configuration includes data flow logic 451 coupled to buses 405 and 461, which performs the function of distributing data and control parameters among the elements used in the base calling operation.
[0074] Configurable processor 450 is also configured with base call execution logic 452 to execute the multi-cycle neural network. Logic 452 includes a number of multi-cycle execution clusters (e.g., 453), which in this example include multi-cycle cluster 1 through multi-cycle cluster X. The number of multi-cycle clusters can be selected according to tradeoffs involving the desired throughput of operation and available resources on the configurable processor.
[0075] The multi-cycle clusters are coupled to data flow logic 451 by data flow paths 454, implemented using configurable interconnect and memory resources on a configurable processor, and by control paths 455, implemented using, for example, configurable interconnect and memory resources on a configurable processor, that provide control signals indicating available clusters, readiness to provide input units to available clusters for execution of neural network operations, readiness to provide trained parameters for the neural network, readiness to provide output patches of base call classification data, and other control data used in the execution of the neural network.
[0076] The configurable processor is configured to use the trained parameters to perform a multi-cycle neural network operation to generate classification data for a sensing cycle of the base calling operation. The neural network operation is performed to generate classification data for a subject sensing cycle of the base calling operation. The neural network operation operates on an array including a number N of arrays of tile data from each sensing cycle of the N sensing cycles, which in the examples described herein provide sensor data for different base calling operations for one base position per operation in the time series. Optionally, some of the N sensing cycles can be out of the array as needed according to the particular neural network model being performed. The number N can be any number greater than 1. In some examples described herein, the sensing cycle of the N sensing cycles represents a set of sensing cycles for at least one sensing cycle preceding the subject sensing cycle and at least one sensing cycle following the subject sensing cycle in the time series. Examples described herein are in which the number N is an integer greater than or equal to 5.
[0077] The data flow logic 451 is configured to use an input unit for a given operation that includes tile data of an array of N spatially aligned patches to move the tile data and at least some trained parameters of the model from the memory 460 to the configurable processor for operation of the neural network. The input unit can be moved by direct memory access operations in a single DMA operation, or in smaller units that move during available time slots in coordination with the execution of the deployed neural network.
[0078] The tile data of the sensing cycle described herein can include an array of sensor data having one or more features. For example, the sensor data can include two images that are analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The tile data can also include metadata about the images and the sensor. For example, in an embodiment of a base calling operation, the tile data can include information about the alignment of the images with the clusters, such as distance from center information indicating the distance of each pixel in the array of sensor data from the center of the group of genetic material on the tile.
[0079] During execution of the multi-cycle neural network as described below, the tile data may also include data produced during execution of the multi-cycle neural network, referred to as intermediate data, which may be reused rather than recomputed during execution of the multi-cycle neural network. For example, during execution of the multi-cycle neural network, the data flow logic may write the intermediate data to memory 460 in place of the sensor data for a given patch of the array of tile data. Such embodiments are described in more detail below.
[0080] As shown, a system for analysis of base calling sensor output is described that includes a memory (e.g., 460) accessible by a runtime program that stores tile data including sensor data for tiles from sensing cycles of a base calling operation. The system also includes a neural network processor, such as a configurable processor 450, having access to the memory. The neural network processor is configured to perform operations of the neural network using trained parameters to generate classification data for the sensing cycles. As described herein, the operations of the neural network operate on an arrangement of N arrays of tile data from each sensing cycle of the N sensing cycles that comprise a subject cycle to generate classification data for the subject cycle. Data flow logic 451 is provided to move the tile data and trained parameters from the memory to the neural network processor for execution of the neural network using input units including data for the spatially aligned patches of the N arrays from each sensing cycle of the N sensing cycles.
[0081] Also described is a system in which the neural network processor has access to a memory and includes a plurality of execution clusters, and execution logic clusters in the plurality of execution clusters are configured to execute the neural network. The data flow logic includes access to the memory and executes a cluster in the plurality of execution clusters to provide an input unit of tile data to an available execution cluster in the plurality of execution clusters, the input unit including a number N of spatially aligned patches of the array of tile data from a respective sensing cycle, and causing the execution cluster to apply the N spatially aligned patches to the neural network, including a subject sensing cycle, to produce an output patch of classification data for the spatially aligned patches of the subject sensing cycle, where N is greater than 1.
[0082] FIG. 5 is a simplified diagram illustrating aspects of a base calling operation, including the functionality of a runtime program executed by a host processor. In this diagram, the output of an image sensor from a flow cell (such as that shown in FIG. 1 and FIG. 2) is provided on line 500 to an image processing thread 501, which can perform processes on the image such as resampling, alignment and positioning of the array of sensor data for individual tiles, which can be used by a process to calculate a tile cluster mask for each tile in the flow cell, which can be used by a process to identify pixels in the array of sensor data that correspond to clusters of genetic material on the corresponding tile of the flow cell. To calculate the cluster mask, one exemplary algorithm is based on a process that detects unreliable clusters in early sequencing cycles using a metric derived from the softmax output, and then data from those wells / clusters is discarded and no output data for those clusters is produced. For example, the process can identify clusters that are highly reliable during the first N1 (e.g., 25) base calls and reject other clusters. Rejected clusters can be polyclonal or very weak intensity or unclear according to the criteria. This procedure can be executed on the host CPU. Alternate implementations could potentially use this information to identify necessary clusters that are to be returned to the CPU, thereby limiting the storage required for intermediate data.
[0083] The output of image processing thread 501 is provided on line 502 to dispatch logic 510 in the CPU, which routes the array of tile data according to the state of the base calling operation either on high speed bus 503 to a data cache 504 or on high speed bus 505 to hardware 520, such as the configurable processor of Figure 4. Hardware 520 may be a multi-cluster neural network processor for executing a neural network based base caller, or may be hardware for executing a non-neural based base caller, as discussed later in this specification.
[0084] Hardware 520 returns classification data (e.g., output by the neural network based caller and / or the non-neural network based caller) to dispatch logic 510, which passes the information to data cache 504 or on line 511 to thread 502, which can use the classification data to perform base calling and quality score calculations and place the data in a standard format for base called reads. The output of thread 502, which performs base calling and quality score calculations, is provided on line 512 to thread 503, which aggregates the base called reads, performs other operations such as data compression, and writes the resulting base calling output to a specified destination for consumption by the customer.
[0085] In some embodiments, the host may include a thread (not shown) that performs final processing of the output of the hardware 520 supporting the neural network. For example, the hardware 520 may provide an output of classification data from a final layer of a multi-cluster neural network. The host processor may perform output activation functions, such as a softmax function, over the classification data to populate the data used by the base calling and quality score thread 502. The host processor may also perform input operations (not shown), such as resampling, batch normalization, or other adjustments of the tile data before inputting it to the hardware 520.
[0086] FIG. 6 is a simplified diagram of a configuration of a configurable processor, such as the configurable processor of FIG. 4. In FIG. 6, the configurable processor comprises an FPGA with multiple high-speed PCIe interfaces. The FPGA is configured with a wrapper 600 including the data flow logic described with reference to FIG. 1. The wrapper 600 manages interfacing and coordinating with a runtime program in the CPU via a CPU communication link 609, and manages communication with an on-board DRAM 602 (e.g., memory 460) via a DRAM communication link 610. The data flow logic in the wrapper 600 provides patch data obtained by traversing an array of tile data on the on-board DRAM 602 to a cluster 601 for a number N of cycles, and obtains process data 615 from the cluster 601 and delivers it to the on-board DRAM 602. The wrapper 600 also manages the transfer of data between the on-board DRAM 602 and the host memory for both the input array of tile data and the output patch of classification data. The wrapper transfers the patch data on line 613 to the assigned cluster 601. The wrapper provides cluster 601 with trained parameters such as weights and biases obtained from on-board DRAM 602 on line 612. The wrapper provides cluster 601 with configuration and control data provided from or generated in response to a runtime program on the host over CPU communication link 609 on line 611. The cluster can also provide status signals to wrapper 600 on line 616 that are used in conjunction with control signals from the host to manage traversal of an array of tile data to provide spatially aligned patch data, and to perform operations for multi-cycle neural network and / or non-neural network based base calling for base calling on the patch data using resources of cluster 601.
[0087] As described above, there may be multiple clusters on a single configurable processor managed by wrapper 600 configured to run on corresponding ones of the multiple patches of tile data. Each cluster may be configured to provide classification data for base calls in a subject sensing cycle using the tile data of multiple sensing cycles as described herein.
[0088] In an example system, model data including kernel data such as filter weights and biases can be sent from the host CPU to the configurable processor, so that the model can be updated as a function of cycle number. The base calling operation can include, in a representative example, on the order of hundreds of sensing cycles. The base calling operation can include paired end reading in some embodiments. For example, the model trained parameters may be updated every 20 cycles (or other number of cycles) or according to an update pattern implemented for a particular system. In some embodiments where the sequence for a given string in a genetic cluster on a tile includes a first portion extending down (or up) from a first end of the string and a second portion extending up (or down) from a second end of the string, the trained parameters can be updated at the transition from the first portion to the second portion.
[0089] In some implementations, image data for multiple cycles of sensor data for a tile may be sent from the CPU to the wrapper 600. The wrapper 600 may optionally perform some pre-processing and conversion of the sensor data and write the information to the on-board DRAM 602. The input tile data for each sensing cycle may include an array of sensor data including 4000x3000 pixels or more per sensing cycle per tile, with two features representing the colors of the two images of the tile, and including one or two bytes per pixel. In an embodiment where the number N is three sensing cycles used in each operation of the multi-cycle neural network, the array of tile data for each operation of the multi-cycle neural network may consume several hundred megabytes per number. In some embodiments of the system, the tile data also includes an array of DFC data stored once per tile, or other types of metadata about the sensor data and the tile.
[0090] In operation, if a multi-cycle cluster is available, the wrapper assigns the patch to the cluster. The wrapper fetches the next patch of tile data for the cross section of the tile and sends it to the assigned cluster along with appropriate control and configuration information. The cluster can be configured with enough memory on the configurable processor to have enough memory to hold the patch of data, including the patch, from multiple cycles in some systems being processed in place, and in various embodiments are processed using a ping-pong buffer technique or a raster scan technique.
[0091] When the assigned cluster completes its operation of the neural network of the current patch and produces an output patch, it signals the wrapper. The wrapper will either read the output patch from the assigned cluster or the assigned cluster will push the data to the wrapper. The wrapper will then assemble the output patch for the processed tile in DRAM 602. Once the processing of the entire tile is complete and the output patch of data is transferred to DRAM, the wrapper will send the processed output array back to the host / CPU in a specific format. In some embodiments, the on-board DRAM 602 is managed by memory management logic in the wrapper 600. The runtime program can control the sequencing operations to complete the analysis of the array of all tile data for every cycle running in a continuous flow to provide real-time analysis.
[0092] Sharpening mask generation FIG. 7 illustrates a system 700 for generating and / or updating a sharpening mask 706 by training a base caller 704. The system 700 includes a trainer 714 that trains the base caller 704, for example, using least squares estimation. As used herein, a "sharpening mask" maximizes the signal-to-noise ratio of a signal that is corrupted by noise. A sharpening mask may be a value or function that is applied to data to modify the data in a desired manner. For example, the data may be modified to increase its accuracy, relevance, or applicability with respect to a particular situation. A sharpening mask may be applied to data by any of a variety of mathematical operations, including, but not limited to, addition, subtraction, division, multiplication, or combinations thereof. A sharpening mask may be a mathematical formula, a logical function, a computer-implemented algorithm, or the like. The data may be image data, electrical data, or combinations thereof. In one implementation, the sharpening mask is an equalizer (e.g., a spatial equalizer). The equalizer can be trained (e.g., using least squares estimation, adaptive equalization algorithms) to improve and / or maximize the signal-to-noise ratio of the cluster intensity data in the sequencing image. In some implementations, the equalizer includes coefficients that are learned from training. In one implementation of the convolution operation, the training produces equalizer coefficients that are configured to mix / combine the intensity values of pixels representing the intensity radiation from the base-called target cluster and the intensity radiation from one or more adjacent clusters to maximize the signal-to-noise ratio. The signal that is maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that is minimized in the signal-to-noise ratio is the intensity radiation from the adjacent clusters, i.e., spatial crosstalk, plus some random noise (e.g., to account for background intensity radiation). The equalizer coefficients are used as weights, and the mixing / combining includes performing an element-wise multiplication between the equalizer coefficients and the intensity values of the pixels to calculate a weighted sum of the intensity values of the pixels, i.e., the convolution operation. Furthermore, if the image data spans multiple color channels, a set of equalizer coefficients is generated for each color channel (eg, one channel, three channels, four channels, etc.).
[0093] The sequencing image 702 is generated during a sequencing operation performed by a sequencing instrument, such as a sequencing instrument including the biosensor 100 discussed with respect to FIG. 1. Examples of such sequencing instruments include Illumina's iSeq, HiSeqX, HiSeq3000, HiSeq4000, HiSeq2500, NovaSeq6000, NextSeq550, NextSeq1000, NextSeq2000, NextSeqDx, MiSeq, and MiSeqDx. In one embodiment, the Illumina sequencer uses cyclic reversible termination (CRT) chemistry for base calling. This process relies on extending a nascent strand complementary to a template strand with fluorescently labeled nucleotides while tracking the emission signal of each newly added nucleotide. The fluorescently labeled nucleotides have a 3' removable block that anchors the fluorophore signal of the nucleotide type.
[0094] Sequencing is performed in repeated cycles, each of which includes three steps: (a) extension of the emerging strand by adding fluorescently labeled nucleotides; (b) excitation of fluorophores using one or more lasers in the sequencing instrument's optical system and generation of a sequencing image by imaging through different filters in the optical system; and (c) cleavage of the fluorophores and removal of the 3' block in preparation for the next sequencing cycle. The capture and imaging cycles are repeated for a specified number of sequencing cycles, which defines the read length. Using this approach, each cycle locates a new position along the template strand.
[0095] The enormous power of Illumina sequencers comes from their ability to simultaneously perform and sense the CRT reactions of millions or even billions of analytes (e.g., clusters). A cluster contains about 1000 identical copies of a template strand, but the size and shape of the clusters vary. Clusters are grown from the template strands by bridge or exclusion amplification of the input library before a sequencing run. The purpose of amplification and cluster extension is to increase the intensity of the emitted signal, since imaging devices cannot reliably sense the fluorophore signal of a single strand. However, the physical distance of the strands within a cluster is small, so the imaging device perceives the cluster of strands as a single spot.
[0096] Sequencing occurs in a flow cell, a small glass slide that holds the input strands (see, e.g., FIG. 2). The flow cell is connected to an optical system that includes microscope imaging, an excitation laser, and a fluorescence filter. The flow cell contains multiple chambers, called lanes. The lanes are physically separated from one another and can contain different tagged sequencing libraries that are distinguishable without cross-contamination of samples. In some embodiments, the flow cell includes a patterned surface. "Patterned surface" refers to an arrangement of different regions in or on an exposed layer of a solid support. For example, one or more of the regions can be features in which one or more amplification primers are present. The features can be separated by interstitial regions in which no amplification primers are present. In some embodiments, the pattern can be an xy format of features in rows and columns. In some embodiments, the pattern can be a repeating sequence of features and / or interstitial regions. In some embodiments, the pattern can be a random arrangement of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Pat. No. 8,778,849, U.S. Pat. No. 9,079,148, U.S. Pat. No. 8,778,848, and U.S. Patent Application Publication No. 2014 / 0243224, each of which is incorporated herein by reference.
[0097] In some embodiments, the flow cell comprises an array of wells or depressions in a surface, which can be fabricated as commonly known in the art using a variety of techniques, including but not limited to photolithography, stamping techniques, molding techniques, and microetching techniques. As will be understood in the art, the technique used will depend on the composition and shape of the array substrate.
[0098] The features within the patterned surface may be wells in an array of wells (e.g., microwells or nanowells) on glass, silicon, plastic, or other suitable solid support with a patterned covalently attached gel, such as poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication Nos. 2013 / 184796, WO 2016 / 066586, and 2015-002813, each of which is incorporated herein by reference in its entirety). This process creates a gel pad used for sequencing, which may be stable over a large number of cycles of sequencing operations. Covalently attaching a polymer to the wells is useful to maintain the gel in the structured features throughout the life of the structured substrate during various applications. However, in many embodiments, the gel does not need to be covalently attached to the wells. For example, in some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Pat. No. 8,563,477, which is incorporated herein by reference in its entirety) that is not covalently bonded to any portion of the structured substrate can be used as the gel material.
[0099] In certain other embodiments, a structured substrate can be made by patterning a solid support material with wells (e.g., microwells or nanocells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or chemically modified variants thereof), such as an azido version of SFA (azido-SFA), and polishing the gel-coated support, for example by chemical or mechanical polishing, thereby retaining the gel in the wells but removing or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the wells. A primer nucleic acid can be attached to the gel material. A solution of target nucleic acid (e.g., a fragmented human genome) can then be contacted with the polished substrate such that individual target nucleic acids seed individual wells via interaction with primers attached to the gel material. However, due to the absence or inactivity of the gel material, the target nucleic acid does not occupy the interstitial regions. Amplification of the target nucleic acid will be restricted to the wells because the absence or inactivity of the gel in the interstitial regions prevents outward migration of growing nucleic acid colonies. The process is manufacturable and scalable, utilizing conventional micro- or nano-fabrication methods.
[0100] The sequencing instrument's imaging device (e.g., a solid-state imager such as a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS) sensor) takes snapshots at multiple locations along the lane, in a series of non-overlapping regions called tiles. For example, there may be 64 or 96 tiles per lane. A tile holds hundreds of thousands to millions of clusters.
[0101] The output of a sequencing run are sequencing images, each showing the intensity emission of a cluster and its surrounding background. A sequencing image shows the intensity emission generated as a result of incorporating a nucleotide into a sequence during sequencing. The intensity emission arises from the associated analytes / clusters and their surrounding background.
[0102] Sequencing images 702 are provided from multiple sequencing instruments, sequencing operations, cycles, flow cells, tiles, wells, and clusters. In one embodiment, the sequencing images are processed by base caller 704 on an imaging channel basis. A sequencing run produces m images per sequencing cycle corresponding to m imaging channels. In one embodiment, each imaging channel (also called color channel) corresponds to one of a number of filter wavelength bands. In another embodiment, each imaging channel corresponds to one of a number of imaging events in a sequencing cycle. In yet another embodiment, each imaging channel corresponds to a combination of illumination by a particular laser and imaging through a particular optical filter. In different embodiments, such as 4-channel chemistry, 2-channel chemistry, and 1-channel chemistry, m is 4 or 2. In other embodiments, m is greater than 1, 3, or 4.
[0103] In another embodiment, the input data is based on pH changes induced by the release of hydrogen ions during molecular elongation. The pH change is detected and converted to a voltage change proportional to the number of bases incorporated (e.g., in the case of Ion Torrent). In yet another embodiment, the input data is constructed from nanopore sensing, using a biosensor to measure the disruption of current as an analyte passes through or near the opening of the nanopore. For example, Oxford Nanopore Technologies (ONT) sequencing is based on the following concept: a single strand of DNA (or RNA) is passed through a membrane via a nanopore and a potential difference is applied across the membrane. Nucleotides present within the pore affect the electrical resistance of the pore, so that current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal (due to its appearance of "squishing" when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer data acquisition (DAC) values taken at a 4 kHz frequency (e.g.). With a DNA strand speed of about 450 base pairs per second, this gives, on average, about 9 raw observations per base. The signals are then processed to identify breaks in the aperture signal that correspond to individual reads. Extension of these raw signals is a process that base calls and converts DAC values into sequences of DNA bases. In some embodiments, the input data includes normalized or scaled DAC values. Additional information regarding non-image-based sequence data can be found in U.S. Provisional Patent Application No. 62 / 849,132, entitled "Base Calling Using Convolutions," filed May 16, 2020; U.S. Provisional Patent Application No. 62 / 849,133, entitled "Base Calling Using Compact Convolutions," filed May 16, 2019; and U.S. Non-Provisional Patent Application No. 16 / 826,168, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2019.
[0104] Spatially-varying sharpening mask A particular sharpening mask / mask / convolution kernel may be configured / trained to improve and / or improve and / or maximize the signal-to-noise ratio of a particular category / type / configuration / characteristic / class / bin of data. Similarly, each sharpening mask may be configured to improve and / or improve and / or maximize the signal-to-noise ratio of a respective instance / category / type / configuration / characteristic / class / bin of data. Various sharpening masks are disclosed. For example, a "surface-specific specialist sharpening mask" is configured / trained to improve and / or maximize the signal-to-noise ratio of sequencing data of clusters located on a particular surface or a particular surface type / category / class (e.g., top or bottom surface or surfaces 1-N of a flow cell). Similarly, a "lane-specific specialist sharpening mask" is configured / trained to improve and / or maximize the signal-to-noise ratio of sequencing data of clusters located on a particular lane or a particular lane type / category / class (e.g., central or peripheral lane or lanes 1-N of a flow cell). Also, a "tile-specific specialist sharpening mask" is configured / trained to improve and / or maximize the signal-to-noise ratio of sequencing data of clusters located on a particular tile or a particular tile type / category / class (e.g., central or peripheral tiles of a flow cell or tiles 1-N). Also, a "subtile-specific specialist sharpening mask" is configured / trained to improve and / or maximize the signal-to-noise ratio of sequencing data of clusters located on a particular subtile or a particular subtile type / category / class (e.g., central or peripheral subtiles of a flow cell or subtiles 1-N). In some implementations, a single sharpening mask can include multiple specialist coefficient sets such that each specialist coefficient set is configured / trained to improve and / or maximize the signal-to-noise ratio of a particular category / type / configuration / characteristic / class / bin of data. In some implementations, a single sharpening mask can include various specialist coefficient sets.For example, a "surface-specific specialist coefficient set" is configured / trained to improve and / or maximize the signal-to-noise ratio of sequencing data of clusters located on a particular surface or a particular surface type / category / class (e.g., top or bottom surface or surfaces 1-N of a flow cell). Similarly, a "lane-specific specialist coefficient set" is configured / trained to improve and / or maximize the signal-to-noise ratio of sequencing data of clusters located on a particular lane or a particular lane type / category / class (e.g., central or peripheral lanes or lanes 1-N of a flow cell). Also, a "tile-specific specialist coefficient set" is configured / trained to improve and / or maximize the signal-to-noise ratio of sequencing data of clusters located on a particular tile or a particular tile type / category / class (e.g., central or peripheral tiles or tiles 1-N of a flow cell). Also, the "subtile specific specialist coefficient sets" are configured / trained to improve and / or maximize the signal to noise ratio of sequencing data for clusters located on specific subtiles or specific subtile types / categories / classes (e.g., central or peripheral subtiles or subtiles 1-N of a flow cell). The disclosed specialist sharpening masks are applicable to clusters located on both patterned and unpatterned surfaces of a flow cell. In an unpatterned surface, the clusters are randomly distributed on the flow cell. The randomly distributed clusters and the data therefor (e.g., images) may be binned spatially, temporally, signal-wise, or by any combination thereof. Thus, the specialist sharpening masks may be configured and trained for different configurations of differently binned and randomly distributed clusters. In a patterned surface, the clusters are located on patterned wells with fixed locations. The patterned wells and component clusters may be binned spatially, temporally, signal-wise, or by any combination thereof. Therefore, specialist sharpening masks can be constructed and trained for different configurations of differently binned and patterned clusters.The disclosed specialist sharpening masks are configuration-specific sharpening masks that are trained to improve and / or maximize the signal-to-noise ratio of image data generated for different configurations of a sequencing operation. These configurations can be spatial configurations associated with different regions on the flow cell, temporal configurations associated with different sequencing / imaging cycles of a sequencing operation, signal distribution configurations associated with different distributions / patterns of signal profiles observed / encoded in the imaging data, or combinations thereof. Other examples of configurations encompassed by the present disclosure include segmenting the sequencing data by imaging type, color channel type, laser type, optics type, lens type, optical filter type, illumination type, library type, sample type, index type (first index read vs. second index read), read type (forward read vs. reverse read), physical characteristics of the sample, noise type (e.g., bubbles), and reagent type, and training the corresponding specialist sharpening masks.
[0105] FIG. 8A shows multiple sharpening masks 820 used for corresponding sections of a sequencing image generated for corresponding regions of a flow cell, where each tile of the flow cell is divided into 3×3 subtile regions, and each subtile region is assigned one or more corresponding sharpening masks.
[0106] For example, in FIG. 8A, two exemplary tiles 812 and 814 of a flow cell are shown (see FIG. 2 for further discussion of tiles and flow cells), which generates the sequencing image 702 of FIG. 7. Tile 812 is divided into 3×3 subtile regions 812a, 812b, ..., 812i, as shown. Similarly, tile 814 is divided into 3×3 subtile regions 814a, 814b, ..., 814i, as shown. Similarly, other tiles of the flow cell may be divided into corresponding 3×3 subtile regions. By way of example only, if a tile has 9000×9000 pixels in the corresponding image, the image is divided into subtile regions such that each subtile region has 3000×3000 pixels.
[0107] Each tile contains multiple clusters, for example each 3000x3000 pixel subtile region of an image contains images of corresponding multiple clusters.
[0108] Each subtile region of a tile is assigned one or more corresponding sharpening masks. For example, in the example of FIG. 8A, two color channels 802A, 802B are assumed as an example only, but there may be any different number of color channels. For example, sharpening mask 820Ax corresponds to color channel 802A, sharpening mask 820Bx corresponds to color channel 802B, and "A" in sharpening mask 820Ax indicates that these masks are for processing images of color channel 802A, and "B" in sharpening mask 820Bx indicates that these masks are for processing images of color channel 802B.
[0109] Additionally, the index "x" in masks 820Ax and 820Bx is associated with the corresponding subtile 812x, 814x for which the mask is to be used. For example, mask 820Aa is used for the section of sequencing image 702 generated from subtile 812a of tile 812 and also used for subtile 814a of tile 814, mask 820Ba is used for the section of sequencing image 702 generated from subtile 812a of tile 812 and also used for subtile 814a of tile 814, mask 820Ab is used for the section of sequencing image 702 generated from subtile 812b of tile 812 and also used for subtile 814b of tile 814, etc.
[0110] Thus, in summary, for example, mask 820Aa is used for the section of sequencing image 702 corresponding to color channel 802A and for subtile regions 812a and 814a, mask 820Ba is used for the section of sequencing image 702 corresponding to color channel 802B and for subtile regions 812a and 814a, mask 820Ab is used for the section of sequencing image 702 corresponding to color channel 802A and for subtile regions 812b and 814b, mask 820Bb is used for the section of sequencing image 702 corresponding to color channel 802B and for subtile regions 812b and 814b, etc.
[0111] Note that the same sharpening mask is used for corresponding subtile regions of multiple tiles, for example, sharpening masks 802Aa and 802Ba are used for the top left subtile of multiple or all tiles of the flow cell, sharpening masks 802Ae and 802Be are used for the center subtiles of multiple or all tiles of the flow cell, etc.
[0112] Thus, in the example of Figure 8A, where each tile is divided into 3 × 3 subtile regions and two color channels are assumed, * There are 2, or 18 sharpening masks. In general, if each tile is divided into N subtile regions and M color channels are assumed, there are MxN sharpening masks.
[0113] In one example, a k×k (e.g., 3×3) subdivision of the tile may be used for a scenario in which a fully automated image capture system is used to capture the sequencing images. For example, in a fully automated image capture system, the center of the tile may be captured slightly differently than the edges of the tile due to, for example, distortion effects, different focusing on different sections of the tile, etc. Thus, the edges of the tile may have a different sharpening mask than the center of the tile, as shown in FIG. 8A. Furthermore, images from different edges of the tile may also be slightly different (i.e., each edge may not be represented similarly in the image) due to factors such as tilt of the optics relative to the flow cell. Thus, in the example of FIG. 8A, each of the nine subtiles may have a different associated sharpening mask.
[0114] FIG. 8B shows multiple sharpening masks 840 used for corresponding sections of a sequencing image generated for corresponding regions of the flow cell, where each tile of the flow cell is divided into 1 x 9 subtile regions, and each subtile region is assigned one or more corresponding sharpening masks.
[0115] For example, in Figure 8B, two exemplary tiles 832 and 834 of a flow cell are shown (see Figure 2 for further discussion of tiles and flow cells) that generate the sequencing image 702 of Figure 7. Tile 832 is divided into 1 x 9 subtile regions 832a, 832b, ..., 832i as shown. Similarly, tile 834 is divided into 1 x 9 subtile regions 834a, 834b, ..., 834i as shown. Similarly, other tiles of the flow cell may be divided into corresponding 1 x 9 subtile regions.
[0116] By way of example only, if a tile has 9000 x 9000 pixels in the corresponding image, the image is divided into subtile regions such that each subtile region has 9000 x 1000 pixels. Each 9000 x 1000 pixel subtile region of the image contains images of a corresponding number of clusters.
[0117] Each subtile region of a tile is assigned one or more corresponding sharpening masks. For example, in the example of FIG. 8B (similar to the example of FIG. 8A), two color channels 804A, 804B are assumed as an example only, but there may be any different number of color channels. For example, sharpening mask 840Ax corresponds to color channel 804A, sharpening mask 840Bx corresponds to color channel 804B, and "A" in sharpening mask 840Ax indicates that these masks are for processing images of color channel 804A, and "B" in sharpening mask 840Bx indicates that these masks are for processing images of color channel 804B.
[0118] Further, the index "x" in masks 840Ax and 840Bx is associated with the corresponding subtile 832x, 834x for which the mask is to be used. For example, mask 840Aa is used for the section of sequencing image 702 generated from subtile 832a of tile 832 and is also used for subtile 834a of tile 834, mask 840Ba is used for the section of sequencing image 702 generated from subtile 832a of tile 832 and is also used for subtile 834a of tile 834, and similarly, masks 840Ab and 840Bb are used for the section of sequencing image 702 generated from subtile 832b of tile 832 and is also used for subtile 834b of tile 834.
[0119] Mask 840Aa is used for the section of the sequencing image 702 corresponding to color channel 804A and for subtile regions 832a and 834a, mask 840Ba is used for the section of the sequencing image 702 corresponding to color channel 804B and for subtile regions 832a and 834a, mask 840Ab is used for the section of the sequencing image 702 corresponding to color channel 804A and for subtile regions 832b and 834b, mask 840Bb is used for the section of the sequencing image 702 corresponding to color channel 804B and for subtile regions 832b and 834b, and so on.
[0120] Thus, in the example of Figure 8B where each tile is divided into 1 × 9 subtile regions and two color channels are assumed, * There are 2, or 18 sharpening masks. In general, if each tile is divided into N subtile regions and M color channels are assumed, there are MxN sharpening masks.
[0121] In one example, a 1×k (e.g., 1×9) subdivision of tiles may be used for a scenario in which a line-scan image capture system is used to capture the sequencing image. For example, in a line-scan image capture system, various vertical sub-regions of an image may be captured differently. Thus, the image is divided into different vertical sub-regions, as shown in FIG. 8B, and each sub-region is assigned its own corresponding sharpening mask.
[0122] FIG. 8C shows multiple sharpening masks 860 used for corresponding sections of a sequencing image generated for corresponding regions of a flow cell, where each tile of the flow cell is divided into multiple subtile regions and similar subregions occurring periodically within a tile are assigned one or more corresponding sharpening masks.
[0123] For example, in Figure 8C, two exemplary tiles 852 and 854 of a flow cell are shown that generate the sequencing image 702 of Figure 7. Tile 852 is divided into 3x3 subtile regions, with the corner regions of each subtile shown using shades of grey. Shaded regions in the various subtiles of tiles 852 and 858 are labeled as shaded regions 855a, and non-shaded regions in the various subtiles of tiles 852 and 858 are labeled as non-shaded regions 855b.
[0124] In the example of FIG. 8C, the shaded regions 855a occur with a particular periodicity (e.g., the top left corner of each subtile), but this is by way of example only, and the shaded regions 855a can occur with any other type of periodicity. For example, two horizontal lines of pixels of a tile may be included in the shaded region 855a, followed by a non-shaded region 855b that includes five horizontal lines of pixels, and this pattern may be repeated. Thus, in this example, the two lines of pixels of the shaded region 855a and the five lines of pixels of the non-shaded region 855b are interleaved and occur in a repeating pattern. Any other patterns of shaded regions 855a and non-shaded regions 855b may also be possible. As merely an example, the intersection of every other pixel (fourth row) with a pixel (fifth and sixth columns) may be included in the shaded region 855a, and this pattern of shaded regions may be repeated throughout the image.
[0125] In one example, the use of the repeating pattern of shaded and non-shaded regions shown in FIG. 8C can be used in a scenario where a CMOS (complementary metal oxide semiconductor) image capture sensor is used to capture the sequencing image. For example, some sequencing platforms use a flow cell with an embedded CMOS sensor. The sequencing chemistry is performed directly on top of the CMOS sensor, which is then imaged with the aid of LEDs that excite fluorescent molecules on the sensor. In one example (e.g., due to design and cost requirements to meet both imaging and chemistry), the CMOS sensor readout circuitry is embedded in the sensor itself as repeating rows and columns of "dark pixels," with periodic patches of such dark pixels symbolically represented as shaded regions 855a in FIG. 8C. This design pattern creates a unique intensity extraction challenge that requires the use of different extraction kernels with specific periodicities, as discussed with respect to FIG. 8C. The use of CMOS sensors embedded within flow cells may be found in International Publication No. WO 2020 / 236945, which is incorporated by reference as if fully set forth herein.
[0126] Each shaded region 855a of a tile is assigned one or more corresponding sharpening masks. For example, in the example of FIG. 8C (similar to the example of FIG. 8A), two color channels 806A, 806B are assumed by way of example only, but there may be any different number of color channels. For example, sharpening masks 860Ax correspond to color channel 806A, sharpening masks 860Bx correspond to color channel 806B, and "A" in sharpening masks 840Bx indicates that these masks are for processing images of color channel 806A, and "B" in sharpening masks 860Bx indicates that these masks are for processing images of color channel 806B.
[0127] Additionally, the index "x" in masks 860Ax and 860Bx is associated with the corresponding shaded / non-shaded region 855x for which the mask is to be used. For example, masks 860Aa and 860Ba are used for sections of sequencing image 702 generated from shaded regions 855a of various tiles. Similarly, masks 860Ab and 860Bb are used for sections of sequencing image 702 generated from shaded regions 855b of various tiles.
[0128] Thus, mask 860Aa is used for the sections of the sequencing image 702 corresponding to color channel 806A and for shaded regions 855a, mask 860Ba is used for the sections of the sequencing image 702 corresponding to color channel 806B and for shaded regions 855a, mask 860Ab is used for the sections of the sequencing image 702 corresponding to color channel 806A and for non-shaded regions 855b, and mask 860Bb is used for the sections of the sequencing image 702 corresponding to color channel 806B and for non-shaded regions 855b.
[0129] Thus, in the example of FIG. 8C where each tile is divided into shaded and non-shaded regions and two color channels are assumed, * There are 2 or 4 sharpening masks.
[0130] training Referring again to FIG. 7, the base caller 704 generates one or more sharpening masks 706 (such as those discussed with respect to FIGS. 8A-8C) that are used to sharpen the sequencing image 702 (the sharpening operation is discussed in more detail with respect to FIGS. 10A-10K and 11). The sharpening operation includes intensity extraction from the sequencing image to generate a corresponding feature map, and a subsequent interpolation operation to assign weighted feature values for various clusters based on the sub-pixel locations of the clusters, as discussed in more detail later in this specification. The clusters with corresponding assigned weighted feature values are then base called.
[0131] In one embodiment, the number of sharpening masks 706 generated by the base caller 704 may be implementation specific, as discussed with respect to Figures 8A-8C. For example, each color channel may have a corresponding sharpening mask 706. In another example, a tile of a flow cell from which a sequencing image 702 is generated may be divided into two or more sections, with dedicated sharpening masks for each section of the tile, as discussed in more detail with respect to Figures 8A-8C.
[0132] As discussed in more detail (e.g., in FIG. 10F later herein), the sharpening masks 706 act as convolution kernels, where the sharpening masks are convolved with corresponding sections of the image. By way of example only, and referring to FIG. 8A, the sharpening masks 820Aa are convolved with the sections of the sequencing image 702 generated by subtiles 812a of color channel 802A. In one training implementation, the coefficients of each sharpening mask 706 are determined using a least squares estimate on a corresponding subset of data from the corresponding sections of the image. Thus, referring again to FIGS. 7 and 8A, for example, data from subtiles 812a for color channel 802A is used to generate and / or train the sharpening masks 820a.
[0133] As shown in FIG. 7, the input to the base caller 704 are the raw sensor pixels of the sequencing images from various tiles of the flow cell. Each sharpening mask 706 has multiple coefficients learned from training. In one embodiment, the number of coefficients in the sharpening mask corresponds to the number of sensor pixels used to base call the clusters. In one example, the sharpening mask is a square matrix with k×k coefficients, where k is a suitable positive integer, such as 3, 5, 7, 9, etc. Thus, each sharpening mask 706 has k 2 It has coefficients.
[0134] The training produces sharpening mask coefficients that are configured to mix / combine the intensity values of pixels representing intensity radiation from the base-called cluster and intensity radiation from one or more adjacent clusters to maximize the signal-to-noise ratio. The signal that is maximized in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise that is minimized in the signal-to-noise ratio is the intensity radiation from the adjacent clusters, i.e., spatial crosstalk plus some random noise (e.g., to account for background intensity radiation). The sharpening mask coefficients are used as weights, and the mixing / combining involves performing element-wise multiplication between the sharpening mask coefficients and the intensity values of the pixels to calculate a weighted sum of the intensity values of the pixels (e.g., features in a feature map, see Figures 10E and 10F).
[0135] During training, the base caller 704 learns to maximize the signal-to-noise ratio by least-squares estimation, according to one embodiment. Using least-squares estimation, the base caller 704 is trained to estimate the shared sharpening mask coefficients from the pixel intensities around the wells of interest and the desired output. Least-squares estimation is suitable for this purpose because it minimizes the squared error and outputs coefficients that take into account the effects of noise amplification.
[0136] The desired output is an impulse at a well (i.e., cluster) location (point source) when the intensity channel is on, and a background level when the intensity channel is off. In some implementations, ground truth 712 is used to generate the desired output. In one example, ground truth 712 includes ground truth base calls. Additionally or alternatively, in some examples, ground truth includes the center (or mean) of the cloud for each base, as shown in FIG. 9A and discussed in more detail herein.
[0137] In some embodiments, the ground truth 712 is modified to account for the DC offset per well, the amplification factor, the degree of polyclonality, and the gain offset parameters included in the least squares estimate. In one embodiment, during training, a DC offset, i.e., a fixed offset, is calculated as part of the least squares estimate. During inference, the DC offset is added as a bias to each sharpening mask calculation.
[0138] In one embodiment, the desired output is estimated using Illumina's Real-time Analysis (RTA) base caller. Details regarding RTA can be found in U.S. Patent Application No. 13 / 006,206, which is incorporated by reference as if fully set forth herein. The base calling error is averaged over many training examples. In another embodiment, the ground truth 712 is provided using aligned genomic data, which has better quality because it can use reference genomes and truth information that incorporate knowledge gained from multiple sequencing platforms and sequencing operations to average out noise.
[0139] The ground truth 712 are base-specific intensity values (or feature values discussed later herein) that reliably represent the intensity profiles of bases A, C, G, and T, respectively. A base caller such as RTA processes the sequencing image 702 and base calls clusters by producing per-color intensity values / output for each base call. The per-color intensity values can be considered as per-base intensity values, since depending on the type of chemistry (e.g., 2-color chemistry or 4-color chemistry), a color is mapped to each of the bases A, C, G, and T. The base with the closest intensity profile match is called.
[0140] FIG. 9A illustrates one embodiment of a per-base Gaussian fit centered on a per-base target that is used as a ground truth value for error calculation during training. The per-base intensity output produced by the base caller for a large number of base calls in the training data (e.g., tens, hundreds, thousands, or millions of base calls) is used to create a per-base intensity distribution. FIG. 9A illustrates a chart of four Gaussian clouds that are probability distributions of the per-base intensity output for bases A, C, G, and T, respectively. The intensity values at the centers of the four Gaussian clouds are used as the ground truth intensity targets (or feature value targets) of the ground truth 712 for bases A, C, G, and T, respectively, and are referred to herein as targets (e.g., intensity or feature value targets).
[0141] Consider that during training, the input image data provided to the base caller 704 is annotated with base "A" as the ground truth base call. The ground truth 712 also includes base-specific intensity values that reliably represent the intensity profiles of bases A, C, G, and T, respectively. Thus, for example, the ground truth 712 also includes, for base A, the coordinates of the average intensity or average feature value for base A (i.e., the center of the green cloud in FIG. 9A) as shown in FIG. 9A (feature values are discussed later in this specification). The target / desired output of the base caller 704 is then the intensity value or feature value at the center of the green cloud in FIG. 9A, i.e., the intensity target for base A. Similarly, for base "C", the ground truth includes the intensity value or feature value at the center of the blue cloud in FIG. 9A, i.e., the intensity target (or feature value target) for base C with coordinates (Cx,Cy). Similarly, for base "T", the ground truth includes the intensity value or feature value at the center of the red cloud in Figure 9A, i.e., an intensity target (or feature value target) for base T with coordinates (Tx, Ty). And for base "G", the ground truth includes the intensity value or feature value at the center of the brown cloud in Figure 9A, i.e., an intensity target (or feature value target) for cardinality G with coordinates (Gx, Gy).
[0142] Thus, the targets or desired outputs during training of the base caller 704 are the average intensities (or average feature values) for each base A, C, G, and T after averaging over the training data. In one embodiment, the trainer 714 uses least squares estimation to fit the coefficients of the sharpening mask 706 to minimize the output error towards these intensity targets.
[0143] In one embodiment, during training, the base caller 704 applies the coefficients in a given sharpening mask to the pixels of the sequencing image labeled with a given base. This involves element-wise multiplication of the coefficients with the intensity values of the pixels to generate a weighted sum of the intensity values of the feature map, where the coefficients act / act / are used as weights. The feature map includes various features with corresponding feature values. Note that the centers of the clusters may not be aligned with the centers of the pixels of the sequencing image 702. To account for such misalignment, in the feature map generated from the sequencing image 702 (the feature map is generated by convolving the sharpening mask with the corresponding section of the image), the weighted feature values assigned to the clusters are generated by bilinear interpolation, e.g., neighboring features are interpolated to generate the weighted feature value corresponding to the cluster, which will be discussed in more detail herein. The interpolated feature value corresponding to the cluster then becomes the predicted output of the base caller 704 for that cluster. Next, the error (e.g., least square error, minimum mean square error) between the interpolated weighted feature value (e.g., from the center of the intensity Gaussian fit corresponding to the average intensity observed for a given base) and the intensity target determined for a given base is calculated based on a cost / error function (e.g., sum of squared errors, SSE). Cost functions such as SSE are differentiable functions used to estimate the sharpening mask coefficients using an adaptive approach, so the derivatives of the error with respect to the coefficients can be evaluated, and these derivatives are used to update the coefficients with values that minimize the error. This process is repeated until the updated coefficients no longer reduce the error. In another embodiment, a batch least squares method is used to train the base caller 704.
[0144] For example, assume that the center of the green cloud in FIG. 9A, i.e., the intensity target for base A, is (Ax, Ay), which is the target or desired output (e.g., target feature value) for base A base calling. Assume that during a sequencing operation, a cluster 904 has weighted feature values represented by coordinates (Ix, Iy). In one embodiment, the base caller 704 updates the coefficients in a given sharpening mask, so that the intensity of the cluster 904 is transposed from coordinates (Ix, Iy) to coordinates (Ax, Ay). Thus, training aims to minimize or reduce the distance between coordinates (Ax, Ay) and coordinates (Ix, Iy).
[0145] In another example, the per-base intensity distribution / Gaussian cloud shown in Figure 9A can be generated for each well and noise corrected by adding DC offsets, amplification factors, and / or phasing parameters. In this way, depending on the well location of a particular well, the corresponding per-base Gaussian cloud can be used to generate a target intensity value for that particular well (or cluster corresponding to the well).
[0146] In one implementation, a bias term is added to the dot product that produces the output of the base caller 704. During training, the bias parameter can be estimated using a similar approach used to learn the sharpening mask, i.e., least squares or least mean squares (LMS). In some implementations, the value of the bias parameter is a constant value equal to 1, i.e., a value that does not change with the input pixel intensity. There is one bias per coefficient set. The bias is learned during training and then fixed for use during inference. The learned bias, along with the learning coefficient for each sharpening mask, represents a DC offset that is used in all calculations during inference. This bias accounts for random noise caused by different cluster sizes, different background intensities, changing stimulus responses, changing focus, changing sensor sensitivity, and changing lens aberrations.
[0147] In yet another decision-directed implementation, the output of the basecaller 704 is presumed to be correct for training purposes.
[0148] A trainer 714 can train the base caller 704 and generate trained coefficients for the sharpening mask 706 using a number of training techniques. Examples of training techniques include least squares estimation, least squares, least mean squares, and recursive least squares. In the least squares technique, the parameters of a function are adjusted to best fit the data set such that the sum of squared residuals is minimized. In other implementations, other estimation and fitting algorithms can be used to train the base caller 704.
[0149] The base caller 704 can be trained in an offline or online mode of adaptation. According to one embodiment, the trained coefficients of the base caller 704 are generated and / or updated using the following batch least squares algorithm:
[0150]
number
[0151] In the above formula, the sharpening mask coefficient is the beta hat
[0152]
number
[0153]
number
[0154]
number
[0155] X is a matrix of pixel values of size m×k×k, i.e., with m rows and (k×k) columns, where m is a suitable positive integer. Each row of the matrix X corresponds to one cluster, and each column is the value of an image pixel after adjustment for sub-pixel interpolation.
[0156] y is a vector of size m corresponding to the centroid location of each cluster, i.e., y is the target output for all training examples, i.e., each value is the intensity center of the on / off cloud depending on the truth of the training examples, and beta is the set of coefficients that minimizes the sum of squared residuals.
[0157] In one example, the base caller 704 may be trained in an online mode to adapt the coefficients of the sharpening mask 706 to track changes in temperature (e.g., optical distortion), focus, chemistry, machine-specific variations, etc., for example, as the sequencing machine operates and the sequencing operation proceeds cyclically. In the online mode, the trained coefficients of the sharpening mask 706 are updated using an adaptation technique. In the online mode, a least mean square method, which is a form of stochastic gradient descent, is used as the training algorithm. Further details regarding the online adaptation of the coefficients of the sharpening mask 706 are discussed later in this specification, for example, with respect to Figures 12 and 13.
[0158] In the least mean square method, the gradient of the squared error for each coefficient is used to move the coefficient in the direction that minimizes the cost function, which is the expected value of the squared error. It has a very low computational cost, and only multiplication and accumulation operations per coefficient are performed. No long-term storage is required, except for the coefficients. The least mean square method is suitable for processing large amounts of data (e.g., parallel processing of data from billions of clusters). Extensions of the least mean square method include the normalized least mean square method and the frequency domain least mean square method, which can also be used here. In some embodiments, the least mean square method can be applied in a decision-directed manner that assumes that our decisions are correct, i.e., our error rate is very low, and small mu values filter out the hindered updates due to inaccurate base calls.
[0159] FIG. 9B illustrates one embodiment of an adaptation technique that may be used to train the base caller 104, for example, using an offline or online mode. Here, the logic is y=x.h+d, where x is the input pixel intensity, h is the sharpening mask coefficient, and d is a DC offset. In one embodiment, x and h are row and column vectors, respectively, with length 81. This vector model corresponds to the dot product of a 9×9 matrix representing the input pixel and the coefficients. The cost is the expected squared error. The gradient update moves each coefficient in the direction that reduces the expected squared error. This results in the next update.
[0160]
number
[0161] In most systems, the expectation function E{x(n)e * (n)}. This can be done using the following unbiased estimator:
[0162]
number
[0163] N denotes the number of samples to estimate. In the simplest case, N=1.
[0164]
number
[0165] For this simple case, the update algorithm is as follows:
[0166]
number
[0167] In effect, this constitutes the update algorithm for the LMS filter.
[0168] In the above equation, h is a vector of sharpening mask coefficients, x is a vector of input intensities, and e is the error of the calculation performed using the values at x, i.e., only one error term per output.
[0169] Applying this update produces new estimates of the coefficients that move the coefficients (on average) in the direction that reduces the mean squared error (MSE). In some implementations, Mu is a small constant that is used to change the adaptation rate / convergence speed. The DC term update can be calculated in a similar manner. The gain term update can be calculated in a similar manner.
[0170] In some implementations, since linear interpolation is applied to the coefficient set, the updates are applied slightly differently in the following manner. h(q,n+1)=h(q,n)+lambda_q.mu.x(n).e(n)
[0171] where h(q,n) is the weight q in cycle n, and lambda_q is the linear interpolation weight for a particular set of coefficients, which may include four updates per output by linear interpolation in two dimensions.
[0172] Recursive least squares is an extension of least squares to a recursive algorithm.
[0173] Spatial Crosstalk Attenuator Figures 10A-10K jointly illustrate various implementations of using the trained sharpening mask 706 of Figures 7-8C to attenuate spatial crosstalk from sensor pixels and base call clusters using the crosstalk corrected sensor data. Specifically, Figure 10A illustrates a section 1000 of a sequencing image 702 from a subtile of a tile (e.g., subtile 812a of tile 812, see Figure 8A), where various cluster centers are offset relative to the centers of corresponding pixels.
[0174] Although a subtile likely generates a large number of pixels in the sequencing image 706, the section 1000 of FIG. 10A corresponding to a subtile includes only a few pixels for simplicity.
[0175] FIG. 10A further illustrates the centers of multiple clusters within a subtile, which are superimposed on a section 1000 of the sequencing image 706. Also assume that the section 1000 shown in FIG. 10 is for a particular color channel. Consider an optical system of a sequencer that uses two different imaging channels, a red channel and a green channel (although the sequencer may generate any different number of color channels, such as 1, 3, 4, or more). Then, at each sequencing cycle, the optical system creates a red image with a red channel intensity and a green image with a green channel intensity, which together form a single sequencing image (like the RGB channels of a typical color image). In one example, the pixels shown in FIG. 10 are for a particular color channel.
[0176] In Figure 10A, some of the clusters are shown and labeled with their centers using black dots. For example, in the XY coordinate plane, cluster 1011 has a center located at location (x1,y1), cluster 1012 has a center located at location (x2,y2), cluster 1013 has a center located at location (x3,y3), cluster 1014 has a center located at location (x4,y4), and cluster 1015 has a center located at location (x5,y5).
[0177] In one example, the location (e.g., coordinates) of the clusters on the tiles are identified using fiducial markers. The solid support on which the biological specimen is imaged may include such fiducial markers to facilitate the determination of the orientation of the specimen or its image relative to the probe attached to the solid support. Exemplary fiducials include, but are not limited to, beads (with or without a moiety such as a fluorescent moiety or a nucleic acid to which a labeled probe can bind), fluorescent molecules bound to a known or determinable feature, or structures that combine a morphological shape with a fluorescent moiety. Exemplary fiducials are described in U.S. Patent Application Publication No. 2002 / 0150909, which is incorporated herein by reference. Thus, in one example, fiducial markers are used to determine the location of the clusters relative to section 1000 of sequencing image 706, and the coordinates of the clusters shown in FIG. 10A.
[0178] Note that the centers of the clusters do not have to coincide with the centers of the corresponding pixels: for example, the center of cluster 1011 is within pixel 1001 but is off-center, the center of cluster 1012 is within pixel 1002 but is off-center, the center of cluster 1013 is within pixel 1003 but is off-center, the center of cluster 1014 is within pixel 1004 but is off-center, and the center of cluster 1015 is within pixel 1005 but is off-center.
[0179] FIG. 10B visualizes an example of cluster-to-pixel signal 1033. In one embodiment, the sensor pixels are in the pixel plane. The spatial crosstalk is caused by a periodic distribution 1037 of clusters in the sample plane (e.g., flow cell). In one embodiment, the clusters are periodically distributed on the flow cell in a diamond shape and immobilized on the wells of the flow cell. In another embodiment, the clusters are periodically distributed on a hexagonal shaped flow cell and immobilized on the wells of the flow cell. The signal cones 1035 from the clusters are optically coupled to a local grid of sensor pixels via at least one lens (e.g., one or more lenses of an overhead or adjacent CCD camera).
[0180] In addition to diamonds and hexagons, the clusters can be arranged in other regular shapes, such as squares, diamonds, triangles, etc. In yet other embodiments, the clusters are arranged on the sample plane in a random, non-periodic arrangement. One of skill in the art will appreciate that the clusters can be arranged on the sample plane in any arrangement, as required by a particular sequencing implementation.
[0181] 10C visualizes an example of cluster-to-pixel signal overlap. Signal cones 1035 (see FIG. 10B) overlap and impinge on sensor pixels, creating spatial crosstalk 1037.
[0182] 10D visualizes an example of a cluster signal pattern. In one implementation, the cluster signal pattern follows an attenuation pattern 1039, where the cluster signal is strongest at the cluster center and attenuates as it propagates away from the cluster center.
[0183] FIG. 10E illustrates a convolution operation 1030Aa, in which the sharpening mask 820Aa is convolved with a corresponding section of the sequencing image to generate a corresponding feature map. In the example of FIG. 10E, a k×k (k=3 in this example, but k can be another suitable positive integer) sharpening mask 820Aa (see FIG. 8A) is convolved with a section 1000 of the sequencing image 702 from a subtile 812a of tile 812 for color channel 802A (see FIG. 10A showing section 1000). As in FIG. 10A, cluster centers of black dots are superimposed on the section 1000 of the sequencing image 702.
[0184] Feature map 1042Aa is generated as a result of the convolution operation. Note that feature map 1042Aa is specific to subtile 812a of tile 812 and is specific to color channel 802A. Again, cluster centers of black dots are superimposed on feature map 1042Aa.
[0185] Section 1000 has dimensions w×h, where w (width) and h (height) can be, for example, 100,000 or more, depending on the size of subtile 812a. Thus, w and h are based on partitioning the tile into different subtiles. In one implementation, due to convolution 1030Aa, the dimensionality of feature map 1042Aa can be different (e.g., smaller) than the dimensionality of section 1000. In another implementation, the dimensionality can be preserved, for example, by appropriately padding section 1000 before convolution 1030Aa or by appropriately padding feature map 1042A after the convolution operation.
[0186] The feature map 1042Aa includes a plurality of features, each corresponding to a respective pixel in the section 1000 of the sequencing image 702. By way of example only, feature 1051 of the feature map 1042Aa corresponds to pixel 1001 of the section 1000. For example, during the convolution 1030Aa, the sharpening mask 820Aa is moved across the section 1000, and multiplication and addition operations are performed at each location of the sharpening mask 820Aa. The feature 1051 is generated due to the multiplication and addition operations, for example, when the sharpening mask 820Aa is convolved with a patch of the section 1000 centered on pixel 1001, and thus the feature 1051 corresponds to pixel 1001. Similarly, other features of the feature map 1042Aa correspond to respective pixels of the section 1000 (i.e., a one-to-one positional mapping between the pixels of the section 1000 and the features of the feature map 1042Aa).
[0187] In the example of Figure 10E, the locations of the clusters are also superimposed on the features of feature map 1042Aa. For example, as shown in Figure 10A, the cluster centers of one or more clusters are off-center with respect to the centers of the corresponding pixels. Similarly, in Figure 10E, the cluster centers of one or more clusters are also off-center with respect to the centers of the corresponding features.
[0188] FIG. 10F illustrates a plurality of convolution operations, in which each of a plurality of sharpening masks is convolved with a corresponding one of a plurality of sections of the sequencing image 702 to generate a corresponding one of a plurality of feature maps. For example, referring to FIG. 8A and FIG. 10F, the sharpening mask 820Aa is convolved with the section 1000 of the sequencing image 702 corresponding to the subtile 812a for the color channel 802A to generate a corresponding feature map 1042Aa, which convolution operation is discussed in more detail with respect to FIG. 10E. Similarly, the sharpening mask 820Ab is convolved with the respective section of the sequencing image 702 corresponding to the subtile 812b for the color channel 802A to generate a corresponding feature map 1042Ab. Similarly, the sharpening mask 820Ai is convolved with the respective section of the sequencing image 702 corresponding to the subtile 812i for the color channel 802A to generate a corresponding feature map 1042Ai. Generally speaking, the sharpening mask 820Ax is convolved with each section of the sequencing image 702 corresponding to the subtile 812x for the color channel 802A to generate a corresponding feature map 1042Ax, x=a,...,i. The convolution operation 1030Ax (x=a,...,i) on the left side of Figure 10F is for the example color channel 802A.
[0189] The convolution operations 1030By, y=a,...,i, on the right side of FIG. 10F, are for the exemplary color channel 802B. For example, the sharpening mask 820Ba is convolved with the respective sections of the sequencing image 702 corresponding to the subtile 812a for the color channel 802B to generate the corresponding feature map 1042Ba. Similarly, the sharpening mask 820Bb is convolved with the respective sections of the sequencing image 702 corresponding to the subtile 812b for the color channel 802B to generate the corresponding feature map 1042By, y=a,...,i. Generally speaking, the sharpening mask 820Bx is convolved with the respective sections of the sequencing image 702 corresponding to the subtile 812y for the color channel 802B to generate the corresponding feature map 1042By, y=a,...,i.
[0190] Also, as previously discussed, the two color channels 802A and 802B are merely examples, and the sequencer may include any different number of color channels, such as one color channel, or three or more color channels.
[0191] Figure 10G shows feature map 1042Aa of Figure 10E in more detail, with some of the features and cluster centers labeled. For example, cluster 1011 has its center at location (x1,y1) and is within feature 1051, cluster 1012 has its center at location (x2,y2) and is within feature 1052, etc. (see also Figure 10A for cluster center coordinates in section 1000 of the sequencing image).
[0192] 10E and 10G, where the portion 1029 of the feature map 1042Aa that contains the target cluster 1011 is shown in more detail in an enlarged view. For example, the view of the portion 1029 of the feature map 1042Aa is enlarged or zoomed in, and the center of the cluster 1011 at location (x1, y1) is superimposed on the portion 1029 of the feature map 1042Aa.
[0193] As discussed, cluster 1011 is within feature 1051 (labeled 1051e in FIG. 10H), but is off-center relative to the center of feature 1051e. Eight neighboring features 1051a, ..., 1051d, 1051f, ..., 1051i surrounding feature 1051e are also labeled.
[0194] The center of each feature is represented using a black square in Figure 10H and several subsequent figures. As shown in Figure 10H, the center of feature 1051a has coordinates (xa,ya), the center of feature 1051b has coordinates (xb,yb), etc., and the center of feature 1051i has coordinates (xi,yi).
[0195] As discussed with respect to FIG. 10E, each feature 1051(a,e) in FIG. 10H has a corresponding feature value generated by the convolution 1030Aa. Referring to FIG. 10H, in one example, the cluster 1011 is assigned a weighted feature value, and the weighted feature value is assigned based on a suitable interpolation technique. For example, if the center of the cluster 1011 coincides with the center of the feature 1051e, the feature value of the feature 1051e may be assigned to the cluster 1011. However, in the example of FIG. 10H, since the center of the cluster 1011 does not coincide with the center of the feature 1051e, the weighted feature value assigned to the cluster 1011 is influenced not only by the feature 1051e, but also by one or more features adjacent to the feature 1051e.
[0196] In one embodiment, a suitable interpolation technique is used to assign a weighted feature value to a cluster 1011 based, for example, on (i) the feature value of the feature 1051e in which the center of the cluster 1011 resides, (ii) the feature values of one or more neighboring features that are within a threshold distance from the center of the cluster 1011, (iii) the center-to-center distance between the cluster center and the feature center, (iv) the center-to-center distance between the cluster center and the pixel center, and (v) the center-to-center distance associated with the cluster.
[0197] Note that FIG. 10H shows feature map 1042A in the feature map domain, i.e., cluster 1011 is superimposed on the feature map. The coordinates of the feature center and the center of cluster 1011 are also shown. As the name suggests, the center-to-center distance between cluster center and feature center refers to the distance between the center of the cluster and the center of the feature, and is also called the center-to-center distance between the cluster and the feature. For example, the center-to-center distance d1 between cluster 1011 and feature 1051e is the distance between coordinates (x1, y1) and (xe, ye), and is determined, for example, as follows:
[0198]
number
[0199] Similarly, the center-to-center distance between cluster 1011 and any other features may be determined.
[0200] Meanwhile, the center-to-center distance between the cluster center and the pixel center (also called the center-to-center distance between the cluster and the pixel) refers to the distance between the center of the cluster and the center of the pixel. For example, referring to FIG. 10A, a section 1000 of the sequencing image 702 is shown. As in FIG. 10H, the coordinates of the centers of the various pixels can be determined, and thus the center-to-center distance between the cluster 1011 and the various pixels can also be determined.
[0201] For example, FIG. 10I illustrates the convolution operation 1030Aa of FIG. 10E, further illustrating the center-to-center distance d2 between cluster 1011 and pixel 1011, as well as the center-to-center distance d1 between cluster 1011 and feature 1051e. Note that feature 1051e corresponds to pixel 1011, as discussed with respect to FIG. 10E. For example, during convolution 1030Aa, sharpening mask 820Aa is moved across section 1000, and multiplication and addition operations are performed at each location of sharpening mask 820Aa. Feature 1051e is generated due to multiplication and addition operations, for example, when sharpening mask 820Aa is convolved with a patch of section 1000 centered on pixel 1001, and thus feature 1051 corresponds to pixel 1001. Thus, the location of cluster 1011 relative to the center of pixel 1001 is the same as the location of cluster 1011 relative to the center of feature 1051e. That is, the distance d1 and the distance d2 are the same.
[0202] At least some of the interpolation operations discussed later in this specification may use either (i) the center-to-center distance between the cluster and the pixel, or (ii) the center-to-center distance between the cluster and the feature. For example, one implementation may use the center-to-center distance between the cluster and the pixel, and another implementation may use the center-to-center distance between the cluster and the feature, where these two center-to-center distances are numerically the same.
[0203] Some of the interpolation examples discussed later in this specification discuss the center-to-center distance between clusters and features; however, as will be readily understood by one of ordinary skill in the art, the center-to-center distance between clusters and pixels may be used instead.
[0204] For purposes of this disclosure, unless otherwise stated, the center-to-center distance associated with a cluster implies the center-to-center distance between the cluster and the corresponding pixel, or the center-to-center distance between the cluster and the corresponding feature.
[0205] In one example, the subpixel location of a cluster includes the location of the center of the cluster relative to the boundary or center of the pixel that the cluster is located in. For example, if pixel 1001 in Figure 10I is divided into a 3x3 grid of subpixels, cluster 1011 is likely to fall within the top right subpixel of pixel 1001.
[0206] In one example, the sub-feature location of a cluster includes the location of the center of the cluster relative to the boundary or center of the feature in which the cluster is located. For example, if feature 1051e in Figure 10I is divided into a 3x3 grid of sub-features, then cluster 1011 is likely contained within the top right sub-feature of pixel 1001.
[0207] Interpolation to determine weighted feature values for target clusters As discussed above with respect to Figure 10H, any suitable interpolation technique may be used to assign weighted feature values to clusters 1011, which may be based, for example, on (i) the feature value of feature 1051e in which the center of cluster 1011 resides, and (ii) the feature values of one or more neighboring features that are within a threshold distance from the center of cluster 1011. Several such interpolation techniques are discussed herein below. It should be noted that the list of interpolation techniques discussed below is not exhaustive and other suitable interpolation techniques known to those skilled in the art may also be used.
[0208] A. Nearest Neighbor Interpolation In this interpolation technique, the closest feature to cluster 1011 is determined and the feature value of the closest feature is assigned to cluster 1011. As shown in Figure 10H, the center of feature 1051e at location (xe,ye) is closest to the center (x1,y1) of cluster 1011. Therefore, cluster 1011 is assigned the feature value of feature 1051e.
[0209] Thus, this technique involves determining a center-to-center distance, for example between the center of the cluster 1011 (i.e., coordinates (x1,y1)) and the center of a nearby feature (although the center-to-center distance between the cluster and the pixel can also be used). The feature corresponding to the closest center-to-center distance is selected as the nearest neighbor, and the feature value of the nearest feature is assigned to the cluster. Note that the interpolation is also based on sub-pixel or sub-feature locations of the cluster.
[0210] B. Average of nearest neighbor interpolations Another exemplary interpolation technique involves averaging the feature values of the n nearest neighboring pixels, where n is a suitable integer such as 1, 4, 9, etc. For example, assuming n=4, the weighted feature value assigned to the cluster 1011 is the average of the feature values of the four nearest neighboring features, which in the example of FIG. 10H are features 1051b, 1051c, 1051e, and 1051f. Thus, this technique involves determining the center-to-center distance between the center of the cluster 1011 (i.e., coordinates (x1, y1)) and the centers of the neighboring features (although the center-to-center distance between the cluster and the neighboring pixels can also be used). The four nearest features are selected and their intensities are averaged to determine the weighted feature value assigned to the cluster 1011. Thus, the interpolation is also based on the sub-pixel or sub-feature location of the cluster. Note that n=4 is merely an example, and n can be any other suitable value, as would be readily understood by one of ordinary skill in the art based on the teachings of this disclosure.
[0211] C. Bilinear Interpolation In one embodiment, bilinear interpolation may be used to determine a weighted feature value to be assigned to a cluster 1011 based on the feature values of neighboring features.
[0212] Bilinear interpolation is an extension of linear interpolation to interpolate functions of two variables (e.g., x and y) on a rectilinear 2D grid. Bilinear interpolation is performed using linear interpolation, first in one direction and then again in the other direction. Each step is linear in the sampled values and positions, but the interpolation as a whole is not linear, but rather quadratic in the sample locations. Bilinear interpolation is one of the fundamental resampling techniques in computer vision and image processing, and is also called bilinear filtering or bilinear texture mapping.
[0213] Figure 10J shows an example scheme illustrating bilinear interpolation, where four features 1051b, 1051e, 1051c, and 1051f are the four closest features to the center of cluster 1011 (see Figure 10H for further details), and the feature values of features 1051b, 1051e, 1051c, and 1051f are to be bilinearly interpolated to generate a weighted feature value for cluster 1011.
[0214] Assume that the coordinates of the center of feature 1051b are (x1,y2), the coordinates of the center of feature 1051e are (x1,y1), the coordinates of the center of feature 1051c are (x2,y2), the coordinates of the center of feature 1051f are (x2,y1), and the coordinates of the center of cluster 1011 are (x,y), as shown in Figure 10J. Note that such labeling of the coordinates is opposite to the labeling in Figure 10H. The coordinates of the centers are so labeled in Figure 10J for simplicity.
[0215] Assume that features 1051b, 1051e, 1051c, and 1051f are labeled Q12, Q11, Q22, and Q21, respectively, based on the coordinates discussed above. Thus, the feature values of features 1051b, 1051e, 1051c, and 1051f are labeled f(Q12), f(Q11), f(Q22), and f(Q21), respectively, and are known. For example, during the convolution operation discussed with respect to FIG. 10E, the feature values f(Q12), f(Q11), f(Q22), and f(Q21) are determined.
[0216] Bilinear interpolation aims to interpolate feature values f(Q12), f(Q11), f(Q22), and f(Q21) to the cluster centers at (x,y) to assign weighted feature values to cluster 1011.
[0217] First, linear interpolation in the x direction is performed on the coordinates (x, y1) and (x, y2) as follows.
[0218]
number
[0219] Then, linear interpolation in the y direction is performed on the coordinates (x,y) as follows:
[0220]
number
[0221] Thus, f(x,y) provides the weighted feature at coordinate (x,y) that is the center of cluster 1011, using bilinear interpolation. Thus, f(x,y) is the weighted feature assigned to cluster 1011.
[0222] D. Bicubic Interpolation In mathematics, bicubic interpolation is an extension of cubic interpolation to interpolate data points on a two-dimensional regular grid. The interpolated surface is smoother than the corresponding surface obtained by bilinear or nearest neighbor interpolation. Bicubic interpolation can be achieved using either Lagrange polynomials, cubic splines, or cubic convolution algorithms. In one example, bicubic interpolation may be chosen over bilinear or nearest neighbor interpolation in image resampling when processing speed is not an issue.
[0223] In contrast to the bilinear interpolation discussed above, which takes into account four neighboring features when determining a weighted feature value for the cluster 1011, bicubic interpolation takes into account 16 feature values (such as in a 4×4 grid of features surrounding the center of the cluster 1011). For example, to select the 4×4 grid of features closest to the cluster center, the center-to-center distance between the cluster center and the feature center (or pixel center, as discussed herein above) is considered. The feature values of the 4×4 grid of features are then used to determine a weighted feature value for the cluster 1011, for example according to bicubic interpolation.
[0224] E. Interpolation based on weighted area coverage Another interpolation technique assigns weighted feature values to cluster 1011 based on the coverage area around the central cluster, as shown in FIG. 10K. For example, as shown in FIG. 10K, a coverage A area is drawn around cluster 1011 such that the center of cluster 1011 and the center of coverage A area coincide. In one example, coverage A area has a square shape. In one example, coverage A area has a size equal to the size of the feature, as an example only. For example, assume that coverage A area covers Wb% of feature 1051b, Wc% of feature 1051c, Wf% of feature 1051f, and We% of feature 1051e. Then, the weighted feature value assigned to cluster 1011 is as follows:
[0225]
number
[0226] Note that Figure 10K assumes that the region of coverage A has a size equal to the size of the feature. In another example, the region of coverage A may have a size equal to, for example, two or three times the size of the feature, or may be a non-integer multiple of the feature (e.g., 1.5 times the size of the feature). In such an example, the weighted feature values of cluster 1011 may be based on feature values of more than four features, as would be readily understood by one of ordinary skill in the art based on the teachings of this disclosure.
[0227] F. Other Exemplary Interpolation Techniques Some examples of interpolation techniques are discussed herein above. In an embodiment, any other suitable interpolation technique may be used. For example, Lanczos resampling or Lanczos interpolation may be used for interpolation to determine the weighted feature values to be assigned to the clusters 1011. Lanczos filtering and Lanczos resampling are two applications of mathematical formulas that can be used to smoothly interpolate values of a digital signal between its samples. For example, this technique maps each sample of a given signal to a transformed and scaled copy of a Lanczos kernel, which is a sinc function windowed by the central lobe of a second, longer sinc function. The sum of these transformed and scaled kernels is then evaluated at the desired point. This filter is named after its inventor, Cornelius Lanczos.
[0228] Another exemplary type of interpolation technique uses a Hanning window, which may be used for interpolation to determine the weighted feature values to be assigned to the clusters 1011. In signal processing and statistics, a window function is a mathematical function that is zero-valued outside of some selected interval, usually symmetric around the center of the interval, usually near a central maximum, and usually becomes smaller and smaller the further away from the center. The Hanning window, also known as a raised cosine because of its zero-phase version, is one example of a window function. Unlike the Hamming window, the endpoints of the Hanning window only touch zero. In an embodiment, another suitable window function may be used for the interpolation.
[0229] Base calling Following the interpolation discussed herein above, the weighted feature values for a cluster are provided as input to a base caller 704 to generate a base call for that cluster. The base caller 704 can be a non-neural network based base caller or a neural network based base caller, examples of both are described in applications incorporated herein by reference, such as U.S. Patent Application No. 62 / 821,766 and U.S. Patent Application No. 16 / 826,168.
[0230] As discussed, the weighted feature value assignment to the clusters maximizes or increases the signal-to-noise ratio and reduces the spatial crosstalk between adjacent clusters. For example, by convolution (see FIG. 10E) and interpolation, the spatial crosstalk between adjacent clusters is reduced or eliminated. For example, the coefficients of the sharpening mask 820 are tuned to maximize or increase the signal-to-noise ratio. The signal maximized or increased in the signal-to-noise ratio is the intensity radiation from the target cluster, and the noise minimized or reduced in the signal-to-noise ratio is the intensity radiation from the adjacent clusters, i.e., the spatial crosstalk plus some random noise (e.g., to account for background intensity radiation).
[0231] Once the weighted feature values are assigned to the clusters, base calls are made to the clusters by the base caller based on the weighted feature values assigned to the clusters. Thus, for a sequencing operation that includes multiple sequencing cycles, a sequencing image 702 is generated for each sequencing cycle. The sequencing image 702 includes images for multiple clusters and one or more color channels for a given sequencing cycle.
[0232] For example, as discussed, for a particular sequencing cycle, a first weighted feature value can be assigned to a particular cluster for a first color channel, and a second weighted feature value can be assigned to a particular cluster for a second color channel (e.g., assuming there are two color channels, there may be one, three, or any other greater number of color channels). In such an example, a base call for a particular cluster and a particular sequencing cycle can be based on the first weighted feature and the second weighted feature. Further details of base calling are described in applications incorporated herein by reference, such as U.S. Patent Application No. 62 / 821,766 and U.S. Patent Application No. 16 / 826,168.
[0233] Base calling methods and performance results using convolution and interpolation FIG. 11A shows a method 1100 of base calling based on convolution and subsequent interpolation of at least one section of a sequencing image to assign one or more weighted feature values to a cluster, and base calling of the cluster based on the assigned one or more weighted feature values.
[0234] At 1104 of method 1100, for a particular sequencing cycle of a sequencing operation, a sequencing image (e.g., sequencing image 702 of FIG. 7) output by a flow cell (e.g., a flow cell discussed with respect to FIG. 1) during the corresponding sequencing cycle is accessed by a base caller, such as base caller 704 of FIG. 7.
[0235] At 1108, the sequencing image is partitioned into a number of sections based on color channels and / or spatial portions of the flow cell, with each section of the sequencing image including a number of clusters for a corresponding color channel.
[0236] For example, in Figure 8A, each tile of the flow cell is divided into 3x3 spatial portions, and thus the sequencing image generated from the tile for a particular color channel is partitioned into corresponding 3x3 sections. Further, in Figure 8A, two color channels are assumed, merely by way of example, without limiting the scope of the present disclosure. Thus, for a particular tile, the sequencing image is partitioned into a first 3x3 section for a first color channel and a second 3x3 section for a second color channel.
[0237] Similarly, in the example of FIG. 8B, each tile of the flow cell is divided into 1×9 sections, and therefore a sequencing image generated from the tile for a particular color channel is partitioned into a corresponding first 1×9 section for the first color channel and a second 1×9 section for the second color channel (i.e., assuming two color channels).
[0238] Those skilled in the art may envision other exemplary partitions of the sequencing image based on the teachings of this disclosure, for example, different exemplary divisions of tiles and different numbers of color channels.
[0239] The method 1100 then proceeds to 1112, where each section of the sequencing image is convolved with a corresponding sharpening mask to generate a corresponding feature map for the corresponding section, such that multiple feature maps are generated for multiple sections. For example, as discussed with respect to Figures 8A and 8B, each section of the sequencing image has a corresponding sharpening mask. As shown in Figure 10F, each section of the sequencing image is convolved with a corresponding sharpening mask to generate a corresponding feature map. Figure 10E illustrates the convolution operation for a particular section of the sequencing image.
[0240] It should be noted that each section of the sequencing image has a corresponding plurality of clusters. For example, assume that the first and second sections of the sequencing image are generated for the first and second color channels, respectively, and are generated from the same first subtile portion of the tile. Thus, both the first and second sections will have the same first plurality of clusters. In another example, assume that the third and fourth sections of the sequencing image are generated for the first and second color channels, respectively, and are generated from the same second subtile portion of the tile. Thus, both the third and fourth sections will have the same second plurality of clusters that are different from the first plurality of clusters.
[0241] The method 1100 then proceeds to 1116, where for each cluster in each feature map, a weighted feature value is assigned based on a suitable interpolation technique, so that each cluster has one or more corresponding weighted feature values for one or more color channels. For example, assuming an example of two color channels, each cluster is assigned two weighted feature values corresponding to the two color channels. Some exemplary interpolation techniques have been previously discussed herein, but any other interpolation technique not discussed herein may also be used.
[0242] Method 1100 then proceeds to 1120, where a base caller base calls each cluster based on the corresponding one or more weighted feature values for the corresponding cluster. For example, the weighted feature values of a cluster are provided as input to base caller 704 to generate a base call for that cluster. Base caller 704 can be a non-neural network based base caller or a neural network based base caller, examples of both of which are described in applications incorporated by reference herein, such as U.S. Patent Application No. 62 / 821,766 and U.S. Patent Application No. 16 / 826,168.
[0243] The method 1100 then proceeds to 1124, where the method 1100 proceeds to the next sequencing cycle of the sequencing operation, and the method 1100 loops back to 1104. This iteration of the method 1100 continues until all sequencing cycles of the sequencing operation have been completed.
[0244] FIG. 11B shows a comparison of performance results of the disclosed intensity extraction technique using a sharpening mask with various other intensity extraction techniques associated with base calling. The X-axis of the plot in FIG. 11B represents sequencing cycles, and the Y-axis of the plot represents error rates for base calling. For example, the red line in the plot is for base calling without using a sharpening mask for intensity extraction, and the green line in the plot is for base calling with a sharpening mask using an equalizer technique as disclosed in co-pending U.S. patent application Ser. No. 17 / 308,035, entitled "Equalization-Based Image Processing and Spatial Crosstalk Attenuator," which is incorporated by reference for all purposes as if fully set forth herein. The blue line in the plot is for base calling with a sharpening mask using the technique discussed herein with respect to FIGS. 7-11A. As can be seen, the blue line in the plot (for base calling with a sharpening mask using the techniques discussed) has a substantially lower error rate than the red line in the plot (for base calling without using a sharpening mask for intensity extraction).
[0245] FIG. 11B also shows a table showing the average error rate and average pass filter percentage for base calling. The pass filter percentage represents the percentage of clusters that have good quality base calls (e.g., base calls with a confidence level above a threshold percentage) and are base called. Thus, a higher pass filter percentage improves throughput. As can be seen, base calling with a sharpening mask using the techniques discussed herein (represented in the third column of the table) has a lower error rate and a better pass filter percentage than the scenario without a sharpening mask (represented in the first column of the table). Furthermore, base calling with sharpening masks using the techniques discussed herein (represented in the third column of the table) has a slightly lower error rate and a slightly higher pass filter percentage relative to a scenario using sharpening masks using an equalizer technique, as disclosed in co-pending U.S. patent application Ser. No. 17 / 308,035, entitled "Equalization-Based Image Processing and Spatial Crosstalk Attenuator," also referred to herein as "Scenario Using Sharpening Masks Using Equalizer Techniques." Note that, as discussed herein with respect to FIG. 11C, base calling with sharpening masks using the techniques discussed herein uses a smaller number of sharpening masks and has a faster execution time compared to a scenario using sharpening masks using an equalizer technique.
[0246] Figure 11C shows another comparison of performance results between the disclosed technique using a sharpening mask and various other techniques for base calling. Specifically, in Figure 11C, the base calling speed (or base calling execution time) for various scenarios is compared.
[0247] Two plots are shown, plot 1100c1 and 1100c2. Plot 1100c1 is generated using sequencing data from a new sequencing platform under development, and plot 1100c2 is generated using sequencing data from an Illumina NextSeq 1000 / NextSeq 2000 sequencer. Furthermore, in plot 1100c1, the number of wells or clusters per pixel is 0.3 and a kernel (or sharpening mask) size of 7×7 is used. In plot 1100c2, the number of wells or clusters per pixel is 0.1 and a kernel (or sharpening mask) size of 9×9 is used. Thus, plot 1100c1 has a higher cluster density compared to plot 1100c2.
[0248] As seen in plot 1100c2, base calling with a sharpening mask using the techniques discussed herein (represented in green) is 12.5% faster than the scenario using a sharpening mask using the equalizer technique. The performance improvement is even more noticeable in plot 1100c1, which has a higher cluster density. For example, as seen in plot 1100c1, base calling with a sharpening mask using the techniques discussed herein (represented in green) is 49.8% faster than the scenario using a sharpening mask using the equalizer technique.
[0249] Regular Cache Access A scenario using sharpening masks using an equalizer technique (such as disclosed in co-pending U.S. patent application Ser. No. 17 / 308,035, entitled "Equalization-Based Image Processing and Spatial Crosstalk Attenuator") would use different sharpening masks for different clusters, for example, depending on the sub-pixel location of the cluster relative to the center of the pixel. Thus, for example, three adjacent clusters on a tile of a flow cell could arguably use three different sharpening masks.
[0250] In contrast, in the case of the intensity extraction technique disclosed in the present disclosure (e.g., with respect to Figs. 7-11A), clusters on the entire subtile region of a tile use the same sharpening mask. For example, referring to Fig. 8A, all clusters on subtile 812a use the same sharpening mask 820Aa for color channel 802A. Thus, in one example, when processing clusters on subtile 812a for color channel 802A, the corresponding sharpening mask 820Aa is loaded into a cache, and the same sharpening mask 820Aa is repeatedly accessed from the cache during the convolution operation 1030Aa of Fig. 10F. In another example, once the sharpening mask 820Aa is loaded from the cache into the processing unit, the same sharpening mask 820Aa is used throughout the convolution operation 1030Aa. This improves the cache access pattern, which is relatively more regular (i.e., a regular cache access pattern), resulting in fewer or no cache misses.
[0251] In contrast, as discussed, in a scenario using a sharpening mask using an equalizer technique (as disclosed in co-pending U.S. patent application Ser. No. 17 / 308,035 entitled "Equalization-Based Image Processing and Spatial Crosstalk Attenuator"), different adjacent clusters on a tile of a flow cell may use correspondingly different sharpening masks, which results in relatively irregular cache access patterns and more cache misses. Thus, the intensity extraction technique disclosed in the present disclosure (e.g., with respect to Figs. 7-11A) is relatively faster than the equalizer-based intensity extraction technique disclosed in co-pending U.S. patent application Ser. No. 17 / 308,035, as also reflected in Fig. 11C.
[0252] On-line adaptation of sharpening mask coefficients. Note that each sharpening mask used in the convolution of the corresponding section of the sequencing image is a k×k matrix, where k is a suitable positive integer, such as 3, 5, 7, 9, or more. Assuming there are “m” color channels (where m is a positive integer, such as 1, 2, or more), there are m×k×k coefficients for each subtile of the tile. Assuming that the tile is subdivided into “n” parts (see, for example, FIG. 8A and FIG. 8B), there are n×m×k×k coefficients to be updated during the training process. Due to the relatively low values of n, m, and k, the number of coefficients to be updated is not significantly high. As a mere example, if two color channels are assumed, the sharpening masks are assumed to have dimensions of 3×3, and each tile is divided into 3×3 or 9 subtiles, the total number of coefficients of all sharpening masks is 2×9×3×3=162.
[0253] In addition to the offline training previously discussed herein, in one embodiment, the coefficients of the sharpening mask are also adaptively updated during the sequencing operation. For example, in the example discussed above, all the sharpening masks have only 162 coefficients, and it is relatively easy to adapt the 162 coefficients online, for example, when the sequencing operation is in progress (note, however, that the number 162 is merely an example). In contrast, a sharpening mask using an equalizer technique (as disclosed in co-pending U.S. patent application Ser. No. 17 / 308,035 entitled "Equalization-Based Image Processing and Spatial Crosstalk Attenuator") may have a larger number of sharpening mask parameters (such as 4050 in one example).
[0254] FIG. 12 shows a method 1200 of base calling based on convolution and subsequent interpolation of at least one section of a sequencing image to assign one or more weighted feature values to clusters, and base calling of clusters based on feature values attached to the assigned one or more weighted feature values, where coefficients of a sharpening mask are adaptively updated during sequencing operations.
[0255] In one example, online adaptation of the coefficients of the sharpening mask allows the coefficients to track changes in the operating parameters of the sequencing operation, such as changes in temperature, focus (e.g., optical distortion), chemistry, machine-specific variations, etc., while the sequencing machine is operating and the sequencing operation is cyclically progressing. For example, temperature (e.g., optical distortion), focus, chemistry, and / or machine-specific variations may at least partially invalidate offline training of the sharpening mask coefficients. For example, as the sequencing operation is cyclically progressing, online adaptation of the coefficients can bring the coefficients back on track to adapt to any changes to any parameters that affect the sequencing operation.
[0256] Method 1200 and method 1100 share various common operations that are labeled using the same labels in the two figures. For example, blocks 1104, 1108, 1112, 1116, 1120, and 1124 in both figures are the same and labeled similarly, and the operations of these blocks will not be discussed again with respect to FIG.
[0257] After completing the operations discussed with respect to blocks 1104-1120 (discussed with respect to FIG. 11A), the method 1200 of FIG. 12 proceeds to 1204, where it is determined whether the coefficients of the sharpening mask should be updated / trained using data from the current sequencing cycle. For example, the coefficients of the sharpening mask may not be updated every sequencing cycle of the sequencing operation. Rather, in one example, the coefficients of the sharpening mask may be updated during one or more selected sequencing cycles (not necessarily all) of the sequencing operation (although in another example, the coefficients may be updated during each sequencing cycle).
[0258] For example, the sequencing cycle in which the coefficients are to be updated may be implementation specific and may be a user configurable parameter. For example, as seen in Figure 14 later in this specification, results are presented for a scenario in which the sharpening mask coefficients are updated during sequencing cycles 10 and 30.
[0259] If the answer is "No" at 1204 (ie, the coefficients are not updated during the current sequencing operation), the method 1200 proceeds to 1124 and then loops back to 1104 as discussed with respect to the method 1100 of FIG. 11A.
[0260] If "yes" at 1204 (i.e., the coefficients should be updated during the current sequencing operation), the method 1200 proceeds to 1208 where the coefficients of the sharpening mask are updated or adapted using data from the current sequencing cycle C. In one example, the updated coefficients of the sharpening mask are applied for intensity extraction during sequencing cycle (C+2) and subsequent sequencing cycles. The adaptation or update process is discussed in more detail herein above with respect to Equation 1 and Figures 9A and 9B. The method then proceeds to block 1124 and then loops back to block 1104.
[0261] 12, updating or adapting the coefficients of the sharpening mask using data from sequencing cycle C occurs at least in part during a next iteration of at least some of the operations of blocks 1104-1102. That is, while the base caller is processing data from sequencing cycle (C+1), the base caller may in parallel perform updating of the coefficients using data from sequencing cycle C. Thus, in one example, the updated coefficients may not be applied to the images of sequencing cycle (C+1) but may be applied to the images of sequencing cycle (C+2).
[0262] Note that to base call the current sequencing cycle C, the intensities of sequencing cycle (C+1) must first be extracted. For example, the intensities of sequencing cycle (C+1) are used to correct for prephasing / phasing of sequencing cycle C. Further details regarding phasing and prephasing are discussed in co-pending U.S. Provisional Patent Application No. 63 / 228,954, entitled "Base Calling Using Multiple Base Caller Models," which is incorporated by reference for all purposes as if fully set forth herein.
[0263] 13 illustrates the adaptation of the sharpening mask coefficients used for intensity extraction. For example, at 1304, the base caller receives a sequencing image from the flow cell for sequencing cycle (C+1) and extracts the intensities using techniques disclosed herein (e.g., using convolution followed by interpolation). Note that intensity extraction for earlier cycles, such as sequencing cycle C, is assumed to be already completed when operation 1304 is performed for sequencing cycle (C+1).
[0264] At 1308, the base caller corrects the phase error of cycle (C+1). At 1312, the base caller corrects the pre-phasing error of sequencing cycle C, for example, using the extracted (and phasing corrected) intensities of sequencing cycle (C+1). At 1316, the base caller calls bases of various clusters for sequencing cycle C. At 1320, the base caller uses data from sequencing cycle C to adapt or update the coefficients of the sharpening mask. Finally, the updated coefficients of the sharpening mask are used for sequencing cycle (C+2) onwards. The actual adaptation or update process has been discussed previously herein with respect to Equation 1 and Figures 9A and 9B.
[0265] FIG. 14 shows a comparison of performance results of the disclosed intensity extraction technique using a sharpening mask and adaptation to another intensity extraction technique without adaptation. The plot and table in FIG. 14 were generated based on sequencing data from a NextSeq 1000 / NextSeq 2000 sequencer. The X-axis of the plot in FIG. 14 represents the sequencing cycle, and the Y-axis of the plot represents the error rate for the base call. For example, the red dotted line in the plot is for the base call without using the adaptation of the sharpening mask for intensity extraction, and the blue line in the plot is for the base call with the adaptation of the sharpening mask as disclosed in the present disclosure (see FIG. 12). The table in FIG. 14 compares the error rate and the pass filter percentage for the two scenarios. As can be seen in the table, when the adaptation to the sharpening mask is used, the average error rate improves by about 9.4%. In the example of FIG. 14, the adaptation is performed for sequencing cycles 10 and 30 of the non-indexed reads. The discontinuities in the graphs for sequencing cycle 150 and thereafter are due to index reads that occur during those sequencing cycles. Further details about index reads can be found in U.S. Provisional Patent Application No. 62 / 979,384, entitled "Artificial Intelligence-Based Base Calling of Index Sequences," filed February 20, 2020, which is incorporated herein by reference.
[0266] FIG. 15 shows a comparison of performance results of the disclosed intensity extraction technique using a sharpening mask and adaptation to another intensity extraction technique that does not use adaptation. The plot and table of FIG. 15 were generated based on sequencing data from a new sequencing platform under development by Illumina, Inc. (San Diego, Calif.). The X-axis of the plot of FIG. 15 represents sequencing cycles, and the Y-axis of the plot represents error rate for base calling. For example, the red line in the plot is for base calling using adaptation of the sharpening mask for intensity extraction as disclosed in this disclosure (see, for example, FIG. 12), and the blue line in the plot is for base calling without adaptation of the sharpening mask. The table of FIG. 15 compares the error rate and pass filter percentage for the two intensity extraction techniques. As can be seen in the table, when adaptation to the sharpening mask is used, the average error rate improves by about 23%, and the pass filter percentage also improves somewhat. The discontinuities in the graphs for sequencing cycle 150 and thereafter are due to index reads that occur during those sequencing cycles. Further details about index reads can be found in U.S. Provisional Patent Application No. 62 / 979,384, entitled "Artificial Intelligence-Based Base Calling of Index Sequences," filed February 20, 2020, which is incorporated herein by reference.
[0267] In this application, the terms "cluster", "well", "sample" and "fluorescent sample" are used interchangeably, as a well contains the corresponding cluster / sample / fluorescent sample. As defined herein, "sample" and its derivatives are used in the broadest sense and include any specimen, culture, etc. suspected of containing a target. In some embodiments, the sample includes DNA, RNA, PNA, LNA, chimeric or hybrid forms of nucleic acid. A sample can include any biological sample, clinical sample, surgical sample, agricultural sample, air sample or water sample that contains one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, fresh frozen or formalin-fixed paraffin-embedded nucleic acid sample. It is also envisioned that the sample can be derived from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples from a single individual such as a tumor sample and a normal tissue sample (matched), or a sample from a single source containing two different forms of genetic material such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample containing plant or animal DNA. In some embodiments, the source of nucleic acid material can include nucleic acid obtained from a newborn, such as that typically used for newborn screening.
[0268] The nucleic acid sample can include high molecular weight material such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another embodiment, the low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some embodiments, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resection, and other clinical or laboratory obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some embodiments, the sample can include nucleic acid molecules obtained from animals, such as human or mammalian sources. In another embodiment, the sample can include nucleic acid molecules obtained from non-mammalian sources, such as plants, bacteria, viruses, or fungi. In some embodiments, the source of the nucleic acid molecule can be a preserved or extinct sample or species.
[0269] Additionally, the methods and compositions disclosed herein may be useful for amplifying nucleic acid samples having low quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample may include nucleic acid obtained from a crime scene, from a missing persons DNA database, from a laboratory associated with a forensic investigation, or may include a forensic sample obtained by a law enforcement agency, one or more military forces, or such personnel. The nucleic acid sample may be crude DNA, including purified samples or lysates, for example, from buccal swabs, paper, cloth, or other substrates that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may include small or fragmented portions of DNA, such as genomic DNA. In some embodiments, the target sequence may be present in one or more bodily fluids, including, but not limited to, blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from hair, skin, tissue samples, autopsies, or remains of victims. In some embodiments, the nucleic acid comprising one or more target sequences may be obtained from deceased animals or humans. In some embodiments, the target sequence can comprise nucleic acid obtained from a non-human, such as a microorganism, a plant cell, or an entomological. In some embodiments, the target sequence or the amplified target sequence is for human identification. In some embodiments, the present disclosure generally relates to a method for identifying features of a forensic sample. In some embodiments, the present disclosure generally relates to a human identification method using one or more target-specific primers disclosed herein or one or more target-specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic sample or human identification sample comprising at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.
[0270] As used herein, the term "adjacent" when used in reference to two reaction sites means that there is no other reaction site between the two reaction sites. The term "adjacent" can have a similar meaning when used in reference to adjacent detection paths and adjacent photodetectors (e.g., adjacent photodetectors have no other photodetectors between them). In some cases, a reaction site may not be adjacent to another reaction site, but may still be in close proximity to another reaction site. A first reaction site may be in close proximity to a second reaction site if a fluorescent emission signal from the first reaction site is detected by a photodetector associated with the second reaction site. More specifically, a first reaction site may be in close proximity to a second reaction site if a photodetector associated with the second reaction site detects, for example, crosstalk from the first reaction site. Adjacent reaction sites may be contiguous so as to be adjacent to each other, or adjacent sites may be non-contiguous with an intervening space between them.
[0271] Upsampled Implementation In one embodiment, the image can be upsampled, for example, by using one or more interpolation or transposed convolution techniques to generate an upsampled image. In some embodiments, the image can have pixel resolution and the upsampled image can have sub-pixel resolution. In one embodiment, the convolution kernel / sharpening mask / mask can be upsampled, for example, by using one or more interpolation or transposed convolution techniques to generate an upsampled convolution kernel / sharpening mask / mask. In some embodiments, the convolution kernel / sharpening mask / mask can have pixel resolution and the upsampled convolution kernel / sharpening mask / mask can have sub-pixel resolution. The upsampled convolution kernel / sharpening mask / mask is then applied to the upsampled image to generate upsampled features. In some embodiments, the features can have pixel resolution and the upsampled features can have sub-pixel resolution. The upsampled features can then be analyzed for pixel-by-pixel correspondence to base call target clusters. In other embodiments, the upsampled features can then be analyzed for cluster-by-cluster correspondence to base call target clusters.
[0272] Technical Improvements and Terminology All literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and web pages, are expressly incorporated by reference in their entirety, regardless of the form of such literature and similar materials. In the event that one or more of the incorporated literature and similar materials differs or contradicts this application, including but not limited to, in terms defined, term usage, techniques described, etc., this application controls. Further information regarding terminology can be found in U.S. Nonprovisional Patent Application No. 16 / 826,168, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2019, and U.S. Provisional Patent Application No. 62 / 821,766, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2020.
[0273] The disclosed technology uses neural networks to improve the quality and quantity of nucleic acid sequence information that can be obtained from a nucleic acid template or its complement, e.g., a nucleic acid sample, such as a DNA or RNA polynucleotide or other nucleic acid sample. Thus, certain implementations of the disclosed technology provide higher throughput polynucleotide sequencing, e.g., higher rates of collection of DNA or RNA sequence data, greater efficiency in collecting sequence data, and / or lower costs of obtaining such sequence data, as compared to previously available methods.
[0274] The disclosed technology uses neural networks to identify centers of solid-phase nucleic acid clusters and analyze optical signals generated during sequencing of such clusters to unambiguously distinguish between adjacent, neighboring, or overlapping clusters and assign sequencing signals to single, discrete source clusters. These and related embodiments thus enable the recovery of meaningful information, such as sequence data, from regions of high-density cluster arrays where useful information may not have been previously obtained from such regions due to confounding effects of overlapping or closely spaced neighboring clusters, including the effects of overlapping signals (e.g., as used in nucleic acid sequencing).
[0275] As described in more detail below, in certain embodiments, a composition is provided that comprises a solid support immobilized with one or more nucleic acid clusters as provided herein.Each cluster comprises a plurality of immobilized nucleic acids of the same sequence, and has a distinguishable center with a detectable central label as provided herein, and the distinguishable center is distinguishable from the nucleic acids immobilized in the surrounding area within the cluster.Also described herein are methods for making and using such clusters with distinguishable centers.
[0276] Embodiments of the present disclosure will find use in many contexts where benefits derive from the ability to identify, determine, annotate, record, or otherwise assign the location of a substantially central location within a cluster, such as high throughput nucleic acid sequencing, development of image analysis algorithms for assigning optical or other signals to individual source clusters, and other applications where recognition of the centers of immobilized nucleic acid clusters is desirable and beneficial.
[0277] In certain embodiments, the present invention contemplates methods related to high-throughput nucleic acid analysis, such as nucleic acid sequencing (e.g., "sequencing"). Exemplary high-throughput nucleic acid analyses include, but are not limited to, de novo sequencing, resequencing, whole genome sequencing, gene expression analysis, gene expression monitoring, epigenetics analysis, genome methylation analysis, allele specific primer extension (APSE), genetic diversity profiling, whole genome polymorphism discovery and analysis, single nucleotide polymorphism analysis, hybridization-based sequencing, and the like. One of skill in the art will appreciate that a variety of different nucleic acids can be analyzed using the methods and compositions of the present invention.
[0278] Although the implementations of the present invention are described in the context of nucleic acid sequencing, they are applicable in any field where image data acquired at different times, spatial locations, or other temporal or physical aspects are analyzed. For example, the methods and systems described herein are useful in the fields of molecular and cell biology, where image data from microarrays, biological specimens, cells, organisms, etc. are acquired and analyzed at different times or perspectives. Images can be obtained using any number of techniques known in the art, including, but not limited to, fluorescent microscopy, optical microscopy, confocal microscopy, optical imaging, magnetic resonance imaging, tomographic scanning, etc. As another example, the methods and systems described herein can be applied when image data acquired, such as by surveillance, aerial, or satellite imaging techniques, are acquired and analyzed at different times or perspectives. The methods and systems are particularly useful for analyzing images acquired in a field of view, where the analytes observed remain in the same location relative to each other in the field of view. However, the analytes may have different properties in separate images, e.g., the analytes may appear different in separate images of the field of view. For example, the analytes may appear different in color for a given analyte detected in different images, which may indicate a change in the intensity of the signal detected for a given analyte in different images, or even the appearance of a signal for a given analyte in one image and the disappearance of the signal for that analyte in another image.
[0279] As used herein, the term "analyte" is intended to mean a point or region of a pattern that can be distinguished from other points or regions according to their relative location. An individual analyte can include one or more molecules of a particular type. For example, an analyte can include a single target nucleic acid molecule with a particular sequence, or an analyte can include several nucleic acid molecules with the same sequence (and / or its complementary sequence). Different molecules that are different analytes of a pattern can be differentiated from each other according to the location of the analyte in the pattern. Exemplary analytes include wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, pads of gel material on a substrate, or channels in a substrate.
[0280] Any of a variety of target analytes to be detected, characterized, or identified can be used in the devices, systems, or methods described herein. Exemplary analytes include, but are not limited to, nucleic acids (e.g., DNA, RNA, or analogs thereof), proteins, polysaccharides, cells, antibodies, epitopes, receptors, ligands, enzymes (e.g., kinases, phosphatases, or polymerases), small molecule drug candidates, cells, viruses, organisms, and the like.
[0281] The terms "analyte", "nucleic acid", "nucleic acid molecule", and "polynucleotide" are used interchangeably herein. In various embodiments, a nucleic acid may be used as a template (e.g., a nucleic acid template, or a nucleic acid complement complementary to a nucleic acid template) as provided herein for a particular type of nucleic acid analysis, including, but not limited to, nucleic acid amplification, nucleic acid expression analysis, and / or nucleic acid sequencing, or a suitable combination thereof. Nucleic acids in certain implementations include, for example, linear polymers of deoxyribonucleotides in 3'-5' phosphodiester, or deoxyribonucleic acid (DNA), such as single-stranded and double-stranded DNA, genomic DNA, copy DNA or complementary DNA (cDNA), recombinant DNA, or any form of synthetic or modified DNA. In other embodiments, the nucleic acid may be, for example, a linear polymer of ribonucleotides in 3'-5' phosphodiester or other combinations such as ribonucleic acid (RNA), e.g., single-stranded and double-stranded RNA, messenger (mRNA), copy RNA or complementary RNA (cRNA), alternatively spliced mRNA, ribosomal RNA, small nuclear RNA (snoRNA), microRNA (miRNA), small interfering RNA (sRNA), piwi RNA (piRNA), or any form of synthetic or modified RNA. The nucleic acids used in the compositions and methods of the invention may vary in length and may be intact or full-length molecules or fragments, or smaller portions of larger nucleic acid molecules. In certain embodiments, the nucleic acid may bear one or more detectable labels, as described elsewhere herein.
[0282] The terms "analyte", "cluster", "nucleic acid cluster", "nucleic acid colony", and "DNA cluster" are used interchangeably and refer to multiple copies of a nucleic acid template and / or its complement bound to a solid support. Typically, in certain preferred embodiments, a nucleic acid cluster comprises multiple copies of a template nucleic acid and / or its complement bound to a solid support via their 5' ends. The copies of the nucleic acid strands that make up a nucleic acid cluster may be in single-stranded or double-stranded form. The copies of the nucleic acid template present in a cluster may have nucleotides at corresponding positions that differ from each other due to, for example, the presence of a label moiety. The corresponding positions may also include analog structures with different chemical structures but similar Watson-Crick base pairing properties, such as in the case of uracil and thymine.
[0283] Colonies of nucleic acids may also be referred to as "nucleic acid clusters." Nucleic acid colonies may optionally be generated by cluster amplification or bridge amplification techniques, as described in more detail elsewhere herein. Multiple repeats of a target sequence may be present in a single nucleic acid molecule, such as a disrupting agent generated using a rolling circle amplification procedure.
[0284] The nucleic acid clusters of the present invention can have different shapes, sizes, and densities depending on the conditions used. For example, the clusters can have a substantially circular, multi-sided, donut-shaped, or ring-shaped shape. The diameter of the nucleic acid clusters can be designed to be about 0.2 μm to about 6 μm, about 0.3 μm to about 4 μm, about 0.4 μm to about 3 μm, about 0.5 μm to about 2 μm, about 0.75 μm to about 1.5 μm, or any intervening diameter. In certain embodiments, the diameter of the nucleic acid clusters is about 0.5 μm, about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The diameter of the nucleic acid clusters can be influenced by a number of parameters, including, but not limited to, the number of amplification cycles performed in the production of the clusters, the length of the nucleic acid template, or the density of the primers attached to the surface on which the clusters are formed. The density of the nucleic acid clusters is typically less than 0.1 / mm 2 , 1 / mm 2 , 10 / mm 2 , 100 / mm 2 , 1,000 / mm 2 , 10,000 / mm 2 ~100,000 / mm 2 The present invention, in part, relates to higher density nucleic acid clusters, e.g., up to 100,000 / mm 2 ~1,000,000 / mm 2 , and 1,000,000 / mm 2 ~10,000,000 / mm 2 The following are further intended:
[0285] As used herein, an "analyte" is a specimen or region of interest within a field of view. When used in connection with a microarray device or other molecular analysis device, an analyte refers to a region occupied by similar or identical molecules. For example, an analyte can be an amplification oligonucleotide, or any other group of polynucleotides or polypeptides having the same or similar sequence. In other embodiments, an analyte can be any element or group of elements that occupies a physical area on a specimen. For example, an analyte can be a parcel of land, a body of water, etc. When analytes are imaged, each analyte has some area. Thus, in many embodiments, an analyte is not simply a pixel.
[0286] The distance between analytes can be described in any number of ways. In some embodiments, the distance between analytes can be described from the center of one analyte to the center of another analyte. In other embodiments, the distance can be described from the edge of one analyte to the edge of another analyte, or between the outermost identifiable points of each analyte. The edge of the analyte can be described as a theoretical or actual physical boundary on the chip, or some point within the boundary of the analyte. In other embodiments, the distance can be described with respect to a fixed point on the specimen, or an image of the specimen.
[0287] Generally, some embodiments are described herein with respect to the analysis method. It will be understood that a system for performing the method in an automated or semi-automated manner is also provided. Thus, the present disclosure provides a neural network-based template generation and base calling system, the system includes a processor, a storage device, and a program for image analysis, the program includes instructions for performing one or more of the methods described herein. Thus, the methods described herein can be performed, for example, on a computer having components described herein or known in the art.
[0288] The methods and systems described herein are useful for analyzing any of a variety of objects. Particularly useful objects are solid supports or solid surfaces with attached analytes. The methods and systems described herein provide advantages when used with objects having repeating patterns of analytes in the xy plane. One example is a microarray having a collection of cells, viruses, nucleic acids, proteins, antibodies, carbohydrates, small molecules (such as drug candidates), biologically active molecules, or other analytes of interest.
[0289] There has been an increase in the number of applications of arrays with analytes having biological molecules such as nucleic acids and polypeptides. Such microarrays typically include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) probes. These are specific for nucleotide sequences present in humans and other organisms. In certain applications, for example, individual DNA or RNA probes can be attached to individual analytes of the array. Test samples, such as those from known humans or organisms, can be exposed to the array such that target nucleic acids (e.g., gene fragments, mRNA, or amplicons) hybridize to complementary probes at each analyte in the array. The probes can be labeled by target-specific processes (e.g., due to a label present on the target nucleic acid or due to an enzyme label of the probe or target present in hybridized form in the analyte). The analytes can then be examined by scanning a specific frequency of light over them to identify which target nucleic acids are present in the sample.
[0290] Biological microarrays can be used for gene sequencing and similar applications. In general, gene sequencing involves determining the order of nucleotides of a length of target nucleic acid, such as a fragment of DNA or RNA. A relatively short sequence is typically sequenced in each analyte, and the resulting sequence information may be used in a variety of bioinformatics methods to reliably determine the sequence of many broad lengths of genetic material from which the fragments are derived. Automated computer-based algorithms of signature fragments have been developed and have been used more recently in genome mapping, identification of genes and their functions, and the like. Microarrays are particularly useful for characterizing genome content, since there are many variants, which is an alternative to performing many experiments on individual probes and targets. Microarrays are an ideal format for carrying out such studies in a practical manner.
[0291] Any of a variety of analyte arrays (also referred to as "microarrays") known in the art can be used in the methods or systems described herein. A typical array includes analytes, each with an individual probe or a population of probes. In the latter case, the population of probes in each analyte is typically homogenous, with a single type of probe. For example, in the case of nucleic acid sequences, each analyte can have multiple nucleic acid molecules, each with a common sequence. However, in some embodiments, the population in each analyte of the array can be heterogeneous. Similarly, protein sequences can have analytes with a single protein or a population of proteins, typically, but not necessarily, with the same amino acid sequence. The probes can be attached to the surface of the array, for example, by covalently binding the probes to the surface or through non-covalent interactions between the probes and the surface. In some embodiments, the probes, such as nucleic acid molecules, can be attached to the surface through a gel layer, as described, for example, in U.S. Patent Application Serial No. 13 / 784,368 and U.S. Patent Application Publication No. 2011 / 0059865(A1), each of which is incorporated herein by reference.
[0292] Exemplary arrays include, but are not limited to, BeadChip arrays available from Illumina, Inc. (San Diego, Calif.) or others, such as those described in U.S. Pat. Nos. 6,266,459, 6,355,431, 6,770,441, 6,859,570, or 7,622,294, or WO 00 / 63437, in which the probes are attached to beads present on a surface (e.g., beads in wells on a surface), each of which is incorporated herein by reference. Further examples of commercially available microarrays that can be used include, for example, Affymetrix® GeneChip® microarrays or other microarrays synthesized according to a technique sometimes referred to as VLSIPS™ (Very Large Scale Immobilized Polymer Synthesis) technology. Spotted microarrays can also be used in the methods or systems according to some embodiments of the present disclosure. An exemplary spotted microarray is the CodeLink™ Array available from Amersham Biosciences. Another useful microarray is one manufactured using inkjet printing techniques, such as SurePrint™ Technology available from Agilent Technologies.
[0293] Other useful arrays include those used in nucleic acid sequencing applications.For example, arrays with amplicons of genome fragments (often referred to as clusters) are described in Bentley et al., Nature 456:53-59 (2008), International Publication No. WO 04 / 018497, International Publication No. WO 91 / 06678, International Publication No. WO 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281 or U.S. Patent No. 7,057,026, or U.S. Patent Application Publication No. 2008 / 0108082 (A1), each of which is incorporated herein by reference.Another type of array that is useful for nucleic acid sequencing is the array of particles produced from emulsion PCR technology. Examples are described in Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), WO 05 / 010145, U.S. Patent Application Publication No. 2005 / 0130173, or U.S. Patent Application Publication No. 2005 / 0064460, each of which is incorporated herein by reference in its entirety.
[0294] Arrays used for nucleic acid sequencing often have random spatial patterns of nucleic acid analytes. For example, the HiSeq or MiSeq sequencing platforms available from Illumina Inc (San Diego, Calif.) utilize flow cells in which nucleic acid sequences are formed by random seeding followed by bridge amplification. However, patterned arrays can also be used for nucleic acid sequencing or other analytical applications. Examples of patterned arrays, their manufacture and use are described in U.S. Patent Application Nos. 13 / 787,396, 13 / 783,043, 13 / 784,368, U.S. Patent Application Publication Nos. 2013 / 0116153(A1), and 2012 / 0316086(A1), each of which is incorporated herein by reference. Analytes in such patterned arrays can be used to capture single nucleic acid template molecules for subsequent formation of homogenous colonies, for example, via bridge amplification. Such patterned arrays are particularly useful for nucleic acid sequencing applications.
[0295] The size of the analytes on an array (or other object used in the methods or systems herein) can be selected to suit a particular application. For example, in some embodiments, the analytes of the array can have a size that accommodates only a single nucleic acid molecule. Surfaces with multiple analytes in this size range are useful for constructing arrays of molecules for detection with single molecule resolution. Analytes in this size range are also useful for use in arrays with analytes each comprising a colony of nucleic acid molecules. Thus, the analytes of the array can each be approximately 1 mm 2 Below, approximately 500μm 2 Below, approximately 100μm 2 Below, approximately 10μm 2 Below, approximately 1μm 2 Below, about 500nm 2 Less than or equal to about 100 nm 2 Below, approximately 10nm 2 Below, approximately 5nm 2 Less than or equal to 1 nm 2Alternatively or additionally, the analytes of the array can have an area of about 1 mm 2 or more, about 500μm 2 More than approximately 100μm 2 or more, about 10μm 2 or more, approximately 1μm 2 or more, about 500nm 2 More than 100 nm 2 or more, about 10nm 2 or more, about 5nm 2 More than or about 1 nm 2 And that's it. In practice, the analytes can have a size within a range between upper and lower limits selected from those exemplified above. Although some size ranges of surface analytes have been exemplified with respect to nucleic acids and nucleic acid scales, it will be understood that analytes in these size ranges can be used in applications that do not involve nucleic acids. It will be further understood that the size of the analytes need not necessarily be limited to the scale used for nucleic acid applications.
[0296] In embodiments involving objects having multiple analytes, such as arrays of analytes, the analytes can be distinct, separated by a space between each other. Arrays useful in the present invention can have analytes separated by an edge-to-edge distance of at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm or less. Alternatively or additionally, arrays can have analytes separated by an edge-to-edge distance of at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm or more. These ranges can apply to the average edge-to-edge spacing and edge-to-edge spacing of the analytes, as well as the minimum or maximum spacing.
[0297] In some embodiments, the analytes of the array do not have to be distinct, instead, adjacent analytes can abut each other. Whether the analytes are distinct or not, the size of the analytes and / or the pitch of the analytes can vary so that the array can have a desired density. For example, the average analyte pitch in a regular pattern can be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm or less. Alternatively or additionally, the average analyte pitch in a regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm or more. These ranges can also apply to the maximum or minimum pitch of a regular pattern. For example, the maximum analyte pitch in the regular pattern can be 100 μm or less, 50 μm or less, 10 μm or less, 5 μm or less, 1 μm or less, 0.5 μm or less, and / or the minimum analyte pitch in the regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more.
[0298] The density of analytes in an array can also be understood in terms of the number of analytes present per unit area. For example, the average density of analytes for an array is at least about 1×10 3 Analytes / mm 2 , 1×10 4 Analytes / mm 2 , 1×10 5 Analytes / mm 2 , 1×10 6 Inspection / mm 2 , 1×10 7 Analytes / mm 2 , 1×10 8 Analytes / mm 2 , or 1 × 10 9 Analytes / mm 2 Alternatively or additionally, the average density of analytes on the array can be at most about 1×10 9 Analytes / mm 2 , 1×10 8 Analytes / mm 2 , 1×10 7 Analytes / mm 2 , 1×10 6Analytes / mm 2 , 1×10 5 Analytes / mm 2 , 1×10 4 Analytes / mm 2 , or 1 × 10 3 Analytes / mm 2 It can be the following:
[0299] The above ranges can apply, for example, to all or a portion of a regular pattern that includes all or a portion of an array of analytes.
[0300] The analytes in the pattern can have any of a variety of shapes. For example, when viewed in a two-dimensional plane, such as on the surface of an array, the analytes may appear rounded, circular, elliptical, rectangular, square, symmetrical, asymmetrical, triangular, polygonal, etc. The analytes can be arranged in a regular repeating pattern, including, for example, a hexagonal or rectilinear pattern. The pattern can be selected to achieve a desired level of packing. For example, circular analytes are optimally packed in a hexagonal arrangement. Of course, other packing configurations can also be used for circular analytes, and vice versa.
[0301] A pattern can be characterized in terms of the number of analytes present in a subset that forms the smallest geometric unit of the pattern. A subset can include, for example, at least about 2, 3, 4, 5, 6, 10 or more analytes. Depending on the size and density of the analytes, a geometric unit can be on the order of 1 mm 2 , 500μm 2 , 100μm 2 , 50μm 2 , 10μm 2 , 1 μm 2 , 500 nm 2 , 100 nm 2 , 50 nm 2 , 10 nm 2 Alternatively or additionally, the geometric unit may be 10 nm 2 , 50 nm 2 , 100 nm 2 , 500 nm 2, 1 μm 2 , 10μm 2 , 50μm 2 , 100μm 2 , 500μm 2 , 1mm 2 The properties of the analytes in the geometric units, such as shape, size, pitch, etc., can be selected from those described herein more generally with respect to the analytes in the array or pattern.
[0302] An array with a regular pattern of analytes may be ordered with respect to the relative location of the analytes, but random with respect to one or more other characteristics of each analyte. For example, in the case of nucleic acid sequences, the nucleic acid analytes may be regular with respect to their relative location, but random with respect to the knowledge of the sequence regarding the nucleic acid species present in any particular analyte. As a more specific example, a nucleic acid sequence formed by seeding a repeating pattern of analytes with a template nucleic acid and amplifying the template in each analyte to form copies of the template in the analyte (e.g., via cluster amplification or bridge amplification) will have a regular pattern of nucleic acid analytes, but will be random with respect to the distribution of the sequence of the nucleic acid across the array. Thus, detection of the presence of nucleic acid material on an array can result in a repeating pattern of analytes, whereas sequence-specific detection can result in a non-repeating distribution of signals across the array.
[0303] It will be understood that descriptions of patterns, orders, randomness, etc. herein relate not only to analytes on an object, such as analytes on an array, but also to analytes in an image. Thus, the patterns, orders, randomness, etc. can be present in any of a variety of formats used to store, manipulate, or communicate image data, including, but not limited to, computer readable media or computer components such as graphical user interfaces or other output devices.
[0304] As used herein, the term "image" is intended to mean a representation of all or a part of an object. The representation may be an optically detected reproduction. For example, an image may be obtained from fluorescent, luminescent, scattering, or absorption signals. The part of the object present in the image may be the surface or other xy plane of the object. Typically, an image is a two-dimensional representation, but in some cases, the information in the image may be derived from three or more dimensions. An image need not include optically detected signals. Non-optical signals may be present instead. An image may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.
[0305] As used herein, an "image" refers to a reproduction or representation of at least a portion of a specimen or other object. In some embodiments, the reproduction is an optical reproduction, for example, produced by a camera or other optical detector. The reproduction can be a non-optical reproduction, for example, a representation of an electrical signal obtained from an array of nanopore analytes, or a representation of an electrical signal obtained from an ion-sensitive CMOS detector. In certain embodiments, non-optical reproductions can be excluded from the methods or devices described herein. The image can have a resolution that can distinguish analytes of a specimen that are present at any of a variety of intervals, including, for example, those that are less than 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm apart.
[0306] As used herein, "acquisition," "capture," and like terms refer to any part of the process of acquiring an image file. In some embodiments, data acquisition can include generating an image of the specimen, looking for a signal in the specimen, directing a detection device to look for or generate an image of the signal, providing instructions for further analysis or transformation of the image file, and instructions for any number of transformations or manipulations of the image file.
[0307] As used herein, the term "template" refers to a representation of the location or relationship between signals or analytes. Thus, in some embodiments, the template is a physical grid with a representation of signals corresponding to analytes in a specimen. In some embodiments, the template can be a chart, table, text file, or other computer file that indicates locations corresponding to analytes. In the embodiments presented herein, a template is generated to track the location of analytes of a specimen across a set of images of the specimen captured at different reference points. For example, the template can be a set of x, y coordinates or a set of values that describe the direction and / or distance of one analyte relative to another analyte.
[0308] As used herein, the term "analyte" can refer to an object or area of an object from which an image is captured. For example, in an embodiment where an image is taken from the surface of soil, a parcel of land can be an analyte. In other embodiments where biomolecule analysis is performed in a flow cell, the flow cell can be divided into any number of subdivisions, each of which can be an analyte. For example, the flow cell can be divided into various channels or lanes, and each lane can be further divided into 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 140, 160, 180, 200, 400, 600, 800, 1000 or more separate areas to be imaged. One example of a flow cell has eight lanes, with each lane divided into 120 analytes or tiles. In another embodiment, an analyte can be created in multiple tiles, or even the entire flow cell. Thus, the image of each specimen can represent a larger area of the surface being imaged.
[0309] It will be understood that references to ranges and sequential lists of numbers herein include not only the numbers recited but all real numbers between the recited numbers.
[0310] As used herein, a "reference point" refers to any temporal or physical distinction between images. In another preferred embodiment, the reference point is a time point. In a more preferred embodiment, the reference point is a time point or cycle during the sequencing reaction. However, the term "reference point" can include other aspects that distinguish or separate images, such as angle, rotation, time, or other aspects that can distinguish or separate the images.
[0311] As used herein, "subset of images" refers to a group of images in a set. For example, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or any number of images selected from a set of images. In certain other embodiments, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or less or any number of images selected from a set of images. In another preferred embodiment, the images are obtained from one or more sequencing cycles, with four images correlating to each cycle. Thus, for example, a subset may be a group of 16 images acquired over four cycles.
[0312] Base refers to a nucleotide base or nucleotide, (adenine), C (cytosine), T (thymine), or G (guanine). This application uses "base" and "nucleotide" interchangeably.
[0313] The term "chromosome" refers to the genetic carrier of the present invention of a living cell, derived from a chromatin strand containing DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is used herein.
[0314] The term "site" refers to a unique location (e.g., chromosome ID, chromosomal location and orientation) on a reference genome. In some embodiments, a site may be a residue, a sequence tag, or the location of a segment on a sequence. The term "locus" may be used to refer to a specific location of a nucleic acid sequence or polymorphism on a reference chromosome.
[0315] The term "sample" as used herein typically refers to a sample derived from a biological fluid, cell, tissue, organ, or organism containing the nucleic acid to be sequenced and / or phased, or a sample derived from a mixture of nucleic acids containing at least one nucleic acid sequence to be sequenced and / or phased. Such samples include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, fine needle biopsy samples (e.g., surgical biopsy, needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, tissue explants, organ cultures, and any other tissue or cell preparations thereof, or fractions or derivatives thereof. Samples are often taken from human subjects (e.g., patients), but samples can be taken from any organism that has chromosomes, including, but not limited to, dogs, cats, horses, goats, sheep, cows, pigs, etc. Samples can be used directly as obtained from a biological source, or after pretreatment to modify the characteristics of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, and the like.
[0316] The term "sequence" includes or refers to a chain of nucleotides linked together. The nucleotides can be based on DNA or RNA. It is understood that one sequence may include multiple subsequences. For example, a single sequence (e.g., a PCR amplicon) may have 350 nucleotides. A sample read may include multiple subsequences within these 350 nucleotides. For example, a sample read may include a first and a second flanking subsequence, e.g., having 20-50 nucleotides. The first and second flanking subsequences may be located on either side of a repeat segment with a corresponding subsequence (e.g., 40-100 nucleotides). Each of the flanking subsequences may include (or may include a portion of) a primer subsequence (e.g., 10-30 nucleotides). For ease of reading, the term "subsequence" is referred to as "sequence", but it is understood that the two sequences need not be separate from each other on a common strand. To distinguish the various sequences described herein, the sequences may be given different labels (e.g., target sequence, primer sequence, flanking sequence, reference sequence, etc.). Other terms, such as "allele," may be given different labels to distinguish similar entities. The application uses "read" and "sequence read" interchangeably.
[0317] The term "paired end sequencing" refers to a sequencing method in which both ends of a target fragment are sequenced. Paired end sequencing can facilitate the detection of genomic rearrangements and repeated segments, as well as the detection of gene fusions and novel transcripts. Methods of paired end sequencing are described in International Publication No. WO 07010252, International Application No. GB2007 / 003798, and US Patent Publication No. 2009 / 0088327, each of which is incorporated herein by reference. In one example, the sequence of operations can be performed as follows: (a) generating clusters of nucleic acids, (b) linearizing the nucleic acids, (c) hybridizing a first sequencing primer and repeatedly performing cycles of extension, scanning and deblocking as described above, (d) "flipping" the target nucleic acid on the flow cell surface by synthesizing a complementary copy, (e) linearizing the resynthesized strand, and (f) hybridizing a second sequencing primer and repeatedly performing cycles of extension, scanning and deblocking as described above. The flipping operation can deliver the reagents described above for a single cycle of bridge amplification.
[0318] The term "reference genome" or "reference sequence" refers to a specific known genomic sequence, either partial or complete, of any organism that can be used to reference an identified sequence from a subject. For example, reference genomes used for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information (ncbi.nlm.nih.gov). "Genome" refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences. A genome includes both genetic and non-coding sequences of DNA. A reference sequence may be larger than the reads aligned to it. For example, a reference sequence may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105 times larger, or at least about 106 times larger, or at least about 107 times larger. In one example, the reference genome sequence is of a full-length human genome. In another example, the reference genome sequence is limited to a specific human chromosome, such as chromosome 13. In some embodiments, the reference chromosome is a chromosomal sequence from the human genome version hg19. Such sequences are sometimes referred to as chromosomal reference sequences, although the term reference genome is intended to encompass such sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes of any species, partial chromosomal regions (strands, etc.), etc. In various embodiments, the reference genome is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a specific individual. In other embodiments, "genome" also covers so-called "graph genomes," which use a particular storage format and representation of genome sequences. In one embodiment, a graph genome stores data in a linear file. In another embodiment, a graph genome refers to a representation in which alternative sequencing (e.g., different copies of a chromosome with small differences) are stored as different paths in a graph.Additional information regarding the implementation of graph genomes can be found at https: / / www.biorxiv.org / content / biorxiv / early / 2018 / 03 / 20 / 194530.full.pdf, the contents of which are incorporated by reference in their entirety.
[0319] The term "read" refers to a collection of sequence data describing a fragment of a nucleotide sample or reference. The term "read" may refer to a sample read and / or a reference read. Typically, but not necessarily, a read represents a short sequence of consecutive base pairs in a sample or reference. A read may be symbolically represented by the base pair sequence (ATCG) of the sample or reference fragment. A read may be stored in a memory device and appropriately processed to determine whether the read matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing device or indirectly from stored sequence information about the sample. In some cases, the read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region that can be aligned and specifically assigned to, for example, a chromosome or genomic region or gene.
[0320] Next generation sequencing methods include, for example, sequencing by synthesis technology (Illumina), pyrosequencing (454), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences), and sequencing by ligation (SOLiD sequencing). Depending on the sequencing method, the length of each read can vary from about 30 bp to over 10,000 bp. For example, DNA sequencing using a SOLiD sequencer generates nucleic acid reads of about 50 bp. In another example, Ion Torrent Sequencing generates nucleic acid reads of up to 400 bp, and 454 pyrosequencing generates nucleic acid reads of about 700 bp. In yet another example, single molecule real-time sequencing can generate reads of 10,000 bp to 15,000 bp. Thus, in certain embodiments, the nucleic acid sequence reads have a length of 30 to 100 bp, 50 to 200 bp, or 50 to 400 bp.
[0321] The term "sample read", "sample sequence" or "sample fragment" refers to sequence data relating to a genomic sequence of interest from a sample. For example, a sample read includes sequence data from a PCR amplicon having forward and reverse primer sequences. The sequence data can be obtained from any selected sequence methodology. A sample read can be, for example, a sequencing-by-synthesis (SBS) reaction, a sequencing-ligation reaction, or any other suitable sequencing methodology in which it is desired to determine the length and / or identity of repetitive elements. A sample read can be a consensus (e.g., average or weighted) sequence derived from multiple sample reads. In certain embodiments, providing a reference sequence includes identifying a locus of interest based on primer sequences of a PCR amplicon.
[0322] The term "raw fragment" refers to sequence data of a portion of a genomic sequence of interest that at least partially overlaps a designated or secondary location of interest within a sample read or sample fragment. Non-limiting examples of raw fragments include double stitched fragments, simple stitched fragments, and simple non-stitched fragments. The term "raw" is used to indicate that the raw fragment includes sequence data that has some relationship to the sequence data in the sample read, regardless of whether the raw fragment shows supporting variants that correspond to and authenticate or confirm potential variants in the sample read. The term "raw fragment" does not indicate that the fragment necessarily includes supporting variants that validate the variant call in the sample read. For example, when a sample read is determined by a variant calling application to exhibit a first variant, the variant calling application can determine that one or more raw fragments lack a corresponding type of "supporting" variant that would otherwise be expected to occur in light of the variants in the sample read.
[0323] The terms "mapped", "aligned", "alignment", or "aligning" refer to the process of comparing a read or tag to a reference sequence, thereby determining whether the reference sequence contains the read sequence. If the reference sequence is read, the read may be mapped to the reference sequence, or in certain alternative embodiments, may be mapped to a specific location within the reference sequence. In some cases, alignment simply tells whether the read is a member of a particular reference sequence (i.e., whether the read is present or absent in the reference sequence). For example, alignment of a read to a reference sequence for human chromosome 13 tells whether the read is present in the reference sequence for chromosome 13. Tools that provide this information are sometimes called set membership testers. In some cases, alignment also indicates where in the reference sequence the read or tag map is. For example, if the reference sequence is the entire human genome sequence, alignment may indicate that the read is present on chromosome 13, and may further indicate that the read is on a specific strand and / or site of chromosome 13.
[0324] The term "indel" refers to the insertion and / or deletion of bases in the DNA of an organism. Microindels refer to indels that result in a net change of 1-50 nucleotides. Unless the length of the indel is a multiple of three, a frameshift mutation occurs in coding regions of the genome. Indels can be contrasted with point mutations. An indel insertion deletes a nucleotide from the sequence, whereas a point mutation is a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels can also be contrasted with Tandem Base Mutations (TBM), which can be defined as substitutions at adjacent nucleotides (mostly two adjacent nucleotides are replaced, although substitutions at three adjacent nucleotides have been observed).
[0325] The term "variant" refers to a nucleic acid sequence that differs from a nucleic acid reference. Exemplary nucleic acid sequence variants include, but are not limited to, single nucleotide polymorphisms (SNPs), short deletion and insertion polymorphisms (Indels), copy number variations (CNVs), microsatellite markers, or short tandem repeats and structural variants. Somatic variant calling is an effort to identify variants that are present at low frequency in a DNA sample. Calling somatic variants is of interest in the context of cancer treatment. Cancer is caused by the accumulation of mutations in DNA. DNA samples from tumors are generally heterogeneous, containing some normal cells, early stages of cancer progression (with fewer mutations), and some late-stage cells (with more mutations). Due to this heterogeneity, somatic mutations often appear at low frequency when sequencing tumors (e.g., from FFPE samples). For example, an SNV may be found in only 10% of the reads that cover a given base. A variant that is classified as somatic or germline by the variant classifier is also referred to herein as the "variant under test."
[0326] The term "noise" refers to erroneous variant calls that arise from one or more errors in the sequencing process and / or variant calling application.
[0327] The term "variant frequency" refers to the relative frequency of an allele (variant of a gene) at a particular locus in a population, expressed as a fraction or percentage. For example, the fraction or percentage may be the percentage of all chromosomes in a population that carry that allele. As an example, a sample variant frequency refers to the relative frequency of an allele / variant at a particular locus / position along a genomic sequence of interest across a "population" corresponding to the number of reads and / or samples obtained for the genomic sequence of interest from an individual. As another example, a baseline variant frequency refers to the relative frequency of an allele / variant at a particular locus / position along one or more baseline genomic sequences, where the relative frequency of an allele / variant at a particular locus / position along one or more baseline genomic sequences obtained for one or more baseline genomic sequences.
[0328] The term "Variant Allele Frequency (VAF)" refers to the proportion of sequenced reads that carry a variant divided by the overall coverage at the target position. VAF is a measure of the proportion of sequenced reads that carry a variant.
[0329] The terms "position," "designated position," and "locus" refer to the location or coordinates of one or more nucleotides within a nucleotide sequence. The terms "position," "designated position," and "locus" also refer to the location or coordinates of one or more base pairs in a sequence of nucleotides.
[0330] The term "haplotype" refers to a combination of alleles at adjacent sites on a chromosome that are inherited together. A haplotype, if present, may be one locus, several loci, or an entire chromosome, depending on the number of recombination events that have occurred between a given set of loci.
[0331] The term "threshold" herein refers to a number or values used as a cutoff to characterize a sample, a nucleic acid, or a portion thereof (e.g., a read). The threshold may vary based on empirical analysis. The threshold may be compared to a measured or calculated value to determine whether a source giving rise to such a value should be classified in a particular manner. The threshold may be identified empirically or analytically. The choice of threshold depends on the confidence with which a user desires the classification to be made. The threshold may be selected for a particular purpose (e.g., for a balance of sensitivity and selectivity). As used herein, the term "threshold" refers to a point at which the course of an analysis may change and / or an action may be triggered. A threshold does not have to be a predetermined number. Instead, the threshold may be, for example, a function based on multiple factors. The threshold may be adapted to the situation. Furthermore, the threshold may indicate an upper limit, a lower limit, or a range between limits.
[0332] In some embodiments, an index or score based on the sequencing data may be compared to a threshold. As used herein, the term "metric" or "score" may include a value or result determined from the sequencing data, or may include a function based on a value or result determined from the sequencing data. As with a threshold, an index or score may be adapted to the context. For example, an index or score may be a normalized value. As an example of a score or metric, one or more embodiments may use a count score in analyzing the data. The count score may be based on the number of sample reads. The sample reads may have been through one or more filtering stages such that the sample reads have at least one common characteristic or quality. For example, each of the sample reads used to determine the count score may be aligned with a reference sequence or may be assigned as a potential allele. The number of sample reads with a common characteristic may be counted to determine a read count. The count score may be based on the read count. In some embodiments, the count score may be a value equal to the read count. In other examples, the count score may be based on the read count and other information. For example, the counting score may be based on the read counts of a particular allele of a locus and the total number of reads of the locus. In some embodiments, the counting score may be based on the read counts of a locus and previously obtained data. In some embodiments, the counting score may be a normalized score between predetermined values. The counting score may also be a function of read counts from other loci of the sample, or read counts from other samples run simultaneously with the sample of interest. For example, the counting score may be a function of the read counts of a particular allele and read counts of other loci in the sample, and / or read counts from other samples. As an example, the read counts from other loci and / or read counts from other samples may be used to normalize the counting score for a particular allele.
[0333] The term "coverage" or "fragment coverage" refers to a count or other measure of multiple sample reads for the same fragment of a sequence. A read count may represent a count of the number of reads that cover the corresponding fragment. Alternatively, coverage may be determined by multiplying the read count by a specified factor based on historical knowledge, knowledge of the sample, knowledge of the locus, etc.
[0334] The term "read depth" (conventionally a number followed by "x") refers to the number of sequenced reads with overlapping alignment at the target position. It is often expressed as an average or percentage above a cutoff for a set of intervals (such as an exon, gene, or panel). For example, a clinical report may state that the panel average coverage is 1,105x with 98% of the targeted bases covered >100x.
[0335] The term "base call quality score" or "Q score" refers to a PHRED-scaled probability ranging from 0-50 that is inversely proportional to the probability that a single sequenced base is correct. For example, a T base call with a Q of 20 is considered 99.99% likely to be correct. Any base call with a Q<20 should be considered low quality, and any variant identified where a significant proportion of sequenced reads supporting the variant is low should be considered a potential false positive.
[0336] The term "variant read" or "variant read number" refers to the number of sequenced reads that support the presence of a variant.
[0337] With regard to "strandedness" (or DNA strandedness), the genetic message in DNA can be represented as the letters A, G, C, and T, e.g., 5'-AGGACA-3'. Often the sequence is written in the orientation shown here, i.e., 5' end to the left and 3' end to the right. Although DNA can occur as a single-stranded molecule (like certain viruses), we usually find DNA as a double-stranded unit. It has a double helix structure with two antiparallel strands. In this case, the word "antiparallel" means that the two strands run parallel but have opposite polarity. Double-stranded DNA is held together by base pairing, and pairing is always maintained such that adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). This pairing is called complementarity, and one DNA strand is said to be the complement of the other. Thus, double-stranded DNA can be represented as two strings, such as 5'-AGGACA-3' and 3'-TCCTGT-5'. Note that the two strands have opposite polarity. Thus, the strandedness of the two DNA strands can be referred to as a reference strand and its complement, a forward and reverse strand, a top and bottom strand, a sense and an antisense strand, or a Watson and Crick strand.
[0338] Read alignment (also called read mapping) is the process by which sequences in a genome are referred to as derived. Once the alignment is performed, the "mapping quality" or "mapping quality score (MAPQ)" of a given read quantifies the probability that its location on the genome is correct. Mapping quality is encoded on a phase scale, where P is the probability that the alignment is incorrect. The probability is P=10 (-MAQ / 10)MAPQ is calculated as: MAPQ = MAPQ ... The MAPQ value can be used as a quality control for the alignment results: the percentage of aligned reads with a MAPQ higher than 20 is usually for downstream analysis.
[0339] As used herein, a "signal" refers to a detectable event, such as, for example, an emission in an image, preferably an emission. Thus, in another preferred embodiment, a signal can represent any detectable emission (i.e., a "spot") captured in an image. Thus, as used herein, a "signal" can refer to both an actual emission from an analyte of a specimen, and a spurious emission that does not correlate with an actual analyte. Thus, a signal can result from noise and can be subsequently discarded as not representative of an actual analyte of a specimen.
[0340] As used herein, the term "clamp" refers to a group of signals. In certain embodiments, the signals are from different analytes. In another preferred embodiment, the signal clump is a group of signals that cluster together. In a more preferred embodiment, the signal clump represents the physical area covered by one amplification oligonucleotide. Each signal clump should ideally be observed as several signals (one per template cycle, possibly more due to crosstalk). Thus, overlapping signals are detected, where two (or more) signals are included in the template from the same signal clump.
[0341] As used herein, terms such as "minimum," "maximum," "minimize," "maximize," and grammatical variations thereof may include values that are not absolute maximums or minimums. In some implementations, the values include near maximums and minimums. In other examples, the values may include local maximums and / or local minimums. In some implementations, the values may include only absolute maximums or minimums.
[0342] As used herein, "crosstalk" refers to the detection of a signal in one image that is also detected in a separate image. In another preferred embodiment, crosstalk may occur when an emitted signal is detected in two separate detection channels. For example, if an emitted signal occurs in one color, the emission spectrum of that signal may overlap with another emitted signal in another color. In a preferred embodiment, the fluorescent molecules used to indicate the presence of nucleotide bases A, C, G, and T are detected in separate channels. However, because the emission spectra of A and C overlap, a portion of the C color signal may be detected during detection using a color channel. Thus, crosstalk between the A and C signals allows a signal from one color image to appear in the other color image. In some embodiments, there is G and T crosstalk. In some embodiments, the amount of crosstalk between channels is asymmetric. It will be understood that the amount of crosstalk between channels can be controlled by, among other things, the selection of signal molecules with appropriate emission spectra, as well as the selection of the size and wavelength range of the detection channel.
[0343] As used herein, "register," "registering," "registration," and similar terms refer to any process for correlating signals in an image or dataset with signals in an image or dataset from another time or perspective. For example, registration can be used to align signals from a set of images to form a template. In another example, registration can be used to align signals from other images to a template. One signal may be directly or indirectly registered to another signal. For example, a signal from image "S" may be directly registered to image "G." As another example, a signal from image "N" may be directly registered to image "G," or alternatively, a signal from image "N" may be registered to image "S" that was previously registered to image "G." That is, the signal from image "N" is indirectly registered to image "G."
[0344] As used herein, the term "fiducial" is intended to mean a distinguishable reference point in or on an object. The fiducial point may be, for example, a mark, a second object, a shape, an edge, an area, an irregularity, a channel, a pit, a post, etc. The fiducial point may be present in an image of the object or in another data set derived from detecting the object. The fiducial point may be specified by an x and / or y coordinate in the plane of the object. Alternatively or additionally, the fiducial point may be specified by a z coordinate orthogonal to the xy plane, for example defined by the relative location of the object and the detector. One or more coordinates for the fiducial point may be specified relative to one or more other analytes of the object, or an image or other data set derived from the object.
[0345] As used herein, the term "optical signal" is intended to include, for example, fluorescent, luminescent, scattering, or absorption signals. Optical signals may be detected in the ultraviolet (UV) range (approximately 200-390 nm), visible (VIS) range (approximately 391-770 nm), infrared (IR) range (approximately 0.771-25 micrometers), or other ranges of the electromagnetic spectrum. Optical signals may be detected in a manner that excludes all or part of one or more of these ranges.
[0346] As used herein, the term "signal level" is intended to mean an amount or quantity of detected energy or encoded information having a desired or predetermined characteristic. For example, optical signals can be quantified by one or more of intensity, wavelength, energy, frequency, power, brightness, etc. Other signals can be quantified according to characteristics such as voltage, current, electric field strength, magnetic field strength, frequency, power, temperature, etc. The absence of a signal is understood to be a signal level of zero, or a signal level that is not significantly distinguished from noise.
[0347] As used herein, the term "simulate" is intended to mean to create a representation or model of an object or action that predicts the characteristics of the object or action. The representation or model may often be distinguishable from the object or action. For example, the representation or model may be distinguishable with respect to one or more characteristics, such as color, texture, size, or strength of signal detected from all or a portion of the shape. In certain implementations, the representation or model may be idealized, exaggerated, muted, or incomplete compared to the object or action. Thus, in some implementations, the representation of the model may be representative of, for example, at least one of the above characteristics. The representation or model may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.
[0348] As used herein, the term "specific signal" is intended to mean a detected energy or encoded information that is selectively observed over other energies or information, such as background energy or information. For example, a specific signal can be an optical signal detected at a particular intensity, wavelength, or color; an electrical signal detected at a particular frequency, power, or field strength; or other signals known in the art related to spectroscopy and analytical detection.
[0349] As used herein, the term "swath" is intended to mean a rectangular portion of an object. A swath may be an elongated strip that is scanned by relative movement between the object and the detector in a direction parallel to the longest dimension of the strip. Generally, the width of the rectangular portion or strip is constant along its entire length. Multiple swaths of an object may be parallel to one another. Multiple swaths of an object may overlap one another, be adjacent to one another, or be separated from one another by interstitial regions.
[0350] As used herein, the term "variance" is intended to mean the expected and observed differences, or the difference between two or more observations. For example, variance can be the discrepancy between expected and measured values. Statistical functions such as standard deviation, squared standard deviation, coefficient of variation, etc. can be used to express variance.
[0351] As used herein, the term "xy coordinates" is intended to mean information that specifies a location, size, shape, and / or orientation within an xy plane. The information may be, for example, numerical coordinates in a Cartesian coordinate system. The coordinates may be provided relative to one or both of the x and y axes, or may be provided relative to another location within the xy plane. For example, the coordinates of an analyte of an object may specify the location of the analyte relative to a reference or other analyte locations of the object.
[0352] As used herein, the term "xy-plane" is intended to mean a two-dimensional region defined by linear axes x and y. When used with reference to a detector and an object observed by the detector, the region may be further specified as being orthogonal to the observation direction between the detector and the object being detected.
[0353] As used herein, the term "z-coordinate" is intended to mean information that specifies the location of a point, line, or region along an axis orthogonal to the xy plane. In certain embodiments, the z-axis is orthogonal to the region of the object observed by the detector. For example, the direction of the focus of an optical system may be specified along the z-axis.
[0354] In some embodiments, the acquired signal data is transformed using an affine transformation. In some such embodiments, the generation of the template uses the fact that the affine transformation between color channels is consistent between runs. Because of this consistency, a set of default offsets can be used in determining the coordinates of the analytes in the sample. For example, the default offset file can include relative transformations (shifts, scales, skews) for different channels relative to one channel, such as the A channel. However, in other embodiments, offsets between color channels drift during and / or between runs, making offset-driven template generation difficult. In such examples, the methods and systems provided herein can utilize offset template generation, which is further described below.
[0355] In some aspects of the above embodiments, the system may include a flow cell. In some aspects, the flow cell includes lanes or other configurations of tiles, at least some of the tiles including one or more analyte groups. In some aspects, the analytes include a plurality of molecules, such as nucleic acids. In certain aspects, the flow cell is configured to extend a primer that hybridizes to a nucleic acid in the analyte to deliver a labeled nucleotide base to a sequence of the nucleic acid, thereby creating a signal corresponding to the analyte including the nucleic acid. In preferred embodiments, the nucleic acids in the analyte are identical or substantially identical to each other.
[0356] In some of the image analysis systems described herein, each image in the set of images includes a color signal, with different colors corresponding to different nucleotide bases. In some embodiments, each image in the set of images includes a signal having a single color selected from at least four different colors. In some embodiments, each image in the set of images includes a signal having a single color selected from four different colors. In some of the systems described herein, nucleic acids can be sequenced by providing four different labeled nucleotide bases to the sequence of molecules to produce four different images, each image includes a signal having a single color, with the signal color being different for each of the four different images, thereby producing a cycle of four color images corresponding to the four possible nucleotides present at a particular position in the nucleic acid. In certain embodiments, the system includes a flow cell configured to deliver additional labeled nucleotide bases to the sequence of molecules, thereby producing a cycle of multiple color images.
[0357] In preferred embodiment forms, the methods provided herein may include determining whether a processor is actively collecting data or whether the processor is in a low activity state. Collecting and storing a large number of high quality images typically requires a large amount of storage capacity. Furthermore, once collected and stored, analysis of image data may be resource intensive and may impede processing power for other functions, such as collection and storage of additional image data. Thus, as used herein, the term "low activity state" refers to the processing power of a processor at a given time. In some embodiments, a low activity state occurs when a processor is not collecting and / or storing data. In some embodiments, a low activity state occurs when some data collection and / or storage occurs, but additional processing power remains such that image analysis can occur simultaneously without interfering with other functions.
[0358] As used herein, "identifying a conflict" refers to identifying a situation in which multiple processes are competing for a resource. In some such implementations, one process is given priority over another process. In some implementations, the conflict may relate to the need to give priority to the allocation of time, processing power, storage power, or any other resource that is given priority. Thus, in some implementations, when processing time or capacity is distributed between two processes, such as either analyzing a data set and acquiring and / or storing a data set, a discrepancy between the two processes exists that can be resolved by giving priority to one of the processes.
[0359] Also provided herein is a system for performing image analysis. The system may include a processor, a storage capacity, and a program for image analysis, the program including instructions for processing a first data set for storage and a second data set for analysis, the processing including acquiring and / or storing the first data set on the storage device, and analyzing the second data set when the processor is not acquiring the first data set. In certain aspects, the program includes instructions for identifying at least one instance of a conflict between acquiring and / or storing the first data set and analyzing the second data set, and priority is given to acquiring and / or storing the image data such that acquiring and / or storing the first data set is given priority. In certain aspects, the first data set includes an image file collected from an optical imaging device. In certain aspects, the system further includes an optical imaging device. In some aspects, the optical imaging device includes a light source and a detection device.
[0360] As used herein, the term "program" refers to instructions or commands for performing a task or process. The term "program" may be used interchangeably with the term "module." In certain implementations, a program may be a compilation of various instructions executed under the same command set. In other implementations, a program may refer to a separate batch or file.
[0361] The following are some of the surprising effects of utilizing the method and system for performing image analysis described herein. In some sequencing implementations, an important measure of the usefulness of a sequencing system is its overall efficiency. For example, the amount of mappable data produced per day, as well as the total cost of installing and operating the equipment, are important aspects of an economical sequencing solution. To reduce the time to generate mappable data and increase the efficiency of the system, real-time base calling can be enabled on the equipment computer and can run in parallel with sequencing chemistry and imaging. This allows data processing and analysis to be completed before sequencing chemistry finalization. In addition, it can reduce the storage required for intermediate data and limit the amount of data that needs to be moved across the network.
[0362] While the sequence output is increasing, the data per operation transferred from the system provided herein to the network and secondary analysis processing hardware is substantially reduced. By converting the data on the instrument computer (acquisition computer), the network load is dramatically reduced. Without these on-instrument, off-network data reduction techniques, the image output of the DNA sequencing instrument FRET would cripple most networks.
[0363] The widespread adoption of high-throughput DNA sequencing instruments has been driven in part by their ease of use, their support for a range of applications, and their suitability for virtually any laboratory environment. The highly efficient algorithms presented herein allow for the addition of significant analytical capabilities to a simple workstation capable of controlling the sequencing instrument. This reduction in computational hardware requirements has several practical advantages that will become even more important as sequencing output levels continue to increase. For example, by implementing image analysis and base calling in a simple tower, heat generation, laboratory footprint, and power consumption are kept to a minimum. In contrast, other commercial sequencing technologies have recently ramped up their computing infrastructure, with up to five times the processing power for primary analysis, before beginning to increase heat output and power consumption. Thus, in some embodiments, the computational efficiency of the methods and systems provided herein allows for their sequencing throughput to be increased while minimizing server hardware.
[0364] Thus, in some embodiments, the methods and / or systems presented herein function as a state machine, keeping track of the individual state of each specimen, and upon detecting that a specimen is ready to progress to the next state, taking appropriate action to advance the specimen to that state. A more detailed example of how a state machine monitors the file system to determine when a specimen is ready to progress to the next state according to a preferred embodiment is described below.
[0365] In preferred embodiments, the methods and systems provided herein are multi-threaded and can work with a configurable number of threads. Thus, for example, in the context of nucleic acid sequencing, the methods and systems provided herein can operate in the background during live sequencing operations for real-time analysis, or can operate using existing image data sets for offline analysis. In certain preferred embodiments, the methods and systems handle multi-threading by giving each thread its own subset of the specimens it is involved in. This minimizes the possibility of thread retention.
[0366] The disclosed method may include using a detection device to acquire a target image of an object, the image including a repeating pattern of analytes on the object. Detection devices capable of high resolution imaging of a surface are particularly useful. In certain embodiments, the detection device will have sufficient resolution to distinguish analytes at the densities, pitches, and / or analyte sizes described herein. Detection devices capable of acquiring images or image data from a surface are particularly useful. Exemplary detectors are those configured to acquire area images while maintaining a static relationship between the object and the detector. Scanning devices may also be used. For example, devices that acquire continuous area images (e.g., referred to as "step and shot" detectors) may be used. Also useful are devices that continuously scan points or lines on the surface of an object to accumulate data to build an image of the surface. Point scanning detectors may be configured to scan points (i.e., small detection areas) on the surface of an object via a raster motion in the xy plane of the surface. Line scanning detectors may be configured to scan a line along the y dimension of the surface of the object, with the longest dimension of the line occurring along the x dimension. It will be appreciated that the detection device, the object, or both may be moved to accomplish scanning detection. For example, detection devices that are particularly useful in nucleic acid sequencing applications are described in U.S. Patent Application Publication Nos. 2012 / 0270305(A1), 2013 / 0023422(A1), and 2013 / 0260372(A1), as well as U.S. Patent Nos. 5,528,050, 5,719,391, 8,158,926, and 8,241,573, each of which is incorporated herein by reference.
[0367] The embodiments disclosed herein may be implemented as a method, apparatus, system, or article of manufacture using programming or engineering techniques to create software, firmware, hardware, or any combination thereof. As used herein, the term "article of manufacture" refers to code or logic implemented in hardware or computer readable media, such as optical storage devices, as well as volatile or non-volatile memory devices. Such hardware includes, but is not limited to, field programmable gate arrays (FPGAs), coarse grained reconfigurable architectures (CGRAs), application-specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microprocessors, or other similar processing devices. In certain embodiments, the information or algorithms described herein reside in a non-transitory storage medium.
[0368] In certain embodiments, the computer-implemented methods described herein can occur in real-time while multiple images of an object are being acquired. Such real-time analysis is particularly useful for nucleic acid sequencing applications where nucleic acid sequences are subjected to repeated cycles of fluidic and detection steps. While analysis of sequencing data can often be beneficial to perform the methods described herein in real-time or in the background, it can be beneficial to perform the methods described herein while other data collection or analysis algorithms are in process. Examples of real-time analysis methods that can be used in the present methods are those commercially available from Illumina, Inc. (San Diego, Calif) and / or used in the MiSeq and HiSeq sequencing instruments described in U.S. Patent Application Publication No. 2012 / 0020537(A1), which is incorporated herein by reference.
[0369] An exemplary data analysis system, formed by one or more programmed computers, having programming stored on one or more machine-readable media with code executed to perform one or more steps of the methods described herein. In one embodiment, for example, the system includes an interface designed to enable networking of the system to one or more detection systems (e.g., optical imaging systems) configured to acquire data from a target object. The interface can receive and condition the data, if appropriate. In certain embodiments, the detection system outputs image data representing, for example, individual image elements or pixels that together form an image of an array or other object. The processor processes the received detection data according to one or more routines defined by the processing code. The processing code may be stored in various types of memory circuits.
[0370] According to currently contemplated embodiments, the processing code executed on the detection data includes data analysis routines designed to analyze the detection data to determine the locations of individual analytes visible or encoded in the data, as well as locations where no analyte is detected (i.e., where no analyte is present or no significant signal is detected from an existing analyte) and metadata. In certain embodiments, analyte locations in the array typically appear brighter than non-analyte locations due to the presence of a fluorescent dye attached to the imaged analyte. It will be understood that if an analyte is not present in the array, for example where a probe's target in the analyte is being detected, the analyte need not appear brighter than their surrounding areas. The color in which the individual analytes appear can be a function of the dye used, as well as the wavelength of light used by the imaging system for imaging purposes. Analytes without bound targets or without a specific label can be identified according to other characteristics, such as their expected location in the microarray.
[0371] Once the data analysis routine has located the individual analytes in the data, a value assignment may be performed. Generally, the value assignment assigns a digital value to each analyte based on the characteristics of the data represented by the detector element (e.g., pixel) at the corresponding location. That is, for example, when the imaging data is processed, the value assignment routine may be designed to recognize that a particular color or wavelength of light has been detected at a particular location. In a typical DNA imaging application, for example, the four common nucleotides are represented by four separate and distinguishable colors. Each color may then be assigned a value corresponding to that nucleotide.
[0372] As used herein, the terms "module," "system," or "system controller" may include hardware and / or software systems and circuits that operate to perform one or more functions. For example, a module, system, or system controller may include a computer processor, controller, or other log-based device that performs operations based on instructions stored on a tangible and non-transitory computer-readable storage medium, such as a computer memory. Alternatively, a module, system, or system controller may include a hardwired device that performs operations based on hardwired logic and circuitry. The modules, systems, or system controllers shown in the accompanying drawings may represent hardware and circuitry that operates based on software or hardwired instructions, software that instructs hardware to perform operations, or a combination thereof. A module, system, or system controller may include or represent hardware circuits or circuitry that includes and / or is connected to one or more processors, such as a computer microprocessor.
[0373] As used herein, the terms "software" and "firmware" are used interchangeably and include any computer program stored in memory that is executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely examples and are not intended to be limiting of the types of memory that can be used to store computer programs.
[0374] In the field of molecular biology, one of the processes for nucleic acid sequencing in use is sequence synthesis. This technique can be applied to highly parallel sequencing projects. For example, by using an automated platform, it is possible to perform millions of sequencing reactions simultaneously. Therefore, one embodiment of the present invention relates to an apparatus and method for collecting, storing, and analyzing image data generated during nucleic acid sequencing.
[0375] The enormous gain in the amount of data that can be collected and stored makes streamlined image analysis methods even more beneficial. For example, the image analysis methods described herein allow both designers and end users to make efficient use of existing computer hardware. Thus, methods and systems are presented herein that reduce the computational complexity of processing data in terms of rapidly increasing data output. For example, in the field of DNA sequencing, yields have been scaled up 15-fold in recent processes, and can reach hundreds of gigases in a single run of a DNA sequencing device. Large-scale genome-scale experiments are out of reach for most researchers if the computational infrastructure requirements increase proportionately. Thus, the generation of more raw sequence data increases the need for secondary analysis and data storage, making optimization of data transport and storage highly beneficial. Some embodiments of the methods and systems presented herein can reduce the time, hardware, networking, and laboratory infrastructure requirements required to produce usable sequence data.
[0376] The present disclosure describes various methods and systems for carrying out the methods. Some examples of the methods are described as a series of steps. However, it should be understood that the implementations are not limited to the specific steps and / or order of steps described herein. Steps may be omitted, steps may be modified, and / or other steps may be added. Furthermore, the steps described herein may be combined, steps may be performed simultaneously, steps may be divided into multiple substeps, steps may be performed in a different order, or a step (or a series of steps) may be re-performed iteratively. In addition, although different methods are described herein, it should be understood that in other implementations, different methods (or steps of different methods) may be combined.
[0377] In some embodiments, a processing unit, processor, module, or computing system that is "configured" to perform a task or operation may be understood to be specifically structured to perform the task or operation (e.g., having one or more programs or instructions that are adapted or intended to perform a task or operation, and / or having an arrangement of processing circuitry that is adapted or intended to perform a task or operation). For clarity and avoidance of doubt, a general purpose computer (which may be "configured to perform a task or operation when suitably programmed) is not configured as "configured" to perform a task or operation unless specifically programmed or structurally modified to perform the task or operation).
[0378] Moreover, the operations of the methods described herein may be sufficiently complex such that the operations cannot be performed by the average person or by a person skilled in the art in a commercially reasonable period of time. For example, the methods may rely on relatively complex calculations such that such a person cannot complete the methods in a commercially reasonable period of time.
[0379] Throughout this application, various publications, patents, or patent applications are referenced. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which this invention pertains.
[0380] The term "comprising," as used herein, is intended to be open ended, encompassing not only the recited elements, but any additional elements as well.
[0381] As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set. Exceptions may occur where express disclosure or context clearly dictates otherwise.
[0382] Although the invention has been described with reference to the above embodiments, it will be understood that various modifications can be made without departing from the invention.
[0383] The modules of the present application may be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some may be implemented on different processors or computers, or may be spread among many different processors or computers. In addition, it will be understood that some of the modules may be operated in parallel or in a different order than that shown in the figures, without affecting the functionality achieved. Also, as used herein, the term "module" may include "sub-modules," which may be considered herein to constitute a module. The blocks of the figures designated as modules may also be considered as flow chart steps in a method.
[0384] As used herein, "identification" of an item of information does not necessarily require direct specification of that item of information. Information may be "identified" within a field by simply referencing the actual information through one or more layers in one direction, or by identifying one or more items of different information that are sufficient to determine the actual item of information. Additionally, the term "designate" is used herein to mean the same thing as "identify."
[0385] As used herein, a given signal, event, or value is "dependent on a pre-decessor signal, event, or value of a pre-decessor signal, an event, or value that is affected by the given signal, event, or value. If an intervening processing element, step, or time period is present, then the given signal, event, or value may "exist" depending on a "pre-decessor signal, event, or value." If an intervening processing element or step combines two or more signals, events, or values, then the signal output of the processing element or step is considered to be "dependent" on each of the signal, event, or value inputs. If a given signal, event, or value is the same as a pre-decessor signal, event, or value, then this is simply considered to mean that the given signal, event, or value is "dependent" or "depends" on or "depends" on a "pre-decessor signal, event, or value." The "responsiveness" of a given signal, event, or value to another signal, event, or value is defined similarly.
[0386] As used herein, "concurrently" or "parallel" does not require exact simultaneity. It is sufficient if the assessment of one of the individuals begins before the assessment of another of the individuals is completed.
[0387] Computer Systems 16 is a computer system 1600 that can be used to implement the disclosed techniques. Computer system 1600 includes at least one central processing unit (CPU) 1672 that communicates with a number of peripheral devices via a bus subsystem 1655. These peripheral devices can include, for example, a storage subsystem 1610 including memory devices and a file storage subsystem 1636, user interface input devices 1638, user interface output devices 1676, and a network interface subsystem 1674. The input and output devices enable user interaction with computer system 1600. Network interface subsystem 1674 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0388] In one embodiment, the basecaller 704 is communicatively linked to the storage subsystem 1610 and the user interface input device 1638 .
[0389] User interface input devices 1638 can include pointing devices such as a keyboard, a mouse, a trackball, a touch pad, or a graphics tablet, a scanner, a touch screen integrated into a display, audio input devices such as a voice recognition system and a microphone, as well as other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and manners for inputting information into computer system 1600.
[0390] The user interface output devices 1676 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide non-visual displays such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and manners for outputting information from the computer system 1600 to a user or to another machine or computer system.
[0391] The storage subsystem 1610 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 1678.
[0392] The processor 1678 can be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 1678 can be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processors 1678 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX16 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.
[0393] The memory subsystem 1622 used in the storage subsystem 1610 may include multiple memories including a main random access memory (RAM) 1632 for storing instructions and data during program execution, and a read only memory (ROM) 1634 in which fixed instructions are stored. The file storage subsystem 1636 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular embodiment may be stored by the file storage subsystem 1636 in the storage subsystem 1610 or in another machine accessible by the processor.
[0394] Bus subsystem 1655 provides a mechanism for allowing the various components and subsystems of computer system 1600 to communicate with each other as intended. Although bus subsystem 1655 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0395] The computer system 1600 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 1600 shown in Figure 16 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 1600 can have more or fewer components than the computer system shown in Figure 16.
[0396] item The disclosed technology attenuates spatial crosstalk from sensor pixels using a sharpening mask-based image processing technique. The disclosed technology can be implemented as a system, method, or product. One or more features of the embodiments can be combined with the base embodiments. Non-mutually exclusive embodiments are taught as combinable. One or more features of the embodiments can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of the enumeration of repeating these options should not be interpreted as limiting the combinations taught in the preceding section. These descriptions are incorporated herein by reference into each of the following embodiments.
[0397] In one embodiment, the disclosed technology proposes a computer-implemented method for attenuating spatial crosstalk from sensor pixels.
[0398] The disclosed technology can be implemented as a system, a method, or a product. One or more features of the embodiments can be combined with the base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of the embodiments can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of the enumeration of repeating these options should not be interpreted as limiting the combinations taught in the preceding section. These descriptions are incorporated herein by reference into each of the following embodiments.
[0399] One or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of a computer product including a non-transitory computer-readable storage medium with computer usable program code for performing the illustrated method steps. Furthermore, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform the illustrated method steps. Furthermore, in another aspect, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of a means for performing one or more of the method steps described herein, which means can include (i) a hardware module, (ii) a software module running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of (i)-(iii) implementing a particular technique described herein, and the software module is stored in a computer-readable storage medium (or multiple such media).
[0400] The clauses described in this section can be combined as features. For the sake of brevity, combinations of features are not listed separately and are not repeated for each base set of features. The reader will understand how the features identified in the clauses described in this section can be easily combined with the sets of basic features identified as embodiments in other sections of this application. These clauses are not meant to be mutually exclusive, exhaustive, or restrictive, and the disclosed technology is not limited to these clauses, but rather encompasses all possible combinations, modifications, and variations within the scope of the technology described in the claims and their equivalents.
[0401] Other implementations of the provisions described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to implement any of the provisions described in this section. Yet another implementation of the provisions described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to implement any of the provisions described in this section.
[0402] The present inventors disclose the following items: 1. A computer-implemented method for base calling, the method comprising: accessing a section of an image output by the biosensor, the section of the image including a plurality of pixels indicative of intensity radiation values from a plurality of clusters within the biosensor and from locations within the biosensor adjacent to the plurality of clusters, the plurality of clusters including a target cluster; convolving a section of the image with a convolution kernel to generate a feature map including a plurality of features having a corresponding plurality of feature values; assigning a weighted feature value to a target cluster, the weighted feature value being based on one or more feature values of the feature map; and processing the weighted feature values assigned to the target clusters to base call the target clusters. 2. The section of the image is a first section generated from a first portion of a flow cell of a biosensor, the convolution kernel is a first convolution kernel, the plurality of clusters is a first plurality of clusters, the plurality of pixels is a first plurality of pixels, the feature map is a first feature map, the plurality of feature values is a first plurality of feature values, the target cluster is a first target cluster, and the weighted feature value is a first weighted feature value, and the method comprises: accessing a second section of an image output by a second portion of the flow cell of the biosensor, the second section of the image including a second plurality of pixels indicative of intensity radiation values from a second plurality of clusters within the biosensor and from locations within the biosensor adjacent to the second plurality of clusters, the second plurality of clusters including a second target cluster; convolving a second section of the image with a second convolution kernel different from the first convolution kernel to generate a second feature map including a second plurality of features having a corresponding second plurality of feature values; assigning a second weighted feature value to a second target cluster, the second weighted feature value being based on one or more feature values of a second plurality of feature values of the second feature map; 2. The method of claim 1, further comprising processing the second weighted feature values assigned to the second target cluster to base call the second target cluster. 3. The method of claim 2, wherein a tile of a flow cell of the biosensor is divided into k x k parts, k is a positive integer, and the first part and the second part are two parts of the k x k parts of the tile. 4. The method of claim 3, wherein k is one of 3, 5, or 9. 5. The method of any one of clauses 1-4, further comprising capturing an image within the biosensor using a fully automated image capture system. 6. The method of claim 2, wherein a tile of a flow cell of the biosensor is divided into 1 x k parts, k is a positive integer, and the first part and the second part are two parts of the 1 x k parts of the tile. 7. 7. The method of any one of clauses 1-6, further comprising capturing an image within the biosensor using a line-scan image capture system. 8. a tile of a flow cell of the biosensor is divided into a plurality of portions, the plurality of portions including a first type of portion and a second type of portion, the second type of portion being periodically interleaved within the first type of portion; the first moiety is one of a first type of moiety; 8. The method of any one of clauses 2 to 7, wherein the second moiety is one of a second type of moiety. 9. 9. The method of any one of clauses 1 to 8, further comprising capturing an image within the biosensor using one or more CMOS (complementary metal oxide semiconductor) sensors. 10. a tile of a flow cell of the biosensor is divided into a plurality of portions including a first portion, a second portion, and a third portion; a first section of the image generated from a first portion of the tile of the flow cell is convolved with a first convolution kernel; a second section of the image generated from a second portion of the flow cell tile is convolved with a second convolution kernel; 10. The method of any one of clauses 2 to 9, wherein a third section of the image generated from a third portion of the tile of the flow cell is convolved with a third convolution kernel different from each of the first convolution kernel and the second convolution kernel. 11. A method comprising: the section of the image is a first section generated for a first color channel from a first portion of a flow cell; the convolution kernel is a first convolution kernel; the plurality of pixels is a first plurality of pixels; the feature map is a first feature map; the plurality of feature values is a first plurality of feature values; and the weighted feature value is a first weighted feature value; and accessing a second section of the image generated for a second color channel from the first portion of the flow cell, the second section of the image including a second plurality of pixels indicative of intensity radiation values from a plurality of clusters within the biosensor and from locations within the biosensor adjacent to the plurality of clusters; convolving a second section of the image with a second convolution kernel different from the first convolution kernel to generate a second feature map including a second plurality of features having a corresponding second plurality of feature values; assigning a second weighted feature value to the target cluster, the second weighted feature value being based on one or more feature values of the second plurality of feature values of the second feature map; 10. The method of any one of clauses 1-9, further comprising processing the first weighted feature value and the second weighted feature value assigned to the target cluster to base call the target cluster. 12. a first section of the image for a first color channel is convolved with a first convolution kernel; A second section of the image for a second color channel is convolved with a second convolution kernel; 12. The method of claim 11, wherein a third section of the image for a third color channel is convolved with a third convolution kernel different from each of the first convolution kernel and the second convolution kernel. 13. Assigning weighted feature values to target clusters 13. The method of any one of clauses 1-12, comprising assigning weighted feature values to target clusters based on sub-pixel or sub-feature locations of the target clusters. 14. The method of claim 13, wherein the sub-pixel location of the target cluster includes ...
Claims
1. One or more processors connected to a memory, and computer instructions, a system comprising, when the computer instructions are executed by the one or more processors, the system is caused to, access an image section of an image drawing signal from a cluster, filter a set of feature values regarding a feature map corresponding to the image section from the image section, determine corresponding feature values regarding a target cluster based on the set of feature values from the feature map, generate a base call regarding the target cluster by applying the corresponding feature values to a subset of signals regarding the target cluster, A system.
2. when executed by the one or more processors, the system is caused to, determine the corresponding feature values regarding the target cluster by determining weighted feature values regarding the target cluster based on neighboring feature values from the feature map corresponding to a neighboring cluster near the target cluster, and further includes computer instructions for causing, The system according to claim 1.
3. when executed by the one or more processors, the system is caused to, access the image section of the image by accessing an image in an imaging channel of an array determination device, and generate the feature map regarding the image section in the imaging channel, and further includes computer instructions for causing, The system according to claim 1.
4. when executed by the one or more processors, the system is caused to, determine the corresponding feature values regarding the target cluster by complementing one or more feature values from the set of feature values from the feature map regarding the image section of the image captured during a single array determination cycle, and generate the base call regarding the target cluster by applying the corresponding feature values to a subset of the signals from the image captured during the single array determination cycle, and further includes computer instructions for causing, The system according to claim 1.
5. The system according to claim 1, wherein the image section of the image includes a sub - tile area related to a sub - tile of a flow cell.
6. When executed by the one or more processors, the system is caused to access an additional image section of the image rendering signal from one or more clusters, filter from the additional image section a set of additional feature values related to an additional feature map corresponding to the additional image section, determine an additional corresponding feature value related to an additional target cluster based on the set of additional feature values from the additional feature map, generate an additional base call related to the additional target cluster by applying the additional corresponding feature value to a signal related to the additional target cluster, The system according to claim 1, further comprising computer instructions to cause the above to be executed.
7. When executed by the one or more processors, the system is caused to filter the set of feature values related to the feature map by extracting one or more of the signals using a sharpening operation to generate one or more feature values of the set of feature values related to the feature map. The system according to claim 1, further comprising computer instructions.
8. When executed by the one or more processors, the system is caused to generate the feature map by convolving the image section with a corresponding convolution kernel in a set of convolution kernels. The system according to claim 1, further comprising computer instructions.
9. When executed by the one or more processors, the system is caused to filter the set of feature values related to the feature map corresponding to the image section using a convolutional neural network, and determine the corresponding feature value related to the target cluster based on the set of feature values from the feature map using the convolutional neural network. The system according to claim 1, further comprising computer instructions to cause the above to be executed.
10. When executed by one or more processors, the system is caused to filter the set of feature values related to the feature map corresponding to the image section using a convolutional neural network, and determine the corresponding feature value related to the target cluster based on the set of feature values from the feature map using the convolutional neural network. The system according to claim 1, further comprising computer instructions to cause the above to be executed.
11. When executed by one or more processors, the system is caused to
10. When executed by one or more processors, the system is caused to Accessing an image section of an image rendering signal from a cluster, Filtering, from the image section, a set of feature values regarding a feature map corresponding to the image section, Determining corresponding feature values regarding a target cluster based on the set of feature values from the feature map, Generating a base call regarding the target cluster by applying the corresponding feature values to a subset of signals regarding the target cluster, A non-transitory computer-readable storage medium storing computer instructions. [
11. ] When executed by the one or more processors, the system is caused to Further store computer instructions for determining the corresponding feature values regarding the target cluster by determining weighted feature values regarding the target cluster based on neighboring feature values from the feature map corresponding to a neighboring cluster in the vicinity of the target cluster, The non-transitory computer-readable storage medium according to claim 10. [
12. ] When executed by the one or more processors, the system is caused to Access the image section of the image by accessing an image in an imaging channel of an array determination device, Further store computer instructions for generating the feature map regarding the image section in the imaging channel, The non-transitory computer-readable storage medium according to claim 10. [
13. ] When executed by the one or more processors, the system is caused to Determine the corresponding feature values regarding the target cluster by complementing one or more feature values from the set of feature values from the feature map regarding the image section of the image captured during a single array determination cycle, Further store computer instructions for generating the base call regarding the target cluster by applying the corresponding feature values to a subset of the signals from the image captured during the single array determination cycle, The non-transitory computer-readable storage medium according to claim 10. [
14. ] The non-transitory computer-readable storage medium according to claim 10, wherein the image section of the image includes a sub-tile region regarding a sub-tile of a flow cell.
15. When executed by the one or more processors, cause the system to access an additional image section of the image rendering signal from one or more clusters; filter from the additional image section a set of additional feature values regarding an additional feature map corresponding to the additional image section; determine additional corresponding feature values regarding an additional target cluster based on the set of additional feature values from the additional feature map; generate an additional base call regarding the additional target cluster by applying the additional corresponding feature values to a signal regarding the additional target cluster; further store computer instructions for causing execution; The non-transitory computer-readable storage medium according to claim 10.
16. A method executed by a computer, comprising: accessing an image section of an image rendering signal from a cluster; filtering from the image section a set of feature values regarding a feature map corresponding to the image section; determining corresponding feature values regarding a target cluster based on the set of feature values from the feature map; generating a base call regarding the target cluster by applying the corresponding feature values to a subset of signals regarding the target cluster.
17. The method executed by a computer according to claim 16, wherein filtering the set of feature values regarding the feature map includes extracting one or more signals among the signals by using a sharpening operation to generate one or more feature values of the set of feature values regarding the feature map. The method executed by a computer according to claim 16.
18. The method executed by a computer according to claim 16, further comprising generating the feature map by convolving the image section with a corresponding convolution kernel in a set of convolution kernels. The method executed by a computer according to claim 16.
19. Filtering the set of feature values regarding the feature map corresponding to the image section by using a convolutional neural network; using the convolutional neural network to determine corresponding feature values for the target cluster based on the set of feature values from the feature map, and The method executed by a computer according to claim 16. [
20. ] Accessing additional image sections of the image rendering signal from one or more clusters; filtering, from the additional image sections, a set of additional feature values for an additional feature map corresponding to the additional image sections; determining additional corresponding feature values for an additional target cluster based on the set of additional feature values from the additional feature map; generating an additional base call for the additional target cluster by applying the additional corresponding feature values to a signal for the additional target cluster; further comprising; The method executed by a computer according to claim 16.