Increasing sequencing flux in next generation sequencing of three-dimensional samples
By acquiring the flow cell image stack in three-dimensional samples and segmenting high-density samples using barcode sequences, the problems of limited sequencing throughput and low primer use efficiency in traditional sequencing techniques are solved, achieving more efficient sequencing and cost reduction.
Patent Information
- Application Number
- CN202380080822.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-06
- Filing Date
- 2023-09-22
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional three-dimensional sequencing techniques cannot generate reliable sequencing results in spatially overlapping clusters or communities, and it takes extra time and effort to sequence different clusters or communities using different sequencing primers.
By acquiring the stack of flow cell images in three-dimensional samples, using optical systems and computational methods for accurate base recognition, improving sequencing throughput, and dividing high-density samples into subsets by using barcode sequences to reduce primer usage, cost and time consumption.
A higher sequencing throughput than traditional methods is achieved, improving the efficiency of the sequencing system, reducing costs, and being able to process three-dimensional samples with higher spatial density.
Smart Images

Figure CN120303414A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 409,546, filed on September 23, 2022, and U.S. Provisional Patent Application No. 63 / 413,864, filed on October 6, 2022, which are hereby incorporated by reference in their entirety. Field of the Invention
[0003] The present disclosure generally relates to three-dimensional sequencing, and more particularly to sequencing high-density 3D samples that spatially overlap during DNA sequencing. Background Art
[0004] In next-generation sequencing (NGS) or NGS-like applications such as sequencing-by-synthesis, sequencing-by-ligation, or affinity sequencing, to identify the sequence of a target nucleic acid, new strands are synthesized one nucleotide base at a time. During each sequencing cycle, one base is attached to any given strand. During the imaging step of each cycle, an image is recorded. A base identification algorithm is applied to the image to "read" the consecutive signals from each cluster or colony, and the optical signals are converted into an identification of the nucleotide base sequence added to each DNA fragment.
[0005] When clusters or colonies spatially overlap, traditional 3D sequencing may not be able to generate reliable sequencing results. Thus, the sequencing throughput is limited by the number of clusters or colonies that can be spatially separated. Additionally, sequencing different clusters or colonies using different sequencing primers requires additional time and effort to hybridize and block certain colonies and to hybridize the primers to other colonies. Summary of the Invention
[0006] Provided herein are embodiments of systems, devices, methods, and / or computer program products, and / or combinations and sub-combinations thereof, that are capable of performing 3D sequencing of samples such as in situ cells or tissues with increased sequencing throughput compared to existing systems and methods.
[0007] As such specific applications, embodiments of methods, systems, and media for performing 3D sequencing of in situ samples are disclosed herein such that a higher sequencing throughput than traditional sequencing throughput can be achieved.
[0008] Other embodiments of these aspects include corresponding computer systems, devices, and computer program products recorded on computer storage devices, which are configured, alone or in combination, to perform the actions of the method. For a computer system that has been configured or is to be configured to perform operations or actions, software, firmware, hardware, or a combination thereof that causes the computer system to perform the operations or actions has been installed on the computer system. For a computer program product that has been configured or is to be configured to perform operations or actions, the computer program product includes instructions that cause a hardware processor to perform the operations or actions when executed by the hardware processor.
[0009] Additional embodiments, features, and advantages of the present disclosure, as well as the structure and operations of the embodiments of the present disclosure, are described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings incorporated herein and forming a part of this specification illustrate embodiments of the present disclosure and, together with the specification, further serve to explain the principles of the present disclosure and enable those skilled in the art to make and use the embodiments.
[0011] Figure 1 A block diagram of a system for performing 3D base identification of a flow cell image according to some embodiments is shown.
[0012] Figures 2A - 2C Exemplary flow cell images, processed images, and filtered images for 3D base identification according to some embodiments are shown.
[0013] Figures 3A - 3C Exemplary flow cell images, processed images, and filtered images for 3D base identification according to some embodiments are shown.
[0014] Figures 3D - 3E Exemplary flow cell images for 3D base identification according to some embodiments and their corresponding filtered images are shown.
[0015] Figure 4 A block diagram of a computer system for performing sequencing analysis and / or base identification according to some embodiments is shown.
[0016] Figure 5 Exemplary projection images of flow cell images taken at different axial positions of a 3D sample according to some embodiments are shown.
[0017] Figure 6A A flowchart of an exemplary method for performing 3D base identification of a flow cell image according to some embodiments is shown.
[0018] Figure 6BFlowchart of an exemplary method for performing 3D base identification of a flow cell image according to some embodiments.
[0019] Figures 7A - 7B Shows an exemplary registration of a sequencing image of a community with a cell staining image.
[0020] Figure 8A Shows a schematic diagram of a flow cell image, sub - tiles, and regions of a community according to some embodiments.
[0021] Figure 8B Shows a schematic diagram of a portion of a flow cell with multiple tiles according to some embodiments.
[0022] Figure 9 Displays a flowchart of a method for performing image registration of a flow cell image according to some embodiments.
[0023] Figures 10A - 10B Shows a schematic diagram of an image transformation and corresponding 2D displacement according to some embodiments.
[0024] Figures 11A - 11C Shows exemplary barcode sequences according to some embodiments, each of which can be used to uniquely identify nucleotides corresponding to DNA or RNA fragments.
[0025] Figures 12A - 12C Shows, according to some embodiments, a community being contacted in an exemplary sequencing cycle prior to imaging with 4 different types of affinity agents ( Figure 12A ), 3 different types of affinity agents ( Figure 12B ), and 4 different types of affinity agents (including one type with a "dark" fluorescent dye)( Figure 12C ).
[0026] Figure 13 Shows exemplary barcode sequences of 4 subsets of a community and imaging of individual subsets at different sequencing cycles according to some embodiments.
[0027] Figure 14 Is a schematic diagram showing an exemplary linear single - stranded library molecule (1101), which includes: a surface - pinned primer binding site (1121); an optional left unique identifier sequence (1181); a left index sequence (1161); a forward sequencing primer binding site sequence (1141); an insertion region with a sequence of interest (1111); a reverse sequencing primer binding site sequence (1151); a right index sequence (1171); and a surface capture primer binding site (1131).
[0028] Figure 15Is a schematic diagram showing an exemplary linear single-stranded library molecule (1101), which includes: a surface-pinned primer binding site (1121); a left index sequence (1161); a forward sequencing primer binding site sequence (1141); an insertion region having a sequence of interest (1111); a reverse sequencing primer binding site sequence (1151); a right index sequence (1171); an optional right unique identifier sequence (1190); and a surface capture primer binding site (1131).
[0029] Figure 16 Is a schematic diagram of various exemplary configurations of multivalent molecules. Left (Class I): A schematic diagram of a multivalent molecule having a "starburst" or "helter-skelter" configuration. Middle figure (Class II): A schematic diagram of a multivalent molecule having a dendrimer configuration. Right figure (Class III): A schematic diagram of multiple multivalent molecules formed by the reaction of streptavidin with a 4-arm or 8-arm PEG-NHS having biotin and dNTP. Nucleotide units are designated as 'N', biotin is designated as 'B', and streptavidin is designated as 'SA'.
[0030] Figure 17 Is a schematic diagram of an exemplary multivalent molecule comprising a universal core attached to multiple nucleotide arms.
[0031] Figure 18 Is a schematic diagram of an exemplary multivalent molecule comprising a dendrimer core attached to multiple nucleotide arms.
[0032] Figure 19 Shows a schematic diagram of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, wherein the nucleotide arms comprise biotin, a spacer, a linker, and nucleotide units.
[0033] Figure 20 Is a schematic diagram of an exemplary nucleotide arm comprising a core attachment portion, a spacer, a linker, and nucleotide units.
[0034] Figure 21 Shows the chemical structure of an exemplary spacer (top) and the chemical structures of various exemplary linkers, including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker (bottom).
[0035] Figure 22 Shows the chemical structures of various exemplary linkers, including Linkers 1 to 9.
[0036] Figure 23 Shows the chemical structures of various exemplary linkers conjugated / attached to nucleotide units.
[0037] Figure 24Shows the chemical structures of various exemplary linkers conjugated / attached to nucleotide units.
[0038] Figure 25 Shows the chemical structures of various exemplary linkers conjugated / attached to nucleotide units.
[0039] Figure 26 Shows the chemical structure of an exemplary biotinylated nucleotide arm. In this example, the nucleotide unit is linked to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base.
[0040] Figure 27 Provides a schematic diagram of an embodiment of the low-binding solid support of the present disclosure, wherein the support comprises an alternating layer of a glass substrate and a hydrophilic coating covalently or non-covalently attached to the glass, and the support further comprises a chemically reactive functional group serving as an attachment site for oligonucleotide primers.
[0041] Figure 28 Is a schematic diagram of a guanine tetrad (e.g., G-tetrad).
[0042] Figure 29 Is a schematic diagram of an exemplary intramolecular G-quadruplex structure.
[0043] Figure 30A (i) Is a schematic diagram of an exemplary support having a plurality of nucleic acid capture primers arranged on the support in a non-predetermined and random manner.
[0044] Figure 30A (ii) Is Figure 30A A schematic diagram of the same support as shown in, wherein individual nucleic acid capture primers are attached to nucleic acid template molecules having one of four different batch sequences. The different batch sequences of the template molecules are represented by horizontal stripes, vertical dotted lines, bricks, or solid black.
[0045] Figure 30B (iii) Is a schematic diagram of an exemplary support having a plurality of nucleic acid template molecules fixed to the support (e.g., via attachment to capture primers), wherein the template molecules are arranged on the support in a predetermined manner.
[0046] Figure 30B (iv) Is a schematic diagram of an exemplary support having a plurality of nucleic acid template molecules fixed to the support (e.g., via attachment to capture primers), wherein the template molecules are arranged on the support in a predetermined manner.
[0047] Figure 31AA schematic diagram showing an exemplary workflow for generating circularized padlock probes, including hybridizing a first target-specific padlock probe and a second target-specific padlock probe with a first target molecule and a second target molecule, respectively, to generate a first circularized padlock probe and a second circularized padlock probe having a nick or gap, respectively, and closing the nick or gap to generate a circularized padlock probe. In some embodiments, the first padlock probe comprises: (i) a batch-specific barcode sequence corresponding to a first sequence of interest (batch BC-1); (ii) a batch-specific sequencing primer binding site sequence corresponding to a first sequence of interest (e.g., batch Seq-1); (iii) a capture primer binding site; and (iv) a compaction oligonucleotide binding site. In some embodiments, the second padlock probe comprises: (i) a batch-specific barcode sequence corresponding to a second sequence of interest (batch BC-2); (ii) a batch-specific sequencing primer binding site sequence corresponding to a second sequence of interest (e.g., batch Seq-2); (iii) a capture primer binding site; and (iv) a compaction oligonucleotide binding site.
[0048] Figure 31B A schematic diagram showing an exemplary workflow in which Figure 31A the circularized padlock probes shown in are subjected to rolling circle amplification (RCA) to generate a first concatemeric template molecule and a second concatemeric template molecule, which are immobilized on a support having one type of immobilized capture primer.
[0049] Figure 32 A schematic diagram of an exemplary workflow in which the circularized padlock probes are subjected to rolling circle amplification and batch sequencing.
[0050] Figure 33 A schematic diagram of an exemplary workflow in which the circularized padlock probes are subjected to rolling circle amplification and batch sequencing.
[0051] Figure 34 A schematic diagram of an exemplary workflow in which the circularized padlock probes are subjected to rolling circle amplification and batch sequencing.
[0052] Figure 35 A schematic diagram of an exemplary workflow in which the circularized padlock probes are subjected to rolling circle amplification and batch sequencing.
[0053] Figure 36 A schematic diagram of an exemplary workflow in which the circularized padlock probes are subjected to rolling circle amplification and batch sequencing.
[0054] Figure 37 A schematic diagram of an exemplary workflow in which a linear single-stranded library molecule (100) hybridizes with a single-stranded splint molecule / strand (200), thereby circularizing the library molecule to form a library-splint complex (300) having a nick.
[0055] Figure 38 Schematic diagram of an exemplary workflow in which a linear single-stranded library molecule (100) hybridizes with a single-stranded splint molecule / strand (200), thereby circularizing the library molecule to form a nicked library-splint complex (300).
[0056] Figure 39A Schematic diagram of an exemplary workflow in which a linear single-stranded library molecule-1 (100) hybridizes with a single-stranded splint molecule / strand (200), thereby circularizing the library molecule to form a nicked library-splint complex (300) having a nick that can be enzymatically ligated.
[0057] Figure 39B Schematic diagram of an exemplary workflow in which a linear single-stranded library molecule-2 (100) hybridizes with a single-stranded splint molecule / strand (200), thereby circularizing the library molecule to form a nicked library-splint complex (300).
[0058] Figure 40A Schematic diagram of an exemplary workflow in which Figure 39A the nick in the library-splint complex (300) shown in Figure 40A is ligated to generate the first covalently closed circular library molecule (400) shown in
[0059] Figure 40B Schematic diagram of an exemplary workflow in which Figure 39B the nick in the library-splint complex (300) shown in Figure 40B is ligated to generate the second covalently closed circular library molecule (400) shown in
[0060] Figure 41A Schematic diagram of an exemplary workflow in which a linear single-stranded library molecule-1 (100) hybridizes with a single-stranded splint molecule / strand (200), thereby circularizing the library molecule to form a nicked library-splint complex (300) having a nick that can be enzymatically ligated.
[0061] Figure 41B Schematic diagram of an exemplary workflow in which a linear single-stranded library molecule-2 (100) hybridizes with a single-stranded splint molecule / strand (200), thereby circularizing the library molecule to form a nicked library-splint complex (300).
[0062] Figure 42A Schematic diagram of an exemplary workflow in which Figure 41A the nick in the library-splint complex (300) shown in Figure 42A is ligated to generate the first covalently closed circular library molecule (400) shown in
[0063] Figure 42B is a schematic diagram of an exemplary workflow, where Figure 41B the nicks in the library-clamp complex (300) shown in Figure 42B are ligated to generate the second covalently closed circular library molecule (400) shown in
[0064] Figure 43 is a schematic diagram of an exemplary workflow in which a linear single-stranded library molecule (100) hybridizes with a double-stranded adaptor (500), thereby circularizing the library molecule to form a library-clamp complex (800) having two nicks.
[0065] Figure 44 is a schematic diagram of an exemplary workflow in which a linear single-stranded library molecule (100) hybridizes with a double-stranded adaptor (500), thereby circularizing the library molecule to form a library-clamp complex (800) having two nicks.
[0066] Figure 45 is a schematic diagram of an exemplary workflow in which a linear single-stranded library molecule (100) hybridizes with a double-stranded adaptor (500), thereby circularizing the library molecule to form a library-clamp complex (800) having two nicks.
[0067] Figure 46A is a schematic diagram of an exemplary workflow in which a linear single-stranded library molecule-1 (100) hybridizes with a double-stranded adaptor (500), thereby circularizing the library molecule to form a library-clamp complex (800) having two nicks that can be enzymatically ligated.
[0068] Figure 46B is a schematic diagram of an exemplary workflow in which a linear single-stranded library molecule-2 (100) hybridizes with a double-stranded adaptor (500), thereby circularizing the library molecule to form a library-clamp complex (800) having two nicks that can be enzymatically ligated.
[0069] Figure 47A is a schematic diagram of an exemplary workflow, where Figure 46A the two nicks in the library-clamp complex (800) shown in Figure 47A are ligated to generate the first covalently closed circular library molecule (900) shown in
[0070] Figure 47B is a schematic diagram of an exemplary workflow, where Figure 46B the nicks in the library-clamp complex (800) shown in Figure 47B are ligated to generate the second covalently closed circular library molecule (900) shown in
[0071] Figure 48Is a schematic diagram showing an exemplary linear single-stranded library molecule (100) hybridized with a single-stranded splint molecule / strand (200), thereby circularizing the library molecule to form a nicked library-splint complex (300).
[0072] Figure 49 Is a schematic diagram showing an exemplary linear single-stranded library molecule (100) hybridized with a double-stranded splint molecule (200), thereby circularizing the library molecule to form a library-splint complex (500) with two nicks.
[0073] Figure 50 Shows a sequencing image of a community (e.g., DNA nanoballs) immobilized on a support at high density.
[0074] Figures 51A - 51C Shows a flowchart of different embodiments of a method for sequencing and analyzing a high spatial density sample in 3D.
[0075] In the drawings, like reference numerals generally denote the same or similar elements. Additionally, generally, the leftmost digit of a reference numeral may identify the drawing in which the reference numeral first appears. Detailed Description
[0076] Embodiments of systems, devices, methods, and / or computer program products, and / or combinations and sub-combinations thereof are provided herein that are capable of performing sequencing runs and sequencing analysis on high spatial density three-dimensional (3D) samples where communities or clusters may overlap. The techniques disclosed herein can be used for accurate and reliable base calling in next-generation sequencing (NGS), and NGS will be used as the main example for describing the applications of these techniques herein. However, such techniques may also be useful in other applications.
[0077] Conventional two-dimensional (2D) flow cell images can show clusters or communities from 2D samples, and base calling can be performed using their corresponding image intensities. The optical system can be adjusted to image the clusters or communities in focus. However, in situ samples such as cells or tissues may have a thickness along the axial or z-direction that may not be in focus in a single 2D image. Additionally, 3D samples may have spatially overlapping communities or clusters. The techniques disclosed herein can be used to acquire a stack of flow cell images of a 3D sample and generate accurate and reliable base calling for communities or clusters within the 3D sample, even when the spatial density of the communities or clusters is higher than what conventional 2D or 3D sequencing systems can typically handle.
[0078] Due to limitations including but not limited to the number of kits restricting the available quantity of different primers, the optical design and structure, and / or the chemical characteristics of the sequencing reaction, existing sequencing systems and methods may have limited sequencing throughput. The optical system (e.g., design and structure) can control the image resolution as well as the image quality, thereby also restricting the spatial density of samples that existing systems can accurately and reliably process. The techniques disclosed herein are capable of sequencing samples with a higher spatial density in 3D (e.g., a spatial density that is 2 times, 3 times, 5 times, 10 times, 20 times, 30 times, or higher than the spatial density that existing sequencing systems can process), thereby advantageously increasing the sequencing throughput by multiple folds without the need to improve the existing cartridges, optical designs, and / or chemical properties in the sequencing system. The techniques disclosed herein also advantageously eliminate the need for different subsets of hybridization and / or dehybridization communities associated with using different primers and can exceed the sequencing throughput achievable by using only different primers. The techniques disclosed herein can also advantageously reduce the cost of goods sold (COGS) by using reagents (e.g., affimers) associated with only a subset rather than all 4 types of nucleotides. The techniques disclosed herein can also save imaging time by acquiring images from only a subset of channels (e.g., 3 out of 4 channels in a 4-channel sequencing system). The techniques herein can be used to sequence 3D samples with a higher complexity (e.g., exceeding 20 or 30) and / or lower diversity than what existing sequencing systems may be able to sequence.
[0079] Sequencing System
[0080] Figure 1 FIG. shows a block diagram of a computer-implemented system 1000 according to one or more embodiments disclosed herein. The system 1000 has a sequencing system 1100, which includes a flow cell 1120, a sequencer 1140, an imager 1160, a data storage device 1220, and a user interface 1240. The sequencing system 1100 can be connected to a cloud 1300. The sequencing system 1100 can include one or more of the following: a dedicated processor 1180, a field programmable gate array (FPGA) 1200, and a computer system 1260.
[0081] In some embodiments, the flow cell 1120 is configured to capture DNA fragments and form a DNA sequence for base calling on the flow cell. The flow cell 1120 can include the support disclosed herein. The carrier can be a solid-phase carrier. As disclosed herein, the carrier can include a surface coating thereon. The surface coating can be a polymer coating as disclosed herein.
[0082] Flow cell 1120 may include a plurality of tiles or imaging regions thereon, and each tile may be divided into a grid of sub-tiles. Each sub-tile may include a plurality of clusters or colonies thereon. As a non-limiting example, the flow cell may have 424 tiles, and each tile may be divided into a 6x9 grid, thus having 54 sub-tiles. A flow cell image disclosed herein may be an image of signals including a plurality of clusters or colonies. The flow cell image may include one or more signal tiles or one or more signal sub-tiles. In some embodiments, the flow cell image may be an image including all tiles and substantially all signals thereon. The flow cell image may use imager 1160 during an imaging or sequencing cycle. In some embodiments, each tile may include millions of colonies or clusters. As a non-limiting example, a tile may include from about 1 to 10 million clusters or colonies. Each colony may be a collection of many copies of DNA fragments.
[0083] In the case of sequencing a three-dimensional (3D) sample, such as a cell or tissue immobilized on a flow cell, flow cell images may be acquired at a plurality of z-levels orthogonal to the image plane of the flow cell image to cover the volume of the 3D sample. The z-axis may extend from the objective lens of the optical system disclosed herein to a carrier, e.g., a flow cell device. Each z-level of the flow cell image may be parallel to and separated from an adjacent z-level by a predetermined distance, e.g., from about 0.1 μm to about 15 μm. Each z-level of the flow cell image may be separated from an adjacent level by 1 μm to 10 μm. At each z-level, a flow cell image may be acquired from one or more sequencing cycles and / or one or more channels. Each flow cell image may include at least a portion of one or more tiles or sub-tiles of the flow cell in its field of view. Figure 8B A portion of flow cell 1120 having a plurality of tiles 2900 is shown. The image plane is defined by the x-axis and the y-axis. And the z-axis is orthogonal to the x-y plane. Although the flow cell image, the sample, and the z-axis are described in a Cartesian coordinate system, any other coordinate system may be used to define the spatial positions and relationships of the colonies or clusters and their images herein. Other coordinate systems may include, but are not limited to, polar coordinate systems, cylindrical coordinate systems, or spherical coordinate systems.
[0084] Sequencer 1140 can be configured to flow a nucleotide mixture onto flow cell 1120, cleave blockers from the nucleotides between flow steps, and perform other steps for forming a DNA sequence on flow cell 1120. The nucleotides can have attached fluorescent elements that emit light or energy at wavelengths indicative of the nucleotide type. Each type of fluorescent element can correspond to a specific nucleobase (e.g., A, G, C, T). The fluorescent elements can emit light at visible wavelengths. In some embodiments, sequencer 1140 and flow cell 1120 can be configured to perform the various sequencing methods disclosed herein, e.g., affinity sequencing.
[0085] For example, each nucleobase can be assigned a color. Different types of nucleotides can have different colors. For example, adenine (A) can be red, cytosine (C) can be blue, guanine (G) can be green, and thymine (T) can be yellow. The color or wavelength of the fluorescent element for each nucleotide can be selected such that the nucleotides can be distinguished from one another based on the wavelength of the light emitted by the fluorescent element.
[0086] Imager 1160 can be configured to capture an image of flow cell 1120 after each flow step. In one embodiment, imager 1160 is a camera configured to capture digital images, such as a CMOS or CCD camera. The camera can be configured to capture an image of the wavelength of the fluorescent element bound to the nucleotide. These images can be referred to as flow cell images.
[0087] In some embodiments, imager 1160 can include one or more of the optical systems disclosed herein. The optical system can be configured to capture an optical signal from the flow cell and generate a corresponding digital image. The digital image can then be used for base calling.
[0088] In an embodiment, images of the flow cell can be captured in groups, where each image in the group is taken at a wavelength or spectrum that matches or includes only one of the fluorescent elements. In another embodiment, the image can be captured as a single image that captures all wavelengths of the fluorescent elements.
[0089] The resolution of imager 1160 can control the level of detail in the flow cell image, including pixel size. In existing systems, this resolution is very important because it controls the accuracy of the dot-finding algorithm in identifying the center of a community. In some embodiments, the image resolution of the flow cell images disclosed herein can be from about 10 nanometers (nm) to several hundred nm or higher. One way to improve the accuracy of dot finding is to improve the resolution of imager 1160, or to improve the processing of the images captured by imager 1160. It is possible to perform detection of the center of a community in pixels other than those detected by the dot-finding algorithm. These methods can allow for an increase in the accuracy of community center detection without increasing the resolution of imager 1160. The resolution of the imager can even be lower than that of existing systems with comparable performance, which can reduce the cost of sequencing system 1100.
[0090] The image quality of the flow cell image can control the base calling quality. One way to improve the accuracy of base calling is to improve imager 1160, or to improve the processing of the images captured by imager 1160, to obtain better image quality. The methods described herein can register the flow cell images into a common coordinate system so that base calling regarding clusters or communities is more accurate than when such registration is not performed. These methods can allow for accurate and efficient base identification of 3D samples at a higher system throughput than existing sequencing systems. Some or all of the operations disclosed herein can be advantageously performed by an FPGA, and data can be transferred between the CPU and the FPGA to reduce the total operation time of methods that do not use FPGA operations. Additionally, instead of directly registering multiple flow cell images, which may require saving images before and / or after registration, the image intensity and corresponding localization of selected communities are extracted to estimate the transformation of the entire flow cell image.
[0091] Sequencing system 1100 can be configured to perform 3D sequencing runs and sequencing data analysis. The operations or actions disclosed herein can be performed by dedicated processor 1180, FPGA 1200, computer system 1260, or a combination thereof. One or more operations or actions in method 6000 disclosed herein can be performed by dedicated processor 1180, FPGA 1200, computer system 1260, or a combination thereof. In some embodiments, which operations or actions will be performed by dedicated processor 1180, FPGA 1200, computing system 1260, or a combination thereof can be determined based on one or more of the following: the computation time for a particular operation, the complexity of the computation in a particular operation, the need for data transfer between hardware devices, or a combination thereof.
[0092] Computing system 1260 can include one or more general-purpose computers that provide, for example, in Windows TM or LinuxTM interfaces for running various programs in an operating system such as this. Such operating systems typically provide users with a great deal of flexibility.
[0093] In some embodiments, the dedicated processor 1180 may be configured to perform the operations in the methods disclosed herein. The dedicated processor may not be a general-purpose processor, but rather a custom processor with specific hardware or instructions for performing these steps. The dedicated processor runs specific software directly without an operating system. The lack of an operating system reduces overhead, but at the cost of the flexibility of what the processor can execute. The dedicated processor may use a custom programming language, which can be designed to operate more efficiently than software running on a general-purpose computer. This can increase the speed of performing the steps and allow for real-time processing.
[0094] In some embodiments, the dedicated processor 1180 or the computing system 1260 may include reconfigurable logic devices such as artificial intelligence (AI) chips, neural processing units (NPUs), application-specific integrated circuits (ASICs), or combinations thereof. The reconfigurable logic device may be configured to perform one or more of the operations herein. The reconfigurable logic device may be configured to perform one or more of the operations herein and accelerate the operations by allowing parallel data processing compared to a CPU.
[0095] In some embodiments, the FPGA 1200 may be configured to perform some or all of the operations in the methods herein. The FPGA is programmed to be hardware that will only perform a specific task. A special programming language can be used to convert software steps into hardware components. Once the FPGA is programmed, the hardware directly processes the digital data provided to it without running software. Instead, the FPGA can use logic gates and registers to process digital data. Since no operating system overhead is required, the FPGA typically processes data faster than a general-purpose computer. Similar to the dedicated processor, this is at the cost of flexibility.
[0096] The lack of software overhead can also allow the FPGA to operate faster than the dedicated processor, although this will depend on the exact processing to be performed and the specific FPGA and dedicated processor.
[0097] A group of FPGAs 1200 may be configured to perform these steps in parallel. For example, many FPGAs 1200 may be configured to perform processing steps for an image, a collection of images, sub-blocks, or selected regions in one or more images. Each FPGA 1200 can simultaneously perform its own portion of the processing steps, thus reducing the time required to process the data. This can allow the processing steps to be completed in real time. Further discussion of the use of FPGAs is provided below.
[0098] Performing processing steps in real time can allow the system to use less storage because data can be processed as it is received. This is an improvement over conventional systems that may need to store data before processing it, which may require more storage or access to a computer system located in the cloud 1300.
[0099] In some embodiments, the data storage device 1220 is used to store information used in the methods herein. This information may include the flow cell image itself or information and / or images derived from the flow images captured by the imager 1160. The DNA sequence determined according to base identification may be stored in the data storage device 1220. Predetermined barcode sequences may be stored in the data memory. Parameters for identifying the community location may also be stored in the data storage device 1220. The raw and / or processed image intensities for each community may be stored in the data storage device. The regions and / or sub-tiles corresponding to each community may also be stored in the data storage device 1220. Transformation matrices for each region and / or sub-tile of different cycles and / or channels may also be stored in the data storage device 1220. The pool image may be stored in the data memory. The flow cell image, the processed image, and / or the filtered image may be stored in the data memory. Other information or images that contribute to the 3D base identification of the sample may be stored in the data memory.
[0100] The user interface 1240 can be used by the user to operate the sequencing system or access data stored in the data storage device 1220 or the computer system 1260.
[0101] The computer system 1260 can control the general operation of the sequencing system and can be coupled to the user interface 1240. It can also perform steps in image processing, base identification, its previous operations, and / or subsequent operations (including but not limited to image registration). In some embodiments, the computer system 1260 is the computer system 4000, as Figure 4 described in more detail below. The computer system 1260 can store information about the operation of the sequencing system 1100, such as configuration information, instructions for operating the sequencing system 1100, or user information. The computer system 1260 can be configured to transfer information between the sequencing system 1100 and the cloud 1300.
[0102] As discussed above, the sequencing system 1100 can have a dedicated processor 1180, an FPGA 1200, or a computer system 1260. The sequencing system can use one, both, or all of these components to perform the necessary processing described above. In some embodiments, when these components are present together, the processing tasks are divided among them. For example, the FPGA 1200 can be used to perform some or all of the following: preprocessing operations, image processing, image registration, base calling, and any subsequent operations, while the computer system 1260 can perform other processing functions of the sequencing system 1100, such as registering an image for base calling with a cell staining image. Those skilled in the art will understand that various combinations of these components will allow for various system embodiments that balance the efficiency and speed of processing with the cost of the processing components.
[0103] The cloud 1300 can be a network, a remote storage device, or some other remote computing system separate from the sequencing system 1100. The connection to the cloud 1300 can allow access to data stored external to the sequencing system 1100 or allow for updating software in the sequencing system 1100.
[0104] Method for Sequencing a High Spatial Density Sample
[0105] During DNA sequencing of a 3D sample, flow cell images of the sample may be acquired, and bright spots in the flow cell images may represent different communities or clusters, where different light frequencies represent different colors. When the spatial density of the sample is high, the bright spots representing different communities and clusters may overlap with each other. Therefore, it may be difficult to distinguish communities and their corresponding colors using existing sequencing and analysis methods, and thus errors may occur when performing base calling on high-spatial-density 3D samples using existing sequencing and analysis methods.
[0106] Computer-implemented methods (e.g., 6000, 9000, 5200) for sequencing high-spatial-density samples and / or performing 3D sequencing analysis are disclosed herein, where communities or clusters may overlap and result in errors when performing base calling using existing sequencing methods. The methods herein can include some or all of the operations disclosed herein. The operations can be performed in an order that is not limited to that described herein.
[0107] The methods of the present disclosure (e.g., 6000, 9000, 5200) can be executed by one or more processors disclosed herein. In some embodiments, the processor may include one or more of the following: a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit may include a central processing unit (CPU), a graphics processing unit (GPU), and / or a neural processing unit (NPU). The integrated circuit may include a chip such as a field programmable gate array (FPGA). In some embodiments, the processor may include the computer system 4000.
[0108] In some embodiments, some or all of the operations in methods 6000, 9000, 5200 may be executed by an FPGA and / or other devices (e.g., an AI chip or an NPU). In an embodiment, when some operations are executed by the FPGA, the data after the operations executed by the FPGA may be transmitted by the FPGA to other devices (e.g., the CPU or the NPU) such that the other devices may use such data to execute subsequent operations in the method. Similarly, data may also be transmitted from other devices to the FPGA for processing by the FPGA. In some embodiments, all of the operations in method 6000 may be executed by the CPU. Alternatively, the operations executed by the CPU may be executed by other processors (such as a dedicated processor) or the NPU. In some embodiments, all of the operations in the method may be executed by the FPGA. In some embodiments, some operations in methods 6000, 9000, and 5200 may be executed by the FPGA, while some other operations in the method are executed by an AI chip or an NPU to improve the energy consumption, heat dissipation, and / or computational time required for sequencing analysis.
[0109] Flow cell images can be obtained from one, two, three, four, or more channels of the imager 1160 using the optical system disclosed herein. In some embodiments, multiple flow cell images are obtained during a single flow cycle or multiple flow cycles of a sequencing run. Each flow cell image may include one or more tiles 2900 (imaging regions), and each tile may be divided into multiple sub-tiles. Each sub-tile may include multiple colonies or clusters. Each sub-tile may include multiple regions, where each region includes a number of colonies. For example, colonies may be extracted or otherwise identified from corresponding regions of flow cell images from 4 different channels in a given cycle. As another example, colonies may be extracted from flow cell images from a single channel. The flow cell images disclosed herein may be images obtained from an imaging sample immobilized on the flow cell 1120, as Figure 8B shown.
[0110] Flow cell 1120 may include a sample immobilized thereon. The sample may include a plurality of nucleic acid template molecules. The sample may include a two-dimensional (2D) sample or a three-dimensional (3D) volume sample. The nucleic acid template molecules may be distributed randomly or in various patterns on the flow cell 1120. In some embodiments, a plurality of communities or clusters herein may be extracted from a specific region of a tile (e.g., each sub-tile). In the case of each sub-tile, communities may be extracted in a predetermined pattern or randomly.
[0111] In some embodiments, the communities or clusters sequenced in the flow cycle may have a certain nucleotide diversity, e.g., in base calling. Even if the communities or clusters are of low diversity or unbalanced diversity in the sequencing cycle, the method may allow color registration of the flow cell image. The nucleotide diversity of a population of nucleotide acid molecules (e.g., communities or clusters) may refer to the relative proportions of nucleotides A, G, C, and T / U present in each flow cycle. The relative proportions of nucleotides may be within a region of the field of view or within the entire flow cell image. Optimal high or balanced diversity data typically may have approximately equal proportions of all four nucleotides represented in each flow cycle of the sequencing run. Low or unbalanced diversity data typically may include high proportions of certain nucleotides and low proportions of other nucleotides in some flow cycles of the sequencing run, e.g., less than 10% of the total of all 4 nucleotides. Thus, an image corresponding to a high proportion of certain nucleotides may have more signal points (communities or clusters) compared to an image corresponding to a low proportion of certain nucleotides. As an example of low or unbalanced diversity data, in a certain flow cycle, bases A, T, C, G may be approximately 1%, approximately 2%, approximately 1%, and approximately 95% of the total community, respectively. Subsequently, in this particular flow cycle, the flow cell image from the channels corresponding to A, T, and C is darker and has fewer polarities or clusters compared to the flow cell image corresponding to nucleotide G. As another example of low or unbalanced diversity data, in a plurality of flow cycles, bases A, T, C, G in the polarities are approximately 2%, approximately 5%, approximately 10%, and approximately 83%, respectively. In embodiments where low or unbalanced diversity data is present in a particular cycle and imaged for sequencing analysis, image registration using prior art may fail because the images from one or more channels are too dark compared to the images obtained from other channels (e.g., the signal points at the poles are too sparse and / or dim), causing problems in subsequent color correction. Additionally, in embodiments where low or unbalanced diversity data is present in a particular cycle, correction of channel crosstalk using prior art may fail because the images from one or more channels are too dark (e.g., the signal points at the poles are too sparse and / or dim). In some embodiments, even if the communities or clusters are of low diversity, the method (e.g., 6000, 9000, 5200) is configured to perform color correction of the flow cell image.
[0112] The methods 6000, 9000, 5200 in this document can achieve the sequencing and analysis of 3D samples. For example, in-situ samples (such as single cells or tissues) whose spatial density exceeds the range that existing sequencing systems can process at a predetermined quality level. For example, the 3D sample may partially or completely have overlapping communities or clusters in the x-y plane or in 3D. Partial or complete overlap may lead to inaccurate base identification of the overlapping communities, resulting in inaccurate DNA sequencing analysis. In some embodiments, such high-density 3D samples can be divided into "batch" groups, and each batch can use different sequencing primers. Details of exemplary embodiments for sequencing communities or clusters in different batches are described in International Patent Application No. PCT / US23 / 65972 (the content of which is hereby incorporated by reference in its entirety). However, using different sequencing primers to bind communities or clusters may require consuming corresponding reagents and buffers, spending time and effort in the hybridization and / or dehybridization of communities in different "batches", and the number of available primers may be limited by the number and volume of cartridges in the sequencing system. Compared with existing methods, the methods in this document advantageously allow the system throughput to be further increased by 3 times, 5 times, 10 times or more. The existing methods perform batch sequencing of samples without changing the sequencing primers, cartridges and optical systems.
[0113] In some embodiments, by using different barcode sequences to identify each different subset, such high-density 3D samples to be imaged can be divided into multiple subsets, such as a first community subset, a second subset or a third subset. For example, overlapping communities or clusters (such as overlapping communities or clusters of high-spatial-density samples) can be divided into different subsets. In some embodiments, the communities in different subsets can use the same sequencing primers, so the subsets are regarded as "sub-batches" within a single "batch". During sequencing, one or more subsets can be made "bright", while other subsets overlapping with one or more subsets are made "dark", so that the spatial density of the "bright" communities in the flow cell image is lower than the spatial density of all communities. Different subsets can be "bright" and "dark" in different sequencing cycles, so that each subset of communities can be "bright" in at least some sequencing cycles and optionally not "bright" in all sequencing cycles. In some embodiments, the dark communities or clusters are 5 times, 8 times, 10 times, 12 times, 15 times, 20 times, 30 times, 40 times, 50 times or more darker than the bright communities. In some embodiments, the dark communities are 5 times, 8 times, 10 times, 12 times, 15 times, 20 times or 30 times darker than the bright communities. In some embodiments, the intensity of the dark communities is about 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times or 5 times the average background noise level of the corresponding flow cell image.
[0114] The total number of subsets can be customized according to the number of barcode sequences. Exemplary barcode sequences are shown in Figures 11A to 11C . In some embodiments, the number of subsets can be customized to ensure that the spatial density of "bright" clusters does not reduce the accuracy and reliability of sequencing analysis. Each subset may have a spatial density manageable by existing 3D sequencing systems. For example, each subset can have a spatial density equal to or less than the maximum spatial density that can be managed by existing 3D sequencing systems with a predetermined quality control in base calling, such as Q40 or Q30. As another example, each subset can have a spatial density of not less than about 0.01 to about 0.5 clusters / μm 3 . As another example, each subset can have a spatial density of not less than 0.01 to 0.5 clusters / μm 3 . In some embodiments, each subset can have a spatial density of not less than about 0.1 to about 1 cluster or cluster / μm^3. In some embodiments, each subset can have a spatial density of not less than 0.1 to 1 cluster / μm 3 . In some embodiments, each subset can have a spatial density of not less than 0.002 to 50 clusters / μm 3 . In some embodiments, each subset can have a spatial density of not less than 0.005 to 30 clusters / μm 3 . In some embodiments, the density of the clusters or clusters is at least in the image plane or the x-y plane. In some embodiments, the density of the clusters or clusters is in 3D.
[0115] The total number of subsets can be customized according to the spatial density of the sample. The total number of subsets can be an integer in the range of 3 to 8. The total number of subsets can be an integer in the range of 2 to 30. The total number of subsets can be an integer in the range of 2 to 100. For example, using 10 different sequencing primers and 4 different barcodes for each sequencing primer, the system throughput can be increased by 40 times using the methods disclosed herein. The spatial density of the samples processed by the sequencing system can be 40 times higher than the existing spatial density of 3D samples using existing sequencing systems and methods.
[0116] Sequencing without Imaging in the "Dark" Channel
[0117] In some embodiments, method 5200 herein includes obtaining a flow cell image of a sample from some but not all channels, such as Figure 51A . In some embodiments, method 5200 herein includes obtaining a flow cell image of a sample from "bright" channels but not from "dark" channels, such as Figure 51A . In some embodiments, flow cell images from dark channels can be obtained / acquired, but they are not used for generating base calls, such as as Figures 51B - 51CAs shown. To save the imaging time and cost of flowing a reagent corresponding to the "dark" channel (e.g., an affinity body with a "dark" dye) to the sample, it may be preferred not to acquire / obtain any flow cell images in the dark channel and / or not to contact the community with an affinity body with a dark fluorescent dye or without a dye corresponding to the dark channel in some sequencing cycles.
[0118] The sample can be in 3D. Flow cell images can be acquired at different positions along the axial axis (i.e., the z-axis as shown in Figure 8B ). Flow cell images can be acquired in a certain sequencing cycle of the sequence run. Flow cell images can be acquired from one or more channels.
[0119] In some embodiments, the method includes imaging only from some but not all color channels of the sequencing system. The "dark" channel may not be used for imaging in at least some cycles, e.g., the cycles covering barcode sequences. The "dark" channel can correspond to a dark fluorescent dye attached to a corresponding type of nucleotide base (e.g., adenine (A)). In some embodiments, the "dark" channel can correspond to a lack of any fluorescent dye attached to a corresponding type of nucleotide base (e.g., adenine (A)). In other words, the fluorescent dyes applied in such cycles can be some bright dyes and some dark dyes, or only bright dyes, and the number of different types of bright fluorescent dyes is less than the total number of channels.
[0120] There may be a single dark channel. In some embodiments, there may be multiple dark channels.
[0121] In some embodiments, the dark channel corresponds to a fluorescent dye attached to a nucleotide of adenine (A), thymine (T), guanine (G), or cytosine (C) that emits light below a predetermined threshold in the first, second, and / or third plurality of sequencing cycles. The predetermined threshold can be customized based on different sequencing applications. For example, the predetermined threshold can be a relative intensity level with respect to the brightest image intensity in the same flow cell image. As another example, the predetermined threshold can be a relative intensity level with respect to the noise intensity in the same flow cell image.
[0122] In some embodiments, the dark channel corresponds to a lack of any fluorescent dye attached to a nucleotide of adenine (A), thymine (T), guanine (G), or cytosine (C) in the first plurality of sequencing cycles or the second plurality of sequencing cycles.
[0123] In some embodiments, the methods herein can use one or more identical dark channels corresponding to the same type of nucleobase in a first plurality of flow cycles, a second plurality of flow cycles, and / or a third plurality of flow cycles (e.g., in 15 cycles corresponding to sequencing a 15-base barcode sequence). For example, a dark channel (e.g., the channel corresponding to nucleotide A) remains the same in the first plurality of flow cycles, the second plurality of flow cycles, and / or the third plurality of flow cycles. A dark channel (e.g., the channel corresponding to nucleotide A) remains the same in the first plurality of flow cycles and the second plurality of flow cycles.
[0124] In some embodiments, method 5200 herein can include operation 5210 of obtaining a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from a first subset of channels; operation 5220 of obtaining a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from the first subset of channels; operation 5230 of generating, by a processor, a first set of base calls for a first subset of communities of the sample based on the first plurality of flow cell images; and operation 5240 of generating, by the processor, a second set of base calls for a second subset of communities of the sample based on the second plurality of flow cell images. In some embodiments, the first subset of channels includes only some but not all of the color channels of the sequencing system, and wherein the second plurality of sequencing cycles is after the first plurality of sequencing cycles.
[0125] In some embodiments, the image intensity of the second subset of communities in the first plurality of flow cell images in the first plurality of sequencing cycles is lower than a predetermined threshold. The predetermined threshold can be customized within various ranges according to different sequencing applications. The second subset of communities may appear 2 times, 5 times, 8 times, 10 times, 12 times, 15 times, 18 times, 20 times or more darker than the first subset of communities in the first plurality of cycles. In some embodiments, the image intensity of the second subset of communities is 0.2 times, 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times, 4 times, 5 times or more of the average background noise intensity in the corresponding flow cell images in the first plurality of sequencing cycles. In some embodiments, the image intensity of the second subset of communities is less than 15%, 10%, 8%, 5%, 4%, 2% or 1% of the highest signal intensity in the corresponding flow cell images in the first plurality of sequencing cycles, e.g., the average of the top 1%, 2%, 3%, 4% or 5% signal intensity. The individual intensity of each community or the average intensity of the communities within the second subset of communities can be used for comparison with the predetermined threshold, the average background noise, and / or the intensity of the first subset of communities.
[0126] In some embodiments, the image intensity of the first subset of colonies is below a predetermined threshold in the second plurality of flow cell images. The predetermined threshold can be customized to various ranges according to different sequencing applications. In some embodiments, the first subset of colonies appears 5 times, 10 times, 15 times, or 20 times darker than the second subset of colonies in the second plurality of flow cell images. In some embodiments, the image intensity of the first subset of colonies is 0.2 times, 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times, 4 times, 5 times, or more of the average background noise intensity in the corresponding flow cell images in the second plurality of sequencing cycles. In some embodiments, the image intensity of the first subset of colonies is less than 15%, 10%, 8%, 5%, 4%, 2%, or 1% of the highest signal intensity in the corresponding flow cell images in the second plurality of sequencing cycles, such as the average value of the top 1%, 2%, 3%, 4%, or 5% signal intensity. The individual intensity of each colony or the average intensity of the colonies within the first subset of colonies can be used for comparison with the predetermined threshold, the average background noise, and / or the intensity of the second subset of colonies.
[0127] In some embodiments, in the first plurality of flow cell images, the image intensity of the second subset of colonies or any other colonies not in the first subset of colonies is below a predetermined threshold, so they appear "dark" in the first plurality of flow cell images. In some embodiments, the dark colonies are 5 times, 8 times, 10 times, 12 times, 15 times, 20 times, 30 times, 40 times, 50 times, or more times darker than the bright colonies (e.g., the average intensity of the bright colonies). In some embodiments, the dark colonies are 5 times, 8 times, 10 times, 12 times, 15 times, 20 times, or 30 times darker than the bright colonies (e.g., the average intensity of the bright colonies). In some embodiments, the intensity of the dark colonies is about 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times, or 5 times the average background noise level of the corresponding flow cell images. The individual intensity of each colony or the average intensity of the dark colonies can be used for comparison with the predetermined threshold, the average background noise, and / or the intensity of the bright colonies.
[0128] In some embodiments, the first subset of colonies is different from the second subset of colonies. In some embodiments, the first subset of colonies and the second subset of colonies at least partially spatially overlap in the x-y plane or in 3D. In some embodiments, the first subset of colonies and the second subset of colonies contain the same batch-specific sequencing binding sites, which are configured to bind to the same sequencing primers, such that they are "sub-batches" within the same batch of colonies and clusters. In some embodiments, each colony in the first subset of colonies and the second subset of colonies contains the same batch-specific sequencing binding sites, which are configured to bind to the same sequencing primers.
[0129] In some embodiments, each colony in the first colony subset is configured to bind to a first sequencing primer, and each colony in the second colony subset is configured to bind to a second sequencing primer different from the first sequencing primer. Thus, the first colony subset and the second colony subset are not within the same "batch".
[0130] In some embodiments, the first set of base identifications corresponds to the first colony subset in the first plurality of sequencing cycles. In some embodiments, the first set of base identifications includes only one, two, or three types of nucleotide bases, but not all types of nucleotide bases. In some embodiments, the second set of base identifications corresponds to the second set of colonies in the second plurality of sequencing cycles. In some embodiments, the second set of base identifications includes only one, two, or three types of nucleotide bases, but not all types of bases.
[0131] In some embodiments, the operation 5210 of obtaining the first plurality of flow cell images of the sample in the first plurality of sequencing cycles from the first subset of channels includes: obtaining the first plurality of flow cell images of the sample in the first plurality of sequencing cycles by the optical system of the sequencing system only from the first subset of channels and not from any dark channels. In some embodiments, the operation 5210 of obtaining the first plurality of flow cell images of the sample in the first plurality of sequencing cycles from the first subset of channels includes: controlling the optical system by the processor to avoid collecting data from one or more image sensors of the dark channels in the first plurality of sequencing cycles. In some embodiments, the operation 5210 includes: controlling the optical system by the processor to avoid irradiating light within a predetermined frequency range corresponding to the dark channels onto the sample in the first plurality of sequencing cycles.
[0132] In some embodiments, the operation 5220 of obtaining the second plurality of flow cell images of the sample in the second plurality of sequencing cycles from the first subset of channels includes: obtaining the second plurality of flow cell images of the sample in the second plurality of sequencing cycles by the optical system of the sequencing system only from the first subset of channels and not from any dark channels. In some embodiments, the operation 5220 of obtaining the second plurality of flow cell images of the sample in the second plurality of sequencing cycles from the first subset of channels includes: controlling the optical system by the processor to avoid collecting any data from one or more image sensors of any dark channels in the second plurality of sequencing cycles. In some embodiments, the operation 5220 of obtaining the second plurality of flow cell images of the sample in the second plurality of sequencing cycles from the first subset of channels includes: controlling the optical system by the processor to avoid irradiating light within a predetermined frequency range corresponding to the dark channels onto the sample in the second plurality of sequencing cycles.
[0133] In some embodiments, operation 5230 of generating a first set of base calls for a first subset of a sample based on a first plurality of flow cell images includes: generating a first set of base calls for a first subset of a sample based only on the first plurality of flow cell images obtained from a first subset of channels. In some embodiments, operation 5240 of generating a second set of base calls for a second subset of a sample based on a second plurality of flow cell images includes: generating a second set of base calls for a second subset of a sample based only on the second plurality of flow cell images from a first subset of channels.
[0134] In some embodiments, the first plurality of sequencing cycles includes the same number of cycles as the second plurality of sequencing cycles. For example, each of the first plurality of cycles or the second plurality of cycles may include 2, 3, 4, 5, 6, 7, 8, or more consecutive cycles in a sequencing run. In some embodiments, each of the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 2 to 50 consecutive cycles. In some embodiments, each of the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 2 to 20 consecutive cycles. In some embodiments, each of the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 3 to 8 consecutive cycles.
[0135] A sequence run may include any non-zero integer number of sequencing cycles, such as 150, 200, or 300. For example, a sequence run has 9, 10, 11, 12, 13, 14, 15, 18, 25, or more consecutive sequencing cycles corresponding to barcode sequences, for example, 15 cycles as shown in Figures 11A - 11C which may be located at any position in the sequencing run.
[0136] Reference Figure 11A , in a particular embodiment, a sample is divided into 3 subsets of communities. The first plurality of flow cell images may be from the nth cycle to the n + 5th cycle and obtained from a subset of channels corresponding to nucleotides T, G, and C, but not from the "dark" channel corresponding to nucleotide A (in the community of subset 3), where n may be any non-zero integer but less than the total number of cycles in the sequence run. The second plurality of flow cell images are from the n + 6 to n + 10th cycles and are obtained from a subset of channels corresponding to nucleotides T and C, rather than from the dark channel corresponding to nucleotide A (in the community of subset 1). The second plurality of flow cell images may be obtained from a subset of channels corresponding to nucleotides T, C, and G.
[0137] See Figure 11A, in this particular embodiment, A is a "dark" base, and the channel corresponding to the fluorescent dye that detects the affinity body attached to this type of A is a "dark" channel. The other three channels corresponding to the bases T, C, and G are bright channels. In this embodiment, the first plurality of flow cell images are flow cell images including bright signals from the first community subset (i.e., "community 3") in the nth to n+5th cycles. In this particular embodiment, n is equal to one. The communities in the other subsets, namely "community 2" and "community 1", appear dark in the channels corresponding to the nucleotides T, G, and C in the nth to n+5th cycles. Therefore, the spatial density of the sample in the first plurality of flow cell images can be reduced by darkening community 1 and community 2 in the nth to n+5th cycles. The first plurality of flow cell images may be from the channels corresponding to the bases T, C, and G. The other communities in "community 2" and "community 1" may also appear dark in the dark channel corresponding to the nucleotide A.
[0138] Similarly, in the second plurality of flow cell images, the image intensity of the first community subset or any other community that does not belong to the second community subset is lower than a predetermined threshold, so they appear dark. Continuing to refer to Figure 11A , the second plurality of flow cell images are flow cell images from the subset "community 1" in the n+6th to n+10th cycles. In these sequencing cycles, the other communities in "community 2" and "community 3" appear dark in the channels corresponding to the nucleotides T and C. Although the channel corresponding to the nucleotide G is a "bright" channel in this embodiment, there may be no bright signal in the n+6th to n+10th cycles, while there may be a signal in the n+1st to n+5th cycles. In contrast, the dark channel corresponding to A remains dark throughout the n+1st to n+15th cycles.
[0139] In some embodiments, method 5200 herein further includes an operation of obtaining a third plurality of flow cell images of the sample in a third plurality of sequencing cycles from a first subset of channels.
[0140] Continuing to refer to Figure 11A , the third plurality of flow cell images correspond to the community subset labeled "community 2" in the n+11th to n+15th cycles. In this embodiment, the first subset of channels includes 3 channels corresponding to the bases C, G, and T but not corresponding to A.
[0141] The first subset of channels may not include a dark channel. The first subset of channels may not include all channels of the sequencing system, such as 4 channels. Operations 5210, 5220, and the operation of obtaining the third plurality of flow cell images may use some or all of the channels in the first subset of channels. The channels for obtaining one of the first plurality of flow cell images, the second plurality of flow cell images, the third plurality of flow cell images, or other plurality of flow cell images may be the same as or different from the channels for obtaining another of the first plurality of flow cell images, the second plurality of flow cell images, the third plurality of flow cell images, or other plurality of flow cell images, as Figures 11A - 11B shown. For example, the first plurality of flow cell images are obtained from channels corresponding to T, G, and C, and the second plurality of flow cell images are obtained from channels corresponding to T, G, and C, but more specifically, only from the channels of C and T.
[0142] In some embodiments, the subset of channels for obtaining the first plurality of flow cell images in the first plurality of sequencing cycles may not include any dark channels. Similarly, the subset of channels for obtaining the second plurality of flow cell images or the third plurality of flow cell images in their corresponding sequencing cycles may not include any dark channels. As Figures 11A - 11B shown, the subset of channels for obtaining the first plurality of flow cell images, the second plurality of flow cell images, and the third plurality of flow cell images does not include the dark channel corresponding to "A". The dark channel herein may be a channel for detecting light emitted from a fluorescent dye attached to a specific nucleotide (e.g., A), but the emitted light is below a predetermined threshold, so the flow cell images from the dark channel appear dark.
[0143] In some embodiments, operations 5230 and 5240 are based only on the flow cell images from the first subset of channels. The time and computation for obtaining flow cell images in the dark channels and / or performing base identification using the flow cell images from the dark channels can be saved, thereby reducing the total time required for sequencing and analysis.
[0144] Sequencing Using All Channels
[0145] In some embodiments, methods 6000, 9000, 5200 may include an operation of obtaining flow cell images from all channels of the sequencing system, including "dark" channels.
[0146] In some embodiments, as Figure 51BAs shown, the method of the present disclosure includes the operation of obtaining flow cell images from all channels of a sequencing system. In some embodiments, the method includes the operation 5210' of obtaining, by a processor, a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from one or more channels; the operation 5220' of obtaining, by the processor, a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from one or more channels; the operation 5230 of generating, by the processor, a first set of base calls for a first subset of communities of the sample based on the first plurality of flow cell images; and the operation 5240 of generating, by the processor, a second set of base calls for a second subset of communities of the sample based on the second plurality of flow cell images. In some embodiments, one or more channels include dark channels, wherein the image intensity of some of the flow cell images in the first plurality of flow cell images or the second plurality of flow cell images is lower than a predetermined threshold, and wherein the second plurality of sequencing cycles is after the first plurality of sequencing cycles.
[0147] In some embodiments, operations 5210' and 5220' are similar to operations 5210 and 5220, except that operations 5210' and 5220' are from one or more channels.
[0148] In some embodiments, operation 5210' includes: obtaining, by an optical system of the sequencing system, a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from one or more channels. In some embodiments, operation 5220' includes obtaining, by the optical system of the sequencing system, a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from one or more channels. In some embodiments, one or more channels include at least a dark channel and at least a channel that is not a dark channel. In some embodiments, one or more channels include only a single dark channel and two or three channels different from the dark channel.
[0149] As Figure 51B shown, method 5200 may include operations 5230 and 5240 as Figure 51A shown and disclosed herein in connection with Figure 51A the present disclosure.
[0150] In some embodiments, the image intensity of the second subset of colonies in the first plurality of flow cell images in the first plurality of sequencing cycles is below a predetermined threshold. The second subset of colonies may appear 2 times, 5 times, 8 times, 10 times, 12 times, 15 times, 18 times, 20 times or more dimmer than the first subset of colonies in the first plurality of cycles. In some embodiments, the image intensity of the second subset of colonies is 0.2 times, 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times, 4 times, 5 times or more of the average background noise intensity in the corresponding flow cell images in the first plurality of sequencing cycles. In some embodiments, the image intensity of the second subset of colonies is less than 15%, 10%, 8%, 5%, 4%, 2% or 1% of the highest signal intensity in the corresponding flow cell images in the first plurality of sequencing cycles, such as the average of the top 1%, 2%, 3%, 4% or 5% signal intensities.
[0151] In some embodiments, the image intensity of the first subset of colonies is below a predetermined threshold in the second plurality of flow cell images. In some embodiments, the first subset of colonies appears 5 times, 10 times, 15 times or 20 times dimmer than the second subset of colonies in the second plurality of flow cell images. In some embodiments, the image intensity of the first subset of colonies is 0.2 times, 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times, 4 times, 5 times or more of the average background noise intensity in the corresponding flow cell images in the second plurality of sequencing cycles. In some embodiments, the image intensity of the first subset of colonies is less than 15%, 10%, 8%, 5%, 4%, 2% or 1% of the highest signal intensity in the corresponding flow cell images in the second plurality of sequencing cycles, such as the average of the top 1%, 2%, 3%, 4% or 5% signal intensities.
[0152] In some embodiments, in the first plurality of flow cell images, the image intensity of the second subset of colonies or any other colony not in the first subset of colonies is below a predetermined threshold, so they appear "dim" in the first plurality of flow cell images. In some embodiments, the dim colonies are 5 times, 8 times, 10 times, 12 times, 15 times, 20 times, 30 times, 40 times, 50 times or more dimmer than the bright colonies (e.g., the average intensity of the bright colonies). In some embodiments, the dim colonies are 5 times, 8 times, 10 times, 12 times, 15 times, 20 times or 30 times dimmer than the bright colonies (e.g., the average intensity of the bright colonies). In some embodiments, the intensity of the dim colonies is about 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times or 5 times the average background noise level of the corresponding flow cell images.
[0153] In some embodiments, the first community subset is different from the second community subset. In some embodiments, the first community subset and the second community subset at least partially spatially overlap in the x-y plane or in 3D. In some embodiments, the first community subset and the second community subset comprise the same batch-specific sequencing binding sites, which are configured to bind to the same sequencing primers, such that they are "sub-batches" within the same batch of communities and clusters. In some embodiments, each community in the first community subset and the second community subset comprises the same batch-specific sequencing binding sites, which are configured to bind to the same sequencing primers.
[0154] In some embodiments, each community in the first community subset is configured to bind to a first sequencing primer, and each community in the second community subset is configured to bind to a second sequencing primer different from the first sequencing primer. Thus, the first community subset and the second community subset are not within the same "batch".
[0155] In some embodiments, a first set of base identifications corresponds to a first set of communities in a first plurality of sequencing cycles. In some embodiments, the first set of base identifications includes all four types of nucleotide bases. In some embodiments, a second set of base identifications corresponds to a second set of communities in a second plurality of sequencing cycles. In some embodiments, the second set of base identifications includes all four types of nucleotide bases.
[0156] In some embodiments, the operation 5210' of obtaining a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from one or more channels includes: obtaining, by an optical system of the sequencing system, the first plurality of flow cell images of the sample in the first plurality of sequencing cycles from one or more channels, which may include a dark channel. In some embodiments, the operation 5210' of obtaining a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from one or more channels includes: controlling, by a processor, the optical system to collect data from one or more image sensors of one or more channels in the first plurality of sequencing cycles; controlling, by the processor, the optical system to irradiate the sample with light within a predetermined frequency range in the first plurality of sequencing cycles, the predetermined frequency range corresponding to the dark channel.
[0157] In some embodiments, the operation 5220' of obtaining the second plurality of flow cell images of the sample in the second plurality of sequencing cycles from one or more channels includes: obtaining, by an optical system of the sequencing system, the second plurality of flow cell images of the sample in the second plurality of sequencing cycles from one or more channels, where the one or more channels may or may not include a dark channel. In some embodiments, the operation 5220' of obtaining the second plurality of flow cell images of the sample in the second plurality of sequencing cycles from one or more channels includes: controlling, by a processor, the optical system to collect data from one or more image sensors corresponding to one or more channels in the second plurality of sequencing cycles. In some embodiments, the operation 5220' of obtaining the second plurality of flow cell images of the sample in the second plurality of sequencing cycles from one or more channels includes: controlling, by a processor, the optical system to irradiate light within a predetermined frequency range corresponding to a dark channel onto the sample during the second plurality of sequencing cycles.
[0158] In some embodiments, the operation 5230 of generating a first set of base calls for a first subset of colonies of the sample based on the first plurality of flow cell images includes: generating the first set of base calls for the first subset of colonies of the sample based only on the first plurality of flow cell images obtained from one or more channels. In some embodiments, the operation 5240 of generating a second set of base calls for a second subset of colonies of the sample based on the second plurality of flow cell images includes generating the second set of base calls for the second subset of colonies of the sample based only on the second plurality of flow cell images from one or more channels.
[0159] In some embodiments, the first plurality of sequencing cycles includes the same number of cycles as the second plurality of sequencing cycles. For example, each of the first plurality of cycles or the second plurality of cycles may include 2, 3, 4, 5, 6, 7, 8 or more consecutive cycles in a sequencing run. In some embodiments, each of the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 2 to 60 cycles. In some embodiments, the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 2 to 40 cycles. In some embodiments, the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 3 to 15 cycles.
[0160] A sequencing run may include any non - zero integer number of sequencing cycles. For example, a sequence run has 10, 12, 15, 24 or more consecutive sequencing cycles corresponding to barcode sequences, e.g., 15 cycles as shown in Figures 11A - 11C which may be located at any position in the sequencing run.
[0161] In some embodiments, the method herein further includes an operation of obtaining a third plurality of flow cell images of the sample in a third plurality of sequencing cycles from one or more channels.
[0162] One or more channels may include some or all of the dark channels. One or more channels may not include all of the channels of the sequencing system, such as four channels. Operations 5210', 5220', and the operation of obtaining the third plurality of flow cell images may use one or more channels. The channels for obtaining one of the first plurality of flow cell images, the second plurality of flow cell images, the third plurality of flow cell images, or other pluralities of flow cell images may be the same as or different from the channels for obtaining another of the first plurality of flow cell images, the second plurality of flow cell images, the third plurality of flow cell images, or other pluralities of flow cell images, as Figures 11A - 11B shown. For example, the first plurality of flow cell images are obtained from channels corresponding to T, G, and C, and the second plurality of flow cell images are obtained from channels corresponding to T, G, and C, but more specifically, only from the channels of C and T.
[0163] In some embodiments, the channels for obtaining the first plurality of flow cell images in the first plurality of sequencing cycles may include one or more dark channels. Similarly, the channels for obtaining the second plurality of flow cell images or the third plurality of flow cell images in their corresponding sequencing cycles may include one or more dark channels. For example, the channels for obtaining the first plurality of flow cell images, the second plurality of flow cell images, and the third plurality of flow cell images include all four channels. Each dark channel herein may be a channel for detecting the emitted light from a fluorescent dye attached with a specific nucleotide (e.g., A), but the emitted light is below a predetermined threshold, so the flow cell images from the dark channels appear dark.
[0164] In some embodiments, the dark channels correspond to the channels in which no fluorescent emission above a predetermined threshold and within the frequency range corresponding to the dark channels is generated from the sample in the first plurality of sequencing cycles. In some embodiments, the signal intensity of the flow cell images obtained from the dark channels may be much darker, for example, 10 times darker than the bright spots in the flow cell images from other channels. In some embodiments, the signal intensity of the flow cell images obtained from the dark channels may be comparable to the background noise, such that the flow cell images obtained or acquired from the dark channels may not be used for base calling to avoid possible errors in base calling. In addition, the time and computation for using the flow cell images to perform base calling can be reduced. In other words, the flow cell images from the bright channels (e.g., 2 or 3 bright channels) can be used to perform base calling.
[0165] Sequencing Using Different "Dark" Channels
[0166] In some embodiments, methods 6000, 9000, 5200 may include the operation of obtaining or acquiring flow cell images from one or more channels of a sequencing system. One or more channels may not include a dark channel. In some embodiments, one or more channels may include all channels, including the dark channel. The dark channel in the first plurality of sequencing cycles may be different from the dark channel in the second plurality of sequencing cycles. In other words, the dark channel may alternate but not consistently in multiple consecutive cycles of a sequencing run, e.g., consecutive cycles covering the nucleotide bases of a barcode.
[0167] In some embodiments, method 5200 herein includes operation 5210 of obtaining, by a processor, a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from a first subset of channels; operation 5220 of obtaining, by the processor, a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from a second subset of channels; operation 5230 of generating, by the processor, a first set of base calls for a first subset of communities of the sample based on the first plurality of flow cell images; and operation 5240 of generating, by the processor, a second set of base calls for a second subset of communities of the sample based on the second plurality of flow cell images. In some embodiments, the first subset and the second subset of channels are at least partially different, and one or more channels do not include a first dark channel, and the second subset of channels does not include a second dark channel different from the first dark channel.
[0168] In embodiments where flow cell images are obtained or acquired only from bright channels and not from dark channels, the first subset and the second subset of channels are at least partially different, and one or more channels do not include a first dark channel, e.g., corresponding to nucleotide A in the first plurality of cycles, and the second subset of channels does not include a second dark channel different from the first dark channel, e.g., corresponding to nucleotide C in the second plurality of cycles.
[0169] In embodiments where flow cell images are obtained or acquired from dark channels, the first subset and the second subset of channels are the same, and one or more channels may include a dark channel. In such embodiments, the first subset and the second subset of one or more channels may each include all of one or more channels, e.g., all four channels of a sequencing system.
[0170] When the operation includes obtaining flow cell images from a dark channel, an alternating dark channel (e.g., from a channel corresponding to nucleotide A in the first plurality of cycles to a channel corresponding to nucleotide C in the second plurality of cycles) may be compatible with the operations in method 5200. In some embodiments, when the operation includes obtaining flow cell images from some or all bright channels and not from a dark channel, an alternating dark channel may be compatible with method 5200.
[0171] In some embodiments, operations 5210” and 5220” are similar to operations 5210 and 5220, except that operations 5210” and 5220” correspond to different subsets of one or more channels, namely a first subset and a second subset of the channels. The first subset may not include the first dark channel, while the second subset may not include a second dark channel different from the first channel.
[0172] As Figure 51C shown, method 5200 may include operations 5230 and 5240 as shown in FIGS. 52A-52B and is disclosed herein in connection with Figures 51A - 51B this.
[0173] In some embodiments, the image intensity of a second subset of communities in a first plurality of flow cell images in a first plurality of sequencing cycles is below a predetermined threshold. The predetermined threshold may be customized within various ranges based on different sequencing applications. The second subset of communities may appear 2 times, 5 times, 8 times, 10 times, 12 times, 15 times, 18 times, 20 times or more darker than a first subset of communities in the first plurality of cycles. In some embodiments, the image intensity of the second subset of communities is 0.2 times, 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times, 4 times, 5 times or more of the average background noise intensity in the corresponding flow cell images in the first plurality of sequencing cycles. In some embodiments, the image intensity of the second subset of communities is less than 15%, 10%, 8%, 5%, 4%, 2% or 1% of the highest signal intensity in the corresponding flow cell images in the first plurality of sequencing cycles, such as the average of the top 1%, 2%, 3%, 4% or 5% signal intensities. The individual intensity of each community or the average intensity of the communities within the second subset of communities may be used for comparison with the predetermined threshold, the average background noise, and / or the intensity of the first subset of communities.
[0174] In some embodiments, the image intensity of the first subset of communities is below a predetermined threshold in the second plurality of flow cell images. The predetermined threshold can be customized to various ranges based on different sequencing applications. In some embodiments, the first subset of communities appears 5 times, 10 times, 15 times, or 20 times darker than the second subset of communities in the second plurality of flow cell images. In some embodiments, the image intensity of the first subset of communities is 0.2 times, 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times, 4 times, 5 times, or more of the average background noise intensity in the corresponding flow cell images in the second plurality of sequencing cycles. In some embodiments, the image intensity of the first subset of communities is less than 15%, 10%, 8%, 5%, 4%, 2%, or 1% of the highest signal intensity in the corresponding flow cell images in the second plurality of sequencing cycles, such as the average of the top 1%, 2%, 3%, 4%, or 5% signal intensities. The individual intensity of each community or the average intensity of the communities within the first subset of communities can be used for comparison with the predetermined threshold, the average background noise, and / or the intensity of the second subset of communities.
[0175] In some embodiments, in the first plurality of flow cell images, the image intensity of the second subset of communities or any other community not in the first subset of communities is below a predetermined threshold, so they appear "dark" in the first plurality of flow cell images. In some embodiments, the dark communities are 5 times, 8 times, 10 times, 12 times, 15 times, 20 times, 30 times, 40 times, 50 times, or more times darker than the bright communities (e.g., the average intensity of the bright communities). In some embodiments, the dark communities are 5 times, 8 times, 10 times, 12 times, 15 times, 20 times, or 30 times darker than the bright communities (e.g., the average intensity of the bright communities). In some embodiments, the intensity of the dark communities is about 0.5 times, 0.8 times, 1 time, 1.2 times, 1.5 times, 2 times, or 5 times the average background noise level of the corresponding flow cell images.
[0176] In some embodiments, the first subset of communities is different from the second subset of communities. In some embodiments, the first subset of communities and the second subset of communities at least partially spatially overlap in the x - y plane or in 3D. In some embodiments, the first subset of communities and the second subset of communities contain the same batch - specific sequencing binding sites, which are configured to bind to the same sequencing primers, such that they are "sub - batches" within the same batch of communities and clusters. In some embodiments, each community in the first subset of communities and the second subset of communities contains the same batch - specific sequencing binding sites, which are configured to bind to the same sequencing primers.
[0177] In some embodiments, each colony in the first subset of colonies is configured to bind to a first sequencing primer, and each colony in the second subset of colonies is configured to bind to a second sequencing primer different from the first sequencing primer. Thus, the first subset of colonies and the second subset of colonies are not within the same "batch".
[0178] In some embodiments, a first set of base identifications corresponds to a first set of colonies in a first plurality of sequencing cycles. In some embodiments, the first set of base identifications includes only 1, 2, or 3 types of nucleotide bases. In some embodiments, a second set of base identifications corresponds to a second set of colonies in a second plurality of sequencing cycles. In some embodiments, the second set of base identifications includes only 1, 2, or 3 types of nucleotide bases. For example, when the alternating dark channel is not imaged, the first set of base identifications may include A, T, C, but not G corresponding to the first dark channel, while the second set of base identifications may include T, C, and G, but not A corresponding to the second dark channel.
[0179] In some embodiments, the first set of base identifications includes all four types of nucleotide bases. In some embodiments, a second set of base identifications corresponds to a second set of colonies in a second plurality of sequencing cycles. In some embodiments, the second set of base identifications includes all four types of nucleotide bases. For example, each of the first set of base identifications and the second set of base identifications may include all four types of nucleotide bases, where G is the dark channel in the first plurality of cycles, and A is the second dark channel in the second plurality of cycles.
[0180] In some embodiments, the operation of "obtaining a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from one or more channels" 5210" includes: obtaining, by an optical system of the sequencing system, a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from a first subset of channels, where the first subset of channels may or may not include the first dark channel. In some embodiments, the operation of "obtaining a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from a first subset of channels" 5210" includes: controlling, by a processor, the optical system to avoid collecting data from one or more image sensors of the dark channel in the first plurality of sequencing cycles; controlling, by a processor, the optical system to avoid irradiating light within a predetermined frequency range corresponding to the first dark channel to the sample in the first plurality of sequencing cycles.
[0181] In some embodiments, the operation 5220 of obtaining a second plurality of flow cell images of a sample in a second plurality of sequencing cycles from a second subset of channels includes: obtaining, by an optical system of the sequencing system, the second plurality of flow cell images of the sample in the second plurality of sequencing cycles from the second subset of channels, where the second subset of channels may or may not include a second dark channel. In some embodiments, the operation 5220 of obtaining a second plurality of flow cell images of a sample in a second plurality of sequencing cycles from one or more channels includes: controlling, by a processor, the optical system to avoid collecting any data from one or more image sensors of any dark channels in the second plurality of sequencing cycles. In some embodiments, the operation 5220 of obtaining a second plurality of flow cell images of a sample in a second plurality of sequencing cycles from a second subset of channels includes: controlling, by a processor, the optical system to avoid irradiating light within a predetermined frequency range to the sample in the second plurality of sequencing cycles, where the predetermined frequency range corresponds to the second dark channel.
[0182] In some embodiments, the operation 5230 of generating a first set of base calls for a first subset of communities of a sample based on a first plurality of flow cell images includes: generating a first set of base calls for a first subset of communities of the sample based on the first plurality of flow cell images obtained from a first subset of channels. In some embodiments, the operation 5240 of generating a second set of base calls for a second subset of communities of a sample based on a second plurality of flow cell images includes: generating a second set of base calls for a second subset of communities of the sample based only on the second plurality of flow cell images from a second subset of channels.
[0183] In some embodiments, the first plurality of sequencing cycles includes the same number of cycles as the second plurality of sequencing cycles. For example, each of the first plurality of cycles or the second plurality of cycles may include 2, 3, 4, 5, 6, 7, 8 or more consecutive cycles in a sequencing run. In some embodiments, each of the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 2 to 60 cycles. In some embodiments, the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 2 to 40 cycles. In some embodiments, the first plurality of sequencing cycles or the second plurality of sequencing cycles includes 3 to 15 cycles.
[0184] A sequencing run may include any non - zero integer number of sequencing cycles. For example, a sequence run has 12, 15, 24, 30 or more consecutive sequencing cycles corresponding to barcode sequences, e.g., 15 cycles as shown in Figures 11A - 11C which may be located at any position in the sequencing run.
[0185] In some embodiments, the method 5200 herein further includes an operation of obtaining a third plurality of flow cell images of a sample in a third plurality of sequencing cycles from a third subset of channels.
[0186] One or more channels may include some or all of a first dark channel, a second dark channel, and / or a third dark channel. One or more channels may not include all of the channels of the sequencing system, such as 4 channels. The operations "5210", "5220", and the operation of obtaining a third plurality of flow cell images may use different subsets of the channels. The channels for obtaining one of the first plurality of flow cell images, the second plurality of flow cell images, the third plurality of flow cell images, or other plurality of flow cell images may be the same as or may not be the same as the channels for obtaining another of the first plurality of flow cell images, the second plurality of flow cell images, the third plurality of flow cell images, or other plurality of flow cell images. For example, in the first 5 cycles, a first subset of the channels corresponding to nucleotide bases A, T, and C is used to obtain the flow cell images, and the dark channel corresponds to nucleotide base G. In the next 5 cycles, a second subset of the channels corresponds to A, G, C, and the dark channel may be changed and corresponds to nucleotide base T. In the third 5 cycles, a third subset of the channels corresponds to nucleotide bases A, G, and T, and the dark channel may be changed and corresponds to nucleotide base C.
[0187] In some embodiments, the channels for obtaining the first plurality of flow cell images in the first plurality of sequencing cycles may or may not include one or more dark channels. Similarly, the channels for obtaining the second plurality of flow cell images or the third plurality of flow cell images in their corresponding sequencing cycles may or may not include one or more dark channels.
[0188] In some embodiments, the signal intensity of the flow cell images may be much darker, e.g., 10 times smaller than the bright spots in the flow cell images from other channels and comparable to the background noise, such that the flow cell images obtained from or obtained by the dark channels may not be used for base calling to avoid possible errors in base calling. Additionally, the time and computation for using the flow cell images to perform base calling may be reduced. In other words, the flow cell images from bright channels (e.g., 2 or 3 bright channels) may be used to perform base calling.
[0189] In some embodiments, there is only 1 dark channel. In some embodiments, there may be multiple dark channels, and each dark channel or each combination of dark channels corresponds to a specific set of different cycles. In some embodiments, the dark channels alternate in different cycles. In other words, the first dark channel in the first plurality of cycles may not be dark in the second plurality of cycles, while the second dark channel that is not dark in the first plurality of cycles may be dark in the second plurality of channels. In some embodiments, the first dark channel corresponds to a fluorescent dye attached to a first nucleotide of adenine (A), thymine (T), guanine (G), or cytosine (C) that emits light below a predetermined threshold in the first plurality of sequencing cycles, and the second dark channel corresponds to a second fluorescent dye attached to a second nucleotide of A, T, G, or C that emits light below a predetermined threshold in the second plurality of sequencing cycles, and the first nucleotide and the second nucleotide are different. When the fluorescent dye corresponding to the first nucleotide is not applied to the sample, the light emitted from the sample corresponding to the first nucleotide may be zero. When the fluorescent dye corresponding to the dark nucleotide is applied, the light emitted corresponding to the first nucleotide may be darker than other dyes corresponding to "bright" nucleotides, but the concentration of the fluorescent dye is less than that of the other dyes per volume, or the emitted light may be darker or darker than the dark nucleotide with the same concentration among different dyes.
[0190] In some embodiments, the first dark channel corresponds to a first channel from which no flow cell image or only a dark flow cell image is obtained (e.g., where the image intensity from the community is below a predetermined threshold) in the first plurality of sequencing cycles. The second dark channel may correspond to a second channel from which no flow cell image or only a dark flow cell image is obtained (e.g., where the image intensity from the community is below a predetermined threshold) in the second plurality of sequencing cycles. For example, the barcode sequence can be AAAAATTGTCTTTTT. In the first 5 cycles, the channel corresponding to nucleotide A is dark, while in the last 5 cycles, the channel corresponding to nucleotide T is dark. Other barcode sequences within the same group will have "bright" bases that are not A in the first 5 cycles and different "bright" bases that are not T in the last 5 cycles.
[0191] In some embodiments, two different subsets of channels (e.g., between a first subset, a second subset, or a third subset of channels) are at least partially different from each other. For example, the first subset may correspond to channels for T, C, and G, while the second subset of channels may correspond to T and C. In some embodiments, two different subsets of channels (e.g., a first subset, a second subset, or a third subset of channels) are the same as each other. The channels herein may include 3 or 4 different channels. In some embodiments, two of the 3 or 4 channels may have the same first fluorescent color, while the other one or two channels may have the same second fluorescent color.
[0192] In some embodiments, each of the first set of base identifications or the second set of base identifications includes sequences corresponding to base identifications for a first plurality of sequencing cycles or a second plurality of sequencing cycles. For example, an entry in the first set of base identifications can be "TTGTC", and an entry in the second set of base identifications can be "CTCTT", as Figure 11A shown in
[0193] After determining the first set of base identifications and / or the second set of base identifications, the method 5200 herein (e.g., in FIGS. 52A - 52C) can further include an operation of determining whether one or more base identification sequences in the first set of base identifications and / or the second set of base identifications match at least a portion of a barcode sequence. For example, the sequence "TGGTC" may match the barcode of the community from subset 3, where subset 3 is "TTGTC" with 1 base mismatch. A mismatch rate of 1 / 5 may or may not satisfy the error tolerance rate. The error tolerance rate can be preset. If the error tolerance rate is satisfied and in response to the determination, the method can include an operation of assigning one or more sequences of the first set of base identifications and / or the second set of base identifications to the corresponding barcode sequence, thereby associating the community with a DNA or RNA fragment that uniquely identifies the barcode sequence. In some cases, the sequence determined from the base identification only matches a portion of the barcode, rather than the entire barcode. The matching barcode portion can be a fragment with consecutive bases, e.g., bases 1 - 5 or bases 10 - 15 of the barcode sequence.
[0194] In some embodiments, each barcode sequence is used to uniquely identify a DNA or RNA fragment of a sample. The barcode sequences can be pre - determined. The barcode sequences can be saved and retrieved later as a reference for matching base identifications when needed. The number of barcode sequences can determine how many subsets the community can be divided into.
[0195] Figure 13 An exemplary barcode sequence is shown that can be used to achieve accurate and reliable sequencing and analysis, where the spatial density of the sample is 4 times higher than that of a typical 3D sample that can be sequenced using the existing system with method 5200. The community is divided into 4 subsets, and each subset shares a barcode sequence that can uniquely identify a DNA or RNA fragment of the sample. The same barcode sequence can be shared within the same subset of the community. The first community subset is imaged in the 1st to 5th cycles, while the other communities are dark. Similarly, the communities in subsets 2, 3, and 4 "light up" sequentially in later sequencing cycles. Images from the channel corresponding to nucleotide A are not acquired to save imaging time. Thus, the system throughput is higher than imaging the same community sample using the existing system.
[0196] In some embodiments, different barcode sequences are determined such that only a certain number of barcode sequences have "bright" bases in a particular sequencing cycle. As Figures 11A - 11C and Figure 13 shown, at the first sequencing cycle n+1, only one barcode sequence in its own group of barcode sequences has a "bright" nucleobase, and the remaining barcode sequences in the corresponding group at the n+1-th cycle are all "dark". Figure 11C Shows a set of exemplary barcode sequences that includes 3 different barcode sequences, where the "dark" channel corresponds to A. In some embodiments, when there is a relatively large number of barcode sequences (e.g., 10 or 40), more than one barcode sequence may have a "bright" base in the same sequencing cycle.
[0197] In some embodiments, a fragment of the barcode sequence contains only 1, 2, or 3 bases, and the remaining part of the barcode sequence contains all four different bases. In some embodiments, one or more reference cycles (e.g., the first 5 cycles) correspond to the barcode sequence and correspond to only three bases, and the remaining part of the barcode sequence corresponds to subsequent cycles and contains all four different bases. In some embodiments, the barcode sequence contains only three bases. In some embodiments, the barcode sequence contains all four bases. The barcode sequence can have from about 2 to about 100 nucleobases. In some embodiments, the barcode sequence has from about 3 to 60 nucleobases. In some embodiments, the barcode sequence has from about 3 to 30 nucleobases.
[0198] In some embodiments, the barcode sequences used in correspondence with method 5200 (e.g., in FIGS. 52A-52C) can have some random bases. For example, to avoid homopolymers, i.e., the same base repeated continuously n times, where n is greater than 4, 5, 6, or even larger, after randomly selecting a preset number of consecutive repeats of the same nucleobase from 3 nucleobases other than the same nucleobase, the barcode sequence can have one or more nucleotides. For example, Figure 11B showing "N" as a random base that can be randomly selected from bases T, G, and C. In some embodiments, the barcode sequence contains the same unique base repeated no more than 3, 4, 5, 6, 7, 8, 9, or 10 times.
[0199] In correspondence with the randomized base (e.g., Figure 11BIn the sequencing cycle of the base N), all communities in all subsets (e.g., the first subset, the second subset, and the third subset) are not dark. The communities can have spatial densities where they may overlap with each other, and base identification from the flow cell image may be unreliable. In some embodiments, the flow cell image in the sequencing cycle corresponding to the randomized base is not acquired. Alternatively, the capture step of capturing a specific affinity for imaging can also be skipped to further save sequencing run time.
[0200] In some embodiments, the method 5200 herein (e.g., in FIGS. 52A - 52C) can include an operation by a processor and for a sequencing cycle to determine whether the image intensities from both the first set of communities and the second set of communities are higher than a predetermined threshold. If it is determined to be the case, then two subsets of the communities emit light and they may overlap spatially, which may lead to base identification problems. Accordingly, in response to this determination, the method can include an operation of not obtaining a flow cell image in the sequencing cycle (e.g., the n + 6 and n + 12 cycles as shown in Figure 11B ), which is the randomized base that resolves the isomers in the barcode sequence. This particular sequencing cycle with the randomized base (e.g., N) can be before and / or after the first plurality of sequencing cycles or the second plurality of sequencing cycles. In some embodiments, this particular sequencing cycle is before and after the first plurality of sequencing cycles. In some embodiments, this particular sequencing cycle is before and after the second plurality of sequencing cycles.
[0201] In some embodiments, the method 5200 herein (e.g., in FIGS. 52A - 52C) advantageously allows for sequencing of 3D samples with high spatial density. The prior art may be limited by the cassette design and the characteristics of the optical system such that when the spatial density of the 3D sample increases, the communities and clusters overlap with other base identifications, and subsequent sequencing analysis may be inaccurate or unreliable. The method herein can sequence 3D samples at a spatial density that is n times higher than what existing systems and methods can handle without the need to change the cassette design or the optical system of the sequencing system. The quantity n can depend on different existing systems and samples. In some embodiments, n is in the range of 2 to 100. In some embodiments, n is in the range of 2 to 20. In some embodiments, n is in the range of 2 to 15. In some embodiments, n is the total number of barcode sequences used.
[0202] In some embodiments, method 5200 herein (e.g., in FIGS. 52A - 52C) may include operations for preparing a sample for imaging in each sequencing cycle. In some embodiments, method 5200 herein may include an operation of contacting at least one subset of a first community subset, a second subset, and a third subset of a sample with a first mixture of a plurality of sequencing primers, a first plurality of polymerases, and different types of adaptors. Adaptors will be disclosed in more detail below. An individual adaptor in the first mixture includes a nucleus attached with a plurality of nucleotide arms, and each arm of the individual adaptor includes the same type of nucleotide unit, such as A, T, C, or G.
[0203] Figures 12A - 12C shows the steps in an exemplary sequencing cycle using 4 different types of adaptors ( Figure 12A ), 3 different types of adaptors ( Figure 12B ), and 4 different types of adaptors, including one type of adaptor with a "dark" fluorescent dye ( Figure 12C ). The flow operation of a mixture of differently labeled adaptors can be performed by sequencing system 110. A continuous flow or a single flow can be used to flow different types of labeled adaptors into the flow cell. Imaging can follow each flow or a single flow in the continuous flow to capture the light emitted from the community. The operations of flowing the adaptors and obtaining an image of the flow cell can be performed by the sequencing system in various ways. The disclosure herein does not limit how the reagents, including adaptors, specifically flow into contact with the community and the timing of imaging related to the flow operation.
[0204] As Figure 12B shown, the flow lacks one type of adaptor corresponding to nucleotide A, so at least 1 / 4 of the total COGS and / or adaptors can be saved. Figure 12B , in some embodiments, the flow may include 4 different types of adaptors, including one type of adaptor not labeled with any fluorescent dye. Alternatively, one type of adaptor is labeled with a "dark" fluorescent dye that does not emit any fluorescence when excited, or only emits fluorescence beyond the detectable level of the optical system of sequencing system 110. A continuous flow or a single flow of a mixture of 3 differently labeled adaptors can be applied to the sample on the support. Imaging the sample can follow each individual flow or a single flow, and the imaging can include detecting 3 different labeled adaptors bound to the community such that at least one subset of the community is "dark".
[0205] In some embodiments, a first mixture of different types of binders includes four different types of binders corresponding to four different nucleotide bases. Each type of the different types of binders in the first mixture can be labeled with a type of fluorescent dye corresponding to a nucleotide unit to distinguish the different types of binders in the first mixture. In some embodiments, the fluorescent dyes of each type of binder in the first mixture emit light of different wavelengths when excited. In some embodiments, one type of binder is labeled with a type of dark fluorescent dye that emits light below a predetermined threshold in a channel, while the fluorescent dyes of other types of binders emit light above a second predetermined threshold, which may be different from the first threshold.
[0206] In some embodiments, the first mixture of different types of binders contains only two or three different types of binders. Two or three types of the different types of binders in the first mixture are labeled with a corresponding type of fluorescent dye corresponding to a nucleotide unit to distinguish two or three types of the different types of binders in the first mixture. The first mixture may lack the fourth type of binder, so a type of nucleotide base (e.g., A) may not come into contact with its corresponding binder. Thus, during imaging, a type of nucleotide does not emit light (if any) and appears "dark".
[0207] In some embodiments, the operations of generating a first set of base identifications for a first community subset, second subset, third subset, or other subset of a sample based on the first plurality of flow cell images include one or more operations of the 3D base identification method 6000 disclosed herein. In some embodiments, such base identification operations include: generating, by a processor, a first plurality of processed images of the first plurality of flow cell images; filtering, by the processor, the first plurality of flow cell images based on the first plurality of processed images to generate a first plurality of filtered images; obtaining, by the processor, a 3D community map of the sample; extracting, by the processor, the image intensity of the community from among: a second plurality of flow cell images; a second plurality of processed images; a second plurality of filtered images; or a combination thereof; based on the 3D community map; and performing 3D base identification on the first community subset of the sample based on the extracted image intensity of the community.
[0208] In some embodiments, the operations of generating a first set of base identifications for a first community subset, second subset, third subset, or other subset of a sample based on the first plurality of flow cell images include one or more operations of the image registration method 9000 disclosed herein.
[0209] In some embodiments, the operation of obtaining a plurality of flow cell images includes actively retrieving or passively receiving a plurality of flow cell images of a sample to be processed. In some embodiments, the operation includes using the imager 1160 of the sequencing system to obtain the flow cell images.
[0210] The sample can be in-situ. The sample can be a 3D sample. The sample can be a volumetric sample that may contain different biological information at the same x-y position but different z positions. The sample can include a plurality of cells, tissues, or combinations thereof. The 3D sample can be any biological sample having a thickness greater than a predetermined threshold along the axial axis. For example, the thickness can be greater than 2um, 3um, 4um, 5um, 10um, 20um or greater. The z-axis (e.g., the axial axis) is orthogonal to the image plane defined by the x-axis and the y-axis.
[0211] The flow cell images can be obtained from one, two, three, four, or more channels of the imager 1160 using the optical systems disclosed herein. Each flow cell image can include one or more tiles (e.g., imaging regions), and each tile can be divided into a plurality of sub-tiles. Each tile or sub-tile can include a plurality of colonies. Each sub-tile can include a plurality of regions, where each region includes a number of colonies. For example, colonies can be extracted from corresponding regions of flow cell images from 4 different channels in a given cycle. As another example, colonies can be extracted from a flow cell image from a single channel. The flow cell images disclosed herein can be images obtained from the flow cell 1120 as shown in Figure 1 the image shown in.
[0212] In some embodiments, the flow cell images herein can be images of one or more tiles, one or more sub-tiles, one or more segmented regions within a tile or sub-tile, or a combination thereof. Each flow cell image can contain a field of view (FOV). The FOV can be orthogonal to the axial axis. The FOV can be in the x-y plane. The FOVs of different flow cell images at different axial positions can be the same in the x-y plane. The FOVs of different flow cell images at different axial positions can have at least an overlapping portion in the x-y plane. The image resolutions of different flow cell images at different axial positions can be substantially the same or exactly the same. In some embodiments, the image resolutions of different flow cell images at different axial positions are different. Figure 2A and 3A shows two exemplary flow cell images obtained at two different z positions along the axial axis of the same 3D sample within the same sequencing cycle.
[0213] Each flow cell image at a particular z position includes intensities generated by colonies and clusters at the corresponding z position. As Figures 2A - 3AAs shown, the signals from the colonies and clusters are small bright spots in the image. Each bright spot can have various sizes less than a few pixels, such as less than one pixel, about one pixel, about 2 pixels, 3 pixels, 4 pixels, or 5 pixels. In some embodiments, each signal point of a colony or cluster can be any number of pixels in the range from 0.01 pixel to about 72 pixels. In some embodiments, each signal point of a colony or cluster can be any number of pixels in the range from 0.1 pixel to about 16 pixels.
[0214] Each flow cell image can also include intensities generated by the cells and their structural elements. Such structural elements can be background objects or components.
[0215] In some embodiments, when the depth of field of the optical system includes a range extending along the z-axis, such as 0.1um, 0.2um, 0.3um, 0.5um, 0.6um, 0.8um, 1um, 2um, 3um, 4um, 5um, etc. The colonies and clusters within the depth of field can appear focused or substantially focused in the flow cell image. The flow cell image at a particular z position can also include signals from colonies and clusters that are not within the in-focus range of the image. Such colonies or clusters are out of focus. As Figure 3A shown, the larger and blurred signal points represent out-of-focus colonies or clusters. Some out-of-focus colonies or clusters are circled in Figure 3A shown.
[0216] Each flow cell image at a particular z position can also include noise caused by the optical system and / or unwanted signals from the sample. The unwanted signals can be signals from components of the sample (such as membranes, cytoplasm, and mitochondria). Such background objects can be any object that is relatively larger in size than the colonies or clusters. As Figure 3A shown, there is a blurred cell contour (at the arrow) in the flow cell image, and most of the signal points are contained within the blurred contour. Figure 3D Multiple cells and colonies or clusters are shown because the small bright spots are generally located within the contours of different cells. In some embodiments, the background objects can include any object within the 3D sample, but not colonies or clusters.
[0217] 3D Base Identification
[0218] In some embodiments, the operation 2530 of generating a first set of base identifications for a first subset of colonies of a sample by a processor based on a first plurality of flow cell images; or the operation 2540 of generating a second set of base identifications for a second subset of colonies of a sample by a processor based on a second plurality of flow cell images can include one or more base identification methods disclosed herein (e.g., method 6000).
[0219] In some embodiments, method 6000 may include operations comprising one or more of the following: generating, by a processor, a first plurality of processed images of a first plurality of flow cell images; filtering, by the processor, the first plurality of flow cell images based on the first plurality of processed images to generate a first plurality of filtered images; obtaining, by the processor, a 3D community map of a sample; extracting, by the processor, image intensities of a community from among the following based on the 3D community map: a second plurality of flow cell images; a second plurality of processed images; a second plurality of filtered images; or a combination thereof; and performing, by the processor, 3D base calling on a first community subset of the sample based on the extracted image intensities of the community.
[0220] In some embodiments, method 6000 is performed during cycle N that is different from a reference cycle. A template image may be generated in the reference cycle, and communities from one or more channels within the reference cycle may be included in the template image in a reference coordinate system, while base calling for cycle N has not been performed. In some embodiments, cycle N is the current cycle. N may be any non-zero integer. For example, for short read sequencing, N may be any integer from 1 to 150.
[0221] Figure 8B A schematic diagram showing one or more template images generated in a reference coordinate system in a reference cycle is shown. In some embodiments, template image 2100 includes dimensions that are substantially the same as a single tile 2900 including a 5×5 sub-picture grid. In some embodiments, the template images disclosed herein may be separate regions within a sub-tile. Each template image may include a plurality of communities therein.
[0222] In some embodiments, the template image may have approximately the same size as the flow cell image such that different pictures 2900 from Figure 8B and all communities from multiple channels may be registered with the same template image. However, such template images may contain communities that are not used in at least some of the operations described herein to reduce computational burden without sacrificing accuracy.
[0223] In some embodiments, more than one template image may be generated, and each template image corresponds to at least a portion of a sub-tile of flow cell images from a channel.
[0224] The template images herein may be initialized as virtual images that have a black or dark background and no signal from communities. For example, the template image may be initialized to zero, or include a minimum image intensity at all pixels.
[0225] After determining the coordinates of a community through, for example, image registration of flow cell images across different channels, the intensity of the community can be added to a template image at the position determined by the coordinates, where the size and shape are determined based on the registration. The template image can be a virtual image that combines the image intensities of communities obtained from 2, 3, 4, or more channels in a reference cycle. The pixels of the template image without communities remain black or dark, so that the template image can have a cleaner background without the noise present in the actual flow cell images.
[0226] The community can come from a sub-tile of a flow cell image within a reference cycle, and more specifically, from one or more selected regions of the sub-tile. The flow cell image can come from different channels among 1, 2, 3, 4, or more channels of system 1000. As a non-limiting example, the reference cycle can be any of the first 5 or 6 cycles. In some embodiments, the reference cycle can be any cycle greater than 0. In some embodiments, the reference cycle is the first cycle.
[0227] In some embodiments, method 6000 includes the operation of generating a 2D flow cell image with the intensity of communities or clusters of a 3D sample, such that the intensity can be used for performing 3D base calling for different sequencing cycles. The operation of generating the flow cell image can be performed using the sequencing system 1100 herein.
[0228] Method 6000 can include the operation of obtaining a processed image of a subsequent flow cell image. In some embodiments, the operation of obtaining the processed image can include processing the flow cell image with one or more predetermined processing methods.
[0229] In some embodiments, one or more processing methods can include selecting a kernel and generating a processed image by performing an operation on the flow cell image using the selected kernel. For example, the operation performed can be an opening operation, which can be represented as f ° k, where f is the flow cell image and k is the kernel. The opening operation can be performed in the spatial domain. Alternatively, for faster computation or lower complexity, the opening operation can be performed in a different domain (such as the Fourier domain). The opening operation can be the dilation of the erosion of the image f by the kernel k. The opening operation can remove objects smaller than the kernel, and the subsequent dilation operation can restore the size and shape of the remaining objects. Figure 2B and Figure 3B Shows an exemplary processed image after the opening operation. The processed image can also be referred to as the image after the opening operation, where most of the bright spots of the communities or clusters are removed while other background objects are retained.
[0230] As another example, the operation performed can be convolution, and the flow cell images can be convolved with a selected kernel. As another example, obtaining a plurality of processed images further includes: selecting a first kernel and a second kernel; generating a first image by convolving the flow cell images of the first plurality of flow cell images or the second plurality of flow cell images with the first kernel; and generating a second image by convolving the flow cell images of the first plurality of flow cell images or the second plurality of flow cell images with the second kernel. The first image and the second image can be different blurred images after convolution with different blurred kernels.
[0231] The kernel can be of any size smaller than the size of the flow cell image. For example, through morphological opening, the kernel can be 2×2, 3×3, 4×4, 5×5, 6×6. In some embodiments, the kernel size can be customized to remove at least some noise and unwanted signals larger than the kernel size. In some embodiments, the kernel can be circular. The kernel can be of various other shapes.
[0232] In some embodiments, the kernel is a Gaussian kernel. In embodiments where 2 different kernels are used, the first kernel and the second kernel can be different Gaussian kernels.
[0233] In some embodiments, method 6000 can include an operation 6200 of filtering the flow cell images. The filtering operation 6200 can be based on the processed images generated in its previous operation. The filtering operation 6200 can generate a plurality of filtered images, each filtered image corresponding to the flow cell image at a corresponding z - position along the axial axis.
[0234] In some embodiments, operation 6200 can include subtracting the processed image from the corresponding flow cell image to generate a filtered image. The filtered image can be obtained as fi = f - f ° k, where fi represents the filtered image, f represents the flow cell image, and k represents the kernel, and ° represents morphological opening.
[0235] In some embodiments, the filtered image is the image of the corresponding flow cell image filtered by a top - hat filter. In some embodiments, the filtered image is the image of the corresponding flow cell image filtered by a Difference of Gaussians (DoG) filter. In some embodiments, the filtered image is the image of the corresponding flow cell image filtered by other filters configured to extract small elements and details (e.g., focused colonies) from the flow cell image
[0236] After filtering, at least a portion of the noise and undesired signals in the flow cell image are removed, which can include cellular components and out-of-focus populations. Removing such inevitable noise and undesired signals in the 3D flow cell image can advantageously facilitate generating intensities attributable to populations or clusters rather than other background objects or noise. The filtered intensities can be used for more accurate and reliable base calling than the unfiltered intensities.
[0237] Figure 2C and Figure 3C shows two filtered images from two different axial positions. Figure 2C It shows that at the first axial position of z = 0, a large number of populations or clusters are in focus, and at the second axial position of z = 5, a much smaller number of populations or clusters are in focus, and the second axial position is about 10 um away from the first axial position.
[0238] Figure 3E shows Figure 3D another exemplary filtered image of the flow cell image. Background objects in the flow cell image are filtered out in the filtered image. Out-of-focus populations in the flow cell image are also removed from the filtered image. The filtered image shows in-focus populations or clusters.
[0239] In some embodiments, method 6000 can include an operation of adding an offset to the filtered image. In some embodiments, the offset can be predetermined such that after the offset, the range of the image intensities can be within a predetermined range. In some embodiments, different offsets can be used to bring different filtered images having various intensity ranges within a predetermined range common to all the filtered images.
[0240] Method 6000 can include an operation 6300 of generating a maximum intensity projection (MIP) image based on a plurality of filtered images. Operation 6300 can include using the intensities of the plurality of filtered images at corresponding pixels to calculate the maximum intensity of each pixel of the MIP image. The MIP image can have the same image size, resolution, and / or FOV as the flow cell image. The MIP image can be a flattened 2D image of a stack of flow cell images from different axial positions.
[0241] For example, an axial slide of 10 flow cell images is obtained at 10 different z positions. Each z position is approximately 0.1 um to approximately 20 um from its adjacent z position. The first z position is at z = 0, and the 10th z position is at z = 9. Ten filtered images are generated for each different z position. The MIP image can be initialized to the same size as the flow cell image, such as 1028x1028. All intensities in the initial MIP image can be 0 or any other minimum intensity. For each pixel at pixel (i,j), the image intensity at pixel (i,j) is extracted from the 10 filtered images, and the maximum of the 10 different image intensities is selected for pixel (i,j) in the MIP image. In some embodiments, the MIP image can be generated for each cycle. In some embodiments, the MIP image can be generated for each different channel in 1, 2, 3, 4, 5, 6 or more channels.
[0242] In some embodiments, before obtaining the MIP image using the filtered image, the image intensity in the filtered image is normalized to a predetermined range. In some embodiments, the MIP image can be normalized across different cycles. In some embodiments, the MIP images of different FOVs (e.g., covering different tiles) in the x-y plane can be normalized. The normalization can reach a predetermined range.
[0243] In some embodiments, method 6000 includes an operation of registering the MIP image, such as Figure 9 9000 in. In some embodiments, the MIP image is registered across channels and different cycles before performing any base calling. Various image registration techniques can be used to register the MIP image. 2D registration techniques can be used to register the MIP image. For example, by treating the MIP image as a flow cell image obtained from the sequencing system 110. In some embodiments, the image registration method 9000 disclosed herein can be used to register the MIP image by treating each MIP image as a flow cell image, such as across different channels and / or different cycles. In some embodiments, the MIP image can be registered after performing one or more preprocessing operations disclosed herein.
[0244] Method 6000 can include an operation 6400 of performing 3D base calling using the MIP image. The MIP image can be 2D, and base calling can be performed for the MIP image from different channels at each z position for each cycle.
[0245] In some embodiments, operation 6400 of performing base calling using a MIP image includes performing the main analysis steps herein to adjust the image intensity of colonies in the MIP; and performing base calling on the colonies based on the adjusted image intensity in the MIP. In some embodiments, the main analysis steps include one or more of the following: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; image registration; quality score estimation. In some embodiments, image registration, as a main analysis step herein, is configured to align images from different cycles and different channels, for example, relative to a template image or a reference coordinate system. In some embodiments, image registration, as a main analysis step herein, is configured to register colonies or clusters from different cycles and different channels (e.g., in a MIP image) to a template image or a reference coordinate system.
[0246] For example, after registering MIP images from different channels relative to the template image disclosed herein, base calling can be performed using the MIP images from different channels in cycle N. Various existing 2D base calling algorithms can be used. The base calling results can be saved together with their 3D coordinates. Such 3D coordinates can be used to register base calling across different cycles and at different z positions.
[0247] In some embodiments, method 6000 includes operation 6500 of obtaining a second MIP image based on a plurality of flow cell images. The second MIP is a flattened 2D image of an axial stack of flow cell images. In some embodiments, the flow cell images are raw images acquired by sequencing system 110. The operation of obtaining the second MIP can be different from the operation of obtaining the first MIP image. In a flow cell image (e.g., a raw image), out-of-focus colonies can have a larger full width at half maximum (FWHM), and thus the signal of out-of-focus colonies is more dispersed than that of in-focus colonies or clusters. In some embodiments, the larger FWHM can result in a white ring around the colony. Figure 5 An exemplary second MIP image generated directly from a flow cell image is shown, and there is a ring or halo around the colony in the center of the image. The ring or halo may be an image artifact that can cause errors in base calling. The second MIP can also retain some background information (e.g., background information from unwanted background objects), which may interfere with base calling if the second MIP is used for base calling. Thus, the second MIP includes artifacts and / or unwanted background objects that require additional processing before accurate and reliable base calling based on the intensity in the second MIP.
[0248] In some embodiments, the second MIP can be used for the registration of the first MIP and / or the community or cluster because they share the same FOV, resolution, image size, etc. If the second MIP is used for base identification, the artifacts and / or unwanted background objects in the second MIP may interfere with base identification, but the same artifacts and / or unwanted background objects (not in the first MIP) can facilitate the registration of the first MIP by using the second MIP, for example, for registration with the pooled image with staining.
[0249] Figure 6B A flowchart of a computer-implemented method 6000 for performing 3D base identification from a flow cell image is shown. Method 6000 may include some or all of the operations disclosed herein. The operations may be performed in an order that is not limited to that described herein.
[0250] Method 6000 may be executed by one or more processors disclosed herein. In some embodiments, the processor may include one or more of the following: a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit may include a central processing unit (CPU), a graphics processing unit (GPU), or an NPU. The integrated circuit may include a chip such as a field-programmable gate array (FPGA). In some embodiments, the processor may include a computer system 4000.
[0251] Method 6000 may be executed based on the flow cell image from the current sequencing cycle alone, or in combination with information from previous cycles of the current sequencing cycle.
[0252] In some embodiments, some or all of the operations in method 6000 may be executed by an FPGA. In an embodiment, when some operations are executed by an FPGA, the data after the operations are executed by the FPGA may be transmitted by the FPGA to the CPU so that the CPU can use such data to perform subsequent operations in method 6000. Similarly, data may also be transmitted from the CPU to the FPGA for processing by the FPGA. In some embodiments, all of the operations in method 6000 may be executed by the CPU. Alternatively, the operations executed by the CPU may be executed by other processors (such as a dedicated processor) or a GPU. In some embodiments, all of the operations in method 6000 may be executed by an FPGA.
[0253] Method 6000 may include an operation 6100 of obtaining a plurality of flow cell images of a sample. The flow cell images may be acquired at different positions along the axial axis (i.e., the z-axis). In some embodiments, operation 6100 includes actively retrieving or passively receiving a plurality of flow cell images of the sample to be processed. In some embodiments, operation 6100 includes using the imager 1160 of the sequencing system to acquire the flow cell images. In some embodiments, the flow cell images are acquired by an NGS sequencing system.
[0254] The sample can be in 3D. The sample can be a volumetric sample that may contain different biological information at the same x-y position but different z positions. The sample can be an in-situ sample. The sample can include multiple cells, tissues, or a combination thereof. The sample can be any biological sample having a thickness greater than a predetermined threshold (e.g., depth of field) along an axial axis. The sample can be any biological sample having a thickness greater than the depth of field of an optical system. For example, the thickness can be greater than 1 um, 2 um, 3 um, 4 um, 5 um, 8 um, 10 um, 12 um, 15 um, 20 um or greater. The z-axis (e.g., axial axis) can be orthogonal to the image plane defined by the x-axis and y-axis, as Figure 8B shown.
[0255] In some embodiments, the flow cell images can include a first plurality of flow cell images and a second plurality of flow cell images. Each of the first plurality of flow cell images and the second plurality of flow cell images can be acquired at corresponding positions along the axial axis or z-axis. The axial axis can extend from the objective lens of the sequencing system 1100 to the sample located on the flow cell, which is positioned on the sequencing system. As a non-limiting example, the first plurality of flow cell images can be obtained from a reference cycle, and the second plurality of flow cell images can be obtained from one or more cycles different from the reference cycle. As another example, the first plurality of flow cell images can be the same as the second plurality of flow cell images.
[0256] The flow cell images can be acquired from 1, 2, 3, 4 or more channels of the imager 1160 using the optical system disclosed herein. Each flow cell image can include one or more tiles (imaging regions), and each tile can be divided into multiple sub-tiles. Each sub-tile can include multiple colonies or clusters. Each sub-tile can include multiple regions, where each region includes a number of colonies. For example, colonies can be extracted from corresponding regions of flow cell images from 4 different channels in a given cycle. As another example, colonies can be extracted from flow cell images from a single channel. The flow cell images disclosed herein can be images acquired from the flow cell 1120 as Figure 1 shown.
[0257] In some embodiments, the flow cell images herein can be images of one or more tiles, one or more sub-tiles, one or more segmented regions having tiles or sub-tiles, or a combination thereof. Each flow cell image can contain a field of view (FOV). The FOV can be orthogonal to the axial axis. The FOV can be in the x-y plane. The FOVs of different flow cell images at different axial positions can be the same in the x-y plane. The FOVs of different flow cell images at different axial positions can have at least an overlapping portion in the x-y plane.
[0258] The image resolutions of different flow cell images at different axial positions can be approximately the same or exactly the same. In some embodiments, the image resolutions of different flow cell images at different axial positions are different.
[0259] Figure 2A and 3A shows two exemplary flow cell images acquired at two different z - positions along the axial axis of the same 3D sample within the same sequencing cycle.
[0260] Each flow cell image at a particular z - position includes intensities generated by the communities and clusters at the corresponding z - position. As Figures 2A - 3A shown, the signals from the communities and clusters are small bright spots in the image. Each bright spot can have various sizes less than a few pixels, e.g., less than one pixel, about one pixel, about 2 pixels, 3 pixels, 4 pixels, 5 pixels or more. In some embodiments, each signal point of a community or a cluster can be any number of pixels in the range from 0.01 pixel to about 72 pixels. In some embodiments, each signal point of a community or a cluster can be any number of pixels in the range from 0.1 pixel to about 16 pixels.
[0261] Communities can come from sub - tiles of the flow cell image within a reference cycle, and more specifically, from one or more selected regions of the sub - tile. The flow cell images can come from different channels among 1, 2, 3, 4 or more channels of the system 1000. As a non - limiting example, the reference cycle can be any of the first 5 or 6 cycles. In some embodiments, the reference cycle can be any cycle greater than 0. In some embodiments, the reference cycle is the first cycle.
[0262] In some embodiments, the flow cell images are acquired under one or more reference cycles or cycles different from one or more reference cycles. As non - limiting examples, cycles 1 - 5 or 2 - 5 can be reference cycles. The flow cell images can be acquired within a single cycle or several cycles.
[0263] Each flow cell image can include intensities generated by the cell and its structural elements. Such structural elements can be background objects or components. Figure 3D shows multiple cells and communities or clusters because the small bright spots are typically within the contours of different cells.
[0264] In some embodiments, when the focus of the optical system includes a range extending along the z-axis, such as 0.1um, 0.2um, 0.3um, 0.5um, 0.6um, 0.8um, 1um, 2um, 3um, 4um, 5um, etc. Communities and clusters within the range of the focus may appear focused or approximately focused in the flow cell image. The flow cell image at a specific z-position may also include signals from communities and clusters that are not within the focus range of the image but at different z-positions. Therefore, such communities or clusters are out of focus. As Figure 3A shown, the larger and blurred signal points represent out-of-focus communities or clusters. Some out-of-focus communities or clusters are circled in Figure 3A .
[0265] Each flow cell image at a specific z-position may also include noise caused by the optical system and / or unwanted signals from the sample. Unwanted signals may be signals from components of the sample (such as membranes, cytoplasm, and mitochondria). Such background objects may be any object that is relatively larger in size than the community or cluster. As Figure 3A shown, there is a blurred cell contour (at the arrow) in the flow cell image, and most of the signal points are contained within the blurred contour. In some embodiments, the background object may include any object within the 3D sample but not the community or cluster.
[0266] In some embodiments, base identification from communities includes 4 different bases, and the percentage of communities for each of the 4 different bases may be greater than about 10%, such that the data is relatively diverse. In some other embodiments, the bases identified from multiple communities include 4 or fewer different bases, and the percentage of communities for one or more bases may be less than about 10%, and such data may be considered low-diversity data. In some embodiments, the bases identified from multiple communities include 4 or fewer different bases, and the percentage of communities for some bases may be less than about 5%, about 2%, or even about 1%, and such data may be considered low-diversity data. As an example, the bases identified for bases A, T, C, G in multiple communities may be about 1%, about 2%, about 1%, and about 95%. As another example, the bases identified for bases A, T, C, G in multiple communities may be about 10%, about 10%, about 10%, and about 70% respectively. In addition to base bias affecting diversity, complexity is also a factor. When the complexity is below a certain number (e.g., 8 or 16), the diversity of the signal may be low. Method 6000 is configured to process flow cell images, for example, including intensity processing, registration, and base identification, even if the communities are low-diversity data.
[0267] In some embodiments, method 6000 is performed at least in part during cycle N that is different from a reference cycle. The template image and / or the 3D community map may be generated in the reference cycle, and communities from one or more channels within the reference cycle may be included in the template image and / or the 3D community map in a reference coordinate system, while base calling for cycle N has not been performed. In some embodiments, cycle N is the current cycle. N can be any non-zero integer. For example, for short read sequencing, N can be any integer from 1 to 150.
[0268] Figure 8B A schematic diagram showing one or more template images generated in a reference coordinate system in a reference cycle is shown. In some embodiments, the template image 2100 includes dimensions that are substantially the same as a single tile 2900 that includes a 5×5 sub-picture grid. In some embodiments, the template images disclosed herein may be separate regions within a sub-tile. Each template image may include a plurality of communities therein.
[0269] In some embodiments, the template image may have approximately the same size as the flow cell image, such that Figure 8B the different pictures 2900 from and all communities from multiple channels may be registered with the same template image. However, such template images may contain communities that are not used in at least some of the operations described herein to reduce the computational burden without sacrificing accuracy.
[0270] In some embodiments, more than one template image may be generated for the same axial position, and each template image corresponds to at least a portion of a sub-tile of the flow cell image from a channel.
[0271] The template images herein may be initialized as virtual images that have a black or dark background and no signal from communities. For example, the template image may be initialized to zero or include a minimum image intensity at all pixels.
[0272] After determining the coordinates of the communities by, for example, image registration of flow cell images across different channels, the intensity of the communities may be added to the template image at the positions determined by the coordinates, where the size and shape are determined based on the registration. The template image may be a virtual image that combines the image intensities of communities obtained from 2, 3, 4, or more channels in the reference cycle. The pixels of the template that do not contain communities remain black or dark, such that the template image may have a cleaner background without the noise present in the actual flow cell image.
[0273] In some embodiments, method 6000 includes an operation of generating an axial stack of 2D flow cell images. The flow cell images can include intensities of communities or clusters from a 3D sample. The intensities can be used for 3D base calling for different sequencing cycles. The operation of generating the 2D flow cell images can be performed by the sequencing system 1100 herein.
[0274] Method 6000 can include an operation of generating a processed image of the flow cell image. In some embodiments, the operation of generating the processed image can include processing the flow cell image with one or more predetermined processing methods.
[0275] In some embodiments, one or more processing methods can include selecting a kernel and generating a processed image by performing an operation on the flow cell image using the selected kernel. For example, the operation performed can be an opening operation, which can be represented as f ° k, where f is the flow cell image and k is the kernel. The opening operation can be performed in the spatial domain. Alternatively, for faster computation or lower complexity, the opening operation can be performed in a different domain, such as the Fourier domain. The opening operation can be the dilation of the erosion of the image f by the kernel k. The opening operation can remove objects smaller than the kernel, and the subsequent dilation operation can restore the size and shape of the remaining objects. Figure 2B and Figure 3B An exemplary processed image after the opening operation is shown. The processed image can also be referred to as the image after the opening operation, in which most of the bright spots of the communities or clusters are removed while other background objects are retained.
[0276] As another example, the operation performed can be a convolution, and the flow cell image can be convolved with the selected kernel. As another example, obtaining multiple processed images further includes: selecting a first kernel and a second kernel; generating a first image by convolving multiple flow cell images with the first kernel; and generating a second image by convolving multiple flow cell images with the second kernel. The first image and the second image can be different blurred images after convolution with different blurring kernels.
[0277] The kernel can be of any size smaller than the size of the flow cell image. For example, through the opening operation, the kernel can be 2x2, 3x3, 4x4, 5x5, 6x6 pixels. In some embodiments, the kernel size can be customized to remove at least some noise and unwanted signals larger than the kernel size. In some embodiments, the kernel can be circular. The kernel can be various other shapes.
[0278] In some embodiments, the kernel is a Gaussian kernel. In embodiments where 2 different kernels are used, the first kernel and the second kernel can be different Gaussian kernels.
[0279] In some embodiments, method 6000 may include an operation 6200 of filtering a flow cell image. The filtering operation 6200 may be based on the processed image generated in its previous operation. The filtering operation 6200 may generate a plurality of filtered images, each filtered image corresponding to the flow cell image at a corresponding z - position along the axial axis.
[0280] In some embodiments, operation 6200 may include subtracting the processed image from the corresponding flow cell image to generate a filtered image. The filtered image may be obtained as fi = f - f ° k, where fi represents the filtered image, f represents the flow cell image, and k represents the kernel, and ° represents an opening operation.
[0281] In some embodiments, the filtered image is an image of the corresponding flow cell image filtered by a top - hat filter. In some embodiments, the filtered image is an image of the corresponding flow cell image filtered by a Difference of Gaussians (DoG) filter. In some embodiments, the filtered image is an image of the corresponding flow cell image filtered by various filters that extract small elements and details (such as focused colonies).
[0282] After filtering, at least a portion of the noise and unwanted signals in the flow cell image are removed, which may include cellular components and out - of - focus colonies. Removing such noise and unwanted signals advantageously facilitates generating intensities attributable to colonies or clusters rather than other background objects or noise. The filtered intensities can be used for more accurate and reliable base calling than unfiltered intensities.
[0283] Figure 2C and Figure 3C shows two filtered images from two different axial positions. Figure 2C It shows that at a first axial position of z = 0, a large number of colonies or clusters are in focus, and at a second axial position of z = 5, a much smaller number of colonies or clusters are in focus, and the second axial position is about 10 um away from the first axial position.
[0284] Figure 3E shows Figure 3D another exemplary filtered image of the flow cell image in. Background objects in the flow cell image are filtered out in the filtered image. Out - of - focus colonies in the flow cell image are also removed from the filtered image. The filtered image shows focused colonies or clusters.
[0285] In some embodiments, method 6000 may include the operation of adding an offset to the filtered image. In some embodiments, the offset may be predetermined such that after the offset, the range of the image intensity may be within a predetermined range. In some embodiments, different offsets may be used to bring different filtered images having various intensity ranges into a similar predetermined range as all the filtered images.
[0286] In some embodiments, the processed image includes a first plurality of processed images and a second plurality of processed images, and the filtered image includes a first plurality of filtered images and a second plurality of filtered images. In some embodiments, the first plurality of processed images and the first plurality of filtered images are from one or more reference cycles and different channels. In some embodiments, the second plurality of flow cell images, the second plurality of processed images, and the second plurality of filtered images are from one or more reference cycles and different channels. In some embodiments, the second plurality of flow cell images, the second plurality of processed images, and the second plurality of filtered images are from one or more cycles different from one or more reference cycles and different channels. In some embodiments, the first plurality of flow cell images and the second plurality of flow cell images are the same, the first plurality of processed images and the second plurality of processed images are the same, and the first plurality of filtered images and the second plurality of filtered images are the same.
[0287] Method 6000 may include the operation 6350 of obtaining a 3D community map. The 3D community map may be determined in one or more reference cycles. During a cycle different from the reference cycle, the 3D community map may have been pre-generated in the reference cycle and may be obtained by actively requesting or retrieving or passively receiving the 3D community map. In some embodiments, a single 3D community map is used during the sequencing analysis of the same 3D sample such that in a cycle different from the reference cycle, there is no need to generate a new 3D community map.
[0288] Method 6000 may include the operation of generating a 3D community map. The operation of generating a 3D community map may but is not limited to occur during one or more reference cycles. As a non-limiting example, cycles 1-5 or 2-5 may be reference cycles. In some embodiments, the 3D community map generated in one or more reference cycles may be used for any cycle different from the reference cycle such that the additional computational complexity, cost, and time of generating a 3D community map in each cycle may be avoided.
[0289] The 3D colony map can be generated based on one or more 2D template images. The 2D template images are equivalently used as 2D colony maps herein. Each 2D template image or colony map corresponds to a flow cell image at a specific z position. Each 2D template image corresponds to a flow cell image at a specific z position and at a specific tile or sub-tile. Each 2D template image corresponds to a specific sequencing cycle. Each template image can correspond to one or more channels. For example, if there are 10 different z positions for a 3D sample to be sequenced, during cycle N, there are ten 2D template images corresponding to different z positions for a single sub-tile.
[0290] The colony maps (2D or 3D) herein can be saved as a list of coordinates. Each entry in the list of coordinates can correspond to a colony, such as the center of the colony. Instead of saving a 2D or 3D matrix as a colony map, the list of coordinates can be stored with much less storage space and can be utilized more efficiently in calculations.
[0291] In some embodiments, the operation of generating a 3D colony map can include obtaining 2D template images. Obtaining the template images can include generating the template images or receiving or retrieving the template images. The methods and operations described herein can be used to generate 2D template images. The 2D template images can be generated after filtering, such as by a top-hat filter or a DoG filter. The 2D template images can be generated after filtering the flow cell images and registering the filtered images with a reference coordinate system. The 2D template images can be generated in one or more reference cycles, and the same template images can be used across different cycles and channels. The 2D template images can be a list of coordinates of colonies. For example, each entry in the list can be the 2D or 3D coordinates of a colony.
[0292] In some embodiments, the operation of generating a 3D colony map can include combining one or more 2D template images into a candidate 3D colony map. For example, the lists of coordinates in the 2D template images can be added together.
[0293] In some embodiments, when the 3D colony map is saved as a 3D matrix instead of a list of entries, the operation of generating a 3D colony map can include extracting colonies in one or more template images. Instead of directly combining the 2D template images, colonies in one or more template images can be extracted, and the extracted colonies can be included in the candidate 3D colony map based on their coordinates in the template images. The candidate 3D colony map can be initialized with 0 or other minimum intensity values in all its pixels at different z positions. The extracted intensities can be used to replace the initialization values in the candidate 3D colony map to indicate pixels or voxels that are at least part of a colony. As a non-limiting example, the 3D colony map can be a 3D matrix of 0s and 1s, where each pixel with a 1 indicates that the pixel is part of a colony.
[0294] A single colony may appear at one, two, or even more z-positions such that the same colony can be included in multiple flow cell images at different z-positions. For example, a colony at (x1,y1) at position z1 can be included again at (x1-1,y1-1) at position z2. Thus, a candidate 3D colony map may include duplicate colonies. Duplicate colonies need to be removed for accurate and reliable 3D base identification. The operation 6350 of generating a 3D colony map may include removing duplicate colonies from the candidate 3D colony map.
[0295] To remove duplicate colonies, preliminary base identification may be performed. The position of the colony for base identification can be determined by a 2D template image, while the intensity for performing base identification can be extracted from the filtered image. Filtering herein can advantageously remove intensity interference from out-of-focus colonies and background objects such that the intensity can be used for more accurate and reliable base identification than the unfiltered flow cell image. In some embodiments, the 2D template image contains the coordinates of the colonies in a reference coordinate system such that even if the colonies may shift between cycles, base identification can still be attributed to the same colony.
[0296] After obtaining the preliminary base identification, the operation 6350 may include a repeatability operation of removing duplicate colonies until a stop criterion is met. The repeatability operation may include identifying candidate colonies with the same base identification. In response to identifying candidate colonies with the same base identification, the candidate colonies may contain zero, one, two, or more duplicate colonies. The operation 6350 may further include an operation of determining the 3D distance between each pair of colonies among the candidate colonies. For each pair of colonies (non-repetitive pairs), the 3D distance can be calculated based on the coordinates of the colonies. The coordinates can be 3D. The coordinates may include the 2D coordinates of the colonies (e.g., after registration in the reference coordinate system) and the z-position of the colonies. The 3D distance can be in pixels. The 3D distance can be in other units (e.g., um). The 3D distance can be used to determine whether two colonies with the same base identification are close to each other.
[0297] In response to determining that the 3D distance between two clusters is within a predetermined distance threshold, operation 6350 may include determining the image intensity of each of the two clusters from the filtered image or the template image. In some embodiments, determining the image intensity of each cluster may include normalizing and / or offsetting the image intensities at different z positions to a predetermined range. For example, the normalization and / or offset may be based on the intensity of fiducial markers at different z positions. Subsequently, operation 635 may include removing one of the two clusters having a smaller image intensity. Two clusters within the predetermined distance threshold may be considered as repeated identical clusters. The repeat with the smaller intensity may be more out of focus than the repeat with the larger intensity and may be removed to ensure accurate and reliable base identification. The predetermined distance threshold may be customized based on characteristics of the sample, imaging parameters, clusters, etc. For example, the predetermined distance threshold may be based on the depth of field of the optical system, the distance between two adjacent flow cell images along the axial distance, or a combination thereof. The distance threshold may be adjusted individually or in combination with a stop criterion to balance true and false repeats being removed. For example, cluster p1 has a preliminary base identification of A determined using its filtered image intensity, and its position after registration with the reference coordinate system may be at (xp1, yp1). Candidate clusters may be determined to be those having the same base identification as A. The 3D distance from each candidate cluster to cluster p1 may be calculated and compared with a predetermined threshold (e.g., a predetermined threshold of 0.5 um). A cluster p2 that meets the distance threshold may be considered a possible repeat of p1. The intensities of p1 and p2 are compared, and the cluster p2 with the greater intensity is retained while the coordinates of cluster p1 are removed as a repeat.
[0298] After removing the repeated clusters, a 3D cluster map is generated that includes all the clusters and the corresponding coordinates at which base identification can be performed. The 3D cluster map may include position information of such clusters. For example, the 3D cluster map may include the coordinates of each cluster in the corresponding filtered image. The 3D cluster map may further include the z position of each cluster. The 3D cluster map may include the size and / or shape of the clusters. The 3D cluster map may include a unique identifier for each cluster. The 3D cluster map may include the image intensity of the clusters. Such image intensity may be the filtered intensity obtained from the filtered images disclosed herein.
[0299] The stopping criteria can be customized. For example, the stopping criteria can be based on different types of samples, imaging parameters, the size and shape of the community, etc. The stopping criteria, the distance threshold, or both can be adjusted to balance true and false duplicates being removed. As a non-limiting example, the stopping criteria can be to remove the first 1000 duplicate communities. As a non-limiting example, the stopping criteria can be a selected time window. As yet another example, the stopping criteria can be the absence of duplicates that meet a predetermined distance threshold.
[0300] In some embodiments, method 6000 includes an operation of registering flow cell images, processed images, and / or filtered images. In some embodiments, the images are registered across channels and different cycles. In some embodiments, the images are registered before any base calling is performed. In some embodiments, the images are registered across channels and different cycles before generating a template image or obtaining a 3D community map. In some embodiments, the images are registered across channels and different cycles before generating a filtered image or a processed image.
[0301] Various image registration techniques can be used to register the images. 2D registration techniques can be used to register the images. When registering a processed image or a filtered image, they can be treated as flow cell images during registration. In some embodiments, image registration method 9000 as disclosed herein can be used to register the images by treating each image to be registered as a flow cell image that can be obtained using sequencing system 1100, e.g., across different channels and / or different cycles. In some embodiments, the images can be registered after performing one or more of the preprocessing operations disclosed herein. In some embodiments, the operation of registering flow cell images, processed images, and / or filtered images can occur before the operation 6350 of obtaining a 3D community map.
[0302] In some embodiments, the operation of registering flow cell images, processed images, and / or filtered images is relative to a reference coordinate system. In some embodiments, the operation of registering flow cell images, processed images, and / or filtered images is relative to one or more template images. The operation of registering the images can include generating one or more template images in the reference coordinate system. In some embodiments, the operation of registering the images can include registering the communities with template communities in one or more template images. The operation of registering the images can include determining a plurality of transforms based on one or more template images. Each transform in the plurality of transforms can correspond to a corresponding sub-block of the flow cell image, processed image, or filtered image, and is configured to register the sub-block with one or more template images. Each transform can be used to register the corresponding sub-block or tile with one or more template images. The plurality of transforms can include one or more affine transforms.
[0303] In some embodiments, the operation of registering images may include performing image registration of the colony based on fiducial markers. The fiducial markers may be located on the flow cell. Alternatively, the fiducial markers may be external to the flow cell.
[0304] In some embodiments, image registration, which is the main analysis step herein, is configured to align images from different cycles and / or different channels, for example, relative to a template image or a reference coordinate system. In some embodiments, image registration, which is the main analysis step herein, is configured to register colonies or clusters from different cycles and different channels (e.g., in the filtered images) with a template image or a reference coordinate system.
[0305] For example, after registering the filtered images from different channels relative to the corresponding template images disclosed herein, base calling may be performed using the filtered images from different channels in cycle N.
[0306] Method 6000 may include an operation 6450 of extracting colony intensity based on a 3D colony map. For each colony in the 3D colony map, position information of such colony, such as the 2D coordinates and the z position of the colony, may be obtained from the 3D colony map. Using the 2D coordinates and the z position, the corresponding filtered image and its pixels may be determined. The image intensity of such pixels may be extracted from the corresponding filtered image as the intensity of such pixels for performing base calling.
[0307] Method 6000 may include an operation 6550 of performing 3D base calling using the extracted image intensity. Various existing 2D base calling algorithms may be used. The base calling results may be saved together with their 3D position information (e.g., coordinates). Such 3D coordinates may be used to register base calling across different cycles and at different z positions, and / or to register base calling with the pool images herein.
[0308] In some embodiments, the operation 6550 of performing base calling includes performing the main analysis step herein to adjust the image intensity of the colonies in the filtered image; and performing base calling on the colonies based on the adjusted image intensity in the filtered image. The adjustment of the image intensity in the filtered image may be performed before or after filtering.
[0309] In some embodiments, the main analysis step includes one or more of the following: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and pre-phasing correction; quality score estimation.
[0310] In some embodiments, method 6000 includes an operation of obtaining a second MIP image based on a plurality of flow cell images. The second MIP can be a flattened 2D image of an axial stack of flow cell images. In some embodiments, the flow cell images are raw images acquired by the sequencing system 1100. The operation of obtaining the second MIP can be different from the operation of obtaining the first MIP image. In a flow cell image (e.g., a raw image), out-of-focus colonies can have a larger full width at half maximum (FWHM), and thus the signal of out-of-focus colonies is more dispersed than that of in-focus colonies or clusters. In some embodiments, the larger FWHM can result in a white ring around the colony. Figure 5 An exemplary second MIP image generated directly from a flow cell image is shown, and there is a ring or halo around the colony in the center of the image. The ring or halo may be an image artifact that can cause errors in base calling. The second MIP may also retain some background information (e.g., background information from unwanted background objects), which may interfere with base calling if the second MIP is used for base calling. Thus, the second MIP includes artifacts and / or unwanted background objects that require additional processing before accurate and reliable base calling based on the intensities in the second MIP.
[0311] In some embodiments, the second MIP can be used for registration of filtered images and / or colonies or clusters because it shares the same FOV, resolution, image size, etc. If the second MIP is used for base calling, the artifacts and / or unwanted background objects in the second MIP may interfere with base calling, but the same artifacts and / or unwanted background objects (which are not in the filtered images) can facilitate the registration of the filtered images with the stained flow cell images.
[0312] In some embodiments, instead of the second MIP, the flow cell images can be directly used to register the filtered images with the flow cell images. The flow cell images also contain noise and unwanted background objects for the purpose of base calling, but can facilitate the registration of the flow cell images and the filtered images with the flow cell images.
[0313] Pool Image
[0314] In some embodiments, the methods herein advantageously utilize the second MIP for registering the flow cell images with the flow cell images. Out-of-focus colonies and background objects that may interfere with correct base calling can be used to provide information for registering the flow cell images with the flow cell images.
[0315] Method 6000 may further include an operation 6500 of performing image registration of the flow cell image based on the second MIP image. In some embodiments, image registration of the flow cell image with one or more pool images (e.g., stained images) is supplementary to the image registration that is part of the primary analysis, e.g., image registration across cycles and / or channels. Image registration of the flow cell image is configured to align the colonies or clusters relative to the cellular structure such that base calling can be assigned to the cell nucleus, membrane, or other regions of the cell. In some embodiments, registering the flow cell image based on the second MIP image includes registering or aligning the background objects in the second MIP with the corresponding objects in one or more pool images. For example, membrane information can be obtained from the second MIP image and aligned with the membrane in the pool image.
[0316] In some embodiments, an operation 6600 of performing image registration of multiple flow cell images based on the second MIP image includes registering the second MIP image with a template image or a reference coordinate system. In some embodiments, registering the second MIP image with a template image or a reference coordinate system may rely on the image registration information of the first MIP because the second MIP and the first MIP are taken from the same FOV.
[0317] In some embodiments, method 6000 may include an operation 6700 of registering the first MIP image and / or the 3D base calling based on the first MIP image and the second MIP image with one or more pool images. The background objects in the second MIP can be used to align the second MIP and the first MIP with the pool images. The pool images may also have background objects that are the same as the background objects in the second MIP but are transformed. The transformation may be represented by a single transformation of the entire image or may be divided into multiple transformations, each representing a part of the entire image. After finding the transformation of the background objects between the second MIP and the pool images, the colonies and clusters can be registered with the pool images.
[0318] In some embodiments, the operation 6700 of registering the MIP image and / or the 3D base calling with the pool images may not rely on the second MIP image directly obtained from the flow cell image but only on the first MIP image of the filtered image. In these embodiments, the background information can be obtained from one or more of the following: the first MIP image, the opened image / processed image, the filtered image, and the flow cell image. The background objects can be used to align the first MIP with the pool images by using one or more transformations. The transformation may be represented by a single transformation of the entire image or may be divided into multiple transformations, each representing a part of the entire image. After finding the transformation of the background objects between the first MIP and the pool images, the colonies and clusters can be registered with the pool images.
[0319] In some embodiments, the operation 6700 of registering the MIP image and / or 3D base identification with the tile image may not rely on the second MIP image directly obtained from the flow cell image, but rather on fiducial markers. Signals from the same fiducial markers may be present in one or more of the following: the first MIP image, the open image / processed image, the filtered image, and the flow cell image. Such fiducial markers may also be included in the tile image. Aligning the fiducial markers may generate a transformation between the sequencing images, e.g., a transformation between the first MIP image and the tile image. The transformation may be used to register or align colonies or clusters between the sequencing image and the tile image. Figures 7A - 7B An exemplary registered MIP image superimposed on the corresponding tile image is shown. The MIP image includes bright spots representing colonies or clusters. Some of the bright spots overlap with stained cell nuclei, while some other bright spots appear within the cell membrane but outside the cell nucleus. Figure 7B Segmentation of individual cells is shown such that colonies or clusters can be grouped relative to each individual cell.
[0320] In some embodiments, one or more tile images are images of cells and / or tissues having one or more stains (e.g., fluorescent stains). In some embodiments, one or more images may include staining of cell structures, which helps to localize colonies or clusters relative to the stained structures. For example, the stain may be a stain of cell structures or components including but not limited to membranes, cell nuclei, and mitochondria.
[0321] In some embodiments, the cell membrane may be permeabilized after sequencing analysis and imaging using the sequencing system and reaction. In some embodiments, one or more tile images may include staining of lipids such as those contained in the cell membrane. In some embodiments, instead of labeling the lipids, one or more tile images may include staining of one or more transmembrane proteins. The transmembrane proteins may be proteins embedded in the permeabilized membrane.
[0322] In some embodiments, one or more tile images include fluorescence signals from the cell membrane. One or more tile images may be microscopic images. One or more images may be fluorescence images. In some embodiments, different fluorescence colors may be included in the tile image. For example, the cell nucleus and the cell membrane may be stained with different colors.
[0323] In some embodiments, one or more images may include segmentation of cells, membranes, cell nuclei, or combinations thereof. Figure 7BAn exemplary tiled image with individual cells segmented is shown. In some embodiments, the edges of each segment encompass the entire membrane of the cells within the segment. Only one cell may be present in each segment. Some segments may not have any cells. In some embodiments, adjacent segments do not overlap with each other. In some embodiments, adjacent segments overlap with each other only by sharing one or more edges. In some embodiments, various segmentation algorithms may be used to segment the cells.
[0324] In some embodiments, the tiled images disclosed herein are stained. Staining may be performed after the sequencing image is acquired using the sequencing system 110. In some embodiments, staining may be performed before the sequencing image is acquired. Methods for staining 3D samples such as cells, tissues may include one or more operations disclosed herein. Staining of 3D samples may use various methods that can specifically label one or more cellular proteins that are predominantly located in the membrane but have negligible presence (e.g., amount or concentration less than 10%, 5%, 2%) in other regions of the cell.
[0325] The operation may include selecting one or more primary antibodies, each of the one or more primary antibodies specifically binding to a corresponding protein. The corresponding protein may be a transmembrane protein of one or more cells. In some embodiments, the corresponding transmembrane protein does not exist at a predetermined concentration in other cell regions such as the cytoplasm or nucleus, such that staining of the transmembrane protein does not produce a perceptible signal in cell regions other than the membrane. In some embodiments, one or more different transmembrane proteins may be labeled with primary antibodies. For example, if there are 5 different types of transmembrane proteins, 5 different primary antibodies may be used, and each primary antibody specifically binds to one of the transmembrane proteins but not to the other transmembrane proteins. In some embodiments, the same type of primary antibody may non-specifically bind to different proteins.
[0326] The staining method may include the operation of selecting one or more secondary antibodies that bind to one or more primary antibodies. The staining method may further include the operation of labeling one or more secondary antibodies with a fluorescent label.
[0327] In some embodiments, the staining method may further include the operation of using a support element or a tertiary probe (such as a hydrogel) to link the secondary antibody and the fluorescent label. The support element can be used to retain the mRNA of the membrane for facilitating binding and the generation of fluorescent signals. In some embodiments, the mRNA can be any mRNA in the cell. In some embodiments, the staining method may include using various methods to clear the tissue to remove some or all parts of the cells to reduce the background fluorescence from parts of the cell that are not the membrane. In some embodiments, the fluorescent label includes a fluorophore that re-emits light within a specific wavelength range after photoexcitation. FIG. 8 shows an exemplary staining of a transmembrane protein using the staining method disclosed herein.
[0328] The staining method may further include the operation of generating one or more pool images of the corresponding protein. The one or more images contain the fluorescent signals emitted from the fluorescent label.
[0329] In some embodiments, the method for base calling in sequencing data analysis may include generating a 3D colony map based on a plurality of filtered images. The 3D colony map may include a stack of 2D colony maps, each 2D colony map corresponding to a flow cell image acquired at a corresponding axial position. The 3D colony map may include some or all of the colonies that can be identified in the axial stack of the flow cell images.
[0330] In some embodiments, each 2D colony map may be generated based on the filtered images at the same axial position. In some embodiments, the 2D colony map may be generated from the filtered images, similar to generating a template image from the flow cell image. In some embodiments, the 2D colony map is equivalent to the template image because both of them are virtual images that include all the colonies identified in one or more cycles at a specific axial position.
[0331] As disclosed herein, the filtered images can advantageously exclude background objects that may interfere with the signals from the colonies. The filtered images can also remove some out-of-focus colonies or clusters. The filtering parameters (such as the size and shape of the kernel) can be customized to balance between removing out-of-focus colonies and retaining relatively large colonies or clusters of assembled out-of-focus colonies.
[0332] In some embodiments, each 2D colony map is registered with a template image (3D) or a reference coordinate system (in 3D). The template image or the reference coordinate system can be determined in a reference cycle. For each cycle different from the reference cycle, a 2D colony map can be generated, and the colony maps from different cycles can be registered relative to each other using the template image or the reference coordinate system. In some embodiments, a single colony map can be generated for all channels having the same cycle. In some embodiments, a colony map is generated for each channel for each cycle.
[0333] In some embodiments, the 3D community map is a volumetric community map that stacks all the 2D community maps at different axial positions. In some embodiments, the methods herein include removing duplicates of the community maps from the stacked 2D community maps to generate a 3D community map. The 3D community map without duplicates can be used as a reference to locate individual communities in a sample. In some embodiments, the 3D community map can be saved as a list of 3D coordinates indicating the centers of the communities.
[0334] In some embodiments, the methods herein further include an operation of extracting the image intensity of a community based on the 3D community map. The image intensity can be extracted from one or more of the following: flow cell images; processed images; filtered images. In some embodiments, the image intensity can be extracted from the filtered images. In some embodiments, the image intensity can be extracted from the filtered images after processing the filtered images using one or more of the main analysis steps disclosed herein. For example, phasing and pre-phasing correction can be performed on the filtered images before the image intensity can be extracted.
[0335] In some embodiments, the 3D community map can include duplicates of the communities, and such duplicates can be removed after base calling. For example, all the communities in the 3D community map, including duplicates, can be used to extract image intensity for base calling. Candidate duplicates of a community can be identified as communities at different z positions (e.g., adjacent z positions) and the same x,y position. If the base calls of such candidate duplicates are the same, one of them can be removed as a duplicate. Alternatively, both of them can be removed, and a new community map representing both of them can be added at the z position as the average of both, and located at the same x and y positions.
[0336] In some embodiments, the methods herein further include an operation of performing 3D base calling based on the extracted community image intensity. Various 2D base calling algorithms can be used here. For example, base calling of a community can be performed by comparing the image intensities of the same community from different channels, and identifying the base corresponding to the maximum image intensity in all channels.
[0337] In some embodiments, the operation of filtering the flow cell images to generate a plurality of filtered images may include performing deconvolution on the plurality of flow cell images. The deconvolution may be at least along the axial direction. In some embodiments, the deconvolution may be 3D. The deconvolution may perform its equivalent operation in the spatial domain or in a transformed domain (such as the Fourier domain). The deconvolution is configured to reduce or eliminate the diffusion or blurring effect of the optical system on the community, such that the size and shape of the community appear more accurate in the flow cell images. In some embodiments, the deconvolution operation may be used alone as a filtering operation or in combination with other filtering operations (such as top-hat filtering).
[0338] Image Registration with the Pool Image
[0339] Various methods may be used to register the flow cell images based on fiducial markers. The fiducial markers may be located inside or outside the sample. For example, internal fiducial markers may include at least some communities or clusters or background objects in the sample. As another example, external fiducial markers may be microspheres coated on the flow cell such that the signals from the microspheres can similarly be used as internal fiducial markers for registration. The same fiducial markers may appear in sequencing images such as MIP images, flow cell images, filtered images, and tile images, such that the transformation can be derived by aligning the fiducial markers in different images. Exemplary embodiments of image registration methods are described in PCT patent application number PCT / US2023 / 067931 (wherein the content of this patent is hereby incorporated by reference in its entirety).
[0340] For example, a community or other object (such as a background object) with image intensity I centered at position (x1, y1) in a sequencing image may appear at position (x2, y2) with intensity I' in a tile image, where and Mr is the transformation matrix. Similarly, the inverse transformation matrix Mr -1 can be determined such that The registration of the images may be in 2D and may include translation, scaling, rotation, and / or shearing of the flow cell images across different channels. Multiple points in the sequencing image and their corresponding points in the tile image may be used to determine the transformation. The minimum number of points required may be determined by the degrees of freedom of the transformation. In some embodiments, the image registration may be 3D, where the coordinates are on the x, y, and z axes.
[0341] In some embodiments, the sequencing image may be divided into a plurality of sub-tiles, and a transformation may be determined for each sub-tile to represent the transformation of the entire image. In some embodiments, the image transformation of each sub-tile may be uniquely represented by a transformation matrix. The transformation matrix may be determined as follows:
[0342]
[0343] where n is the number of sub-patches, a1 = x1 + dx1, b1 = y1 + dy1, a2 = x2 + dx2, b2 = y2 + dy2, ... an = xn + dxn, bn = yn + dyn, d1... dn are 2D shifts corresponding to the sub-patches, and where dxn and dyn are the shift components of the 2D shift dn on the x-axis and y-axis respectively, and where M is the 3x3 transformation matrix of the sub-patch.
[0344] In some embodiments, the transformation matrix can be defined as the inverse matrix of M, i.e., M -1 , such that Equation (1) can be alternatively expressed as
[0345]
[0346] In some embodiments, the transformation matrix M is an estimate in Equations (1) and (3) based on the 2D shift. In some embodiments, the value of n may affect the accuracy of the estimate.
[0347] In some embodiments, more than one region can be selected within the sub-patch for cross-correlation calculation, and more than one 2D displacement can be calculated for each sub-patch and used to estimate the transformation of the sub-patch. In these embodiments, n in Equation (1) can be replaced with a larger number. For example, when 2 regions are selected for each sub-patch, it can be replaced with 2*n, and the transformation matrix M can be estimated using Equations (1) and (2).
[0348] In some embodiments, (a1,b1)...(an,bn) in Equations (1)–(3) are the coordinates of the selected regions after transformation (e.g., the coordinates of the central pixel of the corresponding region), and (x1,y1)...(xn,yn) are the coordinates of the selected regions before transformation, e.g., the coordinates of the central pixel.
[0349] In some embodiments, n is a number not less than 3. The larger n is, the more information is used to estimate the transformation matrix M. In some embodiments, n is not greater than 9.
[0350] In some embodiments, the transformation of one or more sub-patches is linear. In some embodiments, the transformation of all sub-patches is linear. In some embodiments, the transformation matrix is a matrix where M31 and M32 are equal to 0 and M33 is 1. In some embodiments, one or more of the transformations in the transformation of each sub-patch are affine transformations, and the transformation matrix of the entire flow cell image is an affine matrix.
[0351] In some embodiments, the transformation matrix M is an estimate in equations (1) and (3) based on the size of the selected region. In some embodiments, the size of the selected region may affect the accuracy of the estimate. In some embodiments, the size of the selected region may be about 128x128. In some embodiments, the size of the selected region may be about 32x32, 48x48, 64x64, 96x96, 160x160, 196x196, 256x256, or various different sizes. As disclosed herein, the transformation of each sub-tile can be calculated using the selected region within the sub-tile, and the selected region can be equal to or smaller than the sub-tile. In either case, considering the inherent characteristics of the image transformation across sequencing cycles, the transformation estimated using the region can be used to estimate the transformation of the entire sub-tile. The image transformation between cycles and / or between adjacent pixels can be relatively small, e.g., less than about 8%, 5%, or less than about 1% of scaling, rotation, and / or shear. In some embodiments, the transformation disclosed herein can include an image translation with a difference between cycles and / or between adjacent pixels greater than about 5%.
[0352] After determining the multiple transformations of the individual sub-tiles, the transformation of the entire flow cell image can be accurately and reliably estimated by transforming the individual sub-tiles using the multiple transformations and combining the transformed sub-tiles into a transformed tiled image. The techniques disclosed herein advantageously estimate the transformation of the flow cell image by determining the multiple transformations of the individual sub-tiles of the flow cell image. The multiple transformations can be linear, and even if the transformations are non-linear, the transformation of the flow cell image can be accurately and reliably estimated. The techniques disclosed herein advantageously eliminate the need to calculate the transformation of the entire flow cell image, which can be computationally more intensive, more time-consuming, and more prone to failure compared to estimating the multiple transformations of the sub-tiles.
[0353] Image Registration across Channels and Cycles
[0354] In some embodiments, the method includes the operation of aligning or registering the flow cell images across different sequencing cycles, from different channels, and / or at different z-levels with a common coordinate system before base calling. The common coordinate system can be the reference coordinate system disclosed herein. The common coordinate system can be predetermined. The common coordinate system can be the reference coordinate system disclosed herein. The common coordinate system can be predetermined. The common coordinate system can be a Cartesian coordinate system. Various other coordinate systems can also be used. Other coordinate systems can include, but are not limited to, polar coordinate systems, cylindrical coordinate systems, or spherical coordinate systems.
[0355] Exemplary embodiments of the image registration method are described in PCT patent application number PCT / US2023 / 067931 (the content of which is hereby incorporated by reference in its entirety).
[0356] Prior to registration with the tile image, the flow cell images can be registered relative to each other so that communities or clusters in different cycles and / or channels can be aligned and base calling can be accurate and reliable for a particular community or cluster.
[0357] Various methods can be used to register sequencing images of different cycles and / or channels, such as flow cell images, filtered images, or MIP images.
[0358] In some embodiments, method 6000 includes an operation of registering MIP images, such as Figure 9 9000 in. In some embodiments, prior to performing any base calling, MIP images are registered across channels and different cycles. 2D registration techniques can be used to register MIP images, for example, by treating the MIP image as a flow cell image acquired from the sequencing system 110. In some embodiments, MIP images can be registered using image registration method 9000 as disclosed herein by treating each MIP image as a flow cell image, for example, across different channels and / or different cycles. In some embodiments, MIP images can be registered after performing one or more preprocessing operations disclosed herein.
[0359] For example, the sequencing images can be registered to a reference coordinate system common to all flow cell images so that sequencing images from different cycles and / or channels can be aligned with each other. The reference coordinate system can be determined in a reference cycle or any other predetermined cycle. For example, the reference coordinate system can be the coordinate system of a flow cell image from one channel. As another example, the reference coordinate system can be based on an external fiducial marker or other object external to the flow cell image.
[0360] In some embodiments, method 6000 can include an operation of generating one or more template images in the reference coordinate system by registering the community with one or more template images using the coordinates of the community. Figure 8A A schematic diagram of one or more template images generated in the reference coordinate system in the reference cycle is shown. In some embodiments, template image 2100 includes dimensions that are substantially the same as a single tile 2900 that includes a 5×5 grid of sub-tiles 2200. A region 2300 is selected in each sub-tile, and the region includes the center pixel of the corresponding sub-tile. In this embodiment, the reference coordinate system has an origin 2120 located at the top left pixel thereof. In some embodiments, the template images disclosed herein can be individual regions, such as region 2300. Each template image can include a plurality of communities 2320 therein.
[0361] In some embodiments, the size of the template image can be substantially the same as the size of the flow cell image such that from Figures 8A - 8BThe different tiles 2900 therein and all the colonies from multiple channels can be registered with the same template image. However, such a template image may contain colonies that are not used in at least some of the operations described herein, to reduce the computational burden without sacrificing accuracy.
[0362] In some embodiments, more than one template image can be generated, and each template image 2300 corresponds to at least a portion of a sub-tile of the flow cell image from a channel.
[0363] The template images herein can be initialized as virtual images that have a black or dark background and no signal from colonies. For example, the template image can be initialized to zero, or include a minimum image intensity at all pixels.
[0364] In operation 9100, after determining the coordinates of colonies by image registration of flow cell images across different channels, the intensity of the colonies can be added to the template image at the positions determined by the coordinates, where the size and shape are determined based on the registration. The template image can be a virtual image that combines the image intensities of colonies obtained from 2, 3, 4, or more channels in a reference cycle. The pixels of the template that do not contain colonies remain black or dark, so that the template image can have a cleaner background without the noise present in the actual flow cell images.
[0365] In some embodiments, method 9000 includes an operation of obtaining the image intensity, size, shape, or a combination thereof of colonies from at least a portion of one or more sub-tiles in a reference cycle, so that such information can be used for polygons included in the template image. In some embodiments, the colonies can have a fixed shape and / or size. In some embodiments, the point spread function determined by the optical system is used to determine the fixed shape and / or size of the colonies. In some embodiments, the colonies have a fixed spot size based on the σ of the Gaussian point spread function. In some embodiments, the size of one or more colonies is 1-9 pixels. In some embodiments, the size of one or more colonies is 1-3 pixels.
[0366] The template image can include colonies from different channels and channel information. As an example, the channel information can be provided in the form of a label, or in a specific order of how the colonies are included.
[0367] In some embodiments having multiple template images, each template image 2300 can cover an area within the sub-tile, and such template images can but do not need to include all the colonies within the sub-tile.
[0368] In some embodiments, method 9000 includes operation 9300 of obtaining a flow cell image in a cycle after a reference cycle. Operation 9300 may include passively receiving from or actively requesting from the optical system a flow cell image after the optical system generates the flow cell image as disclosed herein. The optical system may include an imager 1160 in Figure 1 the
[0369] The flow cell image may include some or all of the same colonies in the template image of the reference cycle. Specifically, the flow cell image may include some or all of the same colonies in the region corresponding to the selected region in the reference cycle.
[0370] Figure 8A A flow cell image obtained in a cycle different from the reference cycle is shown at the bottom. The flow cell image 2400 is obtained with a plurality of sub - tiles 2500. The position of the selected region 2600 in this cycle relative to the new origin 2420 in this cycle may be the same as the position of the selected region 2300 in the reference cycle relative to the origin 2120. In this cycle, the flow cell image 2100 in the reference may have been transformed into a transformed image 2110, and the selected region 2300 is correspondingly transformed into region 2310 and has some overlap with region 2600. The image transformation herein may be 2D and may include translation, scaling, rotation, and / or shear.
[0371] In some embodiments, method 9000 is configured to align the template image 2100 or 2300 in the reference cycle and the transformed image 2110 or 2310 in another cycle with a reference coordinate system.
[0372] In some embodiments, instead of directly using region 2310 or 2110 in image registration, method 9000 may include operations of selecting regions 2300 and 2600 for more simple and convenient determination of image registration. Region 2600 may include at least a portion of the colony 2320 in the template image 2300 as colony 2330.
[0373] In some embodiments, method 9000 includes operation 9400 of determining a plurality of transformations of the flow cell image 2400 based on one or more template images 2100 or 2300. As Figure 8A shown, each transformation in the plurality of transformations may correspond to a sub - tile 2500 of the flow cell image 2400 and is configured to register the sub - tile 2500 of the flow cell 2400 image with the corresponding portion of the template image 2100 (if the template image includes the entire tile) or the corresponding template image 2300 (if there are multiple template images within the tile).
[0374] In some embodiments, operation 6400 may include determining each transformation corresponding to a sub - tile of the flow cell image. More specifically, each transformation may correspond to a selected region of each sub - tile among some or all of the sub - tiles. Regions may be selected from the sub - tiles in various ways to include at least a portion of the sub - tile. The region may be a predefined two - dimensional shape, e.g., rectangular, circular, or square. As a non - limiting example, the selected region may include one or more central pixels of the sub - tile, as shown at 2600 in Figure 8A For example, selecting a 64x64 region may be computationally simpler than selecting a 128x128 region, but may be less accurate. In some embodiments, the selected region includes some or all of the colonies 2320 registered in the template image in the reference cycle so that the same colonies and their relative positions in the template image and the flow cell image can be used to determine the transformation. In some embodiments, the sizes of the template image (e.g., 2300) and the region 2600 may be the same or approximately the same. In some embodiments, the sizes of the template image 2100 or 2200 and the selected region 2600 may be different.
[0375] In some embodiments, the cross - correlation of the selected region and the template image may be computed to determine the 2D displacement of the region relative to the template image. Figure 10A The reference image (left) is shown, which is transformed into a different image (middle) by 2D shear, scale, and rotation. The 2D displacements 6010 at the four corners of the reference image can be determined, for example, using the method of cross - correlation disclosed herein. And the 2D displacements at the four corners can be used to estimate the transformation between the two images.
[0376] In some embodiments, cross-correlation can be calculated in the spatial domain. In some embodiments, cross-correlation can be calculated in the spatial frequency domain after Fourier transform (FT). Method 9000 can include generating a corresponding Fourier-transformed image (FTI) of the template image and the Fourier transform of the selected region. The Fourier transform herein can be calculated using discrete FT (DFT), fast FT (FFT), etc. Cross-correlation can be determined based on the FTI of the selected region and the Fourier transform. As a non-limiting example, cross-correlation can be the element-wise multiplication of the FTI and the FT of the selected region, followed by adding the complex conjugate or rotation of one of them. Then, the inverse FT of the element-wise multiplication can be obtained. In some embodiments, cross-correlation can be a 2D image having a peak intensity at its coordinates [xp, yp]. In some embodiments, the 2D displacement can be determined based on the comparison of the coordinates [xp, yp] with the coordinates of the peak obtained from the cross-correlation of the two original images that are never transformed. The 2D displacement of the selected region 2600 can be used to estimate the 2D displacement of the entire sub-block. In some embodiments, the results of calculating cross-correlation in the spatial domain or the Fourier domain can be equivalent. In some embodiments, the calculation in the Fourier domain is simpler and more efficient than that in the spatial domain.
[0377] In some embodiments, the image transformation of a sub-block can be determined according to the 2D displacements from some or all of the adjacent sub-blocks, regardless of whether there is a 2D displacement in itself. In some embodiments, the 2D displacements from all immediate neighbors can be used. For example, to determine the transformation of sub-block 2530, the 2D displacements from 3 adjacent sub-blocks and its own 2D displacement can be used. For sub-block 2510, a total of 6 2D displacements can be used, including the directly adjacent sub-blocks and its own. For sub-block 2520, a total of 9 2D displacements can be used, including the adjacent sub-blocks and its own. In some embodiments, the 2D displacements from some but not all of the adjacent sub-blocks can be used. In some embodiments, the 2D displacements from all adjacent sub-blocks except for 1-2 outliers can be used to determine the transformation. Outliers can be excluded using a predetermined criterion, for example, differing from other 2D displacements by more than 30% or 50%.
[0378] Figure 10Bis an image of the 2D displacements within a tile that displays a flow cell image. In this embodiment, the tile has a 6x9 grid of sub-tiles, and each sub-tile has a 2D displacement 6010 determined using the techniques disclosed herein. The magnitude of each displacement along the x or y axis is less than about 5 pixels. The pixel size can vary according to the imaging parameters, and exemplary pixels can be from 0.01um to 0.9um. The 2D displacements 6010 can be used to calculate the transformation of the tile by separately calculating the transformation of each sub-tile, e.g., an affine matrix. In this embodiment, the affine matrix can be calculated using the methods disclosed herein.
[0379] In some embodiments, sub-pixel resolution (e.g., about 0.01, 0.02, 0.03, or 0.05 pixels) of the 2D shift 6010 can be achieved using various methods including interpolation, upsampling, etc. In some embodiments, sub-pixel resolution can be achieved by fitting peaks using a selected filter (e.g., a 3x3 or 5x5 Gaussian filter).
[0380] In some embodiments, the image transformation of a sub-tile can be uniquely represented by a transformation matrix. The transformation matrix can be determined as follows:
[0381]
[0382] where n is the number of sub-tiles, a1 = x1 + dx1, b1 = y1 + dy1, a2 = x2 + dx2, b2 = y2 + dy2,... an = xn + dxn, bn = yn + dyn, d1... dn are the 2D shifts corresponding to the sub-tiles, and where dxn and dyn are the shift components of the 2D shift dn along the x and y axes respectively, and where M is the 3x3 transformation matrix of the sub-tile.
[0383] In some embodiments, the transformation matrix can be defined as the inverse matrix of M, i.e., M -1 , such that equation (1) can be alternatively represented as
[0384]
[0385] In some embodiments, the transformation matrix M is an estimate based on the 2D shifts in equations (4) and (6). In some embodiments, the value of n may affect the accuracy of the estimate. In some embodiments, more than one region can be selected within a sub-tile for cross-correlation calculation, and more than one 2D displacement can be calculated for each sub-tile and used to estimate the transformation of the sub-tile. In these embodiments, n in equation (1) can be replaced with a larger number, e.g., when 2 regions are selected for each sub-tile, it can be replaced with 2*n, and the transformation matrix M can be estimated using equations (4) and (5).
[0386] In some embodiments, n is a number not less than 3. The larger n is, the more information is used to estimate the transformation matrix M. In some embodiments, n is not greater than 9.
[0387] In some embodiments, the transformation of one or more sub-tiles is linear. In some embodiments, the transformation of all sub-tiles is linear. In some embodiments, the transformation matrix is a matrix where M31 and M32 are equal to 0 and M33 is 1. In some embodiments, one or more of the transformations in the transformation of each sub-tile are affine transformations, and the transformation matrix is an affine matrix.
[0388] In some embodiments, the transformation matrix M is an estimate in equations (4) and (6) based on the size of the selected region. In some embodiments, the size of the selected region may affect the accuracy of the estimate. In some embodiments, the size of the selected region may be about 128x128. In some embodiments, the size of the selected region may be about 32x32, 48x48, 64x64, 96x96, 160x160, 196x196, or 256x256. As disclosed herein, the transformation of each sub-tile can be calculated using the selected region within the sub-tile, and the selected region can be equal to or smaller than the sub-tile. In either case, considering the inherent characteristics of the image transformation across sequencing cycles, the transformation estimated using the region can be used to estimate the transformation of the entire sub-tile. The image transformation between cycles and / or between adjacent pixels can be relatively small, e.g., less than about 5% of scaling, rotation, and / or shear or less than about 1%. In some embodiments, the transformation disclosed herein can include an image translation with a difference between cycles and / or between adjacent pixels greater than about 5%.
[0389] After determining the multiple transformations of the individual sub-tiles, the transformation of the flow cell image can be accurately and reliably estimated through the multiple transformations. The techniques disclosed herein advantageously estimate the transformation of the flow cell image by determining the multiple transformations of the individual sub-tiles of the flow cell image. The multiple transformations can be linear, and even if the transformation is non-linear, the transformation of the flow cell image can be accurately and reliably estimated. The techniques disclosed herein advantageously eliminate the need to calculate the transformation of the entire flow cell image, which may be computationally more intensive and time-consuming compared to estimating the multiple transformations of the sub-tiles.
[0390] In some embodiments, the computer-implemented method 9000 further includes an operation of saving the multiple transformations by the processor disclosed herein. In some embodiments, the computer-implemented method 9000 further includes an operation of transmitting the multiple transformations to a processing unit (such as a CPU) for subsequent operations.
[0391] In some embodiments, the computer-implemented method 9000 further includes registering the sub-tile images with one or more template images using a plurality of transforms. This operation may be performed by a processing unit such as a CPU. In any given cycle different from the reference cycle, each sub-tile image may be registered or transformed to one or more template images by multiplying the sub-tile image by a transform matrix corresponding to the sub-tile image.
[0392] In some embodiments, the computer-implemented method 9000 may include operating on performing one or more preprocessing steps on the flow cell image of the cycle before registering the images from the reference cycle and / or other cycles.
[0393] In some embodiments, the operation of performing one or more preprocessing steps may be performed by an FPGA. In some embodiments, the data after the operation may be transferred by the FPGA to the CPU such that the CPU may use such data to perform subsequent operations in methods 6000 and 9000.
[0394] In some embodiments, one or more preprocessing steps of the flow cell image in the reference cycle may be performed before operation 9100 or 9200 or after operation 9200. In some embodiments, one or more preprocessing steps of the flow cell image in the reference cycle may be performed after the operation of receiving the flow cell image in the reference cycle from the optical system disclosed herein. In some embodiments, one or more preprocessing steps of the flow cell image in the reference cycle may be performed before the operation of obtaining the image intensity, size, shape, or a combination thereof of the colonies from a plurality of sub-tile images of the flow cell image in the reference cycle.
[0395] In some embodiments, one or more preprocessing steps of the flow cell image in a cycle other than the reference cycle may be performed after operation 9300 or 9400. In some embodiments, one or more preprocessing steps of the flow cell image in a cycle other than the reference cycle may be performed after the operation of registering the sub-tile images of the flow cell image to one or more template images. In some embodiments, one or more preprocessing steps of the flow cell image in a cycle other than the reference cycle may be performed before the operation of extracting the image intensity of a plurality of colonies from the sub-tile images of the flow cell image. In some embodiments, one or more preprocessing steps of the flow cell image in a cycle other than the reference cycle may be performed before the operation of performing base calling using the image intensity of the sub-tile images of the flow cell image.
[0396] One or more preprocessing steps may include background subtraction. The background subtraction is configured to remove at least some of the background signals that may interfere with the signal of interest, i.e., the image intensity of the community. The background signals may be noise caused by multiple sources, including the flow cell 1120, the imager 1160, the sequencer 1140, and other sources. The background subtraction may be adjusted to avoid over-subtraction.
[0397] One or more preprocessing steps may include image sharpening to optimize the image intensity of the community in view of the surroundings of the community in the flow cell image. For example, a Laplacian of Gaussian (LoG) filter may be used for sharpening.
[0398] One or more preprocessing steps may include image registration so that the image intensities of the communities are registered relative to each other. For example, the image intensities may be registered with the templates described herein.
[0399] One or more preprocessing steps may include intensity offset adjustment, which may remove intensity offsets not removed during background subtraction.
[0400] One or more preprocessing steps may include color correction for removing interference from other channels or colors to one channel.
[0401] One or more preprocessing steps may include lag and lead correction, which is configured to correct the image intensity within a specific cycle by eliminating intensity deviations caused by sequencing of DNA fragments that are out of sync with other fragments due to lag or lead.
[0402] One or more preprocessing steps may include intensity normalization to normalize the image intensities of the communities from different channels within a predetermined range.
[0403] One or more preprocessing steps may include background subtraction; image sharpening; or a combination thereof.
[0404] In some embodiments, the computer-implemented method 9000 further includes extracting the image intensities of multiple communities from sub-tiles registered with the template image. This operation may be performed by a processing unit such as a CPU or an FPGA. In some embodiments, the communities and their corresponding intensities are extracted from the flow cell image into a different data format that is easier and more efficient to process. For example, each community may have 4 different intensities, each from a different channel. Such intensities may be extracted into a list, with each entry in the list corresponding to a community. The list may be generated after image registration to reflect the position information of the same community in different cycles. Thus, the image intensities of the same community in different cycles may be located in different lists each corresponding to one cycle.
[0405] In some embodiments, the computer-implemented method 9000 further includes performing base calling using the image intensity of sub-tiles of the flow cell image after registration, such that base calling can be performed accurately across different channels and in different cycles with respect to the same community.
[0406] In some embodiments, method 9000 includes an operation 9400 of determining a plurality of transforms of the flow cell image. Operation 9400 may include determining each transform in the transform without using any adjacent sub-tiles disclosed herein. Instead, more than 2 regions may be selected within the sub-tile, and the 2D displacement of each region in the region may be determined. The 2D displacement obtained from the regions within the same sub-tile may be used to determine the transform of the sub-tile using equations (1) and (2). The regions within the sub-tile may be smaller in size than the regions 2600 in the adjacent sub-tiles. For example, the region 2600 may be about 128x128, and the regions within the sub-tile may be 3, 4, 5 or even more regions, and each region includes a matrix of about 64x64. Other operations of method 9000 may remain the same for image registration using or not using adjacent sub-tiles when generating the transform.
[0407] As used herein, the term "dark" refers to an image intensity or signal intensity below a predetermined threshold such that base calling cannot be performed with a predetermined quality. The predetermined threshold may be customized according to different samples, sequencing parameters, and / or other factors that may affect the signal intensity. In other words, base calling performed at "dark" intensity may be incorrect and unreliable. Alternatively, base calling cannot be performed at "dark" intensity. The image intensity or signal intensity is in the flow cell image. A "dark" flow cell image herein refers to a flow cell image in which substantially all pixels or voxels have an intensity below a predetermined intensity threshold. A "dark" base herein refers to a nucleotide base attached to a "dark" fluorescent dye, and the flow cell image of the channel corresponding to the "dark" base may be a "dark" flow cell image.
[0408] As used herein, the term "bright" refers to an image intensity or signal intensity above a predetermined threshold such that base calling cannot be performed with a predetermined quality. The predetermined threshold may be customized to various numbers or ranges of absolute image intensity, percentage of image intensity, and / or relative image intensity. The predetermined thresholds for determining "bright" and "dark" may not be the same. For example, the lowest 5% of the image intensity in the flow cell image may be "dark", and the highest 15% of the image intensity in the flow cell image may be "bright".
[0409] Computer System
[0410] Each embodiment of methods 6000, 9000, 5200 can be implemented, for example, using one or more computer systems (such as the computer system 4000 shown in Figure 4 ). For example, one or more computer systems 4000 can be used to implement any of the various embodiments described herein, as well as their combinations and sub - combinations.
[0411] The computer system 4000 can include one or more hardware processors 404. The hardware processor 404 can be a central processing unit (CPU), a graphics processing unit (GPU), or a combination thereof. The processor 404 can be connected to a bus or communication infrastructure 406.
[0412] The computer system 4000 can also include user input / output devices 403, such as a display, a keyboard, a pointing device, etc., and the computer system can communicate with the communication infrastructure 406 through a user input / output interface 402. The user input / output devices 403 can be coupled to Figure 1 the user interface 1240 in
[0413] One or more of the processors 404 can be a graphics processing unit (GPU). In one embodiment, the GPU can be a processor that is a dedicated electronic circuit designed to process math - intensive applications. The GPU can have a parallel structure that is effective for parallel processing of large blocks of data, such as the common math - intensive data found in computer graphics applications, images, videos, vector processing, array processing, etc., as well as cryptography (including brute - force cracking), generating cryptographic hashes or hash sequences, solving partial hash inversion problems, and / or generating the results of other proof - of - work calculations for some blockchain - based applications, for example. By virtue of the general - purpose computing on graphics processing units (GPGPU) capabilities, the GPU can be particularly useful at least in the image recognition and machine learning described herein.
[0414] Additionally, one or more of the processors 404 can include a coprocessor or other logical implementations for accelerating cryptographic computations or other specialized mathematical functions, including a hardware - accelerated cryptographic coprocessor. Such an acceleration processor can further include an instruction set for using the coprocessor and / or other logic to facilitate such acceleration.
[0415] The computer system 4000 can also include a data storage device, such as a main memory 408, e.g., a random access memory (RAM). The main memory 408 can include one or more levels of cache. The main memory 408 can store control logic (i.e., computer software) and / or data therein.
[0416] The computer system 4000 may also include one or more auxiliary data storage devices or auxiliary memories 410. The auxiliary memory 410 may include, for example, a main storage drive 412 and / or a removable storage device or drive 414. The main storage drive 412 may be, for example, a hard disk drive or a solid state drive. The removable storage drive 414 may be a floppy disk drive, a tape drive, an optical disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.
[0417] The removable storage drive 414 may interact with a removable storage unit 418.
[0418] The removable storage unit 418 may include a computer-usable or readable storage device having computer software and / or data stored thereon. The software may include control logic. The software may include instructions executable by the hardware processor 404. The removable storage unit 418 may be a floppy disk, a tape, an optical disk, a DVD, an optical storage disk, and / any other computer data storage device. The removable storage drive 414 may read from and / or write to the removable storage unit 418.
[0419] The auxiliary memory 410 may include other components, devices, assemblies, tools, or other means for allowing the computer system 4000 to access computer programs and / or other instructions and / or data. Such components, devices, assemblies, tools, or other means may include, for example, a removable storage unit 422 and an interface 420. Examples of the removable storage unit 422 and the interface 420 may include a program cartridge and a cartridge interface (such as the interface found in a video game device), a removable memory chip (such as an EPROM or a PROM) and an associated socket interface, a memory stick and a USB port, a memory card and an associated memory card slot, and / or any other removable storage unit and associated interface.
[0420] The computer system 4000 may also include a communication or network interface 424. The communication interface 424 may enable the computer system 4000 to communicate and interact with any combination of external devices, external networks, external entities, etc. (collectively and individually referred to by the reference numeral 428). For example, the communication interface 424 may allow the computer system 4000 to communicate with an external or remote device 428 via a communication path 426, which may be wired and / or wireless (or a combination thereof) and may include any combination of a LAN, a WAN, a network, etc. Control logic and / or data may be transmitted to and from the computer system 4000 via the communication path 426. In some embodiments, the communication path 426 is a connection to the cloud 130, as Figure 1as depicted. The external device or the like indicated by reference numeral 428 may be a device, network, entity, etc. in the cloud 1300.
[0421] The computer system 4000 may also be any one of the following or any combination thereof: a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet computer, a smartphone, a smartwatch or other wearable device, an appliance, a part of the Internet of Things (IoT), and / or an embedded system, to name just a few non-limiting examples.
[0422] It should be understood that the framework described herein may be implemented as a method, process, device, system, or article of manufacture, such as a non-transitory computer-readable medium or device. For illustrative purposes, the present framework may be described in the context of a distributed ledger that is publicly available or at least accessible to untrusted third parties. An example of a modern use case is a blockchain-based system. However, it should be understood that the present framework may also be applied to other settings where sensitive or confidential information may need to pass through the hands of untrusted third parties, and this technology is in no way limited to distributed ledger or blockchain use.
[0423] The computer system 4000 may be a client or a server that accesses or hosts any application and / or data through any delivery mode, including but not limited to remote or distributed cloud computing solutions; local or on-premises software (e.g., a “local” cloud-based solution); a “service” model (e.g., content as a service (CaaS), digital content as a service (DCaaS), software as a service (SaaS), managed software as a service (MSaaS), platform as a service (PaaS), desktop as a service (DaaS), framework as a service (FaaS), backend as a service (BaaS), mobile backend as a service (MBaaS), infrastructure as a service (IaaS), database as a service (DBaaS), etc.); and / or a hybrid mode that includes any combination of the foregoing examples or other services or delivery paradigms.
[0424] Any applicable data structure, file format, and schema may be derived from standards including but not limited to: JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible HyperText Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other individually or combinatorially functionally similar representation. Alternatively, proprietary data structures, formats, or schemas may be used alone or in combination with known or open standards.
[0425] Any relevant data, documents, and / or databases can be stored, retrieved, accessed, and / or transmitted in a human-readable format (such as numeric, text, graphical, or multimedia formats, further including various types of markup languages and other possible formats). Alternatively or in combination with the above formats, the data, documents, and / or databases can be stored, retrieved, accessed, and / or transmitted in binary, encoded, compressed, and / or encrypted formats or any other machine-readable format.
[0426] The interface connections or interconnections between various systems and layers can employ any number of mechanisms, such as any number of protocols, programming frameworks, layout plans, or application programming interfaces (APIs), including but not limited to the Document Object Model (DOM), Discovery Service (DS), NSUserDefaults, Web Services Description Language (WSDL), Message Exchange Pattern (MEP), Web Distributed Data Exchange (WDDX), Web Hypertext Application Technology Working Group (WHATWG) HTML5 Web Messaging, Representational State Transfer (REST or RESTful web services), eXtensible User Interface Protocol (XUP), Simple Object Access Protocol (SOAP), XML Schema Definition (XSD), XML Remote Procedure Call (XML-RPC), or any other open or proprietary mechanism that can achieve similar functions and results.
[0427] Such interface connections or interconnections can also utilize Uniform Resource Identifiers (URIs), which can further include Uniform Resource Locators (URLs) or Uniform Resource Names (URNs). Other forms of uniform and / or unique identifiers, locators, or names can be used, either alone or in combination with those forms such as the above.
[0428] Any one of the above protocols or APIs can interface with or be implemented in any programming language (procedural, functional, or object-oriented) and can be compiled or interpreted. Non-limiting examples include C, C++, C#, Objective-C, Java, Scala, Clojure, Elixir, Swift, Go, Perl, PHP, Python, Ruby, JavaScript, WebAssembly, or almost any other language, as well as any other library or pattern in any type of framework, runtime environment, virtual machine, interpreter, stack, engine, or similar mechanism, including but not limited to Node.js, V8, Knockout, jQuery, Dojo, Dijit, OpenUI5, AngularJS, Expressjs, Backbone.js, Ember.js, DHTMLX, Vue, React, Electron, etc., and many other non-limiting examples.
[0429] In some embodiments, a tangible non-transitory device or article of manufacture that includes a tangible non-transitory computer-usable or readable medium having control logic (software) stored thereon may also be referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 4000, main memory 408, secondary memory 410, and removable storage units 418 and 422, and tangible articles embodying any combination of the foregoing. When executed by one or more data processing devices, such as computer system 4000, such control logic may cause such data processing devices to operate as described herein.
[0430] Based on the teachings contained in this disclosure, it will be apparent to those skilled in the relevant art how to make and use the various embodiments of this disclosure using data processing devices, computer systems, and / or computer architectures different from those Figure 4 shown. Specifically, the embodiments may operate with software, hardware, and / or operating system implementations other than those described herein.
[0431] Optical System
[0432] Figure 1 The imager 1160 in [ ] may include one or more optical systems. Further disclosed herein are optical system design guidelines and high-performance fluorescence imaging methods and systems that provide improved optical resolution and image quality for fluorescence imaging-based genomics applications. The disclosed optical imaging system design provides a larger field of view, increased spatial resolution, improved modulation transfer, contrast-to-noise ratio, and image quality, higher spatial sampling frequency, faster transition between image captures when repositioning the sample plane to capture a series of images (e.g., images of different fields of view), and improved imaging system duty cycle, and thus enables higher throughput image acquisition and analysis.
[0433] In some cases, improvements in imaging performance, such as for dual-sided (flow cell) imaging applications, can be achieved by using an electro-optic phase plate in combination with an objective lens to compensate for optical aberrations caused by fluid layers separating the upper (near) and lower (far) inner surfaces of the flow cell. In some cases, this design approach can also compensate for vibrations introduced by, for example, a motion-actuated compensator that moves into or out of the optical path depending on which surface of the flow cell is being imaged.
[0434] In some cases, improvements in imaging performance, such as for dual-sided (flow cell) imaging applications, include using thick flow cell walls (e.g., wall (or coverslip) thickness > 700 μm) and fluid channels (e.g., fluid channel height or thickness of 50 - 200 μm), and can be achieved by using a tube lens design that corrects for optical aberrations caused by the thick flow cell wall and / or intervening fluid layer in combination with the objective lens, even when using off-the-shelf commercially available objective lenses.
[0435] In some cases, improvements in imaging performance, such as for multi-channel (e.g., two-color or four-color) imaging applications, can be achieved by using multiple tube lenses (one tube lens per imaging channel), where each tube lens design has been optimized for the specific wavelength range used in this imaging channel.
[0436] Exemplary embodiments disclosed herein may include a fluorescence imaging system, the system comprising: a) at least one light source configured to provide excitation light within one or more specified wavelength ranges; b) an objective lens configured to collect fluorescence generated from a specific field of view in a sample plane when the sample plane is exposed to the excitation light, where the numerical aperture of the objective lens is at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, or at least 0.9, or a numerical aperture value falling within the range defined by any two of the above; where the working distance of the objective lens is at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, at least 1000 μm, or a working distance falling within the range defined by any two of the above; and where the field of view has an area of at least 0.1 mm 2 、at least 0.2 mm 2 、at least 0.5 mm 2 、at least 0.7 mm 2 、at least 1 mm 2 、at least 2 mm 2 、at least 3 mm 2 、at least 5 mm 2 or at least 10 mm 2 or the field of view falls within the range defined by any two of the above; and c) at least one image sensor, where the fluorescence collected by the objective lens is imaged onto the image sensor, and where the pixel size of the image sensor is selected such that the spatial sampling frequency of the fluorescence imaging system is at least twice the optical resolution of the fluorescence imaging system.
[0437] In some embodiments, the numerical aperture can be at least 0.75. In some embodiments, the numerical aperture is at least 1.0. In some embodiments, the working distance is at least 850 μm. In some embodiments, the working distance is at least 1,000 μm. In some embodiments, the field of view can have an area of at least 2.5 mm 2 In some embodiments, the field of view can have an area of at least 3 mm 2 In some embodiments, the spatial sampling frequency can be at least 2.5 times the optical resolution of the fluorescence imaging system. In some embodiments, the spatial sampling frequency can be at least 3 times the optical resolution of the fluorescence imaging system. In some embodiments, the system can further include an X-Y-Z translation stage such that the system is configured to acquire a series of two or more fluorescence images in an automated manner, where each image in the series is or can be acquired for a different field of view. In some embodiments, the positioning of the sample plane can be adjusted simultaneously in the X, Y, and Z directions to match the positioning of the objective focal plane between acquiring images of different fields of view. In some embodiments, the time required for simultaneous adjustment in the X, Y, and Z directions can be less than 0.3 seconds, less than 0.4 seconds, less than 0.5 seconds, less than 0.7 seconds, or less than 1 second or a time falling within a range defined by any two of the foregoing. In some embodiments, the system further includes an autofocus mechanism configured to adjust the focal plane positioning before acquiring images of different fields of view if an error signal indicates that the positioning difference between the focal plane and the sample plane in the Z direction is greater than a specified error threshold. In some embodiments, the specified error threshold is 100 nm or greater. In some embodiments, the specified error threshold is 50 nm or less. In some embodiments, the system includes three or more image sensors, and wherein the system is configured to image fluorescence in each of three or more wavelength ranges onto different image sensors. In some embodiments, the positioning difference between the focal plane of each of the three or more image sensors and the sample plane is less than 100 nm. In some embodiments, the positioning difference between the focal plane of each of the three or more image sensors and the sample plane is less than 50 nm. In some embodiments, the total time required to reposition the sample plane, adjust the focal length (if necessary), and acquire an image is less than 0.4 seconds per field of view. In some embodiments, the total time required to reposition the sample plane, adjust the focal length (if necessary), and acquire an image is less than 0.3 seconds per field of view.
[0438] The present disclosure also discloses a fluorescence imaging system for performing dual-sided imaging of a flow cell. The fluorescence imaging system includes: a) an objective lens configured to collect fluorescence generated within a specified field of view of a sample plane within the flow cell; b) at least one tube lens positioned between the objective lens and at least one image sensor, wherein the at least one tube lens is configured to correct an imaging performance metric of a combination of the objective lens, the at least two tube lenses, and the at least one image sensor when imaging the inner surface of the flow cell, and wherein the wall thickness of the flow cell is at least 700 μm and the gap between the upper inner surface and the lower inner surface is at least 50 μm; wherein the imaging performance metric is substantially the same for imaging the upper inner surface or the lower inner surface of the flow cell without moving an optical compensator into or out of an optical path between the flow cell and the at least one image sensor, without moving one or more optical elements of a barrel lens along the optical path, and without moving one or more optical elements of the barrel lens into or out of the optical path.
[0439] In some embodiments, the objective lens can be a commercially available microscope objective lens. In some embodiments, the numerical aperture of the commercially available microscope objective lens can be at least 0.3. In some embodiments, the working distance of the objective lens can be at least 700 μm. In some embodiments, the objective lens can be corrected to compensate for the thickness of a cover glass (or the thickness of the flow cell wall) of 0.17 mm or greater than or less than 0.17 mm. In some embodiments, the optical system can be corrected to compensate for the distance between the cover glass thickness, the flow cell thickness, or the desired focal plane. In some embodiments, the correction can be performed by inserting a correction optical device such as a lens or an optical assembly into the optical path of the optical system. In some embodiments, the correction can be performed without inserting a correction optical device such as a lens or an optical assembly into the optical path of the optical system. In some embodiments, the fluorescence imaging system can further include an electro-optic phase plate, which is positioned adjacent to the objective lens and between the objective lens and the tube lens, wherein the electro-optic phase plate can provide correction for optical aberrations caused by a fluid filling the gap between the upper inner surface and the lower inner surface of the flow cell. In some embodiments, at least one tube lens can be a compound lens including three or more optical components. In some embodiments, at least one tube lens is a compound lens including four optical components, and the four optical components can include one or more of the following: a first asymmetric convex-convex lens, a second plano-convex lens, a third asymmetric concave-concave lens, and a fourth asymmetric convex-concave lens, which can be present in the order listed above or in any alternative order. In some embodiments, at least one tube lens is configured to correct the imaging performance metric of the combination of the objective lens, at least one tube lens, and at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 1 mm. In some embodiments, at least one tube lens is configured to correct the imaging performance metric of the combination of the objective lens, at least one tube lens, and at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 100 μm. In some embodiments, at least one tube lens is configured to correct the imaging performance metric of the combination of the objective lens, at least one tube lens, and at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 200 μm. In some embodiments, the system includes a single objective lens, two tube lenses, and two image sensors, and each of the two tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths. In some embodiments, the system includes a single objective lens, three tube lenses, and three image sensors, and each of the three tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths. In some embodiments, the system includes a single objective lens, four tube lenses, and four image sensors, and each of the four tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths.In some embodiments, the design of the objective lens or at least one tube lens is configured to optimize the modulation transfer function in the medium to high spatial frequency range. In some embodiments, the imaging performance metric includes measurements of the modulation transfer function (MTF), defocus, spherical aberration, chromatic aberration, coma, astigmatism, field curvature, image distortion, contrast to noise ratio (CNR), or any combination thereof, at one or more specified spatial frequencies. In some embodiments, the difference in the imaging performance metric for imaging the upper inner surface and the lower inner surface of the flow cell is less than 10%. In some embodiments, the difference in the imaging performance metric for imaging the upper inner surface and the lower inner surface of the flow cell is less than 5%. In some embodiments, compared to a conventional system including an objective lens, a motion actuation compensator, and an image sensor, using at least one tube lens provides at least equivalent or better improvement in the imaging performance metric for bilateral imaging. In some embodiments, compared to a conventional system including an objective lens, a motion actuation compensator, and an image sensor, using at least one tube lens provides at least a 10% improvement in the imaging performance metric for bilateral imaging.
[0440] Disclosed herein is an illumination system for imaging-based solid-phase genotyping and sequencing applications, the illumination system including: a) a light source; and b) a liquid light guide configured to collect light emitted by the light source and convey it to a specified illumination field on a support surface including tethered biological macromolecules.
[0441] In some embodiments, the illumination system further includes a condenser lens. In some embodiments, the specified illumination field has an area of at least 2 mm 2 . In some embodiments, the light delivered to the specified illumination field has uniform intensity across the specified field of view of the imaging system used to acquire an image of the support surface. In some embodiments, the specified field of view has an area of at least 2 mm 2 . In some embodiments, the light delivered to the specified illumination field has uniform intensity across the specified field of view when the coefficient of variation (CV) of the light intensity is less than 10%. In some embodiments, the light delivered to the specified illumination field has uniform intensity across the specified field of view when the coefficient of variation (CV) of the light intensity is less than 5%. In some embodiments, the speckle contrast value of the light delivered to the specified illumination field is less than 0.1. In some embodiments, the speckle contrast value of the light delivered to the specified illumination field is less than 0.05.
[0442] Imaging Module and System
[0443] Those skilled in the art will understand that, in some cases, the disclosed optical systems, imaging systems or modules can be stand-alone optical systems designed to image a sample or substrate surface. In some cases, they can include one or more processors or computers. In some cases, they can include one or more software packages that provide instrument control functions and / or image processing functions. In some cases, in addition to optical components such as light sources (e.g., solid-state lasers, dye lasers, diode lasers, arc lamps, tungsten halogen lamps, etc.), lenses, prisms, mirrors, dichroic reflectors, optical filters, optical band-pass filters, apertures and image sensors (e.g., complementary metal oxide semiconductor (CMOS) image sensors and cameras, charge-coupled device (CCD) image sensors and cameras, etc.), they can also include mechanical and / or optomechanical components such as X-Y translation stages, X-Y-Z translation stages, piezoelectric focusing mechanisms, etc. In some cases, they can serve as modules, components, sub-assemblies or subsystems of a larger system designed for genomics applications (e.g., gene testing and / or nucleic acid sequencing applications). For example, in some cases, they can serve as modules, components, sub-assemblies or subsystems of a larger system that further includes a light-tight and / or other environmental control housing, a temperature control module, a fluid control module, a fluid dispensing robot, a pick-and-place robot, one or more processors or computers, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, image processing software packages, data analysis software packages), a data storage module, a data communication module (e.g., Bluetooth, WiFi, intranet or Internet communication hardware and associated software), a display module or any combination thereof.
[0444] Method for Sequencing
[0445] The present disclosure provides methods for sequencing immobilized template molecules or non-immobilized template molecules. The methods can be operated in system 1000, e.g., in sequencer 1140. In some embodiments, the immobilized template molecules comprise a plurality of nucleic acid template molecules having one copy of a target sequence of interest. In some embodiments, the nucleic acid template molecules having one copy of the target sequence of interest can be generated by bridge amplification using linear library molecules. In some embodiments, the immobilized template molecules comprise a plurality of nucleic acid template molecules, each nucleic acid template molecule having two or more tandem copies (e.g., concatemers) of the target sequence of interest. In some embodiments, the nucleic acid template molecules comprising concatemer molecules can be generated by performing rolling circle amplification of circularized linear library molecules. In some embodiments, the non-immobilized template molecules comprise circular molecules. In some embodiments, the methods for sequencing employ soluble (e.g., non-immobilized) sequencing polymerases or sequencing polymerases immobilized to a support.
[0446] In some embodiments, the sequencing reaction employs detectably labeled nucleotide analogs. In some embodiments, the sequencing reaction employs a two-stage sequencing reaction that includes binding a detectably labeled multivalent molecule and incorporating nucleotide analogs. In some embodiments, the sequencing reaction employs unlabeled nucleotide analogs. In some embodiments, the sequencing reaction employs phosphate-linked nucleotides.
[0447] In some embodiments, the immobilized concatemers each include tandem repeat units (e.g., insertion regions) of the sequence of interest and any linker sequences. For example, the tandem repeat device includes: (i) a left universal linker sequence having a binding sequence for a first surface primer (1121) (e.g., a surface-immobilized primer), (ii) a left universal linker sequence having a binding sequence for a first sequencing primer (1141) (e.g., a forward sequencing primer), (iii) a target sequence (1111), (iv) a right universal linker sequence having a binding sequence for a second sequencing primer (1151) (e.g., a reverse sequencing primer), (v) a right universal linker sequence having a binding sequence for a second surface primer (1131) (e.g., a surface capture primer), and (vii) a left sample index sequence (1161) and / or a right sample index sequence (1171). In some embodiments, the tandem repeat unit further includes a left unique identifier sequence (1181) and / or a right unique identifier sequence (1191). In some embodiments, the tandem repeat unit further includes a binding sequence for at least one compaction oligonucleotide. In some embodiments, FIGS. 11 and 12 show the units of linear library molecules or concatemer molecules.
[0448] The immobilized multimer can collapse on its own into a compact nucleic acid nanosphere. The inclusion of one or more compaction oligonucleotides during the RCA reaction can further compact the size and / or shape of the nanosphere. An increase in the number of tandem repeat units in a given concatemer increases the number of sites along the concatemer available for hybridization with multiple sequencing primers (e.g., sequencing primers having universal sequences), which serve as multiple starting sites for polymerase-catalyzed sequencing reactions. When the sequencing reaction employs detectably labeled nucleotides and / or detectably labeled multivalent molecules (e.g., having nucleotide units), the signals emitted by the nucleotides or nucleotide units participating in parallel sequencing reactions along the concatemer result in increased signal intensity for each concatemer. Multiple portions of a given concatemer can be sequenced simultaneously. In addition, multiple binding complexes can form along a particular concatemer molecule, each binding complex including a sequencing polymerase that binds to a template / primer duplex and to a multivalent molecule, where the multiple binding complexes remain stable and do not dissociate, resulting in increased dwell time, which increases signal intensity and reduces imaging time.
[0449] Method for Sequencing Using Nucleotide Analogs
[0450] The present disclosure provides methods for sequencing any immobilized template molecule described herein, the methods comprising step (a): contacting a sequencing polymerase with (i) a nucleic acid template molecule and (ii) a nucleic acid sequencing primer, wherein the contacting is carried out under conditions suitable for binding the sequencing polymerase to the nucleic acid template molecule hybridized to the nucleic acid primer, wherein the nucleic acid template molecule hybridized to the nucleic acid primer forms a nucleic acid duplex. In some embodiments, the sequencing polymerase comprises a recombinant mutant sequencing polymerase that can bind and incorporate nucleotide analogs.
[0451] In some embodiments, in methods for sequencing concatemer molecules, the sequencing primer comprises a 3'-extensible end or a 3'-nonextensible end. In some embodiments, the plurality of nucleic acid template molecules comprises amplified template molecules (e.g., template molecules amplified in a clonal manner). In some embodiments, the plurality of nucleic acid template molecules comprises one copy of a target sequence of interest. In some embodiments, the plurality of nucleic acid molecules comprises two or more tandem copies (e.g., concatemers) of a target sequence of interest. In some embodiments, the plurality of nucleic acid template molecules comprises the same target sequence of interest or different target sequences of interest. In some embodiments, the plurality of nucleic acid primers are in solution or immobilized to a support. In some embodiments, when the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are immobilized to a support, binding to a first sequencing polymerase produces a plurality of immobilized first complex polymerases. In some embodiments, the plurality of nucleic acid template molecules and / or nucleic acid primers are immobilized to 10 2 –10 15 different sites on the support. In some embodiments, the binding of the plurality of concatemer molecules and nucleic acid primers to the plurality of first sequencing polymerases generates a plurality of first complex polymerases immobilized at 10 2 –10 15 different sites on the support. In some embodiments, the plurality of immobilized first complex polymerases on the support are immobilized at predetermined or random sites on the support. In some embodiments, the plurality of immobilized first complex polymerases are in fluid communication with each other to allow a reagent solution (e.g., an enzyme including a sequencing polymerase, a multivalent molecule, nucleotides, and / or divalent cations) to flow onto the support, such that the plurality of immobilized complex polymerases on the support react with the reagent solution in a massively parallel manner.
[0452] In some embodiments, the method for sequencing further comprises step (b): contacting a sequencing polymerase with a plurality of nucleotides under conditions suitable for binding at least one nucleotide to the sequencing polymerase bound to the nucleic acid duplex and suitable for incorporating polymerase-catalyzed nucleotides, the incorporation of the polymerase-catalyzed nucleotides causing the sequencing primer to be extended by one nucleotide. In some embodiments, the sequencing polymerase is contacted with the plurality of nucleotides in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises at least one nucleotide analogue having a chain terminating moiety at the sugar 2' or 3' position. In some embodiments, the chain terminating moiety can be removed from the sugar 2' or 3' position to convert the chain terminating moiety into an OH or H group. In some embodiments, the plurality of nucleotides comprises at least one nucleotide lacking a chain terminating moiety. In some embodiments, at least one nucleotide is labeled with a detectable reporter moiety (e.g., a fluorophore) that emits a detectable signal. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is linked to the nucleobase. In some embodiments, the fluorophore is linked to the nucleobase with a linker that can be cleaved / removed from the base. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, a particular detectable reporter moiety (e.g., a fluorophore) linked to a nucleotide can correspond to a nucleobase (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to permit detection and identification of the nucleobase. When the incorporated chain-terminating nucleotide is detectably labeled, step (b) further comprises detecting the signal emitted from the incorporated chain-terminating nucleotide. In some embodiments, step (b) further comprises identifying the nucleobase of the incorporated chain-terminating nucleotide.
[0453] In some embodiments, the method for sequencing further comprises step (c): removing the chain terminating moiety from the incorporated chain-terminating nucleotide to generate an extendable 3' OH group. In some embodiments, step (c) further comprises removing the detectable label from the incorporated chain-terminating nucleotide. In some embodiments, the sequencing polymerase remains bound to the template molecule that hybridizes to the sequencing primer extended by one nucleobase.
[0454] In some embodiments, the method for sequencing further comprises step (d): repeating steps (b) and (c) at least once.
[0455] Two - Stage Method for Sequencing Nucleic Acids
[0456] The present disclosure provides a two-stage method for sequencing any of the immobilized template molecules described herein. In some embodiments, the first stage generally comprises binding a multivalent molecule to a complex polymerase to form a multivalent-complex polymerase, and detecting the multivalent-complex polymerase.
[0457] In some embodiments, the first part includes step (a): contacting a plurality of first sequencing polymerases with (i) a plurality of nucleic acid template molecules and (ii) a plurality of nucleic acid sequencing primers, wherein the contacting is carried out under conditions suitable for binding the plurality of first sequencing polymerases to the plurality of nucleic acid template molecules and the plurality of nucleic acid primers, thereby forming a plurality of first complex polymerases, each first complex polymerase comprising a first sequencing polymerase bound to a nucleic acid duplex, wherein the nucleic acid duplex comprises a nucleic acid template molecule hybridized to a nucleic acid primer. In some embodiments, the first polymerase comprises a recombinant mutant sequencing polymerase.
[0458] In some embodiments, in a method for sequencing a template molecule, the sequencing primer comprises an oligonucleotide having a 3'-extensible end or a 3'-nonextensible end. In some embodiments, the plurality of nucleic acid template molecules include amplified template molecules (e.g., template molecules amplified in a clonal manner). In some embodiments, the plurality of nucleic acid template molecules include one copy of a target sequence of interest. In some embodiments, the plurality of nucleic acid molecules include two or more tandem copies (e.g., concatemers) of a target sequence of interest. In some embodiments, the nucleic acid template molecules among the plurality of nucleic acid template molecules include the same target sequence of interest or different target sequences of interest. In some embodiments, the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are in solution or immobilized on a support. In some embodiments, when the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are immobilized on a support, binding to the first sequencing polymerase generates a plurality of immobilized first complex polymerases. In some embodiments, the plurality of nucleic acid template molecules and / or nucleic acid primers are immobilized at 10 2 –10 15 different sites on the support. In some embodiments, the binding of the plurality of concatemer molecules and nucleic acid primers to the plurality of first sequencing polymerases generates a plurality of first complex polymerases immobilized at 10 2 –10 15 different sites on the support. In some embodiments, the plurality of immobilized first complex polymerases on the support are immobilized at predetermined or random sites on the support. In some embodiments, the plurality of immobilized first complex polymerases are in fluid communication with each other to allow a reagent solution (e.g., an enzyme including a sequencing polymerase, a multivalent molecule, a nucleotide, and / or a divalent cation) to flow onto the support, such that the plurality of immobilized complex polymerases on the support react with the reagent solution in a massively parallel manner.
[0459] In some embodiments, the method for sequencing further includes step (b): contacting a plurality of first complex polymerases with a plurality of multivalent molecules to form a plurality of multivalent-complex polymerases (e.g., binding complexes). In some embodiments, an individual multivalent molecule among the plurality of multivalent molecules comprises a nucleus linked to a plurality of nucleotide arms, and each nucleotide arm is linked to a nucleotide (e.g., nucleotide unit) (e.g., Figures 16 - 20 ). In some embodiments, step (b) of contacting is carried out under conditions suitable for binding complementary nucleotide units of the multivalent molecules to at least two of the plurality of first complex polymerases thereby forming a plurality of multivalent-complex polymerases. In some embodiments, the conditions are suitable for inhibiting polymerase-catalyzed incorporation of the complementary nucleotide units into the primers of the plurality of multivalent-complex polymerases. In some embodiments, the plurality of multivalent molecules comprises at least one multivalent molecule having a plurality of nucleotide arms (e.g., Figures 16 - 19 ), each nucleotide arm being linked to a nucleotide analogue (e.g., nucleotide analogue unit), wherein the nucleotide analogue comprises a chain terminating moiety at the sugar 2' and / or 3' position. In some embodiments, the plurality of multivalent molecules includes at least one multivalent molecule comprising a plurality of nucleotide arms, each nucleotide arm being attached to a nucleotide unit lacking a chain terminating moiety. In some embodiments, at least one multivalent molecule among the plurality of multivalent molecules is labeled with a detectable reporter gene moiety that emits a signal. In some embodiments, the detectable reporter gene moiety comprises a fluorophore. In some embodiments, step (b) of contacting is carried out in the presence of at least one non-catalytic cation comprising strontium, barium, and / or calcium.
[0460] In some embodiments, the method for sequencing further includes step (c): detecting the plurality of multivalent-complex polymerases. In some embodiments, the detecting comprises detecting a signal emitted by a multivalent molecule that is bound to a complex polymerase, wherein the complementary nucleotide unit of the multivalent molecule binds to a primer but inhibits incorporation of the complementary nucleotide unit. In some embodiments, the multivalent molecule is labeled with a detectable reporter gene moiety to allow detection. In some embodiments, the labeled multivalent molecule comprises a fluorophore attached to the nucleus, linker, and / or nucleotide unit of the multivalent molecule.
[0461] In some embodiments, the method for sequencing further includes step (d): identifying the nucleobase of the complementary nucleotide unit that is bound to the plurality of first complex polymerases, thereby determining the sequence of the template molecule. In some embodiments, the multivalent molecule is labeled with a detectable reporter gene moiety that corresponds to a specific nucleotide unit linked to a nucleotide arm to allow identification of the complementary nucleotide unit (e.g., nucleobase adenine, guanine, cytosine, thymine, or uracil) that is bound to the plurality of first complex polymerases.
[0462] In some embodiments, the method for sequencing further comprises step (e): dissociating a plurality of multivalent - complex polymerases and removing the plurality of first sequencing polymerases and their bound multivalent molecules, and retaining a plurality of nucleic acid duplexes.
[0463] In some embodiments, the second stage of the two - stage sequencing method generally includes nucleotide incorporation. In some embodiments, the method for sequencing further comprises step (f): contacting the plurality of retained nucleic acid duplexes of step (e) with a plurality of second sequencing polymerases, wherein the contacting is carried out under conditions suitable for binding the plurality of second sequencing polymerases to the plurality of retained nucleic acid duplexes, thereby forming a plurality of second complex polymerases each comprising a second sequencing polymerase bound to a nucleic acid duplex. In some embodiments, the second sequencing polymerase comprises a recombinant mutant sequencing polymerase.
[0464] In some embodiments, the plurality of first sequencing polymerases of step (a) have an amino acid sequence that is 100% identical to the amino acid sequence of the plurality of second sequencing polymerases of step (f). In some embodiments, the plurality of first sequencing polymerases of step (a) have an amino acid sequence different from the amino acid sequence of the plurality of second sequencing polymerases of step (f).
[0465] In some embodiments, the method for sequencing further includes step (g): contacting a plurality of second complex polymerases with a plurality of nucleotides, wherein the contacting is performed under conditions suitable for binding at least two complementary nucleotides from the plurality of nucleotides to the complex polymerase, thereby forming a plurality of nucleotide complex polymerases. In some embodiments, the contacting of step (g) is performed under conditions suitable for promoting the catalytic incorporation of the bound complementary nucleotide polymerase into the primer of the nucleotide complex polymerase, thereby extending the sequencing primer by one nucleobase. In some embodiments, incorporating the nucleotide into the 3'-end of the sequencing primer in step (g) comprises a primer extension reaction. In some embodiments, the contacting of step (g) is performed in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the plurality of nucleotides includes natural nucleotides (e.g., non-analogue nucleotides) or nucleotide analogues. In some embodiments, the plurality of nucleotides includes removable or non-removable 2' and / or 3' chain terminating moieties. In some embodiments, at least one nucleotide of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the plurality of nucleotides is unlabeled. In some embodiments, the plurality of nucleotides includes a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety includes a fluorophore. In some embodiments, the fluorophore is linked to the nucleobase. In some embodiments, the fluorophore is attached to the nucleobase with a linker that can be cleaved / removed from the base or cannot be removed from the base. In some embodiments, a particular detectable reporter moiety (e.g., fluorophore) linked to the nucleotide can correspond to a nucleobase (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of the nucleobase.
[0466] In some embodiments, when the plurality of nucleotides in step (g) are detectably labeled, the method for sequencing further includes step (h): detecting the complementary nucleotides incorporated into the primer of the nucleotide complex polymerase. In some embodiments, the plurality of nucleotides is labeled with a detectable reporter moiety to allow detection. In some embodiments, when the plurality of nucleotides in step (g) are unlabeled, the detection of step (h) is omitted.
[0467] In some embodiments, when the plurality of nucleotides in step (g) are detectably labeled, the method for sequencing further comprises step (i): identifying the base of the complementary nucleotide incorporated into the primer of the nucleotide complex polymerase. In some embodiments, identifying the incorporated complementary nucleotide in step (i) can be used to confirm the identity of the complementary nucleotide of the multivalent molecule bound to the plurality of first complex polymerases in step (d). In some embodiments, the identification of step (i) can be used to determine the sequence of the nucleic acid template molecule. In some embodiments, when the plurality of nucleotides in step (g) are unlabeled, the identification of step (i) is omitted.
[0468] In some embodiments, the method for sequencing further comprises step (j): when step (g) is carried out by contacting a plurality of second complex polymerases with a plurality of nucleotides comprising at least one nucleotide having a 2' and / or 3' chain termination moiety, removing the chain termination moiety from the incorporated nucleotides.
[0469] In some embodiments, the method for sequencing further comprises step (k): repeating steps (a) to (j) at least once. In some embodiments, the sequence of the nucleic acid template molecule can be determined by detecting and identifying the multivalent molecule that binds to the sequencing polymerase but is not incorporated into the 3'-end of the primer at steps (c) and (d). In some embodiments, the sequence of the nucleic acid template molecule can be determined (or confirmed) by detecting and identifying the nucleotide incorporated into the 3'-end of the primer at steps (h) and (i).
[0470] In some embodiments, in any method for sequencing a nucleic acid molecule, the binding of a plurality of first complex polymerases to a plurality of multivalent molecules forms at least one affinity complex, the method comprising the steps of: (a) binding a first nucleic acid primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemer template molecule, thereby forming a first binding complex, wherein the first nucleotide unit of the first multivalent molecule binds to the first sequencing polymerase; and (b) binding a second nucleic acid primer, a second sequencing polymerase, and the first multivalent molecule to a second portion of the same concatemer template molecule, thereby forming a second binding complex, wherein the second nucleotide unit of the first multivalent molecule binds to the second sequencing polymerase, wherein the first and second binding complexes comprising the same multivalent molecule form an affinity complex. In some embodiments, the first sequencing polymerase comprises any wild-type or mutant polymerase described herein. In some embodiments, the second sequencing polymerase comprises any wild-type or mutant polymerase described herein. The concatemer template molecule comprises a tandem repeat sequence of a target sequence and at least one universal sequencing primer binding site. The first nucleic acid primer and the second nucleic acid primer can bind to the sequencing primer binding site along the concatemer template molecule. Exemplary multivalent molecules are shown in Figures 16 - 19 in.
[0471] In some embodiments, in any method of nucleic acid molecule sequencing, the method comprises binding a plurality of first complexing polymerases to a plurality of multivalent molecules to form at least one affinity complex, the method comprising the steps of: (a) contacting a plurality of sequencing polymerases and a plurality of nucleic acid primers with different portions of a tandem nucleic acid tandem molecule to form at least a first complex polymerase and a second complex polymerase on the same tandem template molecule; (b) contacting a plurality of multivalent molecules with the at least first complex polymerase and the second complex polymerase on the same tandem template molecule under conditions suitable for binding a single multivalent molecule from the plurality of multivalent molecules to the first complex polymerase and the second complex polymerase, wherein at least a first nucleotide unit of the single multivalent molecule binds to the first complex polymerase, the first complex polymerase comprising a first primer that hybridizes to a first portion of the tandem template molecule, thereby forming a first binding complex (e.g., a first ternary complex), and wherein at least a second nucleotide unit of the single multivalent molecule binds to the second complex polymerase, the second complex polymerase comprising a second primer that hybridizes to a second portion of the tandem template molecule, thereby forming a second binding complex (e.g., a second ternary complex), wherein the contacting is carried out under conditions suitable for inhibiting polymerase-catalyzed incorporation of the bound first nucleotide unit and second nucleotide unit into the first binding complex and the second binding complex, and wherein the first binding complex and the second binding complex bound to the same multivalent molecule form an affinity complex; and (c) detecting the first binding complex and the second binding complex on the same tandem template molecule; and (d) identifying the first nucleotide unit in the first binding complex, thereby determining the sequence of the first portion of the tandem template molecule, and identifying the second nucleotide unit in the second binding complex, thereby determining the sequence of the second portion of the tandem template molecule. In some embodiments, the plurality of sequencing polymerases comprise any wild-type or mutant sequencing polymerase described herein. The tandem template molecule comprises a tandem repeat sequence of a target sequence and at least one universal sequencing primer binding site. The plurality of nucleic acid primers can bind to the sequencing primer binding sites along the tandem template molecule. Exemplary multivalent molecules are shown in Figures 16 - 19 in.
[0472] The present disclosure provides methods for sequencing any immobilized template molecule described herein, wherein the sequencing method comprises a Sequencing by Binding (SBB) procedure using unlabeled chain-terminating nucleotides. In some embodiments, the Sequencing by Binding (SBB) method comprises the steps of: (a) contacting the primed template nucleic acid with at least two separate mixtures sequentially under ternary complex stabilizing conditions, wherein each of the at least two separate mixtures comprises a polymerase and a nucleotide, whereby the sequential contact causes the primed template nucleic acid to contact homologs of nucleotides of the first, second, and third base types in the template under ternary complex stabilizing conditions; (b) examining the at least two separate mixtures to determine whether a ternary complex is formed; and (c) identifying the next correct nucleotide for the primer-template nucleic acid molecule, wherein if a ternary complex is detected in step (b), the next correct nucleotide is identified as a homolog of the first, second, or third base type, and wherein based on the absence of a ternary complex in step (b), the next correct nucleotide is attributed to a nucleotide homolog of a fourth base type; (d) adding the next correct nucleotide to the primer of the primer-template nucleic acid after step (b), thereby producing an extended primer; and (e) repeating steps (a) through (d) at least once on the primer-template nucleic acid comprising the extended primer. Exemplary methods of sequencing while binding are described in U.S. Patent Nos. 10,246,744 and 10,731,141 (the contents of which are hereby incorporated by reference in their entireties).
[0473] Method for Sequencing Using Nucleotides Labeled with Phosphodiester Linkages
[0474] The present disclosure provides methods for sequencing using an immobilized sequencing polymerase that binds to an unimmobilized template molecule, wherein the sequencing reaction is performed with nucleotides labeled with a phosphodiester chain. In some embodiments, the sequencing method comprises step (a): providing a support on which a plurality of sequencing polymerases are immobilized. In some embodiments, the sequencing polymerase comprises a processive DNA polymerase. In some embodiments, the sequencing polymerase comprises a wild-type or mutant DNA polymerase, including, for example, Phi29 DNA polymerase. In some embodiments, the support comprises a plurality of separate compartments, and the sequencing polymerase is immobilized to the bottom of the compartment. In some embodiments, the separate compartment comprises a silica bottom through which light can penetrate. In some embodiments, the separate compartment comprises a silica bottom configured with a nanoparticle confinement structure comprising pores in a metal-coated film (e.g., an aluminum-coated film). In some embodiments, the small pore diameter of the pores in the metal coating is, for example, about 70 nm. In some embodiments, the height of the nanoparticle confinement structure is about 100 nm. In some embodiments, the nanoparticle confinement structure comprises a zero-mode waveguide (ZMW). In some embodiments, the nanoparticle confinement structure contains a liquid.
[0475] In some embodiments, the sequencing method further comprises step (b): contacting a plurality of immobilized sequencing polymerases with a plurality of single-stranded circular nucleic acid template molecules and a plurality of oligonucleotide sequencing primers under conditions suitable for the binding of individual immobilized sequencing polymerases to single-stranded circular template molecules and for the hybridization of individual sequencing primers to individual single-stranded circular template molecules, thereby generating a plurality of polymerase / template / primer complexes. In some embodiments, the individual sequencing primers hybridize to a universal sequencing primer binding site on the single-stranded circular template molecule.
[0476] In some embodiments, the sequencing method further comprises step (c): contacting the plurality of polymerase / template / primer complexes with a plurality of phosphate-chain-labeled nucleotides, each phosphate-chain-labeled nucleotide comprising an aromatic base, a pentose sugar (e.g., ribose or deoxyribose), and a phosphate chain comprising 3 to 20 phosphate groups, wherein the terminal phosphate group is linked to a detectable reporter moiety (e.g., a fluorophore). The first phosphate group, the second phosphate group, and the third phosphate group may be referred to as the α, β, and γ phosphate groups. In some embodiments, the particular detectable reporter moiety attached to the terminal phosphate group corresponds to a nucleobase (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of the nucleobase. In some embodiments, the plurality of polymerase / template / primer complexes are contacted with the plurality of phosphate-chain-labeled nucleotides under conditions suitable for polymerase-catalyzed nucleotide incorporation. In some embodiments, the sequencing polymerase is capable of binding to a phosphate-chain-labeled nucleotide that is complementary and incorporating the complementary nucleotide opposite the nucleotide in the template molecule. In some embodiments, the polymerase-catalyzed nucleotide incorporation reaction cleaves between the α phosphate group and the β phosphate group, thereby releasing the polyphosphate chain linked to the fluorophore.
[0477] In some embodiments, the sequencing method further comprises step (d): detecting the fluorescence signal emitted by the phosphate-chain-labeled nucleotide that is bound by the sequencing polymerase and incorporated into the end of the sequencing primer. In some embodiments, step (d) further comprises identifying the phosphate-chain-labeled nucleotide that is bound by the sequencing polymerase and incorporated into the end of the sequencing primer.
[0478] In some embodiments, the sequencing method further comprises step (e): repeating steps (c) to (d) at least once. In some embodiments, the sequencing method using phosphate-chain-labeled nucleotides may be according to U.S. Patent Nos. 7,170,050; 7,302,146; and / or 7,405,281.
[0479] Sequencing Polymerase
[0480] The present disclosure provides methods for sequencing nucleic acid molecules, wherein any of the sequencing methods described herein employs at least one type of sequencing polymerase and a plurality of nucleotides, or employs at least one type of sequencing polymerase, a plurality of nucleotides, and a plurality of multivalent molecules. In some embodiments, the sequencing polymerase is capable of incorporating complementary nucleotides opposite the nucleotides in the template molecule. In some embodiments, the sequencing polymerase is capable of binding to the complementary nucleotide units of the multivalent molecule opposite the nucleotides in the template molecule. In some embodiments, the plurality of sequencing polymerases includes recombinant mutant polymerases.
[0481] Examples of suitable polymerases for sequencing with nucleotides and / or multivalent molecules include, but are not limited to: Klenow DNA polymerase; Thermus aquaticus DNA polymerase I (Taq polymerase); KlenTaq polymerase; Candidatus altiarchaeales archaea; Candidatus subterraneus thermophilus; Hadarchaeota archaea; Euryarchaeota archaea; Thermoplasmata archaea; Thermococcus polymerases such as Thermococcus kodakarensis, bacteriophage T7 DNA polymerase; human α, δ, and ε DNA polymerases; bacteriophage DNA polymerases such as T4, RB69, phi29; Pyrococcus furiosus DNA polymerase (Pfu polymerase); Bacillus subtilis DNA polymerase III; Escherichia coli DNA polymerase IIIα and ε; 9°N polymerase; reverse transcriptases such as HIV M-type or O-type reverse transcriptase; avian myeloblastosis virus reverse transcriptase; Moloney murine leukemia virus (MMLV) reverse transcriptase; or telomerase. Additional non-limiting examples of DNA polymerases include those from various archaea genera (such as Aeropyrum, Archaeglobus, Desulfurococcus, Pyrobaculum, Pyrococcus, Pyrolobus, Thermofilum, Thermococcus, Thermoproteus, Sulfolobus, Thermococcus, and Vulcanisaeta, etc. or variants thereof), including such polymerases known in the art such as 9°N, VENT, DEEP VENT, THERMINATOR, Pfu, KOD, Pfx, Tgo, and RB69 polymerases.
[0482] Nucleotide
[0483] The present disclosure provides methods for sequencing nucleic acid molecules, wherein any of the sequencing methods described herein employs at least one nucleotide. A nucleotide includes a base, a sugar, and at least one phosphate group. In some embodiments, at least one nucleotide among a plurality of nucleotides comprises an aromatic base, a pentose sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1 - 10 phosphate groups). The plurality of nucleotides may comprise at least one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. The plurality of nucleotides may comprise a mixture of any combination of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, at least one nucleotide among the plurality of nucleotides is not a nucleotide analogue. In some embodiments, at least one nucleotide among the plurality of nucleotides comprises a nucleotide analogue.
[0484] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, at least one nucleotide among the plurality of nucleotides comprises a chain of one, two, or three phosphorus atoms, wherein the chain is typically attached to the 5'-carbon of the sugar moiety via an ester bond or a phosphoramide bond. In some embodiments, at least one nucleotide among the plurality of nucleotides is an analogue having a phosphorus chain, wherein the phosphorus atoms are linked together by intervening O, S, NH, methylene, or ethylene groups. In some embodiments, the phosphorus atoms in the chain include substituted side groups (including O, S, or BH3). In some embodiments, the chain includes a phosphate group substituted with an analogue, the analogue including phosphoramide, thiophosphate, dithiophosphate, and O-methylphosphoramidite groups.
[0485] In some embodiments, in any of the methods described herein for sequencing nucleic acid molecules, at least one nucleotide of a plurality of nucleotides includes a terminator nucleotide analogue that has a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2'-position, at the sugar 3'-position, or at both the sugar 2'- and 3'-positions. In some embodiments, the chain-terminating moiety can inhibit the polymerase-catalyzed incorporation of subsequent nucleotide units or free nucleotides into the nascent strand during a primer extension reaction. In some embodiments, the chain-terminating moiety is linked to the 3'-sugar moiety, where the sugar comprises a ribose or deoxyribose moiety. In some embodiments, the chain-terminating moiety can be removed / cleaved from the 3'-sugar moiety to produce a nucleotide having a 3'-OH sugar group that can be extended with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some aspects, the chain-terminating moiety comprises an alkyl, alkenyl, alkynyl, allyl, aryl, benzyl, azide group, amine group, amide group, ketone group, isocyanate group, phosphate group, thio group, disulfide group, carbonate group, urea group, silyl group, or acetal group. In some embodiments, the chain-terminating moiety can be cleaved / removed from the nucleotide, e.g., by reacting the chain-terminating moiety with a chemical agent, pH change, light, or heat. In some embodiments, the chain-terminating moieties alkyl, alkenyl, alkynyl, and allyl can be cleaved with tetrakis(triphenylphosphine)-palladium(0) (Pd(PPh3)4), with piperidine, or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the chain-terminating moieties aryl and benzyl can be cleaved with H2 Pd / C. In some embodiments, the chain-terminating moieties amine, amide, ketone, isocyanate, phosphate, thio, disulfide can be cleaved with a phosphine or thiol group (including β-mercaptoethanol or dithiothreitol (DTT)). In some embodiments, the chain-terminating moiety carbonate can be cleaved with potassium carbonate (K2CO3) in MeOH, with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some embodiments, the chain-terminating moieties urea and silyl can be cleaved with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride. In some embodiments, the chain-terminating moiety can be cleaved / removed with nitrous acid. In some embodiments, the chain-terminating moiety can be cleaved / removed using a solution comprising nitrite (e.g., a combination of nitrite with an acid such as acetic acid, sulfuric acid, or nitric acid). In some additional embodiments, the solution can comprise an organic acid.
[0486] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, at least one nucleotide of the plurality of nucleotides includes a terminator nucleotide analog that has a chain terminating moiety (e.g., a blocking moiety) at the sugar 2'-position, at the sugar 3'-position, or at both the sugar 2'- and 3'-positions. In some embodiments, the chain terminating moiety comprises an azide, an azido group, or an azidomethyl group. In some embodiments, the chain terminating moiety comprises 3'-O-azido or 3'-O-azidomethyl. In some embodiments, the chain terminating moieties azide, azido, and azidomethyl are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bissulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleavage agent comprises 4-dimethylaminopyridine (4-DMAP). In some embodiments, a chain terminating moiety comprising one or more of 3'-O-amino, 3'-O-aminomethyl, 3'-O-methylamino, or a derivative thereof can be cleaved with nitrous acid by a mechanism utilizing nitrous acid or using a solution comprising nitrous acid. In some embodiments, a chain terminating moiety comprising one or more of 3'-O-amino, 3'-O-aminomethyl, 3'-O-methylamino, or a derivative thereof can be cleaved using a solution comprising nitrite. In some embodiments, for example, the nitrite can be combined or contacted with an acid such as acetic acid, sulfuric acid, or nitric acid. In some additional embodiments, for example, the nitrite can be combined or contacted with an organic acid (e.g., formic acid, acetic acid, propionic acid, butyric acid, isobutyric acid, etc.). In some embodiments, the chain terminating moiety comprises a 3'-acetal moiety that can be cleaved with a palladium deblocking reagent (e.g., Pd(0)).
[0487] In some embodiments, in any of the methods for nucleic acid molecule sequencing described herein, the nucleotide includes a chain terminating moiety selected from the group consisting of: 3'-deoxynucleotide, 2',3'-dideoxynucleotide, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-mercapto, 3'-aminomethyl, 3'-ethyl, 3'-butyl, 3'-tert-butyl, 3'-fluorenylmethoxycarbonyl, 3'-tert-butoxycarbonyl, 3'-O-alkylhydroxyamino group, 3'-thiophosphate, and 3-O-benzyl and 3'-O-benzyl, 3-acetal moiety, or a derivative thereof.
[0488] In some embodiments, in any of the methods described herein for sequencing nucleic acid molecules, the plurality of nucleotides includes a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety includes a fluorophore. In some embodiments, the fluorophore is linked to a nucleobase. In some embodiments, the fluorophore is linked to the nucleobase with a linker cleavable / removable from the base. In some embodiments, at least one nucleotide of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, a particular detectable reporter moiety (e.g., a fluorophore) linked to a nucleotide can correspond to a nucleobase (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of the nucleobase.
[0489] In some embodiments, in any of the methods described herein for sequencing nucleic acid molecules, the cleavable linker on the nucleobase comprises a cleavable moiety that comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a ketone group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the cleavable linker on the base can be cleaved / removed from the base by reacting the cleavable moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the cleavable moieties alkyl, alkenyl, alkynyl, and allyl can be cleaved with tetrakis(triphenylphosphine)-palladium(0) (Pd(PPh3)4), with piperidine, or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some ...
Claims
1. A computer-implemented method for sequencing a three-dimensional sample, the computer-implemented method comprising: obtaining, by a processor, a first plurality of flow cell images of the sample in a first plurality of sequencing cycles from a first subset of channels; obtaining, by the processor, a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from the first subset of channels; generating, by the processor, a first set of base calls for a first subset of communities of the sample based on the first plurality of flow cell images; and generating, by the processor, a second set of base calls for a second subset of communities of the sample based on the second plurality of flow cell images, wherein the first subset of channels includes only some of the channels in the channel, and wherein the second plurality of sequencing cycles is after the first plurality of sequencing cycles.
2. A computer-implemented method for sequencing a three-dimensional sample, the computer-implemented method comprising: obtaining, by a processor, a first plurality of flow cell images of the sample in a first plurality of sequencing cycles from one or more channels; obtaining, by the processor, a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from the one or more channels; generating, by the processor, a first set of base calls for a first subset of communities of the sample based on the first plurality of flow cell images; and generating, by the processor, a second set of base calls for a second subset of communities of the sample based on the second plurality of flow cell images, wherein the one or more channels include dark channels, and wherein the image intensity of some of the first plurality of flow cell images or the second plurality of flow cell images obtained from the dark channels is lower than a predetermined threshold, and wherein the second plurality of sequencing cycles is after the first plurality of sequencing cycles.
3. A computer-implemented method for sequencing a three-dimensional sample, the computer-implemented method comprising: obtaining, by a processor, a first plurality of flow cell images of the sample in a first plurality of sequencing cycles from a first subset of channels; obtaining, by the processor, a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from a second subset of the channels; generating, by the processor, a first set of base calls for a first subset of communities of the sample based on the first plurality of flow cell images; and generating, by the processor, a second set of base calls for a second subset of communities of the sample based on the second plurality of flow cell images, wherein the first subset and the second subset of channels are at least partially different, and the first subset of channels lacks a first dark channel, and wherein the second subset of channels lacks a second dark channel different from the first dark channel.
4. The computer-implemented method according to any one of the preceding claims, wherein in the first plurality of flow cell images, the image intensity of the second subset of communities is lower than a predetermined threshold.
5. The computer-implemented method according to any one of the preceding claims, wherein in the first plurality of flow cell images, the second subset of communities appears dark.
6. The computer-implemented method according to any one of the preceding claims, wherein in the first plurality of flow cell images, the second subset of communities appears 5 times, 10 times, 15 times, or 20 times darker than the first subset of communities.
7. The computer-implemented method according to any one of the preceding claims, wherein in the second plurality of flow cell images, the image intensity of the first subset of communities is below a predetermined threshold.
8. The computer-implemented method according to any one of the preceding claims, wherein in the second plurality of flow cell images, the first subset of communities appears dark.
9. The computer-implemented method according to any one of the preceding claims, wherein in the second plurality of flow cell images, the first subset of communities appears 5 times, 10 times, 15 times, or 20 times darker than the second subset of communities.
10. The computer-implemented method according to any one of the preceding claims, wherein the first subset of communities is different from the second subset of communities.
11. The computer-implemented method according to any one of the preceding claims, wherein the first subset of communities and the second subset of communities spatially at least partially overlap in 2D or 3D.
12. The computer-implemented method according to any one of the preceding claims, wherein the first subset of communities and the second subset of communities comprise the same batch-specific sequencing binding sites, the same batch-specific sequencing binding sites being configured to bind to the same sequencing primers.
13. The computer-implemented method according to any one of the preceding claims, wherein each of the first subset of communities and the second subset of communities comprises the same batch-specific sequencing binding sites, the same batch-specific sequencing binding sites being configured to bind to the same sequencing primers.
14. The computer-implemented method according to any one of the preceding claims, wherein each of the communities in the first subset of communities is configured to bind to a first sequencing primer, and each of the communities in the second subset of communities is configured to bind to a second sequencing primer.
15. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample in the first plurality of sequencing cycles from the first subset of channels comprises: Obtaining the first plurality of flow cell images of the sample in the first plurality of sequencing cycles from only the first subset of channels and not from a dark channel by an optical system of a sequencing system.
16. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample in the first plurality of sequencing cycles from the one or more channels comprises: Obtaining the first plurality of flow cell images of the sample in the first plurality of sequencing cycles from the one or more channels by an optical system of a sequencing system.
17. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample in the first plurality of sequencing cycles from the first subset of channels comprises: The processor controls the optical system to avoid collecting data from one or more image sensors in the dark channel during the first plurality of sequencing cycles.
18. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample during the first plurality of sequencing cycles from the first subset of channels comprises: The processor controls the optical system to avoid irradiating the sample with light within a predetermined frequency range corresponding to the dark channel during the first plurality of sequencing cycles.
19. The computer-implemented method according to any one of the preceding claims, wherein obtaining the second plurality of flow cell images of the sample during the second plurality of sequencing cycles from the first subset of channels comprises: The optical system of the sequencing system obtains the second plurality of flow cell images of the sample during the second plurality of sequencing cycles only from the first subset of channels and not from the dark channel.
20. The computer-implemented method according to any one of the preceding claims, wherein obtaining the second plurality of flow cell images of the sample during the second plurality of sequencing cycles from the one or more channels comprises: The optical system of the sequencing system obtains the second plurality of flow cell images of the sample during the second plurality of sequencing cycles from the one or more channels.
21. The computer-implemented method according to any one of the preceding claims, wherein the one or more channels include at least a dark channel and at least a channel that is not a dark channel.
22. The computer-implemented method according to any one of the preceding claims, wherein the one or more channels contain only a single dark channel and two or three channels different from the dark channel.
23. The computer-implemented method according to any one of the preceding claims, wherein obtaining the second plurality of flow cell images of the sample during the second plurality of sequencing cycles from the first subset of channels comprises: The processor controls the optical system to avoid collecting any data from one or more image sensors in the dark channel during the second plurality of sequencing cycles.
24. The computer-implemented method according to any one of the preceding claims, wherein obtaining the second plurality of flow cell images of the sample during the second plurality of sequencing cycles from the first subset of channels comprises: The processor controls the optical system to avoid irradiating the sample with light within a predetermined frequency range corresponding to the dark channel during the second plurality of sequencing cycles.
25. The computer-implemented method according to any one of the preceding claims, wherein the first set of base identifications comprises only one, two, or three types of nucleotide bases.
26. The computer-implemented method according to any one of the preceding claims, wherein the first set of base identifications comprises four types of nucleotide bases.
27. The computer-implemented method according to any one of the preceding claims, wherein the second set of base identifications comprises only one, two, or three types of nucleotide bases.
28. The computer-implemented method according to any one of the preceding claims, wherein the second set of base identifications comprises four types of nucleotide bases.
29. The computer-implemented method according to any one of the preceding claims, wherein generating the first set of base identifications for the first subset of the population of the sample based on the first plurality of flow cell images comprises: Generating the first set of base identifications for the first subset of the population of the sample based only on the first plurality of flow cell images.
30. The computer-implemented method according to any one of the preceding claims, wherein generating the second set of base identifications for the second subset of the population of the sample based on the second plurality of flow cell images comprises: Generating the second set of base identifications for the second subset of the population of the sample based only on the second plurality of flow cell images.
31. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequencing cycles comprises the same number of cycles as the second plurality of sequencing cycles.
32. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequencing cycles or the second plurality of sequencing cycles comprises 2 to 40 cycles.
33. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequencing cycles or the second plurality of sequencing cycles comprises 3 to 10 cycles.
34. The computer-implemented method according to any one of the preceding claims, further comprising: Obtaining, by the processor, a third plurality of flow cell images of the sample in a third plurality of sequencing cycles from the one or more channels; And Generating, by the processor, a third set of base identifications for a third subset of the population of the sample based on the third plurality of flow cell images.
35. The computer-implemented method according to any one of the preceding claims, further comprising: Obtaining, by the processor, a third plurality of flow cell images of the sample in a third plurality of sequencing cycles from a first subset of the channels; And Generating, by the processor, a third set of base identifications for a third subset of the population of the sample based on the third plurality of flow cell images.
36. The computer-implemented method according to any one of the preceding claims, further comprising: Obtaining, by the processor, a third plurality of flow cell images of the sample in a third plurality of sequencing cycles from a third subset of the channels; And Generating, by the processor, a third set of base identifications for a third subset of the population of the sample based on the third plurality of flow cell images.
37. The computer-implemented method according to any one of the preceding claims, wherein the third set of channels does not include a third dark channel.
38. The computer-implemented method according to any one of the preceding claims, wherein the dark channel corresponds to a fluorescent dye attached to a nucleotide of adenine (A), thymine (T), guanine (G), or cytosine (C), and the fluorescent dye emits light below a predetermined threshold in the first plurality of sequencing cycles or the second plurality of sequencing cycles.
39. The computer-implemented method according to any one of the preceding claims, wherein the dark channel corresponds to any fluorescent dye of nucleotides attached to adenine (A), thymine (T), guanine (G), or cytosine (C) in the first plurality of sequencing cycles or the second plurality of sequencing cycles.
40. The computer-implemented method according to any one of the preceding claims, wherein the first dark channel corresponds to a fluorescent dye of a first type of nucleotide attached to adenine (A), thymine (T), guanine (G), or cytosine (C), the fluorescent dye emitting light below a predetermined threshold in the first plurality of sequencing cycles, and wherein the second dark channel corresponds to a second fluorescent dye of a second type of nucleotide attached to A, T, G, or C, the second fluorescent dye emitting light below a predetermined threshold in the second plurality of sequencing cycles.
41. The computer-implemented method according to any one of the preceding claims, wherein the first dark channel corresponds to a first channel from which no flow cell image is obtained in the first plurality of sequencing cycles, and wherein the second dark channel corresponds to a second channel from which no flow cell image is obtained in the second plurality of sequencing cycles.
42. The computer-implemented method according to any one of the preceding claims, wherein the first dark channel corresponds to a first channel, in the first plurality of sequencing cycles, from which the sample does not generate a fluorescent emission above a predetermined threshold and within a frequency range corresponding to the first channel; and wherein the second dark channel corresponds to a second channel, in the second plurality of sequencing cycles, from which the sample does not generate a fluorescent emission above a predetermined threshold and within a frequency range corresponding to the second channel.
43. The computer-implemented method according to any one of the preceding claims, wherein the first dark channel corresponds to a first channel, from which only a flow cell image having an image intensity below a predetermined threshold is obtained in the first plurality of sequencing cycles; and wherein the second dark channel corresponds to a second channel, from which only a flow cell image having an image intensity below a predetermined threshold is obtained in the second plurality of sequencing cycles.
44. The computer-implemented method according to any one of the preceding claims, wherein the first subset, the second subset, and the third subset of channels are at least partially different.
45. The computer-implemented method according to any one of the preceding claims, wherein the channels comprise 2, 3, or 4 channels.
46. The computer-implemented method according to any one of the preceding claims, wherein each of the first set of base identifications or the second set of base identifications comprises a base identification sequence corresponding to the first plurality of sequencing cycles or the second plurality of sequencing cycles.
47. The computer-implemented method according to any one of the preceding claims, further comprising: determining whether one or more base identification sequences of the first set of base identifications and / or the second set of base identifications match at least a portion of a barcode sequence; and In response to the determination, assign the one or more sequences of the first set of base identifications and / or the second set of base identifications to a corresponding barcode sequence.
48. The computer-implemented method according to any one of the preceding claims, further comprising: Determine whether one or more sequences of the first set of base identifications and / or the second set of base identifications match only a part of the barcode sequence; And In response to the determination, assign the one or more sequences of the first set of base identifications and / or the second set of base identifications to the corresponding barcode sequence.
49. The computer-implemented method according to any one of the preceding claims, wherein the barcode sequence uniquely identifies a DNA or RNA fragment of the sample.
50. The computer-implemented method according to any one of the preceding claims, wherein generating the first set of base identifications, the second set of base identifications, or the third set of base identifications is not based on any flow cell images from the dark channel, the first dark channel, or the second dark channel.
51. The computer-implemented method according to any one of the preceding claims, wherein generating the first set of base identifications, the second set of base identifications, or the third set of base identifications is not based on any flow cell images from one or more dark channels of the channel.
52. The computer-implemented method according to any one of the preceding claims, wherein generating the first set of base identifications, the second set of base identifications, or the third set of base identifications is not based on any flow cell images from any dark channels of the channel.
53. The computer-implemented method according to any one of the preceding claims, wherein the barcode sequence is pre-determined and uniquely different from other barcode sequences.
54. The computer-implemented method according to any one of the preceding claims, wherein each barcode sequence corresponds to a subset of the communities of the sample.
55. The computer-implemented method according to any one of the preceding claims, wherein each community of the first subset of communities of the sample contains the same barcode sequence.
56. The computer-implemented method according to any one of the preceding claims, wherein the total number of different barcode sequences matches the total number of subsets of communities.
57. The computer-implemented method according to any one of the preceding claims, wherein the first n bases of the barcode sequence contain only two or three types of nucleotide bases, and wherein the remainder of the barcode sequence contains all four different types of nucleotide bases.
58. The computer-implemented method according to any one of the preceding claims, wherein one or more reference cycles correspond to the barcode sequence and correspond to only three types of nucleotide bases, and wherein the remainder of the barcode sequence corresponds to subsequent cycles and contains all four different types of nucleotide bases.
59. The computer-implemented method according to any one of the preceding claims, wherein the barcode sequence contains only three types of nucleotide bases.
60. The computer-implemented method according to any one of the preceding claims, wherein the barcode sequence comprises all four types of nucleotide bases.
61. The computer-implemented method according to any one of the preceding claims, wherein one or more nucleotides after a preset number of consecutive repetitions of the same nucleotide base are randomly selected from 3 types of nucleotide bases other than the same type of nucleotide base.
62. The computer-implemented method according to any one of the preceding claims, wherein the barcode sequence has from about 2 to about 100 nucleotide bases.
63. The computer-implemented method according to any one of the preceding claims, wherein the barcode sequence has from about 3 to 60 nucleotide bases.
64. The computer-implemented method according to any one of the preceding claims, wherein the method increases the sequencing throughput by n times compared to existing 3D sequencing methods, where n is the total number of different types of barcode sequences, and where n is in the range of 2 to 100.
65. The computer-implemented method according to any one of the preceding claims, wherein the method increases the sequencing throughput by n times compared to existing 3D sequencing methods, where n is the total number of different types of barcode sequences, and where n is in the range of 2 to 20.
66. The computer-implemented method according to any one of the preceding claims, wherein the first community subset, the second community subset, or the third community subset has a spatial density of not less than about 0.01 to about 0.5 communities / μm^3.
67. The computer-implemented method according to any one of the preceding claims, wherein the sample contains communities with a spatial density of not less than about 0.1 to about 1 community / μm^3.
68. The computer-implemented method according to any one of the preceding claims, wherein the community density is at least in the image plane.
69. The computer-implemented method according to any one of the preceding claims, wherein the community density is in 3D.
70. The computer-implemented method according to any one of the preceding claims, wherein the channels include 4 channels, and the first subset of the channels includes 2 or 3 channels.
71. The computer-implemented method according to any one of the preceding claims, wherein the channels include 4 channels, and the first subset of the channels includes only 2 or 3 channels.
72. The computer-implemented method according to any one of the preceding claims, wherein the first subset of the channels does not include channels corresponding to fluorescent dyes attached to adenine (A), thymine (T), guanine (G), or cytosine (C).
73. The computer-implemented method according to any one of the preceding claims, wherein the barcode sequence comprises consecutive repetitions of the same unique nucleotide base not exceeding 3, 4, 5, 6, 7, 8, 9, or 10 times.
74. The computer-implemented method according to any one of the preceding claims, wherein the one or more channels include 3 channels.
75. The computer-implemented method according to any one of the preceding claims, wherein the one or more channels comprise 4 channels.
76. The computer-implemented method according to any one of the preceding claims, further comprising: determining, by the processor and for a sequencing cycle, that the image intensities from both the first set of colonies and the second set of colonies are higher than a predetermined threshold; and in response to the determination, not obtaining a flow cell image in the sequencing cycle.
77. The computer-implemented method according to any one of the preceding claims, further comprising: determining, by the processor and for a sequencing cycle, that the optical signals from both the first set of colonies and the second set of colonies are higher than a predetermined threshold; and in response to the determination, not acquiring or storing a flow cell image in the sequencing cycle.
78. The computer-implemented method according to any one of the preceding claims, wherein the sequencing cycle is before the first plurality of sequencing cycles or the second plurality of sequencing cycles, after the first plurality of sequencing cycles or the second plurality of sequencing cycles, or before and after the first plurality of sequencing cycles or the second plurality of sequencing cycles.
79. The computer-implemented method according to any one of the preceding claims, wherein the sequencing cycle is before and after the first plurality of sequencing cycles.
80. The computer-implemented method according to any one of the preceding claims, wherein the sequencing cycle is before and after the second plurality of sequencing cycles.
81. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of sequencing cycles and the second plurality of sequencing cycles are included in a single sequencing run.
82. The computer-implemented method according to any one of the preceding claims, wherein generating the first set of base calls for the first subset of colonies of the sample based on the first plurality of flow cell images comprises: generating, by the processor, a first plurality of processed images of the first plurality of flow cell images; filtering, by the processor, the first plurality of flow cell images based on the first plurality of processed images to generate a first plurality of filtered images; obtaining, by the processor, a 3D colony map of the sample; extracting, by the processor, the image intensities of the colonies from one of: a second plurality of flow cell images; a second plurality of processed images; a second plurality of filtered images; or a combination thereof; and performing, by the processor, 3D base calling on the first subset of colonies of the sample based on the extracted image intensities of the colonies.
83. The computer-implemented method according to any one of the preceding claims, wherein generating the first set of base calls for the first subset of colonies of the sample based on the first plurality of flow cell images comprises: filtering, by the processor, the first plurality of flow cell images or the second plurality of flow cell images to generate a plurality of filtered images; obtaining, by the processor, a 3D colony map based on the filtered images of the first plurality of flow cell images or the second plurality of flow cell images; The processor extracts the image intensity of the community from the following based on the 3D community map: The first plurality of flow cell images or the second plurality of flow cell images; Processed images of the first plurality of flow cell images or the second plurality of flow cell images; Filtered images of the first plurality of flow cell images or the second plurality of flow cell images; Or A combination thereof; And The processor performs base identification based on the extracted image intensity of the community.
84. The computer-implemented method according to any one of the preceding claims, wherein generating the first set of base identifications for the first community subset of the sample based on the first plurality of flow cell images includes: The processor generates a plurality of processed images of the plurality of flow cell images; The processor filters the plurality of flow cell images based on the plurality of processed images to generate a plurality of filtered images; The processor generates a first maximum intensity projection (MIP) image based on the plurality of filtered images; And The processor uses the first MIP image to perform base identification.
85. The computer-implemented method according to any one of the preceding claims, wherein generating the first set of base identifications for the first community subset of the sample based on the first plurality of flow cell images includes: The processor filters the plurality of flow cell images through a top-hat filter, a difference of Gaussians (DoG) filter, or a Mexican hat filter to generate a plurality of filtered images; The processor generates a first maximum intensity projection (MIP) image based on the plurality of filtered images; And The processor uses the first MIP image to perform base identification.
86. The computer-implemented method according to any one of the preceding claims, further comprising: Contacting at least one of the first community subset, the second community subset, and the third community subset of the sample with a first mixture of a plurality of sequencing primers, a first plurality of polymerases, and different types of adaptors.
87. The computer-implemented method according to any one of the preceding claims, wherein a separate adaptor in the first mixture comprises a core attached with a plurality of nucleotide arms, and each arm of the separate adaptor comprises the same type of nucleotide unit.
88. The computer-implemented method according to any one of the preceding claims, wherein the first mixture of different types of adaptors comprises 4 different types of adaptors.
89. The computer-implemented method according to any one of the preceding claims, wherein one type of adaptor is labeled with one type of dark fluorescent dye that emits light below a predetermined threshold in the channel.
90. The computer-implemented method according to any one of the preceding claims, wherein one type of adaptor lacks labeling with any fluorescent dye.
91. The computer-implemented method according to any one of the preceding claims, wherein each type of the different types of affinity bodies in the first mixture is labeled with a type of fluorescent dye corresponding to a nucleotide unit to distinguish the different types of affinity bodies in the first mixture.
92. The computer-implemented method according to any one of the preceding claims, wherein the fluorescent dyes of each type of affinity body in the first mixture emit light of different wavelengths upon excitation.
93. The computer-implemented method according to any one of the preceding claims, wherein the first mixture of different types of affinity bodies contains only two or three different types of affinity bodies.
94. The computer-implemented method according to any one of the preceding claims, wherein two or three types of the different types of affinity bodies in the first mixture are labeled with a corresponding type of fluorescent dye corresponding to the nucleotide unit to distinguish the two or three types of affinity bodies among the different types of affinity bodies in the first mixture.
95. The computer-implemented method according to any one of the preceding claims, wherein the fluorescent dye of at least one type of affinity body emits light below a predetermined threshold, and wherein the fluorescent dyes of other types of affinity bodies emit light above a second predetermined threshold.
96. The computer-implemented method according to any one of the preceding claims, wherein the flow cell images of the first plurality of flow cell images or the second plurality of flow cell images are acquired by a next-generation sequencing (NGS) system.
97. The computer-implemented method according to any one of the preceding claims, wherein the sample is an in-situ sample located on the flow cell.
98. The computer-implemented method according to any one of the preceding claims, wherein the in-situ sample contains one or more cells or tissues.
99. The computer-implemented method according to any one of the preceding claims, wherein the in-situ sample contains a community.
100. The computer-implemented method according to any one of the preceding claims, wherein at least some of the communities in the community overlap partially or completely spatially with other communities in the community.
101. The computer-implemented method according to any one of the preceding claims, wherein the community includes at least the first community subset and the second community subset.
102. The computer-implemented method according to any one of the preceding claims, wherein the community includes n community subsets, where n is an integer greater than 2.
103. The computer-implemented method according to any one of the preceding claims, wherein each of the first plurality of flow cell images or the second plurality of flow cell images includes a corresponding field of view orthogonal to the axial axis.
104. The computer-implemented method according to any one of the preceding claims, wherein the corresponding fields of view are the same in the image plane.
105. The computer-implemented method according to any one of the preceding claims, wherein the field of view of each of the first plurality of flow cell images or the second plurality of flow cell images covers at least a portion of a tile of the flow cell.
106. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images and the second plurality of flow cell images have the same image resolution.
107. The computer-implemented method according to any one of the preceding claims, wherein the axial axis extends from the objective lens to a sample located on the flow cell, and the flow cell is positioned on the sequencing system.
108. The computer-implemented method according to any one of the preceding claims, wherein the axial axis is orthogonal to the image plane, and wherein the field of view is within the image plane.
109. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images or the second plurality of processed images comprises: selecting a kernel; and generating the first plurality of processed images or the second plurality of processed images by performing an opening operation on the flow cell images of the first plurality of flow cell images or the second plurality of flow cell images using the selected kernel.
110. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images or the second plurality of processed images comprises: selecting a kernel; and generating the first plurality of processed images or the second plurality of processed images by convolving the first plurality of flow cell images or the second plurality of flow cell images with the selected kernel.
111. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images or the second plurality of processed images further comprises: selecting a first kernel and a second kernel; generating a first blurred image by convolving the first plurality of flow cell images or the second plurality of flow cell images with the first kernel; and generating a second blurred image by convolving the first plurality of flow cell images or the second plurality of flow cell images with the second kernel.
112. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images or the second plurality of processed images comprises: scaling the first plurality of processed images or the second plurality of processed images.
113. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images or the second plurality of processed images comprises: scaling the first blurred image, the second blurred image, or both.
114. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images or the second plurality of flow cell images based on the first plurality of processed images or the second plurality of processed images comprises: Subtract the second blurred image from the first blurred image to generate the first plurality of filtered images or the second plurality of filtered images.
115. The computer-implemented method according to any one of the preceding claims, wherein the kernel is 2x2, 3x3, 4x4, 5x5 or 6x6 pixels.
116. The computer-implemented method according to any one of the preceding claims, wherein the kernel is a circular kernel.
117. The computer-implemented method according to any one of the preceding claims, wherein the kernel is a Gaussian kernel.
118. The computer-implemented method according to any one of the preceding claims, wherein the first kernel and the second kernel are different Gaussian kernels.
119. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images or the second plurality of flow cell images based on the processed images of the first plurality of flow cell images or the second plurality of flow cell images comprises: Subtract each of the first plurality of processed images or the second plurality of processed images from the corresponding flow cell image of the first plurality of flow cell images or the second plurality of flow cell images to generate the first plurality of filtered images or the second plurality of filtered images.
120. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images or the second plurality of flow cell images based on the first plurality of processed images or the second plurality of processed images further comprises: Adding a predetermined offset to the subtracted image to generate the first plurality of filtered images or the second plurality of filtered images.
121. The computer-implemented method according to any one of the preceding claims, wherein generating the first MIP image based on the first plurality of filtered images or the second plurality of filtered images comprises: Calculating the maximum intensity among the intensities of the first plurality of filtered images or the second plurality of filtered images at each pixel of the first MIP image at the corresponding pixel.
122. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Registering the first MIP image with one or more images of the sample.
123. The computer-implemented method according to any one of the preceding claims, wherein the one or more images include staining of the following: membrane, nucleus, or a combination thereof.
124. The computer-implemented method according to any one of the preceding claims, wherein the one or more images include staining of one or more membrane proteins.
125. The computer-implemented method according to any one of the preceding claims, wherein the one or more images include staining of lipids.
126. The computer-implemented method according to any one of the preceding claims, wherein the one or more images include a fluorescence signal from a cell membrane.
127. The computer-implemented method according to any one of the preceding claims, wherein the one or more images include a segmentation of: cells, membranes, cell nuclei, or a combination thereof.
128. The computer-implemented method according to any one of the preceding claims, wherein performing base calling using the first MIP image comprises: performing one or more primary analysis steps to adjust the image intensity of the colonies in the first MIP image; and performing base calling on the colonies based on the adjusted image intensity; wherein the one or more primary analysis steps include: background subtraction; image sharpening; intensity shift adjustment; color correction; intensity normalization; phasing and pre-phasing correction; image registration; quality score estimation; or a combination thereof.
129. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: performing image registration on the first plurality of flow cell images or the second plurality of flow cell images, the first plurality of processed images or the second plurality of processed images, the first plurality of filtered images or the second plurality of filtered images, the first MIP image, or a combination thereof.
130. The computer-implemented method according to any one of the preceding claims, wherein performing image registration on the first plurality of flow cell images or the second plurality of flow cell images comprises: registering the first MIP image with a template image.
131. The computer-implemented method according to any one of the preceding claims, wherein performing image registration on the first plurality of flow cell images or the second plurality of flow cell images comprises: registering the first plurality of flow cell images or the second plurality of flow cell images, the first plurality of processed images or the second plurality of processed images, the first plurality of filtered images or the second plurality of filtered images, the first MIP image, or a combination thereof with a template image.
132. The computer-implemented method according to any one of the preceding claims, wherein performing image registration on the first plurality of flow cell images or the second plurality of flow cell images comprises: registering the colonies in the first MIP image with template colonies in the template image.
133. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: obtaining, by the processor, a second MIP image based on the first plurality of flow cell images or the second plurality of flow cell images; and performing image registration on the first plurality of flow cell images or the second plurality of flow cell images, the first plurality of processed images or the second plurality of processed images, the first plurality of filtered images or the second plurality of filtered images, or a combination thereof based on the second MIP image.
134. A computer-implemented method according to any of the preceding claims, wherein performing image registration on the first plurality of flow cell images or the second plurality of flow cell images, the first plurality of processed images or the second plurality of processed images, the first plurality of filtered images or the second plurality of filtered images, or a combination thereof, comprises: Registering the second MIP image with a template image.
135. A computer-implemented method according to any of the preceding claims, wherein performing image registration on the first plurality of flow cell images or the second plurality of flow cell images based on the second MIP image comprises: Registering the colonies in the second MIP image with the template colonies in the template image.
136. A computer-implemented method according to any of the preceding claims, wherein performing image registration on the first plurality of flow cell images or the second plurality of flow cell images based on the first MIP image or the second MIP image comprises: Registering the colonies in one or more reference cycles with the one or more template images in a reference coordinate system by using the coordinates of the colonies, and generating one or more template images; Determining, by the processor, a plurality of transformations of the first MIP image or the second MIP image based on the one or more template images, the plurality of transformations corresponding to sub-blocks of the first MIP or the second MIP and configured to register the sub-blocks with the one or more template images; And Registering the sub-blocks with the one or more template images by using the plurality of transformations.
137. A computer-implemented method according to any of the preceding claims, wherein the plurality of transformations includes one or more affine transformations.
138. A computer-implemented method according to any of the preceding claims, wherein each of the plurality of transformations includes an affine transformation.
139. A computer-implemented method according to any of the preceding claims, wherein performing base identification by using the first MIP image comprises: Performing base identification based on the image intensity of the colonies from the first MIP image in the first plurality of flow cell images or the second plurality of flow cell images and the position information of the colonies from the second MIP image.
140. A computer-implemented method according to any of the preceding claims, wherein the method further comprises: Performing image registration on the colonies of the first plurality of flow cell images or the second plurality of flow cell images based on fiducial markers.
141. A computer-implemented method according to any of the preceding claims, wherein the fiducial markers are located on the flow cell.
142. A computer-implemented method according to any of the preceding claims, wherein the fiducial markers are outside the flow cell.
143. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the second plurality of flow cell images are acquired at 2, 3, 4, 5, 6, 7, 8, 9 or 10 different positions along the axial axis.
144. The computer-implemented method according to any one of the preceding claims, wherein two adjacent positions along the axial axis are separated by about 1um, 2um, 3um, 4um, 5um, 6um, 7um, 8um, 9um, 10um, 11um or 12um.
145. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images or the second plurality of flow cell images are acquired from 1, 2, 3, 4, 5 or 6 channels.
146. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more processing units; one or more integrated circuits; or a combination thereof.
147. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); one or more neural processing units (NPUs); one or more artificial intelligence (AI) chips; or a combination thereof.
148. The computer-implemented method according to any one of the preceding claims, further comprising: transmitting the base identification by the processor to a processing unit.
149. The computer-implemented method according to any one of the preceding claims, wherein the processing unit is a central processing unit (CPU).
150. The computer-implemented method according to any one of the preceding claims, wherein the processing unit is configured to register the base identification with one or more images.
151. A computer-implemented system for sequencing a three-dimensional sample, the computer-implemented system comprising: one or more hardware processors; one or more data storage devices that store instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including: obtaining, by a processor, a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from a first subset of channels; obtaining, by the processor, a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from the first subset of channels; generating, by the processor, a first set of base identifications for a first subset of communities of the sample based on the first plurality of flow cell images; and generating, by the processor, a second set of base identifications for a second subset of communities of the sample based on the second plurality of flow cell images, wherein the first subset of channels includes only some of the channels, and wherein the second plurality of sequencing cycles are after the first plurality of sequencing cycles.
152. A computer-implemented system, comprising: one or more hardware processors; One or more data storage devices that store instructions that are executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including any one of the preceding claims.
153. One or more non-transitory computer storage media encoded with instructions that are executable by one or more hardware processors to perform operations including: Obtaining, by a processor, a first plurality of flow cell images of a sample in a first plurality of sequencing cycles from a first subset of channels; Obtaining, by the processor, a second plurality of flow cell images of the sample in a second plurality of sequencing cycles from the first subset of channels; Generating, by the processor, a first set of base calls for a first subset of communities of the sample based on the first plurality of flow cell images; And Generating, by the processor, a second set of base calls for a second subset of communities of the sample based on the second plurality of flow cell images, Wherein the first subset of channels includes only some of the channels in the channel, and wherein the second plurality of sequencing cycles is after the first plurality of sequencing cycles.
154. One or more non-transitory computer storage media encoded with instructions that are executable by one or more hardware processors to perform operations including any one of the preceding claims.
155. A computer-implemented method for base calling in sequencing data analysis, the computer-implemented method including: Obtaining, by a processor, a first plurality of flow cell images of a sample, wherein each flow cell image of the first plurality of flow cell images is acquired at a corresponding position along an axial axis; Generating, by the processor, a first plurality of processed images corresponding to the first plurality of flow cell images; Filtering, by the processor, the first plurality of flow cell images based on the first plurality of processed images to generate a first plurality of filtered images; Obtaining, by the processor, a 3D community map; Extracting, by the processor, image intensities of communities from one of the following based on the 3D community map: A second plurality of flow cell images; A second plurality of processed images; A second plurality of filtered images; or A combination thereof; And Performing, by the processor, base calling based on the extracted image intensities of the communities.
156. A computer-implemented method for base calling in sequencing data analysis, the computer-implemented method including: Obtaining, by a processor, a first plurality of flow cell images of a sample, wherein each of the first plurality of flow cell images or the second plurality of flow cell images is acquired at a corresponding position along an axial axis; Filtering, by the processor, the first plurality of flow cell images or the second plurality of flow cell images to generate a plurality of filtered images; Obtaining, by the processor, a 3D community map based on the filtered images of the first plurality of flow cell images or the second plurality of flow cell images; Extracting, by the processor, image intensities of communities from one of the following based on the 3D community map: The first plurality of flow cell images or the second plurality of flow cell images; The processed image of the first plurality of flow cell images or the second plurality of flow cell images; The filtered image of the first plurality of flow cell images or the second plurality of flow cell images; Or A combination thereof; And The processor performs base identification based on the extracted image intensities of the community.
157. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Performing image registration on: The first plurality of flow cell images; The first plurality of processed images; The first plurality of filtered images; Or A combination thereof.
158. The computer-implemented method according to any one of the preceding claims, wherein performing image registration includes: The first plurality of flow cell images; The first plurality of processed images; The first plurality of filtered images; or A combination thereof is registered with one or more template images.
159. The computer-implemented method according to any one of the preceding claims, wherein registering the first plurality of flow cell images; the first plurality of processed images; the first plurality of filtered images; or a combination thereof with one or more template images includes: The processor generates the one or more template images in a reference coordinate system.
160. The computer-implemented method according to any one of the preceding claims, wherein performing image registration includes: The processor clusters the first plurality of flow cell images; The first plurality of processed images; The first plurality of filtered images; Or A combination thereof is registered with the template clusters in the one or more template images.
161. The computer-implemented method according to any one of the preceding claims, wherein generating the one or more template images in the reference coordinate system includes: Using the coordinates of the community to register the community in the one or more reference cycles with the one or more template images.
162. The computer-implemented method according to any one of the preceding claims, wherein the coordinates of the community include 2D coordinates of the community.
163. The computer-implemented method according to any one of the preceding claims, wherein the coordinates of the community include z positions of the community.
164. The computer-implemented method according to any one of the preceding claims, wherein performing image registration includes: The processor determines a plurality of transforms based on the one or more template images, each transform in the plurality of transforms corresponding to a corresponding sub-block of the first plurality of flow cell images, the first plurality of processed images, or the first plurality of filtered images, and configured to register the sub-block with the one or more template images; And Using the plurality of transforms to register the sub-blocks with the one or more template images.
165. The computer-implemented method according to any one of the preceding claims, wherein each transform in the plurality of transforms corresponds to a corresponding image of: The first plurality of flow cell images; The first plurality of processed images; or The first plurality of filtered images.
166. The computer-implemented method according to any one of the preceding claims, wherein the plurality of transformations includes one or more affine transformations.
167. The computer-implemented method according to any one of the preceding claims, wherein each transformation of the plurality of transformations includes an affine transformation.
168. The computer-implemented method according to any one of the preceding claims, wherein performing base identification based on the extracted image intensities of the colony includes: Performing base identification based on the extracted image intensities of the colony from the first plurality of filtered images.
169. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Performing image registration on the colony of the first plurality of filtered images based on fiducial markers.
170. The computer-implemented method according to any one of the preceding claims, wherein the fiducial markers are located on the flow cell.
171. The computer-implemented method according to any one of the preceding claims, wherein the fiducial markers are outside the flow cell.
172. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired at 2, 3, 4, 5, 6, 7, 8, 9, or 10 different positions along the axial axis.
173. The computer-implemented method according to any one of the preceding claims, further comprising: Generating the 3D colony map based on the first plurality of filtered images.
174. The computer-implemented method according to any one of the preceding claims, wherein generating the 3D colony map based on the first plurality of filtered images includes: Generating the 3D colony map based on the one or more template images.
175. The computer-implemented method according to any one of the preceding claims, wherein the one or more template images are in 2D.
176. The computer-implemented method according to any one of the preceding claims, wherein each of the one or more template images corresponds to a corresponding flow cell image of the first plurality of flow cell images at the corresponding position along the axial axis.
177. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images, the first plurality of processed images, and the first plurality of filtered images are from the one or more reference cycles and different channels.
178. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of flow cell images, the second plurality of processed images, and the second plurality of filtered images are from the one or more reference cycles and different channels.
179. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of flow cell images, the second plurality of processed images, and the second plurality of filtered images are from one or more cycles different from the one or more reference cycles and different channels.
180. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images and the second plurality of flow cell images are the same, the first plurality of processed images and the second plurality of processed images are the same, and the first plurality of filtered images and the second plurality of filtered images are the same.
181. The computer-implemented method according to any one of the preceding claims, wherein performing base identification based on the extracted image intensity of the community is within a cycle different from the one or more reference cycles.
182. The computer-implemented method according to any one of the preceding claims, wherein generating the 3D community map based on the one or more template images comprises: extracting communities in the one or more template images; and removing duplicate communities from the extracted communities.
183. The computer-implemented method according to any one of the preceding claims, wherein generating the 3D community map based on the one or more template images comprises: combining the one or more template images into a candidate 3D community map; removing duplicate communities from the candidate 3D community map.
184. The computer-implemented method according to any one of the preceding claims, wherein removing the duplicate communities comprises: performing preliminary base identification based on the one or more template images; and repeating removing the duplicate communities until a stop criterion is met, including: identifying candidate communities with the same base identification; determining the 3D distance between two of the candidate communities; and in response to determining that the 3D distance between the two communities is within a predetermined distance threshold: determining the image intensity of the two communities from the first plurality of filtered images; removing the community with the smaller image intensity among the two communities.
185. The computer-implemented method according to any one of the preceding claims, wherein the predetermined distance threshold is based on the depth of field of the optical system, the distance between two adjacent flow cell images along the axial direction, or a combination thereof.
186. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired by an NGS sequencing system.
187. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired in a cycle different from the reference cycle.
188. The computer-implemented method according to any one of the preceding claims, wherein the sample is an in-situ sample located on a flow cell.
189. The computer-implemented method according to any one of the preceding claims, wherein the in-situ sample comprises one or more cells or tissues.
190. The computer-implemented method according to any one of the preceding claims, wherein each flow cell image in the first plurality of flow cell images includes a field of view orthogonal to the axial axis.
191. The computer-implemented method according to any one of the preceding claims, wherein the field of view of each flow cell image in the first plurality of flow cell images is the same.
192. The computer-implemented method according to any one of the preceding claims, wherein the field of view of each flow cell image in the first plurality of flow cell images covers at least a portion of a tile of the flow cell.
193. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images have the same image resolution.
194. The computer-implemented method according to any one of the preceding claims, wherein the axial axis extends from the objective lens to a sample located on the flow cell, and the flow cell is positioned on the sequencing system.
195. The computer-implemented method according to any one of the preceding claims, wherein the axial axis is orthogonal to the image plane, and wherein the field of view is within the image plane.
196. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images comprises: selecting a kernel; and generating the first plurality of processed images by performing an opening operation on the first plurality of flow cell images using the selected kernel.
197. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images comprises: selecting a kernel; and generating the first plurality of processed images by convolving the first plurality of flow cell images with the selected kernel.
198. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images further comprises: selecting a first kernel and a second kernel; generating a first blurred image by convolving the first plurality of flow cell images with the first kernel; and generating a second blurred image by convolving the first plurality of flow cell images with the second kernel.
199. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images comprises: scaling the first plurality of processed images.
200. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images comprises: scaling the first blurred image, the second blurred image, or both.
201. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images based on the first plurality of processed images comprises: subtracting the second blurred image from the first blurred image, thereby generating the first plurality of filtered images.
202. The computer-implemented method according to any one of the preceding claims, wherein the kernel is 2x2, 3x3, 4x4, 5x5, 6x6 pixels.
203. The computer-implemented method according to any one of the preceding claims, wherein the kernel is a circular kernel.
204. The computer-implemented method according to any one of the preceding claims, wherein the kernel is a Gaussian kernel. The computer-implemented method according to any one of the preceding claims, wherein the first kernel and the second kernel are different Gaussian kernels. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images based on the first plurality of processed images comprises: Subtracting each processed image in the first plurality of processed images from the corresponding flow cell image in the first plurality of flow cell images, thereby generating the first plurality of filtered images. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images based on the first plurality of processed images further comprises: Adding a predetermined offset to the subtracted images, thereby generating the first plurality of filtered images. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Registering the first plurality of flow cell images, the first plurality of processed images, the first plurality of filtered images, or a combination thereof with one or more images of the sample. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise staining of: a membrane, a cell nucleus, or a combination thereof. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise staining of one or more membrane proteins. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise staining of lipids. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise a fluorescence signal from a cell membrane. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise segmentation of: cells, membranes, cell nuclei, or a combination thereof. The computer-implemented method according to any one of the preceding claims, wherein performing base calling based on the extracted image intensities of the population comprises: Performing one or more primary analysis steps to adjust the image intensities of the population in: The first plurality of flow cell images; The first plurality of processed images; The first plurality of filtered images; Or A combination thereof; And Performing base calling on the population based on the adjusted image intensities, Wherein the one or more primary analysis steps comprise: Background subtraction; Image sharpening; Intensity offset adjustment; Color correction; Intensity normalization; Phasing and pre-phasing correction; Image registration; Quality score estimation; or A combination thereof. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired at the same tile or sub-tile of a flow cell. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired at one or more reference cycles.
217. The computer-implemented method according to any one of the preceding claims, wherein two adjacent positions along the axial axis are separated by about 1um, 2um, 3um, 4um, 5um, 6um, 7um, 8um, 9um, 10um, 11um or 12um.
218. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired from 1, 2, 3, 4, 5 or 6 channels.
219. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more processing units; one or more integrated circuits; or a combination thereof.
220. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); or a combination thereof.
221. The computer-implemented method according to any one of the preceding claims, further comprising: transmitting the base identification to a processing unit by the processor.
222. The computer-implemented method according to any one of the preceding claims, wherein the processing unit is a central processing unit (CPU).
223. The computer-implemented method according to any one of the preceding claims, wherein the processing unit is configured to register the base identification with one or more images.
224. The computer-implemented method according to any one of the preceding claims, wherein the 3D community map comprises a list of 3D coordinates, each entry in the list of 3D coordinates corresponding to the 3D position of the community of the sample.
225. A computer-implemented system for base identification in sequencing data analysis, the computer-implemented system comprising: one or more hardware processors; one or more data storage devices that store instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including: obtaining, by the processor, a first plurality of flow cell images of a sample, wherein each flow cell image of the first plurality of flow cell images is acquired at a corresponding position along an axial axis; generating, by the processor, a first plurality of processed images corresponding to the first plurality of flow cell images; filtering, by the processor, the first plurality of flow cell images based on the first plurality of processed images, thereby generating a first plurality of filtered images; obtaining, by the processor, a 3D community map; extracting, by the processor, image intensities of the community from one of: a second plurality of flow cell images; a second plurality of processed images; a second plurality of filtered images; or a combination thereof; and performing base identification by the processor based on the extracted image intensities of the community.
226. A computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images or the second plurality of flow cell images to generate filtered images of the first plurality of flow cell images or the second plurality of flow cell images comprises: Performing 3D deconvolution on the first plurality of flow cell images or the second plurality of flow cell images.
227. A computer-implemented system for base calling in sequencing data analysis, the computer-implemented system comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations comprising: Obtaining, by the processor, a first plurality of flow cell images of a sample, wherein each of the first plurality of flow cell images or the second plurality of flow cell images is acquired at a corresponding position along an axial axis; Filtering, by the processor, the first plurality of flow cell images or the second plurality of flow cell images to generate a plurality of filtered images; Obtaining, by the processor, a 3D colony map based on the filtered images of the first plurality of flow cell images or the second plurality of flow cell images; Extracting, by the processor, image intensities of colonies from one of the following based on the 3D colony map: The first plurality of flow cell images or the second plurality of flow cell images; Processed images of the first plurality of flow cell images or the second plurality of flow cell images; Filtered images of the first plurality of flow cell images or the second plurality of flow cell images; or combinations thereof; and Performing base calling by the processor based on the extracted image intensities of the colonies.
228. A computer-implemented system for base calling in sequencing data analysis, the computer-implemented system comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations comprising any one of the preceding claims.
229. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations for base calling in sequencing data analysis, the operations comprising: Obtaining, by the processor, a first plurality of flow cell images of a sample, wherein each of the first plurality of flow cell images or the second plurality of flow cell images is acquired at a corresponding position along an axial axis; Filtering, by the processor, the first plurality of flow cell images or the second plurality of flow cell images to generate a plurality of filtered images; Obtaining, by the processor, a 3D colony map based on the filtered images of the first plurality of flow cell images or the second plurality of flow cell images; Extracting, by the processor, image intensities of colonies from one of the following based on the 3D colony map: The first plurality of flow cell images or the second plurality of flow cell images; Processed images of the first plurality of flow cell images or the second plurality of flow cell images; The filtered image of the first plurality of flow cell images or the second plurality of flow cell images; or a combination thereof; and base calling is performed by the processor based on the extracted image intensities of the community.
230. One or more non-transitory computer storage media encoded with instructions that are executable by one or more hardware processors to perform operations including any one of the preceding claims.
231. A computer-implemented method according to any one of the preceding claims, wherein the 3D community map comprises a list of 3D coordinates, each entry of the list of 3D coordinates corresponding to a 3D position of the community of the sample.
Citation Information
Patent Citations
Method and system for sequencing nucleic acids
US10246744B2
Engineered polymerases for improved sequencing
US10731141B2
Improvement in drawers
US133138A
DNA sequencing method using acyclonucleoside triphosphates
US5558991A
Punching-machine.
US682686A