Three-dimensional base recognition in next generation sequencing analysis
By applying filtering and image registration technology in the flow cell image stack of three-dimensional samples, the interference problem of defocused polymerase community and background signals in three-dimensional samples is solved, and accurate base recognition of high-density and unbalanced nucleotide diversity samples is achieved, improving sequencing efficiency and accuracy.
Patent Information
- Application Number
- CN202380084133.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-06
- Filing Date
- 2023-10-05
- Publication Date
- 2025-07-18
AI Technical Summary
When the prior art base recognition of three-dimensional samples, it is difficult to effectively remove background signal interference from defocused polymerase communities and cell components, resulting in inaccurate and unreliable base recognition, especially in the case of high density and unbalanced nucleotide diversity.
By acquiring flow cell image stacks at multiple axial positions, filtering and image registration technology are used to remove defocused polymerase communities and background signals, and an accurate 3D polymerase community map is generated to achieve accurate recognition and base recognition of the focused polymerase community.
It improves the accuracy and reliability of base recognition of three-dimensional samples, can process high-density and unbalanced nucleotide diversity samples, reduces calculation load and storage space requirements, and improves sequencing efficiency.
Smart Images

Figure CN120345030A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 413,864, filed on October 6, 2022, which is hereby incorporated by reference in its entirety. Field of the Invention
[0003] The present disclosure generally relates to base calling in DNA sequencing data analysis, and particularly to three-dimensional (3D) base calling. Background Art
[0004] In next-generation sequencing (NGS) or NGS-like applications such as sequencing-by-synthesis, sequencing-by-ligation, or affinity sequencing, to identify the sequence of a target nucleic acid, new strands are synthesized one nucleotide base at a time. During each sequencing cycle, one base is attached to any given strand. During the imaging step of each cycle, an image is recorded. A base calling algorithm is applied to the image to "read" the consecutive signals from each cluster or polymerase colony, and convert the optical signals into an identification of the nucleotide base sequence added to each DNA fragment. Traditional base calling relies on two-dimensional (2D) flow cell images. In sequencing analysis of in-situ samples such as cells or tissues, the sample has a thickness in the z-direction orthogonal to the image plane. Thus, the flow cell image at a selected z-level may include signals from out-of-focus polymerase colonies located at adjacent z-levels and other unwanted signals, e.g., signals from cell membranes. Three-dimensional (3D) base calling is needed to ensure accurate base calling and sequencing analysis of 3D samples such as cells and tissues. Summary of the Invention
[0005] Embodiments of systems, devices, methods, and / or computer program products and / or combinations and sub-combinations thereof are disclosed herein that are capable of performing 3D base calling using flow cell images of a sample such as an in-situ cell or tissue. The flow cell images may be from different sequencing cycles and / or different channels. The flow cell images may be from traditional two-dimensional samples or in-situ samples. The flow cell images may be from samples with unbalanced nucleotide diversity.
[0006] As such a specific application, embodiments of methods, systems, and media for 3D base calling of flow cell images enable accurate base calling of clusters or polymerase colonies to depend on image intensity, position, and / or size.
[0007] Other embodiments of these aspects include corresponding computer systems, apparatus, and computer program products recorded on computer storage devices, which are configured, individually or in combination, to perform the actions of the methods. For a computer system that has been configured or is to be configured to perform operations or actions, the computer system has installed thereon software, firmware, hardware, or a combination thereof that, in operation, causes the computer system to perform the operations or actions. For a computer program product that has been configured or is to be configured to perform operations or actions, the computer program product includes instructions that, when executed by a hardware processor, cause the hardware processor to perform the operations or actions.
[0008] Additional embodiments, features, and advantages of the present disclosure, as well as the structure and operations of the embodiments of the present disclosure, are described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings incorporated herein and forming a part of this specification illustrate embodiments of the present disclosure and, together with the specification, further serve to explain the principles of the present disclosure and enable one of ordinary skill in the art to make and use the embodiments.
[0010] Figure 1 A block diagram of a system for performing 3D base identification of a flow cell image according to some embodiments is shown.
[0011] Figures 2A - 2C Exemplary flow cell images, processed images, and filtered images for 3D base identification according to some embodiments are shown.
[0012] Figures 3A - 3C Exemplary flow cell images, processed images, and filtered images for 3D base identification according to some embodiments are shown.
[0013] Figures 3D - 3E Exemplary flow cell images for 3D base identification according to some embodiments and their corresponding filtered images are shown.
[0014] Figure 4 A block diagram of a computer system for performing sequencing analysis and / or base identification according to some embodiments is shown.
[0015] Figure 5 Exemplary projection images of flow cell images taken at different axial positions of a 3D sample according to some embodiments are shown.
[0016] Figure 6A A flowchart of an exemplary method for performing 3D base identification of a flow cell image according to some embodiments is shown.
[0017] Figure 6BFlowchart of an exemplary method for performing 3D base identification of a flow cell image according to some embodiments.
[0018] Figures 7A - 7B Exemplary registration of a sequencing image of a polymerase community and a cell staining image is shown.
[0019] Figure 8A Schematic diagrams of a flow cell image, sub-tiles, and regions of the flow cell image with a polymerase community according to some embodiments are shown.
[0020] Figure 8B Schematic diagram of a portion of a flow cell with multiple tiles according to some embodiments is shown.
[0021] Figure 9 Flowchart of a method for performing image registration of a flow cell image according to some embodiments is shown.
[0022] Figures 10A - 10B Schematic diagrams of image transformation and corresponding 2D displacements according to some embodiments are shown.
[0023] Figure 11 Schematic diagram showing an exemplary linear single-stranded library molecule (1100), which includes: a surface pinned primer binding site (1120); an optional left unique identification sequence (1180); a left index sequence (1160); a forward sequencing primer binding site (1140); an insertion region (1110) with a target sequence; a reverse sequencing primer binding site (1150); a right index sequence (1170); and a surface capture primer binding site (1130).
[0024] Figure 12 Schematic diagram showing an exemplary linear single-stranded library molecule (1100), which includes: a surface pinned primer binding site (1120); a left index sequence (1160); a forward sequencing primer binding site (1140); an insertion region (1110) with a target sequence; a reverse sequencing primer binding site (1150); a right index sequence (1170); an optional right unique identification sequence (1190); and a surface capture primer binding site (1130).
[0025] Figure 13 Schematic diagram showing an exemplary embodiment of a padlock probe.
[0026] Figure 14It is a schematic diagram showing the workflow of generating circular padlock probes inside cells. The workflow includes generating first and second cDNAs from first and second target RNA molecules respectively, and hybridizing first and second padlock probes with first and second cDNA molecules respectively to generate first and second circular padlock probes. The first padlock probe includes (i) a first target barcode sequence (target BC-1) that uniquely identifies the first target RNA or the first target cDNA, (ii) a first sequencing primer binding site (or its complementary sequence), (iii) a universal binding site for amplification primers (universal RCA) (or its complementary sequence), and (iv) a universal binding site for compaction oligonucleotides (or its complementary sequence). The second padlock probe includes (i) a second target barcode sequence (target BC-2) that uniquely identifies the second target RNA or the second target cDNA, (ii) a second sequencing primer binding site (or its complementary sequence), (iii) a universal binding site for amplification primers (universal RCA) (or its complementary sequence), and (iv) a universal binding site for compaction oligonucleotides (or its complementary sequence).
[0027] Figure 15 It is a schematic diagram showing the workflow of rolling circle and sequencing inside cells. The workflow includes generating first and second concatemers by performing rolling circle amplification using first and second covalently closed circular molecules respectively. A sequencing workflow is performed on the first and second concatemers using a universal sequencing primer, a sequencing polymerase, and a plurality of nucleotide reagents.
[0028] Figure 16 It is a schematic diagram showing an exemplary workflow for sequencing concatemers generated inside cells. The concatemer contains tandem repeat units, where each unit contains: (i) a universal sequencing primer binding site (Seq), (ii) a universal compaction oligonucleotide binding site (CO), (iii) an insert sequence corresponding to a given target cDNA, and (iv) a target barcode sequence (BC) corresponding to a given target cDNA.
[0029] Figure 17 It is a schematic diagram showing an exemplary workflow for sequencing concatemers generated inside cells. The concatemer contains tandem repeat units, where each unit contains: (i) a universal sequencing primer binding site (Seq), (ii) a universal compaction oligonucleotide binding site (CO), (iii) an insert sequence corresponding to a given target cDNA, and (iv) a target barcode sequence (BC) corresponding to a given target cDNA.
[0030] Figure 18A schematic diagram showing an exemplary workflow for sequencing concatemers generated inside cells. A concatemer contains tandem repeat units, where each unit contains: (i) a universal sequencing primer binding site (Seq), (ii) a universal compaction oligonucleotide binding site (CO), and (iii) an insert sequence corresponding to a given target cDNA.
[0031] Figure 19 A schematic diagram showing an exemplary workflow for sequencing concatemers generated inside cells. A concatemer contains tandem repeat units, where each unit includes: (i) a universal sequencing primer binding site (Seq) and (ii) an insert sequence corresponding to a given target cDNA.
[0032] Figure 20 A schematic diagram showing a workflow for generating circularized padlock probes, which includes generating first and second cDNAs from first and second target RNA molecules (respectively), and hybridizing first and second padlock probes (respectively) to the first and second cDNA molecules to generate first and second circularized padlock probes. The first padlock probe includes (i) a first target barcode sequence (target BC-1) that uniquely identifies the first target RNA, (ii) a first batch-specific sequencing primer binding site (batch Seq-1) (or its complementary sequence), (iii) a universal binding site for amplification primers (universal RCA) (or its complementary sequence), and (iv) a universal binding site for compaction oligonucleotides (or its complementary sequence). The second padlock probe includes (i) a second target barcode sequence (target BC-2) that uniquely identifies the second target RNA, (ii) a second batch-specific sequencing primer binding site (batch Seq-2) (or its complementary sequence), (iii) a universal binding site for amplification primers (universal RCA) (or its complementary sequence), and (iv) a universal binding site for compaction oligonucleotides (or its complementary sequence).
[0033] Figure 21 A schematic diagram showing a rolling circle and sequencing workflow, which includes generating first and second concatemers by rolling circle amplification using first and second covalently closed circular molecules (respectively). A first sequencing workflow is performed on the first and second concatemers using a first batch-specific sequencing primer, a sequencing polymerase, and a plurality of nucleotide reagents. The first concatemer is sequenced repeatedly, but the second concatemer is not. A second sequencing workflow is performed on the first and second concatemers using a second batch-specific sequencing primer, a sequencing polymerase, and a plurality of nucleotide reagents. The second concatemer is sequenced repeatedly, but the first concatemer is not.
[0034] Figure 22Schematic of an exemplary low-binding support that includes alternating layers of a glass substrate and a hydrophilic coating covalently or non-covalently adhered to the glass, and the support further includes chemically reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides). In alternative embodiments, the support can be made of any material such as glass, plastic, or polymeric materials.
[0035] Figure 23 Schematic of various exemplary configurations of multivalent molecules. Left panel (Class I): Schematic of a multivalent molecule with a “starburst” or “helter-skelter” configuration. Middle panel (Class II): Schematic of a multivalent molecule with a dendrimer configuration. Right panel (Class III): Schematic of multiple multivalent molecules formed by the reaction of streptavidin with 4-arm or 8-arm PEG-NHS having biotin and dNTP. Nucleotide units are designated as ‘N’, biotin is designated as ‘B’, and streptavidin is designated as ‘SA’.
[0036] Figure 24 Schematic of an exemplary multivalent molecule that includes a universal core attached to multiple nucleotide arms.
[0037] Figure 25 Schematic of an exemplary multivalent molecule that includes a dendrimeric core attached to multiple nucleotide arms.
[0038] Figure 26 Shows a schematic of an exemplary multivalent molecule that includes a core attached to multiple nucleotide arms, where the nucleotide arms include biotin, spacer, linker, and nucleotide units.
[0039] Figure 27 Schematic of an exemplary nucleotide arm that includes a core attachment portion, spacer, linker, and nucleotide unit.
[0040] Figure 28 Shows the chemical structure of an exemplary spacer (top), and the chemical structures of various exemplary linkers, including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker (bottom).
[0041] Figure 29 Shows the chemical structures of various exemplary linkers, including Linkers 1 to 9.
[0042] Figure 30A Shows the chemical structures of various exemplary linkers that are joined / attached to nucleotide units.
[0043] Figure 30B Shows the chemical structures of various exemplary linkers that are joined / attached to nucleotide units.
[0044] Figure 30C Shows the chemical structures of various exemplary linkers conjugated / attached to nucleotide units.
[0045] Figure 31 Shows the chemical structure of an exemplary biotinylated nucleotide arm. In this example, the nucleotide unit is attached to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base.
[0046] Figure 32 Is a schematic diagram of a guanine tetrad (e.g., G-tetrad).
[0047] Figure 33 Is a schematic diagram of an exemplary intramolecular G-quadruplex structure.
[0048] In the drawings, like reference numerals generally denote the same or similar elements. Additionally, generally, the leftmost digit of a reference numeral can identify the drawing in which the reference numeral first appears. Detailed Description
[0049] Embodiments of systems, devices, methods, and / or computer program products and / or combinations and sub-combinations thereof are provided herein that are capable of performing base identification using flow cell images obtained from three-dimensional (3D) samples such as in situ cells or tissues. The 3D base identification techniques herein can be used for flow cell images obtained from various imaging and / or sequencing techniques. The techniques disclosed herein can be used for base identification in next-generation sequencing, and base identification will be used as the main example for describing the applications of these techniques herein. However, such image analysis techniques can also be used in other applications using point detection and / or CCD imaging.
[0050] For traditional DNA sequencing, the optical system can be adjusted to focus on clusters or polymerase colonies of a two-dimensional (2D) sample. The flow cell image can display the clusters or polymerase colonies as 2D bright spots. Base calling can be performed using the corresponding image intensity of the bright spots. However, the thickness of an in situ sample such as a cell or tissue along the axial axis, i.e., the z direction, may not be in focus in a single 2D image. Thus, a stack of multiple 2D flow cell images at different axial positions can be acquired to cover the clusters or polymerase colonies of the in situ sample. Interferences may occur in the stack of flow cell images, such as interference from out-of-focus polymerase colonies and cell components (such as cell membranes) in the background signal. For example, a polymerase colony located at a first axial position can appear in the first flow cell image, and it may also generate a signal spot in a second 2D flow cell image taken at its out-of-focus adjacent axial position. The signal spot may interfere with the intensity of the polymerase colony at or near the same x-y position in the second flow cell image, thereby reducing the accuracy and reliability of base calling. The techniques disclosed herein can be configured to process a stack of flow cell images of a 3D sample and generate accurate and reliable image intensities for polymerase colonies or clusters, thereby enabling accurate and reliable base calling for 3D samples. Existing algorithms for processing image intensities from volumetric 3D samples may have various drawbacks. For example, flattening the stack of images without removing signal interference from large background components such as membranes or cytosol may result in unreliable image intensities and inaccurate base calling. Additionally, after some 3D-to-2D flattening methods, out-of-focus polymerase colonies or clusters may remain and cause incorrect base calling. In some embodiments, flattening the stack of 2D flow cell images into a single 2D image, for example by projection, may result in the loss of polymerase colonies or clusters that are in focus but whose intensity is mixed into out-of-focus polymerase colonies. Furthermore, existing sequencing and analysis methods for 3D samples may not provide sufficient resolution along the z axis, and may not be able to sequence and analyze when the 3D sample has a high density (e.g., 2 times, 4 times, 5 times or higher than the density that can be processed by traditional sequencing methods at a pre-determined quality (e.g., Q30, Q35 or Q40)) and / or unbalanced nucleotide diversity. Thus, there is a need to generate accurate and reliable image intensities for polymerase colonies or clusters from 3D volume samples such that such image intensities can be used for accurate and reliable 3D base calling.
[0051] In some embodiments, the techniques disclosed herein advantageously filter flow cell images before flattening an axial stack of 2D images into a single 2D image such that out-of-focus polymerase colonies or clusters can be effectively removed without affecting in-focus polymerase colonies or clusters. Flattening of the stack of flow cell images herein facilitates finding the intensity at which each polymerase colony or cluster is in focus. The techniques disclosed herein advantageously utilize images that retain background information for accurate and effective registration of polymerase colonies or clusters relative to cellular components (e.g., the nucleus) in a cell.
[0052] In some embodiments, the techniques disclosed herein advantageously generate 3D polymerase colony maps. The techniques herein can effectively filter out-of-focus polymerase colonies and background objects without affecting in-focus polymerase colonies or clusters in the 3D polymerase colony map. The 3D polymerase colony map can be used to extract polymerase colony intensities for base calling. Compared to a single flattened 2D image, the 3D polymerase colony map advantageously retains information on polymerase colonies and clusters that may be removed in the flattened image. The 3D polymerase colony map can be generated in several early process cycles and used in subsequent process cycles such that the additional computational load of recalculating new polymerase colony maps and the storage space for saving them can be minimized. The techniques disclosed herein advantageously utilize images that retain background information for accurate and effective registration of polymerase colonies or clusters to a cell, and such background information can facilitate sequencing analysis by providing spatial information on polymerase colonies or clusters relative to cellular components. Additionally, the techniques disclosed herein advantageously remove duplicate polymerase colonies and resolve polymerase colonies that may partially overlap, which may lead to errors in accurate and reliable 3D base calling.
[0053] In DNA sequencing, identifying the center of a cluster or polymerase colony is sometimes part of a preliminary analysis. The preliminary analysis can include some or all of the operations and / or steps required to perform base calling and calculate a quality score for the base calling. The preliminary analysis can involve forming a template image for at least a portion of the flow cell. The template image can include the estimated positions of all detected clusters or polymerase colonies in a common coordinate system. The template image is generated by identifying the positions of clusters or polymerase colonies in all images in the first few cycles of the sequencing process.
[0054] Sequencing system
[0055] In some embodiments, sequencing and sequencing analysis of a sample are performed using the computer-implemented systems herein. Figure 1FIG. 0 shows a block diagram of a computer-implemented system 100 according to one or more embodiments disclosed herein. System 100 has a sequencing system 110 that includes a flow cell 112, a sequencer 114, an imager 116, a data memory 122, and a user interface 124. Sequencing system 110 may be connected to a cloud 130. Sequencing system 110 may include one or more of the following: a dedicated processor 118, a field programmable gate array (FPGA) 120, and a computer system 126.
[0056] In some embodiments, the flow cell 112 is configured to capture DNA fragments and form a DNA sequence for base identification on the flow cell. The flow cell 112 may include a support as described herein. The support may be a solid-phase support. As disclosed herein, the support may include a surface coating thereon. The surface coating may be a polymer coating as disclosed herein.
[0057] The flow cell 112 may include a plurality of tiles or imaging regions thereon, and each tile may be divided into a grid of sub-tiles. Each sub-tile may include a plurality of clusters or polymerase colonies thereon. As a non-limiting example, the flow cell may have 424 tiles, and each tile may be divided into a 6x9 grid, thus having 54 sub-tiles. A flow cell image as disclosed herein may be an image of signals including a plurality of clusters or polymerase colonies. A flow cell image may include one or more signal tiles or one or more signal sub-tiles. In some embodiments, a flow cell image may be an image including all tiles and substantially all signals thereon. A flow cell image may be acquired from a channel using imager 116 during an imaging or sequencing cycle. In some embodiments, each tile may include millions of polymerase colonies or clusters. As a non-limiting example, a tile may include from about 1 to 10 million clusters or polymerase colonies. Each polymerase colony may be a collection of many copies of DNA fragments.
[0058] In an embodiment of sequencing a three-dimensional (3D) sample, such as cells or tissue immobilized on a flow cell, flow cell images can be acquired at multiple z-levels that are orthogonal to the image plane of the flow cell image to cover the volume of the 3D sample. The z-axis can extend from the objective lens of the optical system disclosed herein to a support, such as a flow cell device. Each z-level of the flow cell image can be parallel to and separated from an adjacent z-level by a predetermined distance, such as from about 0.1 um to about 15 um. Each z-level can include a predetermined thickness. The thickness can be in the range of 0.01 um to 5 um. In some embodiments, the thickness can be determined such that the pixels have an isotropic size in the x, y, and z directions. In other words, the pixels or voxels are cubes. Each flow cell image can include a thickness of 0.01 um to 0.9 um, such as the depth of focus. In some embodiments, each flow cell image can include a thickness of 0.05 um to 0.5 um. In some embodiments, each flow cell image can include a thickness of 0.1 um to 0.3 um.
[0059] The flow cell images at each z-level can be separated from adjacent levels by 0.01 um to 10 um, such as between the centers of the flow cell images at adjacent levels. Each z-level of the flow cell image can be separated from the center of an adjacent level by 0.1 um to 5 um at its center. In some embodiments, the number of z-levels is predetermined to allow coverage of some or all of the 3D volume of the sample extending along the z-axis. For example, for a sample with a thickness of 10 um, 10, 11, 12, or more z-levels with a thickness of about 1 um and spaced 1 um apart from each other can be used to cover the sample along the z-axis without overlapping coverage along the z-axis. There may be no gap between the thicknesses of the flow cell images at adjacent z-levels. As another example, for a sample with a thickness of 10 um, 20, 21, or 22 z-levels with a thickness of 0.5 um and spaced 0.6 um apart from each other can be used to cover the sample along the z-axis, with a gap of 0.1 um between the flow cell images of adjacent z-levels.
[0060] At each z-level, the flow cell image can be acquired from one or more sequencing cycles and / or one or more channels. Each flow cell image can include at least a portion of one or more tiles or sub-tiles of the flow cell in its field of view. Figure 8B A portion of a flow cell 112 with multiple tiles 210 is shown. The image plane is defined by the x-axis and the y-axis. And the z-axis is orthogonal to the x-y plane. Although the flow cell image, the sample, and the z-axis are described in a Cartesian coordinate system, any other coordinate system can be used to define the spatial positions and relationships of the polymerase colonies or clusters and their images herein. Other coordinate systems can include, but are not limited to, polar coordinate systems, cylindrical coordinate systems, or spherical coordinate systems.
[0061] Sequencer 114 can be configured to flow a nucleotide mixture onto flow cell 112, cleave blockers from the nucleotides between flow steps, and perform other steps for forming a DNA sequence on flow cell 112. The nucleotides can have attached fluorescent elements that emit light or energy at wavelengths indicative of the nucleotide type. Each type of fluorescent element can correspond to a specific nucleobase (e.g., A, G, C, T). The fluorescent elements can emit light at visible wavelengths. In some embodiments, sequencer 114 and flow cell 112 can be configured to perform the various sequencing methods disclosed herein, e.g., affinity sequencing.
[0062] For example, each nucleobase can be assigned a color. Different types of nucleotides can have different colors. For example, adenine (A) can be red, cytosine (C) can be blue, guanine (G) can be green, and thymine (T) can be yellow. The color or wavelength of the fluorescent element for each nucleotide can be chosen such that the nucleotides can be distinguished from one another based on the wavelength of the light emitted by the fluorescent element.
[0063] Imager 116 can be configured to capture an image of flow cell 112 after each flow step. In one embodiment, imager 116 is a camera configured to capture digital images, such as a CMOS or CCD camera. The camera can be configured to capture an image of the wavelength of the fluorescent element bound to the nucleotide. These images can be referred to as flow cell images.
[0064] In some embodiments, imager 116 can include one or more of the optical systems disclosed herein. The optical system can be configured to capture an optical signal from the flow cell and generate a corresponding digital image. The digital image can then be used for base calling.
[0065] In an embodiment, images of the flow cell can be captured in groups, where each image in the group is taken at a wavelength or spectrum that matches or includes only one of the fluorescent elements. In another embodiment, the image can be captured as a single image that captures all wavelengths of the fluorescent elements.
[0066] The resolution of imager 116 can control the level of detail in the flow cell image, including pixel size. In existing systems, this resolution is very important because it controls the accuracy of the dot-finding algorithm in identifying the centers of polymerase colonies. In some embodiments, the image resolution of the flow cell images disclosed herein can be from about 10 nanometers (nm) to several hundred nm or higher. In some embodiments, the image resolution of the flow cell images disclosed herein can be from about 10 nanometers (nm) to several micrometers or higher. One way to improve the accuracy of dot finding is to improve the resolution of imager 116 or the processing of the images captured by imager 116. It is possible to perform detection of polymerase colony centers in pixels other than those detected by the dot-finding algorithm. These methods can allow for an increase in the accuracy of polymerase colony center detection without increasing the resolution of imager 116. The resolution of the imager can even be lower than that of existing systems with comparable performance, which can reduce the cost of sequencing system 110.
[0067] The image quality of the flow cell image can control the base calling quality. One way to improve the accuracy of base calling is to improve imager 116 or the processing of the images captured by imager 116 to obtain better image quality.
[0068] The methods described herein are configured to register flow cell images in a common coordinate system such that base calling with respect to clusters or polymerase colonies will be more accurate than when such registration is not performed. These methods can enable accurate and efficient base calling.
[0069] The methods herein can be advantageously executed in parallel in the computer-implemented system 100 without interfering with or delaying the existing sequencing workflow of system 100. After a flow cell image is acquired in a particular cycle, image registration and other processing of such flow cell images can be performed while the sequencing of the current cycle or a subsequent cycle is in progress. This parallel execution of image processing and base calling operations can advantageously speed up the sequencing analysis process and reduce the total time and corresponding time of sequencing. Base calling can also be performed while the sequencing of the current cycle or a subsequent cycle is in progress. Additionally, some or all of the operations disclosed herein can be advantageously performed by an FPGA and / or NPU, and data can be transferred between the CPU and the FPGA and / or NPU to reduce the total operation time of methods that do not use FPGA operations.
[0070] The methods herein can be advantageously performed with less storage space required than traditional sequencing analysis methods that store flow cell images. Image processing and base calling in parallel with the sequencing reaction can advantageously allow for storing the flow cell image only before base calling is performed in parallel, thereby eliminating the need to store the flow cell image and free up the system's storage space until the end of the sequencing run. Instead of directly storing multiple flow cell images before and / or after image processing (e.g., image registration), the image intensities and corresponding positions of the selected polymerase colonies are saved for base calling. Thus, the methods disclosed herein are less computationally intensive than traditional methods, making the heat dissipation of the computer / processor easier to manage and less likely to cause an undesired interference to the chemistry of the sequencing reaction disclosed herein. Additionally, the transformation matrix can be saved instead of the flow cell image, which can save the required memory space and improve the efficiency of performing 3D base calling operations.
[0071] Sequencing system 110 can be configured to perform image processing of flow cell images across different cycles and / or channels. The operations or actions disclosed herein can be performed by a dedicated processor 118, FPGA 120, computing system 126, or a combination thereof. One or more operations or actions in method 600 disclosed herein can be performed by a dedicated processor 118, FPGA 120, computing system 126, or a combination thereof. In some embodiments, which operations or actions will be performed by the dedicated processor 118, FPGA 120, computing system 126, or combination thereof can be determined based on one or more of the following: the computation time for a particular operation, the complexity of the computation in a particular operation, the need for data transfer between hardware devices, or a combination thereof. The image processing (e.g., image registration) disclosed herein can be performed after acquiring the flow cell image but before base calling of the flow cell image in the cycle.
[0072] Computing system 126 can include one or more general-purpose computers that provide an interface for running various programs in operating systems such as Windows TM or Linux TM and the like. Such operating systems typically provide a great deal of flexibility to the user.
[0073] In some embodiments, dedicated processor 118 can be configured to perform the operations in the methods herein. The dedicated processor may not be a general-purpose processor, but a custom processor with specific hardware or instructions for performing these steps. The dedicated processor runs specific software directly without an operating system. The lack of an operating system reduces the overhead, but at the cost of the flexibility of what the processor can execute. The dedicated processor can use a custom programming language, which can be designed to operate more efficiently than software running on a general-purpose computer. This can increase the speed of performing the steps and allow for real-time processing.
[0074] In some embodiments, the dedicated processor 1180 or the computing system 1260 may include a reconfigurable logic device, such as an artificial intelligence (AI) chip, a neural processing unit (NPU), an application specific integrated circuit (ASIC), or a combination thereof. The reconfigurable logic device may be configured to perform one or more operations herein. The reconfigurable logic device may be configured to perform one or more operations herein and accelerate the operations by allowing parallel data processing compared to a CPU.
[0075] In some embodiments, the FPGA 120 may be configured to perform some or all of the operations of the methods herein. The FPGA is programmed to be hardware that will only perform a specific task. Software steps may be converted to hardware components using a special programming language. Once the FPGA is programmed, the hardware directly processes the digital data provided to it without running software. Instead, the FPGA may use logic gates and registers to process digital data. Since the operating system does not require any overhead, the FPGA generally processes data faster than a general-purpose computer. Similar to a dedicated processor, this comes at the cost of flexibility.
[0076] The lack of software overhead may also allow the FPGA to operate faster than a dedicated processor, although this will depend on the exact processing to be performed and the particular FPGA and dedicated processor.
[0077] A group of FPGAs 120 may be configured to perform these steps in parallel. For example, many FPGAs 120 may be configured to perform processing steps for an image, a collection of images, sub-blocks, or a selected region in one or more images. Each FPGA 120 may simultaneously perform its own portion of the processing steps, thereby reducing the time required to process the data. This may allow the processing steps to be completed in real time. Further discussion of the use of FPGAs is provided below.
[0078] Performing the processing steps in real time may allow the system to use less storage, since the data can be processed as it is received. This is an improvement over conventional systems that may need to store the data before processing it, which may require more storage or access to a computer system located in the cloud 130.
[0079] In some embodiments, the data memory 122 is used to store information used in the methods herein. This information can include the flow cell image itself or information and / or images derived from the flow images captured by the imager 116. The DNA sequences determined according to base calling can be stored in the data storage device 122. The parameters for identifying the polymerase colony locations can also be stored in the data storage device 122. The raw and / or processed image intensities of each polymerase colony can be stored in the data storage device. The regions and / or sub-tiles corresponding to each polymerase colony can also be stored in the data storage device 122. The transformation matrices for each region and / or sub-tile of different cycles and / or channels can also be stored in the data storage device 122. The pool image can be stored in the data memory. The flow cell image, the processed image, and / or the filtered image can be stored in the data memory. Other information or images that contribute to the 3D base identification of the sample can be saved in the data memory.
[0080] The user interface 124 can be used by a user to operate the sequencing system or access data stored in the data storage device 122 or the computer system 126.
[0081] The computer system 126 can control the general operation of the sequencing system and can be coupled to the user interface 124. It can also perform steps in image processing, base calling, its previous operations, and / or subsequent operations (including but not limited to image registration). In some embodiments, the computer system 126 is the computer system 400, as Figure 4 described in more detail below. The computer system 126 can store information about the operation of the sequencing system 110, such as configuration information, instructions for operating the sequencing system 110, or user information. The computer system 126 can be configured to transfer information between the sequencing system 110 and the cloud 130.
[0082] As discussed above, the sequencing system 110 can have a dedicated processor 118, an FPGA 120, or a computer system 126. The sequencing system can use one, both, or all of these elements to complete the necessary processing described above. In some embodiments, when these elements are present together, the processing tasks are divided among them. For example, the FPGA 120 can be used to perform some or all of the following: preprocessing operations, image processing, image registration, base calling, and any subsequent operations, while the computer system 126 can perform other processing functions of the sequencing system 110, such as registering the images for base calling with the cell staining images. Those skilled in the art will understand that various combinations of these elements will allow for various system embodiments that balance the efficiency and speed of processing with the cost of the processing elements.
[0083] The cloud 130 can be a network, a remote storage device, or some other remote computing system separate from the sequencing system 110. The connection to the cloud 130 can allow access to data stored outside the sequencing system 110 or allow for the updating of software in the sequencing system 110.
[0084] 3D base calling based on flattened 2D images
[0085] Figure 6A A flowchart of an exemplary embodiment of a computer-implemented method 600 for performing 3D base calling based on flow cell images is shown. Method 600 can include some or all of the operations disclosed herein. The operations can be performed in an order that is not limited to that described herein.
[0086] Method 600 can be executed by one or more processors disclosed herein. In some embodiments, the processor can include one or more of the following: a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit can include a central processing unit (CPU), an artificial intelligence (AI) chip, a neural processing unit (NPU), and / or a graphics processing unit (GPU). The integrated circuit can include a chip such as a field programmable gate array (FPGA). In some embodiments, the processor can include the computing system 400. In some embodiments, some of the operations in method 600 can be performed by an FPGA, while some other operations in method 600 can be performed by an AI chip or an NPU to improve the energy consumption, heat dissipation, and / or computing time required for sequencing analysis.
[0087] In some embodiments, some or all of the operations in method 600 can be performed by an FPGA. In an embodiment, when some operations are performed by an FPGA, the data after the operations are performed by the FPGA can be transmitted by the FPGA to the CPU such that the CPU can use such data to perform subsequent operations in method 600. Similarly, data can also be transmitted from the CPU to the FPGA for processing by the FPGA. In some embodiments, all of the operations in method 600 can be performed by the CPU. Alternatively, the operations performed by the CPU can be performed by other processors (such as a dedicated processor) or an NPU. In some embodiments, all of the operations in method 600 can be performed by an FPGA and / or an NPU.
[0088] Method 600 can include an operation 610 of obtaining a plurality of flow cell images of one or more samples. The flow cell images can be acquired at different z levels along an axial axis, i.e., the z-axis. In some embodiments, operation 610 includes actively retrieving or passively receiving a plurality of flow cell images of the sample to be processed. In some embodiments, operation 610 includes using the imager 116 of the sequencing system to acquire the flow cell images.
[0089] In some embodiments, operation 610 includes: detecting fluorescence signals and colors emitted by polymerase colonies and clusters by one or more image sensors of an optical system. The one or more sensors correspond to 2, 3, 4, or more color channels of the sequencing system. In some embodiments, a single image sensor can correspond to a single channel or more than one color channel of the sequencing system.
[0090] The sample can be in situ. The sample can be a 3D sample. The sample can be a volume sample that may contain different biological information at the same x-y position but different z positions. The sample can include multiple cells, tissues, or a combination thereof. The 3D sample can be any biological sample having a thickness greater than a predetermined threshold along an axial axis. For example, the thickness can be greater than 2um, 3um, 4um, 5um, 10um, 20um, or greater. The z-axis (e.g., the axial axis) is orthogonal to the image plane defined by the x-axis and the y-axis, as Figure 8B shown.
[0091] Multiple flow cell images can be obtained from 1, 2, 3, 4, or more channels of imager 116 using the optical system disclosed herein. In some embodiments, multiple flow cell images are obtained during a single flow cycle or multiple flow cycles of a sequence run. Each flow cell image can include one or more tiles 210 (imaging regions), and each tile can be divided into multiple sub-tiles. Each sub-tile can include multiple polymerase colonies or clusters. Each sub-tile can include multiple regions, where each region includes a number of polymerase colonies. For example, polymerase colonies can be extracted or otherwise identified from corresponding regions of flow cell images from 4 different channels during a certain cycle. As another example, polymerase colonies can be extracted from a flow cell image from a single channel. The flow cell images disclosed herein can be images obtained from an imaging sample fixed on flow cell 112, as Figure 8B shown.
[0092] Flow cell 112 can contain a sample fixed thereon. The sample can include multiple nucleic acid template molecules. The sample can include a three-dimensional (3D) volume sample or a two-dimensional (2D) sample. The nucleic acid template molecules can be distributed randomly or in various patterns on flow cell 112. In some embodiments, multiple polymerase colonies or clusters herein can be extracted from a specific region (e.g., each sub-tile) of a tile. In the case of each sub-tile, polymerase colonies can be extracted in a predetermined pattern or randomly.
[0093] In some embodiments, the flow cell images herein can be images of one or more tiles, one or more sub - tiles, one or more segmented regions within a tile or sub - tile, or a combination thereof. Each flow cell image can include a field of view (FOV). The FOV can be orthogonal to the axial axis. The FOV can be in the x - y plane. The FOVs of different flow cell images at different axial positions can be the same in the x - y plane. The FOVs of different flow cell images at different axial positions can have at least an overlapping portion in the x - y plane. The image resolutions of different flow cell images at different axial positions can be substantially the same or exactly the same. In some embodiments, the image resolutions of different flow cell images at different axial positions are different. Figure 2A and 3A show two exemplary flow cell images acquired at two different z - levels along the axial axis of the same 3D sample during the same sequencing cycle. In some embodiments, the image resolution of the flow cell image along the x, y, and / or z axes can be in the range of 0.001um to 5um. In some embodiments, the image resolution of the flow cell image along the x, y, and / or z axes can be in the range of 0.01um to 2um. In some embodiments, the image resolution of the flow cell image along the x, y, and / or z axes can be in the range of 0.02um to 1um.
[0094] Each flow cell image at a particular z - level includes intensities generated by the polymerase colonies and clusters at the corresponding z - position. As Figures 2A - 3A shown, the signals from the polymerase colonies and clusters are small bright spots in the image. Each bright spot can have various sizes less than a few pixels, such as less than one pixel, about one pixel, about 2 pixels, 3 pixels, 4 pixels, or 5 pixels. In some embodiments, each signal point of the polymerase colony or cluster can be any number of pixels in the range from 0.01 pixels to about 72 pixels. In some embodiments, each signal point of the polymerase colony or cluster can be any number of pixels in the range from 0.1 pixels to about 16 pixels.
[0095] Each flow cell image can also include intensities generated by cells and their structural elements. Such structural elements can be background objects or components.
[0096] In some embodiments, when the depth of field of the optical system includes a range extending along the z - axis, such as 0.1um, 0.2um, 0.3um, 0.5um, 0.6um, 0.8um, 1um, 2um, 3um, 4um, 5um, etc. The polymerase colonies and clusters within the depth of field can appear in focus or substantially in focus in the flow cell image. The flow cell image at a particular z - level can also include signals from polymerase colonies and clusters that are not within the focal range of the image. Such polymerase colonies or clusters are out of focus. AsFigure 3A As shown, the larger and blurred signal dots represent out-of-focus polymerase colonies or clusters. Some out-of-focus polymerase colonies or clusters are circled in Figure 3A .
[0097] Each flow cell image at a particular z-level may also include noise caused by the optical system and / or unwanted signals from the sample. The unwanted signals may be signals from components of the sample (such as membranes, cytoplasm, and mitochondria). Such background objects may be any object that is relatively larger in size than the polymerase colonies or clusters. As Figure 3A shown, there are blurred cell outlines (at the arrows) in the flow cell image, and most of the signal dots are contained within the blurred outlines. Figure 3D Multiple cells and polymerase colonies or clusters are shown because the small bright dots are typically located within the outlines of different cells. In some embodiments, the background objects may include any object within the 3D sample, but not the polymerase colonies or clusters.
[0098] In some embodiments, the polymerase communities or clusters being sequenced in the flow cycle can have a certain nucleotide diversity, e.g., in base calling. Method 600 allows for 3D base identification of flow cell images even when the polymerase communities or clusters have low or unbalanced diversity during the sequencing cycles at a pre-determined base identification quality level. The nucleotide diversity of a nucleotide acid molecule population (e.g., polymerase communities or clusters) can refer to the relative proportions of nucleotides A, G, C, and T / U present in each flow cycle. The relative proportions of nucleotides can be within a region of the field of view or within the entire flow cell image. Optimal high or balanced diversity data typically can have approximately equal proportions of all four nucleotides represented in each flow cycle of a sequencing run. Low or unbalanced diversity data typically can include high proportions of certain nucleotides and low proportions of other nucleotides in some flow cycles of a sequencing run, e.g., less than 10% of the total of all 4 nucleotides. Thus, an image corresponding to a high proportion of certain nucleotides can have more signal points (polymerase communities or clusters) compared to an image corresponding to a low proportion of certain nucleotides. As an example of low or unbalanced diversity data, in a certain flow cycle, bases A, T, C, G can be approximately 1%, 2%, 1%, and 95% of the total polymerase community, respectively. Subsequently, the flow cell image from the channels corresponding to A, T, and C in this particular flow cycle is darker and / or has far fewer polymerase communities or clusters than the flow cell image corresponding to nucleotide G. As another example of low or unbalanced diversity data, bases A, T, C, G in the polymerase community over multiple flow cycles can be approximately 2%, 5%, 10%, and 83%, respectively. In embodiments where low or unbalanced diversity data is present in a particular cycle and the data is imaged for sequencing analysis, using prior art image registration and subsequent base identification may fail because the image from one or more channels is too dark compared to the images acquired from other channels (e.g., the signal points of the polymerase community are too sparse and / or dim), thus causing problems in sequencing analysis. Additionally, in embodiments where low or unbalanced diversity data is present in a particular cycle, existing methods may not be able to identify dimers or sparser polymerase communities or clusters from other background information (e.g., cell structure). The pre-determined quality level can be customized based on different sequencing applications. For example, the pre-determined quality level can be at least Q30, Q35, Q40 or higher. As another example, in terms of base identification, the pre-determined quality level can be less than 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.001% or less error. In some embodiments, method 600 is configured to perform base identification of 3D samples even when the polymerase communities or clusters have low diversity in some regions of the flow cell during some flow cycles from the 3D samples.
[0099] In some embodiments, method 600 is configured to process flow cell images, e.g., including intensity processing, registration, and base calling, even when polymerase colonies or clusters have unbalanced nucleotide diversity in one or more flow cycles at a pre-determined quality level.
[0100] In some embodiments, method 600 is configured to process flow cell images, e.g., including intensity processing, registration, and base calling, even when the spatial density of polymerase colonies or clusters is higher than the spatial density that can be processed using existing 3D sequencing methods. In some embodiments, the spatial density of the 3D sample can be 2 times, 4 times, 6 times, 8 times, 10 times, 12 times, 15 times or greater than the maximum spatial density manageable by an existing 3D sequencing system having a pre-determined quality (e.g., Q40 or Q30) in base calling. In some embodiments, the 3D sample can have a spatial density greater than about 0.01 to about 0.5 polymerase colonies / μm 3 of the spatial density. As another example, the 3D sample can have a spatial density of not less than 0.01 to about 0.5 polymerase colonies / μm 3 of the spatial density. In some embodiments, the 3D sample can have a spatial density of not less than about 0.1 to about 1 polymerase colonies / μm 3 of the spatial density. In some embodiments, the 3D sample can have a spatial density of not less than 0.1 to 1 polymerase colonies / μm 3 of the spatial density. In some embodiments, the 3D sample can have a spatial density of not less than 0.002 to 50 polymerase colonies / μm 3 of the spatial density. In some embodiments, the 3D sample can have a spatial density of not less than 0.005 to 30 polymerase colonies / μm 3 of the spatial density. In some embodiments, the 3D sample can have a spatial density of not less than 0.01 to 50 polymerase colonies / μm 3 of the spatial density. In some embodiments, the 3D sample can have a spatial density of not less than 0.05 to 100 polymerase colonies / μm 3 of the spatial density.
[0101] In some embodiments, such high-density 3D samples can be divided into groups of "batches", and different sequencing primers can be used for each batch. In some embodiments, by controlling the application of fluorescent dyes to be attached to the polymerase colonies or clusters, such high-density 3D samples can be further divided into "sub-batches" within each batch by selectively "turning off" the sub-batches. Details of exemplary embodiments of sequencing polymerase colonies or clusters of different batches and / or sub-batches are described in PCT patent applications PCT / US23 / 65972 and PCT / US23 / 74933, the contents of which are hereby incorporated by reference in their entirety. The methods herein can be combined with "batch" and / or "sub-batch" sequencing techniques, advantageously increasing the base calling throughput by 3-fold, 5-fold, 10-fold or more compared to existing methods of bulk sequencing of samples without changing the sequencing primers, cartridges and optical systems.
[0102] In some embodiments, method 600 is performed during cycle N, which is different from the reference cycle. A template image can be generated in the reference cycle, and polymerase colonies in one or more channels within the reference cycle can be included in the template image in a reference coordinate system, while base calling for cycle N has not been performed. In some embodiments, cycle N is the current cycle. N can be any non-zero integer. For example, for short read sequencing, N can be any integer from 1 to 150 or from 1 to 300.
[0103] Figure 8B A schematic diagram of one or more template images generated in a reference coordinate system in a reference cycle is shown. In some embodiments, template image 210 has approximately the same size as a single tile 210 that includes a 5×5 grid of sub-tiles. In some embodiments, the template images disclosed herein can be separate regions within a sub-tile. Each template image can include a plurality of polymerase colonies or clusters therein.
[0104] In some embodiments, the template image can have approximately the same size as the flow cell image, such that Figure 8B the different pictures 210 from and all polymerase colonies from multiple channels can be registered with the same template image. However, such a template image can contain polymerase colonies that are not used in at least some of the operations described herein to reduce the computational burden without sacrificing accuracy.
[0105] In some embodiments, more than one template image can be generated, and each template image corresponds to at least a portion of a sub-tile of the flow cell image from a channel.
[0106] The template image of the present disclosure can be initialized as a virtual image, which has a black or dark background and no signal from the polymerase population. For example, the template image can be initialized to zero, or include the minimum image intensity at all pixels.
[0107] After determining the coordinates of the polymerase population by, for example, image registration of flow cell images across different channels, the intensity of the polymerase population can be added to the template image at the positions determined by the coordinates, and its size and shape can be determined based on the registration. The template image can be a virtual image that combines the image intensities of the polymerase populations obtained from 2, 3, 4, or more channels in a reference cycle. The pixels of the template without polymerase populations remain black or dark, so that the template image can have a cleaner background without the noise present in the actual flow cell images.
[0108] The polymerase population can be from sub - tiles of flow cell images within a reference cycle, and more specifically, from one or more selected regions of the sub - tiles. The flow cell images can be from different channels among 1, 2, 3, 4, or more channels of the system 100. As a non - limiting example, the reference cycle can be any of the first 5 or 6 cycles. In some embodiments, the reference cycle can be any cycle greater than 0. In some embodiments, the reference cycle is the first cycle.
[0109] In some embodiments, operation 610 includes generating a 2D flow cell image having the intensity of the polymerase populations or clusters of a 3D sample such that the intensity can be used to perform 3D base calling for different sequencing cycles. The operation of generating the flow cell image can be performed using the sequencing system 110 herein. The operation of generating the flow cell image can be performed using the optical system herein.
[0110] Method 600 can include an operation of obtaining a processed image of the flow cell image. In some embodiments, the operation of obtaining the processed image can include processing the flow cell image using one or more of the predetermined processing methods herein.
[0111] One or more of the predetermined processing methods can include various image processing or operations such as filtering, normalization, spatial frequency analysis, etc. In some embodiments, one or more of the processing methods can include selecting a kernel and generating a processed image by performing an operation on the flow cell image using the selected kernel. For example, the operation performed can be an opening operation, which can be represented Where f is the flow cell image and k is the kernel. The opening operation can be performed in the spatial domain. Alternatively, to make the calculation faster or simpler, the opening operation can be performed in a different domain, such as the Fourier domain. The opening operation can be the dilation of the erosion of the image f by the kernel k. The opening operation can remove objects smaller than the kernel, and the subsequent dilation operation can restore the size and shape of the remaining objects. Figure 2B and Figure 3B shows an exemplary processed image after the opening operation. The processed image can also be referred to as the image after the opening operation, in which most of the bright spots of the polymerase colonies or clusters are removed while other background objects are retained.
[0112] As another example, the operation performed can be convolution, and the flow cell image can be convolved with a selected kernel. As yet another example, obtaining a plurality of processed images further includes: selecting a first kernel and a second kernel; generating a first image by convolving a plurality of flow cell images with the first kernel; and generating a second image by convolving a plurality of flow cell images with the second kernel. The first image and the second image can be different blurred images after convolution with different blur kernels.
[0113] The first or second kernel can be of any size smaller than the size of the corresponding flow cell image. For example, through the opening operation, the kernel can be 2×2, 3×3, 4×4, 5×5, or 6×6. In some embodiments, the kernel size can be customized to remove at least some of the noise and unwanted signals larger than the kernel size. In some embodiments, the kernel can be circular. The kernel can be various other shapes, such as oval, square, rectangle, rhombus, etc.
[0114] In some embodiments, the kernel is a Gaussian kernel. In embodiments where 2 different kernels are used, the first kernel and the second kernel can be different Gaussian kernels.
[0115] In some embodiments, method 600 can include an operation 620 of filtering the flow cell image. The filtering operation 620 can be based on the processed image generated in its previous operation. The filtering operation 620 can be based on a predetermined filter. The filtering operation 620 can generate a plurality of filtered images, each filtered image corresponding to the flow cell image at a corresponding z-level along the axial axis.
[0116] In some embodiments, operation 620 can include subtracting the processed image from the corresponding flow cell image to generate a filtered image. The filtered image can be obtained according to where fi represents the filtered image, f represents the flow cell image, and k represents the kernel, and represents an operation, such as the opening operation.
[0117] In some embodiments, a filtered image is generated in operation 620 based on a corresponding flow cell image before filtering and at least one or more image filters. In some embodiments, the filtered image is an image of the corresponding flow cell image filtered by a predetermined filter. In some embodiments, the predetermined filter is a top hat filter. In some embodiments, the predetermined filter is a Difference of Gaussians (DoG) filter, and the filtered image is an image of the corresponding flow cell image filtered by the Difference of Gaussians (DoG) filter. In some embodiments, the predetermined filter may include various filters that are configured to retain small elements and details (e.g., focused polymerase colonies or clusters) in the flow cell image and optionally remove elements and details larger than the focused polymerase colonies or clusters.
[0118] After filtering, at least a portion of the noise and unwanted signals in the flow cell image are removed, which may include cellular components and out-of-focus polymerase colonies. Removal of such noise and unwanted signals that are inevitable in 3D flow cell images can advantageously facilitate the generation of intensities attributable to polymerase colonies or clusters without other background objects or noise. The filtered intensities can be used for more accurate and reliable base calling than the unfiltered intensities.
[0119] Figure 2C and 3C Shows two filtered images from two different axial positions. Figure 2C Indicates that at a first axial position of z = 0, a large number of polymerase colonies or clusters are in focus, and at a second axial position of z = 5, a much smaller number of polymerase colonies or clusters are in focus, with the second axial position being approximately 10 um from the first axial position.
[0120] Figure 3E Shows Figure 3D Another exemplary filtered image of the flow cell image in. Background objects in the flow cell image are filtered out in the filtered image. Out-of-focus polymerase colonies in the flow cell image are also removed from the filtered image. The filtered image shows focused polymerase colonies or clusters.
[0121] In some embodiments, method 600 may include an operation of adding an offset to the filtered image. In some embodiments, the offset may be predetermined such that after the offset, the range of the image intensity may be within a predetermined range. In some embodiments, different offsets may be used to bring different filtered images having various intensity ranges within a predetermined range common to all the filtered images.
[0122] Method 600 may include operation 630 of generating a maximum intensity projection (MIP) image based on a plurality of filtered images. Operation 630 may include calculating the maximum intensity for each pixel of the MIP image using the intensities at the corresponding pixel of the plurality of filtered images. The MIP image may have the same image size, resolution, and / or FOV as the flow cell image. The MIP image may be a flattened 2D image of a stack of flow cell images from different axial positions. The MIP image may correspond to some or all of the FOV of flow cell images from multiple z-levels in a single color channel and in a single flow cycle. The MIP image may correspond to some or all of the FOV of flow cell images from multiple z-levels in a single color channel and in one or more flow cycles.
[0123] For example, an axial stack of 10 flow cell images is acquired at 10 different z-levels corresponding to the same color channel in a particular flow cycle. Each z-level is approximately 0.1 um to approximately 20 um from its adjacent z-position. The first z-level is at z = 0, and the tenth z-level is at z = 9. A filtered image is generated for each of the different z-positions. The MIP image may be initialized to the same size as the flow cell image, such as 1028 x 1028. All intensities in the initial MIP image may be 0 or any other default initial intensity. For each pixel at pixel (i,j) of the MIP image, the image intensities at pixel (i,j) from the 10 filtered images are extracted, and the maximum intensity among the 10 different image intensities is selected for pixel (i,j) in the MIP image. In some embodiments, an MIP image may be generated for each cycle corresponding to the same color channel. In some embodiments, an MIP image may be generated for each different channel among 1, 2, 3, 4, 5, 6 or more channels.
[0124] In some embodiments, before obtaining the MIP image using the filtered images, the image intensities in the filtered images are normalized to a predetermined range. In some embodiments, the MIP images may be normalized across different cycles. In some embodiments, the MIP images in the x-y plane, such as covering different tiles of different FOVs, may be normalized in different ways to address spatial variations in the sample distribution on the flow cell device. The normalization may reach a predetermined range.
[0125] In some embodiments, method 600 includes an operation of image registration of images across different channels and / or different flow cycles. Image registration can be performed after operation 610 but before operation 620. Image registration can be performed after operation 620 but before operation 630. Image registration can be performed after operation 630 but before operation 640. Image registration can be performed before any base calling. Various image registration methods can be used herein. In some embodiments, the operation of image registration can include Figure 9 one or more operations in method 900 in []. Exemplary embodiments of image registration methods are described in PCT patent application number PCT / US23 / 67931, the content of which is incorporated herein by reference in its entirety.
[0126] In some embodiments, image registration using existing methods may fail due to unbalanced nucleotide diversity, which results in one or more flow cell images from some of the color channels being dimers and / or having fewer bright spots. Various image rescue techniques can be used to allow accurate image registration of such flow cell images. Exemplary embodiments of image registration rescue methods are described in PCT patent application number PCT / US23 / 67931, the content of which is incorporated herein by reference in its entirety.
[0127] In some embodiments, method 600 includes an operation of registering a MIP image, such as Figure 9 900 in []. In some embodiments, the MIP image is registered across channels and / or different cycles before any base calling is performed. Various image registration techniques can be used to register the MIP image. 2D registration techniques can be used to register the MIP image, for example, by treating the MIP image as a flow cell image acquired from sequencing system 110. In some embodiments, the MIP image can be registered using image registration method 900 as disclosed herein by treating each MIP image as a flow cell image, for example, across different channels and / or different cycles. In some embodiments, the MIP image can be registered after one or more preprocessing operations disclosed herein are performed.
[0128] Method 600 can include an operation 640 of performing 3D base calling using the MIP image after image registration of the MIP image. The MIP image can be 2D, and base calling can be performed based on MIP images from different channels for each cycle.
[0129] In some embodiments, method 600 may include operations 630' and 640' for performing 3D base identification instead of operations 630 and 640. In some embodiments, operation 630' may include determining MIPs in the filtered flow cell image and their corresponding z-levels. For example, the MIP of the flow cell image at 10 different z-levels at pixel (10, 10) may be 100 in channel 1 and at z = 5, while the MIP for pixel (11, 11) may be 110 in channel 4 and at z = 6. In some embodiments, operation 640' may include performing 3D base identification based on MIPs from different channels (e.g., for pixel (10, 10), from channel 1 at z = 5 and from channel 3 at z = 6) and having matching z-positions. The base identification may correspond to different polymerase colonies located at different z-levels, which polymerase colonies extend throughout the 3D volume of the example.
[0130] In some embodiments, the MIP-based base identification operation may generate a smaller number of base identifications than the base identification method using the 3D polymerase colony map because some of the polymerase colonies may not be included in the MIP or MIP image.
[0131] In some embodiments, operation 640 of performing base identification using the MIP image includes performing the preliminary analysis steps herein to adjust the image intensity of the polymerase colonies in the MIP; and performing base identification of the polymerase colonies according to the adjusted image intensity in the MIP. In some embodiments, the preliminary analysis steps include one or more of the following: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; trailing and leading correction; image registration; quality score estimation. In some embodiments, image registration as the preliminary analysis step herein is configured to align images from different cycles and different channels, for example, relative to a template image or a reference coordinate system. In some embodiments, image registration as the preliminary analysis step herein is configured to register polymerase colonies or clusters from different cycles and different channels (e.g., in the MIP image) with a template image or a reference coordinate system.
[0132] For example, after registering MIP images from different channels relative to the template image disclosed herein, base identification may be performed using the first MIP image from the first color channel and other MIP images from different color channels in cycle N. Various existing 2D base identification algorithms may be used. For example, the image intensities at the same pixel in MIP images from different channels may be compared, and the maximum intensity and its corresponding channel determine the base identification for that pixel.
[0133] In some embodiments, method 600 may include, before operation 640, an operation of determining whether a plurality of flow cell images have unbalanced diversity. Unbalanced diversity may lead to errors in image registration, color correction, or other preliminary analysis steps, resulting in inaccurate base calling using existing methods. Method 600 advantageously enables 3D base calling for flow cell images with low or unbalanced diversity. Low or unbalanced diversity may occur in certain regions of the flow cell images (e.g., in one microfluidic channel but not in other microfluidic channels) and / or in one or more flow cycles. The operation of determining whether a plurality of flow cell images have unbalanced diversity in one or more flow cycles may include determining the corresponding percentages of: (1) the number of each type of nucleotide base (e.g., A, T, C, or G) and (2) the total number of nucleotide bases in the region of the sample immobilized on the flow cell device in at least one region of the FOV, and determining whether the corresponding percentages are less than a predetermined diversity threshold. The diversity threshold may be customized according to different sequencing applications and / or samples. For example, the diversity threshold may be 20%, 18%, 16%, 15%, 12%, 11%, 10%, 8%, 6%, 5% or less.
[0134] In response to determining that the plurality of flow cell images have unbalanced diversity (e.g., in some regions or in one or more flow cycles), method 600 includes an operation of determining image processing parameters based on the existing values of the corresponding image processing parameters determined in the cycle(s) prior to the one or more cycles. The cycle(s) prior to the one or more cycles may have balanced nucleotide diversity. For example, if cycle 30 has unbalanced diversity and nucleotides A and G account for less than 10% of the total number of nucleotides (polymerase population) in that cycle, the color correction and / or image registration for cycle 30 may be performed using the pre-existing color correction and / or image registration parameters from a cycle with balanced diversity (e.g., cycle 29, cycle 25, or even cycle 20). As another example, a flow cell image has 16 different regions, and each region may have its corresponding image processing parameters to address the spatial variation in color correction. Additionally, in image processing steps such as color correction and / or image registration, one region with unbalanced diversity may not affect other regions with balanced diversity. Instead, only the region with unbalanced diversity may use the pre-existing image processing parameters from the same region of a cycle with balanced diversity, while other regions without unbalanced diversity may still use the image processing parameters determined in the current cycle. In response to determining that the plurality of flow cell images have balanced diversity, method 600 includes an operation of determining image processing parameters for each flow cell image of the plurality of flow cell images based on the image intensity of the current cycle rather than any previous cycle.
[0135] In response to determining that, for example, in some regions or in one or more flow cycles, at least some of the plurality of flow cell images at multiple z - levels have unbalanced diversity, various methods can be used to adjust the range of image intensities of the flow cell images. For example, before obtaining the MIP image of such flow cell images, the same fiducial marker with a pre - determined image intensity can be used as a reference intensity to adjust the image intensities of the flow cell images at multiple z - levels.
[0136] In some embodiments, even if in some regions and / or in one or more flow cycles, the plurality of flow cell images have unbalanced nucleotide diversity, method 600 and its operations can be configured to perform 3D base calling at a pre - determined quality level. The pre - determined quality level can be customized based on different sequencing applications. For example, the pre - determined quality level can be at least Q30, Q35, Q40 or higher. As another example, in terms of base calling, the pre - determined quality level can be less than 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.001% or even fewer errors.
[0137] The base calling results can be saved together with their 3D coordinates. Such 3D coordinates can be used to localize base calling across different cycles and different z - levels to a common coordinate system. Such 3D coordinates can allow base calling to be localized within the 3D volume representing the sample.
[0138] In some embodiments, method 600 includes operation 650 of obtaining a second MIP image based on the plurality of flow cell images. The second MIP is a flattened 2D image of the axial stack of the flow cell images. In some embodiments, the flow cell images are the original images acquired by the sequencing system 110. The original images herein can include flow cell images that are not filtered with various image filters. The original images herein can be flow cell images in which polymerase colonies or clusters and other background information, such as noise, cellular components, are retained. The operation of obtaining the second MIP can be different from the operation of obtaining the first MIP image. In the flow cell images (e.g., original images), out - of - focus polymerase colonies can have a larger full width at half maximum (FWHM), such that the signal of out - of - focus polymerase colonies is more dispersed than that of in - focus polymerase colonies or clusters. In some embodiments, the larger FWHM can result in a white ring around the polymerase colony. Figure 5An exemplary second MIP image generated directly from the original flow cell image is shown, and there is a ring or halo around the polymerase community at the center of the image. The ring or halo may be an image artifact that may cause errors in base calling. The second MIP may also retain some background information (e.g., background information from unwanted background objects), which may interfere with base calling if the second MIP is used for base calling. Thus, the second MIP includes artifacts and / or unwanted background objects that require additional processing before accurate and reliable base calling based on the intensities in the second MIP.
[0139] In some embodiments, the second MIP can be used to facilitate registration of the first MIP and / or the polymerase community or clusters because it shares the same FOV, resolution, image size, etc. with the first MIP image. If the second MIP is used for base calling, the artifacts and / or unwanted background objects in the second MIP may interfere with base calling, but the same artifacts and / or unwanted background objects (not in the first MIP) can be used to facilitate registration of the first MIP by using the second MIP, e.g., registering to a stained cell image.
[0140] In some embodiments, the methods herein advantageously utilize the second MIP for registering the flow cell image with the pool image. In some embodiments, the cell image herein is a microscopic image of a stained sample with some cell features (such as membranes and nuclei). Out-of-focus polymerase communities and background objects that may interfere with correct base calling can be used to provide information for registering the flow cell image with the pool image.
[0141] In some embodiments, operation 660 includes: registering a second MIP image to a cell image. Each color channel may include one or more second MIP images per flow cycle. If image registration of flow cell images across different channels and / or cycles occurs prior to operation 660, then the second MIP images of the registered flow cell images may be aligned. The second MIP image may be registered to the cell image based on background information contained within both the second MIP image and the cell image. Various image registration methods may be used to align or register the second MIP image and the cell image. Method 600 may further include operation 660 of performing image registration of flow cell images based on the second MIP image. In some embodiments, image registration of flow cell images with one or more cell images (e.g., stained images) is complementary to image registration of flow cell images across cycles and / or channels. Image registration of flow cell images may be configured to align polymerase colonies or clusters relative to cell structures such that base calling may be assigned to relevant cell components, such as the nucleus, membrane, or other regions of the cell. In some embodiments, operation 660 of registering a flow cell image based on a second MIP image includes: registering or aligning background objects in the second MIP to corresponding objects in one or more cell images. For example, image processing methods such as contour drawing may be used to trace a membrane in the second MIP image and align the membrane to the membrane in the cell image.
[0142] In some embodiments, operation 660 of performing image registration of multiple flow cell images based on a second MIP image includes: registering the second MIP image to a template image or a reference coordinate system.
[0143] In some embodiments, the second MIP and the first MIP are obtained from flow cell images of the same FOV, and thus the second and first MIP images may share the same FOV. In some embodiments, registering the second MIP image to a template image or a reference coordinate system may rely on the image registration information of the first MIP to the template image. In some embodiments, registering the first MIP image to a cell image may be based on registering the second MIP image to the cell image, since the first and second MIP images share the same FOV.
[0144] In some embodiments, method 600 may include operation 670 of registering 3D base identification to one or more cell images. In some embodiments, operation 670 may include: registering a first MIP image to one or more cell images. In some embodiments, operation 670 may include: registering the first MIP image to one or more cell images based on a second MIP image, because the second MIP image contains background information that can facilitate registration to one or more cell images. In some embodiments, operation 670 may include: registering 3D base identification to one or more cell images based on the registered first MIP image and second MIP image. Background objects in the second MIP can be used to align the second MIP and the first MIP with the tile image. The cell images may have at least some of the same background objects as the background objects in the second MIP, but may have transformations, e.g., translations and / or rotations. The transformation may be represented by a single transformation of the entire image, or may be divided into multiple transformations, each representing a portion of the entire image. After finding the transformation of the background object between the second MIP and the tile image, the polymerase colonies and clusters can be registered with the tile image.
[0145] In some embodiments, operation 670 of registering the first MIP image and / or 3D base identification to the cell image may not rely on the second MIP image directly obtained from the flow cell image, but only on the first MIP image of the filtered image. In some embodiments, the flow cell image can be directly used to register the filtered image to the cell image, rather than using the second MIP generated from the flow cell image. The flow cell image may contain noise and unwanted background objects for base identification purposes, but such noise and background objects can facilitate registering the flow cell image and the filtered image to the cell image.
[0146] In these embodiments, method 600 may include operation 660' of registering to the cell image using background information obtained from one or more of the following: the first MIP image, the opened image / processed image, the filtered image, and the flow cell image, rather than operation 660 of using the second MIP image.
[0147] Background objects can be used to align one or more of the following to the cell image by using one or more transformations: the first MIP image, the opened image / processed image, the filtered image, and the flow cell image. Background objects can be used to align the first MIP image to the cell image by using one or more transformations. The transformation may be represented by a single transformation of the entire image, or may be divided into multiple transformations, each representing a portion of the entire image. After finding the transformation of the background object between the first MIP and the tile image, the polymerase colonies and clusters can be registered with the tile image.
[0148] In some embodiments, the operation 670 of registering the MIP image and / or 3D base identification to the cell image may not rely on the second MIP image, but rather on fiducial markers. In some embodiments, method 600 may include an operation 660 of using fiducial markers external to the sample to register to the cell image, rather than an operation 660 of using the second MIP image. The same fiducial markers may be present in one or more of the following: the first MIP image, the open image / processed image, the filtered image, and the flow cell image. Such fiducial markers may also be included in the cell image. Aligning the same fiducial markers may generate a transformation between the sequencing image (e.g., the first MIP image or the flow cell image) and the cell image. The transformation may be used to register or align the polymerase colonies or clusters between the flow cell image and the cell image. Figures 7A - 7B An exemplary registered first MIP image superimposed on a corresponding cell image having a cell structure is shown. The first MIP image includes bright spots representing polymerase colonies. Some of the bright spots overlap with the stained cell nuclei, while some other bright spots appear within the cell membrane but outside the cell nuclei. Figure 7B A segmentation of individual cells is shown such that polymerase colonies or clusters can be grouped relative to each individual cell.
[0149] Base calling using 3D polymerase colony maps
[0150] Figure 6B A flowchart of an exemplary embodiment of a computer-implemented method 600 for performing 3D base identification based on a 3D polymerase colony map according to a flow cell image is shown. Method 600 may include some or all of the operations disclosed herein. The operations may be performed in an order that is not limited to that described herein.
[0151] Method 600 may be executed by one or more processors disclosed herein. In some embodiments, the processor may include one or more of the following: a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit may include a central processing unit (CPU) and / or a neural processing unit (NPU). The integrated circuit may include a chip such as a field programmable gate array (FPGA). In some embodiments, the processor may include computing system 400.
[0152] Method 600 may be executed based on the flow cell image from the current sequencing cycle alone, or in combination with information from previous cycles of the current sequencing cycle.
[0153] In some embodiments, some or all of the operations in method 600 may be performed by an FPGA. In an embodiment, when some operations are performed by the FPGA, the data after the operations are performed by the FPGA may be transmitted by the FPGA to the CPU, such that the CPU may use such data to perform subsequent operations in method 600. Similarly, data may also be transmitted from the CPU to the FPGA for processing by the FPGA. In some embodiments, all of the operations in method 600 may be performed by the CPU. Alternatively, the operations performed by the CPU may be performed by other processors (such as dedicated processors) or NPUs. In some embodiments, all of the operations in method 600 may be performed by the FPGA and / or NPU.
[0154] Method 600 may include operation 610 of obtaining a plurality of flow cell images of one or more samples. The flow cell images may be acquired at different positions along an axial axis (i.e., the z-axis). In some embodiments, operation 610 includes actively retrieving or passively receiving a plurality of flow cell images of the sample to be processed. In some embodiments, operation 610 includes using imager 116 of the sequencing system to acquire the flow cell images. In some embodiments, the flow cell images are acquired by an NGS sequencing system.
[0155] The sample may be in 3D. The sample may be a volumetric sample, which may contain different biological information at the same x-y position but different z positions. The sample may be an in-situ sample. The sample may include a plurality of cells, tissues, or a combination thereof. The sample may be any biological sample having a thickness greater than a predetermined threshold (e.g., depth of field) along the axial axis. The sample may be any biological sample having a thickness greater than the depth of field of the optical system. For example, the thickness may be greater than 1um, 2um, 3um, 4um, 5um, 8um, 10um, 12um, 15um, 20um or greater. The z-axis (e.g., the axial axis) may be orthogonal to the image plane defined by the x-axis and the y-axis, as Figure 8B shown.
[0156] In some embodiments, the flow cell images may include a first plurality of flow cell images and a second plurality of flow cell images. Each of the first plurality of flow cell images and the second plurality of flow cell images may be acquired at corresponding positions along the axial axis or the z-axis. The axial axis may extend from the objective lens of the sequencing system 110 to the sample located on the flow cell positioned on the sequencing system. As a non-limiting example, the first plurality of flow cell images may be obtained from a reference cycle, and the second plurality of flow cell images may be obtained from one or more cycles different from the reference cycle. As another example, the first plurality of flow cell images may be the same as the second plurality of flow cell images.
[0157] Flow cell images can be acquired from one, two, three, four, or more channels of imager 116 using the optical systems disclosed herein. Each flow cell image can include one or more tiles (imaging regions), and each tile can be divided into multiple sub-tiles. Each sub-tile can include multiple polymerase colonies or clusters. Each sub-tile can include multiple regions, where each region includes a number of polymerase colonies. For example, polymerase colonies can be extracted from corresponding regions of flow cell images from four different channels in a given cycle. As another example, polymerase colonies can be extracted from a flow cell image from a single channel. The flow cell images disclosed herein can be images from a flow cell 112 as shown in Figure 1 as shown.
[0158] In some embodiments, the flow cell images herein can be images of one or more tiles, one or more sub-tiles, one or more segmented regions with tiles or sub-tiles, or combinations thereof. Each flow cell image can contain a field of view (FOV). The FOV can be orthogonal to the axial axis. The FOV can be in the x-y plane. The FOVs of different flow cell images at different axial positions can be the same in the x-y plane. The FOVs of different flow cell images at different axial positions can have at least an overlapping portion in the x-y plane.
[0159] The image resolutions of different flow cell images at different axial positions can be approximately the same or exactly the same. In some embodiments, the image resolutions of different flow cell images at different axial positions are different.
[0160] Figure 2A and 3A show two exemplary flow cell images acquired at two different z-levels along the axial axis of the same 3D sample during the same sequencing cycle.
[0161] Each flow cell image at a particular z-level includes intensities generated by polymerase colonies and clusters at the corresponding z-position. As shown in Figures 2A - 3A , the signals from polymerase colonies and clusters are small bright spots in the image. Each bright spot can have various sizes less than a few pixels, e.g., less than one pixel, about one pixel, about 2 pixels, 3 pixels, 4 pixels, 5 pixels, or more. In some embodiments, each signal point of a polymerase colony or cluster can be any number of pixels in the range from 0.01 pixels to about 72 pixels. In some embodiments, each signal point of a polymerase colony or cluster can be any number of pixels in the range from 0.1 pixels to about 16 pixels.
[0162] The polymerase population can be from sub - tiles of the flow cell image within a reference cycle, and more specifically, from one or more selected regions of the sub - tiles. The flow cell image can be from different channels among 1, 2, 3, 4, or more channels of system 100. As a non - limiting example, the reference cycle can be any of the first 5 or 6 cycles. In some embodiments, the reference cycle can be any cycle greater than 0. In some embodiments, the reference cycle is the first cycle.
[0163] In some embodiments, the flow cell image is acquired under one or more reference cycles or cycles different from one or more reference cycles. As a non - limiting example, cycles 1 - 5 or 2 - 5 can be reference cycles. The flow cell image can be acquired within a single cycle or several cycles.
[0164] Each flow cell image can include intensities generated by the flow cell and its structural elements. Such structural elements can be background objects or components. Figure 3D Multiple cells and polymerase populations or clusters are shown, as small bright spots are typically within the outlines of different cells.
[0165] In some embodiments, when the focus of the optical system includes a range extending along the z - axis, such as 0.1um, 0.2um, 0.3um, 0.5um, 0.6um, 0.8um, 1um, 2um, 3um, 4um, 5um, etc. Polymerase populations and clusters within the range of the focus can appear focused or approximately focused in the flow cell image. The flow cell image at a particular z - level can also include signals from polymerase populations and clusters that are not within the focal range of the image but are at different z - positions. Thus, such polymerase populations or clusters are out - of - focus. As Figure 3A shown, larger and blurred signal points represent out - of - focus polymerase populations or clusters. Some out - of - focus polymerase populations or clusters are circled in Figure 3A the figure.
[0166] Each flow cell image at a particular z - level can also include noise caused by the optical system and / or unwanted signals from the sample. Unwanted signals can be signals from components of the sample (such as membranes, cytoplasm, and mitochondria). Such background objects can be any object that is relatively larger in size than the polymerase population or cluster. As Figure 3A shown in the figure, there are blurred cell outlines (at the arrows) in the flow cell image, and most of the signal points are contained within the blurred outlines. In some embodiments, the background object can include any object within the 3D sample but not the polymerase population or cluster.
[0167] In some embodiments, the polymerase population or clusters being sequenced in a flow cycle can have a certain nucleotide diversity, e.g., in base calling. Method 600 allows for 3D base identification of flow cell images even when the polymerase population or clusters have low diversity or unbalanced diversity during the sequencing cycle. The nucleotide diversity of a nucleotide acid molecule population (e.g., polymerase population or clusters) can refer to the relative proportions of nucleotides A, G, C, and T / U present in each flow cycle. The relative proportions of nucleotides can be within a region of the field of view or within the entire flow cell image. Optimal high or balanced diversity data typically can have approximately equal proportions of all four nucleotides represented in each flow cycle of a sequencing run. Low or unbalanced diversity data typically can include high proportions of certain nucleotides and low proportions of other nucleotides in some flow cycles of a sequencing run, e.g., less than 10% of the total of all 4 nucleotides. Thus, an image corresponding to a high proportion of certain nucleotides can have more signal points (polymerase population or clusters) compared to an image corresponding to a low proportion of certain nucleotides. As an example of low or unbalanced diversity data, in a certain flow cycle, bases A, T, C, G can be approximately 1%, 2%, 1%, and 95% of the total polymerase population, respectively. Subsequently, the flow cell image from the channels corresponding to A, T, and C in that particular flow cycle is darker and / or has far fewer polymerase population or clusters than the flow cell image corresponding to nucleotide G. As another example of low or unbalanced diversity data, bases A, T, C, G in the polymerase population in multiple flow cycles can be approximately 2%, 5%, 10%, and 83%, respectively. In embodiments where low or unbalanced diversity data is present in a particular cycle and the data is imaged for sequencing analysis, using prior art image registration and subsequent base identification can fail because the image from one or more channels is too dark compared to the images acquired from other channels (e.g., the signal points of the polymerase population are too sparse and / or dim), causing problems in sequencing analysis. Additionally, in embodiments where low or unbalanced diversity data is present in a particular cycle, existing methods may not be able to identify dimers or sparser polymerase population or clusters from other background information (e.g., cell structure). In some embodiments, method 600 is configured to perform base identification of a 3D sample even when the polymerase population or clusters have low diversity in some regions of the flow cell during some flow cycles from the 3D sample.
[0168] In some embodiments, method 600 is performed at least in part during cycle N that is different from a reference cycle. A template image and / or a 3D polymerase colony map may be generated in the reference cycle, and the polymerase colonies of one or more channels within the reference cycle may be included in the template image and / or the 3D polymerase colony map in a reference coordinate system, while base calling for cycle N has not been performed. In some embodiments, cycle N is the current cycle. N can be any non-zero integer. For example, for short read sequencing, N can be any integer from 1 to 150.
[0169] Figure 8B A schematic diagram showing one or more template images generated in a reference coordinate system in a reference cycle is shown. In some embodiments, the template image 210 has approximately the same size as a single tile 210 that includes a 5×5 grid of sub-tiles. In some embodiments, the template images disclosed herein may be separate regions within a sub-tile. Each template image may include a plurality of polymerase colonies therein.
[0170] In some embodiments, the template image may have approximately the same size as the flow cell image, such that Figure 8B the different pictures 210 from and all the polymerase colonies from multiple channels can be registered with the same template image. However, such a template image may contain polymerase colonies that are not used in at least some of the operations described herein to reduce the computational burden without sacrificing accuracy.
[0171] In some embodiments, more than one template image may be generated for the same axial position, and each template image corresponds to at least a portion of a sub-tile of the flow cell image from a channel.
[0172] The template images herein may be initialized as virtual images that have a black or dark background and no signal from polymerase colonies. For example, the template image may be initialized to zero, or include a minimum image intensity at all pixels.
[0173] After determining the coordinates of the polymerase colonies by, for example, image registration of flow cell images across different channels, the intensity of the polymerase colonies may be added to the template image at the positions determined by the coordinates, and their size and shape may be determined based on the registration. The template image may be a virtual image that combines the image intensities of polymerase colonies obtained from 2, 3, 4, or more channels in the reference cycle. The pixels of the template that do not contain polymerase colonies remain black or dark, such that the template image may have a cleaner background without the noise present in the actual flow cell image.
[0174] In some embodiments, method 600 includes an operation 610 of generating an axial stack of 2D flow cell images. Figures 6A - 6BIn two different embodiments of method 600, operation 610 may be the same. The flow cell image may include the intensity of polymerase colonies or clusters from a 3D sample. The intensity may be used for 3D base identification for different sequencing cycles. The operation of generating a 2D flow cell image may be performed by the sequencing system 110 herein.
[0175] Method 600 may include an operation of generating a processed image of the flow cell image. In some embodiments, the operation of generating a processed image may include processing the flow cell image with one or more predetermined processing methods.
[0176] In some embodiments, one or more processing methods may include selecting a kernel and generating a processed image by performing an operation on the flow cell image using the selected kernel. For example, the operation performed may be an opening operation, which may be represented where f is the flow cell image and k is the kernel. The opening operation may be performed in the spatial domain. Alternatively, to make the calculation faster or simpler, the opening operation may be performed in a different domain, such as the Fourier domain. The opening operation may be the dilation of the erosion of the image f by the kernel k. The opening operation may remove objects smaller than the kernel, and the subsequent dilation operation may restore the size and shape of the remaining objects. Figure 2B and Figure 3B shows an exemplary processed image after the opening operation. The processed image may also be referred to as the image after the opening operation, in which most of the bright spots of the polymerase colonies or clusters are removed while other background objects are retained.
[0177] As another example, the operation performed may be a convolution, and the flow cell image may be convolved with the selected kernel. As yet another example, obtaining a plurality of processed images further includes: selecting a first kernel and a second kernel; generating a first image by convolving a plurality of flow cell images with the first kernel; and generating a second image by convolving a plurality of flow cell images with the second kernel. The first image and the second image may be different blurred images after convolution with different blurred kernels.
[0178] The first or second kernel may be of any size smaller than the size of the corresponding flow cell image. For example, by the opening operation, the kernel may be 2x2, 3x3, 4x4, 5x5, 6x6 pixels. In some embodiments, the kernel size may be customized to remove at least some noise and unwanted signals larger than the kernel size. In some embodiments, the kernel may be circular. The kernel may be of various other shapes, such as oval, square, rectangle, rhombus, etc.
[0179] In some embodiments, the kernel is a Gaussian kernel. In embodiments where 2 different kernels are used, the first kernel and the second kernel may be different Gaussian kernels.
[0180] In some embodiments, method 600 may include an operation 620 of filtering a flow cell image. Figures 6A - 6B In two different embodiments of method 600, operation 620 may be the same. The filtering operation 620 may be based on the processed image generated in its previous operation. The filtering operation 620 may be based on a predetermined filter. The filtering operation 620 may generate a plurality of filtered images, each filtered image corresponding to a flow cell image at a corresponding z-level along the axial axis.
[0181] In some embodiments, operation 620 may include subtracting the processed image from the corresponding flow cell image to generate a filtered image. The filtered image may be obtained according to where fi represents the filtered image, f represents the flow cell image, and k represents the kernel, and represents an operation, such as an opening operation.
[0182] In some embodiments, a filtered image is generated in operation 620 based on the corresponding flow cell image before filtering and at least one or more image filters.
[0183] In some embodiments, the filtered image is an image of the corresponding flow cell image filtered by a predetermined filter. In some embodiments, the predetermined filter is a top-hat filter. In some embodiments, the predetermined filter is a Difference of Gaussians (DoG) filter, and the filtered image is an image of the corresponding flow cell image filtered by the Difference of Gaussians (DoG) filter.
[0184] In some embodiments, the predetermined filter may include various filters that are configured to retain small elements and details (e.g., focused polymerase colonies or clusters) in the flow cell image and optionally remove elements and details larger than the focused polymerase colonies or clusters.
[0185] After filtering, at least a portion of the noise and undesired signals in the flow cell image are removed, which may include cellular components and out-of-focus polymerase colonies. Removing such noise and undesired signals advantageously facilitates generating intensities attributable to polymerase colonies or clusters rather than other background objects or noise. The filtered intensities may be used for more accurate and reliable base calling than the unfiltered intensities.
[0186] Figure 2C and 3C shows two filtered images from two different axial positions. Figure 2CIt shows that at the first axial position where z = 0, a large number of polymerase communities or clusters are in focus, and at the second axial position where z = 5, a much smaller number of polymerase communities or clusters are in focus. The second axial position is about 10 um away from the first axial position.
[0187] Figure 3E shows Figure 3D Another exemplary filtered image of the flow cell image in []. Background objects in the flow cell image are filtered out in the filtered image. Out-of-focus polymerase communities in the flow cell image are also removed from the filtered image. The filtered image shows focused polymerase communities or clusters.
[0188] In some embodiments, method 600 may include an operation of adding an offset to the filtered image. In some embodiments, the offset may be predetermined such that after the offset, the range of image intensity may be within a predetermined range. In some embodiments, different offsets may be used to bring different filtered images with various intensity ranges into a similar predetermined range as all the filtered images.
[0189] In some embodiments, the processed image includes a first plurality of processed images and a second plurality of processed images, and the filtered image includes a first plurality of filtered images and a second plurality of filtered images. In some embodiments, the first plurality of processed images and the first plurality of filtered images are from one or more reference cycles and different channels. In some embodiments, the second plurality of flow cell images, the second plurality of processed images, and the second plurality of filtered images are from one or more reference cycles and different channels. In some embodiments, the second plurality of flow cell images, the second plurality of processed images, and the second plurality of filtered images are from one or more cycles different from one or more reference cycles and different channels. In some embodiments, the first plurality of flow cell images and the second plurality of flow cell images are the same, the first plurality of processed images and the second plurality of processed images are the same, and the first plurality of filtered images and the second plurality of filtered images are the same.
[0190] Method 600 may include an operation 635 of obtaining a 3D polymerase community map. The 3D polymerase community map includes the corresponding 3D positions of polymerase communities or clusters, and the 3D polymerase community map can be used as a roadmap for locating all possible 3D polymerase communities for 3D base identification. In other words, base identification is only performed on the polymerase communities included in the 3D polymerase community map. The 3D polymerase community map may include all of the polymerase communities in the sample within the 3D FOV. In some embodiments, the 3D polymerase community map may include only some of the polymerase communities in the sample within the 3D FOV.
[0191] A 3D polymerase community map can be determined in one or more reference cycles. During a cycle different from the reference cycle, for example, in the reference cycle, the 3D polymerase community map may have been pre-generated and can be obtained by actively requesting or retrieving or passively receiving the 3D polymerase community map. In some embodiments, a single 3D polymerase community map is used during the sequencing analysis of the same 3D sample such that in a cycle different from the reference cycle, there is no need to generate a new 3D polymerase community map. In some embodiments, the 3D polymerase community map can be updated after multiple cycles during the sequencing of the same 3D sample.
[0192] Continuing reference Figure 6B , method 600 can include operation 635 of generating a 3D polymerase community map. The operation of generating the 3D polymerase community map can occur during one or more reference cycles, but is not limited to this. As a non-limiting example, cycles 1-5 or 2-5 can be reference cycles. In some embodiments, the 3D polymerase community map generated in one or more reference cycles can be used in any cycle different from the reference cycle such that the additional computational complexity, cost, and time of generating a 3D polymerase community map in each cycle can be avoided.
[0193] The 3D polymerase community map can be generated based on one or more 2D template images. Each 2D template image in the 2D template images corresponds to a flow cell image at a specific z position. Each 2D template image in the 2D template images corresponds to a flow cell image at a specific z level and at a specific tile or sub-tile. Each 2D template image in the 2D template images can correspond to one or more sequencing cycles. Each template image can correspond to one or more channels. For example, if there are 10 different z levels covering the 3D sample being sequenced, during cycle N, for a single tile or sub-tile, there can be 10 2D template images corresponding to different z levels. The 2D template images can be in the same coordinate system as the 2D polymerase community map. The 2D template images can contain exactly the same polymerase community as the 2D polymerase community map. In some embodiments, the 2D template images can contain a polymerase community different from the 2D polymerase community map because the template images can cover a smaller area than the 2D polymerase community map to reduce computational complexity and save sequencing analysis time. In some embodiments, the 2D template images can contain a polymerase community different from the 2D polymerase community map because the template images can contain duplicate polymerase communities that may not be included in the 2D polymerase community map.
[0194] The 2D or 3D polymerase colony maps in this article can be saved as a list of coordinates. Each entry in the list of coordinates can correspond to a polymerase colony, such as the center of a polymerase colony. Instead of saving a 2D or 3D matrix as a polymerase colony map, the list of coordinates can be stored with much less storage space and can be utilized more efficiently in calculations.
[0195] In some embodiments, operation 635 of generating a 3D polymerase colony map can include obtaining a 2D template image. Obtaining the template image can include generating the template image or receiving or retrieving the template image. A 2D template image can be generated using various methods (including the methods and operations described herein). The 2D template image can include polymerase colonies with sub-pixel resolution. Exemplary embodiments of obtaining a 2D template image are disclosed in U.S. Patent Application Nos. 18 / 078,797 and 18 / 078,820, the contents of which are hereby incorporated by reference in their entirety. The 2D template image can be generated after filtering, such as by a top-hat filter or a DoG filter. The 2D template image can be generated after filtering the flow cell image and registering the filtered image with a reference coordinate system. The 2D template image can be generated in one or more reference cycles, and the same template image can be used across different cycles and channels. The 2D template image can include a list of coordinates of polymerase colonies. For example, each entry in the list can be the 2D or 3D coordinates of a polymerase colony.
[0196] In some embodiments, the operation of generating a 3D polymerase colony map can include combining one or more 2D template images into a candidate 3D polymerase colony map. For example, the lists of coordinates in the 2D template images can be added together.
[0197] In some embodiments, the operation of generating a 3D polymerase colony map can include combining one or more 2D polymerase colony maps into a candidate 3D polymerase colony map. In some embodiments, the operation of generating a 3D polymerase colony map can include combining one or more 2D template images into a candidate 3D polymerase colony map. In some embodiments, the 2D template image and the 2D polymerase colony map may differ in excluding duplicate polymerase colonies (3D) using only 2D information. In some embodiments, the operation of generating a 3D polymerase colony map can include combining one or more 2D polymerase colony maps or one or more 2D template images into a candidate 3D polymerase colony map, depending on which 2D polymerase colony maps or 2D template images include a more complete set of possible polymerase colonies of the sample to avoid erroneously excluding polymerase colonies before removing duplicate polymerase colonies.
[0198] In some embodiments, the 3D polymerase colony map is saved as a list of entries, each entry representing a polymerase colony and including spatial and intensity information of the polymerase colony. In some embodiments, when the 3D polymerase colony map is saved as a 3D matrix instead of a list of entries, the operation of generating the 3D polymerase colony map may include extracting polymerase colonies in one or more 2D template images. Instead of directly combining the 2D template images, polymerase colonies in one or more template images may be extracted, and the extracted polymerase colonies may be included in a candidate 3D polymerase colony map based on their coordinates in the template images.
[0199] The candidate 3D polymerase colony map may include a 3D resolution determined by the resolution in the image plane (i.e., the x-y plane) and the resolution along the z-axis. The resolution along the z-axis may be determined by the slice thickness of the flow cell image. The resolution along the z-axis may be determined by the depth of field of the optical system such that polymerase colonies within the depth of field can be focused. In some embodiments, sub-pixel resolution along x, y, and / or z may be achieved. In some embodiments, the sub-pixel resolution is 2 times, 4 times, 5 times, 6 times, 8 times, 10 times higher than the pixel resolution along the corresponding axis. For example, the pixel resolution may be 3 um, and the 10-fold sub-pixel resolution is 0.3 um. The sub-pixel resolution may be achieved using the methods disclosed in U.S. Patent Application Nos. 18 / 078,797 and 18 / 078,820, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, various interpolation methods may be used to achieve sub-pixel resolution based on the pixel resolution at different z levels such that the focused polymerase colony position may be between two adjacent z levels, and these two adjacent z levels are determined by interpolation of the pixel intensities of the corresponding pixels at the two adjacent z levels (e.g., the pixel (x1, y1) at z1 and the pixel (x1, y1) at z2), optionally using pixel intensities from other z levels (e.g., at z = 0 and z = 3).
[0200] The candidate 3D polymerase colony map may be initialized with 0 or other predetermined intensity values in all of its pixels at different z positions. The extracted intensities may be used to replace the initial values in the candidate 3D polymerase colony map to indicate pixels or voxels that are at least part of a polymerase colony. As a non-limiting example, the 3D polymerase colony map may be a 3D matrix of 0s and 1s, where each pixel having a 1 indicates that the pixel is part of a polymerase colony.
[0201] Individual polymerase colonies in a 3D sample may appear at one, two, or even more z-levels, such that the same polymerase colony may be included in multiple flow cell images at different z-positions. For example, a polymerase colony at (x1,y1) at position z1 may be included again at (x1-1,y1-1) at position z2. Thus, a candidate 3D polymerase colony map may include duplicate polymerase colonies. Duplicate polymerase colonies need to be removed for accurate and reliable 3D base identification. Operation 635 of generating a 3D polymerase colony map may include removing duplicate polymerase colonies from the candidate 3D polymerase colony map.
[0202] To remove duplicate polymerase colonies, operation 635 may include: performing preliminary base identification. The position of the polymerase colony for base identification may be determined by a 2D template image, and the intensity for performing base identification may be extracted from the filtered image obtained in operation 620. Similarly, the position of the polymerase colony for base identification may be determined based on a candidate 3D polymerase colony map that may include all 2D template images. Filtering herein may advantageously remove intensity interference from out-of-focus polymerase colonies and background objects, such that the intensity can be used for more accurate and reliable base identification than unfiltered flow cell images. In some embodiments, the 2D template image contains the coordinates of the polymerase colony in a reference coordinate system, such that even if the polymerase colony may have been displaced across cycles, base identification can still be attributed to the same polymerase colony.
[0203] After obtaining the preliminary base identification, operation 635 may include a repetition operation of removing duplicate polymerase colonies from the 2D template image or from the candidate 3D polymerase colony map until a stop criterion is met. The repetition operation may include identifying candidate polymerase colonies with the same base identification. In response to identifying at least two candidate polymerase colonies with the same base identification, the candidate polymerase colony may contain zero, one, two, or more duplicate polymerase colonies. Operation 635 may further include an operation of determining the 3D distance between each pair of polymerase colonies among the candidate polymerase colonies. For each pair of polymerase colonies (non-repetitive pairs), the 3D distance may be calculated based on the coordinates of the polymerase colonies. The coordinates may be 3D. The coordinates may include, for example, the 2D coordinates of the polymerase colony after registration in a reference coordinate system and the z-level of the polymerase colony. The 3D distance may be in pixels. The 3D distance may be in other units (e.g., um). The 3D distance may be used to determine whether two polymerase colonies with the same base identification are close to each other.
[0204] In response to determining that the 3D distance between two polymerase communities is within a predetermined distance threshold, operation 635 may include determining the image intensity of each of the two polymerase communities from the filtered image or the template image. In some embodiments, determining the image intensity for each of the polymerase communities may include normalizing and / or offsetting the image intensities at different z-levels to a predetermined range. For example, the normalization and / or offset may be based on the intensity of fiducial markers at different z-positions. If two identical fiducial markers at two different z-levels have different intensities, the normalization and / or offset may be used to bring the two different intensities to the same level. This normalization and / or offset may then be applied to the image intensity of the polymerase community at the corresponding z-level that serves as a fiducial marker. Subsequently, operation 635 may include removing one of the two polymerase communities that has a smaller image intensity. Two polymerase communities within the predetermined distance threshold may be considered as repeated identical polymerase communities. The repeat with the smaller intensity may be more out of focus than the repeat with the larger intensity and may be removed to ensure accurate and reliable base identification. The predetermined distance threshold may be customized based on characteristics of the sample, imaging parameters, polymerase communities, etc. For example, the predetermined distance threshold may be based on the depth of field of the optical system, the distance between two adjacent flow cell images along the axial distance, or a combination thereof. The distance threshold may be adjusted individually or in combination with the stop criterion to balance true and false repeats being removed. For example, polymerase community p1 has a preliminary base identification of A determined using its filtered image intensity, and its position after registration with the reference coordinate system may be at (xp1, yp1). Candidate polymerase communities may be determined to be those that have the same base identification as A. The 3D distance from each candidate polymerase community to polymerase community p1 may be calculated and compared with a predetermined threshold (e.g., a predetermined threshold of 0.5 um). The polymerase community p2 that meets the distance threshold may be considered a possible repeat of p1. The intensities of p1 and p2 are compared, and the polymerase community p2 with the greater intensity is retained while the coordinates of polymerase community p1 are removed as a repeat.
[0205] After removing duplicate polymerase colonies from the 2D template image or candidate 3D polymerase colony map, a 3D polymerase colony map is generated, which includes all polymerase colonies capable of performing base recognition and corresponding coordinates. The 3D polymerase colony map can include location information of such polymerase colonies. For example, the 3D polymerase colony map can include the coordinates of each polymerase colony in the corresponding filtered image. The 3D polymerase colony map can further include the z-level of each polymerase colony, which can be the central position of the polymerase colony extended over several pixels. The 3D polymerase colony map can include the size and / or shape of the polymerase colony. The 3D polymerase colony map can include a unique identifier for each polymerase colony. The 3D polymerase colony map can include the image intensity of the polymerase colony. Such image intensity can be the filtered intensity obtained from the filtered images disclosed herein.
[0206] The stopping criteria herein can be customized. For example, the stopping criteria can be based on different types of samples, imaging parameters, the size and shape of polymerase colonies, etc. The stopping criteria, distance threshold, or both can be adjusted to balance true and false duplicates removed. As a non-limiting example, the stopping criteria can be to remove the top 100 duplicate polymerase colonies. As a non-limiting example, the stopping criteria can be to perform duplicate removal only within a selected time window. As yet another example, the stopping criteria can be the absence of duplicates that meet a predetermined distance threshold.
[0207] Details of exemplary embodiments of methods for generating 3D polymerase colony maps are described in U.S. Patent Application Nos. 18 / 078,797 and 18 / 078,820, the contents of which are incorporated herein by reference in their entirety.
[0208] In some embodiments, method 600 includes the operation of registering the flow cell image, the processed image, and / or the filtered image. In some embodiments, the images are registered across channels and different cycles. In some embodiments, the images are registered before performing any base recognition. In some embodiments, the images are registered across channels and different cycles before generating the template image or obtaining the 3D polymerase colony map. In some embodiments, the images are registered across channels and different cycles before generating the filtered image or the processed image.
[0209] Various image registration techniques can be used to register images. 2D registration techniques can be used to register images. When registering processed or filtered images, they can be treated as flow cell images during registration. In some embodiments, an image registration method 900 as disclosed herein can be used to register images by treating each image to be registered as a flow cell image that can be obtained using a sequencing system 110, e.g., across different channels and / or different cycles. Exemplary embodiments of the image registration method are described in PCT patent application number PCT / US23 / 67931, the content of which is incorporated by reference in its entirety.
[0210] In some embodiments, images can be registered after performing one or more preprocessing operations disclosed herein. In some embodiments, the operation of registering flow cell images, processed images, and / or filtered images can occur before the operation 635 of obtaining a 3D polymerase colony map.
[0211] In some embodiments, the operation of registering flow cell images, processed images, and / or filtered images is with respect to a reference coordinate system. In some embodiments, the operation of registering flow cell images, processed images, and / or filtered images is with respect to one or more template images. The operation of registering images can include generating one or more template images in the reference coordinate system. In some embodiments, the operation of registering images can include registering polymerase colonies with template polymerase colonies in one or more template images. The operation of registering images can include determining a plurality of transforms based on one or more template images. Each transform in the plurality of transforms can correspond to a corresponding sub-tile of a flow cell image, processed image, or filtered image and is configured to register the sub-tile with one or more template images. Each transform can be used to register the corresponding sub-tile or tile with one or more template images. The plurality of transforms can include one or more affine transforms.
[0212] In some embodiments, the operation of registering images can include performing image registration of polymerase colonies based on fiducial markers. The fiducial markers can be located on the flow cell. Alternatively, the fiducial markers can be external to the flow cell.
[0213] In some embodiments, image registration, which is a main analysis step herein, is configured to align images from different cycles and / or different channels, e.g., with respect to a template image or a reference coordinate system. In some embodiments, image registration, which is a main analysis step herein, is configured to register polymerase colonies or clusters from different cycles and different channels (e.g., in a filtered image) with a template image or a reference coordinate system.
[0214] For example, after registering filtered images from different channels relative to corresponding template images disclosed herein, base calling can be performed using the filtered images from different channels in cycle N.
[0215] Method 600 can include operation 645 of extracting polymerase colony intensity based on a 3D polymerase colony map. For each polymerase colony in the 3D polymerase colony map, position information of such polymerase colony can be obtained from the 3D polymerase colony map, such as the 2D coordinates and z position of the polymerase colony. Using the 2D coordinates and z position, the corresponding filtered image and its pixels can be determined. The image intensity of such pixels can be extracted from the corresponding filtered image as the intensity of such pixels for performing base calling.
[0216] In some embodiments, multiple neighboring polymerase colonies may at least partially overlap in a 3D polymerase colony map, and they may not be resolvable along the x, y, and / or z directions using existing methods. In some embodiments, operations 645 and / or 655 herein advantageously utilize pre-determined base identification information to determine the polymerase colony intensities of neighboring polymerase colonies that are too close to each other (e.g., partially overlapping). The pre-determined base identification information may include expected image intensities from two or more channels. The pre-determined base identification information may include expected image intensities from two or more cycles. The pre-determined base identification information can be used to resolve at least partially overlapping polymerase colonies and their intensities by determining a linear combination of the measured intensities of the overlapping polymerase colonies to match the expected pixel intensities for a given possible DNA sequence (e.g., barcode). For example, in a first color channel, the image intensities obtained in 3 consecutive cycles are 0, 1, 3, while at the same pixel (m,n) in the same cycles, the intensities observed in a second color channel are 1, 2, 0. The expected signal intensities include: for a first nucleotide base sequence (e.g., the first barcode sequence), 0, 1, 1 in channel 1 and 1, 0, 0 in channel 2 for the same 3 cycles; and the expected signal intensities also include: for a second nucleotide base sequence (e.g., the second barcode sequence), 0, 0, 1 in channel 1 and 0, 1, 0 in channel 2. Thus, for the pre-determined barcode information (e.g., only the first and / or second barcode may be present in this pixel), the relative image intensities of the first barcode sequence and the second barcode sequence can be determined using the following equations: A*[0 1 1]+B*[0 0 1] = [0 1 3]; and A*[1 0 0]+B*[0 1 0] = [1 2 0], where A and B are numerical values. The values of A and B can determine the relative intensities of base identification from the first barcode sequence and the second barcode sequence. In this particular example, B is twice A, indicating that the contribution of the second barcode sequence to the image intensity of pixel (m,n) is twice that of the first barcode sequence. This determination of the linear combination of possible DNA sequences can be at a single pixel or multiple pixels. The same determination can be repeated at neighboring pixels or sub-pixels to confirm the same relative intensities. In response to determining neighboring pixels where different relative intensities of base identification from the first barcode sequence and the second barcode sequence can be obtained, pixels or sub-pixels closer to the center of the polymerase colony under discussion can be weighted. In response to determining neighboring pixels where different relative intensities of base identification from the first barcode and the second barcode can be obtained, the intensity decomposition can be marked as a failure.
[0217] In some embodiments, the intensities obtained from the flow cell images can be rounded, e.g., to the nearest integer multiple of the expected intensity value in that channel, to discretize the intensities. In some embodiments, before resolving at least partially overlapping base calls, base call sequences that are not possible at a position (i.e., not "on" in the channel), e.g., barcodes, can be removed. In some embodiments, such filtering and removal of impossible barcodes can be performed when the complexity of the data is greater than a predetermined threshold (e.g., 16, 20, 24, 30 or more).
[0218] In some embodiments, the equations for each channel can be solved separately, and all possible combinations of different base call sequences can be determined. Then, the possible solutions from different channels can be combined to determine a solution that satisfies the equations for at least some or all of the channels.
[0219] In some embodiments, image intensity quality metrics (such as sharpness, purity, purification degree, quality score, etc.) can be used to identify pixel positions that may require a decomposition operation to avoid decomposing all pixels (including pixels with only a single base call).
[0220] Method 600 can include an operation 655 of performing 3D base calling using the extracted image intensities. Various existing 2D base calling algorithms can be used. The base calling results can be saved together with their 3D position information (e.g., 3D coordinates). Such 3D coordinates can be used to register base calls across different cycles and at different z positions, and / or to register base calls with the pool images herein.
[0221] In some embodiments, the operation 655 of performing base calling includes performing the preliminary analysis steps herein to adjust the image intensity of the polymerase population in the filtered image; and performing base calling on the polymerase population based on the adjusted image intensity in the filtered image. The adjustment of the image intensity in the filtered image can be performed before or after filtering.
[0222] In some embodiments, the primary analysis steps include one or more of the following: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; trailing and leading correction; quality score estimation.
[0223] In some embodiments, for example, Figures 6A - 6B Method 600 in includes: before operation 610, providing a cell sample containing a plurality of RNAs, the plurality of RNAs including at least a first target RNA molecule and a second target RNA molecule.
[0224] In some embodiments, for example, Figures 6A - 6BThe method 600 in [the above] includes: before operation 610, generating a plurality of concatemer molecules, the plurality of concatemer molecules including at least a first concatemer molecule corresponding to a first target RNA molecule, and the plurality of concatemer molecules including at least a second concatemer molecule corresponding to a second target RNA molecule. In some embodiments, the operation of generating a plurality of concatemer molecules includes one or more of the following: generating a plurality of cDNA molecules inside a sample, the plurality of cDNA molecules including at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules including a second target cDNA molecule corresponding to the second target RNA molecule; contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes, the target-specific padlock probes including at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes; generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe inside the sample by performing an enzymatic reaction to close the nicks or gaps in at least the first and second circularized target-specific padlock probes; using the first and second covalently closed circular padlock probes as template molecules to perform a rolling circle amplification reaction inside the sample, thereby generating a plurality of concatemer molecules, the plurality of concatemer molecules including at least a first concatemer molecule corresponding to the first target RNA molecule, and the plurality of concatemer molecules including at least a second concatemer molecule corresponding to the second target RNA molecule; and sequencing the plurality of concatemer molecules inside the sample, which includes: sequencing the first concatemer molecule by performing no more than 2 to 1000 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second concatemer molecule by performing no more than 2 to 1000 sequencing cycles to generate a plurality of second sequencing read products.
[0225] In some embodiments, the operation of sequencing the plurality of concatemer molecules inside the sample includes: contacting the plurality of concatemer molecules inside the sample with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of universal sequencing primers to their corresponding universal sequencing primer binding sites on the concatemers. The nucleotide reagents may include one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs.
[0226] In some embodiments, sequencing the plurality of concatemer molecules inside the sample further includes removing the plurality of first sequencing read products from the first concatemer molecule and retaining the first concatemer molecule in the sample, and removing the plurality of second sequencing read products from the second concatemer molecule and retaining the second concatemer molecule in the sample.
[0227] In some embodiments, for example, Figures 6A - 6BThe method 600 in [the above context] further includes: before operation 610, providing a sample having a plurality of tandem molecules immobilized on a support. Each tandem molecule can correspond to a target RNA of the cell sample. The support can be included in the flow cell disclosed herein.
[0228] In some embodiments, the operation 610 of obtaining a plurality of flow cell images of the sample includes: by the sequencing system, generating a plurality of flow cell images by performing one or more sequencing reaction cycles on the sample immobilized on the support. The plurality of flow cell images herein can be generated from two or more color channels at two or more different z-levels along the axial axis. The plurality of flow cell images herein can be generated from 3 to 1000 different z-levels, and these z-levels can completely cover the 3D volume of the 3D sample. Given the same 3D sample, the resolution along the z-axis can be determined by different numbers of z-levels.
[0229] In some embodiments, the operation 610 of obtaining a plurality of flow cell images of the sample includes: by the sequencing system, generating a plurality of flow cell images by performing one or more sequencing reaction cycles on a plurality of tandem molecules of the sample immobilized on the support. The sample herein can be a cell sample. The sample can include a polymerase community or cluster immobilized thereon. In some embodiments, the polymerase community corresponds to a plurality of nucleotide template molecules or tandem molecules. In some embodiments, the operation of performing one or more sequencing reaction cycles includes: contacting the plurality of tandem molecules with a plurality of nucleotide reagents containing a mixture of different types of nucleotide bases A, G, C, and T / U. In some embodiments, the operation of performing one or more sequencing reaction cycles includes: contacting the plurality of tandem molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of affinity agents. The mixture of different types of affinity agents can include 2, 3, 4, or more types of affinity agents. The number of different types of affinity agents can match the number of different color channels of the sequencing system. The number of different types of affinity agents can be less than the number of different color channels of the sequencing system. A single affinity agent in the mixture can include a core attached with a plurality of nucleotide arms, and each arm of a single affinity agent contains the same type of nucleotide base. In some embodiments, the operation of performing one or more sequencing reaction cycles includes: in each of one or more cycles, imaging the optical color signals emitted by the nucleotide reagents bound to the plurality of tandem molecules by the optical system of the sequencing system. In some embodiments, the operation of performing one or more sequencing reaction cycles includes: in each of one or more cycles, acquiring a flow cell image containing the optical color signals emitted by the nucleotide reagents bound to the plurality of tandem molecules by the optical system.
[0230] In some embodiments, the flow cell images herein may each include optical signals emitted from nucleotide reagents that incorporate an unbalanced diversity of nucleobases A, G, C, and T / U from multiple templates or tandem molecules immobilized on a support in one or more cycles.
[0231] In some embodiments, for example, a plurality of polymerase colonies corresponding to bright spots in the flow cell image contain an unbalanced diversity of nucleobases A, G, C, and T / U, and wherein the unbalanced diversity comprises the percentage of (1) the number of one or more types of nucleobases within a region of the flow cell image to (2) the total number of nucleobases within that region, and the percentage is less than 20%, 15%, 10%, or 5% in one or more cycles. The region may include all of the FOV of the flow cell image. In some embodiments, the region may include only some of the FOV to separate regions within the FOV having spatial differences in nucleotide diversity.
[0232] Cell images and staining
[0233] In some embodiments, base identification is registered to one or more cell images herein. The one or more cell images may include images of cells and / or tissues with one or more stains (e.g., fluorescent stains). In some embodiments, the one or more images may include staining of cell structures that aids in localizing polymerase colonies or clusters relative to the stained structures. For example, the stain may be a stain of a cell structure or component (including but not limited to membranes, nuclei, and mitochondria). Different staining colors may be used to stain different components of the cell.
[0234] In some embodiments, the cell membrane may be permeabilized after sequencing analysis and imaging using the sequencing system and reaction. In some embodiments, one or more pool images may include staining of lipids (such as lipids contained in the cell membrane). In some embodiments, instead of labeling the lipids, one or more pool images may include staining of one or more transmembrane proteins. The transmembrane protein may be a protein embedded in the permeabilized membrane.
[0235] In some embodiments, one or more pool images include fluorescence signals from the cell membrane. The one or more pool images may be microscopic images. The one or more images may be fluorescence images. In some embodiments, different fluorescence colors may be included in the pool image. For example, the nucleus and the cell membrane may be stained with different colors.
[0236] In some embodiments, the one or more images may include segmentation of cells, membranes, nuclei, or combinations thereof. Figure 7BShows an exemplary pooled image of a segmentation having individual cells. In some embodiments, the edges of each segment encompass the entire membrane of the cells within the segment. There may be only one cell in each segment. There may be no cells in some segments. In some embodiments, adjacent segments do not overlap with each other. In some embodiments, adjacent segments overlap with each other only by sharing one or more edges. In some embodiments, various segmentation algorithms may be used to segment the cells.
[0237] In some embodiments, the pooled images disclosed herein are stained. The staining may occur after the flow cell image is acquired using the sequencing system 110. In some embodiments, the staining may be performed before the sequencing image is acquired. Methods for staining 3D samples (such as cells, tissues) may include one or more operations disclosed herein. The staining of 3D samples may use various methods that can specifically label one or more cellular proteins that are mainly located in the membrane but have negligible presence (e.g., amount or concentration less than 10%, 5%, 2%) in other regions of the cell.
[0238] In some embodiments, the cell image can be acquired using the sequencing system 100 herein without moving the sample from its position during sequencing. It is advantageous to stain the sample and acquire the cell image after sequencing while the sample is fixed on the sample stage of the sequencing system. Some transformations may still occur, such as rotation, translation, shear, so it is necessary to register the flow cell image during sequencing to the cell image acquired after sequencing and staining. In some embodiments, the cell image can be acquired using an optical device external to the sequencing system 100 after the sequencing run is completed and the sample is removed from the sequencing system 100.
[0239] The operations may include selecting one or more primary antibodies, each of the one or more primary antibodies specifically binding to a corresponding protein. The corresponding protein may be a transmembrane protein of one or more cells. In some embodiments, the corresponding transmembrane protein does not exist at a predetermined concentration in other cell regions (such as the cytoplasm or nucleus) such that the staining of the transmembrane protein does not produce a perceptible signal in cell regions other than the membrane. In some embodiments, one or more different transmembrane proteins may be labeled with primary antibodies. For example, if there are 5 different types of transmembrane proteins, 5 different primary antibodies may be used, and each primary antibody specifically binds to one of the transmembrane proteins but not to the other transmembrane proteins. In some embodiments, the same type of primary antibody may non-specifically bind to different proteins.
[0240] The staining method may include the operation of selecting one or more secondary antibodies that bind to one or more primary antibodies. The staining method may further include the operation of labeling one or more secondary antibodies with a fluorescent label.
[0241] In some embodiments, the staining method may further include the operation of using a support element or a tertiary probe (such as a hydrogel) to connect the secondary antibody and the fluorescent label. The support element can be used to retain the mRNA of the membrane for facilitating binding and the generation of fluorescent signals. In some embodiments, the mRNA can be any mRNA in the cell. In some embodiments, the staining method may include using various methods for tissue clearing to remove some or all parts of the cell, thereby reducing the background fluorescence from the non-membrane parts of the cell. In some embodiments, the fluorescent label includes a fluorophore that re-emits light within a specific wavelength range after photoexcitation. FIG. 8 shows an exemplary staining of a transmembrane protein using the staining method disclosed herein.
[0242] The staining method may further include the operation of generating one or more cell images of the corresponding protein. The one or more images include the fluorescent signals emitted from the fluorescent label.
[0243] In some embodiments, the method for base calling in sequencing data analysis may include generating a 3D polymerase colony map based on a plurality of filtered images. The 3D polymerase colony map may include a stack of 2D polymerase colony maps, each 2D polymerase colony map corresponding to a flow cell image acquired at a corresponding axial position. The 3D polymerase colony map may include some or all of the polymerase colonies that can be identified in the axial stack of the flow cell images.
[0244] In some embodiments, each 2D polymerase colony map may be generated based on the filtered images at the same axial position. In some embodiments, the 2D polymerase colony map may be generated from the filtered images, similar to generating a template image from a flow cell image. In some embodiments, the 2D polymerase colony map is equivalent to the template image because both of them are virtual images that include all the polymerase colonies identified in one or more cycles at a specific axial position.
[0245] As disclosed herein, the filtered images can advantageously exclude background objects that may interfere with the signals from the polymerase colonies. The filtered images can also remove some out-of-focus polymerase colonies or clusters. The filtering parameters (such as the size and shape of the kernel) can be customized to balance between removing out-of-focus polymerase colonies and retaining relatively large polymerase colonies or clusters of assembled out-of-focus polymerase colonies.
[0246] In some embodiments, each 2D polymerase colony map is registered to a template image (3D) or a reference coordinate system (3D). The template image or the reference coordinate system can be determined in a reference cycle. For each cycle different from the reference cycle, 2D polymerase colony maps can be generated and the polymerase colony maps from different cycles can be registered relative to each other using the template image or the reference coordinate system. In some embodiments, a single polymerase colony map can be generated for all channels having the same cycle. In some embodiments, polymerase colony maps are generated for each channel for each cycle.
[0247] In some embodiments, the 3D polymerase colony map is a volumetric polymerase colony map that stacks all 2D polymerase colony maps at different axial positions. In some embodiments, the methods herein include removing duplicates of polymerase colony maps from the stacked 2D polymerase colony maps to generate a 3D polymerase colony map. The 3D polymerase colony map without duplicates can be used as a reference to locate individual polymerase colonies in a sample for base identification. In some embodiments, the 3D polymerase colony map can be saved as a list of 3D coordinates indicating the centers of polymerase colonies.
[0248] In some embodiments, the methods herein further include an operation of extracting the image intensity of polymerase colonies based on the 3D polymerase colony map. The image intensity can be extracted from one or more of the following: flow cell image; processed image; filtered image. In some embodiments, the image intensity can be extracted from the filtered image. In some embodiments, the image intensity can be extracted from the filtered image after processing the filtered image using one or more of the main analysis steps disclosed herein. For example, a backward and forward correction can be performed on the filtered image before the image intensity can be extracted.
[0249] In some embodiments, the 3D polymerase colony map can include duplicates of polymerase colonies and such duplicates can be removed after base identification. For example, all polymerase colonies including duplicates in the 3D polymerase colony map can be used to extract image intensity for base identification. Candidate duplicates of polymerase colonies can be identified as polymerase colonies at different z positions (e.g., adjacent z positions) and the same x, y positions. If the base identifications of such candidate duplicates are the same, one of them can be removed as a duplicate. Alternatively, both of them can be removed and a new polymerase colony representing both can be added at the z level that is the average of the two, and at the same x position and y position.
[0250] In some embodiments, the methods herein further include the operation of performing 3D base identification based on the extracted polymerase colony image intensity. Various 2D base identification algorithms can be used here. For example, base identification of polymerase colonies can be performed by comparing the image intensities of the same polymerase colonies from different channels, and the base corresponding to the maximum image intensity in all channels can be identified.
[0251] In some embodiments, the operation of filtering the flow cell images to generate a plurality of filtered images can include performing deconvolution on the plurality of flow cell images. The deconvolution can be at least along the axial direction. In some embodiments, the deconvolution can be 3D. The deconvolution can perform its equivalent operation in the spatial domain or in a transformed domain (such as the Fourier domain). The deconvolution is configured to reduce or eliminate the diffusion or blurring effect of the optical system on the polymerase colonies, such that the size and shape of the polymerase colonies appear more accurate in the flow cell images. In some embodiments, the deconvolution operation can be used alone as a filtering operation or in combination with other filtering operations (such as top-hat filtering).
[0252] Image registration with tile images
[0253] Various methods can be used to register the flow cell images based on fiducial markers. The fiducial markers can be located inside or outside the sample. For example, internal fiducial markers can include at least some of the polymerase colonies or clusters or background objects in the sample. As another example, external fiducial markers can be microspheres coated on the flow cell, such that the signals from the microspheres can similarly be used as internal fiducial markers for registration. The same fiducial markers can appear in the sequencing images, for example, MIP images, flow cell images, filtered images, and cell images, such that a transformation can be derived by aligning the fiducial markers in different images. Exemplary embodiments of image registration methods are described in PCT patent application number PCT / US2023 / 067931 (wherein the content of this patent is hereby incorporated by reference in its entirety).
[0254] For example, a polymerase colony or other object (such as a background object) with an image intensity I centered at position (x1, y1) in the sequence image can appear at position (x2, y2) with intensity I’ in the cell image, where and Mr is the transformation matrix. Similarly, the inverse transformation matrix Mr can be determined -1 , such that Registration of cross-channel images can be 2D and can include translation, scaling, rotation, and / or shearing of flow cell images between different channels. Multiple points in the sequencing image and their corresponding points in the flow cell image can be used to determine the transformation. The minimum number of points required can be determined by the degrees of freedom of the transformation. In some embodiments, the image registration can be 3D, where the coordinates are on the x, y, and z axes.
[0255] In some embodiments, the sequencing image can be divided into multiple sub-tiles, and a transformation can be determined for each sub-tile to represent the transformation of the entire image. In some embodiments, the image transformation of each sub-tile can be uniquely represented by a transformation matrix. The transformation matrix can be determined as follows:
[0256]
[0257] where n is the number of sub-tiles, a1 = x1 + dx1, b1 = y1 + dy1, a2 = x2 + dx2, b2 = y2 + dy2,... an = xn + dxn, bn = yn + dyn, d1... dn are 2D shifts corresponding to the sub-tiles, and where dxn and dyn are the shift components of the 2D shift dn on the x-axis and y-axis respectively, and where M is the 3x3 transformation matrix of the sub-tile.
[0258] In some embodiments, the transformation matrix can be defined as the inverse matrix of M, i.e., M -1 , so Equation (1) can be equivalently expressed as
[0259]
[0260] In some embodiments, the transformation matrix M is estimated based on the 2D shifts in Equations (1) and (3). In some embodiments, the value of n may affect the accuracy of the estimate.
[0261] In some embodiments, more than one region can be selected within a sub-tile for cross-correlation calculation, and more than one 2D displacement can be calculated for each sub-tile and used to estimate the transformation of the sub-tile. In these embodiments, n in Equation (1) can be replaced with a larger number. For example, when 2 regions are selected for each sub-tile, it can be replaced with 2*n, and the transformation matrix M can be estimated using Equations (1) and (2).
[0262] In some embodiments, (a1,b1)...(an,bn) in Equations (1)-(3) are the coordinates of the selected regions after transformation (e.g., the coordinates of the central pixel of the corresponding region), and (x1,y1)...(xn,yn) are the coordinates of the selected regions before transformation, e.g., the coordinates of the central pixel.
[0263] In some embodiments, n is a number not less than 3. The larger n is, the more information is used to estimate the transformation matrix M. In some embodiments, n is not greater than 9.
[0264] In some embodiments, the transformation of one or more sub - tiles is linear. In some embodiments, the transformation of all sub - tiles is linear. In some embodiments, the transformation matrix is a matrix where M31 and M32 are equal to 0 and M33 is 1. In some embodiments, one or more of the transformations in the transformation of each sub - tile are affine transformations, and the transformation matrix of the entire flow cell image is an affine matrix.
[0265] In some embodiments, the transformation matrix M is an estimate in equations (1) and (3) based on the size of the selected region. In some embodiments, the size of the selected region may affect the accuracy of the estimate. In some embodiments, the size of the selected region can be about 128×128. In some embodiments, the size of the selected region can be about 32×32, 48×48, 64×64, 96×96, 160×160, 196×196, 256×256 or various different sizes. As disclosed herein, the transformation of each sub - tile can be calculated using the selected region within the sub - tile, and the selected region can be equal to or smaller than the sub - tile. In either case, considering the inherent characteristics of the image transformation across sequencing cycles, the transformation estimated using the region can be used to estimate the transformation of the entire sub - tile. The image transformation between cycles and / or between adjacent pixels can be relatively small, for example, it can be less than about 8%, 5% or less than about 1% of scaling, rotation and / or shear. In some embodiments, the transformation disclosed herein can include an image translation with a difference between cycles and / or between adjacent pixels greater than about 5%.
[0266] After determining the multiple transformations of individual sub - tiles, the transformation of the entire flow cell image can be accurately and reliably estimated by transforming the individual sub - tiles using the multiple transformations and combining the transformed sub - tiles into a transformed flow cell. The techniques disclosed herein advantageously estimate the transformation of the flow cell image by determining the multiple transformations of individual sub - tiles of the flow cell image. The multiple transformations can be linear, and even if the transformation is non - linear, the transformation of the flow cell image can be accurately and reliably estimated. The techniques disclosed herein advantageously eliminate the need to calculate the transformation of the entire flow cell image, which may be computationally more intensive, more time - consuming and more prone to failure compared to estimating the multiple transformations of sub - tiles.
[0267] Image registration across channels and cycles
[0268] In some embodiments, the method includes the operation of aligning or registering flow cell images across different sequencing cycles, from different channels, and / or at different z-levels with a common coordinate system prior to base calling. The common coordinate system can be the reference coordinate system disclosed herein. The common coordinate system can be predefined. The common coordinate system can be the reference coordinate system disclosed herein. The common coordinate system can be predefined. The common coordinate system can be a Cartesian coordinate system. A variety of other coordinate systems can also be used. Other coordinate systems can include, but are not limited to, polar coordinate systems, cylindrical coordinate systems, or spherical coordinate systems. Exemplary embodiments of image registration methods are described in PCT patent application number PCT / US2023 / 067931 (the content of which is hereby incorporated by reference in its entirety).
[0269] Prior to registration with the pool image, the flow cell images can be registered relative to each other such that polymerase colonies or clusters in different cycles and / or channels can be aligned and base calling can be accurate and reliable for a particular polymerase colony or cluster.
[0270] A variety of methods can be used to register sequencing images of different cycles and / or channels, such as flow cell images, filtered images, or MIP images.
[0271] In some embodiments, method 600 includes the operation of registering the first and / or second MIP images, e.g., Figure 9 900 in. In some embodiments, prior to performing any base calling, the MIP images are registered across channels and different cycles. 2D registration techniques can be used to register the MIP images, e.g., by treating the MIP images as flow cell images acquired from the sequencing system 110. In some embodiments, the MIP images can be registered using the image registration method 900 disclosed herein by treating each MIP image as a flow cell image, e.g., across different channels and / or different cycles. In some embodiments, the MIP images can be registered after performing one or more of the preprocessing operations disclosed herein.
[0272] For example, the flow cell images can be registered to a reference coordinate system common to all flow cell images such that the flow cell images from different cycles and / or channels can be aligned relative to each other. The reference coordinate system can be determined in a reference cycle or any other predefined cycle. For example, the reference coordinate system can be the coordinate system of a flow cell image from one channel. As another example, the reference coordinate system can be based on external fiducial markers or other objects external to the flow cell image.
[0273] In some embodiments, method 600 includes the operation of generating one or more template images in the reference coordinate system by registering polymerase colonies to one or more template images using the coordinates of the colonies. Figure 8AA schematic diagram showing one or more template images generated in a reference coordinate system in a reference cycle is shown. In some embodiments, the template image 210 has approximately the same size as a single tile 210 including a 5×5 grid including sub-tiles 220. Regions 230 are selected in each sub-tile, and the regions include the central pixel of the corresponding sub-tile. In this embodiment, the reference coordinate system has an origin 212 located at its upper left pixel. In some embodiments, the template images disclosed herein may be individual regions, such as region 230. Each template image may include a plurality of polymerase colonies 232 therein.
[0274] In some embodiments, the size of the template image may be approximately the same as the size of the flow cell image, such that Figures 8A - 8B the different tiles 210 from and all polymerase colonies from multiple channels can be registered to the same template image. However, such a template image may contain polymerase colonies that are not used in at least some of the operations described herein to reduce the computational burden without sacrificing accuracy.
[0275] In some embodiments, more than one template image may be generated, and each template image 230 corresponds to at least a portion of the sub-tiles of the flow cell image from a channel.
[0276] The template images herein may be initialized as virtual images having a black or dark background and no signal from polymerase colonies. For example, the template image may be initialized to zero or include a minimum image intensity at all pixels.
[0277] After determining the coordinates of the polymerase colonies by image registration of the flow cell images across different channels in operation 910, the intensity of the polymerase colonies may be added to the template image at the positions determined by the coordinates, and their size and shape may be determined based on the registration. The template image may be a virtual image that combines the image intensities of polymerase colonies obtained from 2, 3, 4, or more channels in a reference cycle. The pixels of the template without polymerase colonies remain black or dark, so that the template image may have a cleaner background without the noise present in the actual flow cell image.
[0278] In some embodiments, method 900 includes an operation of obtaining image intensity, size, shape, or a combination thereof of a polymerase community from at least a portion of one or more subtiles in a reference cycle such that such information can be used to include the polymerase community in a template image. In some embodiments, the polymerase community may have a fixed shape and / or size. In some embodiments, the point spread function determined by the optical system is used to determine the fixed shape and / or size of the polymerase community. In some embodiments, the polymerase community has a fixed spot size based on the σ of the Gaussian point spread function. In some embodiments, the size of one or more polymerase communities is 1-9 pixels. In some embodiments, the size of one or more polymerase communities is 1-3 pixels.
[0279] The template image may include polymerase communities from different channels and channel information. As an example, the channel information may be provided in the form of a label or in a specific order of how the polymerase communities are included.
[0280] In some embodiments having multiple template images, each template image 230 may cover an area within a subtile, and such template images may or may not include all polymerase communities within the subtile.
[0281] In some embodiments, method 900 includes an operation 930 of obtaining a flow cell image in a cycle after the reference cycle. Operation 930 may include passively receiving or actively requesting the flow cell image from the optical system after the optical system generates the flow cell image as disclosed herein. The optical system may include the imager 116 in Figure 1 the.
[0282] The flow cell image may include some or all of the same polymerase communities as in the template image of the reference cycle. Specifically, the flow cell image may include some or all of the same polymerase communities in the area corresponding to the selected area in the reference cycle.
[0283] Figure 8A A flow cell image obtained in a cycle different from the reference cycle is shown at the bottom. The flow cell image 240 is obtained with a plurality of subtiles 250. The position of the selected area 260 in this cycle relative to the new origin 242 in this cycle may be the same as the position of the selected area 230 in the reference cycle relative to the origin 212. In this cycle, the flow cell image 210 in the reference may have been transformed into a transformed image 211, and the selected area 230 is correspondingly transformed into area 231 and has some overlap with area 260. The image transformation herein may be 2D and may include translation, scaling, rotation, and / or shear.
[0284] In some embodiments, method 900 is configured to align the template image 210 or 230 in the reference cycle and the transformed image 211 or 231 in another cycle with the reference coordinate system.
[0285] In some embodiments, instead of directly using regions 231 or 211 in image registration, method 900 may include an operation of selecting regions 230 and 260 for more simple and convenient determination of image registration. Region 260 may include at least a portion of the polymerase colony 232 in the template image 230.
[0286] In some embodiments, method 900 includes an operation 940 of determining a plurality of transformations of the flow cell image 240 based on one or more template images 210 or 230. As Figure 8A shown, each transformation in the plurality of transformations may correspond to a sub-tile 250 of the flow cell image 240 and is configured to register the sub-tile 250 of the flow cell 240 image to the corresponding portion of the template image 210 (if the template image includes the entire tile) or the corresponding template image 230 (if there are multiple template images within the tile).
[0287] In some embodiments, operation 540 may include determining each transformation corresponding to the sub-tiles of the flow cell image. More specifically, each transformation may correspond to a selected region of each of some or all of the sub-tiles. Regions may be selected from the sub-tiles in various ways to include at least a portion of the sub-tile. The region may be a predetermined two-dimensional shape, e.g., rectangular, circular, or square. As a non-limiting example, the selected region may include one or more central pixels of the sub-tile, as Figure 8A shown at 260. The size of the region may be determined to balance the trade-off between computational complexity and the accuracy of image registration. For example, selecting a 64×64 region may be computationally simpler than selecting a 128×128 region, but may be less accurate. In some embodiments, the selected region includes some or all of the polymerase colonies 232 registered in the template image in the reference cycle so that the same polymerase colonies and their relative positions in the template image and the flow cell image can be used to determine the transformation. In some embodiments, the size of the template image (e.g., 230) and region 260 may be the same or substantially the same. In some embodiments, the size of the template image 210 or 220 and the selected region 260 may be different.
[0288] In some embodiments, the cross-correlation between the selected region and the template image may be calculated to determine the 2D displacement of the region relative to the template image. Figure 10AA reference image (left) is shown, which is transformed into a different image (middle) by 2D shear, scaling, and rotation. The 2D displacements 601 at the four corners of the reference image can be determined, for example, using the methods disclosed herein and the calculation of cross-correlation. And the 2D displacements at the four corners can be used to estimate the transformation between the two images.
[0289] In some embodiments, cross-correlation can be calculated in the spatial domain. In some embodiments, cross-correlation can be calculated in the spatial frequency domain after Fourier transform (FT). Method 900 can include generating the corresponding Fourier-transformed image (FTI) of the template image and the Fourier transform of the selected region. The Fourier transform herein can be calculated using discrete FT (DFT), fast FT (FFT), etc. Cross-correlation can be determined based on the FTI of the selected region and the Fourier transform. As a non-limiting example, cross-correlation can be the element-by-element multiplication of the FTI and the FT of the selected region, followed by adding the complex conjugate or rotation of one of them. Then, the inverse FT of the element-by-element multiplication can be obtained. In some embodiments, cross-correlation can be a 2D image having a peak intensity at its coordinates [xp, yp]. In some embodiments, the 2D displacement can be determined based on the comparison of the coordinates [xp, yp] with the coordinates of the peak obtained from the cross-correlation of the two original images that are never transformed. The 2D displacement of the selected region 260 can be used to estimate the 2D displacement of the entire sub-block. In some embodiments, the results of calculating cross-correlation in the spatial domain or the Fourier domain can be equivalent. In some embodiments, the calculation in the Fourier domain is simpler and more efficient than that in the spatial domain.
[0290] In some embodiments, the image transformation of a sub-block can be determined according to the 2D displacements from some or all of the neighboring sub-blocks, regardless of whether there is a 2D displacement in itself. In some embodiments, the 2D displacements from all immediate neighbors can be used. For example, to determine the transformation of sub-block 253, the 2D displacements from 3 neighboring sub-blocks and its own 2D displacement can be used. For sub-block 251, a total of 6 2D displacements can be used, including the directly neighboring sub-blocks and its own. For sub-block 252, a total of 9 2D displacements can be used, including the neighboring sub-blocks and its own. In some embodiments, the 2D displacements from some but not all of the neighboring sub-blocks can be used. In some embodiments, except for one or two outliers, the 2D displacements from all neighboring sub-blocks can be used to determine the transformation. Outliers can be excluded using a pre-determined criterion, for example, differing from other 2D displacements by more than 30% or 50%.
[0291] Figure 10Bis an image of the 2D displacement within the tile that displays the flow cell image. In this embodiment, the tile has a 6x9 grid of sub-tiles, and each sub-tile has a 2D displacement 601 determined using the techniques disclosed herein. The magnitude of each displacement along the x or y axis is less than about 5 pixels. The pixel size can vary depending on the imaging parameters, and an exemplary pixel can be from 0.01um to 0.9um. The 2D displacement 601 can be used to calculate the transformation of the tile by separately calculating the transformation of each sub-tile, for example, an affine matrix. In this embodiment, the affine matrix can be calculated using the methods disclosed herein.
[0292] In some embodiments, various methods (including interpolation, upsampling, etc.) can be used to achieve sub-pixel resolution (e.g., about 0.01, 0.02, 0.03, or 0.05 pixels) of the 2D shift 601. In some embodiments, sub-pixel resolution can be achieved by fitting peaks with a selected filter (e.g., 3×3 or 5×5 Gaussian filter).
[0293] In some embodiments, the image transformation of the sub-tile can be uniquely represented by a transformation matrix. The transformation matrix can be determined as follows:
[0294]
[0295] where n is the number of sub-tiles, a1 = x1 + dx1, b1 = y1 + dy1, a2 = x2 + dx2, b2 = y2 + dy2,... an = xn + dxn, bn = yn + dyn, d1... dn are the 2D shifts corresponding to the sub-tiles, and where dxn and dyn are the shift components of the 2D shift dn along the x-axis and y-axis respectively, and where M is the 3x3 transformation matrix of the sub-tile.
[0296] In some embodiments, the transformation matrix can be defined as the inverse matrix of M, i.e., M -1 , so equation (1) can be represented differently as
[0297]
[0298] In some embodiments, the transformation matrix M is estimated based on the 2D shifts in equations (4) and (6). In some embodiments, the value of n may affect the accuracy of the estimate. In some embodiments, more than one region can be selected within the sub-tile for cross-correlation calculation, and more than one 2D displacement can be calculated for each sub-tile and used to estimate the transformation of the sub-tile. In these embodiments, n in equation (1) can be replaced with a larger number, for example, when 2 regions are selected for each sub-tile, it can be replaced with 2*n, and the transformation matrix M can be estimated using equations (4) and (5).
[0299] In some embodiments, (a1, b1)... (an, bn) in equations (1)–(3) are the coordinates of the selected region after transformation (e.g., the coordinates of the central pixel of the corresponding region), and (x1, y1)... (xn, yn) are the coordinates of the selected region before transformation, e.g., the coordinates of the central pixel.
[0300] In some embodiments, n is a number not less than 3. The larger n is, the more information is used to estimate the transformation matrix M. In some embodiments, n is not greater than 9.
[0301] In some embodiments, the transformation of one or more sub - tiles is linear. In some embodiments, the transformation of all sub - tiles is linear. In some embodiments, the transformation matrix is a matrix where M31 and M32 are equal to 0 and M33 is 1. In some embodiments, one or more of the transformations of each sub - tile are affine transformations, and the transformation matrix is an affine matrix.
[0302] In some embodiments, the transformation matrix M is an estimate in equations (4) and (6) based on the size of the selected region. In some embodiments, the size of the selected region may affect the accuracy of the estimate. In some embodiments, the size of the selected region may be about 128×128. In some embodiments, the size of the selected region may be about 32×32, 48×48, 64×64, 96×96, 160×160, 196×196, or 256×256. As disclosed herein, the transformation of each sub - tile can be calculated using the selected region within the sub - tile, and the selected region can be equal to or smaller than the sub - tile. In either case, considering the inherent characteristics of the image transformation across sequencing cycles, the transformation estimated using the region can be used to estimate the transformation of the entire sub - tile. The image transformation between cycles and / or between neighboring pixels can be relatively small, e.g., less than about 5% or less than about 1% of scaling, rotation, and / or shear. In some embodiments, the transformation disclosed herein can include an image translation with a difference between cycles and / or between neighboring pixels greater than about 5%.
[0303] After determining the multiple transformations of individual sub - tiles, the transformation of the flow cell image can be accurately and reliably estimated through the multiple transformations. The techniques disclosed herein advantageously estimate the transformation of the flow cell image by determining the multiple transformations of individual sub - tiles of the flow cell image. The multiple transformations can be linear, and even if the transformation is non - linear, the transformation of the flow cell image can be accurately and reliably estimated. The techniques disclosed herein advantageously eliminate the need to calculate the transformation of the entire flow cell image, which may be computationally more intensive and time - consuming compared to estimating the multiple transformations of sub - tiles.
[0304] In some embodiments, the computer-implemented method 900 further includes an operation of saving a plurality of transforms by the processor disclosed herein. In some embodiments, the computer-implemented method 900 further includes an operation of transmitting the plurality of transforms to a processing unit (such as a CPU) for subsequent operations.
[0305] In some embodiments, the computer-implemented method 900 further includes registering sub-blocks of an image to one or more template images using the plurality of transforms. This operation can be performed by a processing unit such as a CPU. In any given cycle different from the reference cycle, each sub-block of the image can be registered or transformed to one or more template images by multiplying the sub-block by a transform matrix corresponding to the sub-block.
[0306] In some embodiments, the computer-implemented method 900 can include an operation of performing one or more preprocessing steps on the flow cell image of the cycle before registering the images from the reference cycle and / or other cycles.
[0307] In some embodiments, the operation of performing one or more preprocessing steps can be performed by an FPGA or an NPU. In some embodiments, the data after the operation can be transmitted by the FPGA or the NPU to the CPU, such that the CPU can use such data to perform subsequent operations in method 600 900.
[0308] In some embodiments, the preprocessing steps for one or more flow cell images in the reference cycle can be performed before operation 910, 920 or after operation 920. In some embodiments, the preprocessing steps for the flow cell images in the reference cycle can be performed after the operation of receiving the flow cell images in the reference cycle from the optical system disclosed herein. In some embodiments, the preprocessing steps for the flow cell images in the reference cycle can be performed before the operation of obtaining the image intensity, size, shape, or a combination thereof of the polymerase population from a plurality of sub-blocks of the flow cell image in the reference cycle.
[0309] In some embodiments, the preprocessing steps for the flow cell images in cycles other than the reference cycle can be performed after operation 930 or 940. In some embodiments, the preprocessing steps for the flow cell images in cycles other than the reference cycle can be performed after the operation of registering the sub-blocks of the flow cell image to one or more template images. In some embodiments, the preprocessing steps for the flow cell images in cycles other than the reference cycle can be performed before the operation of extracting the image intensity of a plurality of polymerase populations from the sub-blocks of the flow cell image. In some embodiments, the preprocessing steps for the flow cell images in cycles other than the reference cycle can be performed before the operation of performing base identification using the image intensity of the sub-blocks of the flow cell image.
[0310] One or more preprocessing steps may include background subtraction. Background subtraction is configured to remove at least some background signals that may interfere with the signal of interest, i.e., the image intensity of the polymerase population. The background signal may be noise caused by multiple sources, including the flow cell 112, the imager 115, the sequencer 114, and other sources. The background subtraction may be adjusted to avoid over-subtraction.
[0311] One or more preprocessing steps may include image sharpening so that the image intensity of the polymerase population can be optimized in view of the surrounding environment of the polymerase population in the flow cell image. For example, sharpening may be performed using a Laplacian of Gaussian (LoG) filter.
[0312] One or more preprocessing steps may include image registration so that the image intensities of the polymerase populations can be registered relative to each other. For example, the image intensities may be registered with the templates described herein.
[0313] One or more preprocessing steps may include intensity offset adjustment, which can remove intensity offsets that were not removed during background subtraction.
[0314] One or more preprocessing steps may include color correction for removing interference from other channels or colors to one channel.
[0315] One or more preprocessing steps may include lag and lead correction, which is configured to correct the image intensity within a specific cycle by eliminating intensity biases caused by sequencing of DNA fragments that are out of sync with other fragments due to lag or lead.
[0316] One or more preprocessing steps may include intensity normalization so that the image intensities of the polymerase populations from different channels are normalized within a predetermined range.
[0317] One or more preprocessing steps may include: background subtraction; image sharpening; a combination thereof.
[0318] In some embodiments, the computer-implemented method 900 further includes extracting the image intensities of a plurality of polymerase colonies from sub-tiles registered to a template image. This operation can be performed by a processing unit such as a CPU or an FPGA. In some embodiments, the polymerase colonies and their corresponding intensities are extracted from the flow cell image into a different data format that is easier and more efficient to process. For example, each polymerase colony can have 4 different intensities, each from a different channel. Such intensities can be extracted into a list, with each entry in the list corresponding to a polymerase colony. The list can be generated after image registration to reflect the location information of the same polymerase colony in different cycles. Thus, the image intensities of the same polymerase colony in different cycles can be located in different lists, each corresponding to one cycle.
[0319] In some embodiments, the computer-implemented method 900 further includes performing base calling using the image intensities of sub-tiles of the flow cell image after registration, such that base calling can be accurately performed across different channels and in different cycles with respect to the same polymerase colony.
[0320] In some embodiments, method 900 includes operation 940 of determining a plurality of transforms of a flow cell image. Operation 940 can include determining each transform in the plurality of transforms without using any adjacent sub-tiles disclosed herein. Instead, more than 2 regions can be selected within a sub-tile, and a 2D displacement of each region in the plurality of regions can be determined. The 2D displacements obtained from regions within the same sub-tile can be used to determine the transform of the sub-tile using equations (1) and (2). The regions within the sub-tile can be smaller in size than the regions 260 in neighboring sub-tiles. For example, region 260 can be approximately 128×128, and the regions within the sub-tile can be 3, 4, 5, or even more regions, and each region includes a matrix of approximately 64×64. The other operations of method 900 can remain the same for image registration using or not using adjacent sub-tiles when generating the transform.
[0321] Optical system
[0322] Figure 1 The imager 116 therein can include one or more optical systems. Further disclosed herein are optical system design guidelines and high-performance fluorescence imaging methods and systems that provide improved optical resolution and image quality for fluorescence imaging-based genomics applications. The disclosed optical imaging system design provides a larger field of view, increased spatial resolution, improved modulation transfer, contrast-to-noise ratio and image quality, higher spatial sampling frequency, faster transition between image captures when repositioning the sample plane to capture a series of images (e.g., images of different fields of view), and improved imaging system duty cycle, and thus enables higher throughput image acquisition and analysis.
[0323] In some cases, improvements in imaging performance, e.g., for dual-sided (flow cell) imaging applications, can be achieved by using an electro-optic phase plate in combination with an objective lens to compensate for optical aberrations caused by fluid layers separating the upper (near) and lower (far) inner surfaces of the flow cell. In some cases, this design approach can also compensate for vibrations introduced by, e.g., a motion actuator compensator that moves into or out of the optical path depending on which surface of the flow cell is being imaged.
[0324] In some cases, improvements in imaging performance, e.g., for dual-sided (flow cell) imaging applications, involve using a thick flow cell wall (e.g., wall (or coverslip) thickness > 700 μm) and fluid channels (e.g., fluid channel height or thickness of 50 to 200 μm), and can be achieved by using a tube lens design that corrects for optical aberrations caused by the thick flow cell wall and / or intervening fluid layers in combination with the objective lens, even when using an off-the-shelf commercially available objective lens.
[0325] In some cases, improvements in imaging performance, e.g., for multi-channel (e.g., two-color or four-color) imaging applications, can be achieved by using multiple tube lenses (one tube lens per imaging channel), where each tube lens design has been optimized for the specific wavelength range used in that imaging channel.
[0326] Exemplary embodiments disclosed herein may include a fluorescence imaging system that includes: a) at least one light source configured to provide excitation light within one or more specified wavelength ranges; b) an objective lens configured to collect fluorescence generated from a specific field of view in a sample plane when the sample plane is exposed to the excitation light, where the numerical aperture of the objective lens is at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, or at least 0.9, or a numerical aperture value falling within a range defined by any two of the foregoing; wherein the working distance of the objective lens is at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, at least 1000 μm, or a working distance falling within a range defined by any two of the foregoing; and wherein the field of view has at least 0.1 mm 2 , at least 0.2 mm 2 , at least 0.5 mm 2 , at least 0.7 mm 2 , at least 1 mm 2 , at least 2 mm 2 , at least 3 mm 2 , at least 5 mm 2 or at least 10 mm 2the area, or the field of view falls within the range defined by any two of the above; and c) at least one image sensor, wherein the fluorescence collected by the objective lens is imaged onto the image sensor, and wherein the pixel size of the image sensor is selected such that the spatial sampling frequency of the fluorescence imaging system is at least twice the optical resolution of the fluorescence imaging system.
[0327] In some embodiments, the numerical aperture can be at least 0.75. In some embodiments, the numerical aperture is at least 1.0. In some embodiments, the working distance is at least 850 μm. In some embodiments, the working distance is at least 1,000 μm. In some embodiments, the area of the field of view can be at least 2.5 mm2. In some embodiments, the area of the field of view can be at least 3 mm2. In some embodiments, the spatial sampling frequency can be at least 2.5 times the optical resolution of the fluorescence imaging system. In some embodiments, the spatial sampling frequency can be at least 3 times the optical resolution of the fluorescence imaging system. In some embodiments, the system can further include an X-Y-Z translation stage such that the system is configured to acquire a series of two or more fluorescence images in an automated manner, wherein each image in the series is or can be acquired for a different field of view. In some embodiments, the positioning of the sample plane can be adjusted simultaneously in the X direction, Y direction, and Z direction to match the positioning of the objective focal plane between acquiring images of different fields of view. In some embodiments, the time required for simultaneous adjustment in the X direction, Y direction, and Z direction can be less than 0.3 seconds, less than 0.4 seconds, less than 0.5 seconds, less than 0.7 seconds, or less than 1 second or a time falling within the range defined by any two of the foregoing. In some embodiments, the system further includes an autofocus mechanism configured to adjust the focal plane positioning before acquiring images of different fields of view if an error signal indicates that the positioning difference between the focal plane and the sample plane in the Z direction is greater than a specified error threshold. In some embodiments, the specified error threshold is 100 nm or greater. In some embodiments, the specified error threshold is 50 nm or less. In some embodiments, the system includes three or more image sensors, and wherein the system is configured to image the fluorescence in each of three or more wavelength ranges onto different image sensors. In some embodiments, the positioning difference between the focal plane of each of the three or more image sensors and the sample plane is less than 100 nm. In some embodiments, the positioning difference between the focal plane of each of the three or more image sensors and the sample plane is less than 50 nm. In some embodiments, the total time required to reposition the sample plane, adjust the focal length (if necessary), and acquire an image is less than 0.4 seconds per field of view. In some embodiments, the total time required to reposition the sample plane, adjust the focal length (if necessary), and acquire an image is less than 0.3 seconds per field of view.
[0328] The present disclosure also discloses a fluorescence imaging system for dual-sided imaging of a flow cell, the fluorescence imaging system comprising: a) an objective lens configured to collect fluorescence generated within a designated field of view of a sample plane within the flow cell; b) at least one tube lens positioned between the objective lens and at least one image sensor, wherein the at least one tube lens is configured to correct an imaging performance metric of a combination of the objective lens, at least two tube lenses, and at least one image sensor when imaging an inner surface of the flow cell, and wherein a wall thickness of the flow cell is at least 700 μm and a gap between an upper inner surface and a lower inner surface is at least 50 μm; wherein the imaging performance metrics are substantially the same for imaging an upper inner surface or a lower inner surface of the flow cell without moving an optical compensator into or out of an optical path between the flow cell and the at least one image sensor, without moving one or more optical elements of a barrel lens along the optical path, and without moving one or more optical elements of the barrel lens into or out of the optical path.
[0329] In some embodiments, the objective lens can be a commercially available microscope objective lens. In some embodiments, the numerical aperture of the commercially available microscope objective lens can be at least 0.3. In some embodiments, the working distance of the objective lens can be at least 700 μm. In some embodiments, the objective lens can be corrected to compensate for the thickness of a cover glass (or the thickness of the flow cell wall) of 0.17 mm or greater or less than 0.17 mm. In some embodiments, the optical system can be corrected to compensate for the distance between the cover glass thickness, the flow cell thickness, or the desired focal plane. In some embodiments, the correction can be performed by inserting a correction optical device such as a lens or an optical assembly into the optical path of the optical system. In some embodiments, the correction can be performed without inserting a correction optical device such as a lens or an optical assembly into the optical path of the optical system. In some embodiments, the fluorescence imaging system can further include an electro-optic phase plate, which is positioned adjacent to the objective lens and between the objective lens and the tube lens, wherein the electro-optic phase plate can provide correction for optical aberrations caused by the fluid filling the gap between the upper inner surface and the lower inner surface of the flow cell. In some embodiments, at least one tube lens can be a compound lens including three or more optical components. In some embodiments, at least one tube lens is a compound lens including four optical components, and the four optical components can include one or more of the following: a first asymmetric convex-convex lens, a second plano-convex lens, a third asymmetric concave-concave lens, and a fourth asymmetric convex-concave lens, which can be present in the order listed above or in any alternative order. In some embodiments, the at least one tube lens is configured to correct the imaging performance metrics of the combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging the inner surface of a flow cell with a wall thickness of at least 1 mm. In some embodiments, the at least one tube lens is configured to correct the imaging performance metrics of the combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 100 μm. In some embodiments, at least one tube lens is configured to correct the imaging performance metric of the combination of the objective lens, at least one tube lens, and at least one image sensor when imaging the inner surface of a flow cell with a gap of at least 200 μm. In some embodiments, the system includes a single objective lens, two tube lenses, and two image sensors, and each of the two tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths. In some embodiments, the system includes a single objective lens, three tube lenses, and three image sensors, and each of the three tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths. In some embodiments, the system includes a single objective lens, four tube lenses, and four image sensors, and each of the four tube lenses is designed to provide optimal imaging performance at different fluorescence wavelengths.In some embodiments, the design of the objective lens or at least one tube lens is configured to optimize the modulation transfer function in the medium to high spatial frequency range. In some embodiments, the imaging performance metrics include measurements of the modulation transfer function (MTF) at one or more specified spatial frequencies, defocus, spherical aberration, chromatic aberration, coma, astigmatism, field curvature, image distortion, contrast-to-noise ratio (CNR), or any combination thereof. In some embodiments, the difference in the imaging performance metrics for imaging the upper and lower inner surfaces of the flow cell is less than 10%. In some embodiments, the difference in the imaging performance metrics for imaging the upper and lower inner surfaces of the flow cell is less than 5%. In some embodiments, compared to a conventional system including an objective lens, a motion-actuated compensator, and an image sensor, using at least one tube lens provides at least equivalent or better improvement in the imaging performance metrics for bilateral imaging. In some embodiments, compared to a conventional system including an objective lens, a motion-actuated compensator, and an image sensor, using at least one tube lens provides at least a 10% improvement in the imaging performance metrics for bilateral imaging.
[0330] Disclosed herein is an illumination system for imaging-based solid-phase genotyping and sequencing applications, the illumination system comprising: a) a light source; and b) a liquid light guide configured to collect light emitted by the light source and transmit it to a specified illumination field on a support surface including immobilized biological macromolecules.
[0331] In some embodiments, the illumination system further includes a condenser lens. In some embodiments, the area of the specified illumination field is at least 2 mm2. In some embodiments, the light delivered to the specified illumination field has uniform intensity across a specified field of view of an imaging system for acquiring an image of the support surface. In some embodiments, the area of the specified field of view is at least 2 mm2. In some embodiments, the light delivered to the specified illumination field has uniform intensity across the specified field of view when the coefficient of variation (CV) of the light intensity is less than 10%. In some embodiments, the light delivered to the specified illumination field has uniform intensity across the specified field of view when the coefficient of variation (CV) of the light intensity is less than 5%. In some embodiments, the speckle contrast value of the light transmitted to the specified illumination field is less than 0.1. In some embodiments, the speckle contrast value of the light transmitted to the specified illumination field is less than 0.05.
[0332] Imaging modules and systems
[0333] Those skilled in the art will understand that, in some cases, the disclosed optical systems, imaging systems, or modules can be stand-alone optical systems designed to image a sample or a substrate surface. In some cases, they can include one or more processors or computers. In some cases, they can include one or more software packages that provide instrument control functions and / or image processing functions. In some cases, in addition to optical components such as light sources (e.g., solid-state lasers, dye lasers, diode lasers, arc lamps, tungsten halogen lamps, etc.), lenses, prisms, mirrors, dichroic reflectors, optical filters, optical band-pass filters, apertures, and image sensors (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors and cameras, charge-coupled device (CCD) image sensors and cameras, etc.), they can also include mechanical and / or optomechanical components such as X-Y translation stages, X-Y-Z translation stages, piezoelectric focusing mechanisms, etc. In some cases, they can serve as modules, components, sub-assemblies, or subsystems of a larger system designed for genomics applications (e.g., gene testing and / or nucleic acid sequencing applications). For example, in some cases, they can serve as modules, components, sub-assemblies, or subsystems of a larger system that further includes an opaque and / or other environmental control housing, a temperature control module, a fluid control module, a fluid dispensing robot, a pick-and-place robot, one or more processors or computers, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, image processing software packages, data analysis software packages), a data storage module, a data communication module (e.g., Bluetooth, WiFi, intranet, or Internet communication hardware and associated software), a display module, or any combination thereof.
[0334] Methods for sequencing
[0335] The present disclosure provides methods for sequencing immobilized or non-immobilized template molecules. The methods can be operated in system 100, e.g., in sequencer 114. In some embodiments, the immobilized template molecules comprise a plurality of nucleic acid template molecules having one copy of a target sequence of interest. In some embodiments, nucleic acid template molecules having one copy of a target sequence of interest can be generated by bridge amplification using linear library molecules. In some embodiments, the immobilized template molecules comprise a plurality of nucleic acid template molecules, each nucleic acid template molecule having two or more tandem copies (e.g., concatemers) of a target sequence of interest. In some embodiments, nucleic acid template molecules comprising concatemer molecules can be generated by performing rolling circle amplification of circularized linear library molecules. In some embodiments, the non-immobilized template molecules comprise circular molecules. In some embodiments, the methods for sequencing employ soluble (e.g., non-immobilized) sequencing polymerases or sequencing polymerases immobilized to a support.
[0336] In some embodiments, the sequencing reaction employs detectably labeled nucleotide analogs. In some embodiments, the sequencing reaction employs a two-stage sequencing reaction that includes binding a detectably labeled multivalent molecule and incorporating nucleotide analogs. In some embodiments, the sequencing reaction employs unlabeled nucleotide analogs. In some embodiments, the sequencing reaction employs phosphate backbone-labeled nucleotides.
[0337] Linear library molecules
[0338] In some embodiments, the immobilized concatemers each include tandem repeat units (e.g., an insertion region) of the sequence of interest and any adapter sequences. For example, the tandem repeat unit includes: (i) a left universal adapter sequence having a binding sequence for a first surface primer (1120) (e.g., a surface pinning primer), (ii) a left universal adapter sequence having a binding sequence for a first sequencing primer (1140) (e.g., a forward sequencing primer), (iii) the sequence of interest (1110), (iv) a right universal adapter sequence having a binding sequence for a second sequencing primer (1150) (e.g., a reverse sequencing primer), (v) a right universal adapter sequence having a binding sequence for a second surface primer (1130) (e.g., a surface capture primer), and (vii) a left sample index sequence (1160) and / or a right sample index sequence (1170). In some embodiments, the tandem repeat unit further includes a left unique identifier sequence (1180) and / or a right unique identifier sequence (1190). In some embodiments, the tandem repeat unit further includes a binding sequence for at least one compaction oligonucleotide. In some embodiments, Figure 11 and 12 illustrates a linear library molecule or a concatemer molecule unit.
[0339] Methods for performing in - situ short - read sequencing
[0340] In the methods described herein, RNA is not extracted from the cell sample, and there is no need to track sequencing information and map it back to an image of the cell sample. Instead, the RNA remains inside the cell sample to allow direct imaging of the spatial location of the target RNA within the cell. Additionally, the RNA within the cell sample is not fragmented, and there is no need for enrichment of the target RNA. The use of target-specific and / or random sequence reverse transcription primers enables the detection of poly-A and non-poly-A RNA in a singleplex or multiplex mode.
[0341] In some embodiments, the method includes performing a small number of sequencing cycles repeatedly on the same region of a template molecule (e.g., a concatemer molecule). By performing repeated short sequencing cycles, the RNA content of a cell sample can be discovered. Compared to long-read sequencing workflows, the repeated short sequencing cycles described herein use a much smaller amount of sequencing reagents, thereby reducing costs and saving time. The method for performing repeated short sequencing cycles has many uses, including but not limited to detecting specific RNAs of interest, mutant RNA sequences, splice variants, and their abundance levels.
[0342] The concatemer carries tandem repeat units of cDNA of interest, universal sequencing primer binding sites, and target barcode sequences. The concatemer is sequenced inside a cell sample, where a small number of sequencing cycles are performed per round and multiple rounds of short-read sequencing are carried out. The entire lengths of the target barcode and cDNA regions are not sequenced. Instead, at least a portion of the target barcode region is repeatedly sequenced. In some embodiments, it is not necessary to sequence the cDNA region. In some embodiments, the target barcode and a portion of the cDNA region are repeatedly sequenced. It is not necessary to sequence the entire length of the cDNA region. It is not necessary to assemble the sequencing reads or obtain the full-length sequence of the cDNA of interest. The redundant sequencing information obtained from the short sequencing reads obviates the need to sequence the complementary strand of the concatemer. Thus, paired-end sequencing is not required.
[0343] Additionally, a short portion of the cDNA region in the concatemer is re-sequenced (e.g., repeatedly sequenced) at least once from the same starting position to generate overlapping sequencing reads that can be aligned to a reference sequence. For example, the same portion of the concatemer molecule can be sequenced at least two, three, four, five, or up to 50 times. The starting sequencing site can be any position in the concatemer and is determined by the sequencing primer designed to anneal to the selected position within the concatemer. The repeated short sequencing reads increase the redundancy of the sequencing information for a single base in the cDNA region. Repeatedly sequencing one strand of the concatemer template molecule provides sufficient base coverage to reveal the presence of the target RNA in the cell sample, so paired-end sequencing of the complementary strand is not necessary.
[0344] The concatemer template molecule includes multiple sequencing primer binding sites along the same concatemer molecule, which can be used to generate multiple available sequencing reads to increase the sequencing depth. In summary, compared to sequencing a single-copy template molecule, repeatedly sequencing one strand of the concatemer template can increase the sequencing base coverage and sequencing depth.
[0345] The methods described herein can be carried out in single or multiplex modes. Two or more different target RNAs can be simultaneously detected and imaged within a cellular sample using different reverse transcription primers, different target-specific padlock probes, and universal sequencing primers. For example, any of the repeat short-read sequencing methods described herein can be used to simultaneously detect and image the presence of housekeeping RNA and at least one target RNA in a cellular sample.
[0346] The present disclosure provides methods for in situ detection of at least two different target RNA molecules in a cellular sample, comprising step (a): providing a cellular sample containing a plurality of RNAs, the plurality of RNAs comprising at least a first target RNA molecule and a second target RNA molecule. In some embodiments, the cellular sample is fixed and permeabilized. In some embodiments, the cellular sample contains from 2 to 25 different target RNA molecules, or contains from 25 to 50 different target RNA molecules, or contains from 50 to 75 different target RNA molecules or contains from 75 to 100 different target RNA molecules. In some embodiments, the cellular sample contains more than 100 different target RNA molecules, or more than 250 different target RNA molecules, or more than 500 different target molecules, or more than 1000 different target RNA molecules or more. In some embodiments, the cellular sample contains more than 10,000 different target RNA molecules. In some embodiments, the cellular sample comprises whole cells, a plurality of whole cells, intact tissue, or an intact tumor. In some embodiments, the cellular sample comprises a fresh cellular sample, a fresh frozen cellular sample, a sectioned cellular sample, an FFPE cellular sample, or a sectioned FFPE cellular sample. In some embodiments, the cellular sample is deposited onto a solid support. In some embodiments, the cellular sample is deposited onto a solid support passivated with a coating that promotes cell adhesion. In some embodiments, the cellular sample is deposited onto a support lacking immobilized capture oligonucleotides. In some embodiments, the cellular sample is cultured before or after deposition onto the solid support. In some embodiments, the cellular sample is cultured before performing step (b) described below. In some embodiments, the cellular sample comprises an amplified cellular sample that has been cultured in a simple or complex cell culture medium. In some embodiments, the cellular sample is not cultured or amplified before performing step (b).
[0347] In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further comprises step (b): generating a plurality of cDNA molecules inside the cell sample, the plurality of cDNA molecules comprising at least a first target cDNA molecule corresponding to a first target RNA molecule, and the plurality of cDNA molecules comprising a second target cDNA molecule corresponding to the second target RNA molecule. In some embodiments, the method comprises generating at least 2 to 10,000 different target cDNA molecules corresponding to 2 to 10,000 different target RNA molecules. In some embodiments, the generation of step (b) comprises contacting a plurality of RNAs inside the cell sample with (i) a plurality of reverse transcription primers, (ii) a plurality of reverse transcriptases, and (iii) a plurality of nucleotides under conditions suitable for performing a reverse transcription reaction to generate a plurality of cDNA molecules (e.g., a plurality of first-strand cDNA molecules) in the cell sample (e.g., Figure 14 ).
[0348] In some embodiments, the plurality of reverse transcription primers comprises a first subset of target-specific reverse transcription primers that selectively hybridize to the first target RNA, and comprises a second subset of target-specific reverse transcription primers that selectively hybridize to the second target RNA. In some embodiments, the first and second subsets of target-specific reverse transcription primers have the same sequence or different sequences.
[0349] In some embodiments, the entire length of the first subset of target-specific reverse transcription primers hybridizes to the first target RNA molecule. In some embodiments, the first subset of target-specific reverse transcription primers comprises a tailing primer that has a portion that hybridizes to the first target RNA molecule and a portion that does not hybridize to the first target RNA molecule. In some embodiments, the first subset of target-specific reverse transcription primers comprises at least a portion having a poly-T sequence. In some embodiments, the first subset of target-specific reverse transcription primers comprises at least a portion having a random sequence and / or at least a portion having a target-specific sequence.
[0350] In some embodiments, the entire length of the second subset of target-specific reverse transcription primers hybridizes to the second target RNA molecule. In some embodiments, the second subset of target-specific reverse transcription primers comprises a tailing primer that has a portion that hybridizes to the second target RNA molecule and a portion that does not hybridize to the second target RNA molecule. In some embodiments, the second subset of target-specific reverse transcription primers comprises at least a portion having a poly-T sequence. In some embodiments, the second subset of target-specific reverse transcription primers comprises at least a portion having a random sequence and / or at least a portion having a target-specific sequence.
[0351] In some embodiments, under conditions suitable for degrading RNA in an RNA / DNA duplex, a ribonuclease can be used to enzymatically degrade a target RNA molecule hybridized to a cDNA molecule. In some embodiments, the target RNA molecule hybridized to the cDNA molecule is not enzymatically degraded.
[0352] In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further comprises step (c): contacting a plurality of cDNA molecules in the cell sample with a plurality of target-specific padlock probes, the plurality of target-specific padlock probes including at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes. In some embodiments, the method comprises contacting a plurality of cDNA molecules in the cell sample with at least 2 to 10,000 different target-specific padlock probes.
[0353] In alternative embodiments, the cDNA is not generated from RNA inside the cell sample. In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further comprises contacting the RNA inside the cell with a plurality of target-specific padlock probes and generating circularized padlock probes. In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further comprises step (c): contacting a plurality of RNA molecules in the cell sample with a plurality of target-specific padlock probes, the plurality of target-specific padlock probes including at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes. In some embodiments, the method comprises contacting a plurality of cDNA molecules in the cell sample with at least 2 to 10,000 different target-specific padlock probes. In some embodiments, a ribonuclease can be used to enzymatically degrade the target RNA molecule. In some embodiments, the target RNA molecule is not enzymatically degraded.
[0354] In some embodiments, a single padlock probe among the plurality of first target-specific padlock probes comprises first and second terminal regions (e.g., first and second padlock binding arms), wherein the first terminal region selectively hybridizes to a first region of a first target cDNA molecule (or a first target RNA molecule), and the second terminal region selectively hybridizes to a second region of the first target cDNA molecule (or the first target RNA molecule). In some embodiments, the contacting of step (c) comprises: hybridizing the first and second terminal regions of the first target-specific padlock probe to proximal positions on the first target cDNA molecule (or the first target RNA molecule) to form a circularized first target-specific padlock probe, the circularized first target-specific padlock probe having a nick or gap between the hybridized first and second terminal regions (e.g., Figure 14, on the left). In some embodiments, the first target-specific padlock probe comprises a first target barcode sequence (Target BC-1), which corresponds to and uniquely identifies the first target cDNA sequence (or the first target RNA sequence). In some embodiments, the first target-specific padlock probe comprises a first target barcode sequence located adjacent to one of the regions of the first target-specific padlock probe that selectively hybridizes to the first target cDNA molecule (or the first target RNA sequence). In some embodiments, the first target-specific padlock probe comprises at least one universal adapter sequence, e.g., a universal sequencing primer binding site (or its complementary sequence). In some embodiments, the first target-specific padlock probe comprises a universal primer binding site for a rolling circle amplification primer (or its complementary sequence). In some embodiments, the first target-specific padlock probe comprises a universal compaction oligonucleotide binding site (or its complementary sequence).
[0355] In some embodiments, a single padlock probe among the plurality of second target-specific padlock probes comprises first and second terminal regions (e.g., first and second padlock binding arms), wherein the first terminal region selectively hybridizes to a first region of the second target cDNA molecule (or the second target RNA molecule), and the second terminal region selectively hybridizes to a second region of the second target cDNA molecule (or the second target RNA molecule). In some embodiments, the contacting in step (c) comprises: hybridizing the first and second terminal regions of the second target-specific padlock probe to proximal positions on the second target cDNA molecule (or the second target RNA molecule) to form a circularized second target-specific padlock probe, the circularized second target-specific padlock probe having a nick or gap between the hybridized first and second terminal regions (e.g., Figure 14 , on the right). In some embodiments, the second target-specific padlock probe comprises a second target barcode sequence (Target BC-2), which corresponds to and uniquely identifies the second target cDNA sequence (or the second target RNA sequence). In some embodiments, the second target-specific padlock probe comprises a second target barcode sequence located adjacent to one of the regions of the second target-specific padlock probe that selectively hybridizes to the second target cDNA molecule (or the second target RNA sequence). In some embodiments, the second target-specific padlock probe comprises at least one universal adapter sequence, e.g., a universal sequencing primer binding site (or its complementary sequence). In some embodiments, the second target-specific padlock probe comprises a universal primer binding site for a rolling circle amplification primer (or its complementary sequence). In some embodiments, the second target-specific padlock probe comprises a universal compaction oligonucleotide binding site (or its complementary sequence).
[0356] In some embodiments, the first target barcode sequence (Target BC-1) and the second target barcode sequence (Target BC-2) have different sequences and can be used for multiplex RNA detection and sequencing. In some embodiments, the first target barcode sequence (Target BC-1) and the second target barcode sequence (Target BC-2) have the same sequence and can be used for singleplex RNA detection and sequencing.
[0357] In some embodiments, the first target-specific padlock probe and the second target-specific padlock probe comprise a universal sequencing primer binding site and a target barcode sequence adjacent to each other such that the target barcode region of the concatemer is sequenced first. The target barcode sequence can be of any length, such as 3-15 bases, or 15-25 bases, or 25-40 bases or longer.
[0358] In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further comprises step (d): generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe inside the sample by closing the nicks or gaps in at least the first and second circularized target-specific padlock probes through an enzymatic reaction. In some embodiments, closing the nicks in the first and second circular padlock probes comprises performing an enzymatic ligation reaction. In some embodiments, closing the gaps in the first and second circular padlock probes comprises performing a polymerase-catalyzed fill-in reaction using the first or second target cDNA molecule (or the first or second RNA molecule) as a template, and performing an enzymatic ligation reaction. In some embodiments, the method comprises closing the nicks or gaps in at least 2 to 10,000 circularized target-specific padlock probes through one or more enzymatic reactions to generate at least 2 to 10,000 covalently closed circular padlock probes inside the cell sample.
[0359] In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further comprises step (e): performing a rolling circle amplification reaction inside the cell sample using the first and second covalently closed circular padlock probes as template molecules to generate a plurality of concatemer molecules, the plurality of concatemer molecules comprising at least a first concatemer molecule corresponding to the first target RNA molecule, and the plurality of concatemer molecules comprising at least a second concatemer molecule corresponding to the second target RNA molecule. In some embodiments, the first concatemer molecule comprises tandem repeat units, wherein a unit comprises a sequence corresponding to the first target cDNA (or the first target RNA), the first target barcode sequence, and a universal sequencing primer binding site (or its complementary sequence). In some embodiments, the second concatemer molecule comprises tandem repeat units, wherein a unit comprises a sequence corresponding to the first target cDNA (or the second target RNA), the second target barcode sequence, and a universal sequencing primer binding site (or its complementary sequence).
[0360] In some embodiments, the rolling circle amplification reaction of step (e) includes contacting a covalently closed circular padlock probe with an amplification primer (e.g., a universal rolling circle amplification primer), a strand-displacing DNA polymerase, and a plurality of nucleotides under conditions suitable for hybridizing a single amplification primer to the covalently closed padlock probe and under conditions suitable for primer extension using the covalently closed padlock probe as a template molecule to generate nucleic acid concatemers. In some embodiments, the method includes performing a rolling circle amplification reaction inside a cell sample using at least 2 to 10,000 covalently closed circular padlock probes as template molecules, thereby generating at least 2 to 10,000 concatemer molecules corresponding to at least 2 to 10,000 target RNA molecules. In some embodiments, a plurality of concatemers generated inside the cell sample collapse into DNA nanoballs, which have a more compact shape and size compared to the uncollapsed concatemers.
[0361] In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further includes step (f): sequencing a plurality of concatemer molecules inside the cell sample, the sequencing including sequencing a first concatemer molecule by performing 2 to 1000 sequencing cycles to generate a plurality of first sequencing read products; and sequencing a second concatemer molecule by performing 2 to 1000 sequencing cycles to generate a plurality of second sequencing read products ( Figure 15 ). In some embodiments, the sequencing of step (f) includes sequencing no more than 2 to 30 bases of the first concatemer molecule to generate a plurality of first sequencing read products, and it includes sequencing no more than 2 to 30 bases of the second concatemer molecule to generate a plurality of second sequencing read products. In some embodiments, the method includes sequencing at least 2 to 10,000 concatemer molecules inside the cell sample, the sequencing including performing 2 to 1000 sequencing cycles on 2 to 10,000 concatemer molecules to generate a plurality of sequencing read products.
[0362] In some embodiments, only the first target barcode region of the first concatemer molecule is sequenced (e.g., Figure 15 , top). In some embodiments, at least a part or the full length of the first target barcode of the first concatemer molecule is sequenced (e.g., Figure 15 , top). In some embodiments, the first target barcode is sequenced, and a part of the first cDNA region (or the first RNA region) of the first concatemer molecule is sequenced. In some embodiments, at least a part of the first cDNA region (or the first RNA region) of the first concatemer molecule is sequenced.
[0363] In some embodiments, only the second target barcode region of the second concatemer molecule is sequenced (e.g., Figure 15, bottom). In some embodiments, at least a portion or the full length of the second target barcode of the second concatemer molecule is sequenced (e.g., Figure 15 , bottom). In some embodiments, the second target barcode is sequenced, and a portion of the second cDNA region (or second RNA region) of the second concatemer molecule is sequenced. In some embodiments, at least a portion of the second cDNA region (or second RNA region) of the second concatemer molecule is sequenced.
[0364] In some embodiments, the sequencing of step (f) includes contacting the plurality of concatemer molecules within the cell sample with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridization of the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the concatemer. In some embodiments, the sequencing of step (f) further includes performing 2 to 1000 sequencing cycles by sequencing at least the first target barcode region (Target BC-1) to generate at least the first plurality of sequencing read products, and optionally performing 2 to 1000 sequencing cycles by sequencing at least the second target barcode region (Target BC-2) to generate at least the second plurality of sequencing read products. In some embodiments, the nucleotide reagents include multivalent molecules, nucleotides, and / or nucleotide analogs.
[0365] In some embodiments, the sequencing of step (f) includes sequencing at least a portion of the first and second nucleic acid concatemers using an optical imaging system having a field of view (FOV) greater than 1.0 mm 2 .
[0366] In some embodiments, in the sequencing of step (f), the plurality of first and second sequencing read products are image-detectable, and the sequencing includes decoding the plurality of first and second sequencing read products from images obtained during no more than 2 to 1000 sequencing cycles.
[0367] In some embodiments, in the sequencing of step (f), the plurality of first and second sequencing read products are image-detectable, and the sequencing includes imaging the plurality of first and second detectable sequencing read products in the cell sample simultaneously (co-localization of the first and second sequencing read products).
[0368] In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further includes step (g): removing the plurality of first sequencing read products from the first concatemer molecule and retaining the first concatemer molecule in the cell sample, and removing the plurality of second sequencing read products from the second concatemer molecule and retaining the second concatemer molecule in the cell sample.
[0369] In some embodiments, the method for detecting at least two different target RNA molecules in a cell sample further comprises step (h): performing repeated sequencing on a plurality of concatemers by repeating steps (f) and (g) at least once, wherein the sequences of the plurality of first sequencing read products confirm the presence of a first target RNA molecule in the cell sample, and wherein the sequences of the plurality of second sequencing read products confirm the presence of a second target RNA molecule in the cell sample.
[0370] In some embodiments, performing repeated sequencing on at least one region of a concatemer comprises repeating steps (f) to (g) at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, or at least 10 times.
[0371] In some embodiments, performing repeated sequencing on at least one region of a concatemer comprises repeating steps (f)-(g) up to 10 times, up to 20 times, up to 30 times, up to 40 times, or up to 50 times. Examples of the repeat sequences are shown in Figures 16 - 19 the schematic diagram of.
[0372] In some embodiments, for example, as Figure 17 shown, a universal sequencing primer (solid arrow) hybridizes to a universal sequencing primer binding site, and up to 1000 sequencing cycles are performed to generate a plurality of first sequencing read products (dashed arrows), wherein the first sequencing read product contains only the target barcode sequence. The plurality of first sequencing read products are removed from the concatemer, and sequencing is repeated, wherein up to 1000 sequencing cycles are performed to generate another plurality of first sequencing read products (dashed arrows), wherein the first sequencing read product contains only the target barcode sequence. The plurality of first sequencing read products are removed from the concatemer, and sequencing is repeated again, wherein up to 1000 sequencing cycles are performed to generate another plurality of first sequencing read products (dashed arrows), wherein the first sequencing read product contains only the target barcode sequence. In some embodiments, the repeated sequencing can be performed up to 50 times. The sequences of all the first sequencing read products can be determined and aligned with a first reference sequence (e.g., a reference barcode sequence) to confirm the presence of the first target RNA molecule within the cell sample.
[0373] In some embodiments, for example, in Figure 18In this method, a universal sequencing primer (solid arrow) hybridizes to the universal sequencing primer binding site, and up to 1000 sequencing cycles are performed to generate multiple first sequencing read products (dashed arrows), where the first sequencing read products contain a portion of the inserted sequence. The multiple first sequencing read products are removed from the concatemer, and sequencing is repeated, where up to 1000 sequencing cycles are performed to generate another multiple first sequencing read products (dashed arrows), where the first sequencing read products contain a portion of the inserted sequence. The multiple first sequencing read products are removed from the concatemer, and sequencing is repeated again, where up to 1000 sequencing cycles are performed to generate another multiple first sequencing read products (dashed arrows), where the first sequencing read products contain a portion of the inserted sequence. In some embodiments, the repeated sequencing can be performed up to 50 times. The sequences of all the first sequencing read products can be determined and aligned with a first reference sequence (e.g., the inserted sequence corresponding to the target RNA) to confirm the presence of the first target RNA molecule within the cell sample.
[0374] In some embodiments, for example, Figure 19 a universal sequencing primer (solid arrow) hybridizes to the universal sequencing primer binding site, and up to 1000 sequencing cycles are performed to generate multiple first sequencing read products (dashed arrows), where the first sequencing read products contain a portion of the inserted sequence. The multiple first sequencing read products are removed from the concatemer, and sequencing is repeated, where up to 1000 sequencing cycles are performed to generate another multiple first sequencing read products (dashed arrows), where the first sequencing read products contain a portion of the inserted sequence. The multiple first sequencing read products are removed from the concatemer, and sequencing is repeated again, where up to 1000 sequencing cycles are performed to generate another multiple first sequencing read products (dashed arrows), where the first sequencing read products contain a portion of the inserted sequence. In some embodiments, the repeated sequencing can be performed up to 50 times. The sequences of all the first sequencing read products can be determined and aligned with a first reference sequence (e.g., the inserted sequence corresponding to the target RNA) to confirm the presence of the first target RNA molecule within the cell sample.
[0375] In some embodiments, at least one concatemer is sequenced by performing step (f) (non-repeated sequencing) once. In some embodiments, at least one concatemer is sequenced by performing steps (f) to (g) once. In some embodiments, at least one concatemer is repeatedly sequenced by performing steps (f)–(g) at least twice.
[0376] In some embodiments, multiple universal sequencing primers can hybridize with concatemer template molecules using a hybridization reagent that comprises an SSC buffer (e.g., 2X saline-sodium citrate) buffer having formamide (e.g., 10 to 20% formamide). The hybridization conditions include a temperature of about 20 to 30 °C for about 10 to 60 minutes.
[0377] In some embodiments, at a temperature that promotes nucleic acid denaturation (e.g., 30 to 90 °C), a dehybridization reagent comprising an SSC buffer (e.g., saline-sodium citrate) buffer (with or without formamide) can be used to remove multiple sequencing read products from the concatemer and multiple concatemers can be retained inside the cell sample.
[0378] In some embodiments, the multiple nucleotide reagents of step (f) include multiple nucleotides that are detectably labeled or unlabeled. In some embodiments, a single nucleotide is linked to a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the multiple detectably labeled nucleotide analogs include multiple chain-terminating nucleotides, wherein the chain-terminating moiety is linked to the 3'-nucleotide sugar position to form a 3'-blocked nucleotide analog. In some embodiments, the chain-terminating moiety can be removed to convert the 3'-blocked nucleotide analog to an extendable nucleotide having a 3'-OH group on the sugar. In some embodiments, the labeled nucleotide analogs are linked to different fluorophores corresponding to the nucleobases adenine, cytosine, guanine, thymine, or uracil, wherein the different fluorophores emit fluorescence signals during the sequencing of step (f). In some embodiments, the sequencing cycle includes (1) contacting the concatemer / sequencing primer duplex with a sequencing polymerase and a detectably labeled chain-terminating nucleotide under conditions suitable for polymerase-catalyzed incorporation of the detectably labeled chain-terminating nucleotide at the end of the sequencing primer, (2) detecting and imaging the fluorescence signal and color emitted by the incorporated chain-terminating nucleotide, and (3) removing the chain-terminating moiety (e.g., deblocking) and the fluorophore from the incorporated nucleotide and retaining the concatemer / sequencing primer duplex. In some embodiments, no more than 2 to 1000 sequencing cycles are performed on the multiple concatemers inside the cell sample to generate multiple sequencing read products. In some embodiments, the sequence of a first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of a first target RNA molecule within the cell sample. In some embodiments, the sequence of a second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of a second target RNA molecule within the cell sample.
[0379] In some embodiments, after generating first and second sequencing read products having a length of no more than 30 to 1000 bases per round, or after generating a set of repetitive sequencing read products of first and second sequencing read products having a length of no more than 30 to 1000 bases, the sequences of the first and second sequencing read products can be aligned. In some embodiments, the sequencing reaction is performed on a sequencing device having a detector that captures fluorescence signals from the sequencing reaction inside a cell sample. The sequencing device can be configured to relay the fluorescence signal data captured by the detector to a computer system programmed to display an image of different fluorescent spots co-located in the cell sample, where a single fluorescent spot corresponds to a different target RNA molecule. In some embodiments, when sequencing is performed using different fluorescently labeled nucleotide reagents corresponding to different nucleobases (e.g., A, G, C, T / U), then the image can have different color fluorescent spots co-located in the same cell sample at different sequencing cycles.
[0380] In some embodiments, asynchronous lagging and / or leading events can occur during the synchronous sequencing reaction of the clonally amplified template amplicons, where the sequencing reaction comprises a polymerase-catalyzed sequencing reaction employing chain-terminating nucleotides with detectable labels. In some embodiments, the sequencing reaction on one of the template molecules among the clonally amplified template molecules is advanced (e.g., leading) or lagged (e.g., lagging) compared to the sequencing of other template molecules within the clonally amplified template molecules. During sequencing, fluorescence signals corresponding to the incorporation of the labeled chain-terminating nucleotides are typically detected. Thus, the incorporation of the labeled chain-terminating nucleotides can be used to detect and monitor lagging and leading events.
[0381] In some embodiments, the plurality of nucleotide reagents of step (f) includes a plurality of multivalent molecules, each multivalent molecule comprising a core attached to a plurality of nucleotide arms, wherein the nucleotide arms are attached to nucleotide units. In some embodiments, a single multivalent molecule is labeled with a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the core of the multivalent molecule is labeled with a fluorophore, and wherein the fluorophore attached to a given core of the multivalent molecule corresponds to the nucleobase of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or a nucleobase attached to a fluorophore, and wherein the fluorophore attached to a given nucleobase corresponds to the nucleobase of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, the sequencing cycle includes: (1) contacting the concatemer / sequencing primer duplex with a first sequencing polymerase to form a complex polymerase, (2) contacting the complex polymerase with the detectable-labeled multivalent molecule under conditions suitable for binding complementary nucleotide units of the multivalent molecule to the complex polymerase to form a multivalent binding complex, and the conditions are suitable for inhibiting the incorporation of complementary nucleotide units into the end of the sequencing primer, (3) detecting and imaging the fluorescence signal and color emitted by the bound detectable-labeled multivalent molecule, (4) removing the first sequencing polymerase and the bound detectable-labeled multivalent molecule, and retaining the concatemer / sequencing primer duplex, (5) contacting the retained concatemer / sequencing primer duplex with a second sequencing polymerase and unlabeled chain-terminating nucleotides under conditions suitable for polymerase-catalyzed incorporation of the unlabeled chain-terminating nucleotides into the end of the sequencing primer, and (6) removing the chain-terminating moiety (e.g., deblocking) and retaining the concatemer / sequencing primer duplex. In some embodiments, no more than 2 to 1000 sequencing cycles are performed on the plurality of concatemers inside the cell sample to generate a plurality of sequencing read products. In some embodiments, the sequence of the first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of a first target RNA molecule in the cell sample. In some embodiments, the sequence of the second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of a second target RNA molecule in the cell sample. In some embodiments, after each round of generating the first and second sequencing read products with a length of no more than 30 to 1000 bases, or after a set of repetitive sequencing read products of the first and second sequencing read products with a length of no more than 150 or no more than 1000 bases, the sequences of the first and second sequencing read products can be aligned. In some embodiments, the sequencing reaction is performed on a sequencing device having a detector that captures the fluorescence signal from the sequencing reaction inside the cell sample.The sequencing device can be configured to relay fluorescence signal data captured by a detector to a computer system programmed to display images of different fluorescent spots co-located in a cell sample, where a single fluorescent spot corresponds to a different target RNA molecule. In some embodiments, a single cycle time can be achieved in less than 30 minutes. In some embodiments, the field of view (FOV) can exceed 1 mm. 2 , and the cycle time for scanning a large area (>10 mm 2 ) is less than 5 minutes.
[0382] In some embodiments, when sequencing using a multivalent molecule with a detectable label, the multivalent binding complex is formed in step (2), and the bound multivalent molecule with the detectable label is imaged and detected in step (3). Compared to a sequencing workflow using chain-terminating nucleotides with a detectable label, the conditions are mild. For example, steps (2) and (3) can be carried out at a mild temperature of about 35 to 45 °C or about 39 to 42 °C. Steps (2) and (3) can be carried out at a mild temperature, which can help maintain the compact size and shape of DNA nanoballs during multiple sequencing cycles (e.g., up to 30 cycles), thereby improving the full width at half maximum (FWHM) of the dot images of DNA nanoballs inside the cell sample. In some embodiments, the DNA nanoballs do not undergo strand separation during multiple sequencing cycles. In some embodiments, the dot images of DNA nanoballs do not enlarge during multiple sequencing cycles. In some embodiments, the dot images of DNA nanoballs remain discrete dots during multiple sequencing cycles. The dot images can be represented as Gaussian dots, and the size can be measured as FWHM. A smaller dot size indicated by a smaller FWHM is generally associated with an improved image of the dot. In some embodiments, the FWHM of the nanoball dots can be about 10 µm or less.
[0383] In some embodiments, asynchronous lag and / or leading events may occur during a synchronous polymerase-catalyzed sequencing reaction using a multivalent molecule with a detectable label. During sequencing, a fluorescence signal can be detected that corresponds to binding a complementary nucleotide unit of the multivalent molecule to the complex polymerase, thereby forming a multivalent binding complex. Thus, the binding of the labeled multivalent molecule can be used to detect and monitor lag and leading events. In some embodiments, when performing up to 30 or 150 sequencing cycles with a multivalent molecule with a detectable label, the lag and / or leading rate can be less than about 5%, or less than about 1%, or less than about 0.01%, or less than about 0.001%. In contrast, the lag and / or leading rate for up to 30 sequencing cycles using labeled chain-terminating nucleotides can be about 5%.
[0384] Methods for performing in - situ RNA bulk sequencing
[0385] The present disclosure provides methods for in situ multiplex and multi-omics detection and identification using encoded padlock probes. The padlock probes are designed to selectively detect target RNAs.
[0386] RNA-specific padlock probes selectively hybridize to cDNA corresponding to the target RNA. The RNA-specific probes carry barcodes that uniquely identify the cDNA. In some embodiments, the RNA-specific padlock probes also carry batch-specific sequencing primer binding sites.
[0387] Both types of padlock probes are used to generate concatemers of multiple copies with batch-specific sequencing binding sites and barcodes. The concatemers can collapse into DNA nanoballs with a compact shape and size, which produce increased signal intensity and color discrimination during sequencing.
[0388] For in situ sequencing, the limitation of optical resolution hinders the ability to perform highly multiplexed sequencing. The batch-specific sequencing primer binding sites on the padlock probes enable the use of selected batch-specific sequencing primers to sequence a desired subset of concatemers (e.g., a batch) to reduce overcrowded signals and images. The use of batch-specific sequencing primers can generate clear and resolvable optical images. By performing multiple rounds of sequencing on the same cell sample using different batch-specific sequencing primers, multiplexed sequencing can reveal a large number of target RNAs.
[0389] The batch-specific sequencing methods described herein have multiple uses. For example, the number of imaging and sequencing-related spots can be counted. The counted spots can be used as a measure of the RNA level in the cell sample.
[0390] The present disclosure provides methods for in situ detecting at least two different target RNA molecules, comprising the step (a): providing a cell sample deposited on a solid support, wherein the cell sample contains (i) a first plurality of DNA amplicons (e.g., a first concatemer) corresponding to a first target cDNA or RNA molecule and (ii) a second plurality of DNA amplicons (e.g., a second concatemer) corresponding to a second target cDNA or RNA molecule.
[0391] In some embodiments, the method further comprises the step (b): sequencing the first plurality of DNA amplicons inside the cell sample under conditions that inhibit the sequencing of the second plurality of DNA amplicons, wherein sequencing the first plurality of DNA amplicons inside the cell sample includes generating a plurality of first sequencing read products, and wherein the sequence of the first sequencing read products is aligned with a first target reference sequence to confirm the presence of the first target RNA in the cell sample. In some embodiments, the first amplicons can be repeatedly sequenced by performing 2 to 1000 sequencing cycles or can be repeatedly sequenced by performing 1 to 250 sequencing cycles.
[0392] In some embodiments, the method further comprises step (c): sequencing a second plurality of DNA amplicons within a cell sample under conditions that inhibit sequencing of the first plurality of DNA amplicons, wherein sequencing the second plurality of DNA amplicons within the cell sample comprises generating a plurality of second sequencing read products, and wherein the sequences of the second sequencing read products are aligned to a second target reference sequence to confirm the presence of a second target RNA in the cell sample. In some embodiments, the second amplicon may be repetitively sequenced by performing 2 to 1000 sequencing cycles or may be repetitively sequenced by performing 1 to 250 sequencing cycles.
[0393] The present disclosure provides methods for in situ detection of at least two different target RNA molecules, which comprise step (a): providing a cell sample deposited on a solid support, wherein the cell sample contains a first plurality of target RNAs and a second plurality of target RNAs. In some embodiments, the first plurality of target RNAs encode a first polypeptide. In some embodiments, the second plurality of target RNAs encode a second polypeptide. In some embodiments, the cell sample is fixed and permeabilized.
[0394] In some embodiments, the cell sample contains 2 to 25 different target RNA molecules, or contains 25 to 50 different target RNA molecules, or contains 50 to 75 different target RNA molecules or contains 75 to 100 different target RNA molecules. In some embodiments, the cell sample contains more than 100 different target RNA molecules, or more than 250 different target RNA molecules, or more than 500 different target molecules, or more than 1000 different target RNA molecules or more. In some embodiments, the cell sample contains more than 10,000 different target RNA molecules. In some embodiments, the cell sample comprises whole cells, a plurality of whole cells, intact tissue or an intact tumor. In some embodiments, the cell sample comprises a fresh cell sample, a fresh frozen cell sample, a sectioned cell sample or an FFPE cell sample. In some embodiments, the cell sample is deposited onto a solid support. In some embodiments, the cell sample is deposited onto a solid support passivated with a coating that promotes cell adhesion. In some embodiments, the cell sample is deposited onto a support lacking immobilized capture oligonucleotides. In some embodiments, the cell sample is cultured prior to performing step (b) described below.
[0395] In some embodiments, the cell sample contains from 2 to 25 different target polypeptide molecules, or from 25 to 50 different target polypeptide molecules, or from 50 to 75 different target polypeptide molecules or from 75 to 100 different target polypeptide molecules. In some embodiments, the cell sample contains more than 100 different target polypeptide molecules, or more than 250 different target polypeptide molecules, or more than 500 different target molecules, or more than 1000 different target polypeptide molecules or more. In some embodiments, the cell sample contains more than 10,000 different target polypeptide molecules. The target polypeptide molecules are encoded by target RNA molecules.
[0396] In some embodiments, the method includes step (b): generating a plurality of cDNAs (e.g., Figure 20 ) within the cell sample by (i) generating at least a first plurality of target cDNAs from a first plurality of target RNAs and (ii) generating at least a second plurality of target cDNAs from a second plurality of target RNAs. In some embodiments, the first target cDNA corresponds to a first target RNA molecule. In some embodiments, the second target cDNA corresponds to a second target RNA molecule. In some embodiments, the method includes generating at least 2 to 10,000 different target cDNA molecules corresponding to 2 to 10,000 different target RNA molecules. In some embodiments, the generation of step (b) includes contacting a plurality of RNAs within the cell sample with (i) a plurality of reverse transcription primers, (ii) a plurality of reverse transcriptases, and (iii) a plurality of nucleotides under conditions suitable for performing a reverse transcription reaction to generate a plurality of cDNA molecules (e.g., a plurality of first-strand cDNA molecules) within the cell sample. In some embodiments, the plurality of reverse transcription primers includes a first subset of target-specific reverse transcription primers that selectively hybridize to the first target RNA, and / or includes a second subset of target-specific reverse transcription primers that selectively hybridize to the second target RNA. In some embodiments, the plurality of reverse transcription primers includes a first subset of random-sequence reverse transcription primers that hybridize to the first target RNA, and / or includes a second subset of random-sequence reverse transcription primers that hybridize to the second target RNA.
[0397] In some embodiments, the method includes step (c): generating, inside a cell sample, a plurality of DNA concatemers corresponding to first and second pluralities of target RNA molecules, the generating including: (1) generating a first plurality of covalently closed circular padlock probes by contacting a first plurality of target cDNAs with a first plurality of padlock probes, wherein the contacting is carried out under conditions suitable for hybridizing a first binding arm and a second binding arm of a first padlock probe to proximal positions on their respective first target cDNA molecules to form a first plurality of circular padlock probes, each circular padlock probe having a nick or gap between the hybridized first and second binding arms, wherein the first padlock probe comprises (i) a first target barcode sequence (target BC-1) uniquely identifying the first target RNA or cDNA, (ii) a first batch-specific sequencing primer binding site (batch Seq-1) (or its complementary sequence), and (iii) a universal binding site for an amplification primer (universal RCA) (or its complementary sequence) (e.g., Figure 20 , on the left); (2) enzymatically closing the nick or gap in the first plurality of covalently closed circular padlock probes to form a first plurality of covalently closed padlock probes; and (3) performing rolling circle amplification inside the cell sample using the first covalently closed circular padlock probes as template molecules, thereby generating a first plurality of concatemer molecules corresponding to the first plurality of target RNA or cDNA molecules. In some embodiments, the rolling circle amplification reaction can be carried out in the presence or absence of a plurality of compaction oligonucleotides. In some embodiments, the method includes contacting a plurality of cDNA molecules in the cell sample with at least 2 to 10,000 different target-specific padlock probes. In some embodiments, the first padlock probe further comprises a universal compaction oligonucleotide binding site (or its complementary sequence). In some embodiments, closing the nick in the first circular padlock probe includes performing an enzymatic ligation reaction. In some embodiments, closing the nick in the first circularized padlock probe includes using the first target cDNA molecule as a template for a polymerase-catalyzed fill-in reaction and performing an enzymatic ligation reaction. In some embodiments, the method includes enzymatically closing the nick or nicks in at least 2 to 10,000 circularized target-specific padlock probes inside the cell sample, thereby generating at least 2 to 10,000 covalently closed circular padlock probes. In some embodiments, each concatemer molecule in the first plurality comprises a tandem repeat unit, wherein the unit comprises the sequence of the first target cDNA and (i) a first target barcode sequence (target BC-1) uniquely identifying the first target RNA, (ii) a first batch-specific sequencing primer binding site (batch Seq-1) (or its complementary sequence), and (iii) a universal binding site for an amplification primer (universal RCA) (or its complementary sequence). In some embodiments, the unit comprises a universal compaction oligonucleotide binding site (or its complementary sequence).
[0398] In some embodiments, step (c) further comprises: generating, inside the cell sample, a plurality of DNA concatemers corresponding to a second plurality of target RNA molecules, the generating comprising: (1) generating a second plurality of covalently closed circular padlock probes by contacting the second plurality of target cDNAs with a second plurality of padlock probes, wherein the contacting is carried out under conditions suitable for hybridizing the first and second binding arms of the second padlock probes to proximal positions on their respective second target cDNA molecules to form a second plurality of circular padlock probes, each circular padlock probe having a nick or gap between the hybridized first and second binding arms, wherein the second padlock probe comprises (i) a second barcode sequence (target BC-2) that uniquely identifies the second target cDNA or RNA, (ii) a second batch-specific sequencing primer binding site (batch Seq-2) (or its complementary sequence), wherein the sequence of the second batch-specific sequencing primer binding site is different from the sequence of the first batch-specific sequencing primer binding site, and (iii) a universal binding site for amplification primers (universal RCA) (or its complementary sequence) (e.g., Figure 20 , on the right); (2) enzymatically closing the nick or gap in the second plurality of covalently closed circular padlock probes to form a second plurality of covalently closed padlock probes; and (3) performing rolling circle amplification inside the cell sample using the second covalently closed circular padlock probes as template molecules, thereby generating a second plurality of concatemer molecules corresponding to the second plurality of target RNA or cDNA molecules. In some embodiments, the rolling circle amplification reaction can be carried out in the presence or absence of a plurality of compaction oligonucleotides. In some embodiments, the method comprises contacting a plurality of cDNA molecules in the cell sample with at least 2 to 10,000 different target-specific padlock probes. In some embodiments, the second padlock probe further comprises a universal compaction oligonucleotide binding site (or its complementary sequence). In some embodiments, closing the nick in the second circular padlock probe comprises performing an enzymatic ligation reaction. In some embodiments, closing the nick in the second circularized padlock probe comprises performing a polymerase-catalyzed fill-in reaction using the second target cDNA molecule as a template and performing an enzymatic ligation reaction. In some embodiments, the method comprises enzymatically closing the nick or nicks in at least 2 to 10,000 circularized target-specific padlock probes to generate at least 2 to 10,000 covalently closed circular padlock probes inside the cell sample. In some embodiments, each concatemer molecule in the second plurality comprises a tandem repeat unit, wherein the unit comprises the sequence of the second target cDNA and (i) a second target barcode sequence (target BC-2) that uniquely identifies the second target cDNA or RNA, (ii) a second batch-specific sequencing primer binding site (batch Seq-2) (or its complementary sequence), and (iii) a universal binding site for amplification primers (universal RCA) (or its complementary sequence). In some embodiments, the unit comprises a universal compaction oligonucleotide binding site (or its complementary sequence).
[0399] In some embodiments, the method further comprises step (d): sequencing a first plurality of concatemer molecules within a cell sample under conditions that inhibit sequencing of a second plurality of concatemers (e.g., Figure 21 ). In some embodiments, step (d) comprises sequencing a first plurality of concatemers within a cell sample, the sequencing comprising performing 2 to 1000 sequencing cycles to generate a plurality of first sequencing read products, wherein the sequences of the first sequencing read products are aligned with a first target reference sequence to confirm the presence of a first target RNA in the cell sample. In some embodiments, step (d) comprises sequencing a first plurality of concatemers within a cell sample comprising performing 1 to 250 sequencing cycles to generate a plurality of first sequencing read products, wherein the sequences of the first sequencing read products are aligned with a first target reference sequence to confirm the presence of a first target RNA in the cell sample.
[0400] In some embodiments, in step (d), within the first concatemer molecule, only the first target barcode region (target BC-1) is sequenced. In some embodiments, within the first concatemer molecule, at least a portion or the full length of the first target barcode (target BC-1) is sequenced. In some embodiments, within the first concatemer molecule, the first target barcode (target BC-1) is sequenced and a portion of the first cDNA region is sequenced.
[0401] In some embodiments, sequencing the first concatemer in step (d) comprises step (1) contacting the first plurality of concatemer molecules within the cell sample with (i) a plurality of first batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridization of the plurality of first batch-specific sequencing primers to their respective first batch-specific sequencing primer binding sites on the first concatemer. In some embodiments, the sequencing further comprises step (2) using the first concatemer as a template molecule to perform 2 to 1000 sequencing cycles to generate a first plurality of sequencing read products.
[0402] In some embodiments, the sequencing of step (d) comprises sequencing at least a portion of the first nucleic acid concatemer using an optical imaging system having a field of view (FOV) of greater than 1.0 mm 2 .
[0403] In some embodiments, in the sequencing of step (d), the plurality of first sequencing read products are image-detectable, and wherein the sequencing comprises decoding the plurality of first sequencing read products from images obtained during no more than 2 to 30 sequencing cycles or from images obtained during 1 to 1000 sequencing cycles.
[0404] In some embodiments, the method further includes step (e): removing a plurality of first sequencing read products from the first concatemer molecule and retaining the first concatemer molecule inside the cell sample. In some embodiments, a 3' blocking moiety can be added to the first sequencing read products to inhibit further sequencing reactions. For example, nucleotide analogs can be incorporated, where the nucleotide analog inhibits the incorporation of subsequent nucleotides. Exemplary blocking nucleotide analogs include dideoxynucleotides or nucleotides having 2' or 3' chain terminating moieties.
[0405] In some embodiments, the method further includes step (f): repeating the sequencing of the plurality of first concatemers by repeating steps (d) and (e) at least once. In some embodiments, the repeated sequencing of step (f) is optional.
[0406] In some embodiments, sequencing the first concatemer in step (f) includes step (1) contacting the first plurality of concatemer molecules inside the cell sample with (i) a plurality of first batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of first batch-specific sequencing primers to their corresponding first batch-specific sequencing primer binding sites on the first concatemer. In some embodiments, the sequencing further includes step (2) performing 2 to 1000 sequencing cycles using the first concatemer as a template molecule to generate the first plurality of sequencing read products. In some embodiments, the sequencing further includes step (3) removing the first plurality of sequencing read products from the first concatemer and retaining the plurality of first concatemers inside the cell sample. In some embodiments, the sequencing further includes step (4) repeating steps (1) to (3) at least once (e.g., Figure 21 ). In some embodiments, step (4) includes repeating steps (1) to (3) at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, or at least 10 times. In some embodiments, step (4) includes repeating steps (1) to (3) at most 10 times, at most 20 times, at most 30 times, at most 40 times, or at most 50 times.
[0407] In some embodiments, the repeated sequencing of the first concatemer in step (f) can be performed using a binding sequencing procedure, labeled and / or unlabeled chain terminating nucleotides, or multivalent molecules. Descriptions of these three methods for sequencing are provided below.
[0408] In some embodiments, a plurality of universal sequencing primers can be hybridized to the concatemer template molecule with a hybridization reagent that comprises an SSC buffer (e.g., 2X saline-sodium citrate) buffer having formamide (e.g., 10 to 20% formamide). The hybridization conditions include a temperature of about 20 to 30 °C for about 10 to 60 minutes.
[0409] In some embodiments, at a temperature that promotes nucleic acid denaturation (e.g., 30 to 90 °C), a dehybridization reagent containing an SSC buffer (e.g., saline-sodium citrate buffer) (with or without formamide) can be used to remove multiple sequencing read products from the concatemers and multiple concatemers can be retained inside the cell sample.
[0410] In some embodiments, the method further comprises step (g): sequencing a second plurality of concatemer molecules inside the cell sample under conditions that inhibit sequencing of the first plurality of concatemers (e.g., Figure 21 ). In some embodiments, step (g) comprises sequencing a second plurality of concatemers inside the cell sample, the sequencing comprising performing 2 to 1000 sequencing cycles to generate a plurality of second sequencing read products, wherein the sequences of the second sequencing read products are aligned with a second target reference sequence to confirm the presence of a second target RNA in the cell sample. In some embodiments, step (g) comprises sequencing a second plurality of concatemers inside the cell sample, the sequencing comprising performing 1 to 250 sequencing cycles to generate a plurality of second sequencing read products, wherein the sequences of the second sequencing read products are aligned with a second target reference sequence to confirm the presence of a second target RNA in the cell sample.
[0411] In some embodiments, in step (g), in the second concatemer molecule, only the second target barcode region (target BC-2) is sequenced. In some embodiments, in the second concatemer molecule, at least a portion or the full length of the second target barcode (target BC-2) is sequenced. In some embodiments, in the second concatemer molecule, the second target barcode (target BC-2) is sequenced and a portion of the second cDNA region is sequenced.
[0412] In some embodiments, sequencing of the first concatemer in step (g) comprises step (1) contacting the second plurality of concatemer molecules inside the cell sample with (i) a plurality of first batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridization of the plurality of second batch-specific sequencing primers to their corresponding second batch-specific sequencing primer binding sites on the second concatemer. In some embodiments, the sequencing further comprises step (2): performing 2 to 1000 sequencing cycles using the second concatemer as a template molecule to generate a second plurality of sequencing read products.
[0413] In some embodiments, the sequencing of step (g) comprises sequencing at least a portion of the second nucleic acid concatemer using an optical imaging system having a field of view (FOV) greater than 1.0 mm 2 ).
[0414] In some embodiments, in the sequencing of step (g), multiple second sequencing read products are detectable by imaging, and wherein the sequencing comprises decoding multiple second sequencing read products from images obtained during no more than 2 to 30 sequencing cycles or from images obtained during 1 to 1000 sequencing cycles.
[0415] In some embodiments, the method further comprises step (h): removing multiple second sequencing read products from the second concatemeric molecules and retaining the second concatemeric molecules inside the cell sample. In some embodiments, a 3' blocking moiety can be added to the second sequencing read products to inhibit further sequencing reactions. For example, a nucleotide analogue can be incorporated, wherein the nucleotide analogue inhibits the incorporation of subsequent nucleotides. Exemplary blocking nucleotide analogues include dideoxynucleotides or nucleotides having a 2' or 3' chain terminating moiety.
[0416] In some embodiments, the method further comprises step (i): performing repeated sequencing of the multiple second concatemers by repeating steps (g) and (h) at least once. In some embodiments, the repeated sequencing of step (i) is optional.
[0417] In some embodiments, sequencing the first concatemer in step (i) comprises step (1) contacting the second plurality of concatemeric molecules inside the cell sample with (i) a plurality of first batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of second batch-specific sequencing primers to their corresponding second batch-specific sequencing primer binding sites on the second concatemers. In some embodiments, the sequencing further comprises step (2): performing 2 to 1000 sequencing cycles using the second concatemer as a template molecule to generate a first plurality of sequencing read products. In some embodiments, the sequencing further comprises step (3): removing the first plurality of sequencing read products from the second concatemer and retaining the plurality of second concatemers inside the cell sample. In some embodiments, the sequencing further comprises step (4) repeating steps (1) to (3) at least once (e.g., Figure 21 ). In some embodiments, step (4) comprises repeating steps (1) to (3) at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, or at least 10 times. In some embodiments, step (4) comprises repeating steps (1) to (3) at most 10 times, at most 20 times, at most 30 times, at most 40 times, or at most 50 times.
[0418] In some embodiments, the repeated sequencing of the second concatemer in step (i) can be performed using a binding sequencing procedure, labeled and / or unlabeled chain-terminating nucleotides, or multivalent molecules. Descriptions of these three methods for sequencing are provided below.
[0419] In some embodiments, the plurality of nucleotide reagents of steps (d) and (g) include a plurality of nucleotides that are detectably labeled or unlabeled. In some embodiments, a single nucleotide is linked to a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the plurality of detectably labeled nucleotide analogs comprise a plurality of chain-terminating nucleotides, wherein the chain-terminating moiety is linked to the 3'-nucleotide sugar position to form a 3'-blocked nucleotide analog. In some embodiments, the chain-terminating moiety can be removed to convert the 3'-blocked nucleotide analog into an extendable nucleotide having a 3'-OH group on the sugar. In some embodiments, the labeled nucleotide analogs are linked to different fluorophores corresponding to the nucleobases adenine, cytosine, guanine, thymine, or uracil, wherein the different fluorophores emit fluorescence signals. In some embodiments, a sequencing cycle comprises (1) contacting the tandem / sequencing primer duplex with a sequencing polymerase and a detectably labeled chain-terminating nucleotide under conditions suitable for polymerase-catalyzed incorporation of the detectably labeled chain-terminating nucleotide at the end of the sequencing primer, (2) detecting and imaging the fluorescence signal and color emitted by the incorporated chain-terminating nucleotide, and (3) removing the chain-terminating moiety (e.g., deblocking) and the fluorophore from the incorporated nucleotide while retaining the tandem / sequencing primer duplex. In some embodiments, no more than 2 to 30 or no more than 1000 sequencing cycles are performed on the plurality of tandems inside the cell sample to generate a plurality of sequencing read products. In some embodiments, the sequence of a first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of a first target RNA molecule within the cell sample. In some embodiments, the sequence of a second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of a second target RNA molecule within the cell sample.
[0420] In some embodiments, after each round of generating first and second sequencing read products that are no more than 30 to 150 bases in length, or after a set of repeated sequencing read products of first and second sequencing read products that are no more than 30 to 150 bases in length, the sequences of the first and second sequencing read products can be aligned. In some embodiments, the sequencing reaction is performed on a sequencing device having a detector that captures fluorescence signals from the sequencing reaction inside the cell sample. The sequencing device can be configured to relay the fluorescence signal data captured by the detector to a computer system that is programmed to display an image of different fluorescent spots co-located within the cell sample, wherein a single fluorescent spot corresponds to a different target RNA molecule. In some embodiments, when sequencing is performed using different fluorescently labeled nucleotide reagents corresponding to different nucleobases (e.g., A, G, C, T / U), then the image can have different color fluorescent spots co-located within the same cell sample at different sequencing cycles.
[0421] In some embodiments, asynchronous lagging and / or leading events can occur during the synchronous sequencing reaction of the clonally amplified template amplicons, where the sequencing reaction comprises a polymerase-catalyzed sequencing reaction using chain-terminating nucleotides with detectable labels. In some embodiments, the sequencing reaction on one of the template molecules among the clonally amplified template molecules is ahead (e.g., leading) or behind (e.g., lagging) the sequencing of other template molecules within the clonally amplified template molecules. During sequencing, fluorescence signals corresponding to the incorporation of the labeled chain-terminating nucleotides are typically detected. Thus, the incorporation of the labeled chain-terminating nucleotides can be used to detect and monitor lagging and leading events.
[0422] In some embodiments, the plurality of nucleotide reagents of steps (d) and (g) include a plurality of multivalent molecules, each multivalent molecule comprising a core attached to a plurality of nucleotide arms, wherein the nucleotide arms are attached to nucleotide units. In some embodiments, a single multivalent molecule is labeled with a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the core of the multivalent molecule is labeled with a fluorophore, and wherein the fluorophore attached to a given core of the multivalent molecule corresponds to the nucleobase of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or a nucleobase attached to a fluorophore, and wherein the fluorophore attached to a given nucleobase corresponds to the nucleobase of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, the sequencing cycle includes: (1) contacting the concatemer / sequencing primer duplex with a first sequencing polymerase to form a complex polymerase, (2) contacting the complex polymerase with the detectable-labeled multivalent molecule under conditions suitable for binding complementary nucleotide units of the multivalent molecule to the complex polymerase to form a multivalent binding complex, and the conditions are suitable for inhibiting the incorporation of complementary nucleotide units into the end of the sequencing primer, (3) detecting and imaging the fluorescence signal and color emitted by the bound detectable-labeled multivalent molecule, (4) removing the first sequencing polymerase and the bound detectable-labeled multivalent molecule, and retaining the concatemer / sequencing primer duplex, (5) contacting the retained concatemer / sequencing primer duplex with a second sequencing polymerase and an unlabeled chain-terminating nucleotide under conditions suitable for the polymerase-catalyzed incorporation of the unlabeled chain-terminating nucleotide into the end of the sequencing primer, and (6) removing the chain-terminating moiety (e.g., deblocking) and retaining the concatemer / sequencing primer duplex. In some embodiments, no more than 2 to 30 or no more than 1000 sequencing cycles are performed on the plurality of concatemers inside the cell sample to generate a plurality of sequencing read products. In some embodiments, the sequence of the first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of a first target RNA molecule in the cell sample. In some embodiments, the sequence of the second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of a second target RNA molecule in the cell sample. In some embodiments, after each round of generating the first and second sequencing read products with a length of no more than 30 to 1000 bases, or after generating a set of repeated sequencing read products of the first and second sequencing read products with a length of no more than 30 or no more than 1000 bases, the sequences of the first and second sequencing read products can be aligned. In some embodiments, the sequencing reaction is performed on a sequencing device having a detector that captures the fluorescence signal from the sequencing reaction inside the cell sample.The sequencing device can be configured to relay fluorescence signal data captured by a detector to a computer system programmed to display an image of different fluorescent spots co-located in a cell sample, where a single fluorescent spot corresponds to a different target RNA molecule. In some embodiments, a single cycle time can be achieved in less than 30 minutes. In some embodiments, the field of view (FOV) can exceed 1 mm. 2 , and for scanning a large area (>10 mm 2 ) with a cycle time of less than 5 minutes.
[0423] In any of the methods described herein, multiple RNAs or cDNAs inside a cell sample can be amplified to generate amplicons of the RNA or cDNA, where the amplicons contain concatemers. In some embodiments, multiple RNA or cDNA molecules inside a cell sample can be amplified by performing a padlock probe ligation and rolling circle amplification workflow. In some embodiments, the method includes contacting multiple RNA or cDNA molecules inside a cell sample with multiple padlock probes, the multiple padlock probes including a first plurality of target-specific padlock probes that hybridize to a first target RNA or cDNA molecule and a second plurality of target-specific padlock probes that hybridize to a second target RNA or cDNA molecule.
[0424] In some embodiments, the padlock probe comprises a single-stranded oligonucleotide. In some embodiments, the padlock probe comprises DNA, RNA, or DNA and RNA. In some embodiments, a single padlock probe comprises an internal region between a first and a second end region, where the internal region contains at least one universal adaptor sequence, which includes a sample barcode sequence, an amplification primer binding site, a sequencing primer binding site, a compaction oligonucleotide binding site, and / or a surface capture primer binding site ( Figure 13 ). In some embodiments, the padlock probe comprises at least one target barcode sequence that corresponds to a given target RNA or target cDNA to which the padlock probe binds. In some embodiments, the padlock probe comprises at least one unique identification sequence (e.g., unique molecular index (UMI)). In some embodiments, the padlock probe comprises at least one restriction enzyme recognition sequence.
[0425] In some embodiments, a padlock probe comprises a single-stranded nucleic acid molecule having two end regions (e.g., a first and a second binding arm) and an internal region. In some embodiments, the first end region of a single padlock probe has a first target-specific sequence that selectively hybridizes to a first region of a target RNA or target cDNA molecule, and the second end region of the single padlock probe has a second target-specific sequence that selectively hybridizes to a second region of the same target RNA or target cDNA molecule. In some embodiments, the internal region of the padlock comprises a target barcode sequence corresponding to a given target RNA or target cDNA (e.g., target BC-1 or target BC-2 for the left schematic and the right schematic, respectively). In some embodiments, the target barcode sequence uniquely identifies the target RNA or target cDNA. In some embodiments, the internal region of the padlock comprises a universal primer binding site (or its complementary sequence) for a rolling circle amplification primer. In some embodiments, the internal region of the padlock comprises a universal primer binding site (or its complementary sequence) for a rolling circle amplification primer. In some embodiments, the internal region of the padlock comprises a universal binding site (or its complementary sequence) for a compaction oligonucleotide. In some embodiments, the internal region of the padlock probe comprises a target barcode sequence and at least one universal primer binding site (e.g., for binding a sequencing primer, for binding a rolling circle amplification primer, and / or for binding a compaction oligonucleotide) in any arrangement and orientation ( Figure 13 , top and bottom).
[0426] In some embodiments, a single padlock probe comprises first and second end regions (e.g., a first and a second binding arm) that hybridize to portions of a target RNA or target cDNA molecule to form a plurality of RNA-padlock probe complexes or a plurality of cDNA-padlock probe complexes, wherein a single complex has first and second end probe regions that hybridize to proximal regions of the RNA or cDNA molecule to form a nick or gap between the first and second end probe termini. In some embodiments, the first end region of a single padlock probe has a first target-specific sequence that selectively hybridizes to a first region of a target RNA or cDNA molecule, and the second end region of the single padlock probe has a second target-specific sequence that selectively hybridizes to a second region of the same target RNA or cDNA molecule, wherein a nick or gap is formed between the hybridized first and second end regions, thereby circularizing the padlock probe (e.g., Figure 14 ).
[0427] In some embodiments, the padlock probe comprises standard nucleotides and / or nucleotide analogs. In some embodiments, the padlock probe is modified to confer resistance to nuclease degradation (e.g., ribonuclease degradation). For example, the padlock probe comprises at least one phosphorothioate diester bond at its 5' end, which can render the padlock probe resistant to nuclease degradation. In some embodiments, the padlock probe comprises 2 to 5 or more consecutive phosphorothioate diester bonds at its 5' end. In some embodiments, the padlock probe comprises at least one ribonucleotide and / or at least one 2'-O-methyl, 2'-O-methoxyethyl (MOE), 2'-fluorinated base nucleotide. In some embodiments, the padlock probe comprises a phosphorylated 3' end. In some embodiments, the padlock probe comprises at least one locked nucleic acid (LNA) base. In some embodiments, the padlock probe comprises a phosphorylated 5' end (e.g., using polynucleotide kinase).
[0428] In some embodiments, a single padlock probe in a set of padlock probes (e.g., multiple padlock probes) comprises first and second terminal regions that hybridize to the same target region of a target RNA or cDNA molecule to form multiple RNA-padlock probe complexes or multiple cDNA-padlock probe complexes having the same RNA or cDNA sequence.
[0429] In some embodiments, a set of padlock probes (e.g., multiple padlock probes) comprises at least two sub-sets of padlock probes. In some embodiments, a single padlock probe in the first sub-set of padlock probes comprises first and second terminal regions that hybridize to the same target region (e.g., a first target region) of a target RNA or cDNA molecule to form a first plurality of RNA-padlock probe complexes or a first plurality of cDNA-padlock probe complexes having the same RNA or cDNA sequence. In some embodiments, a single padlock probe in the second sub-set of padlock probes comprises first and second terminal regions that hybridize to the same target region (e.g., a second target region) of a target RNA or cDNA molecule to form a second plurality of RNA-padlock probe complexes or a second plurality of cDNA-padlock probe complexes having the same cDNA sequence. In some embodiments, the first and second sub-sets of padlock probes hybridize to different target regions of the same target RNA or cDNA molecule. In some embodiments, the first and second sub-sets of padlock probes hybridize to different target regions of different target RNA or cDNA molecules. In some embodiments, the set of padlock probes includes 2 to 10 sub-sets of padlock probes, or 10 to 25 sub-sets of padlock probes, or 25 to 50 sub-sets of padlock probes or up to 100 sub-sets of padlock probes. In some embodiments, the set of padlock probes includes at least 100 sub-sets of padlock probes, at least 500 sub-sets of padlock probes, at least 1000 sub-sets of padlock probes, at least 10,000 sub-sets of padlock probes or more sub-sets of padlock probes.
[0430] In some embodiments, the nicks can be ligated enzymatically to generate covalently closed circular padlock probes. In some embodiments, the ligase can distinguish between matched and mismatched hybridized ends to ensure target-specific hybridization. In some embodiments, the ligation reaction comprises using a ligase, including T3, T4, T7 or Taq DNA ligase.
[0431] In some embodiments, the gap size between the first and second terminal regions of the hybridization is 1 to 25 bases. The 3'OH end of the hybridized padlock probe can serve as a starting site for a polymerase-catalyzed fill-in reaction (e.g., gap filling reaction) that uses the target cDNA molecule (or target RNA molecule) as a template. After the fill-in reaction, the remaining nick can be ligated enzymatically to generate a covalently closed circular padlock probe.
[0432] In some embodiments, the gap filling reaction comprises contacting the circular padlock probe with a DNA polymerase and a plurality of nucleotides. In some embodiments, the DNA polymerase comprises Escherichia coli (E. coli) DNA polymerase I, the Klenow fragment of Escherichia coli DNA polymerase I, T7 DNA polymerase or T4 DNA polymerase. In some embodiments, the ligase can distinguish between matched and mismatched hybridized ends to ensure target-specific hybridization. In some embodiments, the ligation reaction comprises using a ligase, including T3, T4, T7 or Taq DNA ligase.
[0433] In any of the methods described herein, a plurality of covalently closed circular padlock probes can be subjected to a rolling circle amplification reaction to generate a plurality of concatemer molecules, each concatemer molecule having two or more tandem copies of a unit, where the unit comprises a target sequence corresponding to the target RNA molecule and any other sequences carried by the padlock probe, including a universal adapter sequence, a unique molecular index sequence, and / or a restriction enzyme recognition sequence.
[0434] In some embodiments, the rolling circle amplification reaction comprises contacting the covalently closed circularized padlock probe with an amplification primer (e.g., a universal rolling circle amplification primer), a strand-displacing DNA polymerase, and a plurality of nucleotides under conditions suitable for hybridizing a single amplification primer to the covalently closed padlock probe and under conditions suitable for primer extension using the covalently closed padlock probe as a template molecule to generate a nucleic acid concatemer. In some embodiments, the plurality of nucleotides in the rolling circle amplification reaction comprises any mixture of two or more of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, any of the rolling circle amplification reactions described herein can be carried out in the presence or absence of a plurality of compaction oligonucleotides.
[0435] In some embodiments, when the rolling circle amplification reaction includes a plurality of nucleotides containing dUTP, the resulting concatemers can be cross-linked to a cross-linking reaction group by treating the cell sample with succinimidyl esters (NHS), maleimides (sulfosuccinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate (sulfo-SMCC)), imidates (dimethyl adipimidate (DMP)), carbodiimides (dicyclohexylcarbodiimide (DCC), 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC)) or phenyl azides. In some embodiments, the polymerization of the cross-linking reaction group can be initiated by light or UV light. In some embodiments, the resulting concatemers can be cross-linked to a matrix by treating the cell sample with cross-linked agarose, cross-linked dextran or cross-linked polyethylene glycol (PEG), polyacrylamide, cellulose alginate or polyamide. In some embodiments, the PEG contains a sulfo-NHS ester moiety at one or both ends, for example, polyethylene glycolylated bis(sulfosuccinimidyl)suberate (e.g., BS(PEG)9 from Thermo Fisher Scientific, catalog number 21582).
[0436] In some embodiments, the rolling circle amplification reaction can be carried out at a constant temperature (e.g., isothermal), where the constant temperature is from room temperature to about 30 °C, or about 30 to 40 °C, or about 40 to 50 °C or about 50 to 65 °C.
[0437] In some embodiments, the DNA polymerase having strand displacement activity can be selected from the group consisting of phi29 DNA polymerase, the large fragment of Bst DNA polymerase, the large fragment of Bsu DNA polymerase and Bca (exo) DNA polymerase, the Klenow fragment of Escherichia coli DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV virus reverse transcriptase or Deep Vent DNA polymerase. In some embodiments, the phi29 DNA polymerase can be a wild-type phi29 DNA polymerase (e.g., MagniPhi from Expedeon) or a variant EquiPhi29 DNA polymerase (e.g., from Thermo Fisher Scientific) or a chimeric QualiPhi DNA polymerase (e.g., from 4basebio).
[0438] In some embodiments, the rolling circle amplification primers can be modified to increase resistance to nuclease degradation. In some embodiments, the rolling circle amplification primers contain at least one phosphorothioate diester bond at their 5' end, which can render the amplification primers resistant to exonuclease degradation. In some embodiments, the rolling circle amplification primers contain 2 to 5 or more consecutive phosphorothioate diester bonds at their 5' end. In some embodiments, the rolling circle amplification primers contain at least one ribonucleotide and / or at least one 2'-O-methyl or 2'-O-methoxyethyl (MOE) nucleotide.
[0439] In some embodiments, the rolling circle amplification reaction can be carried out in the presence of multiple compaction oligonucleotides, which, when hybridized to the concatemer molecule, compact the size and / or shape of the concatemer to form a compact nanosphere. In some embodiments, the compaction oligonucleotides comprise single-stranded oligonucleotides having a first region at one end that hybridizes to a portion of the concatemer molecule and a second region at the other end that hybridizes to another portion of the same concatemer molecule, wherein hybridization of the compaction oligonucleotide to a given concatemer compacts the size and / or shape of the concatemer.
[0440] The compaction oligonucleotide comprises a 5' region, optionally an internal region (intermediate region), and a 3' region...
Claims
1. A computer-implemented method for base calling in sequencing data analysis, comprising: obtaining, by a processor, a plurality of flow cell images of a sample at a plurality of z-levels along an axial axis, wherein each of the plurality of flow cell images is acquired at a corresponding z-level along the axial axis; generating, by the processor, a plurality of processed images of the plurality of flow cell images; filtering, by the processor, the plurality of flow cell images based on the plurality of processed images to generate a plurality of filtered images; generating, by the processor, a first maximum intensity projection (MIP) image based on the plurality of filtered images; and performing base calling by the processor using the first MIP image.
2. The computer-implemented method according to any one of the preceding claims, wherein the plurality of flow cell images are acquired using a next-generation sequencing (NGS) system.
3. The computer-implemented method according to any one of the preceding claims, wherein the plurality of flow cell images are acquired at one or more sequencing cycles different from a reference cycle.
4. The computer-implemented method according to any one of the preceding claims, wherein the plurality of flow cell images are acquired at a single sequencing cycle different from a reference cycle.
5. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels are spaced 0.1 um to 5 um apart from each other.
6. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels cover at least some of the thickness of the sample along the axial axis.
7. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels cover the entire thickness of the sample along the axial axis.
8. The computer-implemented method according to any one of the preceding claims, wherein each of the plurality of flow cell images has an image thickness of 0.1 um to 6 um.
9. The computer-implemented method according to any one of the preceding claims, wherein the sample is an in-situ sample fixed on a support of the flow cell.
10. The computer-implemented method according to any one of the preceding claims, wherein the in-situ sample comprises one or more cells or tissues.
11. The computer-implemented method according to any one of the preceding claims, wherein each of the plurality of flow cell images includes a field of view orthogonal to the axial axis, and wherein the field of view is two-dimensional (2D).
12. The computer-implemented method according to any one of the preceding claims, wherein the plurality of flow cell images are 2D.
13. The computer-implemented method according to any one of the preceding claims, wherein the field of view of each of the plurality of flow cell images is the same in the image plane.
14. The computer-implemented method according to any one of the preceding claims, wherein the field of view of each of the plurality of flow cell images covers at least a portion of a tile of the flow cell.
15. A computer-implemented method according to any of the preceding claims, wherein the plurality of flow cell images include the same image resolution.
16. A computer-implemented method according to any of the preceding claims, wherein the axial axis extends from the objective lens to a sample located on a flow cell positioned on the sequencing system.
17. A computer-implemented method according to any of the preceding claims, wherein the axial axis is orthogonal to the image plane, and wherein the field of view is within the image plane.
18. A computer-implemented method according to any of the preceding claims, wherein obtaining the plurality of flow cell images of the sample from a plurality of z-levels along the axial axis includes: Obtaining the plurality of flow cell images of the sample from a plurality of z-levels along the axial axis from a first color channel.
19. A computer-implemented method according to any of the preceding claims, wherein obtaining the plurality of flow cell images of the sample from a plurality of z-levels along the axial axis includes: Obtaining the plurality of flow cell images of the sample from a plurality of z-levels along the axial axis from 2, 3, or 4 color channels of the sequencing system.
20. A computer-implemented method according to any of the preceding claims, wherein the first MIP image corresponds to a first color channel of the sequencing system.
21. A computer-implemented method according to any of the preceding claims, wherein performing base calling by the processor using the first MIP image includes: Performing base calling by the processor using the first MIP image from the first color channel and one or more MIP images corresponding to one or more color channels different from the first color channel.
22. A computer-implemented method according to any of the preceding claims, wherein performing base calling by the processor using the first MIP image includes: Performing base calling by the processor using the first MIP image from the first color channel and corresponding MIP images for each color channel of the sequencing system different from the first color channel. A computer-implemented method according to any of the preceding claims, wherein obtaining the plurality of processed images includes: Selecting a kernel; And Generating the plurality of processed images by performing an opening operation on the plurality of flow cell images using the selected kernel.
23. A computer-implemented method according to any of the preceding claims, wherein obtaining the plurality of processed images includes: Selecting a kernel; And Generating the plurality of processed images by convolving the plurality of flow cell images with the selected kernel.
24. A computer-implemented method according to any of the preceding claims, wherein obtaining the plurality of processed images further includes: Selecting a first kernel and a second kernel; Generating a first blurred image by convolving the plurality of flow cell images with the first kernel; And Generating a second blurred image by convolving the plurality of flow cell images with the second kernel.
25. The computer-implemented method according to any one of the preceding claims, wherein obtaining the plurality of processed images comprises: Scaling the plurality of processed images.
26. The computer-implemented method according to any one of the preceding claims, wherein obtaining the plurality of processed images comprises: Scaling the first blurred image, the second blurred image, or both.
27. The computer-implemented method according to any one of the preceding claims, wherein filtering the plurality of flow cell images based on the plurality of processed images comprises: Subtracting the second blurred image from the first blurred image, thereby generating the plurality of filtered images.
28. The computer-implemented method according to any one of the preceding claims, wherein the kernel is 2×2, 3×3, 4×4, 5×5, or 6×6 pixels.
29. The computer-implemented method according to any one of the preceding claims, wherein the kernel is a circular kernel.
30. The computer-implemented method according to any one of the preceding claims, wherein the kernel is a Gaussian kernel.
31. The computer-implemented method according to any one of the preceding claims, wherein the first kernel and the second kernel are different Gaussian kernels.
32. The computer-implemented method according to any one of the preceding claims, wherein filtering the plurality of flow cell images based on the plurality of processed images comprises: Subtracting each of the processed images in the plurality of processed images from the corresponding flow cell image of the plurality of flow cell images, thereby generating the plurality of filtered images.
33. The computer-implemented method according to any one of the preceding claims, wherein filtering the plurality of flow cell images based on the plurality of processed images further comprises: Adding a predetermined offset to the subtracted images, thereby generating the plurality of filtered images.
34. The computer-implemented method according to any one of the preceding claims, wherein generating the first MIP image based on the plurality of filtered images comprises: Calculating the maximum intensity among the intensities of the plurality of filtered images at the corresponding pixels for each pixel of the first MIP image.
35. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Registering the first MIP image to one or more images of the sample.
36. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Registering the one or more MIP images corresponding to one or more color channels different from the first color channel to one or more images of the sample.
37. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise image intensities corresponding to cell components or structures.
38. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise staining of the following: membrane, cell nucleus, or a combination thereof.
39. The computer-implemented method according to any one of the preceding claims, wherein the one or more images include staining of one or more membrane proteins.
40. The computer-implemented method according to any one of the preceding claims, wherein the one or more images include staining of lipids.
41. The computer-implemented method according to any one of the preceding claims, wherein the one or more images include fluorescence signals from cell membranes.
42. The computer-implemented method according to any one of the preceding claims, wherein the segmentation of the one or more images includes: cells, membranes, cell nuclei, or combinations thereof.
43. The computer-implemented method according to any one of the preceding claims, wherein performing base calling using the first MIP image includes: Performing one or more preliminary analysis steps to adjust the image intensity of the polymerase colonies in the first MIP image or the one or more MIP images; And Performing base calling for the polymerase colonies based on the adjusted image intensity; Wherein the one or more preliminary analysis steps include: Background subtraction; Image sharpening; Intensity offset adjustment; Color correction; Intensity normalization; Lag and overshoot correction; Image registration; Quality score estimation; or Combinations thereof.
44. The computer-implemented method according to any one of the preceding claims, wherein the method further includes: Performing image registration of the plurality of flow cell images, the plurality of processed images, the plurality of filtered images, the first MIP image, the one or more MIP images, or combinations thereof.
45. The computer-implemented method according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images includes: Registering the first MIP image or the one or more MIP images to a template image.
46. The computer-implemented method according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images includes: Registering the plurality of flow cell images, the plurality of processed images, the plurality of filtered images, the first MIP image, the one or more MIP images, or combinations thereof to a template image.
47. The computer-implemented method according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images includes: Registering the polymerase colonies in the first MIP image to the template polymerase colonies in the template image.
48. The computer-implemented method according to any one of the preceding claims, wherein the method further includes: Obtaining a second MIP image by the processor based on the plurality of flow cell images; And Performing image registration of the plurality of flow cell images, the plurality of processed images, the plurality of filtered images, or combinations thereof based on the second MIP image.
49. The computer-implemented method according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images, the plurality of processed images, the plurality of filtered images, or combinations thereof includes: Register the second MIP image to the template image.
50. The computer-implemented method according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images based on the second MIP image comprises: Registering the polymerase colonies in the second MIP image to the template polymerase colonies in the template image.
51. The computer-implemented method according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images based on the first MIP image or the second MIP image comprises: Generating one or more template images in a reference coordinate system by registering the polymerase colonies in one or more reference cycles to the one or more template images using the coordinates of the polymerase colonies; Determining, by the processor, a plurality of transformations of the first MIP image or the second MIP image based on the one or more template images, the plurality of transformations corresponding to sub-patches of the first MIP or the second MIP and configured to register the sub-patches to the one or more template images; And Registering the sub-patches to the one or more template images using the plurality of transformations.
52. The computer-implemented method according to any one of the preceding claims, wherein the plurality of transformations comprises one or more affine transformations.
53. The computer-implemented method according to any one of the preceding claims, wherein each of the plurality of transformations comprises an affine transformation.
54. The computer-implemented method according to any one of the preceding claims, wherein performing base calling using the first MIP image comprises: Performing base calling based on the image intensity of the polymerase colonies from the first MIP image and the position information of the polymerase colonies of the second MIP image from the plurality of flow cell images.
55. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Performing image registration of the polymerase colonies of the plurality of flow cell images based on fiducial markers.
56. The computer-implemented method according to any one of the preceding claims, wherein the fiducial markers are located on the flow cell.
57. The computer-implemented method according to any one of the preceding claims, wherein the fiducial markers are outside the flow cell.
58. The computer-implemented method according to any one of the preceding claims, wherein the plurality of flow cell images are acquired at 2, 3, 4, 5, 6, 7, 8, 9, or 10 different positions along the axial axis.
59. The computer-implemented method according to any one of the preceding claims, wherein two adjacent positions along the axial axis are separated by approximately 1um, 2um, 3um, 4um, 5um, 6um, 7um, 8um, 9um, 10um, 11um, or 12um.
60. The computer-implemented method according to any one of the preceding claims, wherein the plurality of flow cell images are acquired from 1, 2, 3, 4, 5, or 6 channels.
61. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more processing units; one or more integrated circuits; or a combination thereof.
62. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); one or more neural processing units (NPUs); or a combination thereof.
63. The computer-implemented method according to any one of the preceding claims, further comprising: transmitting the base identification to a processing unit by the processor.
64. The computer-implemented method according to any one of the preceding claims, wherein the processing unit is a central processing unit (CPU).
65. The computer-implemented method according to any one of the preceding claims, wherein the processing unit is configured to register the base identification to one or more images.
66. The computer-implemented method according to any one of the preceding claims, further comprising: providing the sample having a plurality of tandem molecules immobilized on a support, wherein each tandem molecule corresponds to a target RNA of a cell sample.
67. The computer-implemented method according to any one of the preceding claims, wherein obtaining the plurality of flow cell images of the sample comprises: generating the plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on the sample immobilized on the support, wherein the plurality of flow cell images are generated from two or more color channels at two or more different z-levels along an axial axis.
68. The computer-implemented method according to any one of the preceding claims, wherein obtaining the plurality of flow cell images of the sample comprises: generating the plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on a plurality of tandem molecules of the sample immobilized on the support.
69. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a polymerase community immobilized thereon.
70. The computer-implemented method according to any one of the preceding claims, wherein the polymerase community corresponds to the plurality of nucleotide template molecules or tandem molecules.
71. The computer-implemented method according to any one of the preceding claims, further comprising: generating a flow cell image by the sequencing system by performing one or more sequencing reaction cycles on a plurality of tandem molecules of the sample immobilized on the support.
72. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of tandem molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleotide bases A, G, C, and T / U.
73. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: Contact the plurality of tandem molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of affinity bodies.
74. The computer-implemented method according to any one of the preceding claims, wherein the individual affinity bodies in the mixture comprise a core to which a plurality of nucleotide arms are attached, and each arm of the individual affinity bodies comprises the same type of nucleobase.
75. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, imaging, by an optical system, an optical color signal emitted from a nucleotide reagent bound to the plurality of tandem molecules.
76. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, acquiring, by an optical system, the flow cell image, the flow cell image comprising an optical color signal emitted from a nucleotide reagent bound to the plurality of tandem molecules.
77. The computer-implemented method according to any one of the preceding claims, wherein the flow cell image comprises optical signals emitted from nucleotide reagents with an unbalanced diversity of nucleobases A, G, C, and T / U in one or more cycles that are bound to the nucleobases of the plurality of templates or tandem molecules immobilized on the support.
78. The computer-implemented method according to any one of the preceding claims, wherein the plurality of polymerase populations comprise an unbalanced diversity of nucleobases A, G, C, and T / U, and wherein the unbalanced diversity comprises the percentage of (1) the number of one or more types of nucleobases in the region of the flow cell image to (2) the total number of nucleobases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the region.
79. The computer-implemented method according to any one of the preceding claims, further comprising: Provide the sample containing a plurality of RNAs, the plurality of RNAs comprising at least a first target RNA molecule and a second target RNA molecule.
80. The computer-implemented method according to any one of the preceding claims, further comprising: Generate a plurality of cDNA molecules within the sample, the plurality of cDNA molecules comprising at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules comprising a second target cDNA molecule corresponding to the second target RNA molecule.
81. The computer-implemented method according to any one of the preceding claims, further comprising: Contact the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes, the plurality of target-specific padlock probes comprising at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes.
82. The computer-implemented method according to any one of the preceding claims, further comprising: Close the nicks or gaps in the at least first and second circularized target-specific padlock probes by performing an enzymatic reaction, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe within the sample.
83. The computer-implemented method according to any one of the preceding claims, further comprising: Using the first and second covalently closed circular padlock probes as template molecules, a rolling circle amplification reaction is carried out inside the sample to generate a plurality of concatemer molecules, the plurality of concatemer molecules including at least a first concatemer molecule corresponding to a first target RNA molecule, and the plurality of concatemer molecules including at least a second concatemer molecule corresponding to a second target RNA molecule.
84. The computer-implemented method according to any one of the preceding claims, further comprising: Sequencing the plurality of concatemer molecules inside the sample, the sequencing including sequencing the first concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of second sequencing read products.
85. The computer-implemented method according to any one of the preceding claims, wherein sequencing the plurality of tandem molecules inside the sample comprises: Contacting the plurality of concatemer molecules inside the sample with (i) the plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridization of the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the concatemers.
86. The computer-implemented method according to any one of the preceding claims, wherein the nucleotide reagent comprises one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs.
87. The computer-implemented method according to any one of the preceding claims, further comprising: Removing the plurality of first sequencing read products from the first concatemer molecule and retaining the first concatemer molecule in the sample, and removing the plurality of second sequencing read products from the second concatemer molecule and retaining the second concatemer molecule in the sample.
88. A computer-implemented system for base calling in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including: Obtaining, by the processor, a plurality of flow cell images of a sample from a plurality of z levels along an axial axis, wherein each flow cell image of the plurality of flow cell images is acquired at a corresponding position along the axial axis; Generating, by the processor, a plurality of processed images of the plurality of flow cell images; Filtering, by the processor, the plurality of flow cell images based on the plurality of processed images to generate a plurality of filtered images; Generating, by the processor, a first maximum intensity projection (MIP) image based on the plurality of filtered images; and Performing base calling by the processor using the first MIP image.
89. A computer-implemented system for base calling in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations according to any one of the preceding claims.
90. One or more non - transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations for base calling in sequencing data analysis, the operations including: Obtaining, by the processor, a plurality of flow cell images of a sample from a plurality of z - levels along an axial axis, wherein each of the plurality of flow cell images is acquired at a corresponding position along the axial axis; Generating, by the processor, a plurality of processed images of the plurality of flow cell images; Filtering, by the processor, the plurality of flow cell images based on the plurality of processed images to generate a plurality of filtered images; Generating, by the processor, a first maximum intensity projection (MIP) image based on the plurality of filtered images; And Performing base calling by the processor using the first MIP image.
91. One or more non - transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations including any one of the preceding claims.
92. A computer - implemented method for base calling in sequencing data analysis, the method including: Obtaining, by the processor, a plurality of flow cell images of a sample from a plurality of z - levels along an axial axis, wherein each of the plurality of flow cell images is acquired at a corresponding position along the axial axis; Filtering, by the processor, the plurality of flow cell images with a top - hat filter, a difference of Gaussians (DoG) filter, or a Mexican hat filter to generate a plurality of filtered images; Generating, by the processor, a first maximum intensity projection (MIP) image based on the plurality of filtered images; And Performing base calling by the processor using the first MIP image.
93. A computer - implemented method for base calling in sequencing data analysis, the method including any one of the preceding claims.
94. A computer - implemented system for base calling in sequencing data analysis, the system including: One or more hardware processors; One or more data storage devices storing instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations including: Obtaining, by the processor, a plurality of flow cell images of a sample from a plurality of z - levels along an axial axis, wherein each of the plurality of flow cell images is acquired at a corresponding position along the axial axis; Filtering, by the processor, the plurality of flow cell images with a top - hat filter, a difference of Gaussians (DoG) filter, or a Mexican hat filter to generate a plurality of filtered images; Generating, by the processor, a first maximum intensity projection (MIP) image based on the plurality of filtered images; and Performing base calling by the processor using the first MIP image.
95. A computer - implemented system for base calling in sequencing data analysis, the system including: One or more hardware processors; One or more data storage devices that store instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including according to any one of the preceding claims.
96. One or more non-transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations for base calling in sequencing data analysis, the operations including: Obtaining, by a processor, a plurality of flow cell images of a sample from a plurality of z-levels along an axial axis, wherein each flow cell image of the plurality of flow cell images is acquired at a corresponding position along the axial axis; Filtering, by the processor, the plurality of flow cell images with a top-hat filter, a difference of Gaussians (DoG) filter, or a Mexican hat filter to generate a plurality of filtered images; Generating, by the processor, a first maximum intensity projection (MIP) image based on the plurality of filtered images; And Performing base calling by the processor using the first MIP image.
97. One or more non-transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations, the operations including any one of the preceding claims.
98. The computer-implemented method according to any one of the preceding claims, wherein the plurality of flow cell images are acquired using a next-generation sequencing (NGS) system.
99. The computer-implemented system according to any one of the preceding claims, wherein the plurality of flow cell images are acquired at a cycle different from a reference cycle.
100. The computer-implemented method according to any one of the preceding claims, wherein the plurality of flow cell images are acquired at a single sequencing cycle different from a reference cycle.
101. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels are spaced 0.1 um to 5 um apart from each other.
102. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels cover at least some of the thickness of the sample along the axial axis.
103. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels cover the entire thickness of the sample along the axial axis.
104. The computer-implemented method according to any one of the preceding claims, wherein each flow cell image of the plurality of flow cell images has an image thickness of 0.1 um to 6 um.
105. The computer-implemented system according to any one of the preceding claims, wherein the sample is an in-situ sample fixed on a flow cell.
106. The computer-implemented system according to any one of the preceding claims, wherein the in-situ sample comprises one or more cells or tissues.
107. The computer-implemented system according to any one of the preceding claims, wherein each of the plurality of flow cell images includes a field of view orthogonal to the axial axis, and wherein the field of view is two-dimensional.
108. The computer-implemented system according to any one of the preceding claims, wherein the field of view of each of the plurality of flow cell images is the same in the image plane.
109. The computer-implemented system according to any one of the preceding claims, wherein the field of view of each of the plurality of flow cell images covers at least a portion of a tile of the flow cell.
110. The computer-implemented system according to any one of the preceding claims, wherein the plurality of flow cell images have the same image resolution.
111. The computer-implemented system according to any one of the preceding claims, wherein the axial axis extends from the objective lens to a sample on a flow cell positioned on the sequencing system.
112. The computer-implemented system according to any one of the preceding claims, wherein the axial axis is orthogonal to the image plane, and wherein the field of view is within the image plane.
113. The computer-implemented system according to any one of the preceding claims, wherein obtaining the plurality of flow cell images of the sample from a plurality of z-levels along the axial axis includes: Obtaining the plurality of flow cell images of the sample from a plurality of z-levels along the axial axis from a first color channel.
114. The computer-implemented system according to any one of the preceding claims, wherein obtaining the plurality of flow cell images of the sample from a plurality of z-levels along the axial axis includes: Obtaining the plurality of flow cell images of the sample from a plurality of z-levels along the axial axis from 2, 3, or 4 color channels of the sequencing system.
115. The computer-implemented system according to any one of the preceding claims, wherein the first MIP image corresponds to a first color channel of the sequencing system.
116. The computer-implemented system according to any one of the preceding claims, wherein performing base calling by the processor using the first MIP image includes: Performing base calling by the processor using the first MIP image from the first color channel and one or more MIP images corresponding to one or more color channels different from the first color channel.
117. The computer-implemented system according to any one of the preceding claims, wherein performing base calling by the processor using the first MIP image includes: Performing base calling by the processor using the first MIP image from the first color channel and corresponding MIP images for each color channel of the sequencing system different from the first color channel.
118. The computer-implemented system according to any one of the preceding claims, wherein obtaining the plurality of processed images includes: Selecting a kernel; And Generating the plurality of processed images by performing an opening operation on the plurality of flow cell images using the selected kernel.
119. A computer-implemented system according to any of the preceding claims, wherein obtaining the plurality of processed images comprises: selecting a kernel; and generating the plurality of processed images by convolving the plurality of flow cell images with the selected kernel.
120. A computer-implemented system according to any of the preceding claims, wherein obtaining the plurality of processed images further comprises: selecting a first kernel and a second kernel; generating a first blurred image by convolving the plurality of flow cell images with the first kernel; and generating a second blurred image by convolving the plurality of flow cell images with the second kernel.
121. A computer-implemented system according to any of the preceding claims, wherein obtaining the plurality of processed images comprises: scaling the plurality of processed images.
122. A computer-implemented system according to any of the preceding claims, wherein obtaining the plurality of processed images comprises: scaling the first blurred image, the second blurred image, or both.
123. A computer-implemented system according to any of the preceding claims, wherein filtering the plurality of flow cell images based on the plurality of processed images comprises: subtracting the second blurred image from the first blurred image, thereby generating the plurality of filtered images.
124. A computer-implemented system according to any of the preceding claims, wherein the kernel is 2×2, 3×3, 4×4, 5×5, or 6×6 pixels.
125. A computer-implemented system according to any of the preceding claims, wherein the kernel is a circular kernel.
126. A computer-implemented system according to any of the preceding claims, wherein the kernel is a Gaussian kernel.
127. A computer-implemented system according to any of the preceding claims, wherein the first kernel and the second kernel are different Gaussian kernels.
128. A computer-implemented system according to any of the preceding claims, wherein filtering the plurality of flow cell images based on the plurality of processed images comprises: subtracting each of the processed images in the plurality of processed images from the corresponding flow cell image of the plurality of flow cell images, thereby generating the plurality of filtered images.
129. A computer-implemented system according to any of the preceding claims, wherein filtering the plurality of flow cell images based on the plurality of processed images further comprises: adding a predetermined offset to the subtracted images, thereby generating the plurality of filtered images.
130. A computer-implemented system according to any of the preceding claims, wherein generating the first MIP image based on the plurality of filtered images comprises: calculating, for each pixel of the first MIP image, the maximum intensity among the intensities of the plurality of filtered images at the corresponding pixel.
131. A computer-implemented system according to any of the preceding claims, wherein the method further comprises: registering the first MIP image to one or more images of the sample.
132. The computer-implemented system according to any one of the preceding claims, wherein the one or more images include staining of the following: a membrane, a cell nucleus, or a combination thereof.
133. The computer-implemented system according to any one of the preceding claims, wherein the one or more images include staining of one or more membrane proteins.
134. The computer-implemented system according to any one of the preceding claims, wherein the one or more images include staining of lipids.
135. The computer-implemented system according to any one of the preceding claims, wherein the one or more images include fluorescence signals from a cell membrane.
136. The computer-implemented system according to any one of the preceding claims, wherein the one or more images include segmentation of the following: a cell, a membrane, a cell nucleus, or a combination thereof.
137. The computer-implemented system according to any one of the preceding claims, wherein performing base calling using the first MIP image comprises: performing one or more preliminary analysis steps to adjust the image intensity of the polymerase colonies in the first MIP image; and performing base calling for the polymerase colonies based on the adjusted image intensity; wherein the one or more preliminary analysis steps include: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; lag and lead correction; image registration; quality score estimation; or a combination thereof.
138. The computer-implemented system according to any one of the preceding claims, wherein the method further comprises: performing image registration of the plurality of flow cell images, the plurality of processed images, the plurality of filtered images, the first MIP image, or a combination thereof.
139. The computer-implemented system according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images comprises: registering the first MIP image to a template image.
140. The computer-implemented system according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images comprises: registering the plurality of flow cell images, the plurality of processed images, the plurality of filtered images, the first MIP image, or a combination thereof to a template image.
141. The computer-implemented system according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images comprises: registering the polymerase colonies in the first MIP image to template polymerase colonies in the template image.
142. The computer-implemented system according to any one of the preceding claims, wherein the method further comprises: obtaining, by the processor, a second MIP image based on the plurality of flow cell images; and performing image registration of the plurality of flow cell images, the plurality of processed images, the plurality of filtered images, or a combination thereof based on the second MIP image.
143. The computer-implemented system according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images, the plurality of processed images, the plurality of filtered images, or a combination thereof comprises: Registering the second MIP image to a template image.
144. The computer-implemented system according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images based on the second MIP image comprises: Registering the polymerase colonies in the second MIP image to the template polymerase colonies in the template image.
145. The computer-implemented system according to any one of the preceding claims, wherein performing image registration of the plurality of flow cell images based on the first MIP image or the second MIP image comprises: Generating one or more template images in a reference coordinate system by registering the polymerase colonies in one or more reference cycles to the one or more template images using the coordinates of the polymerase colonies; Determining, by the processor, a plurality of transforms of the first MIP image or the second MIP image based on the one or more template images, the plurality of transforms corresponding to sub-blocks of the first MIP or the second MIP and configured to register the sub-blocks to the one or more template images; And Registering the sub-blocks to the one or more template images using the plurality of transforms.
146. The computer-implemented system according to any one of the preceding claims, wherein the plurality of transforms comprises one or more affine transforms.
147. The computer-implemented system according to any one of the preceding claims, wherein each of the plurality of transforms comprises an affine transform.
148. The computer-implemented system according to any one of the preceding claims, wherein performing base calling using the first MIP image comprises: Performing base calling based on the image intensity of the polymerase colonies from the first MIP image and the position information of the polymerase colonies of the second MIP image from the plurality of flow cell images.
149. The computer-implemented system according to any one of the preceding claims, wherein the method further comprises: Performing image registration of the polymerase colonies of the plurality of flow cell images based on fiducial markers.
150. The computer-implemented system according to any one of the preceding claims, wherein the fiducial markers are located on the flow cell.
151. The computer-implemented system according to any one of the preceding claims, wherein the fiducial markers are outside the flow cell.
152. The computer-implemented system according to any one of the preceding claims, wherein the plurality of flow cell images are acquired at 2, 3, 4, 5, 6, 7, 8, 9, or 10 different positions along the axial axis.
153. The computer-implemented system according to any one of the preceding claims, wherein two adjacent positions along the axial axis are separated by approximately 1um, 2um, 3um, 4um, 5um, 6um, 7um, 8um, 9um, 10um, 11um or 12um.
154. The computer-implemented system according to any one of the preceding claims, wherein the plurality of flow cell images are acquired from 1, 2, 3, 4, 5 or 6 channels.
155. The computer-implemented system according to any one of the preceding claims, wherein the processor comprises: one or more processing units; one or more integrated circuits; or a combination thereof.
156. The computer-implemented system according to any one of the preceding claims, wherein the processor comprises: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); one or more neural processing units (NPUs); or a combination thereof.
157. The computer-implemented system according to any one of the preceding claims, wherein the operation further comprises: transferring the base identification to a processing unit by the processor.
158. The computer-implemented system according to any one of the preceding claims, wherein the processing unit is a central processing unit (CPU).
159. The computer-implemented system according to any one of the preceding claims, wherein the processing unit is configured to register the base identification to one or more images.
160. The computer-implemented method according to any one of the preceding claims, wherein the operation further comprises: providing the sample having a plurality of tandem molecules immobilized on a support, wherein each tandem molecule corresponds to a target RNA of a cell sample.
161. The computer-implemented system according to any one of the preceding claims, wherein obtaining the plurality of flow cell images of the sample comprises: generating the plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on the sample immobilized on the support, wherein the plurality of flow cell images are generated from two or more color channels at two or more different z-levels along the axial axis.
162. The computer-implemented system according to any one of the preceding claims, wherein obtaining the plurality of flow cell images of the sample comprises: generating the plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on a plurality of tandem molecules of the sample immobilized on the support.
163. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a polymerase community immobilized thereon.
164. The computer-implemented method according to any one of the preceding claims, wherein the polymerase community corresponds to the plurality of nucleotide template molecules or tandem molecules.
165. The computer-implemented method according to any one of the preceding claims, wherein the operation further comprises: The flow cell image is generated by the sequencing system by performing one or more sequencing reaction cycles on the plurality of tandem molecules immobilized on the support.
166. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of tandem molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleotide bases A, G, C, and T / U.
167. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: contacting the plurality of tandem molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of affinity bodies.
168. The computer-implemented method according to any one of the preceding claims, wherein an individual affinity body in the mixture comprises a core attached with a plurality of nucleotide arms, and each arm of the individual affinity body comprises the same type of nucleotide base.
169. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: imaging, by an optical system, an optical color signal emitted from a nucleotide reagent bound to the plurality of tandem molecules in each of the one or more cycles.
170. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: acquiring, by an optical system, the flow cell image in each of the one or more cycles, the flow cell image comprising an optical color signal emitted from a nucleotide reagent bound to the plurality of tandem molecules.
171. The computer-implemented method according to any one of the preceding claims, wherein the flow cell image comprises optical signals emitted from nucleotide reagents with an unbalanced diversity of nucleotide bases A, G, C, and T / U bound to the plurality of templates or tandem molecules immobilized on the support in one or more cycles.
172. The computer-implemented method according to any one of the preceding claims, wherein a plurality of polymerase populations comprise an unbalanced diversity of nucleotide bases A, G, C, and T / U, and wherein the unbalanced diversity comprises a percentage of (1) the number of one or more types of nucleotide bases in a region of the flow cell image to (2) the total number of nucleotide bases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the region. The computer-implemented system according to any one of the preceding claims, wherein the operation further comprises: Providing the sample containing a plurality of RNAs, the plurality of RNAs comprising at least a first target RNA molecule and a second target RNA molecule. The computer-implemented system according to any one of the preceding claims, wherein the operation further comprises: Generating a plurality of cDNA molecules inside the sample, the plurality of cDNA molecules comprising at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules comprising a second target cDNA molecule corresponding to the second target RNA molecule. The computer-implemented system according to any one of the preceding claims, wherein the operation further comprises: Contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes, the plurality of target-specific padlock probes comprising at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes. The computer-implemented system according to any one of the preceding claims, wherein the operation further comprises: Closing the nicks or gaps in the at least first and second circularized target-specific padlock probes by performing an enzymatic reaction, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe inside the sample.
177. A computer-implemented system according to any one of the preceding claims, wherein the operation further comprises: Using the first and second covalently closed circular padlock probes as template molecules, performing a rolling circle amplification reaction inside the sample, thereby generating a plurality of tandem molecules, the plurality of tandem molecules including at least a first tandem molecule corresponding to a first target RNA molecule, and the plurality of tandem molecules including at least a second tandem molecule corresponding to a second target RNA molecule. The computer-implemented system according to any one of the preceding claims, wherein the operation further comprises: Sequencing the plurality of tandem molecules inside the sample, the sequencing including sequencing the first tandem molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second tandem molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of second sequencing read products.
179. The computer-implemented system according to any one of the preceding claims, wherein sequencing the plurality of concatenated molecules within the sample comprises: Contacting the plurality of tandem molecules inside the sample with (i) the plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridization of the plurality of universal sequencing primers to their corresponding universal sequencing primer binding sites on the tandem.
180. The computer-implemented system according to any one of the preceding claims, wherein the nucleotide reagents comprise one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs.
181. The computer-implemented system according to any one of the preceding claims, wherein the operation further comprises: Removing the plurality of first sequencing read products from the first tandem molecule and retaining the first tandem molecule in the sample, and removing the plurality of second sequencing read products from the second tandem molecule and retaining the second tandem molecule in the sample.
182. A method for staining a cell or tissue, comprising: Selecting one or more primary antibodies, each of the one or more primary antibodies specifically binding to a corresponding protein; Selecting one or more secondary antibodies that bind to the one or more primary antibodies; Labeling the one or more secondary antibodies with a fluorescent label; And Generating one or more images of the corresponding protein, the one or more images containing a fluorescent signal generated from the fluorescent label.
183. The computer-implemented method according to any one of the preceding claims, wherein the corresponding protein is a transmembrane protein of one or more cells.
184. The computer-implemented method according to any one of the preceding claims, wherein the corresponding protein is a transmembrane protein that is not present in the cytosol or nucleus of one or more cells.
185. The computer-implemented method according to any one of the preceding claims, wherein the one or more images are microscopic images.
186. The computer-implemented method according to any one of the preceding claims, wherein the one or more images are fluorescence images.
187. The computer-implemented method according to any one of the preceding claims, wherein labeling the one or more secondary antibodies with the fluorescent label comprises cross-linking the secondary antibody with a fluorophore using a scaffold element.
188. A computer-implemented method for base calling in sequencing data analysis, comprising: obtaining, by a processor, a first plurality of flow cell images of a sample from a plurality of z-levels along an axial axis, wherein each flow cell image of the first plurality of flow cell images is acquired at a corresponding z-level along the axial axis; generating, by the processor, a first plurality of processed images of the first plurality of flow cell images; filtering, by the processor, the first plurality of flow cell images based on the first plurality of processed images, thereby generating a first plurality of filtered images; obtaining, by the processor, a 3D polymerase colony map of the sample; extracting, by the processor, image intensities of polymerase colonies from one of the following based on the 3D polymerase colony map: a second plurality of flow cell images; a second plurality of processed images; a second plurality of filtered images; or a combination thereof; and performing, by the processor, base calling based on the extracted image intensities of the polymerase colonies.
189. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Performing image registration of: the first plurality of flow cell images; the first plurality of processed images; the first plurality of filtered images; or a combination thereof.
190. The computer-implemented method according to any one of the preceding claims, wherein performing image registration comprises registering: the first plurality of flow cell images; the first plurality of processed images; the first plurality of filtered images; or a combination thereof to one or more template images.
191. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images; the first plurality of processed images; Registering the first plurality of filtered images; or a combination thereof to one or more template images comprises: generating, by the processor, the one or more template images in a reference coordinate system.
192. The computer-implemented method according to any one of the preceding claims, wherein performing image registration comprises: registering, by the processor, the polymerase colonies of the first plurality of flow cell images; the first plurality of processed images; the first plurality of filtered images; or a combination thereof to the template polymerase colonies in the one or more template images.
193. The computer-implemented method according to any one of the preceding claims, wherein generating the one or more template images in the reference coordinate system comprises: registering the polymerase colonies in one or more reference cycles to the one or more template images using the coordinates of the polymerase colonies.
194. The computer-implemented method according to any one of the preceding claims, wherein the coordinates of the polymerase colonies comprise 2D coordinates of the polymerase colonies.
195. The computer-implemented method according to any one of the preceding claims, wherein the coordinates of the polymerase colonies comprise the z-level of the polymerase colonies.
196. The computer-implemented method according to any one of the preceding claims, wherein performing image registration comprises: determining, by the processor, a plurality of transformations based on the one or more template images, each transformation in the plurality of transformations corresponding to a corresponding sub-tile of the first plurality of flow cell images, the first plurality of processed images, or the first plurality of filtered images and configured to register the sub-tile to the one or more template images; and registering the sub-tiles to the one or more template images using the plurality of transformations.
197. The computer-implemented method according to any one of the preceding claims, wherein each transformation in the plurality of transformations corresponds to a corresponding image among the following: the first plurality of flow cell images; the first plurality of processed images; or the first plurality of filtered images.
198. The computer-implemented method according to any one of the preceding claims, wherein the plurality of transformations comprises one or more affine transformations.
199. The computer-implemented method according to any one of the preceding claims, wherein each transformation in the plurality of transformations comprises an affine transformation.
200. The computer-implemented method according to any one of the preceding claims, wherein performing base identification based on the extracted image intensity of the polymerase population comprises: performing base identification based on the image intensity of the polymerase population extracted from the first plurality of filtered images.
201. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: performing image registration of the polymerase population of the first plurality of filtered images based on fiducial markers.
202. The computer-implemented method according to any one of the preceding claims, wherein the fiducial markers are located on the flow cell.
203. The computer-implemented method according to any one of the preceding claims, wherein the fiducial markers are outside the flow cell.
204. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired at 2, 3, 4, 5, 6, 7, 8, 9, or 10 different positions along the axial axis.
205. The computer-implemented method according to any one of the preceding claims, which further comprises: generating the 3D polymerase population map based on the first plurality of filtered images.
206. The computer-implemented method according to any one of the preceding claims, wherein generating the 3D polymerase population map based on the first plurality of filtered images comprises: generating the 3D polymerase population map based on the one or more template images.
207. The computer-implemented method according to any one of the preceding claims, wherein the one or more template images are 2D.
208. The computer-implemented method according to any one of the preceding claims, wherein each template image in the one or more template images corresponds to a corresponding flow cell image of the first plurality of flow cell images at a corresponding position along the axial axis.
209. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images, the first plurality of processed images, and the first plurality of filtered images are from the one or more reference cycles and different channels.
210. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of flow cell images, the second plurality of processed images, and the second plurality of filtered images are from the one or more reference cycles and the different channels.
211. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of flow cell images, the second plurality of processed images, and the second plurality of filtered images are from one or more cycles different from the one or more reference cycles and from the different channels.
212. The computer-implemented method according to any one of the preceding claims, wherein the first and second pluralities of flow cell images are the same, the first and second pluralities of processed images are the same, and the first and second pluralities of filtered images are the same.
213. The computer-implemented method according to any one of the preceding claims, wherein performing base identification based on the extracted image intensity of the polymerase population is within a cycle different from the one or more reference cycles.
214. The computer-implemented method according to any one of the preceding claims, wherein generating the 3D polymerase population map based on the one or more template images comprises: extracting the polymerase population in the one or more template images; and removing duplicate polymerase populations from the extracted polymerase population.
215. The computer-implemented method according to any one of the preceding claims, wherein generating the 3D polymerase population map based on the one or more template images comprises: combining the one or more template images into a candidate 3D polymerase population map; and removing duplicate polymerase populations from the candidate 3D polymerase population map.
216. The computer-implemented method according to any one of the preceding claims, wherein removing the duplicate polymerase populations comprises: performing preliminary base identification based on the one or more template images; and repeating removing the duplicate polymerase populations from the candidate 3D polymerase population map until a stop criterion is met, comprising: identifying candidate polymerase populations having the same preliminary base identification; determining the 3D distance between two polymerase populations in the candidate polymerase populations; and in response to determining that the 3D distance between the two polymerase populations meets a predetermined distance threshold: determining the image intensities for the two polymerase populations from the first plurality of filtered images; removing the polymerase population having the smaller image intensity from the two polymerase populations.
217. The computer-implemented method according to any one of the preceding claims, wherein the first or second plurality of flow cell images are acquired at one or more sequencing cycles.
218. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired at one or more reference cycles.
219. The computer-implemented method according to any one of the preceding claims, wherein the second plurality of flow cell images are acquired at one or more sequencing cycles different from the one or more reference cycles.
220. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels are spaced from each other by 0.1 um to 5 um.
221. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels cover at least some of the thickness of the sample along the axial axis.
222. The computer-implemented method according to any one of the preceding claims, wherein the plurality of z-levels cover the entire thickness of the sample along the axial axis.
223. The computer-implemented method according to any one of the preceding claims, wherein each flow cell image in the first or second plurality of flow cell images has an image thickness of 0.1 um to 6 um.
224. The computer-implemented method according to any one of the preceding claims, wherein removing the duplicate polymerase colonies comprises: performing preliminary base identification based on the one or more template images; and repeating removing the duplicate polymerase colonies from the 3D candidate polymerase colony map until a stop criterion is met, including: identifying candidate polymerase colonies having the same preliminary base identification; determining the 3D distance between two polymerase colonies in the candidate polymerase colonies; and in response to determining that the 3D distance between the two polymerase colonies is within a predetermined distance threshold: determining the image intensity for the two polymerase colonies from the first plurality of filtered images; removing from the candidate 3D polymerase colony map the polymerase colony having the smaller image intensity of the two polymerase colonies; and in response to determining that the 3D distance between two polymerase colonies fails to meet the predetermined distance threshold: retaining the two polymerase colonies in the candidate 3D polymerase colony map.
225. The computer-implemented method according to any one of the preceding claims, wherein the predetermined distance threshold is based on the depth of field of the optical system, the distance between two adjacent flow cell images in the axial direction, or a combination thereof.
226. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired by an NGS sequencing system.
227. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired at a cycle different from the reference cycle.
228. The computer-implemented method according to any one of the preceding claims, wherein the sample is an in-situ sample located on a flow cell.
229. The computer-implemented method according to any one of the preceding claims, wherein the in-situ sample comprises one or more cells or tissues.
230. The computer-implemented method according to any one of the preceding claims, wherein each flow cell image of the first plurality of flow cell images includes a field of view orthogonal to the axial axis.
231. The computer-implemented method according to any one of the preceding claims, wherein the field of view of each flow cell image of the first plurality of flow cell images is the same.
232. The computer-implemented method according to any one of the preceding claims, wherein the field of view of each flow cell image of the first plurality of flow cell images covers at least a portion of a tile of the flow cell.
233. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images have the same image resolution.
234. The computer-implemented method according to any one of the preceding claims, wherein the first or second plurality of flow cell images are 2D.
235. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of flow cell images of the sample from a plurality of z-levels along the axial axis includes: Obtaining the first plurality of flow cell images of the sample from a plurality of z-levels along the axial axis from one or more color channels.
236. The computer-implemented method according to any one of the preceding claims, wherein the one or more color channels include 1, 2, 3, or 4 different color channels.
237. The computer-implemented method according to any one of the preceding claims, wherein the axial axis extends from an objective lens to a sample located on a flow cell positioned on a sequencing system.
238. The computer-implemented method according to any one of the preceding claims, wherein the axial axis is orthogonal to an image plane, and wherein the field of view is within the image plane.
239. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images includes: Selecting a kernel; And Generating the first plurality of processed images by performing an opening operation on the first plurality of flow cell images using the selected kernel.
240. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images includes: Selecting a kernel; And Generating the first plurality of processed images by convolving the first plurality of flow cell images with the selected kernel.
241. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images further includes: Selecting a first kernel and a second kernel; Generating a first blurred image by convolving the first plurality of flow cell images with the first kernel; And Generating a second blurred image by convolving the first plurality of flow cell images with the second kernel.
242. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images includes: Scaling the first plurality of processed images.
243. The computer-implemented method according to any one of the preceding claims, wherein obtaining the first plurality of processed images comprises: Scaling the first blurred image, the second blurred image, or both.
244. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images based on the first plurality of processed images comprises: Subtracting the second blurred image from the first blurred image to generate the first plurality of filtered images.
245. The computer-implemented method according to any one of the preceding claims, wherein the kernel is 2×2, 3×3, 4×4, 5×5, 6×6 pixels.
246. The computer-implemented method according to any one of the preceding claims, wherein the kernel is a circular kernel.
247. The computer-implemented method according to any one of the preceding claims, wherein the kernel is a Gaussian kernel.
248. The computer-implemented method according to any one of the preceding claims, wherein the first kernel and the second kernel are different Gaussian kernels.
249. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images based on the first plurality of processed images comprises: Subtracting each processed image in the first plurality of processed images from the corresponding flow cell image in the first plurality of flow cell images to generate the first plurality of filtered images.
250. The computer-implemented method according to any one of the preceding claims, wherein filtering the first plurality of flow cell images based on the first plurality of processed images further comprises: Adding a predetermined offset to the subtracted images to generate the first plurality of filtered images.
251. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Registering the first plurality of flow cell images, the first plurality of processed images, the first plurality of filtered images, or a combination thereof to one or more images of the sample.
252. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise staining of the following: membrane, cell nucleus, or a combination thereof.
253. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise staining of one or more membrane proteins.
254. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise staining of lipids.
255. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise a fluorescence signal from a cell membrane.
256. The computer-implemented method according to any one of the preceding claims, wherein the one or more images comprise segmentation of the following: cells, membrane, cell nucleus, or a combination thereof.
257. The computer-implemented method according to any one of the preceding claims, wherein performing base identification based on the extracted image intensity of the polymerase community comprises: performing one or more preliminary analysis steps to adjust the image intensity of the polymerase community in the following: the first plurality of flow cell images; the first plurality of processed images; the first plurality of filtered images; or a combination thereof; and performing base identification for the polymerase community based on the adjusted image intensity, wherein the one or more preliminary analysis steps comprise: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; lag and overshoot correction; image registration; quality score estimation; or a combination thereof.
258. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired at the same tile or sub-tile of the flow cell.
259. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired at one or more reference cycles.
260. The computer-implemented method according to any one of the preceding claims, wherein two adjacent positions along the axial axis are separated by approximately 1um, 2um, 3um, 4um, 5um, 6um, 7um, 8um, 9um, 10um, 11um or 12um.
261. The computer-implemented method according to any one of the preceding claims, wherein the first plurality of flow cell images are acquired from 1, 2, 3, 4, 5 or 6 channels.
262. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more processing units; one or more integrated circuits; or a combination thereof.
263. The computer-implemented method according to any one of the preceding claims, wherein the processor comprises: one or more central processing units (CPUs); one or more field programmable gate arrays (FPGAs); one or more neural processing units (NPUs); or a combination thereof.
264. The computer-implemented method according to any one of the preceding claims, further comprising: transmitting the base identification by the processor to a processing unit.
265. The computer-implemented method according to any one of the preceding claims, wherein the processing unit is a central processing unit (CPU).
266. The computer-implemented method according to any one of the preceding claims, wherein the processing unit is configured to register the base identification to one or more images.
267. The computer-implemented method according to any one of the preceding claims, wherein the 3D polymerase community map comprises a list of 3D coordinates, each entry of the list of 3D coordinates corresponding to the 3D position of the polymerase community of the sample.
268. The computer-implemented method according to any one of the preceding claims, further comprising: Providing the sample with multiple tandem molecules immobilized on a support, wherein each tandem molecule corresponds to a target RNA of a cell sample.
269. The computer-implemented method according to any one of the preceding claims, wherein obtaining the plurality of flow cell images of the sample comprises: Generating the plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on the sample immobilized on the support, wherein the plurality of flow cell images are generated from two or more color channels at two or more different z-levels along an axial axis.
270. The computer-implemented method according to any one of the preceding claims, wherein obtaining the plurality of flow cell images of the sample comprises: Generating the plurality of flow cell images by the sequencing system by performing one or more sequencing reaction cycles on the multiple tandem molecules of the sample immobilized on the support.
271. The computer-implemented method according to any one of the preceding claims, wherein the sample comprises a polymerase community immobilized thereon.
272. The computer-implemented method according to any one of the preceding claims, wherein the polymerase community corresponds to the plurality of nucleotide template molecules or tandem molecules.
273. The computer-implemented method according to any one of the preceding claims, further comprising: Generating a flow cell image by the sequencing system by performing one or more sequencing reaction cycles on the multiple tandem molecules immobilized on the support.
274. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: Contacting the multiple tandem molecules with a plurality of nucleotide reagents comprising a mixture of different types of nucleobases A, G, C, and T / U.
275. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: Contacting the multiple tandem molecules with a mixture of a plurality of sequencing primers, a plurality of polymerases, and different types of affinity bodies.
276. The computer-implemented method according to any one of the preceding claims, wherein each individual affinity body in the mixture comprises a core attached with a plurality of nucleotide arms, and each arm of the individual affinity body comprises the same type of nucleobase.
277. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, imaging an optical color signal emitted from a nucleotide reagent bound to the multiple tandem molecules by an optical system.
278. The computer-implemented method according to any one of the preceding claims, wherein performing the one or more sequencing reaction cycles comprises: In each of the one or more cycles, acquiring the flow cell image by an optical system, the flow cell image comprising an optical color signal emitted from a nucleotide reagent bound to the multiple tandem molecules.
279. The computer-implemented method according to any one of the preceding claims, wherein the flow cell image comprises optical signals emitted by nucleotide reagents that bind to nucleobases A, G, C, and T / U of the plurality of templates or concatemer molecules immobilized on the support and exhibit unbalanced diversity in one or more cycles.
280. The computer-implemented method according to any one of the preceding claims, wherein the plurality of polymerase colonies comprise unbalanced diversity of nucleobases A, G, C, and T / U, and wherein the unbalanced diversity comprises the percentage of (1) the number of one or more types of nucleobases in the region of the flow cell image to (2) the total number of nucleobases, and wherein in one or more cycles, the percentage is less than 20%, 15%, 10%, or 5% in the region. The computer-implemented method according to any one of the preceding claims, further comprising: Providing the sample containing a plurality of RNAs, the plurality of RNAs comprising at least a first target RNA molecule and a second target RNA molecule. The computer-implemented method according to any one of the preceding claims, further comprising: Generating a plurality of cDNA molecules within the sample, the plurality of cDNA molecules comprising at least a first target cDNA molecule corresponding to the first target RNA molecule, and the plurality of cDNA molecules comprising a second target cDNA molecule corresponding to the second target RNA molecule. The computer-implemented method according to any one of the preceding claims, further comprising: Contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes, the plurality of target-specific padlock probes comprising at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes. The computer-implemented method according to any one of the preceding claims, further comprising: Closing the nicks or gaps in the at least first and second circularized target-specific padlock probes by an enzymatic reaction, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe within the sample. The computer-implemented method according to any one of the preceding claims, further comprising: Using the first and second covalently closed circular padlock probes as template molecules, performing a rolling circle amplification reaction within the sample, thereby generating a plurality of concatemer molecules, the plurality of concatemer molecules comprising at least a first concatemer molecule corresponding to the first target RNA molecule, and the plurality of concatemer molecules comprising at least a second concatemer molecule corresponding to the second target RNA molecule. The computer-implemented method according to any one of the preceding claims, further comprising: Sequencing the plurality of concatemer molecules within the sample, the sequencing comprising sequencing the first concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second concatemer molecule by performing no more than 2 to 150 sequencing cycles to generate a plurality of second sequencing read products.
287. The computer-implemented method according to any one of the preceding claims, wherein sequencing the plurality of tandem molecules within the sample comprises: Contacting the plurality of concatemer molecules within the sample with (i) the plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridization of the plurality of universal sequencing primers to their corresponding universal sequencing primer binding sites on the concatemer.
288. The computer-implemented method according to any one of the preceding claims, wherein the nucleotide reagent comprises one or more of the following: multivalent molecules, nucleotides, and nucleotide analogs. The computer-implemented method according to any one of the preceding claims, further comprising: Remove the plurality of first sequencing read products from the first concatenated molecule and retain the first concatenated molecule in the sample, and remove the plurality of second sequencing read products from the second concatenated molecule and retain the second concatenated molecule in the sample.
290. A computer-implemented system for base calling in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations including: Obtaining, by the processor, a first plurality of flow cell images of a sample, wherein each flow cell image of the first plurality of flow cell images is acquired at a corresponding position along an axial axis; Generating, by the processor, a first plurality of processed images corresponding to the first plurality of flow cell images; Filtering, by the processor, the first plurality of flow cell images based on the first plurality of processed images to generate a first plurality of filtered images; Obtaining, by the processor, a 3D polymerase colony map; Extracting, by the processor, image intensities of polymerase colonies from one of the following based on the 3D polymerase colony map: A second plurality of flow cell images; A second plurality of processed images; A second plurality of filtered images; or A combination thereof; and Performing, by the processor, base calling based on the extracted image intensities of the polymerase colonies.
291. A computer-implemented system for base calling in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations according to any one of the preceding claims.
292. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations for base calling in sequencing data analysis, the operations including: Obtaining, by the processor, a first plurality of flow cell images of a sample, wherein each flow cell image of the first plurality of flow cell images is acquired at a corresponding position along an axial axis; Generating, by the processor, a first plurality of processed images corresponding to the first plurality of flow cell images; Filtering, by the processor, the first plurality of flow cell images based on the first plurality of processed images to generate a first plurality of filtered images; Obtaining, by the processor, a 3D polymerase colony map; Extracting, by the processor, image intensities of polymerase colonies from one of the following based on the 3D polymerase colony map: A second plurality of flow cell images; A second plurality of processed images; A second plurality of filtered images; Or A combination thereof; And Performing, by the processor, base calling based on the extracted image intensities of the polymerase colonies. One or more non - transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations, the operations including any one of the foregoing claims. A computer - implemented method for base calling in sequencing data analysis, comprising: Obtaining, by a processor, a first plurality of flow cell images of a sample, wherein each of the plurality of flow cell images is acquired at a corresponding position along an axial axis; Filtering, by the processor, the plurality of flow cell images to generate a plurality of filtered images; Obtaining, by the processor, a 3D polymerase colony map based on the plurality of filtered images; Extracting, by the processor, image intensities of polymerase colonies from one of the following based on the 3D polymerase colony map: The plurality of flow cell images; The plurality of processed images; The plurality of filtered images; or A combination thereof; And Performing base calling by the processor based on the extracted image intensities of the polymerase colonies. The computer - implemented method according to any one of the foregoing claims, wherein filtering the plurality of flow cell images to generate the plurality of filtered images comprises: Performing 3D deconvolution of the plurality of flow cell images. A computer - implemented system for base calling in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices storing instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations, the operations including: Obtaining, by a processor, a first plurality of flow cell images of a sample, wherein each of the plurality of flow cell images is acquired at a corresponding position along an axial axis; Filtering, by the processor, the plurality of flow cell images to generate a plurality of filtered images; Obtaining, by the processor, a 3D polymerase colony map based on the plurality of filtered images; Extracting, by the processor, image intensities of polymerase colonies from one of the following based on the 3D polymerase colony map: The plurality of flow cell images; The plurality of processed images; The plurality of filtered images; or A combination thereof; and Performing base calling by the processor based on the extracted image intensities of the polymerase colonies. A computer - implemented system for base calling in sequencing data analysis, comprising: One or more hardware processors; One or more data storage devices storing instructions that can be executed by the one or more hardware processors to cause the one or more hardware processors to perform operations including any one of the foregoing claims. One or more non - transitory computer storage media encoded with instructions that can be executed by one or more hardware processors to perform operations for base calling in sequencing data analysis, the operations including: A first plurality of flow cell images of a sample are obtained by a processor, wherein each of the plurality of flow cell images is acquired at a corresponding position along an axial axis; The plurality of flow cell images are filtered by the processor to generate a plurality of filtered images; A 3D polymerase colony map is obtained by the processor based on the plurality of filtered images; The processor extracts image intensities of the polymerase colony from among the following based on the 3D polymerase colony map: The plurality of flow cell images; The plurality of processed images; The plurality of filtered images; or A combination thereof; And Base calling is performed by the processor based on the extracted image intensities of the polymerase colony.
299. One or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors to perform operations including any one of the foregoing claims.
300. A computer-implemented method according to any one of the foregoing claims, wherein the 3D polymerase colony map includes a list of 3D coordinates, each entry of the list of 3D coordinates corresponding to a 3D position of the polymerase colony of the sample.
Citation Information
Patent Citations
Method and system for sequencing nucleic acids
US10246744B2
Expansion microscopy
US10309879B2
Engineered polymerases for improved sequencing
US10731141B2
Improvement in drawers
US133138A
Primary analysis in next generation sequencing
US20230326064A1