Base calling in next generation sequencing analysis
Patent Information
- Application Number
- PCT/US2026/018398
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-10
- Filing Date
- 2026-03-09
- Publication Date
- 2026-09-17
Smart Images

Figure US2026018398_17092026_PF_FP_ABST
Abstract
Description
BASE CALLING IN NEXT GENERATION SEQUENCING ANALYSIS CROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No.63 / 769,481, filed Mar. 10, 2025, which is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] This disclosure relates generally to image processing and base calling in sequencing data analysis of two-dimensional (2D) sample(s) and three-dimensional (3D) sample(s).BACKGROUND
[0003] In next-generation sequencing (NGS) or NGS-like applications such as sequencing by synthesis, sequencing by binding, or sequencing by avidity, in order to identify the sequence of a target nucleic acid, a new strand is synthesized one nucleotide base at a time. During each sequencing cycle, one base attaches to any given strand. At the imaging step of each cycle, image(s) are recorded. A base-calling algorithm is applied to the image(s) to “read” the successive signals from each cluster or polony and convert the optical signals into an identification of the nucleotide base sequence added to each DNA fragment. Traditional sequencing data analysis relies on flow cell images from 4 different color channels. Images acquired from each color channel corresponds to emission signal from a different type of nucleotide base. When it comes to sequencing and analysis using less than 4 color channels, at least two emission signals with different colors or wavelengths may appear in a same image with chromatic aberration. There is a need for fast and accurate image processing and base calling to ensure reliable sequencing analysis for flow cell images with chromatic aberration.BRIEF SUMMARY
[0004] Provided herein are system, apparatus, method, and / or computer program product embodiments, and / or combinations and sub-combinations thereof which enable fast andaccurate flow cell image processing for reliable and accurate base calling of sample(s), such as in situ cells or tissue. The flow cell images can come from different sequencing cycles, different channels, and / or different z levels. The flow cell images can be obtained with chromatic aberration that may cause base calling errors if not corrected.
[0005] As a particular application of such, disclosed herein are embodiments of methods, systems, and media for sequencing analysis of flow cell images of 3D volumetric sample (s), e.g., in situ samples including cells or tissue, in which emission signal with at least two different emission wavelengths are acquired within a same flow cell image and result in chromatic aberration. The methods herein may correct or at least reduce such chromatic aberration in flow cell images and improve base calling accuracy based on such flow cell images. The chromatic aberration correction herein in flow cell images of 2D or 3D samples may be applied in sequencing analysis workflow with various image processing algorithms, for example, it may be applied after high-resolution flow cell images are generated based low resolution flow cell images acquired by the image sensor, and it may be applied after some preprocessing steps, such as polony map generation. The chromatic aberration correction methods herein may advantages reduce or eliminate chromatic aberration without using complicated modeling of spatial offset caused by chromatic aberration, also without computationally intensive training and / or prediction involving artificial intelligence models. The chromatic aberration methods herein advantageously use subsets of polonies and their preliminary intensities and / or locations for determining chromatic aberration for the entire flow cell images to simplify the computation but also ensures the quality of chromatic aberration correction. Additionally, the chromatic aberration correction methods handles the chromatic spatial shifts in different portions of full flow cell images based on the spatial dependency of chromatic aberration in flow cell images. The chromatic aberration correction methods and systems herein advantageously allow flow cell images to be acquired using optical elements that are not optimized for chromatic aberration reduction, thus, alleviate the complexity and cost in making sequencing optical designs. The chromatic aberration correction methods and systems herein are compatible with various sequencing and sequencing analysis methods so that it can be applied in various sequencing and sequencing analysis workflows for different DNA sequencing applications. In some embodiments, the methods herein may function to reverse the imaging process of an optical system and virtually improve the full width half maximum (FWHM) of the optical system. As such,the image processing methods disclosed herein may advantageously increase detectable density of polonies or clusters in 3D samples or traditional 2D samples. The methods herein may advantageously lessen the impact of color mixing of polonies that may be caused by neighboring polonies in 2D or 3D dimensions by computationally increasing the spatial resolution of the flow cell images.
[0006] In some embodiments, a neural network, e.g., a convolutional neural network, is used in generating high-resolution flow cell images of the biological sample(s) from the low-resolution images that has been acquired from the sequencing system, and subsequent primary analysis can be performed based on the high-resolution flow cell images instead of the low-resolution flow cell images. The methods of chromatic aberration correction may be applied before or after such high-resolution flow cell images are generated. In some embodiments, the neural network, e.g., a convolutional neural network, is used in processing of the high-resolution flow cell images of the samples to generate the base calls, after chromatic aberration correction is completed.
[0007] Embodiments of these aspects include corresponding computer systems, apparatus, and computer program product recorded on computer storage device(s), which, alone or in combination, configured to perform the operations of the methods. For a computer system configured or to be configured to perform operations, the computer system has installed on it software, firmware, hardware, or their combinations that in operation cause the computer system to perform the operations or actions. For a computer program product configured or to be configured to perform operations or actions, the computer program product includes instructions that, when executed by a hardware processor, cause the hardware processor to perform the operations or actions.
[0008] Further embodiments, features, and advantages of the present disclosure, as well as the structure and operation of the various embodiments of the present disclosure, are described in detail below with reference to the accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the disclosure and to enable a person skilled in the art(s) to make and use the embodiments.
[0010] FIG. 1 illustrates a block diagram of a system for performing sequencing, flow cell image processing including chromatic aberration correction, and / or primary analysis operations including base calling of flow cell images, according to some embodiments.
[0011] FIGS. 2 shows a schematic of an exemplary sequencing analysis method to generate aberration-corrected base calls using flow cell images acquired with three color sequencing, according to some embodiments.
[0012] FIG. 3 shows a schematic of an exemplary sequencing analysis method to generate aberration-corrected base calls using flow cell images acquired with three color sequencing, according to some embodiments.
[0013] FIG. 4 illustrates a block diagram of a computer system for performing image processing, sequencing analysis, and / or base calling operations, according to some embodiments.
[0014] FIG. 5 shows an exemplary support with multiple tiles for immobilizing sample(s) thereon for sequencing, including the 2D samples and / or 3D cellular sample(s), according to some aspects.
[0015] FIG. 6 is a schematic showing exemplary embodiments of padlock probes.
[0016] FIG. 7 is a schematic showing a workflow for generating inside a cell circularized padlock probes, comprising generating first and second cDNAs from first and second target RNA molecules (respectively), hybridizing first and second padlock probes to the first and second cDNA molecules (respectively) to generate first and second circularized padlock probes (respectively).
[0017] FIG. 8 is a schematic showing a rolling circle and sequencing workflow inside a cell, comprising generating first and second concatemers by conducting rolling circle amplification using first and second covalently closed circular molecules (respectively). The first and second concatemers are subjected to a sequencing workflow using universal sequencing primers, sequencing polymerases, and a plurality of nucleotide reagents.
[0018] FIG. 9 is a schematic showing an exemplary workflow for sequencing a concatemer that is generated inside the cell.
[0019] FIG. 10 is a schematic showing an exemplary workflow for sequencing a concatemer that is generated inside the cell.
[0020] FIG. 11 is a schematic showing an exemplary workflow for sequencing a concatemer that is generated inside the cell.
[0021] FIG. 12 is a schematic showing an exemplary workflow for sequencing a concatemer that is generated inside the cell.
[0022] FIG. 13 is a schematic showing a workflow for generating circularized padlock probes, comprising generating first and second cDNAs from first and second target RNA molecules (respectively), hybridizing first and second padlock probes to the first and second cDNA molecules (respectively) to generate first and second circularized padlock probes (respectively).
[0023] FIG. 14 is a schematic showing a rolling circle and sequencing workflow comprising generating first and second concatemers by conducting rolling circle amplification using first and second covalently closed circular molecules (respectively).
[0024] FIG. 15 is a schematic of an exemplary low binding support comprising a glass substrate and alternating layers of hydrophilic coatings which are covalently or non- covalently adhered to the glass, and which further comprises chemically-reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides).
[0025] FIG. 16 is a schematic of various exemplary configurations of multivalent molecules. Left (Class I): schematics of multivalent molecules having a “starburst” or “helter-skelter” configuration. Center (Class II): a schematic of a multivalent molecule having a dendrimer configuration. Right (Class III): a schematic of multiple multivalent molecules formed by reacting streptavidin with 4-arm or 8-arm PEG-NHS with biotin and dNTPs. Nucleotide units are designated ‘N’, biotin is designated ‘B’, and streptavidin is designated ‘ SA’ .
[0026] FIG. 17 is a schematic of an exemplary multivalent molecule comprising a generic core attached to a plurality of nucleotide-arms.
[0027] FIG. 18 is a schematic of an exemplary multivalent molecule comprising a dendrimer core attached to a plurality of nucleotide-arms.
[0028] FIG. 19 shows a schematic of an exemplary multivalent molecule comprising a core attached to a plurality of nucleotide-arms, where the nucleotide arms comprise biotin, spacer, linker and a nucleotide unit.
[0029] FIG. 20 is a schematic of an exemplary nucleotide-arm comprising a core attachment moiety, spacer, linker and nucleotide unit.
[0030] FIG. 21 shows the chemical structure of an exemplary spacer (top), and the chemical structures of various exemplary linkers, including an 11 -atom Linker, 16-atom Linker, 23-atom Linker and an N3 Linker (bottom).
[0031] FIG. 22 shows the chemical structures of various exemplary linkers, including Linkers 1-9.
[0032] FIG. 23 A shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.
[0033] FIG. 23B shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.
[0034] FIG. 23 C shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.
[0035] FIG. 23D shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.
[0036] FIG. 24 shows the chemical structure of an exemplary biotinylated nucleotide- arm.
[0037] FIG. 25 is a schematic of a guanine tetrad (e.g., G-tetrad).
[0038] FIG. 26 is a schematic of an exemplary intramolecular G-quadruplex structure.
[0039] FIG. 27 shows a flow chart of exemplary sequencing analysis methods of generating aberration-corrected base calls of the flow cell images of biological samples, according to some embodiments.
[0040] In the drawings, like reference numbers generally indicate identical or similar elements. Additionally, generally, the left-most digit(s) of a reference number identifies the drawing in which the reference number first appears.DETAILED DESCRIPTION
[0041] Provided herein are system, apparatus, method, and / or computer program product embodiments, and / or combinations and sub-combinations thereof which enables image processing of flow cell images, e.g., flow cell imaged obtained from in situ sample(s) or traditional 2D sample(s) in a sequencing run, to: generate flow cell images that are free of or with reduced chromatic aberration; and to perform base calling using such flow cell images. The techniques herein can be used while a sequence run is still in progress to improve efficiency of sequencing and sequencing analysis. The techniques herein can beused while a sequence run is still in progress to reduce memory and data storage space required during sequencing and sequencing analysis. The techniques herein can be used on flow cell images obtained from various imaging and / or sequencing techniques of volumetric 3D samples and / or traditional 2D samples. The techniques disclosed herein are useful for base calling in next generation sequencing (NGS), and NGS flow cell images will be used as the primary example herein for describing the application of these techniques. However, such image analysis techniques may also be useful in other applications where spot-detection and / or CCD imaging is used or flow cell images are acquired.
[0042] Existing NGS sequencing utilizes flow cell images of samples from 4 different color channels, and each flow cell image only contain emission signal from fluorescent label(s) of a corresponding nucleotide base in the sample(s) at a single wavelength, i.e., color, or within a single narrow wavelength range spanning only several nanometers, e.g., less than 2, 3, 4, 5 nms. Sequencing using less than 4 colors relies on flow cell images with at least two different emission signal wavelengths, i.e., colors, or narrow wavelength ranges spanning only several nanometers, within a same image. Chromatic aberration may occur within such flow cell images and may result in base calling errors. There is a need for generating aberration-corrected intensities for polonies or clusters so that they can be used for accurate and reliable base calling in sequencing methods with less than 4 colors.
[0043] Aberration correction in a single flow cell image may exhibit spatial dependency and vary across the field of view. Existing methods for aberration correction, such as calibration-based methods that depend on reference fiducials may not be effective where polonies overlaps with each other due to chromatic aberration. Further, existing chromatic aberration correction may assume uniform spatial distortion, and therefore may fail to accurately correct localized, position-dependent chromatic aberration across the field of view. Complicated modeling of spatial dependent chromatic aberration may be computationally expensive and time consuming. There is a need to eliminate or reduce chromatic aberration so that accurate and reliable base calling can be performed using the flow cell images.
[0044] The techniques disclosed herein may advantageously utilize the reconfigurable logic device, e.g., FPGAs, and other integrated circuits, e.g., Al chips or neural processing units (NPUs) for accelerating some or all of the operations disclosed herein.The utilization of the reconfigurable logic device, e.g., FPGAs, and other integrated circuits, e.g., Al chips or neural processing units (NPUs) on-board the sequencing system may advantageously reduce computational time, reduce energy consumption, improve sequencing analysis efficiency, reduce data storage space required, and reduce sequencing system cost in analysis of flow cell images when compared with sequencing analysis using existing sequencing systems. The techniques disclosed herein may advantageously utilize the reconfigurable logic device, e.g., FPGAs, and other integrated circuits, e.g., Al chips or neural processing units (NPUs), to perform one or more operations in the training and / or the prediction.Sequencing systems
[0045] FIG. 1 illustrates a block diagram of a computer-implemented system 100, according to one or more embodiments disclosed herein. The system 100 has a sequencing system 110 that includes a flow cell 112, a sequencer 114, an imager 116, data storage 122, and user interface 124. The sequencing system 110 may be connected to a cloud 130. The sequencing system 110 may include one or more of dedicated processors 118, a first reconfigurable logic device, e.g., Field-Programmable Gate Array(s) (FPGAs) 120, and a computer system 126.
[0046] In some embodiments, the flow cell 112 is configured to capture DNA fragments and form DNA sequences for base-calling on the flow cell. The flow cell 112 can include a support as disclosed herein. The support can be a solid support. The support can include a surface coating thereon as disclosed herein. The surface coating can be a polymer coating as disclosed herein.
[0047] A flow cell 112 can include multiple tiles or imaging areas thereon, and each tile may be separated into a grid of subtiles. Each subtile can include a plurality of clusters or polonies immobilized thereon. As a nonlimiting example, a flow cell can have 424 tiles, and each tile can be divided into a 6 x 9 grid, therefore 54 subtiles. The flow cell image as disclosed herein can be an image including signals of a plurality of clusters or polonies. The flow cell image can include one or more tiles of signals or one or more subtiles of signals. In some embodiments, a flow cell image can be an image that includes all the tiles and approximately all signals thereon. The flow cell image can be acquired from a channel during an imaging or sequencing cycle using the imager 116. In some embodiments, each tile may include millions of polonies or clusters. As a nonlimitingexample, a tile can include about 1 to 10 million of clusters or polonies. Each polony can be a collection of many copies of DNA fragments.
[0048] Depending on the sample(s) immobilized on the support (e.g., a flow cell), the flow cell images may be acquired using the imager 116 at single or multiple z levels along a z axis orthogonal to the image plane of the flow cell images. In particular, for three dimensional samples, e.g., cells, tissues, or other in situ samples, the flow cell images can include multiple z-levels (i.e., z levels) in order to cover the whole sample(s) in 3D. The z axis can extend from the objective lens of the imager 116 disclosed herein to the support, e.g., flow cell. The z axis can be orthogonal to the image plane of the flow cell images. Each z level of flow cell images may be separated from the adjacent z level(s) for a predetermined distance, for example, ranging from about 0.1 um to about 15 urns, or from 0.02 um to 10 urns. Each z level of flow cell images may be separated from the adjacent z level(s) for a distance ranging from 0.5 um to 10 urns, from 0.01 um to 5 urns, or from 0.1 um to 15 urns. At each z level, flow cell images can be acquired from one or more sequencing cycles and / or one or more channels. Each flow cell image may include in its field of view at least part of one or more tiles or subtiles of the flow cell. FIG. 5 shows a portion of a flow cell 2712 with multiple tiles 2710. The image plane is defined by the x and y axis. And the z direction (i.e., z axis) is orthogonal to the x-y plane. Although the flow cell images, samples, and the z axis are described in a Cartesian coordinate system as shown in FIG. 5, any other coordinate systems can be used to define spatial locations and relationships herein. Other coordinate systems can include but are not limited to the polar coordinate system, cylindrical, or spherical coordinate systems.
[0049] The sequencer 114 may be configured to flow a nucleotide mixture onto the flow cell 112, cleave blockers from the nucleotides in between flowing steps, and perform other steps for the formation of the DNA sequences on the flow cell 112. The nucleotides may have fluorescent elements attached that emit light or energy in a wavelength that indicates the type of nucleotide. Each type of fluorescent element may correspond to a particular nucleotide base (e.g., A, G, C, T). The fluorescent elements may emit light in visible wavelengths. In some embodiments, the sequencer 114 and the flow cell 112 may be configured to performing various sequencing methods disclosed herein, for example, sequencing-by-avidite.
[0050] For example, each nucleotide base may be assigned a color. Different types of nucleotides can have different colors. Adenine(A) may be red, cytosine(C) may be blue,guanine(G) may be green, and thymine(T) may be yellow, for example. The color or wavelength of the fluorescent element for each nucleotide may be selected so that the nucleotides are distinguishable from one another based on the wavelengths of light emitted by the fluorescent elements.
[0051] The imager 116 may be configured to capture images of the flow cell 112 after each flowing step. In some embodiment, the imager 116 includes a camera configured to capture digital images, such as a CMOS or a CCD camera. The camera may be configured to capture images at the wavelengths of the fluorescent elements bound to the nucleotides. The images acquired by the imager of the sample(s) immobilized on at least a portion of the flow cell can be called the flow cell images.
[0052] In some embodiments, the imager 116 can include one or more optical systems disclose herein. The optical system(s) can be configured to capture optical signals from the flow cell and generate corresponding flow cell images thereof. The flow cell images can then be used for base calling.
[0053] In an embodiment, the images of the flow cell may be captured in groups, where each image in the group is taken at a wavelength or in a spectrum that matches or includes only one of the fluorescent elements, e.g., a single color. In another embodiment, the images may be captured as single images that captures all of the wavelengths of the fluorescent labels of the sample(s).
[0054] In some embodiments, the imager 116 of the sequencing system may include a single optical system or module which may be used for capturing emission signals at all wavelengths of the fluorescent labels.
[0055] In some embodiments, the imager 116 may include one or more illumination sources, e.g., a first and a second illuminator. The one or more illumination sources may be configured to generation excitation light for the samples in one or more colors. In embodiments with three-color sequencing, the one or more illumination sources are configured to generate excitation light with 3 different wavelengths, i.e., colors. In some embodiments, the single optical system or module may include a single optical path from the illumination source(s) to the one or more sample(s), and then from the sample(s) to the single image sensor. In some embodiments, emission signal of different wavelengths may travel through the same optical path to arrive at the single image sensor. Exemplary embodiments of the imager 116 are disclosed in PCT application No. PCT / US2024 / 12802 in more details and are incorporated herein by reference in its entirety.
[0056] In some embodiments, the imager 116 lacks any emission filtering so that all the emission signal from the sample(s) may be acquired by the image sensor without blockage of any emission signals from reaching the image sensor. In some embodiments, the imager 116 lacks any emission filter that selectively allow some wavelengths to pass through while filtering / blocking some other wavelengths in at least one color channel of the imager.
[0057] In a particular embodiment with three color sequencing (e.g., excitation with three different colors), a flow cell image corresponding to excitation of an individual color channel is acquired to cover at least a portion of the flow cell, e.g., a corresponding FOV. Three flow cell images are acquired in total corresponding to the three color channels used for excitation, e.g., blue, red, and green, of the same corresponding FOV per cycle per z-level. For at least one color channel, e.g., the green channel, with excitation by the green illumination, the sample(s) may be excited to emit signals in red and green. The corresponding flow cell image may include emission signal of such two different wavelengths, e.g., two different colors, red and green, due to the lack of emission filter before detecting the emission signals. The emission signal in either red or green may look like similar bright spots since the flow cell images are in grayscale. Chromatic aberration may occur within such flow cell images with emission signal of more than one wavelength, e.g., more than one color. With chromatic aberration, polonies or clusters with emission signal in red may experience a different spatial shift (in 2D or even in 3D) from the spatial shift of polonies or clusters with emission signal in green. Such difference may vary spatially across the flow cell image. In other words, the difference in spatial shift between two colored emission signal may vary depending on the position of the pixel(s) within the flow cell image. As a nonlimiting example, two polonies of different colors close to the center of the flow cell image may experience less shift than two other polonies of different colors at the edge of the flow cell image. The spatial dependency of the chromatic aberration may be linear or nonlinear. Such chromatic aberration may cause error in identifying image intensities for polonies and result in base calling errors. As an example, chromatic aberration may cause emission signal of a first polony to shift and overlap with the signal of a neighboring polony, resulting in error in base calling of the neighboring polony. Further, base calling of the first polony may also be adversely affected due to the shift of the emission signal. Thus, there is a need to correct chromatic aberration in order to perform accurate and reliable base calling.
[0058] The resolution of the imager 116 can control the level of detail in the flow cell images, including pixel size. In existing systems, this resolution is very important, as it controls the accuracy with which a spot-finding algorithm identifies the polony or cluster centers. In some embodiments, the image resolution of flow cell images disclosed herein can be about 10 nanometers (nms) to a couple of hundreds of nms or greater. In some embodiments, the image resolution of flow cell images can in a range from 0.1 nm to 1000 nms. In some embodiments, the image resolution of flow cell images can be in a range from 1 nm to 500 nms. In some embodiments, the image resolution of flow cell images can in a range from 5 nm to 300 nms. One way to increase the accuracy of polony or cluster finding is to improve the resolution of the imager 116, or improve the processing performed on images taken by imager 116. Detecting polony or cluster centers in pixels other than those detected by a spot-finding algorithm can be performed. These methods can allow for improved accuracy in detection of polony or cluster centers without increasing the resolution of the imager 116. The resolution of the imager 116 may even be better than existing systems with comparable performance, which may reduce the cost of the sequencing system 110.
[0059] The image quality of the flow cell images can control the base calling quality. One way to increase the accuracy of base calling is to improve the imager 116, or improve the processing performed on images taken by imager 116 to result in a better image quality. High resolution versions of the flow cell images (2x, 4x, or more than existing flow cell image resolution as acquired by the imager, in a reference coordinate system) may be generated so that the detectable polony or cluster density can be improved with reduced or eliminated interferences from neighboring polonies, cellular background signal, color mixing, and / or other noises in the flow cell images. As a result, 3D base calling can be more accurate when compared with existing methods without using such high resolution flow cell images. The high resolution flow cell images can be available for making base calling in the current sequencing cycle in the sequencing workflow, e.g., after chromatic aberration correction using the methods disclosed herein. Further, some or all of the operations disclosed herein can be advantageously performed by the first reconfigurable logic device, e.g., FPGA(s) or the integrated circuit, e.g., an application specific integrated circuit (ASIC) chip, neural processing unit (NPU), or artificial intelligence (Al) chip and data can be communicated between the CPU(s) and the first reconfigurablelogic device or integrated circuit to reduce the total operational time from methods operating using only the CPUs.
[0060] The sequencing system 110 may be configured to perform operations or actions for image processing of the flow cell images across different cycles and / or channels. The operations or actions disclosed herein may be performed by the dedicated processors 118, the reconfigurable logic device(s) and / or integrated circuit(s) 120, the computing system 126, or a combination thereof. One or more operations or actions in the methods, e.g., 2800, disclosed herein may be performed by the dedicated processors 118, the reconfigurable logic device(s) and / or integrated circuit(s) 120, the computing system 126, or a combination thereof. In some embodiments, which operations or actions are to be performed by the dedicated processors 118, the reconfigurable logic device(s) and / or integrated circuit(s) 120, the computing system 126, or their combinations can be determined based on one or more of: a computation time for the specific operation(s), the complexity of computation in the specific operation(s), the need for data transmission between the hardware devices, the power required for the specific operation(s), or their combinations. Image processing operations or actions of the flow cell images can be performed after the corresponding flow cell images are acquired but before base calling of the flow cell images is performed.
[0061] In some embodiments, the data storage 122 is used to store information used in the methods herein. This information may include the flow cell images themselves or information and / or images derived from the flow images captured by the imager 116. The DNA sequences determined from the base-calling may be stored in the data storage 122. Parameters identifying polony or cluster locations may also be stored in the data storage 122. Raw and / or processed image intensities of each polony or cluster may be stored in the data storage. The region and / or subtile that each polony or cluster corresponds to may also be stored in the data storage 122. The transformation matrix of each region and / or subtile for different cycle(s) and / or channel(s) may also be stored in the data storage 122. Cell images may be stored in the data storage. The flow cell images, the processed images, and / or the filtered images may be stored in the data storage. Other information or images that can facilitate 3D base calling of the sample can be saved in the data storage.
[0062] The user interface 124 may be used by a user to operate the sequencing system or access data stored in the data storage 122 or the computer system 126.
[0063] The computer system 126 may control the general operation of the sequencing system and may be coupled to the user interface 124. It may also perform steps in image processing, base calling, their preceding operations, and / or subsequent operations including but not limited to operations leading to chromatic aberration correction. In some embodiments, the computer system 126 is a computer system 400, as described in more detail in FIG. 4. The computer system 126 may store information regarding the operation(s) of the sequencing system 110, such as configuration information, instructions for operating the sequencing system 110, or user information. The computer system 126 may be configured to pass information between the sequencing system 110 and the cloud 130.
[0064] The computing system 126 can include one or more general purpose computers that provide interfaces to run a variety of program in an operating system, such as Windows™ or Linux™. Such an operating system typically provides great flexibility to a user. In some embodiments, the computing system 126 may include one or more processors, e.g., CPUs, the CPUs may be configured for artificial intelligence algorithm development and training (e.g., neural network training), either alone or in combination with the reconfigurable logic device and / or integrated circuit 120.
[0065] In some embodiments, the sequencing system may include one or more reconfigurable logic devices 120 and / or one or more other integrated circuits 120. The reconfigurable logic device 120 can include one or more FPGA devices. The integrated circuit 120 herein may or may not be reconfigurable, and it may include an Al chip, an application-specific integrated circuit (ASIC) chip, a neural processing unit (NPU), or a combination thereof. In some embodiments, the reconfigurable logic device and / or integrated circuit 120 may be configured for artificial intelligence algorithm development and training (e.g., training of a neural network), either alone or in combination with the CPU and / or GPU.
[0066] In some embodiments, the reconfigurable logic device and / or integrated circuit 120 include a main unit and an edge unit. For example, the main unit may be a FPGA device and the edge unit may be an ASIC or Al chip. In some embodiments, the edge unit is an additional hardware processing module that may be individually installed and / or uninstalled on the system 110. The edge unit may be configured for artificial intelligence algorithm development and training. The edge unit may be configured for making inferences or predictions using deployed Al algorithm(s), e.g., neural networks. The edgeunit may communicate electronically with the main unit e.g., data communication via DMA connections. The edge unit may communicate electronically for data with other parts of the system 100 via various connections, such as a chip2chip connection. As an example, the edge unit may include a neural processing unit (NPU) chip, an Al chip, or any other integrated circuit(s).
[0067] In some embodiments, the dedicated processors 118 may be configured to perform operations in the methods disclosed herein. The dedicated processors 118 may include one or more reconfigurable logic devices and / or integrated circuits disclosed herein. The dedicated processors 118 may not include general -purpose processors, but instead custom processors with specific hardware or instructions for performing those steps. Dedicated processors directly run specific software without an operating system. The lack of an operating system reduces overhead, at the cost of the flexibility in what the processor may perform. A dedicated processor may make use of a custom programming language, which may be designed to operate more efficiently than the software run on general-purpose computers. This may increase the speed at which the steps are performed and allow for real time processing.
[0068] In some embodiments, the reconfigurable logic device and / or the integrated circuit 120, e.g., FPGA and / or Al chip, may be configured to perform some or all of operations in the methods herein. The reconfigurable logic device and / or the integrated circuit may be programmed as hardware that can perform specific task(s). A special programming language may be used to transform software steps into hardware componentry. Each software step may correspond to at least one operation or action in the methods disclosed herein. Each software step may include at least a part of the operation or action in the methods disclosed herein. Once the reconfigurable logic device is programmed, the hardware directly processes digital data that is provided to it without running software. The reconfigurable logic device and / or integrated circuit instead uses logic gates and registers to process the digital data. Because there is no overhead required for an operating system, the reconfigurable logic device and / or integrated circuit generally processes data faster than a general-purpose computer. Similar to dedicated processors, this may be at the cost of flexibility. The lack of software overhead may also allow the reconfigurable logic device and / or the integrated circuit to operate faster than a dedicated processor, although this will depend on the exact processing to be performed and the specific the reconfigurable logic device and / or integrated circuit and dedicated processor.
[0069] A group of the reconfigurable logic devices and / or integrated circuits 120 may be configured to perform the steps in parallel. In some embodiments, a number of processing engines of the FPGA(s) may be configured to perform one or more identical image processing steps for an image, a set of images, a subtile, or a select region in one or more images. Each FPGA(s) 120 may perform its own part of the image processing step(s) in parallel, reducing the time needed to process data. This may allow the image processing step(s) to be completed in real time. For example, a number of processing engines of a first FPGA may be configured to generate a polony map for a tile of the flow cell. Each processing engine may be responsible for generating a portion, e.g., non-overlapping portion, of the polony map at a different subtile within the tile, e.g., in parallel. A second FPGA may be configured to perform intensity normalization in parallel with the generation of the polony map. As another example, a number of FPGA(s) and integrated circuits, e.g., Al chips, may be configured to perform one or more image processing step(s) for the flow cell images. Each FPGA(s) 120 may perform its own part of the processing step(s) in parallel, reducing the time needed to process data, while each Al chip may perform polony or cluster prediction after receiving data from its corresponding FPGA. This may allow the image processing steps to be completed in real time. For example, a first and second FPGA may be configured to perform intensity registration in parallel for a different subtile or tile of the flow cell. A corresponding Al chip may perform prediction of high resolution flow cell image of the corresponding subtile or tile after image registration is completed by its corresponding FPGA. Further discussion of the use of FPGAs is provided below.
[0070] The reconfigurable logic device and / or the integrated circuit may be configured to perform some or all of the operations or actions in the methods disclosed herein in real time. Performing the operations or actions in real time may allow the system 110 to use less memory and / or data storage, as the data may be processed as it is received. This is an improvement over conventional systems that may need to store the data before it may be processed and consequently require more memory / data storage or accessing a computer system located in the cloud 130. Further, performing the operations or actions in real time may allow more efficient sequencing analysis as it is being performing in parallel while a sequencing run is still in progress. Furthermore, performing the processing steps using the FPGAs and Al chips may allow the system to use less power, e.g., 2x, 5x, lOx, 20x ormore, thus producing less heat than performing the same processing steps using the CPUs and / or GPUs. Further discussion of the use of FPGAs is provided below.
[0071] As discussed above, the sequencing system 110 may have dedicated processors 118, the reconfigurable logic device and / or integrate circuitl20, or the computer system 126. The sequencing system may use one, two, or all of these elements to accomplish one or more operations or actions in the methods disclosed herein. In some embodiments, when these hardware elements are present together, the image processing tasks are split between them. For example, the reconfigurable logic device 120 may be used to perform some or all of the preprocessing operations color correction, polony map generation, image registration, predicting high resolution flow cell images, training a neural network for predicting high resolution flow cell images, generating the training flow cell images, base calling, and any subsequent operations, while the computer system 126 may perform other processing functions for the sequencing system 110 such as intensity normalization, registering images for base calling with cell staining image(s). Those skilled in the art will understand that various combinations of these elements will allow various system embodiments that balance efficiency and speed of processing with cost of processing elements.
[0072] In some embodiments, one or more reconfigurable logic devices and / or integrated circuits 120 can accelerate base calling and / or any primary analysis steps of flow cell images acquired from 2D or 3D sample(s). In some embodiments, the reconfigurable logic devices and / or integrated circuits can accelerate primary analysis of 2D sample(s) or 3D volumetric sample(s) by 2x, 4x, 5x, lOx, 15x, 20x, 25x, 30x, 40x, 50x, lOOx, 200x, 400x, 500x, 800x, lOOOx, or more than traditional primary analysis methods using only CPUs and / or GPUs. In some embodiments, one or more reconfigurable logic devices and / or integrated circuits 120 herein can accelerate sequencing and sequencing analysis (including at least primary analysis) of the flow cell images acquired from 2D or 3D sample(s). In some embodiments, the reconfigurable logic devices and / or integrated circuits herein can accelerate sequencing and sequencing analysis (including at least primary analysis) of the flow cell images acquired from 2D or 3D sample(s) by 2x, 4x, 5x, lOx, 15x, 20x, 25x, 30x, 40x, 50x, lOOx, 200x, 400x, 500x, 800x, lOOOx, or more than traditional sequencing systems with only CPUs and / or GPUs. In some embodiments, making inferences or predictions of high resolution images, of base calls, or of classifications, using the neural network disclosed herein and the reconfigurable logicdevices and / or integrated circuits can be less than 800 ms, 500ms, 400ms, 300ms, 200 ms, 100ms, 50ms, 20 ms, or less per tile per cycle. The tile size can be varied in different flow cells. The tile size may be at least 0.0012mm, 0.01 mm2, 0.05 mm2, 0.1 mm2, 0.5 mm2, 1 mm2, 2 mm2, 3 mm2or more.
[0073] In some embodiments, one or more reconfigurable logic devices and / or integrated circuits 120 can enable primary analysis (base calling) of polonies for flow cell images at multiple z levels. For example, processing time using reconfigurable logic devices can be less than 400 hours for at least 50 flow cell images (e.g., covering 50 tiles and from two or more color channels) with a FOV of at least 1 mm2with a resolution of 1 um or better in three dimensions for one or more flow cycles, e.g., 1-15 cycles. The flow cell images can be from multiple z- levels to cover some or all of the volumetric 3D samples (e.g., completely covering at least two samples).
[0074] In some embodiments, one or more reconfigurable logic devices and / or integrated circuits 120 can be used for accelerating primary analysis of 3D samples involving training neural network(s) and using the trained neural networks for making predictions or inferences. For example, neural network(s) can be used to predict polony locations and / or predict cell boundaries or segmentation thereby identifying polonies within the cell(s). Using the reconfigurable logic device and / or integrated circuits 120 for computations associated with neural networks can reduce the training and / or prediction time needed in comparison with usage of GPUs or other computer processors, thereby accelerate sequence analysis, and enabling sequence analysis of flow cycles while subsequent flow cycles are to be performed or in progress in the sequence run. In some embodiments, the reconfigurable logic device(s) and / or integrated circuits 120 can accelerate training and / or prediction by lOx, 20x, 50x, 80x, lOOx, 200x, 500x, 600x, 800x, lOOOx, or more than training and / or prediction using CPUs and / or GPUs. In some embodiments, the reconfigurable logic devices and / or integrated circuits 120 can be used to achieve optimal acceleration in sequencing analysis. For example, one or more FPGA chips can be used in combination with an integrated circuit specific for computations corresponding to artificial intelligence (Al) algorithms, e.g., a NPU. The integrated circuit(s) can be specific circuits for Al functions. The integrated circuit(s) can include application-specific integrated circuits (ASIC). Computational tasks can be distributed to the FPGA(s) and the integrated circuit(s) to optimize computational time, energy consumption, heat dissipation, etc. For example, the Al chip may be used only forcomputations involving a neuron network (e.g., predicting polony locations in the polony map, predicting high resolution flow cell images, or training the neural network) and the FPGA(s) may be used for the rest of the primary analysis steps. The primary analysis time using dual FPGA chips or single FGPA chip in connection with the Al chip(s) can be less than 400, 300, 200, 100, 50, or 20 hours for at least 50 flow cell images (e.g., covering about 50 tiles of the flow cell and from two or more color channels) with a FOV of at least 1 mm2with a resolution of 1 um or better for each flow cell image in three dimensions for one or more flow cycles, e.g., 1-15 cycles. The flow cell images can be from multiple z-levels to cover some or all of the volumetric 3D samples (e.g., 10 to 20 z- locations to completely cover at least two samples). The primary analysis time may include a total time of image processing from obtaining raw flow cell images acquired using the imager 116 to generating base calls and saving base call results. The 3D samples herein includes polonies or clusters that are centered at different z levels that are spaced apart from each other with at least 0.01 um, 0.05 um, 0.1 um, 0.2 um, 0.5 um, 1 um, or more along the z direction or axial direction.
[0075] The cloud 130 may be a network, remote storage, or some other remote computing system separate from the sequencing system 110. The connection to cloud 130 may allow access to data stored externally to the sequencing system 110 or allow for updating of software in the sequencing system 110.
[0076] In some embodiments, the sequencing system 110 comprises: a first reconfigurable logic device comprising a first plurality of data processing engines configured to perform data processing in parallel; first reconfigurable routing channels connecting at least some of the first plurality of data processing engines; a neural network deployed at least partly on the first reconfigurable logic device; a first processor that selectively activates or deactivates different combinations of the first plurality of data processing engines and the first reconfigurable routing channels to perform one or more operation(s) herein to facilitate generating the sequencing analysis result(s). The sequencing analysis may include operations or steps of primary analysis. Such operation(s) may include one or more of: obtaining sensor data directly from one or more sensors of the sequencing system; processing the sensor data to generate a first set of flow cell images; determining, by the sequencing system, aberration-corrected base calls for a first set of polonies in the first cycle of the sequencing run and optionally forwardingflow cell images and / or base calls to the first reconfigurable logic device, the first processor, or one or more hardware processors of the sequencing system.
[0077] In some embodiments, obtaining sensor data from one or more sensors (in the imager 116) of the sequencing system may be via a direct connection. In some embodiments, the direct connection between the first reconfigurable logic device and the sensor(s) lacks other hardware components that may process or store the sensor data thus causes undesired complexity, delay, and possible errors in sensor data communication. Such hardware components include the first processor , the memory device, or any processors, e.g., CPUs 126 of the sequencing system. Comparing with traditional sequencing systems in which sensor data is communicated to other hardware components before it is communicated to where it is being processed (e.g., communicating to CPU and then to GPU to be processed) the direct sensor data communication herein advantageously improves data transmission efficiency from the sensor to the FPGAs 120, frees-up the other hardware(s), e.g., CPUs, storage devices, for other data processing functions, decreases power consumption from indirect data communication, and reduces time consumption in data communication thus sequencing analysis.
[0078] In some embodiments, the connection between the first reconfigurable logic device and sensor may include other hardware components that may process or store the sensor data. Such hardware components may include the first processor, the memory device, or any processors, e.g., CPUs 126 of the sequencing system. For example, the sensor data may be saved into the memory device, and then it can be accessed by the first reconfigurable logic device using memory controlled s).
[0079] The reconfigurable logic device may include digital logic circuits therein, in a sense that it is also an integrated circuit. However, the integrated circuit herein (e.g., the Al chip, NPU, etc.) may have various difference with the reconfigurable logic device, e.g., the integrated circuit may not be as flexible in reconfiguration as the reconfigurable logic device. For example, the integrated circuit herein, e.g., the Al chip, NPU, etc., may not be reconfigurable.
[0080] In some embodiments, the sequencing system 110 comprises at least one reconfigurable logic device but lacks any integrated circuits, e.g., Al chips, ASIC chips, or NPUs. The reconfigurable logic device may perform one or more operations in sequencing analysis and may forward its output back to the CPU as end results of primary analysis, e.g. base calls. Alternatively, the reconfigurable logic device may forward itsoutput back to the CPU so that subsequent operations may be performed based on its output by the CPU to generate the end results of sequencing analysis.
[0081] In some embodiments, the sequencing system 110 comprises at least one reconfigurable logic device, and at least one integrated circuit. The integrated circuit may perform one or more operations in sequencing analysis and may forward its output back to the reconfigurable logic device so that subsequent operations may be performed based on its output at the reconfigurable logic device.
[0082] In some embodiments, the output of the reconfigurable logic device or the integrated circuit comprises base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises identification of base calling locations in two dimensions. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises identification of base calling locations in three dimensions. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises flow cell images with high resolution and free of chromatic aberration or with reduced level of chromatic aberration.
[0083] In some embodiments, the data communication between any two of the reconfigurable logic device, the integrated circuits, the first processor, and the second processor may be direct such that the direct communication lacks any other hardware components that may process or store the data. Such other hardware components may include memory device(s), and / or other processor(s) of the sequencing system. Such direct communication may include DMA connections. In some embodiments, the data communication the data communication between any two of the reconfigurable logic device, the integrated circuits, the first processor, and the second processor may be direct such the data may not be utilized by other logic circuits or stored before reaching its communication destination, but the data may be stored in a memory device before reach its communication destination.
[0084] The sequencing system may further comprise one or more memory devices electrically connected for data communication with one or more components of the sequencing system, the one or more components may include one or more of: the first reconfigurable logic device, the integrated circuit, the first reconfigurable routing channels; the one or more memory controllers; the first processor; a second processor; and one or more processors of the sequencing system.
[0085] In some embodiments, the first reconfigurable routing channels are configured to allow data communication between the first reconfigurable logic device and one or more memory devices. In some embodiments, the one or more DMA connections and the first reconfigurable routing channels are configured to allow data communication between the first reconfigurable logic device and the integrated circuit.
[0086] In some embodiments, the sequencing system further comprises an integrated circuit that is different from the first reconfigurable logic device,. The integrated circuit herein may not be reconfigurable. The integrated circuit may comprise an application specific integrated circuit (ASIC) chip. In some embodiments, the integrated circuit comprises a neural processing unit (NPU) or an artificial intelligence (Al) chip. The integrated circuit may comprise a second plurality of data processing engines, each data processing engine comprising multiple digital logic circuits. The integrated circuit may further comprise: second plurality of data processing engines and second routing channels, each connecting at least some of the second plurality of data processing engines.
[0087] In some embodiments, the sequencing system further comprises a first processor.The first processor may be configured to selectively activate or deactivate different combinations of the first plurality of data processing engines and the first reconfigurable routing channels to perform the operations disclosed herein. In some embodiments, the sequencing system further comprises a second processor. The second processor may be configured to control digital circuits of the integrated circuit herein.
[0088] In some embodiments, the first processor, or a second processor, e.g., of the integrate circuit, is configured to selectively activate or deactivate different combinations of the second plurality of data processing engines and the second reconfigurable routing channels to perform the operations. The first processor or a second processor may be configured to selectively activate or deactivate different combinations of the second plurality of data processing engines and the second reconfigurable routing channels to perform the operations herein.
[0089] The sequencing system may further comprise a housing that encloses the first reconfigurable logic device, the first reconfigurable routing channels, the one or more DMA connections, the integrated circuit, and the first processor therein. In some embodiments, the sequencing system further comprising: a housing that encloses at leastthe first reconfigurable logic device therein and the integrated circuit is external to the housing.
[0090] In some embodiments, the sequencing system further comprises: a power source that is configured to supply different power levels to the first reconfigurable logic device and the integrated circuit. A first power level supplied by the power source to the first reconfigurable logic device may be higher than a second power level supplied to the integrated circuit while a sequencing run and / or sequencing analysis is in progress. A maximum power output of the power source of the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the maximum power output of the power source of sequencers, e.g., traditional sequencers without the first reconfigurable logic device (e.g., FPGA), the integrated circuit (e.g., Al chip), or both. The time consumption in performing a sequencing run and corresponding sequencing analysis (e.g., primary analysis) thereof using the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the time consumption in performing the same sequencing run using a sequencer without the first reconfigurable logic device, the integrated circuit, or both (e.g., a traditional sequencer without FPGA and / or Al chips). Time consumption in performing a sequencing run and sequencing analysis of the sequencing run (e.g., primary analysis) using the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the time consumption in performing the same sequencing run and analysis using a sequencer without the first reconfigurable logic device, the integrated circuit, or both(e.g., a traditional sequencer without FPGA and / or Al chips). In some embodiments, a maximum power output of the power source to the sequencing system in performing a sequencing run and corresponding sequencing analysis thereof is less than 900 Watts, 800 Watts, 700 Watts, 650 Watts, 600 Watts, 550 Watts, or 500 Watts. The power source may be configured to supply a first power level to the first reconfigurable logic device, the first power level is less than 500 Watts, 400 Watts, 350 Watts, or 300 Watts. The power source may be configured to supply a second power level to the integrated circuit; the second power level is less than 450 Watts, 400 Watts, 350 Watts, or 300 Watts.
[0091] In some embodiments, one or more components of the first reconfigurable logic device and / or integrated circuit may include a computational performance of at least 2, 4, 8, 10, 16, 20, 30, 40, 50, 60, 70, 80, or 100 Giga-operations per second (GOPs) or more. In some embodiments, one or more processing engines of the first reconfigurable logic device and / or integrated circuit may include a computational performance of at least 12,4, 8, 10, 16, 20, 30, 40, 50, 60, 70, 80, or 100 Giga-operations per second (GOPs), or more Giga-operations per second (GOPs), or more. In some embodiments, the first reconfigurable logic device and / or the integrated circuit includes a computational performances of at least 10, 20, 40, 50, 60, 80, or 100 Tera-operations per second (TOPs).
[0092] In some embodiments, one or more components are located on a first printed circuit board (PCB). The one or more components may include: the first reconfigurable logic device the first reconfigurable routing channels; the first processor; and the one or more DMA connections. In some embodiments, the integrated circuit is located on a second printed circuit board (PCB) different from the first printed circuit board. The integrated circuit and the second PCB may be positioned within a same housing of the sequencing system as the first PCB or external to the housing of the sequencing system. Being on a separate PCB makes connecting the first reconfigurable logic device, e.g., FPGA device with various integrated circuit on a chip convenient, efficient, and easily customizable. In some embodiments, the first PCB board may be a main board, and the second PCB board may be a daughter board or edge unit.
[0093] In some embodiments, the sequencing systems lacks any graphic processing units (GPUs) or tensor processing units (TPUs). Instead, the sequencing systems utilizes FPGAs, Al chips, NPUs, or other ASIC chips for performing the operations disclosed herein. The sequencing system disclosed herein advantageously requires less power, generates less heat, and reduces the hardware complexity and costs for performing NGS sequencing runs and corresponding sequencing analysis than sequencers that uses GPUs or TPUs.
[0094] In some embodiments, the sequencing systems include logic devices that are not limited to reconfigurable logic devices (e.g., FPGAs) and / or other integrated circuits (e.g., Al chips, NPUs). In some embodiments, the sequencing systems include various types of processing units or processors configured for reconfigurable parallel processing, In some embodiments, the sequencing systems include various types of logic devices or integrated circuits, e.g., ASIC chips. In some embodiments, the sequencing systems include GPUs, TPUs, or other various types of processing units that are configured to perform one or more operations disclosed herein.
[0095] In some embodiments, the sequencing systems include GPUs, TPUs, or other various types of processing units that are configured to perform one or more operationsthat can be performed by the reconfigurable logic devices (e.g., FPGAs) and / or other integrated circuits (e.g., Al chips, NPUs).
[0096] The first processor may be positioned on the first PCB board together with the reconfigurable logic device for convenient and efficient control of the reconfigurable logic device. In some embodiments, the first processor is a separate processor from one or more processors of the sequencing system configured to control the optical system, the fluidics of the sequencing system, etc. In some embodiments, the first processor can be configured to only control the components on the first PCB board, e.g., the FPGA device, alone or in combination with components on the second PCB board, e.g., the Al chip. In some embodiments, the sequencing system may comprise a second processor that is configured to separately control the Al chip. The first processor or second processor of the sequencing system, e.g., 120_c, may comprise a CPU. The one or more hardware processors of the sequencing system comprises a CPU.
[0097] In some embodiments, the sequencing system may further comprise a heat dissipator configured to maintain a system temperature in a range from 0 degrees to 120 degrees Celsius or less than 120 degrees Celsius.
[0098] In some embodiments, each of the one or more operations performed by the first reconfigurable logic device or the integrated circuit are in real time. In some embodiments, each of the one or more operations performed by the first reconfigurable logic device or the integrated circuit are within the time window of performing sequencing reactions and / or imaging of a single sequencing cycle of the sequencing run. In some embodiments, each of the one or more operations performed by the first reconfigurable logic device or the integrated circuit are within the time window of performing sequencing reactions and / or imaging of a single z-level of a single sequencing cycle.
[0099] The flow cell images herein may be obtained from multiple z levels covering at least partly of an in situ sample, e.g., of cells or tissue(s). The flow cell images may be obtained from one or more color channels at each z level of the multiple z levels covering at least partly of the in situ sample. The flow cell images may be of a first spatial resolution in x, y, and / or z directions, e.g., the first set of flow cell images in method 2800. The extracted intensities (e.g., the preliminary intensities, the first or second set of intensities in method 2800 ) may be generated based on the first set of flow cell images. The flow cell images may be of a second spatial resolution in x, y, and / or z directions.The extracted intensities may be of a second spatial resolution in x, y, and / or z directions or of the same first spatial resolution. The first spatial resolution may be lower than the second spatial resolution, and a higher resolution herein indicate that a pixel size is smaller so that the polonies in the flow cell images are of finer spatial details. The first spatial resolution may be 2x, 4x, 6x, 8x, lOx, 16x, 24x, 32x, or 48x lower than the second spatial resolution in x, y, and / or z directions. The first spatial resolution may be at least 2x, 4x, 6x, 8x, lOx, 16x, 24x, 32x, or 48x lower than the second spatial resolution in x,y, and / or z directions. In some embodiments, the first and second resolution is in 3D. In some embodiments, the first resolution is in a range of 0.1 um to 5 um. In some embodiments, the second resolution is in a range of 0.01 um to 2 um. In some embodiments, the second resolution is at least 4, 6, or 8 times greater than the first resolution in all three dimensions.
[0100] In some embodiments, the sequencing system further comprises one or more image sensors, e.g., 116 in FIG. 1, configured to receive optical signals generated from sequencing reactions of a sample immobilized on a support. The support may comprise a glass or plastic substrate. The support may be comprised in a flow cell device. The one or more image sensors may be configured to generated sensor data based on the optical signals. In some embodiments, the sequencing system further comprises: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations disclosed herein. The one or more data storage devices may include one or more memory devices. The one or more memory devices may be accessible by the one or more processors, the first processor, the second processor, the first reconfigurable logic device, the integrated circuit.
[0101] In some embodiments, the one or more processors are separate from the first or second processors. The operations performed by the one or more processors may include one or more of: 1) recording sensor data generated in the sequencing system in one or more flow cycles; 2) optionally processing the recorded sensor data; 3) sending the recorded sensor data or the optionally processed data to the first reconfigurable logic device or the integrated circuit; 4) receiving outcome from the first reconfigurable logic device or integrated circuit; and 5) generating sequencing analysis results based on the received outcome. The operations performed by the one or more processors may include one or more of: 1) receiving outcome from the first reconfigurable logic device orintegrated circuit; and 2) generating sequencing analysis results based on the received outcome.
[0102] In some embodiments, the sequencing analysis results comprise primary analysis results. In some embodiments, the sequencing analysis results comprise a data file in a predetermined data format. In some embodiments, the sequencing analysis results comprise base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the sequencing analysis results comprises quality measurements of base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the sequencing analysis results comprises quality scores corresponding to base calls of nucleotide bases in a sample immobilized on a support.
[0103] In embodiments where the sample is a 3D volumetric sample, the operations are performed for a single z level in each cycle within a predetermined time window. The predetermined time window is for a single z level in a single sequencing cycle. In some embodiments, the predetermined time window is less than 1000 ms, 900 ms, 800 ms, 700ms, 600 ms, 500 ms, 400 ms, 300 ms, 250 ms, 200 ms, or 100 ms. In some embodiments, each of the one or more operations are performed within the predetermined time window and in parallel while the sequencing run is in progress. In some embodiments, each of the one or more operations are performed in parallel within a time window that sequencing, imaging, or both of a subsequent sequencing cycle is completed.
[0104] The flow cell images herein may be obtained from multiple z levels covering at least partly of an in situ sample, e.g., of cells or tissue(s). The first plurality of flow cell images may be obtained from one or more color channels at each z level of the multiple z levels covering at least partly of the in situ sample. In some embodiments, the first plurality of flow cell images are from a single color channel. The first plurality of flow cell images may be of a first spatial resolution in x, y, and / or z directions. The second plurality of flow cell images may be generated based on the first plurality of flow cell images. The second plurality of flow cell images may be of a second spatial resolution in x, y, and / or z directions. The first spatial resolution may be lower than the second spatial resolution, and a higher resolution herein indicate that a pixel size is smaller so that the polonies in the flow cell images are of finer spatial details. The first spatial resolution may be 2x, 4x, 6x, 8x, lOx, 16x, 24x, 32x, or 48x lower than the second spatial resolution in x, y, and / or z directions. The first spatial resolution may be at least 2x, 4x, 6x, 8x, lOx, 16x, 24x, 32x, or 48x lower than the second spatial resolution in x,y, and / or z directions.In some embodiments, the first and second resolution is in 3D. In some embodiments, the first resolution is in a range of 0.1 um to 5 um. In some embodiments, the second resolution is in a range of 0.01 um to 2 um. In some embodiments, the second resolution is at least 4, 6, or 8 times greater than the first resolution in all three dimensions.
[0105] In some embodiments, the first plurality of flow cell images are from one or more color channels. In some embodiments, the first plurality of flow cell images are of unbalanced nucleotide diversity. In some embodiments, the first plurality of flow cell images comprises: an unbalanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more flow cycles. In some embodiments, the first plurality of flow cell images comprises: a balanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more cycles. In some embodiments, two or more different concatemer molecules among the concatemer molecules have different insert sequences. In some embodiments, different insert sequences correspond to different target RNA molecules or target cDNA molecules. In some embodiments, each location of the determined polonies corresponds to a location of the concatemer molecules. In some embodiments, the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents bound to a balanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support. In some embodiments, the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents bound to a unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support in the one or more subsequent cycles In some embodiments, the unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules comprises: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of bases is less than 20%, 15%, 10%, or 5% in the one or more cycles. In some embodiments, the balanced diversity of nucleotide bases of A, G, C and T / U among the plurality of concatemer molecules comprises: a percentage of (I) a number of each type of nucleotide bases to (2) a total number of bases in the one or more cycles is more than 10%, 15%, or 20%
[0106] In some embodiments, the sample(s) comprises overloaded concatemer molecules with a spatial density in a range of 102-1015per mm2. In some embodiments, the sample(s) comprises overloaded concatemer molecules with a spatial density in a range of 103-1010per mm2.
[0107] In some embodiments, one or more of operations herein are performed while a sequencing run is being performed. In some embodiments, the first plurality of flow cell images are acquired in sequencing cycles ranging from 1 to 500. In some embodiments, the one or more cycles comprises a current cycle N. In some embodiments, wherein N is in a range from 1 to 500. In some embodiments, the one or more cycles comprises a single cycle ranging from 1 to 500. In some embodiments, the one or more cycles comprises multiple cycles ranging from 1 to 500. In some embodiments, one or more of operations are performed while the sequencing reactions in cycles subsequent to the current cycle N is yet to be performed or currently being performed. In some embodiments, the z-axis is orthogonal to image planes of the flow cell images.
[0108] In some embodiments, the operation of providing locations of the first set of polonies to the neural network comprises: generating a 2D or 3D polony map comprising spatial location of the first set of polonies. In some embodiments, the polony map includes spatial locations (e.g., coordinates in a reference coordinate system) of some or all of the first set of polonies. In some embodiments, the polony map may be at the same spatial resolution as the first, second, or third set of flow cell images. In some embodiments, the polony map may be at the same spatial resolution as the low-resolution or high-resolution flow cell images or the extracted intensities (e.g., the preliminary, first, or second intensities in method 2800). In some embodiments, the polony map may be generated using various algorithms. In some embodiments, the polony map may be generated using flow cell images from early cycles of a sequencing run, e.g., cycles 1-5. In some embodiments, the polony map may be generated using flow cell images from 4 color channels. The operation of generating a 2D or 3D polony map comprising spatial location of the first set of polonies may further comprise: deleting duplicate polonies, wherein the duplicate polonies are out-of-focus. In some embodiments, the operation of the operation of providing locations of the first set of polonies to the neural network comprises: superimposing the high resolution flow cell images with corresponding cell staining images; and generating the polony map by only including polonies that are within cell boundaries in the corresponding cell staining images. Exemplary embodiments of methods for generating 2D or 3D polony map are disclosed in U.S. Patent Application No. 18 / 078,820 and PCT Application No. PCT / US2023 / 076125, which are incorporated by reference in their entireties.Aberration correction
[0109] DNA sequencing may utilize detection with less than 4 colors for improved sequencing throughput, optical system simplicity, cost reduction, and other various advantages over existing sequencing method using 4 color channels. In embodiments of DNA sequencing using less than 4 colors, e.g., three color excitation and sequencing, a flow cell image corresponding to an individual color channel is acquired to cover at least a portion of the flow cell, e.g., a corresponding field of view (FOV). Three flow cell images may be acquired in total corresponding to the three color channels, e.g., emission signal in blue, red, and green, of the same corresponding FOV per cycle per z-level. For at least one color channel, e.g., the green channel, with excitation by the one or more illumination sources, two different types of fluorescent labels of the sample(s) may be excited to emit signals in two different wavelengths or wavelength ranges (e.g., a narrow range spanning only a few nanometers of the same color). The corresponding flow cell image may include emission signals of such two different wavelengths or wavelength ranges, e.g., two different colors in red and green. The optical system for DNA sequencing utilizing only less than 4 color channels may remove at least some emission filters that can existing in traditional optical systems for DNA sequencing with four colors, so that emission signal of two different wavelength or wavelength ranges may be collected by a single image sensor of the same color channel, and appear as bright spots in a same flow cell image.
[0110] The emission signals in different colors may look like similar bright spots since the flow cell images are in grayscale. Chromatic aberration may occur within such flow cell images with emission signals of more than one wavelength, e.g., more than one color. With chromatic aberration, polonies or clusters with emission signal in a first color may experience a different spatial shift (in 2D or even in 3D) from the spatial shift of polonies or clusters with emission signal in a second color. Such difference may vary spatially across the flow cell image. In other words, the level of spatial shift between two colored emission signals may vary depending on the position of the pixel(s) within the flow cell image. As a nonlimiting example, two polonies of different colors close to the center of the flow cell image may experience less shift than two other polonies of different colors at the edge of the flow cell image. As an example, chromatic aberration may cause the emission signal of a first polony to shift and overlap with the signal of a neighboring polony, resulting in error in base calling of the neighboring polony. Further, base callingof the first polony may also be adversely affected due to the shift of the emission signal. Such chromatic aberration may cause error in identifying image intensities for polonies and result in base calling errors. Thus, there is a need for correcting chromatic aberration in sequencing analysis of flow cell images acquired with signal of different wavelengths in a single image, e.g., using 3 color sequencing or any other number of colors that is less than 4, which is same number of types of nucleotides.[OHl] In some embodiments, chromatic aberration may be corrected using the methods and systems disclosed herein. FIG. 27 shows a flow chart of an exemplary embodiment of a computer-implemented method 2800 for chromatic aberration correction in flow cell images acquired with signal of different wavelengths or wavelength ranges in a single image, e.g., using sequencing using less than 4 color channels and for generating accurate and reliable base calls with no or reduced level of chromatic aberration using such flow cell images.
[0112] The methods herein, e.g., 2800, can include some or all of the operations disclosed herein. The operations may be performed in but is not limited to the order that is described herein.
[0113] In some embodiments, one or more operations of the methods for chromatic aberration correction, e.g., method 2800, are performed during an active sequencing run such that aberration-corrected base calls are generated in real time in a current cycle, prior to completion of subsequent sequencing cycles.
[0114] The methods herein, e.g., method 2800, can be performed by one or more processors disclosed herein. In some embodiments, the processor can include one or more of: a computing system comprising a processing unit 126, a reconfigurable logic device 120, an integrated circuit that is not reconfigurable 120, or their combinations. For example, the processing unit can include a central processing unit (CPU). The reconfigurable logic device can include one or more FPGA devices. The integrated circuit can include a chip such as an Al chip or an ASIC chip. In some embodiments, the one or more processors can include the computing system 400 disclosed herein.
[0115] In some embodiments, some or all operations in methods herein, e.g., method 2800, can be performed by the reconfigurable logic device, e.g., the FPGA(s), and / or the integrated circuit, e.g., the Al chip. In embodiments when some operations are performed by the reconfigurable logic device and / or integrated circuit, e.g., FPGA(s), the data produced by the reconfigurable logic device and / or integrated circuit, e.g., the FPGA(s)after performing one or more operations, can be communicated to various hardware elements of the system 100, e.g., CPU(s) or GPU(s), so that subsequent operation(s) in methods herein, e.g., method 2800, can be performed by such various hardware using the communicated data. Similarly, data can also be communicated in the opposite direction from various hardware e.g., CPU(s), to the reconfigurable logic device or the integrated circuit for processing. In some embodiments, all the operations in method 2800 can be performed by CPU(s). Alternatively, the operations performed by CPU(s) can be performed by other processors such as the dedicated processors, or GPU(s). In some embodiments, all the operations in methods herein, e.g., method 2800, can be performed by the reconfigurable logic device and / or the integrated circuit, e.g., FPGA(s) and / or the Al chip(s).
[0116] In some embodiments, the sensor data acquired by the imager 116 may be directly communicated to the reconfigurable logic device and / or the integrated circuit, e.g., via DMA connections. In some embodiments, the sensor data acquired by the imager 116 may be directly communicated to the reconfigurable logic device and / or the integrated circuit without being routed first to a CPU, a GPU, or any other processing units before reaching the reconfigurable logic device and / or the integrated circuit.
[0117] In some embodiments, aberration correction using the methods herein, e.g., method 2800, with the reconfigurable logic device, e.g., the FPGA, and / or other integrated circuit, e.g., Al chips, may require at least 2x, 8x, lOx, 15x, 20x, 40x, 50x, or lOOx less power than making the same computations using other computing hardware including but not limited to CPUs or GPUs.
[0118] In some embodiments, the sequencing system herein for performing method 2800 may comprise: a first reconfigurable logic device comprising a first plurality of data processing engines configured to perform data processing in parallel; first reconfigurable routing channels connecting at least some of the first plurality of data processing engines; an aberration correction algorithm at least partly deployed on the first reconfi urable logic device; and a first processor that selectively activates or deactivates different combinations of the first plurality of data processing engines and the first reconfigurable routing channels to perform the one or more operations.
[0119] The sequencing system may further comprise: a power source that is configured to supply identical or different power levels to the reconfigurable logic device and the integrated circuit. In some embodiments, a maximum power output of the power sourceto the sequencing system in performing methods herein, e.g., method 2800, is less than 2000 Watts, 1000 Watts, 900 Watts, 800 Watts, 700 Watts, 650 Watts, 600 Watts, 550 Watts, 500 Watts, 400 Watts, 300 Watts, 200 Watts, or 100 Watts.
[0120] In some embodiments, the sequencing system herein comprises: a first reconfigurable logic device, e.g., a FPGA unit, comprising a plurality of data processing engines configured to perform data processing in parallel; first reconfigurable routing channels, each connecting at least some of the first plurality of data processing engines; a neural network or an aberration correction algorithm deployed at least partly on the first reconfigurable logic device; a first processor to selectively activate or deactivate different combinations of the first plurality of data processing engines and the first reconfigurable routing channels to perform one or more operations in methods herein (e.g,, method 2800).
[0121] In some embodiments, the first reconfigurable logic device and the integrated circuit is within the same physical housing as the other elements of the sequencing system as show in FIG 1. In some embodiments, the first reconfigurable logic device and the integrated circuit is not physically external to the sequencing system 110 as show in FIG 1, e.g., not in the cloud 130.
[0122] The method 2800 may comprise an operation 2810 of obtaining, by the sequencing system, a first set of flow cell images corresponding to sensor data acquired by one or more image sensors of the sequencing system of biological sample(s) immobilized on a flow cell device.
[0123] The sample(s) may comprise concatemer molecules therewithin. The sample(s) may include concatemer molecules from one or more different sample sources, e.g., different organs, or different species. The sample(s) may include a thickness along the z- axis so that the first plurality of flow cell images may be acquired at a z-stack of different z-locations with a first resolution to cover the sample in 3D. The sample (s) may be acquired from a single z-location of a 2D or 3D sample.
[0124] The sample can be in situ. The sample can be a 3D sample. The sample can be a volumetric sample that may contain different biological information at the same x-y location but different z levels. The sample can be a cellular sample including multiple cells, tissue, or their combination. The sample can be any sample that has a thickness that is greater than a predetermined threshold along the z axis. For example, the thickness can be greater than 2 um, 3 um, 4 um, 5 um, 10 um, 20 um, or more. The z axis (e.g., z axis)is orthogonal to the image plane defined by x and y axes. In some embodiments, the sample can be traditional 2D sequencing samples. In embodiments where the sample is 3D, the sample comprises concatemer molecules therewithin. In embodiments where the sample is 2D, the sample comprises template molecules therewithin.
[0125] In some embodiments, the sensor data at one of the one or more image sensors comprises emitted signal from the sample at more than one wavelengths or wavelength ranges. In some embodiments, each wavelength or a wavelength range corresponds to a single color, e.g., red. In some embodiments, the sensor data at one of the one or more image sensors comprises emitted signal from the sample with more than one color, e.g., red and green, or red, green, and yellow. In some embodiments, the sensor data at one of the one or more image sensors comprises emitted signal with two or more colors selected from: blue, red, green, and yellow. In some embodiments, at least one of the one or more image sensors is configured to sense emitted signal of the sample with more than one color.
[0126] In some embodiments, the emitted signal with more than one color sensed by the one or more image sensor may include some level of chromatic aberration. In some embodiments, the chromatic aberration may be cause by the optical element(s) of the imager 116. In some embodiments, the chromatic aberration may cause a polony or cluster to shift spatially in 2D or 3D, e.g., within a x-y plane, relative from its place without chromatic aberration. As a result, a same polony or cluster (e.g., a signal spot) may shift to different positions of the flow cell image at different cycles, resulting in difficulty in aligning the same polony or cluster across cycles and / or across different color channels for base calling. Additionally, the polony or cluster with chromatic aberration may also cause confusion in correctly identifying its neighboring or overlapping polonies or cluster. In some embodiments, a first polony of the first set of polonies in the flow cell image comprises a first level of chromatic aberration, and wherein a second polony of the first set of polonies comprises a second level of chromatic aberration that is different from the first level of chromatic aberration, thus the spatial shift may be of a different distance and / or toward a different 3D direction. In some embodiments, the first or second level of chromatic aberration is less than 0.1 pixels, 0.2 pixels, 0.5 pixels, 0.8 pixels, 1 pixel, 1.5 pixels, or 2 pixels. In some embodiments, the first or second level of chromatic aberration herein may cause a spatial shift, in 2D or 3D, that is less than 0.1 pixels, 0.2 pixels, 0.4 pixels, 0.6 pixels, 0.8 pixels, 1 pixel, 1.2 pixels,1.5 pixels, 1.8 pixels, 2 pixels, 2.5 pixels, 3 pixels or more. In some embodiments, the first or second level of chromatic aberration herein may cause a spatial shift, in 2D or 3D, that is less than 0.01 um, 0.02 um, 0.04 um, 0.06 um, 0.08 um, 0.1 um, 0.12 um, 0.15 um, 0.18 um, 0.2 um, 0.25 um, 0.3 um, 0. 35 um, 0.4 um, 0. 45 um, 0. 5 um, 0.6 um, or more.
[0127] In some embodiments, the chromatic aberration herein is in 2D. In some embodiments, the chromatic aberration herein is within the x-y plane. In some embodiments, the spatial shift of the polony caused by chromatic aberration may be in any directions in 3D, e.g., along x, along y, or along z axis. As an example, a first polony of green emission light at (xl, yl, zl) and a second polony of red emission light may be spaced apart from each other for 0.8 pixels with the second polony at (x2, y2, z2). Due to chromatic aberration, the second polony may be shifted away from the first polony along x axis for 0.3 pixels resulting in the polony center at (x2+0.3, y2, z2). Thus, extract intensity of the second polony at (x2, y2, z2) may be different from the intensity at (x2+0.3, y2, z2), possibly cause error in base calling of the second polony.
[0128] In some embodiments, chromatic aberration may result in base calling errors. In some embodiments, the chromatic aberration may cause at least 0.5%, 1%, 2%, 3%, 5%, 8%, 10%, 12%, 15%, 18%, or 20% base calling errors when compared with base calling free of aberration correction, e.g., the aberration-corrected base calls. In some embodiments, the aberration-corrected base calls comprises at least 0.5%, 1%, 2%, 3%, 5%, 8%, 10%, 12%, 15%, 18%, or 20% less base calling errors than preliminary base calls (e.g., made with no chromatic aberration correction).
[0129] In some embodiments, the first set of flow cell images comprises at least two flow cell images, e.g., 2, 3, or 4, per cycle per z level and per FOV, and wherein each flow cell image corresponds to different emitted signal from the same FOV of the sample(s) in response to a different illumination per cycle per z level. In some embodiments, the first set of flow cell images comprises 2, 3, or 4 flow cell images, and wherein each flow cell image corresponds to a different color channel within a same cycle of a sequencing run. As a nonlimiting example, three different illuminators, e.g., red, green, and blue, may be used for illuminating the sample(s) in a cycle of a sequencing run, e.g., simultaneously or sequentially, and emitted signals in a same FOV in response to a different illumination may be sensed by a single image sensor sequentially, thereby generating three different flow cell images in the same sequencing cycle (e.g., 211-1 in FIGS. 2-3). Such operationmay be repeated in different sequencing cycles generating different sets of flow cell images including the first set of flow cell images .
[0130] In some embodiments, the methods 2800 further comprises an operation of
[0131] acquiring, by of the sequencing system, at least one flow cell image of the first set of flow cell images in a first color channel of the sample immobilized on the flow cell device in the first cycle of the sequencing run.
[0132] In some embodiments, the methods 2800 further comprises an operation of acquiring, by the sequencing system, at least one full flow cell image of a first set of full flow cell images in a first color channel of the sample immobilized on the flow cell device in the first cycle of the sequencing run; and an operation of obtaining, by the sequencing system, the first set of flow cell images by obtaining a portion of the at least one full flow cell image. In some embodiments, the full flow cell image may cover the one or more samples across a tile, a subtile, multiple tiles, multiple subtiles, or at least portion thereof of the flow cell device. In some embodiments, the full flow cell image may cover the one or more samples across at least 10 subtiles, 20 subtiles, 40 subtiles, 60 subtiles, 80 subtiles, 100 subtiles, 140 subtiles, 180 subtiles, 200 subtiles, 300 subtiles, or more of the flow cell device. In some embodiments, the full flow cell image may have a field of view (FOV) of at least 1 mm2, 2 mm2, 3 mm2, 5 mm2, 6 mm2, 8 mm2, 10 mm2, 20 mm2, 30 mm2, 40 mm2, 50 mm2, 60 mm2, or more. Each flow cell image in the first set of flow cell images may cover at least a portion of the full flow cell image, and the full flow cell image may be divided into 1, 2, 4, 8, 16, 64, 128, 256, 400, 420, 512, 1024 or more flow cell images, and each of the divvied flow cell image may be comprised in the first set of flow cell images. In some embodiments, each flow cell image may only cover a portion of the full flow cell image. In some embodiments, each flow cell image in the first set of flow cell images may be a full flow cell image. In some embodiments, it is advantageous to divide the full flow cell images into portions, i.e., the flow cell images in the first set of flow cell images, to reduce the chromatic aberration variation across the flow cell images.
[0133] In some embodiments, the operation of acquiring, by the imager of the sequencing system, the first set of flow cell images of the sample immobilized on the flow cell device in the first cycle of the sequencing run comprises: illuminating or exciting, by a first illuminator of the sequencing system, the sample(s) to cause emission of a signal of at least a first color and a second color simultaneously; allowing the emitted signal of atleast the first color and second color to travel through an identical optical path of the imager of the sequencing system to arrive at the one or more image sensors of the sequencing system; detecting the emitted signal by the one or more image sensors of the sequencing system; and generating, by the sequencing system, a first flow cell image of the first set of flow cell images based on the detected signal. In some embodiments, the first and second color may be green and red.
[0134] In some embodiments, the operation of acquiring, by the imager of the sequencing system, the first set of flow cell images of the sample(s) immobilized on the flow cell device in the first cycle of the sequencing run further comprises: illuminating, by the first illuminator or a second illuminator of the sequencing system, the sample(s) to cause emission of a signal of at least a third color; allowing the emitted signal of at least the third color to travel through the identical optical path of the imager to arrive at the one or more image sensors of the sequencing system; detecting the emitted signal by the one or more image sensors of the sequencing system; and generating, by the sequencing system, a second flow cell image of the first set of flow cell images based on the detected signal. In some embodiments, the third color may be blue.
[0135] In some embodiments, the operation 2810 of obtaining the first set of flow cell images corresponding to the sensor data acquired by the one or more image sensors of the sequencing system is performed in the first cycle or in a second cycle of the sequencing run subsequent to the first cycle. In some embodiments, the first or second cycle is among the first 1, 2, 3, 5, 10, 15, or 20 cycles of the sequencing run.
[0136] In some embodiments, the flow cell images of the first set or the full flow cell image herein can be acquired using the optical system or the imager 116 disclosed herein, from the 1, 2, 3, 4, or more color channels. Each flow cell image can include at least a portion of one or more tiles (e.g., imaging areas or FOV). Each tile can be divided into multiple subfiles. Each tile or subtile can include a plurality of polonies or clusters. Each subtile can include multiple regions with each region including a number of polonies or clusters. The flow cell image as disclosed herein can be an image that is acquired from a flow cell 112 as shown in FIG. 1 or 2712 as shown in FIG. 5. In some embodiments, the flow cell images are acquired from one or more color channels, and at least the flow cell images from at least one channel of the one or more color channels comprises emission signal at two different wavelengths, e.g., with two different colors.
[0137] In some embodiments, the full flow cell image herein can be an image of one or more tiles. In some embodiments, the full flow cell image herein can be an image that covers the entire sample(s). In some embodiments, the flow cell image in the first set of flow cell images herein can be an image of one or more tiles, one or more subtiles, one or more segmented regions within tile(s) or subtile(s), or their combinations. Each flow cell image can comprise a field of view (FOV). The FOV can be orthogonal to the z axis. The FOV can be within the x-y plane. The FOV of different flow cell images at different z levels can be identical within the x-y plane. The FOV of different flow cell images at different z levels can have at least an overlapping portion within the x-y plane. The image resolution of different flow cell images at different z levels can be about identical or exactly identical. In some embodiments, The image resolution of different flow cell images at different z levels is different. The FOV can be in 3D and be of various sizes to cover the volumetric sample to be imaged. The FOV along x, y, and / or z direction can be in a range from 10 um to 5 mm. The FOV along x, y, and / or z direction can be in a range from about 0.05 um to about 2 mm. The FOV along x, y, and / or z direction can be in a range from 0.5 um to 1 mm. For example, the FOV can be about 1 mm by 1 mm by 20 um for certain cellular samples along the x, y, and z direction, respectively.
[0138] The flow cell images herein may be of various sizes. The pixel number along x, y, and / or z axis may be any integer greater than 64 or 128. The flow cell images herein may be of various sizes, the pixel number along x, y, and / or z axis may be in a range from 2 to 65536. A single flow cell image can be separated into different number of regions, for example, 4, 8, 16, or even more regions, and each region may include a size of 256 by 256 by 1, 512 by 512 by 3, or other sizes. In some embodiments, the number of pixels along x, y, and / or z direction may be adjusted to maintain a particular spatial resolution in a given FOV. For example, with a spatial resolution of 0.2 um, to cover a FOV of 0.8 mm, the number of pixels may be 4000.
[0139] Each flow cell image at a specific z level may include intensities generated by polonies or clusters at the corresponding z level. Signals from polonies or clusters can be small bright spots within the images. Each bright spot can be of various sizes that is less than a couple of pixels, e.g., less than a pixel, about a pixel, about 2 pixels, 3 pixels, 4, pixels, 5 pixels, or more. In some embodiments, each signal spot of the polonies or clusters can be any number of pixels in the range from 0.01 pixel to about 100 pixels. Insome embodiments, each signal spot of the polonies or clusters can be any number of pixels in the range from 0.1 pixel to about 16 pixels.
[0140] Each flow cell image can also include intensities generated by the cell and its structural elements. Such structural elements can be background objects or components. Each flow cell images can also include noise and / or artifacts that are not from the polonies or cellular structures. The structural elements may be in focus or out-of-focus in the flow cell images.
[0141] The structural elements may be in focus in at least some regions and out-of-focus in at least some regions.
[0142] In some embodiments, when the depth of field the optical system includes a range, e.g., 0.1 um, 0.2 um, 0.3 um, 0.5 um, 0.6 um, 0.8 um, 1 um, 2 um, 3, um, 4 um, 5 um, etc. expanding along z axis, polonies or clusters that are within the range of depth of field can appear in-focus or about in-focus in the flow cell image. Flow cell images at a specific z level can also include signals from polonies or clusters that are not within the focus range of the image. Such polonies or clusters are out-of-focus. In some embodiments, bigger and blurry signal spots represent out-of-focus polonies or clusters.
[0143] Each flow cell image at a specific z level can also include noises caused by the optical system and / or undesired signal from the sample. The undesired signal can be signal coming from components of the sample such as membrane, cytosol, and mitochondria. Such background objects can be any objects, relatively larger in size than the polonies or clusters. In some embodiments, background objects can include any objects within the sample(s) but are not polonies or clusters.
[0144] In some embodiments, the first set of flow cell images are of unbalanced nucleotide diversity. In some embodiments, the flow cell images comprises: an unbalanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more sequencing cycles. In some embodiments, the flow cell images comprises: a balanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more cycles. In some embodiments, two or more different concatemer molecules among the concatemer molecules have different insert sequences. In some embodiments, different insert sequences correspond to different target RNA molecules or target cDNA molecules. In some embodiments, each location of the determined polonies corresponds to a location of the concatemer molecules. In some embodiments, the flow cell imagescomprises optical signals emitted from nucleotide reagents bound to a balanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support. In some embodiments, the flow cell images comprises optical signals emitted from nucleotide reagents bound to a unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support in the one or more subsequent cycles. In some embodiments, the unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules comprises: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of bases that is less than 20%, 15%, 10%, or 5% in the one or more sequencing cycles. In some embodiments, the balanced diversity of nucleotide bases of A, G, C and T / U among the plurality of concatemer molecules comprises: a percentage of (1) a number of each type of nucleotide bases to (2) a total number of bases in the one or more cycles is more than 10%, 15%, or 20%. As an example, bases calls from the polonies include 4 different bases, and percentage of polonies for each of the 4 different bases can be greater than about 10% so that the data are of balanced diversity. As another example, , bases called from the plurality of polonies includes 4 or less different bases, and percentage of polonies for one or more bases can be less than about 10%, and such data can be considered as unbalanced diversity. In some embodiments, bases called from the plurality of polonies include 4 or less different bases, and percentage of polonies for some of the bases can be less than about 5%, about 2%, or even about 1%, and such data can be considered as unbalanced diversity. As yet another example, the unbalanced diversity data include bases A, T, C, G in the plurality of polonies, and their percentages of the total base calls are about 1%, about 2%, about 1%, and about 95%, respectively. In addition to the base biases affecting diversity, plexity can also be a factor that when plexity is lower than a number, e.g., 8 or 16, the signal is of unbalanced diversity.
[0145] The method 2800 may be configured to correct chromatic aberration in flow cell images even if the polonies or clusters in the acquired flow cell images are of unbalanced diversity in one or more sequencing cycles.
[0146] In some embodiments, the operation 2810 further comprises: generating, by a processor of the sequencing system, the first set of flow cell images with a second resolution higher than the first resolution of the flow cell images that was acquired using the sequencing system. In some embodiments, generating the first set of flow cell images with the second resolution higher than the first resolution may use various algorithms. Insome embodiments, generating the first set of flow cell images with the second resolution higher than the first resolution may use a neural network based algorithm. In some embodiments, the second resolution of the flow cell images is at least 2x, 4x, 6x, 8x, 12x, 16x, or more than the first resolution of the flow cell images. In some embodiments, the first set of flow cell images of the first resolution or second resolution comprises chromatic aberration. Such chromatic aberration may be corrected using the methods and systems disclosed herein.
[0147] In some embodiments, the method 2800 comprises an operation 2820 of determining, by the sequencing system, a set of preliminary intensities of a first set of polonies in the first set of flow cell images. In some embodiments, determining the set of preliminary intensities of the first set of polonies may be based on a preliminary registration of the first set of polonies, e.g., to a reference coordinate system such as a polony map. In some embodiments, the first set of polonies includes some or all of the polonies in the first set of flow cell images. In some embodiments, the first set of polonies includes at least a portion of the polonies in the first set of flow cell images. In some embodiments, the first set of polonies includes some or all of the polonies in the full flow cell images. In some embodiments, the first set of polonies includes some or all of the polonies in the sample(s) immobilized on the flow cell.
[0148] In some embodiments, the set of preliminary intensities are from two, three, or four color channels. In some embodiments, the set of preliminary intensities are comprised in a set of preliminary images, and each preliminary image corresponds to a different type of nucleotide base. In some embodiments, the set of preliminary intensities are comprised in a set of preliminary images, and each preliminary image corresponds to a different color corresponding to the color of emitted signal from the fluorescent label representing individual nucleotide bases. For example, as shown in FIGS. 2-3, the set of preliminary images are 212-1, which include 4 different preliminary images, each image corresponding to a different type of nucleotide base, e.g., A, T, C, and G and a different emission color of a corresponding fluorescent label representing the nucleotide bases, e.g., blue, red, green, and yellow. The set of preliminary intensities may include intensities of the first set of polonies in the set of preliminary images. The set of preliminary intensities may be saved or stored in various forms or formats. In some embodiments, the set of preliminary intensities may be a list of intensities of the first set of polonies, their corresponding pixels / spatial coordinates, and their corresponding colors.In some embodiments, the set of preliminary intensities may be in the form of 2D preliminary images with intensity values at some or all pixels the preliminary images. In some embodiments, the set of preliminary intensities may be preliminary images with intensity values only at pixels of the first set of polonies.
[0149] In some embodiments, the preliminary registration of the first set of polonies is configured to register the first set of polonies in the first set of flow cell images so that the first set of polonies in different flow cell images are aligned in a reference coordinate system without considering chromatic aberration. In some embodiments, the first set of polonies are used as fiducials for aligning the first set of flow cell images, e.g., to the reference coordinate systems. In some embodiments, the polony map is in the reference coordinate system.
[0150] In some embodiments, the preliminary registration comprises determining and applying one or more spatial transformations to the first set of polonies. In some embodiments, each spatial transformation may be linear. In some embodiments, each spatial transformation may be an affine transformation. In some embodiments, each transformation may include spatial translation and / or rotation. In some embodiments, each transformation may include nonlinear transformation or nonlinear distortion correction.
[0151] In some embodiments, the one or more spatial transformation may include a single transformation that can be applied to the first set of polonies. In some embodiments, the one or more spatial transformation may include a single transformation that can be applied to different subsets of the first set of polonies. For example, such subsets may divide the first set of polonies based on their spatial locations.
[0152] In some embodiments, each spatial transform may be determined using a subset of polonies within the images to be transformed. The subset may be selected based on but not limited to one or more criteria, including signal intensity, signal-to-noise ratio, spatial distribution within the field of view, and preliminary base call confidence / quality. The spatial transform may comprise a rigid transform, affine transform, polynomial transform, spline-based transform, or higher-order distortion model. Parameters of the spatial transform may be determined by minimizing an error metric between corresponding polony locations in different flow cell images. Suitable error metrics include least-squares error, weighted least-squares error, robust estimators, or RANS AC -based matching. Insome embodiments, the transform parameters may be iteratively refined until a convergence threshold is satisfied.
[0153] In some embodiments, the operation 2820 of determining, by the sequencing system, a set of preliminary intensities of the first set of polonies based on the preliminary registration of the first set of polonies to the polony map comprises: extracting, by the sequencing system, the set of preliminary intensities from the preliminarily registered first set of flow cell images based on locations of the first set of polonies, e.g., in the polony map or in the reference coordinate system. By registering each flow cell image in the first set of flow cell images to the reference coordinate system or polony map, four preliminary flow cell images can be determined. In some embodiments, the registration of a flow cell image herein comprises determining and applying one or more spatial transforms to the flow cell image via the registration. A single spatial transformation, e.g., linear or non-linear, may be applied to the flow cell image. In some embodiments, multiple spatial transformations may be applied to different portions of the flow cell image. Each preliminary flow cell image corresponding to a different type of nucleotide base, e.g., A, T, C, and G and a different color representing the nucleotide bases, e.g., blue, red, green, and yellow.
[0154] An exemplary label detection scheme for different types of nucleotide bases from the first set of flow cell images is shown in Table 1, wherein “1” represents bright signal for the nucleotide base in the corresponding flow cell image in the first set of flow cell images of the corresponding color channel, and wherein “0” represents dark or no signal for the nucleotide base in the corresponding flow cell image in the first set of flow cell images. Using the labeling and label detection scheme, 3 flow cell images per FOV per cycle per z level may be acquired from the green, red, and blue color channels. For example, in the flow cell image acquired using the green channel, two different types of nucleotides with correspond different fluorescent labels may appear as similar bright spots spanning one or more pxiels. Four preliminary images, e.g., 212-1 in FIG. 2, may be determined from the less than 4 flow cell images in the first set of flow cell images, e.g., 3 flow cell images in 211-1. The three flow cell images may each correspond to a different color channel of green, red, and blue channels.
[0155] Table 1. An exemplary label detection scheme in three color sequencing.
[0156] In some embodiments, the method 2800 may include an operation 2830 of determining, by the sequencing system, the preliminary base calls based on the extracted intensities. Various base calling methods may be used to perform base calling using the extracted intensities of preliminarily registered images, e.g., 212-1. For example, preliminary base calling may consider extracted intensities of preliminarily registered images in all 4 different color channels, and determine the base call as the nucleotide base corresponding to the color channel of the highest intensity, optionally taking into consideration of a quality (e.g., a quantitative quality metric) of the preliminary intensities. FIGS. 2-3 shows exemplary preliminary base calls 213-1. The preliminary base calls may include errors that are caused by chromatic aberration. In some embodiments, chromatic aberration may cause at least 1%, 2%, 5%, 6%, 8%, 10%, 12%, 15%, 18%, 20%, or more errors in base calling in comparison to base calling with at least some level of chromatic aberration correction.
[0157] In some embodiments, the method 2800 may include an operation 2840 of identifying, by the sequencing system, a first subset of polonies of the first set of polonies based on the preliminary base calls.
[0158] In some embodiments, the method 2800 may include an operation of filtering the preliminary base calls, e.g., using a quantitative quality metric, of the preliminary base call, and the operation 2840 of identifying the first subset of polonies is based on the filtered preliminary base calls. In some embodiments, using the filtered preliminary base calls may help reduce the errors in preliminary base calls that are caused by chromatic aberration.
[0159] In some embodiments, the method 2800 may optionally include an operation of identifying, by the sequencing system, a second subset of polonies of the first set of polonies based on the preliminary base calls. The first and second subset of polonies may be non-overlapping subsets of the first set of polonies. The first and second subset may be selected using various methods. For example, the first subset may be polonies that are called a first type of nucleotide base in the preliminary base calls and the second subset may be polonies that are called a second type of nucleotide base in the preliminary base calls. In some embodiments, the operation 2840 of identifying, by the sequencing system, the first subset of polonies of the first set of polonies based on the preliminarybase calls comprises: identifying, by the sequencing system, the first subset of polonies of the first set of polonies based on a first type of base call among the preliminary base calls.
[0160] In some embodiments, the first and second types of nucleotide bases are those whose emission signal are acquired within the same flow cell images. In the first and second types of nucleotide bases are those whose emission signal are acquired within the same flow cell image but are at different wavelengths, wavelength ranges, or colors.
[0161] In some embodiments, the operation 2840 is based on the labeling scheme, e.g., Table 1. In some embodiments, the first subset of polonies may be the polonies whose emission signal are sensed and captured in a first flow cell image of the first set of flow cell images, and the second subset of polonies may be the polonies whose emission signal are sensed and captured in the same first flow cell image but are at a different wavelength or color as those of the fist subset of polonies. In an exemplary embodiment, polonies or clusters are labeled to emit fluorescent signals as shown in Table 1. The first subset of polonies are As and the second subset of polonies are Cs, and both subsets appear in the flow cell image acquired after excitation with a green light illumination or excitation.
[0162] In some embodiments, the method 2800 includes an operation 2850 of determining, by the sequencing system, a first set of intensities of the first set of polonies based on registration of the first subset of polonies to the reference coordinate system or the polony map. In some embodiments, the operation 2850 of determining, by the sequencing system, a first set of intensities of the first set of polonies based on registration of the first subset of polonies to the polony map comprises: registering, by the sequencing system, the first subset of polonies to the polony map, thereby generating a first set of registered flow cell images; and extracting, by the sequencing systems, the first set of intensities of the first set of polonies from the first set of registered flow cell images based on locations of the first set of polonies in the polony map. FIG. 2 shows an exemplary embodiment of the first set of intensities of the first set of polonies as intensities in the first set of registered flow cell images 212-2. In some embodiments, the operation 2850 may be performed using a reference coordinate system for registration instead of a polony map.
[0163] In some embodiments, registering the first subset of polonies to the polony map may ignore other polonies in the first set of polonies in the registration. In some embodiments, each of the first set of registered flow cell images, e.g., 212-2, correspondsto a different colors, and the first set of registered flow cell images may have 4 images corresponding to 4 different colors per FOV per cycle.
[0164] In some embodiments, the method 2800 may optionally include an operation 2850’ of determining, by the sequencing system, the second set of intensities of the first set of polonies based on registration of the second subset of polonies to the polony map. The operation 2850’ may comprise: registering, by the sequencing system, the second subset of polonies to the polony map, thereby generating a second set of registered flow cell image; and extracting, by the sequencing systems, the second set of intensities of the first set of polonies from the second set of registered flow cell images based on locations of the first set of polonies in the polony map. FIG. 2 shows the second set of intensities of the first set of polonies as intensities in the second set of registered flow cell images, e.g., 212-3. As shown in FIG. 2, in some embodiments of method 2800 with operations 2850 and 2850’, the first set of intensities, and the second set of intensities of the differently registered flow cell images, e.g., 212-2 and 212-3, are used to determine the shifted intensities, which may then be used to determine the aberration-corrected base calls.
[0165] In some embodiments, registering the second subset of polonies to the polony map may ignore other polonies in the first set of polonies in the registration. In some embodiments, each of the second set of registered flow cell images, e.g., 212-3 in FIG. 2, corresponds to a different color.
[0166] In some embodiments, the method 2800 may include an operation 2860 of selecting, by the sequencing system, one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities based on evaluation of quality (e.g., quantitative quality metric) of the one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities.
[0167] In embodiments where the second set of intensities are not obtained, the operation 2860 comprises selecting, by the sequencing system, one or more of: the set of preliminary intensities and the first set of intensities based on evaluation of quality (e.g., one or more quantitative quality metrics) of the one or more of: the set of preliminary intensities and the first set of intensities.
[0168] In some embodiments, each subset of the preliminary intensities in the set of preliminary intensities, each subset of the first intensities in the first set of intensities, and each subset of the second intensities in the second set of intensities correspond to a different color or wavelength. In some embodiments, evaluation of quality (e.g., usingone or more quantitative quality metrics) of one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities comprises: evaluation of quality of the set of preliminary intensities, the first set of intensities, and the second set of intensities. In some embodiments, evaluation of quality of one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities comprises: evaluation of quality of the first set of intensities and the second set of intensities.
[0169] In some embodiments, evaluation of quality of one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities comprises: calculating, by the sequencing system, a preliminary quantitative value for one or more quality metrics for each individual polony of the first set of polonies; calculating, by the sequencing system, a first quantitative value for the quality metric(s) for each individual polony of the first set of polonies; and / or calculating, by the sequencing system, a second quantitative value for the quality metrics for each individual polony of the first set of polonies. In some embodiments, the quality metrics comprises one or more of: clarity, a quality score, a chastity, a max_intensity, a low intensity median clarity, a phasing value, and a prephasing value.
[0170] In some embodiments, the quality metric(s), e.g., clarity, may be calculated per polony, per subtile, per tile, or per field of view. In some embodiments, the quality metric(s), e.g., clarity, may be calculated for at least some polonies of the first set of polonies. In some embodiments, clarity values are pre-processed, e.g., normalized, prior to comparison, e.g., among the preliminary intensities, first intensities, or shifted intensities. In some embodiments, the clarity metric is based on a contrast-to-noise ratio, signal-to-noise ratio, peak-to-background ratio, or a combination thereof.
[0171] In embodiments where the quality metric is determined per polony, the quantitative quality matric of a set of intensities, e.g., the preliminary intensities or the first intensities, is based on some or all the values for each polonies, e.g., a maximum, a sum, an average, a median, etc.
[0172] In some embodiments, the clarity of a polony or multiple polonies may be determined based on a ratio of a max intensity to a second max intensity. For one or more polonies or clusters within a cycle, after determining a brightest channel among multiple color channels, e.g., highest average image intensity among flow cell images from 4 color channels. The max intensity of the one or more polonies or clusters may beobtained from the brightest channel. The second max intensity of the same one or more polonies or clusters may be obtained from the second brightest channel.
[0173] The max intensity may be normalized. The normalization may be by a predetermined value, e.g., by 90th percentile brightest intensity within the channel, for flow cell images of each channel. The normalization may be performed during preprocessing operations of the flow cell images disclosed herein before estimation of quality. Before obtaining the max intensity, the flow cell image may be preprocessed using one or more of the image processing operations disclosed herein. The normalization may be performed after background subtraction, correction, and phasing and prephasing correction. After normalization, the signal intensity may be scaled to a pre-determined range, e.g., [0, 2000], [0, 2500], or [0, 3000], The predetermined range may be determined to encode the range of normalized intensities in integers. To make the clarity changing inversely and / or negatively with the quality score, the scaled intensity may be negated and added to an intensity offset.
[0174] The max intensity may be obtained as c- cl*(int / int_norm), wherein c and cl may be identical or different integers, int is the intensity of the polony or cluster, and int norm is the intensity used for normalization of intensity int. For example, int norm may be the brightest intensity of the flow cell image after background subtraction and color correction. The max intensity may be obtained for each polony or a group of polonies or clusters within the flow cell image in a cycle. In other words, the max_intensity may be per polony-cycle.
[0175] The second max intensity may be determined similarly as the max intensity disclosed herein, but with respect to the second brightest channel among multiple channels, e.g., the second highest average intensity among 4 channels. The second max intensity may be obtained for each polony within the flow cell image in a cycle. In other words, the second max_intensity may be per polony-cycle. As such, clarity may be per polony-cycle. The max intensity may also be per polony-cycle. In some aspects, the clarity may be for a group of polonies or clusters per cycle. Similarly, the max intensity may also be for a group of polonies or clusters per cycle.
[0176] Alternatively, the max intensity of the one or more polonies or clusters may be obtained from the brightest intensity of the individual polonies or clusters across four different channels, and the second max intensity of the one or more polonies or clustersmay be obtained from the second brightest intensity of the individual polonies or clusters across 4 different channels.
[0177] The low intensity median clarity may be indicative of density of polonies or clusters in the flow cell images. The low intensity median clarity may comprise a median clarity for a selected number of polonies or cluster, e.g., in a subtitle. A flow cell image may be separated into multiple tiles, and each tile is an imaging area that may have multiple subtiles, and each subtile may have an equal or different number of polonies. A flow cell image disclosed herein may be an image of a single tile. Alternatively, a flow cell image disclosed herein may be an image of multiple tiles. The low intensity median clarity is inversely proportional to the median clarity of selected polonies per tile-cycle in the one or more flow cell images. The selected number of polonies are dim polonies selected from the flow cell image. A dim polony may be a polony with an intensity lower than a pre-determined percentage of the brightest signal intensity in a tile or a subtile. For example, the pre-determined percentage may be 15%, 20%, 25%, 30%, or any other percentages lower than 45%. Alternatively, a dim polony belongs to the darkest population of polonies with a pre-selected percentage. The pre-selected percentage may be 8%, 10%, 12%, or 15% of the total population of polonies within the subtile or tile. A median of all clarity values from the dim polonies or clusters, subtiles, or tiles may be used to generate the low intensity median clarity.
[0178] The median clarity of dim polonies or clusters, subtiles, or tile may be further inversed so that it is smaller when the quality score estimation is more accurate. The low intensity median clarity may be calculated as c*median(max_r2 / max_r), wherein c is a constant, max r is maximal signal intensity of the dim polonies within a selected region of the flow cell image, and max_r2 is the second maximal signal intensity of the dim polonies within the same selected region. The low intensity median clarity may be calculated as c / median(max_r / max_r2), wherein c is a constant, max r is maximal signal intensity of the dim polonies within a selected region of the flow cell image, and max_r2 is the second maximal signal intensity of the dim polonies within the same selected region. In some aspects, the low intensity median clarity may be calculated using certain constants, e.g., as c*median[(max_r2+a) / (max_r+b)], where a or b can be non-zero numbers.
[0179] The phasing and prephasing may be determined as an average percentage of phasing and prephasing in a selected number of polonies in a region of the flow cellimages. The region of the flow cells are tiles or subtiles as described herein. The corresponding value of the phasing and prephasing is based on multiple polonies in the flow cell image. In some aspects, the phasing and prephasing can be per subtile-cycle or per tile-cycle. The average percentage of phasing and prephasing can be indicative about how much image intensity is affected (e.g. decreased) by the phasing and prephasing effect.
[0180] In some embodiments, the method 2800 may comprise an operation 2870 of determining, by the sequencing system, aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the selection, e.g., of operation 2860. In some embodiments, instead of aberration-corrected base calls, aberration-corrected classifications may be determined instead. The classification for each pixel may be one of: the four different types of nucleotide bases and the background.
[0181] For operations that determine the base calls herein, e.g., methods 2800, classifications may be determined as an alternative. In some embodiments, the determination of classification may be based on cell images with cell staining and / or cell segmentation information. The cell images of the sample(s) may be acquired using the same sequencing system without moving the sample(s) before or after the sequencing run.
[0182] In some embodiments, the operation 2870 may be performed on individual polonies. In some embodiments, the operation 2870 of determining, by the sequencing system, the aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the quality evaluation comprises: identifying a maximum value for each individual polony among one or more of: the preliminary quantitative value; the first quantitative value; and second quantitative value; determining an aberration- corrected intensity for each individual polony of the first set of polonies as one of: the preliminary intensity, the first intensity, or the second intensity corresponding to the maximum value; and determining, by the sequencing system, aberration-corrected base calls of each individual polony based on the aberration-corrected intensity. FIG. 2 shows the aberration-corrected base calls 215 that can be determined based on selecting intensities among two or more of the preliminary intensities in images 212-1, first intensities of images in images 212-2, and second intensities in images 212-3 based on quality evaluation thereof. For example, if the second intensities of polony X yields the best quantitative value in clarity comparing to clarities of the preliminary intensities and the first intensities, the aberration-corrected base call may be determined using the secondintensities of polony X. In some embodiments, such quality evaluation can be polonybased such that each polony is evaluated separately for the selection. In some embodiments, such quality evaluation can be based on the FOV of the first set of flow cell images, such that polonies are evaluated together for the selection, e.g., by using summation or average of the individual evaluation. For polonies with multiple pixels, one representative pixel and its corresponding quality metric measurement and intensity may be used, e.g., the pixel with maximum intensity and / or maximum quantitative value of the quality metric among the multiple pixels within the polony. Alternatively, average, median, or other processed value among the multiple pixels may be considered for determining the quantitative value and the corresponding intensity representing the polony.
[0183] In some embodiments, the operation 2870 may be performed on a group of polonies that include either some or all of the polonies in the first set of polonies. In some embodiments, the operation 2870 of determining, by the sequencing system, the aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the quality evaluation comprises: determining a first sum of the preliminary quantitative value; a second sum of the first quantitative value; and a third sum of second quantitative value for some or all of the first set of polonies; identifying a maximum for some or all of first set of polonies among the first, second, and third sums; determining an aberration-corrected intensity for each individual polony of the first set of polonies as one of: the preliminary intensity, the first intensity, or the second intensity corresponding to the maximum; and determining, by the sequencing system, aberration- corrected base calls of each individual polony based on the aberration-corrected intensity.
[0184] In some embodiments, the method 2800 may include an operation 2870’ of determining, by the sequencing system, aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the shifted intensities. The operation 2870’ is similar to operation 2870 except that different intensities from the preliminary intensities, first intensities, and second intensities, i.e., shifted intensities, are used for making the base calls.
[0185] In some embodiments, the method 2800 comprises operations 2810 to 2870, without the operations 2850’, 2850”, or 2870’.
[0186] In some embodiments, the method 2800 comprises operations 2810 to 2850, operation 2850’, and operations 2860 to 2870, without operations 2850” or 2870’.
[0187] In some embodiments, the method 2800 comprises operations 2810 to 2850, operation 2850”, and operation 2870’, without the operations 2850’, 2860, or 2870.
[0188] FIG. 3 shows an exemplary schematic of the method 2800 with operation 2850” and operation 2870’. In methods of chromatic aberration correction with operations 2850” and 2870,’ instead of selecting intensities for the first set of polonies (e.g., from preliminary, first, and second set of intensities) for making aberration corrected base calls, shifted intensities are generated based on such intensities (e.g., preliminary, first, and second set of intensities) and locations corresponding to such intensities for the first set of polonies for making aberration corrected base call.
[0189] In some embodiments, determining the shifted intensities comprises determining intensities at spatial coordinates that are offset from preliminary locations by at least noninteger pixels. In some embodiments, determining the shifted intensities comprises interpolating pixel locations of the first set of flow cell images at subpixel-shifted coordinates corresponding to first locations and second locations. In some embodiments, the interpolating comprises bilinear interpolation, bicubic interpolation, spline interpolation, or kernel-based interpolation.
[0190] In some embodiments, the operation 2850” comprises: determining, by the sequencing system, a set of shift intensities of the first set of polonies based on: the first set of intensities; the first set of intensities and the set of preliminary intensities; first locations corresponding to the first set of intensities, preliminary locations corresponding to the set of preliminary intensities, or a combination thereof. FIG. 3 shows the shifted intensities as 212-4.
[0191] In some embodiments, the operation 2850” comprises: determining, by the sequencing system, preliminary locations of the first set of polonies based on registration of the first set of polonies to the polony map; determining, by the sequencing system, first locations of the first set of polonies based on registration of the first subset of polonies to the polony map; and determining, by the sequencing system, shifted locations based on the preliminary locations and the first locations of the first set of polonies. In this particular embodiment, the shifted intensities can be determined as intensities for locations determined based on the first and preliminary locations, e.g., midpoint locations of first and preliminary locations.
[0192] In some embodiments, the operation 2850” comprises: determining, by the sequencing system, first locations of the first set of polonies based on registration of the first subset of polonies to the polony map; and determining, by the sequencing system, second locations of the first set of polonies based on registration of the second subset of polonies to the polony map; and determining, by the sequencing system, a shifted location of each individual polony based on the first locations and the second locations of the first set of polonies. In such embodiments, the first locations and second locations, instead of intensities, are extracted based on the registered flow cell images, which are registered to the polony map or otherwise a reference coordinate system using different subsets of polonies. The first and second locations then may be used to determine locations for extracting the shifted intensities. FIG. 3 shows the shifted intensities as 212-4, which may be determined using locations based on the first locations and second locations of the first set of polonies. The shifted intensities may include, for each polony, a shifted intensity corresponding to each color channel. In the embodiment shown in FIG. 3, the shifted intensities as 212-4 include 4 intensities for each polony in the first set of polonies.
[0193] In some embodiments, the method 2800 comprises operations 2810 to 2840. In some embodiments, the method 2800 further comprises an operation that is different from 2850”as: determining, by the sequencing system, first locations of the first set of polonies based on registration of the first subset of polonies to the polony map; and determining, by the sequencing system, second locations of the first set of polonies based on registration of the second subset of polonies to the polony map. In such embodiments, instead of determining the set of shift intensities based on the set of preliminary intensities, the first set of intensities, and / or the second set of intensities, the method 2800 may rely on the first locations and the second locations in order to determine the set of shift intensities. In such embodiments, the method 2800 further comprises an operation 2870’ as determining, by the sequencing system, aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the set of shifted intensities.
[0194] As shown in FIG. 3, the aberration-corrected base calls may be generated based on the shifted intensities 212-4. The shifted intensities 212-4 may include intensities corresponding to each individual nucleotide base of the 4 different types of nucleotide bases. In some embodiments, the set of the shift intensities are of the shifted locations determined based on the first and second locations. Various mathematical operations canbe used to determine the shifted locations. For example, each of the shifted locations may be the mid-point between corresponding first and second locations. As another example, a different weighting may be customized to the first and second locations so that the shift locations may be closer to one of the first locations and the second locations.
[0195] In some embodiments, the shifted intensity may be obtained at the shifted locations determined using the first and second locations, embodiments, determining the shifted intensities comprises interpolating image intensity values at subpixel-shifted coordinates corresponding to the first and second locations. In some embodiments, interpolation may comprise bilinear interpolation, bicubic interpolation, spline interpolation, or kernel-based interpolation. In some embodiments, the shifted intensities may represent estimated intensities, at a spatial coordinate corresponding to reduced chromatic aberration of each of the polonies.
[0196] In some embodiments, the shifted intensities may be determined using interpolation techniques, including but not limited to nearest-neighbor interpolation, bilinear interpolation, bicubic interpolation, spline interpolation, or subpixel interpolation methods. In come embodiments, pixel values of the first and second locations may be integrated, averaged, or otherwise combined to determine an updated shifted location. Corresponding, intensity values of the first and second locations may also be processed and combined to determine the shifted intensites. The shifted intensities may then be determined at the shifted locations reflecting a corrected spatial alignment of the fluorescence signal relative to the polony location.
[0197] In some embodiments, each flow cell image of the first set of flow cell images comprises a FOV of greater than 0.01 mm2, 0.02 mm2, 0.04 mm2, 0.06 mm2, 0.08 mm2, 0.1 mm2, 0.2 mm2, 0.25 mm2, 0.3 mm2, 0.4 mm2, 0.5 mm2, 0.6 mm2, 0.7 mm2, 0.8 mm2, 1 mm2, 2 mm2, 3 mm2, 4 mm2, 5 mm2, 6 mm2, 7 mm2, 8 mm2, or 10 mm2. In some embodiments, each flow cell image of the first set of flow cell images comprises at least 1 x 105pixels, 5 x 105pixels, lx 106pixels, 2 x 106pixels, 4 x 106pixels, 8 x 106pixels, 16 x 106pixels, 64 x 106pixels, or 256 x 106pixels. In some embodiments, each flow cell image of the first set of flow cell images comprises an image resolution of at least 0.01 um, 0.02 um, 0.04 um, 0.05 um, 0.08 um, 0.1 um, 0.12 um, 0.15 um, 0.18 um, 0.2 um, 0.4 um, 0.5 um, 0.6 um, 0.8 um, 1 um, 1.5 um, 2 um, 3 um, or 5 um. In some embodiments, each flow cell image of the first set of flow cell images comprises an image resolution of at least 0.04 um, 0.05 um, 0.08 um, 0.09 um, 0.1 um, 0.12 um, 0.15um, 0.16 um, 0.18 um, 0.2 um, 0.3 um or 0.4 um. In some embodiments, each flow cell image of the first set of flow cell images covers a portion of at least O.OOOlx, O.OOlx, 0.002x, 0.005x, O.Olx, 0.02x, 0.05x, or O.lx of each full flow cell image of a first set of full flow cell image.
[0198] In some embodiments, the sequencing system comprises a single image sensor. In some embodiments, the sequencing system lacks an emission filter in an optical pathway from a sample to the one or more image sensors. In some embodiments, the one or more sensors comprises a single sensor. In some embodiments, the sequencing system comprises an imager that comprises a single image sensor, and wherein the imager lacks an emission filter in an optical path from the sample to the single image sensor. In some embodiments, the sequencing system lacks any emission filter that selectively allow some wavelengths to pass through while filtering / blocking some other wavelengths in at least one color channel. In some embodiments, the emitted signals corresponding to different nucleotide types are detected using the single image sensor without separating different emission wavelengths using emission filter(s) prior to detection by the same image sensor.
[0199] In some embodiments, chromatic aberration results at least from wavelengthdependent spatial displacement of emitted signals detected within the same image.
[0200] In some embodiments, the emitted signals corresponding to multiple nucleotide types are captured within a same flow cell image acquired by a same image sensor without the use of emission filters separating the wavelengths prior to detection by the same image sensor.
[0201] In some embodiments, the polony map comprises the first set of polonies and their corresponding locations in a reference coordinate system. In some embodiments, the polony map comprises a list of information of at least some polonies of the first set of polonies of polonies, and wherein the information comprises one or more of: a unique identification of an individual polony, 3D coordinates of the individual polony in a reference coordinate system, and 2D coordinates of the individual polony in a reference coordinate system. In some embodiments, the polony map comprises one or more individual maps, each individual map corresponds to a different FOV covering at least a portion of a full flow cell image. In some embodiments, the polony map comprises one or more individual maps, each individual map corresponds to a different FOV covering at least a portion of the first set of polonies.
[0202] In some embodiments, the sequencing system comprises: a processor, a reconfigurable logic device, an integrated circuit, or a combination thereof. In some embodiments, the sequencing system comprises: a CPU, a GPU, a FPGA, an Al chip, a NPU, a TPU, or a combination thereof. In some embodiments, the sequencing system lacks any GPU.
[0203] In some embodiments, the sample is a 2D sample comprising template molecules.In some embodiments, the sample is a 3D sample comprising concatemer molecules. In some embodiments, the sample is a cellular sample comprising in situ cells or tissue. In some embodiments, at least part of the sample comprises predetermined bases in the one or more cycles. In some embodiments, the sample comprises overloaded concatemer molecules with a spatial density in a range of 102-1015per mm2. In some embodiments, the sample comprises overloaded concatemer molecules with a spatial density in a range of 103-IO10per mm2. In some embodiments, the sample comprises unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules in one or more cycles, and wherein the unbalanced diversity comprises: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of all bases; and the percentage is less than 20%, 15%, 10%, or 5% in the one or more cycles. In some embodiments, the flow cell device comprises primer molecules with a surface density in a range from 100 molecules per um2to 106molecules per um2. In some embodiments, the flow cell device comprises template molecules immobilized thereon at a surface density of 102to 1015sites per mm2. In some embodiments, the sample comprises a polony density in a range from 102to 1015polonies per mm2. In some embodiments, the sample comprises a polony density in a range from 103to 1010polonies per mm2.Cell images and staining
[0204] In some embodiments, the sequencing system is configured to acquire one or more cell images that may include images of the cell and / or tissue with various types of staining, e.g., fluorescent staining, configured to show morphological information of the sample. In some embodiments, the one or more cell images can comprise staining of cellular structures that help locate polonies or clusters relative to the stained structures. For example, staining can be of cellular structures or components including but not limited to membranes, nuclei, and mitochondria. Different staining colors may be used to stain different components of the cell.
[0205] In some embodiments, the cell membrane after sequencing analysis and imaging using the sequencing system and reactions can be permeabilized. In some embodiments, the one or more cell images can comprise staining of lipids, such as lipids comprised in the cell membrane. In some embodiments, instead of labeling the lipids, the one or more cell images can comprise staining of one or more transmembrane proteins. The transmembrane proteins can be proteins embedded in the permeabilized membrane.
[0206] In some embodiments, the one or more cell images comprises fluorescence or luminescence signals from cell membranes. The one or more cell images can be microscopic images. The one or more images can be fluorescent images. In some embodiments, different fluorescent colors can be included in the cell images. For example, the nuclei and the cell membrane can be stained with different colors.
[0207] In some embodiments, the one or more cell images can comprise segments of: cells, membranes, nuclei, and / or other morphological structures. In some embodiments, the edge(s) of each segment encompass the entire membrane of the cell within the segment. There can be only one cell in each segment. Some segments may not have any cell in them. In some embodiments, adjacent segments do not overlap with each other. In some embodiments, adjacent segments only overlap with each other by sharing one or more edges. In some embodiments, various segmentation algorithms can be used for segmenting the cells.
[0208] In some embodiments, the cell images disclosed herein are stained. The staining can occur after acquiring flow cell images using the sequencing system 110. In some embodiments, the staining can occur before acquiring sequencing images. The methods of staining the 3D sample such as the cells and tissue can include one or more operations disclosed herein. The staining of the 3D sample can use various methods that can specifically label one or more cell protein(s) that are located mostly in the membrane but with neglectable occurrence in other regions of the cell (e.g., less than 10%, 5%, 2% in amount or concentration).
[0209] In some embodiments, the cell images may be acquired using the sequencing system 100 herein without moving the sample(s) from its position during sequencing. It is advantageous to stain the sample after sequencing and acquire the cell images while keep the samples immobilized to the sample stage of the sequencing system. Some transformation, e.g., rotation, translation, shearing may still occur so that there is a need to registered the flow cell images during sequencing to the cell images acquired aftersequencing and staining. In some embodiments, the cell images may be acquired using optical device(s) external to the sequencing system 100 after the sequencing run has been completed and after moving the sample away from the sequencing system 110.Samples
[0210] In some embodiments, the sequencing system including optical system advantageously enable sequencing and imaging of target analyte(s) or features while they remain intact inside the cell or tissue. In some embodiments, the sample(s) herein include cell or tissue and the targets (e.g., target analytes, structure elements, organelles, etc.) therewithin remain intact during sequencing and / or imaging. In some embodiments, the one or more samples being imaged using the optical systems herein can be 2D or 3D samples. The 2D sample(s) may include traditional nucleotide acid molecules extracted from various sources. The 3D samples can include various samples in which polonies within the sample does not fit into a single z level while keeping the polonies in focus. The 3D samples may include in situ samples such as cells and / or tissues. In some embodiments, the cells or tissue samples are immobilized on the flow cell device or otherwise substrate for sequencing and / or imaging without modifying the spatial locations of targets within the cells or tissue. In some embodiments, the cells or tissue samples are immobilized on the flow cell device or otherwise substrate for sequencing or imaging without modifying the spatial relationship of targets or target analytes within the cells or tissue. In some embodiments, the cells and / or tissue are immobilized with the morphological features, RNA, mRNA, and protein targets of the samples intact inside the cell(s) or tissue during sequencing and / or imaging. In some embodiments, the spatial locations or relationships of the target analytes or targets remain intact during sequencing and / or imaging. In some embodiments, the spatial locations or relationships of the target analytes or targets during sequencing and / or imaging are not manually reconstructed using artificially added structure or features in the sample. For example, the nucleus, cell membrane, mitochondria, and extracellular matrix can retain their relative spatial relationship to each other in the sample(s) during imaging and / or sequencing.
[0211] In some embodiments, the one or more samples herein may include a cell or cells may be cultured on a support, e.g., a flow cell, or on a surface that is transferred to a support, e.g., the flow cell. In some embodiments, a cell may be an adherent cell. I some embodiments, a cell may be a confluent cell. In some embodiments, a cell may be asuspended cell. In some embodiments, a suspended cell may be adhered to the surface by a specific capture mechanism such as an antigen-antibody interaction, or a receptor-ligand interaction, especially including an interaction of a known surface receptor with a known ligand; an unknown surface receptor with a known ligand, an known ligand with an unknown ligand, or an unknown receptor with an unknown ligand. In some embodiments, a suspended cell may be adhered to a surface by interaction with a specific carbohydrate binding interaction, a specific protein or peptide binding interaction, or a specific lipid-lipid interaction, lipid-peptide interaction, or lipid-carbohydrate interaction. In some embodiments, a suspended cell may be adhered to a surface by a nonspecific interaction with said surface, such as by use of a charged surface (e.g., a polylysine, poly argininine, polyglutamic acid, polyaspartic acid surface or the like, or a charged polymer surface, such as a polyethylenimine surface; or a plasma-treated or ion-treated glass or polystyrene surface, or the like). It will be understood by one of skill in the art that in addition to surfaces disclosed herein, any surface useful for, or conventionally used for, cell culture, will be useful for capture of adherent cells. In particular embodiments, a surface useful for capture of adherent cells will comprise at least one of polyethylene oxide, streptavidin, protein A, or any combination thereof. In some embodiments, a suspended cell may be introduced to a flow cell by flow through the flow cell, by direct pipetting or liquid transfer onto a surface of the flow cell, by gravitational precipitation, by centrifugation, or by any method known in the art for bringing cells into contact with a surface.
[0212] In some embodiments, the one or more samples include target analyte(s) that are located inside the sample(s) or on the membrane of the sample(s). In some embodiments, the one or more samples include target analyte(s) that are on the exterior or interior surface of the cell. In some embodiments, the one or more samples include target analyte(s) that are on the exterior or interior surface of the cell membrane. In some embodiments, In some embodiments, the one or more samples include target analyte(s) that are part of the extracellular matrix. In some embodiments, the one or more samples include target analyte(s) that are part of and / or located on one or more organelles within the cell or tissue. In some embodiments, the one or more samples include target analytes that are on or in the glycocalyx or belong to part of the glycocalyx.
[0213] In some embodiments, the target analyte(s) comprise at least one polypeptide, lipid, nucleic acid or polysaccharide. In some embodiments, the target analyte(s)comprise at least one polypeptide, enzyme or lipid located anywhere in the sample(s) including the cytoplasm and nucleus. In some embodiments, the target analyte(s) comprise at least one polypeptide, enzyme or lipid located in or on a cellular structure including without limits any cellular membrane, nucleus, nucleolus, mitochondria, chloroplast, Golgi apparatus, ribosome, endoplasmic reticulum, microtubules, peroxisome and lysosome.
[0214] The methods, devices, and systems disclosed herein allow sequencing and analysis of various samples and sources. The samples may include nucleic acids extracted from any of a variety of samples, e.g., blood samples, saliva samples, urine samples, cell samples, tissue samples, and the like. In some embodiments, the samples here may include a variety of different cell, tissue, or sample types known to those of skill in the art. For example, the sample(s) may be from eukaryotes (such as animals, plants, fungi, protista), archaebacteria, or eubacteria. In some embodiments, the sample(s) may include prokaryotic or eukaryotic cells, such as adherent or non-adherent eukaryotic cells. In some embodiments, the sample(s) may be from, for example, primary or immortalized rodent, porcine, feline, canine, bovine, equine, primate, or human cell lines. In some embodiments, the sample(s) may include a variety of different cell, organ, or tissue types (e.g., white blood cells, red blood cells, platelets, epithelial cells, endothelial cells, neurons, glial cells, astrocytes, fibroblasts, skeletal muscle cells, smooth muscle cells, gametes, or cells from the heart, lungs, brain, liver, kidney, spleen, pancreas, thymus, bladder, stomach, colon, or small intestine). In some embodiments, the sample(s) may include normal or healthy cells. Alternately or in combination, the sample(s) may include diseased cells, such as cancerous cells, or from pathogenic cells that are infecting a host. In some embodiments, the sample(s) may include a distinct subset of cell types, e.g., immune cells (such as T cells, cytotoxic (killer) T cells, helper T cells, alpha beta T cells, gamma delta T cells, T cell progenitors, B cells, B-cell progenitors, lymphoid stem cells, myeloid progenitor cells, lymphocytes, granulocytes, Natural Killer cells, plasma cells, memory cells, neutrophils, eosinophils, basophils, mast cells, monocytes, dendritic cells, and / or macrophages, or any combination thereof), undifferentiated human stem cells, human stem cells that have been induced to differentiate, rare cells (e.g., circulating tumor cells (CTCs), circulating epithelial cells, circulating endothelial cells, circulating endometrial cells, bone marrow cells, progenitor cells, foam cells, mesenchymal cells, or trophoblasts). Other cells are contemplated and consistent with the disclosure herein.Computer systems
[0215] Various aspects of the method 200 and 300 may be implemented, for example, using one or more computer systems, such as computer system 400 shown in FIG. 4. One or more computer systems 400 may be used, for example, to implement any of the aspects discussed herein, as well as combinations and sub-combinations thereof.
[0216] Computer system 400 may include one or more hardware processors 404. The hardware processor 404 may be central processing unit (CPU), graphic processing units (GPU), or their combination. Processor 404 may be connected to a bus or communication infrastructure 406.
[0217] Computer system 400 may also include user input / output device(s) 403, such as monitors, keyboards, pointing devices, etc., which may communicate with communication infrastructure 406 through user input / output interface(s) 402. The user input / output devices 403 may be coupled to the user interface 124 in FIG. 1.
[0218] One or more units of processors 404 may be a graphics processing unit (GPU). In an aspect, a GPU may be a processor that is a specialized electronic circuit designed to process mathematically intensive applications. The GPU may have a parallel structure that is efficient for parallel processing of large blocks of data, such as mathematically intensive data common to computer graphics applications, images, videos, vector processing, array processing, etc., as well as cryptography (including brute-force cracking), generating cryptographic hashes or hash sequences, solving partial hashinversion problems, and / or producing results of other proof-of-work computations for some blockchain-based applications, for example. With capabilities of general-purpose computing on graphics processing units (GPGPU), the GPU may be particularly useful in at least the image recognition and machine learning aspects described herein.
[0219] Additionally, one or more of processors 404 may include a coprocessor or other implementation of logic for accelerating cryptographic calculations or other specialized mathematical functions, including hardware-accelerated cryptographic coprocessors. Such accelerated processors may further include instruction set(s) for acceleration using coprocessors and / or other logic to facilitate such acceleration.
[0220] Computer system 400 may also include a data storage device such as a main or primary memory 408, e.g., random access memory (RAM). Main memory 408 may include one or more levels of cache. Main memory 408 may have stored therein control logic (i.e., computer software) and / or data.
[0221] Computer system 400 may also include one or more secondary data storage devices or secondary memory 410. Secondary memory 410 may include, for example, a main storage drive 412 and / or a removable storage device or drive 414. Main storage drive 412 may be a hard disk drive or solid-state drive, for example. Removable storage drive 414 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, tape backup device, and / or any other storage device / drive.
[0222] Removable storage drive 414 may interact with a removable storage unit 418.
[0223] Removable storage unit 418 may include a computer usable or readable storage device having stored thereon computer software and / or data. The software may include control logic. The software may include instructions executable by the hardware processor(s) 404. Removable storage unit 418 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / any other computer data storage device. Removable storage drive 414 may read from and / or write to removable storage unit 418.
[0224] Secondary memory 410 may include other means, devices, components, instrumentalities or other approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 400. Such means, devices, components, instrumentalities or other approaches may include, for example, a removable storage unit 422 and an interface 420. Examples of the removable storage unit 422 and the interface 420 may include a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other removable storage unit and associated interface.
[0225] Computer system 400 may further include a communication or network interface 424. Communication interface 424 may enable computer system 400 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referenced by reference number 428). For example, communication interface 424 may allow computer system 400 to communicate with external or remote devices 428 over communication path 426, which may be wired and / or wireless (or a combination thereof), and which may include any combination of LANs, WANs, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 400 via communication path 426. In some aspects, communication path 426 is the connection to the cloud 130, as depicted in FIG. 1. The external devices, etc.referred to by reference number 428 may be devices, networks, entities, etc. in the cloud 130.
[0226] Computer system 400 may also be any of a personal digital assistant (PDA), desktop workstation, laptop or notebook computer, netbook, tablet, smart phone, smart watch or other wearable, appliance, part of the Internet of Things (loT), and / or embedded system, to name a few non-limiting examples, or any combination thereof.
[0227] It should be appreciated that the framework described herein may be implemented as a method, process, apparatus, system, or article of manufacture such as a non-transitory computer-readable medium or device. For illustration purposes, the present framework may be described in the context of distributed ledgers being publicly available, or at least available to untrusted third parties. One example as a modem use case is with blockchainbased systems. It should be appreciated, however, that the present framework may also be applied in other settings where sensitive or confidential information may need to pass by or through hands of untrusted third parties, and that this technology is in no way limited to distributed ledgers or blockchain uses.
[0228] Computer system 400 may be a client or server, accessing or hosting any applications and / or data through any delivery paradigm, including but not limited to remote or distributed cloud computing solutions; local or on-premises software (e.g., “onpremise” cloud-based solutions); “as a service” models (e.g., content as a service (CaaS), digital content as a service (DCaaS), software as a service (SaaS), managed software as a service (MSaaS), platform as a service (PaaS), desktop as a service (DaaS), framework as a service (FaaS), backend as a service (BaaS), mobile backend as a service (MBaaS), infrastructure as a service (laaS), database as a service (DBaaS), etc.); and / or a hybrid model including any combination of the foregoing examples or other services or delivery paradigms.
[0229] Any applicable data structures, file formats, and schemas may be derived from standards including but not limited to JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other functionally similar representations alone or in combination. Alternatively, proprietary data structures, formats or schemas may be used, either exclusively or in combination with known or open standards.
[0230] Any pertinent data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in human-readable formats such as numeric, textual, graphic, or multimedia formats, further including various types of markup language, among other possible formats. Alternatively or in combination with the above formats, the data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in binary, encoded, compressed, and / or encrypted formats, or any other machine-readable formats.
[0231] Interfacing or interconnection among various systems and layers may employ any number of mechanisms, such as any number of protocols, programmatic frameworks, floorplans, or application programming interfaces (API), including but not limited to Document Object Model (DOM), Discovery Service (DS), NSUserDefaults, Web Services Description Language (WSDL), Message Exchange Pattern (MEP), Web Distributed Data Exchange (WDDX), Web Hypertext Application Technology Working Group (WHATWG) HTML5 Web Messaging, Representational State Transfer (REST or RESTful web services), Extensible User Interface Protocol (XUP), Simple Object Access Protocol (SOAP), XML Schema Definition (XSD), XML Remote Procedure Call (XML- RPC), or any other mechanisms, open or proprietary, that may achieve similar functionality and results.
[0232] Such interfacing or interconnection may also make use of uniform resource identifiers (URI), which may further include uniform resource locators (URL) or uniform resource names (URN). Other forms of uniform and / or unique identifiers, locators, or names may be used, either exclusively or in combination with forms such as those set forth above.
[0233] Any of the above protocols or APIs may interface with or be implemented in any programming language, procedural, functional, or object-oriented, and may be compiled or interpreted. Non-limiting examples include C, C++, C#, Objective-C, Java, Scala, Clojure, Elixir, Swift, Go, Perl, PHP, Python, Ruby, JavaScript, WebAssembly, or virtually any other language, with any other libraries or schemas, in any kind of framework, runtime environment, virtual machine, interpreter, stack, engine, or similar mechanism, including but not limited to Node.js, V8, Knockout, j Query, Dojo, Dijit, OpenUI5, AngularJS, Expressjs, Backbone) s, Ember js, DHTMLX, Vue, React, Electron, and so on, among many other non-limiting examples.
[0234] In some aspects, a tangible, non-transitory apparatus or article of manufacture comprising a tangible, non-transitory computer useable or readable medium havingcontrol logic (software) stored thereon may also be referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 400, main memory 408, secondary memory 410, and removable storage units 418 and 422, as well as tangible articles of manufacture embodying any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 400), may cause such data processing devices to operate as described herein.
[0235] Based on the teachings contained in this disclosure, it will be apparent to persons skilled in the relevant art(s) how to make and use aspects of this disclosure using data processing devices, computer systems and / or computer architectures other than that shown in FIG. 4. In particular, aspects may operate with software, hardware, and / or operating system implementations other than those described herein.Methods for conducting in situ short read sequencing
[0236] In the methods described herein, the RNA is not extracted from the sample(s) and sequencing information does not need to be tracked and mapped back to an image of the sample(s). Rather, RNA is retained inside the sample(s) to permit direct imaging of the spatial location of target RNAs within the cells. Additionally, RNA within the sample(s) is not fragmented and enrichment of target RNA is not necessary. Use of target-specific and / or random-sequence reverse transcription primers enables detection of both poly-A and non-poly-A RNAs in either uni-plex or multi-plex modes.
[0237] In some embodiments, the methods comprise repeatedly conducting a short number of sequencing cycles of the same region of the template molecules (e.g., concatemer molecules). By conducting reiterative short sequencing cycles, the RNA content of the sample(s) can be discovered. Compared to long read sequencing workflows, the reiterative short sequencing cycles described herein use a reduced amount of sequencing reagents which reduces cost and saves time. Methods for conducting reiterative short sequencing cycles has many uses including but not limited to detecting specific RNAs of interest, mutant RNA sequences, splice variants, and their abundance levels thereof.
[0238] The concatemers carry tandem repeat units of a cDNA-of-interest, the universal sequencing primer binding site, and the target barcode sequence. The concatemers are sequenced inside the sample(s) where a short number of sequencing cycles are conductedfor each round and multiple rounds of short read sequencing is conducted. The full length of the target barcode and cDNA region are not sequenced. Instead, at least a portion of the target barcode region is reiteratively sequenced. In some embodiments, it is not necessary to sequence the cDNA region. In some embodiments, the target barcode and a portion of the cDNA region are reiteratively sequenced. It is not necessary to sequence the entire length of the cDNA region. It is not necessary to assemble the sequencing reads or to obtain a full length sequence of the cDNAs-of-interest. The redundant sequencing information obtained from the short sequencing reads obviates the need to sequence the complementary strand of the concatemer. Thus pairwise sequencing is not necessary.
[0239] Additionally, a short portion of the cDNA region in the concatemer is resequenced at least once (e.g., reiterative sequencing) from the same start position to generate overlapping sequencing reads that can be aligned to a reference sequence. For example, the same portion of the concatemer molecule can be sequenced at least two, three, four, five, or up to 50 times. The start sequencing site can be any location of the concatemer and is dictated by the sequencing primers which are designed to anneal to a selected position within the concatemer. The reiterative short sequencing reads increase the redundancy of sequencing information for individual bases in the cDNA region. Reiteratively sequencing one strand of the concatemer template molecule provides enough base coverage to reveal the presence of target RNAs in the sample(s) so that pairwise sequencing of the complementary strand is not necessary.
[0240] A concatemer template molecule includes multiple sequencing primer binding sites along the same concatemer molecule which can be used to generate multiple usable sequencing reads for increased sequencing depth. Together, reiteratively sequencing one strand of the concatemer templates increases sequencing base coverage and sequencing depth compared to sequencing a one-copy template molecule.
[0241] The methods described herein can be conducted in uni-plex or multi-plex modes.Two or more different target RNAs can be detected and imaged simultaneously inside a cellular sample using different reverse transcription primers, different target-specific padlock probes, and universal sequencing primers. For example, the presence of a housekeeping RNA and at least one target RNA in a cellular sample can be simultaneously detected and imaged using any of the reiterative short read sequencing methods described herein.
[0242] The present disclosure provides methods for detecting in situ at least two different target RNA molecules in a cellular sample comprising step (a): providing a cellular sample harboring a plurality of RNA which comprises at least a first target RNA molecule and a second target RNA molecule. In some embodiments, the sample(s) is fixed and permeabilized. In some embodiments, the sample(s) harbors 2-25 different target RNA molecules, or harbors 25-50 different target RNA molecules, or harbors SO- 75 different target RNA molecules, or harbors 75-100 different target RNA molecules. In some embodiments, the sample(s) harbors more than 100 different target RNA molecules, or more than 250 different target RNA molecules, or more than 500 different target molecules, or more than 1000 different target RNA molecules, or more. In some embodiments, the sample(s) harbors more than 10,000 different target RNA molecules. In some embodiments, the sample(s) comprises a whole cell, a plurality of whole cells, an intact tissue or an intact tumor. In some embodiments, the sample(s) comprises a fresh cellular sample, a freshly-frozen cellular sample, a sectioned cellular sample, an FFPE cellular sample, or a sectioned FFPE cellular sample. In some embodiments, the sample(s) is deposited onto a solid support. In some embodiments, the sample(s) is deposited onto a solid support which is passivated with a coating that promotes cell adhesion. In some embodiments, the sample(s) is deposited on a support that lacks immobilized capture oligonucleotides. In some embodiments, the sample(s) is cultured before or after depositing the sample(s) onto the solid support. In some embodiments, the sample(s) is cultured prior to conducting step (b) which is described below. In some embodiments, the sample(s) comprises an expanded cellular sample that has been cultured in a simple or complex cell culture media. In some embodiments, the sample(s) is not cultured or expanded prior to conducting step (b).
[0243] In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprise step (b): generating inside the sample(s) a plurality of cDNA molecules which include at least a first target cDNA molecule that corresponds to the first target RNA molecule, and the plurality of cDNA molecules includes a second target cDNA molecule that corresponds to the second target RNA molecule. In some embodiments, the method comprises generating at least 2-10,000 different target cDNA molecules that correspond to 2-10,000 different target RNA molecules. In some embodiments, the generating of step (b) comprises contacting the plurality of RNA inside the sample(s) with (i) a plurality of reverse transcription primers,(ii) a plurality of reverse transcriptase enzymes, and (iii) a plurality of nucleotides, under a condition suitable for conducting a reverse transcription reaction to generate a plurality of cDNA molecules (e.g., a plurality of first strand cDNA molecules) in the sample(s) (e g., FIG. 7).
[0244] In some embodiments, the plurality of reverse transcription primers comprises a first sub-population of target-specific reverse transcription primers that hybridize selectively to the first target RNA, and comprises a second sub-population of targetspecific reverse transcription primers that hybridize selectively to the second target RNA. In some embodiments, the first and second sub-population of target-specific reverse transcription primers have the same sequence or different sequences.
[0245] In some embodiments, the entire length of the first sub-population of targetspecific reverse transcription primers hybridize to a first target RNA molecule. In some embodiments, the first sub-population of target-specific reverse transcription primers comprise tailed primers having a portion that hybridizes to a first target RNA molecule and a portion that does not hybridize to a first target RNA molecule. In some embodiments, the first sub-population of target-specific reverse transcription primers comprise at least a portion having a poly-T sequence. In some embodiments, the first subpopulation of target-specific reverse transcription primers comprise at least a portion having a random sequence and / or at least a portion having a target-specific sequence.
[0246] In some embodiments, the entire length of the second sub-population of targetspecific reverse transcription primers hybridize to a second target RNA molecule. In some embodiments, the second sub-population of target-specific reverse transcription primers comprise tailed primers having a portion that hybridizes to a second target RNA molecule and a portion that does not hybridize to a second target RNA molecule. In some embodiments, the second sub-population of target-specific reverse transcription primers comprise at least a portion having a poly-T sequence. In some embodiments, the second sub-population of target-specific reverse transcription primers comprise at least a portion having a random sequence and / or at least a portion having a target-specific sequence.
[0247] In some embodiments, a target RNA molecule that is hybridized to a cDNA molecule can be subjected to enzymatic degradation using a ribonuclease under a condition suitable for degrading RNA in an RNA / DNA duplex. In some embodiments, a target RNA molecule that is hybridized to a cDNA molecule is not subjected to enzymatic degradation.
[0248] In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprise step (c): contacting the plurality of cDNA molecules in the sample(s) with a plurality of target-specific padlock probes which includes at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes. In some embodiments, the method comprises contacting the plurality of cDNA molecule in the sample(s) with at least 2-10,000 different targetspecific padlock probes.
[0249] In an alternative embodiment, cDNA is not generated from RNA inside the sample(s). In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprise contacting RNA inside the cell with a plurality of target-specific padlock probes and generating circularized padlock probes. In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprise step (c): contacting the plurality of RNA molecules in the sample(s) with a plurality of target-specific padlock probes which includes at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes. In some embodiments, the method comprises contacting the plurality of cDNA molecule in the sample(s) with at least 2-10,000 different target-specific padlock probes. In some embodiments, a target RNA molecule can be subjected to enzymatic degradation using a ribonuclease. In some embodiments, a target RNA molecule is not subjected to enzymatic degradation.
[0250] In some embodiments, individual padlock probes in the plurality of first targetspecific padlock probes comprise first and second terminal regions (e.g., first and second padlock binding arms), wherein the first terminal region selectively hybridizes to a first region of the first target cDNA molecule (or the first target RNA molecule), and the second terminal region selectively hybridizes to a second region of the first target cDNA molecule (or the first target RNA molecule). In some embodiments, the contacting of step (c) comprises: hybridizing the first and second terminal regions of the first target-specific padlock probes to proximal positions on the first target cDNA molecule (or the first target RNA molecule) to form a circularized first target-specific padlock probe having a nick or gap between the hybridized first and second terminal regions (e.g., FIG. 7, left). In some embodiments, the first target-specific padlock probe comprises a first target barcode sequence (target BC-1) that corresponds to and uniquely identifies the first target cDNA sequence (or the first target RNA sequence). In some embodiments, the first target-specific padlock probe comprises a first target barcode sequence that is located adjacent to one of the regions of the first target-specific padlock probe that selectively hybridizes to the first target cDNA molecule (or the first target RNA sequence). In some embodiments, the first target-specific padlock probe comprises at least one universal adaptor sequence, such as for example a universal sequencing primer binding site (or a complementary sequence thereof). In some embodiments, the first target-specific padlock probe comprises a universal primer binding site for a rolling circle amplification primer (or a complementary sequence thereof). In some embodiments, the first target-specific padlock probe comprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof).
[0251] In some embodiments, individual padlock probes in the plurality of second targetspecific padlock probes comprise first and second terminal regions (e.g., first and second padlock binding arms), wherein the first terminal region selectively hybridizes to a first region of the second target cDNA molecule (or the second target RNA molecule), and the second terminal region selectively hybridizes to a second region of the second target cDNA molecule (or the second target RNA molecule). In some embodiments, the contacting of step (c) comprises: hybridizing the first and second terminal regions of the second target-specific padlock probes to proximal positions on the second target cDNA molecule (or the second target RNA molecule) to form a circularized second targetspecific padlock probe having a nick or gap between the hybridized first and second terminal regions (e.g., FIG. 7, right). In some embodiments, the second target-specific padlock probe comprises a second target barcode sequence (target BC-2) that corresponds to and uniquely identifies the second target cDNA sequence (or the second target RNA sequence). In some embodiments, the second target-specific padlock probe comprises a second target barcode sequence that is located adjacent to one of the regions of the second target-specific padlock probe that selectively hybridizes to the second target cDNA molecule (or the second target RNA sequence). In some embodiments, the second targetspecific padlock probe comprises at least one universal adaptor sequence, such as for example a universal sequencing primer binding site (or a complementary sequence thereof). In some embodiments, the second target-specific padlock probe comprises a universal primer binding site for a rolling circle amplification primer (or a complementary sequence thereof). In some embodiments, the second target-specific padlock probecomprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof).
[0252] In some embodiments, the first target barcode sequence (target BC-1) and the second target barcode sequence (target BC-2) have different sequences and can be used to conduct multiplex RNA detection and sequencing. In some embodiments, the first target barcode sequence (target BC-1) and the second target barcode sequence (target BC-2) have the same sequence and can be used to conduct uni-plex RNA detection and sequencing.
[0253] In some embodiments, the first and second target-specific padlock probes comprise a universal sequencing primer binding site and a target barcode sequence that are adjacent to each other so that the target barcode region of the concatemer is sequenced first. The target barcode sequence can be any length, for example 3-15 bases, or 15-25 bases, or 25-40 bases, or longer.
[0254] In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprising step (d): closing the nick or gap in the at least first and second circularized target-specific padlock probes by conducting an enzymatic reaction, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe inside the sample(s). In some embodiments, the closing the nick in the first and second circularized padlock probes comprises conducting an enzymatic ligation reaction. In some embodiments, closing the gap in the first and second circularized padlock probes comprises conducting a polymerase-catalyzed fill-in reaction using the first or second target cDNA molecule (or the first or second RNA molecule) as a template, and conducting an enzymatic ligation reaction. In some embodiments, the method comprises closing the nick or gap in at least 2-10,000 circularized target-specific padlock probes by conducting one or more enzymatic reactions, thereby generating at least 2-10,000 covalently closed circular padlock probes inside the sample(s).
[0255] In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprising step (e): conducting a rolling circle amplification reaction inside the sample(s) using the first and second covalently closed circular padlock probes as template molecules, thereby generating a plurality of concatemer molecules including at least a first concatemer molecule that corresponds to a first target RNA molecule, and the plurality of concatemer molecules includes at least asecond concatemer molecule that corresponds to a second target RNA molecule. In some embodiments, the first concatemer molecule comprises tandem repeat units, wherein a unit comprises a sequence that corresponds to the first target cDNA (or the first target RNA), the first target barcode sequence, and the universal sequencing primer binding site (or a complementary sequence thereof). In some embodiments, the second concatemer molecule comprises tandem repeat units, wherein a unit comprises a sequence that corresponds to the second target cDNA (or the second target RNA), the second target barcode sequence, and the universal sequencing primer binding site (or a complementary sequence thereof).
[0256] In some embodiments, the rolling circle amplification reaction of step (e) comprises contacting the covalently closed circularized padlock probes with an amplification primer (e.g., a universal rolling circle amplification primer), a stranddisplacing DNA polymerase, and a plurality of nucleotides, under a condition suitable for hybridizing individual amplification primers to a covalently closed padlock probe, and under a condition suitable for conducting primer extension using the covalently closed padlock probe as a template molecule to generate a nucleic acid concatemer. In some embodiments, the method comprises conducting a rolling circle amplification reaction inside the sample(s) using the at least 2-10,000 covalently closed circular padlock probes as template molecules, thereby generating at least 2-10,000 concatemer molecules that correspond to at least 2-10,000 target RNA molecules. In some embodiments, the plurality of concatemers that are generated inside the sample(s) collapse into a DNA nanoball having a shape and size that is more compact compared to a non-collapsed concatemer.
[0257] In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprising step (f): sequencing the plurality of concatemer molecules inside the sample(s), which comprises sequencing the first concatemer molecule by conducting no more than 2-30 sequencing cycles to generate a plurality of first sequencing read products, and sequencing the second concatemer molecule by conducting no more than 2-30 sequencing cycles to generate a plurality of second sequencing read products (FIG. 8). In some embodiments, the sequencing of step (f) comprises sequencing no more than 2-30 bases of the first concatemer molecules to generate a plurality of first sequencing read products, and which comprises sequencing no more than 2-30 bases of the second concatemer molecules to generate a plurality ofsecond sequencing read products. In some embodiments, the method comprises sequencing the at least 2-10,000 concatemer molecules inside the sample(s), which comprises conducting no more than 2-30 sequencing cycles on the 2-10,000 concatemer molecules to generate a plurality of sequencing read products.
[0258] In some embodiments, only the first target barcode region of the first concatemer molecules are sequenced (e.g., FIG. 8, top). In some embodiments, at least a portion or the full length of the first target barcode of the first concatemer molecules are sequenced (e.g., FIG. 8, top). In some embodiments, the first target barcode is sequenced and a portion of the first cDNA region (or the first RNA region) of the first concatemer molecules are sequenced. In some embodiments, at least a portion of the first cDNA region (or the first RNA region) of the first concatemer molecules are sequenced.
[0259] In some embodiments, only the second target barcode region of the second concatemer molecules are sequenced (e.g., FIG. 8, bottom). In some embodiments, at least a portion or the full length of the second target barcode of the second concatemer molecules are sequenced (e.g., FIG. 8, bottom). In some embodiments, the second target barcode is sequenced and a portion of the second cDNA region (or the second RNA region) of the second concatemer molecules are sequenced. In some embodiments, at least a portion of the second cDNA region (or the second RNA region) of the second concatemer molecules are sequenced.
[0260] In some embodiments, the sequencing of step (f) comprises contacting the plurality of concatemer molecules inside the sample(s) with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents, under a condition suitable for hybridizing the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the concatemers. In some embodiments, the sequencing of step (f) further comprises conducting no more than 2-30 sequencing cycles to generate at least a first plurality of sequencing read products by sequencing at least the first target barcode region (Target BC-1), and optionally conducting no more than 2-30 sequencing cycles to generate at least a second plurality of sequencing read products by sequencing at least the second target barcode region (Target BC-2). In some embodiments, the nucleotide reagents comprise multivalent molecules, nucleotides and / or nucleotide analogs.
[0261] In some embodiments, the sequencing of step (f) comprises sequencing at least a portion of the first and second nucleic acid concatemers using an optical imaging system comprising a field-of-view (FOV) greater than 1.0 mm2.
[0262] In some embodiments, in the sequencing of step (f), the plurality of first and second sequencing read products are detectable by imaging, and wherein the sequencing comprises decoding the plurality of first and second sequencing read products from the images obtained during the no more than 2-30 sequencing cycles.
[0263] In some embodiments, in the sequencing of step (f), the plurality of the first and second sequencing read products are detectable by imaging, and wherein the sequencing comprises simultaneously imaging the plurality of first and second detectable sequencing read products in the sample(s) (co-localization of the first and second sequencing read products).
[0264] In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprising step (g): removing the plurality of first sequencing read products from the first concatemer molecules and retaining the first concatemer molecules in the sample(s), and removing the plurality of second sequencing read products from the second concatemer molecules and retaining the second concatemer molecules in the sample(s).
[0265] In some embodiments, methods for detecting at least two different target RNA molecules in a cellular sample further comprising step (h): reiteratively sequencing the plurality of concatemers by repeating steps (f) and (g) at least once, wherein the sequences of the plurality of first sequencing read products confirms the presence of the first target RNA molecules in the sample(s), and wherein the sequences of the plurality of second sequencing read products confirms the presence of the second target RNA molecules in the sample(s).
[0266] In some embodiments, reiteratively sequencing at least one region of the concatemer comprises repeating steps (f) - (g) at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, or at least 10 times.
[0267] In some embodiments, reiteratively sequencing at least one region of the concatemer comprises repeating steps (f) - (g) up to 10 times, up to 20 times, up to 30 time, up to 40 times, or up to 50 times. An example of reiterative sequence is shown in a schematic in FIG. 9-12.
[0268] In some embodiments, e.g., in FIG. 9, the concatemer includes tandem repeat units where each unit comprises: (i) a universal sequencing primer binding site (Seq), (ii) universal compaction oligonucleotide binding site (CO), (iii) an insert sequence that corresponds to a given target cDNA, and (iv) a target barcode sequence that corresponds to the given target cDNA (BC). In some embodiments, universal sequencing primers (solid arrows) hybridize to the universal sequencing primer binding sites and no more than 30 sequencing cycles are conducted to generate a plurality of first sequencing read products (dashed arrows), where the first sequencing read products include only the target barcode sequence. The plurality of first sequencing read products are removed from the concatemer, and the sequencing is repeated where no more than 30 sequencing cycles are conducted to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products include only the target barcode sequence. The plurality of first sequencing read products are removed from the concatemer, and the sequencing is once again repeated where no more than 30 sequencing cycles are conducted to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products include only the target barcode sequence. In some embodiments, the reiterative sequencing can be conducted up to 50 times. The sequences of all of the first sequencing read products can be determined and aligned with a first reference sequence (e.g., reference barcode sequence) to confirm the presence of the first target RNA molecules inside the sample(s).
[0269] In some embodiments, e.g., in FIG. 10, the concatemer includes tandem repeat units where each unit comprises: (i) a universal sequencing primer binding site (Seq), (ii) universal compaction oligonucleotide binding site (CO), (iii) an insert sequence that corresponds to a given target cDNA, and (iv) a target barcode sequence that corresponds to the given target cDNA (BC). In some embodiments, universal sequencing primers (solid arrows) hybridize to the universal sequencing primer binding sites and no more than 30 sequencing cycles are conducted to generate a plurality of first sequencing read products (dashed arrows), where the first sequencing read products include the target barcode sequence and a portion of the insert sequence. The plurality of first sequencing read products are removed from the concatemer, and the sequencing is repeated where no more than 30 sequencing cycles are conducted to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products include the target barcode sequence and a portion of the insert sequence. The plurality offirst sequencing read products are removed from the concatemer, and the sequencing is once again repeated where no more than 30 sequencing cycles are conducted to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products include the target barcode sequence and a portion of the insert sequence. In some embodiments, the reiterative sequencing can be conducted up to 50 times. The sequences of all of the first sequencing read products can be determined and aligned with a first reference sequence (e.g., reference barcode sequence and the insert sequence that corresponds to the target RNA) to confirm the presence of the first target RNA molecules inside the sample(s).
[0270] In some embodiments, e.g., in FIG. 11, the concatemer includes tandem repeat units where each unit comprises: (i) a universal sequencing primer binding site (Seq), (ii) universal compaction oligonucleotide binding site (CO), and (iii) an insert sequence that corresponds to a given target cDNA. In some embodiments, universal sequencing primers (solid arrows) hybridize to the universal sequencing primer binding sites and no more than 30 sequencing cycles are conducted to generate a plurality of first sequencing read products (dashed arrows), where the first sequencing read products include a portion of the insert sequence. The plurality of first sequencing read products are removed from the concatemer, and the sequencing is repeated where no more than 30 sequencing cycles are conducted to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products include a portion of the insert sequence. The plurality of first sequencing read products are removed from the concatemer, and the sequencing is once again repeated where no more than 30 sequencing cycles are conducted to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products include a portion of the insert sequence. In some embodiments, the reiterative sequencing can be conducted up to 50 times. The sequences of all of the first sequencing read products can be determined and aligned with a first reference sequence (e.g., the insert sequence that corresponds to the target RNA) to confirm the presence of the first target RNA molecules inside the sample(s).
[0271] In some embodiments, e.g., in FIG. 12, the concatemer includes tandem repeat units where each unit comprises: (i) a universal sequencing primer binding site (Seq) and (ii) an insert sequence that corresponds to a given target cDNA. In some embodiments, universal sequencing primers (solid arrows) hybridize to the universal sequencing primer binding sites and no more than 30 sequencing cycles are conducted to generate a pluralityof first sequencing read products (dashed arrows), where the first sequencing read products include a portion of the insert sequence. The plurality of first sequencing read products are removed from the concatemer, and the sequencing is repeated where no more than 30 sequencing cycles are conducted to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products include a portion of the insert sequence. The plurality of first sequencing read products are removed from the concatemer, and the sequencing is once again repeated where no more than 30 sequencing cycles are conducted to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products include a portion of the insert sequence. In some embodiments, the reiterative sequencing can be conducted up to 50 times. The sequences of all of the first sequencing read products can be determined and aligned with a first reference sequence (e.g., the insert sequence that corresponds to the target RNA) to confirm the presence of the first target RNA molecules inside the sample(s).
[0272] In some embodiments, at least one concatemer is sequenced by conducting step (f) once (non-reiterative sequencing). In some embodiments, at least one concatemer is sequenced by conducting steps (f) - (g) once. In some embodiments, at least one concatemer is reiteratively sequenced by conducting steps (f) - (g) at least twice.
[0273] In some embodiments, the plurality of universal sequencing primers can be hybridized to concatemer template molecules with a hybridization reagent comprising an SSC buffer (e.g., 2X saline-sodium citrate) buffer with formamide (e.g., 10-20% formamide). The hybridization conditions comprise a temperature of about 20-30 °C, for about 10-60 minutes.
[0274] In some embodiments, the plurality of sequencing read products can be removed from the concatemers and the plurality of concatemers can be retained inside the sample(s) using a de-hybridization reagent comprising an SSC buffer (e.g., saline-sodium citrate) buffer, with or without formamide, at a temperature that promotes nucleic acid denaturation such as for example 30 - 90 °C.
[0275] In some embodiments, the plurality of nucleotide reagents of step (f) comprise a plurality of nucleotides that are detectably labeled or non-labeled. In some embodiments, individual nucleotides are linked to a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the plurality of detectably labeled nucleotide analogs comprise a plurality of chainterminating nucleotides, where the chain terminating moiety is linked to the 3’ nucleotide sugar position to form a 3’ blocked nucleotide analog. In some embodiments, the chain terminating moiety can be removed to convert the 3’ blocked nucleotide analog to an extendible nucleotide having a 3’ OH group on the sugar. In some embodiments, the labeled nucleotide analogs are linked to a different fluorophore that corresponds to the nucleo-bases adenine, cytosine, guanine, thymine or uracil, where the different fluorophores emit a fluorescent signal during the sequencing of step (f). In some embodiments, a sequencing cycle comprises (1) contacting the concatemer / sequencing primer duplex with a sequencing polymerase and a detectably labeled chain terminating nucleotide under a condition suitable for polymerase-catalyzed incorporation of the detectably labeled chain terminating nucleotide into the terminal end of the sequencing primer, (2) detecting and imaging the fluorescent signal and color emitted by the incorporated chain terminating nucleotide, and (3) removing the chain terminating moiety (e.g., unblocking) and the fluorophore from the incorporated nucleotide and retaining the concatemer / sequencing primer duplex. In some embodiments, no more than 2-30 sequencing cycles are conducted on the plurality of concatemers inside the sample(s) to generate a plurality of sequencing read products. In some embodiments, the sequence of the first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of the first target RNA molecules inside the sample(s). In some embodiments, the sequence of the second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of the second target RNA molecules inside the sample(s).
[0276] In some embodiments, the sequences of the first and second sequencing read products can be aligned after each round of generating the first and second sequencing read products which are no more than 30 bases in length, or after generating a set of reiterative sequencing read products wherein the first and second sequencing read products which are no more than 30 bases in length. In some embodiments, the sequencing reactions are conducted on a sequencing apparatus having a detector that captures fluorescent signals from the sequencing reactions inside the sample(s). The sequencing apparatus can be configured to relay the fluorescent signal data captured by the detector to a computer system that is programmed to display images of different fluorescent spots which are co-located in the sample(s), where individual fluorescent spots correspond to different target RNA molecules. In some embodiments, when thesequencing is conducted using different fluorescently-labeled nucleotide reagents that correspond to different nucleo-bases (e.g., A, G, C, T / U), then the images can have different color fluorescent spots co-located in the same cellular sample at different sequencing cycles.
[0277] In some embodiments, out-of-sync phasing and / or pre-phasing events can occur during synchronized sequencing reactions on clonally amplified template amplicons, where the sequencing reactions comprise polymerase-catalyzed sequencing reactions employing detectably labeled chain terminator nucleotides. In some embodiments, a sequencing reaction on one template molecule in the clonally-amplified template molecules moves ahead (e.g., pre-phasing) or fall behind (e.g., phasing) of the sequencing of the other template molecules within the clonally-amplified template molecules. During sequencing, a fluorescent signal is typically detected which corresponds to incorporation of a labeled chain terminator nucleotide. Thus, phasing and pre-phasing events can be detected and monitored using incorporation of a labeled chain terminator nucleotide.
[0278] In some embodiments, the plurality of nucleotide reagents of step (f) comprise a plurality of multivalent molecules each comprising a core attached to a plurality of nucleotide-arms, wherein the nucleotide-arms are attached to a nucleotide unit. In some embodiments, individual multivalent molecules are labeled with a detectably reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the core of the multivalent molecule is labeled with a fluorophore, and wherein the fluorophore which is attached to a given core of the multivalent molecule corresponds to the nucleotide base (e.g., adenine, guanine, cytosine, thymine or uracil) of the nucleotide arm. In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or nucleotide base that is attached to a fluorophore, and wherein the fluorophore which is attached to a given nucleotide base corresponds to the nucleotide base (e.g., adenine, guanine, cytosine, thymine or uracil) of the nucleotide arm. In some embodiments, a sequencing cycle comprises (1) contacting the concatemer / sequencing primer duplex with a first sequencing polymerase to form a complexed polymerase, (2) contacting the complexed polymerase with a detectably labeled multivalent molecule under a condition suitable for binding a complementary nucleotide unit of the multivalent molecule to the complexed polymerase thereby forming a multivalent-binding complex, and the condition is suitable for inhibiting incorporation of the complementary nucleotide unit into the terminal end of the sequencing primer, (3)detecting and imaging the fluorescent signal and color emitted by the bound detectably labeled multivalent molecule, (4) removing the first sequencing polymerase and the bound detectably labeled multivalent molecule, and retaining the concatemer / sequencing primer duplex, (5) contacting the retained concatemer / sequencing primer duplex with a second sequencing polymerase and a non-labeled chain terminating nucleotide under a condition suitable for polymerase-catalyzed incorporation of the non-labeled chain terminating nucleotide into the terminal end of the sequencing primer, and (6) removing the chain terminating moiety (e.g., unblocking) and retaining the concatemer / sequencing primer duplex. In some embodiments, no more than 2-30 sequencing cycles are conducted on the plurality of concatemers inside the sample(s) to generate a plurality of sequencing read products. In some embodiments, the sequence of the first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of the first target RNA molecules inside the sample(s). In some embodiments, the sequence of the second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of the second target RNA molecules inside the sample(s). In some embodiments, the sequences of the first and second sequencing read products can be aligned after each round of generating the first and second sequencing read products which are no more than 30 bases in length, or after generating a set of reiterative sequencing read products wherein the first and second sequencing read products which are no more than 30 bases in length. In some embodiments, the sequencing reactions are conducted on a sequencing apparatus having a detector that captures fluorescent signals from the sequencing reactions inside the sample(s). The sequencing apparatus can be configured to relay the fluorescent signal data captured by the detector to a computer system that is programmed to display images of different fluorescent spots which are co-located in the sample(s), where individual fluorescent spots correspond to different target RNA molecules. In some embodiments, individual cycle times can be achieved in less than 30 minutes. In some embodiments, the field of view (FOV) can exceed 1 mm2and the cycle time for scanning large area (> 10 mm2) can be less than 5 minutes.
[0279] In some embodiments, when sequencing with detectably labeled multivalent molecules, step (2) in which multivalent-binding complexes are formed and step (3) in which the bound detectably labeled multivalent molecules are imaged and detected, the conditions are gentle compared to sequencing workflows that employ detectable labeledchain terminating nucleotides. For example, steps (2) and (3) can be conducted at a gentle temperature of about 35 - 45 °C, or about 39 - 42 °C. Steps (2) and (3) can be conducted at a gentle temperature which can help retain the compact size and shape of a DNA nanoball during multiple sequencing cycles (e.g., up to 30 cycles) which can improve FWHM (full width half maximum) of a spot image of the DNA nanoball inside a cellular sample. In some embodiments, the DNA nanoball does not unravel during multiple sequencing cycles. In some embodiments, the spot image of the DNA nanoball does not enlarge during multiple sequencing cycles. In some embodiments, the spot image of the DNA nanoball remains a discrete spot during multiple sequencing cycles. The spot image can be represented as a Gaussian spot and the size can be measured as a FWHM. A smaller spot size as indicated by a smaller FWHM typically correlates with an improved image of the spot. In some embodiments, the FWHM of a nanoball spot can be about 10 um or smaller.
[0280] In some embodiments, out-of-sync phasing and / or pre-phasing events can occur during synchronized polymerase-catalyzed sequencing reactions employing detectably labeled multivalent molecules. During sequencing, a fluorescent signal can be detected which corresponds to binding of complementary nucleotide unit of a multivalent molecule to the complexed polymerase thereby forming a multivalent-binding complex. Thus, phasing and pre-phasing events can be detected and monitored using binding of labeled multivalent molecules. In some embodiments, when conducting up to 30 sequencing cycles with detectably labeled multivalent molecules, the phasing and / or prephasing rate can be less than about 5%, or less than about 1%, or less than about 0.01%, or less than about 0.001%. By contrast, the phasing and / or pre-phasing rates for conducting up to 30 sequencing cycles using labeled chain terminator nucleotides can be about 5%.Methods for conducting in situ RNA batch sequencing
[0281] The present disclosure provides methods for conducting in situ multiplex and multi-omics detection and identification using coded padlocks probes. The padlock probes are designed to selectively detect target RNA.
[0282] The RNA-specific padlock probes selectively hybridize to cDNA that corresponds to target RNA. The RNA-specific probes carry barcodes that uniquely identify the cDNA.In some embodiments, the RNA-specific padlock probes also carry batch-specific sequencing primer binding sites.
[0283] Both types of padlock probes are used to generate concatemers which having multiple copies of batch-specific sequencing binding sites and barcodes. The concatemers can collapse into DNA nanoballs having compact shape and size that produce increased signal intensity and color differentiation during sequencing.
[0284] For in situ sequencing, the limit of optical resolution impedes the ability to perform highly multiplex sequencing. The batch-specific sequencing primer binding sites on the padlock probes enables sequencing a desired subset (e.g., a batch) of the concatemers using selected batch-specific sequencing primers to reduce over-crowding signals and images. The use of batch-specific sequencing primers produces optical images that are intense and resolvable. By conducting multiple rounds of sequencing on the same cellular sample using different batch-specific sequencing primers enables multiplex sequencing to reveal numerous target RNAs.
[0285] The batch-specific sequencing methods described herein have many uses. For example, the number of spots that are imaged and associated with sequencing can be counted. The counted spots can be used as a measure of RNA levels in a cellular sample.
[0286] The present disclosure provides methods for detecting in situ at least two different target RNA molecules, comprising step (a): providing a cellular sample deposited on a solid support, wherein the sample(s) harbors (i) a first plurality of DNA amplicons (e.g., first concatemers) that correspond to a first target cDNA or RNA molecule, and (ii) a second plurality of DNA amplicons (e.g., second concatemers) that correspond to a second target cDNA or RNA molecule.
[0287] In some embodiments, the method further comprises step (b): sequencing the first plurality of DNA amplicons inside the sample(s) under a condition that inhibits sequencing the second plurality of DNA amplicons, wherein sequencing the first plurality of DNA amplicons inside the sample(s) comprises generating a plurality of first sequencing read products, wherein the sequences of the first sequencing read products are aligned with a first target reference sequence to confirm the presence of the first target RNA in the sample(s). In some embodiments, the first amplicons can be reiteratively sequenced by conducting no more than 2-30 sequencing cycles, or can be reiteratively sequenced by conducting 1-250 sequencing cycles.
[0288] In some embodiments, the method further comprises step (c): sequencing the second plurality of DNA amplicons inside the sample(s) under a condition that inhibits sequencing the first plurality of DNA amplicons, wherein sequencing the second plurality of DNA amplicons inside the sample(s) comprises generating a plurality of second sequencing read products, wherein the sequences of the second sequencing read products are aligned with a second target reference sequence to confirm the presence of the second target RNA in the sample(s). In some embodiments, the second amplicons can be reiteratively sequenced by conducting no more than 2-30 sequencing cycles, or can be reiteratively sequenced by conducting 1-250 sequencing cycles.
[0289] The present disclosure provides methods for detecting in situ at least two different target RNA molecules, comprising step (a): providing a cellular sample deposited on a solid support, wherein the sample(s) harbors a first plurality of target RNA and a second plurality of target RNA. In some embodiments, the first plurality of target RNA encode a first polypeptide. In some embodiments, the second plurality of target RNA encode a second polypeptide. In some embodiments, the sample(s) is fixed and permeabilized.
[0290] In some embodiments, the sample(s) harbors 2-25 different target RNA molecules, or harbors 25-50 different target RNA molecules, or harbors 50-75 different target RNA molecules, or harbors 75-100 different target RNA molecules. In some embodiments, the sample(s) harbors more than 100 different target RNA molecules, or more than 250 different target RNA molecules, or more than 500 different target molecules, or more than 1000 different target RNA molecules, or more. In some embodiments, the sample(s) harbors more than 10,000 different target RNA molecules. In some embodiments, the sample(s) comprises a whole cell, a plurality of whole cells, an intact tissue or an intact tumor. In some embodiments, the sample(s) comprises a fresh cellular sample, a freshly-frozen cellular sample, a sectioned cellular sample, or an FFPE cellular sample. In some embodiments, the sample(s) is deposited onto a solid support. In some embodiments, the sample(s) is deposited onto a solid support which is passivated with a coating that promotes cell adhesion. In some embodiments, the sample(s) is deposited on a support that lacks immobilized capture oligonucleotides. In some embodiments, the sample(s) is cultured prior to conducting step (b) which is described below.
[0291] In some embodiments, the sample(s) harbors 2-25 different target polypeptide molecules, or harbors 25-50 different target polypeptide molecules, or harbors 50-75different target polypeptide molecules, or harbors 75-100 different target polypeptide molecules. In some embodiments, the sample(s) harbors more than 100 different target polypeptide molecules, or more than 250 different target polypeptide molecules, or more than 500 different target molecules, or more than 1000 different target polypeptide molecules, or more. In some embodiments, the sample(s) harbors more than 10,000 different target polypeptide molecules. The target polypeptide molecules are encoded by the target RNA molecules.
[0292] In some embodiments, the methods comprise step (b): generating inside the sample(s) a plurality of cDNA by (i) generating at least a first plurality of target cDNA from the first plurality of target RNA, and (ii) generating at least a second plurality of target cDNA from the second plurality of target RNA (e.g., FIG. 13). In some embodiments, the first target cDNAs correspond to the first target RNA molecules. In some embodiments, the second target cDNAs correspond to the second target RNA molecules. In some embodiments, the method comprises generating at least 2-10,000 different target cDNA molecules that correspond to 2-10,000 different target RNA molecules. In some embodiments, the generating of step (b) comprises contacting the plurality of RNA inside the sample(s) with (i) a plurality of reverse transcription primers, (ii) a plurality of reverse transcriptase enzymes, and (iii) a plurality of nucleotides, under a condition suitable for conducting a reverse transcription reaction to generate a plurality of cDNA molecules (e.g., a plurality of first strand cDNA molecules) in the sample(s). In some embodiments, the plurality of reverse transcription primers comprises a first subpopulation of target-specific reverse transcription primers that hybridize selectively to the first target RNA, and / or comprises a second sub-population of target-specific reverse transcription primers that hybridize selectively to the second target RNA. In some embodiments, the plurality of reverse transcription primers comprises a first subpopulation of random-sequence reverse transcription primers that hybridize to the first target RNA, and / or comprises a second sub-population of random-sequence reverse transcription primers that hybridize to the second target RNA.
[0293] In some embodiments, .e.g., in FIG. 13, the first padlock probe comprises (i) a first target barcode sequence (target BC-1) that uniquely identifies the first target RNA, (ii) a first batch-specific sequencing primer binding site (Batch Seq-1) (or a complementary sequence thereof), (iii) a universal binding site for an amplification primer (universal RCA) (or a complementary sequence thereof), and (iv) a universalbinding site for a compaction oligonucleotide (or a complementary sequence thereof). The second padlock probe comprises (i) a second target barcode sequence (target BC-2) that uniquely identifies the second target RNA, (ii) a second batch-specific sequencing primer binding site (Batch Seq-2) (or a complementary sequence thereof), (iii) a universal binding site for an amplification primer (universal RCA) (or a complementary sequence thereof), and (iv) a universal binding site for a compaction oligonucleotide (or a complementary sequence thereof).
[0294] In some embodiments, the methods comprise step (c): generating inside the sample(s) a plurality of DNA concatemers which correspond to the first and second plurality of target RNA molecules, comprising: (1) generating a first plurality of covalently closed circular padlock probes by contacting the first plurality of target cDNA with a first plurality of padlock probes, wherein the contacting is conducted under a condition suitable for hybridizing the first and second binding arms of the first padlock probes to proximal positions on their respective first target cDNA molecules to form a first plurality of circularized padlock probes each having a nick or gap between the hybridized first and second binding arms, wherein the first padlock probes include a (i) a first target barcode sequence (target BC-1) that uniquely identifies the first target RNA or cDNA, (ii) a first batch-specific sequencing primer binding site (Batch Seq-1) (or a complementary sequence thereof), and (iii) a universal binding site for an amplification primer (universal RCA) (or a complementary sequence thereof) (e.g., FIG. 13, left side); (2) enzymatically closing the nick or gap in the first plurality of covalently closed circular padlock probes to form a first plurality of covalently closed padlock probes; and (3) conducting rolling circle amplification inside the sample(s) using the first covalently closed circular padlock probes as template molecules, thereby generating a first plurality of concatemer molecules that correspond to the first plurality of target RNA or cDNA molecules. In some embodiments, the rolling circle amplification reaction can be conducted in the presence or absence of a plurality of compaction oligonucleotides. In some embodiments, the method comprises contacting the plurality of cDNA molecule in the sample(s) with at least 2-10,000 different target-specific padlock probes. In some embodiments, the first padlock probe further comprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof). In some embodiments, the closing the nick in the first circularized padlock probes comprises conducting an enzymatic ligation reaction. In some embodiments, closing the gap in thefirst circularized padlock probes comprises conducting a polymerase-catalyzed fill-in reaction using the first target cDNA molecule as a template, and conducting an enzymatic ligation reaction. In some embodiments, the method comprises closing the nick or gap in at least 2-10,000 circularized target-specific padlock probes by conducting an enzymatic reaction, thereby generating at least 2-10,000 covalently closed circular padlock probes inside the sample(s). In some embodiments, each concatemer molecule in the first plurality comprises tandem repeat units, wherein a unit comprises the sequence of the first target cDNA and (i) the first target barcode sequence (target BC-1) that uniquely identifies the first target RNA, (ii) the first batch-specific sequencing primer binding site (Batch Seq-1) (or a complementary sequence thereof), and (iii) the universal binding site for an amplification primer (universal RCA) (or a complementary sequence thereof). In some embodiments, the unit further comprises the universal compaction oligonucleotide binding site (or a complementary sequence thereof).
[0295] In some embodiments, step (c) further comprises: generating inside the sample(s) a plurality of DNA concatemers which correspond to the second plurality of target RNA molecules, comprising: (1) generating a second plurality of covalently closed circular padlock probes by contacting the second plurality of target cDNA with a second plurality of padlock probes, wherein the contacting is conducted under a condition suitable for hybridizing the first and second binding arms of the second padlock probes to proximal positions on their respective second target cDNA molecules to form a second plurality of circularized padlock probes each having a nick or gap between the hybridized first and second binding arms, wherein the second padlock probes include a (i) a second barcode sequence (target BC-2) that uniquely identifies the second target cDNA or RNA, (ii) a second batch-specific sequencing primer binding site (Batch Seq-2) (or a complementary sequence thereof) wherein the sequence of the second batch-specific sequencing primer binding site differs from the sequence of the first batch-specific sequencing primer binding site, and (iii) the universal binding site for an amplification primer (universal RCA) (or a complementary sequence thereof) (e.g., FIG. 13, right side); (2) enzymatically closing the nick or gap in the second plurality of covalently closed circular padlock probes to form a second plurality of covalently closed padlock probes; and (3) conducting rolling circle amplification inside the sample(s) using the second covalently closed circular padlock probes as template molecules, thereby generating a second plurality of concatemer molecules that correspond to the second plurality of target RNA molecules.In some embodiments, the rolling circle amplification reaction can be conducted in the presence or absence of a plurality of compaction oligonucleotides. In some embodiments, the method comprises contacting the plurality of cDNA molecule in the sample(s) with at least 2-10,000 different target-specific padlock probes. In some embodiments, the second padlock probe further comprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof). In some embodiments, the closing the nick in the second circularized padlock probes comprises conducting an enzymatic ligation reaction. In some embodiments, closing the gap in the second circularized padlock probes comprises conducting a polymerase-catalyzed fill-in reaction using the second target cDNA molecule as a template, and conducting an enzymatic ligation reaction. In some embodiments, the method comprises closing the nick or gap in at least 2-10,000 circularized target-specific padlock probes by conducting an enzymatic reaction, thereby generating at least 2-10,000 covalently closed circular padlock probes inside the sample(s). In some embodiments, each concatemer molecule in the second plurality comprises tandem repeat units, wherein a unit comprises the sequence of the second target cDNA and (i) the second target barcode sequence (target BC-2) that uniquely identifies the second target cDNA or RNA, (ii) the second batch-specific sequencing primer binding site (Batch Seq-2) (or a complementary sequence thereof), and (iii) the universal binding site for an amplification primer (universal RCA) (or a complementary sequence thereof). In some embodiments, the unit further comprises the universal compaction oligonucleotide binding site (or a complementary sequence thereof).
[0296] In some embodiments, the methods further comprise step (d): sequencing the first plurality of concatemer molecules inside the sample(s) under a condition that inhibits sequencing the second plurality of concatemers (e.g., FIG. 14). In some embodiments, step (d) comprises sequencing the first plurality of concatemers inside the sample(s) comprises conducting no more than 2-30 sequencing cycles to generate a plurality of first sequencing read products, wherein the sequences of the first sequencing read products are aligned with a first target reference sequence to confirm the presence of the first target RNA in the sample(s). In some embodiments, step (d) comprises sequencing the first plurality of concatemers inside the sample(s) comprises conducting 1-250 sequencing cycles to generate a plurality of first sequencing read products, wherein the sequences of the first sequencing read products are aligned with a first target reference sequence to confirm the presence of the first target RNA in the sample(s).
[0297] In some embodiments, e.g., in FIG. 14, The first and second concatemers are subjected to a first sequencing workflow using first batch-specific sequencing primers, sequencing polymerases, and a plurality of nucleotide reagents. The first concatemers undergo reiterative sequencing but the second concatemers do not. The first and second concatemers are subjected to a second sequencing workflow using second batch-specific sequencing primers, sequencing polymerases, and a plurality of nucleotide reagents. The second concatemers undergo reiterative sequencing but the first concatemers do not.
[0298] In some embodiments in step (d), in the first concatemer molecules, only the first target barcode region (target BC-1) is sequenced. In some embodiments, in the first concatemer molecules, at least a portion or the full length of the first target barcode (target BC-1) is sequenced. In some embodiments, in the first concatemer molecules, the first target barcode (target BC-1) is sequenced and a portion of the first cDNA region is sequenced.
[0299] In some embodiments, the sequencing the first concatemers of step (d) comprises step (1) contacting the first plurality of concatemer molecules inside the sample(s) with (i) a plurality of first batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents, under a condition suitable for hybridizing the plurality of first batch-specific sequencing primers to their respective first batch-specific sequencing primer binding sites on the first concatemers. In some embodiments, the sequencing further comprises step (2) conducting no more than 2-30 sequencing cycles to generate a first plurality of sequencing read products using the first concatemers as template molecules.
[0300] In some embodiments, the sequencing of step (d) comprises sequencing at least a portion of the first nucleic acid concatemers using an optical imaging system comprising a field-of-view (FOV) greater than 1.0 mm2.
[0301] In some embodiments, in the sequencing of step (d), the plurality of first sequencing read products are detectable by imaging, and wherein the sequencing comprises decoding the plurality of first sequencing read products from the images obtained during the no more than 2-30 sequencing cycles, or from the images obtained during the 1-250 sequence cycles.
[0302] In some embodiments, the methods further comprise step (e): removing the plurality of first sequencing read products from the first concatemer molecules and retaining the first concatemer molecules inside the sample(s). In some embodiments, a 3’blocking moiety can be added to the first sequencing read products to inhibit further sequencing reactions. For example, a nucleotide analog can be incorporated where the nucleotide analog inhibits incorporation of a subsequent nucleotide. Exemplary blocking nucleotide analogs include dideoxynucleotide or a nucleotide having a 2’ or 3’ chain terminating moiety.
[0303] In some embodiments, the methods further comprise step (f): reiteratively sequencing the plurality of first concatemers by repeating steps (d) and (e) at least once. In some embodiments, reiterative sequencing of step (f) is optional.
[0304] In some embodiments, the sequencing the first concatemers of step (f) comprises step (1) contacting the first plurality of concatemer molecules inside the sample(s) with (i) a plurality of first batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents, under a condition suitable for hybridizing the plurality of first batch-specific sequencing primers to their respective first batch-specific sequencing primer binding sites on the first concatemers. In some embodiments, the sequencing further comprises step (2) conducting no more than 2-30 sequencing cycles to generate a first plurality of sequencing read products using the first concatemers as template molecules. In some embodiments, the sequencing further comprises step (3) removing the first plurality of sequencing read products from the first concatemers and retaining the plurality of first concatemers inside the sample(s). In some embodiments, the sequencing further comprises step (4) repeating steps (1) - (3) at least once (e.g., FIG. 14). In some embodiments, step (4) comprises repeating steps (1) - (3) at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, or at least 10 times. In some embodiments, step (4) comprises repeating steps (1) - (3) up to 10 times, up to 20 times, up to 30 time, up to 40 times, or up to 50 times.
[0305] In some embodiments, the reiterative sequencing of the first concatemers of step (f) can be conducting using a sequencing-by-binding procedure, labeled and / or nonlabeled chain-terminating nucleotides, or multivalent molecules. Descriptions of these three sequencing methods is described below.
[0306] In some embodiments, the plurality of universal sequencing primers can be hybridized to concatemer template molecules with a hybridization reagent comprising an SSC buffer (e.g., 2X saline-sodium citrate) buffer with formamide (e.g., 10-20%formamide). The hybridization conditions comprise a temperature of about 20-30 °C, for about 10-60 minutes.
[0307] In some embodiments, the plurality of sequencing read products can be removed from the concatemers and the plurality of concatemers can be retained inside the sample(s) using a de-hybridization reagent comprising an SSC buffer (e.g., saline-sodium citrate) buffer, with or without formamide, at a temperature that promotes nucleic acid denaturation such as for example 30 - 90 °C.
[0308] In some embodiments, the methods further comprise step (g): sequencing the second plurality of concatemer molecules inside the sample(s) under a condition that inhibits sequencing the first plurality of concatemers (e.g., FIG. 14). In some embodiments, step (g) comprises sequencing the second plurality of concatemers inside the sample(s) comprises conducting no more than 2-30 sequencing cycles to generate a plurality of second sequencing read products, wherein the sequences of the second sequencing read products are aligned with a second target reference sequence to confirm the presence of the second target RNA in the sample(s). In some embodiments, step (g) comprises sequencing the second plurality of concatemers inside the sample(s) comprises conducting 1-250 sequencing cycles to generate a plurality of second sequencing read products, wherein the sequences of the second sequencing read products are aligned with a second target reference sequence to confirm the presence of the second target RNA in the sample(s).
[0309] In some embodiments in step (g), in the second concatemer molecules, only the second target barcode region (target BC-2) is sequenced. In some embodiments, in the second concatemer molecules, at least a portion or the full length of the second target barcode (target BC-2) is sequenced. In some embodiments, in the second concatemer molecules, the second target barcode (target BC-2) is sequenced and a portion of the second cDNA region is sequenced.
[0310] In some embodiments, the sequencing the second concatemers of step (g) comprises step (1) contacting the second plurality of concatemer molecules inside the sample(s) with (i) a plurality of second batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents, under a condition suitable for hybridizing the plurality of second batch-specific sequencing primers to their respective second batch-specific sequencing primer binding sites on the second concatemers. In some embodiments, the sequencing further comprises step (2) conductingno more than 2-30 sequencing cycles to generate a second plurality of sequencing read products using the second concatemers as template molecules.
[0311] In some embodiments, the sequencing of step (g) comprises sequencing at least a portion of the second nucleic acid concatemers using an optical imaging system comprising a field-of-view (FOV) greater than 1.0 mm2.
[0312] In some embodiments, in the sequencing of step (g), the plurality of second sequencing read products are detectable by imaging, and wherein the sequencing comprises decoding the plurality of second sequencing read products from the images obtained during the no more than 2-30 sequencing cycles, or from the images obtained during the 1-250 sequencing cycles.
[0313] In some embodiments, the methods further comprise step (h): removing the plurality of second sequencing read products from the second concatemer molecules and retaining the second concatemer molecules inside the sample(s). In some embodiments, a 3’ blocking moiety can be added to the second sequencing read products to inhibit further sequencing reactions. For example, a nucleotide analog can be incorporated where the nucleotide analog inhibits incorporation of a subsequent nucleotide. Exemplary blocking nucleotide analogs include dideoxynucleotide or a nucleotide having a 2’ or 3’ chain terminating moiety.
[0314] In some embodiments, the methods further comprise step (i): reiteratively sequencing the plurality of second concatemers by repeating steps (g) and (h) at least once. In some embodiments, reiterative sequencing of step (i) is optional.
[0315] In some embodiments, the sequencing the second concatemers of step (i) comprises step (1) contacting the second plurality of concatemer molecules inside the sample(s) with (i) a plurality of second batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents, under a condition suitable for hybridizing the plurality of second batch-specific sequencing primers to their respective second batch-specific sequencing primer binding sites on the second concatemers. In some embodiments, the sequencing further comprises step (2) conducting no more than 2-30 sequencing cycles to generate a first plurality of sequencing read products using the second concatemers as template molecules. In some embodiments, the sequencing further comprises step (3) removing the first plurality of sequencing read products from the second concatemers and retaining the plurality of second concatemers inside the sample(s). In some embodiments, the sequencing further comprises step (4)repeating steps (1) - (3) at least once (e.g., FIG. 14). In some embodiments, step (4) comprises repeating steps (1) - (3) at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, or at least 10 times. In some embodiments, step (4) comprises repeating steps (1) - (3) up to 10 times, up to 20 times, up to 30 time, up to 40 times, or up to 50 times.
[0316] In some embodiments, the reiterative sequencing of the second concatemers of step (i) can be conducting using a sequencing-by-binding procedure, labeled and / or nonlabeled chain-terminating nucleotides, or multivalent molecules. Descriptions of these three sequencing methods is described below.
[0317] In some embodiments, the plurality of nucleotide reagents of steps (d) and (g) comprise a plurality of nucleotides that are detectably labeled or non-labeled. In some embodiments, individual nucleotides are linked to a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the plurality of detectably labeled nucleotide analogs comprise a plurality of chain terminating nucleotides, where the chain terminating moiety is linked to the 3’ nucleotide sugar position to form a 3’ blocked nucleotide analog. In some embodiments, the chain terminating moiety can be removed to convert the 3’ blocked nucleotide analog to an extendible nucleotide having a 3’ OH group on the sugar. In some embodiments, the labeled nucleotide analogs are linked to a different fluorophore that corresponds to the nucleo-bases adenine, cytosine, guanine, thymine or uracil, where the different fluorophores emit a fluorescent signal. In some embodiments, a sequencing cycle comprises (1) contacting the concatemer / sequencing primer duplex with a sequencing polymerase and a detectably labeled chain terminating nucleotide under a condition suitable for polymerase-catalyzed incorporation of the detectably labeled chain terminating nucleotide into the terminal end of the sequencing primer, (2) detecting and imaging the fluorescent signal and color emitted by the incorporated chain terminating nucleotide, and (3) removing the chain terminating moiety (e.g., unblocking) and the fluorophore from the incorporated nucleotide and retaining the concatemer / sequencing primer duplex. In some embodiments, no more than 2-30 sequencing cycles are conducted on the plurality of concatemers inside the sample(s) to generate a plurality of sequencing read products. In some embodiments, the sequence of the first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of the first target RNA molecules inside the sample(s). In some embodiments,the sequence of the second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of the second target RNA molecules inside the sample(s).
[0318] In some embodiments, the sequences of the first and second sequencing read products can be aligned after each round of generating the first and second sequencing read products which are no more than 30 bases in length, or after generating a set of reiterative sequencing read products wherein the first and second sequencing read products which are no more than 30 bases in length. In some embodiments, the sequencing reactions are conducted on a sequencing apparatus having a detector that captures fluorescent signals from the sequencing reactions inside the sample(s). The sequencing apparatus can be configured to relay the fluorescent signal data captured by the detector to a computer system that is programmed to display images of different fluorescent spots which are co-located in the sample(s), where individual fluorescent spots correspond to different target RNA molecules. In some embodiments, when the sequencing is conducted using different fluorescently-labeled nucleotide reagents that correspond to different nucleo-bases (e.g., A, G, C, T / U), then the images can have different color fluorescent spots co-located in the same cellular sample at different sequencing cycles.
[0319] In some embodiments, out-of-sync phasing and / or pre-phasing events can occur during synchronized sequencing reactions on clonally amplified template amplicons, where the sequencing reactions comprise polymerase-catalyzed sequencing reactions employing detectably labeled chain terminator nucleotides. In some embodiments, a sequencing reaction on one template molecule in the clonally-amplified template molecules moves ahead (e.g., pre-phasing) or fall behind (e.g., phasing) of the sequencing of the other template molecules within the clonally-amplified template molecules. During sequencing, a fluorescent signal is typically detected which corresponds to incorporation of a labeled chain terminator nucleotide. Thus, phasing and pre-phasing events can be detected and monitored using incorporation of a labeled chain terminator nucleotide.
[0320] In some embodiments, the plurality of nucleotide reagents of steps (d) and (g) comprise a plurality of multivalent molecules each comprising a core attached to a plurality of nucleotide-arms, wherein the nucleotide-arms are attached to a nucleotide unit. In some embodiments, individual multivalent molecules are labeled with a detectably reporter moiety. In some embodiments, the detectable reporter moietycomprises a fluorophore. In some embodiments, the core of the multivalent molecule is labeled with a fluorophore, and wherein the fluorophore which is attached to a given core of the multivalent molecule corresponds to the nucleotide base (e.g., adenine, guanine, cytosine, thymine or uracil) of the nucleotide arm. In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or nucleotide base that is attached to a fluorophore, and wherein the fluorophore which is attached to a given nucleotide base corresponds to the nucleotide base (e.g., adenine, guanine, cytosine, thymine or uracil) of the nucleotide arm. In some embodiments, a sequencing cycle comprises (1) contacting the concatemer / sequencing primer duplex with a first sequencing polymerase to form a complexed polymerase, (2) contacting the complexed polymerase with a detectably labeled multivalent molecule under a condition suitable for binding a complementary nucleotide unit of the multivalent molecule to the complexed polymerase thereby forming a multivalent-binding complex, and the condition is suitable for inhibiting incorporation of the complementary nucleotide unit into the terminal end of the sequencing primer, (3) detecting and imaging the fluorescent signal and color emitted by the bound detectably labeled multivalent molecule, (4) removing the first sequencing polymerase and the bound detectably labeled multivalent molecule, and retaining the concatemer / sequencing primer duplex, (5) contacting the retained concatemer / sequencing primer duplex with a second sequencing polymerase and a non-labeled chain terminating nucleotide under a condition suitable for polymerase-catalyzed incorporation of the nonlabeled chain terminating nucleotide into the terminal end of the sequencing primer, and (6) removing the chain terminating moiety (e.g., unblocking) and retaining the concatemer / sequencing primer duplex. In some embodiments, no more than 2-30 sequencing cycles are conducted on the plurality of concatemers inside the sample(s) to generate a plurality of sequencing read products. In some embodiments, the sequence of the first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of the first target RNA molecules inside the sample(s). In some embodiments, the sequence of the second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of the second target RNA molecules inside the sample(s). In some embodiments, the sequences of the first and second sequencing read products can be aligned after each round of generating the first and second sequencing read products which are no more than 30 bases in length, or after generating a set of reiterative sequencing read products wherein the firstand second sequencing read products which are no more than 30 bases in length. In some embodiments, the sequencing reactions are conducted on a sequencing apparatus having a detector that captures fluorescent signals from the sequencing reactions inside the sample(s). The sequencing apparatus can be configured to relay the fluorescent signal data captured by the detector to a computer system that is programmed to display images of different fluorescent spots which are co-located in the sample(s), where individual fluorescent spots correspond to different target RNA molecules. In some embodiments, individual cycle times can be achieved in less than 30 minutes. In some embodiments, the field of view (FOV) can exceed 1 mm2and the cycle time for scanning large area (> 10 mm2) can be less than 5 minutes.
[0321] In any of the methods described herein, the plurality of RNA or cDNA inside the sample(s) can be amplified to generate amplicons of the RNA or cDNA where the amplicons comprise concatemers. In some embodiments, the plurality of RNA or cDNA molecules inside the sample(s) can be amplified by conducting a padlock probe circularization and rolling circle amplification workflow. In some embodiments, the methods comprise contacting the plurality of RNA or cDNA molecules inside the sample(s) with a plurality of padlock probes, including a first plurality of target-specific padlock probes that hybridize with first target RNA or cDNA molecules, and a second plurality of target-specific padlock probes that hybridize with second target RNA or cDNA molecules.
[0322] In some embodiments, the padlock probes comprise single-stranded oligonucleotides. In some embodiments, the padlock probes comprise DNA, RNA, or DNA and RNA. In some embodiments, individual padlock probes comprise an internal region between the first and second terminal regions, where the internal region comprises at least one universal adaptor sequence including a sample barcode sequence, an amplification primer binding site, a sequencing primer binding site, a compaction oligonucleotide binding site and / or a surface capture primer binding site (FIG. 6). In some embodiments, the padlock probes comprise at least one target barcode sequence that corresponds to a given target RNA or target cDNA to which the padlock probes binds. In some embodiments, the padlock probes comprise at least one unique identification sequence (e.g., unique molecular index (UMI)). In some embodiments, the padlock probes comprise at least one restriction enzyme recognition sequence.
[0323] In some embodiments, a padlock probe comprises a single-stranded nucleic acid molecule having two terminal regions (e.g., first and second binding arms) and an internal region. In some embodiments, the first terminal region of an individual padlock probe has a first target-specific sequence that selectively hybridizes to a first region of a target RNA or target cDNA molecule, and the second terminal region of the individual padlock probe has a second target-specific sequence that selectively hybridizes to a second region of the same target RNA or target cDNA molecule. In some embodiments, the internal region of a padlock comprises a target barcode sequence (e.g., Target BC-1 or Target BC-2, left and right schematics respectively) which corresponds to a given target RNA or target cDNA. In some embodiments, the target barcode sequence uniquely identifies the target RNA or target cDNA. In some embodiments, the internal region of a padlock comprises a universal primer binding site for a sequencing primer (or a complementary sequence thereof). In some embodiments, the internal region of a padlock comprises a universal primer binding site for a rolling circle amplification primer (or a complementary sequence thereof). In some embodiments, the internal region of a padlock comprises a universal binding site for a compaction oligonucleotide binding (or a complementary sequence thereof). In some embodiments, the internal region of a padlock probe includes a target barcode sequence and at least one universal primer binding site (e.g., for binding a sequencing primer, for binding a rolling circle amplification primer and / or for binding a compaction oligonucleotide) in any arrangement and orientation (FIG. 6, top and bottom).
[0324] In some embodiments, individual padlock probes comprise first and second terminal regions (e.g., first and second binding arms) that hybridize to portions of target RNA or target cDNA molecules to form a plurality of RNA-padlock probe complexes or a plurality of cDNA-padlock probe complexes, wherein individual complexes have the first and second terminal probe regions hybridized to proximal regions of an RNA or cDNA molecule to form a nick or gap between the first and second terminal probe ends. In some embodiments, the first terminal region of an individual padlock probe has a first target-specific sequence that selectively hybridizes to a first region of a target RNA or cDNA molecule, and the second terminal region of the individual padlock probe has a second target-specific sequence that selectively hybridizes to a second region of the same target RNA or cDNA molecule, where a nick or gap is formed between the hybridized first and second terminal regions, thereby circularizing the padlock probe (e.g., FIG. 7).
[0325] As shown in FIG. 7, the first padlock probe comprises (i) a first target barcode sequence (target BC-1) that uniquely identifies the first target RNA or the first target cDNA, (ii) a first sequencing primer binding site (or a complementary sequence thereof), (iii) a universal binding site for an amplification primer (universal RCA) (or a complementary sequence thereof), and (iv) a universal binding site for a compaction oligonucleotide (or a complementary sequence thereof). The second padlock probe comprises (i) a second target barcode sequence (target BC-2) that uniquely identifies the second target RNA or the second target cDNA, (ii) a second sequencing primer binding site(or a complementary sequence thereof), (iii) a universal binding site for an amplification primer (universal RCA) (or a complementary sequence thereof), and (iv) a universal binding site for a compaction oligonucleotide (or a complementary sequence thereof).
[0326] In some embodiments, the padlock probes comprise canonical nucleotides and / or nucleotide analogs. In some embodiments, the padlock probes are modified to confer resistance to nuclease degradation (e.g., ribonuclease degradation). For example, the padlock probes comprise at least one phosphorothioate diester bond at their 5’ ends which can render the padlock probes resistant to nuclease degradation. In some embodiments, the padlock probes comprise 2-5 or more consecutive phosphorothioate diester bonds at their 5’ ends. In some embodiments, the padlock probes comprise at least one ribonucleotide and / or at least one 2’-O-methyl, 2’-O-methoxyethyl (MOE), 2’ fluoro-base nucleotide. In some embodiments, the padlock probes comprise phosphorylated 3’ ends. In some embodiments, the padlock probes comprise at least one locked nucleic acid (LNA) base. In some embodiments, the padlock probes comprise a phosphorylated 5’ end (e.g., using a polynucleotide kinase).
[0327] In some embodiments, individual padlock probes in a set of padlock probes (e.g., a plurality of padlock probes) comprise first and second terminal regions that hybridize to the same target regions of the target RNA or cDNA molecules to form a plurality of RNA-padlock probe complexes or a plurality of cDNA-padlock probe complexes having the same RNA or cDNA sequence.
[0328] In some embodiments, a set of padlock probes (e.g., a plurality of padlock probes) comprise at least two sub-sets of padlock probes. In some embodiments, individual padlock probes in a first sub-set of padlock probes comprise first and second terminal regions that hybridize to the same target regions (e.g., a first target region) of the targetRNA or cDNA molecules to form a first plurality of RNA-padlock probe complexes or a first plurality of cDNA-padlock probe complexes having the same RNA or cDNA sequence. In some embodiments, individual padlock probes in a second sub-set of padlock probes comprise first and second terminal regions that hybridize to the same target regions (e.g., a second target region) of the target RNA or cDNA molecules to form a second plurality of RNA-padlock probe complexes or a second plurality of cDNA- padlock probe complexes having the same cDNA sequence. In some embodiments, the first and second sub-sets of padlock probes hybridize to different target regions of the same target RNA or cDNA molecules. In some embodiments, the first and second subsets of padlock probes hybridize to different target regions of different target RNA or cDNA molecules. In some embodiments, the set of padlock probes comprise 2-10 subsets of padlock probes, or 10-25 sub-sets of padlock probes, or 25-50 sub-sets of padlock probes, or up to 100 sub-sets of padlock probes. In some embodiments, the set of padlock probes comprise at least 100 sub-sets of padlock probes, at least 500 sub-sets of padlock probes, at least 1000 sub-sets of padlock probes, at least 10,000 sub-sets of padlock probes, or more sub-sets of padlock probes.
[0329] In some embodiments, the nicks can be enzymatically ligated to generate covalently closed circular padlock probes. In some embodiments, the ligase enzyme can discriminate between matched and mis-matched hybridized ends to ensure target-specific hybridization. In some embodiments, the ligation reaction comprises use of a ligase enzyme, including a T3, T4, T7 or Taq DNA ligase enzyme.
[0330] In some embodiments, the size of the gap between the hybridized first and second terminal regions is 1-25 bases. The 3 ’OH end of hybridized padlock probe can serve as an initiation site for a polymerase-catalyzed fill-in reaction (e.g., gap fill-in reaction) using the target cDNA molecule (or the target RNA molecule) as a template. After the fill-in reaction, the remaining nick can be enzymatically ligated to generate covalently closed circular padlock probes.
[0331] In some embodiments, the gap-filling reaction comprises contacting the circularized padlock probe with a DNA polymerase and a plurality of nucleotides. In some embodiments, the DNA polymerase comprises E. coli DNA polymerase I, Klenow fragment of E. coli DNA polymerase I, T7 DNA polymerase, or T4 DNA polymerase. In some embodiments, the ligase enzyme can discriminate between matched and mismatched hybridized ends to ensure target-specific hybridization. In some embodiments,the ligation reaction comprises use of a ligase enzyme, including a T3, T4, T7 or Taq DNA ligase enzyme.
[0332] In any of the methods described herein, the plurality of covalently closed circular padlock probes can be subjected to a rolling circle amplification reaction to generate a plurality of concatemer molecules each having two or more tandem copies of a unit wherein the unit comprises a target sequence that corresponds to a target RNA molecules and any additional sequence(s) carried by the padlock probes including universal adaptor sequence(s), unique molecular index sequence(s) and / or restriction enzyme recognition sequence(s).
[0333] In some embodiments, the rolling circle amplification reaction comprises contacting the covalently closed circularized padlock probes with an amplification primer (e.g., a universal rolling circle amplification primer), a strand-displacing DNA polymerase, and a plurality of nucleotides, under a condition suitable for hybridizing individual amplification primers to a covalently closed padlock probe, and under a condition suitable for conducting primer extension using the covalently closed padlock probe as a template molecule to generate a nucleic acid concatemer. In some embodiments, the plurality of nucleotides in the rolling circle amplification reaction comprise any mixture of two or more of dATP, dGTP, dCTP, dTTP and / or dUTP. In some embodiments, any of the rolling circle amplification reactions described herein can be conducted in the presence or in the absence of a plurality of compaction oligonucleotides.
[0334] In some embodiments, when the rolling circle amplification reaction includes a plurality of nucleotide which includes dUTP, the resulting concatemer can be cross-linked to a cross-linking reactive group by treating the sample(s) with a succinimide ester (NHS), maleimide (Sulfo-SMCC), imidoester (DMP), carbodiimide (DCC, EDC) or phenyl azide. In some embodiments, polymerization of the cross-linking reactive group can be initiated with light or UV light. In some embodiments, the resulting concatemer can be cross-linked to a matrix by treating the sample(s) with a cross-linked agarose, cross-linked dextran or cross-linked polyethylene glycol (PEG), polyacrylamide, cellulose alginate or polyamide. In some embodiments, the PEG comprises a sulfo-NHS ester moiety at one or both ends, for example a PEGylated bis(sulfosuccinimidyl)suberate) (e.g., BS(PEG)9 from Thermo Fisher Scientific, catalog No. 21582).
[0335] In some embodiments, the rolling circle amplification reaction can be conducted at a constant temperature (e.g., isothermal) wherein the constant temperature is at room temperature to about 30 °C, or about 30 - 40 °C, or about 40 - 50 °C, or about 50 - 65 °C.
[0336] In some embodiments, the DNA polymerase having a strand displacing activity can be selected from a group consisting of phi29 DNA polymerase, large fragment of Bst DNA polymerase, large fragment of Bsu DNA polymerase, and Bea (exo-) DNA polymerase, Klenow fragment of E. coli DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV viral reverse transcriptase, or Deep Vent DNA polymerase. In some embodiments, the phi29 DNA polymerase can be wild type phi29 DNA polymerase (e.g., MagniPhi from Expedeon), or variant EquiPhi29 DNA polymerase (e.g., from Thermo Fisher Scientific), and chimeric QualiPhi DNA polymerase (e.g., from 4basebio).
[0337] In some embodiments, the rolling circle amplification primers can be modified to increase resistance to nuclease degradation. In some embodiments, the rolling circle amplification primers comprise at least one phosphorothioate diester bond at their 5’ ends which can render the amplification primers resistant to exonuclease degradation. In some embodiments, the rolling circle amplification primers comprise 2-5 or more consecutive phosphorothioate diester bonds at their 5’ ends. In some embodiments, the rolling circle amplification primers comprise at least one ribonucleotide and / or at least one 2’-O- methyl or 2’-O-methoxyethyl (MOE) nucleotide.
[0338] In some embodiments, the rolling circle amplification reaction can be conducted in the presence of a plurality of compaction oligonucleotides which, when hybridized to a concatemer molecule, compacts the size and / or shape of the concatemer to form a compact nanoball. In some embodiments, the compaction oligonucleotides comprise single stranded oligonucleotides having a first region at one end that hybridizes to a portion of a concatemer molecule and a second region at the other end that hybridizes to another portion of the same concatemer molecule, where hybridization of the compaction oligonucleotide to a given concatemer compacts the size and / or shape of the concatemer.
[0339] The compaction oligonucleotides include a 5’ region, an optional internal region (intervening region), and a 3’ region. The 5’ and 3’ regions of the compaction oligonucleotide can hybridize to any portions of the concatemer. The 5’ and 3’ regions of the compaction oligonucleotide can hybridize to different portions of the concatemer to pull together distal portions of the concatemer causing compaction of the concatemer to form a DNA nanoball. For example, the 5’ region of the compaction oligonucleotide isdesigned to hybridize to a first portion of the concatemer molecule (e.g., a universal compaction oligonucleotide binding site), and the 3’ region of the compaction oligonucleotide is designed to hybridized to a second portion of the concatemer molecule (e.g., a universal compaction oligonucleotide binding site). Inclusion of compaction oligonucleotides during RCA can promote formation of DNA nanoballs having tighter size and shape compared to concatemers generated in the absence of the compaction oligonucleotides. The compact and stable characteristics of the DNA nanoballs improves in situ sequencing accuracy by increasing signal intensity and the nanoballs retain their shape and size during multiple sequencing cycles.
[0340] In some embodiments, the compaction oligonucleotides comprise single stranded oligonucleotides comprising DNA, RNA, or a combination of DNA and RNA. The compaction oligonucleotides can be any length, including 20-150 nucleotides, or 30-100 nucleotides, or 40-80 nucleotides in length.
[0341] In some embodiments, the compaction oligonucleotides comprises a 5’ region and a 3’ region, and optionally an intervening region between the 5’ and 3’ regions. The intervening region can be any length, for example about 2-20 nucleotides in length. The intervening region comprises a homopolymer having consecutive identical bases (e.g., AAA, GGG, CCC, TTT or UUU). The intervening region comprises a non-homopolymer sequence.
[0342] The 5’ region of the compaction oligonucleotides can be wholly complementary or partially complementary along its length to a first portion of a concatemer molecule. The 3’ region of the compaction oligonucleotides can be wholly complementary or partially complementary along its length to a second portion of a concatemer molecule. The 5’ region of the compaction oligonucleotides can hybridize to a first universal sequence portion of a concatemer molecule. The 3’ region of the compaction oligonucleotides can hybridize to a second universal sequence portion of a concatemer molecule.
[0343] In some embodiments, the 5’ region of the compaction oligonucleotide can have the same sequence as the 3’ region. The 5’ region of the compaction oligonucleotide can have a sequence that is different from the 3’ region. In some embodiments, the 3’ region of the compaction oligonucleotide can have a sequence that is a reverse sequence of the 5’ region. In some embodiments, the 5’ region of the compaction oligonucleotide can have a sequence that is a reverse sequence of the 3’ region.
[0344] In some embodiments, the 3’ region of any of the compaction oligonucleotides can include an additional three bases at the terminal 3’ end which comprises 2’-O-methyl RNA bases (e.g., designated mUmUmU) or the terminal 3’ end lacks additional 2’-O- methyl RNA bases.
[0345] In some embodiments, the compaction oligonucleotides comprise one or more modified bases or linkages at their 5’ or 3’ ends to confer certain functionalities. In some embodiments, the compaction oligonucleotides comprise at least one phosphorothioate linkages at their 5’ and / or 3’ ends to confer exonuclease resistance. In some embodiments, at least one nucleotide at or near the 3’ end comprises a 2’ fluoro base which confers exonuclease resistance. In some embodiments, the 3’ end of the compaction oligonucleotides comprise at least one 2’-O-methyl RNA base which blocks polymerase-catalyzed extension. For example, the 3’ end of the compaction oligonucleotide comprises three bases comprising 2’-O-methyl RNA base (e.g., designated mUmUmU). In some embodiments, the compaction oligonucleotides comprise a 3’ inverted dT at their 3’ ends which blocks polymerase-catalyzed extension. In some embodiments, the compaction oligonucleotides comprise 3’ phosphorylation which blocks polymerase-catalyzed extension. In some embodiments, the internal region of the compaction oligonucleotides comprise at least one locked nucleic acid (LNA) which increases the thermal stability of duplexes formed by hybridizing a compaction oligonucleotide to a concatemer molecule. In some embodiments, the compaction oligonucleotides comprise a phosphorylated 5’ end (e.g., using a polynucleotide kinase).
[0346] In some embodiments, the compaction oligonucleotide comprises the sequence
[0347] 5 ’ -C ATGT AATGC ACGT ACTTTC AGGGT AAAC ATGT AATGC ACGT ACTTT
[0348] CAGGGT-3’ (SEQ ID NO: 1). In some embodiments, the compaction oligonucleotides includes an additional three bases at the terminal 3’ end which comprises 2’-O-methyl RNA bases (e.g., designated mUmUmU) or the terminal 3’ end lacks additional 2’-O-methyl RNA bases.
[0349] In some embodiments, the compaction oligonucleotides can include at least one region having consecutive guanines. For example, the compaction oligonucleotides can include at least one region having 2, 3, 4, 5, 6 or more consecutive guanines. In some embodiments, the compaction oligonucleotides comprise four consecutive guanines which can form a guanine tetrad structure (see FIG. 25). The guanine tetrad structure canbe stabilized via Hoogsteen hydrogen bonding. The guanine tetrad structure can be stabilized by a central cation including potassium, sodium, lithium, rubidium or cesium.
[0350] At least one compaction oligonucleotide can form a guanine tetrad (FIG. 25) and hybridize to the universal binding sequences in a concatemer which can cause the concatemer to fold to form an intramolecular G-quadruplex structure (FIG. 26). The concatemers can self-collapse to form compact nanoballs. Formation of the guanine tetrads and G-quadruplexes in the nanoballs may increase the stability of the nanoballs to retain their compact size and shape which can withstand changes in pH, temperature and / or repeated flows of reagents during sequencing inside the sample(s).
[0351] In some embodiments, the plurality of compaction oligonucleotides in the rolling circle amplification reaction have the same sequence. Alternatively, the plurality of compaction oligonucleotides in the rolling circle amplification reaction comprise a mixture of two or more different populations of compaction oligonucleotides having different sequences.
[0352] In some embodiment, the immobilized concatemer template molecule can selfcollapse into a compact nucleic acid nanoball. The nanoballs can be imaged and a FWHM measurement can be obtained to give the shape / size of the nanoballs.
[0353] In some embodiments, inclusion of compaction oligonucleotides in the rolling circle amplification reaction can promote collapsing of a concatemer into a DNA nanoball. Conducting RCA with compaction oligonucleotides helps retain the compact size and shape of a DNA nanoball during multiple sequencing cycles which can improve FWHM (full width half maximum) of a spot image of the DNA nanoball inside a cellular sample. In some embodiments, the DNA nanoball does not unravel during multiple sequencing cycles. In some embodiments, the spot image of the DNA nanoball does not enlarge during multiple sequencing cycles. In some embodiments, the spot image of the DNA nanoball remains a discrete spot during multiple sequencing cycles. The spot image can be represented as a Gaussian spot and the size can be measured as a FWHM. A smaller spot size as indicated by a smaller FWHM typically correlates with an improved image of the spot. In some embodiments, the FWHM of a nanoball spot can be about 10 um or smaller.
[0354] The single-stranded concatemers collapse into compact DNA nanoballs, where each nanoball carries numerous tandem copies of a polynucleotide unit along their lengths, where the polynucleotide unit includes a sequence-of-interest (e.g., thatcorresponds to target RNA or target cDNA) and at least a universal sequencing primer binding site. Each polynucleotide unit can bind a sequencing primer, a sequencing polymerase and a detectably-labeled nucleotide reagent (e.g., detectably labeled multivalent molecules), to form a detectable sequencing complex (e.g., a detectable ternary complex). Each nanoball carries numerous detectable sequencing complexes. Thus, the compact nature of the nanoballs increases the local concentration of detectably- labeled nucleotide reagents that are used during the sequencing workflow which increases the signal intensity emitted from a nanoball to give a discrete detectable signal which can be imaged as a fluorescent spot inside the sample(s). Each spot corresponds to a concatemer and each concatemer corresponds to a target RNA molecule in the sample(s). Multiple spots can be detected and imaged simultaneously in the sample(s). The DNA nanoballs having compact shape and size that produce increased signal intensity and color differentiation during sequencing.
[0355] In any of the methods described herein, the sample(s) comprises a whole cell, a plurality of whole cells, an intact tissue or an intact tumor. In some embodiments, the sample(s) comprises a fresh cellular sample, a freshly-frozen cellular sample, a sectioned cellular sample, or an FFPE cellular sample. In some embodiments, the sample(s) comprise one or more living cells or non-living cells.
[0356] In some embodiments, the sample(s) can be obtained from a virus, fungus, prokaryote or eukaryote. In some embodiments, the sample(s) can be obtained from an animal, insect or plant. In some embodiments, the sample(s) comprises one or more virally-infected cells.
[0357] In some embodiments, the sample(s) can be obtained from any organism including human, simian, ape, canine, feline, bovine, equine, murine, porcine, caprine, lupine, ranine, piscine, plant, insect or bacteria.
[0358] In some embodiments, the sample(s) can be obtained from any organ including head, neck, brain, breast, ovary, cervix, colon, rectum, endometrium, gallbladder, intestines, bladder, prostate, testicles, liver, lung, kidney, esophagus, pancreas, thyroid, pituitary, thymus, skin, heart, larynx, or other organs.
[0359] In any of the methods described herein, the sample(s) harbors a plurality of RNA which include target RNA and non-target RNA. In some embodiments, cells typically produce RNA by gene expression which includes transcription of DNA (e.g., genomic DNA) into RNA molecules. The transcribed RNA can undergo splicing or may not bespliced. The transcribed RNA can be translated into a polypeptide (e.g., coding RNA), or do not undergo translation but can be processed into tRNA or rRNA (e.g., non-coding RNA).
[0360] In some embodiments, the plurality of RNA harbored by the sample(s) includes target and non-target RNA. In some embodiments, the plurality of RNA harbored by the sample(s) comprises wild type RNA, mutant RNA or splice variant RNA. In some embodiments, the plurality of RNA harbored by the sample(s) comprises pre-spliced RNA, partially spliced RNA, or fully spliced RNA. In some embodiments, the plurality of RNA harbored by the sample(s) comprises coding RNA, non-coding RNA, mRNA, tRNA, rRNA, microRNA (miRNA), mature microRNA, or immature microRNA. In some embodiments, the plurality of RNA harbored by the sample(s) comprises housekeeping RNA, cell-specific RNA, tissue-specific RNA or disease-specific RNA. In some embodiments, the plurality of RNA harbored by the sample(s) comprises RNA expressed by one or more cells in response to a stimulus such as heat, light, a chemical or a drug. In some embodiments, the plurality of RNA harbored by the sample(s) comprises RNA found in healthy cells or diseased cells. In some embodiments, the plurality of RNA harbored by the sample(s) comprises RNA transcribed from transgenic DNA sequences that are introduced into the sample(s) using recombinant DNA procedures. For example, the RNA can be transcribed from a transgenic DNA sequence that is controlled by an inducible or constitutive promoter sequence. In some embodiments, the plurality of RNA harbored by the sample(s) comprises RNA that is transcribed from DNA sequences that are not transgenic.
[0361] In any of the methods described herein, the sample(s) can be cultured on the support. In some embodiments, the methods comprise culturing the sample(s) on the support under a condition suitable for expanding the sample(s) for 2-10 generations or more. The cultured cellular sample can generate a colony of cells. In some embodiments, the methods comprise culturing the sample(s) to confluence or non-confluence. In some embodiments, the methods comprise culturing the sample(s) on the support in a simple or complex cell culture media. For example, the cell culture media comprises D-MEM high glucose (e.g., from Thermo Fisher Scientific, catalog No. 11965118), fetal bovine serum (e.g., 10% FBS; for example from Thermo Fisher Scientific, catalog No. A3160402), MEM non-essential amino acids (e.g., 0.1 mM MEM, for example from Thermo Fisher Scientific, catalog No. 11140050), L-glutamine (e.g., 6 mM L-glutamine, for examplefrom Thermo Fisher Scientific, catalog No. A2916801), MEM sodium pyruvate (e.g., 1 mM sodium pyruvate, for example from Thermo Fisher Scientific, catalog No.11360070), and an antibiotic (e.g., 1% penicillin-streptomycin-glutamine, for example from Thermo Fisher, catalog No. 10378016). In some embodiments, the methods comprise culturing the sample(s) at a humidity and temperature that is suitable for culturing the cell(s) on the support. Exemplary suitable conditions comprise approximately 37 °C with a humidified atmosphere of approximately 5-10% carbon dioxide in air. The sample(s) can be cultured with suitable aeration with oxygen and / or nitrogen.
[0362] In any of the methods described herein, the term “simple cell media” or related terms refers to a cell media that typically lacks ingredients to support cell growth and / or proliferation in culture. Simple cell media can be used for example to wash, suspend, or dilute the sample(s). Simple cell media can be mixed with certain ingredients to prepare a cell media that can support cell growth and / or proliferation in culture. A simple cell media comprises any one or any combination of two or more of a buffer, a phosphate compound, a sodium compound, a potassium compound, a calcium compound, a magnesium compound and / or glucose. In some embodiments, the simple cell media comprises PBS (phosphate buffered saline), DPBS (Dulbecco’s phosphate-buffered saline), HBSS (Hank’s balanced salt solution), DMEM (Dulbecco’s Modified Eagle’s Medium), EMEM (Eagle’s Minimum Essential Medium), and / or EBSS. In some embodiments, the sample(s) can be placed in a simple cell media prior to or during the step of conducting any of the nucleic acid methods described herein.
[0363] In any of the methods described herein, the term “complex cell media” or related terms refers to a cell media that can be used to support cell growth and / or proliferation in culture without supplementation or additives. Complex cell media can include any combination of two or more of a buffering system (e.g., HEPES), inorganic salt(s), amino acid(s), protein(s), polypeptide(s), carbohydrate(s), fatty acid(s), lipid(s), purine(s) and their derivatives (e.g., hypoxanthine), pyrimidine(s) and their derivatives, and / or trace element(s). Complex cell media includes fluids obtained from a fluid or tissue extract. Complex cell media includes artificial cell media. In some embodiments, complex cell media can be a serum-containing media, for example complex cell media includes fluids such as fetal bovine serum, blood plasma, blood serum, lymph fluid, human placental cord serum and amniotic fluid. In some embodiments, complex cell media can be aserum-free media, which are typically (but not necessarily) defined cell culture media. In some embodiments, complex cell media can be a chemically-defined media which typically (but not necessarily) include recombinant polypeptides, and ultra-pure inorganic and / or organic compounds. In some embodiments, complex cell media can be a protein- free media which include for example MEM (minimal essential media) and RPMI-1640 (Roswell Park Memorial Institute). In some embodiments, the complex cell media comprises IMDM (Iscove’s Modified Dulbecco’s Medium. In some embodiments, the complex cell media comprises DMEM (Dulbecco’s Modified Eagle’s Medium). In some embodiments, the sample(s) can be placed in a complex cell media prior to or during the step of conducting any of the nucleic acid methods described herein.
[0364] In any of the methods described herein, the sample(s) comprises a fixed cellular sample. In some embodiments, the sample(s) can be treated with a fixation reagent (e.g., a fixing reagent) that preserves the cell and its contents to inhibit degradation and can inhibit cell lysis. For example, the fixation reagent can preserve RNA harbored by the sample(s). In some embodiments, the fixation reagent inhibits loss of nucleic acids from the sample(s).
[0365] In some embodiments, the fixation reagent can cross-link the RNA to prevent the RNA from escaping the sample(s). In some embodiments, a cross-linking fixation reagent comprises any combination of an aldehyde, formaldehyde, paraformaldehyde, formalin, glutaraldehyde, imidoesters, N-hydroxysuccinimide esters (NHS) and / or glyoxal (a bifunctional aldehyde).
[0366] In some embodiments, the fixation reagent comprises at least one alcohol, including methanol or ethanol. In some embodiments, the fixation reagent comprises at least one ketone, including acetone. In some embodiments, the fixation reagent comprises acetic acid, glacial acetic acid and / or picric acid. In some embodiments, the fixation reagent comprises mercuric chloride. In some embodiments, the fixation reagent comprises a zinc salt comprising zinc sulphate or zinc chloride. In some embodiments, the fixation reagent can denature polypeptides.
[0367] In some embodiments, the fixation reagent comprises 4% w / v of paraformaldehyde to water / PBS. In some embodiments, the fixation reagent comprises 10% of 35% formaldehyde at a neutral pH. In some embodiments, the fixation reagent comprises 2% v / v of glutaraldehyde to water / PBS. In some embodiments, the fixationreagent comprises 25% of 37% formaldehyde solution, 70% picric acid and 5% acetic acid.
[0368] In some embodiments, the sample(s) can be fixed on the support with 4% paraformaldehyde for about 30-60 minutes and washed with PBS.
[0369] In some embodiments, the sample(s) can be stained, de-stained or un-stained.
[0370] In any of the methods described herein, the sample(s) comprises a permeabilized cellular sample. In some embodiments, the methods comprise treating the sample(s) with a permeabilization reagent that alters the cell membrane to permit penetration of experimental reagents into the cells. For example, the permeabilization reagent removes membrane lipids from the cell membrane. In some embodiments, the sample(s) can be treated with a permeabilization reagent which comprises any combination of an organic solvent, detergent, chemical compound, cross-linking agent and / or enzyme. In some embodiments, the organic solvents comprise acetone, ethanol, and methanol. In some embodiments, the detergents comprise saponin, Triton X-100, Tween-20, sodium dodecyl sulfate (SDS), an N-lauroylsarcosine sodium salt solution, or a nonionic polyoxyethylene surfactant (e.g., NP40). In some embodiments, the cross-linking agent comprises paraformaldehyde. In some embodiments, the enzyme comprises trypsin, pepsin or protease (e.g. proteinase K). In some embodiments, the cells can be permeabilized using an alkaline condition, or an acidic condition with a protease enzyme. In some embodiments, the permeabilization reagent comprises water and / or PBS.
[0371] For example, the fixed cells can be permeabilized with 70% ethanol for about 30- 60 minutes, and the permeabilizing reagent can be exchanged with PBS-T (e.g., PBS with 0.05% Tween-20). In some embodiments, the cells can be post-fixed with 3% paraformaldehyde and 0.1% glutaraldehyde for about 30-60 minutes, and washed with PBS-T multiple times.
[0372] In any of the methods described herein, the sample(s) is infused with a swellable polyelectrolyte hydrogel (U.S. patent No. 10,309,879 and Chen 2015 Science 347:543, the contents of these documents are incorporated by reference in their entireties). In some embodiments, a fixed and permeabilized cellular sample can be infused with sodium acrylate, acrylamide and a cross-linker N-N’ -methylenebisacrylamide. In some embodiments, ammonium persulfate (APS) initiator and tetramethylethylenediamine (TEMED) accelerator were infused to achieve polymerization. In some embodiments, the sample(s) can be infused with proteinase K for proteolysis and incubated in a digestionbuffer. In some embodiments, the gel inside the sample(s) can be swelled by addition of water.
[0373] In any of the methods described herein, the plurality of RNAs inside cellular sample can be converted to cDNA. In some embodiments, the methods comprise contacting the plurality of RNA inside the fixed and permeabilized cellular sample with (i) a plurality of reverse transcription primers, (ii) a plurality of reverse transcriptase enzymes, and (iii) a plurality of nucleotides, under a condition suitable for conducting a reverse transcription reaction to generate a plurality of cDNA molecules (e.g., a plurality of first strand cDNA molecules) in the sample(s). In some embodiments, synthesis of second strand cDNA molecules is omitted. In some embodiments, the RNA inside the sample(s) is not converted into cDNA, where the RNA is hybridized to target-specific padlock probes.
[0374] In some embodiments, the reverse transcriptase enzyme exhibits RNA-dependent DNA polymerase activity. In some embodiments, the reverse transcriptase enzyme comprises a reverse transcriptase enzyme from AMV (avian myeloblastosis virus), M- MuLV (moloney murine leukemia virus), or HIV (human immunodeficiency virus). In some embodiment, the reverse transcriptase enzyme comprises a recombinant enzyme that exhibits reduced RNase H activity, for example REVERTAID (e.g., from Thermo Fisher Scientific, catalog No. EP0441). In some embodiments, the reverse transcriptase can be a commercially-available enzyme, including MULTISCRIBE (e.g., from Thermo Fisher Scientific, catalog # 4311235), THERMOSCRIPT (e.g., from Thermo Fisher Scientific, catalog # 12236-014), or ARRAYSCRIPT (e.g., from Ambion, catalog No. AM2048). In some embodiments, the reverse transcriptase enzyme comprises SUPERSCRIPT II (e.g., catalog No. 18064014), SUPERSCRIPT III (e g., catalog No.18080044), or SUPERSCRIPT IV enzymes (e.g., catalog No. 18090010 ) (all SUPERSCRIPT enzymes from Invitrogen). In some embodiments, the reverse transcription reaction can include an RNase inhibitor.
[0375] In some embodiments, the reverse transcription primers comprise a singlestranded oligonucleotide comprising DNA, RNA, or chimeric DNA / RNA. In some embodiments, the reverse transcription primers Any combination of adenine (A), thymine (T), guanine (G), cytosine (C), uracil (U) and / or inosine (I). In some embodiments, the reverse transcription primers can be any length, for example 5-25 bases, or 25-50 bases, or 50-75 bases, or 75-100 bases in length or longer. The reverse transcription primerseach comprise a 5’ end and 3’ end. In some embodiments, the 3’ end of the reverse transcription primers can include a 3’ OH moiety which serves as a nucleotide polymerization initiation site in a polymerase-catalyzed primer extension reaction. In some embodiments, the 3’ end of the reverse transcription primers have a chain terminating moiety which blocks a polymerase-catalyzed primer extension reaction. The chain terminating moiety can be removed to convert the 3’ sugar position to an extendible 3 ’OH.
[0376] In some embodiments, the reverse transcription primers are modified to confer resistance to nuclease degradation (e.g., ribonuclease degradation). For example, the reverse transcription primers comprise at least one phosphorothioate diester bond at their 5’ ends which can render the reverse transcription primers resistant to nuclease degradation. In some embodiments, the reverse transcription primers comprise 2-5 or more consecutive phosphorothioate diester bonds at their 5’ ends. In some embodiments, the plurality of reverse transcription primers comprise at least one ribonucleotide and / or at least one 2’-O-methyl, 2’-O-methoxyethyl (MOE), 2’ fluoro-base nucleotide. In some embodiments, the reverse transcription primers comprise phosphorylated 3’ ends. In some embodiments, the reverse transcription primers comprise locked nucleic acid (LNA) bases. In some embodiments, the reverse transcription primers comprise a phosphorylated 5’ end (e.g., using a polynucleotide kinase).
[0377] In some embodiments, the entire length of a reverse transcription primer can hybridize to a portion of an RNA molecule. In some embodiments, individual reverse transcription primers comprise a 3’ region having a sequence that hybridizes to a portion of an RNA molecule and a 5’ region that carries a tail that does not hybridize to an RNA molecule. In some embodiments, the 5’ tail comprises a universal adaptor sequence including any one or any combination of two or more of a sample barcode sequence, an amplification primer binding site, a sequencing primer binding site, a compaction oligonucleotide binding site and / or a surface capture primer binding site. In some embodiments, the 5’ tail comprises a unique identification sequence (e.g., unique molecular index (UMI). In some embodiments, the 5’ tail comprises a restriction enzyme recognition sequence. In some embodiments, individual reverse transcription primers comprise at least a portion of the 3’ region having a homopolymer sequence, for example poly-A, poly-T, poly-C, poly-G or poly-U. In some embodiments, the reverse- Ill -transcription primers can hybridize to any portion of an RNA molecule, including the 5’ or the 3’ end of the RNA molecule, or an internal portion of the RNA molecule.
[0378] In some embodiments, the plurality of reverse transcription primers comprises a first sub-population of target-specific reverse transcription primers that hybridize selectively to the first target RNA (e.g., targeted transcriptomics). In some embodiments, the plurality of reverse transcription primers further comprise a second sub-population of target-specific reverse transcription primers that hybridize selectively to the second target RNA. In some embodiments, the target-specific reverse transcription primers comprise a pre-determined sequence at the 3’ region which hybridizes to a target RNA molecule. In some embodiments, the pre-determined sequence portion of the reverse transcription primers can be 4-20 bases, or 20-40 bases, or 40-50 bases in length.
[0379] In some embodiments, the first sub-population of target-specific reverse transcription primers can selectively hybridize to an RNA transcribed in the sample(s) by a housekeeping gene. In some embodiments, selection of the housekeeping gene may be dependent upon the type of cellular sample to be used for the in situ methods described herein. Exemplary housekeeping genes include glyceraldehyde-3 -phosphate dehydrogenase (GAPDH), beta-actins (ACTB), tubulins, PPIA (peptidyl-prolyl cis-trans isomerase), NME4 (NME / NM23 nucleoside diphosphate kinase 4), SMARCAL1 (SWI / SNF related matrix associated actin dependent regulator of chromatin, subfamily A like 1), and POMK (protein-O-mannose kinase). The skilled artisan can design the first sub-population of target-specific reverse transcription primers to hybridize to RNA transcripts from any of the numerous housekeeping genes.
[0380] In some embodiments, the second sub-population of target-specific reverse transcription primers can selectively hybridize to an RNA transcribed from a gene that is expressed in the sample(s) being examined (e.g., a cell-specific or tissue-specific RNA).
[0381] In some embodiments, the plurality of reverse transcription primers comprises a first sub-population of random-sequence reverse transcription primers that hybridize to the first target RNA (e.g., whole transcriptomics). In some embodiments, the plurality of reverse transcription primers further comprises a second sub-population of randomsequence reverse transcription primers that hybridize to the second target RNA. In some embodiments, the reverse transcription primers comprise a random and / or degenerate sequence at the 3’ region which hybridizes to an RNA molecule. In some embodiments,the random-sequence or the degenerate-sequence portion of the reverse transcription primers can be 4-20 bases, or 20-40 bases, or 40-50 bases in length.Sequencing Polymerases
[0382] In any of the methods described herein, sequencing polymerases can be used for conducting sequencing reactions. In some embodiments, the sequencing polymerase(s) is / are capable of binding and incorporating a complementary nucleotide opposite a nucleotide in a concatemer template molecule. In some embodiments, the sequencing polymerase(s) is / are capable of binding a complementary nucleotide unit of a multivalent molecule opposite a nucleotide in a concatemer template molecule. In some embodiments, the plurality of sequencing polymerases comprise recombinant mutant polymerases.
[0383] Examples of suitable polymerases for use in sequencing with nucleotides and / or multivalent molecules include but are not limited to: Klenow DNA polymerase; Thermus aquaticus DNA polymerase I (Taq polymerase); KlenTaq polymerase; Candidatus altiarchaeales archaeon; Candidatus Hadarchaeum Yellowstonense; Hadesarchaea archaeon; Euryarchaeota archaeon; Thermoplasmata archaeon; Thermococcus polymerases such as Thermococcus litoralis, bacteriophage T7 DNA polymerase; human alpha, delta and epsilon DNA polymerases; bacteriophage polymerases such as T4, RB69 and phi29 bacteriophage DNA polymerases; Pyrococcus furiosus DNA polymerase (Pfu polymerase); Bacillus subtilis DNA polymerase III; E. coli DNA polymerase III alpha and epsilon; 9 degree N polymerase; reverse transcriptases such as HIV type M or O reverse transcriptases; avian myeloblastosis virus reverse transcriptase; Moloney Murine Leukemia Virus (MMLV) reverse transcriptase; or telomerase. Further non-limiting examples of DNA polymerases include those from various Archaea genera, such as, Aeropyrum, Archaeglobus, Desulfurococcus, Pyrobaculum, Pyrococcus, Pyrolobus, Pyrodictium, Staphylothermus, Stetteria, Sulfolobus, Thermococcus, and Vulcanisaeta and the like or variants thereof, including such polymerases as are known in the art such as 9 degrees N, VENT, DEEP VENT, THERMINATOR, Pfu, KOD, Pfx, Tgo and RB69 polymerases.Sequencing-by-Binding
[0384] In any of the methods described herein, the sequencing comprises conducting sequencing-by-binding (SBB) reactions inside the sample(s), where the cDNA amplicons are the concatemer molecules. In some embodiments, the sequencing-by-binding (SBB) procedure employs non-labeled chain-terminating nucleotides. In some embodiments, a cycle of sequencing-by-binding (SBB) comprises the steps of (a) sequentially contacting a primed concatemer (e.g., a concatemer annealed to a plurality of sequencing primers) with at least two separate mixtures under ternary complex stabilizing conditions, wherein the at least two separate mixtures each include a polymerase and a nucleotide, whereby the sequentially contacting results in the primed concatemer being contacted, under the ternary complex stabilizing conditions, with nucleotide cognates for first, second and third base type base types in the template; (b) examining the at least two separate mixtures to determine whether a ternary complex formed; and (c) identifying the next correct nucleotide for the primed concatemer, wherein the next correct nucleotide is identified as a cognate of the first, second or third base type if ternary complex is detected in step (b), and wherein the next correct nucleotide is imputed to be a nucleotide cognate of a fourth base type based on the absence of a ternary complex in step (b); (d) adding a next correct nucleotide to the primer of the primed concatemer after step (b), thereby producing an extended primer; and (e) repeating steps (a) through (d) at least once on the primed concatemer that comprises the extended primer. Exemplary sequencing-by- binding methods are described in U.S. patent Nos. 10,246,744 and 10,731,141 (where the contents of both patents are hereby incorporated by reference in their entireties).Nucleotides and Chain-Terminating Nucleotides
[0385] In any of the methods described herein, any of the sequencing methods described herein can employ at least one nucleotide. The nucleotides comprise a base, sugar and at least one phosphate group. In some embodiments, at least one nucleotide in the plurality comprises an aromatic base, a five carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1-10 phosphate groups). The plurality of nucleotides can comprise at least one type of nucleotide selected from a group consisting of dATP, dGTP, dCTP, dTTP and dUTP. The plurality of nucleotides can comprise at a mixture of any combination of two or more types of nucleotides selected from a group consisting of dATP, dGTP, dCTP, dTTP and / or dUTP. In some embodiments, at least one nucleotide inthe plurality is not a nucleotide analog. In some embodiments, at least one nucleotide in the plurality comprises a nucleotide analog.
[0386] In some embodiments, in any of the methods for sequencing described herein, at least one nucleotide in the plurality of nucleotides comprise a chain of one, two or three phosphorus atoms where the chain is typically attached to the 5’ carbon of the sugar moiety via an ester or phosphoramide linkage. In some embodiments, at least one nucleotide in the plurality is an analog having a phosphorus chain in which the phosphorus atoms are linked together with intervenin...
Claims
WHAT IS CLAIMED IS:
1. A sequencing analysis method comprising:obtaining, by a sequencing system, a first set of flow cell images corresponding to sensor data acquired by one or more image sensors of the sequencing system of one or more samples immobilized on a flow cell device in a first cycle of a sequencing run; determining, by the sequencing system, a set of preliminary intensities of a first set of polonies in the first set of flow cell images based on a preliminary registration of the first set of polonies to a polony map;determining, by the sequencing system, preliminary base calls of the first set of polonies based on the set of preliminary intensities;identifying, by the sequencing system, a first subset and a second subset of polonies of the first set of polonies based on the preliminary base calls;determining, by the sequencing system, a first set of intensities of the first set of polonies based on registration of the first subset of polonies to the polony map;determining, by the sequencing system, a second set of intensities of the first set of polonies based on registration of the second subset of polonies to the polony map; selecting, by the sequencing system, one of: the set of preliminary intensities, the first set of intensities, and the second set of intensities based on evaluation of quality of one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities; anddetermining, by the sequencing system, aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the selection.
2. A sequencing analysis method comprising:obtaining, by a sequencing system, a first set of flow cell images corresponding to sensor data acquired by one or more image sensors of the sequencing system of one or more samples immobilized on a flow cell device in a first cycle of a sequencing run; determining, by the sequencing system, a set of preliminary intensities of a first set of polonies in the first set of flow cell images based on a preliminary registration of the first set of polonies to a polony map;determining, by the sequencing system, preliminary base calls of the first set of polonies based on the set of preliminary intensities;identifying, by the sequencing system, a first subset of polonies of the first set of polonies based on the preliminary base calls;determining, by the sequencing system, a first set of intensities of the first set of polonies based on registration of the first subset of polonies to the polony map;selecting, by the sequencing system, one of: the set of preliminary intensities and the first set of intensities based on evaluation of quality of the set of preliminary intensities and the first set of intensities; anddetermining, by the sequencing system, aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the selection.
3. A sequencing analysis method comprising:obtaining, by a sequencing system, a first set of flow cell images corresponding to sensor data acquired by one or more image sensors of the sequencing system of one or more samples immobilized on a flow cell device in a first cycle of a sequencing run; determining, by the sequencing system, a set of preliminary intensities of a first set of polonies in the first set of flow cell images based on a preliminary registration of the first set of polonies to a polony map;determining, by the sequencing system, preliminary base calls of the first set of polonies based on the set of preliminary intensities;identifying, by the sequencing system, a first subset of polonies of the first set of polonies based on the preliminary base calls;determining, by the sequencing system, a first set of intensities of the first set of polonies based on registration of the first subset of polonies to the polony map;determining, by the sequencing system, a set of shift intensities of the first set of polonies based on: the first set of intensities; the first set of intensities and the set of preliminary intensities; first locations corresponding to the first set of intensities, preliminary locations corresponding to the set of preliminary intensities, or a combination thereof; anddetermining, by the sequencing system, aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the shifted intensities.
4. A sequencing analysis method comprising:obtaining, by a sequencing system, a first set of flow cell images corresponding to sensor data acquired by one or more image sensors of the sequencing system of one or more samples immobilized on a flow cell device in a first cycle of a sequencing run; determining, by the sequencing system, a set of preliminary intensities of a first set of polonies in the first set of flow cell images based on a preliminary registration to a polony map;determining, by the sequencing system, preliminary base calls based on the set of preliminary intensities of the first set of flow cell images;identifying, by the sequencing system, a first subset and a second subset of polonies of the first set of polonies based on the preliminary base calls;determining, by the sequencing system, first locations of the first set of polonies based on registration of the first subset of polonies to the polony map;determining, by the sequencing system, second locations of the first set of polonies based on registration of the second subset of polonies to the polony map;determining, by the sequencing system, shifted intensities of the first set of polonies based on the first and second locations; anddetermining, by the sequencing system, aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the shifted intensities.
5. The method of any one of the preceding claims, wherein the sensor data at one of the one or more image sensors comprises emitted signal from the one or more samples at more than one wavelength ranges.
6. The method of any one of the preceding claims, wherein the sensor data at one of the one or more image sensors comprises emitted signal from the one or more samples with more than one color.
7. The method of any one of the preceding claims, wherein the sensor data at one of the one or more image sensors comprises emitted signal with two or more colors selected from: blue, red, green, and yellow.
8. The method of any one of the preceding claims, wherein at least one of the one or more image sensors is configured to sense emitted signal of the one or more samples with more than one color.
9. The method of any one of the preceding claims, wherein a first polony of the first set of polonies in the flow cell image comprises a first level of chromatic aberration, and wherein a second polony of the first set of polonies comprises a second level of chromatic aberration that is different from the first level of chromatic aberration.
10. The method of any one of the preceding claims, wherein the first or second level of chromatic aberration is less than 0.1 pixels, 0.2 pixels, 0.5 pixels, 0.8 pixels, 1 pixel, 1.5 pixels, 2 pixels, 3 pixels, 4 pixels, or 5 pixels.
11. The method of any one of the preceding claims, wherein the aberration-corrected base calls comprises at least 1%, 5%, 8%, 10%, 12%, 15%, 18%, or 20% less base calling errors than the preliminary base calls.
12. The method of any one of the preceding claims, wherein the first set of flow cell images comprises 2 or 3 flow cell images per FOV, and wherein each flow cell image corresponds to different emitted signal from the one or more samples in response to a different illumination.
13. The method of any one of the preceding claims, wherein the first set of flow cell images comprises 2 or 3 flow cell images, and wherein each flow cell image corresponds to a different color channel.
14. The method of any one of the preceding claims, wherein obtaining the first set of flow cell images corresponding to the sensor data acquired by the one or more image sensors of the sequencing system is performed while a sequencing run is in progress.
15. The method of any one of the preceding claims, wherein the sequencing analysis method further comprises:acquiring, by of the sequencing system, at least one flow cell image of the first set of flow cell images in a first color channel of the one or more samples immobilized on the flow cell device in the first cycle of the sequencing run.
16. The method of any one of the preceding claims, wherein the sequencing analysis method further comprises:acquiring, by of the sequencing system, at least one full flow cell image of a first set of full flow cell images in a first color channel of the one or more samples immobilized on the flow cell device in the first cycle of the sequencing run; and obtaining, by the sequencing system, the first set of flow cell images by obtaining a portion of the at least one full flow cell image.
17. The method of any one of the preceding claims, wherein acquiring, by the imager of the sequencing system, the first set of flow cell images of the one or more samples immobilized on the flow cell device in the first cycle of the sequencing run comprises: illuminating, by a first illuminator of the sequencing system, the one or more samples to cause emission of a signal of at least a first color and a second color simultaneously;allowing the emitted signal of at least the first color and second color to travel through an identical optical path of the imager to arrive at the one or more image sensors of the sequencing system;detecting the emitted signal by the one or more image sensors of the sequencing system; andgenerating, by the sequencing system, a first flow cell image of the first set of flow cell images based on the detected signal.
18. The method of any one of the preceding claims, wherein acquiring, by the imager of the sequencing system, the first set of flow cell images of the one or more samples immobilized on the flow cell device in the first cycle of the sequencing run further comprises:illuminating, by the first or a second illuminator of the sequencing system, the one or more samples to cause emission of a signal of at least a third color;allowing the emitted signal of at least the third color to travel through the identical optical path of the imager to arrive at the one or more image sensors of the sequencing system;detecting the emitted signal by the one or more image sensors of the sequencing system; andgenerating, by the sequencing system, a second flow cell image of the first set of flow cell images based on the detected signal.
19. The method of any one of the preceding claims, wherein obtaining the first set of flow cell images corresponding to the sensor data acquired by the one or more image sensors of the sequencing system is performed in the first cycle or in a second cycle of the sequencing run subsequent to the first cycle.
20. The method of any one of the preceding claims, wherein the first or second cycle is among the first 1, 2, 3, 5, 10, 15, or 20 cycles of the sequencing run.
21. The method of any one of the preceding claims, wherein the set of preliminary intensities are from two, three, or four color channels.
22. The method of any one of the preceding claims, wherein determining, by the sequencing system, a set of preliminary intensities of the first set of polonies based on the preliminary registration of the first set of polonies to the polony map comprises:extracting, by the sequencing system, the set of preliminary intensities from the preliminarily registered first set of flow cell images based on locations of the first set of polonies in the polony map.
23. The method of any one of the preceding claims, wherein identifying, by the sequencing system, the first subset of polonies of the first set of polonies based on the preliminary base calls comprises:identifying, by the sequencing system, the first subset of polonies of the first set of polonies based on a first type of base call among the preliminary base calls.
24. The method of any one of the preceding claims, wherein determining, by the sequencing system, the first set of intensities of the first set of polonies based on registration of the first subset of polonies to the polony map comprises:registering, by the sequencing system, the first subset of polonies to the polony map, thereby generating a first set of registered flow cell images; andextracting, by the sequencing systems, the first set of intensities of the first set of polonies from the first set of registered flow cell images based on locations of the first set of polonies in the polony map.
25. The method of any one of the preceding claims, wherein determining, by the sequencing system, the second set of intensities of the first set of polonies based on registration of the second subset of polonies to the polony map comprises:registering, by the sequencing system, the second subset of polonies to the polony map, thereby generating a second set of registered flow cell image; andextracting, by the sequencing systems, the second set of intensities of the first set of polonies from the second set of registered flow cell images based on locations of the first set of polonies in the polony map.
26. The method of any one of the preceding claims, wherein each of the first set of registered flow cell images corresponds to a different color channel.
27. The method of any one of the preceding claims, wherein each of the second set of registered flow cell images corresponds to a different color channel.
28. The method of any one of the preceding claims, wherein each subset of preliminary intensities in the set of preliminary intensities, each subset of first intensities in the first set of intensities, and each subset of second intensities in the second set of intensities correspond to a different color channel.
29. The method of any one of the preceding claims, wherein evaluation of quality of one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities comprises:evaluation of quality of the set of preliminary intensities, the first set of intensities, and the second set of intensities.
30. The method of any one of the preceding claims, wherein evaluation of quality of one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities comprises:evaluation of quality of the first set of intensities and the second set of intensities.
31. The method of any one of the preceding claims, wherein evaluation of quality of one or more of: the set of preliminary intensities, the first set of intensities, and the second set of intensities comprises:calculating, by the sequencing system, a preliminary quantitative value for a quality metrics for each individual polony of the first set of polonies;calculating, by the sequencing system, a first quantitative value for the quality metrics for each individual polony of the first set of polonies; and / orcalculating, by the sequencing system, a second quantitative value for the quality metrics for each individual polony of the first set of polonies.
32. The method of any one of the preceding claims, wherein the quality metrics comprises clarity or a quality score.
33. The method of any one of the preceding claims, wherein determining, by the sequencing system, the aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the quality evaluation comprises:identifying a maximum value for each individual polony among one or more of: the preliminary quantitative value; the first quantitative value; and second quantitative value;determining an aberration-corrected intensity for each individual polony of the first set of polonies as one of: the preliminary intensity, the first intensity, or the second intensity corresponding to the maximum value; anddetermining, by the sequencing system, aberration-corrected base calls of each individual polony based on the aberration-corrected intensity.
34. The method of any one of the preceding claims, wherein determining, by the sequencing system, the aberration-corrected base calls of the first set of polonies in the first set of flow cell images based on the quality evaluation comprises:determining a first sum of the preliminary quantitative value; a second sum of the first quantitative value; and a third sum of second quantitative value for the first set of polonies;identifying a maximum for the first set of polonies among the first, second, and third sums;determining an aberration-corrected intensity for each individual polony of the first set of polonies as one of: the preliminary intensity, the first intensity, or the second intensity corresponding to the maximum; anddetermining, by the sequencing system, aberration-corrected base calls of each individual polony based on the aberration-corrected intensity.
35. The method of any one of the preceding claims, wherein determining, by the sequencing system, the shifted intensities of the first set of polonies based on registration of the first subset of polonies to the polony map comprises:determining, by the sequencing system, preliminary locations of the first set of polonies based on registration of the first set of polonies to the polony map;determining, by the sequencing system, first locations of the first set of polonies based on registration of the first subset of polonies to the polony map; and determining, by the sequencing system, shifted locations based on the preliminary locations and the first locations of the first set of polonies.
36. The method of any one of the preceding claims, wherein determining, by the sequencing system, the shifted intensities of the first set of polonies based on registration of the first subset of polonies to the polony map comprises:determining, by the sequencing system, first locations of the first set of polonies based on registration of the first subset of polonies to the polony map; and determining, by the sequencing system, second locations of the first set of polonies based on registration of the second subset of polonies to the polony map; and determining, by the sequencing system, a shifted location of each individual polony based on the first locations and the second locations of the first set of polonies.
37. The method of any one of the preceding claims, wherein the method further comprising: obtaining, by a sequencing system, high resolution images of the first set of flow cell images.
38. The method of any one of the preceding claims, wherein the first set of flow cell images comprises a first resolution and wherein the high resolution images, the preliminary intensities, and the preliminary base calls are at the first resolution or at a second resolution that is higher than the first resolution.
39. The method of any one of the preceding claims, wherein each flow cell image of the first set of flow cell images comprises a field of view of greater than 0.1 mm2, 0.2 mm2, 0.3 mm2, 0.4 mm2, 0.5 mm2, 0.6 mm2, 0.7 mm2, 0.8 mm2, 1 mm2, 2 mm2, 3 mm2, 4 mm2, 5 mm2, 6 mm2, 7 mm2, 8 mm2, or 10 mm2.
40. The method of any one of the preceding claims, wherein each flow cell image of the first set of flow cell images comprises at least 1 x 105pixels, 5 x 105pixels, lx 106pixels, 2 x 106pixels, 4 x 106pixels, 8 x 106pixels, 16 x 106pixels, 64 x 106pixels, or 256 x 106pixels.
41. The method of any one of the preceding claims, wherein each flow cell image of the first set of flow cell images comprises an image resolution of at least 0.05 um, 0.1 um, 0.2 um, 0.4 um, 0.5 um, 0.6 um, 0.8 um, 1 um, 1.5 um, 2 um, 3 um, or 5 um.
42. The method of any one of the preceding claims, wherein the first resolution is at least 0.05 um, 0.1 um, 0.2 um, 0.4 um, 0.5 um, 0.6 um, 0.8 um, 1 um, 1.5 um, 2 um, 3 um, or 5 um.
43. The method of any one of the preceding claims, wherein the second resolution is at least 2x, 4x, 8x, 16x, or 32x in one spatial dimension of the first resolution.
44. The method of any one of the preceding claims, wherein each flow cell image of the first set of flow cell images covers a portion of at least O.OOOlx, O.OOlx, 0.002x, 0.005x,O.Olx, 0.02x, O.O5x, or O.lx of each full flow cell image of a first set of full flow cell image.
45. The method of any one of the preceding claims, wherein the sequencing system comprises a single image sensor.
46. The method of any one of the preceding claims, wherein the sequencing system lacks an emission filter in an optical pathway from a sample to the one or more image sensors.
47. The method of any one of the preceding claims, wherein the one or more sensors comprises a single sensor.
48. The method of any one of the preceding claims, wherein the sequencing system comprises an imager that comprises a single image sensor, and wherein the imager lacks an emission filter in an optical path from the one or more samples to the single image sensor.
49. The method of any one of the preceding claims, wherein the polony map comprises the first set of polonies and their corresponding locations in a reference coordinate system.
50. The method of any one of the preceding claims, wherein the polony map comprises a list of information of at least some polonies of the first set of polonies of polonies, and wherein the information comprises one or more of: a unique identification of an individual polony, a 3D coordinate of the individual polony in a reference coordinate system, and a 2D coordinate of the individual polony in a reference coordinate system.
51. The method of any one of the preceding claims, wherein the polony map comprises one or more individual maps, each individual map corresponds to a different FOV covering at least a portion of a full flow cell image.
52. The method of any one of the preceding claims, wherein the sequencing system comprises: a processor, a reconfigurable logic device, an integrated circuit, or a combination thereof.
53. The method of any one of the preceding claims, wherein the sequencing system comprises: a CPU, a GPU, a FPGA, an Al chip, a NPU, a TPU, or a combination thereof.
54. The method of any one of the preceding claims, wherein the sequencing system lacks any GPU.
55. The method of any one of the preceding claims, wherein the one or more samples comprise a 2D sample comprising template molecules.
56. The method of any one of the preceding claims, wherein the one or more samples comprise a 3D sample comprising concatemer molecules.
57. The method of any one of the preceding claims, wherein the one or more samples comprise a cellular sample comprising in situ cells or tissue.
58. The method of any one of the preceding claims, wherein at least part of the one or more samples comprise predetermined bases in the one or more cycles.
59. The method of any one of the preceding claims, wherein the one or more samples comprise overloaded concatemer molecules with a spatial density in a range of 102-10152per mm .
60. The method of any one of the preceding claims, wherein the one or more samples comprise overloaded concatemer molecules with a spatial density in a range of 103-IO102per mm .
61. The method of any one of the preceding claims, wherein the one or more samples comprises unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules in one or more cycles, and wherein the unbalanced diversity comprises: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of all bases; and the percentage is less than 20%, 15%, 10%, or 5% in the one or more cycles.
62. The method of any one of the preceding claims, wherein the flow cell device comprises primer molecules with a surface density in a range from 100 molecules per um2to 106molecules per um2.
63. The method of any one of the preceding claims, wherein the flow cell device comprises template molecules immobilized thereon at a surface density of 102to 1015sites per mm2.
64. The method of any one of the preceding claims, wherein the one or more samples comprise a polony density in a range from 102to 1015polonies per mm2.
65. The method of any one of the preceding claims, wherein the one or more samples comprise a polony density in a range from 103to IO10polonies per mm2.
66. The method of any one of the preceding claims is performed while the sequencing run is in progress and wherein the first cycle of the sequencing run has been completed.
67. The method of any one of the preceding claims is performed while the sequencing run is in progress and wherein the first cycle is a current cycle in progress and has not been completed.
68. The method of any one of the preceding claims, wherein the first set of flow cell images is of a first resolution, and wherein obtaining, by a sequencing system, a first set of flow cell images corresponding to sensor data acquired by one or more image sensors of the sequencing system of one or more samples immobilized on a flow cell device in a first cycle of a sequencing run comprises:generating, by a processor of the sequencing system, the first set of flow cell images with a second resolution higher than the first resolution.
69. The method of any one of the preceding claims, wherein generating the first set of flow cell images with the second resolution higher than the first resolution is by using a neural network based algorithm.
70. A sequencing system comprising:a flow cell configured to hold one or more samples immobilized thereon;an imager comprising:one or more illuminators configured to illuminate the one or more samples with less than four different colors; andone or more image sensors configured to sense corresponding emitted signal from the one or more samples when excited by the one or more illuminators with each color of the less than four different colors, wherein the imager lacks an emission filter in an optical path from the one or more samples to the one or more image sensors;one or more hardware processors; andone or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform one or more operations of any one of the preceding claims.
71. A sequencing system comprising:a flow cell configured to hold one or more samples immobilized thereon;an imager comprising:one or more illuminators configured to illuminate the one or more samples with less than four different colors; anda single image sensor configured to sequentially sense corresponding emitted signal from the one or more samples when excited by the one or more illuminators with each color of the less than four different colors;one or more hardware processors; andone or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform one or more operations of any one of the preceding claims.