Three-dimensional base calling in next generation sequencing analysis
A neural network-based approach enhances the spatial resolution of flow cell images to improve the detection and accuracy of base calling in three-dimensional samples by increasing polony density and reducing interference, addressing the limitations of traditional two-dimensional analysis.
Patent Information
- Application Number
- PCT/US2025/014022
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-20
- Filing Date
- 2025-01-31
- Publication Date
- 2025-08-07
AI Technical Summary
Traditional sequencing data analysis relies on two-dimensional flow cell images, which fail to accurately differentiate polonies or clusters from background noises and out-of-focus signals in three-dimensional samples like cells or tissue, leading to inaccurate base calling.
Employing a neural network, such as a convolutional neural network, to process high-resolution z-stacks of flow cell images, enhancing spatial resolution and detectable density of polonies or clusters, and reducing the impact of color mixing by computationally increasing spatial resolution.
Improves the accuracy and reliability of base calling in three-dimensional samples by increasing polony or cluster detection density and reducing interference from neighboring clusters, while reducing computational time and energy consumption.
Smart Images

Figure IMGF000155_0001 
Figure 00000367_0000 
Figure 00000368_0000
Abstract
Description
THREE-DIMENSIONAL BASE CALLING IN NEXT GENERATION SEQUENCING ANALYSISCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 549,327, filed Feb. 2, 2024, U.S. Provisional Application No. 63 / 549,333, filed Feb. 2, 2024, U.S. Provisional Application No. 63 / 570,038, filed Mar. 26, 2024, U.S. Provisional Application No. 63 / 661,332 , filed Jun. 18, 2024, U.S. Provisional Application No. 63 / 724,712, filed Nov. 25, 2024, and U.S. Provisional Application No. 63 / 736,743, filed Dec. 20, 2024, each of which is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] Embodiments of this disclosure relate generally to image processing and base calling in sequencing data analysis, and particularly to three-dimensional (3D) images of in situ samples.BACKGROUND
[0003] In next-generation sequencing (NGS) or NGS-like applications such as sequencing by synthesis, sequencing by binding, or sequencing by avidity, in order to identify the sequence of a target nucleic acid, a new strand is synthesized one nucleotide base at a time. During each sequencing cycle, one base attaches to any given strand. At the imaging step of each cycle, image(s) are recorded. A base-calling algorithm is applied to the image(s) to “read” the successive signals from each cluster or polony and convert the optical signals into an identification of the nucleotide base sequence added to each DNA fragment. Traditional sequencing data analysis relies on two-dimensional (2D) flow cell images. When it comes to sequencing analysis of in situ samples such as cells or tissue, the sample has a thickness along the z direction orthogonal to the image plane. As such, flow cell images at a selected z level can include signals from out-of-focus polonies located at adjacent z levels and other undesired signals, e.g., from the cell membrane. There is a need for fast and accurate three-dimensional (3D) flow cell image processing and base calling to ensure reliable base calling and sequencing analysis of 3D samples such as cells and tissue.BRIEF SUMMARY
[0004] Provided herein are system, apparatus, method, and / or computer program product embodiments, and / or combinations and sub-combinations thereof which enables fast and accurate flow cell image processing for reliable and accurate base calling of samples such as in situ cells or tissue. The flow cell images can come from different sequencing cycles and / or different channels.
[0005] As a particular application of such, embodiments of methods, systems, and media for image processing of flow cell images of 3D volumetric samples, e.g., cells or tissue, so that the image intensity, location and / or size of clusters or polonies can be relied upon for accurate base calling. The image processing methods herein may function to reverse the imaging process of an optical system and virtually improve the full width half maximum (FWHM) of the optical system. As such, the image processing methods disclosed herein may advantageously increase detectable density of polonies or clusters in 3D samples or traditional 2D samples. The methods herein may advantageously lessen the impact of color mixing of polonies that may be caused by neighboring polonies in 2D or 3D dimensions by computationally increasing the spatial resolution of the flow cell images.
[0006] In some embodiments, a neural network, e.g., a convolutional neural network, is used in generating a high-resolution z-stack of flow cell images of the 3D sample from the low-resolution z-stack that has been acquired from the sequencing system, and subsequent primary analysis can be performed based on the high-resolution flow cell images instead of the low-resolution flow cell images. In some embodiments, the neural network, e.g., a convolutional neural network, is used in image processing of the high- resolution z-stacks of flow cell images of the samples to generate the base callings.
[0007] Embodiments of these aspects include corresponding computer systems, apparatus, and computer program product recorded on computer storage device(s), which, alone or in combination, configured to perform the operations of the methods. For a computer system configured or to be configured to perform operations, the computer system has installed on it software, firmware, hardware, or their combinations that in operation cause the computer system to perform the operations or actions. For a computer program product configured or to be configured to perform operations or actions, thecomputer program product includes instructions that, when executed, by a hardware processor, cause the hardware processor to perform the operations or actions.
[0008] Further embodiments, features, and advantages of the present disclosure, as well as the structure and operation of the various embodiments of the present disclosure, are described in detail below with reference to the accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the disclosure and to enable a person skilled in the art(s) to make and use the embodiments.
[0010] FIG. 1 illustrates a block diagram of a sequencing system for performing sequencing, flow cell image processing, and / or primary analysis operations including base calling using flow cell images, according to some embodiments.
[0011] FIGS. 2A-2C show an exemplary simulated flow cell image (FIG. 2A, of a in situ cell sample) and two different images (FIGS. 2B-2C) predicted using the systems and methods herein and corresponding to the image in FIG. 2A, according to some embodiments.. The predicted images are at different z levels.
[0012] FIGS. 2D-2E show exemplary simulated flow cell image in the reference set. The simulated flow cell images are generated using the methods herein with the first (FIG. 2D) and second (FIG. 2E) resolutions, according to some embodiments.
[0013] FIGS. 3A-3D show two exemplary flow cell images (FIGS. 3A and 3D) with multiple cells at two different z levels, and two different predicted images at different z levels (FIGS. 3B-3C) generated from the image in FIG. 3A using the systems and methods herein according to some embodiments.
[0014] FIGS. 3E-3F shows improved detection of targets per cell in the same imaging area (FIG. 3E) and fewer false positives (FIG. 3F) using the methods herein when compared with non-artificial intelligence-based methods; in this case, the targets are polonies or clusters within the cells.
[0015] FIG. 3G shows improved detection of targets in simulated flow cell images of sample(s) using the neural network herein which produces higher R2value than a traditional method.
[0016] FIG. 4 illustrates a block diagram of a computer system for performing image processing, sequencing analysis, training of neural network(s), predicting base calls, image intensities, high resolution images, and / or classifications using the pre-trained neural networks, and / or base calling, according to some embodiments.
[0017] FIG. 5A is a flow chart of an exemplary method of predicting 3D flow cell images of sequencing sample(s) and performing base calling using the 3D flow cell images, according to some embodiments.
[0018] FIG. 5B is a flow chart of an exemplary method of training a neural network that can be used to predict higher resolution flow cell images of sequencing sample(s), according to some embodiments.
[0019] FIG. 5C is a schematic showing of an exemplary embodiment of the first reconfigurable logic device, the integrated circuit, and their connection(s) to the processor of the sequencing system.
[0020] FIG. 5D is a schematic showing of an exemplary embodiment of using the first reconfigurable logic device and the integrated circuit in parallel with a sequencing run in progress within a predetermined time window.
[0021] FIG. 5E is a flow chart of an exemplary method of training a neural network, thereby generating a pre-trained neural network that can be used to predict higher resolution flow cell images of sequencing sample(s), base calls, intensities, and / or classifications, according to some embodiments.
[0022] FIG. 5F shows scatter plots for an exemplary embodiment of generating reference intensities from high resolution training flow cell images.
[0023] FIG. 6 is a schematic showing exemplary embodiments of padlock probes.
[0024] FIG. 7 is a schematic showing a workflow for generating inside a cell circularized padlock probes, comprising generating first and second cDNAs from first and second target RNA molecules (respectively), hybridizing first and second padlock probes to the first and second cDNA molecules (respectively) to generate first and second circularized padlock probes (respectively).
[0025] FIG. 8 is a schematic showing a rolling circle and sequencing workflow inside a cell, comprising generating first and second concatemers by conducting rolling circle amplification using first and second covalently closed circular molecules (respectively). The first and second concatemers are subjected to a sequencing workflow using universal sequencing primers, sequencing polymerases, and a plurality of nucleotide reagents.
[0026] FIG. 9 is a schematic showing an exemplary workflow for sequencing a concatemer that is generated inside the cell.
[0027] FIG. 10 is a schematic showing an exemplary workflow for sequencing a concatemer that is generated inside the cell.
[0028] FIG. 11 is a schematic showing an exemplary workflow for sequencing a concatemer that is generated inside the cell.
[0029] FIG. 12 is a schematic showing an exemplary workflow for sequencing a concatemer that is generated inside the cell.
[0030] FIG. 13 is a schematic showing a workflow for generating circularized padlock probes, comprising generating first and second cDNAs from first and second target RNA molecules (respectively), hybridizing first and second padlock probes to the first and second cDNA molecules (respectively) to generate first and second circularized padlock probes (respectively).
[0031] FIG. 14 is a schematic showing a rolling circle and sequencing workflow comprising generating first and second concatemers by conducting rolling circle amplification using first and second covalently closed circular molecules (respectively).
[0032] FIG. 15 is a schematic of an exemplary low binding support comprising a glass substrate and alternating layers of hydrophilic coatings which are covalently or non- covalently adhered to the glass, and which further comprises chemically-reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides).
[0033] FIG. 16 is a schematic of various exemplary configurations of multivalent molecules. Left (Class I): schematics of multivalent molecules having a “starburst” or “helter-skelter” configuration. Center (Class II): a schematic of a multivalent molecule having a dendrimer configuration. Right (Class III): a schematic of multiple multivalent molecules formed by reacting streptavidin with 4-arm or 8-arm PEG-NHS with biotin and dNTPs. Nucleotide units are designated ‘N’, biotin is designated ‘B’, and streptavidin is designated ‘ SA’ .
[0034] FIG. 17 is a schematic of an exemplary multivalent molecule comprising a generic core attached to a plurality of nucleotide-arms.
[0035] FIG. 18 is a schematic of an exemplary multivalent molecule comprising a dendrimer core attached to a plurality of nucleotide-arms.
[0036] FIG. 19 shows a schematic of an exemplary multivalent molecule comprising a core attached to a plurality of nucleotide-arms, where the nucleotide arms comprise biotin, spacer, linker and a nucleotide unit.
[0037] FIG. 20 is a schematic of an exemplary nucleotide-arm comprising a core attachment moiety, spacer, linker and nucleotide unit.
[0038] FIG. 21 shows the chemical structure of an exemplary spacer (top), and the chemical structures of various exemplary linkers, including an 11 -atom Linker, 16-atom Linker, 23-atom Linker and an N3 Linker (bottom).
[0039] FIG. 22 shows the chemical structures of various exemplary linkers, including Linkers 1-9.
[0040] FIG. 23 A shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.
[0041] FIG. 23B shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.
[0042] FIG. 23 C shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.
[0043] FIG. 23D shows the chemical structures of various exemplary linkers joined / attached to nucleotide units.
[0044] FIG. 24 shows the chemical structure of an exemplary biotinylated nucleotide- arm.
[0045] FIG. 25 is a schematic of a guanine tetrad (e.g., G-tetrad).
[0046] FIG. 26 is a schematic of an exemplary intramolecular G-quadruplex structure.
[0047] FIG. 27 shows an exemplary support with multiple tiles for immobilizing 2D or 3D sample(s) thereon for sequencing, including the cellular sample(s), according to some aspects.
[0048] FIG. 28 shows a flow chart of an exemplary method of predicting base calls of the flow cell images (e.g., of in situ samples) using the neural network disclosed herein, according to some embodiments.
[0049] FIG. 29 shows a flow chart of an exemplary method of training the neural network that can be used to predict base calls or high resolution flow cell images, according to some embodiments.
[0050] FIGS. 30A-30B show a flow cell image (FIG. 30A) and its high resolution image predicted using the neural network that is pre-trained using reference base calls. In thiscase, base calls are determined from the high resolution image using non-neural network based algorithm(s).
[0051] FIG. 31 shows a block diagram of an exemplary method of training the neural network(s) and an exemplary method of predicting high resolution flow cell images and / or predicting base calls using such pretrained neural network(s).
[0052] In the drawings, like reference numbers generally indicate identical or similar elements. Additionally, generally, the left-most digit(s) of a reference number identifies the drawing in which the reference number first appears.DETAILED DESCRIPTION
[0053] Provided herein are system, apparatus, method, and / or computer program product embodiments, and / or combinations and sub-combinations thereof which enables image processing of flow cell images, e.g., flow cell images obtained from in situ samples or traditional 2D samples in a sequencing run, to: 1) generate images with improved spatial resolution and improved detectable density of polonies or clusters and perform base calling using flow cell images with such improved spatial resolutions, and the generated images may be used for subsequent sequencing analysis including but not limited to base calling; or 2) to predict intensities, base call(s), or classifications of polonies or clusters. The techniques herein can be used while a sequence run is still in progress to improve efficiency of sequencing and sequencing analysis, reduce data storage required during sequencing and sequencing analysis, and improve accuracy and reliability of sequencing analysis. The techniques herein can be used on flow cell images obtained using various imaging and / or sequencing techniques of volumetric 3D samples and / or traditional 2D samples and / or obtained using various sequencing systems, e.g., next generation sequencing (NGS) systems. The techniques disclosed herein are useful for base calling in NGS, and NGS flow cell images will be used as the primary example herein for describing the application of these techniques. However, such image analysis techniques may also be useful in other applications where spot-detection and / or CCD imaging is used.
[0054] Traditional flow cell images can show clusters or polonies from 2D samples and base calling can be performed using their corresponding image intensities. Existing base calling algorithms may not be able to differentiate polonies or clusters from background noises or other signal spots that may have similar shape or size to polonies or clusters.There is a need for identifying polony or cluster locations for accurate and reliable base calling, preferably with improved spatial resolution than the resolution of flow cell images, especially for samples with increased polony or cluster density than traditional 2D samples. There is also a need for generating accurate and reliable intensities for polonies or clusters so that they can be used for accurate and reliable base callings. Further, there is a need to accurately predict base calling using information included in the flow cell images, such as polony locations, shapes, intensities, etc.
[0055] The techniques herein can be used for processing flow cell images (e.g., 2D or 3D) to generate accurate and reliable image intensities for polonies or clusters with improved spatial resolution thus improved maximum polony or cluster density detected in the sample(s) for accurate and reliable sequencing analysis. The technologies disclosed herein may advantageously function to reverse the imaging process of an optical system and virtually improve the full width half maximum (FWHM) of the imager so that the density of polony locations are not limited by the optical design of the sequencing systems. As such, the disclosed technologies herein may advantageously increase detected density of polonies, e.g., by 2x, 4x, 8x, 16x, 27x, 40x, 50x, lOOx or more than polony density detectable using traditional optical systems and image processing methods. In some embodiments, the disclosed technologies herein may advantageously increase spatial resolution of flow cell images in each of the one or more spatial dimensions by 2x, 4x, 8x, 16x, 27x, 40x, 50x, lOOx or more than flow cell images acquired using traditional optical systems and / or image processing methods. The methods herein may also advantageously lessen the impact of color mixing of polonies that may be caused by neighboring polonies or clusters by computationally increasing the spatial resolution of flow cell images.
[0056] In situ samples such as cells or tissue can have a thickness along the axial or z direction that cannot remain in-focus within a single 2D image. A z-stack of multiple 2D flow cell images may be acquired to cover clusters or polonies at different z levels, e.g., in a 3D cellular sample. Interferences may occur in the z-stack of flow cell images, such as out-of-focus polonies and background signal from cellular components. For example, a polony that locates at a first z level can appear in a first flow cell image at a first z level and it may also generate a blob of signal in a second 2D flow cell image taken at its adjacent z level where it is out-of-focus. The blob of signal may interfere with intensities of polonies at or near the same x-y location in the second flow cell image, thusdeteriorating the accuracy and reliability of base callings. As another example, color mixing from neighboring polonies may interfere with polony intensity or polony density that can be detected for subsequent base calling. There is a need for identifying polony or cluster locations for accurate and reliable 3D base callings. There is also a need for generating accurate and reliable intensities for polonies or clusters from 3D volumetric samples so that they can be used for accurate and reliable base callings.
[0057] The techniques disclosed herein advantageously train a neural network to efficiently and accurately predict polony or cluster locations in the sample(s). The techniques disclosed herein advantageously train a neural network to efficiently and accurately predict high resolution intensities, base calls, and / or classifications for polonies or clusters in the sample(s). The samples herein are not limited to 3D samples, e.g., in situ cells and / or tissue. The samples herein may also include traditional 2D samples.
[0058] The techniques disclosed herein may advantageously utilize the reconfigurable logic device, e.g., FPGAs, and other integrated circuits, e.g., Al chips or neural processing units (NPUs), to: 1) predict high-resolution polony or cluster locations based on low-resolution flow cell images; or 2) to predict intensities, base calls and / or classifications at the high-resolution for the polonies or clusters in the sample(s). The utilization of the reconfigurable logic device, e.g., FPGAs, and other integrated circuits, e.g., Al chips or neural processing units (NPUs), on-board the sequencing system may advantageously reduce computational time, reduce energy consumption, improve sequencing analysis efficiency, reduce data storage space required, and reduce sequencing system cost in analysis of flow cell images when compared with sequencing analysis using existing sequencing systems.
[0059] The techniques disclosed herein advantageously train a neural network based on a loss function that is determined by comparison to reference base calls as ground truth, while the trained neural network may be used to accurately and reliably predict high resolution post image-processing flow cell images based on the flow cell images that are acquired from the sample(s). The techniques disclosed herein advantageously allow a mismatch in the training outputs and the prediction outputs. For example, the neural network may be trained by generating training base calls as training outputs and comparing the training outputs to reference base calls as ground truth. The trained neural network may then be used to predict high resolution flow cell images or to predict base calls. Such mismatching in training and prediction outputs may advantageously allowreference base calls to be considered in training parameters of the neural network and prediction of higher resolution higher quality version of the flow cell images that can be used to improve base calling accuracy and reliability. Such training and prediction advantageously enable utilization of a simplified neural network which requires less computational burden, reduction in computational time, reduction in power consumption, and reduction in making predictions.
[0060] The samples herein are not limited to 3D samples, e.g., in situ cells and / or tissue. The samples herein may also include traditional 2D samples. The techniques disclosed herein may advantageously utilize the reconfigurable logic device, e.g., FPGAs, and other integrated circuits, e.g., Al chips or neural processing units (NPUs), to perform one or more operations in the training and / or the prediction.
[0061] In DNA sequencing, identifying the centers of clusters or polonies is sometimes referred to as part of primary analysis. Primary analysis can include some or all of operations and / or steps needed to perform base calling and compute quality score of the base callings. Primary analysis can involve the formation of a template image for at least part of the flow cell. The template image can include the estimated locations of all detected clusters or polonies in a common coordinate system. The template image can include a polony map that is 2D or 3D. Template images are generated by identifying cluster or polony locations in all images in the first cycle or the first few cycles of the sequencing process. Generation of the template image may need sufficient spatial resolution to differentiate the polonies from background features, neighboring polonies, and / or duplicate polonies that are out-of-focus.Sequencing systems
[0062] FIG. 1 illustrates a block diagram of a computer-implemented system 100, according to one or more embodiments disclosed herein. The system 100 has a sequencing system 110 that includes a flow cell 112, a sequencer 114, an imager 116, data storage 122, and user interface 124. The sequencing system 110 may be connected to a cloud 130. The sequencing system 110 may include one or more of dedicated processors 118, a first reconfigurable logic device, e.g., Field-Programmable Gate Array(s) (FPGAs) 120, and a computing system 126.
[0063] In some embodiments, the flow cell 112 is configured to capture DNA fragments and form DNA sequences for base-calling on the flow cell. The flow cell 112 can includea support as disclosed herein. The support can be a solid support. The support can include a surface coating thereon as disclosed herein. The surface coating can be a polymer coating as disclosed herein.
[0064] A flow cell 112 can include multiple tiles or imaging areas thereon, and each tile may be separated into a grid of subtiles. Each subtile can include a plurality of clusters or polonies immobilized thereon. As a nonlimiting example, a flow cell can have 424 tiles, and each tile can be divided into a 6 x 9 grid, therefore 54 subtiles. The flow cell image as disclosed herein can be an image including signals of a plurality of clusters or polonies. The flow cell image can include one or more tiles of signals or one or more subtiles of signals. In some embodiments, a flow cell image can be an image that includes all the tiles and approximately all signals thereon. The flow cell image can be acquired from a channel during an imaging or sequencing cycle using the imager 116. In some embodiments, each tile may include millions of polonies or clusters. As a nonlimiting example, a tile can include about 1 to 10 million of clusters or polonies. Each polony can be a collection of many copies of DNA fragments.
[0065] Depending on the sample(s) immobilized on the support (e.g., a flow cell), the flow cell images may be acquired using the imager 116 at single or multiple z levels along a z axis orthogonal to the image plane of the flow cell images. In particular, for three dimensional samples, e.g., cells, tissues, or other in situ samples, the flow cell images can include multiple z-levels (i.e., z levels) in order to cover the whole sample(s) in 3D. The z axis can extend from the objective lens of the imager 116 disclosed herein to the support, e.g., flow cell 112. The z axis can be orthogonal to the image plane of the flow cell images. Each z level of flow cell images may be separated from the adjacent z level(s) for a predetermined distance, for example, ranging from about 0.1 um to about 15 urns, or from 0.02 um to 10 urns. Each z level of flow cell images may be separated from the adjacent z level(s) for a distance ranging from 0.5 um to 10 urns, from 0.01 um to 5 urns, or from 0.1 um to 15 urns. At each z level, flow cell images can be acquired from one or more sequencing cycles and / or one or more channels. Each flow cell image may include in its field of view at least part of one or more tiles or subtiles of the flow cell. FIG. 27 shows a portion of a flow cell 2712 with multiple tiles 2710. The image plane is defined by the x and y axis. And the z direction (i.e., z axis) is orthogonal to the x-y plane. Although the flow cell images, samples, and the z axis are described in a Cartesian coordinate system as shown in FIG. 27, any other coordinate systems can be used todefine spatial locations and relationships herein. Other coordinate systems can include but are not limited to the polar coordinate system, cylindrical, or spherical coordinate systems.
[0066] The sequencer 114 may be configured to flow a nucleotide mixture onto the flow cell 112, cleave blockers from the nucleotides in between flowing steps, and perform other steps for the formation of the DNA sequences on the flow cell 112. The nucleotides may have fluorescent elements attached that emit light or energy in a wavelength that indicates the type of nucleotide. Each type of fluorescent element may correspond to a particular nucleotide base (e.g., A, G, C, T). The fluorescent elements may emit light in visible wavelengths. In some embodiments, the sequencer 114 and the flow cell 112 may be configured to perform various sequencing methods disclosed herein, for example, sequencing-by-avidite.
[0067] For example, each nucleotide base may be assigned a color. Different types of nucleotides can have different colors. Adenine(A) may be red, cytosine(C) may be blue, guanine(G) may be green, and thymine(T) may be yellow, for example. The color or wavelength of the fluorescent element for each nucleotide may be selected so that the nucleotides are distinguishable from one another based on the wavelengths of light emitted by the fluorescent elements.
[0068] The imager 116 may be configured to capture images of the flow cell 112 after each flowing step. In some embodiment, the imager 116 includes a camera configured to capture digital images, such as a CMOS or a CCD camera. The camera may be configured to capture images at the wavelengths of the fluorescent elements bound to the nucleotides. The images acquired by the imager of the sample(s) immobilized on at least a portion of the flow cell can be called the flow cell images.
[0069] In some embodiments, the imager 116 can include one or more optical systems disclose herein. The optical system(s) can be configured to capture optical signals from the flow cell and generate corresponding flow cell images thereof. The flow cell images can then be used for base calling.
[0070] In an embodiment, the images of the flow cell may be captured in groups, where each image in the group is taken at a wavelength or in a spectrum that matches or includes only one of the fluorescent elements. In another embodiment, the images may be captured as single images that capture all of the wavelengths of the fluorescent elements.
[0071] The resolution of the imager 116 can control the level of detail in the flow cell images, including pixel size. In existing systems, this resolution is very important, as it controls the accuracy with which a spot-finding algorithm identifies the polony or cluster centers. In some embodiments, the image resolution of flow cell images disclosed herein can be about 10 nanometers (nms) to a couple of hundreds of nms or greater. In some embodiments, the image resolution of flow cell images can be in a range from 0.1 nm to 1000 nms. In some embodiments, the image resolution of flow cell images can be in a range from 1 nm to 500 nms. In some embodiments, the image resolution of flow cell images can be in a range from 5 nm to 300 nms. One way to increase the accuracy of polony or cluster finding is to improve the resolution of the imager 116, or improve the processing performed on images taken by imager 116. Detecting polony or cluster centers in pixels other than those detected by a spot-finding algorithm can be performed. These methods can allow for improved accuracy in detection of polony or cluster centers without increasing the resolution of the imager 116. The resolution of the imager 116 may even be better than existing systems with comparable performance, which may reduce the cost of the sequencing system 110.
[0072] The image quality of the flow cell images can control the base calling quality. One way to increase the accuracy of base calling is to improve the imager 116, or improve the processing performed on images taken by imager 116 to result in a better image quality. The methods described herein may predict high resolution of the flow cell images (2x, 4x, or more than existing flow cell image resolution, in a common coordinate system) so that the detectable polony or cluster density can be improved with reduced or eliminated interferences from neighboring polonies, cellular background signal, color mixing, and / or other noises in the flow cell images. As a result, 3D base calling can be more accurate using the methods herein when compared with existing methods without using such high resolution flow cell images. Such methods herein can allow for accurate and efficient base calling. These methods can be advantageously performed in parallel with a sequencing run in the computer-implemented system 100, without interference with or delay of existing sequencing workflow of the sequencing system 110. The results of predicted high resolution flow cell images can be available for making base calling in the current sequencing cycle in the sequencing workflow. Further, some or all of the operations disclosed herein can be advantageously performed by the first reconfigurable logic device, e.g., FPGA(s) or the integrated circuit, e.g., an application specificintegrated circuit (ASIC) chip, neural processing unit (NPU), or artificial intelligence (Al) chip and data can be communicated between the CPU(s) and the first reconfigurable logic device or integrated circuit to reduce the total operational time from methods operating using only the CPUs.
[0073] The sequencing system 110 may be configured to perform operations or actions for image processing of the flow cell images across different cycles and / or channels. The operations or actions disclosed herein may be performed by the dedicated processors 118, the reconfigurable logic device(s) and / or integrated circuit(s) 120, the computing system 126, or a combination thereof. One or more operations or actions in the methods 500, 600, 700, 2800, 2900 disclosed herein may be performed by the dedicated processors 118, the reconfigurable logic device(s) and / or integrated circuit(s) 120, the computing system 126, or a combination thereof. In some embodiments, which operations or actions are to be performed by the dedicated processors 118, the reconfigurable logic device(s) and / or integrated circuit(s) 120, the computing system 126, or their combinations can be determined based on one or more of a computation time for the specific operation(s), the complexity of computation in the specific operation(s), the need for data transmission between the hardware devices, the power required for the specific operation(s), or their combinations. Image processing operations or actions of the flow cell images can be performed after the corresponding flow cell images are acquired but before base calling of the flow cell images is performed.
[0074] In some embodiments, the data storage 122 is used to store information used in the methods herein. This information may include the flow cell images themselves or information and / or images derived from the flow images captured by the imager 116. The DNA sequences determined from the base-calling may be stored in the data storage 122. Parameters identifying polony or cluster locations may also be stored in the data storage 122. Raw and / or processed image intensities of each polony or cluster may be stored in the data storage 122. The region and / or subtile that each polony or cluster corresponds to may also be stored in the data storage 122. The transformation matrix of each region and / or subtile for different cycle(s) and / or channel(s) may also be stored in the data storage 122. Cell images may be stored in the data storage 122. The flow cell images, the processed images, and / or the filtered images may be stored in the data storage. Other information or images that can facilitate 3D base calling of the sample can be saved in the data storage.
[0075] The user interface 124 may be used by a user to operate the sequencing system or access data stored in the data storage 122 or the computing system 126.
[0076] The computing system 126 may control the general operation of the sequencing system and may be coupled to the user interface 124. It may also perform steps in image processing, base calling, their preceding operations, and / or subsequent operations including but not limited to predicting high resolution flow cell images. In some embodiments, the computing system 126 is a computer system 400, as described in more detail in FIG. 4. The computing system 126 may store information regarding the operation(s) of the sequencing system 110, such as configuration information, instructions for operating the sequencing system 110, or user information. The computing system 126 may be configured to pass information between the sequencing system 110 and the cloud 130.
[0077] The computing system 126 can include one or more general purpose computers that provide interfaces to run a variety of program in an operating system, such as Windows™ or Linux™. Such an operating system typically provides great flexibility to a user. In some embodiments, the computing system 126 may include one or more processors, e.g., CPUs, the CPUs may be configured for artificial intelligence algorithm development and training (e.g., neural network training), either alone or in combination with the reconfigurable logic device and / or integrated circuit 120.
[0078] In some embodiments, the sequencing system may include one or more reconfigurable logic devices 120 and / or one or more other integrated circuits 120. The reconfigurable logic device 120 can include one or more FPGA devices. The integrated circuit 120 herein may or may not be reconfigurable, and it may include an Al chip, an application-specific integrated circuit (ASIC) chip, a neural processing unit (NPU), or a combination thereof. In some embodiments, the reconfigurable logic device and / or integrated circuit 120 may be configured for artificial intelligence algorithm development and training (e.g., training of a neural network), either alone or in combination with the CPU and / or GPU.
[0079] In some embodiments, the reconfigurable logic device and / or integrated circuit 120 include a main unit and an edge unit. For example, the main unit may be a FPGA device and the edge unit may be an ASIC or Al chip. In some embodiments, the edge unit is an additional hardware processing module that may be individually installed and / or uninstalled on the system 110. The edge unit may be configured for artificial intelligencealgorithm development and training. The edge unit may be configured for making inferences or predictions using deployed Al algorithm(s), e.g., neural networks. The edge unit may communicate electronically with the main unit e.g., data communication via DMA connections. The edge unit may communicate electronically for data with other parts of the system 100 via various connections, such as a chip2chip connection. As an example, the edge unit may include a neural processing unit (NPU) chip, an Al chip, or any other integrated circuit(s).
[0080] In some embodiments, the dedicated processors 118 may be configured to perform operations in the methods disclosed herein. The dedicated processors 118 may include one or more reconfigurable logic devices and / or integrated circuits disclosed herein. The dedicated processors 118 may not include general-purpose processors, but instead custom processors with specific hardware or instructions for performing those steps. Dedicated processors directly run specific software without an operating system. The lack of an operating system reduces overhead, at the cost of the flexibility in what the processor may perform. A dedicated processor may make use of a custom programming language, which may be designed to operate more efficiently than the software run on general-purpose computers. This may increase the speed at which the steps are performed and allow for real time processing.
[0081] In some embodiments, the reconfigurable logic device and / or the integrated circuit 120, e.g., FPGA and / or Al chip, may be configured to perform some or all of operations in the methods herein. The reconfigurable logic device and / or the integrated circuit may be programmed as hardware that can perform specific task(s). A special programming language may be used to transform software steps into hardware componentry. Each software step may correspond to at least one operation or action in the methods disclosed herein. Each software step may include at least a part of the operation or action in the methods disclosed herein. Once the reconfigurable logic device is programmed, the hardware directly processes digital data that is provided to it without running software. The reconfigurable logic device and / or integrated circuit instead uses logic gates and registers to process the digital data. Because there is no overhead required for an operating system, the reconfigurable logic device and / or integrated circuit generally processes data faster than a general-purpose computer. Similar to dedicated processors, this may be at the cost of flexibility. The lack of software overhead may also allow the reconfigurable logic device and / or the integrated circuit to operate faster than a dedicatedprocessor, although this will depend on the exact processing to be performed and the specific the reconfigurable logic device and / or integrated circuit and dedicated processor.
[0082] A group of the reconfigurable logic devices and / or integrated circuits 120 may be configured to perform the steps in parallel. In some embodiments, a number of processing engines of the FPGA(s) may be configured to perform one or more identical image processing steps for an image, a set of images, a subtile, or a select region in one or more images. Each FPGA(s) 120 may perform its own part of the image processing step(s) in parallel, reducing the time needed to process data. This may allow the image processing step(s) to be completed in real time. For example, a number of processing engines of a first FPGA may be configured to generate a polony map for a tile of the flow cell. Each processing engine may be responsible for generating a portion, e.g., non-overlapping portion, of the polony map at a different subtile within the tile, e.g., in parallel. A second FPGA may be configured to perform intensity normalization in parallel with the generation of the polony map. As another example, a number of FPGA(s) and integrated circuits, e.g., Al chips, may be configured to perform one or more image processing step(s) for the flow cell images. Each FPGA(s) 120 may perform its own part of the processing step(s) in parallel, reducing the time needed to process data, while each Al chip may perform polony or cluster prediction after receiving data from its corresponding FPGA. This may allow the image processing steps to be completed in real time. For example, a first and second FPGA may be configured to perform intensity registration in parallel for a different subtile or tile of the flow cell. A corresponding Al chip may perform prediction of high resolution flow cell image of the corresponding subtile or tile after image registration is completed by its corresponding FPGA. Further discussion of the use of FPGAs is provided below.
[0083] The reconfigurable logic device and / or the integrated circuit may be configured to perform some or all of the operations or actions in the methods disclosed herein in real time. Performing the operations or actions in real time may allow the system 110 to use less memory and / or data storage, as the data may be processed as it is received. This is an improvement over conventional systems that may need to store the data before it may be processed and consequently require more memory / data storage or accessing a computer system located in the cloud 130. Further, performing the operations or actions in real time may allow more efficient sequencing analysis as it is being performing in parallel while a sequencing run is still in progress. Furthermore, performing the processing steps using theFPGAs and Al chips may allow the system to use less power, e.g., 2x, 5x, lOx, 20x or more, thus producing less heat than performing the same processing steps using the CPUs and / or GPUs. Further discussion of the use of FPGAs is provided below.
[0084] As discussed above, the sequencing system 110 may have dedicated processors 118, the reconfigurable logic device and / or integrate circuit 120, or the computing system 126. The sequencing system may use one, two, or all of these elements to accomplish one or more operations or actions in the methods disclosed herein. In some embodiments, when these hardware elements are present together, the image processing tasks are split between them. For example, the reconfigurable logic device 120 may be used to perform some or all of: the preprocessing operations, color correction, polony map generation, image registration, predicting high resolution flow cell images, training a neural network, generating the training flow cell images, base calling, and any subsequent operations, while the computing system 126 may perform other processing functions for the sequencing system 110 such as intensity normalization and registering images for base calling with cell staining image(s). Those skilled in the art will understand that various combinations of these elements will allow various system embodiments that balance efficiency and speed of processing with cost of processing elements.
[0085] In some embodiments, one or more reconfigurable logic devices and / or integrated circuits 120 can accelerate base calling and / or any primary analysis steps of flow cell images acquired from 2D or 3D sample(s). In some embodiments, the reconfigurable logic devices and / or integrated circuits can accelerate primary analysis of 2D sample(s) or 3D volumetric sample(s) by 2x, 4x, 5x, lOx, 15x, 20x, 25x, 30x, 40x, 50x, lOOx, 200x, 400x, 500x, 800x, lOOOx, or more than traditional primary analysis methods using only CPUs and / or GPUs. In some embodiments, one or more reconfigurable logic devices and / or integrated circuits 120 herein can accelerate sequencing and sequencing analysis (including at least primary analysis) of the flow cell images acquired from 2D or 3D sample(s). In some embodiments, the reconfigurable logic devices and / or integrated circuits herein can accelerate sequencing and sequencing analysis (including at least primary analysis) of the flow cell images acquired from 2D or 3D sample(s) by 2x, 4x, 5x, lOx, 15x, 20x, 25x, 30x, 40x, 50x, lOOx, 200x, 400x, 500x, 800x, lOOOx, or more than traditional sequencing systems with only CPUs and / or GPUs. In some embodiments, making inferences or predictions of high resolution images, of base calls, or of classifications, using the neural network disclosed herein and the reconfigurable logicdevices and / or integrated circuits can be less than 800 ms, 500ms, 400ms, 300ms, 200 ms, 100ms, 50ms, 20 ms, or less per tile per cycle. The tile size can be varied in different flow cells. The title size may be at least 0.0012mm, 0.01 mm2, 0.05 mm2, 0.1 mm2, 0.5 mm21 mm2, 2 mm2, 3 mm2or more.
[0086] In some embodiments, one or more reconfigurable logic devices and / or integrated circuits 120 can enable primary analysis (base calling) of polonies for flow cell images at multiple z levels. For example, processing time using reconfigurable logic devices can be less than 400 hours for at least 50 flow cell images (e.g., covering 50 tiles and from two or more color channels) with a FOV of at least 1 mm2with a resolution of 1 um or better in three dimensions for one or more flow cycles, e.g., 1-15 cycles. The flow cell images can be from multiple z- levels to cover some or all of the volumetric 3D samples (e.g., completely covering at least two samples).
[0087] In some embodiments, one or more reconfigurable logic devices and / or integrated circuits 120 can be used for accelerating primary analysis of 3D samples involving training neural network(s) and using the trained neural networks for making predictions or inferences. For example, neural network(s) can be used to predict polony locations and / or predict cell boundaries thereby identifying polonies within the cell(s). Using the reconfigurable logic device and / or integrated circuits 120 for computations associated with neural networks can reduce the training and / or prediction time needed in comparison with usage of GPUs or other computer processors, thereby accelerating sequence analysis, and enabling sequence analysis of flow cycles while subsequent flow cycles are to be performed or in progress in the sequence run. In some embodiments, the reconfigurable logic device(s) and / or integrated circuits 120 can accelerate training and / or prediction by lOx, 20x, 50x, 80x, lOOx, 200x, 500x, 600x, 800x, lOOOx, or more than training and / or prediction using CPUs and / or GPUs. In some embodiments, the reconfigurable logic devices and / or integrated circuits 120 can be used to achieve optimal acceleration in sequencing analysis. For example, one or more FPGA chips can be used in combination with an integrated circuit specific for computations corresponding to artificial intelligence (Al) algorithms, e.g., a NPU. The integrated circuit(s) can be specific circuits for Al functions. The integrated circuit(s) can include applicationspecific integrated circuits (ASIC). Computational tasks can be distributed to the FPGA(s) and the integrated circuit(s) to optimize computational time, energy consumption, heat dissipation, etc. For example, the Al chip may be used only forcomputations involving a neural network (e.g., predicting polony locations, predicting high resolution flow cell images, or training the neural network) and the FPGA(s) may be used for the rest of the primary analysis steps. The primary analysis time using dual FPGA chips or single FGPA chip in connection with the Al chip(s) can be less than 400, 300, 200, 100, 50, or 20 hours for at least 50 flow cell images (e.g., covering about 50 tiles of the flow cell and from two or more color channels) with a FOV of at least 1 mm2with a resolution of 1 um or better for each flow cell image in three dimensions for one or more flow cycles, e.g., 1-15 cycles. The flow cell images can be from multiple z-levels to cover some or all of the volumetric 3D samples (e.g., 10 to 20 z-locations to completely cover at least two samples). The primary analysis time may include a total time of image processing from obtaining raw flow cell images acquired using the imager 116 to generating base calls and saving base call results. The 3D samples herein includes polonies or clusters that are centered at different z levels that are spaced apart from each other with at least 0.01 um, 0.05 um, 0.1 um, 0.2 um, 0.5 um, 1 um, or more along the z direction or axial direction.
[0088] The cloud 130 may be a network, remote storage, or some other remote computing system separate from the sequencing system 110. The connection to cloud 130 may allow access to data stored externally to the sequencing system 110 or allow for updating of software in the sequencing system 110.Reconfigurable logic devices and integrated circuits
[0089] FIG. 5C shows an exemplary embodiment of the reconfigurable logic device and the integrated circuit(s) of the sequencing system disclosed herein. In some embodiments, the sequencing system 110 may include one or more reconfigurable logic devices 120_a. In the embodiment shown in FIG. 5C, the sequencing system comprises a single reconfigurable logic device, i.e., a first reconfigurable logic device 120_a. In some embodiments, the sequencing system comprises multiple reconfigurable logic devices (not shown). The reconfigurable logic device may comprise data processing engines 5011 configured to perform data processing in parallel. Each data processing engine may include a combination of digital logic circuit to perform its function, e.g., intensity extraction, convolution, registration, etc. The sequencing system 110 may further include reconfigurable routing channels 5013 that may function as connections among the data processing engines 5011 and may also connect the data processing engines to otherstructural elements, e.g., the first processor and the memory device, of the sequencing system 110. In some embodiments, a neural network may be deployed at least partly on the reconfigurable logic device 120_a so that the reconfigurable logic device can be used for at least some computational tasks for generating inferences using the neural network. The neural network may be pretrained using various training methods and data, for example, using the training methods and training data disclosed herein. The sequencing system may further include a first processor 120_c to selectively activate or deactivate different combinations of the of data processing engines 120_a and the reconfigurable routing channels 120_b. The FPGA(s) 120 as shown in FIG. 1 of the sequencing system 110 may include one or more of the reconfigurable logic device 120_a, the integrated circuit 120_b, and the processor 120_c.
[0090] In some embodiments, the FPGA(S) 120 may only include the reconfigurable logic device 120_a and the processor 120_c, but not the integrated circuit(s) 120_b. The different combinations of the of data processing engines 5011 and the reconfigurable routing channels 5013 may be configured to perform operation(s) in sequencing analysis to facilitate generating the sequencing analysis result(s). The sequencing analysis may include operations or steps of primary analysis. Such operation(s) may include one or more of (a) obtaining sensor data from one or more sensors (in the imager 116) of the sequencing system; (b) processing the sensor data to generate a first plurality of flow cell images; (c) predicting a second plurality of flow cell images using the neural network based on the sensor data or the first plurality of flow cell images; (d) determining polonies from the second plurality of flow cell images; and (e) performing a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images.
[0091] In some embodiments, the sensor data includes raw data that has been acquired from the sensor(s) of the imager without any additional image processing. In some embodiments, the sensor data includes raw flow cell images that have not been processed by the computing system 126, the dedicated processors 118, and / or the reconfigurable logic device and integrated circuit(s) 120 of the sequencing system 110.
[0092] In some embodiments, the sequencing system comprises: a first reconfigurable logic device 120_a comprising a first plurality of data processing engines 5011 configured to perform data processing in parallel; first reconfigurable routing channels 5013 connecting at least some of the first plurality of data processing engines 5011; aneural network deployed at least partly on the first reconfigurable logic device 5011; a first processor 120_c that selectively activates or deactivates different combinations of the first plurality of data processing engines 5011 and the first reconfigurable routing channels 5013 to perform operation(s) in sequencing analysis to facilitate generating the sequencing analysis result(s). The sequencing analysis may include operations or steps of primary analysis. Such operation(s) may include one or more of (a) obtaining sensor data directly from one or more sensors of the sequencing system; (b) processing the sensor data to generate a first plurality of flow cell images; (c) performing a first convolution in one or more dimensions on the first plurality of flow cell images, thereby generating a first convolution result; (d) repetitively performing, for one or more times, downsampling operations comprising: (1) performing a second convolution in one or more dimensions on the first convolution result, thereby generating a second convolution result; and (2) performing a down sampling of the second convolution result by a down sampling factor thereby generating a first down-sampled result, wherein in each repetition, the second convolution comprises a corresponding number of filters, thereby generating a third convolution result after (d); (e) performing the second convolution in one or more dimensions on the third convolution result, thereby generating a fourth convolution result; (f) repetitively performing up sampling operations comprising: (3) performing an up sampling of the fourth convolution result by an up sampling factor thereby generating a first up-sampled result; and (4) performing the second convolution in one or more dimensions of the first up-sampled result, thereby generating a fifth convolution result, wherein in each repetition, the second convolution comprises a corresponding number of filters, thereby generating a sixth convolution result after (f); (g) performing the first convolution in one or more dimensions on the sixth convolution result, thereby generating a seventh convolution result; (h) predicting a second plurality of flow cell images based on the seventh convolution result, wherein each of the second plurality of flow cell images corresponds to the corresponding flow cell image of the first plurality of flow cell images with a second resolution that is at least 2, 4, 6, 8, 10, 12, 16, or 32 times greater than the first resolution in one or more spatial dimensions; (i) determining polonies in the second plurality of flow cell images; (j) performing a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images; and (k) optionally forwarding the second plurality of flow cell images, the corresponding basecallings, or both to the first reconfigurable logic device, the first processor, or one or more hardware processors of the sequencing system.
[0093] In some embodiments, obtaining sensor data from one or more sensors (in the imager 116) of the sequencing system may be via a direct connection. In some embodiments, the direct connection between the first reconfigurable logic device (120 and 120_a) and the sensor(s) lacks other hardware components that may process or store the sensor data thus causing undesired complexity, delay, and possible errors in sensor data communication. Such hardware components include the first processor 120_c, the memory device 5030, or any processors, e.g., computing system 126, e.g., CPU, of the sequencing system. Comparing with traditional sequencing systems in which sensor data is communicated to other hardware components before it is communicated to where it is being processed (e.g., communicating to CPU and then to GPU to be processed) the direct sensor data communication herein advantageously improves data transmission efficiency from the sensor to the FPGAs 120, frees-up the other hardware(s), e.g., CPUs, storage devices, for other data processing functions, decreases power consumption from indirect data communication, and reduces time consumption in data communication thus sequencing analysis.
[0094] In some embodiments, the connection between the first reconfigurable logic device (120 and 120_a) and sensor may include other hardware components that may process or store the sensor data. Such hardware components may include the first processor 120_c, the memory device 5030, or any processors, e.g., CPUs 126 of the sequencing system. For example, the sensor data may be saved into the memory device 5030, and then it can be accessed by the first reconfigurable logic device using memory controller(s) 5013.
[0095] The reconfigurable logic device may include digital logic circuits therein, in a sense that it is also an integrated circuit. However, the integrated circuit herein (e.g., the Al chip, NPU, etc.) may have various difference with the reconfigurable logic device, e.g., the integrated circuit may not be as flexible in reconfiguration as the reconfigurable logic device. For example, the integrated circuit herein, e.g., the Al chip, NPU, etc., may not be reconfigurable.
[0096] In some embodiments, the sequencing system 110 comprises at least one reconfigurable logic device but lacks any integrated circuits, e.g., Al chips, ASIC chips, or NPUs. The reconfigurable logic device may perform one or more operations insequencing analysis and may forward its output back to the CPU as end results of primary analysis, e.g. base calls. Alternatively, the reconfigurable logic device may forward its output back to the CPU so that subsequent operations may be performed based on its output by the CPU to generate the end results of sequencing analysis.
[0097] In some embodiments, the sequencing system 110 comprises at least one reconfigurable logic device, and at least one integrated circuit as shown in FIG. 5C. The integrated circuit may perform one or more operations in sequencing analysis and may forward its output back to the reconfigurable logic device so that subsequent operations may be performed based on its output at the reconfigurable logic device.
[0098] In some embodiments, the output of the reconfigurable logic device or the integrated circuit comprises base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises identification of base calling locations in two dimensions. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises identification of base calling locations in three dimensions.
[0099] In some embodiments, the data communication between any two of the reconfigurable logic device, the integrated circuits, the first processor, and the second processor may be direct such that the direct communication lacks any other hardware components that may process or store the data. Such other hardware components may include memory device(s), and / or other processor(s) of the sequencing system. Such direct communication may include DMA connections. In some embodiments, the data communication the data communication between any two of the reconfigurable logic device, the integrated circuits, the first processor, and the second processor may be direct such the data may not be utilized by other logic circuits or stored before reaching its communication destination, but the data may be stored in a memory device before reach its communication destination.
[0100] In some embodiments, the sequencing system 110 may include a first reconfigurable logic device 120_a, e.g., FPGA, comprising a first plurality of data processing engines 5011 configured to perform data processing in parallel; an integrated circuit 120_b, e.g., an Al chip; a neural network deployed at least partly on the integrated circuit; a first processor to selectively activate or deactivate different combinations of the first plurality of data processing engines alone or in combination with the fist routing channels to perform operation(s) in sequencing analysis to facilitate generating thesequencing analysis result(s). The sequencing analysis may include operations or steps of primary analysis. The sequencing analysis may include operations or steps of secondary analysis. Such operation(s) may include one or more of: obtaining sensor data from one or more image sensors of the sequencing system; processing the sensor data to generate a first plurality of flow cell images; and communicating the sensor data, the first plurality of flow cell images, or both to the integrated circuit. The sequencing system may include a second processor or the first processor to control the integrated circuit to perform one or more operations in sequencing analysis to facilitate generating the sequencing analysis result(s). The sequencing analysis may include operations or steps of primary analysis and / or secondary analysis. Such operation(s) may include one or more of: receiving the sensor data, the first plurality of flow cell images, or both from the first reconfigurable logic device; predicting a second plurality of flow cell images using the neural network based on the sensor data, the first plurality of flow cell images, or both; determining polonies from the second plurality of flow cell images; performing a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images; and forwarding the second plurality of flow cell images, the determined polonies; corresponding base callings of polonies in the second plurality of flow cell images to one or more of: the first reconfigurable logic device 120_a, the first 120_c or second processor, and / or one or more processors of the sequencing system 126. In some embodiments, the operation of forwarding the second plurality of flow cell images, the determined polonies; corresponding base callings of polonies in the second plurality of flow cell images comprises forward to a memory device herein, e.g., DDR memory, so that one or more of: the first reconfigurable logic device 120_a, the first 120_c or second processor, and / or one or more processors of the computing system 126 can access the data from the memory. Accessing data from the memory including reading, writing, editing, etc., may be assisted by the memory controllers disclosed herein.
[0101] In some embodiments, the sequencing system 110 comprises at least one reconfigurable logic device, and at least one integrated circuit as shown in FIG. 5C. The integrated circuit may perform one or more operations in sequencing analysis and may generate its output as the end results of primary analysis and forward its output to one or more devices including: the reconfigurable logic device, the first or second processor, the hardware processor of the sequencing system, etc., so that the end results can be saved or presented to a user. In some embodiments, the output of the reconfigurable logic deviceor the integrated circuit comprises base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises identification of base calling locations in two dimensions. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises identification of base calling locations in three dimensions.
[0102] In some embodiments, the integrated circuit may perform one or more operations in sequencing analysis and generate its output as intermediate results of primary analysis, e.g., location of polonies, and may forward its output back to one or more of: the reconfigurable logic device, the first or second processor, the hardware processor of the sequencing system, etc., so that the end results can be determined based on its output.
[0103] In some embodiments, the integrated circuit may forward its output, either intermediate or end results, to be stored in a memory device, so that one or more devices including: the reconfigurable logic device, the first or second processor, and the hardware processor of the sequencing system can access the stored output whenever needed. The access to the output stored in a memory device can be via a memory controller of the sequencing system, e.g., 5013.
[0104] In some embodiments, the output of the reconfigurable logic device or the integrated circuit comprises base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises identification of base calling locations in two dimensions. In some embodiments, the output data of the reconfigurable logic device or the integrated circuit comprises identification of base calling locations in three dimensions.
[0105] In some embodiments, the sequencing system comprises: a first reconfigurable logic device comprising a first plurality of data processing engines configured to perform data processing in parallel with each other; an integrated circuit; a neural network deployed at least partly on the integrated circuit; a first processor to selectively activate or deactivate different combinations of the first plurality of data processing engines. The different combinations of the first plurality of data processing engines may be configured to perform operations comprising: obtaining sensor data from one or more image sensors of the sequencing system to generate the first plurality of flow cell images; and communicating the sensor data, the first plurality of flow cell images, or both to the integrated circuit. The integrated circuit may perform operations comprising: (1) receiving the sensor data, the first plurality of flow cell images, or both from the firstreconfigurable logic device; and (2) predicting a second plurality of flow cell images using the neural network based on the sensor data, the first plurality of flow cell images, or both; and (3) communicating the second plurality of flow cell images to the first reconfigurable logic device or one or more hardware processors of the sequencing system.
[0106] In some embodiments, the sequencing system comprises: a first reconfigurable logic device comprising a first plurality of data processing engines arranged in a first pipeline and configured to perform data processing in parallel with each other; an integrated circuit; a neural network deployed at least partly on the integrated circuit; a first processor of the first reconfigurable logic device to selectively activate or deactivate different combinations of the first plurality of data processing engines to perform operations comprising: (a) obtaining sensor data from one or more sensors of the sequencing system; (b) processing the sensor data to generate a first plurality of flow cell images; and (c) communicating the sensor data, the first plurality of flow cell images, or both to the integrated circuit; wherein the integrated circuit performs operations comprising: (d) receiving the sensor data, the first plurality of flow cell images, or both from the first reconfigurable logic device; (e) performing a first convolution in one or more dimensions on the first plurality of flow cell images, thereby generating a first convolution result; (f) repetitively performing, for one or more times, down-sampling operations comprising: (1) performing a second convolution in one or more dimensions on the first convolution result, thereby generating a second convolution result; and (2) performing a down sampling of the second convolution result by a down sampling factor thereby generating a first down-sampled result, wherein in each repetition, the second convolution comprises a corresponding number of filters, thereby generating a third convolution result; (g) performing the second convolution in one or more dimensions on the third convolution result, thereby generating a fourth convolution result; (h) repetitively performing up sampling operations comprising: (3) performing an up sampling of the fourth convolution result by an up sampling factor thereby generating a first up-sampled result; and (4) performing the second convolution in one or more dimensions of the first up-sampled result, thereby generating a fifth convolution result, wherein in each repetition, the second convolution comprises a corresponding number of filters, thereby generating a sixth convolution result; (i) performing the first convolution in one or more dimensions on the sixth convolution result, thereby generating a seventhconvolution result; and (j) predicting a second plurality of flow cell images based on the seventh convolution result, wherein each of the second plurality of flow cell images corresponds to the corresponding flow cell image of the first plurality of flow cell images with a second resolution that is at least 2, 4, 6, 8, 10, 12, or 16 times greater than the first resolution in one or more spatial dimensions.
[0107] In some embodiments, the first reconfigurable routing channels comprises one or more electronic nodes, and the electronic nodes are programmable. The electronic nodes here may include junction points in the circuit(s). The electronic nodes may include points where two or more circuit elements are connected together. In some embodiments, the first reconfigurable routing channels comprises one or more interconnects. The interconnect may include the physical wiring(s) that connects transistors and other components on an integrated circuit. In some embodiments, reconfigurable routing channels comprises one or more memory controllers, e.g., 5013 in FIG. 5C. In some embodiments, the first reconfigurable routing channels comprises one or more network- on-chips (NoCs), e.g., 5013 in FIG. 5C. In some embodiments, the first reconfigurable routing channels may comprise one or more of: a network-on-chip (NoC), and a memory controller.
[0108] The first reconfigurable routing channels may be configured to passively communicate data between components of the sequencing system. For example, the reconfigurable routing channels may be configured to communicate data bilaterally between the data processing engines, e.g., 5011 in FIG. 5C and the memory device, e.g., 5030 in FIG. 5C. The first reconfigurable routing channels may be configured to allow data communication between the first reconfigurable logic device, e.g.,120_a, and one or more memory devices, e.g., 5030. The first reconfigurable routing channels may be configured to allow data communication between the first reconfigurable logic device e.g., 120_a, and the integrated circuit, e.g., 120_b.
[0109] The reconfigurable logic device herein may each comprise one or more data processing engines, e.g., 5011. Each data processing engine may comprise multiple digital logic circuits.
[0110] The first reconfigurable logic device may be configured to communicate data with one or more memory devices external thereto. The first reconfigurable logic device may be configured to communicate data with one or more memory devices external thereto via the first reconfigurable routing channels. The first reconfigurable logic device maycomprise digital circuits that are integrated and forming a FPGA device. For example, the FPGA device in FIG. 5C includes the first reconfigurable logic device, the DMA connections, the first reconfigurable routing channels (e.g., NoC and memory controllers).[OHl] The sequencing system may further comprise one or more memory devices electrically connected for data communication with one or more components of the sequencing system, the one or more components may include one or more of: the first reconfigurable logic device; the integrated circuit; the first reconfigurable routing channels; the one or more memory controllers; the first processor; a second processor; and one or more processors of the sequencing system.
[0112] In some embodiments, the sequencing system further comprises one or more direct data access (DMA) connections, e.g., 5012 in FIG. 5C, that are in data communication with the plurality of data processing engines and the first reconfigurable routing channels, e.g., 5013 in FIG. 5C. The DMA connections may be configured to actively communicate data between components of the sequencing system. For example, the DMA connections may be configured to fetch data or send data to components that are connected thereto, e.g., the data processing engines, e.g., 5011 in FIG. 5C and the reconfigurable routing channels, e.g., 5013 in FIG. 5C. The DMA connections herein may be configured to actively request data from or actively sending data directly to: the first reconfigurable logic device; the first reconfigurable routing channels; the integrated circuit; or a combination thereof. One or more direct data access (DMA) connections may be in data communication with the first reconfigurable routing channels and the integrated circuit herein. The DMA connections may be configured to allow data communication based on a predetermined protocol, e.g., a PCIe protocol.
[0113] In some embodiments, the first reconfigurable routing channels are configured to allow data communication between the first reconfigurable logic device and one or more memory devices. In some embodiments, the one or more DMA connections and the first reconfigurable routing channels are configured to allow data communication between the first reconfigurable logic device and the integrated circuit.
[0114] In some embodiments, the sequencing system further comprises an integrated circuit that is different from the first reconfigurable logic device, , e.g., 120_b in FIG. 5C. The integrated circuit herein may not be reconfigurable. The integrated circuit may comprise an application specific integrated circuit (ASIC) chip. In some embodiments,the integrated circuit comprises a neural processing unit (NPU) or an artificial intelligence (Al) chip. The integrated circuit may comprise a second plurality of data processing engines, each data processing engine comprising multiple digital logic circuits. The integrated circuit may further comprise: second plurality of data processing engines and second routing channels, each connecting at least some of the second plurality of data processing engines.
[0115] In some embodiments, the sequencing system further comprises a first processor. The first processor may be configured to selectively activate or deactivate different combinations of the first plurality of data processing engines and the first reconfigurable routing channels to perform the operations disclosed herein. In some embodiments, the sequencing system further comprises a second processor. The second processor may be configured to control digital circuits of the integrated circuit herein.
[0116] In some embodiments, the first processor, or a second processor, e.g., of the integrated circuit, is configured to selectively activate or deactivate different combinations of the second plurality of data processing engines and the second reconfigurable routing channels to perform the operations. The first processor or a second processor may be configured to selectively activate or deactivate different combinations of the second plurality of data processing engines and the second reconfigurable routing channels to perform the operations herein.
[0117] The sequencing system may further comprise a housing that encloses the first reconfigurable logic device, the first reconfigurable routing channels, the one or more DMA connections, the integrated circuit, and the first processor therein. In some embodiments, the sequencing system further comprises: a housing that encloses at least the first reconfigurable logic device therein and the integrated circuit is external to the housing.
[0118] In some embodiments, the sequencing system further comprises: a power source that is configured to supply different power levels to the first reconfigurable logic device and the integrated circuit. A first power level supplied by the power source to the first reconfigurable logic device may be higher than a second power level supplied to the integrated circuit while a sequencing run and / or sequencing analysis is in progress. A maximum power output of the power source of the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the maximum power output of the power source of sequencers, e.g., traditional sequencers without the first reconfigurable logic device (e.g., FPGA), theintegrated circuit (e.g., Al chip), or both. The time consumption in performing a sequencing run and corresponding sequencing analysis (e.g., primary analysis) thereof using the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the time consumption in performing the same sequencing run using a sequencer without the first reconfigurable logic device, the integrated circuit, or both (e.g., a traditional sequencer without FPGA and / or Al chips). Time consumption in performing a sequencing run and sequencing analysis of the sequencing run (e.g., primary analysis) using the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the time consumption in performing the same sequencing run and analysis using a sequencer without the first reconfigurable logic device, the integrated circuit, or both (e.g., a traditional sequencer without FPGA and / or Al chips). In some embodiments, a maximum power output of the power source to the sequencing system in performing a sequencing run and corresponding sequencing analysis thereof is less than 900 Watts, 800 Watts, 700 Watts, 650 Watts, 600 Watts, 550 Watts, or 500 Watts. The power source may be configured to supply a first power level to the first reconfigurable logic device, the first power level is less than 500 Watts, 400 Watts, 350 Watts, or 300 Watts. The power source may be configured to supply a second power level to the integrated circuit, the second power level is less than 450 Watts, 400 Watts, 350 Watts, or 300 Watts.
[0119] In some embodiments, one or more components of the first reconfigurable logic device and / or integrated circuit may include a computational performance of at least 2, 4, 8, 10, 16, 20, 30, 40, 50, 60, 70, 80, or 100 Giga-operations per second (GOPs) or more. In some embodiments, one or more processing engines of the first reconfigurable logic device and / or integrated circuit may include a computational performance of at least 12, 4, 8, 10, 16, 20, 30, 40, 50, 60, 70, 80, or 100 Giga-operations per second (GOPs), or more Giga-operations per second (GOPs), or more. In some embodiments, the first reconfigurable logic device and / or the integrated circuit includes a computational performances of at least 10, 20, 40, 50, 60, 80, or 100 Tera-operations per second (TOPs).
[0120] In some embodiments, one or more components are located on a first printed circuit board (PCB). The one or more components may include: the first reconfigurable logic device the first reconfigurable routing channels; the first processor; and the one or more DMA connections. In some embodiments, the integrated circuit is located on a second printed circuit board (PCB) different from the first printed circuit board, e.g., as shown in FIG. 5C. The integrated circuit and the second PCB may be positioned within asame housing of the sequencing system as the first PCB or external to the housing of the sequencing system. Being on a separate PCB makes connecting the first reconfigurable logic device, e.g., FPGA device with various integrated circuit on a chip convenient, efficient, and easily customizable. In some embodiments, the first PCB board may be a main board, and the second PCB board may be a daughter board or edge unit.
[0121] In some embodiments, the sequencing systems lacks any graphic processing units (GPUs) or tensor processing units (TPUs). Instead, the sequencing systems utilizes FPGAs, Al chips, NPUs, or other ASIC chips for performing the operations disclosed herein. The sequencing system disclosed herein advantageously requires less power, generates less heat, and reduces the hardware complexity and costs for performing NGS sequencing runs and corresponding sequencing analysis than sequencers that use GPUs or TPUs.
[0122] In some embodiments, the sequencing systems include logic devices that are not limited to reconfigurable logic devices (e.g., FPGAs) and / or other integrated circuits (e.g., Al chips, NPUs). In some embodiments, the sequencing systems include various types of processing units or processors configured for reconfigurable parallel processing, In some embodiments, the sequencing systems include various types of logic devices or integrated circuits, e.g., ASIC chips. In some embodiments, the sequencing systems include GPUs, TPUs, or other various types of processing units that are configured to perform one or more operations disclosed herein.
[0123] In some embodiments, the sequencing systems include GPUs, TPUs, or other various types of processing units that are configured to perform one or more operations that can be performed by the reconfigurable logic devices (e.g., FPGAs) and / or other integrated circuits (e.g., Al chips, NPUs).
[0124] The first processor may be positioned on the first PCB board together with the reconfigurable logic device for convenient and efficient control of the reconfigurable logic device. In some embodiments, the first processor is a separate processor from one or more processors of the sequencing system configured to control the optical system, the fluidics of the sequencing system, etc. In some embodiments, the first processor can be configured to only control the components on the first PCB board, e.g., the FPGA device, alone or in combination with components on the second PCB board, e.g., the Al chip. In some embodiments, the sequencing system may comprise a second processor that is configured to separately control the Al chip. The first processor or second processor ofthe sequencing system, e.g., 120_c, may comprise a CPU. The one or more hardware processors of the sequencing system comprises a CPU. In some embodiments, the first or second processor, e.g., 120_c, lacks any GPU or TPU. In some embodiments, the first or second processor, e.g., 120_c, comprises only CPU(s).
[0125] In some embodiments, the sequencing system may further comprise a heat dissipator configured to maintain a system temperature in a range from 0 degrees to 120 degrees Celsius or less than 120 degrees Celsius.
[0126] In some embodiments, the operation for processing the sensor data to generate the first plurality of flow cell images comprises one or more of: registering the first plurality of flow cell images to a reference coordinate system; adjusting image intensities of the first plurality of flow cell images; color correction of the first plurality of flow cell images; correcting phasing and prephasing of the first plurality of flow cell images; and subtracting background intensities from the first plurality of flow cell images.
[0127] In some embodiments, each of the one or more operations performed by the first reconfigurable logic device or the integrated circuit are in real time. In some embodiments, each of the one or more operations performed by the first reconfigurable logic device or the integrated circuit are within the time window of performing sequencing reactions and / or imaging of a single sequencing cycle of the sequencing run. In some embodiments, each of the one or more operations performed by the first reconfigurable logic device or the integrated circuit are within the time window of performing sequencing reactions and / or imaging of a single z-level of a single sequencing cycle.
[0128] FIG. 5D shows an exemplary embodiment of performing sequencing analysis in parallel with performing a sequencing run. In this particular embodiment, the sequencing run includes multiple sequencing cycles, only part of a single cycle is shown herein. For each cycle, flow cell images are acquired at multiple z-levels from different color channels of an in situ sample . The sequencing reactions are repeatedly performed for each z-level in each cycle within a time window 5601. The operations of the integrated circuit are performed within a processing window 5602 within the time window 5609 of a single sequencing cycle and also within a time window 5601 for sequencing reactions and imaging at a single z-level 5601. The operations of the first reconfigurable logic device (e.g., on board primary analysis operations) are also performed with a processing window 5603 that is within the time window 5609 of each sequencing cycle. The processingwindows 5602 and 5603 may be of identical or different duration depending on various factors such as sequencing data, primary analysis algorithms, etc. In some embodiments, the operations are not just performed within the processing windows but completed within the processing windows with respect to the data of the current cycle, e.g., of a preceding z-level of the current cycle that sensor data has been acquired. In some embodiments, the operations are completed within the processing windows with respect to the data of a preceding cycle, e.g., the cycle immediately preceding the current cycle.
[0129] In embodiments where the sample is a 3D volumetric sample, the operations are performed for a single z level in each cycle within a predetermined time window, e.g., 5602, 5603. The predetermined time window is for a single z level in a single sequencing cycle. In some embodiments, the predetermined time window is less than 1000 ms, 900 ms, 800 ms, 700ms, 600 ms, 500 ms, 400 ms, 300 ms, 250 ms, 200 ms, or 100 ms. In some embodiments, each of the one or more operations are performed within the predetermined time window and in parallel while the sequencing run is in progress. In some embodiments, each of the one or more operations are performed in parallel within a time window that sequencing, imaging, or both of a subsequent sequencing cycle is completed.
[0130] The first plurality of flow cell images herein may be obtained from a single z level of a 2D or 3D sample. In some embodiments, the first plurality of flow cell images herein may be obtained from multiple z levels covering at least partly of an in situ sample, e.g., of cells or tissue(s). The first plurality of flow cell images may be obtained from one or more color channels at each z level of the multiple z levels covering at least partly of the in situ sample. In some embodiments, the first plurality of flow cell images are from a single color channel. In some embodiments, the first plurality of flow cell images are from multiple color channels. In some embodiments, the first plurality of flow cell images are from a single sequencing cycle. In some embodiments, the first plurality of flow cell images are from multiple sequencing cycles. The first plurality of flow cell images may be of a first spatial resolution in x, y, and / or z directions. The second plurality of flow cell images may be generated based on the first plurality of flow cell images. The second plurality of flow cell images may be of a second spatial resolution in x, y, and / or z directions. The first spatial resolution may be lower than the second spatial resolution, and a higher resolution herein indicates that a pixel size is smaller so that the polonies in the flow cell images are of finer spatial details. The first spatial resolution may be 2x, 4x,6x, 8x, lOx, 16x, 24x, 32x, or 48x lower than the second spatial resolution in x, y, and / or z directions. The first spatial resolution may be at least 2x, 4x, 6x, 8x, lOx, 16x, 24x, 32x, or 48x lower than the second spatial resolution in x,y, and / or z directions. In some embodiments, the first and second resolution is in 3D. In some embodiments, the first resolution is in a range of 0.1 um to 5 um. In some embodiments, the second resolution is in a range of 0.01 um to 2 um. In some embodiments, the second resolution is at least 4, 6, or 8 times greater than the first resolution in all three dimensions.
[0131] In some embodiments, the sequencing system further comprises one or more image sensors configured to receive optical signals generated from sequencing reactions of a sample immobilized on a support. The support may comprise a glass or plastic substrate. The support may be included in a flow cell device. The one or more image sensors may be configured to generate sensor data based on the optical signals. In some embodiments, the sequencing system further comprises: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations disclosed herein. The one or more data storage devices may include one or more memory devices. The one or more memory devices may be accessible by the one or more processors, the first processor, the second processor, the first reconfigurable logic device, and the integrated circuit.
[0132] In some embodiments, the one or more processors are separate from the first or second processors. The operations performed by the one or more processors may include one or more of 1) recording sensor data generated in the sequencing system in one or more flow cycles; 2) optionally processing the recorded sensor data; 3) sending the recorded sensor data or the optionally processed data to the first reconfigurable logic device or the integrated circuit; 4) receiving outcome from the first reconfigurable logic device or integrated circuit; and 5) generating sequencing analysis results based on the received outcome. The operations performed by the one or more processors may include one or more of 1) receiving outcome from the first reconfigurable logic device or integrated circuit; and 2) generating sequencing analysis results based on the received outcome.
[0133] In some embodiments, the sequencing analysis results comprise primary analysis results. In some embodiments, the sequencing analysis results comprise a data file in a predetermined data format. In some embodiments, the sequencing analysis resultscomprise base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the sequencing analysis results comprises quality measurements of base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the sequencing analysis results comprises quality scores corresponding to base calls of nucleotide bases in a sample immobilized on a support.
[0134] In some embodiments, the sequencing system further comprises: a sample immobilized on a support; and an optical system comprising: an illumination system; an objective lens and the one or more image sensors. The optical system is configured to emit light to the sample and to collect optical signals emitted from the sample, thereby generating the first plurality of flow cell images. The support may be comprised in a flow cell device.
[0135] In some embodiments, the operation(s) performed by the first reconfigurable routing channels or the integrated circuit using the neural network comprises one or more of: generating quality measurements of the base callings; and generating a data output file based on the base callings.
[0136] In some embodiments, the neural network herein comprises a convolutional neural network (CNN). In some embodiments, the neural network comprises a U-Net. In some embodiments, the neural network has been pretrained. In some embodiments, the neural network has been trained using the first reconfigurable logic device or the integrated circuit. In some embodiments, the neural network is a 3D neural network.
[0137] In some embodiments, the first convolution comprises a 3D convolution with a convolution kernel. In some embodiments, the convolutional kernel has at least four dimensions. In some embodiments, the convolutional kernel is m x m x m x n, wherein m is an integer in a range from 3 to 30, wherein n is an integer. In some embodiments, n is an integer from 1 to 16384. In some embodiments, the second convolution in operation (1) comprises a corresponding number of n, 2*n, 4*n, and 8*n filters in a first, second, third, and fourth repetition, respectively. In some embodiments, the second convolution in (4) comprises a corresponding number of 2*n, 2*n, 4*n, 8*n filters in a last repetition, last minus one, last minus two, and last minus three repetition, respectively. In some embodiments, n is in a range from 4 to 1024.
[0138] In some embodiments, the neural network has been trained using the first reconfigurable logic device or the integrated circuit. In some embodiments, the neural network is a 2D neural network. In some embodiments, the first convolution comprises a2D convolution with a convolution kernel. In some embodiments, the convolutional kernel has at least three dimensions. In some embodiments, the convolutional kernel is m x m x n, wherein m is an integer in a range from 3 to 30, wherein n is an integer. In some embodiments, n is an integer from 1 to 16384.
[0139] In some embodiments, the second convolution in operation (1) comprises a corresponding number of n, 2*n, 4*n, and 8*n filters in a first, second, third, and fourth repetition, respectively. In some embodiments, the second convolution in (4) comprises a corresponding number of 2*n, 2*n, 4*n, 8*n filters in a last repetition, last minus one, last minus two, and last minus three repetition, respectively. In some embodiments, n is in a range from 4 to 1024.
[0140] In some embodiments, the second convolution in operation (1) comprises a corresponding number of n, 2*n, 4*n filters in a first, second, third repetition, respectively. In some embodiments, the second convolution in (4) comprises a corresponding number of 2*n, 2*n, 4*n, filters in a last repetition, last minus one, last minus two, repetition, respectively. In some embodiments, n is in a range from 4 to 1024.
[0141] In some embodiments, the neural network is pretrained with 2D flow cell images at multiple z-levels that encompass the 3D volume of the volumetric sample(s).Comparing with neural networks trained with 3D volumes of training data, e.g., 3D CNN, the neural networks pretrained with 2D flow cell images, e.g., 2D CNN, are less complex and requires less computational effort in making predictions or inferences, thereby providing higher efficiency and saving time and computational effort in its prediction of polony locations. In some embodiments, the neural network pretrained with 2D flow cell images may predict polony locations per tile per cycle in a time window that is lOx, 50x, 80x, lOOx, 200x, 400x, 600x, 800x, lOOOx, 1500x, 2000x or less than making identical predictions using neural networks trained from 3D volumes of flow cell images.
[0142] In some embodiments, the neural network pretrained with 2D flow cell images may predict polony locations per tile per cycle using the reconfigurable logic device and / or other integrated circuits, e.g., FPGA and / or Al chips, in a time window that is 5x, lOx, 20x, 40x, 50x, 80x, lOOx, 200x, 400x, 600x, 800x, lOOOx or less than identical neural network using CPUs or other processors.
[0143] In some embodiments, the operation (e) performing a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images comprises: performing a corresponding base calling for each of the determined poloniesbased on the second plurality of flow cell images and based on a fourth plurality of flow cell images, wherein the fourth plurality of images are predicted using a second neural network based on a third plurality of flow cell images. In some embodiments, the third plurality of flow cell images are acquired from one or more color channels that is different from the single channel, and wherein the third plurality of flow cell images comprises the first resolution. In some embodiments, the fourth plurality of flow cell images comprises the second resolution.
[0144] In some embodiments, the first plurality of flow cell images are from one or more color channels. In some embodiments, the first plurality of flow cell images are of unbalanced nucleotide diversity. In some embodiments, the first plurality of flow cell images comprises: an unbalanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more flow cycles. In some embodiments, the first plurality of flow cell images comprises: a balanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more cycles. In some embodiments, two or more different concatemer molecules among the concatemer molecules have different insert sequences. In some embodiments, different insert sequences correspond to different target RNA molecules or target cDNA molecules. In some embodiments, each location of the determined polonies corresponds to a location of the concatemer molecules. In some embodiments, the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents bound to a balanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support. In some embodiments, the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents bound to a unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support in the one or more subsequent cycles. In some embodiments, the unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules comprises: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of bases is less than 20%, 15%, 10%, or 5% in the one or more cycles. In some embodiments, the balanced diversity of nucleotide bases of A, G, C and T / U among the plurality of concatemer molecules comprises: a percentage of (1) a number of each type of nucleotide bases to (2) a total number of bases in the one or more cycles is more than 10%, 15%, or 20%.
[0145] In some embodiments, the cellular sample comprises overloaded concatemer molecules with a spatial density in a range of 102-1015per mm2. In some embodiments, the cellular sample comprises overloaded concatemer molecules with a spatial density in a range of 103-IO10per mm2.
[0146] In some embodiments, the down-sampling factor is 2, 4, or 8. In some embodiments, the up-sampling factor is 2, 4, or 8. In some embodiments, the downsampling factor is 2, 4, 8, 16, 32, 64, or more. In some embodiments, the up-sampling factor is 2, 4, 8, 16, 32, 64, or more.
[0147] In some embodiments, one or more of operations of (a) to (k) are performed while a sequencing run is being performed. In some embodiments, the first plurality of flow cell images are acquired in sequencing cycles ranging from 1 to 500. In some embodiments, the one or more cycles comprises a current cycle N. In some embodiments, wherein N is in a range from 1 to 500. In some embodiments, the one or more cycles comprises a single cycle ranging from 1 to 500. In some embodiments, the one or more cycles comprises multiple cycles ranging from 1 to 500. In some embodiments, one or more of operations, e.g., operations (a) to (j), are performed while the sequencing reactions in cycles subsequent to the current cycle N is yet to be performed or currently being performed.
[0148] In some embodiments, the training data set of flow cell images comprises z-stacks of flow cell images taken at different z-locations, and each z-stack is used as a 3D volume for training the neural network. In some embodiments, the training data set of flow cell images comprises 2D flow cell images taken at different z-locations, and individual 2D flow cell images at multiple z-levels are used as 2D images for training the neural network.
[0149] In some embodiments, the training data set of flow cell images comprises simulated flow cell images of in situ samples at different z-locations. In some embodiments, the training data set of flow cell images comprises actual flow cell images acquired from in situ samples at different z-locations.
[0150] In some embodiments, performing the first convolution in one or more dimensions on the first plurality of flow cell images comprises: performing a first convolution in 3D on the first plurality of flow cell images, thereby generating a first convolution result. In some embodiments, performing a second convolution in one or more dimensions on the first convolution result, thereby generating a second convolution result comprises:performing the second convolution in 3D on the first convolution result, thereby generating a second convolution result.
[0151] In some embodiments, performing the first convolution in one or more dimensions on the first plurality of flow cell images comprises: performing a first convolution in 2D on the first plurality of flow cell images, thereby generating a first convolution result. In some embodiments, performing a second convolution in one or more dimensions on the first convolution result, thereby generating a second convolution result comprises: performing the second convolution in 2D on the first convolution result, thereby generating a second convolution result.
[0152] In some embodiments, repetitively performing up sampling operations comprises: (3) performing an up sampling of the fourth convolution result by an up sampling factor thereby generating a first up-sampled result; (4) concatenating the first up-sampled result in a current up-sampling repetition with the first down-sampled result in a previous downsample repetition, wherein the first up-sampled result has a same size as the first down- sampled result in the previous down-sampling repetition; and (5) performing the second convolution in one or more dimensions of the first up-sampled result, thereby generating a fifth convolution result.
[0153] In some embodiments, the different combinations of the first plurality of data processing engines are configured to perform operations further comprising: (a) receiving the second plurality of flow cell images from the integrated circuit; (b) determining polonies from the second plurality of flow cell images; and (c) performing a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images; and (d) forwarding the second plurality of flow cell images, the determined polonies, the corresponding base callings to the first processor or one or more hardware processors of the sequencing system or a combination thereof.
[0154] In some embodiments, the one or more operations performed by the first reconfigurable logic device further comprises: forwarding the second plurality of flow cell images, the determined polonies, the corresponding base callings, or a combination thereof to the first processor or one or more hardware processors of the sequencing system. In some embodiments, the one or more operations performed by the integrated circuit further comprises forwarding the second plurality of flow cell images, the corresponding base callings, or both to the first reconfigurable logic device, the first processor or one or more hardware processors of the sequencing system.
[0155] In some embodiments, the operations performed by the first reconfigurable logic device or the integrated circuit further comprising: registering the second plurality of flow cell images to a common coordinate system.
[0156] In some embodiments, the operations performed by the integrated circuit further comprising one or more of: determining polonies from the second plurality of flow cell images; performing a corresponding base call for each of the determined polonies based on the second plurality of flow cell images; and forwarding the second plurality of flow cell images, the corresponding base callings, or both to the first reconfigurable device, the first processor, or one or more hardware processors of the sequencing system.
[0157] In some embodiments, the operation (d) or (i) of determining polonies from the second plurality of flow cell images comprises: generating a 3D polony map comprising spatial location of polonies based on the determined polonies. The operation of generating a 3D polony map comprising spatial location of polonies based on the determined polonies may further comprise: deleting duplicate polonies from the determined polonies, wherein the duplicate polonies are out-of-focus. In some embodiments, the operation of determining polonies from the second plurality of flow cell images comprises: superimposing the second plurality of flow cell images with corresponding cell staining images; and generating the polony map by only including polonies that are within cell boundaries in the corresponding cell staining images. Exemplary embodiments of methods for generating the polony maps are disclosed in U.S. Patent Application No. 18 / 078,820 and PCT Application No. PCT / US2023 / 076125, which are incorporated by reference in their entireties.
[0158] Disclosed herein, in some embodiments, are sequencing methods comprising operations herein. Such operation may include one or more of: (a) obtaining, by a first reconfigurable logic device of a sequencing system, sensor data from one or more sensors of the sequencing system; (b) processing, by the first reconfigurable logic device, the sensor data to generate a first plurality of flow cell images; (c) predicting, by the first reconfigurable logic device, a second plurality of flow cell images using a neural network at least partly deployed on the first reconfigurable device and based on the sensor data or the first plurality of flow cell images; (d) determining, by the first reconfigurable logic device, polonies from the second plurality of flow cell images; (e) performing, by the first reconfigurable logic device, a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images; and (f) optionally forwarding,by the first reconfigurable logic device, the second plurality of flow cell images, the corresponding base calling, or both to one or more processors of the sequencing system.
[0159] Disclosed herein, in some embodiments, are sequencing methods comprising operations herein. Such operations may include one or more of (a) obtaining, by the first reconfigurable logic device, sensor data from one or more image sensors of the sequencing system; (b) processing, by the first reconfigurable logic device, the sensor data to generate a first plurality of flow cell images; (c) communicating, by the first reconfigurable logic device to an integrated circuit, the sensor data, the first plurality of flow cell images, or both; (d) receiving, by the integrated circuit and from the first reconfigurable logic device, the sensor data, the first plurality of flow cell images, or both; (e) predicting, by the integrated circuit, a second plurality of flow cell images using the neural network based on the sensor data, the first plurality of flow cell images, or both; (f) determining, by the integrated circuit, polonies from the second plurality of flow cell images; and (g) performing, by the integrated circuit, a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images.
[0160] Disclosed herein, in some embodiments , are sequencing methods comprising operations herein. Such operation may include one or more of (a) obtaining, by the first reconfigurable logic device of a sequencing system, sensor data from one or more image sensors of the sequencing system to generate the first plurality of flow cell images; (b) communicating, by the first reconfigurable logic device, the sensor data, the first plurality of flow cell images, or both to the integrated circuit; (c) receiving, by the integrated circuit of the sequencing system, the sensor data, the first plurality of flow cell images, or both from the first reconfigurable logic device; (d) predicting, by the by the integrated circuit, a second plurality of flow cell images using a neural network deployed at least partly on the integrated circuit and based on the sensor data, the first plurality of flow cell images, or both; and (e) communicating, by the integrated circuit, the second plurality of flow cell images to the first reconfigurable logic device or one or more hardware processors of the sequencing system.
[0161] In some embodiments, the first reconfigurable routing channels comprises one or more electronic nodes, and the electronic nodes are programmable. The electronic nodes here may include junction points in the circuit(s). The electronic nodes may include points where two or more circuit elements are connected together. In some embodiments, the first reconfigurable routing channels comprises one or more interconnects. Theinterconnect may include the physical wiring(s) that connects transistors and other components on an integrated circuit. In some embodiments, reconfigurable routing channels comprises one or more memory controllers, e.g., 5013 in FIG. 5C. In some embodiments, the first reconfigurable routing channels comprises one or more network- on-chips (NoCs), e.g., 5013 in FIG. 5C. In some embodiments, the first reconfigurable routing channels may comprise one or more of a network-on-chip (NoC), and a memory controller.
[0162] The first reconfigurable routing channels may be configured to passively communicate data between components of the sequencing system. For example, the reconfigurable routing channels may be configured to communicate data bilaterally between the data processing engines, e.g., 5011 in FIG. 5C and the memory device, e.g., 5030 in FIG. 5C. The first reconfigurable routing channels may be configured to allow data communication between the first reconfigurable logic device and one or more memory devices. The first reconfigurable routing channels may be configured to allow data communication between the first reconfigurable logic device and the integrated circuit.
[0163] The reconfigurable logic device herein may each comprise one or more data processing engines. Each data processing engine may comprise multiple digital logic circuits.
[0164] The first reconfigurable logic device may be configured to communicate data with one or more memory devices external thereto. The first reconfigurable logic device may be configured to communicate data with one or more memory devices external thereto via the first reconfigurable routing channels. The first reconfigurable logic device may comprise a first integrated circuit forming a FPGA device. For example, the FPGA device in FIG. 5C includes the first reconfigurable logic device, the DMA connections, and the first reconfigurable routing channels (e.g., NoC and memory controllers).
[0165] The sequencing system may further comprises one or more memory devices electrically connected for data communication with one or more components of the sequencing system, the one or more components may include one or more of the first reconfigurable logic device; the integrated circuit; the first reconfigurable routing channels; the one or more memory controllers; the first processor; a second processor; and one or more processors of the sequencing system.
[0166] In some embodiments, the sequencing system further comprises one or more direct data access (DMA) connections, e.g., 5012 in FIG. 5C, that are in data communication with the plurality of data processing engines and the first reconfigurable routing channels, e.g., 5013 in FIG. 5C. The DMA connections may be configured to actively communicate data between components of the sequencing system. For example, the DMA connections may be configured to fetch data or send data to components that are connected thereto, e.g., the data processing engines, e.g., 5011 in FIG. 5C and the reconfigurable routing channels, e.g., 5013 in FIG. 5C. The DMA connections herein may be configured to actively request data from or actively sending data directly to: the first reconfigurable logic device; the first reconfigurable routing channels; the integrated circuit; or a combination thereof. One or more direct data access (DMA) connections may be in data communication the first reconfigurable routing channels and the integrated circuit herein. The DMA connections may be configured to allow data communication based on a predetermined protocol, e.g., a PCIe protocol.
[0167] In some embodiments, the first reconfigurable routing channels are configured to allow data communication between the first reconfigurable logic device and one or more memory devices. In some embodiments, the one or more DMA connections and the first reconfigurable routing channels are configured to allow data communication between the first reconfigurable logic device and the integrated circuit.
[0168] In some embodiments, the sequencing system further comprises an integrated circuit that is different from the first reconfigurable logic device, e.g., 120_b in FIG. 5C. The integrated circuit herein may not be reconfigurable. The integrated circuit may comprise an application specific integrated circuit (ASIC) chip. In some embodiments, the integrated circuit comprises a neural processing unit (NPU) or an artificial intelligence (Al) chip. The integrated circuit may comprise a second plurality of data processing engines, each data processing engine comprising multiple digital logic circuits. The integrated circuit may further comprise: second plurality of data processing engines and second routing channels, each connecting at least some of the second plurality of data processing engines.
[0169] In some embodiments, the sequencing system further comprises a first processor. The first processor may be configured to selectively activate or deactivate different combinations of the first plurality of data processing engines and the first reconfigurable routing channels to perform the operations disclosed herein.
[0170] In some embodiments, the first processor or a second processor is configured to selectively activate or deactivate different combinations of the second plurality of data processing engines and the second reconfigurable routing channels to perform the operations. The first processor or a second processor may be configured to selectively activate or deactivate different combinations of the second plurality of data processing engines and the second reconfigurable routing channels to perform the operations herein.
[0171] The sequencing system may further comprise a housing that encloses the first reconfigurable logic device, the first reconfigurable routing channels, the one or more DMA connections, the integrated circuit, and the first processor therein. In some embodiments, the sequencing system further comprises: a housing that encloses at least the first reconfigurable logic device therein and the integrated circuit is external to the housing.
[0172] In some embodiments, the sequencing system further comprises: a power source that is configured to supply different power levels to the first reconfigurable logic device and the integrated circuit. A first power level supplied by the power source to the first reconfigurable logic device may be higher than a second power level supplied to the integrated circuit while a sequencing run and / or sequencing analysis is in progress. A maximum power output of the power source of the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the maximum power output of the power source of sequencers, e.g., traditional sequencers without the first reconfigurable logic device (e.g., FPGA), the integrated circuit (e.g., Al chip), or both. The time consumption in performing a sequencing run and corresponding sequencing analysis (e.g., primary analysis) thereof using the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the time consumption in performing the same sequencing run using a sequencer without the first reconfigurable logic device, the integrated circuit, or both (e.g., a traditional sequencer without FPGA and / or Al chips). Time consumption in performing a sequencing run and sequencing analysis of the sequencing run (e.g., primary analysis) using the sequencing system is 2x, 3x, 5x, 8x, lOx, or 20x lower than the time consumption in performing the same sequencing run and analysis using a sequencer without the first reconfigurable logic device, the integrated circuit, or both(e.g., a traditional sequencer without FPGA and / or Al chips).
[0173] In some embodiments, a maximum power output of the power source to the sequencing system in performing a sequencing run and corresponding sequencinganalysis thereof is less than 900 Watts, 800 Watts, 700 Watts, 650 Watts, 600 Watts, 550 Watts, or 500 Watts.
[0174] In some embodiments, the sequencing system further comprises a power source configured to supply a first power level to the first reconfigurable logic device, the first power level is less than 500 Watts, 400 Watts, 350 Watts, or 300 Watts.
[0175] In some embodiments, the sequencing system further comprises a power source configured to supply a second power level to the integrated circuit, the second power level is less than 450 Watts, 400 Watts, 350 Watts, or 300 Watts.
[0176] In some embodiments, one or more components are located on a first printed circuit board (PCB). The one or more components may include: the first reconfigurable logic device the first reconfigurable routing channels; the first processor; and the one or more DMA connections. In some embodiments, the integrated circuit is located on a second printed circuit board (PCB) different from the first printed circuit board, e.g., as shown in FIG. 5C. The integrated circuit and the second PCB may be positioned within a same housing of the sequencing system as the first PCB or external to the housing of the sequencing system. Being on a separate PCB makes connecting the first reconfigurable logic device, e.g., FPGA device with various integrated circuit on a chip convenient, efficient, and easily customizable. In some embodiments, the first PCB board may be a main board, and the second PCB board may be a daughter board.
[0177] In some embodiments, the sequencing systems lacks any graphic processing units (GPUs) or tensor processing units (TPUs). Instead, the sequencing systems utilizes FPGAs, Al chips, NPUs, or other ASIC chips for performing the operations disclosed herein. The sequencing system disclosed herein advantageously requires less power, generate less heat, and reduces the hardware costs for performing NGS sequencing runs and corresponding sequencing analysis.
[0178] The first processor may be positioned on the first PCB board together with the reconfigurable logic device for convenient and efficient control of the reconfigurable logic device. In some embodiments, the first processor is a separate processor from one or more processors of the sequencing system configured to control the optical system, the fluidics of the sequencing system, etc. In some embodiments, the first processor can be configured to only control the components on the first PCB board, e.g., the FPGA device, alone or in combination with components on the second PCB board, e.g., the Al chip. In some embodiments, the sequencing system may comprise a second processor that isconfigured to separately control the Al chip. The first processor or second processor of the sequencing system, e.g., 120_c, may comprise a CPU. The one or more hardware processors of the sequencing system comprises a CPU.
[0179] Continuing referring to FIG. 5C, in this particular embodiment, the sensor data at the imager 116 can be communicated directly to the data processing engine(s) 5011 of the first reconfigurable logic device 120(a). Alternatively, the sensor data may be saved into a memory device, e.g., 5030 so that it can be accessed by the data processing engine. The first processor 120_c may control operation of the data processing engines and the routing channels to process the sensor data and generate the first plurality of flow cell images. The processing may include operations disclosed herein such as intensity normalization, color correction, phasing and prephasing correction, background subtraction, etc. The first plurality of flow cell images may then be communicated from the processing engines through the routing channels to the memory device 5030 so that the integrated circuit may be controlled by the first processor or a second processor to access the first plurality of flow cell images for subsequent steps in primary analysis. Alternatively, the first plurality of flow cell images may be directly communicated to the integrated circuit 120-b via DMA connections 5012. In this embodiment, the integrated circuit is only used for prediction higher resolution polony locations using a pretrained CNN, thereby generating the second plurality of flow cell images with a resolution that is at least 8 times higher than the resolution of the first plurality of flow cell images. The CNN may be pretrained using simulated images or real flow cell images. The second plurality of flow cell images are transmitted back from the integrated circuit to the first reconfigurable logic device for subsequent processing steps such as base calling. The base calls along with quality information may then be saved into a FastQ data file. Other information including cell segmentation and staining may also be saved in the same file or another FastQ file with compatible data format.
[0180] In some embodiments, the sequencing system may further comprise a heat dissipator configured to maintain a system temperature in a range from 0 degrees to 120 degrees Celsius or less than 120 degrees Celsius.
[0181] In some embodiments, the operation for processing the sensor data to generate the first plurality of flow cell images comprises one or more of registering the first plurality of flow cell images to a reference coordinate system; adjusting image intensities of the first plurality of flow cell images; color correction of the first plurality of flow cellimages; correcting phasing and prephasing of the first plurality of flow cell images; and subtracting background intensities from the first plurality of flow cell images.
[0182] In some embodiments, each of the one or more operations performed by the first reconfigurable logic device or the integrated circuit are performed within the time window of performing a single sequencing cycle of the sequencing run. FIG. 5D shows an exemplary embodiment of performing sequencing analysis in parallel with performing a sequencing run. In this particular embodiment, the sequencing run include multiple sequencing cycles. For each cycle, flow cell images are acquired at multiple z-levels from different color channels. The sequencing reactions are repeatedly performed for each z- level in each cycle within a time window 5601. The operations of the integrated circuit are performed within a processing window 5602 within the time window 5609 of a single sequencing cycle and also within a time window 5601 for sequencing reactions and imaging at a single z-level 5601. The operations of the first reconfigurable logic device (e.g., on board primary analysis operations) are also performed with a processing window 5603 that is within the time window 5609 of each sequencing cycle. The processing windows 5602 and 5603 may be identical or different depending on various factors such as sequencing data, primary analysis algorithms, etc. In some embodiments, the operations are not just performed within the processing windows but completed within the processing windows with respect to the data of the current cycle, e.g., a preceding z- level that sensor data has been acquired. In some embodiments, the operations are completed within the processing windows with respect to the data of a preceding cycle, e.g., the cycle immediately preceding the current cycle.
[0183] In embodiments where the sample is a 3D volumetric sample, the operations are performed for a single z level in each cycle within a predetermined time window, e.g., 5602, 5603. The predetermined time window is for a single z level in a single sequencing cycle. In some embodiments, the predetermined time window is less than 1000 ms, 900 ms, 800 ms, 700 ms, 600 ms, 500 ms, 400 ms, 300 ms, 250 ms, 200 ms, or 100 ms. In some embodiments, each of the one or more operations are performed within the predetermined time window and in parallel while the sequencing run is in progress. In some embodiments, each of the one or more operations are performed in parallel within a time window that sequencing, imaging, or both of a subsequent sequencing cycle is completed.
[0184] The first plurality of flow cell images herein may be obtained from multiple z levels covering at least partly of an in situ sample, e.g., of cells or tissue(s). The first plurality of flow cell images may be obtained from one or more color channels at each z level of the multiple z levels covering at least partly of the in situ sample. In some embodiments, the first plurality of flow cell images are from a single color channel. The first plurality of flow cell images may be of a first spatial resolution in x, y, and / or z directions. The second plurality of flow cell images may be generated based on the first plurality of flow cell images. The second plurality of flow cell images may be of a second spatial resolution in x, y, and / or z directions. The first spatial resolution may be lower than the second spatial resolution, and a higher resolution herein indicate that a pixel size is smaller so that the polonies in the flow cell images are of finer spatial details. The first spatial resolution may be 2x, 4x, 6x, 8x, lOx, 16x, 24x, 32x, or 48x lower than the second spatial resolution in x, y, and / or z directions. The first spatial resolution may be at least 2x, 4x, 6x, 8x, lOx, 16x, 24x, 32x, or 48x lower than the second spatial resolution in x,y, and / or z directions. In some embodiments, the first and second resolution is in 3D. In some embodiments, the first resolution is in a range of 0.1 um to 5 um. In some embodiments, the second resolution is in a range of 0.01 um to 2 um. In some embodiments, the second resolution is at least 4, 6, or 8 times greater than the first resolution in all three dimensions.
[0185] In some embodiments, the sequencing system further comprises one or more image sensors configured to receive optical signals generated from sequencing reactions of a sample immobilized on a support. The support may comprise a glass or plastic substrate. The support may be comprised in a flow cell device. The one or more image sensors may be configured to generated sensor data based on the optical signals. In some embodiments, the sequencing system further comprises: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors to cause the one or more hardware processors to perform operations disclosed herein. The one or more data storage devices may include one or more memory devices. The one or more memory devices may be accessible by the one or more processors, the first processor, the second processor, the first reconfigurable logic device, the integrated circuit.
[0186] In some embodiments, the one or more processors are separate from the first or second processors. The operations performed by the one or more processors may includeone or more of: 1) recording sensor data generated in the sequencing system in one or more flow cycles; 2) optionally processing the recorded sensor data; 3) sending the recorded sensor data or the optionally processed data to the first reconfigurable logic device or the integrated circuit; 4) receiving outcome from the first reconfigurable logic device or integrated circuit; and 5) generating sequencing analysis results based on the received outcome. The operations performed by the one or more processors may include one or more of: 1) receiving outcome from the first reconfigurable logic device or integrated circuit; and 2) generating sequencing analysis results based on the received outcome.
[0187] In some embodiments, the sequencing analysis results comprise primary analysis results. In some embodiments, the sequencing analysis results comprise a data file in a predetermined data format. In some embodiments, the sequencing analysis results comprise base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the sequencing analysis results comprises quality measurements of base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the sequencing analysis results comprises quality scores corresponding to base calls of nucleotide bases in a sample immobilized on a support.
[0188] In some embodiments, the sequencing system further comprises: a sample immobilized on a support; and an optical system comprising: an illumination system; an objective lens and the one or more image sensors. The optical system is configured to emit light to the sample and to collect optical signals emitted from the sample, thereby generating the first plurality of flow cell images. The support may be comprised in a flow cell device.
[0189] In some embodiments, the output data comprises base calls of nucleotide bases in a sample immobilized on a support. In some embodiments, the output data comprises identification of base calling locations in two dimensions. In some embodiments, the output data comprises identification of base calling locations in three dimensions.
[0190] In some embodiments, the operation(s) performed by the first reconfigurable routing channels or the integrated circuit using the neural network comprises one or more of: generating quality measurements of the base callings; and generating a data output file based on the base callings.
[0191] In some embodiments, the neural network comprises a convolutional neural network (CNN). In some embodiments, the neural network comprises a U-Net. In someembodiments, the neural network has been trained using the first reconfigurable logic device or the integrated circuit. In some embodiments, the first convolution comprises a 3D convolution with a convolution kernel. In some embodiments, the convolutional kernel have at least four dimension. In some embodiments, the convolutional kernel is m x m x m x n, wherein m is an integer in a range from 3 to 30, wherein n is an integer. In some embodiments, n is an integer from 1 to 16384. In some embodiments, the second convolution in operation (1) comprises a corresponding number of n, 2*n, 4*n, and 8*n filters in a first, second, third, and fourth repetition, respectively. In some embodiments, the second convolution in (4) comprises a corresponding number of 2*n, 2*n, 4*n, 8*n filters in a last repetition, last minus one, last minus two, and last minus three repetition, respectively. In some embodiments, n is in a range from 4 to 1024.
[0192] In some embodiments, the operation (e) performing a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images comprises: performing a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images and based on a fourth plurality of flow cell images, wherein the fourth plurality of images are predicted using a second neural network based on a third plurality of flow cell images. In some embodiments, the third plurality of flow cell images are acquired from one or more color channels that is different from the single channel, and wherein the third plurality of flow cell images comprises the first resolution. In some embodiments, the fourth plurality of flow cell images comprises the second resolution.
[0193] In some embodiments, the first plurality of flow cell images are from one or more color channels. In some embodiments, the first plurality of flow cell images are of unbalanced nucleotide diversity. In some embodiments, the first plurality of flow cell images comprises: an unbalanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more flow cycles. In some embodiments, the first plurality of flow cell images comprises: a balanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more cycles. In some embodiments, two or more different concatemer molecules among the concatemer molecules have different insert sequences. In some embodiments, different insert sequences correspond to different target RNA molecules or target cDNA molecules. In some embodiments, each location of the determined polonies corresponds to a location of the concatemer molecules. In some embodiments, the firstplurality of flow cell images comprises optical signals emitted from nucleotide reagents bound to a balanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support. In some embodiments, the first plurality of flow cell images comprises optical signals emitted from nucleotide reagents bound to a unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support in the one or more subsequent cycles. In some embodiments, the unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules comprises: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of bases is less than 20%, 15%, 10%, or 5% in the one or more cycles. In some embodiments, the balanced diversity of nucleotide bases of A, G, C and T / U among the plurality of concatemer molecules comprises: a percentage of (1) a number of each type of nucleotide bases to (2) a total number of bases in the one or more cycles is more than 10%, 15%, or 20%.
[0194] In some embodiments, the cellular sample comprises overloaded concatemer molecules with a spatial density in a range of 102-1015per mm2. In some embodiments, the cellular sample comprises overloaded concatemer molecules with a spatial density in a range of 103-1010per mm2.
[0195] In some embodiments, the down-sampling factor is 2, 4, or 8. In some embodiments, the up-sampling factor is 2, 4, or 8. In some embodiments, the downsampling factor is 2, 4, 8, 16, 32 or 64. In some embodiments, the up-sampling factor is 2, 4, 8, 16, 32, or 64.
[0196] In some embodiments, one or more of operations of (a) to (k) are performed while a sequencing run is being performed. In some embodiments, the first plurality of flow cell images are acquired in sequencing cycles ranging from 1 to 500. In some embodiments, the one or more cycles comprises a current cycle N. In some embodiments, wherein N is in a range from 1 to 500. In some embodiments, the one or more cycles comprises a single cycle ranging from 1 to 500. In some embodiments, the one or more cycles comprises multiple cycles ranging from 1 to 500. In some embodiments, one or more of operations, e.g., operations (a) to (j), are performed while the sequencing reactions in cycles subsequent to the current cycle N is yet to be performed or currently being performed. In some embodiments, the z-axis is orthogonal to image planes of the flow cell images.
[0197] In some embodiments, performing the first convolution in one or more dimensions on the first plurality of flow cell images comprises: performing a first convolution in 3D on the first plurality of flow cell images, thereby generating a first convolution result. In some embodiments, performing a second convolution in one or more dimensions on the first convolution result, thereby generating a second convolution result comprises: performing the second convolution in 3D on the first convolution result, thereby generating a second convolution result.
[0198] In some embodiments, performing the first convolution in one or more dimensions on the first plurality of flow cell images comprises: performing a first convolution in 2D on the first plurality of flow cell images, thereby generating a first convolution result. In some embodiments, performing a second convolution in one or more dimensions on the first convolution result, thereby generating a second convolution result comprises: performing the second convolution in 2D on the first convolution result, thereby generating a second convolution result.
[0199] In some embodiments, repetitively performing up sampling operations comprises: (3) performing an up sampling of the fourth convolution result by an up sampling factor thereby generating a first up-sampled result; (4) concatenating the first up-sampled result in a current up-sampling repetition with the first down-sampled result in a previous downsample repetition, wherein the first up-sampled result has a same size as the first down- sampled result in the previous down-sampling repetition; and (5) performing the second convolution in one or more dimensions of the first up-sampled result, thereby generating a fifth convolution result.
[0200] In some embodiments, the different combinations of the first plurality of data processing engines are configured to perform operations further comprising: (a) receiving the second plurality of flow cell images from the integrated circuit; (b) determining polonies from the second plurality of flow cell images; and (c) performing a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images; and (d) forwarding the second plurality of flow cell images, the determined polonies, the corresponding base callings, or a combination thereof to the first processor or one or more hardware processors of the sequencing system.
[0201] In some embodiments, the one or more operations performed by the first reconfigurable logic device further comprises: forwarding the second plurality of flow cell images, the determined polonies, the corresponding base callings, or a combinationthereof to the first processor or one or more hardware processors of the sequencing system. In some embodiments, the one or more operations performed by the integrated circuit further comprises forwarding the second plurality of flow cell images, the corresponding base callings, or both to the first reconfigurable logic device, the first processor or one or more hardware processors of the sequencing system.
[0202] In some embodiments, the operations performed by the integrated circuit further comprising one or more of: determining polonies from the second plurality of flow cell images; performing a corresponding base call for each of the determined polonies based on the second plurality of flow cell images; and forwarding the second plurality of flow cell images, the corresponding base callings, or both to the first reconfigurable device, the first processor, or one or more hardware processors of the sequencing system.
[0203] In some embodiments, the operations performed by the first reconfigurable logic device or the integrated circuit further comprising: registering the second plurality of flow cell images to a common coordinate system.
[0204] In some embodiments, the operation (d) or (i) of determining polonies from the second plurality of flow cell images comprises: generating a 3D polony map comprising spatial location of polonies based on the determined polonies. The operation of generating a 3D polony map comprising spatial location of polonies based on the determined polonies may further comprise: deleting duplicate polonies from the determined polonies, wherein the duplicate polonies are out-of-focus. In some embodiments, the operation of determining polonies from the second plurality of flow cell images comprises: superimposing the second plurality of flow cell images with corresponding cell staining images; and generating the polony map by only including polonies that are within cell boundaries in the corresponding cell staining images. Exemplary embodiments of methods for generating 3D polony map are disclosed in U.S. Patent Application No. 18 / 078,820 and PCT Application No. PCT / US2023 / 076125, which are incorporated by reference in their entireties.
[0205] In some embodiments, the method further comprises: providing the cellular sample harboring a plurality of RNA which comprises the first target RNA molecule and the second target RNA molecule. In some embodiments, the method further comprises: generating inside the cellular sample a plurality of cDNA molecules which include a first target cDNA molecule that corresponds to the first target RNA molecule and a second target cDNA molecule that corresponds to the second target RNA molecule. In someembodiments, the method further comprises: contacting the plurality of cDNA molecules in the cellular sample with a plurality of target-specific padlock probes which includes at least a first plurality of first target-specific padlock probes and a second plurality of second target-specific padlock probes.
[0206] In some embodiments, the method further comprises: contacting the plurality of RNA molecules in the cellular sample with a plurality of target-specific padlock probes which includes at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes.
[0207] In some embodiments, individual padlock probes in the first plurality of first target-specific padlock probes comprise: first and second terminal regions, wherein the first terminal region selectively hybridizes to a first region of the first target cDNA molecule or the first target RNA molecule, and the second terminal region selectively hybridizes to a second region of the first target cDNA molecule or the first target RNA molecule.
[0208] In some embodiments, contacting the plurality of RNA molecules in the cellular sample with the plurality of target-specific padlock probes comprises: hybridizing the first and second terminal regions of the first target-specific padlock probes to proximal positions on the first target cDNA molecule or the first target RNA molecule to form a circularized first target-specific padlock probe having a nick or gap between the hybridized first and second terminal regions. In some embodiments, the first targetspecific padlock probe comprises a first target barcode sequence that corresponds to and uniquely identifies the first target cDNA sequence or the first target RNA sequence. In some embodiments, the first target-specific padlock probe comprises a first target barcode sequence that is located adjacent to one of the regions of the first target-specific padlock probe that selectively hybridizes to the first target cDNA molecule or the first target RNA sequence. In some embodiments, the first target-specific padlock probe comprises at least one universal adaptor sequence. In some embodiments, the first target-specific padlock probe comprises a universal primer binding site for a rolling circle amplification primer or a complementary sequence thereof. In some embodiments, the first target-specific padlock probe comprises a universal compaction oligonucleotide binding site or a complementary sequence thereof. In some embodiments, the method further comprises: closing the nick or gap in the at least first and second circularized target-specific padlock probes by conducting an enzymatic reaction, thereby generating at least a first covalentlyclosed circular padlock probe and a second covalently closed circular padlock probe inside the cellular sample. In some embodiments, the method further comprises: conducting a rolling circle amplification reaction inside the cellular sample using the first and second covalently closed circular padlock probes as template molecules, thereby generating a plurality of concatemer molecules including at least the first concatemer molecule that corresponds to the first target RNA molecule, and the second concatemer molecule that corresponds to the second target RNA molecule. In some embodiments, the first concatemer comprises: tandem repeat units of: a first target barcode sequence that uniquely identifies the first target RNA or the first target cDNA sequence, a first insert sequences that corresponds to the first target RNA or the first target cDNA, and a first sequencing primer binding site or a complementary sequence thereof. In some embodiments, the first concatemer further comprises: a universal binding site for an amplification primer or a complementary sequence thereof, and a universal binding site for a compaction oligonucleotide or a complementary sequence thereof. In some embodiments, the second concatemer comprises: tandem repeat units of: a second target barcode sequence that uniquely identifies the second target RNA or the second target cDNA sequence, a second insert sequences that corresponds to the second target RNA or the second target cDNA, and a second sequencing primer binding site or a complementary sequence thereof. In some embodiments, the second concatemer further comprises: a universal binding site for an amplification primer or a complementary sequence thereof, and a universal binding site for a compaction oligonucleotide or a complementary sequence thereof.
[0209] In some embodiments, conducting the one or more cycles of sequencing reactions comprises: contacting the plurality of concatemer molecules inside the cellular sample with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents, under a condition suitable for hybridizing the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the concatemers. In some embodiments, the plurality of nucleotide reagents comprise: multivalent molecules, nucleotides, nucleotide analogs, or their combinations. In some embodiments, individual nucleotides or nucleotide analogs are detectably labeled or non-labeled. In some embodiments, the detectably labeled individual nucleotides or nucleotide analogs comprises a different detectable color label that corresponds with each different type of nucleotide base of A, G, C, and T / U. In someembodiments, an individual multivalent molecule comprise a core attached with multiple nucleotide arms and each arm of the individual multivalent molecule comprises the same type of nucleotide base.
[0210] In some embodiments, generating the first plurality of flow cell images comprises: in each cycle, imaging, by an optical system, optical color signals emitted from the nucleotide reagents that are bound to the plurality of concatemer molecules. In some embodiments, the first plurality of flow cell images comprises optical color signals emitted from the nucleotide reagents that are bound to the plurality of concatemer molecules. In some embodiments, conducting the one or more cycles of sequencing reactions comprises: sequencing only the first target barcode sequence region of the first concatemer, thereby generating the first sequencing read product. In some embodiments, conducting the one or more cycles of sequencing reactions comprises: sequencing the first target barcode sequence region and at least a portion of the first insert sequence of the first concatemer, thereby generating the first sequencing read product.
[0211] In some embodiments, conducting the one or more cycles of sequencing reactions comprises: sequencing only the second target barcode sequence region of the second concatemer, thereby generating the second sequencing read product. In some embodiments, conducting the one or more cycles of sequencing reactions comprises: sequencing the second target barcode sequence region and at least a portion of the second insert sequence of the second concatemer, thereby generating the second sequencing read product.
[0212] In some embodiments, the method further comprises: removing a first sequencing read product from the first concatemer molecule and retaining the first concatemer molecule in the cellular sample, and removing a second sequencing read product from the second concatemer molecule and retaining the second concatemer molecule in the cellular sample. In some embodiments, the method further comprises: reiteratively sequencing the plurality of concatemers by repeating the following operations for at least once: generating the first plurality of flow cell images of a cellular sample immobilized on a support by conducting one or more cycles of sequencing reactions thereby generating the first sequencing read product and the second sequencing product, the cellular sample comprising a plurality of concatemer molecules therewithin, wherein a first concatemer molecule of the plurality of concatemer molecules corresponds to a first target RNA molecule of the cellular sample, and a second concatemer molecule of the plurality ofconcatemer molecules corresponds to a second target RNA molecule of the cellular sample, wherein the first plurality of flow cell images; and removing a first sequencing read product from the first concatemer molecule and retaining the first concatemer molecule in the cellular sample, and removing a second sequencing read product from the second concatemer molecule and retaining the second concatemer molecule in the cellular sample.
[0213] In some embodiments, the first sequencing read product comprises some or all of: a first target barcode sequence in one or more tandem units of the first concatemer molecule; a first insert sequence in one or more tandem units of the first concatemer molecule; or their combinations.
[0214] In some embodiments, the method further comprises: confirming presence of the first target RNA molecule, the second target RNA molecule, or both molecules in the cellular sample based on the performed base calling of the second plurality of flow cell images at the base calling locations in the base calling template.
[0215] In some embodiments, the method further comprises: generating, by the sequencing system, the second plurality of flow cell images of the cellular sample immobilized on the support by conducting subsequent cycles of sequencing reactions after the one or more cycles. In some embodiments, generating the first plurality of flow cell images of the cellular sample immobilized on the support comprises: sequencing at least the first concatemer inside the cellular sample under a condition that inhibits sequencing the second concatemer. In some embodiments, sequencing at least the first concatemer inside the cellular sample comprises: generating a plurality of first sequencing read products, and wherein the sequences of the first sequencing read products are aligned with a first target reference sequence to confirm presence of the first target RNA in the cellular sample. In some embodiments, generating the first plurality of flow cell images of the cellular sample immobilized on the support comprises: sequencing at least the second concatemer inside the cellular sample under a condition that inhibits sequencing the first concatemer. In some embodiments, sequencing at least the second concatemer inside the cellular sample comprises: generating a plurality of second sequencing read products, and wherein sequences of the second sequencing read products are aligned with a second target reference sequence to confirm presence of the second target RNA in the cellular sample.Predicting high resolution flow cell images
[0216] FIG. 5A shows a flow chart of a computer-implemented method 500 for predicting high resolution flow cell images thereby improving detectable polony density in the flow cell images. The method 500 can include some or all of the operations disclosed herein. The operations may be performed in but is not limited to the order that is described herein.
[0217] The method 500 can be performed by one or more processors disclosed herein. In some embodiments, the processor can include one or more of: a computing system comprising a processing unit 118, a reconfigurable logic device 120, an integrated circuit that is not reconfigurable 120, or their combinations. For example, the processing unit can include a central processing unit (CPU). The reconfigurable logic device can include one or more FPGA devices. The integrated circuit can include a chip such as an Al chip or an ASIC chip. In some embodiments, the one or more processors can include the computer system 400 disclosed herein.
[0218] In some embodiments, some or all operations in method 500, 600, 700, 2800, and 2900 can be performed by the reconfigurable logic device, e.g., the FPGA(s), and / or the integrated circuit, e.g., the Al chip. In embodiments when some operations are performed by the reconfigurable logic device and / or integrated circuit, e.g., FPGA(s), the data produced by the reconfigurable logic device and / or integrated circuit, e.g., the FPGA(s) after performing one or more operations, can be communicated to various hardware elements of the system 100, e.g., CPU(s) or GPU(s), so that subsequent operation(s) in method 500, 600, 700, 2800, and 2900 can be performed by such various hardware using the communicated data. Similarly, data can also be communicated in the opposite direction from various hardware e.g., CPU(s), to the reconfigurable logic device or the integrated circuit for processing. In some embodiments, all the operations in method 500, 600, 700, 2800, and 2900 can be performed by CPU(s). Alternatively, the operations performed by CPU(s) can be performed by other processors such as the dedicated processors, or GPU(s). In some embodiments, all the operations in method 500, 600, 700, 2800, and 2900 can be performed by the reconfigurable logic device and / or the integrated circuit, e.g., FPGA(s) and / or the Al chip(s).
[0219] In some embodiments, the sensor data acquired by the imager 116 may be directly communicated to the reconfigurable logic device and / or the integrated circuit, e.g., via DMA connections. In some embodiments, the sensor data acquired by the imager 116may be directly communicated to the reconfigurable logic device and / or the integrated circuit without being routed first to a CPU, a GPU, or any other processing units before reaching the reconfigurable logic device and / or the integrated circuit.
[0220] In some embodiments, predicting high resolution flow cell images using the methods 500 herein with the reconfigurable logic device, e.g., the FPGA, and / or other integrated circuit, e.g., Al chips, may require at least 2x, 8x, lOx, 15x, 20x, 40x, 50x, or lOOx less power than making the same predict! on(s) using other computing hardware including but not limited to CPUs or GPUs.
[0221] In some embodiments, the sequencing system herein further comprises: a power source that is configured to supply identical or different power levels to the reconfigurable logic device and the integrated circuit. In some embodiments, a maximum power output of the power source to the sequencing system in performing methods 500, 600, 700, 2800, and / or 2900 is less than 2000 Watts, 1000 Watts, 900 Watts, 800 Watts, 700 Watts, 650 Watts, 600 Watts, 550 Watts, 500 Watts, 400 Watts, 300 Watts, 200 Watts, or 100 Watts.
[0222] In some embodiments, the sequencing system herein comprises: a first reconfigurable logic device, e.g., a FPGA unit, comprising a plurality of data processing engines configured to perform data processing in parallel; first reconfigurable routing channels, each connecting at least some of the first plurality of data processing engines; a neural network deployed at least partly on the first reconfigurable logic device; a first processor to selectively activate or deactivate different combinations of the first plurality of data processing engines and the first reconfigurable routing channels to perform one or more operations in methods herein (e.g., methods 500, 2800) to make predictions.
[0223] In some embodiments, the sequencing system herein comprises: a first reconfigurable logic device comprising a first plurality of data processing engines arranged in a first pipeline and configured to perform data processing in parallel with each other; an integrated circuit in data communication with the first reconfigurable logic device; a neural network deployed at least partly on the integrated circuit and / or the first reconfigurable logic device; a first processor of the first reconfigurable logic device to selectively activate or deactivate different combinations of the first plurality of data processing engines to perform one or more operations in methods herein (e.g., methods 500, 2800) to make prediction using the neural network.
[0224] In some embodiments, the first reconfigurable logic device and the integrated circuit is within the same physical housing as the other elements of the sequencing system as show in FIG 1. In some embodiments, the first reconfigurable logic device and the integrated circuit are not physically external to the sequencing system 110 as shown in FIG. 1, e.g., not in the cloud 130.
[0225] The method 500 can comprise an operation 510 of (i) generating, by the sequencing system 110, a first plurality of flow cell images of sample(s) immobilized on a support by conducting one or more cycles of sequencing reactions. The sample(s) may comprise concatemer molecules therewithin. The sample(s) may include concatemer molecules from one or more different sample sources. The sample(s) may include a thickness along the z-axis so that the first plurality of flow cell images may be acquired at a z-stack of different z-locations with a first resolution to cover the sample in 3D. The samples may be acquired from a single z-location of a 2D or 3D sample.
[0226] The sample can be in situ. The sample can be a 3D sample. The sample can be a volumetric sample that may contain different biological information at the same x-y location but different z levels. The sample can be a cellular sample including multiple cells, tissue, or their combination. The sample can be any biological sample that has a thickness that is greater than a predetermined threshold along the z axis. For example, the thickness can be greater than 2 um, 3 um, 4 um, 5 um, 10 um, 20 um, or more. The z axis (e.g., z axis) is orthogonal to the image plane defined by x and y axes. In some embodiments, the sample can be traditional 2D sequencing samples.
[0227] In some embodiments, such computer-implemented method comprises an operation (i) of generating, by a sequencing system, a first plurality of flow cell images of a sample immobilized on a support by conducting one or more cycles of sequencing reactions, wherein the first plurality of flow cell images are acquired with a first resolution. Such operation is similar to operation 510 in FIG. 5 A except that the sample may be 2D or 3D sample. In embodiments where the sample is 3D, the sample comprises concatemer molecules therewithin. In embodiments where the sample is 2D, the sample comprises template molecules therewithin.
[0228] The flow cell images can be acquired using the optical system of the imager 116 disclosed herein, from the 1, 2, 3, 4, or more color channels. Each flow cell image can include at least a portion of one or more tiles (e.g., imaging areas). Each tile can be divided into multiple subfiles. Each tile or subtile can include a plurality of polonies orclusters. Each subtile can include multiple regions with each region including a number of polonies or clusters. The flow cell image as disclosed herein can be an image that is acquired from a flow cell 112 as shown in FIG. 1 or 2712 as shown in FIG. 27. In some embodiments, the flow cell images are acquired from a single color channel, and subsequent prediction is by using a pretrained neural network corresponding to that single channel. In some embodiments, the flow cell images are acquired from 2, 3, 4, or more color channels, and subsequent prediction is by using a pretrained neural network corresponding to the multiple color channels.
[0229] In some embodiments, a flow cell image herein can be an image of one or more tiles, one or more subtiles, one or more segmented regions within tile(s) or subtile(s), or their combinations. Each flow cell image can comprise a field of view (FOV). The FOV can be orthogonal to the z axis. The FOV can be within the x-y plane. The FOV of different flow cell images at different z levels can be identical within the x-y plane. The FOV of different flow cell images at different z levels can have at least an overlapping portion within the x-y plane. The image resolution of different flow cell images at different z levels can be about identical or exactly identical. In some embodiments, The image resolution of different flow cell images at different z levels is different. FIGS. 3A and 3D show two exemplary flow cell images acquired at two different z levels along the z axis of a same 3D sample within a same sequencing cycle. The FOV can be in 3D and be of various sizes to cover the volumetric sample to be imaged. The FOV along x, y, and / or z direction can be in a range from 10 um to 5 mm. The FOV along x, y, and / or z direction can be in a range from about 0.1 um to about 2 mm. The FOV along x, y, and / or z direction can be in a range from 0.5 um to 1 mm. For example, the FOV can be about 0.5 mm by 0.5 mm by 20 um for certain cellular samples along the x, y, and z direction, respectively.
[0230] The flow cell images herein may be of various sizes, the pixel number along x, y, and / or z axis may be any integer greater than 64 or 128. The flow cell images herein may be of various sizes, the pixel number along x, y, and / or z axis may be in a range from 2 to 65536. A single flow cell image can be separated into different number of regions, for example, 4, 8, 16, or even more regions, and each region may include a size of 256 by 256 by 1, 512 by 512 by 3, or other sizes. In some embodiments, the number of pixels along x, y, and / or z direction may be adjusted to maintain a particular spatial resolution ina given FOV. For example, with a spatial resolution of 0.2 um, to cover a FOV of 0.8 mm, the number of pixels may be 4000.
[0231] Each flow cell image at a specific z level may include intensities generated by polonies or clusters at the corresponding z level. As shown in FIGS. 3 A and 3D, signals from polonies or clusters are small bright spots within the images. Each bright spot can be of various sizes that is less than a couple of pixels, e.g., less than a pixel, about a pixel, about 2 pixels, 3 pixels, 4, pixels, 5 pixels, or more. In some embodiments, each signal spot of the polonies or clusters can be any number of pixels in the range from 0.01 pixel to about 100 pixels. In some embodiments, each signal spot of the polonies or clusters can be any number of pixels in the range from 0.1 pixel to about 16 pixels.
[0232] Each flow cell image can also include intensities generated by the cell and its structural elements. Such structural elements can be background objects or components, e.g., in FIG. 3 A. Each flow cell images can also include noise and / or artifacts that are not from the polonies or cellular structures.
[0233] In some embodiments, when the depth of field the optical system includes a range, e.g., 0.1 um, 0.2 um, 0.3 um, 0.5 um, 0.6 um, 0.8 um, 1 um, 2 um, 3, um, 4 um, 5 um, etc. expanding along z axis, polonies or clusters that are within the range of depth of field can appear in-focus or about in-focus in the flow cell image. Flow cell images at a specific z level can also include signals from polonies or clusters that are not within the focus range of the image. So, such polonies or clusters are out-of-focus. As shown in FIG. 3 A, bigger and blurry signal spots represent out-of-focus polonies or clusters. Some of the out-of- focus polonies or clusters are circled in FIG. 3 A.
[0234] Each flow cell image at a specific z level can also include noises caused by the optical system and / or undesired signal from the sample. The undesired signal can be signal coming from components of the sample such as membrane, cytosol, and mitochondria. Such background objects can be any objects, relatively larger in size than the polonies or clusters. As shown in FIG. 3 A, there is a blurry cellular contour (at the arrows) in the flow cell image, and most of the signal spots are contained within the blurry contour. In some embodiments, background objects can include any objects within the 3D sample but are not polonies or clusters.
[0235] In some embodiments, the flow cell images are from multiple color channels. In some embodiments, the flow cell images are of unbalanced nucleotide diversity. In some embodiments, the flow cell images comprises: an unbalanced diversity of nucleotidebases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more sequencing cycles. In some embodiments, the flow cell images comprises: a balanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more cycles. In some embodiments, two or more different concatemer molecules among the concatemer molecules have different insert sequences. In some embodiments, different insert sequences correspond to different target RNA molecules or target cDNA molecules. In some embodiments, each location of the determined polonies corresponds to a location of the concatemer molecules. In some embodiments, the flow cell images comprises optical signals emitted from nucleotide reagents bound to a balanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support. In some embodiments, the flow cell images comprises optical signals emitted from nucleotide reagents bound to a unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules immobilized on the support in the one or more subsequent cycles. In some embodiments, the unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules comprises: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of bases that is less than 20%, 15%, 10%, or 5% in the one or more sequencing cycles. In some embodiments, the balanced diversity of nucleotide bases of A, G, C and T / U among the plurality of concatemer molecules comprises: a percentage of (1) a number of each type of nucleotide bases to (2) a total number of bases in the one or more cycles is more than 10%, 15%, or 20%. As an example, bases calls from the polonies include 4 different bases, and percentage of polonies for each of the 4 different bases can be greater than about 10% so that the data are of balanced diversity. As another example, bases called from the plurality of polonies includes 4 or less different bases, and percentage of polonies for one or more bases can be less than about 10%, and such data can be considered as unbalanced diversity. In some embodiments, bases called from the plurality of polonies include 4 or less different bases, and percentage of polonies for some of the bases can be less than about 5%, about 2%, or even about 1%, and such data can be considered as unbalanced diversity. As yet another example, the unbalanced diversity data include bases A, T, C, G in the plurality of polonies, and their percentages of the total base calls are about 1%, about 2%, about 1%, and about 95%, respectively. In addition to the base biases affecting diversity, plexity can also be a factor that when plexity is lower than a number, e.g., 8 or 16, the signal is of unbalanced diversity.
[0236] The method 500 is configured to predict high resolution flow cell images even if the polonies in the acquired flow cell images are of unbalanced diversity in one or more sequencing cycles.
[0237] In some embodiments, the method 500 comprises an operation 520 of (ii) providing, by a processor, the first plurality of flow cell images as an input to a neural network (e.g., CNN), wherein the neural network is pre-trained using a training data set of training flow cell images using a training method 600 herein. The neural network is pretrained so that the values of parameters of the neural network has been optimized based on the training. The neural network may be retrained when needed, for example, for predicting flow cell images from different cellular samples.
[0238] In some embodiments, the computer-implemented method 500 may be used to predict high resolution flow cell images that are at higher resolution than the first plurality of flow cell images (e.g., 2x, 4x, or 6x along at least one spatial dimension) acquired by the imager 116. In some embodiments, the high resolution flow cell images may be post image-processing images of the first plurality of flow cell images, e.g., by going through the image processing part 3120 of the neural network in FIG. 31. Image processing herein may include various image processing steps including but are not limited to: background removal, background reduction, artifact removal, artifact suppression, adjusting signal to noise ratio, adjusting contrast to noise ratio, intensity normalization, intensity offset correction, noise reduction, color correction, phasing or dephasing correction, image registration, and deconvolution.
[0239] In some embodiments, the neural network in operation 520 is a first neural network that can be trained using method 700 disclosed herein.
[0240] In some embodiments, the method 500 comprises an operation 520’ in replacement of the operation 520. In some embodiments, the operation 520 includes (ii) providing, by a processor or a first reconfigurable logic device, the first plurality of flow cell images as an input to a neural network, wherein the neural network is pre-trained using a training data set of training flow cell images and reference base calls of the training dataset.
[0241] In some embodiments, the operation 520’ is similar to the operation 520, e.g., as shown in FIG. 5 A, with the exception of a different neural network. In some embodiments, the operation 520’ may replace the operation 520 in method 500. In someembodiments, the neural network in operation 520’ is a different neural network from that in operation 520.
[0242] In some embodiments, the neural network in operation 520 is a first neural network, and the neural network in operation 520’ is a second neural network that is different from the first neural network in operation 520. The difference(s) among the first and second neural networks may include but is not limited to: different types of neural networks, differences in values of parameters, number of parameters, number of convolutional layers, number of layers, or a combination thereof.
[0243] In some embodiments, the second neural network in operation 520’ is a different neural network that is pretrained using the same training data set of flow cell images as that used for training the first neural network in operation 520.
[0244] In some embodiments, the second neural network in operation 520’ is a different neural network that is pretrained using a different training data set of flow cell images as that used for training the first neural network in operation 520. In some embodiments, the second neural network in operation 520’ is pre-trained using reference base calls of the training dataset.
[0245] In some embodiments, the first neural network in operation 520 is pretrained using reference images or reference intensities as ground truths, e.g., reference high resolution images or reference intensities in high resolution images, and the second neural network in operation 520 is pre-training using reference base calls of the training flow cell images in the training datasets as ground truths.
[0246] In some embodiments, the reference base calls may be generated using various base calling methods including those methods disclosed herein in relation to training methods for predicting base calls herein. In some embodiments, the reference base calls may be generated using methods that lacks usage of a neural network. Exemplary embodiments of generating base calls from flow cell images are disclosed in U.S. Patent Application No. 18 / 078,820 and PCT Application No. PCT / US2023 / 076125, which are incorporated by reference in their entireties.
[0247] In some embodiments, the second neural network in operation 520’ may be trained using a training method similar to methods 700 in FIG. 5E, In such embodiments, the reference intensities are not used in operations, e.g., operations 725, 730, and 755. Instead, reference base calls are used in such operations, e.g., operations 725’, 730’, and 755.
[0248] In some embodiments, the loss function for training the second neural network in operation 520’ may be different from the loss function used in training the first neural network in the operation 520. In some embodiments, various loss functions may be used for training the second neural network in operation 520’. In some embodiments, the second neural network is pre-trained using one or more loss functions based on comparing training base calls of the training flow cell images to the reference base calls of the training flow cell images. In some embodiments, the loss function for training the second neural network in operation 520’ may be based on comparison of training outputs, e.g., base calls, to the reference base calls. In some embodiments, training of the second neural network in operation 520’ may be completed when the loss function satisfies a predetermined criteria. The predetermined criteria can be customized to include various aspects of training outputs. In some embodiments, the predetermined criteria is determined based on the comparison of training base calls to reference base calls. In some embodiments, the predetermined criteria is based on the correctness of the training base calls in comparison to the reference base calls. In some embodiments, the predetermined criteria is at least based on training time that has been spent.
[0249] FIG. 31 is a block diagram showing an exemplary embodiment of the first and second neural networks and the method for training such neural networks.
[0250] It is worth noting that although neural network 3110 is disclosed in this exemplary embodiment, the neural network 3110 herein may be any artificial intelligence-based or machine learning based model that can include an imaging processing part 3120 and a base calling part 3130. Similarly, for the imaging processing part 3120 and the base calling part 3130, each of them may be any artificial intelligence-based or machine learning based model that may achieve similar functions as the neural network-based equivalent
[0251] In some embodiments, the neural network 3110 may be the first neural network in operation 520 or the second neural network in operation 520’. The method for training the neural network 3110 may be method 700 as an example. The neural network may include two separate parts, the first part is the image processing part 3120, and the second part is the base calling part 3130.
[0252] The image processing part 3120 is configured to perform one or more image processing steps disclosed herein, e.g., in relation to method 500, on the flow cell images herein, e.g., the first or second plurality of flow cell images. The one or more imageprocessing steps may include but are not limited to: background removal, background reduction, artifact removal, artifact suppression, adjusting signal to noise ratio, adjusting contrast to noise ratio, intensity normalization, intensity offset correction, noise reduction, color correction, phasing or dephasing correction, image registration, intensity extraction, and deconvolution.
[0253] In some embodiments, the base calling part 3130 is configured to perform base calling using the output images 3150 from the image processing part 3120 of the neural network. The base calling part 3130 may be configured to perform some image processing steps including but not limited to intensity extraction, color correction, and / or phasing or dephasing correction in embodiments where such image processing steps are not performed in the image processing part 3120 of the neural network.
[0254] The first or second part of the neural network 3120, 3130 may each include one or more structural elements of the neural network such as a convolutional layer. In some embodiments, the first or second part of the neural network 3120, 3130 may include one or more embedding layers of the neural network. The first part of the neural network 3120 may include at least part of an encoder of the neural network, and the second part of the neural network may include at least part of an decoder of the neural network. In some embodiments, the second part of the neural network may include at least part of: a convolutional layer, a pooling layer, a fully connected layer, a SoftMax layer, an input layer, an output layer, an embedding layer, an encoder, and a decoder of the neural network.
[0255] In some embodiments, the base calling part 3130 may lack any structural element of a neural network, e.g., a convolutional layer or a pooling layer. In some embodiments, the base calling part 3130 may lack any artificial-intelligence based algorithm. In some embodiments, the base calling part 3130 may lack any convolutional layers of the neural network. In some embodiments, the base calling part 3130 may lack any part of an embedding layer or a decoder of the neural network. In some embodiments, the base calling part 3130 may lack any part of: a convolutional layer, a pooling layer, a fully connected layer, a SoftMax layer, an input layer, an output layer, an embedding layer, an encoder, and a decoder of the neural network. In some embodiments, the base calling part 3130 may only comprise non-neural network base calling algorithm(s). As an example, the neural network 3110 that generates the output images 3150 is the second neuralnetwork in operation 520’ . As another example, the neural network 3110 that generate the output base calls 3160 is the neural network disclosed in relation to method 2800.
[0256] In embodiments where the base calling part 3130 lack any elements of a neural network or an artificial intelligence based algorithm, training of the neural network may include training of one or more parameters of the base calling part 3130. For example, the one or more parameters may include a feature size. In such embodiments, during training, the back propagation for finding adjustments of values for parameters of the neural network 3110, e.g., originating from the loss function, the references, and the output base calls, goes through the base calling part 3130 with making any adjustment to parameters of the base calling part 3130 and the image processing part 3120 (with adjustment of parameters) as the solid gray line with arrow shown in FIG. 31.
[0257] In some embodiments where the base calling part 3130 lacks any elements of a neural network or an artificial intelligence based algorithm, training of the neural network does not include training of any parameters of the base calling part 3130. During training, the back propagation for finding adjustments of values for parameters of the neural network 3110, e.g., originating from the loss function, the references, and the output base calls, may go through the base calling part 3130 without making any adjustment to parameters of the base calling part 3130 and then the image processing part 3120 (but with adjustment of parameters) as the solid gray line with arrow shown in FIG. 31.
[0258] In embodiments where the base calling part 3130 comprises at least some elements of a neural network or an artificial intelligence based algorithm, training of the neural network may include training of the base calling part 3130 and the image processing part 3120, including adjusting parameters from both parts, as shown in FIG. 31 as the solid gray line with arrow.
[0259] In other embodiments where the base calling part 3130 lack any elements of a neural network or an artificial intelligence based algorithm, training of the neural network may only include training of the image processing part 3120 but not training of any of the parameters in the base calling part 3130. In such embodiments, during training, the back propagation for updating thereby training the parameters of the neural network 3110, e.g., originating from the loss function, the references, and the output base calls, goes directly to the image processing part 3120 without going through the base calling part 3130 as shown in FIG. 31 as the dotted grey line with arrow. In other words, in such embodiments, the base calling part 3130 is not trained and the parameters in the basecalling part 3130 are fixed. In such embodiments, the loss function may be based on the output of the base calling part 3130, and the value of the loss function may be determined based on the output of the base calling part.
[0260] After the neural network is trained, the input 3140 may go through the image processing part 3120 to generate the output images 3150. In some embodiments, the output images 3150 comprise the second plurality of flow cell images, e.g., disclosed herein in relation to methods 500. In some embodiments, the output images 3150 comprise high resolution post-processing images corresponding to the input images 3140. The output images may go through the base calling part 3130 to generate the base calls 3160. Alternatively, the output images 3150 may go through various base calling algorithms, e.g., non-neural network based traditional base calling algorithms, but not the base calling part 3130 for generating the base calls. In such embodiments, after the model is trained, only image processing part 3120 of the neural network is used for making predictions, but not the base calling part 3130 of the neural network. In such embodiments, the neural network 3110 advantageously reduces the time required to make predictions, and reduces the computational burden and power required to make the prediction comparing with existing neural networks that predicts base calls.
[0261] In some embodiments, the input images 3140 comprise raw flow cell images acquired at the imager 116. In some embodiments, the input images 3140 comprise the first plurality of flow cell images disclosed herein. In some embodiments, the input images 3140 may be from multiple color channels and multiple sequencing cycles. In some embodiments, the input images 3140 may be from multiple color channels and a single sequencing cycle. In some embodiments, the input images 3140 may be from a single color channel and multiple sequencing cycles. In some embodiments, the input images 3140 may be from a single z level or multiple z levels.
[0262] Continuing referring to FIG. 31, in embodiments of training the neural network 3110, e.g., using methods 700, references or ground truths 3180 can be used for comparison of the output base calls 3160, and the value of the loss function 3170 can be calculated based on such comparison. The value of the loss function then can be used during training for back propagation into the neural network 3110 for adjusting values of the parameters of the neural network 3110, e.g., gradients. In some embodiments, adjusting parameters of the neural network may include parameters of the base calling part 3130 and the image processing part 3120. In other words, both parts 3120, 3130 aretrained during training of the neural network, e.g., using the training methods herein 700, 2900. In some embodiments, the value of the loss function then can be back propagated into the neural network 3110 for adjusting values of the parameters of only the image processing part 3120, but not the base calling part 3130. In other words, in such embodiments, only the image processing part 3120, but not the base calling part 3130, is trained during training of the neural network, e.g., using the training methods herein 700, 2900. In some embodiments, the neural network 3110 that is trained only on the image processing part 3120 is the second neural network in operation 520’. In some embodiments, the neural network 3110 that is trained on both the image processing part 3120 and the base calling part 3130 is the second neural network in operation 520’ .
[0263] In some embodiments, the second neural network in operation 520’ comprises a convolutional neural network. In some embodiments, the second neural network in operation 520’ comprises a recurrent neural network. In some embodiments, the second neural network in operation 520’ comprises a U-Net, residual U-Net, ResNet (residual neural network), and / or a LSTM (long short-term memory) neural network.
[0264] In some embodiments, the training flow cell images are acquired only from a same color channel. In some embodiments, each of the training flow cell images comprise flow cell images of a same field of view from a plurality of sequencing cycles stacked along a time dimension. The plurality of sequencing cycles may be of a same sequencing run. The plurality of sequencing cycles may be consecutive sequencing cycles in the sequencing run. In some embodiments, each of the training flow cell images comprise flow cell images of a same field of view from one or more sequencing cycles. In some embodiments, each of the training flow cell images comprise flow cell images of the sample at one or more z-levels. In some embodiments, the training flow cell images comprise flow cell images of the sample at multiple different field of views of the same sample. In some embodiments, the training flow cell images comprise flow cell images of the sample at multiple different field of views of one or more sample(s). The different field of views may be at the same x, y, or z location of the same sample. For example, the different field of views may be different subtitles of the sample at the same z location, but different x,y locations. For example, the multiple different views may be adjacent to each other, with none or at least some spatial overlap with other field of views. In some embodiments, each of the training flow cell images comprise flow cell images of differentfield of views (e.g., adjacent FOVs of the same sample) from a plurality of sequencing cycles stacked along one or two spatial dimensions.
[0265] In some embodiments, the training dataset for the neural network, e.g., first or second neural network, may only include flow cell images of the same color channel thereby the neural network is not trained on variations across different color channels that may be caused by differences in optical elements in response to different colored light signals (e.g., emission filter, illumination, etc.), differences in fluorescent dyes, or other factors of the sequencing system, etc. In some embodiments, such variations may cause but is not limited to cause different background levels, different signal to noise ratio, different artifacts in the field of view, different full width at half maximum (FWHM) of emission light signals, point spread function (PSF), etc. Training the neural network using flow cell images of the same color channel may advantageously remove fitting to variations across different color channels, and may simplify and speed up training the neural network and avoid possible errors in training.
[0266] In some embodiments, a different neural network is trained with flow cell images of a corresponding color channel. In other words, the neural network is trained to be a channel-specific neural network. In embodiments where 3 colors are used for sequencing, 3 different neural networks are trained using corresponding flow cell images of the corresponding color channels. Each channel-specific neural network is used for prediction of high resolution flow cell images of the corresponding color channel.
[0267] In embodiments where two or more channels are of the same color but of different wavelength ranges, a single neural network may be trained using flow cell images of such same colors from the two or more channels. Such neural network may be used to predict or make inferences of high resolution flow cell images from the two or more channels of the same color.
[0268] In some embodiments, when two or more channels are of the same color but of different wavelength ranges, a different neural network may be trained using flow cell images of a single channel. Each different neural network is a channel specific neural network that may be used for prediction or inference only of the corresponding channel.
[0269] In some embodiments, subsequent to operation 520’, the method 500 comprises an operation 530 of (iii) predicting, by the first reconfigurable device or an integrated circuit, a second plurality of flow cell images using the neural network, wherein each of the second plurality of flow cell images is with a second resolution and corresponds to acorresponding image of the first plurality of flow cell images, and wherein the second resolution is at least 2 to 32 times greater than the first resolution in one or more spatial dimensions.
[0270] In some embodiments, the operation 530 of (iii) predicting the second plurality of flow cell images using the neural network comprises predicting high resolution postprocessing images corresponding to the first plurality of flow cell images, and wherein the processing comprises various image processing or intensity processing steps. For example, the processing steps may comprise one or more of: noise reduction, background reduction; background removal; artifact removal; artifact suppression; intensity offset correction; intensity normalization; adjusting signal to noise ratio; adjusting contrast to noise ratio; color correction; phasing and / or dephasing; image registration; and deconvolution. In some embodiments, predicting high resolution post-processing images corresponding to the first plurality of flow cell images may advantageously allow a higher resolution and higher image quality version of the first plurality of flow cell images to be generated, and the higher resolution, higher quality version may be used for generating more accurate and reliable base calls. To achieve more accurate and reliable base calling, the second neural network of method 500, e.g., in operation 520’, may be trained using not reference flow cell images or reference intensities, but reference base calls as ground truths. The reference base calls may be generated using various methods including methods disclosed herein in relation to training neural network for predicting base calls herein. In some embodiments, the training, e.g., using methods 700, or other training methods herein, may optimize at least some of the parameters of the neural network for producing training base calls that are similar enough to the reference base calls (e.g., determined by the value of the loss function satisfying a predetermined criteria). As a result, the trained neural network may be used to predict high resolution post-processing images corresponding to the first plurality of flow cell images, and such high resolution post-processing images may be used to produce accurate and reliable base calls. Comparing with embodiments of method 500 with the operation 520 that predicts high resolution flow cell images, or other existing methods that predicts base calls directly, the embodiments of method 500 with the operation 520’ may improve base calling accuracy, reliability, and reduce computation complexity in the prediction, free up storage space, save power and time when compared with methods that predicts base calling directly.
[0271] In some embodiments, training the second neural network in the operation 520’ for each corresponding color channel in comparison with training of first neural network in the operation 520 with flow cell images from multiple color channels may require less computations, require less power consumption, require less memory or data storage, reduce training time, and avoid possible training failures.
[0272] FIG. 30A shows an exemplary flow cell image of the first plurality of flow cell images. In this case, the exemplary flow cell image is of a 2D sequencing sample, and is acquired from one of the 4 different color channels. The image size is 608 pixels by 608 pixels. FIG. 30B is a high resolution image of the flow cell image in FIG. 30A and it is predicted using method 500 with the second neural network in the operation 520 herein. The neural network is pretrained using a training data set comprising training flow cell images. The neural network is pretrained using the training data set and reference base calls instead of reference intensities. The high resolution image has a size of 1216 by 1216 pixels, which provides 2x resolution of the flow cell image in FIG. 30A in x and y direction. The detectable polony density in high resolution image, e.g., FIG. 30B, is increased by at least 2x, 4x, or more than that in the first plurality of flow cell images, e.g., FIG. 30A. In this case, the neural network predicts the high resolution image with less background noise, blurriness, bright artifacts, etc. The high resolution image, in combination with other high resolution images from the other 3 color channels, can then be used together to determine base calls. In this case, the error rate in determining polonies, and the error rate in making base calls can be lower using the high resolution image in FIG. 30B than using the flow cell image in FIG. 30A. The prediction of high resolution images rather than prediction of base calls directly may advantageously reduce computational complexity, computation burden, power consumption, storage usage for performing base calls while maintaining or improving accuracy and reliability.
[0273] In some embodiments, the method 500 include an operation 540 (iv) determining, by the processor, the first reconfigurable logic device, or the integrated circuit, polonies from the second plurality of flow cell images. In some embodiments, determining the polonies comprises determining locations of the polonies, locations of the center of the polonies, size of the polonies, or a combination thereof. In some embodiments, the location of the polonies, or the locations of the center of the polonies may be 2D or 3D. In some embodiments, the polonies excludes duplicate polonies. In some embodiments, the method 500 comprises an operation 550 of (v) performing, by the processor, the firstreconfigurable logic device, or the integrated circuit, a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images.
[0274] In some embodiments, the second neural network in the operation 520’, when it is being trained, may include one or more layers, e.g., convolutional layers, for generating base calls from the high resolution post image processing images. In such embodiments, the second neural network, after it is trained, in operation 520’ may utilize only a subset of the layers of the second neural network being trained since the neural network only predicts the high resolution post image processing images but not the base calls. In such embodiments, the neural network, after it is trained, in operation 520’ may utilize only a subset of the layers of the neural network in operation 520’ since the neural network only predicts the high resolution post image processing images but not the base calls. In some embodiments, the pretrained second neural network using reference base calls may have a first number of layers, while a second number of the layers in the pretrained second neural network is used in operation 530 in predicting the high resolution flow cell images. The second number of layers is less than the first number of layers. For example, the pretrained second neural network may have 5 layers with the first 4 layers for predicting high resolution post image processing images and the last layer for predicting base calls. In the operation of 530, only the first 4 layers of the pretrained second neural network is used. The operation 550 may use the last layer of the pretrained second neural network.
[0275] In some embodiments, the second neural network in operation 520’ utilizes the same number of layers as the second neural network being trained. In such embodiments, the neural network in operation 520’ may lack any neural network layers that is specific for generating base calls based on the high resolution post image-processing images. In such embodiments, the neural network in operation 520’ may rely on other non-neural network based algorithms or software for base calling of the high resolution post imageprocessing images. For example, the pretrained second neural network may have 5 layers with the first 4 layers for predicting high resolution post image processing images and the last layer for predicting base calls. In the operation of 530, only the first 4 layers of the pretrained second neural network is used. The operation 550 may use non-neural network based algorithms or software for base calling.
[0276] In some embodiments, the neural network in operation 520’ has fewer number of convolutional layers than the number of convolutional layers in the neural network inoperation 520. In some embodiments, the second neural network in operation 520’ has the same number of layers as the first neural network in operation 520.
[0277] In some embodiments, the neural network herein has less than or equal to 18, 15, 12, 10, 8, 7, 6, 5, 4, 3, or 2 layers. In some embodiments, the neural network herein has 6, 5, 4, 3, or 2 layers. In some embodiments, the neural network has less than 256, 128, 96, 80, 64, or 32 features.
[0278] In some embodiments, the method 500 is performed during a cycle N, cycle N may be one of the reference cycle(s) for generating the polony map. In some embodiments, cycle N may be a cycle different from the reference cycle(s). The polony map can be generated in the reference cycle(s) as a subsequent operation after the methods herein have improved the detectable polony density in flow cell images. Polonies from one or more channels within the reference cycle(s) can be included in the polony in a reference coordinate system, while base calling of cycle N is yet to be performed. In some embodiments, cycle N is the current cycle. N can be any non-zero integer. For example, for short read sequencing, N can be any integer from 1 to 150, from 1 to 200, or from 1 to 1000.
[0279] In some embodiments, the polony map disclosed herein can include individual regions within a subtile or tile. Each polony map can include a plurality of polonies therein. In some embodiments, the polony map can be of about the same size of a flow cell image so that all the polonies, from different tiles, and from multiple channels, can be registered to the same polony map. However, such polony map may contain polonies that will not be used in at least some operations described herein to reduce computational burden without sacrificing accuracy. In some embodiments, more than one polony map can be generated, and each corresponds to at least part of a subtile of a flow cell image from a channel. The more than one polony map may be tiled together in order to cover the entire sample region of the flow cell device.
[0280] In some embodiments, the polony map disclosed herein can include polonies that are within individual cells or tissue, or on the membrane thereof. In some embodiments, the polony map disclosed herein can exclude polonies or signal spots that are outside cell boundaries. In some embodiments, the polony map disclosed herein can exclude duplicate polonies, such duplication may occur at different z-locations, with one or more in-focus and / or out-of-focus in the flow cell images. The duplicate polonies may be within the same flow cell image or in different flow cell images.
[0281] The polony map herein can be initialized as a virtual image that has a black or dark background with no signals from polonies. For example, the polony map can be initialized to be zero or include otherwise minimal image intensity at all pixels.
[0282] After the coordinates of a polony is determined by image registration of flow cell images, e.g., across different channels, the intensity of the polony can be added to the polony map at the location determined by the coordinates and with the size and shape determined based on registration. The polony map can be a virtual image that combines image intensity from polonies obtained from 2, 3, 4, or even more channels at the reference cycle. The pixels of the template containing no polonies in them remains to be black or dark so that the polony map can have a cleaner background without noise that appear in actual flow cell images. In some embodiments, the polony map includes a list of entries, and each entry corresponding to information for identifying a corresponding polony. For example, each entry can include spatial coordinates of the corresponding polony center in the reference coordinate system, and image intensity of the polony. The entry may also include a unique identification number of the polony.
[0283] The polonies can be from a subtile of flow cell images within a reference cycle, and more specifically, from one or more selected regions of the subtile. The flow cell images can be from different channels of 1, 2, 3, 4, or more channels of the system 100. As a nonlimiting example, a reference cycle can be any cycle of the first 5 or 6 cycles. In some embodiments, the reference cycle can be any cycle that is greater than 0. In some embodiments, the reference cycle is the first cycle.
[0284] In some embodiments, the operation 540 comprises performing image processing step(s) to adjust image intensities of polonies. In some embodiments, the image processing steps comprise one or more of the following: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and prephasing correction; image registration; quality score estimation, or the like. In some embodiments, the image registration is configured to align images from different cycles and / or different channels, for example, with respect to a template image (i.e., a polony map) or a reference coordinate system. In some embodiments, the image registration herein is configured to register polonies or clusters from different cycles and different channels, to a template image or a reference coordinate system.
[0285] In some embodiments, the second plurality of flow cell images may be the output of the neural network. The second resolution may be 2 to 32 times greater than the firstresolution in one or more spatial dimensions. The second resolution may be 4 to 32 times greater than the first resolution in 2D or 3D.
[0286] In some embodiments, the operation 540 is based on a polony map that has been generated. The polony map may be 2D or 3D. In some embodiments, the polony map has the second resolution. In some embodiments, the operation 540 comprises generating a polony map, and determining the polonies based on the generated polony map. The details of generating a 2D or 3D polony map has been disclosed in U.S. Patent Application Nos. 18 / 078,820 and 18 / 078,797, and are incorporated herein by reference in their entirety.
[0287] For example, the base calling can be performed using polony locations in the second plurality of flow cell images from different channels in cycle N, after the second plurality of flow cell images from different channels are registered relative to the polony map disclosed herein. Various existing 2D base calling algorithms can be used. The base calling results can be saved with its 3D coordinates. Such 3D coordinates can be used to register the base calling across different cycles and at different z levels.
[0288] The method 500 can comprise an operation 550 of (v) performing, by the processor, a corresponding base calling for each of the determined polonies. The operation 550 of performing base calling may be based on the second plurality of images generated in operation 530. The operation 540 may be further based on the determined polony map in operation 540. The base calling can be performed using intensity of the polonies from different channels per cycle per z level.
[0289] In some embodiments, the method 500 may include an operation of saving the base calls obtained in operation 550 in a predetermined format, e.g., in a FastQ file compatible with subsequent operations so that subsequent analysis such as adaptor trimming and secondary analysis can be performed.
[0290] In some embodiments, the neural network is a convolutional neural network (CNN). In some embodiments, the neural network is a U-Net. In some embodiments, the neural network comprises a U-Net with a first predetermined repetition of down-sampling and convolution operations and then a second predetermined repetition of up-sampling, concatenation, and convolution operations. The first and second predetermined repetition can have an identical quantity, e.g., 3 or 4. In some embodiments, the neural network is a U-Net with a first predetermined number of filters in each repetition of down sampling, and then a second predetermined number of filters in each repetition of up samplingand / or concatenation. For example, the first predetermined number of filters can be 32, 64, 128, and 256 filters in three repetitions and the second predetermined number can be 128, 64, 64, and 32 filters in the corresponding three repetitions. As another example, the first predetermined number of filters can be 32, 64, 128, and 256 filters in three repetitions and the second predetermined number can be 256, 128, 64, and 32 filters in the corresponding three repetitions.
[0291] In some embodiments, the operation 530 may comprise: performing, by the processor, a first convolution in one or more dimensions on the first plurality of flow cell images, thereby generating a first convolution result; repetitively performing, for one or more times, down-sampling operations comprising: (a) performing, by the processor, a second convolution in one or more dimensions on the first convolution result, thereby generating a second convolution result; and (b) performing, by the processor, a down sampling of the second convolution result by a down sampling factor thereby generating a first down-sampled result. In each repetition, the second convolution may comprises a corresponding number of filters, thereby generating a third convolution result after the repetitions.
[0292] In some embodiments, the operation 530 may further comprise: performing, by the processor, the second convolution in one or more dimensions on the third convolution result, thereby generating a fourth convolution result; repetitively performing, for one or more times, up sampling operations comprising: (c) performing, by the processor, an up sampling of the fourth convolution result by an up sampling factor thereby generating a first up-sampled result; and (d) performing, by the processor, the second convolution in one or more dimensions of the first up-sampled result, thereby generating a fifth convolution result. In each repetition, the second convolution may comprise a corresponding number of filters, thereby generating a sixth convolution result after the repetitions.
[0293] In some embodiments, the first convolution comprises a 3D convolution with a convolution kernel. In some embodiments, the convolutional kernel may have 4 dimensions. In some embodiments, the convolutional kernel is m*m*m for the first three spatial dimensions and the size of its fourth dimension is determined by the filter number in the corresponding repetition. In some embodiments, m can be an integer in the range of 2 to 20. For example, the input can be 512x512 flow cell images, and the z-stack can have 12 slices. The first convolution can include 32 filters and each filter has one kernel that is3x3x3xl. The output from that convolutional block is 512x512x12x32. Then there is a double convolutional block, i.e., the second convolution having two first convolutions with 32 filters. The input to both of those blocks is 512x512x12x32 and the output is 512x512x12x32. Each filter uses a kernel sized 3x3x3x3x32. The number of filters may correspond to features of the input.
[0294] In some embodiments, the second convolution comprises two 3D convolutional layers, e.g., as shown in the pseudo code. In other words, the second convolution comprises two repetition or blocks of the first convolution in 3D, and usage of the output and the number of filters changes, as convolution process will increase the depth of the image. The depth of image may increase as the number of features or filters increases. In some embodiments, the first and second resolution is in 2D or 3D.
[0295] In some embodiments, the first convolution comprises a 2D convolution with a convolution kernel. In some embodiments, the convolutional kernel may have 3 dimensions. In some embodiments, the convolutional kernel is m x m for the first two spatial dimensions and the size of its third dimension is determined by the filter number in the corresponding repetition. In some embodiments, m can be an integer in the range of 2 to 20. For example, the input can be flow cell images with a size of 512x512x1. The first convolution can include 64 filters and each filter has one kernel that is 3x3x1. The output from that convolutional block is 512x512x64. Then there is a double convolutional block, i.e., the second convolution having two first convolutions with 32 filters. The input to both of those blocks is 512x512x64 and the output is 512x512x32. Each filter can use a kernel sized 3x3x32.
[0296] In some embodiments, the second convolution comprises two convolutional layers, e.g., as shown in the pseudo codes. In other words, the second convolution comprises two repetition or blocks of the first convolution, and usage of the output and the number of filters changes, as convolution process will increase the depth of the image. The depth of image may increase as the number of features or filters increases. In some embodiments, the first and second resolution is in 2D or 3D.
[0297] In some embodiments, the second convolution in operation (a) comprises a corresponding number of n, 2*n, 4*n, and 8*n filters in a first, second, third, and fourth repetition, respectively. In some embodiments, the second convolution in operation (c) comprises a corresponding number of 2*n, 2*n, 4*n, 8*n filters in a last repetition, last minus one, last minus two, and last minus three repetition, respectively. In someembodiments, n can be an integer in the range from 8 to 256. For example, operation (a) comprises 32, 64, 128, and 256 filters in three repetitions and operation (c) comprises 128, 64, 64, and 32 filters in the corresponding three repetitions.
[0298] In some embodiments, the second convolution in operation (c) comprises a corresponding number of n, 2*n, 4*n, 8*n filters in a last repetition, last minus one, last minus two, and last minus three repetition, respectively. For example, operation (a) comprises 32, 64, 128, and 256 filters in four repetitions and operation (c) comprises 256, 128, 64, and 32 filters in the corresponding four repetitions.
[0299] In some embodiments, the second convolution in operation (c) comprises a corresponding number of n, 2*n, 4*n filters in a last repetition, last minus one, last minus two, repetition, respectively. For example, operation (a) comprises 32, 64, 128 filters in three repetitions and operation (c) comprises 128, 64, and 32 filters in the corresponding three repetitions.
[0300] In some embodiments, the operation 530 may further comprise: performing, by the processor, the first convolution in one or more dimensions on the sixth convolution result, thereby generating a seventh convolution result; and predicting, by the processing, the second plurality of flow cell images based on the seventh convolution result. Each of the second plurality of flow cell images may correspond to the corresponding flow cell image of the first plurality of flow cell images with a second resolution that is 2, 4, 6, 8, 10, 12, or 16 times greater than the first resolution in one or more spatial dimensions. In some embodiments, the second resolution is at least 4, 6, or 8 times greater than the first resolution in all three dimensions.
[0301] In some embodiments, the first plurality of flow cell images are from a single color channel. In some embodiments, the first plurality of flow cell images are from one or more color channels. In some embodiments, the first plurality of flow cell images are of unbalanced nucleotide diversity in one or more sequencing cycles. In some embodiments, the cellular sample comprises overloaded concatemer molecules with a spatial density in a range of 102-1015per mm2. In some embodiments, the cellular sample comprises overloaded concatemer molecules with a spatial density in a range of 103-10102 per mm .
[0302] In some embodiments, the first resolution is in a range of 0.1 um to 5 um. In some embodiments, the first resolution is in a range of 0.01 um to 10 um. In some embodiments, the second resolution is in a range of 0.02 um to 2 um. In someembodiments, the second resolution is in a range of 0.001 um to 3 um. In some embodiments, the down-sampling factor is 2, 4, 6, 8, 16, or more. In some embodiments, the up-sampling factor is 2, 4, 6, 8, 16, or more.
[0303] In some embodiments, one or more of operations (ii) to (v) are performed while a sequencing run is being performed. In some embodiments, one or more operations (ii) to (v) are performed in parallel as the corresponding sequencing run to reduce sequencing analysis time.
[0304] In some embodiments, the one or more cycles comprises a current cycle N. N may be in a range from 1 to 150, 1 to 300, 1 to 500, or 1 to 1000. In some embodiments, one or more of operations (ii) to (v) are performed while the sequencing reactions in cycles subsequent to the current cycle N is yet to be performed or currently being performed.
[0305] In some embodiments, the training data set of training flow cell images comprises z-stacks of training flow cell images taken at different z-locations. Each z-stack may represent an individual FOV of cellular sample(s). In some embodiments, the z-axis is orthogonal to image planes of the flow cell images.
[0306] In some embodiments, the training data set of training flow cell images comprises flow cell images from multiple sequencing cycles. One or more sequencing cycles may be of unbalanced diversity so that image appear dimmer or the number of polonies are less than images from sequencing cycles of high nucleotide diversity. In other words, the number of polonies in the training flow cell images in a particular cycle may vary from 1% to 99% of a total number of polonies within a FOV of that cycle. When the number of polonies in the training flow cell image of a particular cycle is from 1% to 5% or 1% to 10% of the total number of polonies within that cycle, it is of low or unbalance diversity. When the number of polonies in the training flow cell image of a particular cycle is greater than 10% or 15% of the total number of polonies within that cycle, it is of high or unbalanced diversity.
[0307] In some embodiments, the training data set of training flow cell images comprises flow cell images from multiple samples and multiple sequencing cycles, and the training flow cell images include a subset of flow cell images with unbalanced diversity in multiple sequencing cycles and another subset of flow cell images with balanced diversity in multiple sequencing cycles.
[0308] In some embodiments, the training flow cell images from one or more cycles may be transformed from other training flow cell images from different cycle(s) to simulate the transformation that may occur across cycles within a same color channel.
[0309] In some embodiments, the operation of performing, by the processor, the first convolution in one or more dimensions on the first plurality of flow cell images comprises: performing, by the processor, a first convolution in 3D on the first plurality of flow cell images, thereby generating a first convolution result. In some embodiments, operation (a) comprises performing, by the processor, the second convolution in 3D on the first convolution result, thereby generating a second convolution result.
[0310] In some embodiments, the operation of performing, by the processor, the first convolution in one or more dimensions on the first plurality of flow cell images comprises: performing, by the processor, a first convolution in 2D on the first plurality of flow cell images, thereby generating a first convolution result. In some embodiments, operation (a) comprises performing, by the processor, the second convolution in 2D on the first convolution result, thereby generating a second convolution result.
[0311] In some embodiments, repetitively performing, for one or more times, operations comprising (c) and (d) comprise: repetitively performing, for one or more times, operations comprising (c), (d), and (e), wherein (e) is after operation (c) and before operation (e), and wherein (e) comprises: concatenating, by the processor, the first up- sampled result in a current up-sampling repetition with the first down-sampled result in a previous down-sample repetition, wherein the first up-sampled result has a same size as the first down-sampled result in the previous down-sampling repetition. In some embodiments, operation (e) is in each repetition. In other words, repetitively performing, for one or more times, operations comprising (c) and (d) comprise: repetitively performing operations comprising (c), (d), and (e) in each repetition of one or more repetitions.
[0312] The kernel may take any size that is smaller than the size of the flow cell image undergoing the convolution. For example, with an opening operation, the kernel can be 2 by 2 by 2, 3 by 3 by 3, 4 by 4 by 4, 5 by 5 by 5, or 6 by 6 by 6 in the first three spatial dimensions. In some embodiments, the kernel size can be customized to remove at least some of the noise and unwanted signal that are larger than the kernel size. In some embodiments, the kernel can be circular. The kernel can be in various other shapes.
[0313] In some embodiments, when the focus of the optical system includes a range, e.g., 0.1 um, 0.2 um, 0.3 um, 0.5 um, 0.6 um, 0.8 um, 1 um, 2 um, 3, um, 4 um, 5 um, etc. expanding along z axis. Polonies or clusters that are within the range of focus can appear in-focus or about in-focus in the flow cell image. Flow cell images at a specific z level can also include signals from polonies or clusters that are not within the focus range of the image, but at different z levels. Such polonies or clusters are out-of-focus. As shown in FIG. 3 A, bigger and blurred signal spots represent out-of-focus polonies or clusters. Some of the out-of-focus polonies or clusters are circled in FIG. 3 A.
[0314] Each flow cell image at a specific z level can also include noises caused by the optical system and / or undesired signal from the sample. The undesired signal can be signal coming from components of the sample such as membrane, cytosol, and mitochondria. Such background objects can be any objects, relatively larger in size than the polonies or clusters. As shown in FIG. 3 A, there is a blurry cellular contour (at the arrows) in the flow cell image, and most of the signal spots are contained within the blurry contour. In some embodiments, background objects can include any objects within the 3D sample but are not polonies or clusters.
[0315] In some embodiments, the method 500 include an operation of registering the second plurality of flow cell images. In some embodiments, the images are registered across channels and / or across different cycles. In some embodiments, the images are registered before any base calling are performed in operation 550. In some embodiments, the images are registered across channels and different cycles before generating or obtaining the polony maps. In some embodiments, the images are registered across channels and different cycles before one or more primary analysis steps here. In some embodiments, the images can be registered after one or more preprocessing operations disclosed herein are performed. Various image registration techniques can be used to register the images. Various image registration techniques can be used to register the images. The images can be registered using 2D or 3D registration techniques.
[0316] In some embodiments, the operation of registering the flow cell images is with respect to a reference coordinate system. In some embodiments, the operation of registering the flow cell images is with respect to one or more template images. The operation of registering the images can comprise generating the one or more template images in a reference coordinate system. In some embodiments, the operation of registering the images can comprise registering polonies to template polonies in the oneor more template images. The operation of registering the images can comprise determining a plurality of transformations based on the one or more template images. Each of the plurality of transformations can corresponds to a corresponding subtile of the flow cell images, the processed images, or the filtered images and configured to register the subtile to the one or more template images. Each transformation can be used to register a corresponding subtile or tile to the one or more template images. The plurality of transformations can comprise one or more affine transformations.
[0317] In some embodiments, the operation of registering the images can comprise performing image registration of the polonies based on fiducial markers. The fiducial markers can be located on the flow cell. Alternatively, the fiducial markers can be external to the flow cell.
[0318] In some embodiments, the image registration as an image processing step herein is configured to align images from different cycles and / or different channels, for example, with respect to a template image or a reference coordinate system. In some embodiments, the image registration herein is configured to register polonies or clusters from different cycles and / or different channels, e.g., in the filtered image, to a template image or a reference coordinate system.
[0319] For example, the base calling can be performed using the filtered images from different channels in cycle N after the filtered images from different channels are registered relative to the corresponding template image disclosed herein.
[0320] The operation 540 can comprise an operation of extracting polony intensities based on the polony map. For each polony in the polony map, the location information of such polony can be obtained from the polony map, e.g., 2D coordinates of the polony and the z level. Using the 2D coordinates and the z level, the corresponding flow cell image and its pixel(s) can be determined. Image intensity of such pixels can be extracted from the corresponding processed image after one or more image processing steps as intensity of such pixel for performing base calling.
[0321] In some embodiments, the operation of registering the flow cell images may be based on background objects in the flow cell images. The background objects can be used to align the flow cell image to the cell images by using one or more transformation(s).The cell staining images herein are staining images of the sample(s) immobilized on the support, with possible transformation (e.g., translation) from the sample(s) in the flow cell images. The transformation may be represented by a single transformation of the wholeimage or be separated into multiple transformations, each representing a portion of the whole image. After finding the transformation(s) of the background objects between the flow cell images and the cell staining images, the polonies or clusters can be registered to the cell staining images.
[0322] In some embodiments, the method 500 may include an operation of registering the base calling in 550 to the cell staining images. In some embodiments, such registration may be based on fiducial markers. Such fiducial markers can also be included in the cell staining images. Aligning the fiducial markers can generate the transformation(s) between the flow cell images or between flow cell images and cell staining images. The transformation(s) can be used to register or align polonies or clusters between the sequencing images and the cell images.
[0323] As an example, the simulated z-stack is 2048x2048x3, each cell may include 200 to 2000 polonies per cell. The spatial resolution can be about 0.1 um. Prediction is performed independently for each 512x512 region of the simulated z-stack. The predicted high-resolution z-stack is 8192x8192x12. FIGS. 2A-2C show simulated flow cell images, and two different predicted flow cell images with 4x resolution at different z-locations. FIGS. 3 A and 3D show two actual flow cell images at different z-locations in a 512x512x3 z-stack. The predicted high resolution flow cell images (2048x2048) in FIGS. 3B-3C are at two different z-locations corresponding to the low resolution image in FIG. 3A.Exemplary pseudo code for predicting high resolution flow cell images
[0324] An exemplary neural network is shown below in the pseudo code. The neural network may be used to predict polony locations using z-stack(s) of flow cell images comprising flow cell images from multiple z-levels forming 3d volume(s). In some embodiments, m is in a range from 2 to 10, filters can be in a range from 8 to 1024, k size and be in 4 dimensions, and the fourth dimension of k size can match the number of filters in the corresponding repetition. The input flow cell images can have various sizes in 3D as disclosed herein, e.g., 1024 by 1024 by 4. def conv block(x, filters, k size) x=conv3D (filters, (k size, k size, k size)) (x) x = tf.keras. filters. ReLu()(x) def double conv ():x=conv block(x, filters, k size) x=conv block(x, filter s,k size) inputs= input (dim x, dim y, dim z, 1) bi = conv block (inputs, filters, k size) b2 = double conv (bi, filters* 2, k size) % repeating a first predetermined number of down sampling and convolutions for n= 2:m dn= downsampling3D(bn) bn+i = double conv(dn, filter s*2n+1, k size)% repeating a second predetermined number of down sampling, concatenation, and convolutions for n = l:m-l un= upsampling3D(bm+n) catsn= concatenate (un, bm+i-n) bm+n+1=double conv(catsn, filter s* 2m~n~1, k size) b2m+i = conv block b 2m, filters, k size)Exemplary pseudo code for predicting high resolution flow cell images
[0325] An exemplary neural network is shown below in the pseudo code. The neural network may be used for predicting polony locations based 2D flow cell images at different z-levels. In some embodiments, m is in a range from 2 to 10, filters can be in a range from 8 to 1024, k size can be in 3 dimensions, and the third dimension of k size can match the number of filters in the corresponding repetition. The input flow cell images can have various sizes in 2D as disclosed herein, e.g., 1024 by 1024, and there can be 3, 4, 5, or other numbers of z-levels. def conv block(x, filters, k size) x=conv2D (filters, (k size, k size)) (x) x = tf.keras. filters. ReLu()(x) def double conv (): x=conv block(x, filters, k size) x=conv block(x, filter s,k size) k size = n filter num = input [0]def u_net(): inputs = Input() b2 = double conv(inputs, filter num, k size)% repeating a first predetermined number of down sampling and convolutions for n= 2:m dn-i = MaxPooling2D(bn) bn+i = double conv(dn-i, filter s*2n l, k size)% repeating a second predetermined number of down sampling, concatenation, and convolutions for n = 2:m-l un= upsampling2D(b m+n-2) catsn= concatenate (un+i, bm-n+i) bm+n-l=double conv(catsn, filter s* 2m~n~1, k size) model = tfkeras.Model(inputs, outputs)Predicting base calls using neural networks
[0326] In some embodiments, the methods and systems herein can be used to predict base calls for some or all polonies of the flow cell images. The systems and methods herein advantageously use a neural network that is pretrained for predicting the base calls for polonies of flow cell images. The same neural network may also be advantageously used, without additional training, to generate a polony map or a template image so that the locations of the predicted base calls can be determined. The embodiments herein used convolutional neural network as an example, however, it is understood that various other neural networks or machine learning models may also be used achieve prediction of base calls using the systems and methods herein.
[0327] In some embodiments, the methods for predicting base calls may include one or more operations here. When there are multiple operations involved, such operations may or may not be performed in the order that is described herein.
[0328] FIG. 28 shows a flow chart of a computer-implemented method 2800 for predicting base calls for flow cell images of biological samples, e.g., cellular samples, thereby enabling efficient and accurate primary analysis. The method 2800 can include some or all of the operations disclosed herein. The operations may be performed in but is not limited to the order that is described herein.
[0329] The method 2800 can be performed by one or more processors disclosed herein. In some embodiments, the processor can include one or more of: a processing unit, e.g., a CPU, a reconfigurable logic device, an integrated circuit that is not reconfigurable, or their combinations. For example, the processing unit can include a central processing unit (CPU). The reconfigurable logic device can include one or more FPGA devices. The integrated circuit can include a chip such as an Al chip or an ASIC chip. In some embodiments, the processor can include the computing system 400.
[0330] In some embodiments, some or all operations in method 2800 can be performed by the reconfigurable logic device, e.g., the FPGA(s), and / or the integrated circuit, e.g., the Al chip(s). In embodiments when some operations are performed by the reconfigurable logic device and / or integrated circuit, e.g., FPGA(s), the data produced by the reconfigurable logic device and / or integrated circuit, e.g., the FPGA(s), after performing one or more operations can be communicated to various hardware elements of the system 100, e.g., CPU(s) or GPU(s), so that subsequent operation(s) in method 500, 600, 700, 2800, and 2900 can be performed by such various hardware using the communicated data. Similarly, data can also be communicated in the opposite direction from various hardware e.g., CPU(s), to the reconfigurable logic device or the integrated circuit for processing. In some embodiments, all the operations in the methods herein can be performed by CPU(s). Alternatively, the operations performed by CPU(s) can be performed by other processors such as the dedicated processors, or GPU(s). In some embodiments, all the operations in the methods herein can be performed by the reconfigurable logic device and / or the integrated circuit, e.g., FPGA(s) and / or the Al chip(s).
[0331] In some embodiments, the sensor data acquired by the imager 116 may be directly communicated to the reconfigurable logic device and / or the integrated circuit, e.g., via DMA connections. In some embodiments, the sensor data acquired by the imager 116 may be directly communicated to the reconfigurable logic device and / or the integrated circuit without being routed first to a CPU, a GPU, or any other processing units before reaching the reconfigurable logic device and / or the integrated circuit.
[0332] In some embodiments, making predictions or inferences using the methods 2800 herein with the reconfigurable logic device, e.g., the FPGA, and / or other integrated circuit, e.g., Al chips, may require at least 2x, 8x, lOx, 15x, 20x, 40x, 50x, or lOOx less power than making prediction(s) or interference(s) with the same neural network(s) withidentical training images using other computing hardware including but not limited to CPUs or GPUs.
[0333] In some embodiments, the sequencing system herein further comprises: a power source that is configured to supply identical or different power levels to the reconfigurable logic device and the integrated circuit. In some embodiments, a maximum power output of the power source to the sequencing system in performing methods 500, 600, 700, 2800, and / or 2900 is less than 2000 Watts, 1000 Watts, 900 Watts, 800 Watts, 700 Watts, 650 Watts, 600 Watts, 550 Watts, 500 Watts, 400 Watts, 300 Watts, 200 Watts, or 100 Watts.
[0334] The method 2800 can comprise an operation 2810 of (i) generating, by the sequencing system 110, a first plurality of flow cell images of sample(s) immobilized on a support by conducting one or more cycles of sequencing reactions.
[0335] The sample(s) may be traditional 2D sequencing samples containing biological analytes. The sample(s) may be cellular or tissue samples. The samples may comprise concatemer molecules therewithin. The sample(s) may include concatemer molecules from one or more different sample sources. The sample(s) may include a thickness along the z-axis so that the first plurality of flow cell images may be acquired at a z-stack of different z-locations with a first resolution to cover the cellular sample in 3D.
[0336] The sample can be in situ. The sample can be a 3D sample. The sample can be a volumetric sample that may contain different biological information at the same x-y location but different z level. The sample can include multiple cells, tissue, or their combinations. The 3D sample can be any biological sample that has a thickness that is greater than a predetermined threshold along the z axis. For example, the thickness can be greater than 1 um, 2 um, 3 um, 4 um, 5 um, 10 um, 20 um, or more. The z axis (e.g., z axis) is orthogonal to the image plane defined by x and y axes. In some embodiments, the sample can be traditional 2D sequencing samples.
[0337] The flow cell images can be acquired using the optical system of the imager 116 disclosed herein, from the 1, 2, 3, 4, or more channels. Each flow cell image can include at least a portion of one or more tiles (e.g., imaging areas), and each tile can be divided into multiple subtiles. Each tile or subtile can include a plurality of polonies or clusters. Each subtile can include multiple regions with each region including a number of polonies. The flow cell image as disclosed herein can be an image that is acquired from a flow cell 112 as shown in FIG. 1 or 2712 as shown in FIG. 27. In some embodiments, theflow cell images are acquired from a single color channel, and subsequent prediction is by using a pretrained neural network corresponding to that single channel. In some embodiments, the flow cell images are acquired from 2, 3, 4, or more color channels, and subsequent prediction is by using a pretrained neural network corresponding to the multiple color channels.
[0338] In some embodiments, a flow cell image herein can be an image of one or more tiles, one or more subtiles, one or more segmented regions within tile(s) or subtile(s), or their combinations. Each flow cell image can comprise a field of view (FOV). The FOV can be orthogonal to the z axis. The FOV can be within the x-y plane. The FOV of different flow cell images at different z levels can be identical within the x-y plane. The FOV of different flow cell images at different z levels can have at least an overlapping portion within the x-y plane. The image resolution of different flow cell images at different z levels can be about identical or exactly identical. In some embodiments, The image resolution of different flow cell images at different z levels is different. FIGS. 3A and 3D show two exemplary flow cell images acquired at two different z levels along the z axis of a same 3D sample within a same sequencing cycle. The FOV can be in 3D and be of various sizes to cover the volumetric sample to be imaged. The FOV along x, y, and / or z direction can be in a range from 10 um to 5 mm. The FOV along x, y, and / or z direction can be in a range from about 0.1 um to about 2 mm. The FOV along x, y, and / or z direction can be in a range from 0.5 um to 1 mm. For example, the FOV can be about 0.5 mm by 0.5 mm by 20 um for certain cellular samples along the x, y, and z direction, respectively.
[0339] The flow cell images herein may be of various sizes, the pixel number along x, y, and / or z axis may be any integer greater than 64 or 128. The flow cell images herein may be of various sizes, the pixel number along x, y, and / or z axis may be in a range from 2 to 65536. A single flow cell image can be separated into different number of regions, for example, 4, 8, 16, or even more regions, and each region may include a size of 256 by 256 by 1, 512 by 512 by 3, or other sizes. In some embodiments, the number of pixels along x, y, and / or z direction may be adjusted to maintain a particular spatial resolution in a given FOV. For example, with a spatial resolution of 0.2 um, to cover a FOV of 0.8 mm, the number of pixels may be 4000.
[0340] Each flow cell image at a specific z level may include intensities generated by polonies or clusters at the corresponding z level. As shown in FIGS. 3 A and 3D, signalsfrom polonies or clusters are small bright spots within the images. Each bright spot can be of various sizes that is less than a couple of pixels, e.g., less than a pixel, about a pixel, about 2 pixels, 3 pixels, 4, pixels, 5 pixels, or more. In some embodiments, each signal spot of the polonies or clusters can be any number of pixels in the range from 0.01 pixel to about 100 pixels. In some embodiments, each signal spot of the polonies or clusters can be any number of pixels in the range from 0.1 pixel to about 16 pixels.
[0341] Each flow cell image can also include intensities generated by the cell and its structural elements. Such structural elements can be background objects or components, e.g., in FIG. 3 A. Each flow cell images can also include noise and / or artifacts that are not from the polonies or cellular structures.
[0342] In some embodiments, when the depth of field the optical system includes a range, e.g., 0.1 um, 0.2 um, 0.3 um, 0.5 um, 0.6 um, 0.8 um, 1 um, 2 um, 3, um, 4 um, 5 um, etc. expanding along z axis. Polonies or clusters that are within the range of depth of field can appear in-focus or about in-focus in the flow cell image. Flow cell images at a specific z level can also include signals from polonies or clusters that are not within the focus range of the image. Such polonies or clusters are out-of-focus. As shown in FIG. 3 A, bigger and blurry signal spots represent out-of-focus polonies or clusters. Some of the out-of-focus polonies or clusters are circled in FIG. 3 A.
[0343] Each flow cell image at a specific z level can also include noises caused by the optical system and / or undesired signal from the sample. The undesired signal can be signal coming from components of the sample such as membrane, cytosol, and mitochondria. Such background objects can be any objects, relatively larger in size than the polonies or clusters. As shown in FIG. 3 A, there is a blurry cellular contour (at the arrows) in the flow cell image, and most of the signal spots are contained within the blurry contour. In some embodiments, background objects can include any objects within the 3D sample but are not polonies or clusters.
[0344] In some embodiments, base calls from the polonies include 4 different bases, and percentage of polonies for each of the 4 different bases can be greater than about 10% so that the data are relatively diverse. In some other embodiments, bases called from the plurality of polonies includes 4 or less different bases, and percentage of polonies for one or more bases can be less than about 10%, and such data can be considered as data of unbalanced diversity. In some embodiments, bases called from the plurality of polonies include 4 or less different bases, and percentage of polonies for some of the bases can beless than about 5%, about 2%, or even about 1%, and such data can be considered as data of unbalanced diversity. As an example, the base called for bases A, T / U, C, G in the plurality of polonies can be about 1%, about 2%, about 1%, and about 95%. As another example, the base called for bases A, T / U, C, G in the plurality of polonies can be about 10%, about 10%, about 10%, and about 70%, respectively. In addition to the base biases affecting diversity, plexity can also be a factor that when plexity is lower than a number, e.g., 8 or 16, the signal could be of unbalanced diversity . The method 2800 is configured to predict base calls of flow cell images, e.g., of a first resolution, even if the polonies in the flow cell images are of unbalanced nucleotide diversity in one or more sequencing cycles, and the base calls may be spatially aligned to the polonies of the flow cell images, of a second resolution. The second resolution may be higher than the first resolution.
[0345] In some embodiments, the method 2800 comprises an operation 2802 of (ia) generating, by a processor or a first reconfigurable logic device, a second plurality of flow cell images comprising a second resolution. In some embodiments, each of the second plurality of flow cell images corresponds to a corresponding flow cell image of the first plurality of flow cell images. The second plurality of flow cell images may be generated using various up-sampling algorithms including but not limited to interpolation. The second resolution may be greater than the first resolution in one or more spatial dimensions. The second resolution may be at least 2 times greater than the first resolution in one or more spatial dimensions. The second resolution may be 2 to 32 times greater than the first resolution in one or more spatial dimensions. The second resolution may be 4 to 64 times greater than the first resolution in one or more spatial dimensions, e.g., along x, y, and / or z direction. The second resolution may be at least 2 to 32 times greater than the first resolution in one or more spatial dimensions. The second resolution may be at least 4 to 64 times greater than the first resolution in one or more spatial dimensions.
[0346] In some embodiments, the method 2800 comprises an operation 2804 of (ii) providing, by a processor, the second plurality of flow cell images as an input to a neural network, e.g., a convolutional neural network (CNN), wherein the neural network is pretrained using a training data set of training flow cell images using a training method disclosed herein, e.g., 600, 700, 2900 herein. In some embodiments, the neural network is pre-trained so that the values of parameters (e.g., weights) of the neural network has been optimized based on the training. The neural network may be retrained when needed, for example, for predicting flow cell images from different cellular samples.
[0347] In some embodiments, the method 2800 may include image processing step(s) that can be performed on the first or second plurality of flow cell images, optionally prior to providing any input to the neural network. The processing step(s) may include: intensity normalization, background subtraction, background removal, artifact reduction, artifact removal, adjustment of signal to noise ratio, adjustment of contrast to noise ratio, color correction, adjusting intensity offset, image registration, phasing and prephasing, filtering, segmentation, noise reduction, deconvolution (e.g., to differentiate neighboring or at least partly overlapping signal spots), or a combination thereof.
[0348] In some embodiments, the method 2800 comprises an operation 2804’ of providing, by the processor, the first reconfigurable logical device, or the integrated circuit, the first or the second plurality of flow cell images to a polony map generation algorithm or a base calling algorithm. In some embodiments, the polony map generation algorithm and the base calling algorithm does not include a trained neural network or an artificial intelligence-based algorithm. In some embodiments, the polony map generation algorithm and base calling algorithm does not include a trained neural network or an artificial intelligence-based algorithm. Exemplary polony map generation algorithms for generating 2D or 3D polony maps and base calling algorithms for generating base calls have been disclosed in U.S. Application No. 18 / 078,797 and 18 / 078,820, and U.S. Patent No. 10,266,888, and are incorporated herein by reference in their entireties.
[0349] In some embodiments, the method 2800 comprises an operation 2806 of (iia) of determining, by the first reconfigurable device or the integrated circuit, the polony map based on the second plurality of flow cell images. The operation 2806 can be based on the operation of 2804 in some embodiments, and based on the operation of 2804’ in some other embodiments.
[0350] In some embodiments, the polony map is 3D. In some embodiment, the 3D polony map includes multiple 2D polony maps at different z levels. In some embodiments, the 3D polony map has the second resolution. In some embodiments, generating a polony map using a polony map generation algorithm. In some embodiments, the polony map generation algorithm lacks any neural network or artificial intelligence based algorithms. In some embodiments, the polony map generation algorithm lacks any neural network that has been pretrained and can predict base calls in operation 2812 without additional training for predicting the polony map. In some embodiments, the polony map generation algorithm utilize traditional algorithms that lacks artificial intelligence.
[0351] In some embodiments, the neural network in operation 2804-2806 is the same pretrained neural network used in operation 2812. In some embodiments, the same pretrained neural networks may include identical parameters, layers, and neural network structures therewithin. In some embodiments, the same pretrained neural networks may include an identical number of parameters, an identical number of layers, and neural network structures therewithin. In some embodiments, the method 2800 may further comprise an operation to train the neural network before operations 2804 and 2806. In some embodiments, the neural network is trained before operation 2804 and 2806, e.g., using method 700 or 2900 disclosed herein. In some embodiments, the pretrained neural network may be used to predict polony locations, polony shape and / or size, polony center locations, or equivalently the polony map. In some embodiments, the operation 2800 may further include one or more operations in method 500, e.g., operation 530 and 540, and / or 550 for predicting locations of the polonies, thus predicting the polony map.
[0352] In some embodiments, the same neural network used in operations 2804, 2806, and 2812 may be trained using identical training data including identical flow cell images of samples. The identical training data may also include identical “ground truths” or references in training. As a result, after training, the same neural networks may comprise identical values for parameters, identical number of layers, and identical neural network structures.
[0353] In some embodiments, the same neural network used in operations 2804, 2806, and 2812 may be trained using at least a different portion of the identical training data. As a result, the same neural networks may comprise identical parameters with identical or different values for such parameters, identical layers, and identical neural network structures. Training of the same neural network may be performed before operation 2804 and does not require retraining the neural network after operation 2806 and before operation 2812. The pretrained neural network may then be used in operations 2804-2806 and operation 2812 without retraining to allow fast and efficient prediction of the base calls using methods 2800.
[0354] In some embodiments, the pretrained neural network may be used in operations 2804-2806 to update an existing polony map. The existing polony map may be generated in an earlier cycle of the sequencing run. The predicted polony map using the pretrained neural network may be used to update the existing polony map in a later cycle of the sequencing run. For example, an initial polony map may be generated by a non-neuralnetwork algorithm in the first cycle or first several cycles, e.g., cycles 1-4, of the sequencing run. The neural network may be trained using data of the first cycles or a number of cycles, e.g., cycles 1-4 or cycles 1-5. The pretrained neural network then can be used to predict a second polony map that can be used to update the initial polony map. The second polony map may advantageously reselect more accurate and reliable locations of the polonies for making predictions of base calls, intensities, or classifications, e.g., in operation 2812. Such prediction may be repeated by training the neural network with different cycles that has been completed in the sequencing run to improve reselection of polony locations. For example, the trained neural network may be retrained using data of cycles 1-6 or 1-7 following the training using data from cycle 1-4, and make another prediction of the polony map after the training.
[0355] In some embodiments, the same neural network used in operation 2806 and 2816 may be trained using different reference information as the “ground truth” in training. In some embodiments, the training of the neural network for predicting polony locations may use reference intensities as the “ground truth,” while the training of the neural network for predicting base call may use reference base calls as the “ground truth.”
[0356] In some embodiments, the same neural network used in operations 2804-2806 and 2816 may be trained using identical reference information as the “ground truth” in training. In some embodiments, the training of the neural network for predicting polony locations may use reference intensities as the “ground truth,” and the training of the neural network for predicting base call may use reference base calls that can be determined based on such reference intensities.
[0357] In some embodiments, the same neural network may be trained to predict base calls using various training methods, e.g., method 2900 disclosed herein. Reference base calls may be used for the training of the neural network. The reference base calls used in training, e.g., using method 2900, may include spatial information thereof. The reference base calls used in training, e.g., using method 2900, may be of a first resolution, a second resolution, or a third resolution. In some embodiments, the third resolution can be higher than the first and second resolution.
[0358] In some embodiments, the same neural network may be trained to predict base calls, e.g., using method 2900. After being trained, such neural network may be used to predict base calls of the second plurality of flow cell images. The prediction of base calls can then be processed for determining locations of the polonies, thereby generating thepolony map. For example, the polony map may be determined as the locations at which the base calls are predicted with a probability satisfying a predetermined threshold. As another example, the polony locations, thus the polony map, may be determined as the locations in which one or more quality metrics satisfy a predetermined threshold. Such quality metrics can include but is not limited to maximum, medium, or average intensity of the polony among different color channels, a Q score of the base call, a clarity of the base call, and a purity of the base call.
[0359] As disclosed above, the second plurality of flow cell images may be used for generating the polony map at the second resolution using operations 2804- 2806 or operations 2804’-2806. Alternatively, in some embodiments, the method may include an operation of generating the polony map based on the first plurality of flow cell images at the first resolution, and an operation of up-sampling to generate the polony map at the second resolution after operation 2810. In such embodiment, the first plurality of flow cell images may be provided instead of the second plurality of flow cell images in operation 2804 or operation 2804’ and then the operation 2806 may be replaced by an operation of determining the polony map based on the first plurality of flow cell images.
[0360] In some embodiments, the method 2800 is performed during a cycle N, cycle N may be one of the reference cycle(s) for generating the polony map. In some embodiments, cycle N may be a cycle different from the reference cycle(s). The polony map can be generated in the reference cycle(s) as a subsequent operation after the methods herein have improved the detectable polony density in flow cell images. Polonies from one or more channels within the reference cycle(s) can be included in the polony in a reference coordinate system, while base calling of cycle N is yet to be performed. In some embodiments, cycle N is the current cycle. N can be any non-zero integer. For example, for short read sequencing, N can be any integer from 1 to 150. In some embodiments, N can be any integer from 1 to 20, 1 to 200, 1 to 300, 1 to 500, or 1 to 1000.
[0361] In some embodiments, the polony map disclosed herein can include individual regions within a subtile or subtile. Each polony map can include a plurality of polonies therein. In some embodiments, the polony map can be of about the same size of a flow cell image so that all the polonies, from different tiles, and from multiple channels, can be registered to the same polony map. However, such polony map may contain polonies that will not be used in at least some operations described herein to reduce computationalburden without sacrificing accuracy. In some embodiments, more than one polony map can be generated, and each corresponds to at least part of a subtile of a flow cell image from a channel. The more than one polony map may be tiled together in order to cover the entire sample region of the flow cell device.
[0362] In some embodiments, the polony map disclosed herein can include polonies that are within individual cells or tissue, or on the membrane thereof. In some embodiments, the polony map disclosed herein can exclude polonies or signal spots that are outside cell boundaries. In some embodiments, the polony map disclosed herein can exclude duplicate polonies, such duplication may occur at different z-locations, with one or more in-focus and / or out-of-focus in the flow cell images. The duplicate polonies may be within the same flow cell image or in different flow cell images.
[0363] The polony map herein can be initialized as a virtual image that has a black or dark background with no signals from polonies. For example, the polony map can be initialized to be zero or include otherwise minimal image intensity at all pixels.
[0364] After the coordinates (e.g., 3D coordinates) of a polony is determined by image registration of flow cell images, e.g., across different channels, the intensity of the polony can be added to the polony map at the location determined by the coordinates and with the size and shape determined based on registration. The polony map can be a virtual image that combines image intensity from polonies obtained from 2, 3, 4, or even more channels at the reference cycle. The pixels of the template containing no polonies in them remains to be black or dark so that the polony map can have a cleaner background without noise that appear in actual flow cell images. In some embodiments, the polony map includes a list of entries, and each entry corresponding to information for identifying a corresponding polony. For example, each entry can include spatial coordinates of the corresponding polony center in the reference coordinate system, and image intensity of the polony. The entry may also include a unique identification number of the polony.
[0365] The polonies can be from a subtile of flow cell images within a reference cycle, and more specifically, from one or more selected regions of the subtile. The flow cell images can be from different channels of 1, 2, 3, 4, or more channels of the system 100. As a nonlimiting example, a reference cycle can be any cycle of the first 5 or 6 cycles. In some embodiments, the reference cycle can be any cycle that is greater than 0. In some embodiments, the reference cycle is the first cycle.
[0366] In some embodiments, the processing steps herein comprises performing image processing step(s) herein to adjust image intensities of polonies. In some embodiments, the image processing steps comprise one or more of the following: background subtraction; image sharpening; intensity offset adjustment; color correction; intensity normalization; phasing and prephasing correction; image registration; quality score estimation, or the like . In some embodiments, the image registration is configured to align images from different cycles and / or different channels, for example, with respect to a template image (i.e., a polony map) or a reference coordinate system. In some embodiments, the image registration herein is configured to register polonies or clusters from different cycles and different channels, to a template image or a reference coordinate system.
[0367] The method 2800 can comprise an operation 2812 of: (iii) predicting, by the first reconfigurable device or the integrated circuit, one or more base calls corresponding to one or more polonies of the second plurality of flow cell images using the neural network; or predicting, by the first reconfigurable device or the integrated circuit, one or more classifications corresponding to one or more pixels of the second plurality of flow cell images using the neural network. The operation 2812 of performing base calling may be based on the second plurality of flow cell images. The operation 2812 may be further based on the determined polony map in operation 2804 or 2804’.
[0368] In some embodiments, the second plurality of flow cell images may be from one or more color channels, one or more z levels, and / or one or more cycles. The prediction of base calls in operation 2812 can be performed using intensity of the polonies. In some embodiments, the second plurality of flow cell images may be from a single color channel, a single z level, and / or a single cycle. In some embodiments, the prediction of base calling can be performed using intensity of the polonies from a single color channel and one or more cycles. For example, flow cell images acquired from each color channel of the multiple color channels in multiple cycles may use a different pre-trained neural network for predicting the polony intensity of the corresponding channel. The prediction in operation 2812 of base calling can then be performed using intensities of the polony from different color channels. As another example, flow cell images from a single z level may require a different pre-trained neural network for predicting the base calls from a different z level using operation 2812. In some embodiments, prediction of base calling in operation 2812 can be performed using intensity of the polonies from different colorchannels, multiple z levels, and multiple cycles. In some embodiments, prediction of base calling in operation 2812 can be performed using intensity of the polonies from different color channels, a single z level, and multiple cycles. In some embodiments, prediction of base calling in operation 2812 can be performed using intensity of the polonies from a single color channel, one or more z levels, and one or more cycles. In some embodiments, prediction of base calling in operation 2812 can be performed using intensity of the polonies from one or more color channels, one or more z levels, and one or more cycles.
[0369] In some embodiments, the operation 2812 (iii) may include generating outputs that includes base calls, e.g., A, T, C, G, and / or U for one or more pixels of the second plurality of flow cell images. The one or more pixels may be determined using a polony map or a location list of polonies disclosed herein so that each pixel of the one or more e pixels is comprised in at least one polony in the polony map.
[0370] In some embodiments, the operation 2812 of (iii) may comprise generating outputs that includes base calls, e.g., A, T, C, G, and / or U for one or more pixels of the second plurality of flow cell images. In some embodiments, the operation 2812 of (iii) may comprise generating outputs that includes classifications, e.g., A, T, C, G, U, and / or background for one or more pixels of the second plurality of flow cell images. In some embodiments, the one or more pixels may include pixels that are not included in the polony map or the location list disclosed herein. For example, the one or more pixels may include all pixels within the FOV of the second plurality of flow cell images. In some embodiments, the one or more pixels include at least one pixel that is not comprised in any polony of the polony map. In some embodiments, the one or more pixels include at least one pixel that is comprised in the background of the polonies comprise noise signal(s). In some embodiments, the one or more pixels include at least one pixel that is not comprised in any polony in the polony map and at least one pixel that is comprised in at least one polony in the polony map. In some embodiments, the one or more pixels include at least one pixel that is not within a cell membrane or on the cell membrane.
[0371] FIGS. 3E -3F show comparison of accuracy of identifying transcripts (corresponding to polonies) using the neural network methods herein (“new algorithm”), e.g., 2800, and a traditional non-neural network based algorithm (“POR-YOLO”). In this case, simulated flow cell images of in situ sample with multiple cells are used. Each area may include a number of targets ranging from 0 to 4000. Such targets can be transcripts. The neural network herein, e.g., in method 2800, and a classic non-neural network basedalgorithm are used to predict / detect transcripts in such cells. And the prediction / determination is then compared with ground truths (or equivalently, the reference polony map) for accuracy. The correct number of targets per area is higher using the neural network and method disclosed herein than using the non-neural network based algorithm. The detected targets per area using the neural network and methods herein are much higher (2x or 3x higher) than that detected by the non-neural network based algorithm when the target density per area is greater than 2000 per area. FIG. 3F shows the false negative per cell for both the neural network (“new algorithm”) and non- neural network based algorithm (“POR-YOLO”). The false negative per area using the neural network and methods herein are much lower (lOx or more) than that detected by the non-neural network based algorithm when the target density per area is greater than 1000 per area.
[0372] FIG. 3G shows comparison of accuracy of identifying transcripts using the methods, e.g., 2800, and a traditional non-neural network based algorithm. In this case, simulated flow cell images of in situ sample with multiple cells are used. Each cell may include a number of transcripts ranging from 0 to 6000. The neural network herein, e.g., in method 2800, and a classic non-neural network based algorithm are used to predict / detect transcripts in such cells. And the prediction / determination is then compared with ground truths (or equivalently, the reference polony map) for accuracy. The R2values show correlations of the prediction / determination with the references. The neural network and method herein, e.g., method 2800, showed consistently higher correlation with all the different numbers of transcripts per cell than the correlation using classic non- neural network based algorithm, thereby indicating higher accuracy in identifying polonies or clusters in flow cell images (e.g., transcripts) of in situ samples.
[0373] In some embodiments, the method 500, 2800 may include an operation of determining a biological analyte including but not limited to a morphological feature, a transcript, a RNA, a mRNA, a protein, or their combinations based on the base calling or classification of the polony in one or more sequencing cycles. For example, base calling or classification sequence of a polony in 6 consecutive sequencing cycles of ATTCGA may indicate a cellular protein that may be labeled by the unique barcode of “ATTCGA.”
[0374] In some embodiments, the method 2800 further include an operation (iv) of: in response to determining that a first pixel of the one or more pixels has a predicted classification that is different from a background (e.g., the classifications may include A,T, C, G, U, or background), determining a first morphological feature, a first RNA or mRNA, or a first protein based on the one or more predicted classifications. In some embodiments, the method 2800 further include an operation (v) of in response to determining that a second pixel of the one or more pixels has a predicted classification that is different from the background classification (e.g., the classifications may include A, T, C, G, U, or background), determining a second morphological feature, a second RNA or mRNA, or a second protein based on the one or more predicted classifications.
[0375] In some embodiments, the method 2800 further include an operation (iv) of determining a first morphological feature, a first RNA or mRNA, or a first protein based on predicted base calls of a first pixel in one or more cycles. In some embodiments, the method 2800 further include an operation (v) of determining a second morphological feature, a second RNA or mRNA, or a second protein based on predicted base calls of a second pixel in one or more cycles.
[0376] In some embodiments, the method 2800 further include an operation of determining a spatial relationship of the first pixel and the second pixel which may include one or more of visualizing the first and second pixels within a common coordinate system, calculating a spatial distance in 2D or 3D between the first and second pixels; and determining whether the first and second pixels are within a same polony or not.
[0377] In some embodiments, the method 2800 further comprises: (iv) in response to determining that a first pixel of the one or more pixels has a predicted classification that is different from a background classification, determining at least a first target of a first morphological feature; a first RNA or mRNA; and a first protein based on the one or more predicted classifications; and (v) in response to determining that a second pixel of the one or more pixels has a predicted classification that is different from the background classification, determining at least a second target different from the first target from: the first morphological feature; the first RNA or mRNA; and the first protein based on the one or more predicted classifications. In some embodiments, the second target is of a different type of target from the first target (e.g., a protein vs. a morphological feature) thereby advantageously enable multi-omics analysis and research of the biological analyte(s) of interest using the methods herein. In some embodiments, the first target and the second target correspond to the biological analyte(s) of the sample. In some embodiments, the method 2800 further comprises: spatially aligning the location of thefirst and the second targets based on the one or more predicted classifications; and determining a biological analyte of the sample immobilized on the support based on the spatial alignment.
[0378] In some embodiments, the methods 2800 herein advantageously allow spatial alignment or in other words, co-localization of two or more different biological analytes using the neural network disclosed herein. Such different biological analyte may be of a different type. For example, a first biological analyte may be a morphological feature, and a second biological analyte may be a protein or mRNA. Such different biological analytes may be sequenced within a same sequencing run in same or different sequencing cycles. Exemplary embodiment of staining and sequencing different target analytes within cells or tissue are disclosed in PCT application No. PCT / US2025 / 10310, filed January 3, 2025, the contents of which are incorporated by reference in their entireties. The number of different biological analytes may be limited by the availability of unique barcodes that may be used to differentiate the biological analyte from others. For example, the number of different biological analytes can be in a range from 2 to 100, 4 to 350, 10 to 500, 50 to 1000, or more. For example, protein A may be localized to be within the nucleus of a specific cell type, while protein B may be localized to be adjacent to a certain transcript within the mitochondria but not within the cytosol based on the prediction of intensities, base calling, and / or classification in one or more cycles using methods 500 or 2800. Identification of such different biological analytes may advantageously provide more information, e.g., spatial relationships, which may facilitate biological, physiological, or pathological analysis of the sample(s) being sequenced. In some embodiments, the biological analytes herein may be any physical features of the sample(s) or source of sample(s). The detection, localization, and spatial alignment of the biological analytes may correspond to various physiological, biological, pathological characteristics of cells or tissue which may advantageously provide information that may advance understanding of cellular function, regulation, and interactions which in turn may advance existing biomedical research, including but not limited to, more effective disease modeling and drug discovery efforts.
[0379] In some embodiments, the method 2800 further comprises an operation of (iv) determining a location of one or more of a first morphological feature, a first RNA or mRNA, a first transcript, and a first protein based on the corresponding location of the one or more predicted base calls or predicted classifications. In some embodiments, themethod 2800 further comprises an operation of (v) determining a location of one or more of: a second morphological feature, a second RNA or mRNA, a second transcript, and a second protein based on the corresponding location of one or more second predicted base calls or predicted classifications. In some embodiments, the method 2800 further comprises an operation of (vi) spatially aligning the location of one or more of: a second morphological feature, a second RNA or mRNA, and second protein with the location of one or more of: the first morphological feature, the first RNA or mRNA, and the first protein; and an operation of (vii) determining a biological character of the sample immobilized on the support based on the spatial alignment.
[0380] In some embodiments, the method 2800 may include an operation of saving the base calls obtained in operation 2812 in a predetermined format, e.g., in a FastQ file compatible with subsequent operations so that subsequent analysis such as adaptor trimming and secondary analysis can be performed.Predicting base calls using patches in flow cell images
[0381] In some embodiments, the method 2800 may include an operation 2812 of (iii) performing, by the processor, a corresponding base calling for each of the determined polonies. In some embodiments, the operation 2812 comprises extracting a plurality of patches from the second plurality of flow cell images based on the polony map. The polony map may be generated using various algorithms, for example, from operation 2804 or 2804’. In some embodiments, the operation 2812 further comprises providing input to the neural network, the input comprising the plurality of patches, wherein each patch comprises one or more patch images from the multiple color channels, and wherein each patch comprises at least a portion of the second plurality of flow cell images; and predicting a plurality of base calls using the neural network and based on the input, wherein each base call corresponds to a corresponding patch.
[0382] In some embodiments, each corresponding patch comprises a polony located at or in close vicinity to a center of the corresponding patch. For example, the polony may be no more than 1 to 10 pixels away from the center of the corresponding patch. In some embodiments, each patch comprises 3 to 128 pixels along a spatial dimension, e.g., along x or y direction. The size of the patches are maintained to be relatively small comparing to the size of the flow cell images, e.g., lOx, 20x, 50x, lOOx, 500x, lOOOx or less than the size of the flow cell image. In some embodiments, the plurality of patches comprises 100to 108patches. In some embodiments, two or more different patches may overlap at least partly with each other. In some embodiments, each patch may contain more than one, two, three, five, or ten polonies therewithin, but only the pixel(s)of the single polony at its center is used for generating base call(s) corresponding to the patch. For example, when each patch include a patch image sized to be 32 by 32, a first patch may include pixels 1- 32 in both x and y directions to cover a polony centered at pixels (16, 16) of the flow cell images, a second patch may include pixels 2-33 in both x and y directions to cover a second polony centered at pixels (17, 17.5), and a third patch may include pixels 5-36 in both x and y directions to cover a third polony centered at pixels (19, 19) of the flow cell images. In some embodiments, instead of using only the single polony for generating reference base calls, reference intensities or making predictions, a very limited number of polonies in each patch may be used. The very limited number of polonies can be in a range from 1 to 4, 1 to 8, 1 to 20, 1 to 50, or 1 to 100. The very limited number of polonies can be lOOx, lOOOx, 104x, 105x, 106x, 107x, or 108x less than a total number of polonies in a corresponding flow cell image.
[0383] In some embodiments, the number of pixels within each patch can be optimized to balance the computational complexity and spatial context information to be included for training the neural network(s). The number of patch images within each patch can be optimized to balance the computational complexity and the spatial context information within each patch for accurate and reliable prediction using the neural network. In some embodiments, the number of pixels within each patch can be at least partly based on polony density of the sample being imaged. In some embodiments, each patch may include multiple pixels, but prediction may only be performed for a single polony at or near the center of the patch. In training the neural network, e.g., using methods 2900, for predicting the base call, similarly reference base calls are only for a single polony at or near the center of the patch. In some embodiments, instead of the single polony, a very limited number of polonies in each patch may be used for training the neural network(s) or making predictions. The very limited number of polonies can be in a range from 1 to 4, 1 to 8, 1 to 20, 1 to 50, or 1 to 100. The very limited number of polonies can be lOOx, lOOOx, 104x, 105x, 106x, 107x, or 108x less than a total number of polonies in a corresponding flow cell image.
[0384] In some embodiments, each patch may comprise multiple patch images corresponding to different color channels. For example, each patch may comprise a patchimage covering same pixels within the x-y plane in three different color channels. The same pixels may be pixels determined after registration to correct for the spatial offset across different color channels. In some embodiments, each patch may comprise multiple patch images corresponding to different cycles, e.g., continuous cycles n-1, n, n+1, within a sequencing run. For example, each patch may comprise 3 images, each from a different color channel in 4 adjacent cycles, so that each patch may comprise 12 patch images in total. When the sample is in 3D, e.g., an in situ cell sample, each patch may include 5 different z levels to make the total number of patch images of 60.
[0385] In some embodiments, at least two patches of the plurality of patches comprise at least partially overlapped patch images that comprise some identical pixels. In some embodiments, each patch of the plurality of patches comprise at least partially overlapped pixels with another patch of the plurality of patches.
[0386] In some embodiments, the first plurality of flow cell images are acquired only from a single color channel so that flow cell images acquired from different color channels may require different neural networks for predicting high resolution intensities, base calls, classifications, etc., as disclosed herein.
[0387] In some embodiments, the first plurality of flow cell images are acquired only from a single z level, so that flow cell images acquired at different z levels of 3D sample(s), e.g., in situ cells, may require different neural network for predicting high resolution intensities, base calls, classifications, etc., as disclosed herein.
[0388] In some embodiments, the first plurality of flow cell images are acquired from the one or more cycles. In some embodiments, the one or more cycles comprises a plurality of cycles in a sequencing run. In some embodiments, the one or more cycles comprises a current cycle N, and the first plurality of flow cell images are acquired from at least one cycle prior to the current cycle N. The current cycle N is a cycle in which sequencing is currently being performed in of a sequencing cycle. In some embodiments, the flow cell images may have been acquired in the current cycle N, but no flow cell images have been acquired in the next cycle N+1.
[0389] In some embodiments, the operation 2802 (ii) of providing, by the processor or the first reconfigurable logic device, the second plurality of flow cell images as the input to the neural network comprises: (ii) providing, by the processor or the first reconfigurable logic device, the second plurality of flow cell images as the input to the neural network without providing a polony map or locations of polonies in the secondplurality of flow cell images as the input to the neural network. In other words, the operation (ii) of method 2800 does not require the input of a polony map, a location list of polonies, or the like to be provided as input to the neural network in order to predict the base calls. In some embodiments, the spatial location of the polonies within the flow cell images, e.g., the second plurality of flow cell images are not used in predicting the base calling using the neural network. Instead, each patch may contain relative spatial information of the polony with respect to the rest of the pixels in the same patch(es) that may be used for predicting the base calling using the neural network. The method 2800 may predict base calling, e.g., in operation 2812, without using the input of a polony map, a location list of polonies, or the like. Instead, the polony map, the location list of polonies, or the like may be used to extract the plurality of patches from the second plurality of flow cell images.
[0390] In some embodiments, the operation of predicting the plurality of base calls using the neural network and based on the input, wherein each base call corresponds to a corresponding patch comprises: predicting a probability map for each channel of the multiple color channels corresponding to the corresponding patch; and determining the base call of the corresponding patch based on the probability maps. For example, for flow cell images from 4 different color channels, 4 different probability maps may be generated. Each probability map may have the same size and dimension as the flow cell images or covering at least a portion of the flow cell images. Each pixel in the probability map may a probability value corresponding to the channel. As an example, pixel (12,12) may have a probability value of 0.2, 0.01, 0.2, and 0.59 in 4 different channels representing nucleotides A, T, C, and G, and the base call of pixel (12, 12) may be determined as the largest probability among probabilities of different color channels, which is 0.59 and correspond to nucleotide G for its base calling. The neural network may be trained to predict probability maps. In some embodiments, training of the neural network to predict probability maps can be based on reference polony maps or any equivalent information indicative of polony locations, e.g., a location list of polonies. In some embodiments, the neural network to predict probability maps can be trained by comparing each probability map to a corresponding reference polony map. In some embodiments, the neural network may be trained to minimize a loss function based on the comparison of the probability map and the corresponding reference polony map. For example, a probability map may be initialized to have random values in each pixel, andthe neural network may be trained to produce higher value for pixel(s) corresponding to polonies than pixels corresponding to non-polony structure(s) in the probability map. In some embodiments, the sum of values for each pixel in all probability maps of different color channels may add up to a fixed number, e.g., 1, 10, 100, etc. As an example, pixel (24, 25) in 3 probability maps corresponding to 3 different color channels may be 0.24, 0.51, and 0.25, which adds up to 1. In some embodiments, each base call corresponds to a corresponding patch which includes one or more patch images. In some embodiments, the operation of predicting the plurality of base calls using the neural network and based on the input comprises: generating a first single intensity for a first channel of the multiple color channels corresponding to the corresponding patch; and determining the base call of the corresponding patch based on the single intensity. As an example, a first single intensity of a first color channel may be determined using prediction by the neural network disclosed herein. The first single intensity may or may not be normalized. The first single intensity may correspond to the single polony of the corresponding patch containing one or multiple patch images of the same polony at adjacent cycles of a sequencing run. The first single intensity may correspond to one of the adjacent cycles, e.g., a current cycle. A base call may be determined based on the first single intensity of the current cycle, e.g., by comparing the first single intensity with other intensities of the same polony from other color channels. The other intensities may be predicted similarly using the same or different neural networks.
[0391] In some embodiments, the method further comprises an operation of predicting a second single intensity for a second channel of the multiple color channels corresponding to the corresponding patch using a second neural network; and determining the base call of the corresponding patch based on at least the first single intensity and the second single intensity.
[0392] In some embodiments, the method further comprises an operation of predicting a second single intensity for a second channel of the multiple color channels corresponding to the corresponding patch using a second neural network or the same first neural network; and an operation of predicting a third single intensity for a third channel of the multiple color channels corresponding to the corresponding patch using a third neural network or the same first neural network; and determining the base call of the corresponding patch based on at least the first, second, and third single intensities. For example, at a current cycle N, the first, second, and third intensities may be predictedusing different neural networks (e.g., each of the neural networks may be trained using different training data but with identical neural network layers and numbers of parameters) to be 50, 690, 80 for the same polony. The base call of the polony may correspond to the nucleotide that lights up in the second color channel with an intensity of 690 but not the first or third color channel.
[0393] In some embodiments, the operation (iii) of predicting, by the first reconfigurable device or the integrated circuit, one or more base calls corresponding to one or more polonies of the second plurality of flow cell images using the neural network comprises: determining two or more pixels of the second plurality of flow cell images as duplications of a single polony; and selecting one pixel of the two or more pixels as a center of the single polony. In some embodiments, the two or more pixels may be at a same z level. In some embodiments, the two or more pixels may be at different z levels. Exemplary embodiments of the operation of determining two or more pixels of the second plurality of flow cell images as duplications of a single polony and selecting one pixel of the two or more pixels as a center of the single polony are disclosed in PCT Application No. PCT / US23 / 76125, and is incorporated herein by reference in its entirety.
[0394] Although embodiments herein are disclosed with a focus on using and training neural networks, other artificial intelligence-based models may also be used for similar purposes. In some embodiments, the methods 500 and 2800 herein may be performed using artificial intelligence-based models other than neural networks. In some embodiments, the methods 600, 700 and 2900 may be used to train artificial intelligencebased models other than neural networks for making predictions or inferences using methods 500 or 2800. Some non-limiting examples of the artificial intelligence-based models include: random forest, decision tree, k-mean clustering, and gradient boosted tree. In some embodiments, the artificial intelligence-based models may be used to predict intensities, classifications, or base calls by working on intensities from flow cell images and / or the high resolution flow cell images. In some embodiments, the artificial intelligence-based models other than neural networks may predict intensities, classifications, or base calls using information only including intensities, and such information may lack spatial context of the intensities, shapes of the polonies, background noise, signal from other cellular structures, etc. In some embodiments, the neural networks herein predict intensities, classifications, or base calls by advantageously using the flow cell images or high resolution flow cell images which not only include theintensities but also other information including but not limited to background noise, polony sizes and shapes, spatial relationship among polonies, etc. for more accurate predictions or inferences.
[0395] In some embodiments, the neural network herein is a convolutional neural network (CNN). In some embodiments, the neural network is a 3D CNN. In some embodiments, the neural network is a 2D CNN. In some embodiments, the neural network comprises one or more convolutional layers. In some embodiments, the neural network is a recurrent neural network (RNN). In some embodiments, the neural network is a 3D RNN. In some embodiments, the neural network is a 2D RNN. In some embodiments, the neural network comprises one or more long short-term memory (LSTM) layers. In some embodiments, the neural network is a U-Net. In some embodiments, the neural network includes a residual network (ResNet). In some embodiments, the neural network can include a transformer based model like a vision transformer (ViT). In some embodiments, the neural network comprises a U-Net with a first predetermined repetition of down-sampling and convolution operations and then a second predetermined repetition of up-sampling, concatenation, and convolution operations. The first and second predetermined repetition can have an identical quantity, e.g., 3 or 4. In some embodiments, the neural network is a U-Net with a first predetermined number of filters in each repetition of down sampling, and then a second predetermined number of filters in each repetition of up sampling and / or concatenation. For example, the first predetermined number of filters can be 32, 64, 128, and 256 filters in three repetitions and the second predetermined number can be 128, 64, 64, and 32 filters in the corresponding three repetitions. As another example, the first predetermined number of filters can be 32, 64, 128, and 256 filters in three repetitions and the second predetermined number can be 256, 128, 64, and 32 filters in the corresponding three repetitions.
[0396] In some embodiments, the operation 2812 may comprise: performing, by the processor, a first convolution in one or more dimensions on the first plurality of flow cell images, thereby generating a first convolution result; repetitively performing, for one or more times, down-sampling operations comprising: (a) performing, by the processor, a second convolution in one or more dimensions on the first convolution result, thereby generating a second convolution result; and (b) performing, by the processor, a down sampling of the second convolution result by a down sampling factor thereby generating afirst down-sampled result. In each repetition, the second convolution may comprises a corresponding number of filters, thereby generating a third convolution result after the repetitions.
[0397] In some embodiments, the operation 2812 may further comprise: performing, by the processor, the second convolution in one or more dimensions on the third convolution result, thereby generating a fourth convolution result; repetitively performing, for one or more times, up sampling operations comprising: (c) performing, by the processor, an up sampling of the fourth convolution result by an up sampling factor thereby generating a first up-sampled result; and (d) performing, by the processor, the second convolution in one or more dimensions of the first up-sampled result, thereby generating a fifth convolution result. In each repetition, the second convolution may comprise a corresponding number of filters, thereby generating a sixth convolution result after the repetitions.
[0398] In some embodiments, the first convolution comprises a 3D convolution with a convolution kernel. In some embodiments, the convolutional kernel may have 4 dimensions. In some embodiments, the convolutional kernel is m*m*m for the first three spatial dimensions and the size of its fourth dimension is determined by the filter number in the corresponding repetition. In some embodiments, m can be an integer in the range of 2 to 20. For example, the input can be 512x512 flow cell images, and the z-stack can have 12 slices. The first convolution can include 32 filters and each filter has one kernel that is 3x3x3xl. The output from that convolutional block is 512x512x12x32. Then there is a double convolutional block, i.e., the second convolution having two first convolutions with 32 filters. The input to both of those blocks is 512x512x12x32 and the output is 512x512x12x32. Each filter uses a kernel sized 3x3x3x3x32. The number of filters may correspond to features of the input.
[0399] In some embodiments, the second convolution comprises two 3D convolutional layers, e.g., as shown in the pseudo code. In other words, the second convolution comprises two repetition or blocks of the first convolution in 3D, and usage of the output and the number of filters changes, as convolution process will increase the depth of the image. The depth of image may increase as the number of features or filters increases. In some embodiments, the first and second resolution is in 2D or 3D.
[0400] In some embodiments, the first convolution comprises a 2D convolution with a convolution kernel. In some embodiments, the convolutional kernel may have 3dimensions. In some embodiments, the convolutional kernel is m x m for the first two spatial dimensions and the size of its third dimension is determined by the filter number in the corresponding repetition. In some embodiments, m can be an integer in the range of 2 to 20. For example, the input can be flow cell images with a size of 512x512x1. The first convolution can include 64 filters and each filter has one kernel that is 3x3x1. The output from that convolutional block is 512x512x64. Then there is a double convolutional block, i.e., the second convolution having two first convolutions with 32 filters. The input to both of those blocks is 512x512x64 and the output is 512x512x32. Each filter can use a kernel sized 3x3x32.
[0401] In some embodiments, the second convolution comprises at least two convolutional layers or exactly two convolutional layers, e.g., as shown in the pseudo codes. In other words, the second convolution comprises two repetition or blocks of the first convolution, and usage of the output and the number of filters changes, as convolution process will increase the depth of the image. The depth of image may increase as the number of features or filters increases. In some embodiments, the first and second resolution is in 2D or 3D.
[0402] In some embodiments, the second convolution in operation (a) comprises a corresponding number of n, 2*n, 4*n, and 8*n filters in a first, second, third, and fourth repetition, respectively. In some embodiments, the second convolution in operation (c) comprises a corresponding number of 2*n, 2*n, 4*n, 8*n filters in a last repetition, last minus one, last minus two, and last minus three repetition, respectively. In some embodiments, n can be an integer in the range from 8 to 256. For example, operation (a) comprises 32, 64, 128, and 256 filters in three repetitions and operation (c) comprises 128, 64, 64, and 32 filters in the corresponding three repetitions.
[0403] In some embodiments, the second convolution in operation (c) comprises a corresponding number of n, 2*n, 4*n, 8*n filters in a last repetition, last minus one, last minus two, and last minus three repetition, respectively. For example, operation (a) comprises 32, 64, 128, and 256 filters in four repetitions and operation (c) comprises 256, 128, 64, and 32 filters in the corresponding four repetitions.
[0404] In some embodiments, the second convolution in operation (c) comprises a corresponding number of n, 2*n, 4*n filters in a last repetition, last minus one, last minus two, repetition, respectively. For example, operation (a) comprises 32, 64, 128 filters inthree repetitions and operation (c) comprises 128, 64, and 32 filters in the corresponding three repetitions.
[0405] In some embodiments, the operation 2800 may further comprise: performing, by the processor, the first convolution in one or more dimensions on the sixth convolution result, thereby generating a seventh convolution result; and predicting, by the processing, the second plurality of flow cell images based on the seventh convolution result. Each of the second plurality of flow cell images may correspond to the corresponding flow cell image of the first plurality of flow cell images with a second resolution that is 2, 4, 6, 8, 10, 12, or 16 times greater than the first resolution in one or more spatial dimensions. In some embodiments, the second resolution is at least 4, 6, or 8 times greater than the first resolution in all three dimensions.
[0406] In some embodiments, the first plurality of flow cell images are from a single color channel. In some embodiments, the first plurality of flow cell images are from one or more color channels. In some embodiments, the first plurality of flow cell images are of unbalanced nucleotide diversity in one or more sequencing cycles. In some embodiments, the cellular sample comprises overloaded concatemer molecules with a spatial density in a range of 102-1015per mm2. In some embodiments, the cellular sample comprises overloaded concatemer molecules with a spatial density in a range of 103-10102 per mm .
[0407] In some embodiments, the first resolution is in a range of 0.1 um to 5 um. In some embodiments, the first resolution is in a range of 0.01 um to 10 um. In some embodiments, the second resolution is in a range of 0.02 um to 2 um. In some embodiments, the second resolution is in a range of 0.001 um to 3 um. In some embodiments, the down-sampling factor is 2, 4, 6, 8, 16, or more. In some embodiments, the up-sampling factor is 2, 4, 6, 8, 16, or more.
[0408] In some embodiments, one or more of operations, e.g., operation 2810, 2802, 2804, 2804’, 2806, 2812, are performed while a sequencing run is being performed. In some embodiments, one or more operations are performed in parallel as the corresponding sequencing run to reduce sequencing analysis time.
[0409] In some embodiments, the sequencing analysis time includes a t...
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A computer-implemented method comprising:(i) generating, by a sequencing system, a first plurality of flow cell images of a sample immobilized on a support by conducting one or more cycles of sequencing reactions in one or more color channels, wherein the first plurality of flow cell images are acquired with a first resolution;(ii) providing, by a processor or a first reconfigurable logic device, the first plurality of flow cell images as an input to a neural network, wherein the neural network is pre-trained using a training data set of training flow cell images;(iii) predicting, by the first reconfigurable device or an integrated circuit, a second plurality of flow cell images using the neural network, wherein each of the second plurality of flow cell images is of a second resolution and corresponds to a corresponding image of the first plurality of flow cell images, and wherein the second resolution is at least 2 to 32 times greater than the first resolution in one or more spatial dimensions;(iv) determining, by the processor, the first reconfigurable logic device, or the integrated circuit, polonies from the second plurality of flow cell images; and(v) performing, by the processor, the first reconfigurable logic device, or the integrated circuit, a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images.
2. A computer-implemented method for predicting base calls comprising:(i) generating, by a sequencing system, a first plurality of flow cell images of a sample immobilized on a support by conducting one or more cycles of sequencing reactions in one or more color channels, the first plurality of flow cell images comprising a first resolution;(ia) generating, by a processor, a first reconfigurable logic device, or an integrated circuit, a second plurality of flow cell images comprising a second resolution, wherein the second resolution is at least 2 to 32 times greater than the first resolution in one or more spatial dimensions;(ii) providing, by the processor or the first reconfigurable logic device, or the integrated circuit, the second plurality of flow cell images as an input to a neural network; orproviding, by the processor or the first reconfigurable logic device, or the integrated circuit, the second plurality of flow cell images to a polony map generation algorithm or a base calling algorithm;(iii) predicting, by the first reconfigurable device or the integrated circuit, one or more base calls corresponding to one or more polonies of the second plurality of flow cell images using the neural network; or predicting, by the first reconfigurable device or the integrated circuit, one or more classifications corresponding to one or more pixels of the second plurality of flow cell images using the neural network.
3. The computer-implemented method of claim 2 further comprising:(iia) determining, by the processor, the first reconfigurable device, or the integrated circuit, a polony map based on the second plurality of flow cell images; and(iiia) determining, by the processor, the first reconfigurable logic device, or the integrated circuit, a corresponding location of the one or more predicted base calls or the one or more predicted classifications based on the polony map.
4. The computer-implemented method of claim 2, wherein the one or more pixels include at least one pixel that is not comprised in any polony of the polony map.
5. The computer-implemented method of claim 2, wherein the one or more pixels include at least one pixel that is not comprised in any polony in the polony map and at least one pixel that is comprised in at least one polony in the polony map.
6. The computer-implemented method of claim 2 further comprising:(iv) in response to determining that a first pixel of the one or more pixels has a predicted classification that is different from a background classification, determining a first morphological feature, a first RNA or mRNA, or a first protein based on the one or more predicted classifications; and(v) in response to determining that a second pixel of the one or more pixels has a predicted classification that is different from the background classification, determining a second morphological feature, a second RNA or mRNA, or a second protein based on the one or more predicted classifications.
7. The computer-implemented method of claim 3 further comprising:(iv) determining a location of one or more of: a first morphological feature, a first RNA or mRNA, and a first protein based on the corresponding location of the one or more predicted base calls or predicted classifications.
8. The computer-implemented method of claim 7 further comprising:(v) determining a location of one or more of: a second morphological feature, a second RNA or mRNA, and a second protein based on the corresponding location of one or more second predicted base calls or predicted classifications.
9. The computer-implemented method claim 8 further comprising:(vi) spatially aligning the location of one or more of: a second morphological feature, a second RNA or mRNA, and second protein with the location of one or more of: the first morphological feature, the first RNA or mRNA, and the first protein; and(vii) determining a biological character of the sample immobilized on the support based on the spatial alignment.
10. The computer-implemented method of any one of claims 2-9, wherein the neural network is pre-trained using a training data set of training flow cell images.
11. The computer-implemented method of claim 10, wherein the neural network is pretrained using the training data set of training flow cell images prior to the operation (iia).
12. The computer-implemented method of any one of the preceding claims, wherein the neural network is not trained or retrained after operation (iia) and prior to operation (iii).
13. The computer-implemented method of any one of the preceding claims, wherein the sample comprises concatemer molecules therewithin, wherein the first plurality of flow cell images are acquired at a z-stack of different z-locations with the first resolution along a z direction.
14. The computer-implemented method of any one of the preceding claims, wherein the first plurality of flow cell images are acquired from a single color channel.
15. The computer-implemented method of any one of the preceding claims, wherein the one or more cycles comprises a current cycle N, and the first plurality of flow cell images are acquired from at least one cycle prior to the current cycle N.
16. The computer-implemented method of any one of the preceding claims, wherein the one or more cycles comprises a plurality of cycles in a sequencing run.
17. The computer-implemented method of any one of the preceding claims, wherein (iia) determining, by the first reconfigurable device or the integrated circuit, the polony map based on the second plurality of flow cell images comprises: predicting, using the neural network and by the first reconfigurable device or the integrated circuit, the polony map based on the second plurality of flow cell images.
18. The computer-implemented method of any one of the preceding claims, wherein (iia) comprises: predicting, by the first reconfigurable device or the integrated circuit, a base call corresponding to each polony of the second plurality of flow cell images using the neural network at the second resolution or a third resolution; and determining the polony map based on the predicted base calls and a corresponding quality index of each predicted base call at the second or third resolution.
19. The computer-implemented method of any one of the preceding claims, wherein the third resolution is at least 2 to 32 times greater than the first or second resolution in one or more spatial dimensions.
20. The computer-implemented method of any one of the preceding claims, wherein the third resolution is greater than the first and second resolution in one or more spatial dimensions.
21. The computer-implemented method of any one of the preceding claims, wherein the polony map comprises a spatial coordinate for each of at least some of polonies in the second plurality of flow cell images.
22. The computer-implemented method of any one of claims 2-21, wherein (ii) providing, by the processor or the first reconfigurable logic device, the second plurality of flow cell images as the input to the neural network comprises:(ii) providing, by the processor or the first reconfigurable logic device, the second plurality of flow cell images as the input to the neural network without providing a polony map or locations of polonies in the second plurality of flow cell images as the input to the neural network.
23. The computer-implemented method of any one of claims 2-22, wherein (iii) predicting, by the first reconfigurable device or an integrated circuit, base calls corresponding to oneor more polonies of the second plurality of flow cell images using the neural network, comprises: extracting a plurality of patches from the second plurality of flow cell images based on the polony map; providing an input to the neural network, the input comprising the plurality of patches, wherein each patch comprises one or more patch images from the one or more color channels, and wherein each patch comprises at least a portion of the second plurality of flow cell images; and predicting a plurality of base calls using the neural network and based on the input, wherein each base call corresponds to a corresponding patch.
24. The computer-implemented method of claim 23, wherein each corresponding patch comprises a polony located at or in close vicinity to a center of the corresponding patch.
25. The computer-implemented method of any one of claims 23-24, wherein each corresponding patch comprises 2 to 128 pixels along a spatial dimension.
26. The computer-implemented method of any one of claims 23-25, the plurality of patches comprises 100 to 106patches.
27. The computer-implemented method of any one of claims 23-26, wherein at least two patches of the plurality of patches comprise at least partially overlapped patch images that comprise identical pixels.
28. The computer-implemented method of any one of claims 23-27, wherein each patch of the plurality of patches comprise at least partially overlapped pixels with another patch of the plurality of patches.
29. The computer-implemented method of any one of the preceding claims, the one or more color channels comprises 2, 3, or 4 color channels.
30. The computer-implemented method of any one of claims 2-29, the method further comprising: up-sampling the first plurality of flow cell images to generate the second plurality of flow cell images.
31. The computer-implemented method of any one of claims 23-30, wherein each patch comprises one or more patch images from the one or more color channels and one or more cycles.
32. The computer-implemented method of any one of the preceding claims, wherein the one or more cycles comprise continuous cycles of a sequencing run.
33. The computer-implemented method of any one of claims 23-32, wherein predicting the plurality of base calls using the neural network and based on the input, wherein each base call corresponds to a corresponding patch comprises: predicting a probability map for each channel of the one or more color channels corresponding to the corresponding patch, wherein each probability map comprise probability values of a base calling for each pixel of the corresponding patch; and determining the base call of the corresponding patch based on the probability maps of the one or more channels.
34. The computer-implemented method of any one of claims 23-33, wherein predicting the plurality of base calls using the neural network and based on the input, wherein each base call corresponds to a corresponding patch comprises: generating a first single intensity for a first channel of the one or more color channels corresponding to the corresponding patch; and determining the base call of the corresponding patch based on the single intensity.
35. The computer-implemented method of any one of claims 23-34, wherein predicting the plurality of base calls using the neural network and based on the input, wherein each base call corresponds to a corresponding patch comprises: generating a first single intensity for a first channel of the one or more color channels corresponding to the corresponding patch; and determining the base call of the corresponding patch based on the single intensity.
36. The computer-implemented method of claim 35 further comprising: predicting a second single intensity for a second channel of the one or more color channels corresponding to the corresponding patch using a second neural network; and determining the base call of the corresponding patch based on at least the first single intensity and the second single intensity.
37. The computer-implemented method of claim 36 further comprising: predicting a second single intensity for a second channel of the one or more color channels corresponding to the corresponding patch using a second neural network; andpredicting a third single intensity for a third channel of the one or more color channels corresponding to the corresponding patch using a third neural network; and determining the base call of the corresponding patch based on at least the first, second, and third single intensities.
38. The method of claim 2, wherein (iii) predicting, by the first reconfigurable device or the integrated circuit, one or more base calls corresponding to one or more polonies of the second plurality of flow cell images using the neural network comprises: determining two or more pixels of the second plurality of flow cell images as duplications of a single polony; and selecting one pixel of the two or more pixels as a center of the single polony.
39. The computer-implemented method of claim 1, wherein the sample comprises concatemer molecules therewithin, and wherein the first plurality of flow cell images are acquired at a z-stack of different z-locations.
40. A computer-implemented method comprising:(i) generating, by a sequencing system, a first plurality of flow cell images of a sample immobilized on a support by conducting one or more cycles of sequencing reactions, wherein the first plurality of flow cell images are of a first resolution;(ii) providing, by a processor or a first reconfigurable logic device, the first plurality of flow cell images as an input to a neural network, wherein the neural network is pre-trained using a training data set of training flow cell images and reference base calls of the training dataset;(iii) predicting, by the first reconfigurable device or an integrated circuit, a second plurality of flow cell images using the neural network, wherein each of the second plurality of flow cell images is with a second resolution and corresponds to a corresponding image of the first plurality of flow cell images, and wherein the second resolution is at least 2 to 32 times greater than the first resolution in one or more spatial dimensions;(iv) determining, by the processor, the first reconfigurable logic device, or the integrated circuit, polonies from the second plurality of flow cell images; and(v) performing, by the processor, the first reconfigurable logic device, or the integrated circuit, a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images.
41. The method of any one of the preceding claims, wherein the sample is a cellular sample comprising in situ cells, tissue, both.
42. The method of any one of the preceding claims, wherein the neural network is a convolutional neural network.
43. The method of any one of the preceding claims, wherein the sample is a 3D sample, and the neural network is a 3D neural network.
44. The method of any one of the preceding claims, wherein the sample is 2D sample and wherein the neural network is a 2D neural network.
45. The computer-implemented method of any one of the preceding claims, wherein the neural network is a convolutional neural network.
46. The computer-implemented method of any one of the preceding claims, wherein the neural network is pre-trained using one or more loss functions based on comparing training base calls of the training flow cell images to the reference base calls of the training flow cell images.
47. The computer-implemented method of any one of the preceding claims, wherein the training flow cell images are only from a single color channel.
48. The computer-implemented method of any one of the preceding claims, wherein each of the training flow cell images comprise flow cell images of a same field of view from a plurality of sequencing cycles stacked along a time dimension.
49. The computer-implemented method of any one of the preceding claims, wherein each of the training flow cell images comprise flow cell images of a same field of view from one or more cycles.
50. The computer-implemented method of any one of the preceding claims, wherein each of the training flow cell images comprise flow cell images of the sample at one or more z- levels.
51. The computer-implemented method of any one of the preceding claims, wherein (iii) predicting the second plurality of flow cell images using the neural network comprises predicting high resolution post-processing images of the first plurality of flow cell images, and wherein the processing comprises one or more of: noise reduction,background reduction; intensity offset correction; intensity normalization; color correction; phasing and / or dephasing; image registration; and deconvolution.
52. The computer-implemented method of any one of the preceding claims, wherein (v) performing the corresponding base calling for each of the determined polonies based on the second plurality of flow cell images lacks usage of a neural network.
53. The computer-implemented method of any one of the preceding claims, wherein (iv) determining the polonies from the second plurality of flow cell images lacks usage of a neural network.
54. The computer-implemented method of any one of the preceding claims, wherein (iv) determining the polonies from the second plurality of flow cell images comprises determining a location for a center of each of the polonies from the second plurality of flow cell images.
55. The computer-implemented method of any one of the preceding claims, wherein (iv) determining the polonies from the second plurality of flow cell images lacks usage of a neural network.
56. The computer-implemented method of any one of the preceding claims, wherein the neural network comprises a first image processing part and a second base calling part.
57. The computer-implemented method of any one of the preceding claims, wherein predicting, by the first reconfigurable device or the integrated circuit, the second plurality of flow cell images using the neural network comprises: generating output images from the first image processing part of the neural network as the second plurality of flow cell images without going through the second base calling part of the neural network.
58. The computer-implemented method of any one of the preceding claims, wherein the first image processing part comprises at least part of one or more of: an input layer, a convolutional layer, a pooling layer, an embedding layer, an output layer, and an encoder of the neural network, and wherein the second base calling part lacks any one of: a convolutional layer, a pooling layer, an embedding layer, an output layer, an encoder, and a decoder of the neural network.
59. The computer-implemented method of any one of the preceding claims, wherein the first image processing part comprises at least part of one or more of: an input layer, aconvolutional layer, a pooling layer, an embedding layer, and an encoder of the neural network, and wherein the second base calling part comprises at least part of one or more of: an input layer, an output layer, a convolutional layer, a pooling layer, an embedding layer, an encoder, and a decoder of the neural network.
60. The computer-implemented method of any one of the preceding claims, wherein the processor comprises one or more of: a CPU, a GPU, a TPU, a NPU, a FPGA, and an Al chip.
61. The computer-implemented method of any one of the preceding claims, wherein the processor comprises one or more of: a CPU, a NPU, and a FPGA.
62. The computer-implemented method of any one of the preceding claims, wherein the processor lacks any NPU, FPGA, or Al chips.
63. The computer-implemented method of any one of the preceding claims, wherein the first reconfigurable logic device comprises one or more FPGA units.
64. The computer-implemented method of any one of the preceding claims, wherein the integrated circuit comprises one or more NPUs, Al chips, or both.
65. The computer-implemented method of any one of the preceding claims, wherein one or more of the processor, the first reconfigurable logic device, and the integrated circuit is comprised in the sequencing system within a single housing of the sequencing system.
66. The computer-implemented method of any one of the preceding claims, wherein the integrated circuit is in data communication with the first reconfigurable logic device.
67. The computer-implemented method of any one of the preceding claims, wherein the convolutional network comprises a U-Net.
68. The computer-implemented method of any one of the preceding claims, wherein the first convolution comprises a 3D convolution with a convolution kernel.
69. The sequencing system of any one of the preceding claims, wherein the first convolution comprises a 2D convolution with a convolution kernel.
70. The sequencing system of any one of the preceding claims, wherein the convolutional kernel have at least three dimension.
71. The sequencing system of any one of the preceding claims, wherein the neural network is a 3D convolutional neural network.
72. The sequencing system of any one of the preceding claims, wherein the neural network is a 2D convolutional neural network.
73. The computer-implemented method of any one of the preceding claims, wherein the first and second resolution is in 3D.
74. The computer-implemented method of any one of the preceding claims, wherein the first plurality of flow cell images are from a single color channel.
75. The computer-implemented method of any one of the preceding claims, wherein (v) performing, by the processor, a corresponding base calling for each of the determined polonies based on the second plurality of flow cell images comprises: performing, by the processor, a corresponding base calling for each of the determined polonies based on a fourth plurality of flow cell images, wherein the fourth plurality of images are predicted using a second neural network based on a third plurality of flow cell images.
76. The computer-implemented method of any one of the preceding claims, wherein the third plurality of flow cell images are acquired from one or more color channels that is different from the single channel, and wherein the third plurality of flow cell images comprises the first resolution.
77. The computer-implemented method of any one of the preceding claims, wherein the fourth plurality of flow cell images comprises the second resolution.
78. The computer-implemented method of any one of the preceding claims, wherein the first plurality of flow cell images are from one or more color channels.
79. The computer-implemented method of any one of the preceding claims, wherein the first plurality of flow cell images comprises: an unbalanced diversity of nucleotide bases of A, G, C and T / U among concatemer molecules immobilized on the support in one or more cycles.
80. The computer-implemented method of any one of the preceding claims, wherein two or more different concatemer molecules among the concatemer molecules have differentinsert sequences that correspond to different target RNA molecules or target cDNA molecules.
81. The computer-implemented method of any one of the preceding claims, wherein each location of the determined polonies corresponds to a location of the concatemer molecules.
82. The computer-implemented method of any one of the preceding claims, wherein the unbalanced diversity of nucleotide bases of A, G, C and T / U among the concatemer molecules comprises: a percentage of (1) a number of one or more types of nucleotide bases to (2) a total number of bases is less than 20%, 15%, 10%, or 5% in the one or more cycles.
83. The computer-implemented method of any one of the preceding claims, wherein the sample comprises overloaded concatemer molecules with a spatial density in a range of 103-IO10per mm2.
84. The computer-implemented method of any one of the preceding claims, wherein the first resolution is in a range of 0.1 um to 5 um.
85. The computer-implemented method of any one of the preceding claims, wherein the down-sampling factor is 2, 4, or 8.
86. The computer-implemented method of any one of the preceding claims, wherein one or more of operations (ii) to (v) are performed while a sequencing run is being performed.
87. The computer-implemented method of any one of the preceding claims, wherein the one or more cycles comprises a current cycle N.
88. The computer-implemented method of any one of the preceding claims, wherein N is in a range from 1 to 500.
89. The computer-implemented method of any one of the preceding claims, wherein one or more of operations (ii) to (v) are performed while the sequencing reactions in cycles subsequent to the current cycle N is yet to be performed or currently being performed.
90. The computer-implemented method of any one of the preceding claims, wherein the training data set of flow cell images comprises z-stacks of flow cell images taken at different z-locations.
91. The computer-implemented method of any one of the preceding claims, wherein the second resolution is at least 4, 6, or 8 times greater than the first resolution in all three dimensions.
92. The computer-implemented method of any one of the preceding claims further comprising: registering the second plurality of flow cell images to a common coordinate system.
93. The computer-implemented method of any one of the preceding claims, wherein the first plurality of flow cell images are acquired from a single color channel of the sequencing system.
94. The computer-implemented method of any one of the preceding claims, wherein (vi) determining, by the processor, polonies from the second plurality of flow cell images comprises: generating a polony map comprising spatial location of polonies based on the determined polonies in (iv).
95. The computer-implemented method of any one of the preceding claims, wherein generating the polony map comprising spatial location of polonies based on the determined polonies in (iv) further comprises: deleting duplicate polonies from the determined polonies, wherein the duplicate polonies are out-of-focus.
96. The computer-implemented method of any one of the preceding claims, wherein determining, by the processor, polonies from the second plurality of flow cell images comprises: superimposing the second plurality of flow cell images with corresponding cell staining images; and generating the polony map by only including polonies that are within cell boundaries in the corresponding cell staining images.
97. The computer-implemented method of any one of the preceding claims, wherein the support comprises a glass or plastic substrate.
98. The computer-implemented method of any one of the preceding claims, wherein the support is comprised in a flow cell device.
99. The computer-implemented method of any one of the preceding claims further comprising: providing the sample harboring a plurality of RNA which comprises the first target RNA molecule and the second target RNA molecule.
100. The computer-implemented method of any one of the preceding claims further comprising: generating inside the sample a plurality of cDNA molecules which include a first target cDNA molecule that corresponds to the first target RNA molecule and a second target cDNA molecule that corresponds to the second target RNA molecule.
101. The computer-implemented method of any one of the preceding claims further comprising: contacting the plurality of cDNA molecules in the sample with a plurality of target-specific padlock probes which includes at least a first plurality of first targetspecific padlock probes and a second plurality of second target-specific padlock probes.
102. The computer-implemented method of any one of the preceding claims further comprising: contacting the plurality of RNA molecules in the sample with a plurality of targetspecific padlock probes which includes at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes.
103. The computer-implemented method of any one of the preceding claims, wherein individual padlock probes in the first plurality of first target-specific padlock probes comprise: first and second terminal regions, wherein the first terminal region selectively hybridizes to a first region of the first target cDNA molecule or the first target RNA molecule, and the second terminal region selectively hybridizes to a second region of the first target cDNA molecule or the first target RNA molecule.
104. The computer-implemented method of any one of the preceding claims, wherein the first target-specific padlock probe comprises a first target barcode sequence that corresponds to an uniquely identifies the first target cDNA sequence or the first target RNA sequence.
105. The computer-implemented method of any one of the preceding claims, wherein the first target-specific padlock probe comprises a first target barcode sequence that is locatedadjacent to one of the regions of the first target-specific padlock probe that selectively hybridizes to the first target cDNA molecule or the first target RNA sequence.
106. The computer-implemented method of any one of the preceding claims, wherein the first target-specific padlock probe comprises at least one universal adaptor sequence.
107. The computer-implemented method of any one of the preceding claims, wherein the first target-specific padlock probe comprises a universal primer binding site for a rolling circle amplification primer or a complementary sequence thereof.
108. The computer-implemented method of any one of the preceding claims, wherein the first target-specific padlock probe comprises a universal compaction oligonucleotide binding site or a complementary sequence thereof.
109. The computer-implemented method of any one of the preceding claims, wherein the plurality of nucleotide reagents comprise: multivalent molecules, nucleotides, nucleotide analogs, or their combinations.
110. The computer-implemented method of any one of the preceding claims, wherein individual nucleotides or nucleotide analogs are detectably labeled or non-labeled.
111. The computer-implemented method of any one of the preceding claims, wherein the detectably labeled individual nucleotides or nucleotide analogs comprises a different detectable color label that corresponds with each different type of nucleotide base of A, G, C, and T / U.
112. The computer-implemented method of any one of the preceding claims, wherein an individual multivalent molecule comprise a core attached with multiple nucleotide arms and each arm of the individual multivalent molecule comprises the same type of nucleotide base.
113. The computer-implemented method of any one of the preceding claims, wherein generating the first plurality of flow cell images comprises: in each cycle, imaging, by an optical system, optical color signals emitted from the nucleotide reagents that are bound to the plurality of concatemer molecules.
114. The computer-implemented method of any one of the preceding claims, wherein the first plurality of flow cell images comprises optical color signals emitted from the nucleotide reagents that are bound to the plurality of concatemer molecules.
115. The computer-implemented method of any one of the preceding claims further comprising: removing a first sequencing read product from the first concatemer molecule and retaining the first concatemer molecule in the sample, and removing a second sequencing read product from the second concatemer molecule and retaining the second concatemer molecule in the sample.
116. The computer-implemented method of any one of the preceding claims further comprising: reiteratively sequencing the plurality of concatemers by repeating the following operations for at least once: generating the first plurality of flow cell images of a sample immobilized on a support by conducting one or more cycles of sequencing reactions thereby generating the first sequencing read product and the second sequencing product, the sample comprising a plurality of concatemer molecules therewithin, wherein a first concatemer molecule of the plurality of concatemer molecules corresponds to a first target RNA molecule of the sample, and a second concatemer molecule of the plurality of concatemer molecules corresponds to a second target RNA molecule of the sample, wherein the first plurality of flow cell images; and removing a first sequencing read product from the first concatemer molecule and retaining the first concatemer molecule in the sample, and removing a second sequencing read product from the second concatemer molecule and retaining the second concatemer molecule in the sample.
117. The computer-implemented method of any one of the preceding claims, wherein the first sequencing read product comprises some or all of: a first target barcode sequence in one or more tandem units of the first concatemer molecule; a first insert sequence in one or more tandem units of the first concatemer molecule; or their combinations.
118. The computer-implemented method of any one of the preceding claims further comprising: confirming presence of the first target RNA molecule, the second target RNA molecule, or both molecules in the sample based on the performed base calling of the second plurality of flow cell images at the base calling locations in the base calling template.
119. The computer-implemented method of any one of the preceding claims further comprising: generating, by the sequencing system, the second plurality of flow cell images of the sample immobilized on the support by conducting subsequent cycles of sequencing reactions after the one or more cycles.
120. The computer-implemented method of any one of the preceding claims, wherein generating the first plurality of flow cell images of the sample immobilized on the support comprises: sequencing at least the first concatemer inside the sample under a condition that inhibits sequencing the second concatemer.
121. The computer-implemented method of any one of the preceding claims, wherein sequencing at least the first concatemer inside the sample comprises: generating a plurality of first sequencing read products, and wherein the sequences of the first sequencing read products are aligned with a first target reference sequence to confirm presence of the first target RNA in the sample.
122. The computer-implemented method of any one of the preceding claims, wherein generating the first plurality of flow cell images of the sample immobilized on the support comprises: sequencing at least the second concatemer inside the sample under a condition that inhibits sequencing the first concatemer.
123. The computer-implemented method of any one of the preceding claims, wherein sequencing at least the second concatemer inside the cellular sample comprises: generating a plurality of second sequencing read products, and wherein sequences of the second sequencing read products are aligned with a second target reference sequence to confirm presence of the second target RNA in the sample.
Citation Information
Patent Citations
High resolution imaging fountain flow cytometry
US20050036139A1
Systems, methods and computer-accessible medium which provide microscopic images of at least one anatomical structure at a particular resolution
US20110218403A1
Methods and systems for processing polynucleotides
US20180112266A1
Flow cell device and use thereof
US20210121882A1
Systems And Methods For Applying Machine Learning to Analyze Microcopy Images in High-Throughput Systems
US20210303818A1
Cited By
Spatially resolved surface capture of nucleic acids
US12540350B2