Primary Analysis in Next Generation Sequencing
Patent Information
- Application Number
- JP2024534541
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-11
- Filing Date
- 2022-12-09
- Publication Date
- 2025-12-16
AI Technical Summary
Existing DNA sequencing technologies face challenges in accurately identifying cluster centers due to limitations in imager resolution, leading to improper sequence identification and increased processing time, especially in high-density sequencing environments.
The method involves capturing flow cell images at sub-pixel resolution using interpolation techniques and determining purity values for candidate cluster centers, allowing for the identification of actual cluster centers without increasing imager resolution, and utilizing specialized processors like FPGAs for real-time processing.
This approach enhances the accuracy of base calling by correctly assigning light signals to the correct DNA fragments, reducing processing time and costs by identifying more clusters with lower error rates, even in high-density sequencing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation-in-part of U.S. Patent Application No. 17 / 854,042, filed June 30, 2022, which is a continuation-in-part of U.S. Patent Application No. 17 / 547,602, filed December 10, 2021, which is a continuation-in-part of U.S. Patent Application No. 17 / 219,556, filed March 31, 2021, which claims priority to U.S. Provisional Patent Application No. 63 / 072,649, filed August 31, 2020. This application also claims priority to U.S. Provisional Patent Application No. 63 / 388,183, filed July 11, 2022. This application claims priority to U.S. Provisional Patent Application No. 63 / 349,421, filed June 6, 2022. The above-referenced patent applications are incorporated herein by reference in their entireties.
[0002] The present disclosure relates generally to image data analysis, and more particularly to identifying cluster or polony locations for performing base calling within digital images of flow cells during DNA sequencing. [Background technology]
[0003] Next-generation sequencing by synthesis using a flow cell can be used to identify DNA sequences. When single-stranded DNA fragments from a sequencing library are flooded across the flow cell, the fragments are typically randomly attached to the surface of the flow cell due to complementary oligomers or beads present on the surface of the flow cell. The DNA fragments are then subjected to an amplification process, resulting in copies of a given fragment forming clusters or polonies of denatured cloned nucleotide chains. In some embodiments, a single bead can comprise the cluster, and the bead can be attached to the flow cell at a random position.
[0004] To identify the sequence of a strand, the strand pair is reassembled one nucleotide base at a time. During each base assembly cycle, a mixture of single nucleotides, each attached to a fluorescent label (or tag) and a blocker, is flooded across the flow cell. The nucleotides are attached to complementary positions on the strands. A blocker is included so that only one base is attached to any given strand during a single cycle. The flow cell is exposed to excitation light, which excites the label to fluoresce. As cloned strands cluster together, the fluorescent signal of any one fragment is amplified by the signal from its cloned counterpart, and the resulting fluorescence of the cluster can be recorded by the imager. After the flow cell is imaged, the blocker is cleaved and washed from the flowed-in nucleotides, and more nucleotides are flooded over the flow cell, and the cycle is repeated. One or more images are recorded with each sequencing cycle.
[0005] A base-calling algorithm is applied to the recorded image to "read" the sequential signals from each cluster and convert the optical signals into the identification of the nucleotide base sequence attached to each fragment. Accurate base-calling requires accurate identification of the cluster centers to ensure that the sequential signals belong to the correct fragment. Summary of the Invention
[0006] Aspects of systems, apparatus, articles of manufacture, methods and / or computer program products, and / or combinations and sub-combinations thereof are provided herein that computationally improve the resolution of an imager beyond its physical resolution limit and / or provide more accurate source location within an image.
[0007] As a particular application of such, embodiments of methods and systems for identifying a set of base calling positions within a flow cell are described. These include capturing flow cell images after each sequencing cycle and identifying candidate cluster centers in at least one of the flow cell images. Intensity is determined for each candidate cluster center. Purity is determined for each candidate cluster center based on the intensity. In some embodiments, intensity and / or purity are determined at the sub-pixel level. Each candidate cluster center that has a purity higher than the purity of surrounding candidate cluster centers within a distance threshold is added to the set of base calling positions. The set of base calling positions may be referred to herein as a template.
[0008] In some embodiments, identifying candidate cluster centers comprises labeling each pixel of the flow cell image as a candidate cluster center.
[0009] In some aspects, identifying candidate cluster centers includes using a spot search algorithm to detect a set of potential cluster center locations and then identifying additional cluster locations around each potential cluster center location.
[0010] Further aspects, features, and advantages of the present disclosure, as well as the structure and operation of various embodiments of the present disclosure, are described in detail below with reference to the accompanying drawings.
[0011] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate aspects of the present disclosure and together with the description further serve to explain the principles of the present disclosure and to enable one skilled in the art(s) to make and use the aspects. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 shows a block diagram of a system for identifying cluster locations on a flow cell, according to some embodiments. [Figure 2]1 illustrates an exemplary flow cell image with candidate cluster centroids, according to some embodiments. [Figure 3] 1 is a flowchart illustrating a method for identifying positions to perform base calling, according to some embodiments. [Figure 4] FIG. 1 illustrates a block diagram of a computer that may be used to implement various aspects of the disclosure, according to some aspects. [Figure 5] 1 is a flowchart illustrating a method for identifying positions to perform base calling, according to some embodiments. [Figure 6] 1 is a schematic diagram illustrating an exemplary embodiment of a padlock probe. [Figure 7] FIG. 1 is a schematic diagram showing a workflow for generating circularized padlock probes in a cell, including generating first and second cDNAs from first and second target RNA molecules (respectively), and hybridizing first and second padlock probes to the first and second cDNA molecules (respectively) to generate first and second circularized padlock probes (respectively). [Figure 8] FIG. 1 is a schematic diagram showing an intracellular rolling circle and sequencing workflow that involves generating first and second concatemers by performing rolling circle amplification using first and second covalently closed circular molecules (respectively). [Figure 9] FIG. 1 is a schematic diagram showing an exemplary workflow for sequencing concatemers generated within cells. [Figure 10] FIG. 1 is a schematic diagram showing an exemplary workflow for sequencing concatemers generated within cells. [Figure 11] FIG. 1 is a schematic diagram showing an exemplary workflow for sequencing concatemers generated within cells. [Figure 12] FIG. 1 is a schematic diagram showing an exemplary workflow for sequencing concatemers generated within cells. [Figure 13]FIG. 1 is a schematic diagram showing a workflow for generating circularization padlock probes, including generating first and second cDNAs from first and second target RNA molecules (respectively), and hybridizing the first and second padlock probes to the first and second cDNA molecules (respectively) to generate first and second circularization padlock probes (respectively). [Figure 14] FIG. 1 is a schematic diagram showing a rolling circle and sequencing workflow that involves generating first and second concatemers by performing rolling circle amplification using first and second covalently closed circular molecules (respectively). [Figure 15] Schematic of an exemplary low-binding support comprising a glass substrate and alternating layers of hydrophilic coating, the alternating layers of hydrophilic coating covalently or non-covalently adhered to the glass and further comprising chemically reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides). [Figure 16] 1A-1D are schematic diagrams of various exemplary configurations of multivalent molecules. [Figure 17] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a generic core attached to multiple nucleotide arms. [Figure 18] FIG. 1 is a schematic diagram of an exemplary multivalent molecule comprising a dendrimer core attached to multiple nucleotide arms. [Figure 19] 1 shows a schematic diagram of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, the nucleotide arms comprising biotin, spacers, linkers, and nucleotide units. [Figure 20] FIG. 1 is a schematic diagram of an exemplary nucleotide arm comprising a core attachment moiety, a spacer, a linker, and a nucleotide unit. [Figure 21] Shown are the chemical structures of an exemplary spacer (top) and various exemplary linkers (bottom), including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker. [Figure 22]1 shows the chemical structures of various exemplary linkers, including linkers 1-9. [Figure 23A] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 23B] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 23C] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 23D] 1 shows the chemical structures of various exemplary linkers linked / attached to nucleotide units. [Figure 24] 1 shows the chemical structure of an exemplary nucleotide arm. [Figure 25] 1 is a schematic diagram of a G-quadruplex (e.g., a G-quadruplex). [Figure 26] FIG. 1 is a schematic diagram of an exemplary intramolecular G-quadruplex structure. [Figure 27A] 1 illustrates an exemplary support having a plurality of tiles for immobilizing a cell sample, according to some embodiments. [Figure 27B] FIG. 1 is a schematic diagram showing flow cell images at multiple axial positions with base calling positions at some of the axial positions and base calling positions between two predetermined axial positions, according to some embodiments. [Figure 28] 1 illustrates an exemplary data file containing data of cellular features and in situ sequencing results, and a reconstructed image combining the in situ sequencing results and cellular features, according to some embodiments. [Figure 29A] 29A is a schematic diagram of a cell (top) with its major structural elements. The corresponding ASCI II markers for each structural element are shown. Such ASCI II markers can be used in the data files disclosed herein, according to some embodiments, such that 3D cell or tissue morphology from in situ sequencing can be reconstructed from the data files. Figure 29A also shows an exemplary list (bottom) of decoders or indicators that can be used in the data files, according to some embodiments. [Figure 29B] 29B illustrates an exemplary data file utilizing the indicators listed in FIG. 29A, according to some embodiments.
[0013] In the drawings, like reference numbers generally indicate the same or similar elements. Additionally, the left-most digit(s) of a reference number generally identifies the drawing in which the reference number first appears. DETAILED DESCRIPTION OF THE INVENTION
[0014] Aspects of systems, devices, products, methods, and / or computer program products, and / or combinations and subcombinations thereof, are provided herein that computationally improve the resolution of an imager beyond its physical resolution limit and / or provide more accurate source location within an image. The image processing techniques described herein are particularly useful for base calling in next-generation sequencing, and base calling is used as the primary example herein to illustrate the application of these techniques. However, such image analysis techniques may also be particularly useful in other applications in which spot detection and / or CCD imaging is used. For example, identifying the actual center (e.g., source location) of a perceived light signal has utility in many other fields, such as location detection and tracking, astrophotography, heat mapping, and the like. In addition, techniques such as those described herein may be useful in any other application that benefits from computationally increasing resolution once the physical resolution limit of the imager is reached.
[0015] In DNA sequencing, identifying the centers of clusters or polonies (often formed on beads) is sometimes called primary analysis. Primary analysis involves creating a template for the flow cell. The template contains the estimated locations of all detected clusters in a common coordinate system. The template is generated by identifying cluster locations in all images in the first few flows of the sequencing process. The images can be aligned across all images to provide a common coordinate system. Cluster locations from different images can be merged based on proximity in the coordinate system. Once the template is generated, all further images are aligned to it, and sequencing is performed based on the cluster locations in the template.
[0016] Various algorithms exist for identifying cluster centers in an image. These existing algorithms have several drawbacks. As discussed above, cluster centers may appear merged if they are close to each other. The proximity may be due to precision or alignment issues. Thus, different clusters may be treated as a single cluster, resulting in improper identification of sequences or even missing sequences.
[0017] Additionally, the algorithm may need to find clusters across several images to identify cluster locations for the template, which may require excessive processing time.
[0018] 1 shows a block diagram of a system 100 for identifying cluster locations on a flow cell, according to one embodiment. The system 100 has a sequencing system 110, which may include a flow cell 112, a sequencer 114, an imager 116, data storage 122, and a user interface 124. The sequencing system 110 may be connected to a cloud 130. The sequencing system 110 may include one or more of a dedicated processor 118, a field programmable gate array(s) (FPGA(s)) 120, and a computer system 126.
[0019] In some embodiments, the flow cell 112 is configured to capture DNA fragments and form a DNA sequence for base calling on the flow cell. The sequencer 114 may be configured to flow a nucleotide mixture over the flow cell 112, cleave blockers from the nucleotides during the flowing step, and perform other steps to form a DNA sequence on the flow cell 112. The nucleotides may have fluorescent elements that emit light or energy at wavelengths indicative of the type of nucleotide. Each type of fluorescent element may correspond to a specific nucleotide base (e.g., A, G, C, T). The fluorescent elements may emit light in visible wavelengths.
[0020] For example, each nucleotide base may be assigned a color. Adenine may be, for example, red, cytosine may be blue, guanine may be green, and thymine may be yellow. The color or wavelength of the fluorescent moiety of each nucleotide may be selected so that the nucleotides are distinguishable from one another based on the wavelength of light emitted by the fluorescent moiety.
[0021] Imager 116 can be configured to capture an image of flow cell 112 after each flowing step. In one embodiment, imager 116 is a camera configured to capture digital images, such as a CMOS or CCD camera. The camera can be configured to capture images at the wavelength of the fluorescent moiety attached to the nucleotide.
[0022] The resolution of the imager 116 controls the level of detail in the flow cell image, including pixel size. In existing systems, this resolution is very important because it controls the accuracy with which the spot-finding algorithm identifies cluster centers. One way to increase the accuracy of spot-finding is to improve the resolution of the imager 116 or to improve the processing performed on the images captured by the imager 116. The methods described herein can detect cluster centers in pixels other than those detected by the spot-finding algorithm. These methods allow for improved cluster center detection accuracy without increasing the resolution of the imager 116. The imager resolution may be smaller than existing systems with comparable performance, which may reduce the cost of the sequencing system 110.
[0023] In one aspect, images of the flow cell can be captured in groups, with each image in the group taken at wavelengths or spectra that correspond to or include only one of the fluorescent elements, or in another aspect, the images can be captured as a single image that captures all wavelengths of the fluorescent elements.
[0024] The sequencing system 100 may be configured to identify cluster locations on the flow cell 112 based on the flow cell image. The processing to identify the clusters may be performed by a dedicated processor 118, FPGA(s) 120, a computing system 126, or a combination thereof. Identifying or determining the cluster locations may involve performing traditional cluster detection in combination with the cluster detection methods more specifically described herein.
[0025] General-purpose processors provide an interface for running a variety of programs within an operating system such as Windows™ or Linux™, which typically offer great flexibility to the user.
[0026] In some embodiments, dedicated processors 118 may be configured to perform the steps of the cluster detection methods described herein. These may not be general-purpose processors, but rather custom processors with specific hardware or instructions for performing those steps. Dedicated processors execute specific software directly without an operating system. The lack of an operating system reduces overhead at the expense of flexibility in what the processor can execute. Dedicated processors may utilize custom programming languages that may be designed to operate more efficiently than software running on general-purpose processors. This may increase the speed at which steps are executed and enable real-time processing.
[0027] In some embodiments, the FPGA(s) 120 can be configured to perform the steps of the cluster detection method described herein. The FPGA is programmed as hardware that performs only specific tasks. Special programming languages may be used to translate software steps into hardware components. Once the FPGA is programmed, the hardware directly processes the provided digital data without executing software. Instead, the FPGA uses logic gates and registers to process the digital data. Without the overhead required for an operating system, FPGAs generally process data faster than general-purpose processors. As with special-purpose processors, this comes at the expense of flexibility.
[0028] Also, without software overhead, FPGAs can run faster than dedicated processors, although this depends on the exact processing being performed and the particular FPGA and dedicated processor.
[0029] A group of FPGA(s) 120 may be configured to perform steps in parallel. For example, several FPGA(s) 120 may be configured to perform a processing step on an image, a set of images, or cluster locations within one or more images. Each FPGA(s) 120 may perform its own portion of the processing step simultaneously, reducing the time required to process the data. This allows the processing step to be completed in real time. Further discussion of the use of FPGAs is provided below.
[0030] By performing processing steps in real time, the system may use less memory because data may be processed as it is received, which is an improvement over conventional systems that may need to store data before it can be processed, which may require more memory or access to computer systems located in the cloud 130.
[0031] In some embodiments, data storage 122 is used to store information used in identifying cluster locations. This information may include the image itself or information derived from the image captured by imager 116. DNA sequences determined from base calling may be stored in data storage 122. Parameters identifying cluster locations may also be stored in data storage 122.
[0032] The user interface 124 may be used by a user to operate the sequencing system or to access data stored in the data storage 122 or the computer system 126 .
[0033] The computer system 126 may control the general operation of the sequencing system and may be coupled to the user interface 124. It may also perform the steps of cluster location identification and base calling. In some embodiments, the computer system 126 is computer system 400, as described in more detail in FIG. 4. The computer system 126 may store information related to the operation of the sequencing system 110, such as configuration information, instructions for operating the sequencing system 110, or user information. The computer system 126 may be configured to pass information between the sequencing system 110 and the cloud 130.
[0034] As discussed above, sequencing system 110 may have a dedicated processor 118, FPGA(s) 120, or computer system 126. The sequencing system may use one, two, or all of these elements to accomplish the necessary processing described above. In some embodiments, when these elements are present together, processing tasks are divided between them. For example, FPGA(s) 120 may be used to perform the cluster center detection methods described herein, while computer system 126 may perform other processing functions for sequencing system 110. Those skilled in the art will understand that various combinations of these elements enable various system embodiments that balance efficiency and speed of processing with cost of processing elements.
[0035] Cloud 130 may be a network, remote storage, or some other remote computing system separate from sequencing system 110. Connection to cloud 130 may allow access to data stored outside of sequencing system 110 or may allow updates to software within sequencing system 110.
[0036] FIG. 3 is a flowchart illustrating a method 300 for identifying actual cluster center locations for performing base calling. Cluster centers in a flow cell image are locations in the image that correspond to the locations of cloned clusters on the physical flow cell. The wavelength of the optical signal detected at the cluster center correlates to the nucleotide base added to the fragment on the flow cell at that location. For the DNA sequence to be correctly determined, successively detected optical signals must consistently be attributed to the correct DNA fragment. Thus, accurately identifying the location of a cluster center improves the accuracy of base calling for that fragment. In some embodiments, once the actual cluster center is identified, such location can be mapped onto a template for use in subsequent base calling cycles using the same flow cell. Method 300 can be performed by a dedicated processor 118, FPGA(s) 120, or computer system 126.
[0037] In step 310, a flow cell image is captured. As discussed above, the flow cell image may be captured by imager 116. Step 310 may involve capturing images one at a time to be processed by the following steps, or may involve capturing a set of images for simultaneous processing. In examples where a set of images is captured, each image in the set of images may correspond to a different detected wavelength. For example, considering the color designations associated with nucleotides, an image set may include four images, each corresponding to a signal captured at one of the red, blue, green, and yellow wavelengths. In examples where a single image is captured, the image may include all detected wavelengths of interest. Each image or set of images may be captured for a single flow step on the flow cell. In some embodiments, the flow cell images are captured with reference to a coordinate system.
[0038] Depending on the sample(s) immobilized on the support (e.g., flow cell), the flow cell image may include a single or multiple axial positions along an axial axis orthogonal to the image plane of the flow cell image. In particular, for three-dimensional samples, such as cells, tissues, or other in situ samples, the flow cell image may include multiple z-levels (i.e., axial positions) to provide 3D coverage of the entire sample(s). The axial axis may extend from the objective lens of the optical system disclosed herein to the support, e.g., the flow cell. The axial axis may be orthogonal to the image plane of the flow cell image. Each axial position(s) of the flow cell image may be separated from adjacent axial positions by a predetermined distance, e.g., about 0.1 um to about 15 um. Each axial position of the flow cell image may be separated from adjacent level(s) by 0.5 um to 10 um. At each axial position, the flow cell image may be acquired from one or more sequencing cycles and / or one or more channels. Each flow cell image may include at least a portion of one or more tiles or subtiles of the flow cell within its field of view. Figure 27A shows a portion of a flow cell 2712 having multiple tiles 2710. The image plane is defined by the x-axis and y-axis, and the axial axis (i.e., the z-axis) is orthogonal to the xy-plane. The flow cell image, sample, and axial axis are described in a Cartesian coordinate system as shown in Figures 27A-27B, although any other coordinate system may be used to define spatial locations and relationships herein. Other coordinate systems may include, but are not limited to, polar, cylindrical, or spherical coordinate systems.
[0039] 2 shows a schematic diagram of a flow cell image 200 in which signals from clusters are present. Flow cell image 200 is made up of pixels 210, such as pixels 210A, 210B, and 210C. During step 310, the imager records the light signals received from the flow cell after, for example, excitation of fluorescent elements bound to fragments on the flow cell, such fragments being located within a cloned cluster of fragments.
[0040] In step 320, the locations of potential cluster centers within the flow cell image are identified. For example, in some embodiments, the optical signal imaged in step 310 can be input to a spot-finding algorithm, such that the spot-finding algorithm outputs a set of potential cluster centers. In some embodiments, potential cluster centers can be identified using only a single flow cell image (e.g., one image including all wavelengths of interest). In some other embodiments, potential cluster centers can be identified from a set of images from a single flow cycle on the flow cell (e.g., one image at each wavelength of interest). Using only a flow cell image or set of flow cell images from a single flow cycle advantageously reduces the amount of processing time because the spot-finding algorithm does not need to wait for additional images from future flow cycles to be acquired.
[0041] In yet other embodiments, the spot-finding algorithm can be applied to images from two or more sequencing cycles, and potential cluster centers can be searched for using some combination of those images. For example, potential cluster centers can be identified by the presence of spots occupying the same location in images from two or more sequencing cycles.
[0042] Potential cluster center locations identified by the spot-finding algorithm are indicated by "Xs" in FIG. 2 , such as potential cluster center locations 220A, 220B, and 220C. Due to the random nature of fragment attachment to the flow cell, some of the cloned clusters may be close to each other, while other clusters may be further apart or may be isolated. As a result, some "Xs" in FIG. 2 are located closer to each other than others. In addition, some pixels may be identified as containing potential cluster centers, while other pixels are not. For example, pixel 210A may be identified as containing potential cluster center location 225A, while pixel 210B is not initially identified as containing potential cluster center location 220A.
[0043] In some embodiments, the spot finding algorithm may identify potential cluster center locations 220 at sub-pixel resolution by interpolating across pixels 210. For example, potential cluster center location 225A is located on the lower right side of pixel 210A rather than at the center of pixel 210A. Other potential cluster center locations may be located in different regions of their respective pixels 210. For example, potential cluster center location 225B is located on the upper right side of the pixel, and potential cluster center location 225C is located on the upper left side of the pixel. The interpolation may be performed by an interpolation function.
[0044] In some embodiments, the interpolation function is a Gaussian interpolation function known to those skilled in the art. Sub-pixel resolution may allow potential cluster locations to be determined at, for example, one-tenth pixel resolution, although other resolutions are also contemplated. In embodiments, for example, the resolution may be one-quarter pixel resolution, one-half pixel resolution, etc. The interpolation function may be configured to determine this resolution.
[0045] An interpolation function may be used to fit the light intensity at one or more pixels 210. This interpolation allows for the identification of sub-pixel locations. The interpolation function may be applied across a set of pixels 210 that include potential cluster center locations 220. In one aspect, the interpolation function may be fitted to a pixel 210 that has a potential cluster center location 220 within it, and to surrounding pixels 210 that touch the edge of that pixel 210.
[0046] In some embodiments, the interpolation function may be determined at several points in the image. The resolution determines the number of points located at each pixel 210. For example, if the resolution is one-tenth of a pixel, then nine points are calculated across the pixel 210 and one point at each edge along lines perpendicular to the pixel edges, dividing the pixel 210 into ten portions. In some embodiments, the interpolation function is calculated at each point, and the difference between the interpolation function and the pixel intensity at each point is determined. The center of the interpolation function is shifted to minimize the difference between the interpolation function and the intensity of each pixel 210. This sub-pixel interpolation allows the system to achieve higher resolution with lower resolution imagers, reducing the cost and / or complexity of the system.
[0047] In some embodiments, the interpolation may be performed on a 5x5 grid. The grid may be centered on pixel centers 210 within pixels having potential cluster center locations.
[0048] Some embodiments of step 320 use a spot-finding algorithm to identify potential cluster center locations, while other embodiments of step 320 first identify each pixel in the captured flow cell image as a potential cluster center location. For example, in FIG. 2, all pixels 210 may be identified as potential cluster center locations 220. This approach eliminates the need for a spot-finding algorithm, which may simplify the type of processing required to perform method 300. This approach may be advantageous when massively parallel processing is available, because each potential cluster center location may be processed in parallel. This may reduce processing time, although potentially at the cost of additional hardware, such as an increased dedicated processor 118 or FPGA(s) 120. In some embodiments, an interpolation function may then be used as described above to identify intensities at sub-pixel resolution throughout the flow cell image.
[0049] As discussed above, cluster centers identify locations in an image, such as pixels 210, that correspond to the location of cloned clusters on the physical flow cell. Potential cluster center locations 220 are locations in an image where light of one or more wavelengths is detected by the imager. In some cases, while the physical location of a cluster corresponds to one set of pixels 210, it is possible for the light signal from that cluster to overflow into additional pixels 210 adjacent to that set, for example, due to saturation (also referred to as "blooming") of the corresponding sensor in the camera. In addition, when clusters are closely spaced, the light signals from those clusters may overlap even if the clusters themselves do not. Identifying true cluster centers allows the detected signals to be attributed to the correct DNA fragments, thus improving the accuracy of the base-calling algorithm.
[0050] Thus, in step 325, additional cluster signal locations 225 are identified around each potential cluster center location 220. These are indicated by "+" in Figure 2, e.g., 225A, 225B, and 225C. These additional cluster signal locations 225 correspond to other locations within the flow cell that may constitute cluster centers, instead of or in addition to locations already identified as potential cluster center locations 220.
[0051] In some embodiments, additional cluster locations 225 are positioned around potential cluster center locations 220. In some embodiments, these additional cluster locations 225 are not initially identified by the spot search algorithm, but are positioned in a pattern around each potential cluster center location 220 identified by the spot search algorithm. The additional cluster locations 225 do not represent actual detected cluster centers, but rather represent potential locations for checking cluster centers that might otherwise go undetected. Such cluster centers may not be detected due to mixing between signals from nearby cluster centers, errors in the spot search algorithm, or other effects.
[0052] As an example, additional cluster location 225A may be placed at pixel 210A based on the location of potential cluster center location 220A. In this context, this may mean that additional cluster location 225A is placed a pixel width away from potential cluster center location 220A. Other additional cluster locations 225 are also placed around potential cluster center location 220A. It should be understood that additional cluster locations 225 are not placed where no pixel 210 exists, such as when potential cluster center location 220 is near an edge of flow cell image 200.
[0053] In some embodiments, the additional cluster signal locations 225 are arranged in a grid centered on the potential cluster center location 220. The additional cluster locations 225 may be separated from each other and from the potential cluster center location 220 by a pixel width. In some embodiments, the additional cluster locations 225 and the potential cluster center location 220 form a square grid. The grid may have an area of 5 pixels by 5 pixels, 9 by 9, 15 by 15, or other dimensions.
[0054] In some embodiments, potential cluster center locations 220 may be close enough to one another to overlap corresponding grids of additional cluster locations 225. This may result in the same pixel 210 including additional cluster locations 225 from more than one potential cluster center location 220, and more than one additional cluster location 225 being attributed to the same pixel 210. For example, pixel 210C includes additional cluster location 225B (identified based on potential cluster center location 220B) as well as additional cluster location 225C (identified based on potential cluster center location 220C).
[0055] In some embodiments, if additional cluster locations 225 within the same pixel 210 are sufficiently close to one another, one of the additional cluster locations 225 is discarded and the other is used to represent both. Two additional cluster locations 225 may be considered sufficiently close to one another for such processing if they are within, for example, but not limited to, two-tenths of a pixel, one-tenth of a pixel, or some other sub-pixel distance.
[0056] Those skilled in the art will understand that if all pixel locations are identified as potential cluster center locations 220 in step 320, then step 325 may be skipped because no other pixels remain in the flow cell image to consider in addition to the identified potential cluster center locations.
[0057] In some embodiments, the potential cluster center locations 220 (and surrounding additional cluster locations 225, if identified) together constitute a set of all candidate cluster centers, which can be processed to identify actual cluster centers for each captured flow cell image.
[0058] Thus, once the potential cluster center locations 220 and their surrounding grid of additional cluster locations 225 (i.e., candidate cluster centers) have been identified, they can be used as a starting point for determining the actual locations of the cluster centers, which may or may not be the same as the originally identified potential cluster center locations 220.
[0059] 3, a purity value for each candidate cluster center on each captured flow cell image is determined in step 330. The purity value may be determined based on the wavelength of the fluorescent element attached to the nucleotide and the intensity of the pixel in the captured flow cell image.
[0060] At each candidate cluster center, the pixel intensity is the combination of energy or light emitted across the spectral bandwidth of the imager. In some embodiments, the amount of energy or light corresponding to the fluorescence spectral bandwidth of each nucleotide base can be determined. The purity of each signal corresponding to a particular nucleotide base can be determined as the ratio of the amount of energy for one nucleotide base signal to the total amount of energy for each other nucleotide base signal (e.g., the purity of a "red" signal can be determined based on the relative intensity of the red wavelength detected for that pixel or subpixel compared to each of the blue, green, and yellow wavelengths detected). The overall purity of a pixel can be the maximum ratio, minimum ratio, average ratio, or median ratio. The calculated purity can then be assigned to that pixel or subpixel.
[0061] As described above, a set of flow cell images can be captured for a single sequencing cycle. Each image in the set is captured at a different wavelength, with each wavelength corresponding to one of the fluorescent elements attached to the nucleotides. The purity of a given cluster center across the set of images can be the highest, lowest, median, or average purity from the set of purities for the set of images.
[0062] In some embodiments, the purity of a candidate cluster center can be determined as the ratio of the second highest intensity or energy from a pixel's wavelength minus the highest intensity or energy from that pixel's wavelength. A threshold can be set for what constitutes high purity or low purity. For example, the highest purity may be 1, and low purity pixels may have a purity value near zero. A threshold can be set somewhere in between.
[0063] In some embodiments, the ratio of two intensities can be corrected by adding an offset to both the second-highest intensity and the highest intensity. The offset can provide improved accuracy of the quality score. For example, in some cases, two ratio intensities may differ by only a small amount of the absolute maximum intensity (which is also a high percentage of the highest intensity). As a specific, non-limiting example, the highest intensity may be 10, and the lowest intensity may be the maximum possible intensity of 1000. In this case, the ratio would be 0.1, resulting in a purity of 0.9. If nothing more, this could be read as a high quality score. This contrasts with, for example, an intensity of 500 for the highest intensity and 490 for the second-highest intensity. This example has nearly the same absolute difference, but the ratio is close to 1 and the purity is close to zero. In the first case, the purity is misleading because the low overall intensity suggests the absence of polony. In the second case, the purity is more accurate, indicating that the pixel displays intensities or energies from two different cluster centers located nearby.
[0064] An offset can be a value added to the strength of the ratio to solve such problems. For example, if the offset is 10% of the maximum amplitude, in the example above, the offset would be 100 and the first ratio would be 101 to 110, which is much closer to 1 and results in a purity near zero that accurately reflects the small delta between the two wavelength intensities. For the second ratio, the ratio would be 600 to 590, which is still closer to 1 and also results in a purity near zero.
[0065] As another example incorporating an offset, if the highest intensity is 800 and the lowest intensity is 1, the purity without the offset is close to 1 because the ratio is near zero. If the offset is again 100, the ratio becomes 101 for 900. This slightly reduces the purity from 1 to approximately 0.89. While this may reduce the purity, the calculated purity is still high. The offset value may be set to reduce this effect. For example, the offset in another case could be 10. Using the previous example where the highest intensity is 10 and the lowest intensity is 1, the purity would be 1-10 / 11, or approximately 0.09, which accurately reflects the small difference between the intensities. In the example where the highest intensity is 600 and the lowest intensity is 590, the purity would be 1-600 / 610, or approximately 0.016, which also reflects the small difference between the intensities. In the example where the highest intensity is 800 and the lowest intensity is 1, the purity is 1-11 / 810, or about 0.99, which is a much smaller drop in purity and reflects the large difference between the intensities.
[0066] In step 340, actual cluster centers are identified based on the purity values calculated for the candidate cluster centers for the flow cell. The actual cluster centers may be a subset of the candidate cluster locations identified in steps 320 and 325.
[0067] In some embodiments, the actual cluster center is identified by comparing the purity of each candidate cluster center across the entire flow cell image with nearby candidate cluster centers within the same image. In some embodiments, when considering two candidate cluster centers being compared, the candidate cluster center with the higher purity is retained. In some embodiments, the candidate cluster center is only compared with other candidate cluster centers within a certain distance. For example, this distance threshold can be based on pixel size and the size of a cluster of a given nucleotide.
[0068] For example, if the average size of a cluster is four pixels wide / height, the distance threshold could be two pixels wide / height, since any candidate cluster centers within two pixels wide / height of each other are likely to belong to the same cluster or have higher intensity (indicating that the candidate cluster center is actually on the edge of two separate clusters).
[0069] In some embodiments where purity is calculated over multiple flow cell cycles, determining that a candidate cluster center consistently has higher purity than surrounding candidate cluster centers across multiple flow cell images may further increase the likelihood that the location is an actual cluster center. Lower purity may indicate that the signal detected at the candidate cluster center is not an actual cluster center, but rather noise, a mixture of other signals, or some other phenomenon.
[0070] In step 350, the actual cluster centers are used to perform base calling on the flow cell image. For example, the wavelength detected from the actual cluster center may be determined, which correlates to a particular nucleotide base (e.g., A, G, C, T). That nucleotide base is then recorded as being added to the sequence corresponding to the actual cluster center.
[0071] By successive repetitions of the sequencing cycle and fluorescence wavelength discrimination at the actual cluster centers in successive flow cell images, the sequence of the DNA fragment corresponding to each actual cluster center on the flow cell can be constructed.
[0072] In some embodiments, a template is formed from the actual cluster centers identified for a single sequencing cycle. The actual cluster locations in the template can then be used to identify locations to perform base calling in images from subsequent sequencing cycles.
[0073] Flow cell images captured at different sequencing cycles may have alignment issues due to shifts in the position of the flow cell or imager between sequencing cycles. Therefore, in some embodiments, step 350 may include an alignment step to properly align consecutive images. This ensures that actual clusters accurately map the centers of the templates to the same locations on each flow cell image, thus improving the accuracy of base calling.
[0074] In some embodiments where templates are used to identify actual cluster centers in subsequent images, only data corresponding to relevant locations in those subsequent images need to be maintained and / or processed. This reduction in the amount of data processed increases the speed and / or efficiency of processing, such that accurate results can be obtained more quickly than with legacy systems. Additionally, the reduction in the amount of data stored reduces the amount of storage required for the sequencer, thus reducing the amount of resources required and / or costs.
[0075] In addition, some legacy systems require comparing different images to each other to identify cluster locations. This comparison may involve applying a spot finding algorithm to images from multiple flow cell cycles and then comparing the spot finding results across images. This may require storing images or spot finding results for each of multiple sequencing cycles. Because images do not need to be compared directly, method 300 may improve processing and storage efficiency of cluster finding. Instead, images may be processed in real time, and only purity information and / or final template locations need to be stored.
[0076] The sequencing cycles and image creation processes often run faster than the spot-finding and base-calling programs that analyze the images. This difference in execution time may require storing the flow cell image after each sequencing cycle, or delaying the sequencing cycle while waiting for some or all of the image analysis processes to complete.
[0077] The use of FPGAs can increase processing speed without sacrificing accuracy. Implementing some or all of the processes described herein on an FPGA can reduce processor overhead and enable parallel processing on the FPGA. For example, each possible cluster location may be processed by a different FPGA, or by a single FPGA configured to process all possible cluster locations simultaneously in parallel. When properly implemented, this can enable real-time processing. Real-time processing has the advantage that images can be processed as they are generated. The FPGA is ready to process the next image by the time the sequencing system prepares the flow cell. The sequencing system does not need to wait for post-processing, and the entire process of primary analysis can be completed in a matter of seconds. Additionally, because the entire image is processed as it is received, the only information that needs to be stored is the data needed to perform base calling. Instead of storing the entire image, only the purity or intensity of specific pixels needs to be stored. This significantly reduces the need for data storage in the sequencing system or remote storage of images.
[0078] In some embodiments, the entire process, including image registration, intensity extraction, purity calculation, base calling, and other steps, is performed by an FPGA, which may provide the most compact implementation and may offer the speed and throughput required for real-time processing.
[0079] In some embodiments, processing responsibilities are shared, such as between the FPGA and an associated computer system. For example, in some embodiments, the FPGA may handle image registration, intensity extraction, and purity calculations. The FPGA then passes information to the computer system for base calling. This approach balances the load between the FPGA and computer system resources, including scaling down communication between the FPGA and the computer system. It also provides flexibility for software on the computer system that handles base calling with rapid algorithm tune-up capabilities. Such an approach may provide real-time processing.
[0080] Those skilled in the art will recognize that different configurations of FPGAs, special purpose processors, and computer systems can be used to perform the various steps. The selection of a given configuration may be based on the flow cell image size, imager resolution, number of images to process, desired accuracy, and required speed. Implementation and hardware costs of the FPGAs, special purpose processors, and computer systems may also influence the selection of a configuration.
[0081] As a non-limiting example of a performance comparison between existing methods and embodiments of the methods described herein, tests were performed on two exemplary flow cells: one with a low density of clusters and one with a high density of clusters. For comparison, tests were performed using each method to target a specific average error rate of false positives on the identified clusters.
[0082] For the low-density flow cell, the average error rate was 0.3%. The existing method identified approximately 78,000 cluster centers, while the method described herein identified approximately 98,000 cluster centers. For the high-density flow cell, the average error rate was 1.1%. The existing method identified approximately 63,000 cluster centers, while the method described herein identified approximately 170,000 clusters.
[0083] The results suggest that the methods described herein effectively identify more clusters than existing methods. Furthermore, even at the same error rate, when the density of clusters on the flow cell increases, the methods disclosed herein perform even better, identifying nearly three times as many clusters. In some embodiments, this may allow the flow cell to run at higher densities without the performance loss typically experienced with existing methods.
[0084] In situ sequencing In some embodiments, the method, system, and medium for determining base calling positions can be used, for example, in in situ sequencing of 3D volumetric samples. In the in situ sequencing described herein, RNA is not extracted from the cell sample, and there is no need to track and map sequencing information back to an image of the cell sample. Rather, RNA is retained in the cell sample to enable direct imaging of the spatial location of target RNA within the cell. In addition, RNA within the cell sample is not fragmented, and target RNA enrichment is not required. The use of target-specific and / or random sequence reverse transcription primers allows for the detection of both polyA RNA and non-polyA RNA in either uniplex or multiplex mode.
[0085] In some embodiments, in situ sequencing herein involves repeatedly sequencing the same region of a template molecule (e.g., a concatemeric molecule) multiple times. By performing repetitive sequencing cycles (e.g., fewer than 30 or 40 cycles compared to over 100 cycles), the RNA content of a cell sample can be discovered. Compared to long-read sequencing workflows, the short repetitive sequencing cycles described herein use reduced amounts of sequencing reagents, which reduces costs and saves time. Methods for performing short (e.g., 2-40 cycles) repetitive sequencing cycles have many applications, including, but not limited to, detecting specific RNAs of interest, mutant RNA sequences, splice variants, and their abundance levels.
[0086] Concatemers herein contain a cDNA of interest, a universal sequencing primer binding site, and tandem repeat units of a target barcode sequence. The concatemers are sequenced within a cell sample, with multiple rounds of short-read sequencing, with multiple sequencing cycles (e.g., 2 to 40 cycles) performed for each round. The entire length of the target barcode and cDNA region is not sequenced. Instead, at least a portion of the target barcode region is sequenced repeatedly. In some embodiments, the cDNA region does not need to be sequenced. In some embodiments, the target barcode and a portion of the cDNA region are sequenced repeatedly. In some embodiments, the entire length of the cDNA region does not need to be sequenced. In some embodiments, it is not necessary to assemble sequencing reads or obtain the full-length sequence of the cDNA of interest. The redundant sequencing information obtained from the short sequencing reads herein can advantageously eliminate the need to sequence the complementary strands of the concatemers. Thus, pairwise sequencing is not required.
[0087] In addition, in some embodiments, a short portion of the cDNA region in the concatemer is re-sequenced at least once (e.g., repeated sequencing) from the same starting position to generate overlapping sequencing reads that can be aligned to a reference sequence.For example, the same portion of the concatemer molecule can be sequenced at least twice, three times, four times, five times, or up to 50 times.The starting sequencing position can be any position in the concatemer, and is determined by a sequencing primer designed to anneal to a selected position within the concatemer.The short repeated sequencing reads increase the redundancy of the sequencing information for each base in the cDNA region.By repeatedly sequencing one strand of the concatemer template molecule, sufficient base coverage can be obtained to reveal the presence of target RNA in a cell sample, thereby eliminating the need for pairwise sequencing of complementary strands.
[0088] A concatemer template molecule can contain multiple sequencing primer binding sites along the same concatemer molecule, which can be used to generate multiple usable sequencing reads for increased sequencing depth. Together, repeated sequencing of one strand of a concatemer template can increase sequencing base coverage and sequencing depth compared to sequencing a one-copy template molecule.
[0089] The in situ sequencing described herein can be carried out in uniplex mode or multiplex mode.By using different reverse transcription primers, different target-specific padlock probes, and universal sequencing primers, two or more different target RNAs can be simultaneously detected and imaged in a cell sample.For example, the presence of housekeeping RNA and at least one target RNA in a cell sample can be simultaneously detected and imaged using any of the short repeat read sequencing methods described herein.
[0090] 3D Base Calling Template In some embodiments, disclosed herein is a method for generating a base calling template for a cell sample(s) in in situ sequencing.The base calling template can include base calling positions from a single 2D plane, multiple 2D planes, or 3D space.The base calling template can be generated using a flow cell image in a single 2D plane, multiple 2D planes, or 3D space.The base calling template can be used to align the flow cell image acquired in a single 2D plane, multiple 2D planes, or 3D space with a common coordinate system.
[0091] A conventional 2D flow cell image can show clusters or polonies from a flattened sample, and base calling can be performed using their corresponding image intensities. The optical system can be adjusted to image the clusters or polonies while they are in focus ("in focus" herein can include substantially in focus and / or within the depth of field to acquire corresponding flow cell image(s). However, in situ samples such as cells or tissues can have thicknesses along the axial or z-direction that cannot be focused in a single 2D image. To cover clusters or polonies at different axial positions, multiple 2D flow cell image stacks may be required. Flow cell image stacks can introduce interference, such as from out-of-focus polonies and background signals from cellular components. For example, a polonie located at a first axial position can appear in the first flow cell image and generate a blob of signal in a second flow cell image taken at an adjacent axial position of the first flow cell image, where the polonie is out of focus. Signal blobs can interfere with the intensity of polony at or near the same xy location in the second flow cell image, thus reducing the accuracy and reliability of base calling for the in situ sample. Furthermore, the same polony may be included in both polony maps corresponding to two adjacent axial locations, thereby causing inaccurate base calling results. Therefore, there is a need to generate accurate and reliable base calling templates for polony or clusters from 3D volumetric samples so that they can be used for 3D base calling.
[0092] The techniques disclosed herein advantageously generate 3D base-calling templates (i.e., polony maps) for in situ sequencing analysis. Based on the characteristics of polony or clusters in flow cell images, the techniques disclosed herein efficiently remove out-of-focus polony and background objects without affecting in-focus polony or clusters in the flow cell image and ultimately in the 3D polony map. The 3D polony map can be used to extract polony intensities for base-calling. Compared to a single, flattened 2D polony map, the 3D polony map advantageously retains polony and cluster information that may be attenuated or completely removed within the 2D polony map. In some embodiments, the 3D polony map disclosed herein is not limited to generating base-calling positions only at predetermined axial positions for generating flow cell images. The 3D polony map may include subpixel resolution along the axial axis so that base-calling positions can be between two adjacent axial positions for the flow cell image. The 3D polony map can be generated in several early sequencing cycles and used in subsequent sequencing cycles to minimize the additional computational burden of recalculating a new polony map and the storage space required to store the new polony map. In some embodiments, the techniques disclosed herein advantageously utilize images that retain background information for accurate and efficient alignment of polony or clusters to images showing one or more cellular features.
[0093] 5 shows a flowchart of an exemplary embodiment of a method 500 for generating a base calling template for in situ sequencing by identifying base calling positions to be included in the base calling template. Method 500 may include some or all of the operations disclosed herein. The operations may be performed in the order described herein, but are not limited to such.
[0094] Method 500 may be performed by one or more processors (e.g., 404 of FIG. 4 ) disclosed herein. In some aspects, the processor may include one or more of a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit may include a central processing unit (CPU) and / or a graphics processing unit (GPU). The integrated circuit may include a chip such as a field programmable gate array (FPGA). In some aspects, the processor may include computing system 400.
[0095] In some aspects, some or all of the operations in method 500 may be performed by FPGA(s). In aspects in which some operations are performed by FPGA(s), data following the operations performed by the FPGA(s) may be communicated by the FPGA(s) to the CPU(s) so that the CPU(s) can use such data to perform subsequent operation(s) in method 500. Similarly, data may also be communicated from the CPU(s) to the FPGA(s) for processing by the FPGA(s). In some aspects, all of the operations in method 500 may be performed by CPU(s). Alternatively, the operations performed by the CPU(s) may be performed by dedicated processors or other processors, such as GPU(s). In some aspects, all of the operations in method 500 may be performed by FPGA(s).
[0096] In some embodiments, some or all of the operations of method 500 are performed during or before sequencing cycle N in a sequencing run. A base calling template, e.g., a polony map, can be generated using information obtained from some or all of cycles 1 through N. Polonies or clusters from one or more channels within such a cycle can be included in the template in a reference coordinate system, but flow cell images for cycle N and / or subsequent cycles have not yet been captured or are not currently being captured. In some embodiments, cycle N is the current cycle. N can be any non-zero integer, e.g., 3, 4, 5, 6, or 10.
[0097] As disclosed herein, the terms "base calling template" and "polony map" are equivalent and can be used interchangeably. A base calling template or polony map can include one or more base calling positions as disclosed herein. A base calling template or polony map disclosed herein can be either 2D or 3D. If all base calling positions in a base calling template or polony map are in the same 2D plane, the base calling template or polony map is 2D. If some of the base calling positions are not in the same 2D plane but are in 3D space, the base calling template or polony map is 3D.
[0098] Various coordinate systems can be used to define spatial locations and relationships herein. Although used herein to define spatial locations, various alternative coordinate systems can be used to define spatial locations and relationships. Some example coordinate systems other than Cartesian systems can include, but are not limited to, polar, cylindrical, or spherical coordinate systems. Other coordinate systems can include homogeneous or non-homogeneous coordinate systems.
[0099] In some embodiments, the method 500 can include providing a cell sample carrying a plurality of RNAs, including a first target RNA molecule and a second target RNA molecule.
[0100] In some embodiments, the cell sample is fixed and permeabilized. In some embodiments, the cell sample possesses 2-25 different target RNA molecules, or 25-50 different target RNA molecules, or 50-75 different target RNA molecules, or 75-100 different target RNA molecules. In some embodiments, the cell sample possesses more than 100 different target RNA molecules, or more than 250 different target RNA molecules, or more than 500 different target molecules, or more than 1000 different target RNA molecules, or more. In some embodiments, the cell sample possesses more than 10,000 different target RNA molecules. In some embodiments, the cell sample comprises whole cells, a plurality of whole cells, an intact tissue, or an intact tumor. In some embodiments, the cell sample comprises a fresh cell sample, a fresh frozen cell sample, a sectioned cell sample, an FFPE cell sample, or a sectioned FFPE cell sample. In some embodiments, the cell sample is deposited on a solid support. In some embodiments, the cell sample is deposited on a solid support passivated with a coating that promotes cell adhesion. In some embodiments, the cell sample is deposited on a support lacking immobilized capture oligonucleotides. In some embodiments, the cell sample is cultured before or after depositing the cell sample on the solid support. In some embodiments, the cell sample is cultured before performing step (b) described below. In some embodiments, the cell sample comprises an expanded cell sample cultured in a simple or complex cell culture medium. In some embodiments, the cell sample is not cultured or expanded before generating a plurality of cDNA molecules in the cell sample.
[0101] In some embodiments, the method 500 can include generating, within the cell sample, a plurality of cDNA molecules, including a first target cDNA molecule corresponding to a first target RNA molecule and a second target cDNA molecule corresponding to a second target RNA molecule.
[0102] In some embodiments, the act of generating the plurality of cDNA molecules comprises generating at least 2-10,000 different target cDNA molecules corresponding to 2-10,000 different target RNA molecules. In some embodiments, the act of generating the cDNA molecules comprises contacting a plurality of RNAs in the cellular sample with (i) a plurality of reverse transcription primers, (ii) a plurality of reverse transcriptases, and (iii) a plurality of nucleotides under conditions suitable for performing a reverse transcription reaction to generate a plurality of cDNA molecules (e.g., a plurality of first strand cDNA molecules) in the cellular sample (e.g., FIG. 7).
[0103] In some embodiments, the plurality of reverse transcription primers comprises a first subpopulation of target-specific reverse transcription primers that selectively hybridize to a first target RNA and a second subpopulation of target-specific reverse transcription primers that selectively hybridize to a second target RNA, In some embodiments, the first and second subpopulations of target-specific reverse transcription primers have the same sequence or different sequences.
[0104] In some embodiments, the entire length of the first subpopulation of target-specific reverse transcription primers hybridizes to the first target RNA molecule. In some embodiments, the first subpopulation of target-specific reverse transcription primers comprises a tailed primer having a portion that hybridizes to the first target RNA molecule and a portion that does not hybridize to the first target RNA molecule. In some embodiments, the first subpopulation of target-specific reverse transcription primers comprises at least a portion having a poly-T sequence. In some embodiments, the first subpopulation of target-specific reverse transcription primers comprises at least a portion having a random sequence and / or at least a portion having a target-specific sequence.
[0105] In some embodiments, the entire length of the second subpopulation of target-specific reverse transcription primers hybridizes to the second target RNA molecule. In some embodiments, the second subpopulation of target-specific reverse transcription primers comprises a tailed primer having a portion that hybridizes to the second target RNA molecule and a portion that does not hybridize to the second target RNA molecule. In some embodiments, the second subpopulation of target-specific reverse transcription primers comprises at least a portion having a poly-T sequence. In some embodiments, the second subpopulation of target-specific reverse transcription primers comprises at least a portion having a random sequence and / or at least a portion having a target-specific sequence.
[0106] In some embodiments, the target RNA molecule hybridized to the cDNA molecule can be subjected to enzymatic degradation using ribonuclease under conditions suitable for degrading the RNA in the RNA / DNA duplex. In some embodiments, the target RNA molecule hybridized to the cDNA molecule is not subjected to enzymatic degradation.
[0107] In some embodiments, the method 500 further comprises contacting the plurality of cDNA molecules in the cell sample with a plurality of target-specific padlock probes, comprising at least a first plurality of first target-specific padlock probes and a second plurality of second target-specific padlock probes. In some embodiments, the method comprises contacting the plurality of cDNA molecules in the cell sample with at least 2 to 10,000 different target-specific padlock probes.
[0108] In alternative embodiments, the cDNA is not generated from RNA within the cell sample. In some embodiments, method 500 includes contacting RNA within a cell with a plurality of target-specific padlock probes to generate circularized padlock probes. In some embodiments, the method includes contacting a plurality of RNA molecules in the cell sample with a plurality of target-specific padlock probes, including at least a first plurality of target-specific padlock probes and a second plurality of target-specific padlock probes. In some embodiments, the method includes contacting a plurality of cDNA molecules in the cell sample with at least 2 to 10,000 different target-specific padlock probes. In some embodiments, the target RNA molecules may be subjected to enzymatic degradation using ribonucleases. In some embodiments, the target RNA molecules are not subjected to enzymatic degradation.
[0109] In some embodiments, each padlock probe in the plurality of first target-specific padlock probes comprises a first and a second terminal region (e.g., a first and a second padlock binding arm), wherein the first terminal region selectively hybridizes to a first region of the first target cDNA molecule (or the first target RNA molecule), and the second terminal region selectively hybridizes to a second region of the first target cDNA molecule (or the first target RNA molecule). In some embodiments, contacting the plurality of cDNA molecules comprises hybridizing the first and the second terminal regions of the first target-specific padlock probe to proximal positions on the first target cDNA molecule (or the first target RNA molecule) to form a circularized first target-specific padlock probe having a nick or a gap between the hybridized first and the second terminal regions (e.g., Figure 7, left). In some embodiments, the first target-specific padlock probe comprises a first target barcode sequence (target BC-1) that corresponds to and uniquely identifies the first target cDNA sequence (or the first target RNA sequence). In some embodiments, the first target-specific padlock probe comprises a first target barcode sequence located adjacent to one of the regions of the first target-specific padlock probe that selectively hybridizes to the first target cDNA molecule (or the first target RNA sequence). In some embodiments, the first target-specific padlock probe comprises at least one universal adapter sequence, e.g., a universal sequencing primer binding site (or a complementary sequence thereof). In some embodiments, the first target-specific padlock probe comprises a universal primer binding site for a rolling circle amplification primer (or a complementary sequence thereof). In some embodiments, the first target-specific padlock probe comprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof).
[0110] In some embodiments, each padlock probe in the plurality of second target-specific padlock probes comprises first and second terminal regions (e.g., first and second padlock binding arms), wherein the first terminal region selectively hybridizes to a first region of the second target cDNA molecule (or second target RNA molecule), and the second terminal region selectively hybridizes to a second region of the second target cDNA molecule (or second target RNA molecule). In some embodiments, the contacting in step (c) comprises hybridizing the first and second terminal regions of the second target-specific padlock probe to proximal positions on the second target cDNA molecule (or second target RNA molecule) to form a circularized second target-specific padlock probe having a nick or gap between the hybridized first and second terminal regions (e.g., Figure 7, right). In some embodiments, the second target-specific padlock probe comprises a second target barcode sequence (target BC-2) that corresponds to and uniquely identifies the second target cDNA sequence (or the second target RNA sequence). In some embodiments, the second target-specific padlock probe comprises a second target barcode sequence located adjacent to one of the regions of the second target-specific padlock probe that selectively hybridizes to the second target cDNA molecule (or the second target RNA sequence). In some embodiments, the second target-specific padlock probe comprises at least one universal adapter sequence, e.g., a universal sequencing primer binding site (or its complementary sequence). In some embodiments, the second target-specific padlock probe comprises a universal primer binding site for a rolling circle amplification primer (or its complementary sequence). In some embodiments, the second target-specific padlock probe comprises a universal compaction oligonucleotide binding site (or its complementary sequence).
[0111] In some embodiments, the first target barcode sequence (target BC-1) and the second target barcode sequence (target BC-2) have different sequences and can be used to perform multiplex RNA detection and sequencing. In some embodiments, the first target barcode sequence (target BC-1) and the second target barcode sequence (target BC-2) have the same sequence and can be used to perform uniplex RNA detection and sequencing.
[0112] In some embodiments, the first and second target-specific padlock probes comprise a universal sequencing primer binding site and a target barcode sequence adjacent to each other, such that the target barcode region of the concatemer is sequenced first. The target barcode sequence can be any length, for example, 3-15 bases, or 15-25 bases, or 25-40 bases, or longer.
[0113] In some embodiments, contacting the plurality of RNA molecules in the cell sample with the plurality of target-specific padlock probes comprises hybridizing first and second terminal regions of a first target cDNA molecule or a first target RNA molecule to proximal positions to form a circularized first target-specific padlock probe having a nick or gap between the hybridized first and second terminal regions.
[0114] Figure 6 is a schematic diagram illustrating an exemplary embodiment of a padlock probe. In some embodiments, a padlock probe comprises a single-stranded nucleic acid molecule having two terminal regions (e.g., first and second binding arms) and an internal region. In some embodiments, the first terminal region of each padlock probe has a first target-specific sequence that selectively hybridizes to a first region of a target RNA or target cDNA molecule, and the second terminal region of each padlock probe has a second target-specific sequence that selectively hybridizes to a second region of the same target RNA or target cDNA molecule. In some embodiments, the internal region of the padlock contains a target barcode sequence (e.g., target BC-1 or target BC-2, left and right schematics, respectively) corresponding to a given target RNA or target cDNA. In some embodiments, the target barcode sequence uniquely identifies the target RNA or target cDNA. In some embodiments, the internal region of the padlock contains a universal primer binding site for a sequencing primer (or its complementary sequence). In some embodiments, the internal region of the padlock comprises a universal primer binding site for a rolling circle amplification primer (or its complementary sequence). In some embodiments, the internal region of the padlock comprises a universal binding site for a compaction oligonucleotide binding (or its complementary sequence). In some embodiments, the internal region of the padlock probe comprises a target barcode sequence and at least one universal primer binding site (e.g., for binding a sequencing primer, for binding a rolling circle amplification primer, and / or for binding a compaction oligonucleotide) in any configuration and orientation (Figure 6, top and bottom).
[0115] 7 is a schematic diagram showing a workflow for generating circularization padlock probes in a cell, including generating first and second cDNAs from first and second target RNA molecules, respectively, and hybridizing first and second padlock probes to the first and second cDNA molecules, respectively, to generate first and second circularization padlock probes, each of which contains (i) a first target barcode sequence (target BC-1) that uniquely identifies the first target RNA or the first target cDNA, (ii) a first sequencing primer binding site (or its complementary sequence), (iii) a universal binding site for an amplification primer (universal RCA) (or its complementary sequence), and (iv) a universal binding site for a compaction oligonucleotide (or its complementary sequence). The second padlock probe comprises (i) a second target barcode sequence (target BC-2) that uniquely identifies the second target RNA or the second target cDNA, (ii) a second sequencing primer binding site (or its complementary sequence), (iii) a universal binding site for an amplification primer (universal RCA) (or its complementary sequence), and (iv) a universal binding site for a compaction oligonucleotide (or its complementary sequence).
[0116] In some embodiments, the method 500 includes performing an enzymatic reaction to close nicks or gaps in at least a first and a second circularized target-specific padlock probe, thereby generating at least a first covalently closed circular padlock probe and a second covalently closed circular padlock probe in the cell sample. In some embodiments, closing the nicks in the first and second circularized padlock probes includes performing an enzymatic ligation reaction. In some embodiments, closing the gaps in the first and second circularized padlock probes includes performing a polymerase-catalyzed fill-in reaction using the first or second target cDNA molecule (or the first or second RNA molecule) as a template and performing an enzymatic ligation reaction. In some embodiments, the method includes performing one or more enzymatic reactions to close nicks or gaps in at least 2-10,000 circularized target-specific padlock probes, thereby generating at least 2-10,000 covalently closed circular padlock probes in the cell sample. FIG. 8 shows an exemplary covalently closed circular padlock probe.
[0117] In some embodiments, the method 500 includes performing a rolling circle amplification reaction in a cell sample using first and second covalently closed circular padlock probes as template molecules, thereby generating a plurality of concatemeric molecules including at least a first concatemeric molecule corresponding to a first target RNA molecule, and the plurality of concatemeric molecules including at least a second concatemeric molecule corresponding to a second target RNA molecule. In some embodiments, the first concatemeric molecule includes a tandem repeat unit, the unit including a sequence corresponding to the first target cDNA (or the first target RNA), the first target barcode sequence, and a universal sequencing primer binding site (or a complementary sequence thereof). In some embodiments, the second concatemeric molecule includes a tandem repeat unit, the unit including a sequence corresponding to the second target cDNA (or the second target RNA), the second target barcode sequence, and a universal sequencing primer binding site (or a complementary sequence thereof).
[0118] In some embodiments, the rolling circle amplification reaction described herein comprises contacting a covalently closed circularized padlock probe with an amplification primer (e.g., a universal rolling circle amplification primer), a strand-displacing DNA polymerase, and a plurality of nucleotides under conditions suitable for hybridizing each amplification primer to the covalently closed circularized padlock probe and for primer extension using the covalently closed circularized padlock probe as a template molecule to generate nucleic acid concatemers. In some embodiments, the method comprises performing a rolling circle amplification reaction in a cellular sample using at least 2 to 10,000 covalently closed circularized padlock probes as template molecules, thereby generating at least 2 to 10,000 concatemer molecules corresponding to at least 2 to 10,000 target RNA molecules. In some embodiments, the multiple concatemers generated in the cellular sample are collapsed into DNA nanoballs having a more compact shape and size than the uncollapsed concatemers.
[0119] In some embodiments, the method 500 includes an operation 510 of generating a plurality of flow cell images of a cell sample immobilized on a support by performing one or more sequencing reaction cycles.
[0120] The plurality of flow cell images can be a first plurality of flow cell images acquired in one or more cycles of a sequencing reaction. The one or more cycles can be any cycles during a sequencing run. The one or more cycles can be earlier cycles during a sequencing run, such as cycles 1 to 10 or cycles 1 to 5.
[0121] In some embodiments, the cell sample comprises a plurality of concatemer molecules therein. A first concatemer molecule(s) of the plurality of concatemer molecules may correspond to a first target RNA molecule of the cell sample. A second concatemer molecule(s) of the plurality of concatemer molecules may correspond to a second target RNA molecule of the cell sample.
[0122] In some embodiments, the cell sample can be cultured on a support. In some embodiments, method 500 includes culturing the cell sample on a support under conditions suitable for expanding the cell sample for 2-10 generations or more. The cultured cell sample can produce colonies of cells. In some embodiments, the cell sample includes one or more in situ samples. In some embodiments, the cell sample includes one or more cells or tissues. In some embodiments, the method can include culturing the cell sample to confluence or non-confluence. In some embodiments, the method includes culturing the cell sample on a support in a simple or complex cell culture medium. For example, the cell culture medium may include D-MEM high glucose (e.g., from Thermo Fisher Scientific, catalog number 11965118), fetal bovine serum (e.g., 10% FBS; e.g., from Thermo Fisher Scientific, catalog number A3160402), MEM non-essential amino acids (e.g., 0.1 mM MEM, e.g., from Thermo Fisher Scientific, catalog number 11140050), L-glutamine (e.g., 6 mM L-glutamine, e.g., from Thermo Fisher Scientific, catalog number A2916801), MEM sodium pyruvate (e.g., 1 mM sodium pyruvate, e.g., from Thermo Fisher Scientific, catalog number 11360070), and antibiotics (e.g., 1% penicillin-streptomycin-glutamine, e.g., from Thermo Fisher Scientific, catalog number 10378016). In some embodiments, the method includes culturing the cell sample at a humidity and temperature suitable for culturing the cell(s) on the support. Exemplary suitable conditions include approximately 37°C with a humidified atmosphere of approximately 5-10% carbon dioxide in air. The cell sample can be cultured with suitable aeration with oxygen and / or nitrogen.
[0123] In any of the methods described herein, a cell sample possesses a plurality of RNAs, including target RNA and / or non-target RNA. In some embodiments, cells typically produce RNA through gene expression, which involves transcription of DNA (e.g., genomic DNA) into RNA molecules. The transcribed RNA may or may not be spliced. The transcribed RNA may be translated into a polypeptide (e.g., coding RNA), or may not be translated but processed into tRNA or rRNA (e.g., non-coding RNA).
[0124] In some embodiments, a cell sample can contain a plurality of RNAs therein. The RNAs can be within individual cells of the cell sample. The plurality of RNAs can be carried by the cell sample. The RNAs can include target and non-target RNAs. In some embodiments, the plurality of RNAs carried by the cell sample includes wild-type RNA, mutant RNA, or splice variant RNA. In some embodiments, the plurality of RNAs carried by the cell sample includes pre-spliced RNA, partially spliced RNA, or fully spliced RNA. In some embodiments, the plurality of RNAs carried by the cell sample includes coding RNA, non-coding RNA, mRNA, tRNA, rRNA, microRNA (miRNA), mature microRNA, or immature microRNA. In some embodiments, the plurality of RNAs carried by the cell sample includes housekeeping RNA, cell-specific RNA, tissue-specific RNA, or disease-specific RNA. In some embodiments, the plurality of RNAs carried by the cell sample includes RNA expressed by one or more cells in response to a stimulus such as heat, light, a chemical, or a drug. In some embodiments, the plurality of RNAs carried by the cell sample includes RNA found in healthy cells or diseased cells. In some embodiments, the plurality of RNAs carried by the cell sample comprises RNAs transcribed from transgenic DNA sequences introduced into the cell sample using recombinant DNA procedures. For example, the RNAs can be transcribed from transgenic DNA sequences controlled by an inducible or constitutive promoter sequence. In some embodiments, the plurality of RNAs carried by the cell sample comprises RNAs transcribed from DNA sequences that are not transgenic.
[0125] In some embodiments, operation 510 includes imaging optical color signals emitted from nucleotide reagents bound to the plurality of concatemeric molecules in each of one or more cycles. The imaging operation can be performed by the optical system or imager 116 of sequencing system 100 herein. The first plurality of flow cell images can include optical color signals emitted from nucleotide reagents bound to the plurality of concatemeric molecules.
[0126] In some embodiments, operation 510 includes sequencing a plurality of concatemer molecules in a cell sample, which includes sequencing the first concatemer molecules by performing a predetermined number of sequencing cycles (e.g., 2 to 30 cycles) or less to generate a plurality of first sequencing read products, and sequencing the second concatemer molecules by performing a predetermined number of sequencing cycles or less to generate a plurality of second sequencing read products (Figure 8). In some embodiments, the sequencing operation includes sequencing a predetermined number of bases (e.g., 2 to 30 bases) or less of the first concatemer molecules to generate a plurality of first sequencing read products, which includes sequencing a predetermined number of bases of the second concatemer molecules for a predetermined number of sequencing cycles or less to generate a plurality of second sequencing read products. In some embodiments, the method comprises sequencing at least 2-10,000 concatemeric molecules in a cellular sample, comprising performing no more than a predetermined number of sequencing cycles on the 2-10,000 concatemeric molecules to generate a plurality of sequencing read products.
[0127] In some embodiments, only the first target barcode region of the first concatemer molecule is sequenced (e.g., Figure 8, top). In some embodiments, at least a portion or the entire length of the first target barcode of the first concatemer molecule is sequenced (e.g., Figure 8, top). In some embodiments, the first target barcode is sequenced, and a portion of the first cDNA region (or first RNA region) of the first concatemer molecule is sequenced. In some embodiments, at least a portion of the first cDNA region (or first RNA region) of the first concatemer molecule is sequenced.
[0128] In some embodiments, only the second target barcode region of the second concatemer molecule is sequenced (e.g., Figure 8, bottom). In some embodiments, at least a portion or the entire length of the second target barcode of the second concatemer molecule is sequenced (e.g., Figure 8, bottom). In some embodiments, the second target barcode is sequenced, and a portion of the second cDNA region (or second RNA region) of the second concatemer molecule is sequenced. In some embodiments, at least a portion of the second cDNA region (or second RNA region) of the second concatemer molecule is sequenced.
[0129] In some embodiments, the sequencing operation comprises contacting a plurality of concatemer molecules in a cell sample with (i) a plurality of universal sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of universal sequencing primers to their respective universal sequencing primer binding sites on the concatemers. In some embodiments, the sequencing operation further comprises performing a predetermined number or less of sequencing cycles (e.g., 2 to 30 cycles) to generate at least a first plurality of sequencing read products by sequencing at least a first target barcode region (target BC-1), and optionally performing a predetermined number or less of sequencing cycles to generate at least a second plurality of sequencing read products by sequencing at least a second target barcode region (target BC-2). In some embodiments, the nucleotide reagents comprise multivalent molecules, nucleotides, and / or nucleotide analogs.
[0130] In some embodiments, the sequencing operation is performed at 1.0 mm 2 sequencing at least a portion of the first and second nucleic acid concatemers using an optical imaging system comprising a field of view (FOV) of greater than or equal to 100 nm;
[0131] In some embodiments, the plurality of first and second sequencing read products are detectable by imaging, and sequencing comprises decoding the plurality of first and second sequencing read products from images obtained during a predetermined number of sequencing cycles or less.
[0132] In some embodiments, the plurality of first and second sequencing read products are detectable by imaging, and the sequencing comprises simultaneously imaging the plurality of first and second detectable sequencing read products in the cell sample (co-localization of the first and second sequencing read products).
[0133] Exemplary sequencing read products are shown as dashed arrows in Figures 8-10. The first sequencing read product may include a first target barcode sequence and, optionally, a portion of a first insert sequence. The first sequencing read product may include a complement or reverse complement of the target barcode sequence and, optionally, a portion of the insert sequence.
[0134] 8 is a schematic diagram illustrating an intracellular rolling circle and sequencing workflow that includes performing rolling circle amplification using first and second covalently closed circular molecules to generate first and second concatemers, which are then subjected to a sequencing workflow using a universal sequencing primer, a sequencing polymerase, and multiple nucleotide reagents.
[0135] Figure 9 is a schematic diagram showing an exemplary workflow for sequencing concatemers generated in cells. The concatemers contain tandem repeat units, each of which contains (i) a universal sequencing primer binding site (Seq), (ii) a universal compaction oligonucleotide binding site (CO), (iii) an insert sequence corresponding to a given target cDNA, and (iv) a target barcode sequence corresponding to a given target cDNA (BC). In some embodiments, a universal sequencing primer (solid arrow) hybridizes to the universal sequencing primer binding site, and 30 or fewer sequencing cycles are performed to generate multiple first sequencing read products (dashed arrows), each of which contains only the target barcode sequence. A plurality of first sequencing read products are removed from the concatemer, and sequencing is repeated, and 30 or less sequencing cycles are performed to generate another plurality of first sequencing read products (dashed arrows), and the first sequencing read products contain only the target barcode sequence. A plurality of first sequencing read products are removed from the concatemer, and sequencing is repeated again, and 30 or less sequencing cycles are performed to generate another plurality of first sequencing read products (dashed arrows), and the first sequencing read products contain only the target barcode sequence. In some embodiments, the repeated sequencing can be performed up to 50 times. The sequences of all of the first sequencing read products can be determined and aligned with a first reference sequence (e.g., a reference barcode sequence) to confirm the presence of the first target RNA molecule in the cell sample.
[0136] In some embodiments, the method 500 includes removing a plurality of first sequencing read products from the first concatemer molecules, retaining the first concatemer molecules in the cellular sample, and removing a plurality of second sequencing read products from the second concatemer molecules, retaining the second concatemer molecules in the cellular sample. In some embodiments, the first and second concatemer molecules remain in the cellular sample, but the complementary or reverse complementary strands of the sequencing read products are removed.
[0137] In some embodiments, method 500 includes iteratively sequencing the plurality of concatemers by repeating the operation of 510, removing a plurality of first sequencing read products from the first concatemer molecules, retaining the first concatemer molecules in the cellular sample, and removing a plurality of second sequencing read products from the second concatemer molecules, and retaining the second concatemer molecules in the cellular sample, n times.
[0138] In some embodiments, method 500 includes repeating the steps of: sequencing a first concatemer molecule(s) and a second concatemer molecule(s) by performing a predetermined number of sequencing cycles or less to generate a plurality of first and second sequencing read products; removing a plurality of first sequencing read products from the first concatemer molecules and retaining the first concatemer molecules in the cellular sample; and removing a plurality of second sequencing read products from the second concatemer molecules and retaining the second concatemer molecules in the cellular sample, the repeating steps comprising n times repetitively sequencing the plurality of concatemers.
[0139] In some embodiments, the sequences of the plurality of first sequencing read products confirm the presence of a first target RNA molecule in the cellular sample, and the sequences of the plurality of second sequencing read products confirm the presence of a second target RNA molecule in the cellular sample.
[0140] In some embodiments, repeatedly sequencing at least one region of the concatemer comprises repeating the sequencing and removal operations n times, where n can range from 1 to 20, 1 to 30, 1 to 100, or an even larger integer range.
[0141] Examples of iterative sequencing are shown in the schematic diagrams in Figures 9 to 12. Iterative sequencing can help reduce sequencing errors and improve sequencing accuracy and reliability.
[0142] In some embodiments, at least one concatemer is sequenced by non-iterative sequencing. In some embodiments, at least one concatemer is sequenced by repeating the sequencing and removal operations once or twice.
[0143] 10 is a schematic diagram showing an exemplary workflow for sequencing concatemers generated in cells. The concatemers comprise tandem repeat units, each comprising (i) a universal sequencing primer binding site (Seq), (ii) a universal compaction oligonucleotide binding site (CO), (iii) an insert sequence corresponding to a given target cDNA, and (iv) a target barcode sequence corresponding to a given target cDNA (BC). In some embodiments, a universal sequencing primer (solid arrow) hybridizes to the universal sequencing primer binding site, and 30 or fewer sequencing cycles are performed to generate a plurality of first sequencing read products (dashed arrows), each of which comprises the target barcode sequence and a portion of the insert sequence. A plurality of first sequencing read products are removed from the concatemers, and sequencing is repeated, performing up to 30 sequencing cycles to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products comprise the target barcode sequence and a portion of the insert sequence. A plurality of first sequencing read products are removed from the concatemers, and sequencing is repeated again, performing up to 30 sequencing cycles to generate another plurality of first sequencing read products (dashed arrows), where the first sequencing read products comprise the target barcode sequence and a portion of the insert sequence. In some embodiments, the repeated sequencing can be performed up to 50 times. The sequences of all of the first sequencing read products can be determined and aligned with a first reference sequence (e.g., a reference barcode sequence and an insert sequence corresponding to the target RNA) to confirm the presence of the first target RNA molecule in the cell sample.
[0144] Figure 11 is a schematic diagram showing an exemplary workflow for sequencing concatemers generated in cells. The concatemers contain tandem repeat units, each containing (i) a universal sequencing primer binding site (Seq), (ii) a universal compaction oligonucleotide binding site (CO), and (iii) an insert sequence corresponding to a given target cDNA. In some embodiments, a universal sequencing primer (solid arrow) hybridizes to the universal sequencing primer binding site, and 30 or fewer sequencing cycles are performed to generate a plurality of first sequencing read products (dashed arrows), each of which contains a portion of the insert sequence. The plurality of first sequencing read products are removed from the concatemers, and sequencing is repeated, followed by 30 or fewer sequencing cycles to generate another plurality of first sequencing read products (dashed arrows), each of which contains a portion of the insert sequence. A plurality of first sequencing read products are removed from concatemers, and sequencing is repeated again, and 30 or less sequencing cycles are carried out to generate another plurality of first sequencing read products (dashed arrows), and the first sequencing read products contain a portion of the inserted sequence.In some embodiments, repeat sequencing can be carried out up to 50 times.The sequences of all of the first sequencing read products can be determined and aligned with the first reference sequence (for example, the inserted sequence corresponding to the target RNA), to confirm the existence of the first target RNA molecule in the cell sample.
[0145] 12 is a schematic diagram showing an exemplary workflow for sequencing concatemers generated in cells. The concatemers contain tandem repeat units, each containing (i) a universal sequencing primer binding site (Seq) and (ii) an insert sequence corresponding to a given target cDNA. In some embodiments, a universal sequencing primer (solid arrow) hybridizes to the universal sequencing primer binding site, and 30 or fewer sequencing cycles are performed to generate a plurality of first sequencing read products (dashed arrows), each of which contains a portion of the insert sequence. The plurality of first sequencing read products are removed from the concatemers, and sequencing is repeated, followed by 30 or fewer sequencing cycles to generate another plurality of first sequencing read products (dashed arrows), each of which contains a portion of the insert sequence. A plurality of first sequencing read products are removed from concatemers, and sequencing is repeated again, and 30 or less sequencing cycles are carried out to generate another plurality of first sequencing read products (dashed arrows), and the first sequencing read products contain a portion of the inserted sequence.In some embodiments, repeat sequencing can be carried out up to 50 times.The sequences of all of the first sequencing read products can be determined and aligned with the first reference sequence (for example, the inserted sequence corresponding to the target RNA), to confirm the existence of the first target RNA molecule in the cell sample.
[0146] In some embodiments, operation 510 comprises sequencing only the first target barcode sequence region of the first concatemer molecule in each of one or more cycles, thereby generating a first sequencing read product. In some embodiments, operation 510 can comprise sequencing some or all of the first target barcode sequence region and at least a portion of the first insert sequence of the first concatemer in each of one or more cycles, thereby generating a first sequencing read product.
[0147] In some embodiments, operation 510 can include sequencing only the second target barcode sequence region of the second concatemer in each of one or more cycles, thereby generating a second sequencing read product. In some embodiments, operation 510 can include sequencing the second target barcode sequence region and at least a portion of the second insert sequence of the second concatemer in each of one or more cycles, thereby generating a second sequencing read product.
[0148] In some embodiments, operation 510 comprises sequencing only the first target barcode sequence region of the first concatemer molecule and the second target barcode sequence region of the second concatemer, but not any portion of the first insert sequence of the first concatemer or any portion of the second insert sequence of the second concatemer, in each of one or more cycles, thereby generating a first sequencing read product and a second sequencing read product. In some embodiments, operation 510 comprises sequencing some or all of the first target barcode sequence region and at least a portion of the first insert sequence of the first concatemer, in each of one or more cycles, thereby generating a first sequencing read product.
[0149] In some embodiments, each of the plurality of concatemer molecules immobilized on the support corresponds to a polony. In some embodiments, each of the plurality of concatemer molecules immobilized on the support corresponds to a base calling position of a base calling template. Two or more different concatemer molecules of the plurality of concatemer molecules can have different insert sequences. Exemplary insert sequences are shown in Figure 8. The different insert sequences correspond to different target RNA molecules or target cDNA molecules.
[0150] Nucleotide diversity of a population of immobilized concatemeric molecules can refer to the relative proportions of nucleotides A, G, C, and T / U present in a sequencing cycle. High or balanced diversity data can generally include barcode and / or insertion regions in which all four nucleotides are represented in approximately equal proportions in each cycle of a sequencing run. Low or unbalanced diversity data can generally include barcode and / or sequence-of-interest (insertion) regions in which certain nucleotides are overrepresented and other nucleotides are underrepresented. To overcome the problem of low-diversity libraries, a small amount of a high-diversity library prepared from PhiX bacteriophage is typically mixed with a library of interest (e.g., a PhiX spike-in library) and sequenced together on the same flow cell. While the PhiX library spike-in library can provide nucleotide diversity, it also occupies space on the flow cell, thereby displacing the target library carrying the sequence of interest and reducing the amount of sequencing data obtained from the target library (e.g., reducing sequencing throughput). Another way to overcome the problem of low-diversity libraries may be to prepare target library molecules with several index and / or barcode sequences designed to be color-balanced and therefore have high diversity. However, it may be desirable to design concatemer sets for multiple index or barcode sets, for example, 16-plex, 24-plex, 96-plex, or greater plexity levels. It may be difficult to design index or barcode sequences for a large sample index set in which all sample index or barcode sequences are color-signal-balanced.
[0151] A flow cell image herein may include light signals emitted from nucleotide reagents bound to A, G, C, and T / U nucleotide bases of balanced diversity among a plurality of concatemeric molecules immobilized on a support. A plurality of flow cell images herein may include light signals emitted from nucleotide reagents bound to A, G, C, and T / U nucleotide bases of unbalanced diversity among a plurality of concatemeric molecules immobilized on a support. The balanced or unbalanced diversity of bases can be determined in a particular sequencing cycle or multiple sequencing cycles, for example, cycles encompassing a barcode sequence herein.
[0152] Having a balanced diversity of bases can be advantageous over using an unbalanced diversity of bases in one or more sequencing cycles, and can facilitate more accurate and reliable generation of base calling templates. In some embodiments, the barcode sequence herein can be determined to provide a balanced diversity of bases in one or more sequencing cycles.
[0153] In some embodiments, the second plurality of flow cell images comprises light signals emitted from nucleotide reagents bound to A, G, C, and T / U nucleotide bases of unbalanced diversity among the plurality of concatemeric molecules immobilized on the support in one or more subsequent cycles.
[0154] The unbalanced diversity of A, G, C, and T / U nucleotide bases among multiple concatemeric molecules can comprise a percentage of the number of one or more types of nucleotide bases (1) relative to the total number of bases of all polonies within the field of view (2), which percentage is less than 20%, 15%, 10%, 5%, 3%, 2%, or 1% in one or more cycles.
[0155] The balanced diversity of A, G, C, and T / U nucleotide bases among multiple concatemeric molecules includes a percentage of the number of each type of nucleotide base (1) relative to the total number of bases in all polonies within the field of view (2) in one or more cycles, where the percentage is greater than 5%, 8%, 10%, 12%, 15%, 18%, or 20%. For example, the unbalanced diversity of nucleotide bases includes nucleotide base A, which is about 5% of the total number of all nucleotide bases in the concatemeric molecules of the cycle. As another example, the unbalanced diversity of nucleotide bases includes nucleotide base C, which is about 8% of the total number of all nucleotide bases among the concatemeric molecules of the cycle, and nucleotide base T, which is about 1%.
[0156] During a single sequencing run, base-calling positions, i.e., polonies or clusters, can shift, rotate, or otherwise spatially translate within flow cell images obtained from different cycles and / or across channels. As a result, templates are required to ensure that base-calling positions in a sequencing run are spatially aligned and that base calls are accurately assigned to the corresponding polonies or clusters, template molecules, and samples. It can be advantageous to generate templates early in a sequencing run to enable alignment of base-calling positions in all subsequent sequencing cycles after the template is generated. Early template generation can also advantageously reduce delays in primary analysis of flow cell images in subsequent cycles, enabling real-time analysis while the sequencing run is still in progress. For example, if templates are generated from the first 3-5 cycles, primary analysis can be performed in real time after flow cell images are acquired in cycle 6, and in cycle 7 in parallel with sequencing and imaging operations. Similar analysis for each cycle after cycle 6 can be performed while subsequent sequencing cycles are still in progress. Thus, the primary analysis can be quite soon after, if not when, the sequencing run is completed.
[0157] Methods for performing in situ batch-specific sequencing In some embodiments, the method 500 is configured to perform in situ sequencing and analysis thereof, including in situ multiplexed and multi-omics detection and discrimination, using coded padlock probes as disclosed herein. Padlock probes can be designed to selectively detect target RNAs.
[0158] The RNA-specific padlock probe can selectively hybridize with the cDNA corresponding to the target RNA.The RNA-specific probe can have a barcode that uniquely identifies the cDNA.In some embodiments, the RNA-specific padlock probe can also have a batch-specific sequencing primer binding site.
[0159] Different types (e.g., two or more) of padlock probes can be used to generate concatemers with batch-specific sequencing binding sites and multiple copies of barcodes. The concatemers can be collapsed into DNA nanoballs with compact shapes and sizes, which result in increased signal intensity and color differentiation during sequencing.
[0160] In the case of in situ sequencing, optical resolution limitations can hinder the ability to perform highly multiplexed sequencing. Batch-specific sequencing primer binding sites on padlock probes may allow sequencing of a desired subset (e.g., a batch) of concatemers using selected batch-specific sequencing primers to reduce signal and image overcrowding. As a result, cell samples can be prepared with a higher cell density per unit imaging space (than existing in situ sequencing methods use), and sequencing throughput may be improved using batch sequencing compared to existing in situ sequencing methods. The use of batch-specific sequencing primers produces clear, resolvable optical images. Performing multiple rounds of sequencing on the same cell sample using different batch-specific sequencing primers may enable multiplex sequencing to reveal multiple target RNAs.
[0161] The batch-specific sequencing described herein can have many applications. For example, the number of spots that are imaged and associated with sequencing can be counted. The counted spots can be used as a measure of RNA levels in cell samples.
[0162] In some embodiments, method 500 can be configured to determine base calling templates in in situ batch-specific sequencing applications. In some embodiments, a method in the context of in situ batch-specific sequencing includes providing a cellular sample deposited on a solid support, the cellular sample carrying (i) a first plurality of DNA amplicons (e.g., first concatemers) corresponding to a first target cDNA or RNA molecule, and (ii) a second plurality of DNA amplicons (e.g., second concatemers) corresponding to a second target cDNA or RNA molecule.
[0163] In some embodiments, operation 510 in the in situ batch-specific sequencing includes sequencing a first plurality of DNA amplicons in the cellular sample under conditions that inhibit sequencing of a second plurality of DNA amplicons, wherein sequencing the first plurality of DNA amplicons in the cellular sample includes generating a plurality of first sequencing read products, wherein sequences of the first sequencing read products are aligned with a first target reference sequence to confirm the presence of the first target RNA in the cellular sample. In some embodiments, the first amplicons can be iteratively sequenced by performing 2 to 30 or fewer sequencing cycles, or can be iteratively sequenced by performing 1 to 250 sequencing cycles.
[0164] In some embodiments, operation 510 in the in situ batch-specific sequencing may further include sequencing a second plurality of DNA amplicons in the cellular sample under conditions that inhibit sequencing of the first plurality of DNA amplicons, wherein sequencing the second plurality of DNA amplicons in the cellular sample includes generating a plurality of second sequencing read products, wherein the sequences of the second sequencing read products are aligned with a second target reference sequence to confirm the presence of the second target RNA in the cellular sample. In some embodiments, the second amplicons may be iteratively sequenced by performing 2 to 30 or fewer sequencing cycles, or may be iteratively sequenced by performing 1 to 250 sequencing cycles.
[0165] In some embodiments, the method 500 for in situ batch-specific sequencing includes preparing a cell sample deposited on a solid support, the cell sample carrying a first plurality of target RNAs and a second plurality of target RNAs. In some embodiments, the first plurality of target RNAs encodes a first polypeptide. In some embodiments, the second plurality of target RNAs encodes a second polypeptide. In some embodiments, the cell sample is fixed and permeabilized.
[0166] In some embodiments, the cell sample contains 2-25 different target RNA molecules, or 25-50 different target RNA molecules, or 50-75 different target RNA molecules, or 75-100 different target RNA molecules. In some embodiments, the cell sample contains more than 100 different target RNA molecules, or more than 250 different target RNA molecules, or more than 500 different target molecules, or more than 1000 different target RNA molecules, or more. In some embodiments, the cell sample contains more than 10,000 different target RNA molecules. In some embodiments, the cell sample comprises whole cells, a plurality of whole cells, an intact tissue, or an intact tumor. In some embodiments, the cell sample comprises a fresh cell sample, a fresh frozen cell sample, a sectioned cell sample, or an FFPE cell sample. In some embodiments, the cell sample is deposited on a solid support. In some embodiments, the cell sample is deposited on a solid support passivated with a coating that promotes cell adhesion. In some embodiments, the cell sample is deposited on a support lacking immobilized capture oligonucleotides. In some embodiments, the cell sample is cultured prior to performing step (b) described below.
[0167] In some embodiments, the cell sample possesses 2-25 different target polypeptide molecules, or possesses 25-50 different target polypeptide molecules, or possesses 50-75 different target polypeptide molecules, or possesses 75-100 different target polypeptide molecules. In some embodiments, the cell sample possesses more than 100 different target polypeptide molecules, or more than 250 different target polypeptide molecules, or more than 500 different target molecules, or more than 1000 different target polypeptide molecules, or more. In some embodiments, the cell sample has more than 10,000 different target polypeptide molecules. The target polypeptide molecules are encoded by target RNA molecules.
[0168] In some embodiments, a method 500 for in situ batch-specific sequencing includes generating a plurality of cDNAs in a cellular sample by (i) generating at least a first plurality of target cDNAs from a first plurality of target RNAs and (ii) generating at least a second plurality of target cDNAs from a second plurality of target RNAs (e.g., FIG. 13). In some embodiments, the first target cDNA corresponds to a first target RNA molecule. In some embodiments, the second target cDNA corresponds to a second target RNA molecule. In some embodiments, the method includes generating at least 2-10,000 different target cDNA molecules corresponding to 2-10,000 different target RNA molecules. In some embodiments, generating a plurality of cDNAs in a cellular sample includes contacting a plurality of RNAs in the cellular sample with (i) a plurality of reverse transcription primers, (ii) a plurality of reverse transcriptases, and (iii) a plurality of nucleotides under conditions suitable for performing a reverse transcription reaction to generate a plurality of cDNA molecules (e.g., a plurality of first strand cDNA molecules) in the cellular sample. In some embodiments, the plurality of reverse transcription primers comprises a first subpopulation of target-specific reverse transcription primers that selectively hybridize to a first target RNA and / or a second subpopulation of target-specific reverse transcription primers that selectively hybridize to a second target RNA. In some embodiments, the plurality of reverse transcription primers comprises a first subpopulation of random-sequence reverse transcription primers that hybridize to the first target RNA and a second subpopulation of random-sequence reverse transcription primers that hybridize to the second target RNA.
[0169] In some embodiments, the method 500 for in situ batch-specific sequencing includes generating, within a cell sample, a plurality of DNA concatemers corresponding to a first and second plurality of target RNA molecules, the operation comprising: (1) contacting the first plurality of target cDNAs with a first plurality of padlock probes to generate a first plurality of covalently closed circular padlock probes, the contacting being performed under conditions suitable for hybridizing a first binding arm and a second binding arm of the first padlock probe to proximal positions on their respective first target cDNA molecules to form a first plurality of circular padlock probes each having a nick or gap between the hybridized first binding arm and second binding arm; and the first padlock probes (i) bind to the first target cDNAs. (ii) a first target barcode sequence (Target BC-1) that uniquely identifies a target RNA or cDNA molecule; (ii) a first batch-specific sequencing primer binding site (Batch Seq-1) (or its complementary sequence); and (iii) a universal binding site for an amplification primer (Universal RCA) (or its complementary sequence) (e.g., Figure 13, left); (2) enzymatically closing nicks or gaps in the first plurality of covalently closed circular padlock probes to form a first plurality of covalently closed circular padlock probes; and (3) performing rolling circle amplification in a cell sample using the first covalently closed circular padlock probes as template molecules, thereby generating a first plurality of concatemeric molecules corresponding to the first plurality of target RNA or cDNA molecules. In some embodiments, the rolling circle amplification reaction can be performed in the presence or absence of a plurality of compaction oligonucleotides. In some embodiments, the method comprises contacting a plurality of cDNA molecules in a cell sample with at least 2 to 10,000 different target-specific padlock probes. In some embodiments, the first padlock probe further comprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof). In some embodiments, closing the nick in the first circularization padlock probe comprises performing an enzymatic ligation reaction.In some embodiments, closing the gaps in the first circularized padlock probe comprises performing a polymerase-catalyzed fill-in reaction using the first target cDNA molecule as a template and performing an enzymatic ligation reaction. In some embodiments, the method comprises performing an enzymatic reaction to close nicks or gaps in at least 2-10,000 circularized target-specific padlock probes, thereby generating at least 2-10,000 covalently closed circular padlock probes in the cell sample. In some embodiments, each concatemeric molecule in the first plurality comprises a tandem repeat unit, the unit comprising the sequence of the first target cDNA and (i) a first target barcode sequence (Target BC-1) that uniquely identifies the first target RNA, (ii) a first batch-specific sequencing primer binding site (Batch Seq-1) (or a complementary sequence thereof), and (iii) a universal binding site for amplification primers (Universal RCA) (or a complementary sequence thereof). In some embodiments, the unit further comprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof).
[0170] In some embodiments, the method 500 for in situ batch-specific sequencing includes generating, within a cell sample, a plurality of DNA concatemers corresponding to a second plurality of target RNA molecules, the operation comprising: (1) contacting the second plurality of target cDNAs with a first plurality of padlock probes to generate a second plurality of covalently closed circular padlock probes, the contacting being performed under conditions suitable for hybridizing a first binding arm and a second binding arm of the second padlock probe to proximal positions on their respective second target cDNA molecules to form a second plurality of circular padlock probes having a nick or gap between the hybridized first binding arm and second binding arm, respectively; and wherein the second padlock probe contains: (i) a second target barcode sequence (target BC-2) that uniquely identifies the second target cDNA or RNA; (ii) a second target barcode sequence (target BC-3) that uniquely identifies the second target cDNA or RNA; (ii) generating a batch-specific sequencing primer binding site (Batch Seq-2) (or a complementary sequence thereof), wherein the sequence of the second batch-specific sequencing primer binding site is different from the sequence of the first batch-specific sequencing primer binding site, and (iii) a universal binding site for amplification primers (universal RCA) (or a complementary sequence thereof) (e.g., FIG. 13, right); (ii) enzymatically closing nicks or gaps in the second plurality of covalently closed circular padlock probes to form a second plurality of covalently closed circular padlock probes; and (iii) performing rolling circle amplification in the cell sample using the second covalently closed circular padlock probes as template molecules, thereby generating a second plurality of concatemeric molecules corresponding to a second plurality of target RNAs. In some embodiments, the rolling circle amplification reaction can be performed in the presence or absence of a plurality of compaction oligonucleotides. In some embodiments, the method comprises contacting a plurality of cDNA molecules in a cell sample with at least 2-10,000 different target-specific padlock probes.In some embodiments, the second padlock probe further comprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof). In some embodiments, closing the nick in the second circularized padlock probe comprises performing an enzymatic ligation reaction. In some embodiments, closing the gap in the second circularized padlock probe comprises performing a polymerase-catalyzed fill-in reaction using the second target cDNA molecule as a template and performing an enzymatic ligation reaction. In some embodiments, the method comprises performing an enzymatic reaction to close nicks or gaps in at least 2-10,000 circularized target-specific padlock probes, thereby generating at least 2-10,000 covalently closed circular padlock probes in the cell sample. In some embodiments, each concatemeric molecule of the second plurality comprises a tandem repeat unit, the unit comprising a sequence of a second target cDNA and (i) a second target barcode sequence (Target BC-2) that uniquely identifies the second target cDNA or RNA, (ii) a second batch-specific sequencing primer binding site (Batch Seq-2) (or a complementary sequence thereof), and (iii) a universal binding site for an amplification primer (Universal RCA) (or a complementary sequence thereof). In some embodiments, the unit further comprises a universal compaction oligonucleotide binding site (or a complementary sequence thereof).
[0171] In some embodiments, performing one or more sequencing reaction cycles in in situ batch-specific sequencing comprises sequencing a first plurality of concatemer molecules in the cellular sample under conditions that inhibit sequencing of a second plurality of concatemers (e.g., FIG. 14). In some embodiments, sequencing the first plurality of concatemers in the cellular sample comprises performing 2 to 30 or fewer sequencing cycles to generate a plurality of first sequencing read products, and aligning the sequences of the first sequencing read products with a first target reference sequence to confirm the presence of the first target RNA in the cellular sample. In some embodiments, sequencing the first plurality of concatemers in the cellular sample comprises performing 1 to 250 sequencing cycles to generate a plurality of first sequencing read products, and aligning the sequences of the first sequencing read products with a first target reference sequence to confirm the presence of the first target RNA in the cellular sample.
[0172] In some embodiments, in the first concatemer molecule, only the first target barcode region (target BC-1) is sequenced. In some embodiments, in the first concatemer molecule, at least a portion or the entire length of the first target barcode (target BC-1) is sequenced. In some embodiments, in the first concatemer molecule, the first target barcode (target BC-1) is sequenced, and a portion of the first cDNA region is sequenced.
[0173] In some embodiments, sequencing the first concatemer in in situ batch-specific sequencing comprises contacting a first plurality of concatemer molecules in a cell sample with (i) a plurality of first batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of first batch-specific sequencing primers to their respective first batch-specific sequencing primer binding sites on the first concatemers (1). In some embodiments, sequencing further comprises performing 2 to 30 or fewer sequencing cycles using the first concatemers as template molecules to generate a first plurality of sequencing read products.
[0174] In some embodiments, the sequencing operation in the in situ batch-specific sequencing herein is performed at a speed of 1.0 mm 2 sequencing at least a portion of the first nucleic acid concatemer using an optical imaging system comprising a field of view (FOV) of greater than 1000 nm.
[0175] In some embodiments, the plurality of first sequencing read products are detectable by imaging, and the sequencing comprises decoding the plurality of first sequencing read products from images acquired during 2 to 30 or fewer sequencing cycles, or from images acquired during 1 to 250 sequencing cycles.
[0176] In some embodiments, the in situ batch-specific sequencing method 500 includes removing a plurality of first sequencing read products from a first concatemer molecule and retaining the first concatemer molecule in a cell sample. In some embodiments, a 3' blocking moiety can be added to the first sequencing read product to inhibit further sequencing reactions. For example, a nucleotide analog can be incorporated if it inhibits the incorporation of a subsequent nucleotide. Exemplary blocking nucleotide analogs include dideoxynucleotides or nucleotides with 2' or 3' chain terminating moieties.
[0177] In some embodiments, the in situ batch-specific sequencing method further comprises repeating steps (d) and (e) at least once to repeatedly sequence the plurality of first concatemers. In some embodiments, the repeated sequencing is optional.
[0178] In some embodiments, sequencing a first concatemer in in situ batch-specific sequencing comprises contacting a first plurality of concatemer molecules in a cellular sample with (i) a plurality of first batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the plurality of first batch-specific sequencing primers to their respective first batch-specific sequencing primer binding sites on the first concatemers (1). In some embodiments, sequencing further comprises performing 2 to 30 or fewer sequencing cycles using the first concatemers as template molecules to generate a first plurality of sequencing read products. In some embodiments, sequencing further comprises removing the first plurality of sequencing read products from the first concatemers and retaining the first plurality of concatemers in the cellular sample (3). In some embodiments, the sequencing further comprises step (4) repeating steps (1)-(3) at least once (e.g., FIG. 14). In some embodiments, step (4) comprises repeating steps (1)-(3) at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, or at least 10 times. In some embodiments, step (4) comprises repeating steps (1)-(3) up to 10 times, up to 20 times, up to 30 times, up to 40 times, or up to 50 times.
[0179] In some embodiments, the iterative sequencing of the first concatemer can be performed using a sequencing-by-binding procedure, labeled and / or unlabeled chain-terminating nucleotides, or multivalent molecules. Descriptions of these three sequencing methods are provided herein.
[0180] In some embodiments, multiple universal sequencing primers can be hybridized to concatemeric template molecules using a hybridization reagent containing formamide (e.g., 10-20% formamide) in an SSC buffer (e.g., 2x saline-sodium citrate) buffer. Hybridization conditions include a temperature of about 20-30°C for about 10-60 minutes.
[0181] In some embodiments, multiple sequencing read products can be removed from concatemers, and the multiple concatemers can be retained in the cell sample using a dehybridization reagent comprising an SSC buffer (e.g., saline-sodium citrate) buffer with or without formamide at a temperature that promotes nucleic acid denaturation, such as 30-90°C.
[0182] In some embodiments, the in situ batch-specific sequencing method 500 further comprises sequencing a second plurality of concatemer molecules in the cellular sample under conditions that inhibit sequencing of the first plurality of concatemers (e.g., FIG. 14). In some embodiments, sequencing the second plurality of concatemers in the cellular sample comprises performing 2 to 30 or fewer sequencing cycles to generate a plurality of second sequencing read products, and aligning the sequences of the second sequencing read products with a second target reference sequence to confirm the presence of the second target RNA in the cellular sample. In some embodiments, sequencing the second plurality of concatemers in the cellular sample comprises performing 1 to 250 sequencing cycles to generate a plurality of second sequencing read products, and aligning the sequences of the second sequencing read products with a second target reference sequence to confirm the presence of the second target RNA in the cellular sample.
[0183] In some embodiments, in the second concatemer molecule, only the second target barcode region (target BC-2) is sequenced. In some embodiments, in the second concatemer molecule, at least a portion or the entire length of the second target barcode (target BC-2) is sequenced. In some embodiments, in the second concatemer molecule, the second target barcode (target BC-2) is sequenced and a portion of the second cDNA region is sequenced.
[0184] In some embodiments, sequencing the second concatemers comprises contacting a second plurality of concatemer molecules in the cell sample with (i) a plurality of second batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the second batch-specific sequencing primers to their respective second batch-specific sequencing primer binding sites on the second concatemers (1). In some embodiments, sequencing further comprises performing 2 to 30 or fewer sequencing cycles using the second concatemers as template molecules to generate a second plurality of sequencing read products.
[0185] In some embodiments, sequencing the second plurality of concatemeric molecules comprises sequencing at least a portion of the second nucleic acid concatemers using an optical imaging system comprising a field of view (FOV) of greater than 1.0 mm.
[0186] In some embodiments, in sequencing the second plurality of concatemeric molecules, the plurality of second sequencing read products are detectable by imaging, and the sequencing comprises decoding the plurality of second sequencing read products from images obtained during 2 to 30 or fewer sequencing cycles, or from images obtained during 1 to 250 sequencing cycles.
[0187] In some embodiments, the batch-specific in situ sequencing method 500 further comprises removing a plurality of second sequencing read products from the second concatemer molecules and retaining the second concatemer molecules in the cell sample. In some embodiments, a 3' blocking moiety can be added to the second sequencing read products to inhibit further sequencing reactions. For example, a nucleotide analog can be incorporated if it inhibits the incorporation of a subsequent nucleotide. Exemplary blocking nucleotide analogs include dideoxynucleotides or nucleotides with 2' or 3' chain terminating moieties.
[0188] In some embodiments, the batch-specific in situ sequencing method 500 further comprises repeatedly sequencing the plurality of second concatemers by repeating at least once: (1) sequencing the second plurality of concatemer molecules in the cellular sample under conditions that inhibit sequencing of the first plurality of concatemers; and (2) removing the plurality of second sequencing read products from the second concatemer molecules and retaining the second concatemer molecules in the cellular sample. In some embodiments, the repeated sequencing is optional.
[0189] In some embodiments, sequencing the second concatemers comprises contacting a second plurality of concatemer molecules in the cellular sample with (i) a plurality of second batch-specific sequencing primers, (ii) a plurality of sequencing polymerases, and (iii) a plurality of nucleotide reagents under conditions suitable for hybridizing the second batch-specific sequencing primers to their respective second batch-specific sequencing primer binding sites on the second concatemers (step (1)). In some embodiments, sequencing further comprises performing 2 to 30 or fewer sequencing cycles using the second concatemers as template molecules to generate a first plurality of sequencing read products. In some embodiments, sequencing further comprises removing the second plurality of sequencing read products from the first concatemers and retaining the second plurality of concatemers in the cellular sample (step (3)). In some embodiments, the sequencing further comprises step (4) repeating steps (1)-(3) at least once (e.g., FIG. 14). In some embodiments, step (4) comprises repeating steps (1)-(3) at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, or at least 10 times. In some embodiments, step (4) comprises repeating steps (1)-(3) up to 10 times, up to 20 times, up to 30 times, up to 40 times, or up to 50 times.
[0190] In some embodiments, the iterative sequencing of the second concatemers in step (i) can be performed using a sequencing-by-binding procedure, labeled and / or unlabeled chain-terminating nucleotides, or multivalent molecules. Descriptions of these three sequencing methods are provided below.
[0191] In some embodiments, the plurality of nucleotide reagents used in the procedure comprise a plurality of nucleotides that are detectably labeled or unlabeled. In some embodiments, each nucleotide is linked to a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the plurality of detectably labeled nucleotide analogs comprise a plurality of chain-terminating nucleotides, where the chain-terminating moiety is linked to the 3' nucleotide sugar position to form a 3' blocked nucleotide analog. In some embodiments, the chain-terminating moiety can be removed to convert the 3' blocked nucleotide analog to an extendable nucleotide bearing a 3' OH group on the sugar. In some embodiments, the labeled nucleotide analogs are linked to different fluorophores corresponding to the nucleobases adenine, cytosine, guanine, thymine, or uracil, where the different fluorophores emit fluorescent signals. In some embodiments, a sequencing cycle includes (1) contacting a concatemer / sequencing primer duplex with a sequencing polymerase and a detectably labeled chain-terminating nucleotide under conditions suitable for polymerase-catalyzed incorporation of the detectably labeled chain-terminating nucleotide onto the end of the sequencing primer; (2) detecting and imaging the fluorescent signal and color emitted by the incorporated chain-terminating nucleotide; and (3) removing the chain-terminating moiety (e.g., unblocking) and fluorophore from the incorporated nucleotide, preserving the concatemer / sequencing primer duplex. In some embodiments, 2 to 30 or fewer sequencing cycles are performed on multiple concatemers in a cellular sample to generate multiple sequencing read products. In some embodiments, the sequence of a first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of a first target RNA molecule in the cellular sample. In some embodiments, the sequence of a second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of a second target RNA molecule in the cellular sample.
[0192] In some embodiments, the sequences of the first and second sequencing read products can be aligned after each round that generates first and second sequencing read products of 30 bases or less in length, or after generating a set of iterative sequencing read products in which the first and second sequencing read products read products of 30 bases or less in length. In some embodiments, the sequencing reaction is performed on a sequencing device having a detector that captures fluorescent signals from the sequencing reaction within the cell sample. The sequencing device can be configured to relay the fluorescent signal data captured by the detector to a computer system programmed to display an image of different fluorescent spots coexisting in the cell sample, each corresponding to a different target RNA molecule. In some embodiments, when sequencing is performed using different fluorescently labeled nucleotide reagents corresponding to different nucleobases (e.g., A, G, C, T / U), the image can have fluorescent spots of different colors coexisting in the same cell sample at different sequencing cycles.
[0193] In some embodiments, asynchronous phase and / or pre-phase events can occur during a synchronous sequencing reaction on clonally amplified template amplicons, where the sequencing reaction involves a polymerase-catalyzed sequencing reaction using detectably labeled chain terminator nucleotides. In some embodiments, the sequencing reaction on one template molecule within the clonally amplified template molecule advances (e.g., pre-phasing) or retards (e.g., phasing) the sequencing of other template molecules within the clonally amplified template molecule. During sequencing, typically, a fluorescent signal corresponding to the incorporation of labeled chain terminator nucleotides is detected. Thus, phasing and pre-phasing events can be detected and monitored using the incorporation of labeled chain terminator nucleotides.
[0194] In some embodiments, the plurality of nucleotide reagents in steps (d) and (g) comprise a plurality of multivalent molecules each comprising a core attached to a plurality of nucleotide arms, wherein the nucleotide arms are attached by nucleotide units. In some embodiments, each multivalent molecule is linked to a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the core of the multivalent molecule is labeled with a fluorophore, wherein the fluorophore attached to a given core of the multivalent molecule corresponds to the nucleotide base of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or nucleotide base attached to a fluorophore, wherein the fluorophore attached to a given nucleotide base corresponds to the nucleotide base of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, a sequencing cycle comprises: (1) contacting the concatemer / sequencing primer duplex with a first sequencing polymerase to form a multiplexed polymerase; (2) contacting the multiplexed polymerase with a detectably labeled multivalent molecule under conditions suitable for binding complementary nucleotide units of the multivalent molecule to the multiplexed polymerase, thereby forming a multivalent binding complex, wherein the conditions are suitable for inhibiting incorporation of complementary nucleotide units onto the end of the sequencing primer; and (3) contacting the multiplexed polymerase with the detectably labeled multivalent molecule under conditions suitable for inhibiting incorporation of complementary nucleotide units onto the end of the sequencing primer. imaging the emitted fluorescent signal and color; (4) removing the first sequencing polymerase and the bound detectably labeled multivalent molecule, retaining the concatemer / sequencing primer duplex; (5) contacting the retained concatemer / sequencing primer duplex with a second sequencing polymerase and an unlabeled chain-terminating nucleotide under conditions suitable for incorporation of an unlabeled chain-terminating nucleotide onto the end of the sequencing primer; and (6) removing (e.g., unblocking) the chain-terminating portion, retaining the concatemer / sequencing primer duplex.In some embodiments, 2 to 30 or fewer sequencing cycles are performed on multiple concatemers within a cellular sample to generate multiple sequencing read products. In some embodiments, the sequence of a first sequencing read product can be determined and aligned with a first reference sequence to confirm the presence of a first target RNA molecule within the cellular sample. In some embodiments, the sequence of a second sequencing read product can be determined and aligned with a second reference sequence to confirm the presence of a second target RNA molecule within the cellular sample. In some embodiments, the sequences of the first and second sequencing read products can be aligned after each round that generates first and second sequencing read products of 30 bases or less in length, or after the first and second sequencing read products generate a set of iterative sequencing read products that read products of 30 bases or less in length. In some embodiments, the sequencing reaction is performed on a sequencing device having a detector that captures fluorescent signals from the sequencing reaction within the cellular sample. The sequencing device can be configured to relay the fluorescent signal data captured by the detector to a computer system that is programmed to display images of different fluorescent spots coexisting in a cell sample, each corresponding to a different target RNA molecule. In some embodiments, the individual cycle time can be achieved in less than 30 minutes. In some embodiments, the field of view (FOV) is 1 mm. 2 It can exceed the limit for large areas (>10mm 2 The cycle time for scanning the ) can be less than 5 minutes.
[0195] 13 is a schematic diagram showing a workflow for generating circularization padlock probes, including generating first and second cDNAs (respectively) from first and second target RNA molecules, and hybridizing first and second padlock probes (respectively) to the first and second cDNA molecules to generate first and second circularization padlock probes (respectively). The first padlock probe contains (i) a first target barcode sequence (Target BC-1) that uniquely identifies the first target RNA, (ii) a first batch-specific sequencing primer binding site (Batch Seq-1) (or its complementary sequence), (iii) a universal binding site for an amplification primer (Universal RCA) (or its complementary sequence), and (iv) a universal binding site for a compaction oligonucleotide (or its complementary sequence). The second padlock probe contains (i) a second target barcode sequence (Target BC-2) that uniquely identifies the second target RNA, (ii) a second batch-specific sequencing primer binding site (Batch Seq-2) (or its complementary sequence), (iii) a universal binding site for an amplification primer (Universal RCA) (or its complementary sequence), and (iv) a universal binding site for a compaction oligonucleotide (or its complementary sequence).
[0196] 14 is a schematic diagram showing a rolling circle and sequencing workflow, which includes generating first and second concatemers by rolling circle amplification using first and second covalently closed circular molecules (respectively). The first and second concatemers are subjected to a first sequencing workflow using a first batch-specific sequencing primer, a sequencing polymerase, and multiple nucleotide reagents. The first concatemers undergo iterative sequencing, but the second concatemers do not. The first and second concatemers are subjected to a second sequencing workflow using a second batch-specific sequencing primer, a sequencing polymerase, and multiple nucleotide reagents. The second concatemers undergo iterative sequencing, but the first concatemers do not.
[0197] Method 500 may include operation 520 of determining pixel intensities for pixels of the first plurality of flow cell images and respective color purities of each of the pixel intensities. Operation 520 may be performed by processor 404 disclosed herein. Each of the pixel intensities may include an intensity of the pixel and / or one or more subpixel intensities of a subpixel corresponding to the pixel. For example, the subpixel may be within a pixel. The respective color purities of each of the pixel intensities may include one or more color purities corresponding to one of the color purities of the pixel and / or one or more subpixel intensities of the corresponding pixel. The respective color purities of each of the individual pixel intensities of the pixel intensities may include respective color purities of one or more color channels.
[0198] In in situ batch-specific sequencing, operation 520 may be repeated for multiple flow cell images corresponding to each batch. For example, if there are four batches, operation 520 may be performed four times, and each time only for flow cell images from the same batch.
[0199] In some embodiments, operation 520 of determining pixel intensity and respective color purity may include determining pixel intensity and respective color purity for each such pixel at a first axial position and determining pixel intensity and respective color purity for each such pixel at a second axial position. The first and second axial positions may be at different predetermined axial positions for generating the flow cell image. For example, the first axial position may be at z=0, and the second axial position may be at an adjacent z level, such as z=0.1, 0.2, or 0.3 μm.
[0200] In some embodiments, a "pixel" can be a 2D spatial element of a digital image (e.g., a flow cell image) known to those skilled in the art. In some embodiments in the context of in situ sequencing, a "pixel" is a 3D spatial element equivalent to a "voxel" known to those skilled in the art. In embodiments where a "pixel" is a 3D spatial element, the flow cell image can still be 2D, but has a depth of field greater than 0.3 um to 2 um. In embodiments where a "pixel" is a 3D spatial element, the flow cell image can still be 2D, but has a depth of field that is 0.5 um, 0.6 um, 0.7 um, 0.8 um, 0.9 um, 1 um, 1.1 um, 1.2 um, 1.3 um, 1.5 um, 1.8 um, 2 um, 2.2 um, 2.5 um, 2.8 um, 3 um, 3.5 um, 3.8 um, 4 um, 4.5 um, 4.8 um, 5 um, 5.5 um, 6 um, 6.5 um, 7 um, 8 um, 9 um or more. The 2D flow cell images can be stacked in a third dimension orthogonal to the image plane to cover a 3D volume. In some embodiments, the flow cell images are in a single 2D plane or multiple 2D planes, with each flow cell image having a depth of focus in the third dimension. In some embodiments, a pixel is two-dimensional. A pixel can be in the image plane. In some embodiments, a pixel is three-dimensional. A pixel can have two dimensions in the image plane and a third dimension parallel to the axial axis (i.e., the z-axis).
[0201] Determining the pixel intensity can include determining each channel intensity in a set of channel intensities, each channel corresponding to a different fluorescence wavelength, and such determination can be based on a comparison of the set of channel intensities at corresponding pixel or one or more sub-pixel locations.
[0202] Determining the color purity of each of the pixel intensities can include determining (1) a constant for the signal corresponding to a particular type of nucleotide base and an optional first constant, and (2) a constant for the total amount of signal for other types of nucleotide bases and an optional second constant. The first and second constants can be different or the same. In some embodiments, determining the color purity of each of the pixel intensities can include, at least in part, operation 330 disclosed herein.
[0203] Method 500 may include an operation 530 of determining a base-calling template including base-calling positions based on the pixel intensities and respective color purity of the pixel intensities determined in operation 520 .
[0204] In some embodiments, a flow cell image herein is acquired at a single axial position. In some embodiments, a flow cell image is acquired at a plurality of predetermined axial positions. The plurality of predetermined axial positions can include 2 to 500 predetermined axial positions. Figure 27B shows three exemplary predetermined axial positions. In some embodiments, the plurality of predetermined axial positions includes 2 to 30 predetermined axial positions. In some embodiments, the plurality of predetermined axial positions includes 2 to 20 predetermined axial positions. Each of the plurality of predetermined axial positions may be spaced 0.1 to 500 um from its adjacent axial position. The predetermined axial positions may be spaced the same distance apart within a range of 0.1 to 500 um. The predetermined axial positions may be spaced different distances apart within a range of 0.1 to 500 um. Each of the plurality of predetermined axial positions may be spaced 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 um from its adjacent axial position.
[0205] The flow cell image may include a first and / or second plurality of flow cell images. In some embodiments, the base calling positions are at the same axial position. In other words, the base calling positions may be 2D. In some embodiments, at least some of the base calling positions are at different axial positions or z-levels. In other words, the base calling positions are 3D. The different axial positions of the base calling positions may be at some or all of the predetermined axial positions for the flow cell image. In some embodiments, at least some of the axial positions of the base calling positions are different from any of the plurality of predetermined axial positions for the flow cell image. In some embodiments, at least some of the different axial positions of the base calling positions are between two adjacent predetermined axial positions of the plurality of predetermined axial positions. The first or second plurality of flow cell images may be from two, three, four, five, or six different color channels.
[0206] The one or more cycles in which flow cell images are acquired to generate a base calling template can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more cycles. The one or more cycles include 1 to 50 cycles. The one or more cycles include 1 to 6 cycles, 2 to 5 cycles, or 2 to 6 cycles.
[0207] Each of the first or second plurality of flow cell images comprises a field of view perpendicular to the axial axis (i.e., z-axis). In some embodiments, the field of view of each of the first or second plurality of flow cell images is identical within the image plane. The field of view of each of the first or second plurality of flow cell images may cover at least a portion of a tile of the flow cell. In some embodiments, the first or second plurality of flow cell images comprises the same image resolution. The axial axis may extend from the objective lens to the support. The axial axis is perpendicular to the image plane, and the field of view is within the image plane.
[0208] The density of the multiple concatemer molecules on the support is 1 mm 2 10 per 1 ~10 12 10. The method of any one of the preceding claims, wherein:
[0209] The density of the multiple concatemer molecules on the support is 1 mm 2 10 per 2 ~10 8 10. The method of any one of the preceding claims, wherein:
[0210] As disclosed herein, operation 530 of determining a base calling template based on pixel intensities and the color purity of each of the pixel intensities may include determining, for each pixel or subpixel, whether the respective color purity is greater than the color purity of other pixels or subpixels within a distance threshold. In response to determining that the respective color purity is greater than the color purity of other pixels or subpixels within the distance threshold, adding or confirming the location of the corresponding pixel or subpixel in the base calling template. In response to determining that the respective color purity is not greater than the color purity of other pixels or subpixels within the distance threshold, making no changes to the base calling template and proceeding to the next pixel, and repeating the determining operation until all pixels or at least a subset of pixels have passed similar determining operations.
[0211] In some embodiments, operation 530 is configured to remove other light signals that interfere with the intensity of the focused polony, such as overlapping polony, out-of-focus polony, and / or background signals from cellular components.
[0212] The distance threshold can be customized to optimize the effectiveness of rejecting interfering signals (e.g., out-of-focus and / or overlapping polonies) while preserving weaker or larger polonies that are neither overlapping nor out-of-focus. The distance threshold can be determined as the distance between the centers of two polonies. In some embodiments, the distance threshold can be determined between the centers of two pixels or subpixels. When the two polonies are in a single 2D plane, the distance threshold is 2D. In embodiments where the two pixels or subpixels are in multiple 2D planes or in three 3D spaces (e.g., in situ sequencing), the distance threshold is 3D. For example, as shown in Figure 27B, polonie p2 at pixel (x1, y1, z2) may have a second polonie p1 at pixel (x1, y1, z1) and a third polonie p3 at pixel (x1, y1, z3) that are within the 3D distance threshold. The 3D distance threshold can determine a cylinder around polonie p2. z1, z2, and z3 are predetermined axial positions for acquiring flow cell images and are approximately 2 μm apart. Polony p1' is outside the 3D distance threshold. Of the three polonies within the distance threshold, p1 has the lowest purity, and as a result, p1 can be eliminated as either a duplicated polonie or an out-of-focus polonie. The base-calling position is determined to be between the predetermined axial positions z3 and z2 and z3_1. The axial position z3_1 can be determined by linear interpolation or weighting based on the respective purity of multiple polonies, for example, p2 and p3. If p2 has a higher color purity than p1, the axial position z3_1 can be close to p2. In this embodiment, each pixel is 3D and may include a thickness on the axial axis. In some embodiments, if two or more polonies are within the distance threshold along an axial position, only one of them is selected. The unique polonie can be selected based on weighting, interpolation, averaging, or various statistical or mathematical functions.
[0213] The 3D distance threshold may include one, two, three, or more distance elements, each corresponding to a distance in the x, y, or z direction. For example, the 3D distance threshold may include three identical distance elements in the x, y, and z directions, such that the 3D distance threshold determines a spherical region, and a polony within the sphere falls within the distance threshold. As another example, the distance element in the z direction may be different from the distance element in the x or y direction, such that a polony within a cylinder, ellipse, or sphere falls within the distance threshold. The distance threshold may include various numbers of elements (e.g., in different non-Cartesian coordinate systems) that can be converted to three distance elements in the x, y, or z directions in a Cartesian coordinate system, as shown in FIGS. 27A-B.
[0214] In some embodiments, the distance threshold can be customized based on the image resolution in the x, y, and / or z directions. In some embodiments, the distance threshold can be customized based on the image resolution in the x, y, and / or z directions and the size of the polony or cluster. The image resolution in the z direction can be the distance between flow cell images at two adjacent z levels. For example, the flow cell images at two adjacent z levels can be spaced 1 um to 10 um apart from each other, and the z resolution can be determined as the gap between them.
[0215] In some embodiments, the 3D distance threshold comprises a first element distance along the axial axis (i.e., z-axis) and a second element distance in a plane orthogonal to the axial axis. In some embodiments, the 3D distance threshold comprises a first element distance along the axial axis (i.e., z-axis) and a second element distance and a third element distance in a plane orthogonal to the axial axis. In some embodiments, the first element distance is different from the second element and / or the third element distance. In some embodiments, the first element distance is the same as the second element distance and / or the third element distance.
[0216] The operation of determining the base calling template may include, at least in part, the operation of 340.
[0217] After determining the base calling template, the flow cell image, for example, a second plurality of flow cell images of the flow cell device at one or more cycles following one or more cycles, can be configured to be registered. In some embodiments, the second plurality of flow cell images are of the in situ sample at one or more cycles following the cycle for generating the base calling template. The second plurality of flow cell images can be acquired at multiple axial positions.
[0218] In some embodiments, the method 500 may further include generating a second plurality of flow cell images, which may be generated in one or more cycles subsequent to the one or more cycles corresponding to the generation of the base calling template.
[0219] In some embodiments, method 500 can further comprise registering or aligning a second plurality of flow cell images from one or more subsequent sequencing cycles with the base calling template. In some embodiments, registering the second plurality of flow cell images can include generating coordinates of polonies in the second plurality of flow cell images in a common coordinate system. The base calling template is also in the common coordinate system. In some embodiments, registering the second plurality of flow cell images can utilize multiple transformations corresponding to subtiles of the flow cell image (e.g., at each individual axial position or in a 2D plane) to estimate the image transformation of the entire flow cell image. Registering the flow cell images can include generating a transformation of the subtile that provides an estimate of the image transformation of the flow cell image in 2D. Information in neighboring subtiles can be used when determining each individual transformation of the subtile. In some embodiments, if translation along the axial axis is negligible (e.g., 0.05 um, 0.02 um, 0.001 um, 0.0001 um, or more), the 2D translation can reliably register or align the flow cell image with the base-calling template. In some embodiments, if translation along the axial axis needs to be included (e.g., 0.05 um, 0.02 um, 0.001 um, 0.0001 um, or more), the 2D translation in the xy plane can be combined with a shift along the axial axis (e.g., the z axis) to register or align the second flow cell image with the base-calling template. The translation along the axial axis can be reliably represented by a shift based on the physical properties of the optical system.
[0220] In some embodiments, the operation of registering the second plurality of flow cell images can utilize multiple 3D transformations (e.g., at multiple individual axial positions or 2D planes) corresponding to subvolumes of the volumetric sample to estimate the image transformation of the flow cell image. The operation of registering the flow cell images can include generating a transformation that provides an estimate of the image transformation of the flow cell image in 3D. Information in neighboring volumes or subvolumes can be used when determining the individual transformations.
[0221] In some embodiments, the coordinates of the polony may be stored in a one-dimensional vector or list, where each entry in the vector or list may contain a unique identification of the polony, its coordinates, and other relevant information such as pixel intensities in one or more channels of the cycle.
[0222] In some embodiments, method 500 can further include performing base calling of the second plurality of flow cell images at base calling positions within the base calling template using signals from the aligned second plurality of flow cell images. In some embodiments, method 500 can further include performing base calling of the second plurality of flow cell images only at some or all of the base calling positions within the base calling template. In other words, the base calling template serves as a polony map that identifies all accurate polony positions on the flow cell and ignores other signal positions that are unlikely to be polony, so as to avoid base calls from overlapping polony or other interfering signals. Base calls in the primary analysis herein can be based on aligned image intensities in a common coordinate system. As disclosed herein, base calling positions can be at sub-pixel positions in the x, y, and / or z directions.
[0223] Performing base calling on the second plurality of flow cell images can include, at least in part, operation 350 of FIG. 3 .
[0224] In some embodiments, base-calling positions in the base-calling template that are not at any of the predetermined axial positions of the flow cell image can be obtained as subpixel base-calling positions in the xy plane. In some embodiments, interpolation along the axial axis can be used to determine the subpixel positions. The subpixel positions along the axial axis can be any axial position between two adjacent z levels of the flow cell image. Interpolation along the axial axis can use any of the interpolation functions disclosed for 2D interpolation in the xy plane. In some embodiments, various other interpolation functions can be used. In some embodiments, linear interpolation along the axial axis can be used based on the optical properties of the microscope. In some embodiments, interpolation is not used, and the base-calling positions can be at predetermined axial positions of the flow cell image. In some embodiments, the predetermined axial positions can be adjusted for the size of the cell sample so that interpolation is not required.
[0225] The subpixel intensities herein may have 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 17, 18, 19, or 20 different axial positions along the axial axis. The subpixel intensities herein may have 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 17, 18, 19, or 20 different axial positions between two adjacent axial positions of the plurality of predetermined axial positions. The subpixel intensities herein may have 2 to 15 different axial positions along the axial axis. The subpixel intensities herein may have 1 to 15 different axial positions between two adjacent axial positions of the plurality of predetermined axial positions. Each of the subpixel intensities may be along the axial axis and have 1 to 80 different axial positions between two adjacent axial positions of the plurality of predetermined axial positions. In some embodiments, the sub-pixel resolution along the axial direction may be based on the size of the cell sample and the density of concatemers along the axial axis.
[0226] The density of concatemer molecules on the support is approximately 10 per mm. 1 ~10 12 The density of the concatemer molecules on the support can be 1 mm 2 Approximately 10 per 2 ~10 8 The density of the concatemer molecules on the support can be 1 mm 2 Approximately 10 per 4 ~10 8 The density of the concatemer molecules on the support can be 1 mm 2 Approximately 10 per 2 ~10 5 The density of the concatemer molecules on the support can be 1 mm 2 10 per 1 ~10 12 The density of the concatemer molecules on the support can be 1 mm 2 10 per 2 ~10 8 The density of the concatemer molecules on the support can be 1 mm 2 10 per 4 ~10 8 The density of the concatemer molecules on the support can be 1 mm 2 10 per 2 ~10 5 It could be.
[0227] In some embodiments, method 500 may include operation 520' instead of operation 520. Operation 520' may include identifying base calling locations within the first plurality of flow cell images by using a spot finding algorithm based on pixel intensities and the color purity of each of the pixel intensities. Differences between operations 520 and 520' are described in connection with method 300. In some embodiments, operation 520 may determine base calling locations among the candidate base calling locations by treating some or all pixels of the flow cell image as candidate base calling locations and adding each determined base calling location to a base calling template. In some embodiments, operation 520' includes inputting the optical signals of the flow cell image into a spot finding algorithm, such that the spot finding algorithm outputs a set of potential candidate base calling locations. The spot finding algorithm may function in 3D, such that it can acquire flow cell images from multiple z-levels and output a set of potential candidate base calling locations. The set of potential candidate base calling locations may be at least two different z-locations.
[0228] Sequencing results and cell feature data files Conventional data files of sequencing results, such as FastQ files, generally contain base calls and their corresponding quality scores. For example, in situ cell or tissue sequencing requires more challenging and complex sequencing systems and sequencing analysis than conventional 2D sequencing. Sequencing of in situ samples relies on conventional spatial biology to determine the spatial information of sequencing results. Currently, sequencing data is collected using one device, such as a sequencing system, while spatial biology data may be collected using a different device, such as a conventional microscope with various cell stains. Therefore, spatial information and sequencing results are stored in different data file formats, thereby requiring multiple software and / or processes to analyze the sequencing results in relation to the spatial information. Furthermore, images containing sequencing and cellular features may be stored in complex data formats, which can pose storage space challenges. To facilitate in situ sequencing analysis, there is a need for spatial biology data that seamlessly integrates with sequencing data.
[0229] Embodiments of the present disclosure provide data files. The data files herein can include sequencing results. The data files herein can include multiple base calls, similar to conventional data files, such as conventional FastQ files. Base calls can be generated from a primary analysis of a flow cell image at a single axial position (i.e., z-position). Figure 28 shows a simulated flow cell image (top right) of an in situ cell with a polony as a small bright spot (e.g., the size of two pixels). In some embodiments, base calls can be obtained from sequencing analysis of a flow cell image at multiple axial positions or in 3D space.
[0230] In embodiments of the present disclosure, the data files and corresponding formats described herein advantageously contain spatial biological information content and can compress complex image datasets into a single data format. The data files disclosed herein can advantageously contain information specific to 3D sequencing, making such data files convenient and more efficient for 3D sequencing analysis. Other advantages of the systems and methods disclosed herein include, but are not limited to, enabling efficient reconstruction of 3D cellular features, such as cell or tissue morphology, from data files; simplifying biomarker information in data files and their data processing; enabling easier and more convenient data learning from a single file type, allowing seamless integration of information from various assay types; and enabling data compression while preserving important information content across different assay types.
[0231] In some embodiments, base calls are generated from polonies in one or more sequencing cycles, hi some embodiments, base calls are generated from polonies in short sequencing reads that include only 1 to 50 cycles, 2 to 30 cycles, 2 to 20 cycles, or 2 to 60 cycles.
[0232] In some embodiments, the base call may include one or more base calls of the barcode sequence(s). In some embodiments, the base call may include one or more base calls of the index sequence(s). In some embodiments, the base call may include one or more base calls of the insert sequence(s) of interest. In some embodiments, the base call may include only base calls from the barcode sequence(s). In some embodiments, the base call may include only base calls from the index sequence(s). In some embodiments, the base call lacks a base of the insert sequence(s) of interest.
[0233] The data files herein may include quality indicators corresponding to some or all of the base calls. The quality indicators may include quality scores for the base calls.
[0234] The data file may include spatial coordinates of the base calls on the support. In addition, the data file may include data of one or more cellular features. The cellular features may be of a cellular sample immobilized on the support. The base calls may be generated from sequencing of the cellular sample immobilized on the support. The data for the cellular features may be obtained from the same in situ sample sequenced to generate the base calls in the data file. The data for each of the one or more cellular features may include a cellular feature indicator. The cellular feature indicator may be configured to indicate a cellular element selected from a nucleus, a membrane, a cytoplasm, and a mitochondria. The cellular feature indicator may be configured to indicate a cellular element selected from a nuclear boundary and a membrane boundary. For example, "m" may indicate a mitochondrion or a membrane, "c" may indicate a cytosol, and "n" may indicate a nucleus. The data file may include spatial coordinates of the cellular features on the support. The spatial coordinates of the base calls and the spatial coordinates of the cellular features may be in the same coordinate system. The spatial coordinates of each of the cellular features may include multiple sets of spatial coordinates. Data for one or more cell features may be divided into multiple entries, with each entry in the multiple entries corresponding to a spatial location on the support. The spatial location may be 2D or 3D. Data for one or more cell features may be divided into multiple entries, with each entry in the multiple entries corresponding to a set of spatial coordinates. The set of spatial coordinates may include an x-coordinate, a y-coordinate, a z-coordinate, or a combination thereof. Each set of spatial coordinates is configured to indicate a unique location in 2D or 3D. For example, the cell feature may be the cell membrane(s) of one or more cells. Each pixel in the image that is on the cell membrane may be an entry, with all the multiple entries together determining the 3D spatial location of the cell membrane(s). Although 2D and / or 3D Cartesian coordinate systems are used to define spatial locations in 2D or 3D space herein, various coordinate systems may be used to define spatial locations and relationships herein. Some exemplary coordinate systems other than Cartesian systems may include, but are not limited to, a polar coordinate system, a cylindrical coordinate system, or a spherical coordinate system. The other coordinate systems may include homogeneous or non-homogeneous coordinate systems.
[0235] The spatial coordinates of the base call or the cellular feature can be 3D. The spatial coordinates of the base call and the spatial coordinates of the one or more cellular features can include x coordinates, y coordinates, z coordinates, or a combination thereof.
[0236] In some embodiments, the base call, the quality indicator, the one or more cellular features, the spatial coordinate of the base call, or the spatial coordinate of the cellular feature comprises one or more letters, one or more numbers, one or more symbols, or a combination thereof.
[0237] The data file may be a text file. The data file may be a FastQ file. The data file may be encoded using various encoding formats for text files. The data file may be ASCII encoded. The data file may be in an 8-bit, 12-bit, 16-bit, 18-bit, 32-bit, or 36-bit encoding format.
[0238] In some embodiments, the data file includes one or more sequencing parameters and their values corresponding to cell features. The one or more sequencing parameters may include a target sequence. The one or more sequencing parameters may include a sequencing system name. The one or more sequencing parameters may include a sequencing run identification. The one or more sequencing parameters may include a flow cell identification. The one or more sequencing parameters may include a flow cell lane number or lane identification. The one or more sequencing parameters may include a tile number or subtile number.
[0239] In some embodiments, the data file includes one or more delimiters. Each of the one or more delimiters may be configured to separate one or more cellular features and / or separate base calls from one or more cellular features. Each of the one or more delimiters may include letters, numbers, symbols, or combinations thereof. In some embodiments, data for one or more cellular features is separated by one or more delimiters based on the spatial coordinates of the cellular features. In some embodiments, each entry of a plurality of entries for one or more cellular features is separated by one or more delimiters.
[0240] In some embodiments, the data files herein are generated by a next-generation sequencing (NGS) system and include a data file containing base calls, such as a conventional FastQ file, and include a data format that is the same as or compatible with that of the data file. In some embodiments, at least a portion of the data files herein can be analyzed or processed by software or a computer program configured to process conventional sequencing data files, such as FastQ files. For example, the sequencing result portion of the data files herein, including at least base calls and optionally their corresponding quality scores, can be analyzed or processed by software or a computer program configured to process conventional sequencing data files, such as FastQ files.
[0241] In some embodiments, for an individual cell sample, e.g., a sample having multiple in situ cells immobilized on a flow cell, the sequencing results of the sample can be included in a first data file, e.g., a conventional FastQ file, and the corresponding cellular feature data of the sample can be included in the same first data file or in a second, different data file. The second data file can still have a compatible or identical data format to the first data file, making it easy for users to combine the information in the first and second data files without having to reformat, convert, or otherwise process the data file(s). For example, the conventional data of cellular features can be in JPEG, PNG, TIFF, or other digital image encoding format, ensuring data processing to extract the spatial coordinates of the cellular features. Furthermore, the extracted spatial coordinates may need to be converted to 8-bit or other encoding compatible with conventional FastQ. The data files and formats disclosed herein can advantageously contain spatial biological information and in situ sequencing results. The data files disclosed herein can advantageously include information specific to in situ sequencing, so that such data files can be conveniently and more efficiently utilized for in situ sequencing analysis that combines in situ sequencing information with corresponding cellular features. Other advantages of the systems and methods disclosed herein include, but are not limited to, enabling efficient reconstruction of in situ cell or tissue morphology from data files, simplifying biomarker information in data files and their data processing, enabling easier and more convenient data learning from a single file type, allowing seamless integration of information from various assay types, and enabling significant data compression while preserving significant information content across different assay types.
[0242] In some embodiments, the data file is configured for reconstruction of an image including one or more cellular features and base calling results, where the one or more cellular features and base calling results are aligned or registered with each other. Figure 28 shows an exemplary reconstructed image (bottom right) aligning the cell membrane and nuclear membrane segmentation with sequencing results. Thus, the data file can advantageously enable users of conventional sequencing results, e.g., base calling, to perform sequencing analysis based on cellular features.
[0243] Figure 29A shows an exemplary cell and some of its structural elements. Unique cell feature indicators can be used in the data files disclosed herein to represent the nucleus, membrane, cytoplasm, and mitochondria. The data files disclosed herein can also include one or more first parameters and their values associated with a sequencing system and one or more second parameters and their values associated with imaging parameters. Such first and second parameters and their values can be used to uniquely identify sequencing hardware, software, or other sequencing-related processes. As shown in Figure 29B, the letter "n" and its ASCII code can be used to indicate that the base calling result following the letter "n" was obtained from the cell's nucleus. The letter "m" can indicate that the information contained after this letter pertains to cell morphology. Such information can end when the next separate "@" appears, as shown in Figure 29B.
[0244] For each cell feature indicator, the data file may further include sub-feature indicators and may further include separate features within the same cell feature indicator. For example, as shown in FIG. 29B, the cell feature indicator for cytoplasm is "c." For different elements within the cytoplasm, such as proteins, the data file may include sub-feature indicators. The sub-feature indicators may be different letters or various unique combinations of letters and numbers. For example, the data file in FIG. 29B uses protein markers that are different nucleotide combinations. The nucleotide combinations may be variable or fixed length, e.g., 5. Similarly, the data file in FIG. 29B uses transcription markers that are different nucleotide combinations.
[0245] An exemplary list of indicators that the data file may use is listed in Figure 29 A. In some embodiments, each cell feature indicator, parameter, or parameter value disclosed herein can include one or more letters, one or more numbers, one or more symbols, or a combination thereof.
[0246] The data file may also include one or more delimiters. Each delimiter can be used to separate two different entries in the data file. For example, the conventional delimiter "+" can be used to separate the base calling results and the quality scores, as shown in Figure 29B. As another example, delimiters can include "@", ":", and one or more spaces, as shown in Figure 29B.
[0247] The data file may further include an index array as shown in Figures 29A-29B.
[0248] Disclosed herein is a method for generating a data file, which may include some or all of the operations disclosed herein.
[0249] The method for generating a data file may include some or all of the operations disclosed herein, and the operations may be performed in the order described herein, but are not limited to this.
[0250] The method for generating a data file may be performed by one or more processors (e.g., 404 of FIG. 4 ) disclosed herein. In some aspects, the processor may include one or more of a processing unit, an integrated circuit, or a combination thereof. For example, the processing unit may include a central processing unit (CPU) and / or a graphics processing unit (GPU). The integrated circuit may include a chip such as a field programmable gate array (FPGA). In some aspects, the processor may include computing system 400.
[0251] In some embodiments, some or all of the operations in the method of generating a data file may be performed by FPGA(s). In embodiments in which some operations are performed by FPGA(s), data after operations performed by the FPGA(s) may be communicated by the FPGA(s) to the CPU(s), so that the CPU(s) can use such data to perform subsequent operation(s) in the method of generating a data file. Similarly, data may also be communicated from the CPU(s) to the FPGA(s) for processing by the FPGA(s). In some embodiments, all of the operations in the method of generating a data file may be performed by CPU(s). Alternatively, the operations performed by the CPU(s) may be performed by a dedicated processor or other processor, such as GPU(s). In some embodiments, all of the operations in the method of generating a data file may be performed by FPGA(s). In some embodiments, some or all of the operations in the method of generating a data file are performed during or before sequencing cycle M in a sequencing run. In some embodiments, M is less than 5, 10, 20, 30, 40, 50, or 60.
[0252] The operations may be performed in the order described herein, but are not limited to this. The method for generating a data file herein may include generating one or more images containing one or more cellular features using an optical system. The one or more images may include optical signals from one or more cellular features. The one or more images may include increased contrast of one or more cellular features. For example, the one or more images may be microscopic images of staining of a cell sample on a support. In some embodiments, the one or more images are acquired after a sequencing read is completed. In some embodiments, the one or more images are acquired after a flow cell image for generating base calls using a sequencing system is acquired, as shown in FIG. 1.
[0253] FIG. 28 shows a simulated image (top center) of a cell sample with simulated staining of the cells and nuclei.
[0254] The method for generating the data file may include determining the spatial coordinates of each of one or more cellular features in one or more images. The determining operation may use various image processing algorithms for segmentation, contour tracing, etc. Figure 28 shows an exemplary segmentation (top left) derived from an image showing cellular features such as the nucleus and cell membrane (top center). The spatial coordinates can be determined using the segmentation image.
[0255] The method for generating a data file herein may include generating data for one or more cellular features based on the determined spatial coordinates. For each of the one or more cellular features, generating data based on the determined spatial coordinates includes generating a cellular feature indicator and inserting the spatial coordinates of the cellular feature. Such generating and inserting operations may be for multiple entries of the cellular feature, each entry corresponding to a non-repeated spatial location in 2D or 3D. Such generating and inserting operations may be iterative to include some or all of the multiple entries of a particular cellular feature. Figure 28 shows two exemplary entries for a boundary, one from the membrane at lines 1 and 134549, respectively, and the other from the nuclear boundary. An exemplary format for including entries is also shown in Figure 28: Boundary format: @Instrument:RunID:FlowcellID:Lane:Tile:X:Y:_:Paired:Filter:ControlBits:Feature(n or m):_
[0256] The data file may contain a data format that is the same as or compatible with a data file generated by a next-generation sequencing (NGS) system and that contains base calls.
[0257] The methods for generating data files herein may include operations for generating data for one or more cellular features based on the determined spatial coordinates in a data format identical to or compatible with the data format generated by a next-generation sequencing (NGS) system, including base calls. In some embodiments, such operations for generating data for one or more cellular features include performing several iterations of the operations for each of the one or more cellular features, including (i) inserting a cellular feature indicator for the cellular feature, (ii) inserting a set of spatial coordinates for the cellular feature, (iii) inserting one or more sequencing parameters and their values corresponding to the cellular feature, and (iv) inserting at least one of one or more delimiters, or a combination thereof. In some embodiments, such operations for generating data for one or more cellular features include performing several iterations of the operations for each of the one or more cellular features, including (v) inserting a base call, (vi) inserting a quality indicator corresponding to the base call, and (vii) inserting the spatial coordinate of the base call, or a combination thereof. Figure 28 shows an exemplary entry of a base call from at least a portion of a barcode sequence in lines 249307-249309. "PACRG" is the target protein / gene of the barcode. An exemplary barcode format may be @Instrument:RunID:FlowcellID:Lane:Tile:X:Y:Z:Paired:Filter:ControlBits:Feature(n or c):AssignedTarget.
[0258] The quality score may be inserted in the subsequent line after the delimiter "+".
[0259] In some embodiments, the method for generating a data file herein may include generating, by a sequencing system, a first plurality of flow cell images of a cell sample immobilized on a support, similar to operation 510.
[0260] In some embodiments, the methods for generating a data file herein may include generating base calls based on a first plurality of flow cell images of a cell sample as disclosed herein.
[0261] In some embodiments, the exact format for inserting barcodes or other base calls and the exact format for inserting boundaries or cellular features include, but are not limited to, the exemplary formats disclosed herein, and can be customized based on a variety of factors, including, for example, the type of cell sample being sequenced, the coordinate system used to define spatial locations, the characteristics of the flow cell, etc.
[0262] Computer Systems Various aspects may be implemented using one or more computer systems, such as, for example, computer system 400 shown in Figure 4. One or more computer systems 400 may be used, for example, to implement any of the aspects discussed herein, as well as combinations and subcombinations thereof.
[0263] Computer system 400 may include one or more processors (also referred to as central processing units, or CPUs), such as processor 404. Processor 404 may be connected to a bus or communication infrastructure 406.
[0264] Computer system 400 may also include user input / output device(s) 403, such as a monitor, keyboard, pointing device, etc., that may communicate with a communications infrastructure 406 via user input / output interface(s) 402. User input / output device 403 may be coupled to user interface 124 of FIG.
[0265] One or more of the processors 404 may be a graphics processing unit (GPU). In one aspect, a GPU may be a processor that is a specialized electronic circuit designed to handle mathematically intensive applications. Due to the power of general-purpose computing on a graphics processing unit (GPGPU), GPUs may be particularly useful in at least the image recognition and machine learning aspects described herein.
[0266] Additionally, one or more of the processors 404 may include a coprocessor or other implementation of logic for accelerating cryptographic calculations or other specialized mathematical functions, including hardware-accelerated cryptographic coprocessors. Such accelerated processors may further include instruction set(s) for acceleration using the coprocessor and / or other logic to facilitate such acceleration.
[0267] Computer system 400 may also include a main or primary memory 408, such as random access memory (RAM). Main memory 408 may include one or more levels of cache. Main memory 408 may store control logic (i.e., computer software) and / or data therein.
[0268] Computer system 400 may also include one or more secondary storage or secondary memories 410. Secondary memories 410 may include, for example, a main storage drive 412 and / or a removable storage device or drive 414. Main storage drive 412 may be, for example, a hard disk drive or a solid state drive. Removable storage drive 414 may be, for example, a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.
[0269] The removable storage drive 414 may interface with a removable storage unit 418 .
[0270] Removable storage unit 418 may include a computer-usable or computer-readable storage device that stores computer software (control logic) and / or data. Removable storage unit 418 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / or any other computer data storage device. Removable storage drive 414 may read from and / or write to removable storage unit 418.
[0271] Secondary memory 410 may include other means, devices, components, instruments, or other approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 400. Such means, devices, components, equipment, or other approaches may include, for example, removable storage unit 422 and interface 420. Examples of removable storage unit 422 and interface 420 may include a program cartridge and cartridge interface (such as found in a video game device), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other removable storage unit and associated interface.
[0272] Computer system 400 may further include a communications or network interface 424. Communications interface 424 may enable computer system 400 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referenced by reference numeral 428). For example, communications interface 424 may enable computer system 400 to communicate with external or remote devices 428 via communications path 426, which may be wired and / or wireless (or a combination thereof) and may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 400 via communications path 426. In some aspects, communications path 426 is a connection to cloud 130, as shown in FIG. 1 . The external devices, etc. referenced by reference numeral 428 may be devices, networks, entities, etc. within cloud 130.
[0273] Computer system 400 may also be any of a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet, a smartphone, a smartwatch or other wearable, an appliance, part of the Internet of Things (IoT), and / or an embedded system, to name a few non-limiting examples, or any combination thereof.
[0274] It should be understood that the framework described herein may be implemented as a method, process, apparatus, system, or article of manufacture, such as a non-transitory computer-readable medium or device. For illustrative purposes, the framework may be described in the context of a distributed ledger being publicly available, or at least available to untrusted third parties. A current use case includes a blockchain-based system. However, it should be understood that the framework may also be applied to other settings where sensitive or confidential information may need to pass through the hands of untrusted third parties, and that the technology is in no way limited to the use of distributed ledgers or blockchains.
[0275] Computer system 400 may be a client or server accessing or hosting any application and / or data through any delivery paradigm, including, but not limited to, remote or distributed cloud computing solutions, local or on-premise software (e.g., an “on-premise” cloud-based solution), an “as a service” model (e.g., Content as a Service (CaaS), Digital Content as a Service (DCaaS), Software as a Service (SaaS), Managed Software as a Service (MSaaS), Platform as a Service (PaaS), Desktop as a Service (DaaS), Framework as a Service (FaaS), Backend as a Service (BaaS), Mobile Backend as a Service (MBaaS), Infrastructure as a Service (IaaS), Database as a Service (DBaaS), etc.), and / or a hybrid model including any combination of the above examples or other service or delivery paradigms.
[0276] Any applicable data structures, file formats, and schemas may be derived from standards, including, but not limited to, JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or other functionally similar representations, alone or in combination. Alternatively, proprietary data structures, formats, or schemas may be used exclusively or in combination with known or open standards.
[0277] Any associated data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in a human-readable format, such as a numeric, textual, graphical, or multimedia format, further including various types of markup languages, among other possible formats. Alternatively, or in combination with the above formats, the data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in binary, coded, compressed, and / or encrypted format, or any other machine-readable format.
[0278] The interfaces or interconnections between the various systems and layers may use any number of mechanisms, such as any number of protocols, programmatic frameworks, floorplans, or application programming interfaces (APIs), including, but not limited to, the Document Object Model (DOM), Discovery Service (DS), NSUserDefaults, Web Services Description Language (WSDL), Message Exchange Patterns (MEP), Web Distributed Data Exchange (WDDX), Web Hypertext Applications Technology Working Group (WHATWG), HTML5 Web Messaging, Representational State Transfer (REST or RESTful Web Services), Extensible User Interface Protocol (XUP), Simple Object Access Protocol (SOAP), XML Schema Definition (XSD), XML Remote Procedure Call (XML-RPC), or any other mechanism that can achieve similar functionality and results.
[0279] Such interfaces or interconnections may also utilize Uniform Resource Identifiers (URIs), which may further include Uniform Resource Locators (URLs) or Uniform Resource Names (URNs). Other forms of uniform and / or unique identifiers, locators, or names may be used exclusively or in combination with forms such as those described above.
[0280] Any of the above protocols or APIs may interface with or be implemented in, and may be compiled or interpreted in, any programming language, procedural, functional, or object-oriented, including, but not limited to, C, C++, C#, Objective-C, Java, Scala, Clojure, Elixir, Swift, Go, Perl, PHP, Python, Ruby, JavaScript, WebAssembly, or virtually any other language (with any other libraries or schema, with any kind of framework, runtime environment, virtual machine, interpreter, stack, engine, or similar mechanism (including, but not limited to, Node.js, V8, Knockout, jQuery, Dojo, Dijit, OpenUI5, AngularJS, Expressjs, Backbone.js, Ember.js, DHTMLX, Vue, React, Electron, etc., among many other non-limiting examples)).
[0281] In some aspects, a tangible, non-transitory apparatus or article of manufacture comprising a tangible, non-transitory computer-usable or readable medium having control logic (software) stored thereon may be referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 400, main memory 408, secondary memory 410, and removable storage units 418 and 422, as well as tangible articles of manufacture embodying any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 400), may cause such data processing devices to operate as described herein.
[0282] Based on the teachings contained herein, it will be apparent to one skilled in the relevant art(s) how to make and use aspects of the present disclosure using data processing devices, computer systems, and / or computer architectures other than those shown in Figure 4. In particular, aspects may operate with software, hardware, and / or operating system implementations other than those described herein.
[0283] It is understood that the Detailed Description section, and not the other sections, is intended to be used to interpret the claims, which may describe one or more example aspects as contemplated by the inventor(s), but are not all, and are therefore not intended to limit the scope of the disclosure or the appended claims in any way.
[0284] While this disclosure describes exemplary embodiments for exemplary fields and applications, it should be understood that the disclosure is not limited thereto. Other embodiments and modifications thereof are possible and are within the scope and spirit of the present disclosure. For example, without limiting the generality of this paragraph, the embodiments are not limited to the software, hardware, firmware, and / or entities shown in the drawings and / or described herein. Moreover, the embodiments (whether or not explicitly described herein) have significant utility for fields and applications beyond the examples described herein.
[0285] Aspects are described above using functional building blocks that illustrate implementations of particular functions and relationships thereof. The boundaries of these functional building blocks are arbitrarily defined herein for convenience of description. Alternative boundaries may be defined so long as the specified functions and relationships (or equivalents thereof) are appropriately performed. Also, alternative aspects may implement functional blocks, steps, operations, methods, etc. using an order different from that described herein.
[0286] References herein to "one embodiment," "embodiment," "exemplary embodiment," "some embodiments," or similar phrases indicate that the described embodiment may include a particular feature, structure, or characteristic, but not all embodiments may necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an embodiment, it would be within the knowledge of one of ordinary skill in the relevant art(s) to incorporate such feature, structure, or characteristic into other embodiments, whether or not explicitly mentioned or described herein.
[0287] Additionally, some aspects may be described using the terms "coupled" and "connected," along with their derivatives. These terms are not necessarily intended as synonyms for each other. For example, some aspects may be described using the terms "connected" and / or "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other, but yet still cooperate or interact with each other.
[0288] In some embodiments, when sequencing using detectably labeled multivalent molecules, step (2) in which a multivalent binding complex is formed, and step (3) in which the bound, detectably labeled multivalent molecules are imaged and detected, the conditions are milder than those for sequencing workflows using detectably labeled chain-terminating nucleotides. For example, steps (2) and (3) can be performed at a moderate temperature of about 35-45°C or about 39-42°C. Steps (2) and (3) can be performed at a moderate temperature, which can help maintain the compact size and shape of the DNA nanoballs during multiple sequencing cycles (e.g., up to 30 cycles), which can improve the FWHM (full width at half maximum) of the DNA nanoball spot images within the cell sample. In some embodiments, the DNA nanoballs do not resolve during multiple sequencing cycles. In some embodiments, the DNA nanoball spot images do not expand during multiple sequencing cycles. In some embodiments, the DNA nanoball spot images remain discrete spots during multiple sequencing cycles. The spot image can be represented as a Gaussian spot, and the size can be measured as FWHM. A smaller spot size, indicated by a smaller FWHM, typically correlates with an improved image of the spot. In some embodiments, the FWHM of the nanoball spot can be about 10 μm or less.
[0289] In some embodiments, asynchronous phasing and / or prephasing events can occur during a synchronous polymerase-catalyzed sequencing reaction using detectably labeled multivalent molecules. During sequencing, a fluorescent signal corresponding to the binding of the complementary nucleotide units of the multivalent molecules to the hybridized polymerase can be detected, thereby forming a multivalent binding complex. Thus, phasing and prephasing events can be detected and monitored using the binding of labeled multivalent molecules. In some embodiments, when performing up to 30 sequencing cycles using detectably labeled multivalent molecules, the phasing and / or prephasing rate can be less than about 5%, or less than about 1%, or less than about 0.01%, or less than about 0.001%. In contrast, when performing up to 30 sequencing cycles using labeled chain terminator nucleotides, the phasing and / or prephasing rate can be about 5%.
[0290] In any of the methods described herein, a plurality of RNAs or cDNAs in a cellular sample can be amplified to generate RNA or cDNA amplicons, where the amplicons comprise concatemers. In some embodiments, a plurality of RNAs or cDNA molecules in a cellular sample can be amplified by circularizing padlock probes and performing a rolling circle amplification workflow. In some embodiments, the method includes contacting a plurality of RNAs or cDNA molecules in the cellular sample with a plurality of padlock probes, including a first plurality of target-specific padlock probes that hybridize to a first target RNA or cDNA molecule and a second plurality of target-specific padlock probes that hybridize to a second target RNA or cDNA molecule.
[0291] In some embodiments, the padlock probe comprises a single-stranded oligonucleotide. In some embodiments, the padlock probe comprises DNA, RNA, or DNA and RNA. In some embodiments, each padlock probe comprises an internal region between a first terminal region and a second terminal region, the internal region comprising at least one universal adapter sequence comprising a sample barcode sequence, an amplification primer binding site, a sequencing primer binding site, a compaction oligonucleotide binding site, and / or a surface capture primer binding site (Figure 6). In some embodiments, the padlock probe comprises at least one target barcode sequence corresponding to a given target RNA or target cDNA to which the padlock probe binds. In some embodiments, the padlock probe comprises at least one unique identification sequence (e.g., a unique molecular index (UMI)). In some embodiments, the padlock probe comprises at least one restriction enzyme recognition sequence.
[0292] In some embodiments, each padlock probe comprises first and second terminal regions (e.g., first and second binding arms) that hybridize to portions of a target RNA or target cDNA molecule to form multiple RNA-padlock probe complexes or multiple cDNA-padlock probe complexes, each complex having first and second terminal probe regions that hybridize to proximal regions of the RNA or cDNA molecule to form a nick or gap between the first and second terminal probe ends. In some embodiments, the first terminal region of each padlock probe has a first target-specific sequence that selectively hybridizes to a first region of the target RNA or cDNA molecule, and the second terminal region of each padlock probe has a second target-specific sequence that selectively hybridizes to a second region of the same target RNA or cDNA molecule, forming a nick or gap between the hybridized first and second terminal regions, thereby circularizing the padlock probe (e.g., FIG. 7).
[0293] In some embodiments, padlock probes comprise canonical nucleotides and / or nucleotide analogs. In some embodiments, padlock probes are modified to confer resistance to nuclease degradation (e.g., ribonuclease degradation). For example, padlock probes comprise at least one phosphorothioate diester bond at their 5' end, which can render the padlock probe resistant to nuclease degradation. In some embodiments, padlock probes comprise two to five or more consecutive phosphorothioate diester bonds at their 5' end. In some embodiments, padlock probes comprise at least one ribonucleotide and / or at least one 2'-O-methyl, 2'-O-methoxyethyl (MOE), 2'-fluoro-based nucleotide. In some embodiments, padlock probes comprise a phosphorylated 3' end. In some embodiments, padlock probes comprise at least one locked nucleic acid (LNA) base. In some embodiments, padlock probes comprise a phosphorylated 5' end (e.g., using polynucleotide kinase).
[0294] In some embodiments, each padlock probe within a set of padlock probes (e.g., a plurality of padlock probes) comprises first and second terminal regions that hybridize to the same target region of a target RNA or cDNA molecule to form multiple RNA-padlock probe complexes or multiple cDNA-padlock probe complexes with the same RNA or cDNA sequence.
[0295] In some embodiments, a set of padlock probes (e.g., a plurality of padlock probes) comprises at least two subsets of padlock probes. In some embodiments, individual padlock probes in a first subset of padlock probes comprise first and second terminal regions that hybridize to the same target region (e.g., a first target region) of a target RNA or cDNA molecule to form a first plurality of RNA-padlock probe complexes or a first plurality of cDNA-padlock probe complexes having the same RNA or cDNA sequence. In some embodiments, individual padlock probes in a second subset of padlock probes comprise first and second terminal regions that hybridize to the same target region (e.g., a second target region) of a target RNA or cDNA molecule to form a second plurality of RNA-padlock probe complexes or a second plurality of cDNA-padlock probe complexes having the same cDNA sequence. In some embodiments, the first and second subsets of padlock probes hybridize to different target regions of the same target RNA or cDNA molecule. In some embodiments, the first and second subsets of padlock probes hybridize to different target regions of different target RNA or cDNA molecules. In some embodiments, the set of padlock probes comprises 2 to 10 subsets of padlock probes, or 10 to 25 subsets of padlock probes, or 25 to 50 subsets of padlock probes, or up to 100 subsets of padlock probes. In some embodiments, the set of padlock probes comprises at least 100 subsets of padlock probes, at least 500 subsets of padlock probes, at least 1000 subsets of padlock probes, at least 10,000 subsets of padlock probes, or more subsets of padlock probes.
[0296] In some embodiments, the nicks can be enzymatically ligated to generate a covalently closed circular padlock probe. In some embodiments, a ligase enzyme can distinguish between matched and mismatched hybridized ends to ensure target-specific hybridization. In some embodiments, the ligation reaction involves the use of a ligase enzyme, including T3, T4, T7, or Taq DNA ligase enzyme.
[0297] In some embodiments, the size of the gap between the hybridized first and second terminal regions is 1 to 25 bases. The 3' OH end of the hybridized padlock probe can serve as an initiation site for a polymerase-catalyzed filling reaction (e.g., a gap-filling reaction) using the target cDNA molecule (or target RNA molecule) as a template. After the filling reaction, the remaining nick can be enzymatically ligated to generate a covalently closed circular padlock probe.
[0298] In some embodiments, the gap-filling reaction involves contacting the circularized padlock probe with a DNA polymerase and a plurality of nucleotides. In some embodiments, the DNA polymerase comprises E. coli DNA polymerase I, the Klenow fragment of E. coli DNA polymerase I, T7 DNA polymerase, or T4 DNA polymerase. In some embodiments, the ligase enzyme can distinguish between matched and mismatched hybridized ends to ensure target-specific hybridization. In some embodiments, the ligation reaction involves the use of a ligase enzyme, including T3, T4, T7, or Taq DNA ligase enzyme.
[0299] In any of the methods described herein, a plurality of covalently closed circular padlock probes can be subjected to a rolling circle amplification reaction to generate a plurality of concatemeric molecules each having two or more tandem copies of a unit comprising a target sequence corresponding to a target RNA molecule and any additional sequence(s) carried by the padlock probe, including universal adapter sequence(s), unique molecular index sequence(s), and / or restriction enzyme recognition sequence(s).
[0300] In some embodiments, a rolling circle amplification reaction comprises contacting a covalently closed circularization padlock probe with an amplification primer (e.g., a universal rolling circle amplification primer), a strand-displacing DNA polymerase, and a plurality of nucleotides under conditions suitable for hybridization of each amplification primer to the covalently closed padlock probe and for primer extension using the covalently closed padlock probe as a template molecule to generate nucleic acid concatemers. In some embodiments, the plurality of nucleotides in the rolling circle amplification reaction comprises a mixture of any two or more of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, the rolling circle amplification reactions described herein can be performed in the presence or absence of a plurality of compaction oligonucleotides.
[0301] In some embodiments, when the rolling circle amplification reaction includes multiple nucleotides, including dUTP, the resulting concatemers can be crosslinked to crosslinking reactive groups by treating the cell sample with succinimide ester (NHS), maleimide (Sulfo-SMCC), imidoester (DMP), carbodiimide (DCC, EDC), or phenylazide. In some embodiments, polymerization of the crosslinking reactive groups can be initiated with light or UV light. In some embodiments, the resulting concatemers can be crosslinked to a matrix by treating the cell sample with crosslinked agarose, crosslinked dextran, or crosslinked polyethylene glycol (PEG), polyacrylamide, alginate cellulose, or polyamide. In some embodiments, the PEG contains a sulfo-NHS ester moiety at one or both ends, e.g., PEGylated bis(sulfosuccinimidyl)suberate (e.g., BS(PEG)9 from Thermo Fisher Scientific, catalog number 21582).
[0302] In some embodiments, the rolling circle amplification reaction can be carried out at a constant temperature (e.g., isothermal), where the constant temperature is room temperature to about 30°C, or about 30 to 40°C, or about 40 to 50°C, or about 50 to 65°C.
[0303] In some embodiments, the DNA polymerase with strand displacement activity can be selected from the group consisting of phi29 DNA polymerase, the large fragment of Bst DNA polymerase, the large fragment of Bsu DNA polymerase, and Bca (exo-) DNA polymerase, the Klenow fragment of E. coli DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV viral reverse transcriptase, or Deep Vent DNA polymerase. In some embodiments, the phi29 DNA polymerase can be a wild-type phi29 DNA polymerase (e.g., MagniPhi from Expedeon), or a variant EquiPhi29 DNA polymerase (e.g., from Thermo Fisher Scientific), or a chimeric QualiPhi DNA polymerase (e.g., from 4basebio).
[0304] In some embodiments, rolling circle amplification primers can be modified to increase resistance to nuclease degradation. In some embodiments, rolling circle amplification primers contain at least one phosphorothioate diester bond at their 5' ends, which can make the amplification primers resistant to exonuclease degradation. In some embodiments, rolling circle amplification primers contain two to five or more consecutive phosphorothioate diester bonds at their 5' ends. In some embodiments, rolling circle amplification primers contain at least one ribonucleotide and / or at least one 2'-O-methyl or 2'-O-methoxyethyl (MOE) nucleotide.
[0305] In some embodiments, the rolling circle amplification reaction can be performed in the presence of a plurality of compaction oligonucleotides that, when hybridized to the concatemer molecules, compact the size and / or shape of the concatemers to form compact nanoballs. In some embodiments, the compaction oligonucleotides comprise single-stranded oligonucleotides having a first region at one end that hybridizes to a portion of the concatemer molecule and a second region at the other end that hybridizes to another portion of the same concatemer molecule, such that hybridization of the compaction oligonucleotide to a given concatemer compacts the size and / or shape of the concatemer.
[0306] The compaction oligonucleotide comprises a 5' region, an optional internal region (intervening region), and a 3' region. The 5' and 3' regions of the compaction oligonucleotide can hybridize to any portion of the concatemer. The 5' and 3' regions of the compaction oligonucleotide can hybridize to different portions of the concatemer, bringing the distal portions of the concatemer together and causing compaction of the concatemer to form DNA nanoballs. For example, the 5' region of the compaction oligonucleotide is designed to hybridize to a first portion of the concatemer molecule (e.g., a universal compaction oligonucleotide binding site), and the 3' region of the compaction oligonucleotide is designed to hybridize to a second portion of the concatemer molecule (e.g., a universal compaction oligonucleotide binding site). The inclusion of a compaction oligonucleotide in RCA can promote the formation of DNA nanoballs with a more compact size and shape compared to concatemers generated in the absence of the compaction oligonucleotide. The compact and stable characteristics of DNA nanoballs improve in situ sequencing accuracy by increasing signal intensity, and the nanoballs retain their shape and size during multiple sequencing cycles.
[0307] In some embodiments, the compaction oligonucleotide comprises a single-stranded oligonucleotide comprising DNA, RNA, or a combination of DNA and RNA. The compaction oligonucleotide can be any length, including 20 to 150 nucleotides, 30 to 100 nucleotides, or 40 to 80 nucleotides in length.
[0308] In some embodiments, the compaction oligonucleotide comprises a 5' region and a 3' region, and optionally an intermediate region between the 5' and 3' regions. The intervening region can be any length, for example, 2 to 20 nucleotides. The intervening region comprises a homopolymer having consecutive identical bases (e.g., AAA, GGG, CCC, TTT, or UUU). The intervening region comprises a non-homopolymer sequence.
[0309] The 5' region of the compaction oligonucleotide may be fully or partially complementary along its length to the first portion of the concatemer molecule. The 3' region of the compaction oligonucleotide may be fully or partially complementary along its length to the second portion of the concatemer molecule. The 5' region of the compaction oligonucleotide may hybridize to the first universal sequence portion of the concatemer molecule. The 3' region of the compaction oligonucleotide may hybridize to the second universal sequence portion of the concatemer molecule.
[0310] In some embodiments, the 5' region of the compaction oligonucleotide can have the same sequence as the 3' region. The 5' region of the compaction oligonucleotide can have a sequence that is different from the 3' region. In some embodiments, the 3' region of the compaction oligonucleotide can have a sequence that is the reverse of the 5' region. In some embodiments, the 5' region of the compaction oligonucleotide can have a sequence that is the reverse of the 3' region.
[0311] In some embodiments, the 3' region of either of the compaction oligonucleotides can include an additional three bases at the terminal 3' end that include 2'-O-methyl RNA bases (e.g., designated mUmUmU), or the terminal 3' end can lack the additional 2'-O-methyl RNA bases.
[0312] In some embodiments, compaction oligonucleotides contain one or more modified bases or linkages at their 5' or 3' ends to confer certain functionality. In some embodiments, compaction oligonucleotides contain at least one phosphorothioate linkage at their 5' and / or 3' ends to confer exonuclease resistance. In some embodiments, at least one nucleotide at or near the 3' end contains a 2' fluoro base, which confers exonuclease resistance. In some embodiments, the 3' end of the compaction oligonucleotide contains at least one 2'-O-methyl RNA base that blocks polymerase-catalyzed extension. For example, the 3' end of the compaction oligonucleotide contains three bases containing a 2'-O-methyl RNA base (e.g., designated mUmUmU). In some embodiments, the compaction oligonucleotide contains a 3' inverted dT at its 3' end to block polymerase-catalyzed extension. In some embodiments, the compaction oligonucleotide contains a 3' phosphorylation that blocks polymerase-catalyzed extension. In some embodiments, the internal region of the compaction oligonucleotide comprises at least one locked nucleic acid (LNA), which increases the thermal stability of the duplex formed by hybridizing the compaction oligonucleotide to the concatemer molecule. In some embodiments, the compaction oligonucleotide comprises a phosphorylated 5' end (e.g., using polynucleotide kinase).
[0313] In some embodiments, the compaction oligonucleotide comprises: 5'-CATGTAATGCACGTACTTTCAGGGTAAACATGTAATGCACGTACTTTCAGGGT-3' (SEQ ID NO: 1). In some embodiments, the compaction oligonucleotide comprises an additional three bases at the terminal 3' end that include 2'-O-methyl RNA bases (e.g., designated mUmUmU), or the terminal 3' end lacks the additional 2'-O-methyl RNA bases.
[0314] In some embodiments, the compaction oligonucleotide can include at least one region having consecutive guanines. For example, the compaction oligonucleotide can include at least one region having 2, 3, 4, 5, or more consecutive guanines. In some embodiments, the compaction oligonucleotide includes four consecutive guanines that can form a G-quadruplex structure (see Figure 25). The G-quadruplex structure can be stabilized via Hoogsteen hydrogen bonding. The G-quadruplex structure can be stabilized by a central cation, including potassium, sodium, lithium, rubidium, or cesium.
[0315] At least one compaction oligonucleotide can form a G-quadruplex (Figure 25) and hybridize to the universal binding sequence within the concatemer, allowing the concatemer to fold and form an intramolecular G-quadruplex structure (Figure 26). The concatemer can self-collapse to form a compact nanoball. The formation of the G-quadruplex and G-quadruplex in the nanoball can increase the stability of the nanoball, allowing it to retain a compact size and shape that can withstand changes in pH, temperature, and / or repeated flow of reagents during sequencing within a cell sample.
[0316] In some embodiments, the multiple compaction oligonucleotides in the rolling circle amplification reaction have the same sequence. Alternatively, the multiple compaction oligonucleotides in the rolling circle amplification reaction comprise a mixture of two or more different populations of compaction oligonucleotides having different sequences.
[0317] In some embodiments, the immobilized concatemeric template molecules can self-collapse into compact nucleic acid nanoballs, which can be imaged and FWHM measurements obtained to obtain the shape / size of the nanoballs.
[0318] In some embodiments, the inclusion of compaction oligonucleotides in rolling circle amplification reactions can facilitate the collapse of concatemers into DNA nanoballs. Performing RCA using compaction oligonucleotides helps preserve the compact size and shape of DNA nanoballs during multiple sequencing cycles, which can improve the FWHM (full width at half maximum) of spot images of DNA nanoballs within a cell sample. In some embodiments, DNA nanoballs do not resolve during multiple sequencing cycles. In some embodiments, DNA nanoball spot images do not expand during multiple sequencing cycles. In some embodiments, DNA nanoball spot images remain discrete spots during multiple sequencing cycles. Spot images can be represented as Gaussian spots, and size can be measured as FWHM. Smaller spot sizes, indicated by smaller FWHMs, typically correlate with improved spot images. In some embodiments, the FWHM of nanoball spots can be about 10 μm or less.
[0319] The single-stranded concatemers collapse into compact DNA nanoballs, each carrying multiple tandem copies of a polynucleotide unit along its length, where the polynucleotide unit contains a sequence of interest (e.g., corresponding to a target RNA or target cDNA) and at least a universal sequencing primer binding site. Each polynucleotide unit can bind to a sequencing primer, a sequencing polymerase, and a detectably labeled nucleotide reagent (e.g., a detectably labeled multivalent molecule) to form a detectable sequencing complex (e.g., a detectable ternary complex). Each nanoball carries multiple detectable sequencing complexes. The compact nature of the nanoballs thus increases the local concentration of the detectably labeled nucleotide reagent used during the sequencing workflow, increasing the signal intensity emitted from the nanoballs and providing distinct detectable signals that can be imaged as fluorescent spots within a cell sample. Each spot corresponds to a concatemer, and each concatemer corresponds to a target RNA molecule in the cell sample. Multiple spots can be simultaneously detected and imaged within the cell sample. DNA nanoballs with compact shape and size that generate increased signal intensity and color differentiation during sequencing.
[0320] In any of the methods described herein, the cell sample comprises a whole cell, a plurality of whole cells, an intact tissue, or an intact tumor. In some embodiments, the cell sample comprises a fresh cell sample, a fresh frozen cell sample, a sectioned cell sample, or an FFPE cell sample. In some embodiments, the cell sample comprises one or more viable or non-viable cells.
[0321] In some embodiments, the cell sample can be obtained from a virus, a fungus, a prokaryote, or a eukaryote. In some embodiments, the cell sample can be obtained from an animal, an insect, or a plant. In some embodiments, the cell sample comprises one or more virus-infected cells.
[0322] In some embodiments, the cell sample can be obtained from any organism, including a human, monkey, ape, dog, cat, cow, horse, mouse, pig, goat, wolf, frog, fish, plant, insect, or bacterium.
[0323] In some embodiments, the cell sample can be obtained from any organ, including the head, neck, brain, breast, ovary, cervix, colon, rectum, endometrium, gallbladder, intestine, bladder, prostate, testicle, liver, lung, kidney, esophagus, pancreas, thyroid, pituitary, thymus, skin, heart, larynx, or other organ.
[0324] In any of the methods described herein, the term "simple cell medium" or related terms typically refers to a cell medium lacking components that support the growth and / or proliferation of cells in culture. Simple cell medium can be used, for example, to wash, suspend, or dilute a cell sample. Simple cell medium can be mixed with certain components to prepare a cell medium capable of supporting the growth and / or proliferation of cells in culture. Simple cell medium includes any one or any combination of two or more of a buffer, a phosphate compound, a sodium compound, a potassium compound, a calcium compound, a magnesium compound, and / or glucose. In some embodiments, the simple cell medium includes PBS (phosphate-buffered saline), DPBS (Dulbecco's phosphate-buffered saline), HBSS (Hank's balanced salt solution), DMEM (Dulbecco's modified Eagle's medium), EMEM (Eagle's minimal essential medium), and / or EBSS. In some embodiments, a cell sample can be placed in simple cell medium before or during the step of performing any of the nucleic acid methods described herein.
[0325] In any of the methods described herein, the term "complex cell medium" or related terms refers to a cell medium that can be used to support the growth and / or proliferation of cells in culture without supplements or additives. Complex cell media can include any combination of two or more of the following: a buffer system (e.g., HEPES), inorganic salt(s), amino acid(s), protein(s), polypeptide(s), carbohydrate(s), fatty acid(s), lipid(s), purine(s) and their derivatives (e.g., hypoxanthine), pyrimidine(s) and their derivatives, and / or trace element(s). Complex cell media include fluids obtained from fluids or tissue extracts. Complex cell media include artificial cell media. In some embodiments, complex cell media can be serum-containing media, for example, complex cell media include fluids such as fetal bovine serum, plasma, serum, lymph, human placental umbilical cord serum, and amniotic fluid. In some embodiments, the complex cell medium can be a serum-free medium, which is typically (but not necessarily) a defined cell medium. In some embodiments, the complex cell medium can be a chemically defined medium, which typically (but not necessarily) includes recombinant polypeptides and ultra-pure inorganic and / or organic compounds. In some embodiments, the complex cell medium is a protein-free medium including, for example, MEM (Minimum Essential Medium) and RPMI-1640 (Roswell Park Memorial Institute). In some embodiments, the complex cell medium includes IMDM (Iscove's Modified Dulbecco's Medium). In some embodiments, the complex cell medium includes DMEM (Dulbecco's Modified Eagle's Medium). In some embodiments, the cell sample can be placed in the complex cell medium before or during the step of performing any of the nucleic acid methods described herein.
[0326] In any of the methods described herein, the cell sample comprises a fixed cell sample. In some embodiments, the cell sample can be treated with a fixation reagent (e.g., a fixation reagent) that can preserve cells and their contents, inhibiting degradation and inhibiting cell lysis. For example, the fixation reagent can preserve RNA carried by the cell sample. In some embodiments, the fixation reagent inhibits loss of nucleic acids from the cell sample.
[0327] In some embodiments, the fixation reagent can cross-link the RNA to prevent it from escaping from the cell sample. In some embodiments, the cross-linking fixation reagent comprises any combination of aldehydes, formaldehyde, paraformaldehyde, formalin, glutaraldehyde, imidoesters, N-hydroxysuccinimide esters (NHS), and / or glyoxal (a bifunctional aldehyde).
[0328] In some embodiments, the fixation reagent comprises at least one alcohol, wherein the at least one alcohol comprises methanol or ethanol. In some embodiments, the fixation reagent comprises at least one ketone, including acetone. In some embodiments, the fixation reagent comprises acetic acid, glacial acetic acid, and / or picric acid. In some embodiments, the fixation reagent comprises mercury chloride. In some embodiments, the fixation reagent comprises a zinc salt, including zinc sulfate or zinc chloride. In some embodiments, the fixation reagent is capable of denaturing the polypeptide.
[0329] In some embodiments, the fixation reagent comprises 4% w / v paraformaldehyde in water / PBS. In some embodiments, the fixation reagent comprises 10% of 35% formaldehyde at neutral pH. In some embodiments, the fixation reagent comprises 2% v / v glutaraldehyde in water / PBS. In some embodiments, the fixation reagent comprises 25% of a 37% formaldehyde solution, 70% picric acid, and 5% acetic acid.
[0330] In some embodiments, cell samples can be fixed on a support with 4% paraformaldehyde for about 30-60 minutes and washed with PBS.
[0331] In some embodiments, the cell sample may be stained, destained, or unstained.
[0332] In any of the methods described herein, the cell sample comprises a permeabilized cell sample. In some embodiments, the method comprises treating the cell sample with a permeabilization reagent that alters the cell membrane to allow experimental reagents to penetrate the cells. For example, the permeabilization reagent removes membrane lipids from the cell membrane. In some embodiments, the cell sample can be treated with a permeabilization reagent comprising any combination of an organic solvent, a detergent, a compound, a crosslinking agent, and / or an enzyme. In some embodiments, the organic solvent comprises acetone, ethanol, and methanol. In some embodiments, the detergent comprises saponin, Triton X-100, Tween-20, sodium dodecyl sulfate (SDS), N-lauroyl sarcosine sodium salt solution, or a non-ionic polyoxyethylene detergent (e.g., NP40). In some embodiments, the crosslinking agent comprises paraformaldehyde. In some embodiments, the enzyme comprises trypsin, pepsin, or a protease (e.g., proteinase K). In some embodiments, the cells can be permeabilized using alkaline conditions or acidic conditions with a protease enzyme. In some embodiments, the permeabilization reagent comprises water and / or PBS.
[0333] For example, fixed cells can be permeabilized with 70% ethanol for about 30-60 minutes, and the permeabilization reagent can be exchanged for PBS-T (e.g., PBS with 0.05% Tween-20). In some embodiments, cells can be post-fixed with 3% paraformaldehyde and 0.1% glutaraldehyde for about 30-60 minutes and washed multiple times with PBS-T.
[0334] In any of the methods described herein, a cell sample is injected with a swellable polyelectrolyte hydrogel (U.S. Pat. No. 10,309,879 and Chen 2015 Science 347:543, the contents of which are incorporated by reference in their entireties). In some embodiments, fixed and permeabilized cell samples can be injected with sodium acrylate, acrylamide, and the crosslinker N-N'-methylenebisacrylamide. In some embodiments, polymerization is achieved by injection with ammonium persulfate (APS) initiator and tetramethylethylenediamine (TEMED) accelerator. In some embodiments, cell samples can be injected with proteinase K for proteolysis and incubated in a digestion buffer. In some embodiments, the gel within the cell sample can be swelled by the addition of water.
[0335] In any of the methods described herein, a plurality of RNAs in a cell sample can be converted to cDNA. In some embodiments, the method includes contacting a plurality of RNAs in a fixed cell sample and a permeabilized cell sample with (i) a plurality of reverse transcription primers, (ii) a plurality of reverse transcriptases, and (iii) a plurality of nucleotides under conditions suitable for performing a reverse transcription reaction to generate a plurality of cDNA molecules (e.g., a plurality of first-strand cDNA molecules) in the cell sample. In some embodiments, synthesis of second-strand cDNA molecules is omitted. In some embodiments, the RNA in the cell sample is not converted to cDNA, and the RNA is hybridized to a target-specific padlock probe.
[0336] In some embodiments, the reverse transcriptase exhibits RNA-dependent DNA polymerase activity. In some embodiments, the reverse transcriptase comprises a reverse transcriptase from AMV (avian myeloblastosis virus), M-MuLV (Moloney murine leukemia virus), or HIV (human immunodeficiency virus). In some embodiments, the reverse transcriptase comprises a recombinant enzyme exhibiting reduced RNase H activity, such as REVERTAID (e.g., from Thermo Fisher Scientific, catalog number EP0441). In some embodiments, the reverse transcriptase can be a commercially available enzyme, including MULTISCRIBE (e.g., from Thermo Fisher Scientific, catalog number 4311235), THERMOSCRIPT (e.g., from Thermo Fisher Scientific, catalog number 12236-014), or ARRAYSCRIPT (e.g., from Ambion, AM2048). In some embodiments, the reverse transcriptase comprises a superscript II (e.g., catalog number 18064014), a superscript III (e.g., catalog number 18080044), or a superscript IV enzyme (e.g., catalog number 18090010) (all superscript enzymes from Invitrogen). In some embodiments, the reverse transcription reaction can include an RNase inhibitor.
[0337] In some embodiments, the reverse transcription primer comprises a single-stranded oligonucleotide comprising DNA, RNA, or chimeric DNA / RNA. In some embodiments, the reverse transcription primer is any combination of adenine (A), thymine (T), guanine (G), cytosine (C), uracil (U), and / or inosine (I). In some embodiments, the reverse transcription primer can be any length, for example, 5 to 25 bases, 25 to 50 bases, 50 to 75 bases, 75 to 100 bases, or longer. Each reverse transcription primer comprises a 5' end and a 3' end. In some embodiments, the 3' end of the reverse transcription primer comprises a 3' OH moiety that functions as a nucleotide polymerization initiation site in a polymerase-catalyzed primer extension reaction. In some embodiments, the 3' end of the reverse transcription primer has a chain-terminating moiety that blocks the polymerase-catalyzed primer extension reaction. The chain-terminating moiety can be removed to convert the 3' sugar position to an extendible 3' OH.
[0338] In some embodiments, the reverse transcription primers are modified to confer resistance to nuclease degradation (e.g., ribonuclease degradation). For example, the reverse transcription primers contain at least one phosphorothioate diester bond at their 5' ends, which can make the reverse transcription primers resistant to nuclease degradation. In some embodiments, the reverse transcription primers contain two to five or more consecutive phosphorothioate diester bonds at their 5' ends. In some embodiments, the multiple reverse transcription primers contain at least one ribonucleotide and / or at least one 2'-O-methyl, 2'-O-methoxyethyl (MOE), 2'-fluoro-based nucleotide. In some embodiments, the reverse transcription primers contain a phosphorylated 3' end. In some embodiments, the reverse transcription primers contain locked nucleic acid (LNA) bases. In some embodiments, the reverse transcription primers contain a phosphorylated 5' end (e.g., using polynucleotide kinase).
[0339] In some embodiments, the entire length of the reverse transcription primer can hybridize to a portion of the RNA molecule. In some embodiments, each reverse transcription primer comprises a 3' region having a sequence that hybridizes to a portion of the RNA molecule and a 5' region carrying a tail that does not hybridize to the RNA molecule. In some embodiments, the 5' tail comprises a universal adapter sequence comprising any combination of one or more of a sample barcode sequence, an amplification primer binding site, a sequencing primer binding site, a compaction oligonucleotide binding site, and / or a surface capture primer binding site. In some embodiments, the 5' tail comprises a unique identification sequence (e.g., a unique molecular index (UMI)). In some embodiments, the 5' tail comprises a restriction enzyme recognition sequence. In some embodiments, each reverse transcription primer comprises at least a portion of the 3' region having a homopolymer sequence, e.g., polyA, polyT, polyC, polyG, or polyU. In some embodiments, the reverse transcription primer can hybridize to any portion of the RNA molecule, including the 5' or 3' end of the RNA molecule or an internal portion of the RNA molecule.
[0340] In some embodiments, the plurality of reverse transcription primers comprises a first subpopulation of target-specific reverse transcription primers that selectively hybridize to a first target RNA (e.g., targeted transcriptomics). In some embodiments, the plurality of reverse transcription primers further comprises a second subpopulation of target-specific reverse transcription primers that selectively hybridize to a second target RNA. In some embodiments, the target-specific reverse transcription primer comprises a predetermined sequence in the 3' region that hybridizes to the target RNA molecule. In some embodiments, the predetermined sequence portion of the reverse transcription primer can be 4 to 20 bases, or 20 to 40 bases, or 40 to 50 bases in length.
[0341] In some embodiments, the first subpopulation of target-specific reverse transcription primers can selectively hybridize to RNA transcribed in a cell sample by a housekeeping gene. In some embodiments, the selection of the housekeeping gene can depend on the type of cell sample used in the in situ method described herein. Exemplary housekeeping genes include glyceraldehyde-3-phosphate dehydrogenase (GAPDH), beta-actin (ACTB), tubulin, PPIA (peptidyl-prolyl cis-trans isomerase), NME4 (NME / NM23 nucleoside diphosphate kinase 4), SMARCAL1 (SWI / SNF-related matrix actin-dependent chromatin regulator, subfamily A, e.g., 1), and POMK (protein-O-mannose kinase). Those skilled in the art can design a first subpopulation of target-specific reverse transcription primers that hybridize to RNA transcripts from any of a number of housekeeping genes.
[0342] In some embodiments, the second subpopulation of target-specific reverse transcription primers is capable of selectively hybridizing to RNA transcribed from genes expressed in the cell sample being tested (e.g., cell-specific or tissue-specific RNA).
[0343] In some embodiments, the plurality of reverse transcription primers comprises a first subpopulation of random-sequence reverse transcription primers that hybridize to a first target RNA (e.g., total transcriptomics). In some embodiments, the plurality of reverse transcription primers further comprises a second subpopulation of random-sequence reverse transcription primers that hybridize to a second target RNA. In some embodiments, the reverse transcription primers comprise random and / or degenerate sequences in the 3' region that hybridizes to the RNA molecule. In some embodiments, the random or degenerate sequence portion of the reverse transcription primer can be 4 to 20 bases, or 20 to 40 bases, or 40 to 50 bases in length.
[0344] Sequencing Polymerase In any of the methods described herein, sequencing polymerase can be used to carry out sequencing reaction.In some embodiments, sequencing polymerase(s) can bind and incorporate complementary nucleotides opposite the nucleotide in concatemeric template molecule.In some embodiments, sequencing polymerase(s) can bind complementary nucleotide units of multivalent molecules opposite the nucleotide in concatemeric template molecule.In some embodiments, the plurality of sequencing polymerases comprises recombinant mutant polymerases.
[0345] Examples of polymerases suitable for use in sequencing with nucleotides and / or polyvalent molecules include Klenow DNA polymerase; Thermus aquaticus DNA polymerase I (Taq polymerase); KlenTaq polymerase; Candidatus altiarchaeales archaea; Candidatus Hadarchaeum Yellowstonense; Hadesarchaea archaea; Euryarchaeota archaea; Thermoplasmata archaea; Thermococcus polymerases, e.g., Thermococcus litoralis, bacteriophage T7 DNA polymerase; human alpha, delta, and epsilon DNA polymerases; bacteriophage polymerases, e.g., T4, RB69, and phi29 bacteriophage DNA polymerases; Pyrococcus furiosus DNA polymerase (Pfu polymerase); Bacillus subtilis DNA polymerase III; E. coli DNA polymerase III alpha and epsilon; 9 degree Examples of DNA polymerases include, but are not limited to, N polymerase; reverse transcriptases such as HIV type M or O reverse transcriptase; avian myeloblastosis virus reverse transcriptase; Moloney murine leukemia virus (MMLV) reverse transcriptase; or telomerase. Further non-limiting examples of DNA polymerases include those from various archaeal genera, such as Aeropyrum, Archaeglobus, Desulfurococcus, Pyrobaculum, Pyrococcus, Pyrolobus, Pyrodictium, Staphylothermus, Stetteria, Sulfolobus, Thermococcus, and Vulcanisaeta, or variants thereof, including such polymerases known in the art, such as 9 degrees N, VENT, DEEP VENT, THERMINATOR, Pfu, KOD, Pfx, Tgo, and RB69 polymerases.
[0346] Sequencing by ligation In any of the methods described herein, sequencing comprises performing a sequencing-by-binding (SBB) reaction in a cell sample, and the cDNA amplicons are concatemer molecules. In some embodiments, the sequencing-by-binding (SBB) procedure uses unlabeled chain-terminating nucleotides. In some embodiments, sequencing-by-binding (SBB) comprises: (a) sequentially contacting primed concatemers (e.g., concatemers annealed to multiple sequencing primers) with at least two separate mixtures under ternary complex-stabilizing conditions, each of the at least two separate mixtures comprising a polymerase and a nucleotide, thereby resulting in primed template nucleic acids in which the primed concatemers are contacted with nucleotide analogs of the first, second, and third base types in the template under ternary complex-stabilizing conditions; and (b) examining the at least two separate mixtures to determine the ternary complex. (c) identifying the next correct nucleotide of the primed concatemer, where the next correct nucleotide is identified as a congener of the first, second, or third base type if a ternary complex is detected in step (b), and the next correct nucleotide is presumed to be a nucleotide congener of the fourth base type based on the absence of a ternary complex in step (b), (d) adding the next correct nucleotide to the primed concatemer after step (b), thereby producing an extension primer, and (e) repeating steps (a)-(d) on the primed concatemer containing the extension primer. Exemplary sequencing-by-ligation methods are described in U.S. Patent Nos. 10,246,744 and 10,731,141 (the entire contents of both patents are incorporated herein by reference).
[0347] Nucleotides and Chain-Terminating Nucleotides In any of the methods described herein, any of the sequencing methods described herein can use at least one nucleotide. A nucleotide comprises a base, a sugar, and at least one phosphate group. In some embodiments, at least one nucleotide in the plurality of nucleotides comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1 to 10 phosphate groups). The plurality of nucleotides can comprise at least one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. The plurality of nucleotides can comprise a mixture of any combination of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, at least one nucleotide in the plurality of nucleotides is not a nucleotide analog. In some embodiments, at least one nucleotide in the plurality of nucleotides comprises a nucleotide analog.
[0348] In some embodiments, in any of the sequencing methods described herein, at least one nucleotide of the plurality of nucleotides comprises a chain of one, two, or three phosphorus atoms, typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramide bond. In some embodiments, at least one nucleotide in the plurality of nucleotides is an analog having a phosphorus chain, in which the phosphorus atoms are linked together with intervening O, S, NH, methylene, or ethylene. In some embodiments, the phosphorus atoms in the chain comprise a substituted side chain group, including O, S, or BH3. In some embodiments, the chain comprises a phosphate group substituted with an analog, including phosphoramidate, phosphorothioate, phosphorodithioate, and O-methylphosphoramidite groups.
[0349] In some embodiments, in any of the methods for sequencing described herein, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog, wherein the terminator nucleotide analog has a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain-terminating moiety can inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide in the nascent chain during a primer extension reaction. In some embodiments, the chain-terminating moiety is attached to the 3' sugar hydroxyl position, where the sugar comprises a ribose or deoxyribose sugar moiety. In some embodiments, the chain-terminating moiety is removable / cleavable from the 3' sugar hydroxyl position to generate a nucleotide having a 3' OH sugar group that is extendable with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the chain-terminating moiety is cleavable / removable from the nucleotide, for example, by reacting the chain-terminating moiety with a chemical agent, a pH change, light, or heat. In some embodiments, alkyl, alkenyl, alkynyl, and aryl chain-terminating moieties are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, aryl and benzyl chain-terminating moieties are cleavable with HPd / C. In some embodiments, the chain-terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphines or thiol groups, including beta-mercaptoethanol or dithiothritol (DTT).In some embodiments, carbonate chain terminating moieties can be cleaved with potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn(AcOH) in acetic acid. In some embodiments, urea and silyl chain terminating moieties can be cleaved with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0350] In some embodiments, in any of the methods for sequencing described herein, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog, wherein the terminator nucleotide analog has a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group. In some embodiments, the chain-terminating moiety comprises a 3'-O-azido or 3'-O-azidomethyl group. In some embodiments, the chain-terminating azide, azido, and azidomethyl groups are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), or bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0351] In some embodiments, in any of the methods for sequencing described herein, the nucleotide analog comprises a chain-terminating moiety selected from the group consisting of 3'-deoxynucleotides, 2',3'-dideoxynucleotides, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-sulfhydral, 3'-aminomethyl, 3'-ethyl, 3'butyl, 3'-tertbutyl, 3'-fluorenylmethyloxycarbonyl, 3'tert-butyloxycarbonyl, 3'-O-alkylhydroxylamino groups, 3'-phosphorothioates, and 3-O-benzyl, or derivatives thereof.
[0352] In some embodiments, in any of the methods for sequencing described herein, the plurality of nucleotides comprises a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base with a linker, the linker being cleavable / removable from the base. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is not labeled with a detectable reporter moiety. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the nucleotide can correspond to a nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to enable detection and identification of the nucleotide base.
[0353] In some embodiments, in any of the methods for sequencing nucleic acid molecules described herein, the cleavable linker on the nucleotide base comprises a cleavable moiety comprising an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the cleavable linker on the base is cleavable / removable from the base by reacting the cleavable moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the cleavable moieties alkyl, alkenyl, alkynyl, and aryl are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the cleavable moieties aryl and benzyl are cleavable with HPd / C. In some embodiments, the cleavable moieties amine, amide, keto, isocyanate, phosphate, thio, and disulfide are cleavable with phosphines or thiol groups, including beta-mercaptoethanol or dithiothritol (DTT). In some embodiments, the cleavable moiety carbonate is cleavable with potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn(AcOH) in acetic acid. In some embodiments, the cleavable moieties urea and silyl are cleavable with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0354] In some embodiments, in any of the methods for sequencing described herein, the cleavable linker on the nucleotide base comprises a cleavable moiety including azide, azido, and azidomethyl groups. In some embodiments, the cleavable moieties azide, azido, and azidomethyl groups are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0355] In some embodiments, in any of the methods for sequencing described herein, the chain-terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the cleavable linker on the nucleotide base have the same or different cleavable moieties. In some embodiments, the chain-terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the detectable reporter moiety attached to the base are chemically cleavable / removable with the same chemical agent. In some embodiments, the chain-terminating moiety (e.g., at the sugar 2' and / or sugar 3' positions) and the detectable reporter moiety attached to the base are chemically cleavable / removable with different chemical agents.
[0356] Multivalent molecules In any of the methods described herein, sequencing uses at least one multivalent molecule comprising a core and a plurality of nucleotide arms having any configuration, including starburst, helter-skelter, or bottlebrush configurations (e.g., Figure 16). The multivalent molecule comprises (1) a core and (2) a plurality of nucleotide arms, each comprising (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, the spacer is attached to the linker, and the linker is attached to the nucleotide unit. In some embodiments, the nucleotide unit comprises a base, a sugar, and at least one phosphate group, and the linker is attached to the nucleotide unit via the base. In some embodiments, the linker comprises an aliphatic chain or an oligoethylene glycol chain, and both linker chains have 2 to 6 subunits. In some embodiments, the linker also comprises an aromatic moiety. Exemplary nucleotide arms are shown in Figure 20. Exemplary multivalent molecules are shown in Figures 16-20. An exemplary spacer is shown in Figure 21 (top), and an exemplary linker is shown in Figure 21 (bottom) and Figure 22. Exemplary nucleotides attached to linkers are shown in Figures 23A-23D. An exemplary biotinylated nucleotide arm is shown in Figure 24.
[0357] In some embodiments, the multivalent molecule comprises a core bound to a plurality of nucleotide arms, the plurality of nucleotide arms having the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.
[0358] In some embodiments, a multivalent molecule comprises a core to which multiple nucleotide arms are attached, each arm comprising a nucleotide unit. In some embodiments, the nucleotide unit comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1 to 10 phosphate groups). The multiple multivalent molecules can comprise one type of multivalent molecule having one type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. The multiple multivalent molecules can be comprised in a mixture of any combination of two or more types of multivalent molecules, where each individual multivalent molecule in the mixture comprises a nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP.
[0359] In some embodiments, the nucleotide unit comprises a chain of one, two, or three phosphorus atoms, typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramido linkage. In some embodiments, at least one nucleotide unit is a nucleotide analog having a phosphorus chain in which the phosphorus atoms are linked together with intervening O, S, NH, methylene, or ethylene. In some embodiments, the phosphorus atoms in the chain comprise substituted side chain groups including O, S, or BH3. In some embodiments, the chain comprises phosphate groups substituted with analogs including phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoramidite groups.
[0360] In some embodiments, a multivalent molecule comprises a core attached to multiple nucleotide arms, each of which comprises a nucleotide unit that is a nucleotide analog having a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the nucleotide unit comprises a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain-terminating moiety can inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide into a nascent chain during a primer extension reaction. In some embodiments, the chain-terminating moiety is attached to the 3' sugar hydroxyl position, where the sugar comprises a ribose or deoxyribose sugar moiety. In some embodiments, the chain-terminating moiety is removable / cleavable from the 3' sugar hydroxyl position to generate a nucleotide having a 3' OH sugar group that is extendable with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the chain-terminating moiety is cleavable / removable from the nucleotide unit, for example, by reacting the chain-terminating moiety with a chemical agent, a pH change, light, or heat. In some embodiments, alkyl, alkenyl, alkynyl, and aryl chain-terminating moieties are cleavable with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, aryl and benzyl chain-terminating moieties are cleavable with HPd / C. In some embodiments, the chain-terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable with phosphines or thiol groups, including beta-mercaptoethanol or dithiothritol (DTT).In some embodiments, carbonate chain terminating moieties can be cleaved with potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn(AcOH) in acetic acid. In some embodiments, urea and silyl chain terminating moieties can be cleaved with tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0361] In some embodiments, the nucleotide units comprise a chain-terminating moiety (e.g., a blocking moiety) at the 2' sugar position, the 3' sugar position, or the 2' and 3' sugar positions. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group. In some embodiments, the chain-terminating moiety comprises a 3'-O-azido or 3'-O-azidomethyl group. In some embodiments, the chain-terminating azide, azido, and azidomethyl groups are cleavable / removable with a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bis-sulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0362] In some embodiments, the nucleotide units comprise a chain-terminating moiety selected from the group consisting of 3'-deoxynucleotides, 2',3'-dideoxynucleotides, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-sulfhydral, 3'-aminomethyl, 3'-ethyl, 3'butyl, 3'-tertbutyl, 3'-fluorenylmethyloxycarbonyl, 3'tert-butyloxycarbonyl, 3'-O-alkylhydroxylamino groups, 3'-phosphorothioates, and 3-O-benzyl, or derivatives thereof.
[0363] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, the nucleotide arms comprising spacers, linkers, and nucleotide units, and the core, linkers, and / or nucleotide units are labeled with a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the multivalent molecule can correspond to the base of the nucleotide unit (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and discrimination of the nucleotide base.
[0364] In some embodiments, at least one nucleotide arm of the multivalent molecule has a nucleotide unit attached to a detectable reporter moiety. In some embodiments, the detectable reporter moiety is attached to a nucleotide base. In some embodiments, the detectable reporter moiety comprises a fluorophore. In some embodiments, the particular detectable reporter moiety (e.g., fluorophore) attached to the multivalent molecule can correspond to the base of the nucleotide unit (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow for detection and discrimination of the nucleotide base.
[0365] In some embodiments, the core of the multivalent molecule comprises an avidin-like or streptavidin-like moiety and the core-attached moiety comprises biotin. In some embodiments, the core comprises a streptavidin-type or avidin-type moiety, including avidin protein and any derivatives, analogs, and other non-natural forms of avidin that can bind to at least one biotin moiety. Other forms of avidin moieties include natural and recombinant avidin and streptavidin, as well as derivatized molecules such as non-glycosylated avidin and truncated streptavidin. For example, avidin moieties include deglycosylated forms of avidin, bacterial streptavidin produced by Streptomyces (e.g., Streptomyces avidinii), as well as derivatized forms, such as N-acylavidins, e.g., N-acetyl, N-phthalyl, and N-succinyl avidin, and the commercially available products EXTRAVIDIN, CAPTAVIDIN, NEUTRAVIDIN, and NEUTRALITE AVIDIN.
[0366] In some embodiments, any of the methods for sequencing nucleic acid molecules described herein may include forming a binding complex, wherein the binding complex comprises (i) a polymerase, a nucleic acid concatemer molecule duplexed with a primer, and a nucleotide, or the binding complex comprises (ii) a polymerase, a nucleic acid concatemer molecule duplexed with a primer, and a nucleotide unit of a multivalent molecule. In some embodiments, the binding complex has a duration of greater than about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. The binding complex has a duration of greater than about 0.1-0.25 seconds, or about 0.25-0.5 seconds, or about 0.5-0.75 seconds, or about 0.75-1 second, or about 1-2 seconds, or about 2-3 seconds, or about 3-4 seconds, or about 4-5 seconds, and / or the method is or can be performed at 15°C or higher, 20°C or higher, 25°C or higher, 35°C or higher, 37°C or higher, 42°C or higher, 55°C or higher, 60°C or higher, 72°C or higher, or 80°C or higher, or within a range defined by any of the foregoing. The binding complex (e.g., ternary complex) remains stable until subjected to conditions that cause dissociation of interactions between the polymerase, template molecule, primer, and / or any of the nucleotide units or nucleotides. For example, dissociation conditions include contacting the binding complex with any one of detergent, EDTA, and / or water, or any combination thereof. In some embodiments, the disclosure provides such methods, wherein the binding complex is deposited on, attached to, or hybridized to a surface that, in the detecting step, exhibits a contrast to noise ratio of greater than 20. In some embodiments, the disclosure provides such methods, wherein the contacting is performed under conditions that stabilize the binding complex when the nucleotide or nucleotide unit is complementary to the next base of the template nucleic acid and destabilize the binding complex when the nucleotide or nucleotide unit is not complementary to the next base of the template nucleic acid.
[0367] In some embodiments, in any of the sequencing methods using multivalent molecules, binding a plurality of first multiplexed polymerases to a plurality of multivalent molecules forms at least one avidity complex, the method comprising the steps of: (a) binding a first nucleic acid primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemeric template molecule, thereby forming a first binding complex, wherein a first nucleotide unit of the first multivalent molecule binds to the first sequencing polymerase; and (b) binding a second nucleic acid primer, a second sequencing polymerase, and a first multivalent molecule to a second portion of the same concatemeric template molecule, thereby forming a second binding complex, wherein a second nucleotide unit of the first multivalent molecule binds to the second sequencing polymerase; and the first and second binding complexes comprising the same multivalent molecule form an avidity complex. In some embodiments, the first sequencing polymerase comprises any wild-type or mutant polymerase described herein. In some embodiments, the second sequencing polymerase comprises any wild-type or mutant polymerase described herein. The concatemeric template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal sequencing primer binding site. First and second nucleic acid primers can bind to the sequencing primer binding sites along the concatemeric template molecule. Exemplary multivalent molecules are shown in Figures 16-20.
[0368] In some embodiments, in any of the sequencing methods using multivalent molecules, the method comprises combining a plurality of first multiplexed polymerases with a plurality of multivalent molecules to form at least one avidity complex, the method comprising the steps of: (a) contacting a plurality of sequencing polymerases and a plurality of nucleic acid primers with different portions of concatemeric nucleic acid concatemeric molecules to form at least first and second multiplexed polymerases on the same concatemeric molecule; and (b) contacting the plurality of multivalent molecules with at least first and second multiplexed polymerases on the same concatemeric template molecule under conditions suitable for binding of a single multivalent molecule from the plurality of multivalent molecules to the first and second multiplexed polymerases, wherein at least a first nucleotide unit of the single multivalent molecule comprises a first primer that hybridizes to a first portion of the concatemeric template molecule, thereby forming a first binding complex (e.g., a first ternary complex). (c) contacting a first composite polymerase with a second composite polymerase comprising a second primer, wherein at least a second nucleotide unit of the single multivalent molecule hybridizes to a second portion of the concatemeric template molecule, thereby forming a second binding complex (e.g., a second ternary complex), under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second binding complexes, such that the first and second binding complexes bound to the same multivalent molecule form an avidity complex; (d) identifying the first nucleotide unit in the first binding complex, thereby determining the sequence of the first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second binding complex, thereby determining the sequence of the second portion of the concatemeric template molecule. In some embodiments, the plurality of sequencing polymerases comprises any wild-type or mutant sequencing polymerase described herein.The concatemer template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal sequencing primer binding site. Multiple nucleic acid primers can bind to the sequencing primer binding sites along the concatemer template molecule. Exemplary multivalent molecules are shown in Figures 16-20.
[0369] Figure 16 is a schematic diagram of various exemplary configurations of multivalent molecules. Left (Class I): Schematic of a multivalent molecule with a "starburst" or "helter-skelter" configuration. Middle (Class II): Schematic of a multivalent molecule with a dendrimer configuration. Right (Class III): Schematic of multiple multivalent molecules formed by reacting streptavidin with a 4-arm or 8-arm PEG-NHS bearing biotin and dNTPs. Nucleotide units are represented as "N," biotin is represented as "B," and streptavidin is represented as "SA."
[0370] FIG. 17 is a schematic diagram of an exemplary multivalent molecule comprising a generic core attached to multiple nucleotide arms.
[0371] FIG. 18 is a schematic diagram of an exemplary multivalent molecule comprising a dendrimer core attached to multiple nucleotide arms.
[0372] FIG. 19 shows a schematic diagram of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, the nucleotide arms comprising biotin, spacers, linkers, and nucleotide units.
[0373] FIG. 20 is a schematic diagram of an exemplary nucleotide arm comprising a core attachment moiety, a spacer, a linker, and a nucleotide unit.
[0374] FIG. 21 shows the chemical structures of exemplary spacers (top) and various exemplary linkers (bottom), including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker.
[0375] FIG. 22 shows the chemical structures of various exemplary linkers, including linkers 1-9.
[0376] 23A shows the chemical structures of various exemplary linkers linked / attached to nucleotide units.
[0377] FIG. 23B shows the chemical structures of various exemplary linkers linked / attached to nucleotide units.
[0378] FIG. 23C shows the chemical structures of various exemplary linkers linked / attached to nucleotide units.
[0379] FIG. 23D shows the chemical structures of various exemplary linkers linked / attached to nucleotide units.
[0380] Figure 24 shows the chemical structure of an exemplary nucleotide arm. In this example, the nucleotide unit is connected to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base.
[0381] FIG. 25 is a schematic diagram of a G-quadruplex (eg, a G-quadruplex).
[0382] FIG. 26 is a schematic diagram of an exemplary intramolecular G-quadruplex structure.
[0383] Flow cell In any of the methods described herein, the cell sample can be deposited on a solid support (e.g., a flow cell). In some embodiments, the cell sample is deposited on a flow cell having walls (e.g., a top or first wall and a bottom or second wall) and a gap therebetween, the gap can be filled with a fluid, and the flow cell is placed in a fluorescence optical imaging system. Because the cell sample has a thickness, when using a conventional imaging system, it may be necessary to separately focus the imaging system on the first and second surfaces of the flow cell. For improved imaging of concatemer sequencing reactions in the cell sample, the flow cell can be placed in a high-performance fluorescence imaging system including two or more tube lenses designed to provide optimal imaging performance on the first and second surfaces of the flow cell at two or more fluorescence wavelengths. In some embodiments, the high-performance imaging system further comprises a focusing mechanism configured to refocus the optical system between acquiring images of the first and second surfaces of the flow cell. In some embodiments, the high-performance imaging system is configured to image two or more fields of view on at least one of the first flow cell surface or the second flow cell surface.
[0384] Substrates and Coatings In any of the methods described herein, the solid support comprises a flow cell having a coating that promotes cell adhesion. In some embodiments, the flow cell comprises a support, which can be a planar or non-planar support. The support can be solid or semi-solid. In some embodiments, the support can be porous, semi-porous, or non-porous. The support can be made of any material, such as glass, plastic, or a polymeric material. In some embodiments, the surface of the support can be coated with one or more compounds to create a passivation layer on the support (Figure 15). In some embodiments, the passivation layer forms a porous or semi-porous layer. In some embodiments, the support is coated with a lysine compound, a poly-lysine compound, an arginine compound, or an amino-terminal compound. The support can be coated with an unbranched compound, a branched compound, or a mixture of unbranched and branched compounds. In some embodiments, the support is coated with a surface primer for capturing nucleic acids from a cell sample. Alternatively, the support lacks a surface primer.
[0385] 15 is a schematic diagram of an exemplary low-binding support comprising a glass substrate and alternating layers of hydrophilic coating, the alternating layers of hydrophilic coating covalently or non-covalently adhered to the glass and further comprising chemically reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides and circularization oligonucleotides). In alternative embodiments, the support can be made from any material, such as glass, plastic, or a polymeric material.
[0386] The support may include one or more substrates. The support may include a glass or plastic substrate. The support may include a transparent upper substrate closest to the objective lens of the optical system. The support may include one or more microfluidic channels and concatemer molecules, and the cell sample is immobilized on the surface of the microfluidic channel. In some embodiments, the support is included in a flow cell device.
[0387] The imager 116 of Figure 1 can include one or more optical systems. Further disclosed herein are optical system design guidelines and high-performance fluorescence imaging methods and systems that provide improved optical resolution and image quality for fluorescence imaging-based genomics applications. The disclosed optical imaging system designs provide larger fields of view, increased spatial resolution, improved modulation transfer, contrast-to-noise ratio, and image quality, higher spatial sampling frequencies, faster transitions between image captures when repositioning the sample plane to capture a series of images (e.g., of different fields of view), and improved imaging system duty cycles, thus enabling higher throughput image acquisition and analysis.
[0388] In some cases, for example, improved imaging performance for dual-sided (flow cell) imaging applications can be achieved by using an electro-optic phase plate in combination with an objective lens to compensate for optical aberrations caused by the fluid layer separating the upper (near) and lower (far) inner surfaces of the flow cell. In some cases, this design approach can also compensate for vibrations introduced by, for example, a motion-activated compensator that is moved in or out of the optical path depending on which surface of the flow cell is being imaged.
[0389] In some cases, for example, for dual-sided (flow cell) imaging applications involving the use of thick flow cell walls (e.g., wall (or coverslip) thickness greater than 700 µm) and fluidic channels (e.g., fluidic channel height or thickness between 50 and 200 µm), improved imaging performance can be achieved even when using commercially available, off-the-shelf objectives by using a tube lens design that, in combination with the objective, corrects for optical aberrations caused by the thick flow cell walls and / or intermediate fluidic layers.
[0390] In some cases, for example, improved imaging performance for multi-channel (e.g., two-color or four-color) imaging applications can be achieved by using multiple tube lenses (one for each imaging channel), with each tube lens design optimized for the particular wavelength range used in that imaging channel.
[0391] Exemplary embodiments disclosed herein may include a fluorescence imaging system comprising: a) at least one light source configured to provide excitation light within one or more specified wavelength ranges; and b) an objective configured to collect fluorescence arising from within a specified field of view of the sample surface upon exposure of the sample surface to the excitation light, wherein the numerical aperture of the objective is at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or a numerical aperture falling within a range defined by any two of the foregoing; and wherein the working distance of the objective is at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, at least and c) at least one objective lens having a working distance of at least 1000 μm, or within a range defined by any two of the foregoing, and a field of view of at least 0.1 mm, at least 0.2 mm, at least 0.5 mm, at least 0.7 mm, at least 1 mm, at least 2 mm, at least 3 mm, at least 5 mm, or at least 10 mm, or a field of view within a range defined by any two of the foregoing; and c) at least one image sensor, wherein fluorescence collected by the objective lens is imaged onto the image sensor, and the pixel dimension of the image sensor is selected such that the spatial sampling frequency of the fluorescence imaging system is at least twice the optical resolution of the fluorescence imaging system.
[0392] In some embodiments, the numerical aperture may be at least 0.75. In some embodiments, the numerical aperture is at least 1.0. In some embodiments, the working distance is at least 850 μm. In some embodiments, the working distance is at least 1,000 μm. In some embodiments, the field of view may have an area of at least 2.5 mm. In some embodiments, the field of view may have an area of at least 3 mm. In some embodiments, the spatial sampling frequency may be at least 2.5 times the optical resolution of the fluorescence imaging system. In some embodiments, the spatial sampling frequency may be at least 3 times the optical resolution of the fluorescence imaging system. In some embodiments, the system may further comprise an XYZ translation stage such that the system is configured to acquire a series of two or more fluorescence images in an automated manner, each image in the series being acquired or can be acquired for a different field of view. In some embodiments, the position of the sample plane may be simultaneously adjusted in the X, Y, and Z directions to coincide with the position of the objective focal plane during acquisition of images of different fields of view. In some embodiments, the time required for simultaneous adjustment in the X, Y, and Z directions can be less than 0.3 seconds, less than 0.4 seconds, less than 0.5 seconds, less than 0.7 seconds, or less than 1 second, or a time within a range defined by any two of the foregoing. In some embodiments, the system further comprises an autofocus mechanism configured to adjust the position of the focal plane before acquiring images of different fields of view if the error signal indicates that the difference in position of the focal plane and the sample plane in the Z direction is greater than a specified error threshold. In some embodiments, the specified error threshold is 100 nm or greater. In some embodiments, the specified error threshold is 50 nm or less. In some embodiments, the system comprises three or more image sensors, and the system is configured to image fluorescence in each of three or more wavelength ranges onto a different image sensor. In some embodiments, the difference in the position of the focal plane between each of the three or more image sensors and the sample plane is less than 100 nm. In some embodiments, the difference in the position of the focal plane between each of the three or more image sensors and the sample plane is less than 50 nm.In some embodiments, the total time required to reposition the sample plane, adjust the focus as needed, and acquire an image is less than 0.4 seconds per field of view, hi some embodiments, the total time required to reposition the sample plane, adjust the focus as needed, and acquire an image is less than 0.3 seconds per field of view.
[0393] The present disclosure also discloses a fluorescence imaging system for dual-sided imaging of a flow cell, comprising: a) an objective lens configured to collect fluorescence arising from within a specific field of view of a sample surface within the flow cell; and b) at least one tube lens disposed between the objective lens and at least one image sensor, wherein the at least one tube lens is configured to compensate for an imaging performance metric for the combination of the objective lens, the at least one tube lens, and the at least one image sensor when imaging an inner surface of the flow cell, wherein the flow cell has a wall thickness of at least 700 μm and a gap between its upper and lower inner surfaces of at least 50 μm, and wherein the imaging performance metric is substantially the same for imaging the upper or lower inner surface of the flow cell without moving an optical compensator into or out of the optical path between the flow cell and the at least one image sensor, without moving one or more optical elements of the tube lens along the optical path, and without moving one or more optical elements of the tube lens into or out of the optical path.
[0394] In some embodiments, the objective lens can be a commercially available microscope objective lens. In some embodiments, the commercially available microscope objective lens can have a numerical aperture of at least 0.3. In some embodiments, the objective lens can have a working distance of at least 700 μm. In some embodiments, the objective lens can be corrected to compensate for a coverslip thickness (or flow cell wall thickness) of 0.17 mm, or a coverslip thickness (or flow cell wall thickness) greater than or less than 0.17 mm. In some embodiments, the optical system can be corrected to correct for coverslip thickness, flow cell thickness, or the distance between the desired focal planes. In some embodiments, the correction can be made by inserting a corrective optical system, such as a lens or optical assembly, into the optical path of the optical system. In some embodiments, the correction can be made without inserting a corrective optical system, such as a lens or optical assembly, into the optical path of the optical system. In some embodiments, the fluorescence imaging system may further include an electro-optic phase plate disposed adjacent to the objective lens between the objective lens and the tube lens, where the electro-optic phase plate may provide correction for optical aberrations caused by fluid filling the gap between the upper and lower inner surfaces of the flow cell. In some embodiments, the at least one tube lens may be a compound lens including three or more optical components. In some embodiments, the at least one tube lens is a compound lens including four optical components, which may include one or more of a first asymmetric convex-convex lens, a second convex-plano lens, a third asymmetric concave-concave lens, and a fourth asymmetric convex-concave lens, which may be present in the above order or in any alternate order. In some embodiments, the at least one tube lens is configured to correct imaging performance metrics for the combination of the objective lens, the at least one tube lens, and the at least one imaging element when imaging the inner surface of a flow cell having a wall thickness of at least 1 mm.In some embodiments, the at least one tube lens is configured to correct an imaging performance metric of the combination of the objective lens, the at least one tube lens, and the at least one imaging element when imaging the inner surface of a flow cell having a gap of at least 100 μm. In some embodiments, the at least one tube lens is configured to correct an imaging performance metric of the combination of the objective lens, the at least one tube lens, and the at least one imaging element when imaging the inner surface of a flow cell having a gap of at least 200 μm. In some embodiments, the system includes a single objective lens, two tube lenses, and two image sensors, each of the two tube lenses designed to provide optimal imaging performance at a different fluorescent wavelength. In some embodiments, the system includes a single objective lens, three tube lenses, and three image sensors, each of the three tube lenses designed to provide optimal imaging performance at a different fluorescent wavelength. In some embodiments, the system includes a single objective lens, four tube lenses, and four image sensors, each of the four tube lenses designed to provide optimal imaging performance at a different fluorescent wavelength. In some embodiments, the design of the objective lens or at least one tube lens is configured to optimize the modulation transfer function in the mid- to high-spatial frequency range. In some embodiments, the imaging performance metric includes a measurement of the modulation transfer function (MTF) at one or more specified spatial frequencies, defocus, spherical aberration, chromatic aberration, coma, astigmatism, field curvature, image distortion, contrast-to-noise ratio (CNR), or any combination thereof. In some embodiments, the difference in the imaging performance metric for imaging the upper and lower inner surfaces of the flow cell is less than 10%. In some embodiments, the difference in the imaging performance metric for imaging the upper and lower inner surfaces of the flow cell is less than 5%.In some embodiments, the use of at least one tube lens improves imaging performance metrics for dual-sided imaging by at least the same or better than that for a conventional system comprising an objective lens, a motion-activated compensator, and an image sensor. In some embodiments, the use of at least one tube lens improves imaging performance metrics for dual-sided imaging by at least 10% compared to that for a conventional system comprising an objective lens, a motion-activated compensator, and an image sensor.
[0395] Disclosed herein is an illumination system for use in imaging-based solid-phase genotyping and sequencing applications, the illumination system comprising: a) a light source; and b) a liquid light guide configured to collect light emitted by the light source and deliver it to a designated illumination field on a support surface containing a tethered biological macromolecule.
[0396] In some embodiments, the illumination system further comprises a focusing lens. In some embodiments, the designated illumination field has an area of at least 2 mm. In some embodiments, the light delivered to the designated illumination field is of uniform intensity across a field of view designated for an imaging system used to acquire an image of the substrate surface. In some embodiments, the designated field of view has an area of at least 2 mm. In some embodiments, the light delivered to the designated illumination field is of uniform intensity across the designated field of view when the coefficient of variation (CV) of the light intensity is less than 10%. In some embodiments, the light delivered to the designated illumination field is of uniform intensity across the designated field of view when the coefficient of variation (CV) of the light intensity is less than 5%. In some embodiments, the light delivered to the designated illumination field has a speckle contrast value of less than 0.1. In some embodiments, the light delivered to the designated illumination field has a speckle contrast value of less than 0.05.
[0397] Imaging Modules and Systems: It will be understood by those skilled in the art that the disclosed optical systems, imaging systems, or modules may, in some cases, be standalone optical systems designed to image sample or substrate surfaces. In some cases, they may include one or more processors or computers. In some cases, they may include one or more software packages providing instrument control and / or image processing functions. In some cases, in addition to optical components such as light sources (e.g., solid-state lasers, dye lasers, diode lasers, arc lamps, tungsten-halogen lamps, etc.), lenses, prisms, mirrors, dichroic reflectors, optical filters, optical bandpass filters, apertures, and image sensors (e.g., complementary metal-oxide semiconductor (CMOS) image sensors and cameras, charge-coupled device (CCD) image sensors and cameras, etc.), they may also include mechanical and / or opto-mechanical components such as XY translation stages, XYZ translation stages, piezoelectric focusing mechanisms, etc. In some cases, they may function as modules, components, subassemblies, or subsystems of larger systems designed for genomics applications (e.g., genetic testing and / or nucleic acid sequencing applications). For example, in some cases, they may function as modules, components, subassemblies, or subsystems of a larger system that further includes a light-tight and / or other environmental control housing, a temperature control module, a fluid control module, fluid dispensing robotics, pick-and-place robotics, one or more processors or computers, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, image processing software packages, data analysis software packages), data storage modules, data communication modules (e.g., Bluetooth, WiFi, intranet, or internet communication hardware and associated software), a display module, or any combination thereof.
[0398] Methods for sequencing Some embodiments of the present disclosure provide methods for sequencing immobilized or non-immobilized template molecules. The methods can be operated in the system 100, for example, in the sequencer 114. In some embodiments, the immobilized template molecules comprise a plurality of nucleic acid template molecules each having one copy of a target sequence of interest. In some embodiments, the nucleic acid template molecules each having one copy of a target sequence of interest can be generated by performing bridge amplification using linear library molecules. In some embodiments, the immobilized template molecules comprise a plurality of nucleic acid template molecules (e.g., concatemers), each having two or more tandem copies of a target sequence of interest. In some embodiments, nucleic acid template molecules comprising concatemer molecules can be generated by performing rolling circle amplification of circularized linear library molecules. In some embodiments, the non-immobilized template molecules comprise circular molecules. In some embodiments, the methods for sequencing use a soluble (e.g., non-immobilized) sequencing polymerase or a sequencing polymerase immobilized on a support.
[0399] In some embodiments, the sequencing reaction uses detectably labeled nucleotide analogs. In some embodiments, the sequencing reaction uses a two-step sequencing reaction, including binding to detectably labeled polyvalent molecules and incorporating nucleotide analogs. In some embodiments, the sequencing reaction uses unlabeled nucleotide analogs. In some embodiments, the sequencing reaction uses phosphate-chain labeled nucleotides.
[0400] In some embodiments, the immobilized concatemers each comprise a tandem repeat unit of a sequence of interest (e.g., an insert region) and an optional adapter sequence. For example, the tandem repeat unit comprises: (i) a left universal adapter sequence (e.g., a surface pinning primer) having a binding sequence for a first surface primer (720), (ii) a left universal adapter sequence having a binding sequence for a first sequencing primer (740) (e.g., a forward sequencing primer), (iii) a sequence of interest (710), (iv) a right universal adapter sequence having a binding sequence for a second sequencing primer (750) (e.g., a reverse sequencing primer), (v) a right universal adapter sequence having a binding sequence for a second surface primer (730) (e.g., a surface capture primer), and (vii) a left sample index sequence (760) and / or a right sample index sequence (770). In some embodiments, the tandem repeat unit further comprises a left unique identifier sequence (780) and / or a right unique identifier sequence (790). In some embodiments, the tandem repeat unit further comprises at least one binding sequence for a compaction oligonucleotide. In some embodiments, Figures 6 and 7 show units of a linear library molecule or a concatemer molecule.
[0401] Immobilized concatemers can self-collapse into compact nucleic acid nanoballs. Including one or more compaction oligonucleotides in an RCA reaction can further compact the size and / or shape of the nanoballs. Increasing the number of tandem repeat units in a given concatemer increases the number of sites along the concatemer for hybridizing to multiple sequencing primers (e.g., sequencing primers with universal sequences) that serve as multiple initiation sites for polymerase-catalyzed sequencing reactions. When sequencing reactions use detectably labeled nucleotides and / or detectably labeled multivalent molecules (e.g., with nucleotide units), signals emitted by nucleotides or nucleotide units participating in parallel sequencing reactions along the concatemer result in increased signal intensity for each concatemer. Multiple portions of a given concatemer can be sequenced simultaneously. Furthermore, multiple binding complexes can form along a particular concatemeric molecule, each containing a sequencing polymerase bound to a template / primer duplex and a multivalent molecule, and the multiple binding complexes remain stable without dissociation, resulting in increased duration, which increases signal intensity and reduces imaging time.
[0402] Methods for sequencing using nucleotide analogs Some aspects of the present disclosure provide methods for sequencing any of the immobilized template molecules described herein, comprising: (a) contacting a sequencing polymerase with (i) a nucleic acid template molecule and (ii) a nucleic acid sequencing primer, wherein the contacting is performed under conditions suitable for binding the sequencing polymerase to the nucleic acid template molecule that hybridizes to the nucleic acid primer, such that the nucleic acid template molecule that hybridizes to the nucleic acid primer forms a nucleic acid duplex. In some aspects, the sequencing polymerase comprises a recombinant mutant sequencing polymerase that can bind and incorporate nucleotide analogs.
[0403] In some embodiments, in the method for sequencing a template molecule, the sequencing primer comprises a 3' extendable end or a 3' non-extendable end. In some embodiments, the plurality of nucleic acid template molecules comprises amplified template molecules (e.g., clonally amplified template molecules). In some embodiments, the plurality of nucleic acid template molecules comprises one copy of a target sequence of interest. In some embodiments, the plurality of nucleic acid molecules comprises two or more tandem copies of a target sequence of interest (e.g., concatemers). In some embodiments, the plurality of nucleic acid template molecules comprises the same target sequence of interest or different target sequences of interest. In some embodiments, the plurality of nucleic acid primers is in solution or immobilized on a support. In some embodiments, when the plurality of nucleic acid template molecules and / or the plurality of nucleic acid primers are immobilized o...
Claims
1. 1. A method comprising: generating, by a sequencing system, a first plurality of flow cell images of a cellular sample immobilized on a support by performing one or more sequencing reaction cycles, the cellular sample comprising a plurality of concatemeric molecules therein, a first concatemeric molecule of the plurality of concatemeric molecules corresponding to a first target RNA molecule of the cellular sample, and a second concatemeric molecule of the plurality of concatemeric molecules corresponding to a second target RNA molecule of the cellular sample; determining, by a processor, pixel intensities and respective color purity of each pixel intensities for pixels of the first plurality of flow cell images; determining, by the processor and based on the pixel intensities and the respective color purities of the pixel intensities, a base calling template including base calling positions of the cell sample at different axial positions along an axial axis, wherein the base calling template is configured to align a second plurality of flow cell images of the cell sample in one or more subsequent cycles of the one or more cycles.
2. The method of claim 1 , wherein the first or second plurality of flow cell images of the cell sample are generated at different axial positions along an axial axis.
3. The method of claim 2 , wherein the axial axis is perpendicular to an image plane and the field of view of the first or second plurality of flow cell images is within the image plane.
4. 2. The method of claim 1, further comprising: registering, by the processor, the second plurality of flow cell images of the one or more subsequent cycles of the one or more cycles to the base calling template.
5. 2. The method of claim 1, further comprising: performing, by the processor, base calling of the second plurality of flow cell images based on the base calling positions of the base calling template.
6. the first or second plurality of flow cell images are acquired at a plurality of predetermined axial positions; The method of claim 1 , wherein at least some of the different axial locations of the base calling position are different from any of the plurality of predetermined axial locations.
7. 2. The method of claim 1, wherein each of the plurality of concatemeric molecules immobilized on the support corresponds to a base-calling position.
8. The method of claim 1 , wherein the first or second plurality of flow cell images comprises the same image resolution.
9. 2. The method of claim 1, wherein the first plurality of flow cell images or the second plurality of flow cell images comprise light signals emitted from nucleotide reagents bound to A, G, C, and T / U nucleotide bases of unbalanced diversity among the plurality of concatemeric molecules immobilized on the support.
10. 10. The method of claim 9, wherein the unbalanced diversity of A, G, C, and T / U nucleotide bases among the plurality of concatemeric molecules comprises a percentage of the number of nucleotide bases of at least one type (1) relative to the total number of bases of all four types (2) in the one or more cycles is less than 20%, 15%, 10%, or 5%.
11. The method of claim 1 , wherein each pixel intensity comprises one or more sub-pixel intensities of the corresponding pixel.
12. The method of claim 1 , wherein the respective color purity of each of the pixel intensities comprises one or more color purity values corresponding to one or more sub-pixel intensities of the corresponding pixel.
13. The method of claim 1 , wherein the respective color purity of each of the pixel intensities comprises the respective color purity for one or more color channels.
14. 14. The method of claim 13, wherein determining the respective color purity of each of the pixel intensities comprises determining a ratio of a signal corresponding to (1) a particular type of nucleotide base to (2) a total amount of signals for other types of nucleotide bases.
15. determining the pixel intensities for the pixels of the first plurality of flow cell images; 14. The method of claim 13, comprising determining each channel intensity in the set of channel intensities corresponding to a respective different fluorescence wavelength based on a comparison of the sets of channel intensities at corresponding pixel or sub-pixel locations.
16. determining the base calling template based on the pixel intensities and the respective color purity of the pixel intensities; determining for each of said pixels or sub-pixels whether said respective color purity is greater than the color purity of other pixels or sub-pixels within a threshold distance; responsive to determining that the respective color purity is greater than the color purity of the other pixels or sub-pixels within the threshold distance, adding the corresponding pixel or sub-pixel location to the base calling template; and making no changes to the base-calling template in response to determining that the respective color purity is not greater than the color purity of the other pixels or sub-pixels within the threshold distance.
17. The method of claim 16 , wherein the threshold distance is three-dimensional.
18. The method of claim 1 , wherein the support comprises a flow cell.
19. 1. A system comprising: one or more hardware processors; one or more data storage devices storing instructions executable by the one or more hardware processors, the instructions, when executed, causing the one or more hardware processors to perform operations, the operations including: generating, by a sequencing system, a first plurality of flow cell images of a cellular sample immobilized on a support by performing one or more sequencing reaction cycles, the cellular sample comprising a plurality of concatemeric molecules therein, a first concatemeric molecule of the plurality of concatemeric molecules corresponding to a first target RNA molecule of the cellular sample, and a second concatemeric molecule of the plurality of concatemeric molecules corresponding to a second target RNA molecule of the cellular sample; determining, by a processor, pixel intensities and respective color purity of each pixel intensities for pixels of the first plurality of flow cell images; and determining, by the processor and based on the pixel intensities and the respective color purities of the pixel intensities, a base calling template including base calling positions, wherein the base calling template is configured to align a second plurality of flow cell images of the cell sample in one or more subsequent cycles of the one or more cycles.
20. one or more non-transitory computer storage media encoded with instructions executable by one or more hardware processors, the instructions, when executed, causing the one or more hardware processors to perform operations, the operations including: generating, by a sequencing system, a first plurality of flow cell images of a cellular sample immobilized on a support by performing one or more sequencing reaction cycles, the cellular sample comprising a plurality of concatemeric molecules therein, a first concatemeric molecule of the plurality of concatemeric molecules corresponding to a first target RNA molecule of the cellular sample, and a second concatemeric molecule of the plurality of concatemeric molecules corresponding to a second target RNA molecule of the cellular sample; determining, by a processor, pixel intensities and respective color purity of each pixel intensities for pixels of the first plurality of flow cell images; and determining, by the processor and based on the pixel intensities and the respective color purities of the pixel intensities, a base calling template including base calling positions, wherein the base calling template is configured to align a second plurality of flow cell images of the cell sample in one or more subsequent cycles of the one or more cycles.