Spatial analysis of analytes
By using reference markers and capture points on the substrate, combined with microscopy techniques and alignment algorithms, the problem of difficult sample image alignment was solved, enabling efficient and accurate spatial analysis of analytes and reducing the need for training and labor.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 10X GENOMICS INC
- Filing Date
- 2020-11-18
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies suffer from sample quality defects during tissue imaging, making it difficult to align sample images with barcode capture probes. Furthermore, traditional methods require extensive training and labor, leading to human error and hindering efficient and accurate analysis of spatial analytes.
Using a substrate containing reference markers and capture points, sample images are acquired via transmission or fluorescence microscopy. The reference markers and capture probe arrays are aligned, and a local alignment algorithm and a heuristic classifier are combined to achieve reproducible alignment of the sample and efficient analyte localization.
It achieves accurate alignment between sample images and capture probes, reduces human error, provides a cost-effective and user-friendly spatial analysis tool, and improves the accuracy and efficiency of analyte analysis.
Smart Images

Figure CN117078725B_ABST
Abstract
Description
[0001] This application is a divisional application of application number 202080094371.4, filed on November 18, 2020, entitled "Spatial Analysis of Analytes".
[0002] Cross-reference to related applications
[0003] This application claims priority to U.S. Provisional Patent Application No. 63 / 041,825, entitled "Pipeline for Spatial Analysis of Analytes," filed June 20, 2020; U.S. Provisional Patent Application No. 62 / 980,073, entitled "Pipeline for Analysis of Analytes," filed February 21, 2020; and U.S. Provisional Patent Application No. 62 / 938,336, entitled "Pipeline for Analysis of Analytes," filed November 21, 2019, each of which is incorporated herein by reference in its entirety. Technical Field
[0004] This specification describes techniques for processing observed analyte data (such as spatially arranged next-generation sequencing data) in large, complex datasets and using that data to visualize patterns. Background Technology
[0005] Spatial resolution of analytes in complex tissues provides new insights into fundamental processes of biological function and morphology, such as cell fate and development, disease progression and detection, and regulatory networks at the cellular and tissue levels. See Satija et al., 2015, “Spatial reconstruction of single-cell gene expression data,” *Nature Biotechnology* 33, 495-502, doi:10.1038.nbt.3192, and Achim et al., 2015, “High-throughput spatial mapping of single-cell RNA-seq data to tissue of origin,” *Nature Biotechnology* 33:503-509, doi:10.1038 / nbt.3209, each incorporated herein by reference in its entirety. Understanding spatial patterns or other forms of relationships between analytes can provide information about differential cellular behavior. This, in turn, helps to elucidate complex symptoms, such as complex diseases. For example, determining the association between the abundance of an analyte (e.g., a gene) and a tissue subpopulation of a specific tissue class (e.g., diseased tissue, healthy tissue, the boundary between diseased and healthy tissues, etc.) provides inferential evidence that the analyte is associated with a condition such as a complex disease. Similarly, determining the association between the abundance of an analyte and a specific subpopulation of a heterogeneous cell population in a complex 2D or 3D tissue (e.g., the mammalian brain, liver, kidney, heart, tumor, or developing embryo of a model organism) provides inferential evidence that the analyte is associated with a specific subpopulation.
[0006] Therefore, spatial analysis of analytes can inform the early detection of diseases by identifying risk regions in complex tissues and characterizing the analyte maps present in these regions through spatial reconstruction (e.g., gene expression, protein expression, DNA methylation, and / or single nucleotide polymorphisms). High-resolution spatial mapping of analytes to their specific locations within regions or subregions reveals spatial expression patterns of analytes, provides relevant data, and further suggests analyte network interactions associated with diseases or other morphologies or phenotypes of interest, leading to a holistic understanding of cells within their morphological context. See, 10X, 2019, "Spatially-Resolved Transcriptomics," 10X, 2019, "Inside Visium Spatial Technology," and 10X, 2019, "Visium Spatial Gene Expression Solution," each incorporated herein by reference in its entirety.
[0007] Spatial analysis of analytes can be performed by capturing the analyte and / or analyte capture agents or analyte binding domains and mapping them to known locations using a reference image indicating the tissue or region of interest corresponding to a known location (e.g., using a barcode capture probe attached to a substrate). For example, in some embodiments of spatial analysis, a sample is prepared (e.g., freshly frozen tissue sections are placed on a slide, fixed, and / or stained for imaging). Imaging of the sample provides a reference image for spatial analysis. Analyte detection is then performed using, for example, analyte or analyte ligand capture via a barcode capture probe, library construction, and / or sequencing. The resulting barcode analyte data and reference image can be combined during data visualization for spatial analysis. See, 10X, 2019, *Internal Visium Spatial Technology*.
[0008] One challenge in this analysis is ensuring proper alignment of the sample or its image (e.g., a tissue section or image of a tissue section) with the barcode capture probe (e.g., using reference alignment). The frequent occurrence of sample quality defects in conventional wet laboratory methods for tissue sample preparation and sectioning further complicates the technical limitations in this field. These problems arise from the nature of the tissue sample itself (particularly including interstitial regions, vacuoles, and / or general grain size that is often difficult to interpret after imaging) or from improper handling or sample degradation, which results in gaps or pores in the sample (e.g., torn samples or obtaining only partial samples, as in biopsies). Additionally, wet laboratory methods for imaging can lead to other defects, including but not limited to bubbles, debris, crystalline staining particles deposited on the substrate or tissue, inconsistent or poor contrast staining, and / or microscopic limitations resulting in blurred images, overexposure or underexposure, and / or poor resolution. See Uchida, 2013, “Image processing and recognition for biological images,” Develop.GrowthDiffer. 55, 523-549, doi:10.1111 / dgd.12054, which is incorporated herein by reference in its entirety. This deficiency makes alignment more difficult.
[0009] Therefore, there is a need in the art for improved systems and methods for the analysis of spatial analytes (e.g., nucleic acids and proteins). Such systems and methods would allow for reproducible identification and alignment of tissue samples in images without requiring significant training and labor costs, and would further improve identification accuracy by eliminating human error caused by subjective alignment. Such systems and methods would also provide practitioners with cost-effective, user-friendly tools to reliably perform spatial analyte analyses. Summary of the Invention
[0010] This disclosure provides technical solutions (e.g., computing systems, methods, and non-transitory computer-readable storage media) for solving the above-mentioned problems.
[0011] The following summary of the invention is presented to provide a basic understanding of some aspects of this disclosure. This summary is not an extensive overview of the invention. It is not intended to identify key / essential elements of the invention or to define its scope. Its sole purpose is to present some concepts of the invention in a simplified form as a prelude to the more detailed description that follows.
[0012] One aspect of this disclosure provides a spatial analysis method for an analyte, comprising: A) placing a sample (e.g., a tissue section) on a substrate, wherein the substrate includes a plurality of reference markers and a set of capture points. In some embodiments, the set of capture points includes at least 1,000, 2,000, 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, 50,000, 55,000, 60,000, 65,000, 70,000, 75,000, 80,000, 85,000, 90,000, 95,000, or 100,000 capture points. The reference markers are not directly or indirectly associated with the analyte. Instead, the reference markers serve to provide a reference frame for the substrate. In some embodiments, there are more than 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 500, or 1000 reference markers. In some embodiments, there are fewer than 1000 reference markers.
[0013] Obtain one or more images of a biological sample on a substrate. Each of the one or more images contains a corresponding plurality of pixels in the form of a pixel value array. In some embodiments, the pixel value array contains at least 100, 10,000, 100,000, or 1x10^6 pixels. 6 2x10 6 3x10 6 5x10 6 8x10 6 10x10 6 Or 15x10 6 Pixel values. In some embodiments, one or more images are acquired using a transmission optical microscope. In some embodiments, one or more images are acquired using a fluorescence microscope. After A), multiple sequence readings are obtained electronically from the set of capture points. In some embodiments, the multiple sequence readings include more than 100, 1000, 50,000, 100,000, 500,000, 1x10^6 pixels. 6 2x10 6 3x10 6 Or 5x10 6Each sequence readout. For each given image in one or more images, each corresponding capture probe group in a set of capture probe groups (i) is located at a different capture point in that set of capture points, and (ii) is associated directly or indirectly (e.g., by an analyte capture agent) with one or more analytes (e.g., nucleic acids, proteins, and / or metabolites, etc.) from a biological sample from the slice. In some embodiments, each corresponding capture probe group in the set of capture probe groups is characterized by at least one unique spatial barcode from a plurality of spatial barcodes.
[0014] In some embodiments, the substrate may have two or more capture points with the same spatial barcode. That is, there is no unique spatial barcode between two capture points. In some such embodiments, these capture points with repeating spatial barcodes are considered a single capture point. In other embodiments, capture points without unique spatial barcodes are not considered part of a capture point group used to locate corresponding sequence readings to capture a particular set of capture points.
[0015] In some embodiments, at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, or at least 40% of the capture points on the substrate may not have a unique spatial barcode. That is, for each such capture point and for each corresponding spatial barcode, there is at least one other capture point on the substrate that has a corresponding spatial barcode. In some such embodiments, these capture points without unique spatial barcodes are not considered part of the capture point set used to locate the corresponding sequence reading to capture a particular set of capture points.
[0016] In some embodiments, at least ten, at least 100, at least 1000, at least 10,000, at least 100,000, or at least 1,000,000 capture points on the substrate may not have unique spatial barcodes. That is, for each corresponding spatial barcode of each such capture point, there is at least one other capture point on the substrate with a corresponding spatial barcode. In some such embodiments, these capture points without unique spatial barcodes are not considered part of a capture point set used to locate corresponding sequence readings to capture a particular set of capture points.
[0017] Multiple sequence readings comprise all or part of the sequence readings corresponding to one or more analytes. Each corresponding sequence reading in the multiple sequence readings includes a spatial barcode for a group of capture probes corresponding to that set of capture probes. The multiple spatial barcodes are used to locate the corresponding sequence reading in the multiple sequence readings to a corresponding capture point in that set of capture points, thereby dividing the multiple sequence readings into multiple sequence reading subsets, each sequence reading subset corresponding to a different capture point in that set of capture points. For each corresponding image in one or more images, multiple reference markers are used to provide a corresponding composite representation, the composite representation comprising (i) the corresponding image aligned with that set of capture points on the substrate and (ii) a representation of each sequence reading subset mapped to a corresponding capture point on the substrate at a corresponding position within the corresponding image.
[0018] In some embodiments, a corresponding composite representation of images in one or more images provides the relative abundance of nucleic acid fragments of each of a plurality of genes at each of a plurality of capture points.
[0019] In some embodiments, a process comprising the following steps is used to align a corresponding image with a set of capture points on a substrate: analyzing an array of pixel values to identify multiple derived reference points of the corresponding image; selecting a first template from a plurality of templates using a substrate identifier uniquely associated with the substrate, wherein each template contains reference positions and corresponding coordinate systems for the multiple reference reference points; aligning the multiple derived reference points of the corresponding image with the corresponding multiple reference reference points of the first template using an alignment algorithm to obtain a transformation between the multiple derived reference points of the corresponding image and the corresponding multiple reference reference points of the first template; and using the transformation and the coordinate system of the first template to locate the corresponding position of each capture point in the set of capture points in the corresponding image.
[0020] In some embodiments, the alignment algorithm is a local alignment that uses a scoring system to align a corresponding sequence read with a reference sequence. This scoring system (i) penalizes mismatches between nucleotides in the corresponding sequence read and their corresponding nucleotides in the reference sequence based on a permutation matrix, and (ii) penalizes gaps introduced into the alignment of the sequence read with the reference sequence. In some such embodiments, the local alignment is a Smith-Waltman alignment. In some such embodiments, the reference sequence is all or part of a reference genome. In some embodiments, one or more sequence reads that do not cover any of the multiple loci are removed from the multiple sequence reads. In some embodiments, the multiple sequence reads for a given image are RNA sequence reads, and removal includes removing one or more sequence reads that overlap with splice sites in the reference sequence. In some embodiments, the multiple loci include one or more loci on a first chromosome and one or more loci on a second chromosome different from the first chromosome.
[0021] In some embodiments, the coordinate system of the transformation and the first template is used to locate and measure one or more optical properties of each of the set of capture points by assigning each corresponding pixel in the plurality of pixels to a first class or a second class. The first class represents a biological sample on a substrate, and the second class represents the background. In some embodiments, this is accomplished by a process comprising the following steps: (i) defining bounding boxes within the corresponding image using a plurality of reference markers, (ii) removing corresponding pixels from the plurality of pixels that fall outside the bounding boxes, (iii) after removing (ii), running a plurality of heuristic classifiers on the plurality of pixels (e.g., in a grayscale space), wherein, for each corresponding pixel in the plurality of pixels, each of the plurality of heuristic classifiers votes between the first class and the second class for the corresponding pixel, thereby forming a corresponding aggregate score for each corresponding pixel in the plurality of pixels, and (iv) applying the aggregate score and intensity of each corresponding pixel in the plurality of pixels to a segmentation algorithm (e.g., image cutting) to independently assign a probability as tissue or background to each corresponding pixel in the plurality of pixels.
[0022] In some embodiments, each corresponding aggregate score is one of a set of classes including obvious first class, possible first class, possible second class and obvious second class.
[0023] In some embodiments, the method further includes executing a procedure for each corresponding locus among a plurality of loci. In such embodiments, the method includes aligning each corresponding sequence readout among a plurality of sequence readouts mapped to the corresponding locus to determine the morphological identity of the corresponding sequence readouts from a set of corresponding morphologies of the corresponding locus. Each corresponding sequence readout among the plurality of sequence readouts mapped to the corresponding locus is classified by spatial barcoding of the corresponding sequence readouts and by morphological identity to determine the spatial distribution of one or more morphologies in the biological sample. For each capture point in the set of capture points on the substrate, the spatial distribution includes the abundance of each morphology in the morphology set of each locus among the plurality of loci.
[0024] In some such embodiments, the method further includes using spatial distribution to characterize the subject's biological status. For example, in some embodiments, the biological status is the absence or presence of a disease. In some embodiments, the biological status is cancer. In some embodiments, the biological status is a stage of a disease. In some embodiments, the biological status is a stage of cancer.
[0025] In some embodiments, an organization mask is overlaid on the corresponding image. The organization mask assigns a first attribute to each of the plurality of pixels in the corresponding image that is more likely to be assigned as organization, and assigns a second attribute to each of the plurality of pixels that is more likely to be assigned as background. In some embodiments, the first attribute is a first color (e.g., one of red and blue), and the second attribute is a second color (e.g., the other of red and blue).
[0026] In some embodiments, the first attribute is a first level of brightness or opacity, and the second attribute is a second level of brightness or opacity.
[0027] In some embodiments, based on the allocation of pixels near the corresponding representation of a capture point in the composite representation, the corresponding representation of each of the plurality of capture points in the composite representation is assigned a first attribute or a second attribute.
[0028] In some embodiments, the capture points in the set of capture points include capture domains. In some embodiments, the capture points in the set of capture points include cutting domains. In some embodiments, each capture point in the set of capture points is attached directly or indirectly to a substrate.
[0029] In some embodiments, one or more analytes comprise five or more analytes, ten or more analytes, fifty or more analytes, one hundred or more analytes, five hundred or more analytes, one thousand or more analytes, two thousand or more analytes, or two thousand or more analytes to one hundred,000,000 analytes.
[0030] In some embodiments, the unique spatial barcode pair is selected from the sets {1,…,1024}, {1,…,4096}, {1,…,16384}, {1,…,65536}, {1,…,262144}, {1,…,1048576}, {1,…,4194304}, {1,…,16777216}, {1,…,67108864}, or {1,…,1x10}. 12 The unique predefined value selected in} is encoded.
[0031] In some embodiments, each corresponding capture probe group in the set of capture probe groups includes 1000 or more capture probes, 2000 or more capture probes, 10,000 or more capture probes, 100,000 or more capture probes, 1x10 6 or more capture probes, 2x10 6 More or more capture probes, or 5x10 6 Or more capture probes.
[0032] In some embodiments, each capture probe in a corresponding capture probe group includes a poly-A sequence or a poly-T sequence and a unique spatial barcode characterizing the corresponding capture probe group. In some embodiments, each capture probe in a corresponding capture probe group includes a spatial barcode identical to the plurality of spatial barcodes. In some embodiments, each capture probe in a corresponding capture probe group includes a spatial barcode different from the plurality of spatial barcodes.
[0033] In some embodiments, the biological sample is a tissue slice sample with a depth of 100 micrometers or less. In some embodiments, each corresponding slice in a plurality of tissue slices of the sample is considered a “spatial projection”, and multiple co-aligned images are taken for each slice.
[0034] In some embodiments, one or more analytes are multiple analytes, and each capture point in the set of capture points includes multiple capture probes. Each capture probe includes a capture domain characterized by one of multiple capture domain types, and each of the multiple capture domain types is configured to combine different analytes from the multiple analytes. In some such embodiments, the multiple capture domain types include 5 to 15,000 capture domain types, and for each of the multiple capture domain types, the corresponding set of capture probes includes at least five, at least 10, at least 100, or at least 1,000 capture probes.
[0035] In some embodiments, one or more analytes are multiple analytes, and each capture point in the set of capture points includes multiple capture probes, and each of the multiple capture probes includes a capture domain characterized by a single capture domain type configured to combine each of the multiple analytes in an unbiased manner.
[0036] In some embodiments, each corresponding capture point in the set of capture points is contained within a 100 μm × 100 μm square on the substrate. In some embodiments, each corresponding capture point in the set of capture points is contained within a 50 μm × 50 μm square on the substrate. In some embodiments, each corresponding capture point in the set of capture points is contained within a 10 μm × 10 μm square on the substrate. In some embodiments, each corresponding capture point in the set of capture points is contained within a 1 μm × 1 μm square on the substrate. In some embodiments, each corresponding capture point in the set of capture points is contained within a 500 nm × 500 nm square on the substrate. In some embodiments, each corresponding capture point in the set of capture points is contained within a 300 nm × 300 nm square on the substrate. In some embodiments, each corresponding capture point in the set of capture points is contained within a 200 nm × 200 nm square on the substrate.
[0037] In some embodiments, the distance between the center of each corresponding capture point in the set of capture points on the substrate and the adjacent capture point is between 40 micrometers and 300 micrometers. In some embodiments, the distance between the center of each corresponding capture point in the set of capture points on the substrate and the adjacent capture point is between 300 nanometers and 5 micrometers, between 400 nanometers and 4 micrometers, between 500 nanometers and 3 micrometers, between 600 nanometers and 2 micrometers, or between 700 nanometers and 1 micrometer.
[0038] In some embodiments, each of the trap points in the set has a diameter of 80 micrometers or less. In some embodiments, the diameter of each trap point in the set is between 25 micrometers and 65 micrometers, between 5 micrometers and 50 micrometers, between 2 micrometers and 7 micrometers, or between 800 nanometers and 1.5 micrometers.
[0039] In some embodiments, the distance between the center of each respective capture point in the set of capture points on the substrate and the adjacent capture point is between 40 micrometers and 100 micrometers, between 300 nanometers and 15 micrometers, between 400 nanometers and 10 micrometers, between 500 nanometers and 8 micrometers, between 600 nanometers and 6 micrometers, between 700 nanometers and 5 micrometers, or between 800 nanometers and 4 micrometers.
[0040] In some embodiments, the plurality of heuristic classifiers includes a first heuristic classifier that identifies a single intensity threshold that divides a plurality of pixels into a first class and a second class, thereby causing the first heuristic classifier to vote for either the first class or the second class for each corresponding pixel among the plurality of pixels. The single intensity threshold represents minimizing the intra-class intensity variance between the first class and the second class or maximizing the inter-class variance between the first class and the second class. In some embodiments, the plurality of heuristic classifiers includes a second heuristic classifier that identifies local neighborhoods of pixels having the same class identified using the first heuristic method, and applies a smoothing measure of the maximum difference in intensity between pixels in the local neighborhood, thereby causing the second heuristic classifier to vote for either the first class or the second class for each corresponding pixel among the plurality of pixels. In some embodiments, the plurality of heuristic classifiers includes a third heuristic classifier that performs edge detection on the plurality of pixels to form a plurality of edges in the image, morphologically closes the plurality of edges to form a plurality of morphologically closed regions in the image, and assigns pixels within the morphologically closed regions to the first class, and assigns pixels outside the morphologically closed regions to the second class, thereby causing the third heuristic classifier to vote for either the first class or the second class for each corresponding pixel among the plurality of pixels. In some embodiments, the plurality of heuristic classifiers consist of a first heuristic classifier, a second heuristic classifier, and a third heuristic classifier. Each corresponding pixel assigned to a second class by each of the plurality of classifiers is labeled as obviously belonging to the second class, while each corresponding pixel assigned to a first class by each of the plurality of classifiers is labeled as obviously belonging to the first class.
[0041] In some embodiments, the segmentation algorithm is a graph cutting and segmentation algorithm, such as GrabCut.
[0042] In some embodiments, the cleavage domain comprises a sequence that is recognized and cleaved by uracil-DNA glycosylase and / or endonuclease VIII.
[0043] In some embodiments, the capture probe group in the capture probe group does not contain a cutting domain and is not cut from the array.
[0044] In some embodiments, one or more analytes comprise DNA or RNA.
[0045] In some embodiments, each of the capture probe groups in the group is directly or indirectly attached to the substrate.
[0046] In some embodiments, obtaining sequence reads includes in situ sequencing of the set of capture sites on the substrate. In some embodiments, obtaining sequence reads includes high-throughput sequencing of the set of capture sites on the substrate.
[0047] In some embodiments, the corresponding locus among the plurality of loci is a biallelic locus, and the corresponding haplotype set of the corresponding locus consists of a first allele and a second allele. In some such embodiments, the corresponding locus includes a heterozygous single nucleotide polymorphism (SNP), a heterozygous insertion, or a heterozygous deletion.
[0048] In some embodiments, the plurality of sequence readings includes 10,000 or more sequence readings, 50,000 or more sequence readings, 100,000 or more sequence readings, or 1x10 6 One or more sequence readings.
[0049] In some embodiments, the plurality of loci comprises two to 100 loci, more than 10 loci, more than 100 loci, or more than 500 loci.
[0050] In some embodiments, a unique spatial barcode in each sequence readout is located within a set of consecutive oligonucleotides in the corresponding sequence readout. In some such embodiments, the set of consecutive oligonucleotides is an N-mer, where N is an integer selected from the set {4,…,20}.
[0051] In some embodiments, the corresponding plurality of sequence readings include sequence readings paired at the 3' end or the 5' end.
[0052] In some embodiments, one or more analytes are multiple analytes, and a corresponding capture probe group in the set of capture probe groups includes multiple capture probes, and each of the multiple capture probes includes a capture domain characterized by a single capture domain type configured to combine each of the multiple analytes in an unbiased manner.
[0053] Another aspect of this disclosure provides a computer system comprising one or more processors and memory. One or more programs are stored in the memory and configured to be executed by the one or more processors. It should be understood that the memory may be on a single computer, a computer network, one or more virtual machines, or in a cloud computing architecture. The one or more programs are used for spatial analysis of the analyte. The one or more programs include instructions for obtaining one or more images of a biological sample (e.g., a tissue section sample, each slice of the tissue section sample, etc.) respectively on a substrate, wherein each instance of the substrate includes a plurality of reference markers and a set of capture points, and wherein each corresponding image contains a plurality of pixels in the form of a pixel value array. In some embodiments, the pixel value array contains at least 100, 10,000, 100,000, or 1x10^6 pixels. 6 2x10 6 3x10 6 5x10 6 8x10 610x10 6 Or 15x10 6 Each set of capture probes is a pixel value. Sequence reads are obtained from the set of capture points in multiple electronic forms after the biological sample is placed on the substrate. Each corresponding capture probe group in the set of capture probe groups (i) is located at a different capture point in the set of capture points, and (ii) is associated with one or more analytes from the biological sample. Each corresponding capture probe group in the set of capture probe groups is characterized by at least one unique spatial barcode from a plurality of spatial barcodes. The plurality of sequence reads includes all or part of the sequence reads corresponding to one or more analytes. Furthermore, each corresponding sequence read in the plurality of sequence reads includes a spatial barcode of the corresponding capture probe group in the set of capture probes. The plurality of spatial barcodes are used to locate the corresponding sequence reads in the plurality of sequence reads to the corresponding capture point in the set of capture points, thereby dividing the plurality of sequence reads into a plurality of sequence read subsets, each sequence read subset corresponding to a different capture point in the plurality of capture points. A plurality of reference markers are used to provide a composite representation comprising (i) an image aligned with the set of capture points on the substrate and (ii) a representation of each subset of sequence reads mapped to the corresponding capture point on the substrate at the corresponding position within the image.
[0054] Another aspect of this disclosure provides a computer-readable storage medium storing one or more programs. The one or more programs contain instructions, when executed by an electronic device having one or more processors and a memory, that cause the electronic device to perform spatial analysis of an analyte by acquiring an image of a biological sample (e.g., a tissue section) on a substrate. The substrate includes a plurality of reference markers and a set of capture points, and the image contains a plurality of pixels in the form of an array of pixel values. After the biological sample is on the substrate, sequence readings are obtained from the set of capture points in a plurality of electronic forms. Each corresponding capture probe group in a set of capture probe groups (i) is located at a different capture point in the set of capture points, and (ii) is associated with one or more analytes from the biological sample. Each corresponding capture probe group in the set of capture probe groups is characterized by at least one unique spatial barcode from a plurality of spatial barcodes. The plurality of sequence readings include sequence readings corresponding to all or part of one or more analytes. Furthermore, each corresponding sequence reading in the plurality of sequence readings includes a spatial barcode of the corresponding capture probe group in the set of capture probes. The multiple spatial barcodes are used to locate corresponding sequence readings from the multiple sequence readings to corresponding capture points in the set of capture points, thereby dividing the multiple sequence readings into multiple subsets of sequence readings, each subset corresponding to a different capture point in the multiple capture points. Multiple reference markers are used to provide a composite representation comprising (i) an image aligned with the set of capture points on the substrate and (ii) a representation of each subset of sequence readings mapped to the corresponding capture point on the substrate at the corresponding position within the image.
[0055] Another aspect of this disclosure provides a computing system comprising one or more processors and memory storing one or more programs for spatial nucleic acid analysis. It should be understood that the memory may reside on a single computer, a computer network, one or more virtual machines, or within a cloud computing architecture. The one or more programs are configured to be executed by the one or more processors. The one or more programs include instructions for performing any of the methods disclosed above.
[0056] Another aspect of this disclosure provides a computer-readable storage medium for storing one or more programs to be executed by an electronic device. The one or more programs include instructions for the electronic device to perform spatial nucleic acid analysis by any of the methods disclosed above. It should be understood that the computer-readable storage medium can exist as a single computer-readable storage medium or as any number of component computer-readable storage media that are physically separate from each other.
[0057] Other embodiments relate to systems, portable consumer devices, and computer-readable media associated with the methods described herein.
[0058] As disclosed herein, any embodiments disclosed herein may be applied in any way where applicable.
[0059] Various embodiments of the systems, methods, and apparatuses within the scope of the appended claims each have several aspects, with no single aspect solely responsible for the desired properties described herein. Without limiting the scope of the appended claims, some prominent features are described herein. After considering this discussion, and particularly after reading the section entitled "Detailed Description," one will understand how to use the features of the various embodiments.
[0060] By incorporating via reference
[0061] All publications, patents, patent applications, and information available on the Internet and mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent, patent application, or information item is specifically and individually indicated to be incorporated by reference. Where any Internet-available publication, patent, patent application, or information item incorporated by reference contradicts the disclosure contained in this specification, the specification is intended to supersede and / or take precedence over any such contradictory material. Attached Figure Description
[0062] The following figures illustrate certain embodiments of the features and advantages of this disclosure. These embodiments are not intended to limit the scope of the appended claims in any way. In several views of this patent application, the same reference numerals denote the same elements.
[0063] Figure 1An exemplary spatial analysis workflow according to embodiments of the present disclosure is illustrated.
[0064] Figure 2 An exemplary spatial analysis workflow according to embodiments of the present disclosure is shown, wherein optional steps are indicated by dashed boxes.
[0065] Figure 3A and 3B An exemplary spatial analysis workflow according to embodiments of the present disclosure is illustrated, wherein in Figure 3A In the diagram, optional steps are indicated by dashed boxes.
[0066] Figure 4 An exemplary spatial analysis workflow according to embodiments of the present disclosure is shown, wherein optional steps are indicated by dashed boxes.
[0067] Figure 5 An exemplary spatial analysis workflow according to embodiments of the present disclosure is shown, wherein optional steps are indicated by dashed boxes.
[0068] Figure 6 This is a schematic diagram illustrating an example of a barcode capture probe as described herein, according to an embodiment of the present disclosure.
[0069] Figure 7 This is a schematic diagram illustrating a severable capture probe according to an embodiment of the present disclosure.
[0070] Figure 8 This is a schematic diagram of the capture points of spatial markers for exemplary multiplexing according to embodiments of the present disclosure.
[0071] Figure 9 Details of the spatial capture point and capture probe according to embodiments of the present disclosure are illustrated.
[0072] Figure 10A , 10B Figures 10C, 10D, and 10E, 10F illustrate non-limiting methods for spatial nucleic acid analysis according to some embodiments of the present disclosure, wherein optional steps are shown by dashed boxes.
[0073] Figure 11A and 11B This is an example block diagram illustrating a computing device according to some embodiments of the present disclosure.
[0074] Figure 12 This is a schematic diagram illustrating the arrangement of barcode capture points within an array according to some embodiments of the present disclosure.
[0075] Figure 13 This is a schematic diagram illustrating a side view of an anti-diffusion medium (e.g., a cover) according to some embodiments of the present disclosure.
[0076] Figure 14 A substrate is illustrated with an image of a biological sample (e.g., a tissue sample) on a substrate according to an embodiment of the present disclosure.
[0077] Figure 15 A substrate having multiple capture areas and a substrate identifier is illustrated according to an embodiment of the present disclosure.
[0078] Figure 16 A base having a plurality of reference markers and a set of capture points according to an embodiment of the present disclosure is illustrated.
[0079] Figure 17 Images of biological samples (e.g., tissue samples) on a substrate according to embodiments of the present disclosure are illustrated, wherein the biological samples are located within a plurality of reference markers.
[0080] Figure 18 A template illustrating the reference positions of a plurality of corresponding reference points and a corresponding coordinate system according to an embodiment of the present disclosure is shown.
[0081] Figure 19 The illustration shows how a template according to an embodiment of the present disclosure uses a corresponding coordinate system to specify the position of the set of capture points of the base relative to a reference point of the base.
[0082] Figure 20 The illustration depicts a substrate design for an image according to an embodiment of the present disclosure, including a plurality of reference markers and a set of capture points, the image including corresponding derived reference points.
[0083] Figure 21 The illustration depicts the use of a coordinate system of transformations and templates to register an image to a base according to an embodiment of the present disclosure, thereby registering the image to the set of capture points of the base.
[0084] Figure 22 The illustration illustrates the analysis of an image that identifies tissue-covered capture points on the substrate after image registration with the substrate, using template transformations and coordinate systems according to embodiments of the present disclosure.
[0085] Figure 23 The illustration depicts capture points on a substrate that has been covered by an organization, according to an embodiment of the present disclosure.
[0086] Figure 24 The illustration depicts the extraction of barcodes and UMIs from each sequence reading in nucleic acid sequencing data associated with a substrate, according to embodiments of the present disclosure.
[0087] Figure 25 Alignment of sequence reads with a reference genome according to an embodiment of the present disclosure is illustrated.
[0088] Figure 26 The illustration illustrates that, according to embodiments of the present disclosure, due to random fragments occurring during workflow steps, even though sequence readings share barcodes and UMIs, not all sequence readings are mapped to the exact same location.
[0089] Figure 27 The illustration depicts how, according to embodiments of the present disclosure, the barcode for each sequence reading is verified against a whitelist of actual barcodes (e.g., in some embodiments, the whitelist corresponds to chromium single-cell 3'v3 chemigel beads with approximately 3.6 million different barcodes, and thus the whitelist has 3.6 million barcodes).
[0090] Figure 28 The illustration shows how unique molecular identifiers (UMIs) that share cell barcodes and genes with a sequence read that has a higher count of UMIs are corrected to that UMI, according to some embodiments of the present disclosure.
[0091] Figure 29 The illustration depicts how, according to some embodiments of the present disclosure, UMI counts are formed from the original feature barcode matrix using only trusted mapping readings with valid barcodes and UMIs.
[0092] Figure 30 The illustration depicts how a secondary analysis is performed on a barcode (filtered feature barcode matrix) referred to as a cell according to an embodiment of the present disclosure, wherein principal component analysis of a normalized filtered gene-cell matrix is used to reduce G genes to the top 10 metagenes, t-SNE is run in PCA space to generate a two-dimensional projection, graph-based (Louvain) and k-means clustering (k = 2…10) is performed in PCA space to identify cell clusters, and the sSeq (negative binomial test) algorithm is used to find the gene that most uniquely defines each cluster.
[0093] Figure 31 A pipeline for analyzing images (e.g., tissue images) by combining nucleic acid sequencing data associated with each of a plurality of capture points is illustrated, thereby performing spatial nucleic acid analysis according to this disclosure.
[0094] Figure 32 The illustration shows how the analysis of tissue images based on nucleic acid sequencing data, according to this disclosure, can be used to observe captured clusters within the context of the image.
[0095] Figure 33 The illustration shows how the analysis of tissue images using combined nucleic acid sequencing data, according to some embodiments of the present disclosure, can be included in the context of the image by zooming in to a cover map of the captured point clusters in order to see more detail.
[0096] Figure 34The illustrations illustrate how the analysis of tissue images incorporating nucleic acid sequencing data, according to some embodiments of the present disclosure, can be used to create custom classes and clusters for differential expression analysis.
[0097] Figure 35 The illustration shows how the analysis of tissue images incorporating nucleic acid sequencing data, according to some embodiments of the present disclosure, can be used to view genes expressed in the context of the tissue image.
[0098] Figure 36A , 36B Images 36C, 36D, 36E, 36F, 36G, 36H, and 36I illustrate image input of tissue sections on a substrate according to some embodiments of the present disclosure. Figure 36A Outputs of various heuristic classifiers Figure 36B , 36C 36D, 36E, 36F, and 36G, as well as the output of the segmentation algorithm. Figure 36H and 36I .
[0099] Figure 37 The diagram illustrates reaction schemes for preparing sequence reads for spatial analysis according to some embodiments of the present disclosure.
[0100] Figure 38A An embodiment according to the present disclosure is illustrated in which all images of the spatial projection are fluorescent images and are displayed.
[0101] Figure 38B The illustration depicts embodiments according to this disclosure. Figure 38A The spatial projection, which only displays the CD3 channel fluorescence image of the spatial projection.
[0102] Figure 38C The illustration depicts embodiments according to this disclosure. Figure 38B The image shows CD3 quantification based on measured intensity.
[0103] Figure 39 The illustrations depict immunofluorescence images according to some embodiments of the present disclosure, representations of all or part of each subset of sequence readings at each corresponding location within one or more images mapped to corresponding capture points corresponding to corresponding locations, and composite representations.
[0104] Figure 40 This is a schematic diagram of an exemplary analyte capture agent according to some embodiments of the present disclosure.
[0105] Figure 41A This is a schematic diagram depicting an exemplary interaction between a feature-fixed capture probe and an analyte capture agent according to some embodiments of the present disclosure.
[0106] Figure 41B This is an exemplary schematic diagram showing an analyte binding moiety containing an oligonucleotide having a capture binding domain (represented by a poly(A) sequence) that hybridizes with a blocking domain (represented by a poly(T) sequence).
[0107] Figure 41C This is an exemplary schematic diagram showing an analyte binding moiety comprising an oligonucleotide containing a hairpin sequence positioned between a blocking domain (represented by a poly(U) sequence) and a capture-binding domain (represented by a poly(A) sequence). As shown, the blocking domain hybridizes with the capture-binding domain.
[0108] Figure 41D This is an exemplary schematic diagram showing the blocking domain released by RNAse H.
[0109] Figure 41E This is an exemplary schematic diagram showing an analyte binding portion comprising an oligonucleotide containing a capture-binding domain blocked using cage-like nucleotides (represented by pentagons).
[0110] Figure 42 This is an exemplary schematic diagram illustrating a spatially labeled analyte trap according to some embodiments of the present disclosure, wherein the analyte trapping sequence is blocked by a blocking probe, and wherein the blocking probe can be removed, for example, by treatment with RNAse.
[0111] Figure 43 This is a schematic diagram illustrating exemplary, non-limiting, non-exhaustive steps for spatial analyte identification in biological samples following antibody staining, according to some embodiments of the present disclosure, wherein the sample is immobilized, stained with fluorescent antibodies and spatially labeled analyte trapping agents, and imaged to detect the spatial location of target analytes within the biological sample.
[0112] Figure 44 Exemplary multiplex imaging results according to some embodiments of the present disclosure are shown, wherein immunofluorescence images show immunofluorescence staining of CD29 and CD4 in tissue sections of mouse spleen (far left), while a series of images in the right figures show the results of a multiplexed, spatially labeled analyte capture workflow, wherein the spatial locations of target proteins CD29, CD3, CD4, CD8, CD19, B220, F4 / 80, and CD169 are visualized by sequencing the analytes-corresponding analyte binding partial barcodes.
[0113] Figure 45 An exemplary workflow for spatial proteomics and genome analysis according to some embodiments of this disclosure is shown.
[0114] Figure 46A A schematic diagram of the analyte trapping agent and the spatial gene expression slide is shown.
[0115] Figure 46B This image shows a merged fluorescence image of DAPI staining on a section of human cerebellar tissue.
[0116] Figure 46C Showing the coverage Figure 46B On Figure 46B Spatial transcriptome analysis of the human cerebellum.
[0117] Figure 46D The t-SNE projection of the sequencing data is shown, plotting the data from... Figure 46C Clustering of cerebellar cell types.
[0118] Figure 46E The spatial gene expression (top) and protein staining (bottom) of the astrocyte marker glutamine synthase (produced by hybridoma clone O91F4) are shown, each covering [the area / area]. Figure 46B superior.
[0119] Figure 46F The spatial gene expression (top) and protein staining (bottom) of the oligodendrocyte marker myelin CNPase (produced by hybridoma clone SMI91) are shown, each covering [the area / area]. Figure 46B superior.
[0120] Figure 46G The spatial gene expression (top) and protein staining (bottom) of myelin basic protein (produced by hybridoma clone P82H9), a marker of oligodendrocytes, are shown, each covering [the area of the image]. Figure 46B superior.
[0121] Figure 46H The image shows the spatial gene expression (top) and protein staining (bottom) of the stem cell marker SOX2 (produced by hybridoma clone 14A6A34), each covering... Figure 46B superior.
[0122] Figure 46I The image shows the spatial gene expression (top) and protein staining (bottom) of the neuronal marker SNAP-25 (produced by the hybridoma clone SMI81), each covering... Figure 46B superior.
[0123] Figure 47 This is an exemplary workflow for acquiring tissue samples and performing analyte capture, as described herein. Detailed Implementation
[0124] I. Introduction
[0125] This disclosure describes apparatus, systems, methods, and compositions for spatial analysis of biological samples. This section specifically describes certain general terms, analytes, sample types, and preparation steps that are referenced in the later parts of this disclosure.
[0126] (a) Spatial analysis.
[0127] Tissues and cells can be obtained from any source. For example, tissues and cells can be obtained from single-celled or multicellular organisms (e.g., mammals). Tissues and cells obtained from mammals (e.g., humans) often have varying levels of analytes (e.g., gene and / or protein expression) that can lead to differences in cell morphology and / or function. The location of cells or subsets of cells (e.g., neighboring and / or non-neighboring cells) within a tissue can influence, for example, cell fate, behavior, morphology, signal transduction, and crosstalk with other cells in the tissue. Information about differences in analyte levels (e.g., gene and / or protein expression) within different cells in mammalian tissues can also help physicians select or administer effective treatments and can allow researchers to identify and elucidate differences in cell morphology and / or cell function in single-celled or multicellular organisms (e.g., mammals) based on the detection of differences in analyte levels within different cells in a tissue. Differences in analyte levels within different cells in mammalian tissues can also provide information about how tissues (e.g., healthy and diseased tissues) function and / or develop. Differences in analyte levels within different cells in mammalian tissues can also provide information about different mechanisms of disease pathogenesis in tissues and the mechanisms of action of therapeutic treatments within tissues. Differences in analyte levels within different cells of mammalian tissues can also provide information about drug resistance mechanisms and their development in mammalian tissues. Differences in the presence or absence of analytes within different cells of multicellular organisms (e.g., mammals) can also provide information about drug resistance mechanisms and their development in multicellular organisms.
[0128] The spatial analysis method described in this paper is used to detect differences in analyte levels (e.g., gene and / or protein expression) within different cells in mammalian tissues or within single cells derived from mammals. For example, the spatial analysis method can be used to detect differences in analyte levels (e.g., gene and / or protein expression) within different cells in histological slide samples, from which data can be recombined to generate a three-dimensional atlas (e.g., with a degree of spatial resolution, such as single-cell resolution) of analyte levels (e.g., gene and / or protein expression) in tissue samples (e.g., tissue samples) obtained from mammals.
[0129] Spatial heterogeneity in developmental systems is typically investigated using RNA hybridization, immunohistochemistry, purification or induction of fluorescent reporter molecules or predefined subsets, and subsequent genome mapping (e.g., RNA-seq). However, this approach relies on a relatively small set of predefined markers, thus introducing selection bias that limits discovery. These existing methods also depend on prior knowledge. Spatial RNA analysis has traditionally relied on staining a limited number of RNA species. In contrast, single-cell RNA sequencing allows for in-depth analysis of cellular gene expression, including non-coding RNA, but established methods separate cells from their native spatial environment.
[0130] The spatial analysis methods described herein provide a wealth of analyte level and / or expression data for a wide variety of analytes within a sample at high spatial resolution, while preserving the natural spatial environment. These methods include, for example, the use of capture probes comprising a spatial barcode (e.g., a nucleic acid sequence) and a capture domain. The spatial barcode provides information about the location of the capture probe within a cell or tissue sample (e.g., a mammalian cell or mammalian tissue sample), and the capture domain is capable of binding analytes (e.g., proteins and / or nucleic acids) produced by and / or present in the cells. As described herein, the spatial barcode can be a nucleic acid having a unique sequence, a unique fluorophore, a unique combination of fluorophores, a unique amino acid sequence, a unique heavy metal, or a unique combination of heavy metals, or any other uniquely detectable reagent. The capture domain can be any reagent capable of binding analytes produced by and / or present in the cells (e.g., nucleic acids capable of hybridizing with nucleic acids from cells (e.g., mRNA, genomic DNA, mitochondrial DNA, or miRNA), a substrate comprising the analyte, a binding partner of the analyte, or an antibody that specifically binds the analyte). The capture probe may also comprise a nucleic acid sequence complementary to the sequence of a universal forward and / or universal reverse primer. Capture probes may also include cleavage sites (e.g., cleavage recognition sites of restriction endonucleases) or light-labile or thermosensitive bonds.
[0131] A variety of different methods can be used to detect the binding of analytes to capture probes, such as nucleic acid sequencing, fluorophore detection, nucleic acid amplification, nucleic acid ligation detection, and / or nucleic acid cleavage product detection. In some instances, this detection is used to associate a specific spatial barcode with a specific analyte produced by and / or present in cells (e.g., mammalian cells).
[0132] The capture probe may be attached to a surface, such as a solid array, beads, or coverslip. In some instances, the capture probe is not attached to a surface. In some instances, the capture probe is encapsulated within, embedded therein, or laminated thereon in a permeable composition (e.g., any substrate described herein). For example, the capture probe may be encapsulated or disposed within a permeable bead (e.g., a gel bead). In some instances, the capture probe is encapsulated within, embedded therein, or laminated thereon in the surface of a substrate (e.g., any exemplary substrate described herein, such as a hydrogel or porous membrane).
[0133] In some instances, cells or tissue samples comprising cells are brought into contact with a capture probe attached to a substrate (e.g., the surface of the substrate), and the cells or tissue sample is permeated to allow analytes to be released from the cells and bind to the capture probe attached to the substrate. In some instances, various methods (e.g., electrophoresis, chemical gradients, pressure gradients, fluid flow, or magnetic fields) can be used to actively guide analytes released from the cells to the capture probe attached to the substrate.
[0134] In other instances, various methods can be used to guide the capture probe to interact with cell or tissue samples, such as including lipid anchors in the capture probe, including reagents that specifically bind to or form covalent bonds with membrane proteins in the capture probe, fluid flow, pressure gradients, chemical gradients, or magnetic fields.
[0135] Non-restrictive aspects of spatial analysis methods are described in WO 2011 / 127099, WO 2014 / 210233, WO2014 / 210225, WO 2016 / 162309, WO 2018 / 091676, WO 2012 / 140224, WO 2014 / 060483, U.S. Patent No. 10,002,316, U.S. Patent No. 9,727,810, U.S. Patent Application Publication No. 2017 / 0016053, Rodriques et al., Science 363(6434):1463-1467, 2019; WO 2018 / 045186, Lee et al., Nat. Protoc. 10(3):442-458, 2015; WO 2016 / 007839, WO 2018 / 045181, WO 2014 / 163886, Trejo et al., PLoS ONE 14(2):e0212031, 2019, U.S. Patent Application Publication No. 2018 / 0245142, Chen et al., Science 348(6233):aaa6090, 2015, Gao et al., BMC Biology 15:50, 2017, WO 2017 / 144338, WO 2018 / 107054, WO 2017 / 222453, WO 2019 / 068880, WO 2011 / 094669, US Patent No. 7,709,198, US Patent No. 8,604,182, US Patent No. 8,951,726, US Patent No. 9,783,841, US Patent No. 10,041,949, WO2016 / 057552, WO 2017 / 147483, WO 2018 / 022809, WO 2016 / 166128, WO 2017 / 027367, WO2017 / 027368, WO 2018 / 136856, WO 2019 / 075091, US Patent No. 10,059,990, WO 2018 / 057999, WO 2015 / 161173, Gupta et al., *Nature Biotechnology* Biotechnol.This document references U.S. Patent Application No. 16 / 992,569, filed August 13, 2020, entitled "Systems and Methods for Using Spatial Distribution of Haplotypes to Determine a Biological Condition," and may be used herein in any combination. Other non-limiting aspects of spatial analysis methods are described herein.
[0136] (b) General Terminology
[0137] Throughout this disclosure, specific terms are used to explain various aspects of the described apparatus, systems, methods, and compositions. This section includes explanations of certain terms that appear in later sections of this disclosure. In the event of a clear conflict between the descriptions in this section and their usage in other sections of this disclosure, the definitions in this section shall prevail.
[0138] (i) Subjects
[0139] "Subject" is an animal, such as a mammal (e.g., a human or a non-human ape), or a bird (e.g., a bird), or other organisms such as a plant. Examples of subjects include, but are not limited to, mammals such as rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cattle, cats, dogs, and primates (e.g., human or non-human primates); plants such as Arabidopsis thaliana, maize, sorghum, oats, wheat, rice, rapeseed, or soybeans; algae such as Chlamydomonas reinhardtii; nematodes such as Caenorhabditis elegans; insects such as Drosophila melanogaster, mosquitoes, fruit flies, bees, or spiders; fish such as zebrafish; reptiles; amphibians such as frogs or Xenopus laevis; Dictyostelium discoideum; and fungi such as Pneumocystis carinii and Takifugu. rubripes), yeast, Saccharomyces cerevisiae or Schizosaccharomyces pombe; or Plasmodium falciparum.
[0140] (ii) Nucleic acids and nucleotides
[0141] The terms “nucleic acid” and “nucleotide” are intended to be consistent with their use in the art and include naturally occurring species or their functional analogs. Particularly useful nucleic acid functional analogs are capable of hybridizing with nucleic acids in a sequence-specific manner (e.g., capable of hybridizing with two nucleic acids such that ligation can occur between the two hybridized nucleic acids) or can be used as templates for the replication of specific nucleotide sequences. Naturally occurring nucleic acids typically have a backbone containing phosphodiester bonds. Analog structures may have alternating backbone linkages, including any kind of backbone linkage known in the art. Naturally occurring nucleic acids typically have deoxyribose (e.g., present in deoxyribonucleic acid (DNA)) or ribose (e.g., present in ribonucleic acid (RNA)).
[0142] Nucleic acids may comprise any of a variety of analogs having these sugar moieties known in the art. Nucleic acids may include natural or non-natural nucleotides. In this respect, natural deoxyribonucleic acid may have one or more bases selected from the group consisting of adenine (A), thymine (T), cytosine (C), or guanine (G), and ribonucleic acid may have one or more bases selected from the group consisting of uracil (U), adenine (A), cytosine (C), or guanine (G). Useful non-natural bases that may be included in nucleic acids or nucleotides are known in the art.
[0143] (iii) Probe and target
[0144] When used to refer to nucleic acids or nucleic acid sequences, "probe" or "target" is intended in the context of a method or composition to serve as a semantic identifier for the nucleic acid or sequence and does not limit the structure or function of the nucleic acid or sequence beyond what is explicitly stated.
[0145] (iv) Oligonucleotides and Polynucleotides
[0146] The terms "oligonucleotide" and "polynucleotide" are used interchangeably and refer to single-stranded polymers of nucleotides with a length of about 2 to about 500 nucleotides. Oligonucleotides can be synthetic, enzymatically prepared (e.g., by polymerization), or produced using a "split pooling" method. Oligonucleotides can include ribonucleotide monomers (e.g., oligoribonucleotides) and / or deoxyribonucleotide monomers (e.g., oligodeoxyribonucleotides). In some instances, oligonucleotides comprise combinations of deoxyribonucleotide and ribonucleotide monomers (e.g., random or ordered combinations of deoxyribonucleotide and ribonucleotide monomers). The length of an oligonucleotide can be, for example, 4 to 10, 10 to 20, 21 to 30, 31 to 40, 41 to 50, 51 to 60, 61 to 70, 71 to 80, 80 to 100, 100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, 350 to 400, or 400 to 500 nucleotides. An oligonucleotide may include one or more functional moieties attached to (e.g., covalently or non-covalently) a multimeric structure. For example, an oligonucleotide may include one or more detectable markers (e.g., radioisotopes or fluorophores).
[0147] (v) Barcode
[0148] A barcode is a label or identifier that conveys or is able to convey information (e.g., information about an analyte in a sample, bead, and / or capture probe). A barcode can be part of the analyte or independent of it. A barcode can be affixed to the analyte. A particular barcode can be unique relative to other barcodes.
[0149] Barcodes can have many different formats. For example, barcodes can include non-random, semi-random, and / or random nucleic acid and / or amino acid sequences, as well as synthetic nucleic acid and / or amino acid sequences.
[0150] Barcodes can have a variety of different formats. For example, barcodes can include polynucleotide barcodes, random nucleic acid and / or amino acid sequences, and synthetic nucleic acid and / or amino acid sequences. Barcodes can be attached to an analyte or another part or structure in a reversible or irreversible manner. Barcodes can be added to fragments of samples, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), before or during sample sequencing. Barcodes can allow for the identification and / or quantification of individual sequencing reads (e.g., a barcode can be or can include a unique molecular identifier or "UMI").
[0151] Barcodes can, for example, spatially resolve molecular components found in biological samples at single-cell resolution (e.g., a barcode can be or may include a "spatial barcode"). In some embodiments, a barcode includes a UMI and a spatial barcode. In some embodiments, a barcode includes two or more sub-barcodes that function together as a single barcode. In some embodiments, a barcode includes a UMI and a spatial barcode. In some embodiments, a barcode includes two or more sub-barcodes that function together as a single barcode (e.g., a polynucleotide barcode). For example, a polynucleotide barcode may include two or more polynucleotide sequences (e.g., sub-barcodes) separated by one or more non-barcode sequences.
[0152] (vi) Capture point
[0153] The term “capture point” (or, alternatively, “feature” or “capture probe cluster”) is used herein to describe an entity that serves as a support or library for various molecular entities used in sample analysis. Examples of capture points include, but are not limited to, beads, points of any two- or three-dimensional geometry (e.g., inkjet dots, masking dots, squares on a grid), pores, and hydrogel pads. In some embodiments, a capture point is a region on a substrate on which capture probes, marked with a spatial barcode, are aggregated. Specific, non-limiting embodiments of capture points and substrates are further described below in this disclosure.
[0154] Other definitions commonly associated with the spatial analysis of analytes can be found in U.S. Patent Application No. 16 / 992,569, filed August 13, 2020, entitled "Systems and Methods for Using the Spatial Distribution of Haplotypes to Determine a Biological Condition," which are incorporated herein by reference.
[0155] (vii) base
[0156] As used in this article, “substrate” is any surface on which the capture probe can be attached (e.g., chip, solid-state array, bead, cover glass, etc.).
[0157] (viii) Genome
[0158] A “genome” generally refers to genomic information from a subject, which may be at least part or all of the genetic information encoded by the subject’s genes. A genome may include coding regions (e.g., coding regions that encode proteins) and non-coding regions. A genome may include sequences of some or all of the subject’s chromosomes. For example, the human genome typically has a total of 46 chromosomes. Some or all of these sequences can constitute the genome.
[0159] (ix) Connectors, linkers, and tags
[0160] "Linker," "adapter," and "tag" are terms that may be used interchangeably in this disclosure and refer to a species that can be coupled to a polynucleotide sequence (in a process called "tag") using any of a number of different techniques, including (but not limited to) ligation, hybridization, and tagging. Linkers can also be functionally enhanced nucleic acid sequences, such as spacer sequences, primer sequences / sites, barcode sequences, or unique molecular identifier sequences.
[0161] (x) antibody
[0162] Antibodies are polypeptide molecules that recognize and bind to complementary target antigens. Antibodies typically have a Y-shaped molecular structure or are polymers thereof. Naturally occurring antibodies, called immunoglobulins, belong to the immunoglobulin class IgG, IgM, IgA, IgD, and IgE. Antibodies can also be synthesized. For example, recombinant antibodies, which are monoclonal antibodies, can be synthesized using a synthetic gene by recovering the antibody gene from a source cell, amplifying it into a suitable vector, and introducing the vector into a host to induce the host to express the recombinant antibody. Typically, recombinant antibodies can be cloned from any antibody-producing animal species using suitable oligonucleotide primers and / or hybridization probes. Recombinant technology can be used to generate antibodies and antibody fragments, including in non-endogenous species.
[0163] Synthetic antibodies can be derived from non-immunoglobulin sources. For example, antibodies can be generated from nucleic acids (e.g., aptamers) and non-immunoglobulin protein scaffolds (such as peptide aptamers), in which hypervariable loops are inserted into the scaffold to form antigen-binding sites. Synthetic antibodies based on nucleic acid or peptide structures can be smaller than immunoglobulin-derived antibodies, resulting in greater tissue penetration.
[0164] Antibodies may also include affinity proteins, which are affinity agents typically having a molecular weight of about 12 to 14 kDa. Affinity proteins typically bind to targets (e.g., target proteins) with high affinity and specificity. Examples of such targets include, but are not limited to, ubiquitin chains, immunoglobulins, and C-reactive proteins. In some embodiments, affinity proteins are derived from cysteine protease inhibitors and include a peptide ring and a variable N-terminal sequence that provides a binding site. Antibodies may also include single-domain antibodies (VHH domain and VNAR domain), scFv, and Fab fragments.
[0165] (c) Analytes
[0166] The apparatus, systems, methods, and compositions described in this disclosure can be used to detect and analyze a wide variety of analytes. For the purposes of this disclosure, "analyte" can include any biological substance, structure, part, or component to be analyzed. The term "target" can be used similarly to refer to the analyte of interest.
[0167] Analytes can be broadly classified into one of two groups: nucleic acid analytes and non-nucleic acid analytes. Examples of non-nucleic acid analytes include, but are not limited to, lipids, carbohydrates, peptides, proteins, glycoproteins (N-linked or O-linked), lipoproteins, phosphoproteins, specific phosphorylated or acetylated variants of proteins, amidated variants of proteins, hydroxylated variants of proteins, methylated variants of proteins, ubiquitinated variants of proteins, sulfated variants of proteins, viral capsid proteins, extracellular and intracellular proteins, antibodies, and antigen-binding fragments. In some embodiments, the analyte is an organelle (e.g., the cell nucleus or mitochondria).
[0168] Cell surface features corresponding to the analyte may include, but are not limited to, receptors, antigens, surface proteins, transmembrane proteins, differentiation protein clusters, protein channels, protein pumps, carrier proteins, phospholipids, glycoproteins, glycolipids, cell-cell interaction protein complexes, antigen-presenting complexes, major histocompatibility complexes, engineered T cell receptors, T cell receptors, B cell receptors, chimeric antigen receptors, extracellular matrix proteins, post-translational modifications (e.g., phosphorylation, glycosylation, ubiquitination, nitrosation, methylation, acetylation, or lipidation) of cell surface proteins, gap junctions, and adhesion junctions.
[0169] Analytes can originate from specific cell types and / or specific subcellular regions. For example, analytes can originate from the cytosol, nucleus, mitochondria, microsomes, and more generally, from any other compartment, organelle, or part of the cell. Permeabilizers that specifically target certain cellular compartments and organelles can be used to selectively release analytes from cells for analysis. Tissue permeabilization, such as… Figure 37 As shown.
[0170] Examples of nucleic acid analytes include DNA analytes such as genomic DNA, methylated DNA, specifically methylated DNA sequences, fragmented DNA, mitochondrial DNA, in situ synthesized PCR products, and RNA / DNA hybrids.
[0171] Examples of nucleic acid analytes also include RNA analytes, such as various types of coding and non-coding RNA. Examples of different types of RNA analytes include messenger RNA (mRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), and viral RNA. RNA can be a transcript (e.g., present in tissue sections). RNA can be small (e.g., less than 200 nucleic acid bases in length) or large (e.g., RNA longer than 200 nucleic acid bases in length). Small RNAs mainly include 5.8S ribosomal RNA (rRNA), 5S rRNA, transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), tRNA-derived small RNA (tsRNA), and rDNA-derived small RNA (srRNA). RNA can be double-stranded or single-stranded RNA. RNA can be circular RNA. RNA can be bacterial rRNA (e.g., 16S rRNA or 23S rRNA).
[0172] Other examples of analytes include mRNA and cell surface features (e.g., using the labelers described herein), mRNA and intracellular proteins (e.g., transcription factors), mRNA and cellular methylation status, mRNA and accessible chromatin (e.g., ATAC-seq, DNase-seq, and / or MNase-seq), mRNA and metabolites (e.g., using the labelers described herein), barcode labelers (e.g., oligonucleotide-labeled antibodies described herein) and V(D)J sequences of immune cell receptors (e.g., T cell receptors), mRNA and interfering agents (e.g., CRISPR crRNA / sgRNA, TALEN, zinc finger nucleases, and / or antisense oligonucleotides as described herein). In some embodiments, the interfering agent is a small molecule, antibody, drug, aptamer, miRNA, physical environment (e.g., temperature change), or any other known interfering agent.
[0173] The analyte may include a nucleic acid molecule having at least a portion of a V(D)J sequence encoding an immune cell receptor (e.g., TCR or BCR). In some embodiments, the nucleic acid molecule is cDNA first generated from the reverse transcription of the corresponding mRNA using a poly(T)-containing primer. The generated cDNA can then be barcoded using a capture probe characterized by a barcoded sequence (and optionally, a UMI sequence) that hybridizes to at least a portion of the generated cDNA. In some embodiments, a template-converting oligonucleotide hybridizes to a poly(C) tail added to the 3' end of the cDNA via reverse transcriptase. The original mRNA template and template-converting oligonucleotide can then be denatured from the cDNA, and the barcoded capture probe can then hybridize to the cDNA and a complement of the generated cDNA. Other methods and compositions suitable for barcoding cDNA generated from mRNA transcripts (including those encoding the V(D)J region of immune cell receptors) and / or barcoding methods and compositions (including template-converting oligonucleotides) are described in PCT patent application PCT / US2017 / 057269, filed October 18, 2017, and U.S. patent application No. 15 / 825,740, filed November 29, 2017, both of which are incorporated herein by reference in their entirety. V(D)J analysis can also be performed using one or more markers that bind to specific surface features of immune cells and associate with barcoded sequences. One or more markers may include MHC or MHC multimers.
[0174] As described above, the analyte may include nucleic acids that can be used as components of a gene editing reaction, such as gene editing using short palindromic repeats (CRISPR) with regular intervals based on clustering. Therefore, the capture probe may include a nucleic acid sequence complementary to the analyte (e.g., a sequence that can hybridize to CRISPR RNA (crRNA), a single guide RNA (sgRNA), or an adaptor sequence engineered into crRNA or sgRNA).
[0175] In some embodiments, analytes are extracted from live cells. Processing conditions can be adjusted to ensure the biological sample remains viable during analysis, and the analyte is extracted (or released) from the live cells of the sample. Live cell-derived analytes can be obtained only once from the sample, or they can be obtained periodically from a sample that continues to remain viable.
[0176] Typically, systems, apparatus, methods, and compositions can be used to analyze any number of analytes. For example, the number of analytes being analyzed can be at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 25, at least about 30, at least about 40, at least about 50, at least about 100, at least about 1,000, at least about 10,000, at least about 100,000, or more different analytes present in a region of the sample or within a single capture point of the substrate. Methods for performing multiple determinations to analyze two or more different analytes will be discussed in subsequent sections of this disclosure.
[0177] (d) Biological samples
[0178] (i) Types of biological samples
[0179] "Biosamples" are obtained from subjects for analysis using any of a variety of techniques, including but not limited to biopsy, surgery, and laser capture microscopy (LCM), and typically include cells and / or other biological material from the subject. In addition to subjects, biosamples can also be obtained from non-mammalian organisms (e.g., plants, insects, ticks, nematodes, pufferfish, amphibians, and fish). Biosamples can be obtained from prokaryotes, such as bacteria like *Escherichia coli*, *Staphylococci*, or *Mycoplasma pneumoniae*; archaea; viruses such as hepatitis C virus or human immunodeficiency virus; or viroids. Biosamples can also be obtained from eukaryotes, such as patient-derived organoids (PDOs) or patient-derived xenografts (PDXs). Biosamples can include organoids, which are miniaturized and simplified forms of organs generated in vitro in three dimensions, displaying realistic microscopic anatomy. Organoids can be generated from one or more cells derived from tissues, embryonic stem cells, and / or induced pluripotent stem cells, and due to their self-renewal and differentiation capabilities, they can self-organize in three-dimensional culture. In some embodiments, organoids are brain organoids, intestinal organoids, gastric organoids, tongue organoids, thyroid organoids, thymus organoids, testicular organoids, liver organoids, pancreatic organoids, epithelial organoids, lung organoids, kidney organoids, gastrula, heart organoids, or retinal organoids. Subjects from whom biological samples can be obtained can be healthy or asymptomatic individuals, individuals with or suspected of having a disease (e.g., cancer) or susceptible to a disease, and / or individuals requiring or suspected of requiring treatment.
[0180] Biological samples can include any number of macromolecules, such as cellular macromolecules and organelles (e.g., mitochondria and the nucleus). Biological samples can be nucleic acid samples and / or protein samples. Biological samples can be carbohydrate samples or lipid samples. Biological samples can be obtained as tissue samples, such as tissue sections, biopsies, core biopsies, needle aspiration, or fine needle aspiration. Samples can be fluid samples, such as blood samples, urine samples, or saliva samples. Samples can be skin samples, colon samples, buccal swabs, histological samples, histopathological samples, plasma or serum samples, tumor samples, live cells, cultured cells, clinical samples (e.g., whole blood or blood-derived products), blood cells, or cultured tissues or cells, including cell suspensions.
[0181] Cell-free biological samples can include extracellular polynucleotides. Extracellular polynucleotides can be isolated from body samples such as blood, plasma, serum, urine, saliva, mucosal secretions, sputum, feces, and tears.
[0182] Biological samples may be derived from homogeneous cultures or populations of the subjects or organisms mentioned herein, or alternatively from collections of several different organisms, such as in a community or ecosystem.
[0183] Biological samples may include one or more diseased cells. Diseased cells may have altered metabolic properties, gene expression, protein expression, and / or morphological characteristics. Examples of diseases include inflammatory diseases, metabolic diseases, neurological diseases, and cancer. Cancer cells may be derived from solid tumors, hematologic malignancies, cell lines, or obtained as circulating tumor cells.
[0184] Biological samples can also include fetal cells. For example, procedures such as amniocentesis can be performed to obtain fetal cell samples from the maternal circulation. Sequencing of fetal cells can be used to identify any of a variety of genetic disorders, including, for example, aneuploidy, such as Down syndrome, Edwards syndrome, and Patau syndrome. Furthermore, cell surface characteristics of fetal cells can be used to identify any of a variety of conditions or diseases.
[0185] Biological samples can also include immune cells. Sequence analysis of the entire immune system of these cells (including the genome, proteome, and cell surface features) can provide a wealth of information that helps in understanding the state and function of the immune system. For example, determining the status of minimal residual disease (MRD) (e.g., negative or positive) in patients with multiple myeloma (MM) after autologous stem cell transplantation is considered a predictor of MRD in MM patients (see, for example, U.S. Patent Publication No. 2018 / 0156784, the entire contents of which are incorporated herein by reference).
[0186] Examples of immune cells in biological samples include, but are not limited to, B cells, T cells (e.g., cytotoxic T cells, natural killer T cells, regulatory T cells, and T helper cells), natural killer cells, cytokine-induced killer (CIK) cells, bone marrow cells (such as granulocytes (basophils, eosinophils, neutrophils / multisegmented neutrophils), monocytes / macrophages, mast cells, platelets / megakaryocytes, and dendritic cells).
[0187] As described above, a biological sample may include a single analyte of interest, or more than one analyte of interest. Methods for performing multiple assays to analyze two or more different analytes in a single biological sample will be discussed in later sections of this disclosure.
[0188] (ii) Preparation of biological samples
[0189] Various steps can be performed to prepare biological samples for analysis. Unless otherwise stated, the preparation steps described below can generally be combined in any way to appropriately prepare a specific sample for analysis.
[0190] (1) Tissue section
[0191] Biological samples can be harvested from a subject (e.g., via surgical biopsy, whole subject section, in vitro growth as a cell population on a growth substrate or culture dish, or prepared for analysis as a tissue section or multiple tissue sections). The grown samples can be thin enough for analysis without further processing steps. Alternatively, grown samples and samples obtained via biopsy or sectioning can be prepared into thin tissue sections using mechanical cutting equipment such as a vibratory microtome. As another alternative, in some embodiments, thin tissue sections can be prepared by applying a tactile imprint of the biological sample onto a suitable substrate material.
[0192] The thickness of a tissue section can be a fraction of the largest cross-sectional size of the cells (e.g., less than 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, or 0.1). However, tissue sections thicker than the largest cross-sectional cell size can also be used. For example, frozen sections can be used, which can be, for example, 10 to 20 micrometers thick.
[0193] More generally, the thickness of tissue sections typically depends on the method used to prepare the sections and the physical properties of the tissue, thus sections with a wide variety of thicknesses can be prepared and used. For example, tissue section thicknesses can be at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.7, 1.0, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, or 50 micrometers. Thicker sections, such as at least 70, 80, 90, or 100 micrometers or more, can also be used if desired or convenient. Typically, tissue section thicknesses are between 1 and 100 micrometers, 1 and 50 micrometers, 1 and 30 micrometers, 1 and 25 micrometers, 1 and 20 micrometers, 1 and 15 micrometers, 1 and 10 micrometers, 2 and 8 micrometers, 3 and 7 micrometers, or 4 and 6 micrometers, but as mentioned above, sections with thicknesses greater than or less than these ranges can also be analyzed.
[0194] Multiple sections can also be obtained from a single biological sample. For example, multiple tissue sections can be obtained from a surgical biopsy sample by sequentially slicing it with a slicing blade. Spatial information within the sequential sections can be preserved in this way, and the sections can be analyzed sequentially to obtain three-dimensional information about the biological sample.
[0195] (2) Freezing
[0196] In some embodiments, biological samples (e.g., tissue sections as described above) can be prepared by deep freezing at temperatures suitable for maintaining or preserving the integrity (e.g., physical properties) of the tissue structure. Such temperatures can be, for example, below -20°C, or below -25°C, -30°C, -40°C, -50°C, -60°C, -70°C, -80°C, -90°C, -100°C, -110°C, -120°C, -130°C, -140°C, -150°C, -160°C, -170°C, -180°C, -190°C, or -200°C. Many suitable methods can be used to slice (e.g., cut into thin slices) the frozen tissue sample onto a substrate surface. For example, tissue samples can be prepared using a cryostat (e.g., a cryostat) set at temperatures suitable for maintaining the structural integrity of the tissue sample and the chemical properties of the nucleic acids in the sample. Such temperatures can be, for example, below -15°C, below -20°C, or below -25°C. Samples can be rapidly frozen in isopentane and liquid nitrogen. Frozen samples can be stored in sealed containers before embedding.
[0197] (3) Formalin fixation and paraffin embedding
[0198] In some embodiments, formalin fixation and paraffin embedding (FFPE) can be used to prepare biological samples, which is an established method. In some embodiments, formalin fixation and paraffin embedding can be used to prepare cell suspensions and other non-tissue samples. After fixing and embedding the sample in paraffin or resin blocks, the sample can be sectioned as described above. Before analysis, the paraffin embedding material can be removed from the tissue sections (e.g., dewaxing) by incubating the tissue sections in a suitable solvent (e.g., xylene) followed by rinsing (e.g., 99.5% ethanol for 2 minutes, 96% ethanol for 2 minutes, and 70% ethanol for 2 minutes).
[0199] (4) Fixed
[0200] As an alternative to formalin fixation, biological samples can be fixed in any of a variety of other fixatives to preserve the biological structure of the sample prior to analysis. For example, samples can be fixed by immersion in ethanol, methanol, acetone, formaldehyde (e.g., 2% formaldehyde), paraformaldehyde-Triton, glutaraldehyde, or combinations thereof.
[0201] In some embodiments, acetone fixation is used in conjunction with freshly frozen samples, which may include, but are not limited to, cortical tissue, mouse olfactory bulbs, human brain tumors, human necropsy brains, and breast cancer samples. In some embodiments, a compatible fixation method is selected and / or optimized based on the desired workflow. For example, formaldehyde fixation may be chosen to be compatible with workflows using IHC / IF protocols for protein visualization. As another example, methanol fixation may be chosen for workflows that emphasize RNA / DNA library quality. In some applications, acetone fixation may be chosen to permeabilize the tissue. When performing acetone fixation, a pre-permeabilization step (described below) may be omitted. Alternatively, acetone fixation may be performed in conjunction with a permeabilization step.
[0202] (5) Embedding
[0203] As an alternative to paraffin embedding, biological samples can be embedded in any of a variety of other embedding materials to provide a substrate for the sample prior to sectioning and other processing steps. Typically, the embedding material is removed before analyzing tissue sections obtained from the sample. Suitable embedding materials include, but are not limited to, wax, resins (e.g., methacrylates), epoxy resins, and agar.
[0204] (6) Staining
[0205] To facilitate visualization, a variety of staining agents and techniques can be used to stain biological samples. In some embodiments, for example, any number of biological staining agents can be used to stain samples, including but not limited to acridine orange, Bismarck brown, carmine, Coomassie blue, cresol purple, DAPI, eosin, ethidium bromide, acid fuchsin, hematoxylin, Hurst stain, iodine, methyl green, methylene blue, neutral red, Nile blue, Nile red, osmium tetroxide, propidium iodide, rhodamine, or safranin.
[0206] Samples can be stained using known staining techniques, including Can-Grunwald, Giemsa, Hematoxylin and Eosin (H&E), Jenner's, Leishman, Masson's trichrome, Papanicolaou, Romanowsky, silver, Sudan, Wright's, and / or periodic acid Schiff (PAS) staining. PAS staining is typically performed after fixation in formalin or acetone.
[0207] In some embodiments, samples are stained using detectable markers (e.g., radioisotopes, fluorophores, chemiluminescent compounds, bioluminescent compounds, and dyes) as described elsewhere herein. In some embodiments, only one type of staining agent or one technique is used to stain biological samples. In some embodiments, staining includes biological staining techniques such as H&E staining. In some embodiments, staining includes identifying analytes using fluorescently conjugated antibodies. In some embodiments, biological samples are stained using two or more different types of staining agents or two or more different staining techniques. For example, a biological sample can be prepared by staining and imaging using one technique (e.g., H&E staining and bright-field imaging) and then staining and imaging the same biological sample using another technique (e.g., IHC / IF staining and fluorescence microscopy).
[0208] In some embodiments, biological samples may be destaining. Methods for destaining or fading biological samples are known in the art and generally depend on the nature of the staining agent applied to the sample. For example, H&E staining can be destaining by washing the sample in HCl or any other low-pH acid (e.g., selenic acid, sulfuric acid, hydroiodic acid, benzoic acid, carbonic acid, malic acid, phosphoric acid, oxalic acid, succinic acid, salicylic acid, tartaric acid, sulfurous acid, trichloroacetic acid, hydrobromic acid, hydrochloric acid, nitric acid, orthophosphoric acid, arsenic acid, selenite, chromic acid, citric acid, hydrofluoric acid, nitrous acid, isocyanate, formic acid, hydrogen selenide, molybdate, lactic acid, acetic acid, carbonic acid, hydrogen sulfide, or combinations thereof). In some embodiments, destaining may include washing the sample 1, 2, 3, 4, 5 or more times in a low-pH acid (e.g., HCl). In some embodiments, destaining may include adding HCl to a downstream solution (e.g., a permeation solution). In some embodiments, destaining may include dissolving an enzyme (e.g., pepsin) used in the disclosed methods in a low-pH acid (e.g., HCl) solution. In some embodiments, after decolorizing hematoxylin with a low-pH acid, other reagents may be added to the decolorizing solution to raise the pH for other applications. For example, SDS may be added to a low-pH acid decolorizing solution to raise the pH compared to a low-pH acid decolorizing solution alone. As another example, in some embodiments, one or more immunofluorescent staining agents are applied to a sample via antibody conjugation. These staining agents can be removed using techniques such as washing with reducing agents and detergents, treatment with dissociative salts, treatment with antigen retrieval solutions, and treatment with acidic glycine buffer to cleave disulfide bonds. Methods for multiple staining and destaining are described, for example, in Bolognesi et al., 2017, J. Histochem. Cytochem. 65(8):431-444; Lin et al., 2015, Nat Commun. 6:8390; Pirici et al., 2009, J. Histochem. Cytochem. 57:567-75; and Glass et al., 2009, J. Histochem. Cytochem. 57:899-905, the full contents of which are incorporated herein by reference.
[0209] (7) Hydrogel embedding
[0210] In some embodiments, hydrogel formation occurs within a biological sample. In some embodiments, the biological sample (e.g., a tissue section) is embedded in the hydrogel. In some embodiments, hydrogel subunits are infused into the biological sample, and polymerization of the hydrogel is initiated by external or internal stimuli. The term "hydrogel" as used herein can include a cross-linked 3D network of hydrophilic polymer chains. "Hydrogel subunits" can be hydrophilic monomers, molecular precursors, or polymers that can be polymerized (e.g., cross-linked) to form a three-dimensional (3D) hydrogel network.
[0211] Hydrogels can swell in the presence of water. In some embodiments, the hydrogel comprises natural materials. In some embodiments, the hydrogel comprises synthetic materials. In some embodiments, the hydrogel comprises a mixture of materials, for example, the hydrogel material comprises components of synthetic and natural polymers. Any of the materials described herein for hydrogels or hydrogels comprising peptide-based materials can be used. Embedding a sample in this manner typically involves contacting the biological sample with the hydrogel such that the biological sample is surrounded by the hydrogel. For example, the sample can be embedded by contacting the sample with a suitable polymeric material and activating the polymeric material to form a hydrogel. In some embodiments, the formation of the hydrogel results in the internalization of the hydrogel within the biological sample.
[0212] In some embodiments, biological samples are immobilized in a hydrogel by crosslinking the polymeric material forming the hydrogel. Crosslinking can be performed chemically and / or photochemically, or alternatively by any other hydrogel-forming method known in the art. For example, biological samples can be immobilized in a hydrogel by polyacrylamide crosslinking. Furthermore, analytes from the biological sample can be immobilized in the hydrogel by crosslinking (e.g., polyacrylamide crosslinking).
[0213] The composition and application of hydrogel matrices in biological samples typically depend on the nature and preparation of the biological sample (e.g., section, unsection, fresh-frozen, fixation type). The hydrogel can be any suitable hydrogel in which the biological sample is anchored or embedded when formed on the biological sample. Non-limiting examples of hydrogels are described herein or are known in the art. As an example, when the biological sample is a tissue section, the hydrogel may comprise a monomer solution and an ammonium persulfate (APS) initiator / tetramethylethylenediamine (TEMED) promoter solution. As another example, when the biological sample consists of cells (e.g., cultured cells or cells isolated from a tissue sample), the cells may be incubated with the monomer solution and the APS / TEMED solution. For cells, the hydrogel is formed in compartments, including but not limited to devices for culturing, maintaining, or transporting cells. For example, the hydrogel can be formed by adding a monomer solution plus APS / TEMED to the compartments to a depth ranging from about 0.1 μm to about 2 mm.
[0214] Other methods and aspects of hydrogel embedding of biological samples are described, for example, in Chen et al., 2015, Science 347(6221):543–548 and PCT Publication 202020176788A1 entitled “Profiling of biological analytes with spatially barcoded oligonucleotide arrays”, the entire contents of which are incorporated herein by reference.
[0215] (8) Transfer of biological samples
[0216] In some embodiments, a hydrogel is used to transfer a biological sample immobilized on a substrate (e.g., a biological sample prepared using methanol fixation or formalin fixation and paraffin embedding (FFPE)) to a spatial array. In some embodiments, a hydrogel is formed on top of the biological sample on a substrate (e.g., a glass slide). For example, hydrogel formation may occur in a manner sufficient to anchor (e.g., embed) the biological sample to the hydrogel. After hydrogel formation, the biological sample is anchored (e.g., embedded) in the hydrogel, wherein separating the hydrogel from the substrate results in the biological sample separating from the substrate along with the hydrogel. The biological sample can then be contacted with the spatial array, thereby allowing spatial analysis of the biological sample. In some embodiments, the hydrogel is removed after contacting the biological sample with the spatial array. For example, the methods described herein may include event-dependent (e.g., light- or chemical) depolymerization of the hydrogel, wherein the hydrogel depolymerizes upon application of an event (e.g., an external stimulus). In one instance, the biological sample may be anchored to a DTT-sensitive hydrogel, wherein the addition of DTT may cause the hydrogel to depolymerize and release the anchored biological sample. The hydrogel can be any suitable hydrogel in which the biological sample is anchored or embedded when the hydrogel is formed on the biological sample. Non-limiting examples of hydrogels are described herein or are known in the art. In some embodiments, the hydrogel includes a connector that allows the biological sample to be anchored to the hydrogel. In some embodiments, the hydrogel includes a connector that allows the bioanalyte to be anchored to the hydrogel. In this case, the connector may be added to the hydrogel before, during, or after hydrogel formation. Non-limiting examples of connectors for anchoring nucleic acids to hydrogels may include 6-((acryloyl)amino)hexanoic acid (acryloyl-X SE) (available from Thermo Fisher Scientific, Waltham, MA), Label-ITamine (available from Mirus Bio, Madison, Wisconsin), and Label X (Chen et al., Nature Methods 13:679-684, 2016). Any kind of feature can determine the transfer conditions required for a given biological sample. Non-limiting examples of features that may affect transfer conditions include the sample (e.g., thickness, fixation, and crosslinking) and / or the analyte of interest (different conditions for preserving and / or transferring different analytes (e.g., DNA, RNA, and proteins)). In some embodiments, hydrogel formation can occur in a manner sufficient to anchor (e.g., embed) the analyte in the biological sample into the hydrogel. In some embodiments, the hydrogel can implode (e.g., shrink) together with the anchored analyte present in the biological sample (e.g., embedded in the hydrogel). In some embodiments, the hydrogel can swell (e.g., expand at equal volumes) with the anchored analyte present in the biological sample (e.g., embedded in the hydrogel).In some embodiments, the hydrogel may be imploded (e.g., contracted) and subsequently expanded with an anchoring analyte present in the biological sample (e.g., embedded in the hydrogel).
[0217] (9) Isovolume expansion
[0218] In some embodiments, the biological sample embedded in the hydrogel may be isovolumetrically expanded. Isovolumetric expansion methods that can be used include hydration, a preparative step in expansion microscopy, as described in Chen et al., 2015, *Science* 347(6221)543-548; Asano et al., 2018, *Current Protocols* 80:1, doi:10.1002 / cpcb.56; Gao et al., 2017, *BMC Biology* 15:50, doi:10.1186 / s12915-017-0393-3; and Wassie et al., 2018, "Expansion Microscopy: Principles and Applications in Biological Research," *Nature Methods* 16(1):33-41, each of which is incorporated herein by reference in its entirety.
[0219] Typically, the steps used to perform isovolitional expansion of biological samples can depend on the characteristics of the sample (e.g., the thickness of the tissue section, fixation, cross-linking) and / or the analyte of interest (e.g., different conditions for anchoring RNA, DNA, and proteins to the gel).
[0220] Isovolume expansion can be achieved by anchoring one or more components of a biological sample to a gel, followed by gel formation, protein hydrolysis, and swelling. Isovolume expansion of the biological sample can occur before or after the biological sample is fixed to the substrate. In some embodiments, the isovolume-expanded biological sample can be removed from the substrate before contacting it with a spatial barcode array (e.g., a spatial barcode capture probe on the substrate).
[0221] In some embodiments, proteins in a biological sample are anchored to a swellable gel, such as a polyelectrolyte gel. Antibodies may be directed to the proteins before, after, or simultaneously with anchoring to the swellable gel. DNA and / or RNA in a biological sample may also be anchored to the swellable gel via suitable adapters. Examples of such adapters include, but are not limited to, 6-((acryloyl)amino)hexanoic acid (acryloyl-X SE) (available from Thermo Fisher Scientific, Waltham, MA), Label-ITamine (available from MirusBio, Madison, Wisconsin), and Label X (described, for example, in Chen et al., Nature Methods 13:679-684, 2016, the entire contents of which are incorporated herein by reference).
[0222] Isovolous expansion of a sample can increase the spatial resolution of subsequent analyses. For example, isovolous expansion of biological samples may lead to increased resolution in spatial analyses (e.g., single-cell analyses). This increased resolution can be determined by comparing isovolous expanded samples with non-isovolous expanded samples.
[0223] Isovolous expansion enables three-dimensional spatial resolution for subsequent analysis of samples. In some embodiments, isovolous expansion of biological samples can occur in the presence of spatial analytical reagents (e.g., analyte traps or trap probes). For example, a swellable gel may include an analyte trap or trap probe anchored to the swellable gel via a suitable connector. In some embodiments, spatial analytical reagents can be delivered to specific locations within the isovolous expanded biological sample.
[0224] In some embodiments, the biological sample is expanded by an equal volume to at least 2x, 2.1x, 2.2x, 2.3x, 2.4x, 2.5x, 2.6x, 2.7x, 2.8x, 2.9x, 3x, 3.1x, 3.2x, 3.3x, 3.4x, 3.5x, 3.6x, 3.7x, 3.8x, 3.9x, 4x, 4.1x, 4.2x, 4.3x, 4.4x, 4.5x, 4.6x, 4.7x, 4.8x, or 4.9x of its unexpanded volume. In some embodiments, the sample is expanded by an equal volume to at least 2x and less than 20x of its unexpanded volume.
[0225] In some embodiments, the biological sample embedded in the hydrogel is expanded by an equal volume to at least 2x, 2.1x, 2.2x, 2.3x, 2.4x, 2.5x, 2.6x, 2.7x, 2.8x, 2.9x, 3x, 3.1x, 3.2x, 3.3x, 3.4x, 3.5x, 3.6x, 3.7x, 3.8x, 3.9x, 4x, 4.1x, 4.2x, 4.3x, 4.4x, 4.5x, 4.6x, 4.7x, 4.8x, or 4.9x of its unexpanded volume. In some embodiments, the biological sample embedded in the hydrogel is expanded by an equal volume to at least 2x and less than 20x of its unexpanded volume.
[0226] (10) Substrate adhesion
[0227] In some embodiments, biological samples may be attached to a substrate (e.g., a microarray). Examples of substrates suitable for this purpose are described in detail below. Attachment of biological samples can be irreversible or reversible, depending on the nature of the sample and subsequent steps in the analytical method.
[0228] In some embodiments, the sample can be reversibly attached to the substrate by applying a suitable polymer coating to it and bringing the sample into contact with the polymer coating. The sample can then be separated from the substrate using an organic solvent that at least partially dissolves the polymer coating. Hydrogels are examples of polymers suitable for this purpose.
[0229] More generally, in some embodiments, the substrate may be coated or functionalized with one or more substances to promote sample adhesion to the substrate. Suitable substances that can be used to coat or functionalize the substrate include, but are not limited to, lectins, polylysine, antibodies, and polysaccharides.
[0230] (11) Cells did not aggregate
[0231] In some embodiments, the biological sample corresponds to cells (e.g., derived from cell cultures or tissue samples). In a cell sample having multiple cells, individual cells may be naturally non-aggregated. For example, cells may be derived from cell suspensions and / or dissociated or deaggregated cells from tissues or tissue sections.
[0232] Alternatively, cells in the sample can aggregate and can be deaggregated into individual cells using, for example, enzymatic or mechanical techniques. Examples of enzymes used for enzymatic deaggregation include, but are not limited to, dispersases, collagenases, trypsin, or combinations thereof. Mechanical deaggregation can be performed, for example, using a tissue homogenizer.
[0233] In some embodiments involving unaggregated or deaggregated cells, the cells are distributed on a substrate such that at least one cell occupies a distinct spatial feature on the substrate. The cells may be immobilized on the substrate (e.g., to prevent lateral cell diffusion). In some embodiments, a cell immobilizer may be used to immobilize unaggregated or deaggregated samples onto a spatial barcode array prior to analyte capture. "Cell immobilizer" may refer to an antibody attached to the substrate that can bind to a cell surface marker. In some embodiments, the distribution of multiple cells on the substrate follows Poisson statistics.
[0234] In some embodiments, cells from a plurality of cells are fixed to a substrate. In some embodiments, the cells are fixed to prevent lateral diffusion, for example by adding a hydrogel and / or by applying an electric field.
[0235] (12) Suspended and adherent cells
[0236] In some embodiments, the biological sample may be derived from an in vitro grown cell culture. The sample derived from the cell culture may include one or more suspension cells that are independent of adherence within the cell culture. Examples of such cells include, but are not limited to, cell lines derived from hematopoietic cells and cell lines derived from: Colo205, CCRF-CEM, HL-60, K562, MOLT-4, RPMI-8226, SR, HOP-92, NCI-H322M, and MALME-3M.
[0237] Samples derived from cell cultures may include one or more adhesive cells grown on the surface of a container containing culture medium. Other non-limiting examples of suspension cells and adhesive cells can be found in U.S. Patent Application No. 16 / 992,569, entitled "Systems and Methods for Using the Spatial Distributions on Haplotypes to Determine a Biological Condition," filed August 13, 2020, and PCT Publication No. 202020176788A1, entitled "Analysis of Bioanalytes Using Spatial Barcode Oligonucleotide Arrays," the entire contents of which are incorporated herein by reference.
[0238] In some embodiments, biological samples may be permeabilized to facilitate the transfer of analytes from the sample and / or to facilitate the transfer of species (such as capture probes) into the sample. If the sample is not sufficiently permeabilized, the amount of analyte captured from the sample may be too low for adequate analysis. Conversely, if the tissue sample is too permeabilized, the relative spatial relationships of analytes within the tissue sample may be lost. Therefore, a balance is desired between permeating the tissue sample sufficiently to obtain good signal intensity and maintaining the spatial resolution of the analyte distribution within the sample.
[0239] Typically, biological samples can be permeated by exposing them to one or more permeabilizing agents. Suitable reagents for this purpose include, but are not limited to, organic solvents (e.g., acetone, ethanol, and methanol), crosslinking agents (e.g., paraformaldehyde), and detergents (e.g., saponins, Triton X-100). TM Tween-20 TMOr sodium dodecyl sulfate (SDS)) and enzymes (e.g., trypsin, protease (e.g., proteinase K). In some embodiments, the detergent is an anionic detergent (e.g., SDS or N-lauroyl sarcosinate sodium salt solution). In some embodiments, the biological sample may be permeated with any method described herein (e.g., using any detergent described herein, such as SDS and / or N-lauroyl sarcosinate sodium salt solution) before or after enzyme treatment (e.g., treatment with any enzyme described herein, such as trypsin, protease (e.g., pepsin and / or proteinase K)).
[0240] In some embodiments, biological samples can be permeated by exposing the sample to greater than about 1.0 w / v% (e.g., greater than about 2.0 w / v%, greater than about 3.0 w / v%, greater than about 4.0 w / v%, greater than about 5.0 w / v%, greater than about 6.0 w / v%, greater than about 7.0 w / v%, greater than about 8.0 w / v%, greater than about 9.0 w / v%, greater than about 10.0 w / v%, greater than about 11.0 w / v%, greater than about 12.0 w / v%, or greater than about 13.0 w / v%) of sodium dodecyl sulfate (SDS) and / or N-lauroyl sarcosine or sodium N-lauroyl sarcosine. In some embodiments, biological samples can be exposed (e.g., about 5 minutes to about 1 hour, about 5 minutes to about 40 minutes, about 5 minutes to about 30 minutes, about 5 minutes to about 20 minutes, or about 5 minutes to about 10 minutes) at about 1.0 w / v% to about 14.0 w / v% (e.g., about 2.0 w / v% to about 14.0 w / v%, about 2.0 w / v% to about 12.0 w / v%, about 2.0 w / v% to about 10.0 w / v%, about 4.0 w / v% to about 14.0 w / v%, about 4.0 w / v% to about 12.0 w / v%, about 4.0 w / v% to about 10.0 w / v%, about 6.0 w / v% to about 14.0 w / v%, about 6.0 w / v% to about 12.0 w / v%, about 6.0 w / v% to about 10.0 w / v%, about 8.0 w / v%). SDS and / or N-lauroyl sarcosinate solution and / or proteinase K (v% to about 14.0 w / v%, about 8.0 w / v% to about 12.0 w / v%, about 8.0 w / v% to about 10.0 w / v%, about 10.0% w / v% to about 14.0 w / v%, about 10.0 w / v% to about 12.0 w / v%, or about 12.0 w / v% to about 14.0 w / v%) For example, it is permeated at temperatures of about 4°C to about 35°C, about 4°C to about 25°C, about 4°C to about 20°C, about 4°C to about 10°C, about 10°C to about 25°C, about 10°C to about 20°C, about 10°C to about 15°C, about 35°C to about 50°C, about 35°C to about 45°C, about 35°C to about 40°C, about 40°C to about 50°C, about 40°C to about 45°C, or about 45°C to about 50°C.
[0241] In some embodiments, biological samples may be incubated with a permeabilizing agent to promote sample permeation. Other methods for sample permeation are described, for example, Jamur et al., 2010, *Method Mol. Biol.* 588:63-66, 2010, the entire contents of which are incorporated herein by reference.
[0242] lysis reagent
[0243] In some embodiments, biological samples can be permeabilized by adding one or more lysis reagents to the sample. Examples of suitable lysis reagents include, but are not limited to, bioactive agents such as lysozymes for lysing different cell types (e.g., Gram-positive or Gram-negative bacteria, plants, yeast, mammals), such as lysozyme, colorless peptidase, lysostaphin, labiase, cell lysin, cytolysin, and a variety of other commercially available lysins.
[0244] Additional or alternative lysis agents can be added to biological samples to facilitate permeation. For example, surfactant-based lysis solutions can be used to lyse sample cells. Lysis solutions may include ionic surfactants such as sodium dodecyl sarcosinate and sodium dodecyl sulfate (SDS). More generally, chemical lysis agents may include, but are not limited to, organic solvents, chelating agents, detergents, surfactants, and liquid release agents.
[0245] In some embodiments, biological samples can be permeated using non-chemical permeation methods. Non-chemical permeation methods are known in the art. For example, non-chemical permeation methods that can be used include, but are not limited to, physical lysis techniques (such as electroporation), mechanical permeation methods (e.g., beading with a homogenizer and grinding balls to mechanically disrupt the tissue structure of the sample), acoustic permeation (e.g., sonication), and thermal lysis techniques (such as heating) to induce thermal permeation of the sample.
[0246] protease
[0247] In some embodiments, the culture medium, solution, or permeation solution may contain one or more proteases. In some embodiments, treatment of biological samples with a protease capable of degrading histones may result in the generation of fragmented genomic DNA. The same capture domain used for capturing mRNA (e.g., a capture domain having a p-(T) sequence) can be used to capture fragmented genomic DNA. In some embodiments, treatment of biological samples with a protease capable of degrading histone proteins and an RNA protectant prior to spatial analysis facilitates the capture of genomic DNA and mRNA.
[0248] In some embodiments, a biological sample is permeabilized by exposing the sample to a protease capable of degrading histone proteins. As used herein, the term "histone" generally refers to adaptor histones (e.g., H1) and / or core histones (e.g., H2A, H2B, H3, and H4). In some embodiments, the protease degrades the adaptor histone, the core histone, or both. Any suitable protease capable of degrading histones in a biological sample can be used. Non-limiting examples of proteases capable of degrading histones include proteases inhibited by leucopeptide and TLCK (toluenesulfonyl-L-lysyl-chloromethane hydrochloride), proteases encoded by the EUO gene from Chlamydia trachomatisserovar A, granzyme A, serine proteases (e.g., trypsin or trypsin-like proteases, neutral serine proteases, elastase, cathepsin G), aspartic proteases (e.g., cathepsin D), peptidase family C1 enzymes (e.g., cathepsin L), pepsin, proteinase K, proteases inhibited by the diazomethane inhibitor Z-Phe-Phe-CHN(2) or the epoxide inhibitor E-64, lysosomal proteases, or azurophiles (e.g., cathepsin G, elastase, proteinase 3, neutral serine proteases).In some embodiments, the serine protease is a trypsin, a trypsin-like enzyme, or a functional variant or derivative thereof (e.g., P00761; C0HK48; Q8IYP2; Q8BW11; Q6IE06; P35035; P00760; P06871; Q90627; P16049; P07477; P00762; P35031; P19799; P350). 36; Q29463; P06872; Q90628; P07478; P07146; P00763; P35032; P70059; P29786; P3503 7; Q90629; P35030; P08426; P35033; P35038; P12788; P29787; P35039; P35040; Q8NHM4; P35041;P35043;P35044;P54624;P04814;P35045;P32821;P54625;P35004;P35046;P 32822;P35047;C0HKA5;C0HKA2;P54627;P35005;C0HKA6;C0HKA3;P52905;P83348;P0 0765; P35042; P81071; P35049; P51588; P35050; P35034; P35051; P24664; P35048; P00764; P00775; P54628; P42278; P54629; P42279; Q91041; P54630; P42280; COHKA4) or combinations thereof. In some embodiments, the trypsin is P00761, P00760, Q29463 or combinations thereof. In some embodiments, the protease capable of degrading one or more histones comprises an amino acid sequence having at least 80% sequence identity with P00761, P00760 or Q29463. In some embodiments, the protease capable of degrading one or more histones comprises an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with P00761, P00760, or Q29463. A protease may be considered a functional variant if it exhibits at least 50% of the activity of a protease under optimal conditions, for example, at least 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the activity of a protease under optimal conditions.In some embodiments, enzyme treatment with pepsin or pepsin-like enzymes may include: P03954 / PEPA1_MACFU; P28712 / PEPA1_RABIT; P27677 / PEPA2_MACFU; P27821 / PEPA2_RABIT; P0DJD8 / PEPA3_HUMAN; P27822 / PEPA3_RABIT; P0DJD7 / PEPA4_HUMAN; P27678 / PEPA4_MACFU; P28713 / PEPA4_RABIT; P0DJD9 / PEPA5_HUMAN; Q9D106 / PEPA5_MOUSE; P27823 / PEPAF_RABIT; P00792 / PEPA_BOVIN; Q9N2D4 / PEPA_CALJA; Q9GMY6 / PEPA_CANLF; P00793 / PEPA_CHICK; P11489 / PEPA_MACMU; P00791 / PEPA_PIG; Q9GMY7 / PEPA_RHIFE; Q9GMY8 / PEPA_SORUN; P81497 / PEPA_SUNMU; P13636 / PEPA_URSTH and their functional variants and derivatives, or combinations thereof. In some embodiments, the pepsinase may include: P00791 / PEPA_PIG; P00792 / PEPA_BOVIN, their functional variants, derivatives, or combinations thereof.
[0249] Additionally, the protease may be included in the reaction mixture (solution), which also includes other components (e.g., buffers, salts, chelating agents (e.g., EDTA) and / or detergents (e.g., SDS, N-lauroyl sarcosinate sodium solution)). The reaction mixture may be buffered to have a pH of about 6.5 to 8.5, for example, about 7.0 to 8.0. Furthermore, the reaction mixture can be used at any suitable temperature, such as about 10 to 50°C, for example, about 10 to 44°C, 11 to 43°C, 12 to 42°C, 13 to 41°C, 14 to 40°C, 15 to 39°C, 16 to 38°C, 17 to 37°C, for example, about 10°C, 12°C, 15°C, 18°C, 20°C, 22°C, 25°C, 28°C, 30°C, 33°C, 35°C, or 37°C, preferably about 35 to 45°C, for example, about 37°C.
[0250] Other reagents
[0251] In some embodiments, the permeation solution may contain additional reagents, or the biological sample may be treated with additional reagents to optimize permeation of the biological sample. In some embodiments, the additional reagent is an RNA protectant. As used herein, the term "RNA protectant" generally refers to a reagent that protects RNA from RNA nucleases (e.g., RNase). Any suitable RNA protectant that protects RNA from degradation may be used. Non-limiting examples of RNA protectants include organic solvents (e.g., at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% v / v), including but not limited to ethanol, methanol, propanol, acetone, trichloroacetic acid, propanol, polyethylene glycol, acetic acid, or combinations thereof. In some embodiments, the RNA protectant comprises ethanol, methanol, and / or propanol, or combinations thereof. In some embodiments, the RNA protectant comprises RNAlaterICE (Thermo Fisher Scientific). In some embodiments, the RNA protectant comprises at least about 60% ethanol. In some embodiments, the RNA protectant comprises about 60 to 95% ethanol, about 0 to 35% methanol, and about 0 to 35% propanol, wherein the total amount of organic solvents in the culture medium does not exceed about 95%. In some embodiments, the RNA protectant comprises about 60 to 95% ethanol, about 5 to 20% methanol, and about 5 to 20% propanol, wherein the total amount of organic solvents in the culture medium does not exceed about 95%.
[0252] In some embodiments, the RNA protectant comprises a salt. The salt may include ammonium sulfate, ammonium bisulfate, ammonium chloride, ammonium acetate, cesium sulfate, cadmium sulfate, ferric(II)cesium sulfate, chromium(III) sulfate, cobalt(II) sulfate, copper(II) sulfate, lithium chloride, lithium acetate, lithium sulfate, magnesium sulfate, magnesium chloride, manganese sulfate, manganese chloride, potassium chloride, potassium sulfate, sodium chloride, sodium acetate, sodium sulfate, zinc chloride, zinc acetate, and zinc sulfate. In some embodiments, the salt is a sulfate, such as ammonium sulfate, ammonium bisulfate, cesium sulfate, cadmium sulfate, ferric(II)cesium sulfate, chromium(III) sulfate, cobalt(II) sulfate, copper(II) sulfate, lithium sulfate, magnesium sulfate, manganese sulfate, potassium sulfate, sodium sulfate, or zinc sulfate. In some embodiments, the salt is ammonium sulfate. Salt can be present at a concentration of about 20 g / 100 ml of culture medium or lower, such as about 15 g / 100 ml, 10 g / 100 ml, 9 g / 100 ml, 8 g / 100 ml, 7 g / 100 ml, 6 g / 100 ml, 5 g / 100 ml or lower, such as about 4 g, 3 g, 2 g or 1 g / 100 ml.
[0253] Additionally, the RNA protectant may be contained in a culture medium that further includes a chelating agent (e.g., EDTA), a buffer (e.g., sodium citrate, sodium acetate, potassium citrate, or potassium acetate, preferably sodium acetate) and / or a pH buffered to about 4 to 8 (e.g., about 5).
[0254] In some embodiments, the biological sample is treated with one or more RNA protectants before, during, or after permeabilization. For example, the biological sample is treated with one or more RNA protectants before treatment with one or more permeabilization agents (e.g., one or more proteases). In another example, the biological sample is treated with a solution comprising one or more RNA protectants and one or more permeabilization agents (e.g., one or more proteases). In yet another example, the biological sample is treated with one or more RNA protectants after treatment with one or more permeabilization agents (e.g., one or more proteases). In some embodiments, the biological sample is treated with one or more RNA protectants before fixation.
[0255] In some embodiments, identifying the location of an analyte captured in a biological sample includes a nucleic acid extension reaction. In some embodiments, the nucleic acid extension reaction includes a DNA polymerase when the capture probe captures a fragmented genomic DNA molecule. For example, the nucleic acid extension reaction includes extending the capture probe using a DNA polymerase, the capture probe hybridizing with the captured analyte (e.g., fragmented genomic DNA) as a template. The product of the extension reaction includes a spatially barcoded analyte (e.g., spatially barcoded fragmented genomic DNA). The spatially barcoded analyte (e.g., spatially barcoded fragmented genomic DNA) can be used to identify the spatial location of an analyte in a biological sample. Any DNA polymerase capable of extending a capture probe using a captured analyte as a template can be used in the methods described herein. Non-limiting examples of DNA polymerases include T7 DNA polymerase; Bsu DNA polymerase; and E. coli DNA polymerase pol I.
[0256] Anti-diffusion medium
[0257] In some embodiments, an antidiffusion medium, typically used to limit analyte diffusion, may include at least one permeation reagent. For example, an antidiffusion medium (e.g., a hydrogel) may include pores (e.g., micropores, nanopores, or lenticels or pores) containing a permeation buffer or reagent. In some embodiments, the antidiffusion medium (e.g., the hydrogel) is soaked in a permeation buffer before contacting the hydrogel with a sample. In some embodiments, the hydrogel or other antidiffusion medium may contain a dried reagent or monomer to deliver the permeation reagent when the antidiffusion medium is applied to a biological sample. In some embodiments, the antidiffusion medium (e.g., the hydrogel) is covalently attached to a solid substrate (e.g., an acrylated glass slide).
[0258] In some embodiments, the hydrogel can be modified to deliver a permeation reagent and contain a capture probe. For example, a hydrogel film can be modified to include a spatial barcode capture probe. The spatial barcode hydrogel film is then immersed in a permeation buffer before contacting the sample. In another example, the hydrogel can be modified to include a spatial barcode capture probe and is designed to function as a porous membrane (e.g., a permeable hydrogel) when exposed to a permeation buffer or any other biological sample preparation reagent. The permeation reagent diffuses through the permeable hydrogel of the spatial barcode and permeates the biological sample on the other side of the hydrogel. Then, upon exposure to the permeation reagent, the analyte diffuses into the spatial barcode hydrogel. In such cases, the spatial barcode hydrogel (e.g., a porous membrane) facilitates the diffusion of bioanalytes from the biological sample into the hydrogel. In some embodiments, the bioanalyte diffuses into the hydrogel before exposure to the permeation reagent (e.g., when secreted analytes are present outside the biological sample, or when the biological sample is otherwise lysed or permeated before the permeation reagent is added). In some embodiments, the permeation reagent flows through the hydrogel at a variable flow rate (e.g., any flow rate that promotes diffusion of the permeation reagent through the spatial barcode hydrogel). In some embodiments, the permeation reagent flows through a microfluidic chamber or channel above the spatial barcode hydrogel. In some embodiments, after the permeation reagent is introduced into the biological sample using flow, a biosample preparation reagent may flow through the hydrogel to further promote the diffusion of the bioanalyte into the spatial barcode hydrogel. Thus, the spatial barcode hydrogel film delivers the permeation reagent to the sample surface in contact with the spatial barcode hydrogel, thereby enhancing analyte migration and capture. In some embodiments, the spatial barcode hydrogel is applied to the sample and placed in a permeation bulk solution. In some embodiments, a hydrogel film soaked in the permeation reagent is sandwiched between the sample and the spatial barcode array. In some embodiments, the target analyte is capable of diffusing through the permeation reagent-soaked hydrogel and hybridizing or binding to a capture probe on the other side of the hydrogel. In some embodiments, the thickness of the hydrogel is proportional to the resolution loss. In some embodiments, the pores (e.g., micropores, nanopores, or pilopores) may contain spatial barcode capture probes and permeation reagent and / or buffer. In some embodiments, the spatial barcode capture probe and the permeation reagent are held between gaps. In some embodiments, the sample is perforated, cut, or transferred into the well, wherein the target analyte diffuses through the permeation reagent / buffer to the spatial barcode capture probe. In some embodiments, the resolution loss may be proportional to the gap thickness (e.g., the amount of permeation buffer between the sample and the capture probe). In some embodiments, the anti-diffusion medium (e.g., hydrogel) is approximately 50 to 500 micrometers thick, including 500, 450, 400, 350, 300, 250, 200, 150, 100, or 50 micrometers thick, or any thickness within 50 to 500 micrometers.
[0259] In some embodiments, a biological sample is exposed to a porous membrane (e.g., a permeable hydrogel) to aid permeation and limit analyte loss through diffusion, while allowing the permeation reagent to reach the sample. Membrane chemistry and pore volume can be controlled to minimize analyte loss. In some embodiments, the porous membrane may be made of glass, silicon, paper, hydrogel, polymer monolith, or other materials. In some embodiments, the material may be naturally porous. In some embodiments, the material may have pores or holes etched into a solid material. In some embodiments, the permeation reagent flows through a microfluidic chamber or channel above the porous membrane. In some embodiments, flow control brings the sample close to the permeation reagent. In some embodiments, the porous membrane is a permeable hydrogel. For example, the hydrogel is permeable when the permeation reagent and / or biosample preparation reagent can diffuse through it. Any suitable permeation reagent and / or biosample preparation reagent described herein can be used under conditions sufficient to release analytes (e.g., nucleic acids, proteins, metabolites, lipids, etc.) from the biological sample. In some embodiments, one side of the hydrogel is exposed to the biological sample, and the other side is exposed to the permeation reagent. The permeation reagent diffuses through the permeable hydrogel and permeates the biological sample on the other side of the hydrogel. In some embodiments, the permeabilizing agent flows through the hydrogel at a variable flow rate (e.g., any flow rate that facilitates diffusion of the permeabilizing agent through the hydrogel). In some embodiments, the permeabilizing agent flows through a microfluidic chamber or channel above the hydrogel. Allowing the permeabilizing agent to flow through the hydrogel enables control over the concentration of the agent. In some embodiments, the hydrogel chemistry and pore volume can be tuned to enhance permeability and limit the loss of diffused analytes.
[0260] In some embodiments, a porous membrane is sandwiched between a spatial barcode array and a sample, wherein a permeation solution is applied to the porous membrane. The permeation reagent diffuses through the pores of the membrane and into the biological sample. In some embodiments, the biological sample may be placed on a substrate (e.g., a glass slide). The bioanalyte then diffuses through the porous membrane and into the space containing the capture probe. In some embodiments, the porous membrane is modified to include the capture probe. For example, the capture probe can be attached to the surface of the porous membrane using any of the methods described herein. In another instance, the capture probe may be embedded in the porous membrane at any depth that allows interaction with the bioanalyte. In some embodiments, the porous membrane is placed on the biological sample in a configuration that allows the capture probe on the porous membrane to interact with the bioanalyte from the biological sample. For example, the capture probe is located on the side of the porous membrane closer to the biological sample. In such cases, the permeation reagent on the other side of the porous membrane diffuses through the porous membrane to the location containing the biological sample and the capture probe to facilitate permeation of the biological sample (e.g., also facilitates capture of the bioanalyte by the capture probe). In some embodiments, the porous membrane is located between the sample and the capture probe. In some embodiments, the permeation reagent flows through a microfluidic chamber or channel above the porous membrane.
[0261] Selective permeation / selective lysis
[0262] In some embodiments, biological samples can be processed according to established methods to selectively release analytes from subcellular regions of cells. In some embodiments, the methods provided herein may include detecting at least one bioanalyte present in subcellular regions of cells within a biological sample. As used herein, "subcellular region" may refer to any subcellular region. For example, a subcellular region may refer to cytosol, mitochondria, nucleus, nucleolus, endoplasmic reticulum, lysosomes, vesicles, Golgi apparatus, plastids, vacuoles, ribosomes, cytoskeleton, or combinations thereof. In some embodiments, a subcellular region comprises at least one of cytosol, nucleus, mitochondria, and microsomes. In some embodiments, a subcellular region is cytosol. In some embodiments, a subcellular region is nucleus. In some embodiments, a subcellular region is mitochondria. In some embodiments, a subcellular region is microsomes.
[0263] For example, bioanalytes can be selectively released from subcellular regions of a cell via selective permeabilization or selective lysis. In some embodiments, "selective permeabilization" can refer to a permeabilization method that can permeate a membrane of a subcellular region while leaving the different subcellular regions substantially intact (e.g., the bioanalyte is not released from the subcellular region due to the applied permeabilization method). Non-limiting examples of selective permeabilization methods include using electrophoresis and / or applying a permeabilizing agent. In some embodiments, "selective lysis" can refer to a lysis method that lyses a membrane of a subcellular region while leaving the different subcellular regions substantially intact (e.g., the bioanalyte is not released from the subcellular region due to the applied lysis method). Several methods of selective permeation or cleavage are known to those skilled in the art, including those described in Lu et al., Lab Chip., January 2005; 5(1):23-9; Niklas et al., 2011, Anal Biochem, 416(2):218-27; Cox and Emili, 2006, Handbook of Experiments in Nature, 1(4):1872-8; Chiang et al., 2000, J Biochem.Biophys.Methods., 20; 46(1-2):53-68; and Yamauchi and Herr et al., 2017, Microsystems and Nanoengineering, 3.pii:16079, each of which is incorporated herein by reference in its entirety.
[0264] In some embodiments, "selective permeabilization" or "selective lysis" refers to the selective permeabilization or selective lysis of a specific cell type. For example, "selective permeabilization" or "selective lysis" can refer to lysing one cell type while leaving the different cell types substantially intact (e.g., the bioanalyte is not released from the cell due to the permeabilization or lysis method applied). Cells of a "different cell type" from another cell can refer to cells from different taxa, prokaryotic and eukaryotic cells, cells from different tissue types, etc. Many methods for selectively permeabilizing or lysing different cell types are known to those skilled in the art. Non-limiting examples include the application of permeabilizing agents, electroporation, and / or sonication. See, for example, International Application No. WO 2012 / 168003; Han et al., 2019, Microsystems & Nanoengineering 5:30; Gould et al., 2018, Oncotarget. 20; 9(21):15606-15615; Oren and Shai, 1997, Biochemistry 36(7), 1826-35; Algayer et al., 2019, Molecules. 24(11).pii:E2079; Hipp et al., 2017, Leukemia 10, 2278; International Application No. WO 2012 / 168003; and U.S. Patent No. 7,785,869; all of these references are incorporated herein by reference in their entirety.
[0265] In some embodiments, applying a selective permeabilizing or lysis reagent involves contacting a biological sample with a hydrogel containing the permeabilizing or lysis reagent.
[0266] In some embodiments, the biological sample is contacted with two or more arrays (e.g., flexible arrays, as described herein). For example, after the subcellular regions are permeated and bioanalytes from the subcellular regions are captured on a first array, the first array can be removed, and bioanalytes from different subcellular regions can be captured on a second array.
[0267] (13) Selective enrichment of RNA species
[0268] In some embodiments, when RNA is the analyte, one or more RNA analyte species of interest may be selectively enriched (e.g., Adiconis et al., 2013, Comparative analysis of RNA sequencing methods for degraded and low-input samples, Nature 10, 623-632, which is incorporated herein by reference in full). For example, one or more RNAs may be selected by adding one or more oligonucleotides to the sample. In some embodiments, the additional oligonucleotide is a sequence for initiating a reaction via a polymerase. For example, one or more primer sequences having sequence complementarity with one or more RNAs of interest may be used to amplify one or more RNAs of interest, thereby selectively enriching these RNAs. In some embodiments, oligonucleotides having sequence complementarity with the complementary strand of the captured RNA (e.g., cDNA) may bind to the cDNA. For example, biotinylated oligonucleotides having sequences complementary to one or more cDNAs of interest bind to the cDNA and may be selected using biotinylation-streptavitin affinity (e.g., streptavitin beads) using any of a variety of methods known in the art.
[0269] Alternatively, any of a variety of methods can be used to downselect (e.g., remove, deplete) one or more RNAs (e.g., ribosomal and / or mitochondrial RNA). Non-limiting examples of hybridization and capture methods for ribosomal RNA deletion include RiboMinus. TM RiboCop TM and Ribo-Zero TM Another non-restrictive RNA deletion method involves hybridizing complementary DNA oligonucleotides with unwanted RNA, followed by degradation of the RNA / DNA hybrid using RNase H. Non-restrictive examples of hybridization and degradation methods include... rRNA depletion, NuGEN AnyDeplete, or RiboZero Plus. Another non-restrictive ribosomal RNA deletion method includes ZapR. TM The ribosomal RNA is digested, for example, by SMARTer. In the SMARTer method, a random nucleic acid adaptor hybridizes to RNA for first-strand synthesis, followed by tailing via reverse transcriptase, template conversion, and extension via reverse transcriptase. Additionally, a first-round PCR amplification adds a full-length Illumina sequencing adaptor (e.g., Illumina index). The ribosomal RNA is cleaved by ZapR v2 and R probe v2. A second-round PCR is performed to amplify non-rRNA molecules (e.g., cDNA). Some or all steps of these ribosomal depletion protocols / kits can be further combined with the methods described herein to optimize protocols for specific biological samples.
[0270] In depletion protocols, probes can be applied to the sample, selectively hybridizing with ribosomal RNA (rRNA) to reduce the pooling and concentration of rRNA in the sample. Probes can also be applied to biological samples, selectively hybridizing with mitochondrial RNA (mtRNA) to reduce the pooling and concentration of mtRNA in the sample. In some embodiments, probes complementary to mitochondrial RNA can be added during cDNA synthesis, or probes complementary to both ribosomal and mitochondrial RNA can be added during cDNA synthesis. The subsequent application of capture probes to the sample, due to the reduction in nonspecific RNA (e.g., downselected RNA) present in the sample, may lead to improved capture of other types of RNA. Additionally and alternatively, double-stranded nuclease (DSN) treatment can remove rRNA (see, e.g., Archer et al., 2014, “Selective and flexible removal of problematic sequences from RNA-seq libraries at the cDNA stage,” BMC Genomics 15 401, the entire contents of which are incorporated herein by reference). In addition, hydroxyapatite chromatography can remove abundant species (e.g., rRNA) (see, for example, Vandernoot, 2012, “Normalization of cDNA by hydroxyapatite chromatography to enrich transcriptomic diversity in RNA-seq applications,” Biotechniques, 53(6):373-80, the entire contents of which are incorporated herein by reference).
[0271] (14) Other reagents
[0272] Additional reagents can be added to biological samples to perform various functions prior to sample analysis. In some embodiments, nuclease inhibitors (such as DNase and RNase inactivators) or protease inhibitors (such as proteinase K), and / or chelating agents (such as EDTA) can be added to the sample. In other embodiments, nucleases (such as DNase or RNase), or proteases (such as pepsin or proteinase K) can be added to the sample. In some embodiments, the additional reagents can be dissolved in solution or applied to the sample as a medium. In some embodiments, the additional reagents (e.g., pepsin) can be dissolved in HCl before application to the sample. For example, hematoxylin from H&E staining can optionally be removed from the biological sample by washing in dilute HCl (0.001M to 0.1M) prior to further processing. In some embodiments, pepsin can be dissolved in dilute HCl (0.001M to 0.1M) prior to further processing. In some embodiments, the biological sample can be washed an additional number of times (e.g., 2, 3, 4, 5 or more times) in dilute HCl before incubation with a protease (e.g., pepsin) but after treatment with proteinase K.
[0273] In some embodiments, the sample may be treated with one or more enzymes. For example, one or more endonucleases for fragmenting DNA, DNA polymerases, and dNTPs for amplifying nucleic acids may be added. Other enzymes that may also be added to the sample include, but are not limited to, polymerases, transposases, ligases, DNase, and RNase.
[0274] In some embodiments, reverse transcriptase may be added to the sample, including an enzyme with terminal transferase activity, primers, and template-changing oligonucleotides (TSOs). Template conversion can be used to increase the length of cDNA, for example, by appending a predetermined nucleic acid sequence to the cDNA. This reverse transcription step is as follows: Figure 37 As shown. In some embodiments, the additional nucleic acid sequence comprises one or more ribonucleotides.
[0275] In some embodiments, additional reagents may be added to improve the recovery rate of one or more target molecules (e.g., cDNA molecules, mRNA transcripts). For example, adding vector RNA to the RNA sample workflow can increase the yield of RNA / DNA hybrids extracted from biological samples. In some embodiments, vector molecules are useful when the concentration of the input or target molecule is low compared to the remaining molecules. Typically, a single target molecule does not form a precipitate, and the addition of a vector molecule can facilitate precipitate formation. Some target molecule recovery protocols use vector RNA to prevent the irreversible binding of small amounts of target nucleic acids present in the sample. In some embodiments, vector RNA may be added immediately before the second-strand synthesis step. In some embodiments, vector RNA may be added immediately before the synthesis of second-strand cDNA on oligonucleotides released from the array. In some embodiments, vector RNA may be added immediately before the in vitro post-transcriptional clearance step. In some embodiments, vector RNA may be added before the purification and quantification of amplified RNA. In some embodiments, vector RNA may be added before RNA quantification. In some embodiments, vector RNA may be added immediately before the second-strand cDNA synthesis and in vitro post-transcriptional clearance steps.
[0276] (15) Capture probe interaction
[0277] In some embodiments, the analyte in the biological sample may be pretreated before interacting with the capture probe. For example, a polymerase-catalyzed polymerization reaction (e.g., DNA polymerase or reverse transcriptase) may be performed in the biological sample before interaction with the capture probe. In some embodiments, the primers used for the polymerization reaction include functional groups that enhance hybridization with the capture probe. The capture probe may include a suitable capture domain to capture the bioanalyte of interest (e.g., capture a poly-dT sequence of poly(A)mRNA).
[0278] In some embodiments, bioanalytes are pretreated using next-generation sequencing to generate libraries. For example, the analytes can be pretreated by adding modifications (e.g., ligation of sequences that allow interaction with capture probes). In some embodiments, the analytes (e.g., DNA or RNA) are fragmented using fragmentation techniques (e.g., using transposases and / or fragmentation buffers).
[0279] After fragmentation, the analyte can be modified. For example, modification can be added by attaching an adaptor sequence that allows hybridization with a capture probe. In some embodiments, when the analyte of interest is RNA, poly(A) tailing is performed. Adding a poly(A) tail to RNA without a poly(A) tail can promote hybridization with a capture probe that includes a capture domain with a poly(dT) sequence having a functional amount.
[0280] In some embodiments, a ligase-catalyzed ligation reaction is performed in the biological sample prior to interaction with the capture probe. In some embodiments, ligation can be performed via chemical ligation. In some embodiments, click chemistry, as further described below, can be used for ligation. In some embodiments, the capture domain comprises a DNA sequence complementary to an RNA molecule, wherein the RNA molecule is complementary to a second DNA sequence, and wherein the RNA-DNA sequence complementarity is used to ligate the second DNA sequence to the DNA sequence in the capture domain. In these embodiments, direct detection of the RNA molecule is possible.
[0281] In some embodiments, a target-specific reaction is performed in the biological sample prior to interaction with the capture probe. Examples of target-specific reactions include, but are not limited to, the ligation of target-specific linkers, probes, and / or other oligonucleotides; target-specific amplification using primers specific to one or more analytes; and target-specific detection using in situ hybridization, DNA microscopy, and / or antibody detection. In some embodiments, the capture probe includes a capture domain that targets a target-specific product (e.g., amplified or ligated).
[0282] II. A General Analysis Method Based on Spatial Arrays
[0283] This section of the disclosure describes methods, apparatus, systems, and compositions for spatial array-based analysis of biological samples.
[0284] (a) Spatial analysis methods
[0285] Array-based spatial analysis methods involve transferring one or more analytes from a biological sample to an array of capture points on a substrate, each capture point associated with a unique spatial location on the array. Subsequent analysis of the transferred analytes involves determining the identity of the analytes and the spatial location of each analyte in the sample. The spatial location of each analyte in the sample is determined based on the capture point associated with each analyte in the array and the relative spatial locations of the capture points within the array.
[0286] There are at least two common methods for associating spatial barcodes with one or more adjacent cells, such that the spatial barcode identifies one or more cells and / or the contents of one or more cells as associated with a specific spatial location. One common method is to induce the analyte to exit from the cell and face the spatial barcode array. Figure 1 An exemplary embodiment of this general method is depicted. Figure 1 In this configuration, a spatial barcode array (as further described herein) filled with capture probes comes into contact with the sample 101, and the sample is permeated 102, thereby allowing the target analyte to leave the sample and migrate toward the array 102. The target analyte interacts with the capture probes on the spatial barcode array. Once the target analyte has hybridized / bound with the capture probes, the sample is optionally removed from the array and the capture probes are analyzed to obtain spatially resolved analyte information 103.
[0287] Another common approach is to cut spatial barcode capture probes from the array and push and / or advance them into or onto the sample. Figure 2 An exemplary embodiment of this general method is depicted, in which a spatial barcode array (as further described herein) filled with capture probes can contact a sample 201. The spatial barcode capture probes are cut and then interact with cells in the provided sample 202. The interaction can be a covalent or non-covalent cell-surface interaction. The interaction can be an intracellular interaction facilitated by a delivery system or a cell-penetrating peptide. Once the spatial barcode capture probes associate with specific cells, the sample can optionally be removed for analysis. The sample can optionally be dissociated prior to analysis. Once the labeled cells associate with the spatial barcode capture probes, the capture probes can be analyzed to obtain spatial resolution information about the labeled cells 203.
[0288] Figure 3A and 3B An exemplary workflow is shown, including the preparation of sample 301 on a spatial barcode array. Sample preparation may include placing the sample on a substrate (e.g., a chip, a slide, etc.), fixing the sample, and / or staining the sample for imaging. Bright-field imaging (e.g., staining with hematoxylin and eosin) or fluorescence imaging (e.g., imaging of capture points) is then used. Figure 3B (as shown in Figure 302 above) and / or emission imaging modes (such as...) Figure 3B (As shown in Figure 304 below) Image the sample (stained or unstained) on the array 302.
[0289] Brightfield images are transmission microscopy images in which a broad-spectrum white light is placed on one side of a sample mounted on a substrate, a camera objective is placed on the other side, and the sample itself filters the light to produce a color or grayscale intensity image 1124, similar to a stained glass window viewed from the inside on a bright day.
[0290] In some embodiments, emission imaging, such as fluorescence imaging, is used in addition to or instead of bright-field imaging. In emission imaging methods, a sample on a substrate is exposed to light in a specific narrow band (a first wavelength band), and then the light re-emitted from the sample at a slightly different wavelength (a second wavelength band) is measured. This absorption and re-emission is due to the presence of a fluorophore, which is sensitive to the excitation used and can be a natural property of the sample or a reagent to which the sample has been exposed during imaging preparation. As an example, in immunofluorescence experiments, an antibody that binds to a protein or class of proteins and is labeled with a fluorophore is added to the sample. When this is done, the site on the sample containing that protein or class of proteins will emit the second wavelength band. In fact, multiple antibodies with multiple fluorophores can be used to label multiple proteins in a sample. Each such fluorophore needs to be excited with light of a different wavelength and further emits light of a different, unique wavelength. To spatially resolve each different emission wavelength, the sample is exposed to light of different wavelengths, which will excite multiple fluorophores on a sequential basis, and images of each of these exposures are saved as images, thus generating multiple images. For example, an image is captured by exciting a first wavelength at a second wavelength to excite a first fluorophore, and a first image of the sample is captured while the sample is exposed to the first wavelength. Then, exposure to the first wavelength is stopped, and the sample is exposed to a third wavelength (different from the first wavelength), which excites a second fluorophore at a fourth wavelength (different from the second wavelength), and a second image of the sample is captured while the sample is exposed to the third wavelength. This process is repeated for each different fluorophore among multiple fluorophores (e.g., two or more fluorophores, three or more fluorophores, four or more fluorophores, five or more fluorophores). In this way, a series of images of the tissue is obtained, each image depicting a spatial arrangement of some different parameters, such as a specific protein or protein class. In some embodiments, more than one fluorophore is imaged simultaneously. In this method, a combination of excitation wavelengths is used, each excitation wavelength is used for one of more than one fluorophore, and a single image is collected.
[0291] In some embodiments, each image collected by emission imaging is grayscale. To distinguish these grayscale images, in some embodiments, each image is assigned a color (red shading, blue shading, etc.) and combined into a composite color image for viewing. This fluorescence imaging allows for spatial analysis of protein abundance in a sample (e.g., spatial proteomics). In some embodiments, this spatial abundance is analyzed alone. In other embodiments, this spatial abundance is analyzed in conjunction with transcriptomics.
[0292] In some embodiments, when analyzing a sample using transcriptomics along with bright-field and / or emission imaging (e.g., fluorescence imaging), a capture probe that releases the target analyte from the sample and forms a spatial barcode array hybridizes with or binds to the released target analyte 303. Optionally, the sample can be removed from the array 304, and the capture probe can optionally be cut from the array 305. The sample and array are then optionally imaged a second time in two modes 305B, while the analyte is reverse transcribed into cDNA, an amplicon library is prepared 306, and sequencing is performed 307. The images are then spatially overlaid to correlate spatially identified sample information 308. When the sample and array are not imaged a second time, 305B, a point coordinate file is provided instead. The point coordinate file replaces the second imaging step 305B. Furthermore, amplicon library preparation 306 and sequencing 307 can be performed using a unique PCR adaptor.
[0293] Figure 4 Another exemplary workflow is illustrated using a spatial barcode array on a substrate (e.g., a chip), where spatial barcode capture probes cluster at regions called capture points. Spatially labeled capture probes may include cleavage domains, one or more functional sequences, spatial barcodes, unique molecular identifiers, and capture domains. Spatially labeled capture probes may also include 5' end modifications for reversible attachment to the substrate. The spatial barcode array is brought into contact with a sample 401, and the sample is permeated 402 by applying a permeation reagent. The permeation reagent can be applied by placing the array / sample assembly in a bulk solution. Alternatively, the permeation reagent can be applied to the sample via an anti-diffusion medium and / or a physical barrier (such as a cap), wherein the sample is sandwiched between the anti-diffusion medium and / or barrier and the array-containing substrate. Using any number of techniques disclosed herein, analytes migrate to the spatial barcode capture array. For example, anti-diffusion medium caps and passive migration can be used for analyte migration. As another example, analyte migration can be active migration, for example, using an electrophoretic transfer system. Once the analyte is very close to the spatial barcode capture probe, the capture probe can hybridize or otherwise bind to the target analyte 403. Sample 404 can be optionally removed from the array.
[0294] The capture probe can optionally be cut 405 times from the array, and the captured analyte can be spatially barcoded by performing a first-strand cDNA reaction using reverse transcriptase. The first-strand cDNA reaction can optionally be performed using template-converting oligonucleotides. For example, template-converting oligonucleotides can be hybridized to a poly(C) tail added to the 3' end of the cDNA using reverse transcriptase. Template conversion is as follows... Figure 37 As shown. The original mRNA template can then be denatured from cDNA, and oligonucleotides can be converted from the template. A spatial barcode capture probe can then hybridize with the cDNA, and a cDNA complement can be generated. The first-strand cDNA can then be purified and collected for downstream amplification steps. Optionally, PCR amplification of the first-strand cDNA 406 can be used, with forward and reverse primers side-attached to the spatial barcode of interest and the target analyte region, thereby generating a library 407 associated with a specific spatial barcode. In some embodiments, library preparation can be quantified and / or quality controlled to verify the success of library preparation step 408. In some embodiments, the cDNA contains sequencing-by-synthesis (SBS) primer sequences. The library amplicon is sequenced and analyzed to decode the spatial information 407, with an additional library quality control (QC) step 408.
[0295] Using the methods, compositions, systems, kits, and devices described herein, RNA transcripts present in biological samples (e.g., tissue samples) can be used for spatial transcriptome analysis. Specifically, in some cases, barcoded oligonucleotides can be configured to initiate, replicate, and thus produce barcoded extension products from an RNA template or its derivatives. For example, in some cases, barcoded oligonucleotides may include mRNA-specific initiation sequences, such as poly-T primer segments or other targeted initiation sequences that allow initiation and replication of mRNA in a reverse transcription reaction. Alternatively or additionally, random RNA initiation can be performed using random N-mer primer segments of barcoded oligonucleotides. Reverse transcriptase (RT) can use an RNA template and a primer complementary to the 3' end of the RNA template to direct the synthesis of first-strand complementary DNA (cDNA). Many RTs can be used for this reverse transcription reaction, including, for example, avian myeloblastomavirus (AMV) reverse transcriptase, Moloney murine leukemia virus (M-MuLV or MMLV), and other variants. Some recombinant M-MuLV reverse transcriptases, such as... II reverse transcriptases, compared to their wild-type counterparts, may exhibit reduced RNase H activity and increased thermostability, and provide higher specificity, higher cDNA yield, and more full-length cDNA products up to 12 kilobases (kb). In some embodiments, the reverse transcriptase is a mutant reverse transcriptase, such as, but not limited to, a mutant MMLV reverse transcriptase. In another embodiment, the reverse transcriptase is a mutant MMLV reverse transcriptase, such as, but not limited to, one or more variants described in U.S. Patent Publication No. 20180312822 and U.S. Provisional Patent Application No. 62 / 946,885, filed December 11, 2019, both of which are incorporated herein by reference in their entirety.
[0296] Figure 5 An exemplary workflow is described, in which a sample is removed from a spatial barcode array, and a spatial barcode capture probe is retrieved from the array for barcode analyte amplification and library preparation. Another embodiment includes first-strand synthesis on the spatial barcode array using template-converting oligonucleotides without cleaving the capture probe. In this embodiment, sample preparation 501 and permeabilization 502 are performed as described elsewhere herein. Once the capture probe captures the target analyte, the first-strand cDNA 503 generated by template conversion and reverse transcriptase is subsequently denatured, and the second strand is subsequently extended 504. The second-strand cDNA is then denatured, neutralized, and transferred from the first-strand cDNA into a tube 505. cDNA quantification and amplification can be performed using the standard techniques discussed herein. Library preparation 506 and indexing 507 can then be performed on the cDNA, including fragmentation, end repair, α-tailing, and indexing PCR steps. Optionally, quality control (QC) of the library can also be tested 508.
[0297] In a non-limiting example of the above workflow, biological samples (e.g., tissue sections) may be fixed with methanol, stained with hematoxylin and eosin, and imaged. Optionally, the sample may be destained prior to permeabilization. The images may be used to map spatial analyte abundance (e.g., gene expression) patterns back to the biological sample. Permeabilizing enzymes may be used to directly permeabilize biological samples on a slide. Analytes released from overlying cells of the biological sample (e.g., polyadenylated mRNA) may be captured by capture probes within capture zones on the substrate. Reverse transcription (RT) reagents may be added to the permeabilized biological sample. Incubation with RT reagents may generate spatially barcoded full-length cDNA from the captured analytes (e.g., polyadenylated mRNA). Second-strand reagents (e.g., second-strand primers, enzymes) may be added to the biological sample on the slide to initiate second-strand synthesis. The resulting cDNA may be denatured from the capture probe template and transferred (e.g., to a clean tube) for amplification and / or library construction. The spatially barcoded full-length cDNA may be amplified by PCR prior to library construction. cDNA fragments can then be fragmented and their sizes selected to optimize cDNA amplicon size. P5, P7, i7, and i5 can be used as sample indexes and TruSeq reads 2 can be added via end repair, A-tailing, linker ligation, and PCR. The cDNA fragments are then sequenced using paired-end sequencing with TruSeq reads 1 and 2 as sequencing primer sites. See Illumina, Indexed Sequencing Overview Guides, February 2018, document 15057455v04; and Illumina linker sequences, May 2019, document #1000000002694v11, each incorporated herein by reference for information on P5, P7, i7, i5, TruSeq reads 2, indexing sequencing, and other reagents described herein.
[0298] In some embodiments, correlation analysis of the data generated by this workflow and other workflows described herein can produce a correlation of more than 95% (e.g., 95% or higher, 96% or higher, 97% or higher, 98% or higher, or 99% or higher) for genes expressed across two capture regions. When the described workflow is performed using single-cell RNA sequencing of the cell nucleus, in some embodiments, correlation analysis of the data can produce a correlation of more than 90% (e.g., more than 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) for genes expressed across two capture regions.
[0299] In some embodiments, cDNA can be amplified directly on the substrate surface after generation (e.g., by reverse transcription). Amplifying multiple copies of the generated cDNA (e.g., cDNA synthesized from captured analytes) directly on the substrate surface increases the complexity of the final sequencing library. Therefore, in some embodiments, cDNA can be amplified directly on the substrate surface via isothermal nucleic acid amplification. In some embodiments, isothermal nucleic acid amplification can amplify either RNA or DNA.
[0300] In some embodiments, isothermal amplification can be faster than a standard PCR reaction. In some embodiments, isothermal amplification can be linear amplification (e.g., asymmetric with a single primer) or exponential amplification (e.g., with two primers). In some embodiments, isothermal nucleic acid amplification can be performed using template-converting oligonucleotide primers. In some embodiments, the template-converting oligonucleotide adds a common sequence to the 5' end of reverse-transcribed RNA. For example, after the capture probe interacts with the analyte (e.g., mRNA) and undergoes reverse transcription, additional nucleotides are added to the end of the cDNA, creating a 3' overhang as described herein. In some embodiments, the template-converting oligonucleotide hybridizes with untemplated poly(C)nucleotides added via reverse transcriptase to continue replication to the 5' end of the template-converting oligonucleotide, thereby generating full-length cDNA ready for further amplification. In some embodiments, the template-converting oligonucleotide adds a common 5' sequence (e.g., the reverse complement of the template-converting oligonucleotide) to the full-length cDNA used for cDNA amplification.
[0301] In some embodiments, once a full-length cDNA molecule is generated, template-converting oligonucleotides can be used as primers in a cDNA amplification reaction (e.g., with DNA polymerase). In some embodiments, double-stranded cDNA (e.g., first-strand cDNA and second-strand reverse complementary cDNA) can be amplified isothermally with a helicase or recombinase, followed by amplification with strand displacement DNA polymerase. The strand displacement DNA polymerase generates a substituted second strand, thereby producing the amplification product.
[0302] In any of the isothermal amplification methods described herein, barcode exchange (e.g., spatial barcodes) may occur after the first amplification cycle, where unused capture probes remain on the substrate surface. In some embodiments, the free 3'OH end of the unused capture probe can be blocked by any suitable 3'OH blocking method. In some embodiments, the 3'OH can be blocked by hairpin connection.
[0303] In addition to, or as an alternative to, standard PCR reactions (e.g., PCR reactions that require heating to approximately 95°C to denature double-stranded DNA), isothermal nucleic acid amplification can be used. Isothermal nucleic acid amplification typically does not require the use of a thermal cycler; however, in some embodiments, isothermal amplification can be performed in a thermal cycler. In some embodiments, isothermal amplification can be performed at approximately 35°C to approximately 75°C. In some embodiments, isothermal amplification can be performed at approximately 40°C, approximately 45°C, approximately 50°C, approximately 55°C, approximately 60°C, approximately 65°C, or approximately 70°C, or any temperature in between, depending on the polymerase and coenzyme used.
[0304] Isothermal nucleic acid amplification techniques are known in the art and can be used alone or in combination with any spatial method described herein. For example, non-limiting examples of suitable isothermal nucleic acid amplification techniques include transcription-mediated amplification, nucleic acid sequence-based amplification, signal-mediated amplification using RNA technology, strand substitution amplification, rolling circle amplification, loop-mediated isothermal amplification of DNA (LAMP), isothermal multiple substitution amplification, recombinase polymerase amplification, helicase-dependent amplification, single-primer isothermal amplification, and loop helicase-dependent amplification (see, for example, Gill and Ghaemi, “Isothermal amplification of nucleic acids: a review,” Nucleosides, Nucleotides, & Nucleic Acids, 27(3), 224–43, doi:10.1080 / 15257770701845204 (2008), which is incorporated herein by reference in its entirety).
[0305] In some embodiments, isothermal nucleic acid amplification is helicase-dependent nucleic acid amplification. Helicase-dependent isothermal nucleic acid amplification is described in Vincent et al., 2004, “Helicase-dependent isothermal DNA amplification,” EMBO Report, 795–800, and U.S. Patent No. 7,282,328, both of which are incorporated herein by reference in their entirety. Furthermore, helicase-dependent nucleic acid amplification on a substrate (e.g., a chip) is described in Andresen et al., 2009, “Helicase-dependent amplification: for chip amplification and potential point-of-care diagnostics,” Expert Rev Mol Diagn. 9, 645–650, doi:10.1586 / erm.09.46, the entire contents of which are incorporated herein by reference. In some embodiments, isothermal nucleic acid amplification is recombinase polymerase nucleic acid amplification. Recombinase polymerase amplification of nucleic acids is described in Piepenburg et al., “DNA detection using recombinant proteins”, PLOS Biol., 4, 7e204, and Li et al., “Review: A comprehensive overview of a decade of development in recombinase polymerase amplification”, Analyst, 144, 31-67, doi:10.1039 / C8AN01621F (2019), both of which are incorporated herein by reference in full.
[0306] Typically, isothermal amplification techniques use standard PCR reagents known in the art (e.g., buffers, dNTPs, etc.). Some isothermal amplification techniques may require additional reagents. For example, helicase-dependent nucleic acid amplification uses single-strand binding proteins and accessory proteins. In another example, recombinase polymerase nucleic acid amplification uses recombinases (e.g., T4 UvsX), recombinase loading factors (e.g., TF UvsY), single-strand binding proteins (e.g., T4 gp32), congestants (e.g., PEG-35K), and ATP.
[0307] Following isothermal amplification of the full-length cDNA using any of the methods described herein, the isothermally amplified cDNA (e.g., single-stranded or double-stranded) can be recovered from the substrate and optionally subsequently amplified using typical cDNA PCR in a microcentrifuge tube. The sample can then be used with any of the spatial methods described herein.
[0308] Immunohistochemistry and immunofluorescence
[0309] In some embodiments, immunofluorescence or immunohistochemistry protocols (direct and indirect staining techniques) are performed as part of or in addition to the exemplary spatial workflows presented herein. For example, tissue sections can be fixed according to the methods described herein. Biological samples can be transferred to an array (e.g., a capture probe array) in which an analyte (e.g., a protein) is detected using an immunofluorescence protocol. For example, the sample can be rehydrated, blocked, and permeabilized (3XSSC, 2% BSA, 0.1% Triton X, 1 U / μl RNase inhibitor, at 4°C for 10 min), and then stained with a fluorescent primary antibody (at a 1:100 ratio in 3XSSC, 2% BSA, 0.1% Triton X, 1 U / μl RNase inhibitor, at 4°C for 30 min). Biological samples can be washed, covered with coverslips (in glycerol + 1 U / μl RNase inhibitor), imaged (e.g., using a confocal microscope or other equipment capable of fluorescence detection), washed, and processed according to the analyte capture or spatial workflows described herein.
[0310] As used herein, “antigen retrieval buffer” can improve antibody capture in IF / IHC protocols. An exemplary protocol for antigen retrieval may be to preheat the antigen retrieval buffer (e.g., to 95°C), immerse the biological sample in the heated antigen retrieval buffer for a predetermined time, and then remove the biological sample from the antigen retrieval buffer and wash the biological sample.
[0311] In some embodiments, optimized permeabilization can be used to identify intracellular analytes. Permeabilization optimization may include selecting a permeabilizing agent, its concentration, and the duration of permeabilization. Tissue permeabilization is discussed elsewhere in this document.
[0312] In some embodiments, blocking the array and / or biological sample during the preparation of the labeled biological sample reduces non-specific binding of the antibody to the array and / or biological sample (reducing background). Some embodiments provide blocking buffers / blocking solutions that can be applied before and / or during labeling, wherein the blocking buffer may include an inhibitor and optionally a surfactant and / or a salt solution. In some embodiments, the inhibitor may be bovine serum albumin (BSA), serum, gelatin (e.g., fish gelatin), milk (e.g., skim milk powder), casein, polyethylene glycol (PEG), polyvinyl alcohol (PVA) or polyvinylpyrrolidone (PVP), biotin blocking agent, peroxidase blocking agent, levamisole, carnosine solution, glycine, lysine, sodium borohydride, benzoyl peroxide, Sudan Black, trypan blue, FITC inhibitor, and / or acetic acid. The blocking buffer / blocking solution may be applied to the array and / or biological sample before and / or during labeling (e.g., applying a fluorophore-conjugated antibody) to the biological sample.
[0313] In some embodiments, additional steps or optimizations may be included when performing the IF / IHC protocol in conjunction with a space array. Additional steps or optimizations may be included when performing the space-tagged analyte capture workflow discussed herein.
[0314] In some embodiments, this document provides a method for spatially detecting an analyte (e.g., detecting the location of an analyte (e.g., a bioanalyte)) from a biological sample (e.g., an analyte present in a biological sample, such as a tissue section), comprising: (a) providing a biological sample on a substrate; (b) staining the biological sample on the substrate, imaging the stained biological sample, and selecting the biological sample or a portion of the biological sample (e.g., a region of interest) for analysis; (c) providing an array comprising one or more capture probes on the substrate; (d) contacting the biological sample with the array to allow one or more capture probes to capture the analyte of interest; and (e) analyzing the captured analyte to spatially detect the analyte of interest. Any kind of staining and imaging technique described herein or known in the art may be used according to the methods described herein. In some embodiments, staining includes optical markers described herein, including but not limited to fluorescent, radioactive, chemiluminescent, calorimetric, or colorimetrically detectable markers. In some embodiments, staining includes fluorescent antibodies against a target analyte (e.g., a cell surface or intracellular protein) in the biological sample. In some embodiments, staining includes immunohistochemical staining against a target analyte in the biological sample (e.g., cell surface or intracellular proteins). In some embodiments, staining includes chemical staining, such as hematoxylin and eosin (H&E) or periodic acid-Schiff (PAS). In some embodiments, a considerable amount of time (e.g., days, months, or years) may pass between staining and / or imaging the biological sample and performing analysis. In some embodiments, reagents for analysis are added to the biological sample before, simultaneously with, or after the array contacts the biological sample. In some embodiments, step (d) includes placing the array on the biological sample. In some embodiments, the array is a flexible array in which multiple spatial barcode features (e.g., a substrate with capture probes, beads with capture probes) are attached to a flexible substrate. In some embodiments, measures are taken to slow the reaction before the array contacts the biological sample (e.g., cooling the temperature of the biological sample or using an enzyme that preferentially performs its primary function at temperatures lower or higher than its optimal functional temperature). In some embodiments, step (e) is performed without removing the biological sample from the array. In some embodiments, step (e) is performed after the biological sample is no longer in contact with the array. In some embodiments, the biological sample is labeled with an analyte trapping agent before, during, or after staining and / or imaging. In this case, a considerable amount of time may pass between staining and / or imaging and the analysis (e.g., days, months, or years). In some embodiments, the array is adapted to facilitate the migration of bioanalytes from the stained and / or imaged biological sample to the array (e.g., using any of the materials or methods described herein). In some embodiments, the biological sample is permeabilized before contact with the array.In some embodiments, the permeation rate is slowed before the biological sample is contacted with the array (e.g., to limit the diffusion of analytes away from their original location in the biological sample). In some embodiments, the permeation rate can be modulated (e.g., the activity of the permeation reagent) by adjusting the conditions to which the biological sample is exposed (e.g., adjusting temperature, pH, and / or light). In some embodiments, the permeation rate can be modulated using an external stimulus (e.g., a small molecule, enzyme, and / or activating agent). For example, a permeation reagent, which is inactive, can be provided to the biological sample before contact with the array until the conditions (e.g., temperature, pH, and / or light) change or an external stimulus (e.g., a small molecule, enzyme, and / or activating agent) is provided.
[0315] In some embodiments, this document provides a method for spatially detecting an analyte (e.g., detecting the location of an analyte (e.g., a bioanalyte)) from a biological sample (e.g., a biological sample present in a tissue section), comprising: (a) providing the biological sample on a substrate; (b) staining the biological sample on the substrate, imaging the stained biological sample, and selecting the biological sample or a portion of the biological sample (e.g., a region of interest) for spatial transcriptome analysis; (c) providing an array on the substrate comprising one or more capture probes; (d) contacting the biological sample with the array to allow one or more capture probes to capture the bioanalyte of interest; and (e) analyzing the captured bioanalyte to spatially detect the bioanalyte of interest.
[0316] (b) Capture probe
[0317] The term “capture probe”, which may also be used interchangeably as “probe” herein, refers to any molecule capable of capturing (directly or indirectly) and / or labeling an analyte (e.g., the analyte of interest) in a biological sample. In some embodiments, the capture probe is a nucleic acid or peptide. In some embodiments, the capture probe is a conjugate (e.g., an oligonucleotide-antibody conjugate). In some embodiments, the capture probe includes a barcode (e.g., a spatial barcode and / or a unique molecular identifier (UMI)) and a capture domain.
[0318] Figure 6 This is a schematic diagram illustrating an example of a capture probe as described herein. As shown, the capture probe 602 is optionally coupled to the capture point 601 via a cutting domain 603 (such as a disulfide connector).
[0319] The capture probe 602 may include functional sequences for subsequent processing, such as functional sequence 604, which may include sequencer-specific mobile cell attachment sequences, such as the P5 sequence, and functional sequence 606, which may include sequencing primer sequences, such as R1 primer binding sites and R2 primer binding sites. In some embodiments, sequence 604 is the P7 sequence and sequence 606 is the R2 primer binding site.
[0320] Spatial barcode 605 can be included within the capture probe for barcoding the target analyte. Functional sequences compatible with various sequencing systems can be selected, such as 454 sequencing, ion-fluid proton or PGM sequencing, enomina sequencing instruments, PacBio, Oxford nanopore sequencing, and their requirements. In some embodiments, functional sequences compatible with non-commercial sequencing systems can be selected. Examples of such sequencing systems and techniques that can use suitable functional sequences include (but are not limited to) ion-fluid proton or PGM sequencing, enomina sequencing, PacBio SMRT sequencing, and Oxford nanopore sequencing. Furthermore, in some embodiments, functional sequences compatible with other sequencing systems, including non-commercial sequencing systems, can be selected.
[0321] In some embodiments, the spatial barcode 605, functional sequence 604 (e.g., a mobile cell attachment sequence), and 606 (e.g., a sequencing primer sequence) may be common to all probes attached to a given capture point. The spatial barcode may also include a capture field 607 to facilitate the capture of the target analyte.
[0322] (i) Capture domain.
[0323] As described above, each capture probe includes at least one capture domain 607. A “capture domain” is an oligonucleotide, peptide, small molecule, or any combination thereof that specifically binds to a desired analyte. In some embodiments, the capture domain can be used to capture or detect the desired analyte.
[0324] In some embodiments, the capture domain is a functional nucleic acid sequence configured to interact with one or more analytes, such as one or more different types of nucleic acids (e.g., RNA molecules and DNA molecules). In some embodiments, the functional nucleic acid sequence may include an N-mer sequence (e.g., a random N-mer sequence) configured to interact with multiple DNA molecules. In some embodiments, the functional sequence may include a poly(T) sequence configured to interact with messenger RNA (mRNA) molecules via a poly(A) tail of an mRNA transcript. In some embodiments, the functional nucleic acid sequence is a binding target for a protein (e.g., a transcription factor, a DNA-binding protein, or an RNA-binding protein), wherein the analyte of interest is a protein.
[0325] The capture probe may include ribonucleotides and / or deoxyribonucleotides and synthetic nucleotide residues capable of participating in Wattssen-Crick-type or similar base pair interactions. In some embodiments, the capture domain is capable of initiating a reverse transcription reaction to generate cDNA complementary to the captured RNA molecule. In some embodiments, the capture domain of the capture probe may initiate a DNA extension (polymerase) reaction to generate DNA complementary to the captured DNA molecule. In some embodiments, the capture domain may template a ligation reaction between the captured DNA molecule and a surface probe directly or indirectly immobilized on a substrate. In some embodiments, the capture domain may be ligated to one strand of the captured DNA molecule. For example, SplintR ligase can be used to ligate single-stranded DNA or RNA to the capture domain along with an RNA or DNA sequence (e.g., degenerate RNA). In some embodiments, ligases with RNA template ligase activity, such as SplintR ligase, T4 RNA ligase 2, or KOD ligase, can be used to ligate single-stranded DNA or RNA to the capture domain. In some embodiments, the capture domain includes a splint oligonucleotide. In some embodiments, the capture domain captures the splint oligonucleotide.
[0326] In some embodiments, the capture domain is located at the 3' end of the capture probe and includes a free 3' end that can be extended, for example, by template-dependent polymerization to form an extended capture probe as described herein. In some embodiments, the capture domain includes a nucleotide sequence capable of hybridizing with nucleic acids (e.g., RNA or other analytes) present in cells of a tissue sample in contact with the array. In some embodiments, the capture domain may be selected or designed to selectively or specifically bind to target nucleic acids. For example, the capture domain may be selected or designed to capture mRNA by hybridizing with the poly(A) tail of mRNA. Thus, in some embodiments, the capture domain includes a poly(T)DNA oligonucleotide, such as a series of consecutive deoxythymidine residues linked by phosphodiester bonds, capable of hybridizing with the poly(A) tail of mRNA. In some embodiments, the capture domain may include nucleotides that are functionally or structurally similar to the poly(T) tail. For example, a poly-U oligonucleotide or an oligonucleotide comprising a deoxythymidine analog. In some embodiments, the capture domain includes at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides. In some embodiments, the capture domain includes at least 25, 30, or 35 nucleotides.
[0327] In some embodiments, the capture probe includes a capture domain having a sequence capable of binding mRNA and / or genomic DNA. For example, the capture probe may include a capture domain comprising a poly(A) tail capable of binding mRNA and / or a nucleic acid sequence (e.g., a poly(T) sequence) present in genomic DNA with a poly(A) homopolymeric sequence. In some embodiments, a terminal transferase is used to add the homopolymeric sequence to an mRNA molecule or a genomic DNA molecule to produce an analyte having a poly(A) or poly(T) sequence. For example, a poly(A) sequence may be added to an analyte (e.g., a fragment of genomic DNA) such that the analyte can be captured by a poly(T) capture domain.
[0328] In some embodiments, a random sequence, such as a random hexamer or similar sequence, may be used to form all or part of the capture domain. For example, the random sequence may be used in conjunction with a poly(T) (or poly(T) analog) sequence. Therefore, when the capture domain includes a poly(T) (or “poly(T)-like”) oligonucleotide, it may also include a random oligonucleotide sequence (e.g., a “poly(T)-random sequence” probe). This may be located, for example, at the 5' or 3' of the poly(T) sequence, such as at the 3' end of the capture domain. The poly(T)-random sequence probe can facilitate the capture of the mRNA poly(A) tail. In some embodiments, the capture domain may be a completely random sequence. In some embodiments, a degenerate capture domain may be used.
[0329] In some embodiments, a set of two or more capture probes forms a mixture, wherein the capture domain of one or more capture probes comprises a poly(T) sequence, and the capture domain of one or more capture probes comprises a random sequence. In some embodiments, a set of two or more capture probes forms a mixture, wherein the capture domain of one or more capture probes comprises a poly(T)-random sequence, and the capture domain of one or more capture probes comprises a random sequence. In some embodiments, a set of two or more capture probes forms a mixture, wherein the capture domain of one or more capture probes comprises a poly(T)-random sequence, and the capture domain of one or more capture probes comprises a random sequence. In some embodiments, a probe having a degenerate capture domain may be added to any of the foregoing combinations listed herein. In some embodiments, a probe having a degenerate capture domain may replace one of the probes in each pair described herein.
[0330] The capture domain can be based on a specific gene sequence or motif sequence or common / conserved sequence designed for capture (i.e., a sequence-specific capture domain). Therefore, in some embodiments, the capture domain can selectively bind to a desired nucleic acid isotype or subgroup, such as a specific type of RNA, such as mRNA, rRNA, tRNA, SRP RNA, tmRNA, snRNA, snoRNA, SmYRNA, scaRNA, gRNA, RNase P, RNase MRP, TERC, SL RNA, aRNA, cis-NAT, crRNA, lncRNA, miRNA, piRNA, siRNA, shRNA, tasiRNA, rasiRNA, 7SK, eRNA, ncRNA, or other types of RNA. In a non-limiting example, the capture domain can selectively bind to a desired ribonucleic acid subgroup, such as microbiome RNA, like 16S rRNA.
[0331] In some embodiments, the capture domain includes an "anchor" or "anchor sequence," which is a nucleotide sequence designed to ensure hybridization of the capture domain with a intended bioanalyte. In some embodiments, the anchor sequence includes a nucleotide sequence, including a 1-mer, 2-mer, 3-mer, or longer sequence. In some embodiments, the short sequence is random. For example, a capture domain including a poly(T) sequence can be designed to capture mRNA. In such embodiments, the anchor sequence may include a random 3-mer (e.g., GGG), which helps ensure hybridization of the poly(T) capture domain with mRNA. In some embodiments, the anchor sequence may be VN, N, or NN. Alternatively, a specific nucleotide sequence can be designed. In some embodiments, the anchor sequence is at the 3' end of the capture domain. In some embodiments, the anchor sequence is at the 5' end of the capture domain.
[0332] In some embodiments, the capture domain of the capture probe is blocked before the biological sample is contacted with the array, and the blocking probe is used when nucleic acids in the biological sample are modified before they are captured on the array. In some embodiments, the blocking probe is used to block or modify the free 3' end of the capture domain. In some embodiments, the blocking probe may hybridize with the capture probe to mask the free 3' end of the capture domain, such as a hairpin probe, a partially double-stranded probe, or a complementary sequence. In some embodiments, the free 3' end of the capture domain may be blocked by chemical modification, such as adding an azidomethyl group as a chemically reversible capping portion, such that the capture probe does not include the free 3' end. Blocking or modifying the capture probe, particularly the free 3' end of the capture domain, before contacting the biological sample with the array prevents modification of the capture probe, for example, preventing the addition of a poly(A) tail to the free 3' end of the capture probe.
[0333] Non-limiting examples of 3' modifications include dideoxy-3' (3'-ddC), 3' reverse dT, 3' C3 spacer, 3' amino group, and 3' phosphorylation. In some embodiments, nucleic acids in a biological sample can be modified such that they can be captured by a capture domain. For example, a linker sequence (including a binding domain capable of binding the capture domain of a capture probe) can be added to the end of a nucleic acid (e.g., fragmented genomic DNA). In some embodiments, this is achieved by ligating a linker sequence or extending the nucleic acid. In some embodiments, an enzyme is used to incorporate additional nucleotides, such as poly(A) tails, into the ends of the nucleic acid sequence. In some embodiments, the capture probe can be reversibly masked or modified such that the capture domain of the capture probe does not include a free 3' end. In some embodiments, the 3' end is removed, modified, or made inaccessible such that the capture domain is insensitive to processes (e.g., ligation or extension) used to modify the nucleic acid in the biological sample.
[0334] In some embodiments, the capture domain of the capture probe is modified to allow removal of any modifications to the capture probe that occur during the modification of nucleic acid molecules in a biological sample. In some embodiments, the capture probe may include an additional sequence downstream of the capture domain, namely the 3' of the capture domain, i.e., the blocking domain.
[0335] In some embodiments, the capture domain of the capture probe may be a non-nucleic acid domain. Examples of suitable capture domains that are not entirely based on nucleic acids include, but are not limited to, proteins, peptides, aptamers, antigens, antibodies, and functional molecular analogs that mimic any capture domain described herein.
[0336] (ii) Cutting domain.
[0337] Each capture probe may optionally include at least one cleaving domain. A cleaving domain represents a portion of the probe used to reversibly attach the probe to the array capture point, as will be described further below. Furthermore, one or more segments or regions of the capture probe may optionally be released from the array capture point by cleaving the cleaving domain. As an example, spatial barcodes and / or universal molecular identifiers (UMIs) may be released by cleaving the cleaving domain.
[0338] Figure 7 This is a schematic diagram illustrating a cleavable capture probe, where the cleaved capture probe can enter non-permeable cells and bind to target analytes in the sample. Capture probe 602 includes a cleavage domain 603, a cell-penetrating peptide 703, a reporter molecule 704, and a disulfide bond (-SS-). 705 represents all other parts of the capture probe, such as spatial barcodes and capture domains.
[0339] In some embodiments, the cleavage domain 603 that attaches the capture probe to the capture point is a covalent bond that can be cleaved by an enzyme. An enzyme may be added to cleave the cleavage domain, causing the capture probe to be released from the capture point. As another example, heating may also cause the degradation of the cleavage domain and the release of the attached capture probe from the array of capture points. In some embodiments, laser radiation is used to heat and degrade the cleavage domain of the capture probe at a specific location. In some embodiments, the cleavage domain is a photosensitive chemical bond (e.g., a chemical bond that dissociates when exposed to light such as ultraviolet light). In some embodiments, the cleavage domain may be an ultrasonic cleavage domain. For example, ultrasonic cleavage may depend on the nucleotide sequence, length, pH, ionic strength, temperature, and ultrasonic frequency (e.g., 22 kHz, 44 kHz) (Grokhovsky, SL, "Specificity of Sonication of DNA," Molecular Biology, 40(2), 276-283 (2006)).
[0340] Other examples of cleavage domain 603 include unstable chemical bonds, such as, but not limited to, ester bonds (e.g., cleavable by acid, base or hydroxylamine), tandem diol bonds (e.g., cleavable by sodium periodate), Diels-Alder bonds (e.g., cleavable by heat), sulfone bonds (e.g., cleavable by base), silyl ether bonds (e.g., cleavable by acid), glycosidic bonds (e.g., cleavable by amylase), peptide bonds (e.g., cleavable by protease) or phosphodiester bonds (e.g., cleavable by nuclease (e.g., DNAase)).
[0341] In some embodiments, the cleavage domain 603 includes a sequence recognized by one or more enzymes capable of cleaving nucleic acid molecules, such as breaking phosphodiester bonds between two or more nucleotides. These bonds can be cleaved by other nucleic acid molecule-targeting enzymes, such as restriction enzymes (e.g., restriction endonucleases). For example, the cleavage domain may include a restriction endonuclease (restriction enzyme) recognition sequence. Restriction enzymes cleave double-stranded or single-stranded DNA at a specific recognition nucleotide sequence called a restriction site. In some embodiments, rare-cleavage restriction enzymes, such as those with long recognition sites (at least 8 base pairs in length), are used to reduce the likelihood of cleavage at other locations on the capture probe.
[0342] Oligonucleotides with photosensitive chemical bonds (e.g., photocleavable linkers) offer various advantages. They can be cleaved efficiently and rapidly (e.g., within nanoseconds and milliseconds). In some cases, photomasks can be used so that only specific regions of the array are exposed to the cleavable stimulus (e.g., exposure to UV light, exposure to light, exposure to laser-induced heat). When photocleavable linkers are used, the cleavage reaction is light-triggered and can be highly selective for the linker, thus being biorthogonal. Typically, the wavelength absorption of photocleavable linkers lies in the near-UV range of the spectrum. In some embodiments, the λmax of the photocleavable linker is about 300 nm to about 400 nm, or about 310 nm to about 365 nm. In some embodiments, the λmax of the photocleavable linker is about 300 nm, about 312 nm, about 325 nm, about 330 nm, about 340 nm, about 345 nm, about 355 nm, about 365 nm, or about 400 nm. Non-limiting examples of photosensitive chemical bonds that can be used to cleave domains are disclosed in PCT Publication 202020176788A1 entitled “Analysis of bioanalytes using spatial barcode oligonucleotide arrays”, the entire contents of which are incorporated herein by reference.
[0343] In some embodiments, the cleavage domain includes a region that can be cleaved by uracil DNA glycosylase (UDG) and DNA glycosylase-lyase endonuclease VIII (commercially known as USER). TM A mixture of enzymes cleaves the poly-U sequence. Once released, the releasable capture probe can be used for the reaction. Thus, for example, the activatable capture probe can be activated by releasing the capture probe from the capture site.
[0344] In some embodiments, when the capture probe is indirectly attached to the substrate, for example via a surface probe, the cleavage domain comprises one or more mismatched nucleotides such that the complementary portions of the surface probe and the capture probe are not 100% complementary (e.g., the number of mismatched base pairs can be one, two, or three base pairs). This mismatch is recognized by, for example, MutY and T7 endonuclease I, which result in the cleavage of the nucleic acid molecule at the mismatch site. As described herein, a “surface probe” can be any portion present on the substrate surface capable of attaching to a reagent (e.g., a capture probe). In some embodiments, the surface probe is an oligonucleotide. In some embodiments, the surface probe is part of the capture probe.
[0345] In some embodiments, when the capture probe is attached to the capture site indirectly (e.g., immobilized) via a surface probe, the cleavage domain includes a nickase recognition site or sequence. A nickase is a single-stranded endonuclease that cleaves only the DNA duplex. Therefore, the cleavage domain may include a nickase recognition site near the 5' end of the surface probe (and / or the 5' end of the capture probe), such that cleavage of the surface probe or the capture probe destabilizes the duplex between the surface probe and the capture probe, thereby releasing the capture probe from the capture site.
[0346] Cleavage enzymes can also be used in some embodiments where the capture probe is directly attached (e.g., immobilized) to the capture site. For example, the substrate can contact a nucleic acid molecule that hybridizes with the cleavage domain of the capture probe to provide or reconstruct a cleavage enzyme recognition site, such as a cleavage helper probe. Thus, contact with the cleavage enzyme will result in cleavage of the cleavage domain, thereby releasing the capture probe from the capture site. Such cleavage helper probes can also be used to provide or reconstruct cleavage recognition sites for other cleavage enzymes (e.g., restriction enzymes).
[0347] Some nickases introduce single-strand cleavage only at specific sites on a DNA molecule by binding to and recognizing specific nucleotide recognition sequences. Many naturally occurring nickases have been discovered, and sequence recognition properties have been identified in at least four of them. Nickases are described in U.S. Patent No. 6,867,028, which is incorporated herein by reference in its entirety. Generally, any suitable nickase can be used to bind to a complementary nickase recognition site of the cleavage domain. After use, the nickase can be removed from the assay or inactivated after the release of the capture probe to prevent unwanted cleavage of the capture probe.
[0348] In some embodiments, the capture probe does not have a cleavage domain. For example, an example of a substrate with an attached capture probe lacking a cleavage domain is described in Macosko et al., (2015) Cell 161, 1202-1214, the entire contents of which are incorporated herein by reference.
[0349] Examples of suitable capture domains that are not entirely based on nucleic acids include, but are not limited to, proteins, peptides, aptamers, antigens, antibodies, and functional molecular analogs that mimic any capture domain described herein.
[0350] In some embodiments, the capture probe region corresponding to the cleavage domain can be used for other functions. For example, an additional region may be included for nucleic acid extension or amplification, where the cleavage domain is typically located. In such embodiments, the region may complement the functional domain or even exist as an additional functional domain. In some embodiments, a cleavage domain is present, but its use is optional.
[0351] (iii) Functional domains
[0352] Each capture probe may optionally include at least one functional domain. Each functional domain typically includes a functional nucleotide sequence for downstream analytical steps throughout the analytical process.
[0353] Further details of the functional domains that may be used in conjunction with this disclosure can be found in U.S. Patent Application No. 16 / 992,569, entitled “System and Method for Determining Biological Status Using Spatial Distribution of Monotypes,” filed August 13, 2020, and PCT Publication 202020176788A1, entitled “Analysis of Bioanalytes Using Spatial Barcode Oligonucleotide Arrays,” both of which are incorporated herein by reference.
[0354] (iv) Spatial barcodes.
[0355] As described above, the capture probe may include one or more spatial barcodes (e.g., two or more, three or more, four or more, five or more). A “spatial barcode” is a continuous nucleic acid segment or two or more non-contiguous nucleic acid segments used as a marker or identifier to convey or be able to convey spatial information. In some embodiments, the capture probe includes spatial barcodes with spatial characteristics, wherein the barcodes are associated with specific locations within an array or on a substrate.
[0356] Spatial barcodes can be part of the analyte or independent of the analyte (i.e., part of a capture probe). In addition to endogenous characteristics of the analyte (e.g., size or end sequences), a spatial barcode can be a tag or combination of tags attached to the analyte (e.g., a nucleic acid molecule). Spatial barcodes can be unique. In some embodiments where the spatial barcode is unique, it serves both as a spatial barcode and as a unique molecular identifier (UMI) associated with a particular capture probe.
[0357] Spatial barcodes can have a variety of different formats. For example, spatial barcodes may include polynucleotide spatial barcodes; random nucleic acid and / or amino acid sequences; and synthetic nucleic acid and / or amino acid sequences. In some embodiments, spatial barcodes are attached to analytes in a reversible or irreversible manner. In some embodiments, spatial barcodes are added to fragments of, for example, DNA or RNA samples before, during, and / or after sample sequencing. In some embodiments, spatial barcodes allow for the identification and / or quantification of individual sequencing reads. In some embodiments, spatial barcodes are used as fluorescent barcodes, with fluorescently labeled oligonucleotide probes hybridizing to the spatial barcodes.
[0358] In some embodiments, a spatial barcode is a nucleic acid sequence that substantially does not hybridize with the analyte nucleic acid molecule in the biological sample. In some embodiments, the spatial barcode has less than 80% sequence identity with the nucleic acid sequence over a large portion (e.g., 80% or more) of the nucleic acid molecule in the biological sample (e.g., less than 70%, 60%, 50%, or less than 40% sequence identity).
[0359] The spatial barcode sequence may comprise about 6 to about 20 or more nucleotides within the sequence of the capture probe. In some embodiments, the length of the spatial barcode sequence may be about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of the spatial barcode sequence may be at least about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of the spatial barcode sequence is at most about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or shorter.
[0360] These nucleotides can be completely continuous, for example, within a segment of adjacent nucleotides, or they can be divided into two or more separate subsequences separated by one or more nucleotides. The length of the separated spatial barcode subsequences can be from about 4 to about 16 nucleotides. In some embodiments, the spatial barcode subsequences can be about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the spatial barcode subsequences can be at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the spatial barcode subsequences can be at most about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter.
[0361] For multiple capture probes attached to a common array capture point, one or more spatial barcode sequences of the multiple capture probes may include the same sequence for all capture probes coupled to the capture point, and / or different sequences for all capture probes coupled to the capture point.
[0362] Figure 8 This is a schematic diagram of the capture points of an exemplary spatial marker for multiplexing. Figure 8In this configuration, capture point 601 can be coupled to a spatial barcode capture probe, wherein the spatial barcode probe for a specific capture point may have the same spatial barcode but different capture domains, and is designed to associate the spatial barcode of the capture point with more than one target analyte. For example, the capture point can be coupled to four different types of spatial barcode capture probes, each type having a spatial barcode 605. One type of capture probe associated with the capture point includes a spatial barcode 605 that binds to a poly(T) capture domain 803, which is designed to capture mRNA target analytes. A second type of capture probe associated with the capture point includes a spatial barcode 605 that binds to a random N-mer capture domain 804 for gDNA analysis. A third type of capture probe associated with the capture point includes a spatial barcode 605 that binds to a capture domain complementary to the capture domain on the analyte capture agent 805. A fourth type of capture probe that associates with a capture site includes a spatial barcode 605 that binds to a capture probe that can specifically bind to nucleic acid molecules 806 that can function in CRISPR assays (e.g., CRISPR / Cas9). Although in Figure 8 Only four different capture probe barcode configurations are shown, but capture probe barcode configurations can be customized for analyzing any given analyte that associates with nucleic acids and can bind to such a configuration. For example, Figure 8 The illustrated scheme can also be used for parallel analysis of other analytes disclosed herein, including but not limited to: (a) mRNA, lineage tracer constructs, cell surface or intracellular proteins and metabolites, and gDNA; (b) mRNA, accessible chromatin (e.g., ATAC-seq, DNase-seq, and / or MNase-seq), cell surface or intracellular proteins and metabolites, and interfering agents (e.g., CRISPR crRNA / sgRNA, TALEN, zinc finger nucleases, and / or antisense oligonucleotides as described herein); (c) V(D)J sequences of mRNA, cell surface or intracellular proteins and / or metabolites, barcode markers (e.g., MHC multimers as described herein), and immune cell receptors (e.g., T cell receptors). In some embodiments, the interfering agent may be a small molecule, antibody, drug, aptamer, miRNA, physical environment (e.g., temperature change), or any other known interfering agent.
[0363] Capture probes attached to a single array capture point may include the same (or common) spatial barcode sequence, different spatial barcode sequences, or a combination of both. Capture probes attached to a capture point may include multiple sets of capture probes. A given set of capture probes may include the same spatial barcode sequence. The same spatial barcode sequence may differ from the spatial barcode sequence of another set of capture probes.
[0364] Multiple capture probes may include spatial barcode sequences (e.g., nucleic acid barcode sequences) associated with specific locations on a spatial array. For example, a first plurality of capture probes may be associated with a first region based on a spatial barcode sequence common to capture probes within a first region, and a second plurality of capture probes may be associated with a second region based on a spatial barcode sequence common to capture probes within a second region. The second region may or may not be associated with the first region. Additional plurality of capture probes may be associated with spatial barcode sequences common to capture probes in other regions. In some embodiments, the spatial barcode sequence may be identical on multiple capture probe molecules.
[0365] In some embodiments, multiple distinct spatial barcodes are combined into a single array capture probe. For example, a mixed but known set of spatial barcode sequences can provide a stronger address or attribution of a spatial barcode to a given point or location by providing repeated or independent verification of location identity. In some embodiments, multiple spatial barcodes represent an increased specificity for a particular array point location.
[0366] (v) Unique molecular identifier.
[0367] Capture probes may include one or more (e.g., two or more, three or more, four or more, five or more) unique molecular identifiers (UMIs). A unique molecular identifier is a continuous nucleic acid segment or two or more non-contiguous nucleic acid segments that serve as a marker or identifier for a particular analyte or a capture probe that binds to a particular analyte (e.g., through a capture domain).
[0368] Further details of the UMI that can be used with the systems and methods disclosed herein are provided in U.S. Patent Application No. 16 / 992,569, entitled “System and Method for Determining Biological Status Using Spatial Distribution of Monotypes,” filed August 13, 2020, and PCT Publication 202020176788A1, entitled “Analysis of Bioanalytes Using Spatial Barcode Oligonucleotide Arrays,” each of which is incorporated herein by reference.
[0369] (vi) Other aspects of the capture probe.
[0370] For a capture probe attached to an array capture point, a single array capture point may include one or more capture probes. In some embodiments, a single array capture point includes hundreds or thousands of capture probes. In some embodiments, a capture probe is associated with a specific single capture point, wherein the single capture point contains a capture probe that includes a spatial barcode unique to a defined area or location on the array.
[0371] In some embodiments, a particular capture point includes a capture probe comprising more than one spatial barcode (e.g., one capture probe at a particular capture point may include a spatial barcode different from the spatial barcode included in another capture probe at the same particular capture point, while both capture probes include a second common spatial barcode), wherein each spatial barcode corresponds to a specific defined region or location on the array. For example, multiple spatial barcode sequences associated with a particular capture point on the array can provide a stronger address or attribute of a given location by providing repeating or independent confirmation of the location. In some embodiments, multiple spatial barcodes represent an increased specificity of the location of a particular array point. In a non-limiting example, a particular array point can be encoded with two different spatial barcodes, wherein each spatial barcode identifies a specific defined region within the array, and the array point having both spatial barcodes identifies a sub-region where the two defined regions overlap, such as the overlapping portion of a Venn diagram.
[0372] In another non-limiting example, a particular array point can be encoded using three different spatial barcodes, wherein a first spatial barcode identifies a first region within the array, a second spatial barcode identifies a second region, wherein the second region is a sub-region entirely within the first region, and a third spatial barcode identifies a third region, wherein the third region is a sub-region entirely within the first and second sub-regions.
[0373] In some embodiments, the capture probe attached to the array capture point is released from the array capture point for sequencing. Alternatively, in some embodiments, the capture probe remains attached to the array capture point, and the probe is sequenced while still attached to the array capture point (e.g., by in situ sequencing). Other aspects of capture probe sequencing are described in subsequent sections of this disclosure.
[0374] In some embodiments, the array capture point may include different types of capture probes attached to the capture point. For example, the array capture point may include a first type of capture probe and a second type of capture probe, the first type of capture probe having a capture domain designed to bind a type of analyte, and the second type of capture probe having a capture domain designed to bind a second type of analyte. Typically, the array capture point may include one or more (e.g., two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, 12 or more, 15 or more, 20 or more, 30 or more, 50 or more) different types of capture probes attached to a single array capture point.
[0375] In some embodiments, the capture probe is a nucleic acid. In some embodiments, the capture probe is attached to an array capture site via its 5' end. In some embodiments, the capture probe includes, from its 5' to 3' end, one or more barcodes (e.g., spatial barcodes and / or UMIs) and one or more capture domains. In some embodiments, the capture probe includes, from its 5' to 3' end, a barcode (e.g., a spatial barcode or a UMI) and a capture domain. In some embodiments, the capture probe includes, from its 5' to 3' end, a cleavage domain, a functional domain, one or more barcodes (e.g., spatial barcodes and / or UMIs), and a capture domain. In some embodiments, the capture probe includes, from its 5' to 3' end, a cleavage domain, a functional domain, one or more barcodes (e.g., spatial barcodes and / or UMIs), a second functional domain, and a capture domain. In some embodiments, the capture probe includes, from its 5' to 3' end, a cleavage domain, a functional domain, a spatial barcode, a UMI, and a capture domain. In some embodiments, the capture probe does not include a spatial barcode. In some embodiments, the capture probe does not include a UMI. In some embodiments, the capture probe includes a sequence for initiating a sequencing reaction.
[0376] In some embodiments, the capture probe is fixed to the capture point at its 3' end. In some embodiments, the capture probe from its 3' to 5' ends includes: one or more barcodes (e.g., spatial barcodes and / or UMIs) and one or more capture fields. In some embodiments, the capture probe from its 3' to 5' ends includes: a barcode (e.g., a spatial barcode or a UMI) and a capture field. In some embodiments, the capture probe from its 3' to 5' ends includes: a cut field, a functional field, one or more barcodes (e.g., spatial barcodes and / or UMIs), and a capture field. In some embodiments, the capture probe from its 3' to 5' ends includes: a cut field, a functional field, a spatial barcode, a UMI, and a capture field.
[0377] In some embodiments, the capture probe comprises an in-situ synthesized oligonucleotide. The in-situ synthesized oligonucleotide may be attached to a substrate or a feature attached to a substrate. In some embodiments, the in-situ synthesized oligonucleotide comprises one or more constant sequences, one or more of which serve as initiating sequences (e.g., primers for amplifying target nucleic acids). The in-situ synthesized oligonucleotide may, for example, include a constant sequence at its 3' end, which is attached to a substrate or a feature attached to a substrate. Additionally or alternatively, the in-situ synthesized oligonucleotide may include a constant sequence at its free 5' end. In some embodiments, the one or more constant sequences may be cleavable sequences. In some embodiments, the in-situ synthesized oligonucleotide comprises a barcode sequence, such as a variable barcode sequence. The barcode may be any barcode described herein. The length of the barcode may be approximately 8 to 16 nucleotides (e.g., 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides). The length of the in-situ synthesized oligonucleotide can be less than 100 nucleotides (e.g., less than 90, 80, 75, 70, 60, 50, 45, 40, 35, 30, 25, or 20 nucleotides). In some cases, the length of the in-situ synthesized oligonucleotide is about 20 to about 40 nucleotides. Exemplary in-situ synthesized oligonucleotides are manufactured by Affymetrix, Inc. In some embodiments, the in-situ synthesized oligonucleotide is attached to the capture sites of the array.
[0378] Additional oligonucleotides can be ligated to in-situ synthesized oligonucleotides to generate capture probes. For example, a primer complementary to a portion of the in-situ synthesized oligonucleotide (e.g., a constant sequence within the oligonucleotide) can be used to hybridize the additional oligonucleotide and extend it (using the in-situ synthesized oligonucleotide as a template, e.g., a primer extension reaction) to form a double-stranded oligonucleotide and further generate a 3' overhang. In some embodiments, the 3' overhang can be generated by a template-independent ligase (e.g., terminal deoxynucleotidyl transferase (TdT) or poly(A) polymerase). Additional oligonucleotides containing one or more capture domains can be ligated to the 3' overhang using suitable enzymes (e.g., ligases) and splice oligonucleotides to generate capture probes. Thus, in some embodiments, the capture probe is the product of two or more oligonucleotide sequences linked together (e.g., the in-situ synthesized oligonucleotide and the additional oligonucleotide). In some embodiments, one of the oligonucleotide sequences is the in-situ synthesized oligonucleotide.
[0379] In some embodiments, the capture probe comprises a splint oligonucleotide. Two or more oligonucleotides can be ligated together using the splint oligonucleotide and any type of ligase known in the art or described herein (e.g., SplintR ligase).
[0380] In some embodiments, one of the oligonucleotides includes: a constant sequence (e.g., a sequence complementary to a portion of the splint oligonucleotide), a degenerate sequence, and a capture domain (e.g., as described herein). In some embodiments, the capture probe is generated by causing an enzyme to add a polynucleotide to the end of the oligonucleotide sequence. The capture probe may include a degenerate sequence, which can be used as a unique molecular identifier.
[0381] The capture probe may include a degenerate sequence, which is a sequence of nucleotides in which some positions contain a number of possible bases. The degenerate sequence may be a degenerate nucleotide sequence comprising about or at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, or 50 nucleotides. In some embodiments, the nucleotide sequence contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 0, 10, 15, 20, 25, or more degenerate positions within the nucleotide sequence. In some embodiments, the degenerate sequence is used as a UMI.
[0382] In some embodiments, the capture probe comprises a restriction endonuclease-recognized sequence or a nucleotide sequence cleavable by specific enzymatic activity. For example, a uracil DNA glycosylase (UDG) or a uracil-specific excision reagent (USER) can be used to enzymatically cleave the uracil sequence from the nucleotide sequence. As another example, other modified bases (e.g., modified by methylation) can be recognized and cleaved by specific endonucleases. The capture probe can be enzymatically cleaved to remove the blocking domain and any additional nucleotides added to the 3' end of the capture probe during the modification process. Removal of the blocking domain reveals and / or restores the free 3' end of the capture domain of the capture probe. In some embodiments, additional nucleotides can be removed to reveal and / or restore the 3' end of the capture domain of the capture probe.
[0383] In some embodiments, the blocking domain may be incorporated into the capture probe during or after its synthesis. The terminal nucleotide of the capture domain is a reversible terminator nucleotide (e.g., a 3'-O-blocking reversible terminator and a 3'-unblocking reversible terminator) and may be included in the capture probe during or after probe synthesis.
[0384] (vii) Extend the capture probe
[0385] "Extended capture probe" is a capture probe with an expanded nucleic acid sequence. For example, in cases where the capture probe comprises nucleic acids, "extended 3' end" means that additional nucleotides are added to the last 3' nucleotide of the capture probe to extend its length, for example, by standard polymerization reactions used to extend nucleic acid molecules, including template polymerization catalyzed by a polymerase (e.g., DNA polymerase or reverse transcriptase).
[0386] In some embodiments, extending the capture probe includes generating cDNA from the captured (hybridized) RNA. The method includes the synthesis of complementary strands of the hybridized nucleic acids, for example, generating cDNA based on a captured RNA template (RNA hybridized to the capture domain of the capture probe). Thus, in the initial step of extending the capture probe, such as cDNA generation, the captured (hybridized) nucleic acid, such as RNA, serves as a template for the extension (e.g., reverse transcription) step.
[0387] In some embodiments, reverse transcription is used to extend the capture probe. For example, reverse transcription includes synthesizing cDNA (complementary or copy DNA) from RNA (e.g., messenger RNA) using a reverse transcriptase. In some embodiments, reverse transcription is performed while the tissue is still in situ to generate an analyte library, wherein the analyte library includes spatial barcodes from adjacent capture probes. In some embodiments, one or more DNA polymerases are used to extend the capture probe.
[0388] In some embodiments, the capture domain of the capture probe includes primers for generating a complementary strand of a nucleic acid that hybridizes to the capture probe, such as primers for DNA polymerase and / or reverse transcription. The nucleic acid molecule (e.g., DNA and / or cDNA) generated by the extension reaction is incorporated into the sequence of the capture probe. The extension of the capture probe, such as by DNA polymerase and / or reverse transcription, can be performed using a variety of suitable enzymes and protocols.
[0389] In some embodiments, a full-length DNA, such as a cDNA molecule, is generated. In some embodiments, a "full-length" DNA molecule refers to the entire captured nucleic acid molecule. However, if the nucleic acid (e.g., RNA) is partially degraded in the tissue sample, the captured nucleic acid molecule will differ in length from the initial RNA in the tissue sample. In some embodiments, the 3' end of the extension probe is modified, for example, by a first-strand cDNA molecule. For example, an adapter or linker can be attached to the 3' end of the extension probe. This can be done using a single-stranded ligase (such as T4 RNA ligase or Circligase). TM (Purchased from Lucigen, Middleton, Wisconsin)) to achieve this. In some embodiments, template-converting oligonucleotides are used to extend cDNA to generate full-length cDNA (or as close to full-length cDNA as possible). In some embodiments, a double-stranded ligase (such as T4 DNA ligase) can be used to ligate a second-stranded synthetic helper probe (a portion of the double-stranded DNA molecule capable of hybridizing to the 3' end of the extension capture probe) to the 3' end of the extension probe, such as a first-stranded cDNA molecule. Other enzymes suitable for the ligation step are known in the art, including, for example, Tth DNA ligase, Taq DNA ligase, and Thermococcus sp. (strain 9°N) DNA ligase (9°N). TMDNA ligase, New England Biolabs, Ampligase TM (Purchased from Lucigen, Middleton, Wisconsin) and SplintR (Purchased from New England Biolabs, Ipswich, Massachusetts). In some embodiments, a polynucleotide tail, such as a poly(A) tail, is incorporated into the 3' end of the extended probe molecule. In some embodiments, a terminal transferase-active enzyme is used to incorporate the polynucleotide tail.
[0390] In some embodiments, the double-stranded extended capture probe is treated to remove any unextended capture probes prior to amplification and / or analysis, such as sequence analysis. This can be achieved through various methods, such as enzymatic degradation of unextended probes, like exonucleases, or purification columns.
[0391] In some embodiments, the extended capture probe is amplified to produce an amount sufficient for analysis, for example, by DNA sequencing. In some embodiments, the first strand of the extended capture probe (e.g., DNA and / or cDNA molecules) is used as a template for an amplification reaction (e.g., polymerase chain reaction).
[0392] In some embodiments, the amplification reaction uses primers comprising an affinity group to incorporate the affinity group into an extension capture probe (e.g., an RNA-cDNA hybrid). In some embodiments, the primers comprise an affinity group, and the extension capture probe comprises an affinity group. The affinity group may correspond to any of the aforementioned affinity groups.
[0393] In some embodiments, an extension capture probe including an affinity group can be coupled to an affinity group-specific array feature. In some embodiments, the substrate may include an antibody or an antibody fragment. In some embodiments, the array feature includes avidin or streptoavidin, and the affinity group includes biotin. In some embodiments, the array feature includes maltose, and the affinity group includes maltose-binding protein. In some embodiments, the array feature includes maltose-binding protein, and the affinity group includes maltose. In some embodiments, an amplified extension capture probe can be used to release the extended probe from the array feature, provided that a copy of the extended probe is not attached to the array feature.
[0394] In some embodiments, an extended capture probe or its complement or amplicon is released from the array feature. The step of releasing the extended capture probe or its complement or amplicon from the array feature can be implemented in a variety of ways. In some embodiments, the extended capture probe or its complement is released from the feature by nucleic acid cleavage and / or denaturation (e.g., by heating to denature double-stranded molecules).
[0395] In some embodiments, an extension capture probe or its complement or amplicon is released from an array feature by physical methods. For example, methods for inducing physical release include denaturing double-stranded nucleic acid molecules. Another method for releasing the extension capture probe is using a solution that interferes with the hydrogen bonds of the double-stranded molecules. In some embodiments, the extension capture probe is released by applying hot water (such as water or a buffer) at at least 85°C, such as at least 90°C, 91°C, 92°C, 93°C, 94°C, 95°C, 96°C, 97°C, 98°C, or 99°C. In some embodiments, a solution comprising salts, surfactants, etc., is added to release the extension capture probe from the array feature, which may further destabilize the interactions between nucleic acid molecules. In some embodiments, a formamide solution may be used to destabilize the interactions between nucleic acid molecules to release the extension capture probe from the array feature.
[0396] (viii) Amplification of the capture probe
[0397] In some embodiments, this document provides a method for amplifying a capture probe attached to a spatial array, wherein the amplification of the capture probe increases the number of capture domains and spatial barcodes on the spatial array. In some embodiments, amplification of the capture probe is performed by rolling circle amplification. In some embodiments, the capture probe to be amplified includes a sequence capable of rolling circle amplification (e.g., a docking sequence, a functional sequence, and / or a primer sequence). In one instance, the capture probe may include a functional sequence capable of binding to a primer used for amplification. In another instance, the capture probe may include one or more docking sequences (e.g., a first docking sequence and a second docking sequence) capable of hybridizing with one or more oligonucleotides (e.g., padlock probes) used for rolling circle amplification. In some embodiments, additional probes are immobilized to a substrate, wherein the additional probes include sequences capable of rolling circle amplification (e.g., docking sequences, functional sequences, and / or primer sequences). In some embodiments, the spatial array is contacted with the oligonucleotide (e.g., the padlock probe). As used herein, a "padlock probe" refers to an oligonucleotide having sequences at its 5' and 3' ends complementary to adjacent or nearby target sequences (e.g., docking sequences) on the capture probe. After hybridization with the target sequence (e.g., the docking sequence), the two ends of the padlock probe are brought into contact or extended until they meet, thereby allowing the padlock probe to be circularized by ligation (e.g., ligation using any of the methods described herein). In some embodiments, after oligonucleotide circularization, rolling circle amplification can be used to amplify the ligation product, which includes at least a capture domain and a spatial barcode from the capture probe. In some embodiments, using padlock oligonucleotides and rolling circle amplification to amplify the capture probe increases the number of capture domains and spatial barcodes on the spatial array.
[0398] In some embodiments, a method for improving the capture efficiency of a spatial array includes amplifying all or part of a capture probe fixed to a substrate. For example, amplifying all or part of a capture probe fixed to a substrate can increase the capture efficiency of a spatial array by increasing the number of capture domains and spatial barcodes. In some embodiments, a method for determining the location of an analyte in a biological sample includes using a spatial array with increased capture efficiency (e.g., a spatial array in which capture probes have been amplified as described herein). For example, the capture efficiency of a spatial array can be improved by amplifying all or part of the capture probes before contact with the biological sample. Amplification results in an increased number of capture domains compared to a spatial array in which capture probes are not amplified before contact with the biological sample, thereby enabling the capture of more analytes. In some embodiments, a method for generating a spatial array with increased capture efficiency includes amplifying all or part of the capture probes. In some embodiments, when a spatial array with increased capture efficiency is generated by amplifying all or part of the capture probes, amplification increases the number of capture domains and spatial barcodes on the spatial array. In some embodiments, a method for determining the location of a capture probe (e.g., a capture probe on a feature) on a spatial array includes amplifying all or part of the capture probes. For example, amplifying capture probes immobilized on a substrate can increase the number of spatial barcodes at the location of the capture probe for direct decoding (e.g., direct decoding using any of the methods described herein, including but not limited to in situ sequencing).
[0399] (ix) Analyte trapping agent
[0400] This disclosure also provides methods and materials for spatial analysis of bioanalytes (e.g., mRNA, genomic DNA, accessible chromatin, and cell surface or intracellular proteins and / or metabolites) using analyte trapping agents. As used herein, an analyte trapping agent (previously sometimes also referred to as a "cell marker") is a reagent that interacts with an analyte (e.g., an analyte in a sample) and a trapping probe (e.g., a trapping probe attached to a substrate) to identify the analyte. In some embodiments, the analyte trapping agent includes an analyte binding portion and a trapping agent barcode field.
[0401] Figure 40This is a schematic diagram of an exemplary analyte capture agent 4002 for capturing an analyte. The analyte capture agent includes an analyte binding portion 4004 and a capture agent barcode domain 4008. The analyte binding portion 4004 is a molecule capable of binding analyte 4006 and interacting with a spatial barcode capture probe. The analyte binding portion can bind analyte 4006 with high affinity and / or high specificity. The analyte capture agent 4002 may include the capture agent barcode domain 4008 and a nucleotide sequence (e.g., an oligonucleotide) capable of hybridizing with at least a portion or all of the capture domain of the capture probe. The analyte binding portion 4004 may include a polypeptide and / or an aptamer (e.g., an oligonucleotide or peptide molecule that binds to a specific target analyte). The analyte binding portion 4004 may include an antibody or antibody fragment (e.g., an antigen-binding fragment).
[0402] As used herein, the term "analyte-binding moiety" refers to a molecule or portion capable of binding a macromolecular component (e.g., an analyte, such as a bioanalyte). In some embodiments of any spatial analysis method described herein, the analyte-binding moiety 4004 of the analyte trap 4002 binding to the bioanalyte 4006 may include, but is not limited to, antibodies or epitope-binding fragments thereof, cell surface receptor-binding molecules, receptor ligands, small molecules, bispecific antibodies, bispecific T-cell conjugates, T-cell receptor conjugates, B-cell receptor conjugates, precursors, aptamers, monomers, affinity molecules, darpins, and protein scaffolds, or any combination thereof. The analyte-binding moiety 4004 may bind macromolecular components (e.g., analytes) with high affinity and / or high specificity. The analyte-binding moiety 4004 may include nucleotide sequences (e.g., oligonucleotides) that may correspond to at least a portion or all of the analyte-binding moiety. The analyte-binding moiety 4004 may include peptides and / or aptamers (e.g., peptides and / or aptamers that bind to a specific target molecule (e.g., an analyte). The analyte binding portion 4004 may include an antibody or antibody fragment (e.g., an antigen-binding fragment) that binds to a specific analyte (e.g., a polypeptide).
[0403] In some embodiments, the analyte binding portion 4004 of the analyte trap 4002 comprises one or more antibodies or antigen-binding fragments thereof. The antibody or antigen-binding fragment comprising the analyte binding portion 4004 can specifically bind to a target analyte. In some embodiments, the analyte 4006 is a protein (e.g., a protein on the surface of a biological sample, such as a cell, or an intracellular protein). In some embodiments, multiple analyte traps comprising multiple analyte binding portions bind multiple analytes present in a biological sample. In some embodiments, the multiple analytes comprise a single type of analyte (e.g., a single type of polypeptide). In some embodiments, where the multiple analytes comprise a single type of analyte, the analyte binding portions of the multiple analyte traps are identical. In some embodiments, where the multiple analytes comprise a single type of analyte, the analyte binding portions of the multiple analyte traps are different (e.g., members of the multiple analyte traps may have two or more analyte binding portions, each of which binds a single type of analyte, e.g., at different binding sites). In some embodiments, the multiple analytes comprise multiple different types of analytes (e.g., multiple different types of polypeptides).
[0404] The analyte trapping agent 4002 may include an analyte binding portion 4004. The analyte binding portion 4004 may be an antibody. Analyte binding moiety 4004 that can be used in analyte capture agent 4002 or exemplary non-limiting antibodies that can be used in the applications disclosed herein include any of the following, including variants thereof: A-ACT, A-AT, ACTH, actin-muscle-specific, actin-smooth muscle (SMA), AE1, AE1 / AE3, AE3, AFP, AKT phosphate, ALK-1, amyloid A, androgen receptor, annexin A1, B72.3, BCA-225, BCL-1 (cyclin D1), BCL-1 / CD20, BCL-2, BCL-2 / BCL-6, BCL-6, Ber-EP4, β-amyloid, β-catenin, BG8 (Lewis Y), BOB-1, CA 19.9, CA125, CAIX, calcitonin, calmodulin-binding protein, calmodulin, calreticulin, CAM 5.2, CAM 5.2 / AE1, CD1a, CD2, CD3(M), CD3(P), CD3 / CD20, CD4, CD5, CD7, CD8, CD10, CD14, CD15, CD20, CD21, CD22, CD23, CD25, CD30, CD31, CD33, CD34, CD35, CD43, CD45(LCA), CD45RA, CD56, CD57, CD61, CD68, CD71, CD74, CD79a, CD99, CD117(c-KIT), CD123, CD138, CD163, CDX-2, CDX-2 / CK-7, CEA(M), CEA(P), Chromogranin A, Chrominase, CK-5, CK-5 / 6, CK-7, CK-7 / TTF-1, CK-14, CK-17, CK-18, CK -19, CK-20, CK-HMW, CK-LMW, CMV-IH, COLL-IV, COX-2, D2-40, DBA44, myodermal dermcin, DOG1, EBER-ISH, EBV (LMP1), E-cadherin, EGFR, EMA, ER, ERCC1, Factor VIII (vWF), Factor XIIIa, myofascitis, FLI-1, FHS, galactoprotein-3, gastrin, GCDFP-15, GFAP, glucagon, blood group glycoprotein A, phosphatidylinositol glycan-3, granzyme B, growth hormone (GH), GST, HAM 56. HMBE-1, HBP, HCAg, HCG, Hemoglobin A, HEP B CORE (HBcAg), HEP B SURF (HBsAg), HepPar1, HER2, Herpes I, Herpes II, HHV-8, HLA-DR, HMB45. HPL, HPV-IHC, HPV(6 / 11)-ISH, HPV(16 / 18)-ISH, HPV(31 / 33)-ISH, HPV WSS-ISH, High HPV-ISH, Low HPV-ISH, High & Low HPV-ISH, IgA, IgD, IgG, IgG4, IgM, Inhibin, Insulin, JC Virus-ISH, Kappa-ISH, KER PAN, Ki-67, λ-IHC, λ-ISH, LH, lipase, lysozyme (MURA), mammary globin, MART-1, MBP, M-cell trypsin, MEL-5, Melan-A, Melan-A / Ki-67, mesothelin, MiTF, MLH-1, MOC-31, MPO, MSH-2, MSH-6, MUC1, MUC2, MUC4, MUC5AC, MUM-1, MYO D1, myopoietin, myoglobin, myoin heavy chain, neoaspartic protease A, NB84a, NEW-N, NF, NK1-C3, NPM, NSE, OCT-2, OCT-3 / 4, OSCAR, p16, p21, p27 / Kip1, p53, p57, p63, p120, P504S, pan-melanoma, PANC.POLY, parvovirus B19, PAX-2, PAX-5, PAX-5 / CD43, PAX=5 / CD5, PAX-8, PC, PD1, perforin, PGP 9.5, PLAP, PMS-2, PR, prolactin, PSA, PSAP, PSMA, PTEN, PTH, PTS, RB, RCC, S6, S100, serotonin, somatostatin, surfactant (SP-A), synaptic proteins, synuclein, TAU, TCL-1, TCRβ, TdT, coagulation regulatory protein, thyroglobulin, TIA-1, TOXO, TRAP, TriView TM Breast, TriView TM Prostate, trypsin, TS, TSH, TTF-1, tyrosinase, ubiquitin, Uroplakin, VEGF, chorionic villi, vimentin (VIM), VIP, VZV, WT1(M)N-terminus, WT1(P)C-terminus, and ZAP-70.
[0405] Furthermore, exemplary non-limiting antibodies that can be used as the analyte binding portion 4004 in the analyte trapping agent 4002 or in the applications disclosed herein include any of the following antibodies (and their variants): cell surface proteins, intracellular proteins, kinases (e.g., the AGC kinase family, such as AKT1, AKT2, PDK1, protein kinase C, ROCK1, ROCK2, SGK3), the CAMK kinase family (e.g., AMPK1, AMPK2, CAMK, Chk1, Chk2, Zip), and the CK1 kinase family. TK kinase family (e.g., Abl2, AXL, CD167, CD246 / ALK, c-Met, CSK, c-Src, EGFR, ErbB2 (HER2 / neu), ErbB3, ErbB4, FAK, Fyn, LCK, Lyn, PKT7, Syk, Zap70), STE kinase family (e.g., ASK1, MAPK, MEK1, MEK2, MEK3, MEK4, MEK5, PAK1, PAK2, PAK4, PAK6), CMGC Kinase families (e.g., Cdk2, Cdk4, Cdk5, Cdk6, Cdk7, Cdk9, Erk1, GSK3, Jnk / MAPK8, Jnk2 / MAPK9, JNK3 / MAPK10, p38 / MAPK) and TKL kinase families (e.g., ALK1, ILK1, IRAK1, IRAK2, IRAK3, IRAK4, LIMK1, LIMK2, M3K11, RAF1, RIP1, RIP3, VEGFR1, VEGFR2, VEGF) R3), Aurora kinase A, Aurora kinase B, IKK, Nemo-like kinase, PINK, PLK3, ULK2, WEE1, transcription factors (e.g., FOXP3, ATF3, BACH1, EGR, ELF3, FOXA1, FOXA2, FOX01, GATA), growth factor receptors, and tumor inhibitors (e.g., anti-p53, anti-BLM, anti-Cdk2, anti-Chk2, anti-BRCA-1, anti-NBS1, anti-BRCA-2, anti-WRN, anti-PTEN, anti-WT1, anti-p38).
[0406] In some embodiments, analyte trapping agent 4002 is capable of binding analyte 4006 present within the cell. In some embodiments, the analyte trapping agent is capable of binding cell surface analytes, which may include, but are not limited to, receptors, antigens, surface proteins, transmembrane proteins, differentiation protein clusters, protein channels, protein pumps, carrier proteins, phospholipids, glycoproteins, glycolipids, cell-cell interaction protein complexes, antigen-presenting complexes, major histocompatibility complexes, engineered T cell receptors, T cell receptors, B cell receptors, chimeric antigen receptors, extracellular matrix proteins, post-translational modifications (e.g., phosphorylation, glycosylation, ubiquitination, nitrosation, methylation, acetylation, or lipidation) states, gap junctions, and adhesion junctions of cell surface proteins. In some embodiments, analyte trapping agent 4002 is capable of binding post-translational modified cell surface analytes. In such embodiments, the analyte trapping agent can be specific to cell surface analytes based on a given post-translational modification state (e.g., phosphorylation, glycosylation, ubiquitination, nitrosation, methylation, acetylation, or lipidation), such that the cell surface analyte map can include post-translational modification information of one or more analytes.
[0407] In some embodiments, the analyte capture agent 4002 includes a capture agent barcode field 4008 conjugated to or otherwise attached to the analyte binding portion. In some embodiments, the capture agent barcode field 4008 is covalently linked to the analyte binding portion 4004. In some embodiments, the capture agent barcode field 4008 is a nucleic acid sequence. In some embodiments, the capture agent barcode field 4008 includes an analyte binding portion barcode and an analyte capture sequence 4114, or covalently bound thereto.
[0408] As used herein, the term "analyte binding portion barcode" refers to a barcode associated with or otherwise identifying the analyte binding portion 4004. In some embodiments, the analyte 4006 bound to the analyte binding portion 4004 may also be identified by identifying the analyte binding portion 4004 and its associated analyte binding portion barcode. The analyte binding portion barcode may be a nucleic acid sequence of a given length and / or a sequence associated with the analyte binding portion 4004. The analyte binding portion barcode may generally include any of the various aspects of barcodes described herein. For example, an analyte capture agent 4002 specific to one type of analyte may have a first capture agent barcode field associated with it (e.g., which includes a first analyte binding portion barcode), while an analyte capture agent specific to different analytes may have a different capture agent barcode field associated with it (e.g., which includes a second barcode analyte binding portion barcode). In some aspects, this capture agent barcode field may include an analyte binding portion barcode, which allows identification of the analyte binding portion 4004 associated with the capture agent barcode field. The selection of the capture agent barcode field 4008 can allow for significant diversity in sequence, while also being easy to attach to most analyte binding portions (e.g., antibodies or aptamers) and easy to detect (e.g., using sequencing or array technologies).
[0409] In some embodiments, the capture agent barcode field of the analyte capture agent 4002 includes an analyte capture sequence. As used herein, the term "analyte capture sequence" refers to a region or portion configured to hybridize, bind, conjugate, or otherwise interact with a capture domain of a capture probe. In some embodiments, the analyte capture sequence includes a nucleic acid sequence complementary or substantially complementary to the capture domain of a capture probe, such that the analyte capture sequence hybridizes with the capture domain of the capture probe. In some embodiments, the analyte capture sequence includes a poly(A) nucleic acid sequence hybridizing with a capture domain containing a poly(T) nucleic acid sequence. In some embodiments, the analyte capture sequence includes a poly(T) nucleic acid sequence hybridizing with a capture domain containing a poly(A) nucleic acid sequence. In some embodiments, the analyte capture sequence includes a heterogeneous nucleic acid sequence hybridizing with a capture domain containing a heterogeneous nucleic acid sequence complementary (or substantially complementary) to the heterogeneous nucleic acid sequence of the analyte capture region.
[0410] In some embodiments of any spatial analysis method employing analyte trap 4002 described herein, trap barcode domains may be directly coupled to the analyte binding portion 4004, or they may be attached to beads, molecular lattices, such as linear, spherical, cross-linked, or other polymers, or attached to or otherwise associated with other frames of the analyte binding portion, allowing multiple trap barcode domains to be attached to a single analyte binding portion. The attachment (coupling) of the trap barcode domains to the analyte binding portion 4004 can be achieved through any of a variety of direct or indirect, covalent or non-covalent binding or attachment methods. For example, in cases where a trap barcode domain is coupled to an analyte binding portion 4004 comprising an antibody or antigen-binding fragment, such trap barcode domains may be used using chemical conjugation techniques (e.g., LIGHTNING-C, available from Innova Biosciences). The capture agent barcode domain is covalently attached to a portion of the antibody or antigen-binding fragment. In some embodiments, a non-covalent attachment mechanism can be used to couple the capture agent barcode domain to the antibody or antigen-binding fragment (e.g., using biotinylated antibodies and oligonucleotides or beads comprising one or more biotinylated linkers, coupled to oligonucleotides with avidin or streptavidin linkers). Antibody and oligonucleotide biotinylation techniques can be used and are described, for example, in Fang et al., 2003, Nucleic Acids Res. 31(2):708-715, the entire contents of which are incorporated herein by reference. Similarly, protein and peptide biotinylation techniques have been developed and are available for use and are described, for example, in U.S. Patent No. 6,265,552, the entire contents of which are incorporated herein by reference. Furthermore, click reaction chemistry such as the methyltetraazine-PEG5-NHS ester reaction and the TCO-PEG4-NHS ester reaction can be used to couple the capture agent barcode domain to the analyte binding portion 4004. The reactive portion on the analyte binding moiety may also include an amine for targeting aldehydes, an amine for targeting maleimides (e.g., free thiols), an azide for targeting click chemistry compounds (e.g., alkynes), biotin for targeting streptavidin, or a phosphate ester for targeting EDC, which in turn targets the active ester (e.g., NH2). The reactive portion on the analyte binding moiety 4004 may be a chemical compound or group bound to the reactive portion. Exemplary strategies for conjugating the analyte binding moiety 4004 to the capture agent barcode domain include conjugation using commercial kits (e.g., Solulink, Thunderlink), mild reduction of the hinge region and maleimide labeling, click chemistry reactions facilitated by staining of labeled amides (e.g., copper-free), and conjugation by periodate oxidation of sugar chains and amine conjugation. In the case where the analyte binding moiety 4004 is an antibody, the antibody may be modified before or simultaneously with oligonucleotide conjugation. For example, antibodies can be glycosylated using a β-1,4-galactosyltransferase-based chemical substrate allowing the mutant GalT (Y289L) and the azide-containing uridine diphosphate-N-acetylgalactosamine analog uridine diphosphate-GalNAz. The modified antibody can be conjugated to an oligonucleotide having a dibenzocyclooctylene-PEG4-NHS group. In some embodiments, certain steps (e.g., COOH activation (such as EDC) and homobifunctional cross-linking agents) can be avoided to prevent the analyte-binding moiety from conjugating to itself.In some embodiments of any spatial analysis method described herein, the analyte trap (e.g., the analyte binding portion 4004 coupled to an oligonucleotide) may be delivered to the cell, for example, by transfection (e.g., using transfected amines, cationic polymers, calcium phosphate, or electroporation), by transduction (e.g., using phages or recombinant viral vectors), by mechanical delivery (e.g., magnetic beads), by lipids (e.g., 1,2-dioleoyl-sn-glycero-3-phosphocholine (DOPC)), or by transport proteins.
[0411] The analyte trap 4002 can be delivered into cells using exosomes. For example, a first cell containing the analyte trap can be generated. The analyte trap can attach to the exosome membrane. The analyte trap can be contained in the cytosol of the exosome. The released exosome can be harvested and provided to a second cell, thereby delivering the analyte trap into the second cell. The analyte trap can be released from the exosome membrane before, during, or after delivery to the cell. In some embodiments, the cell is permeabilized to allow the analyte trap 4002 to couple with intracellular components, such as, but not limited to, intracellular proteins, metabolites, and nuclear membrane proteins. Following intracellular delivery, the analyte trap 4002 can be used to analyze the intracellular components described herein.
[0412] In some embodiments of any of the spatial analysis methods described herein, the capture agent barcode domain coupled to analyte capture agent 4002 may include modifications that prevent it from being extended by polymerase. In some embodiments, the capture agent barcode domain may be used as a template instead of a primer when bound to a capture domain of a capture probe or nucleic acid in a sample for primer extension. When the capture agent barcode domain also includes a barcode (e.g., an analyte binding partial barcode), such a design can increase the efficiency of molecular barcoding by increasing the affinity between the capture agent barcode domain and the unbarcoded sample nucleic acid, and eliminate the potential formation of linker artifacts. In some embodiments, the capture agent barcode domain 4008 may include a random N-mer sequence that is modified to cap and prevent it from being extended by polymerase. In some cases, the composition of the random N-mer sequence may be designed to maximize binding efficiency with free, unbarcoded ssDNA molecules. This design may include a random sequence composition with a high GC content, a partially random sequence with fixed G or C at a specific position, the use of guanosine, the use of locked nucleic acids, or any combination thereof.
[0413] Modifications used to block primer extension by polymerase can be carbon spacer groups of varying lengths or dideoxynucleotides. In some embodiments, the modification can be a debasement site of an analogue of a depurinyl or depyrimidine structure, a base analogue, or a phosphate backbone, such as an N-(2-aminoethyl)-glycine backbone linked by an amide bond, tetrahydrofuran, or 1',2'-dideoxyribose. Modifications can also be uracil bases, 2'OMe-modified RNA, C3-18 spacers (e.g., structures with 3 to 18 consecutive carbon atoms, such as C3 spacers), ethylene glycol polymeric spacers (e.g., spacer 18 (hexaethylene glycol spacer)), biotin, dideoxynucleotide triphosphates, ethylene glycol, amines, or phosphate esters.
[0414] In some embodiments of any of the spatial distribution methods described herein, the capture agent barcode domain 4008 coupled to the analyte binding portion 4004 includes a cleavable domain. For example, after the analyte capture agent binds to the analyte (e.g., a cell surface analyte), the capture agent barcode domain can be cleaved and collected for downstream analysis according to the methods described herein. In some embodiments, the cleavable domain of the capture agent barcode domain includes a U-removal element that allows the species to be released from the bead. In some embodiments, the U-removal element may include a single-stranded DNA (ssDNA) sequence containing at least one uracil. The species can attach to the bead via the ssDNA sequence. The species can be released via a combination of a uracil-DNA glycosylase (e.g., to remove uracil) and a nuclease (e.g., to induce ssDNA breakage). If the nuclease generates a 5' phosphate group from the cleavage, additional enzymatic treatment can be included in downstream processing to remove the phosphate group, for example, before ligating additional sequencing processing elements (e.g., enominar full P5 sequence, partial P5 sequence, full R1 sequence, and / or partial R1 sequence).
[0415] In some embodiments, multiple different types of analytes (e.g., peptides) from a biological sample can subsequently be associated with one or more physical properties of the biological sample. For example, multiple different types of analytes can be associated with the location of the analytes within the biological sample. Such information (e.g., proteomic information when the analyte-binding moiety recognizes the peptide) can be used in conjunction with other spatial information (e.g., genetic information from the biological sample, such as DNA sequence information, transcriptomic information, such as transcript sequences, or both). For example, cell surface proteins can be associated with one or more physical properties of the cell (e.g., cell shape, size, activity, or type). One or more physical properties can be characterized by imaging the cell. The cell can be bound by an analyte trapping agent comprising an analyte-binding moiety bound to a cell surface protein and an analyte-binding moiety barcode identifying the analyte-binding moiety, and the cell can undergo spatial analysis (e.g., any of the various spatial analysis methods described herein). For example, an analyte capture agent 4002 bound to a cell surface protein can bind to a capture probe (e.g., a capture probe on an array) comprising a capture domain that interacts with an analyte capture sequence present on the capture agent barcode domain of the analyte capture agent 902. All or part of the capture agent barcode domain (including the analyte-binding portion barcode) can be replicated with a polymerase using the 3' end of the capture domain as a trigger site to generate an extended capture probe comprising copies of all or part of the complementary sequence to the capture probe (including the spatial barcode present on the capture probe) and the analyte-binding portion barcode. In some embodiments, an analyte capture agent having an extended capture agent barcode domain comprising a sequence complementary to the spatial barcode of the capture probe is referred to as a "spatially labeled analyte capture agent".
[0416] In some embodiments, a spatial array of spatially labeled analyte traps can contact a sample, wherein the analyte traps bound to the spatial array trap the target analyte. The analyte trap can then be denatured from the trap probes of the spatial array, comprising an extended trap probe including a sequence complementary to the spatial barcode of the trap probe and the analyte-binding portion barcode. This allows for the reuse of the spatial array. The sample can be dissociated into non-aggregated cells (e.g., single cells) and analyzed using the single-cell / droplet method described herein. The spatially labeled analyte trap can be sequenced to obtain nucleic acid sequences of the spatial barcode of the trap probe and the analyte-binding portion barcode of the analyte trap. Thus, the nucleic acid sequence of the extended trap probe can be associated with an analyte (e.g., a cell surface protein) and, consequently, with one or more physical properties of the cell (e.g., shape or cell type). In some embodiments, the nucleic acid sequence of the extended trap probe can be associated with an intracellular analyte in a nearby cell, wherein the intracellular analyte is released using any of the cell permeation or analyte migration techniques described herein.
[0417] In some embodiments of any of the spatial analysis methods described herein, the capture agent barcode domains released from the analyte capture agent can then be sequenced to identify which analyte capture agents bind to the analyte. Based on the presence of capture agent barcode domains and analyte binding portion barcode sequences associated with capture points on the spatial array (e.g., capture points at specific locations), analyte maps can be created for biological samples. Maps of individual cells or cell populations can be compared to maps from other cells (e.g., “normal” cells) to identify changes in the analyte that can provide diagnostically relevant information. In some embodiments, these maps can be used to diagnose a variety of diseases characterized by changes in cell surface receptors, such as cancer and other conditions.
[0418] Figure 41A The figure above is a schematic diagram depicting an exemplary interaction between a feature-fixed capture probe 602 and an analyte capture agent 4002 (where the terms "feature" and "capture point" are used interchangeably). As described elsewhere herein, the feature-fixed capture probe 602 may include a spatial barcode 605 and one or more functional sequences 604 and 606. The capture probe 602 may also include a capture domain 607 capable of binding the analyte capture agent 4002. In some embodiments, the analyte capture agent 4002 includes a functional sequence 4118, a capture agent barcode domain 4008, and an analyte capture sequence 4114. In some embodiments, the analyte capture sequence 4114 is capable of binding the capture domain 607 of the capture probe 602. The analyte capture agent 4002 may also include a blocking probe 4120 that allows the capture agent barcode domain 4008 (4114 / 4008 / 4118) to couple with the analyte binding portion 4004.
[0419] Figure 41A The figure below further illustrates the spatially labeled analyte trap 4002, wherein the analyte trap sequence 4114 (poly-A sequence) of the trap barcode domain 4118 / 4008 / 4114 can be blocked by a blocking probe (poly-T oligonucleotide).
[0420] In some embodiments, the capture binding domain may include a sequence that is at least partially complementary to the sequence of the capture domain of the capture probe (e.g., any of the exemplary capture domains described herein). Figure 41B An exemplary capture-binding domain is shown attached to an analyte-binding moiety used for detecting proteins in biological samples. For example... Figure 41B As shown, the analyte binding portion 4004 comprises an oligonucleotide including a primer (e.g., readout 2) sequence 4118, an analyte binding portion barcode 4008, a capture-binding domain (e.g., exemplary polyA) having a first sequence (e.g., a capture-binding domain) 4114, and a blocking probe or a second sequence 4120 (e.g., polyT or polyU), wherein the blocking sequence blocks hybridization of the capture-binding domain with a capture domain on the capture probe. In some cases, the blocking probe 4120 is referred to as a blocking probe as disclosed herein. In some cases, the blocking probe is... Figure 41B The example shown is a poly-T sequence.
[0421] In some cases, such as Figure 41A As shown, the blocking probe sequence is not located on a continuous sequence having a capture-binding domain. In other words, in some cases, the capture-binding domain (also referred to herein as the first sequence) and the blocking sequence are independent polynucleotides. In some cases, it will be apparent to those skilled in the art that the terms "capture-binding domain" and "first sequence" are used interchangeably in this disclosure.
[0422] In a non-limiting example, when the capture domain sequence of the capture probe on the substrate is a poly(T) sequence, the first sequence may be a poly(A) sequence. In some embodiments, the capture binding domain includes a capture binding domain that is substantially complementary to the capture domain of the capture probe. Substantially complementary means that the first sequence of the capture binding domain is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to the sequence in the capture domain of the capture probe, which is also a random sequence. In another example, the first sequence of the capture binding domain may be a random sequence (e.g., a random hexamer) that is at least partially complementary to the capture domain sequence of the capture probe, which is also a random sequence. In another example, when the capture domain sequence of the capture probe is also a sequence comprising both homopolymeric sequences (e.g., poly(A) sequences) and random sequences, the capture binding domain may be a mixture of homopolymeric sequences (e.g., poly(T) sequences) and random sequences (e.g., random hexamers). In some embodiments, the capture binding domain comprises ribonucleotides, deoxyribonucleotides, and / or synthetic nucleotides capable of participating in Wattssen-Crick-type or similar base pair interactions. In some embodiments, the first sequence of the capture binding domain comprises at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, or at least 24 nucleotides. In some embodiments, the first sequence of the capture binding domain comprises at least 25 nucleotides, at least 30 nucleotides, or at least 35 nucleotides.
[0423] In some embodiments, the capture-binding domain (e.g., a first sequence) and the blocking probe for the capture-binding domain (e.g., a second sequence) are located on the same contiguous nucleic acid sequence. When the capture-binding domain and the blocking probe are located on the same contiguous nucleic acid sequence, the second sequence (e.g., the blocking probe) is located at the 3' of the first sequence. When the first sequence and the second sequence (e.g., the blocking probe) for the capture-binding domain are located on the same contiguous nucleic acid sequence, the second sequence (e.g., the blocking probe) is located at the 5' of the first sequence. As used herein, the terms second sequence and blocking probe are used interchangeably.
[0424] In some cases, the second sequence of the capture-binding domain (e.g., a blocking probe) comprises a nucleic acid sequence. In some cases, the second sequence is also referred to as a blocking probe or a blocking domain, and each term is used interchangeably. In some cases, the blocking domain is a DNA oligonucleotide. In some cases, the blocking domain is an RNA oligonucleotide. In some embodiments, the blocking probe of the capture-binding domain comprises a sequence complementary to or substantially complementary to the first sequence of the capture-binding domain. In some embodiments, the blocking probe, when present, prevents the first sequence of the capture-binding domain from binding to the capture domain of the capture probe. In some embodiments, the blocking probe is removed before the first sequence of the capture-binding domain (e.g., present in a linked probe) binds to the capture domain on the capture probe. In some embodiments, the blocking probe of the capture-binding domain comprises a polyuridine sequence, a polythymidine sequence, or both. In some cases, the blocking probe (or second sequence) is part of a hairpin structure that specifically binds to the capture-binding domain and prevents the capture-binding domain from hybridizing with the capture domain of the capture probe. See, for example, Figure 41C .
[0425] In some embodiments, the second sequence of the capture-binding domain (e.g., a blocking probe) includes a sequence configured to hybridize with the first sequence of the capture-binding domain. When the blocking probe hybridizes with the first sequence, the blocking first sequence hybridizes with the capture domain of the capture probe. In some embodiments, the blocking probe includes a sequence complementary to the first sequence. In some embodiments, the blocking probe includes a sequence substantially complementary to the first sequence. In some embodiments, the blocking probe includes a sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to the first sequence of the capture-binding domain.
[0426] In some embodiments, the blocking probe for capturing the binding domain includes a homopolymeric sequence substantially complementary to a first sequence of the capturing binding domain. In some embodiments, the blocking probe is configured to hybridize with a poly(A), poly(T), or poly-rU sequence. In some embodiments, the blocking probe includes a poly(A), poly(T), or poly(U) sequence. In some embodiments, the first sequence includes a homopolymeric sequence. In some embodiments, the first sequence includes a poly(A), poly(U), or poly(T) sequence.
[0427] In some embodiments, the capture binding domain further includes a hairpin sequence (such as...) Figure 41C (As shown). Figure 41C An exemplary capture-binding domain is shown attached to an analyte-binding moiety used for detecting proteins in biological samples. For example... Figure 41CAs shown, the analyte binding portion 4004 comprises an oligonucleotide including a primer (e.g., readout 2) sequence 4118, an analyte binding portion barcode 4008, a capture-binding domain (e.g., exemplary polyA) having a first sequence 4114, a blocking probe 4120, and a third sequence 4140, wherein the second and / or third sequences can be polyT or polyU or combinations thereof, wherein the blocking probe produces a hairpin structure, and the third sequence blocks hybridization of the first sequence with the capture domain on the capture probe. In some cases, the third sequence 4140 is referred to as the blocking sequence. Furthermore, 4150 exemplifies a nuclease capable of digesting the blocking sequence. In this example, 4150 can be a nuclease or mixture of nucleases capable of digesting uracil, such as UDG, or a uracil-specific excision mixture, such as USER (NEB).
[0428] Another embodiment of the hairpin blocking agent solution is in Figure 41D Examples are provided. For instance... Figure 41D As shown, the analyte binding portion 4004 comprises an oligonucleotide including a primer (e.g., readout 2) sequence 4118, an analyte binding portion barcode 4008, a capture-binding domain (e.g., exemplary polyA) having a first sequence (e.g., a capture-binding domain) 4114, a second hairpin sequence 4170, and a third sequence 4180, wherein the third sequence (e.g., a blocking probe) blocks hybridization of the first sequence with the capture domain on the capture probe. In this example, 4190 illustrates an RNase H nuclease capable of digesting uracil-blocking sequencing from a DNA:RNA hybrid formed by blocking the first sequence with a third sequence containing uracil.
[0429] In some embodiments, the hairpin sequence 4170 is located at 5' of the blocking probe within the capture-binding domain. In some embodiments, the hairpin sequence 4170 is located at 5' of the first sequence within the capture-binding domain. In some embodiments, the capture-binding domain from 5' to 3' includes a first sequence substantially complementary to the capture domain of the capture probe, a hairpin sequence, and a blocking probe substantially complementary to the first sequence. Alternatively, the capture-binding domain from 3' to 5' includes a first sequence substantially complementary to the capture domain of the capture probe, a hairpin sequence, and a blocking probe substantially complementary to the first sequence.
[0430] In some embodiments, the hairpin sequence 4170 comprises a sequence of about three nucleotides, about four nucleotides, about five nucleotides, about six nucleotides, about seven nucleotides, about eight nucleotides, about nine nucleotides, or about ten or more nucleotides. In some cases, the hairpin is at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, or more nucleotides.
[0431] In some embodiments, the hairpin sequence comprises DNA, RNA, a DNA-RNA hybrid, or includes modified nucleotides. In some cases, the hairpin is a poly(U) sequence. In some cases, the RNA hairpin sequence is digested by USER and / or RNAse H using the methods disclosed herein. In some cases, the poly(U) hairpin sequence is digested by USER and / or RNAse H using the methods disclosed herein. In some cases, the hairpin is a poly(T) sequence. It should be understood that the hairpin sequence (whether it comprises DNA, RNA, a DNA-RNA hybrid, or includes modified nucleotides) can be virtually any nucleotide sequence, as long as it forms a hairpin, and in some cases, as long as it is digested by USER and / or RNAse H.
[0432] In some embodiments, the methods provided herein require the release of a second sequence (e.g., a blocking probe) of a capture-binding domain that has hybridized with a first sequence of the capture-binding domain from the first sequence. In some embodiments, the release of the blocking probe (or the second sequence) from the first sequence is performed under conditions where the blocking probe dehybridizes from the first sequence.
[0433] In some embodiments, releasing the blocking probe from the first sequence includes cleaving the hairpin sequence. In some embodiments, the hairpin sequence includes a cleavable adapter. For example, the cleavable adapter may be a photocleavable adapter, a UV-cleavable adapter, or an enzyme-cleavable adapter. In some embodiments, the enzyme that cleaves the enzymatically cleavable domain is a restriction endonuclease. In some embodiments, the hairpin sequence includes a target sequence of a restriction endonuclease.
[0434] In some embodiments, releasing a blocking probe (or a second sequence) of a capture-binding domain that hybridizes to a first sequence of the capture-binding domain includes contacting the blocking probe with a restriction endonuclease. In some embodiments, releasing the blocking probe from the first sequence includes contacting the blocking probe with a ribonuclease. In some embodiments, when the blocking probe is an RNA sequence (e.g., a sequence containing uracil), the ribonuclease is one or more of RNase H, RNase A, RNase C, or RNase I. In some embodiments, the ribonuclease is RNase H. In some embodiments, RNase H includes RNase H1, RNase H2, or RNase H1 and RNase H2.
[0435] In some embodiments, the hairpin sequence comprises a homopolymeric sequence. In some embodiments, the hairpin sequence 4170 comprises a poly(T) or poly(U) sequence. For example, the hairpin sequence comprises a poly(U) sequence. In some embodiments, this document provides a method for releasing a blocking probe by contacting the hairpin sequence with a uracil-specific excision reagent (USER) enzyme.
[0436] In some embodiments, releasing the blocking probe from the first sequence includes denaturing the blocking probe under conditions that allow it to dehybridize from the first sequence. In some embodiments, denaturation includes using chemical or physical denaturation. For example, physical denaturation (e.g., temperature) is used to release the blocking probe. In some embodiments, denaturation includes temperature regulation. For example, the first sequence and the blocking probe have a predetermined annealing temperature based on the known composition (A, G, C, or T) within the sequence. In some embodiments, the temperature is regulated to be at most 5°C, at most 10°C, at most 15°C, at most 20°C, at most 25°C, at most 30°C, or at most 35°C above the predetermined annealing temperature. In some embodiments, the temperature is adjusted to be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35°C above a predetermined annealing temperature.In some embodiments, once the temperature is adjusted to a temperature higher than the predetermined annealing temperature, the annealing is performed at a rate of approximately 0.1°C / second to approximately 1.0°C / second (e.g., approximately 0.1°C / second to approximately 0.9°C / second, approximately 0.1°C / second to approximately 0.8°C / second, approximately 0.1°C / second to approximately 0.7°C / second, approximately 0.1°C / second to approximately 0.6°C / second, approximately 0.1°C / second to approximately 0.5°C / second, approximately 0.1°C / second to approximately 0.4°C / second, approximately 0.1°C / second to approximately 0.3°C / second, approximately 0.1°C / second to approximately 0.2°C / second, approximately 0.2°C / second to approximately 1.0°C / second, approximately...). 0.2℃ / sec to about 0.9℃ / sec, about 0.2℃ / sec to about 0.8℃ / sec, about 0.2℃ / sec to about 0.7℃ / sec, about 0.2℃ / sec to about 0.6℃ / sec, about 0.2℃ / sec to about 0.5℃ / sec, about 0.2℃ / sec to about 0.4℃ / sec, about 0.2℃ / sec to about 0.3℃ / sec, about 0.3℃ / sec to about 1.0℃ / sec, about 0.3℃ / sec to about 0.9℃ / sec, about 0.3℃ / sec to about 0.8℃ / sec, about 0.3℃ / sec to about 0.7℃ / sec, about 0.3℃ / sec to about 0.6℃ / sec, about 0 0.3℃ / sec to about 0.5℃ / sec, about 0.3℃ / sec to about 0.4℃ / sec, about 0.4℃ / sec to about 1.0℃ / sec, about 0.4℃ / sec to about 0.9℃ / sec, about 0.4℃ / sec to about 0.8℃ / sec, about 0.4℃ / sec to about 0.7℃ / sec, about 0.4℃ / sec to about 0.6℃ / sec, about 0.4℃ / sec to about 0.5℃ / sec, about 0.5℃ / sec to about 1.0℃ / sec, about 0.5℃ / sec to about 0.9℃ / sec, about 0.5℃ / sec to about 0.8℃ / sec, about 0.5℃ / sec to about 0.7℃ / sec, about 0. The temperature is cooled to a predetermined annealing temperature at heating rates of 5°C / second to about 0.6°C / second, about 0.6°C / second to about 1.0°C / second, about 0.6°C / second to about 0.9°C / second, about 0.6°C / second to about 0.8°C / second, about 0.6°C / second to about 0.7°C / second, about 0.7°C / second to about 1.0°C / second, about 0.7°C / second to about 0.9°C / second, about 0.7°C / second to about 0.8°C / second, about 0.8°C / second to about 1.0°C / second, about 0.8°C / second to about 0.9°C / second, or about 0.9°C / second to about 1.0°C / second. In some embodiments, the denaturation includes temperature cycling. In some embodiments, the denaturation includes alternating between denaturing conditions (e.g., denaturing temperature) and non-denaturing conditions (e.g., annealing temperature).
[0437] It should be understood that, despite any specific function in the embodiments, the hairpin sequence can be any sequence configuration as long as it forms a hairpin. Thus, in some cases, it can be, for example, a degenerate sequence, a random sequence, or others (containing any polynucleotide sequence).
[0438] In some embodiments, the hairpin sequence 4170 further includes a sequence capable of binding to the capture domain of the capture probe. For example, releasing the hairpin sequence from the capture-binding domain may require cleaving the hairpin sequence, wherein the hairpin sequence portion remaining after cleavage includes a sequence capable of binding to the capture domain of the capture probe. In some embodiments, all or part of the hairpin sequence is substantially complementary to the capture domain of the capture probe. In some embodiments, the sequence substantially complementary to the capture domain of the capture probe is located at the free 5' or free 3' end after hairpin sequence cleavage. In some embodiments, hairpin cleavage produces a single-stranded sequence capable of binding to the capture domain of the capture probe on a spatial array. While release of the hairpin sequence may enable hybridization with the capture domain of the capture probe, it is anticipated that release of the hairpin will not significantly affect the capture of the target analyte by the analyte-binding portion or the probe oligonucleotide (e.g., a second probe oligonucleotide).
[0439] In some cases, one or more blocking methods disclosed herein comprise multiple cage-like nucleotides. In some embodiments, methods are provided herein wherein the capture-binding domain comprises multiple cage-like nucleotides. The cage-like nucleotides prevent the capture-binding domain from interacting with the capture domain of the capture probe. The cage-like nucleotides comprise cage-like portions that block Wattssen-Crick hydrogen bonding, thereby preventing interaction until activated, for example, by photolysis of the cage-like portions that release the cage-like portions and restore the ability of the cage-like nucleotides to perform Wattssen-Crick base pairing with complementary nucleotides.
[0440] Figure 41E This demonstrates the use of cage-like nucleotides to block the capture-binding domain. (Example:) Figure 41E As shown, the analyte binding portion 4004 comprises an oligonucleotide including a primer (e.g., readout 2) sequence 4118, an analyte binding portion barcode 4008, and a capture-binding domain (e.g., exemplary polyA) having sequence 4114. A cage-like nucleotide 4130 blocks sequence 4114, thereby blocking the interaction between the capture-binding domain and the capture domain of the capture probe. In some embodiments, the capture-binding domain comprises a plurality of cage-like nucleotides, wherein one of the cage-like nucleotides includes a cage portion capable of preventing interaction between the capture-binding domain and the capture domain of the capture probe. A non-limiting example of a cage-like nucleotide (also known as a photosensitive oligonucleotide) is described in Liu et al., 2014, *Acc. Chem. Res.*, 47(1):45-55 (2014), which is incorporated herein by reference in its entirety. In some embodiments, the cage-like nucleotide includes a cage-like portion selected from the group consisting of 6-nitropiperyloxymethyl (NPOM), 1-(o-nitrophenyl)-ethyl (NPE), 2-(o-nitrophenyl)propyl (NPP), diethylaminocoumarin (DEACM), and nitrodibenzofuran (NDBF).
[0441] In some embodiments, the cage-binding nucleotide comprises a non-naturally occurring nucleotide selected from the group consisting of 6-nitropiperyloxymethyl (NPOM)-cage adenosine, 6-nitropiperyloxymethyl (NPOM)-cage guanosine, 6-nitropiperyloxymethyl (NPOM)-cage uridine, and 6-nitropiperyloxymethyl (NPOM)-cage thymidine. For example, the capture-binding domain comprises one or more cage-binding nucleotides, wherein the cage-binding nucleotides comprise one or more 6-nitropiperyloxymethyl (NPOM)-cage guanosine. In another example, the capture-binding domain comprises one or more cage-binding nucleotides, wherein the cage-binding nucleotides comprise one or more nitropiperyloxymethyl (NPOM)-cage uridine. In yet another example, the capture-binding domain comprises one or more cage-binding nucleotides, wherein the cage-binding nucleotides comprise one or more 6-nitropiperyloxymethyl (NPOM)-cage thymidine.
[0442] In some embodiments, the capture-binding domain comprises a combination of at least two or more cage-like nucleotides described herein. For example, the capture-binding domain may comprise one or more 6-nitropiperyloxymethyl (NPOM)-cage guanosine and one or more nitropiperyloxymethyl (NPOM)-cage uridine. It should be understood that the capture-binding domain may comprise any combination of any cage-like nucleotides described herein.
[0443] In some embodiments, the capture-binding domain includes one, two, three, four, five, six, seven, eight, nine, or ten or more cage nucleotides.
[0444] In some embodiments, the capture-binding domain includes a cage-like nucleotide at its 3' end. In some embodiments, the capture-binding domain includes two cage-like nucleotides at its 3' end. In some embodiments, the capture-binding domain includes at least three cage-like nucleotides at its 3' end.
[0445] In some embodiments, the capture-binding domain includes a cage-like nucleotide at its 5' end. In some embodiments, the capture-binding domain includes two cage-like nucleotides at its 5' end. In some embodiments, the capture-binding domain includes at least three cage-like nucleotides at its 5' end.
[0446] In some embodiments, the capture-binding domain includes a cage-like nucleotide at each odd-numbered position starting from the 3' end of the capture-binding domain. In some embodiments, the capture-binding domain includes a cage-like nucleotide at each odd-numbered position starting from the 5' end of the capture-binding domain. In some embodiments, the capture-binding domain includes a cage-like nucleotide at each even-numbered position starting from the 3' end of the capture-binding domain. In some embodiments, the capture-binding domain includes a cage-like nucleotide at each even-numbered position starting from the 5' end of the capture-binding domain.
[0447] In some embodiments, the capture-binding domain includes a sequence comprising at least 10%, at least 20%, or at least 30% cage-like nucleotides. In some cases, the percentage of cage-like nucleotides in the capture-binding domain is about 40%, about 50%, about 60%, about 70%, about 80%, or higher. In some embodiments, the capture-binding domain includes a sequence wherein each nucleotide is a cage-like nucleotide. It should be understood that the restriction of cage-like nucleotides is based on the sequence of the capture-binding domain and on the spatial restriction that produces cage-like nucleotides adjacent to each other. Thus, in some cases, a specific nucleotide (e.g., guanine) is replaced by a cage-like nucleotide. In some cases, all guanine in the capture-binding domain is replaced by cage-like nucleotides. In some cases, a portion of the guanine in the capture-binding domain (e.g., about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 95%) is replaced by cage-like nucleotides. In some cases, a specific nucleotide (e.g., uridine or thymine) is replaced by a cage-like nucleotide. In some cases, all uridine or thymine in the capture-binding domain is replaced by cage-like nucleotides. In other cases, a portion of the uridine or thymine in the capture-binding domain (e.g., about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 95%) is replaced by cage-like nucleotides. Cage-like nucleotides are disclosed in Govan et al., 2013, Nucleic Acid Research 41; 22, 10518-10528, the entire contents of which are incorporated herein by reference.
[0448] In some embodiments, the capture-binding domain comprises cage-like nucleotides uniformly distributed throughout the capture-binding domain. For example, the capture-binding domain may comprise a sequence comprising at least 10% cage-like nucleotides, wherein the cage-like nucleotides are uniformly distributed throughout the capture-binding domain. In some embodiments, the capture-binding domain comprises a sequence comprising at least 10% cage-like nucleotides, wherein the 10% cage-like nucleotides are located at the 3' end of the capture-binding domain. In some embodiments, the cage-like nucleotides are comprised in every third, fourth, fifth, and sixth nucleotides of the capture-binding domain sequence, or combinations thereof.
[0449] In some embodiments, this document provides a method for releasing a caged portion from a caged nucleotide. In some embodiments, releasing the caged portion from a caged nucleotide includes activating the caged portion. In some embodiments, releasing the caged portion from a caged nucleotide restores the ability of the caged nucleotide to hybridize with a complementary nucleotide via Wattssen-Crick hydrogen bonding. For example, restoring the ability of the caged nucleotide to hybridize with a complementary nucleotide enables / restores the ability of the capture-binding domain to interact with the capture domain. When the caged portion is released from a caged nucleotide, the caged nucleotide is no longer “cage-like” because the caged portion is no longer linked to the caged nucleotide (e.g., covalently or non-covalently). As used herein, the term “cage-like nucleotide” can refer to a nucleotide linked to a caged portion or a nucleotide linked to a caged portion but no longer linked due to activation of the caged portion.
[0450] In some embodiments, this document provides methods for activating cage-like moieties to release cage-like moieties from cage-like nucleotides. In some embodiments, activating the cage-like moieties includes photolytically removing the cage-like moieties from the nucleotides. As used herein, “photolytic” can refer to the process of removing or separating cage-like moieties from cage-like nucleotides using light. In some embodiments, activating (e.g., photolytically removing) the cage-like moieties includes exposing the cage-like moieties to light pulses (e.g., two or more, three or more, four or more, or five or more light pulses) that are collectively sufficient to release the cage-like moieties from the cage-like nucleotides. In some embodiments, activating the cage-like moieties includes exposing the cage-like moieties to light pulses sufficient to release the cage-like moieties from the cage-like nucleotides (e.g., a single light pulse). In some embodiments, activating the cage-like moieties includes exposing the cage-like moieties to multiple pulses (e.g., one or two or more light pulses) wherein the wavelength of the light is less than about 360 nm. In some embodiments, the light source with a wavelength less than 360 nm is UV light. UV light can be derived from a fluorescence microscope, a UV laser, or a UV flash lamp, or any UV light source known in the art.
[0451] In some embodiments, once the cage-like portion is released from the capture-binding domain, the oligonucleotide, probe oligonucleotide, or ligation product comprising the capture-binding domain is capable of hybridizing with the capture domain of the capture probe. Finally, to identify the location of an analyte or to determine the interaction between two or more analyte-binding moieties, all or part of the sequence of the oligonucleotide, probe oligonucleotide, or ligation product or its complement can be determined.
[0452] For further disclosure of embodiments in which the analyte capture sequence is blocked, see International Patent Application No. PCT / US2020 / 059472 entitled “Enhancing Specificity of Analyte Binding”, filed November 6, 2020, which is incorporated herein by reference.
[0453] Figure 42 This illustration depicts how a blocking probe is added to a spatially labeled analyte trap 4002 to prevent non-specific binding to trap domains on an array. In some embodiments, blocking oligonucleotides and antibodies are delivered to tissues, wherein the blocking oligonucleotides can be subsequently removed (e.g., digested with RNase) after binding to the tissue target. Figure 42 In the example shown, cleavage of the linker between the oligonucleotide and the antibody allows the oligonucleotide to migrate into the capture domain on the array. See Examples 3 and 4 below.
[0454] In some embodiments of any of the spatial analysis methods described herein, the method is used to identify immune cell atlases. Immune cells express various adaptive immune receptors associated with immune function, such as T-cell receptors (TCRs) and B-cell receptors (BCRs). T-cell and B-cell receptors play a role in the immune response by specifically recognizing and binding antigens and assisting in their destruction. More information on this application of the disclosed methods is provided in PCT Publication 202020176788A1, entitled “Analysis of bioanalytes using spatial barcode oligonucleotide arrays,” the entire contents of which are incorporated herein by reference.
[0455] (c) Base
[0456] For the spatial array-based analytical methods described in this section, the substrate (e.g., a chip) serves to support the capture probes to attach directly or indirectly to the capture points of the array. Additionally, in some embodiments, a substrate (e.g., the same substrate or different substrates) is used to support biological samples, particularly thin tissue sections, for example. Therefore, a “substrate” is a support that is insoluble in aqueous liquids and allows the biological sample, analyte, capture points, and / or capture probes to be positioned on the substrate.
[0457] A variety of different substrates can be used for the aforementioned purposes. Typically, the substrate can be any suitable supporting material. Exemplary substrates include, but are not limited to, glass, modified and / or functionalized glass, hydrogels, films, membranes, and plastics (including, for example, acrylic resins, polystyrene, copolymers of styrene and other materials, polypropylene, polyethylene, polybutene, polyurethane, Teflon, etc.). TMMaterials include cycloolefins, polyimides, nylon, ceramics, resins, Zeonor, silica or silica-based materials (including silicon and modified silicon), carbon, metals, inorganic glass, fiber bundles and polymers (such as polystyrene, cycloolefin copolymers (COC), cycloolefin polymers (COP), polypropylene, polyethylene and polycarbonate).
[0458] The substrate can also correspond to a flow cell. The flow cell can be formed from any of the aforementioned materials and can include channels that allow reagents, solvents, trapping points, and molecules to pass through the flow cell.
[0459] Among the examples of substrate materials discussed above, polystyrene is a suitable hydrophobic material for binding negatively charged macromolecules because it typically contains very few hydrophilic groups. For nucleic acids immobilized on glass slides, increasing the hydrophobicity of the glass surface can enhance nucleic acid immobilization. This enhancement can allow for relatively denser stacking formation (e.g., providing improved specificity and resolution).
[0460] In some embodiments, the substrate is coated with a surface treatment agent (such as poly-L-lysine). Additionally or alternatively, the substrate may be treated by silanization (e.g., with epoxy-silane, amino-silane, and / or with polyacrylamide).
[0461] The substrate can typically have any suitable form or format. For example, the substrate can be flat or curved, such as convex or concave curves toward the area where the biological sample (e.g., tissue sample) interacts with the substrate. In some embodiments, the substrate is a flat (e.g., planar) chip or slide. The substrate may contain one or more patterned surfaces (e.g., channels, pores, protrusions, ridges, grooves, etc.) within it.
[0462] The substrate can have any desired shape. For example, a substrate can typically be a thin, flat shape (e.g., square or rectangular). In some embodiments, the substrate structure has rounded corners (e.g., for increased security or robustness). In some embodiments, the substrate structure has one or more cut-off corners (e.g., for slide holders or cross-shaped stages). In some embodiments, where the substrate structure is flat, the substrate structure can be any suitable type of support with a flat surface (e.g., a chip or slide, such as a microscope slide).
[0463] The substrate may optionally include various structures, such as, but not limited to, protrusions, ridges, and channels. The substrate may be micropatterned to limit lateral diffusion (e.g., to prevent overlap of spatial barcodes). Substrates modified with such structures may be modified to allow analytes, capture points (e.g., beads), or probes to correlate at a single site. For example, sites on substrates modified with various structures may be adjacent to or not adjacent to other sites.
[0464] In some embodiments, the surface of the substrate may be modified to form discrete sites that have or accommodate only a single capture point. In some embodiments, the surface of the substrate may be modified such that the capture point adheres to random sites.
[0465] In some embodiments, techniques such as (but not limited to) stamping, micro-etching, and molding are used to modify the surface of a substrate to include one or more holes. In embodiments where the substrate includes one or more holes, the substrate may be a recessed sheet or a cavity sheet. For example, the holes may be formed by one or more shallow recesses on the substrate surface. In some embodiments where the substrate includes one or more holes, the holes may be formed by attaching a cartridge (e.g., a cartridge containing one or more chambers) to the surface of the substrate structure.
[0466] In some embodiments, the substrate structures (e.g., holes) may each carry different capture probes. The different capture probes attached to each structure can be identified based on the location of the structure in or on the substrate surface. Exemplary substrates include arrays of separate structures located on the substrate, including, for example, those with holes that accommodate capture points.
[0467] In some embodiments, the substrate includes one or more markings on the substrate surface, for example, to provide guidance for associating spatial information with the characterization of the object of interest. For example, the substrate may be marked with a grid of lines (e.g., to allow easy estimation of the size of an object seen at magnification and / or to provide a reference area for counting objects). In some embodiments, reference markings may be included on the substrate. Such markings may be made using techniques including, but not limited to, printing, sandblasting, and deposition on a surface.
[0468] In some embodiments where the substrate is modified to include one or more structures (including, but not limited to, pores, protrusions, ridges, or markers), the structure may include physically altered sites. For example, substrates modified with various structures may include physical properties, including but not limited to physical configuration, magnetic or compressive forces, chemically functionalized sites, chemically altered sites, and / or electrostatically altered sites.
[0469] In some embodiments where the substrate is modified to include various structures (including but not limited to holes, protrusions, ridges, or markings), these structures are applied in a patterned manner. Alternatively, these structures may be randomly distributed.
[0470] In some embodiments, the substrate is treated to minimize or reduce nonspecific analyte hybridization within or between capture points. For example, the treatment may include coating the substrate with a hydrogel, thin film, and / or membrane that creates a physical barrier against nonspecific hybridization. Any suitable hydrogel can be used. For example, a hydrogel matrix prepared according to the methods described in U.S. Patent Nos. 6,391,937, 9,512,422, and 9,889,422, and U.S. Patent Application Publication Nos. 2017 / 0253918 and 2018 / 0052081 can be used. The entire contents of each of the foregoing documents are incorporated herein by reference.
[0471] Treatment may include adding reactive or activatable functional groups, making the substrate reactive upon receiving a stimulus (e.g., photoreactive). Treatment may include treatment with a polymer having one or more physical properties (e.g., mechanical, electrical, magnetic, and / or thermal) that minimize nonspecific binding (e.g., activating the substrate at certain sites to allow the analyte to hybridize at those sites).
[0472] The substrate (e.g., beads or trapping sites on an array) may include tens of thousands to hundreds of thousands or millions of individual oligonucleotide molecules (e.g., at least about 10,000, 50,000, 100,000, 500,000, 1,000,000, 100,000,000, 1,000,000,000, or 10,000,000,000,000).
[0473] In some embodiments, the surface of the substrate is coated with a cell-allowing coating to allow live cell adhesion. A “cell-allowing coating” is a coating that allows or helps cells maintain cell viability (e.g., remain viable) on the substrate. For example, a cell-allowing coating can enhance cell adhesion, cell growth, and / or cell differentiation; for instance, a cell-allowing coating can provide nutrients to live cells. Cell-allowing coatings can include biomaterials and / or synthetic materials. Non-limiting examples of cell-allowing coatings include those incorporating one or more extracellular matrix (ECM) components (e.g., proteoglycans and fibrous proteins such as collagen, elastin, fibronectin, and laminin), polylysine, poly-L-ornithine, and / or biocompatible silicones (e.g., [missing information]). A coating characterized by [missing information]. For example, a cell-permitting coating comprising one or more extracellular matrix components may include type I collagen, type II collagen, type IV collagen, elastin, fibronectin, laminin, and / or fibronectin. In some embodiments, the cell-permitting coating comprises [missing information] derived from Engelbreth-Holm-Swarm (EHS) mouse sarcoma (e.g., [missing information]). The dissolved base membrane product is extracted from the cell. In some embodiments, the cell-permitted coating includes collagen.
[0474] In cases where the substrate comprises a gel (e.g., a hydrogel or gel matrix), oligonucleotides within the gel can adhere to the substrate. The terms "hydrogel" and "hydrogel matrix" are used interchangeably herein to refer to a macromolecular polymer gel comprising a network. Within the network, some polymer chains may optionally be cross-linked, although cross-linking does not always occur.
[0475] Further details and non-limiting embodiments of hydrogels and hydrogel subunits that may be used in this disclosure are described in U.S. Patent Application No. 16 / 992,569, filed August 13, 2020, entitled “Systems and methods for determining biological conditions using spatial distribution of unit types,” which is incorporated herein by reference.
[0476] Other examples of substrates, including, for example, reference markers on such substrates, are disclosed in PCT Publication 202020176788A1 entitled “Analysis of bioanalytes using spatial barcode oligonucleotide arrays”, which is incorporated herein by reference.
[0477] (d) Array
[0478] In many of the methods disclosed herein, the capture points coexist on a substrate. An “array” is a specific arrangement of multiple capture points (also called “features”) that are irregular or form a regular pattern. The individual capture points in an array are distinct from each other based on their relative spatial positions. Typically, at least two of the multiple capture points in an array comprise different capture probes (e.g., any instance of the capture probes described herein).
[0479] Arrays can be used to simultaneously measure a large number of analytes. In some embodiments, oligonucleotides are used at least in part to generate the array. For example, one or more copies of a single type of oligonucleotide (e.g., a capture probe) may correspond to or be directly or indirectly attached to a given capture point in the array. In some embodiments, a given capture point in the array comprises two or more oligonucleotides (e.g., capture probes). In some embodiments, the two or more oligonucleotides (e.g., capture probes) directly or indirectly attached to a given capture point on the array include a common (e.g., identical) spatial barcode.
[0480] As defined above, a “capture point” is an entity that serves as a support or reservoir for various molecular entities used in sample analysis. Examples of capture points include, but are not limited to, beads, points of any two-dimensional or three-dimensional geometry (e.g., inkjet dots, masking dots...
Claims
1. A spatial analysis method for an analyte, comprising: A) Obtain one or more images of a sample on a substrate, wherein the substrate comprises a plurality of reference markers and a set of capture points, and wherein each of the one or more images comprises a plurality of corresponding pixels in the form of an array of corresponding pixel values; B) Obtain multiple sequence readings in electronic form from the set of capture points, wherein: The substrate further comprises a set of capture probes, wherein each respective capture probe (i) is located at a different capture point in the set of capture points, and (ii) is directly or indirectly associated with one or more analytes from the sample. Each of the corresponding capture probe groups in the set of capture probe groups is characterized by at least one unique spatial barcode from a plurality of spatial barcodes. The plurality of sequence readings includes sequence readings corresponding to all or part of the one or more analytes, and Each of the plurality of sequence readings is obtained from a corresponding capture point in the set of capture points and includes a spatial barcode or its complement in at least one unique spatial barcode of the capture probe group in the set of capture probe groups at the corresponding capture point. C) Using all or a subset of the plurality of spatial barcodes to locate the corresponding sequence readings in the plurality of sequence readings to the corresponding capture point in the set of capture points, thereby dividing the plurality of sequence readings into a plurality of sequence reading subsets, each corresponding sequence reading subset corresponding to a different capture point in the set of capture points; as well as D) Using the plurality of reference markers to provide a composite representation comprising (i) one or more images aligned with the set of capture points on the substrate and (ii) a representation of all or a subset of each sequence reading at each corresponding location within each of the one or more images, mapped to a corresponding capture point corresponding to a corresponding location of the one or more analytes in the sample, wherein the use of the plurality of reference markers to provide the composite representation aligns a first image of the one or more images with the set of capture points through a process comprising the following steps: Analyze the corresponding pixel value array of the first image to identify multiple derived reference markers of the first image; A base identifier uniquely associated with the base is used to select a first template among a plurality of templates, wherein each template contains the reference position of a plurality of corresponding reference datum markers and a corresponding coordinate system; An alignment algorithm is used to align the plurality of derived reference marks of the first image with the plurality of corresponding reference reference marks of the first template to obtain the transformation between the plurality of derived reference marks of the first image and the plurality of corresponding reference reference marks of the first template. as well as The corresponding position of each capture point in the first image is located using the transformation and the coordinate system of the first template, as follows: Multiple heuristic classifiers are run on the plurality of pixels, wherein, for each corresponding pixel among the plurality of pixels, each corresponding heuristic classifier votes on the corresponding pixel between a first class and a second class, thereby forming a corresponding aggregate score for each corresponding pixel among the plurality of pixels, and the aggregate score and intensity of each corresponding pixel among the plurality of pixels are applied to a segmentation algorithm to independently assign a probability as a sample or background to each corresponding pixel among the plurality of pixels.
2. The method of claim 1, wherein the composite representation provides the relative abundance of nucleic acid fragments of each of a plurality of analytes mapped to each of the set of capture points.
3. The method of claim 1, wherein each corresponding aggregate score is one of a set of classes comprising an obvious first class, a possible first class, a possible second class, and an obvious second class.
4. The method of claim 1, wherein the method further comprises, for each corresponding locus among a plurality of loci, performing a process comprising the following steps: i) Align each corresponding sequence read that maps to the corresponding locus among the plurality of sequence reads, thereby determining the haplotype identity of the corresponding sequence reads from the corresponding haplotype set of the corresponding locus, and ii) Classify each corresponding sequence reading mapped to the corresponding locus among the plurality of sequence readings by means of the spatial barcode and the typological identity of the corresponding sequence reading. This determines the spatial distribution of each haplotype in each corresponding haplotype group in the sample, wherein for each capture point in the set of capture points on the substrate, the spatial distribution includes the abundance of each haplotype in the haplotype group of the corresponding locus.
5. The method of claim 4, wherein the method further comprises using the spatial distribution to characterize the biological status of the subject.
6. The method according to claim 2, wherein the method further comprises: A mask is applied over the first image, wherein the mask assigns a first attribute to each of the plurality of pixels in the first image that has a greater probability of being assigned as a sample, and assigns a second attribute to each of the plurality of pixels that has a greater probability of being assigned as background.
7. The method of claim 6, wherein the first attribute is a first color and the second attribute is a second color.
8. The method of claim 7, wherein the first color is one of red and blue, and the second color is the other of red and blue.
9. The method of claim 6, wherein the first attribute is a first level of brightness or opacity, and the second attribute is a second level of brightness or opacity.
10. The method according to claim 6, wherein the method further comprises: Based on the independent allocation of pixels near the corresponding representation of the capture point in the composite representation, each corresponding representation of the capture point in the set of capture points in the composite representation is assigned the first attribute or the second attribute.
11. The method of claim 1, wherein the capture point in the set of capture points comprises a capture domain.
12. The method of claim 1, wherein the capture points in the set of capture points comprise a cutting domain.
13. The method of claim 1, wherein each of the set of capture points is directly or indirectly attached to the substrate.
14. The method of claim 1, wherein the one or more analytes comprises five or more analytes, ten or more analytes, fifty or more analytes, one hundred or more analytes, five hundred or more analytes, 1,000 or more analytes, 2,000 or more analytes, or 2,000 to 100,000 analytes.
15. The method of claim 1, wherein the unique spatial barcode encoding is from the set {1, …, 1024}, {1, …, 4096}, {1, …, 16384}, {1, …, 65536}, {1, …, 262144}, {1, …, 1048576}, {1, …, 4194304}, {1, …, 16777216}, {1, …, 67108864} or {1, …, 1 x 10 12 The only predefined value selected in}.
16. The method according to any one of claims 1 to 15, wherein a corresponding capture probe group in the set of capture probe groups comprises 1,000 or more capture probes, 2,000 or more capture probes, 10,000 or more capture probes, 100,000 or more capture probes, 1 x 10 6 One or more capture probes, 2 x 10 6 One or more capture probes, or 5 x 10 6 One or more capture probes.
17. The method of claim 16, wherein each capture probe in the respective capture probe group comprises a poly-T sequence and the unique spatial barcode characterizing the different capture points.
18. The method of claim 16, wherein each capture probe in the respective capture probe group comprises the same spatial barcode as the plurality of spatial barcodes.
19. The method of claim 16, wherein each capture probe in the respective capture probe group comprises a spatial barcode different from the plurality of spatial barcodes.
20. The method of claim 1, wherein the sample is a slice of tissue with a depth of 100 micrometers or less.
21. The method of claim 20, wherein the one or more images comprise a plurality of images, and a first image of the plurality of images is obtained using a first slice of the sample, and a second image of the plurality of images is obtained using a second slice of the sample.
22. The method of claim 1, wherein The one or more analytes mentioned are multiple analytes. The corresponding capture probe group in the set of capture probe groups includes multiple capture probes, each of the multiple capture probes including a capture domain characterized by one of a variety of capture domain types, and Each of the multiple capture domain types is configured to combine different analytes from the multiple analytes.
23. The method of claim 22, wherein the plurality of capture domain types comprises 2 to 15,000 capture domain types, and for each of the plurality of capture domain types, the corresponding capture probe group comprises at least 5, at least 10, at least 100, or at least 1,000 capture probes.
24. The method of claim 1, wherein The one or more analytes mentioned are multiple analytes. The respective capture point in the set of capture points includes multiple capture probes, each of the multiple capture probes including a capture domain characterized by a single capture domain type configured to combine each of the multiple analytes in an unbiased manner.
25. The method of claim 1, wherein each of at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the set of capture points is contained within a corresponding 100 μm × 100 μm square on the substrate.
26. The method of claim 1, wherein the distance between the center of each respective capture point in the set of capture points on the substrate and the adjacent capture point is between 50 micrometers and 300 micrometers.
27. The method of claim 1, wherein at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the set of capture points have a diameter of 80 micrometers or less.
28. The method of claim 1, wherein at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the set of capture points have a diameter between 25 micrometers and 65 micrometers.
29. The method of claim 1, wherein the distance between the center of each respective capture point in the set of capture points on the substrate and the adjacent capture point is between 40 micrometers and 100 micrometers.
30. The method of claim 2, wherein the plurality of heuristic classifiers includes a first heuristic classifier, the first heuristic classifier identifying a single intensity threshold for dividing the plurality of pixels into a first class and a second class, such that the first heuristic classifier votes for either the first class or the second class for each corresponding pixel among the plurality of pixels, and wherein the single intensity threshold represents minimizing the intra-class intensity variance between the first class and the second class or maximizing the inter-class variance between the first class and the second class.
31. The method of claim 30, wherein the plurality of heuristic classifiers includes a second heuristic classifier that identifies local neighborhoods of pixels having the same class identified by the first heuristic classifier, and applies a smoothing metric of the maximum difference in intensity between pixels in the local neighborhood, such that the second heuristic classifier votes for either the first class or the second class for each corresponding pixel among the plurality of pixels.
32. The method of claim 31, wherein the plurality of heuristic classifiers includes a third heuristic classifier that performs edge detection on the plurality of pixels to form a plurality of edges in the image, morphologically closes the plurality of edges to form a plurality of morphologically closed regions in the image, and assigns pixels in the morphologically closed regions to a first class and assigns pixels outside the morphologically closed regions to a second class, thereby causing the third heuristic classifier to vote for either the first class or the second class for each corresponding pixel among the plurality of pixels.
33. The method according to claim 32, wherein: Each corresponding pixel assigned to the second class by each of the multiple classifiers from the heuristic classifiers is labeled as the obvious second class, and Each corresponding pixel assigned to the first class by each of the plurality of heuristic classifiers is labeled as the obvious first class.
34. The method according to claim 33, wherein the graphic clipping and segmentation algorithm is the GrabCut segmentation algorithm.
35. The method of claim 12, wherein the cleavage domain comprises a sequence recognized and cleaved by uracil-DNA glycosylase and / or endonuclease VIII.
36. The method of claim 1, wherein the group of capture probes in the group of capture probes does not contain a cutting domain, and each capture probe in the group of capture probes is not cut from the substrate.
37. The method of claim 1, wherein one or more analytes comprise DNA or RNA.
38. The method of claim 1, wherein the one or more analytes comprise proteins.
39. The method of claim 1, wherein each of the respective groups of capture probes in the group of capture probes is directly or indirectly attached to the substrate.
40. The method of claim 1, wherein C) involves in situ sequencing of the set of capture sites on the substrate.
41. The method of claim 1, wherein C) obtains high-throughput sequencing.
42. The method of claim 4, wherein the corresponding locus among the plurality of loci is a biallelic gene, and the corresponding haplotype group of the corresponding locus is composed of a first allele and a second allele.
43. The method of claim 42, wherein the corresponding locus comprises a heterozygous single nucleotide polymorphism (SNP), a heterozygous insertion, or a heterozygous deletion.
44. The method according to any one of claims 1 to 15, wherein the plurality of sequence readings comprises 50,000 or more sequence readings, 100,000 or more sequence readings, or 1 x 10 6 One or more sequence readings.
45. The method of claim 4, wherein the plurality of loci comprises 2 to 100 loci, more than 10 loci, more than 100 loci, or more than 500 loci.
46. The method of claim 1, wherein the unique spatial barcode in the corresponding sequence reading is located within a set of consecutive nucleotides in the corresponding sequence reading.
47. The method of claim 46, wherein the set of consecutive nucleotides is an N-mer, where N is an integer selected from the set {4, …, 20}.
48. The method of claim 4, wherein the method further comprises retrieving the plurality of loci from a lookup table, file, or data structure.
49. The method of claim 4, wherein the alignment algorithm is a local alignment that uses a scoring system to align the corresponding sequence reading with a reference sequence, the scoring system (i) penalizing mismatches between nucleotides in the corresponding sequence reading and corresponding nucleotides in the reference sequence according to a permutation matrix, and (ii) penalizing gaps that introduce into the alignment of the sequence reading with the reference sequence.
50. The method of claim 49, wherein the local alignment is a Smith-Waltman alignment.
51. The method of claim 49, wherein the reference sequence is all or part of a reference genome.
52. The method of claim 4, the method further comprising removing one or more sequence reads from the plurality of sequence reads that do not cover any of the plurality of loci.
53. The method of claim 4, wherein the plurality of loci include one or more loci on a first chromosome and one or more loci on a second chromosome other than the first chromosome.
54. The method according to any one of claims 1 to 15, wherein the plurality of sequence readings comprises sequence readings paired at the 3' end or the 5' end.
55. The method of claim 5, wherein the biological condition is the absence or presence of a disease.
56. The method of claim 5, wherein the biological condition is a cancer type.
57. The method of claim 5, wherein the biological condition is a stage of disease.
58. The method of claim 5, wherein the biological condition is a stage of cancer.
59. The method according to any one of claims 1 to 15, wherein the one or more images comprise a bright-field image or a fluorescence image of the sample.
60. The method according to any one of claims 1 to 15, wherein the one or more images are multiple images.
61. The method of claim 60, wherein: The first image among the plurality of images is a bright-field image of the sample, and The second image among the plurality of images is a fluorescence image of the sample.
62. The method of claim 61, wherein the fluorescence image is an immunofluorescence image.
63. The method according to any one of claims 1 to 15, wherein the one or more images are multiple images, and the multiple images comprise two or more fluorescent images.
64. The method according to any one of claims 1 to 15, wherein the representation of all or part of each subset of sequence readings at a corresponding location within the one or more images conveys a mapping to a plurality of unique molecules of a particular analyte or combination of analytes in the sample represented by the subset of sequence readings, which in turn maps to the corresponding capture point.
65. The method of claim 64, wherein the plurality of unique molecules mapped to a specific analyte or combination of analytes in the sample represented by the subset of sequence reads are communicated using a color scale or intensity scale, the subset of sequence reads being mapped to the corresponding capture point.
66. The method according to any one of claims 1 to 15, wherein a corresponding capture probe group in the set of capture probe groups directly associates with the analyte from the sample.
67. The method according to any one of claims 1 to 15, wherein a respective group of capture probes in the group of capture probes associates indirectly with an analyte from the sample via an analyte capture agent.
68. The method of claim 1, further comprising placing the sample on a substrate prior to obtaining (A).
69. A spatial analysis method for an analyte, comprising: A) Obtain one or more images of the sample on the substrate, wherein: The base comprises multiple reference markers and a set of capture points. Each of the one or more images contains a plurality of corresponding pixels in the form of an array of corresponding pixel values. The corresponding pixel value array contains at least 100,000 pixel values, and The set of capture points includes at least 1000 capture points; B) Obtain multiple sequence readings in electronic form from the set of capture points, wherein: The substrate further comprises a set of capture probes, wherein each respective capture probe (i) is located at a different capture point in the set of capture points, and (ii) is associated directly or indirectly with one or more analytes from the sample. Each of the corresponding capture probe groups in the set of capture probe groups is characterized by at least one unique spatial barcode from a plurality of spatial barcodes. The plurality of sequence readings includes all or part of the sequence readings corresponding to one or more of the analytes. The plurality of sequence readings includes at least 10,000 sequence readings, and Each of the plurality of sequence readings is obtained from a corresponding capture point in the set of capture points and includes a spatial barcode or its complement in at least one unique spatial barcode of the capture probe group in the set of capture probe groups at the corresponding capture point; and C) Using all or a subset of the plurality of spatial barcodes, the corresponding sequence reading from the plurality of sequence readings is located to the corresponding capture point in the set of capture points. This divides the multiple sequence readings into multiple sequence reading subsets, with each corresponding sequence reading subset corresponding to a different capture point in the set of capture points; D) Using the plurality of reference markers to provide a composite representation comprising (i) one or more images aligned with the set of capture points on the substrate and (ii) a representation of all or a subset of each sequence reading at each corresponding location within each of the one or more images, mapped to a corresponding capture point corresponding to a corresponding location of the one or more analytes in the sample. The method further includes, for each corresponding locus among a plurality of loci, performing a process comprising the following steps: i) Align each corresponding sequence read that maps to the corresponding locus among the plurality of sequence reads, thereby determining the haplotype identity of the corresponding sequence reads from the corresponding haplotype set of the corresponding locus, and ii) Classify each corresponding sequence reading mapped to the corresponding locus among the plurality of sequence readings by means of the spatial barcode and the typological identity of the corresponding sequence reading. This determines the spatial distribution of each haplotype in each corresponding haplotype group in the sample, wherein for each capture point in the set of capture points on the substrate, the spatial distribution includes the abundance of each haplotype in the haplotype group of the corresponding locus.
70. The method of claim 69, further comprising placing the sample on a substrate prior to obtaining (A).
71. The method of claim 69, wherein the method further comprises using the spatial distribution to characterize the biological status of the subject.
72. The method of claim 69, wherein the capture point in the set of capture points comprises a capture domain.
73. The method of claim 69, wherein the capture point in the set of capture points comprises a cutting domain.
74. The method of claim 69, wherein the one or more analytes comprise five or more analytes, ten or more analytes, fifty or more analytes, one hundred or more analytes, five hundred or more analytes, 1,000 or more analytes, 2,000 or more analytes, or 2,000 to 100,000 analytes.
75. The method of claim 69, wherein the unique spatial barcode encoding is from the set {1, …, 1024}, {1, …, 4096}, {1, …, 16384}, {1, …, 65536}, {1, …, 262144}, {1, …, 1048576}, {1, …, 4194304}, {1, …, 16777216}, {1, …, 67108864} or {1, …, 1 x 10 12 The only predefined value selected in}.
76. The method of claim 69, wherein a corresponding capture probe group in the set of capture probe groups comprises 1,000 or more capture probes, 2,000 or more capture probes, 10,000 or more capture probes, 100,000 or more capture probes, 1 x 10 6 One or more capture probes, 2 x 10 6 One or more capture probes, or 5 x 10 6 One or more capture probes.
77. The method of claim 69, wherein one or more analytes comprise DNA or RNA.
78. The method of claim 69, wherein one or more analytes comprise proteins.
79. The method of claim 69, wherein the plurality of sequence readings comprises 50,000 or more sequence readings, 100,000 or more sequence readings, or 1 x 10 6 One or more sequence readings.
80. A spatial analysis method for an analyte, comprising: A) Aligning a first image of a sample on a substrate with a set of capture points, wherein the substrate comprises a plurality of reference markers and a set of capture points, and wherein the first image comprises a plurality of corresponding pixels in the form of an array of corresponding pixel values, the alignment comprising: Analyze the corresponding pixel value array of the first image to identify multiple derived reference markers of the first image; A base identifier uniquely associated with the base is used to select a first template among a plurality of templates, wherein each template contains the reference position of a plurality of corresponding reference datum markers and a corresponding coordinate system; An alignment algorithm is used to align the plurality of derived reference markers of the first image with the corresponding plurality of reference reference markers of the first template, thereby obtaining the transformation between the plurality of derived reference markers of the first image and the corresponding plurality of reference reference markers of the first template. The transformation and the coordinate system of the first template are used to locate the corresponding position of each capture point in the first image in the set of capture points; as well as B) Provide a composite representation comprising (i) a first image aligned with the set of capture points on the substrate and (ii) a representation of all or a portion of each of a plurality of sequence read subsets at each corresponding location within the first image, mapped to corresponding capture points at corresponding locations corresponding to one or more analytes in the sample, wherein: Each of the plurality of sequence reading subsets comprises (i) all or part of the sequence readings corresponding to one or more of the analytes, and (ii) different capture points within the set of capture points. Each of the plurality of sequence readings includes a spatial barcode in a plurality of spatial barcodes, each of the set of capture probe groups is characterized by at least one spatial barcode in the plurality of spatial barcodes, each of the plurality of sequence readings is obtained from a corresponding capture point in the set of capture points and includes a spatial barcode or its complement in the at least one unique spatial barcode of the capture probe group in the set of capture probe groups at the corresponding capture point, wherein each of the set of capture probe groups (i) is located at a different capture point in the set of capture points and (ii) is directly or indirectly associated with one or more analytes from the sample.
81. The method of claim 80, further comprising, prior to the alignment A): Place the sample on the substrate; Obtain at least the first image of the sample on the substrate; Multiple sequence readings are obtained electronically from the set of capture points; as well as Using all or a subset of the plurality of spatial barcodes to locate the corresponding sequence readings among the plurality of sequence readings to the corresponding capture point in the set of capture points, thereby dividing the plurality of sequence readings into a plurality of sequence reading subsets.
82. The method according to claim 81, wherein: The set of capture points includes at least 1000 capture points. The corresponding pixel value array contains at least 100,000 pixel values, and The plurality of sequence readings comprises at least 10,000 sequence readings.
83. The method of claim 80, wherein the composite representation provides the relative abundance of nucleic acid fragments of each of a plurality of analytes mapped to each of the set of capture points.
84. The method of claim 80, wherein the method further comprises, for each corresponding locus of a plurality of loci, performing a process comprising the following steps: i) Align each corresponding sequence read that maps to the corresponding locus among the plurality of sequence reads, thereby determining the haplotype identity of the corresponding sequence reads from the corresponding haplotype set of the corresponding locus, and ii) Classify each corresponding sequence reading mapped to the corresponding locus among the plurality of sequence readings by means of the spatial barcode and the typological identity of the corresponding sequence reading. This determines the spatial distribution of each haplotype in each corresponding haplotype group in the sample, wherein for each capture point in the set of capture points on the substrate, the spatial distribution includes the abundance of each haplotype in the haplotype group of the corresponding locus.
85. The method of claim 84, wherein the method further comprises using the spatial distribution to characterize the subject's biological status.
86. The method of claim 83, further comprising: A mask is applied over the first image, wherein the mask assigns a first attribute to each of the plurality of pixels in the first image that has a greater probability of being assigned as a sample, and assigns a second attribute to each of the plurality of pixels that has a greater probability of being assigned as background.
87. The method of claim 80, wherein the capture point in the set of capture points comprises a capture domain.
88. The method of claim 80, wherein the capture point in the set of capture points comprises a cutting domain.
89. The method of claim 80, wherein the one or more analytes comprises five or more analytes, ten or more analytes, fifty or more analytes, one hundred or more analytes, five hundred or more analytes, 1,000 or more analytes, 2,000 or more analytes, or 2,000 to 100,000 analytes.
90. The method of claim 80, wherein the unique spatial barcode encoding is from the set {1, …, 1024}, {1, …, 4096}, {1, …, 16384}, {1, …, 65536}, {1, …, 262144}, {1, …, 1048576}, {1, …, 4194304}, {1, …, 16777216}, {1, …, 67108864} or {1, …, 1 x 10 12 The only predefined value selected in}.
91. The method of claim 80, wherein a corresponding capture probe group in the set of capture probe groups comprises 1,000 or more capture probes, 2,000 or more capture probes, 10,000 or more capture probes, 100,000 or more capture probes, 1 x 10 6 One or more capture probes, 2 x 10 6 One or more capture probes, or 5 x 10 6 One or more capture probes.
92. The method of claim 80, wherein one or more analytes comprise DNA or RNA.
93. The method of claim 80, wherein one or more analytes comprise proteins.
94. The method of claim 81, wherein the plurality of sequence readings comprises 50,000 or more sequence readings, 100,000 or more sequence readings, or 1 x 10 6 One or more sequence readings.
95. A computer system, comprising: One or more processors; Memory; as well as One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being used for spatial analysis of an object, the one or more programs including instructions for performing the method of any one of claims 1 to 94.
96. A computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by an electronic device having one or more processors and a memory, cause said electronic device to perform a spatial analysis of an analyzeant comprising the method of any one of claims 1 to 94.
97. A computer system, comprising: One or more processors; Memory; as well as One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being used for spatial analysis of an object, the one or more programs including instructions for: A) Obtain one or more images of a sample on a substrate, wherein the substrate comprises a plurality of reference markers and a set of capture points, and wherein each of the one or more images comprises a plurality of corresponding pixels in the form of an array of corresponding pixel values; B) Obtain multiple sequence readings in electronic form from the set of capture points, wherein: The substrate further comprises a set of capture probes, wherein each respective capture probe (i) is located at a different capture point in the set of capture points, and (ii) is associated directly or indirectly with one or more analytes from the sample. Each of the corresponding capture probe groups in the set of capture probe groups is characterized by at least one unique spatial barcode from a plurality of spatial barcodes. The plurality of sequence readings includes sequence readings corresponding to all or part of the one or more analytes, and Each of the plurality of sequence readings is obtained from a corresponding capture point in the set of capture points and includes a spatial barcode or its complement in at least one unique spatial barcode of the capture probe group in the set of capture probe groups at the corresponding capture point. C) Using all or a subset of the plurality of spatial barcodes to locate corresponding sequence readings from the plurality of sequence readings to corresponding capture points in the set of capture points, thereby dividing the plurality of sequence readings into multiple sequence reading subsets, each corresponding sequence reading subset corresponding to a different capture point in the set of capture points; and D) Using the plurality of reference markers to provide a composite representation comprising (i) one or more images aligned with the set of capture points on the substrate and (ii) a representation of all or a subset of each sequence reading at each corresponding location within each of the one or more images, mapped to a corresponding capture point corresponding to a corresponding location of the one or more analytes in the sample, wherein the use of the plurality of reference markers to provide the composite representation aligns a first image of the one or more images with the set of capture points through a process comprising the following steps: Analyze the corresponding pixel value array of the first image to identify multiple derived reference markers of the first image; A base identifier uniquely associated with the base is used to select a first template among a plurality of templates, wherein each template contains the reference position of a plurality of corresponding reference datum markers and a corresponding coordinate system; An alignment algorithm is used to align the plurality of derived reference markers of the first image with the corresponding plurality of reference reference markers of the first template, thereby obtaining a transformation between the plurality of derived reference markers of the first image and the corresponding plurality of reference reference markers of the first template; and By running multiple heuristic classifiers on the plurality of pixels, the corresponding position of each capture point in the first image is located using the transformation and the coordinate system of the first template. For each corresponding pixel among the plurality of pixels, each corresponding heuristic classifier votes for the corresponding pixel between a first class and a second class, thereby forming a corresponding total score for each corresponding pixel among the plurality of pixels. The total score and intensity of each corresponding pixel among the plurality of pixels are then applied to a segmentation algorithm to independently assign a probability as either a sample or background to each corresponding pixel among the plurality of pixels.
98. A computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by an electronic device having one or more processors and a memory, cause the electronic device to perform spatial analysis of an analyte by a method comprising the following steps: A) Obtain one or more images of a sample on a substrate, wherein the substrate comprises a plurality of reference markers and a set of capture points, and wherein each of the one or more images comprises a plurality of corresponding pixels in the form of an array of corresponding pixel values; B) Obtain multiple sequence readings in electronic form from the set of capture points, wherein: The substrate further comprises a set of capture probes, wherein each respective capture probe (i) is located at a different capture point in the set of capture points, and (ii) is associated directly or indirectly with one or more analytes from the sample. Each of the corresponding capture probe groups in the set of capture probe groups is characterized by at least one unique spatial barcode from a plurality of spatial barcodes. The plurality of sequence readings includes sequence readings corresponding to all or part of the one or more analytes, and Each of the plurality of sequence readings is obtained from a corresponding capture point in the set of capture points and includes a spatial barcode or its complement in at least one unique spatial barcode of the capture probe group in the set of capture probe groups at the corresponding capture point. C) Using all or a subset of the plurality of spatial barcodes to locate the corresponding sequence readings in the plurality of sequence readings to the corresponding capture point in the set of capture points, thereby dividing the plurality of sequence readings into a plurality of sequence reading subsets, each corresponding sequence reading subset corresponding to a different capture point in the set of capture points; as well as D) Using the plurality of reference markers to provide a composite representation comprising (i) one or more images aligned with the set of capture points on the substrate and (ii) a representation of all or a subset of each sequence reading at each corresponding location within each of the one or more images, mapped to a corresponding capture point corresponding to a corresponding location of the one or more analytes in the sample, wherein the use of the plurality of reference markers to provide the composite representation aligns a first image of the one or more images with the set of capture points through a process comprising the following steps: Analyze the corresponding pixel value array of the first image to identify multiple derived reference markers of the first image; A base identifier uniquely associated with the base is used to select a first template among a plurality of templates, wherein each template contains the reference position of a plurality of corresponding reference datum markers and a corresponding coordinate system; An alignment algorithm is used to align the plurality of derived reference marks of the first image with the plurality of corresponding reference reference marks of the first template to obtain the transformation between the plurality of derived reference marks of the first image and the plurality of corresponding reference reference marks of the first template. as well as The corresponding position of each capture point in the first image is located using the transformation and the coordinate system of the first template as follows: Multiple heuristic classifiers are run on the plurality of pixels, wherein, for each corresponding pixel among the plurality of pixels, each corresponding heuristic classifier votes on the corresponding pixel between a first class and a second class, thereby forming a corresponding aggregate score for each corresponding pixel among the plurality of pixels, and the aggregate score and intensity of each corresponding pixel among the plurality of pixels are applied to a segmentation algorithm to independently assign a probability as a sample or background to each corresponding pixel among the plurality of pixels.
99. A kit for performing spatial analysis of analytes, comprising: The base, comprising a plurality of reference markers and a set of capture points, wherein: The set of capture points contains at least 1000 capture points. Each of the set of capture points contains a corresponding capture probe group from a set of capture probe groups. Each corresponding capture probe group is characterized by at least one unique spatial barcode from a plurality of spatial barcodes, and The base includes (i) a base identifier uniquely associated with the base and (ii) a first template among a plurality of templates, wherein the first template among the plurality of templates includes reference positions of a plurality of corresponding reference datum marks and a corresponding coordinate system of the first template; as well as Used for the following instructions: The sample is placed on a substrate, wherein each of the set of capture probes associates with one or more analytes from the sample. Obtain one or more images of the sample on the substrate, wherein each of the one or more images contains a plurality of corresponding pixels in the form of an array of corresponding pixel values. Multiple sequence readings are obtained, each of the multiple sequence readings being a corresponding sequence reading: (i) all or part of the corresponding analyte in the sample, and (ii) Contains a spatial barcode of the capture probe group in the set of capture probe groups associated with the corresponding analyte. Using all or a subset of the plurality of spatial barcodes, the corresponding sequence readings from the plurality of sequence readings are located to the corresponding capture point in the set of capture points. The plurality of sequence readings are divided into a plurality of sequence reading subsets, and each corresponding sequence reading subset corresponds to a different capture point in the set of capture points; as well as The plurality of reference markers are used to provide a composite representation comprising (i) one or more images aligned with the set of capture points on the substrate and (ii) a representation of all or a subset of each sequence reading at each corresponding location within each of the one or more images, mapped to the corresponding capture point corresponding to the corresponding location of the one or more analytes in the sample by aligning a first image of the one or more images with the set of capture points.
100. The kit of claim 99, wherein aligning the first image of the one or more images with the set of capture points comprises: Analyze the corresponding pixel value array of the first image to identify multiple derived reference markers of the first image; A base identifier uniquely associated with the base is used to select a first template among a plurality of templates, wherein each template contains the reference position of a plurality of corresponding reference datum markers and a corresponding coordinate system; An alignment algorithm is used to align the plurality of derived reference marks of the first image with the plurality of corresponding reference reference marks of the first template to obtain the transformation between the plurality of derived reference marks of the first image and the plurality of corresponding reference reference marks of the first template. as well as By running multiple heuristic classifiers on the plurality of pixels, the corresponding position of each capture point in the first image is located using the transformation and the coordinate system of the first template. For each corresponding pixel among the plurality of pixels, each corresponding heuristic classifier votes for the corresponding pixel between a first class and a second class, thereby forming a corresponding total score for each corresponding pixel among the plurality of pixels. The total score and intensity of each corresponding pixel among the plurality of pixels are then applied to a segmentation algorithm to independently assign a probability as either a sample or background to each corresponding pixel among the plurality of pixels.
101. The kit of claim 99, wherein the instructions further comprise, for each corresponding locus of the plurality of loci, performing a process comprising the following steps: i) Align each corresponding sequence read that maps to the corresponding locus among the plurality of sequence reads, thereby determining the haplotype identity of the corresponding sequence reads from the corresponding haplotype set of the corresponding locus, and ii) Classify each corresponding sequence reading mapped to the corresponding locus among the plurality of sequence readings by means of the spatial barcode and the typological identity of the corresponding sequence reading. This determines the spatial distribution of each haplotype in each corresponding haplotype group in the sample, wherein for each capture point in the set of capture points on the substrate, the spatial distribution includes the abundance of each haplotype in the haplotype group of the corresponding locus.
102. The kit of claim 99, wherein the capture point in the set of capture points comprises a capture domain.
103. The kit of claim 99, wherein the capture point in the set of capture points comprises a cutting domain.
104. The kit of claim 99, wherein each of the set of capture points is directly or indirectly attached to the substrate.
105. The kit according to claim 99, wherein the unique spatial barcode encoding is from the set {1, …, 1024}, {1, …, 4096}, {1, …, 16384}, {1, …, 65536}, {1, …, 262144}, {1, …, 1048576}, {1, …, 4194304}, {1, …, 16777216}, {1, …, 67108864} or {1, …, 1 x 10 12 The only predefined value selected in}.
106. The kit of claim 99, wherein a corresponding capture probe group in the set of capture probe groups comprises 1000 or more capture probes, 2000 or more capture probes, 10,000 or more capture probes, 100,000 or more capture probes, 1 x 10 6 One or more capture probes, 2 x 10 6 One or more capture probes, or 5 x 10 6 One or more capture probes.
107. The kit of claim 99, wherein each capture probe in the respective capture probe group comprises a poly-T sequence and the unique spatial barcode characterizing the different capture points.
108. The kit of claim 99, wherein each capture probe in the respective capture probe group comprises a spatial barcode identical to the plurality of spatial barcodes.
109. The kit of claim 99, wherein each capture probe in the respective capture probe group comprises a spatial barcode different from the plurality of spatial barcodes.
110. The kit of claim 99, wherein each of at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the group of capture points is contained within a corresponding 100 μm × 100 μm square on the substrate.
111. The kit of claim 99, wherein the distance between the center of each respective capture point in the set of capture points on the substrate and the adjacent capture point is between 50 micrometers and 300 micrometers.
112. The kit of claim 99, wherein at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the set of capture points have a diameter of 80 micrometers or less.
113. The kit of claim 99, wherein at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the group of capture points have a diameter between 25 micrometers and 65 micrometers.
114. The kit of claim 99, wherein the distance between the center of each respective capture point in the set of capture points on the substrate and the adjacent capture point is between 40 micrometers and 100 micrometers.