Spatially oriented quantum barcoding of cellular targets

The method of assembling unique nucleic acid barcodes on targets in tissues or cells addresses the limitations of current labeling techniques by enabling efficient, high-throughput detection and localization of targets with an essentially unlimited number of labels, overcoming the constraints of fluorescent and isotopic labeling.

JP2025124697AInactive Publication Date: 2025-08-26F HOFFMANN LA ROCHE & CO AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025082357
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-12-04
Filing Date
2025-05-16
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current methods for obtaining spatial (2D or 3D) information about tissues, biofilms, or cells require time-consuming microscopic examination with limited fluorescent labels or isotopic labeling, which is costly and inefficient for high-throughput imaging.

Method used

A method for spatially labeling targets using unique nucleic acid barcodes assembled from subcodes, which can be read by sequencing or mass spectrometry, allowing for an essentially unlimited number of labels at low cost.

Benefits of technology

Enables high-throughput, cost-effective detection and localization of targets in tissues or cells without the need for microscopy, providing qualitative and quantitative analysis of DNA, RNA, and protein targets with high spatial resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025124697000001_ABST
    Figure 2025124697000001_ABST
Patent Text Reader

Abstract

To provide a method for labeling and detecting a target on a 2D surface or in a 3D space, and a method for assembling a unique nucleic acid or mass barcode on each target in a sample.SOLUTION: The invention is a method of simultaneously detecting the presence and spatial location of a target in a tissue sample by attaching an anchor to the target and assembling unique positional barcodes on the anchor. The method enables analyzing cellular targets in 3D.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the fields of tissue staining and molecular pathology, and more particularly to the field of labeling and barcoding cellular targets within three-dimensional tissue samples. [Background technology]

[0002] Proper tissue function requires the spatial positioning and orientation of cells. Cell contact allows for the transmission of information from one cell to another in a predictable manner. Diseased tissues can often be characterized by inaccurate spatial localization of cells or the disruption of specific cell types necessary for the development and function of other cells. Technologies that enable 2D and 3D spatial localization provide information that can be used to understand the normal function of cells and cellular components and why the absence of such function leads to clinical symptoms of disease. Currently, the primary means of obtaining spatial (2D or 3D) information about tissues, biofilms, or cells requires microscopic examination of cells bound with labeled antibodies or other types of probes. Time, cost, and availability of labeling reagents are significant obstacles to high-throughput imaging of biological samples. The number of distinguishable fluorescent labels is very limited, typically three or four at a time, up to 40 or 50 with repeated staining. Isotopic labeling allows for up to 40 options, albeit at a very high cost. Technologies using fluorophores and isotopes for cell imaging are still under development. However, spatial analysis can be taken to a new level by utilizing the use of nucleic acid sequence tags or chemical adducts (detected by mass spectrometry). The use of these labels could potentially enable imaging of tissues or cells without a microscope. This brings imaging from the field of microscopy into the field of tagging and spatial identification. The present invention is a method with an essentially unlimited number of labels (codes) that can be generated at little cost for both nucleic acid and protein targets in tissue or other 2D or 3D samples. Summary of the Invention

[0003] The present invention is a method for labeling and detecting targets on a 2D surface or in 3D space. The method assembles a unique nucleic acid or mass barcode on each target in a sample. The unique barcode encodes both the target's identity and its location. To assemble each barcode from subcodes, a portion of the sample is irradiated, allowing the subcodes to bind to only a portion of the sample. The sample is subdivided into as many regions as necessary, with each region having a unique barcode associated with it. The barcodes can be read and interpreted by sequencing or mass spectrometry.

[0004] In some embodiments, the present invention provides a method for simultaneously detecting the presence and spatial location of a target in a tissue sample, comprising covalently binding an anchor to a target in the tissue sample via a reactive group; assembling a code on the anchor from a set of subcodes by a method comprising: contacting the sample with a first subcode and allowing the first subcode to covalently bind to the anchor in a first portion of the tissue sample to form a code thereon; contacting the sample with a second subcode and allowing the second subcode to covalently bind to the anchor in a second portion of the tissue sample that does not overlap with the first portion to form a code thereon; and repeating a pair of steps i-ii one or more times, wherein in each repetition, the portion of the tissue sample contacted in the first step does not overlap with the portion of the tissue sample contacted in the second step; and wherein the added subcodes bind to and extend the existing subcodes, thereby forming a code marking each portion of the tissue sample. and reading the code assembled on the anchor using the anchor to detect the presence and location of the target in the tissue sample. Only two or more sub-codes may be used in the code assembling step.

[0005] In some embodiments, the anchor is a nucleic acid that can include one or more of a polyA binding sequence, a random sequence, and a target-specific sequence. In some embodiments, the anchor is an aptamer.

[0006] The subcodes can also be nucleic acids. In some embodiments, the anchors are covalently linked to the target via crosslinking. The subcodes can also be covalently linked to the anchors and other subcodes via crosslinking. In some embodiments, prior to covalent linking, the subcodes hybridize to the anchors and other subcodes via regions of complementarity to the anchors and other subcodes. Uncrosslinked subcodes can be removed, for example, by washing in a solution containing one or more of a salt buffer, a detergent, and a solvent. The crosslinking can be by irradiation with radiation selected from acoustic radiation, a light beam (laser), a terahertz frequency beam, and X-rays. The laser can be operated by a computer executing code that references the time of irradiation of the portion of the tissue sample being irradiated and the sequence of the subcode that contacts the tissue sample at that time.

[0007] In some embodiments, the subcodes are attached to anchors and other subcodes via sonication, which promotes a chemical reaction.

[0008] In some embodiments, the target is a protein. The reactive group in the anchor can be a thymidine attached to the target protein via a thymidine-lysine addition.

[0009] In some embodiments, the subcodes are linked to a common linker or to an existing subcode via an annealing primer. The subcodes can be covalently linked by ligation, which can be preceded by chain extension with a polymerase, such as an error-prone polymerase or Taq polymerase, in the presence of manganese ions.

[0010] In some embodiments, reading the code comprises amplifying the code and sequencing the code. In some embodiments, reading the code comprises binding of a specific antibody to the target, the antibody being linked to a primer for reading the code. The antibody and primer can be linked by being bound to the same solid support. Reading can utilize a primer that is at least partially complementary to the anchor or the last subcode.

[0011] In some embodiments, multiple anchors containing different reactive groups are attached to the tissue sample. The anchors can include reactive groups that react with the target in the presence of an electric field. The subcodes comprise non-nucleotide entities, and the codes can be read by mass spectrometry.

[0012] The assembled cord can be a linear or branched polymer.

[0013] In some embodiments, the tissue sample is embedded in a stabilizing matrix, such as agarose or a hydrogel or another matrix that is transparent to the wavelengths used for cross-linking.

[0014] In some embodiments, covalently binding the subcode to the portion of the tissue sample comprises masking the remainder of the tissue sample. [Brief explanation of the drawings]

[0015] [Figure 1] 1 illustrates a workflow for detecting the presence and location of a target in a tissue sample. [Figure 2] 1 shows the sequential assembly of the spatial code in situ in a tissue sample. [Figure 3] 1 illustrates the process of forming a unique address within a sample. [Figure 4] 1 shows detection of spatially labeled protein targets using antibodies. [Figure 5]It shows how to use splints to assemble code from subcodes. DETAILED DESCRIPTION OF THE INVENTION

[0016] definition The following definitions will aid in understanding the present disclosure. The term "sample" refers to any composition that contains or is suspected to contain the target to be analyzed. The term "tissue sample" refers to a sample having a three-dimensional structure. This includes solid tissue samples isolated from an individual, such as organ or tumor biopsies. The term also includes environmental samples, such as microbial biofilms, and samples of in vitro cultures established from cells taken from an individual, including formalin-fixed, paraffin-embedded tissue (FFPET).

[0017] The term "nucleic acid" refers to a polymer of nucleotides (e.g., ribonucleotides and deoxyribonucleotides, both natural and unnatural), including DNA, RNA, and subcategories thereof, such as cDNA and mRNA. Nucleic acids can be single-stranded or double-stranded and generally contain 5'-3' phosphodiester bonds, although in some cases, nucleotide analogs can have other linkages. Nucleic acids can include natural bases (adenosine, guanosine, cytosine, uracil, and thymidine) and unnatural bases. Some examples of unnatural bases include those described, for example, in Seela et al. (1999) Helv. Chim. Acta 82:1640. Unnatural bases can have specific functions, such as increasing the stability of nucleic acid duplexes, inhibiting nuclease digestion, or blocking primer extension or chain polymerization.

[0018] The terms "polynucleotide" and "oligonucleotide" are used interchangeably. A polynucleotide is a single-stranded or double-stranded nucleic acid. Oligonucleotide is a term sometimes used to describe shorter polynucleotides. Oligonucleotides are prepared by any suitable method known in the art, including direct chemical synthesis, as described, for example, in Narang et al. (1979) Meth. Enzymol. 68:90-99; Brown et al. (1979) Meth. Enzymol. 68:109-151; Beaucage et al. (1981) Tetrahedron Lett. 22:1859-1862; Matteucci et al. (1981) J. Am. Chem. Soc. 103:3185-3191. Oligonucleotides may also be used as described in U.S. Patent Application No. 15 / 135,434, filed April 21, 2016, entitled "Devices for Producing and Producing Oligonucleotides." The oligonucleic acid libraries can be prepared by the inkjet printing method described in U.S. Patent Application No. 15 / 015,059, filed February 3, 2016, entitled "Methods and devices for de novo oligonucleic acid assembly."

[0019] The term "primer" refers to a single-stranded oligonucleotide capable of hybridizing to a sequence (the "primer binding site") in a target nucleic acid and acting as a point of initiation of synthesis along a complementary strand of nucleic acid under conditions suitable for such synthesis.

[0020] The term "ligation" refers to a condensation reaction that joins two nucleic acid strands, in which the 5'-phosphate group of one molecule reacts with the 3'-hydroxyl group of another molecule. Ligation is an enzymatic reaction typically catalyzed by ligase or topoisomerase. Ligation can join two single strands to produce one single-stranded molecule. Ligation can also join two strands that each belong to a double-stranded molecule, thus joining two double-stranded molecules. Ligation can also join both strands of a double-stranded molecule to both strands of another double-stranded molecule, thus joining two double-stranded molecules. Ligation can also join the two ends of a strand within a double-stranded molecule, thus repairing a nick in a double-stranded molecule.

[0021] The term "code" as used herein refers to a sequence of two or more oligonucleotides or non-nucleotide entities ("subcodes") that are assembled on a target molecule and mark the target molecule in a sample. The term "subcode" as used herein refers to an oligonucleotide or non-nucleotide entity that can be combined with other such entities to assemble a code. The terms "barcode" and "nucleic acid barcode" refer to a code composed of nucleic acids. Barcode sequences can be detected and identified. Barcodes can be incorporated into various nucleic acids. Barcodes are sufficiently long, e.g., 2, 5, 20 nucleotides, so that barcodes and the nucleic acids that incorporate them can be distinguished from each other in a sample.

[0022] The term "crosslinking" refers to a chemical reaction that forms a covalent bond between two polymers, requiring an external stimulus such as radiation energy of various wavelengths (heat, light, ultrasound) or a change in pH.

[0023] The term "quantum barcoding" or "QBC" refers to a process by which one or more targets in individual cells in a mixture of cells can be labeled with a unique nucleic acid code. The process comprises the step of stepwise assembly of a unique code from subcodes in situ on (or within) a cell or tissue. An example of QBC is described in U.S. Patent Application Publication No. 13 / 981,711, filed April 15, 2016.

[0024] The terms "target" and "target molecule" refer to a molecule of interest to be detected or quantified by the methods described herein. A target can be a nucleic acid sequence or a protein. The term target includes all variants of a target molecule, such as one or more variants of a mutant and a wild-type variant. The term "target sequence" refers to a nucleic acid sequence in a sample to be detected or quantified.

[0025] The term "amplification" refers to the process of making additional copies of a target nucleic acid. Amplification can have two or more cycles, e.g., multiple cycles of exponential amplification. Amplification can have only one cycle (making a single copy of the target nucleic acid). The copies can have additional sequences, e.g., sequences present in the primers used for amplification. Amplification can also produce copies that are only single-stranded (linear amplification) or that are predominantly single-stranded (asymmetric PCR).

[0026] The term "sequencing" refers to any method of determining the sequence of nucleotides in a nucleic acid.

[0027] The present invention is a process that iteratively encodes the location information of targets in a two-dimensional or three-dimensional sample so that the content of that region can be determined in subsequent analysis. can be encoded in the form of a nucleic acid code, e.g., a DNA sequence that can be determined by conventional sequencing methods. Positional information can also be encoded in the form of a chemical code, e.g., as a series of chemical adducts that can be decoded by mass spectrometry.

[0028] Currently, the primary means of obtaining spatial (2D or 3D) information about tissues, biofilms, or cells requires microscopic examination of cells bound with labeled probes or antibodies. The number of distinguishable fluorescent labels is very limited, typically three or four at a time, up to 40 or 50 with repeated staining. Isotopic labeling allows for up to 40 isotopes, albeit at a very high cost. The present invention is a method for spatially labeling targets with novel modular codes. This method provides an essentially unlimited number of labels (codes) that can be detected in a single experiment. This method allows for the qualitative and quantitative detection of DNA, RNA, and protein targets.

[0029] One alternative method for measuring the spatial distribution of targets (e.g., mRNA or protein) involves directly annealing fixed tissue samples to an array of barcoded reverse transcription (RT) primers followed by transcription, sequencing, and computational reconstruction. See Stahl, PL et al. (2016) Visualization and analysis of gene expression in tissue sections by spatial transcriptomics, Science 353:78-82. A major drawback of that method is bleeding, e.g., annealing of primers outside their grid locations. The present invention overcomes this problem by utilizing a laser that can be focused on the desired region with micron-level resolution.

[0030] The tissue sample used in the novel methods is a fragment of tissue derived from an organism, subject, or patient. In some embodiments, the sample can include a fragment of solid tissue or solid tumor derived from an organism or patient, e.g., by biopsy or surgical resection. In some embodiments, to facilitate the methods of the present invention, the tissue sample can be captured in an inert matrix (e.g., an agarose gel matrix). See Andersson et al. (2006) Analysis of protein expression in cell microarrays: a tool for antibody-based proteomics, J. Histochem. Cytochem. 54(12):1413-23. Epub 2006 Sept. 6. Tissue samples can also be embedded in hydrogels, which are three-dimensional networks composed of hydrophilic polymers crosslinked via covalent bonds or held together through physical intra- and intermolecular attractions. See Hoffman AS, (2001) Hydrogels for biomedical applications, Ann NY Acad Sci., 944:62-73. In some embodiments, the hydrogel is a 3D structure consisting of a network of polymers (e.g., acrylamide or bisacrylamide) linked to cellular molecules via formaldehyde. In some embodiments, polymerization is initiated by adding an initiator to the tissue sample. See, e.g., Chung, K. et al. (2013). Structural and molecular interrogation of intact biological systems, Nature, 497(7449), 332. In some embodiments, after the polymerization process, non-protein molecules (e.g., lipids) are washed away or eluted (e.g., by electrophoresis), leaving the target molecules (e.g., proteins and nucleic acids) in place. The gel or matrix must be transparent to the frequencies used to crosslink or otherwise bind the codes by methods described further below.In some embodiments, the matrix is ​​opaque to visible light but is transparent to other types of radiation used in the method.

[0031] In some embodiments, the tissue sample is preserved as a clinical patient sample according to current medical practice. In some embodiments, the sample is fresh frozen at -20°C or below, e.g., -80°C. In other embodiments, the clinical tissue sample is preserved by fixation in formalin and embedding in paraffin (FFPE). In such embodiments, the sample requires deparaffinization by heat or detergent according to methods known in the art.

[0032] In some embodiments, the tissue sample is a microbial colony or biofilm, the only requirement being that the sample be transparent to the electromagnetic frequency of the radiation used to cross-link the cords, as described further below.

[0033] In some embodiments, the anchor is a nucleic acid that specifically binds (hybridizes) to a nucleic acid target. In these embodiments, the anchor comprises a region of complementarity with the target nucleic acid (DNA or RNA). In some embodiments, the anchor is at least partially complementary to the nucleic acid target in the tissue sample to enable specific target recognition. In some embodiments, the target is mRNA. In some embodiments, the anchor comprises a poly-T or poly-dT, poly-U or poly-dU sequence, or any other homopolymer sequence capable of forming a stable hybrid with a poly-A sequence. In some embodiments, the anchor comprises a random sequence. In some embodiments, the anchor is a combination of two or more of a poly-A binding sequence, a random sequence, and a target-specific sequence.

[0034] In some embodiments, the anchor is a nucleic acid aptamer selected to specifically bind to a non-nucleic acid target (see Oliphant, AR; et al., (1989) "Defining the sequence specificity of DNA-binding proteins by selecting binding sites from random-sequence oligonucleotides: analysis of yeast GCN4 proteins," Mol. Cell. Biol. 9(7):2944-2949). Methods for producing and improving nucleic acid aptamers through a process called SELEX are described, for example, in U.S. Patent No. 5,475,096, U.S. Patent No. 5,270,163, U.S. Patent No. 5,567,588, U.S. Patent No. 5,660,985, U.S. Patent No. 5,580,737, U.S. Patent No. 5,496,938, U.S. Patent No. 9,382,533, U.S. Patent No. 8,975,026, U.S. Patent No. 8,975,388, U.S. Patent No. 8,404,840, U.S. Patent No. 7,964,356 and U.S. Patent No. 7,947,447. Aptamers can be chemically linked to targets through the chemi-SELEX process described in U.S. Patent No. 5,705,337. Aptamers can contain modified nucleotides with substitutions at ribose, phosphate and base positions (U.S. Patent No. 5,580,737). Aptamers can be engineered to contain photoreactive functional groups that can bind and photocrosslink to their targets (photo-SELEX) See U.S. Patent Nos. 5,763,177, 6,001,577, 6,291,184, 6,458,539, and 8,409,795.

[0035] The nucleic acid anchor further comprises a region of complementarity that allows for code assembly on the anchor. In some embodiments, the region of complementarity is to a first subcode, such that the first subcode anneals directly to the anchor and additional subcodes are added, as described herein. In other embodiments, the region of complementarity is to an annealing primer that anneals to both the anchor and the first subcode, and additional subcodes are added. Additional subcodes are added as described herein.

[0036] In some embodiments, the anchor is not complementary to the target, i.e., the target is not a nucleic acid. In some embodiments, the anchor binds to all available targets in a tissue sample. For example, the anchor can be a nucleic acid that binds to multiple or all proteins in the sample without specific binding. In the absence of specific recognition, binding is achieved by a facilitated chemical reaction between nucleotides and amino acids. In some embodiments, the nucleic acid is directly bound to the protein via irradiation. In some embodiments, UV irradiation (e.g., wavelengths at or near 250 nm) results in photoaddition between thymidine and the ε-amino group of lysine. See Saito I. and Matsuura T. (1985) Chemical aspects of UV-induced crosslinking of proteins to nucleic acids. Photoreactions with lysine and tryptophan, Acc. Chem. Res., 1985, 18(5), pp. 134-141. In some embodiments, the anchor oligonucleotide contains photoactivatable nucleotides, which allow for highly efficient crosslinking with lasers emitting at various wavelengths. Hafner, M. et al. (2010) Transcriptome-wide Identification of RNA-Binding Protein and MicroRNA Target Sites by PAR-CLIP, Cell 141-129. In some embodiments, the anchor oligonucleotide comprises one or more photoactivatable nucleotides having a modified base selected from 4-thiouridine, 5-bromouridine, 5-iodouridine, and 6-thioguanosine.

[0037] The present invention includes the use of a code composed of subcodes. In some embodiments, the subcodes are nucleic acids. The present invention provides a library of synthetic nucleic acid subcodes, each having a unique sequence that is distinguishable from other subcodes. The subcodes are sequences that do not form stable bonds with any nucleic acid sequence in a tissue sample. Thus, in some embodiments, the subcode sequences are selected so that they are not complementary to any region of a genome of interest, for example, the genome of the organism from which the tissue sample is derived or the genome of a target infectious pathogen for the presence of which the tissue sample is being interrogated.

[0038] After a first subcord is bonded to the anchor, one or more additional subcords are joined together in an ordered manner to form a cord bonded to the anchor. One or more steps of adding the subcords to the cord include masking and cross-linking to mark the spatial location of the subcord. As described herein, the bonding step includes weak bonding, followed by cross-linking in the unmasked portions, and washing away the uncross-linked subcords in the masked portions.

[0039] In some embodiments, annealing primers hybridize and bridge two oligonucleotide subcodes in each round, allowing the subcodes to be linked and then joined by ligation. Enzymatic ligation can utilize DNA or RNA ligases, such as T4 DNA ligase, T4 RNA ligase, Thermus thermophilus (Tth) ligase, Thermus aquaticus (Taq) DNA ligase, or Pyrococcus furiosus (Pfu) ligase. Non-enzymatic or chemical ligation can utilize activators and reducers, such as carbodiimide, cyanogen bromide (BrCN), imidazole, 1-methylimidazole / carbodiimide / cystamine, N-cyanoimidazole, dithiothreitol (DTT), and ultraviolet light.

[0040] In other embodiments, the subcodes are joined together using CLICK chemistry. (El-Sagheer et al. (PNAS, 108:28, 11338-11343, 2011).

[0041] In some embodiments, binding to the annealing primer leaves a gap of one or more nucleotides between the subcodes that requires fill-in by a nucleic acid polymerase before ligation. In some embodiments, a phosphorylation step is performed before ligation. In some embodiments, the polymerase is an error-prone or low-fidelity DNA polymerase. Error-prone polymerases incorporate nucleotide variations (errors) that differ between codes and include additional diversity between codes. In some embodiments, the polymerase is a polymerase as described by Ohmori et al. (2001) The Y-family of DNA polymerases, Mol. Cell 8:7-8. Y-family polymerases are characterized by the lack of detectable 3' to 5' proofreading exonuclease activity and replicate undamaged DNA in vitro with low fidelity and poor processivity. In some embodiments, the polymerase is a bacterial, archaeal, or eukaryotic error-prone polymerase. In some embodiments, the polymerase is Taq polymerase. In some embodiments, the fidelity of Taq polymerase is enhanced by the Mn 2+ or by the presence of Mg 2+ It is further reduced by increasing the concentration of the ion.

[0042] In some embodiments, the subcode is added to a splint oligonucleotide linked to an anchor (Figure 5). The splint is linked to an anchor (A) and, optionally, to an epitope-specific barcode (ESB). The splint provides a landing pad (large open square) for subcode (SC) annealing. Because the subcodes have distinct sequences, the entire subcode is not annealed to the splint. The subcodes have one or more unique sequences (small open squares). After the annealing, crosslinking, and washing steps described herein, the subcodes are joined by ligation or CLICK chemistry, as described herein. In some embodiments, before joining, the subcodes must be extended by a nucleic acid polymerase. In some embodiments, the polymerase is an error-prone polymerase, which generates additional diversity within the code. In some embodiments, the final subcode is joined to an amplification primer binding site (AMP).

[0043] In some embodiments, the subcodes are non-nucleic acid entities. Any chemical moieties that can be linked together in an ordered manner by electromagnetic radiation are within the scope of the present invention. Electromagnetic radiation includes microwaves, terahertz, X-rays, and gamma rays. In some embodiments, sound waves are used. In some embodiments, sound waves are used to promote the conversion of monomers to polymers, for example, by heat generation. In some embodiments, the subcodes are assembled into linear polymer codes. In other embodiments, the subcodes are assembled into branched polymer codes.

[0044] Examples of non-nucleotide subcodes include sugar entities, small, definable organic molecules with side groups of discriminatory mass, and monomeric subunits of any of several different, well-understood polymers. Any small organic moiety (subunit) in which intersubunit bonds are more easily broken than intrasubunit bonds can be used as a subcode. While it is possible to work with monomers that do not possess this latter characteristic, efficiency is improved if intersubunit monomers are more easily connected to each other than broken down in the mass spectrometry process. The process of connecting the monomers to the polymer is achieved by photon-based accelerated bond formation.

[0045] In some embodiments, the anchors and subcords are free to diffuse throughout the tissue sample. In other embodiments, the anchor and / or subcord are introduced into the tissue sample using an electric field.

[0046] In some embodiments, binding of anchors to targets, binding of subcodes to anchors, or binding of subcodes to each other involves crosslinking after hybridization. In some embodiments, hybridization allows the anchor to be positioned on the target, but is not sufficient to form a bond, e.g., a bond that can withstand stringent washing. In some embodiments, after the hybridization and crosslinking steps (described herein below), merely hybridized but uncrosslinked anchors or subcodes, and any unhybridized anchors of the subcodes, are washed away in a washing step. In some embodiments, the washing step comprises contacting the sample with a washing buffer comprising one or more of a salt buffer, e.g., saline sodium citrate (SSC), a detergent, e.g., SDS, and an aprotic solvent, e.g., DMSO.

[0047] In some embodiments, attachment of anchors to tissue samples, attachment of subcodes to anchors, or attachment of subcodes to subcodes in existing codes utilizes photoreactive groups. In some embodiments, two or more photoreactive groups are used, each group being reactive to a unique wavelength, including light beams, terahertz frequency beams, and X-rays. In some embodiments, attachment is by irradiation with acoustic waves, which promote a chemical reaction between the reactive groups.

[0048] In some embodiments, an LED light source can be used. For example, a light source emitting light of an appropriate wavelength (e.g., 365 nm (XeLED-Ni1UV-R3-365, Xenopus Electronix)) can be used to irradiate the tissue. The appropriate duration and distance of irradiation can be experimentally determined, for example, from 2 cm for 30 seconds. The non-irradiated portion of the tissue can be blocked, for example, by aluminum foil or any other light-blocking material. In some embodiments, UV light emitted by a metal halide lamp (X-Cite, Lumen Dynamics) passed through a 20x microscope objective can be used. The UV energy can range from 50 to 10 mW for 10 seconds to 1 minute, depending on the chemical bonds formed.

[0049] In some embodiments, in mass spectrometry-based measurement systems, the sequence of the polymer-protein attachment point is determined by mass spectrometry-based sequencing.

[0050] In some embodiments, the crosslinking is photocrosslinking. Photocrosslinking can be achieved using a laser of a specific wavelength. In embodiments in which the anchors include reactive groups that react with the target in the presence of laser irradiation, multiple anchors are attached to multiple targets using different reactive groups that react in the presence of irradiation with light of different wavelengths.

[0051] In some embodiments, portions of the tissue sample are isolated by masking. A portion of the tissue sample is masked while a subcode is attached to the remaining (unmasked) tissue sample. Masking can be achieved by standard masking systems used in photolithography. Virtual masking or photolithography can be performed via addressable laser systems that are well established in the technical community.

[0052] In some embodiments, the laser light is directed by a maskless method utilizing a digital micromirror device (DMD). Singh-Gasson et al. (1999) Maskless fabrication of light-directed oligonucleotide microarrays using a digital micromirror device (DMD). al micromirror array, Nature Biotechnology 17:974. In this embodiment, the tissue sample forms an addressable array in which subcodes are bridged to addressable regions in the 2D sample.

[0053] In some embodiments, the method utilizes a laser instrument. Lasers offer the advantage of precisely focusing on an area of ​​the tissue sample so that the barcoding is precisely associated with the area and does not have the bleeding experienced with prior art methods. Typical lasers available in the art allow for micron-level resolution, allowing for the differentiation of up to 1,000,000 to 2,500,000 spots on a microscope slide.

[0054] In some embodiments, the laser is a programmable laser. Using the programmable laser, the tissue sample can be fabricated into a virtual addressable array, where each location correlates with the timing of its illumination by the programmable laser and the sequence of the subcode that is currently attached (crosslinked) to the code. Subsequent analysis of the nucleic acid sequence of the barcode reveals the subcode strand, which translates into the location of the barcode within the virtual addressable array formed by the programmable laser on the tissue sample.

[0055] Furthermore, each barcode is directly or indirectly associated with a target-specific sequence. In one embodiment, the barcode is directly assembled on an anchor complementary to the target, which is a nucleic acid. In this embodiment, sequencing the barcode also comprises sequencing the target nucleic acid. In another embodiment, the target, which is a protein, is recognized by an antibody associated with an epitope-specific barcode (ESB). The ESB is a nucleic acid sequence associated with the target-specific antibody. In this embodiment, sequencing the barcode also comprises sequencing the ESB substantially associated with the target protein. Thus, the use of an addressable array with barcodes allows for the determination of the coordinates of each barcoded target within the array and within a tissue sample.

[0056] The codes correlate with their location within the tissue sample. The codes are chronologically bound to targets in each portion of the tissue sample via irradiation by a programmable laser guided by executable code, correlating the time a particular code was bound and the sequence of the code with the portion of the sample processed (e.g., irradiated by a laser) at that time. (Figure 1: Subcode 1 is added first in step 1 to anchor the "left" half, thus marking the left half; then subcode 2 is added first in step 2 to anchor the "right" half, thus marking the right half.) Codes are assembled from two or more subcodes, as described in more detail below. Figure 2 shows a diagram of the assembly of codes 1-1, 1-2, 2-1, and 2-2, as well as subsequent codes 1-1-1, 1-1-2, etc. At the end of the labeling process, all targets (e.g., proteins or nucleic acids) within the sample are labeled with a code unique to the target's location within the sample. The assembled code is unique to each portion of the tissue sample and serves as a positional or spatial marker for that portion of the tissue sample. The length of the code reflects the resolution (size of the smallest distinguishable area) within the sample.

[0057] In some embodiments, the code is a nucleic acid, and reading the code comprises sequencing the nucleic acid. In some embodiments, the 5'-end of the assembled code is the proximal end attached to the anchor, and the 3'-end is the distal end. In other embodiments, the 3'-end of the assembled code is the proximal end attached to the anchor, and the 5'-end is the distal end.

[0058] In some embodiments, the assembled code comprises a universal domain that contains elements for downstream analysis of the code, hi some embodiments, the universal domain comprises binding sites for primers, e.g., amplification primers or sequencing primers.

[0059] In some embodiments, the code is amplified prior to sequencing. In such embodiments, amplification primers bind to primer binding sites within the assembled code. The primer binding site is a sequence that is at least partially shared among all subcodes. In some embodiments, the primer binding site is at the distal end of the code (farthest from the target). The distal end can be the 3'-end to which an amplification primer can bind and initiate a first round of amplification. In some embodiments, the primer binding site is at the proximal end of the code (closest to the target). The proximal end can be the 3'-end to which an amplification primer can bind and initiate a first round of amplification. In some embodiments, amplification uses a primer that is at least partially complementary to an anchor (the anchor comprises at least a portion of the primer binding site).

[0060] In some embodiments, amplification serves as the target recognition step of the method. In some embodiments, all targets in a sample (i.e., all proteins in the sample) are spatially labeled. Detection of specific protein targets can be achieved using antibodies. Antibodies can be obtained from any suitable source, including recombinantly expressed antibodies and antibodies from various animal species. A wide variety of pre-made and custom-made antibodies can be obtained from commercial sources. In some embodiments, the antibody is linked to a nucleic acid. In some embodiments (FIG. 4), both the antibody and amplification primers are conjugated to the same particle on the slide support. In some embodiments, the primers are directly conjugated to the antibody. Any suitable method for binding a nucleic acid to a protein, including an antibody, is encompassed by the methods of the present invention. See, for example, Gullberger et al., PNAS 101(22):228420-8424 (2004); Boozer et al., Analytical Chemistry, 76(23):6967-6972 (2004) and Kozlov et al., Biopolymers 5:73(5):621-630 (2004). In some embodiments, the antibody-anchor is attached to the nucleic acid using a tadpole, as described in Nolan, Nature Methods 2, 11-12 (2005). In some embodiments, the antibody is conjugated to the nucleic acid using SpyTag-SpyCatcher technology, wherein the antibody and nucleic acid comprise SpyTag and SpyCatcher. See Reddington and Howarth (2015) Secrets of a covalent interaction for biomaterials and Biotechnology, SpyTag and SpyCatcher, Current Opinion in Chemical Biology 2015, 29:94-99. In yet other embodiments, the primer is directly conjugated to the antibody and also directly conjugated to the particle or capture moiety (e.g., biotin) of the solid support.

[0061] Antibody binding functions to deliver amplification primers and allow amplification of the spatial code. The code (or anchor) contains a primer binding site that is at least partially complementary to the primer, but the complementarity is not sufficient for both primer annealing and primer extension to occur. For primer extension to occur, the antibody must bind to the target (Figure 4). Without antibody binding, the code is not amplified and cannot be detected. Therefore, only spatial codes bound to the target of interest are amplified and detected.

[0062] The present invention comprises a step of reading the code, anchor and any information related to the target directly or indirectly by nucleic acid sequencing. This can be performed by any method known in the art. High-throughput single-molecule sequencing that can read circular target nucleic acids is particularly advantageous. Examples of such technology include the SOLiD platform (ThermoFisher Scientific, Foster City, California), the Helioscope fluorescence-based sequencing device (Helicos Biosciences, Cambridge, Massachusetts), the Pacific Bioscience platform (Pacific Biosciences, Menlo Park, California) that utilizes SMRT, or the platform that utilizes nanopore technology, such as that manufactured by Oxford Nanopore Technologies (Oxford, UK), or Roche Sequencing Solutions (Roche Genia, Santa Clara, California), reversible terminator sequencing by synthesis (SBS) (Illumina, San Diego, California), and any other currently existing or future DNA sequencing technology, with or without sequencing by synthesis. The sequencing step can utilize platform-specific sequencing primers. The binding site of these primers can be introduced into the code, for example, by being part of the final subcode or amplification primer used to amplify the code before sequencing.

[0063] In some embodiments, the code is a combination of two or more non-nucleic acid chemical entities, and reading the code comprises analysis by mass spectrometry via time-of-flight (TOF) determination. In this embodiment, encoding is achieved through photon-activated polymerization of subunits. While nucleic acids can be read by high-throughput sequencing, other methods are required to determine the sequence of polymerized or conjugated non-nucleotide chemical moieties. This is a well-characterized process in mass spectrometry sequencing of proteins. The process enabled here simply generalizes this to reading polymer subunits. An advantage of other polymer subunits is that it can unexpectedly reveal polymer subunits that are more easily distinguishable by mass spectrometry than amino acid-based polymers.

[0064] In some embodiments, the present invention provides for simultaneous detection of multiple target molecules (multiplex assays). For example, multiple antibody anchors can be added to a tissue sample. Similarly, multiple nucleic acid anchors can be added to a tissue sample. The multiple nucleic acid anchors may not share regions of complementarity to prevent in vitro interactions between the anchors. Complementarity, including partial complementarity, can be excluded experimentally or using software such as BLAST.

[0065] This method can detect and distinguish potentially millions of unique locations within a tissue sample. Longer barcodes (formed by more assembly rounds) can distinguish more locations. The use of multiple reactive groups available for cross-linking allows for even more locations to be distinguished within a tissue sample. For each reactive group, the intensity can vary, thus effectively dividing each set into up to three separate sets. The following calculations allow one skilled in the art to determine the number of barcodes (and rounds of code assembly) required to achieve the desired spatial resolution within a given tissue sample.

[0066] If five different wavelengths are used, then in a single round, the number of positions distinguished in 3D is 5 3 = 125. If two rounds of code assembly are used, the number of distinct positions in 3D is 125 2 = 15,725. For 3 rounds, the number of positions is 125 3 =1.95×10 6 If five different wavelengths are used, each with up to three distinguishable intensities, then in a single round the number of distinguishable positions in 3D is (5 × 3) 3 = 3,375. If two-round assembly is used, it is distinguished in 3D. The number of positions is 3,375 2 =1.1×10 7 For 3 rounds, the number of positions is 3,375. 3 =3.84×10 10 A typical cell is 10 3 It has a volume of cubic microns. An exemplary 1 mm 3 Tissue sections were 10 9 In this example, two or three rounds of code assembly may be sufficient to obtain resolution at the cellular level.

[0067] The present invention comprises correlating a code sequence to a portion of a tissue sample to obtain the location of the target within the tissue sample. In some embodiments, the tissue sample is oriented to allow maximum information content retrieval from the region of the tissue that is of most interest to the researcher. The tissue is oriented within a gel matrix that allows laser addressing and interrogation of the sample. In some embodiments, the tissue sample is marked with a positional marker. For example, a positional (fiducial) marker can be introduced via a bead conjugated to a known nucleic acid code. The code serves as a preprogrammed address. The marker can be introduced either by hand or by robotic placement. In other embodiments, such a positional marker can be introduced by targeting a known sequence with a complementary probe that also contains a known nucleic acid code. Cross-linking the code to a known spot in the sample forms the positional marker.

[0068] In the context of the present invention, positional information regarding a particular code is obtained from the point in time when a portion of the sample is irradiated, resulting in binding of the code to the sample (Figure 2).

[0069] In some embodiments, the present invention is a method for simultaneously detecting the presence and spatial location of a target in a three-dimensional or two-dimensional sample.The tissue sample can be a 2D or 3D eukaryotic (animal, human, plant, or fungal) tissue sample (for example, on a microscope slide), or a prokaryotic sample such as a microbial biofilm with a 2D or 3D structure.The sample must be transparent or made transparent so that it can be addressed by the laser system used in the present invention.

[0070] In the first step, the sample is contacted with an anchor that binds to a target in the tissue sample. The target can be a protein or a nucleic acid (RNA or DNA). Thus, the anchor is a nucleic acid probe, a non-specific nucleic acid, or a protein-specific nucleic acid aptamer. In some embodiments, the anchor binds to one target with specificity (e.g., an aptamer, a nucleic acid probe, or an antibody). In other embodiments, the anchor binds to multiple targets non-specifically. Aptamer binding conditions are described, for example, in Deng et al. (2014) Aptamer binding assays for proteins: The thrombin example—A review, Analytica Chim. Acta, 837:11-15. Nucleic acid probe binding conditions are those used in in-situ hybridization (ISH). See Wilkinson, E.G. et al. (1999) In Situ Hybridization: A Practical Approach (Practical Approach Series) 2nd Edition, Oxford University Press.

[0071] The anchor contains a reactive group, e.g., a photoactive group or another type of group that can be activated by radiation, and is crosslinked to the target via irradiation having a wavelength that activates the reactive group.

[0072] The anchors serve as sites for assembly of the code. The code is assembled in situ on each anchor molecule. The code is assembled from two or more subcodes through two or more rounds of assembly. Assembly involves the binding and cross-linking of the subcodes, followed by This involves removing (washing) unbound subcodes. During each round of assembly, a portion of the tissue sample is masked to prevent crosslinking in the masked portion (Figure 1). The sample is contacted with a first subcode (e.g., 1), which allows subcode 1 to bind to anchors in the first portion of the tissue sample. The second portion is masked. Next, the sample is contacted with a second subcode (e.g., 2), which allows subcode 2 to bind to anchors in the second portion of the tissue sample that do not overlap with the first portion (Figure 1). Here, the first portion is masked. For example, the first and second portions can be the left and right halves of a microscope slide. In the next round, the sample is contacted with the next subcode (e.g., 1 or 2), which allows the subcode to bind to subcodes A and B in the third portion of the tissue sample that partially overlaps the first and second portions, forming a two-part code (e.g., 11 and 21) on it (Figure 2). The fourth portion is masked. The sample is then contacted with a subsequent subcode (e.g., 1 or 2) that allows the subcode to bind to subcodes 1 and 2 of a fourth portion of the tissue sample, which partially overlaps the first and second portions but not the third portion, forming a two-part code (12 and 22) thereon. The third portion is masked. For example, the third and fourth portions can be the top and bottom halves of a microscope slide. This divides the tissue sample into four regions, each with a unique address. If necessary, more subcodes can be added and the tissue sample can be further divided into regions, and the steps can be repeated so that each region has a unique address marked with a unique code (i.e., a unique combination of subcodes). For example, in subsequent steps, portions of the sample are exposed and masked to allow the addition of longer codes (e.g., 111, 211, 112, 212, etc.) that each correspond to a smaller portion of the sample (Figures 2-3). The assembled code is read to determine the spatial location of each anchor within the tissue sample.

[0073] Figure 1 shows a workflow for detecting the presence and location of a target in a tissue sample. In this example, an anchor is conjugated to a first subcode. The anchor comprises a reactive group (in this illustration, a photoreactive group) so that the anchor can be crosslinked to the target. A tissue sample in the form of a tissue slide is contacted with an anchor molecule (α) conjugated to the reactive group so that each target (e.g., each protein) in the tissue sample is modified by the anchor molecule.

[0074] Subcodes can be linked to each other through regions of complementarity to other subcodes (i.e., subcodes from previous rounds). In some embodiments, subcodes do not anneal to each other, but rather to the annealing primer to which both adjacent subcodes anneal. In other embodiments, subcodes anneal to the splint oligo. However, in all embodiments, cross-linking is required to form a stable bond between the two subcodes, the annealing primer or subcode and splint. Uncross-linked subcodes are washed away in a wash step.

[0075] Figure 2 shows the sequential assembly of a spatial code in situ in a tissue sample. Each time, a portion of the slide is masked, while the next subcode molecule is added in the remaining active area. As shown, the masked and active areas are not contiguous. As shown, at each address, a code consisting of two subcodes, then three subcodes, is assembled. The attachment of nucleic acid subcodes can include one or more of nucleic acid chain extension, gap fill, and ligation. For example, when annealing primers are used, the nucleic acid ends can anneal directly adjacent to each other, allowing ligation of adjacent nucleic acids, such as the 5' and 3' ends of the subcodes. In other embodiments, gaps exist between the ends of the nucleic acids. The gaps are formed by the addition of a nucleic acid polymer. The 3' end of the extended strand is ligated to the 5' end of an adjacent nucleic acid, e.g., a subcode.

[0076] In some embodiments, the anchor binds nonspecifically to the target. For example, all proteins in a sample may be bound to the oligonucleotide anchor via a nucleotide-amino acid bridge. Detecting the protein of interest in this setting requires a separate detection step. For example, the anchor-conjugated protein target can be detected with an antibody. The antibody is conjugated to a nucleic acid that interacts with the code constructed on the target molecule (anchor-conjugated protein) to detect the target molecule. As shown in Figure 4, the antibody is conjugated to an extendible oligonucleotide (γ) complementary to the anchor region (α). The extension of the oligonucleotide allows for the duplication or copying and amplification of the code (bar region). Copying and amplification only occurs when the antibody binds to its target, ensuring the specificity of detection. The antibody binding conditions are those that apply to staining tissues with antibodies in either flow cytometry or immunohistochemistry (S. Hockfield et al., Selected Methods for Antibody and (See, Nucleic Acid Probes, Cold Spring Harbor Lab Press (1993)). In some embodiments, detection is qualitative, such that the presence of an amplicon indicates the presence of the target. In some embodiments, detection is quantitative, e.g., the number of different unique barcodes amplified indicates the number of cells in the tissue sample that contain (express) the target protein. In some embodiments, detection is quantitative, e.g., the amount of amplicons with the same barcode indicates the amount (expression level) of the target protein in a particular cell in the tissue sample.

[0077] Although the present invention has been described in detail with reference to specific embodiments, it will be apparent to those skilled in the art that various modifications can be made within the scope of the invention. Accordingly, the scope of the present invention should not be limited by the embodiments described herein, but rather by the claims set forth below.

Claims

1. 1. A method for simultaneously detecting the presence and spatial location of a target in a tissue sample, comprising: a. covalently attaching an anchor to a target in said tissue sample via a reactive group; b. assembling a code from a set of subcodes on said anchor, i. contacting the sample with a first sub-code and allowing the first sub-code to covalently bond to the anchors in a first portion of the tissue sample to form a code thereon; ii. contacting the sample with a second sub-code and allowing the second sub-code to covalently bond to the anchors in a second portion of the tissue sample that does not overlap the first portion to form a code thereon; iii. assembling by a method including repeating said pair of steps i-ii one or more times, wherein in each repetition, said portion of said tissue sample contacted in said first step does not overlap said portion of said tissue sample contacted in said second step, and said added sub-codes join and extend said existing sub-codes, thereby forming a code marking each portion of said tissue sample; c. reading the code assembled on the anchor in step iii, thereby detecting the presence and location of the target in the tissue sample; A method comprising:

2. The method of claim 1 , wherein the subcode is a nucleic acid.

3. The method of claim 2 , wherein the anchor is covalently attached to the target via a crosslink.

4. 4. The method of claim 3, wherein prior to covalent attachment, the subcode hybridizes to the anchor and the other subcode via regions of complementarity to the anchor and the other subcode.

5. The method of claim 3 , wherein the subcodes are covalently linked to the anchors and other subcodes via crosslinks.

6. 4. The method of claim 3, wherein the subcodes are attached to the anchors and the other subcodes via sonication, which promotes a chemical reaction.

7. The method of claim 1 , wherein the target is a protein.

8. 8. The method of claim 7, wherein the reactive group in the anchor is a thymidine attached to the target protein via a thymidine-lysine addition.

9. 2. The method of claim 1, wherein the subcodes in step ii are attached to a common linker.

10. 2. The method of claim 1, wherein the subcode in step ii is linked to an existing subcode via an annealing primer.

11. The method of claim 1 , wherein the subcodes are covalently linked by ligation.

12. 12. The method of claim 11, wherein ligation is preceded by chain extension with a polymerase.

13. The method of claim 1 , wherein reading the code in step c comprises amplifying the code.

14. The method of claim 1 , wherein reading the code in step c comprises sequencing the code.

15. The method of claim 1 , wherein reading the code in step c comprises binding of a specific antibody to the target.

16. 16. The method of claim 15, wherein the antibody is linked to a primer for reading the code.

17. 17. The method of claim 16, wherein the antibody and the primer are linked by being attached to the same solid support.

18. 2. The method of claim 1, wherein reading the code in step c utilizes a primer that is at least partially complementary to the anchor.

19. 2. The method of claim 1, wherein reading the code in step c utilizes a primer that is at least partially complementary to the last subcode.

20. The method of claim 1 , wherein the anchor is an aptamer.

21. The method of claim 1 , wherein multiple anchors with different reactive groups are attached to the tissue sample.

22. The method of claim 1 , wherein the anchor comprises a reactive group that reacts with the target in the presence of an electric field.

23. 10. The method of claim 1, wherein the subcode comprises non-nucleotide entities and the code is read by mass spectrometry.

24. The method of claim 1 , wherein covalently binding the subcode to the portion of the tissue sample comprises masking the remaining portion of the tissue sample.