Target probes with landmark sequences for localizing gene transcripts by fluorescence in situ hybridization

By employing target probes with landmark sequences and in situ amplification techniques, along with neural network image processing, the method addresses the challenges of sensitivity and accuracy in spatial transcriptomics, achieving enhanced localization and identification of gene transcripts.

WO2024167929A9PCT designated stage expired Publication Date: 2025-08-28APPLIED MATERIALS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/014627
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-06
Filing Date
2024-02-06
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing spatial transcriptomics techniques face challenges in achieving high sensitivity and accuracy in the localization and identification of gene transcripts within tissue samples, particularly due to issues with signal-to-noise ratio and non-specific binding of probes.

Method used

The use of target probes with a common landmark sequence, followed by amplification and verification with labeled detector readout probes, and the implementation of in situ ligation and amplification methodologies to enhance signal intensity and accuracy, combined with a convolutional neural network for image processing to refine spot localization.

Benefits of technology

This approach significantly improves the signal-to-noise ratio and enhances the accuracy of gene transcript localization, allowing for precise identification and characterization of nucleic acid species in cell and tissue samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024014627_28082025_PF_FP_ABST
    Figure US2024014627_28082025_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure provides a technology for improving images that are obtained by in situ hybridization of gene transcripts. A set of target probes specific for different target nucleic acids in a sample are adapted to include a common landmark sequence. After binding to the sample, the target probes are optionally amplified, and then specifically identified using labeled detector readout probes. The copy of the landmark sequence in each of the target probes, or amplification products generated therefrom, is identified using a labeled landmark identification probe. Image spots generated from the labeled detector readout probes are compared with image spots generated from the landmark identification probe to see if they correspond. Detector spots having no corresponding landmark spot are removed from the final image. When implemented with LISH Lock'n'Roll hybridization methodology, the probes generate a higher signal intensity with a better signal-to-noise ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Target probes with landmark sequences for localizing gene transcripts by fluorescence in situ hybridizationPRIORITY APPLICATIONS

[0001] This patent application claims the priority benefit of U.S. provisional patent applications 63 / 443,687; 63 / 443,694; and 63 / 443,692, all filed on February 6, 2023. The three priority applications are hereby incorporated herein in their entireties for all purposes.FIELD

[0002] The technology disclosed and claimed below relates generally to the fields of nucleic acid biology and the analysis and characterization of tissue sections. More specifically, it provides reagents and techniques for multimeric spatial transcriptomics.BACKGROUND

[0003] The spatial distribution of gene transcript in a tissue provides a research window for following interactive cell biology: specifically, what each specialized cell is doing in a heterogeneous population, and how the interaction of individual cells affects tissue biology, tissue homeostasis, and the emerging pathology of adverse conditions.

[0004] Spatial transcriptomics determines the location of newly transcribed mRNA’s with intercellular and subcellular resolution. It can be implemented using in situ hybridization and synthesis techniques. Subcellular transcriptome imaging allows quantitative measurements of both the gene expression profiles of individual, spatially localized cells and the intracellular distributions of transcription. At the tissue level, cell- specific gene expression defines cell types and cell states, the spatial organization of which is tightly coupled to both the development and function of normal tissues and to the pathogenesis and prognosis of tissue pathology from patients.

[0005] Multiplexed fluorescence in situ hybridization (mFISH) is a powerful imaging technique for subcellular localization of mRNA. A sample is exposed to multiple oligonucleotide probes that bind particular RNA species of interest to the user. These target- specific oligonucleotide probes have different labeling schemes that allow the user to distinguish different RNA species when the labeled probes are introduced to the sample and bind to complementary sequences on the target-specific oligonucleotide probes. Sequential rounds of fluorescence images are acquired with exposure to excitation light of differentwavelengths. For each given pixel, fluorescence intensities from the different images for the different wavelengths of excitation light form a signal sequence.

[0006] The detected signal sequence is then compared to a library of reference codes that includes reference codes that correspond to particular genes. The best matching reference code is used to identify an associated gene that is expressed at that pixel in the image. Spatial transcriptomics permits visualization of expression of genes of interest spatially within cells and tissues. Depending on the technology used, mFISH is capable of profiling hundreds to thousands of RNA molecules in single cells. Spatially resolved RNA profiling of individual cells can be done for a range of gene transcripts with high accuracy and high detection efficiency.

[0007] The value of the analysis depends on sensitivity and accuracy of the hybridization reaction, and interpretation of the images obtained therefrom. The description that follows provides a combination of technologies that can be used separately or together to facilitate the process of spatial transcriptomics and improve the results obtained.SUMMARY

[0008] This disclosure provides a technology for improving images that are obtained by in situ hybridization of gene transcripts. A set of probes specific for different target nucleic acids in a sample are adapted to include a common landmark sequence. After binding to the sample, the target probes are optionally amplified, and then specifically identified using labeled detector readout probes. The copy of the landmark sequence in each of the target probes acts as a ground truth reference, identified using a labeled landmark identification probe. Image spots generated from the labeled detector readout probes are compared with image spots generated from the landmark identification probe to see if they correspond. Detector spots having no corresponding landmark spot are removed from the final image. When implemented with a particular form of in situ ligation and amplification methodology, the probes generate a higher signal intensity with a better signal -to-noise ratio.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG. 1 A depicts a workflow for spacial transcriptomic analysis using VeraLISH™ lock-and-roll (LnR) technology. An LnR split probe is annealed in a sequence specific manner to each target RNA in the tissue (Step 1). The sample is then contacted with a bridge probe (Step 2) that locks the split probe together (Step 3). This forms a circularized DNA annealed to mRNA sequence in the tissue, bearing a bar code that corresponds to the mRNA sequence. The circularized DNA is amplified to form a concatemer (Step 4), and analyzed using fluorescent probes to identify the bar code in each amplification product (Step 5).

[0010] FIG. IB is a detailed view of an LnR split probe. When hybridized to the RNA target, the probes are annealed adjacent to each other, and bridged together. A common landmark sequence can be incorporated as a non-specific reference point for each mRNA in the tissue.

[0011] FIG. 2 illustrates a nearest neighbor search for amplification products of target nucleic acids A, B, C, and D. Several probes complementary to each bar code bind to amplification products in the tissue (open stars), and form spots of varying intensity. Another fluorescent probe is used to label the common landmark sequence present in all amplification products (solid stars).

[0012] FIG. 3 shows how to obtain images for comparison of target-specific signals with common landmark sequences. This provides a second level of verification for the landmark sequences.

[0013] FIG. 4 depicts a workflow for image processing and spatial transcriptomic analysis. The outcomes of two preprocessing steps are integrated to identify spots and fine- tune their locations.

[0014] FIG. 5 shows how detector probe spots are used to identify particular transcripts at each location in the tissue. After extraction and localization, a codeword vector is constructed for each spot assembly, and compared with all the binary codes in a dictionary. The code that yields the shortest distance and the corresponding sequence is assigned to each spot as its gene identity.

[0015] FIG. 6 compares spots observed from landmark sequences and detector probes in a sample of human brain tissue. The left-most panel shows spots observed using the landmark identification probe (LMP). The other four panels show images obtained using different detector probes corresponding to the emission wavelengths of the labels used. The signals detected from both LMPs and detector probes are comparable.

[0016] FIG. 7A shows images obtained from human A549 cells. The spots for each of the detector probes colocalize with one of the LMP spots. FIG. 7B shows results of a parallel experiment in which the probes RP1 and RP2 were negative control detector probes having sequences that do not correspond to any of the bar codes in the sample.

[0017] FIG. 8 shows results obtained from a single A549 cell. The left image has multiple spots imaged with the landmark identification probe. The middle image was generated by overlaying multiple images. The panel on the right shows a merged image of both panels.

[0018] FIG. 9A is a depiction of a simplified in situ hybridization methodology. The sample is contacted with target hybridization probes, followed by direct ligation of labeled probes. FIG. 9B is a depiction of an in situ hybridization methodology that incorporates a DNA amplification step. mRNA in the sample is reverse transcribed, and then contacted with a target-specific padlock probe containing bar code that uniquely corresponds to the target nucleic acid being detected in the sample. The padlock probe is closed, then amplified and identified. FIG. 9C shows data in which a single nucleotide difference was determined between P-actin transcripts in cocultured human and mouse fibroblast cells, (scale bar, 20 microns). FIG. 9D illustrates an alternative in situ hybridization methodology (PLISH) that incorporates both split probes and in situ amplification in a different configuration.

[0019] FIG. 10A is a flowchart showing how to prepare paired images to create a training image database. The database will be used to train a convolutional neural network (CNN) to convert a raw image to a refined image with no background. FIG. 10B shows a detail of FIG. 10 A, showing a hybridization raw image paired with the reconstructed image obtained therefrom.

[0020] FIG. 11 shows graphically the improvement of accuracy of spot recognition by the CNN through each training cycle, between the “ground truth” image and the predicted image.

[0021] FIGS. 12A and 12B are images that illustrate the CNN image processing technology in action. FIG. 12A is the raw image. FIG. 12B is the corresponding processed image obtained from the trained CNN. There is less spot size variation and a much cleaner background.

[0022] FIG. 13 illustrates how sequential slices of a tissue sample can be used as a further means of identifying and characterizing fluorescent markers in a processed tissue sample. Colocalized gene spots in adjacent slices are presumed to come from the same gene product within the sample. A weighting factor for each colocalized spot is generated by calculating the ratio of each spot’s respective fluorescence intensity, which is applied to the centroid coordinates of each colocalized gene producer. The process refines spot placement in the image and removes false negatives.

[0023] FIG. 14 shows a workflow for combining the convolutional neural network technology explained earlier with the sequential spot colocalization technology.DETAILED DESCRIPTION

[0024] Included in this disclosure is technology for improving the quality and accuracy of images obtained for spatial transcriptomics, and the specific identification and localization of nucleic acid species in cell and tissue samples. Different aspects of the technology include the use of a landmark sequence in the reagents that confirm the validity and location of each and every targeted nucleic acid in the sample. The technology includes the use of artificial intelligence — in particular, a convolutional neural network model — to resolve the point spread function of each spot in two dimensions, and thereby to reconstruct and improve the image. The technology also includes the use of images from sequential slices to improve and validate spot location and intensity in three dimensions.

[0025] These aspects of the technology can be used by themselves or in any effective combination. The technology is developed in part to take advantage of in situ hybridization technology that incorporates strand-displacement amplification in the detection process, which intensifies the image and improves the signal -to-noise ratio. An exemplary methodology is LISH Lock’n ’Roll, which combines circularization of split probes with rolling circle amplification. However, any of the technological features put forth in this disclosure can be implemented for image improvement or for any other purpose in other forms of spatial proteomics analysis. More detailed explanations and illustrations are provided in the description that follows.Terminology used in this disclosure

[0026] Various aspects of the technologies put forward in this disclosure are exemplified in a partial type of fluorescent in situ hybridization referred to as LISH Lock’n’Roll (“LnR”). Except where explicitly stated or required, the technology of this disclosure is not limited tothis particular implementation, except where explicitly stated or otherwise required. Other implementations of the claimed technology are elaborated in separate sections of this disclosure, or the reader may design their own implementation without departing from the fundamental principles and elements of what is claimed.

[0027] The term “sample” or “biological sample” as used in this disclosure is anything that may contain nucleic acid of interest to be detected, quantified, and / or imaged in accordance with this technology. Such samples are typically referred to in this disclosure as a “cell or tissue sample”, but this designation should not be considered limiting. By way of illustration, the sample may be a biopsy specimen or other collection from a human or animal subject for any reasons of interest (for example, diagnosis of a condition or research), it may be a sample generated in the laboratory, such as the products of a tissue culture of eukaryotic or prokaryotic organisms, or it may be a sample collected from the environment that is not cellular in nature, but may contain remnants of an animal, botanical, or microbial cell worthy of assessment according to this technology.

[0028] A “target nucleic acid” refers to an RNA or DNA species present in a tissue or cell sample that is being assessed by the technology provided in this disclosure. For spatial transcriptomics, the target nucleic acid is an mRNA (a gene transcript). The technology can be used to assess other types of nucleic acids in a sample, such as small interfering RNA (siRNA), micro RNA (miRNA), small nuclear RNA (snRNA), long non-coding RNA (IncRNA), or an encoding or non-encoding DNA molecule, mutatis mutandis. The target nucleic acid may be nuclear, cytoplasmic, or compartmentalized in other subcellular organelles.

[0029] A “probe” in the context of this disclosure is a reagent (usually a DNA or mixed nucleic acid or analog thereof) that hybridizes or otherwise binds specifically to a target — either an analyte that may be present in a biological sample being analyzed, or another reagent.

[0030] A primary probe or reagent binds specifically to a particular analyte that may or may not be present in a sample (such as a target mRNA). The primary reagent may be referred to as a “target probe” or “sequence specific probe”. Primary probes bear a “target recognition sequence” or a “target sequence” that hybridizes to the analyte, plus either a detectable label (such as a fluorescent moiety) and / or one or more additional sequences that specifically bind to or receive a secondary probe. These additional sequence(s) may be referred to as a “bar code”, a “detector sequence”, a “detector binding sequence” or a “detection sequence”.

[0031] A secondary probe or reagent binds specifically to a primary probe or reagent to reveal, amplify, and / or quantify its presence in a sample. It may or may not bear a directly attached detectable label, and / or bind to a further reagent.

[0032] A “split probe” is a primary reagent that consists essentially of a first and a second linear nucleic acid, each of which bears at least a part of a target recognition sequence. The first and second part of the split probe typically bind to adjacent sites on a target analyte in a sample, and thereafter may be ligated together.

[0033] A “bridge probe” is a secondary reagent that is also a linear nucleic acid. As explained in more detail below, a bridge probe typically bears a sequence that is complementary to both ends of a split probe that are not hybridized to an analyte in a sample. The bridge probe is thereby configured to bridge the exposed ends of the split probe, whereupon the exposed ends may be ligated together.

[0034] The term “detector sequence” refers to a particular nucleotide sequence of 5 to 50 nucleotides that acts as a bar code corresponding to a particular mRNA or target nucleotide in a tissue or cell sample being analyzed. Primary reagents that are specific for different target nucleotides have different detector sequences. Detector sequences typically are contained in a primary probe or reagent that contacts the sample. They are optionally amplified, and then assessed by a secondary reagent referred to in this disclosure as a “detection probe”, a “readout probe”, or a “detector readout probe” which bears a detectable label that can be measured and / or imaged, as exemplified below.

[0035] A “Hyb probe” as referred to in some of the drawings and particular illustrations of the invention is a detector readout probe used in an implementation of this technology as applied to LISH Lock’n’Roll transcriptomics. The illustrations are exemplary, and do not imply any limitation of the claimed invention. They can be applied to other forms of in situ hybridization or sequencing as described below.

[0036] The term “landmark sequence” refers to a particular nucleotide sequence of 5 to 50 nucleotides that is shared between reagents targeting different nucleic acids in the sample. A common landmark sequence is typically shared between reagents recognizing different target analytes, where it is used for spot verification and characterization. In some instances, the landmark sequence may alternatively be referred to as a “bridge probe” sequence, since this is where it is typically located. The location of a landmark sequence hybridized to an analyte in a sample is determined using a “landmark probe” or a “landmark identification probe”. This can be done before, during, or after the measuring or imaging of detector readout probes.

[0037] An “optical label” is a chemical or large molecule moiety that generates a signal that can be measured optically. Such labels include but are not limited to fluorescent labels, and enzymes that generate a colored reaction product when contacted with an appropriate substrate. Optical labels of both types are exemplified below.

[0038] For ease of understanding and interpretation, nucleic acid sequences that play a functional role in this disclosure include both the originating sequence and the complement thereof. For example, a particular detector sequence in a split probe is referred to as the same detector sequence in the amplification product thereof, which is in turn referred to as the same detector sequence in the detection reagent or detector readout probe to which it hybridizes — even though these sequences are successive complements of each other, in accordance with standard Watson-Crick base pairing chemistry. Each sequence may be DNA, its RNA equivalent, combinations thereof, analogs thereof, and complements thereof as logically implied in the context of the description.

[0039] When two nucleic acids are indicated as binding or hybridizing specifically to each other, this means that the opposing sequences are sufficiently complementary (but not necessarily identical) to each other such that they hybridize to each other, and not to other nucleic acid reagents that may be in the reaction mixture or unrelated nucleic acid species that may be in the sample being tested. The user may adjust the sequence of the nucleic acid reagents and / or hybridization conditions (such as salt concentration and temperature) as a matter of routine reagent design and testing to achieve the degree of specificity that is needed.

[0040] “Stringent” hybridization conditions are conditions that enhance the specificity of hybridization. Particular hybridization conditions are sequence dependent, and will be different with different reaction parameters (such as salt concentrations and presence of organics such as formamide). Generally, stringent conditions are selected to be about 5°C to 20°C lower than the thermal melting point (Tm) for the specific nucleic acid sequence at a defined ionic strength and pH. Preferably, stringent conditions are about 5°C to 10°C lower than the thermal melting point for a specific nucleic acid bound to a complementary nucleic acid. Stringent hybridization conditions may include the presence of 5-40% formamide (or another organic reagent) that lowers the Tm. The Tmis the temperature (under defined ionic strength and pH) at which 50% of a nucleic acid hybridizes to a perfectly matched probe. Depending on the reagents used, stringent hybridization conditions may be 6 * sodium chloride / sodium citrate buffer (SSC) at approximately 45°C, followed by one or more wash steps in 0.2 x SSC, 0.1% SDS at 50° to 65°C.

[0041] VeraLISH™ and PISCES® are trademarks of Applied Materials Inc., Santa Clara, California.Methods for determining spatial transcriptomics

[0042] The technology described and claimed below can be used in any manner on any sample for which it is suitable. Spatial transcriptomics is often performed on fresh frozen sections or slices of a tissue sample, or a preparation of cells obtained from tissue or from tissue culture. To perform in situ hybridization on slices of paraffin embedded samples, the paraffin may be removed using a hydrocarbon such as hexadecane, followed by ethanol exchange into an aqueous solvent compatible with the reagents being used, with surfactants and / or nuclease inhibitors in the solvent as required. Slices are typically mounted on a suitable surface such as a glass slide or coverslip, placed in a reaction chamber to conduct the hybridization reactions, and moved to an imaging apparatus if needed to detect optical labels in the processed sample.

[0043] The technology of this disclosure has been exemplified in a particular in situ hybridization and amplification technology known as LISH Lock’n’Roll (“LnR”), described in more detail below. In brief, LISH LnR refers to in situ hybridization and ligation of probes that form and lock a circular DNA bearing a detector sequence in place in the tissue at the location of a target nucleic acid being detected. The circular DNA is amplified by rolling circle amplification (RCA) before detection using labeled probes and imaging. To simplify explanation of the technology provided in this disclosure, it will be discussed specifically in the context of LISH LnR. Other types of in situ hybridization and in situ sequencing that are compatible with the technology of this disclosure are put forth in a later section of this disclosure.LISH Lock’n’Roll

[0044] The LISH Lock’n’Roll (“LnR”) methodology exemplified in this disclosure is an in situ hybridization technique that hybridizes split probes to target nucleic acids in a tissue or cell sample of interest. The split probes are then circularized, amplified, labeled using labeled detection reagents, and imaged as image spots in the sample.

[0045] FIG. 1A depicts an exemplary LnR workflow. In Step 1, an LnR split probe (comprising a first and second linear portion probe) is annealed (hybridized) to the target RNA. Features of the acceptor and donor probes are discussed in detail below in reference to FIG. IB. In step 2, the sample is exposed to a bridge probe (primer) that is complementary to a bridge sequence and which hybridizes thereto, thereby forming a circularized three-probe complex annealed to the target RNA sequence (Step 2). A ligation step, (using, for example,T4 DNA ligase and T4 RNA ligase 2) is then performed to ligate the ends of the acceptor and donor probes, thereby forming a closed-circular DNA molecule (Step 4).

[0046] Due to the twist of the double helix, the closed-circular DNA molecule is locked into place around the target mRNA. A suitable DNA polymerase is added to the tissue (shown as a grey oval) to initiate rolling circle amplification (RCA) in situ, which is primed by the bridge probe annealed to the closed-circular DNA molecule (Step 5). The DNA polymerase used for rolling circle amplification is preferably an enzyme capable of multiple displacement amplification of DNA: such as 029 (Phi 29) DNA Polymerase or its equivalent (available, for example, from NxGen or ThermoFisher). Desirable DNA polymerase features are high processivity and strand displacement activity, capable of synthesizing DNA up to 70 kb long, highly accurate DNA synthesis, high yields of amplified DNA even from minute amounts of template, and amplification products suitable for hybridizing detector readout probes.

[0047] The RCA product formed from the circular DNA is referred to in this disclosure as a “rolony”. It is in essence a concatemer of single-stranded (ss) DNA containing multiple copies of the closed-circular DNA molecule, which is comprised of the acceptor and donor probe sequences. The rolony can be considered a nanoball of ssDNA and, as discussed later in this disclosure, it is located at a position that approximates the position of the target RNA sequence in the sample. Following completion of RCA step, fluorescently labeled oligonucleotides (detector readout probes) are annealed to complementary detector sequences (originating from the acceptor and / or donor probes as described below), of which there are now many spatially localized copies (Step 6). The tissue sample is now ready to be processed for imaging.

[0048] FIG. IB is a detailed view of the LnR split probe, comprising a first and second linear portion (optionally referred to as an acceptor probe and a donor probe). The 3’ terminus of the acceptor probe is composed of two ribonucleic acid bases at the 3’ end, which foster high efficiency ligation by the T4 RNA ligase 2, Rnl2. The donor probe is phosphorylated at the 5’ end. The acceptor and donor probes have targeting sequences complementary to adjacent target sequences of the RNA target. As such, when hybridized to the RNA target, the probe set are annealed adjacent to each other, with the donor probe annealed at a position 5’ along the target RNA sequence relative to the acceptor probe. The donor probe may be referred to as the 5’ probe and the acceptor probe may be referred to as the 3 ’ probe. The targeting sequences are typically about 20 nucleotides in length, but may also be substantially longer or shorter (such as, for example, 10 to 40 or 18 to 24 nucleotides). Only when the LnR probes are annealed adjacent to one another on a target sequence canthey be ligated together via Rnl2. This requirement that the ligation probes anneal to adjacent sequences provides a high level of specificity.

[0049] There may be one or more detector sequences of 10 to 500 or 20 to 50 nucleotides on either or both of the acceptor and donor probes. As illustrated in FIG. 1A, each acceptor and donor probe feature one or two 30-nucleotide detector sequences and a 17- nucleotide bridge sequence. The detector sequences are positioned between the target specific portion and the bridge probe binding portion of a probe and may be some or all of the nucleotide sequence between these two portions.

[0022] In some instances, there can be an intervening sequence between the target specific, detector specific, and / or bridge probe specific portions of at least one, two, five, ten, twenty, fifty, or one hundred nucleotides. If either the acceptor and donor probe does not contain a detector sequence, typically there is an intervening sequence between the bridge probe binding portion and the target specific portion of the probe. These intervening sequences are referred to collectively in this disclosure as such or as the framework of the split probes. Typically, the sequence of the various portions of the framework are selected to be substantially inert, which means that they do not cross-hybridized with any of the other nucleic acid reagents, or with nucleic acids or other components that may be present in the tissue. Said another way, the intervening sequences are designed to not have complementarity to other reagents or sequences present in the tissue sample so as to avoid non-target specific hybridization of the probes under typical assay conditions. In some instances, the framework or intervening portions may be constructed of nucleic acid analogs or non-nucleic acid materials that are inert to reactions with other nucleic acids.

[0050] The probe sets are typically designed to have 1 to 4 unique detector sequences (or detection sequences). Multiplexing (the detection of multiple target nucleic acids in a single image) is done via color barcoding, where a probe set has two or more distinct detector sequences for simultaneous binding of two or more distinctly labeled detector readout probes (also referred to in this disclosure as detector readout probes).

[0051] For determining a large number of target nucleic acids in a sample, the technology is multiplexed by having multiple cycles of staining, imaging and destaining. Each cycle typically comprises the use of 2 to 10, typically 3 to 6, or more precisely 4 labels that can be distinguished from each other when measured concurrently. By way of illustration, 6 cycles using 4 labels can be used to detect 20 different target nucleic acids in a sample, one landmark sequence, and several controls. Alternatively or in addition, a plurality of different primary reagents bearing different detector sequences can be designed to hybridize specifically to the same target nucleic acid in the same or different cycles (forexample, to different partial sequences of the target nucleic acid). The progression of different labels binding to the same target nucleic acid constitutes “color barcoding” for that target. This achieves a greater level of combinatorial multiplexing and / or error-correcting analyte identification.

[0052] As an example, a panel of LnR probe sets with two different detector sequences per probe set, and five uniquely colored detector readout probes, can be used to simultaneously measure more than 15 targets during a single cycle of imaging. The bridge probe oligonucleotide binds to the bridge probe (BP) recognition regions in both probes in LnR probe set. Bridge probes can be designed as single probe, spanning bridge probe recognition regions on both LnR probes. Or the bridge probe can be split into two, each recognizing bridge probe recognition regions on either of two LnR probes.

[0053] Detector readout probes are labelled oligonucleotide probes used to identify LnR probe sets for different gene transcripts. They are generally DNA probes and may have an optical or fluorescent label conjugated thereto. Exemplary labels include SYBR green, SYBR gold, a CAL Fluor dye such as CAL Fluor Gold 540, CAL Fluor Orange 560, CAL Fluor Red 590, CAL Fluor Red 610, and CAL Fluor Red 635, a Quasar dye such as Quasar 570, Quasar 670, and Quasar 705, an Alexa Fluor such as Alexa Fluor 350, Alexa Fluor 488, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 594, Alexa Fluor 647, and Alexa Fluor 784, a cyanine dye such as Cy 3, Cy3.5, Cy5, Cy5.5, and Cy7, fluorescein, 2’, 4’, 5’, 7’-tetrachloro- 4-7-dichlorofluorescein (TET), carboxyfluorescein (FAM), 6-carboxy-4’,5’-dichloro-2’,7’- dimethoxyfluorescein (JOE), hexachlorofluorescein (HEX), rhodamine, carboxy-X- rhodamine (ROX), tetramethyl rhodamine (TAMRA), FITC, dansyl, umbelliferone, dimethyl acridinium ester (DMAE), and Texas Red.

[0054] Although fluorescent labels are used in the illustrations below, other types of labels may be used. For example, detector and / or landmark identification probes may bear an enzyme capable of generating a fluorescent or colored signal upon supplying a substrate for the enzyme. Such enzymes include horse radish peroxidase, soybean peroxidase, alkaline phosphatase, beta-galactosidase, hematin, acacia, iron porphyrins, and peroxidase-mimicking nanomaterials (nanozymes). The soluble substrate for a peroxidase could be for example, 3,3’,5,5’-tetramethylbenzidine (TMB), luminol (5 -amino-2, 3 -dihydrophthalazine- 1,4-dione), ABTS (2,2’-azinobis [3-ethylbenzothiazoline-6-sulfonic acid]-diammonium salt), or a combination thereof. Alternatively, the reporter may be an enzyme that produces a fluorescent or light-emitting reaction product, such as a luciferin.

[0055] Alternatively or in addition, detector and / or landmark identification probes may be labeled with radioisotopes such as32P,33P,35S,3H,125I, "mTc,95Tc,n iIn,62Cu,64Cu, Ga,68Ga, and153Gd; or with paramagnetic metal ions such as Gd(III), Dy(III), Fe(III), and Mn(II). Small molecule labeling means include biotin or digoxigenin. Biotin, which can be detected by specific and high affinity binding to avidin or streptavidin structurally coupled to an enzyme catalyzing a colorimetric reaction, such as phosphatase, luciferase, or peroxidase. Other suitable labels are nanoparticles such as gold and silver particles, which, when irradiated with angled monochromatic light, scatter the light with high intensity, at a wavelength that depends on the size of the particle. Other suitable labels are quantum dots, such as CdSe, ZnSe, InP, and InAs. These are fluorescing crystals 1 to 5 nm in diameter that emit monochromatic light, with a wavelength dependent on their chemical composition and size.

[0056] Current working models of this technology use 24 different detector readout probes, each labelled with a fluorescent dye: in particular, Cy3, Atto Rhol2, Cy5, and Alexafluor750 fluorescent tags. These dyes are excited by lasers at 530 nm, 577 nm, 640 nm and 750 nm, respectively, and imaged using corresponding filter sets, objective lenses, and cameras.Use of a common landmark sequence to validate fluorescent spots

[0057] Signals or spots obtained from imaging analysis of tissue samples using labeled hybridization probes can be verified and characterized by incorporating a technology - a common landmark sequence - that confirms that each spot is a bone fide representation of a target nucleic acid at or in the vicinity of the location of the spot. Said another way, the common landmark sequence can be used to confirm that a spot is not a non-specific reaction biproduct or optical signal that reflects an intrinsic property of the sample and / or a biproduct of one or more of the probes or reagents that is not obtained by specific hybridization to a target nucleic acid.

[0058] A central feature of the landmark sequence is that it is the same for all of the target-specific probes being used, or at least a subset thereof.

[0059] Referring again to FIG. IB, a common landmark sequence can be incorporated into LISH-Lock’n’Roll in any position within the probe set that is not sequence specific for the target analyte or mRNA (the “target specific” portion) or the detector binding portion. For example, the common landmark sequence can be represented in the framework of either or both parts of the split probe. Preferably, the landmark sequence is present either in the bridge probe binding portion of one of the split probes, or is created from the split probes upon ligation of the two parts of the split probe together across the bridge probe binding portions. Said another way, the landmark sequence may overlap with the bridge probebinding portions of the split probe. This increases the error-checking function of the landmark sequence by requiring that it be present only for split probes that have completed the sequence of reactions required to form a circular DNA ready for amplification. Incorporation of landmark sequences into LISH-Lock’n’Roll probes and other types of hybridization probes is described further in the sections that follow.

[0060] FIG. 2 illustrates a nearest neighbor search for amplification products of target nucleic acids A, B, C, and D. Spots corresponding to gene transcript sequences (open stars) in each implication product are generated by hybridizing with a series of fluorescently labeled target-specific Hyb readout or detector probes. Spots corresponding to a common landmark sequence present in all amplification products (solid stars) are generated by hybridizing with a fluorescently labeled landmark identification probe (LMP).

[0061] Proximity analysis of fluorescence images of a tissue sample is done using a spot recognition algorithm. Each LMP spot serves as a seed or reference location for identifying Hyb probe spots present at that location. For each LMP spot, the nearest neighbor distance is calculated as the linear distance between the centroid of the LMP spot, and all adjacent Hyb signal spots in the image. A Hyb signal spot is characterized as being at the same location as an LMP spot if it falls within the defined radius of that LMP spot (dashed circle). The defined radius used for this characterization may be predetermined by the user before analyzing a particular image or series of images. Alternatively, it may be determined and / or adjusted during the course of analysis, for example, to maximize signal accuracy and / or to minimize the proportion of false negatives.

[0062] Following the proximity analysis, a codeword vector is determined for each LMP spot as the stored brightness of all Hyb signal spots (open stars) at the same location. Each gene transcript is characterized by a unique combination of Hyb probe spots at that location. Gene transcript sequences are identified by comparing the codeword determined for each LMP spot with entries in a preestablished codeword dictionary.

[0063] FIG. 3 shows a series of steps to obtain images for comparison of target-specific signals with common landmark sequences. The location of landmark sequences can be determined at any time during the analysis. In this example, the landmark sequence is determined before any of the target specific spots (Hyb signal spots), and then again after the series is complete. This provides a second level of verification for the landmark sequences themselves, taking into account any effect that the hybridization and detection reactions may have on the sample during the reaction series.

[0064] Each bit shown in the table in FIG. 3 is an individual fluorescence imaging channel that is captured in a given Hyb probe imaging cycle. There is a single bit per Hybprobe imaging cycle where a spot is detected at the left-most position. It corresponds to Gene C in both the Bit 1 and Bit 2 of the Hyb probe 1 and 2 imaging cycles. As both spots appear within the defined radius of a landmark spot, both spots are presumed to be valid Hyb spots — i.e., each spot is considered to represent the binding of a detector readout probe to a rolony at the site of a specific target gene transcript. This will occur where there are two target probes that hybridize to different parts of the same gene transcript.

[0065] Once valid Hyb spots are identified and localized, tGene C can be identified based on the sequence of detector readout probes that have formed Hyb spots within the defined radius of the same landmark sequence.Reaction procedure for LISH Lock’n’Roll

[0066] A suitable protocol and suitable reagents for LnR processing of tissue or cell samples is shown in TABLE 1.Image processing

[0067] FIG. 4 shows a generic workflow of image processing and spatial transcriptomics analysis according to the various aspects of this disclosure. The raw images are first sent through two preprocessing procedures: first, spot-like signals are identified through a deep learning (DL) spot recognition pipeline; second, images from different rounds of channels are aligned to eliminate feature shifts if there are any. The term “feature shifts” refers to changes to a sample or imaging field that may alter the physical location of a gene relative to other Hyb imaging cycles (e.g., tissue stretching, variation in image detector alignment). The outcomes of the two preprocessing steps are integrated to identify spots and fine-tune their locations. These spots are further filtered by their attributes such as brightness and size, and low-quality ones with either trivial size (for example, 1-2 pixels) or poor intensity are rejected. Each site with an assembly of spots is converted to a codeword based on spot identities (i.e., their bits) and brightness, then the binary code in the dictionary that is close to the codeword is selected and its associated gene assigned to the spot assembly. Additional cleanup to further refipne decoded results (for example, using a different LMP probe to generate seed spots) can be applied on the decoded results, if desired.

[0068] FIG. 5 shows how detector probe spots that have been validated are used to identify the particular target nucleic acid at each particular location. After extraction and localization, a codeword vector is constructed for each spot assembly from spot brightness of the occupied bits. For instance, spots with brightness [b1,b3, b4, b5], are found on bits 1, 3, 4, 5 among 16 bits within a neighborhood. The codeword vector of this assembly is thus b = [b-^, b2... . b16] with bj = 0 for j 1, 3, 4, 5. Each codeword vector is then compared with all the binary codes in the dictionary and its distance to each binary code is computed. The binary code that yields the shortest distance and its associated gene will be assigned to the spot assembly as its gene identity. The minimum distance value will be preserved as a quality metric of the codeword for further cleanup steps (optional).Error correction and multiplexing

[0069] LISH Lock’n’Roll and other types of in situ hybridization reagents can incorporate a plurality of different target specific probes into one assay. For example, multiple probes can be configured to hybridize specifically to the same target, for example to amplify the signal and / or to provide an error-correction function (for example, wherein a plurality of different probes must be observed at or about the same location to confirm the presence of the common target). Multiple probes can also be configured to hybridize each to a different target, or wherein a plurality bind to each of several targets. This provides a means of observing different targets in the same tissue sample, either to broaden the survey or to determine to what extent the expression of different gene products are correlated.

[0070] In any of these combinations, a small plurality (about 2 to 8, typically 4 or 5) can be coded with a fluorescent tag and imaged concurrently, depending on the availability of tags of different colors. The plurality can be expanded by having multiple cycles of detection and imaging. Any operative combination of reagents and steps can be used. A particularly convenient way of cycling is to hybridize all target specific probes to the sample at once, and lock and amplify them together (in the context of LISH LnR). The amplification products can then be probed in multiple cycles of different detector readout probes each bearing a different label. After the labels have been imaged in each cycle, the fluorescent tags (with or without the rest of the detector readout probe) can be removed, quenched, or otherwise inactivated as discussed below, rendering the tissue dark for the next cycle of detector probes.

[0071] By way of illustration, if there are 25 different target probes (targeting 25 different mRNAs, or redundantly targeting a lower number of mRNAs), the tissue can be contacted with 25 different LnR probe sets at once, the probe sets closed to form circular DNAs with a bridge probe containing a landmark sequence (i.e. a plurality of the same bridge probe binding to each probe set), and amplification performed resulting in the formation of 25 different rolonies at or in the vicinity of the respective target nucleic acid in the sample, all having one or more different detector sequences corresponding to the target nucleic acid. Different rolonies have different combination of detector probe recognition sequences based on the target nucleic acid from which they arise (i.e. LnR probe sets that bind specifically for a given target nucleic acid have a distinct combination of detector probe recognition sequences compared to LnR probe sets that bind specifically to a different target nucleic acid), but they all contain the same landmark sequence.

[0072] The sample is then hybridized sequentially with 16 or 24 different detector readout probes, optionally in an imaging device. The sample is hybridized sequentially, for example, using about 4 detector probes per cycle, until all the detector probes are imaged, for example four hybridization cycles is required for 16 different detector probes. Image analysis includes decoding of spots using information obtained from images from each hybridization cycle. The landmark sequence can be identified using a landmark identification probe before or after analysis of the rolonies with the detector readout probes, or during one of the detector sequence readout cycles.

[0073] Labels can (in effect) be removed between cycles, for example, by photobleaching. Fluorescent detector readout probes are hybridized on the instrument, imaged and bleached / quenched using lasers at high power. This is followed by next round of hybridization of detector probes. With this technique, detector probes are not removed, but its fluorescence is quenched by high intensity laser. Alternatively, a chemical bleach can be used to denature and remove the detector probes from their hybridization site on the rolonies. Fluorescent detector probes can be hybridized on the instrument, imaged, and removed using a denaturing reagent, such as formamide. Since both the label and the detector probe are removed, this has some advantages in relation to the photo-bleach method, allowing repetitive usage of same detector probes.

[0074] Labels can also be removed between cycles by chemical cleavage of the tag / fluor ophore from detector probes. The cleavage can be performed using different techniques based on the linker between the detector probe and its fluorescent tag. One technique uses a linker that breaks when exposed to UV light. With this technique, each hybridization and imaging cycle is followed by a UV exposure to the tissues / cells, which removes the fluorescent tag, allowing next round of hybridizations. Another technique uses chemically cleavable linker, such as disulfide bridge, which can be broken using TCEP (tris(2-carboxyethyl)phosphine), a reducing agent. Each detector probe hybridization is followed by imaging and running TCEP to remove fluorescent tags. Either of these cleavage techniques do not remove detector probes from their binding regions on rolonies.Preserving spot location in tissue samples

[0075] An advantage of the LISH Lock’n’Roll methodology is that formation of the closed-circular DNA molecule from the split probe becomes locked into place around the target mRNA. In many circumstances, most of the rolonies remain in place where they are formed by the DNA polymerase, and do not move outside substantially throughout multiple hybridization cycles. This is presumably due to the proteinaceous mesh structure inside the cells and tissues formed as a result of fixation.

[0076] Depending on the conditions used, the nature of the sample, and the properties of the target nucleic acids of interest to the user, the rolonies can be treated to fix them in place adjacent to the respective target nucleic acid molecules before hybridizing with the detector probes. Paraformaldehyde is a common fixative that covalently attaches proteins to each other, thereby stabilizing the sample for successive cycles of hybridization.

[0077] Optionally, the tissue in the sample can be treated to form a more rigid matrix. Matrix forming materials include polyacrylamide, cellulose, alginate, polyamide, cross-linked agarose, cross-linked dextran or cross-linked polyethylene glycol. The matrix forming materials can form a matrix by polymerization and / or crosslinking of the matrix forming materials using methods specific for the matrix forming materials and methods, reagents and conditions known to those of skill in the art. The tissue can be treated to form a matrix before hybridizing with target probes if the matrix does not interfere with integrity or accessibility of the target probes. Alternatively, the tissue can be treated to form a matrix after the target probes have been hybridized, closed, and amplified to form rolonies, if the matrix does not interfere with integrity or accessibility of the rolonies to detector readout probes and landmark identification probes.

[0078] Optionally, nucleic acid reagents and / or products can be modified to incorporate a functional moiety for attachment to the matrix. The functional moiety can be covalently cross-linked, copolymerize with or otherwise non-covalently bound to the matrix. The functional moiety can react with a cross-linker. The functional moiety can be part of a ligand- ligand binding pair. dNTP or dUTP can be modified with the functional group, so that the function moiety is introduced into the DNA during amplification. The functional group may be an amine, acrydite, alkyne, biotin, azide, or thiol. In the case of crosslinking, the functional moiety is cross-linked to modified dNTP or dUTP or both. Suitable exemplary cross-linker reactive groups include imidoester (DMP), succinimide ester (NHS), maleimide (sulfo-SMCC), carbodiimide (DCC, EDC) and phenyl azide.Images obtained

[0079] Exemplary results from LISH Lock’n’Roll methodology with landmark sequence validation are shown in FIGS. 6 to 8.

[0080] FIG. 6 compares spots observed from landmark sequences and detector probes in a sample of human brain tissue. The LnR assay protocol was applied using a neuro gene panel of target probes for 126 genes. The sample was imaged on the instrument with multiple cycles. The images depicted were obtained from a single cycle. The left-most panel shows spots observed using the landmark identification probe (LMP). DAPI staining for nuclei and some autofluorescence appears in the background. The other four panels show images obtained using different detector probes (or bit, B), at 530, 577, 640, and 750 nm, corresponding to the emission wavelengths of the labels used. These images show that the signals detected from both LMPs and detector probes are comparable.

[0081] FIG. 7A shows results obtained from A549 cells (a human cell line of lung carcinoma epithelial cells). LnR protocol was applied. LMP was hybridized to see all spots (bottom left). Other channels were added to see the compatibility of detectors probe labeling in relation to LMP labeling. Detector probes RP1, RP2, RP5, RP6, RP7, and RP8 recognize different parts of the same LnR target probes, thereby amplifying that signal. The spots for each of these detector probes colocalize with one of the LMP spots. Detector probes RP3 and RP4 were negative controls, having detector sequences not found on the LnR target probes used in this experiment. No spots were observed for RP3 and RP4, as expected. All of the spots obtained using the detector probes have a corresponding LMP spot, confirming that LMPs label all of the rolonies obtained for the entire set of target probes used in two sequential cycles of staining.

[0082] FIG. 7B is a similar experiment, except that the probes RP1 and RP2 were the negative control detector probes having no corresponding target hybridization probes in the sample. Again, all of the spots obtained using the detector probes have a corresponding LMP spot, confirming that LMPs label all of the rolonies obtained for the entire set of target probes used in two sequential cycles of staining.

[0083] FIG. 8 shows results from another LnR protocol performed on A549 cells. Each image shows the same single cell. The left image shows many spots imaged with the landmark identification probe. Middle image shows an image that is generated by overlayingmultiple images from all the bits, or detector (Hyb) probes, obtained from the series of hybridizations. The panel on the right shows a merged image of two panels, permitting identification of false positive spots and false negative spots. Some spots are identified by both LMPs and detector probes, further increasing the specificity of the LnR assay.Advantages of this technology

[0084] The use of LISH Lock’n’Roll methodology provides several improvements compared with standard fluorescence in situ hybridization probes used without amplification. In the context of multiplexed imaging:• LISH probes have higher signal intensity per labeled RNA• there are fewer constraints in probe design and construction• high specificity as a result of the split probe design• probes can target shorter RNA species in a tissue sample• the technology requires lower laser power and exposure time (thereby reducing sample processing costs)• higher signal -to-background ration (SBR > 3.0) compared with state-of-the-art multiplexed FISH probes (typical SBR 1.5 to 2.5)• smaller bit-to-bit signal variation, reducing the frequency of false negatives and false positives

[0085] The use of landmark sequences for pre- or post-multiplex image collection has additional advantages:• landmark sequence spots have higher signal-to-background ratio and high brightness• the pre-hybridization LMP images can be used to check labeling quality• post-hybridization LMP images can be used to check tissue and assay integrity through multiple probe cycles

[0086] Before the development of this technology, a typical Veranome workflow used 50-100 target probes for RNA detection. Incorporating so many probes for the same target nucleotide in a tissue sample increases signal but limits the range of transcript coverage. It also limits the number of detector readout probes bound at the same time or concurrently. LISH LnR includes a DNA polymerase amplification step, which considerably improves the signal to noise ratio. In a typical implementation, LISH LnR uses one to five sites per nucleic acid target in the sample, hybridized to one to five LnR probes per target. The low number ofsites per target enables the technology to detect small differences that may be important: such as single nucleotide variants and differences in exon-exon junctions.Other types of in situ hybridization or in situ sequencing that are compatible with the technology in this disclosure

[0087] Although various aspects of the technology of this disclosure are exemplified in the context of LISH Lock’n’Roll in situ hybridization, they may be implemented or adapted into other forms of in situ hybridization analysis. Other ways of doing in situ hybridization include the following:

[0088] Fluorescence in situ hybridization (FISH) is a molecular cytogenetic technique in which fluorescent oligonucleotide probes hybridize to nucleic acid sequences in a tissue sample to detect and localize specific RNA targets (mRNA, IncRNA, miRNA) in tissue samples and cells. Multiplexed fluorescence in situ hybridization (mFISH) is a technique that uses a range of probes that each bind specifically to different nucleic acid targets having different sequences, with a sequential or simultaneous labeling strategy that separately identifies mRNA species having different sequences.

[0089] Single-molecule fluorescent in situ hybridization (smFISH) implements short (50 bp) oligonucleotide probes conjugated with five fluorophores to obtain quantitative information about expression of certain genes in the cell. RNAscope employs probes of the specific Z-shaped design to simultaneously amplify hybridization signals and suppress background noise.

[0090] Single-molecule RNA detection at depth by hybridization chain reaction (smHCR) is an advanced seqFISH technique in which a set of short DNA probes attach to a defined subsequence of the target, followed by fluorophore-labeled DNA HCR hairpins that penetrate the sample and assemble into fluorescent amplification polymers attached to the initiating probes. Cyclic-ouroboros smFISH (osmFISH) visualizes transcripts in the manner of smFISH, and an image is acquired before the probe is stripped and the sample is reprobed. Multiplexed error-robust fluorescence in situ hybridization (MERFISH) is a single-cell transcriptome imaging method that encodes RNA target molecules with error-robust binary bar codes. The readout sequences are detected, for example, using orophore-labelled secondary probe, and the fluorescence signal is extinguished via photobleaching before subsequent rounds of imaging.

[0091] DNA microscopy generates cDNA in fixed tissue, following which randomized nucleotides are used to tag and amplify target cDNAs in situ, thereby generating unique labels for each molecule.

[0092] SeqFISH Plus resolves optical issues related to spatial crowding using a primary probe anneals to targeted mRNA, followed by detector readout probes that bind to flanking regions in a way that can be captured as an image and collapsed into a super-resolved image.

[0093] FIG. 9A is a general depiction of multiplexed in situ hybridization methodology that is done using target hybridization probes followed by direct ligation of labeled probes. A series target hybridization probes (primary reagents) for different nucleic acids in the sample are depicted as dish shapes hybridized to a tissue sample below. There are two detector sequences or bar codes on each probe (the left and right side), which between the two of them uniquely correspond to the target hybridization sequence between them, and hence the target nucleic acid being identified in the sample. Each of the hybridization probes is detected on the right side using a corresponding labeled detector readout probe. The labels on the detector readout probes are cleavable using TCEP (a reducing agent), which prepares each target hybridization probe for detection on the left side using a second labeled detector readout probe. The optical signals from the right side and left side detector readout probe for each reagent are used to identify which target nucleic acid is at each location.

[0094] FIG. 9B is a general depiction of an in situ hybridization methodology that incorporates a DNA amplification step. mRNA in the sample is reverse transcribed and then contacted with a target-specific padlock probe. The padlock probe hybridizes to each target nucleic acid by way of complementary sequences at the 5’ and 3’ ends. In between is a target-specific detector sequence or bar code that uniquely corresponds to the target nucleic acid being detected in the sample. The padlock probe is then closed using a ligase, and amplified to form a rolling circle amplification product (RCP). Specific labeled detector readout probes are used to identify the detection sequence in the RCP, and hence the underlying target nucleic acid.

[0095] FIG. 9C provides a working example, in which a single nucleotide difference is determined between P-actin transcripts in cocultured human and mouse fibroblast cells. The cell on the lower left shows spots generated from a mouse-specific detection sequence, whereas the cell on the upper right shows spots generated from a human-specific detection sequence. Scale bar, 20 microns. The amplification step improves the signal -to-noise ratio.

[0096] FIG. 9D illustrates an in situ hybridization methodology that incorporates both split probes and in situ amplification, but is different from LnR. It is referred to as molecularprofiling using proximity ligation (PLISH). The split probe is hybridized to the target DNA, but is not ligated to form a circle. Instead, both a bridge probe and a padlock probe are hybridized to the other ends of the split probe, which are then ligated together to form a closed circular DNA. In this example, the detector or bar code sequence corresponding to the target may be located in the bridge probe or the padlock probe. The closed circular DNA is then amplified and identified using a labeled detector readout probe.

[0097] To incorporate the landmark sequence image improvement technology of this disclosure into the in situ methodology shown in FIG. 9A, the landmark sequence is located or included in one of the ends of the hybridization probe. Between them, the two ends comprise one or more detection sequences that corresponds to the target recognized by that probe plus a landmark sequence that is the same for all target specific probe. To incorporate the landmark sequence image improvement technology of this disclosure into the in situ methodology shown in FIG. 9B, the landmark sequence is located or included in the padlock probe adjacent to the detection sequence, or elsewhere in the framework. To incorporate the landmark sequence image improvement technology of this disclosure into the in situ methodology shown in FIG. 9D, the landmark sequence is included in either the bridge probe or the padlock probe.

[0098] Image analysis for any of these methodologies using the landmark sequence technology and / or the neural network technology and / or sequential slice analysis technology is performed according to the same principles put forward in this disclosure in the context of LISH Lock’n’Roll, mutatis mutandis.Features of the landmark technology of this disclosure as implemented in LISH Lock’n’Roll

[0099] Implemented in the context of LISH Lock’n’Roll methodology, the use of landmark sequences in accordance with this disclosure may include the following features in any combination.

[0100] To identify locations of a plurality of different target nucleic acids in a cell or tissue sample, the user first obtains for each of the target nucleic acids in the sample a split probe that comprises a first linear part and a second linear part. The 5’ end of the first part and the 3’ end of the second part of each split probe contain target recognition sequences configured to hybridize specifically to the respective target nucleic acid at positions adjacent to each other. The 3’ end of the first part and the 5’ end of the second part of each split probe are configured to hybridize specifically to a common bridge probe at positions adjacent toeach other. At least one of the first part and the second part of each split probe comprises a detector sequence that is different for each different target nucleic acid.

[0101] The sample is contacted with the split probes under conditions where the first part and the second part of each split probe hybridize specifically to the probe’s target nucleic acid. For each of the split probes that have a first part and a second part hybridized adjacent to each other on a target nucleic acid, the 3’ end of the first part is ligated to the 5’ end of the second part of the split probe. The sample is contacted with the common bridge probe under conditions where the bridge probe hybridizes specifically to the first part and the second part of each split probe that are adjacent on a target probe in the sample. For each of the split probes that have a first part and a second part hybridized adjacent to each other on a copy of the common bridge probe, the 5’ end of the first part is ligated to the 3’ end of the second part. The two ligating steps forms a circular DNA from the respective split probe, wherein the circular DNA contains at least the detector sequence(s) from the split probe, plus a copy of a common landmark sequence.

[0102] Each of the circular DNA’s formed from split probes is then amplified in situ (typically with a non-processive DNA polymerase) to form an amplification product (a rolony or a nanoball). Locations of copies of the common landmark sequence and of each of the detector sequences is determined in the amplification products in the sample. The image is processed or characterized such that each of the different target nucleic acids in the sample as being the locations where the respective detector sequences are each within a predefined radius from a copy of the landmark sequence.

[0103] Depending on how the technology is implemented, the common landmark sequence may be contained in the bridge probe in its entirety, or it may be formed by ligation of the two parts of the split probes. The landmark sequence may also appear in its entirety in either part of the split probe, which is then reproduced in the circular DNA and in the amplification product. Requiring that the landmark be included in the bridge probe or formed upon ligation provides a further verification that the ligation reaction has gone to completion in the preferred configuration.

[0104] Depending on how the technology is implemented, the hybridizing of the bridge probe and the two ligation steps may be performed sequentially in any operative order. Alternatively, the two ligation steps and optionally the bridge probe hybridization can all be initiated at the same time. As depicted in FIG. IB, one end of one of the split probes may be RNA rather than DNA (1 to 5, typically 2 base pairs in length). In this configuration, different ligating enzymes are required to perform the two ligation steps. The user can titratethe amount of each ligation enzyme so as to perform the two ligation steps simultaneously or sequentially at optimal rates.Features of the landmark technology of this disclosure as implemented more generally into spatial transcriptomics

[0105] More generically, the implementation of landmark sequence technology and the benefits thereof into in situ hybridization and synthesis methodology may incorporate any of the following features in any combination.

[0106] Target probes are used that contains a target recognition sequence that hybridizes specifically to the target nucleic acids of interest to the user in a sample. They contain a detector sequence that is different for each different target nucleic acid, and a copy of a landmark sequence that is common for each target nucleic acid. The sample is contacted with the target probes under conditions wherein the target recognition sequence of each target probe hybridizes specifically to the respective target nucleic acid in the sample. The sample is processed (for example, by ligation) to form circular DNAs from each target probe hybridized to a target nucleic acid in the sample, whereby each circular DNA comprises at least the detector sequence for the respective target nucleic acid to which the target probe is hybridized, plus a copy of the common landmark sequence. Locations of copies of the landmark sequence and locations of each of the different target nucleic acids in the sample are determined. The image is processed or characterized such that the respective target nucleic acid in the sample is located where the corresponding detector sequences are within a predefined radius from a copy of the landmark sequence. Detector spots appearing more than the predefined radius away from a landmark sequence may be designated as false positives, and are optionally filtered out or suppressed from the image.

[0107] Each of the target probes may be a linear probe that contains the target recognition sequence, the detector sequence, and the landmark sequence. Optionally, a first portion and a second portion of the landmark sequence are at the 5’ and 3’ ends of the target probes, respectively, which are assembled into a complete landmark sequence upon formation of the circular DNA.

[0108] Alternatively, each of the target probes may be a padlock probe that contains the target recognition sequence, the detector sequence, and the landmark sequence, wherein a first portion and a second portion of the target recognition sequence are at the 3’ and 5’ ends of the padlock probes, respectively, which are assembled upon formation of the circularDNA. Alternatively, each of the target probe is a split probe that comprises a first part and a second part that hybridize specifically to adjacent positions on the respective target nucleic acid, wherein a first portion and a second portion of the target recognition sequence are at the 3’ and 5’ ends of the first part and the second part of the split probe, respectively, which are assembled upon formation of the circular DNA; and a first portion and a second portion of the landmark sequence are at the 5’ and 3’ ends of the first and second part of the split probe, respectively, which are assembled into a complete landmark sequence upon formation of the circular DNA.

[0109] Optionally, to increase signal intensity, the circular DNAs are amplified in situ by rolling circle amplification before the locations of the detector sequences and copies of the landmark sequences in the sample are determined. Locations of copies of the landmark sequences in the sample may be determined directly or in the amplification product by contacting the tissue with a landmark identification probe bearing an optical label and containing a sequence that hybridizes specifically to the landmark sequence. The location of each of the detector sequences in the sample may be determined by contacting the tissue with a plurality of different detector readout probes each bearing an optical signal, wherein each of the detector readout probes contains a sequence that hybridizes specifically to one of the detector sequences bearing an optical label, wherein the optical label is different for each of the different target nucleic acids. Alternatively, the locations of each of the detector sequences in the sample is determined by sequencing by synthesis of each detector sequence in situ.

[0110] When fluorescent labels are used, each optical label emits fluorescence at an emission frequency when activated by light at an activation frequency. In the image obtained, different target nucleic acids may be represented at locations of the respective detector sequences by different colors. Locations of detector sequences that are not colocalized with a copy of the landmark sequence may be filtered out.[OHl] In situ hybridization procedures may be conducted in a plurality of cycles in which the locations of different target nucleic acids are identified in each cycle. For example, the sample is contacted at the same time with split probes or target probes for each of the nucleic acids being identified. All the probes may be ligated and / or amplified at the same time. Multiple readout cycles can be done for different detector sequences in each of the cycles.Reagents and kits

[0112] The various probes, enzymes, and other reagents used for sample processing put forth in this disclosure can be sold or distributed separately or together in any useful combination. Separate components may be provided in separate containers in a package or kit, or combined as a single reagent in instances where they may be employed effectively when together. Optionally, any such product or product combination may be packaged with or distributed in conjunction with information on the use of the product or combination in the practice of any of the methodology put forth in this disclosure.

[0113] A kit for identifying the locations of a plurality of different target nucleic acids in a sample using LISH Lock’n’Roll may contain one or more of the following reagents in any combination. For each of the different target nucleic acids in the sample the kit contains at least one split probe that comprises a first linear part and a second linear part. The 3 ’ end of the first part and the 5’ end of the second part of each split probe are configured to hybridize specifically to the respective target nucleic acid at positions adjacent to each other, and wherein the 5’ end of the first part and the 3’ end of the second part of each split probe are configured to hybridize specifically to a common bridge probe at positions adjacent to each other. The common bridge probe comprises a sequence that is complementary to a landmark sequence, and at least one of the first part and the second parts of each split probe comprises a detector sequence that is different for each of the target nucleic acids

[0114] The kit further contains said common bridge probe, a ligase that ligates the 3’ end of the first part to the 5’ end of the second part of each split probe when they are hybridized to adjacent positions in the target nucleic acid, and a ligase that ligates the 3’ end of the first part to the 5’ end of the second part of each split probe when they are hybridized to adjacent positions on the common bridge probe. Where the first and second part of the split probe are both DNA, then the ligase may be the same enzyme. The two ligases and the bridge probe may be supplied separately, or as a combined reagent, optionally including the DNA polymerase. To perform an amplification step to increase signal intensity, the kit also contains a non-processive DNA polymerase and dNTPs, formulated to amplify circular DNAs formed from the split probes in situ. Depending on the implementation, the ligating components and the DNA amplification components may be combined as one reagent, which may be reformulated based on the enzymes’ requirements.

[0115] A kit for identifying the locations of a plurality of different target nucleic acids in a sample using other types of in situ hybridization or sequencing methodology may contain one or more of the following reagents in any combination. For each of the different target nucleic acids in the sample at least one target probe (for example, a linear or padlock probe).The target probe for each of the target nucleic acids contains a target sequence that hybridizes specifically to the respective target nucleic acid, a detector sequence that is different for each of the different target nucleic acids, and a landmark sequence that is common to each of the target probes.

[0116] Any of these kits may also contain a means for determining the locations of each copy of the landmark sequence in the sample, and a means for determining the location of each of the detector sequences in the sample. The means for determining the locations of each copy of the landmark sequence may be, for example, a landmark identification probe that bears an optical label and hybridizes specifically to the copies of the landmark sequence. The means for determining the locations of each of the detector sequences may be a plurality of detector readout probes that hybridize specifically to each of the detector sequences. At least some of the readout probes have different optical labels and / or are used in different detection cycles.Machine learning for spot identification

[0117] The high signal-to-background ratio produced by LISH Lock’n’Roll provides an opportunity to acquire additional information using a spot-finding algorithm. This generally gives higher accuracy than pixel-based image processing. The amplified signal from each rolony should be uniform, which means that the bit-to-bit signal variation is smaller. Consequently, the point spread function (PSF) for each spot contains additional information that can be used to remove false positive signals and duplicates.

[0118] For this aspect of the disclosure, a convolutional neural network (CNN) model is built to perform raw image to spotty, zero-background image conversion. The CNN is trained by interactively preparing a ground truth database in which LnR raw images are paired with reconstructed (“ground truth”) images.

[0119] FIG. 10A is a flowchart showing how the system operator may prepare the paired images for the database. FIG. 10B is a detail with a hybridization raw image paired with the reconstructed image obtained therefrom.

[0120] First, the user obtains experimental images from hybridization experiments done with LISH Lock’n’Roll. Preferably, a diverse range of contexts is used for the training data set (such as sample type, noise level, laser wavelength and power settings). For images acquired in sequential slices (Z-stacks), analysis is performed across adjacent slices to generate 2D frames.

[0121] Next, optical spots on the image are identified and characterized interactively by the operator using image processing software such as Fiji, a distribution of ImageJ2 software, bundling a variety of plugins to facilitate scientific image analysis. I mage J a Java-based image processing program developed by the NIH and LOCI (University of Wisconsin). ImageJ supports image processing functions such as logical and arithmetical operations between images, contrast manipulation, convolution, Fourier analysis, sharpening, smoothing, edge detection, and median filtering. The plugin TrackMate can be used to compute numerical features for each spot: for example, the mean, max, min and median intensity, the estimated radius and orientation for each spot, and how these features change through adjacent optical slices. Once the operator is satisfied with the reconstructed image, the location and attributes of each of the analyzed spots are stored, for example, in csv files. Spotty images with zero background are reconstructed electronically by drawing spots on blank images using information saved in the spot-feature csv files. The raw images and the reconstructed spotty images are then sliced, for example, into 128 x 128 small patches as the training data.

[0122] The CNN is then trained using the database of attributes of the paired raw and reconstructed patches. The number of cycles needed to complete training can be determined empirically. A subset of the paired images may be withheld (for example, 10 to 30% of the dataset) and kept as validation data to follow the accuracy of spot analysis through each training cycle (epoch). Theoretically, a model should show comparative performance on the training set and the validation set following training. A significant discrepancy between the two may indicate overfitting.

[0123] FIG. 11 shows the improvement of accuracy of spot recognition by the CNN through each training cycle. The Y axis is the accuracy metricA0, which represents the consistency between the ground truth image and the predicted image. It is calculated as follows:The summation runs over all the pixels in each image, and a small positive value e is added in the denominator to prevent zero division. The two curves represent accuracies evaluated from the training set and the validation set.

[0124] Training of the CNN model is considered complete when the validation accuracy converges and stops improving over more cycles. This can be reflected by a plateau in the training curve. Usually 30 to 50 cycles (epochs) can accomplish this goal.

[0125] Once the CNN model is trained, the operator can apply it to analyze new experimentally obtained image patches that have the same size as the training patches. For example, LnR images are sliced into 128 x 128 crops. The trained model is applied to the raw image patches to acquire 128 x 128 spotty image crops. The spotty image crops are then stitched together to create a full-sized image. Spot locations are extracted using conventional image processing steps (filtering and thresholding), which give reliable performance on zero- background images.

[0126] FIGS. 12A and 12B provide an illustration. FIG. 12A is the raw image.FIG. 12B is the corresponding a multi-channel output image from the spot recognition CNN model. In both images, different excitation channels are represented by pseudo colors.FIG. 12B shows excellent agreement with FIG. 12A on spot locations across all the channels — but with much less spot size variation, and much cleaner context (zero background).Spot filtering by comparing adjacent tissue slices

[0127] FIG. 13 illustrates how analysis of sequential slices of a tissue sample can be used as a further means of identifying and characterizing spots and removing false negatives.

[0128] In this example, for all spots of a given gene located in an optical slice (Zl, Z2 or Z3 in this case), their centroid locations (X and Y coordinates) are compared with the centroid locations of all spots of the same gene in an adjacent slice by pairwise comparison. Neighboring spots are identified in a slice-by-slice manner (Zl with Z2, and then Z2 with Z3). More specifically, if slice Zl has m number of spots and slice Z2 has n number of spots, a pairwise comparison between the coordinates of the gene spots in slice Zl and Z2 results in m times n number of pairwise comparisons. A threshold of maximum pairwise distance is then applied to identify fluorescent spots in adjacent Z-slices that are considered colocalized. In FIG. 13, Al colocalizes with A2, and B2 colocalizes with B3, but C2 stands alone.

[0129] Colocalized gene spots in adjacent Z-slices are then assumed to be generated from the same endogenous gene location within the sample. To avoid over-counting of genes and to improve the accuracy of the analysis results, the colocalized gene spots are merged into a single imputed gene spot. Gene spots located only in a single Z-slice are preserved tobe included in the final dataset after the spot merging process of colocalized spots is completed.

[0130] The process by which the characteristics of colocalized spots are merged to impute merged spot characteristics is as follows. A weighting factor for each colocalized spot is generated by calculating the ratio of each spot’s respective maximum fluorescence intensity to the sum of all spot maximum fluorescence intensities. After multiplying the maximum fluorescence intensity of the colocalized spots that are being merged by the weighting factor, the maximum fluorescence intensity of the imputed spot is calculated as the sum of all weighted maximum fluorescence intensities. The weighting factor is then also applied to the centroid coordinates of each spot in a colocalized set. The sum of all weighted coordinates per dimension (X, Y, and Z) is then calculated to define the imputed spot’s centroid location. The spot area for the imputed spot is defined as the maximum spot area amongst the colocalized spots undergoing the merging process. For example, if colocalized spots Al and A2 have spot areas of 10 and 5, respectively, the imputed, merged spot will have a spot area of 10.

[0131] After colocalized gene spots have been merged into a single imputed gene spot, the preserved data of gene spots located in a single Z-slice is combined with the imputed spot data. The final resulting dataset will be comprised of spot data from imputed, merged spots as well as data from spots located in a single Z-slice for each gene species. Alternatively, data from spots located in a single Z slice can be disqualified and omitted from the final image, as indicated in FIG. 13.

[0132] The CNN model described in a previous section of this disclosure is 2D based. It can be combined with the sequential slice analysis described in this section to obtain the benefits of both techniques. This is illustrated by the flow chart shown in FIG. 14. To identify the 3D location of LnR spots, the spots from different Z planes are merged, and their lateral positions cross checked along the depth.Devices for carrying out this technology

[0133] Veranome Biosystems LLC, an Applied Materials company, has developed the PISCES™ system, an apparatus and computerized control system for conducting mFISH. Incorporated herein by reference in its entirety as part of this disclosure is US 2021 / 0181111 Al. The system comprises the following components:• a flow cell to contain a sample to be exposed to fluorescent probes in a reagent;• a valve to control flow from one of a plurality of reagent sources the flow cell;• a pump to cause fluid flow through the flow cell;• a fluorescence microscope including a variable frequency excitation light source and a camera positioned to receive fluorescently emitted light from the sample;• an actuator to cause relative vertical motion between the flow cell and the fluorescence microscope;• a motor to cause to cause relative lateral motion between the flow cell and the fluorescence microscope; and• a control system.

[0134] The control system operates the components of the apparatus by performing the following actions as nested loops:• cause the valve to sequentially couple the flow cell to a plurality of different reagent sources to expose the sample to a plurality of different reagents,• for each reagent of the plurality of different reagents, cause the motor to sequentially position the fluorescence microscope relative to sample at a plurality of different fields of view,• for each field of view of the plurality of different fields of view, cause the variable frequency excitation light source to sequentially emit a plurality of different wavelengths,• for each wavelength of the plurality of different wavelengths, cause the actuator to sequentially position the fluorescence microscope relative to sample at a plurality of different vertical heights, and• for each vertical height of the plurality of different vertical heights, obtain an image at the respective vertical height covering the respective field of view of the sample having respective fluorescent probes of the respective regent as excited by the respective wavelength.EXAMPLE

[0135] Reagents and reaction conditions that have been used to perform the technologies put forward in this disclosure include the following:

[0136] Pretreatment of frozen tissue samples:1. Take coverslips seeded with cells out of the freezer and thaw at RT for 3-5 min2. Add 3 mL lx PBS and incubated for 1-3 min3. Permeabilize cells with 0.1% Triton™ X-100 Ix-PBS for 20 min at RT (room temperature)4. Wash twice with Ix-PBS and once with diethyl pyrocarbonate (DEPC) treated water5. Add 1-3 mL 0.1 N HCL; incubated for 10 min at RT6. Wash 3x with Ix-PBS

[0137] Additional treatment may be needed depending on the tissue type. For example, for brain tissue:7. Add ~3 mL 1% SDS in Ix-PBS buffer and incubate at 30°C for 15 min8. Wash 3x with Ix-PBS

[0138] LnR protocol1. Wash twice for 5 min each in 200μL Hyb wash buffer with RNase inhibitors (10μL / mL)2. Incubate in 30 μL of 200 nM LnR probe hybridization cocktail on hybridization gasket for 3 to 4 h at 45°C in a padded Petri dish3. Wash twice in 200μL hybridization wash buffer4. Add 30μL 200 nM bridge reaction mixture on hybridization gasket at 37°C for 1 h in a padded Petri dish5. Wash twice in 200μL hybridization wash buffer6. Buffer exchange in 200μL 029 buffer (lx)7. Incubate in 30 μL RNL2 - T4 DNA ligase - 029 polymerase reaction mixture on hybridization gasket at 30°C overnight in a padded Petri dish with RNase inhibitors (lOuL / mL)8. Transfer the coverslip to a clean Petri dish9. Wash three times with 3 mL 2X-SSC

[0139] Target probe hybridization cocktail (30μL )• Probe library (final cone 200 nM / probe)• Protector inhibitor: 0.3μL• Formamide (final 20%): 6μL• 20x SSC: 3μL• Water to 30μL volume

[0140] Bridge reaction mixture (30μL )• Bridge probe (2 pM stock, final cone 200 nM): 6μL• Protector inhibitor: 0.3μL• lOx PBS: 3μL• Water: 20.7μL

[0141] Triple enzyme reaction mixture (30μL )• T4 RNA ligase 2: 3μL• T4 DNA ligase: 1 μL• 029 DNA polymerase: 6μL• 100 mM dNTP mix (final: 1 mM): 0.3μL• 10 mM ATP (final: 1 mM): 3μL• Protector inhibitor: 0.3μL• 029 buffer lOx: 3μL• Water: 13.1μL

[0142] Hybridization wash buffer (1 mL)• Formamide (final: 20%): 200μL• 20x SSC (final: 2x SSC): 100μL• Water: 700μL

[0143] Cell permeabilization buffer (1 mL)• lOx PBS: 100μL• 1% Triton™ X-100 (final: 0.1%): 100 μL• Water: 800 μL

[0144] Suppliers:• T4 RNA ligase2: New England Biolabs• T4 DNA ligase: New England Biolabs• 029 polymerase: Lucigen Corp.• dNTP: Biosystems GeneAmp™• Protector RNase inhibitor: Roche SUPERase.In™• ATP: New England Biolabs* * * * *Practice of the invention

[0145] The technology provided in this disclosure and its use are described within a hypothetical understanding of general principles of nucleic acid chemistry and tissue analysis, in silico image processing, and machine learning. These discussions are provided for the edification and interest of the reader, and are not intended to limit the practice of the claimed invention. All of the products and methods claimed in this application may be used for any suitable purpose without restriction, unless otherwise indicated or required.

[0146] While the invention has been described with reference to the specific examples and illustrations, changes can be made and equivalents can be substituted to adapt the technology to a particular context or intended use as a matter of routine development and optimization and within the purview of one of ordinary skill in the art, thereby achieving benefits of the invention without departing from the scope of what is claimed and their equivalents.

Claims

CLAIMSThe invention claimed is:

1. A method of identifying locations of a plurality of different target nucleic acids in a cell or tissue sample, the method comprising: a) obtaining for each of the target nucleic acids in the sample a split probe that comprises a first part and a second part, each of the first and second parts being a first and a second linear oligonucleotide not attached to the other, wherein a 3 ’ end of the first part and a 5 ’ end of the second part of each split probe, or both, contain target recognition sequences configured to hybridize specifically to the respective target nucleic acid at positions adjacent to each other, and wherein a 5’ end of the first part and a 3’ end of the second part of each split probe are configured to hybridize specifically to a common bridge probe at positions adjacent to each other, wherein at least one of the first part and the second part of each split probe comprises a detector sequence that is different for each different target nucleic acid; b) contacting the sample with the split probes under conditions where the first part and the second part of each split probe hybridize specifically to the target nucleic acid for the split probe; c) for each of the split probes that have a first part and a second part hybridized adjacent to each other on a target nucleic acid, ligating the 3’ end of the first part to the 5’ end of the second part of the split probe; d) contacting the sample with the common bridge probe under conditions where the bridge probe hybridizes specifically to the first part and the second part of each split probe that are adjacent on a target probe in the sample; e) for each of the split probes that have a first part and a second part hybridized adjacent to each other on a copy of the common bridge probe, ligating the 5’ end of the first part to the 3 ’ end of the second part, wherein the ligating in step (c) and the ligating in step (e) forms a circular single- stranded DNA molecule from the respective split probe parts, wherein the circular single-stranded DNA molecule contains at least the detector sequence(s) from the split probe parts and a copy of a common landmark sequence;f) amplifying each of the circular single-stranded DNA molecules in situ to form an amplification product; g) determining locations of copies of the common landmark sequence in the amplification products in the sample; h) determining locations of each of the detector sequences in the amplification products in the sample; and i) identifying the locations of each of the different target nucleic acids in the sample as being the locations where the respective detector sequences are each within a predefined radius from a copy of the common landmark sequence.

2. The method of claim 1, wherein the bridge probe contains the common landmark sequence.

3. The method of claim 1, wherein the common landmark sequence is formed from a 3’ portion of the first part and a 5’ portion of the second part of the split probe when the 3’ end of the first part and 5’ end of the second part are ligated together.

4. The method of claim 1, wherein the ligating in step (c) and the ligating in step (e) are performed concurrently.

5. A method of identifying locations of a plurality of different target nucleic acids in a cell or tissue sample, the method comprising: a) obtaining for each of the target nucleic acids a target probe, wherein the target probe for each of the target nucleic acids contains a target recognition sequence that hybridizes specifically to the respective target nucleic acid, a detector sequence that is different for each different target nucleic acid, and a copy of a landmark sequence that is common for each target nucleic acid; b) contacting the sample with the target probes under conditions wherein the target recognition sequence of each target probe hybridizes specifically to the respective target nucleic acid in the sample; c) processing the sample to form circular single-stranded DNA molecules from each target probe hybridized to a target nucleic acid in the sample, whereby each circular single-stranded DNA molecule comprises at least the detector sequence forthe respective target nucleic acid to which the target probe is hybridized and a copy of the common landmark sequence; d) determining locations of copies of the common landmark sequence in the circular single-stranded DNA molecules in the sample; e) determining locations of each of the detector sequences in the circular single- stranded DNA molecules in the sample; and f) identifying the locations of each of the different target nucleic acids in the sample as being the locations of the respective detector sequences only where the detector sequences are within a predefined radius from a copy of the common landmark sequence.

6. The method of claim 5, wherein each of the target probes is a linear probe that contains the target recognition sequence, the detector sequence, and the common landmark sequence, wherein a first portion and a second portion of the common landmark sequence are at a 5’ and a 3’ ends of each of the target probes, respectively, which are assembled into a complete landmark sequence upon formation of the circular single-stranded DNA molecule.

7. The method of claim 5, wherein each of the target probes is a padlock probe that contains the target recognition sequence, the detector sequence, and the landmark sequence, wherein a first portion and a second portion of the target recognition sequence are at the 3’ and 5’ ends of each of the padlock probes, respectively, which are assembled upon formation of the circular single-stranded DNA molecule.

8. The method of claim 5, wherein each of the target probe is a split probe that comprises a first part and a second part that hybridize specifically to adjacent positions on the respective target nucleic acid, with the first part annealed at a 5’ position with respect to the second part, wherein a first portion and a second portion of the target recognition sequence are at the 5’ end of the first part and the 3’ end of the second part of the split probe, respectively, which are assembled upon formation of the circular single-stranded DNA molecule; and a first portion and a second portion of the landmark sequence are at the 3’ end of the first part and the 5’ end of the second part of the split probe,respectively, which are assembled into a complete landmark sequence upon formation of the circular single-stranded DNA molecule.

9. The method of any of claims 5 to 8, wherein the circular single-stranded DNA molecules are amplified in situ by rolling circle amplification (RCA) before the locations of the detector sequences and copies of the landmark sequences in the sample are determined.

10. The method of claim 9, wherein the RCA is primed from a sequence contained in the bridge probe.

11. The method of any of the preceding claims, wherein the locations of copies of the landmark sequences in the sample are determined by contacting the tissue with a landmark identification probe bearing an optical label and containing a sequence that hybridizes specifically to the landmark sequence.

12. The method of any of the preceding claims, wherein the location of each of the detector sequences in the sample is determined by contacting the tissue with a plurality of different detector readout probes each bearing an optical signal, wherein each of the detector readout probes contains a sequence that hybridizes specifically to one of the detector sequences bearing an optical label, wherein the optical label is different for each of the different target nucleic acids.

13. The method of claim 11 and / or claim 12, wherein each optical label emits fluorescence at an emission frequency when activated by light at an activation frequency.

14. The method of any of the preceding claims, further comprising generating an image of the sample in which different target nucleic acids are represented at locations of the respective detector sequences by different colors, but locations of detector sequences that are not colocalized with a copy of the landmark sequence are filtered out.

15. The method of any of claims 1 to 13, wherein the locations of each of the detector sequences in the sample is determined by sequencing by synthesis of each detector sequence in situ.

16. The method of any of the preceding claims, conducted in a plurality of cycles wherein the locations of the same or different target nucleic acids are identified in each cycle.

17. The method of any of the preceding claims, wherein the sample is contacted concurrently with split probes or target probes for each of the target nucleic acids being identified, following which the locations of the detector sequences in the sample are determined in multiple cycles using different detector readout probes for different detector sequences in each of the cycles.

18. The method of any of the preceding claims, wherein the different target nucleic acids are mRNA transcripts of different genes.

19. A kit for identifying the locations of a plurality of different target nucleic acids at a plurality of locations in a cell or tissue sample in a method according to the method of any of claims 1 to 4, the kit comprising: a) for each of the different target nucleic acids in the sample at least one split probe that comprises a first linear part and a second linear part, wherein the 3’ end of the first part and the 5’ end of the second part of each split probe are configured to hybridize specifically to the respective target nucleic acid at positions adjacent to each other, and wherein the 5’ end of the first part and the 3’ end of the second part of each split probe are configured to hybridize specifically to a common bridge probe at positions adjacent to each other, wherein the common bridge probe comprises a copy of a landmark sequence, and at least one of the first part and the second parts of each split probe comprises a detector sequence that is different for each of the target nucleic acids; b) said common bridge probe; c) one or more reagents comprising a ligase, formulated to ligate the 3’ end of the first part to the 5’ end of the second part of each split probe when they are hybridized to adjacent positions in the target nucleic acid, and to ligate the 3’ end ofthe first part to the 5’ end of the second part of each split probe when they are hybridized to adjacent positions on the common bridge probe; e) optionally an amplification reagent comprising a DNA polymerase and dNTPs, formulated to amplify circular single-stranded DNA molecules formed from the split probes in situ by rolling circle amplification; f) a means for determining the locations of each copy of the landmark sequence in the sample, and g) a means for determining the location of each of the detector sequences in the sample.

20. A kit for identifying the locations of a plurality of different target nucleic acids at a plurality of locations in a cell or tissue sample in a method according to the method of any of claims 1 to 15, the kit comprising: a) for each of the different target nucleic acids in the sample at least one target probe, wherein the target probe for each of the target nucleic acids contains a target sequence that hybridizes specifically to the respective target nucleic acid, a detector sequence that is different for each of the different target nucleic acids, and a landmark sequence that is common to each of the target probes; b) a means for determining the locations of each copy of the landmark sequence in the sample, and c) a means for determining the location of each of the detector sequences in the sample.

21. The kit of claim 19 or 20, wherein the means for determining the locations of each copy of the landmark sequence is a landmark identification probe that bears an optical label and hybridizes specifically to the copies of the landmark sequence; and the means for determining the locations of each of the detector sequences is a plurality of detector readout probes that hybridize specifically to each of the detector sequences and each bearing an optical label, wherein at least some of the detector readout probes have different optical labels.

22. An improvement in spatial transcriptomics,wherein the spacial transcriptomics is a process that comprises contacting a cell or tissue sample with a plurality of different target probes that each hybridize in a sequence specific manner to a target nucleic acid in the sample, thereby producing a target specific spot in one or more images of the sample at or around the location of each corresponding target nucleic acid; wherein the improvement comprises including in each target probe a copy of a landmark sequence that is common to at least some of the target probes and does not hybridize to nucleic acids in the sample, wherein each copy of the landmark sequence forms a landmark spot that is distinguishable from the target specific spots in one or more images of the sample at or around the location of each of the target probes in the sample; and wherein the images are analyzed by a process that includes identifying the locations of target specific spots that are each within a predefined radius of one of the landmark spots and / or excluding target specific spots that each are not within a predefined radius of a landmark spot.