Methods for mapping immune cell clones in situ

The method addresses the challenge of distinguishing closely related RNA sequences by designing specific nucleic acid probes based on a phylogenetic tree hierarchy, enhancing the ability to map immune cell clones with high resolution and accuracy within tissues.

WO2025174923A1PCT designated stage Publication Date: 2025-08-21CHILDRENS MEDICAL CENT CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/015663
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-13
Filing Date
2025-02-13
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing methods such as MERFISH struggle to distinguish between closely related RNA sequences, particularly BCR and TCR sequences of immune cells, limiting the ability to map clonal dynamics and spatial organization within tissues.

Method used

A method is developed to design nucleic acid probes that discriminate between closely related RNA sequences by identifying a set of targets within a similarity threshold, defining a hierarchy based on sequence similarity, and designing probes that can specifically hybridize to one node while minimizing hybridization to other nodes, using a phylogenetic tree approach and iterative processes to ensure probe specificity.

Benefits of technology

The method enables the discrimination of two to three orders of magnitude more B-cell clones with single-cell spatial resolution, allowing for accurate mapping of clonal dynamics and interaction within tissues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025015663_21082025_PF_FP_ABST
    Figure US2025015663_21082025_PF_FP_ABST
Patent Text Reader

Abstract

The technology described herein is directed to methods of preparing nucleic acid probes that can discriminate between closely related RNAs for use in fluorescent in situ hybridization techniques, i.e., FISH and MERFISH.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 701039-192190WOPT METHODS FOR MAPPING IMMUNE CELL CLONES IN SITU CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No.63 / 552,813 filed February 13, 2024, the contents of which are incorporated herein by reference in their entirety. SEQUENCE LISTING

[0002] The instant application contains a Sequence Listing which has been submitted in XML format via Patent Center and is hereby incorporated by reference in its entirety. Said XML copy, created on February 10, 2025, is named 701039-192190WOPT_SL.xml and is 32,907 bytes in size. TECHNICAL FIELD

[0003] The technology described herein relates to methods for identifying closely related cell types within the same family, genus, or germline. GOVERNMENT SUPPORT

[0004] This invention was made with government support under Grant Number GM143277, awarded by the National Institutes of Health. The Government has certain rights in the invention. BACKGROUND

[0005] The function of adaptive immune cells, such as B and T cells, can be differentiated by the specific sequence of the cognate B or T cell receptor (BCR or TCR) that they express. Critically important to the function of these cells within tissues is not just the BCR or TCR they express but the specific cells and environment in which they interact, which is determined by their location in the tissue. Groups of B cells and T cells that carry the same or similar sequences of their BCR or TCR are known as clones, and there are fluorescence microscopy methods that can provided limited discrimination between B cell clones, for example, by leveraging genetically modified animals that have B cells that express different combinations of fluorophores (see, e.g., Degn et al., Cell 170: 913-926 (2017); Tas et al., Science 351 (6277): 1048-1054 (2016)). However, these methods are limited in the number of clones they can 1 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT discriminate, their ability to link clonality to sequence features of the corresponding BCR, and their ability to jointly profile the expression of other molecules within these cells and other cells within the tissue that can report on features such as cell type, cell state, or cell activation.

[0006] By contrast, single-cell sequencing methods can distinguish large numbers of B-cell or T-cell clones by measuring small differences in BCR or TCR sequence. In some cases, these methods can also profile large numbers of nucleic acid molecules and provide information on cell type and state. Commercially available products are designed to facilitate this process. However, such methods must first dissociate cells from tissues before sequencing; thus, they are unable to map the spatial organization of individual B- or T-cell clones within tissues and determine the specific cells with which these clones are interacting.

[0007] Spatial transcriptomic methods that capture RNA molecules and then spatially barcode the cDNA molecules constructed from them (e.g., Visium, Slide-Seq) have been adopted to provide limited discrimination of different B and T cell clones (see, e.g., Liu et al., Immunity 55: 1940-1952 (2022); Engblom et al., Science 382 (6675): eadf8486 (2023)). In these methods, the spatial distribution of B or T cell clones can be placed in the spatial context of tissue slices. However, these methods do not have single-cell resolution and thus cannot unambiguously place specific T or B cells in the context of the specific cells that surround them.

[0008] By contrast, Multiplexed-error robust fluorescence in situ hybridization (MERFISH) allows users to distinguish different types of cells, such as immune cells, based on the types and levels of nucleic acids present or expressed within those cells and can do so with single- cell (and sub-cellular) spatial resolution. However, a current limitation with MERFISH is that it cannot distinguish between highly similar RNA sequences. As BCR and TCR sequences can be similar in sequence, MERFISH, is currently unable to distinguish B-cell and T-cell clones and map their expression within intact tissues. SUMMARY

[0009] A new variant of MERFISH described herein permits the discrimination of different, although closely related RNAs, including, but not limited to discrimination of different B cells and T cells based upon the sequence of the BCR or TCR encoded or expressed by these cells, and as a result map the clonal dynamics of such cells in situ. This approach can detect two to three orders of magnitude more B-cell clones than prior methods while retaining spatial resolution. 2 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT

[0010] The central challenge that is addressed in the technology disclosed herein is the use of in situ hybridization (ISH), fluorescence in situ hybridization (FISH), or variants such as MERFISH to distinguish different nucleic acids from families of nucleic acids that share high degrees of sequence identity or homology. For example, BCR variants are comprised of a heavy and a light chain, and both the sequence of the heavy and light chain are, in part, determined by the V(D)J recombination process in which individual specific genomic elements, i.e. V, D, or J regions, are selected and combined to form the variable sequences of the heavy and / or light chain of the receptor. Each element, e.g. the V region, is drawn from a pre-existing set of possible elements that often share various and sometimes high degrees of sequence similarity.

[0011] ISH and ISH-based methods often leverage nucleic acid oligonucleotide probes that are identical or nearly identical in sequence to the reverse complement of one region of the molecule of interest. Base pairing of these probes to this region of complementary is what directs a given probe to its specific binding partner in the sample of interest. The specific binding of a probe to its target versus non-specific binding to other elements in the sample is often promoted by adjusting the conditions of hybridization to specifically favor proper over non-proper binding. Factors that affect nucleic acid hybridization specificity and considerations for optimizing the specificity of nucleic acid hybridization are discussed, for example, in Zhang et al., Nat. Chem.4: 208-2014 (2012), which is incorporated herein by reference. However, the difference in the specificity or propensity of a probe to bind to a sequence that represents its exact target versus one that differs by a small or modest number of changes in the nucleic acid sequence is often modest, even under the optimal conditions for specific binding. In this context, it has been challenging to design nucleic acid probes that can discriminate individual members of families of homologous targets, as many possible nucleic acid probes designed to bind to one specific target will have a high propensity to bind to multiple possible targets from this family.

[0012] It should be understood that in situ hybridization can be used to detect nucleic acids present in a sample, and that such nucleic acids can be DNA or RNA. Detection of DNA identifies sequences encoded by a given cell, while detection of RNA provides information regarding what genes are expressed at the transcriptional level. While description and Examples herein may refer to detection or quantitation of RNA, it should be understood that the approaches described herein for designing, preparing and using probes that distinguish closely-related RNA sequences via in situ hybridization can apply equally well to the design, preparation and use of probes that distinguish closely-related DNA sequences. For 3 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT conciseness, embodiments are described in the following that refer to RNAs as the targets, but such reference and the probe design and use referred to should be viewed as applying to DNA as well.

[0013] In one aspect, described herein is a method of preparing a nucleic acid probe that discriminates between its target molecule and a closely related non-target molecule in a sample in a ISH assay, the method comprising: a) identifying the set of targets in the sample that are within a chosen threshold for similarity to the target molecule sequence; b) defining a hierarchy for the set of targets identified in (a) that classifies targets in the set on the basis of similarity to each other, such that all targets in the set are grouped into nodes by similarity; c) identifying a sequence for a nucleic acid probe that can discriminate in a FISH assay between targets in different nodes identified in (b), wherein a probe that can discriminate between targets in different nodes will hybridize, under ISH hybridization conditions, to one or more targets in that node but with limited propensity to any target in any other node; and d) preparing a nucleic acid probe comprising the probe sequence identified in (c).

[0014] In one embodiment described herein where the targets are, for example, RNA, the step (a) of identifying the set of RNAs in the sample that are within a chosen threshold for similarity to the target RNA molecule sequence comprises i) aligning target RNA sequence pairwise with each other RNA sequence expressed in the sample, wherein RNAs with target sequence identity greater than the threshold are selected for the set; or ii) dividing each RNA sequence expressed in the sample into a series of all sub-sequences of a given length k (k- mers), wherein the degree of similarity between the target RNA and any non-target RNA expressed in the sample is defined by the fraction of k-mers found in either the target RNA sequence or the non-target RNA sequence that are also found in both sequences, and wherein RNAs with a degree of similarity greater than the threshold are selected for the set.

[0015] In another embodiment described herein, step (b) of defining a hierarchy for the set of, for example, RNAs identified in (a) that classifies RNAs in the set (a) on the basis of similarity to each other comprises generating a phylogenetic tree wherein individual RNAs are grouped into different branches and nodes on the tree.

[0016] In another embodiment described herein, the phylogenetic tree is computed by defining pairwise similarity between individual, for example, RNA sequences expressed in the sample, and applying an agglomerative algorithm to systematically group RNAs into branches and nodes based on the RNAs, or nodes comprised of RNAs, that are most similar.

[0017] In another embodiment described herein, the agglomerative algorithm employs an iterative process comprising: i) computing the pairwise similarity between all RNAs or 4 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT existing nodes; ii) selecting the pair with the greatest similarity, and grouping these together to form a new node, wherein the pair can represent two RNAs, two existing nodes, or an RNA and a node; and iii) repeating steps (i) and (ii) until all RNAs or nodes have been grouped together.

[0018] In another embodiment described herein, in step (i), a) the similarity, for example, between a given RNA and a node is computed by averaging the pairwise similarity between each RNA within a node and the given RNA; b) the similarity between two nodes is computed by averaging the pairwise similarity between all RNAs in the first node and all RNAs in the second node; and c) a node can consist of multiple other nodes, wherein an RNA is considered to be part of a first node if it is part of a second or third node comprised by the first node, and if the second or third nodes themselves contain nodes, then any RNAs within the second or third nodes or their sub-nodes is considered to be part of the first node.

[0019] In another embodiment described herein, the first iteration of the iterative process uses only, for example, RNAs, creating one node by grouping the most similar RNAs, and each subsequent iteration creates another node by grouping two RNAs or the node created in the previous round to another RNA, until all nodes and RNAs have been grouped.

[0020] In another embodiment described herein, step (c) of identifying a sequence for an ISH probe that can discriminate in a ISH assay between target, for example, RNAs in different nodes identified in (b) comprises: a) for each RNA target, consider all possible probe sequences that are of a given length and satisfy a given set of free energy of binding parameters for a probe to its target sequence; b) determine the specificity of each possible probe sequence of (a) for the RNA, wherein specificity is determined by considering all possible fragments of length k2for the possible probe sequence and computing the fraction of k-mers within that possible probe sequence that are found in any other RNA sequence expressed in the sample, wherein the probe is considered specific if this fraction is below a given threshold; c) sum the number of unique specific probes determined in (b), wherein: when the number of specific probes determined in (b) is greater than a predetermined threshold, the target RNA is designated as distinguishable from all others expressed in the sample via one or more of such probes; and when the number of specific probes is below the predetermined threshold, the target RNA is considered not distinguishable from others expressed in the sample via one or more of such probes, d) grouping all RNAs designated not distinguishable with the RNA or nodes to which they are assigned based on the hierarchy defined in this aspect (i.e., at step (b) of paragraph

[0011] ), and repeating steps (a) – (c), but considering the RNAs grouped into a node as a single object, and e) considering each 5 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT possible probe within a given RNA, computing its specificity by calculating the fraction of k- mers within that probe found in all other RNA except the RNAs found within the node to which that RNA has been assigned, wherein a probe is considered specific if less than a given fraction of the k-mers within it are found in RNAs outside of that node, and if all RNAs within a node have a number of specific probes greater than a predetermined threshold, the node is considered to be distinguishable from all others in the sample, and if any RNA within the node does not have a number of specific probes greater than the predetermined threshold, the node is considered indistinguishable and grouped with the next nearest RNA or node based on the hierarchy defined in this aspect (i.e., at step (b) of paragraph

[0011] ), whereby a set of nodes or RNAs that are designated distinguishable is produced, with a set of ISH nucleic acid probes to each RNA that can discriminate that RNA from all other sequences not found within the node in which it has been grouped. This approach is akin to redefining the hierarchical tree in a manner that permits distinguishing targets in a group from targets in another group with ISH.

[0021] In another embodiment described herein, each ISH probe further comprises a barcode element.

[0022] In another embodiment described herein, all targets within a node designated distinguishable are assigned the same barcode and some or all of the ISH probes directed at sequences within that node or to that RNA contain the barcode element.

[0023] In another embodiment described herein, each barcode element comprises a MERFISH barcode comprised of concatenated readout sub-sequences that can be specifically detected in a FISH assay, wherein detection of a given readout sub-sequence in the FISH assay is read as a “1” in a binary barcode.

[0024] In another embodiment described herein, ISH probes are designed by a) assigning a unique binary barcode to each distinguishable target nucleic acid or node; and for each distinguishable target nucleic acid or node, b) creating a set of MERFISH probe sequences by concatenating to all or a subset of the ISH probes against that target nucleic acid the sequences of the barcode readout elements associated with the bits in which there is a value of “1” in the barcode assigned to the node.

[0025] In another aspect, described herein is a method of detecting a target nucleic acid, for example, RNA or set of target RNAs in a fluorescence in situ hybridization assay wherein the target RNA or set of target RNAs is discriminated from one or more closely related non- target RNAs in the same sample, the method comprising hybridizing an ISH probe or set of 6 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT ISH probes prepared according to any one of the embodiments described herein to a sample, and detecting signal from hybridized probe.

[0026] In one embodiment described herein, detection is preformed via iterative hybridization and detection of members of a set of fluorescently labelled probes specific for barcode elements on the set of ISH probes.

[0027] In another embodiment described herein, fluorescent labels are removed between iterative rounds of hybridization and detection.

[0028] In another embodiment described herein, closely related sequences have at least 50% sequence identity.

[0029] In another embodiment described herein, in (i) the threshold is 50% sequence identity.

[0030] In another embodiment described herein, in (ii) the threshold degree of similarity is that 50% of the k-mers found in either sequence are found in both sequences.

[0031] In another embodiment as described herein, the closely related sequences are T cell receptor (TCR) variable (V) region or B cell receptor (BCR) V region coding sequences.

[0032] In another embodiment described herein, the free energy parameters comprise predicted melting temperature and GC content.

[0033] In another embodiment described herein, considered probes have 40% to 60% GC content, inclusive.

[0034] In another embodiment described herein, considered probes have a predicted melting temperature between 55 and 75oC, inclusive.

[0035] In another embodiment described herein, in step (e), no more than one k-mer is found in RNAs outside that node.

[0036] In another embodiment described herein, the considered probe length is at least 20 nt. In another embodiment described herein, the considered probe length is at least 30 nt. In another embodiment described herein, the considered probe length is at least 40 nt.

[0037] In another embodiment as described herein, the predetermined threshold number of specific probes in (c) is 10.

[0038] In another aspect described herein is a computer implemented method of preparing a nucleic acid probe according to the method of any one of the embodiments described herein.

[0039] In another aspect described herein is a computer-readable medium having computer- readable signals stored thereon that define instructions which, as a result of being executed in a computer system having a processor and a user interface including a display and an input device, instruct the computer system to perform a method according to any one of the embodiments described herein. 7 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT BRIEF DESCRIPTION OF THE DRAWINGS

[0040] This patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0041] FIG.1 shows the process by which MERFISH builds fluorescent optical barcodes through multiple rounds of single-molecule FISH imaging in order to localize and identify large numbers of different RNAs; an illustrative example of mapping more than one hundred different RNA molecules simultaneously within a single slice of tissue; and the ability to leverage these measurements to define many different cell types and states while mapping their organization in tissues.

[0042] FIG.2 shows a schematic that depicts how MERFISH can be used to distinguish different B cell clones within intact tissues by identifying the pair of variable heavy (VH) and variable light (VL or VK) chains expressed within individual B cells. By also performing MERFISH targeting other RNAs in the same sample, it is also possible to identify and map the surrounding, non-B cells. A similar strategy can be used to also identify and map T cell clones.

[0043] FIG.3 shows a schematic illustrating how a phylogenetic-like tree created on the degree of sequence homology can sort different RNA molecules into a hierarchy defined by shared and non-shared sequences.

[0044] FIG.4 shows two schematic models of how a hierarchical tree based on sequence similarity can be used to identify nucleic acid sequences or groups of nucleic acid sequences that share sufficient dissimilarity from all other groups such that sufficient FISH probes can be designed. A node can represent a single nucleic acid sequence or a group of sequences, and the similarity tree is based on a measure of the sequence similarity between sets of sequences within nodes. If it is determined that the sequences within a given node are not sufficiently different from all sequences in other nodes for the design of FISH probes, then the similarity tree can be used to identify the most similar node to the given node. These two nodes can then be grouped to effectively form a new node. This node can then be re- evaluated to determine if the sequences within this node are sufficiently different from all sequences in all other nodes for sufficient probe design. This process can be iterated to identify a set of nodes, which each contain a set of sequences that can be sufficiently discriminated from all other sequences in all other nodes. This iterative process can be 8 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT conceptualized, in method 1, as walking up branches in the tree to identify the appropriate nodes, or, in an alternative conception, method 2, to reconstruct the tree at each stage after sequences are grouped into a new node.

[0045] FIG.5 shows an example of a similarity tree calculated for a set of mouse VH genes. The bars show the number of predicted FISH probes that could be calculated for each VH gene, prior to grouping into nodes. Assuming an example threshold of 10 probes per gene, multiple VH genes with insufficient probes can be identified. These genes would then be grouped into a set of nodes such that at least 10 probes per gene within each node can be designed, in the process described above.

[0046] FIG.6 shows an example of the partial output of the proposed method for designing FISH probes to discriminate sets of VH genes. Individual nodes (e.g. node 317, 275, or 88) contain multiple VH genes after the grouping approach discussed above. Exemplary FISH homology regions, termed Target Regions, are listed and correspond to SEQ ID NOs: 1-16, as are MERFISH barcodes assigned to individual nodes, which correspond to SEQ ID NOs: 17-28. MERFISH primary probes that would label each of these nodes with the listed barcodes are also provided and correspond to SEQ ID NOs: 29-35.

[0047] FIG.7 shows a model where expression plasmids that contain pre-arranged BCR are transfected into cell cultures, cells are fixed and permeabilized, FISH probes allowing for the discrimination of different heavy-chain V region nodes are bound to the sample, and then the readout sequences associated with these probes are measured through sequential rounds of FISH. The images illustrate the imaging rounds used to detect each of the bits in the listed barcode, and the heatmap shows the intensity measured in each bit (columns) for each of several hundred different cells (rows) measured in this experiment. The intensity pattern detected is consistent with the node in which the single expressed VH was grouped.

[0048] FIG.8 shows an example of simultaneously detecting multiple known VH genes, each expressed in different cells within cell culture using the procedure described in Fig.7. The heat map shows the intensity measured in each bit (columns) for different cells (rows). The tSNE representation reveals that a large diversity of different nodes can be detected in this experiment, consistent with the nodes associated with the known VH genes expressed.

[0049] FIG.9 shows the intensity profiles observed for one of the clusters identified in Fig.8 consistent with the expected intensity pattern for a node containing the IGHV 5-9*01 gene, which was included in the set of VH genes expressed in this experiment.

[0050] FIG.10 illustrates in vivo verification of detecting a specific B cell clone in situ with this approach. Ileal slices were taken from a transgenic mouse (564Igi; C57BL / 6 background) 9 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT with preferential expression of a specific VH and VL gene. These slices were stained with probes against the expected VH and VL genes.

[0051] FIG.11 illustrates in vivo verification of detecting a specific B cell clone in situ using probes designed to detect all mouse VH and VL / VK nodes. The images illustrate one image associated with a single MERFISH bit associated with the barcodes that define these nodes. The heatmap illustrates that the node that contains the expected VH gene (564H) is the dominant detected node, as expected for this transgenic mouse (564Igi; C57BL / 6 background).

[0052] FIG.12 illustrates the in vivo detection of VH signals from naïve B cells in a mouse Peyer’s patch. Naïve B cells express much lower levels of the BCR than plasma cells, illustrating that this approach can faithfully detect even single molecules of the BCR. Images of two of the four bits associated with the expected VH are shown in gray and white.

[0053] FIG.13 illustrates the series of 14 images used to detect the full 14-bit barcode that discriminates all VH nodes in a slice of the mouse ileum. Included is a DAPI stain and a set of FISH probes that target the constant region associated with IgA, illustrating that the observed signals are from IgA+ plasma cells.

[0054] FIG.14 illustrates that the detection of VH and VL / VK nodes can be performed in the same sample in which MERFISH against mRNAs are measured. A slice of the mouse ileum was simultaneously stained and imaged to detect all VH nodes, all VL / VK nodes, and ~600 mRNAs. The top row identifies plasma cells within this slice based on the VH node (left), the VL / VK node (middle left), or the clone (middle right) identified, where clone is defined by the unique combination of VH and VK / VL nodes found within each plasma cell. VH and VK / VL nodes identified by MERFISH correlate strongly with the VH and VK / VL gene segments identified by BCR-sequencing of paired slices (right). The spatial location of all measured RNAs is mapped in dark gray while the distribution of six example mRNAs are various shades of lighter gray (bottom left and middle) in the same slice where B cell clones were identified. The bottom right panel shows a UMAP representation of the diversity of cell types identified via the MERFISH measurements in this example slice.

[0055] FIG.15 shows an example of how BCR-MERFISH can profile the abundance of plasma cells in the mammalian gastrointestinal tract. The fraction of total plasma cells measured across multiple replicate slices of the ileum from mice harboring a specific pathogen-free (SPF) microbiome or harboring no microbiome (germ-free, GF) associated with specific plasma cell clones determined by BCR-MERFISH (top). Lined circles mark plasma cell clones that are classified as public clonotypes that commonly arise in unrelated 10 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT GF mice. BCR-MERFISH can be used to measure the richness (bottom left) and Shannon diversity (bottom middle) for the VH, VK / VL, and clones observed within plasma cells in the ileum of SPF or GF mice. Bars represent averages over individual mice (markers). SPF is the leftmost bar in each pair while GF is the rightmost. These measurements can determine the relative abundance of public B cell clonotypes (bottom right).

[0056] FIG.16 illustrates the ability of BCR-MERFISH to define the distribution of plasma cells across long stretches of the mammalian gut. The spatial distribution of B cell clones are determined by BCR-MERFISH (clones are shaded differently) in the proximal (top, left), medial (top, middle), or distal (top, right) portions of the mouse ileum. BCR-MERFISH can be used to identify non-uniform distributions of specific clones within each of these regions as highlighted for five different B cell clones determined via BCR in this same slice (middle row). In addition, BCR-MERFISH can be used to identify spatial co-occurrence between specific B cell clones as highlighted by the spatial distribution of five different clones in the same ileum slice (bottom left) or in a zoom in of a portion of this slice (bottom right). DETAILED DESCRIPTION

[0057] The various aspects as described herein relate to methods to discriminate between closely related nucleotide sequences using in situ hybridization. Applying methods known in the prior art, it is not possible to distinguish via in situ hybridization between different, but closely related nucleic acid sequences, exemplified herein but not limited to different heavy or light chain variable region sequences encoded by different B cell or T cell clones.

[0058] Described herein is a method of preparing a nucleic acid probe that discriminates between its nucleic acid target molecule and a closely related non-target nucleic acid molecule in a sample in a fluorescence in situ hybridization (FISH) assay, the method comprising: a) identifying the set of nucleic acids in the sample that are within a chosen threshold for similarity to the target nucleic acid molecule sequence; b) defining a hierarchy for the set of nucleic acids identified in (a) that classifies nucleic acids in the set on the basis of similarity to each other, such that all nucleic acids in the set are grouped into nodes by similarity; c) identifying a sequence for an in situ hybridization probe that can discriminate in a FISH assay between nucleic acids in different nodes identified in (b), wherein a probe that can discriminate between nucleic acids in different nodes will hybridize, under FISH hybridization conditions, to a target sequence in one node, but not to any non-target sequence 11 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT in any other node; and d) preparing an in situ hybridization (ISH) probe comprising the probe sequence identified in (c).

[0059] The following describes various considerations for making and using nucleic acid probes to distinguish closely related nucleic acids, e.g., RNAs, in a sample, e.g., a cell or tissue sample. As noted above, the methods described herein represent an improvement upon the MERFISH technique for detecting target nucleic acids in multiplex, in situ. MERFISH often hybridizes primary nucleic acid probes to fixed cell or tissue samples, where the primary probes comprise a sequence element complementary to a target nucleic acid, and a barcode sequence comprising a plurality of sequence elements that vary for each target sequence. The hybridized primary probes are detected using iterative hybridization and imaging of labelled secondary nucleic acid probes that selectively bind to individual read elements of the barcodes on the primary nucleic acid probes. The selection or design of barcodes, as well as the techniques for the hybridization of primary nucleic acid probes and subsequent readout of barcoded probe sequences using iterative hybridization of secondary nucleic acid probes is similar to that often used in standard MERFISH. Those techniques for the design of barcodes and for hybridizing primary probes to fixed cell or tissue samples, and reading barcodes encoded on the primary probes using iterative rounds of: hybridization with fluorescently labelled secondary probes that bind to the various barcode elements; detection of bound fluorescent label, quenching or removal of fluorescent label; and hybridization with another fluorescently labelled secondary probe are described, for example, in U.S. application Serial No. 17 / 374,000, published as US 2022 / 0025442, which is incorporated herein by reference.

[0060] In somewhat more detail, MERFISH is a method for single-cell transcriptome imaging and has been demonstrated to profile hundreds to thousands of nucleic acids, such as but not limited to RNAs, in single cells. In MERFISH, nucleic acids are identified via a combinatorial labelling approach that encodes, for example RNA species with error-robust binary barcodes followed by sequential rounds of single-molecule fluorescence in situ hybridization (smFISH) to read out these barcodes.

[0061] The nucleic acids can then be detected, for example, using exposures to nucleic acid probes, e.g., repeatedly, which can be used to determine the distribution of nucleic acids within the material. For example, in MERFISH, a sample is exposed to different rounds of nucleic acid probes, and binding of the nucleic acids can be determined using fluorescence or other techniques. In certain cases, a relatively large number of different targets can be identified using a relatively small number of labels, e.g., by using various combinatorial 12 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT approaches. Identification can also be enhanced in some embodiments using error-checking and / or error-correcting codes. See, e.g., U.S. patent application Ser. No.15 / 329,683, entitled “Systems and Methods for Determining Nucleic Acids,” by Zhuang, et al., published as U.S. Patent Application Publication No.2017 / 0220733 on Aug.3, 2017 and U.S. patent application Ser. No.16 / 348,071, entitled “Multiplexed imaging using merfish, expansion microscopy, and related technologies” by Zhuang et al., published as U.S. Patent Application Publication No. 2019 / 0276881 on September 12, 2019, both incorporated herein by reference in their entireties.

[0062] In one embodiment, a series of nucleic acid probes are used to determine nucleic acids within a cell or other sample, e.g., qualitatively or quantitatively. For example, nucleic acids can be identified as being present or absent, and / or the numbers or concentrations of certain nucleic acids can be determined within the cell or other sample. In some cases, the positions of the probes within the cell or other sample can be determined at relatively high resolutions, and in some cases, at resolutions better than the wavelength of visible light.

[0063] Such embodiment is generally directed to spatially detecting nucleic acids within a cell or other sample, e.g., at relatively high resolutions. For example, the nucleic acids can be RNAs, DNAs, or other nucleic acids. In one set of embodiments, the nucleic acids within the cell can be determined by delivering or applying nucleic acid probes to the cell. In some cases, by using combinatorial approaches, a relatively large number of nucleic acids can be determined using a relatively small number of different labels on the nucleic acid probes. Thus, for example, a relatively small number of experiments can be used to determine a relatively large number of nucleic acids in a sample, e.g., due to simultaneous binding of the nucleic acid probes to different nucleic acids in the sample.

[0064] In one set of embodiments, a population of primary nucleic acid probes are applied to the cell (or other sample) that is able to bind nucleic acids suspected of being present within the cell. Afterwards, sequentially, secondary nucleic acid probes that can bind to or otherwise interact with some of the primary nucleic acid probes are added and determined, e.g., using imaging techniques such as fluorescence microscopy (e.g., conventional fluorescence microscopy), STORM (stochastic optical reconstruction microscopy) or other imaging techniques. After imaging, the secondary nucleic acid probes are inactivated or removed, and different secondary nucleic acid probes are added to the sample. This can be repeated multiple times with multiple different secondary nucleic acid probes. The pattern of binding of the various secondary nucleic acid probes can be used to determine the primary nucleic acid probes at locations within the cell or other sample, which can be used to determine 13 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT mRNA or other nucleic acids that are present.

[0065] The primary nucleic acid probes contain, for instance, a target sequence that can recognize a nucleic acid (e.g., a sequence within nucleic acid). Probes can contain the same or different targeting sequences, which can bind to or hybridize with the same or different nucleic acids. As an example, probe X contains a first targeting sequence X’ which targets the probe to nucleic acid X, while probe Y contains a second targeting sequence Y’, not identical to the first targeting sequence X’ and which targets the probe to nucleic acid Y. The target sequence is substantially complementary to at least a portion of a target nucleic acid, and enough of the target sequence is present such that specific binding of the nucleic acid probe to the target nucleic acid can occur.

[0066] Primary nucleic acid probes also contain a set of “read” sequences that make up the elements of a barcode. The read sequences can all independently be the same or different. In addition, in one set of embodiments, different nucleic acid probes can use one or more common read sequences. For example, more than one read sequence can be combinatorically present on different nucleic acid probes, thereby producing a relatively large number of different nucleic acid probes that can be separately identified, even though only a relatively small number of read sequences are used.

[0067] After a set of primary nucleic acid probes have been introduced to the sample and allowed to interact with nucleic acids that are present, one or more secondary nucleic acid probes are applied to the sample to determine the primary nucleic acid probes. The secondary nucleic acid probes each contain a recognition sequence able to recognize one of the read sequences present within the population of primary nucleic acid probes. For instance, the recognition sequence can be substantially complementary to at least a portion of the read sequence, such that the secondary nucleic acid probe is able to bind to or hybridize with the corresponding portion of that read sequence on the primary nucleic acid probe. In addition, the secondary nucleic acid probes contain one or more signaling entities. For example, a signaling entity can be a fluorescent entity attached to the probe, or a certain sequence of nucleic acids that can be determined in some fashion.

[0068] The location of the secondary nucleic acid probes can be determined by determining signal from the signaling entity. For example, if the signaling entity is fluorescent, then fluorescence microscopy can be used to determine the signaling entity. In some embodiments, imaging of a sample to determine the signaling entity can be used at relatively high resolutions, and in some cases, super - resolution imaging techniques (e.g., resolutions better than the wavelength of visible light or the diffraction limit of light) can be used. 14 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT Examples of super-resolution imaging techniques include STORM, or other techniques as known in the art. In some cases, e.g., with certain super - resolution imaging techniques such as STORM, more than one image of the sample can be acquired.

[0069] More than one type of secondary nucleic acid probe can be applied to a cell or other sample. For example, a first secondary nucleic acid probe can be applied that can recognize a first read sequence, then it or its attached signaling entity can be inactivated or removed, and a second secondary nucleic acid probe can be applied that can recognize a second read sequence. This process can be repeated multiple times, each with a different secondary nucleic acid probe, e.g., to determine the read sequences that were present in the various primary nucleic acid probes. Thus, primary nucleic acids within the sample can be determined on the basis of the binding pattern of secondary nucleic acid probes.

[0070] For example, a first location within the cell or other sample can exhibit binding of a first secondary probe and a third secondary probe, but not the binding of a second or a fourth secondary probe, while a second location can exhibit a different pattern of binding of various secondary probes. The primary nucleic acid probe that the secondary probes are able to bind to or hybridize with can be determined by considering the pattern of binding of various secondary probes. For instance, if a first secondary probe is able to determine read sequence A, a second secondary probe is able to determine read sequence B, and a third secondary probe is able to determine read sequence C, then primary nucleic acid X can be determined through the binding of the first and third secondary probes, A and C, (but not the second secondary probe B), while primary nucleic acid Y can be determined through the binding of the first and second secondary probes A and B (but not the third secondary probe C). Similarly, if it is known that first probe X contains target sequence X’ while second probe Y contains target sequence Y’, then nucleic acids X and Y can also be determined within the sample, e.g., spatially, based on the binding pattern of the various secondary nucleic acid probes. In addition, it should be noted that due to the presence of more than one read sequence on the primary nucleic acid probes, even though first probe X and second probe Y contain a common read sequence, these probes can be distinguished in the sample due to the different binding patterns of the various secondary nucleic acid probes.

[0071] A key difference for the technology described herein relative to standard MERFISH lies in the design of the primary probes. Methods are described herein below to permit the selection of target sequences in the target nucleic acids that permit sensitive distinction of closely-related nucleic acids using in situ hybridization.

[0072] As described herein, a nucleic acid probe can be at least 5 nucleotides (nt), at least 15 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT 10nt, at least 15nt, at least 20nt, at least 25nt, at least 30nt, at least 35nt, at least 40nt, or more.

[0073] As described herein, the % sequence identity can be at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100%.

[0074] As described herein, the predetermined threshold number of specific probes in (C) is at least 1, the predetermined threshold number of specific probes in (C) is at least 2, the predetermined threshold number of specific probes in (C) is at least 3, the predetermined threshold number of specific probes in (C) is at least 4, the predetermined threshold number of specific probes in (C) is at least 5, the predetermined threshold number of specific probes in (C) is at least 6, the predetermined threshold number of specific probes in (C) is at least 7, the predetermined threshold number of specific probes in (C) is at least 8, the predetermined threshold number of specific probes in (C) is at least 9, the predetermined threshold number of specific probes in (C) is at least 10, the predetermined threshold number of specific probes in (C) is at least 11, the predetermined threshold number of specific probes in (C) is at least 12, the predetermined threshold number of specific probes in (C) is at least 13, the predetermined threshold number of specific probes in (C) is at least 14, the predetermined threshold number of specific probes in (C) is at least 15, the predetermined threshold number of specific probes in (C) is at least 16, the predetermined threshold number of specific probes in (C) is at least 17, the predetermined threshold number of specific probes in (C) is at least 18, the predetermined threshold number of specific probes in (C) is at least 19, or the predetermined threshold number of specific probes in (C) is at least 20 or more. V(D)J Recombination

[0075] In some embodiments, the methods described herein permit the detection of different B or T cell clones using probes that can detect RNAs encoding B cell receptors or T cell 16 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT receptors that essentially vary only in the sequences of their antigen-binding variable regions, which are characterized by rearranged sequences that define the complementarity determining regions (CDRs) for a given antibody or T cell receptor. The diversity of the V regions is generated by recombination between so-called variable (V), diversity (D) and joining (J) gene segments in the bone marrow (for B cells) and thymus (for T cells).

[0076] B cell receptors and antibodies are composed of heavy and light chains which have constant and variable regions genetically located on 3 loci in humans.

[0077] Each heavy or light chain gene contains multiple copies of three different types of gene segments-variable, diversity, and joining gene segments. DNA rearrangement causes one copy of each type of gene segment to go in any given lymphocyte, generating an enormous antibody repertoire.

[0078] Most T-cell receptors are composed of a variable alpha chain and a beta chain. The T cell receptor genes are similar to immunoglobulin genes in that they too contain multiple V, D, and J gene segments in their beta chains (and V and J gene segments in their alpha chains) that are rearranged during the development of the lymphocyte to provide that cell with a unique antigen receptor. The T cell receptor in this sense is the topological equivalent to an antigen-binding fragment of the antibody, both being part of the immunoglobulin superfamily.

[0079] In the developing B cell, the first recombination event to occur is between one D and one J gene segment of the heavy chain locus. Any DNA between these two gene segments is deleted. This D-J recombination is followed by the joining of one V gene segment, from a region upstream of the newly formed DJ complex, forming a rearranged VDJ gene segment. All other gene segments between V and D segments are now deleted from the cell's genome. Primary transcript (unspliced RNA) is generated containing the VDJ region of the heavy chain and both the constant mu and delta chains (Cμ and Cδ). (i.e. the primary transcript contains the segments: V-D-J-Cμ-Cδ). The primary RNA is processed to add a polyadenylated (poly-A) tail after the Cμ chain and to remove sequence between the VDJ segment and this constant gene segment. Translation of this mRNA leads to the production of the IgM heavy chain protein.

[0080] The kappa (κ) and lambda (λ) chains of the immunoglobulin light chain loci rearrange in a very similar way, except that the light chains lack a D segment. In other words, the first step of recombination for the light chains involves the joining of the V and J chains to give a VJ complex before the addition of the constant chain gene during primary transcription. Translation of the spliced mRNA for either the kappa or lambda chains results in formation 17 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT of the Ig κ or Ig λ light chain protein. Assembly of the Ig μ heavy chain and one of the light chains results in the formation of membrane bound form of the immunoglobulin IgM that is expressed on the surface of the immature B cell.

[0081] During thymocyte development, the T cell receptor (TCR) chains undergo essentially the same sequence of ordered recombination events as that described for immunoglobulins. D-to-J recombination occurs first in the β-chain of the TCR. This process can involve either the joining of the Dβ1 gene segment to one of six Jβ1 segments or the joining of the Dβ2 gene segment to one of six Jβ2 segments. DJ recombination is followed with Vβ-to-DβJβ rearrangements. All gene segments between the Vβ-Dβ-Jβ gene segments in the newly formed complex are deleted and the primary transcript is synthesized that incorporates the constant domain gene (Vβ-Dβ-Jβ-Cβ). mRNA transcription splices out any intervening sequence and allows translation of the full length protein for the TCR β-chain. The rearrangement of the alpha (α) chain of the TCR follows β chain rearrangement and resembles V-to-J rearrangement described for Ig light chains. The assembly of the β- and α- chains results in formation of the αβ-TCR that is expressed on a majority of T cells.

[0082] As discussed, the heavy chain genes can be highly similar to each other and the light chain genes can be highly similar to each other, only differing substantially for different B and T cell clones in the rearranged gene segments that make up the complementarity determining regions. Thus, discriminating between their mRNAs using standard in-situ hybridization approaches is difficult. The problem becomes more acute when trying to simultaneously distinguish populations comprising a number of different B and T cell clones or to map clone lineages. Methods of designing in situ hybridization probes that efficiently discriminate such closely related sequences and their use in detecting cells of different lineages in various populations are provided herein. Classification of Nucleic Acids

[0083] A hierarchy for a set of nucleic acids can be defined in multiple ways. One common implementation of such a hierarchy is a phylogenetic tree, where individual nucleic acids are grouped into different ‘branches’ or ‘nodes’ based on their respective similarity and then individual branches or nodes themselves are grouped based on a measure of the collective similarity of the sequences within a branch or node to the sequences within a different branch or node. A phylogenetic tree is a branching diagram or a tree showing, for example, the relationships among various nucleic acids of interest or other entities based upon similarities and differences in their physical or genetic characteristics. A branch is defined as where the 18 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT target nucleic acids of interest are shown at the tips of the tree's branches. The branches themselves connect up in a way that represents, for example, the evolutionary history or other sequence-based relationships of the target nucleic acids of interest. As used herein, a node of a phylogenetic tree is where one branch splits into two. Each node represents the last common ancestor of the two lineages descended from that node. The methods described herein permit distinguishing between RNA sequences in different nodes of a phylogenetic tree using hybridization probes designed using the approaches described in the Examples herein below. Although the Examples below distinguish between RNA sequences, one who is skilled in the art can use the methodology as described herein to distinguish between two highly similar nucleic acids, such as, but not limited to DNA sequences.

[0084] In addition, one who is skilled in the art will recognize that alternative ways exist for constructing trees in which an individual node can represent a place where a branch splits into two or more different sub-elements of the tree. It should be understood that all Examples herein below can be generalized to include trees that have such generalized structures.

[0085] After contacting the primary nucleic acid probes with a cell or other sample, the nucleic acid probes can be directly determined by using one or more secondary nucleic acid probes, in accordance with certain aspects of the invention. As mentioned, in some cases, the determination can be spatial, e.g., in two or three dimensions. In addition, in some cases, the determination can be quantitative, e.g., the amount or concentration of a primary nucleic acid probe (and therefore of a corresponding target nucleic acid) can be determined. Additionally, the secondary probes can comprise any of a variety of entities able to hybridize with a nucleic acid, e.g., DNA, RNA, LNA, and / or PNA, etc., depending on the application. As noted above, further detail regarding the design and use of secondary probes is provided in U.S. application Serial No. 17 / 374,000, published as US 2022 / 0025442, which is incorporated herein by reference. Signaling entities are discussed in more detail below. Signaling Entities / Fluorophores

[0086] In some embodiments, the spatial positions of the entities (and thus, nucleic acid probes that the entities may be associated with) can be determined at relatively high resolution. For instance, the positions can be determined at spatial resolutions of better than about 100 micrometers, better than about 30 micrometers, better than about 10 micrometers, better than about 3 micrometers, better than about 1 micrometer, better than about 800 nm, better than about 600 nm, better than about 500 nm, better than about 400 nm, better than about 300 nm, better than about 200 nm, better than about 100 nm, better than about 90 nm, 19 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT better than about 80 nm, better than about 70 nm, better than about 60 nm, better than about 50 nm, better than about 40 nm, better than about 30 nm, better than about 20 nm, or better than about 10 nm, etc.

[0087] There are a variety of techniques able to determine or image the spatial positions of entities optically, e.g., using fluorescence microscopy. In some cases, the spatial positions can be determined at super resolutions, or at resolutions better than the wavelength of light or the diffraction limit. Non - limiting examples include STORM (stochastic optical reconstruction microscopy), STED (stimulated emission depletion microscopy), NSOM (Near - field Scanning Optical Microscopy), 4Pi microscopy, SIM (Structured Illumination Microscopy), SM (Spatially Modulated Illumination) microscopy, RESOLFT (Reversible Saturable Optically Linear Fluorescence Transition Microscopy), GSD (Ground State Depletion Microscopy), SSIM (Saturated Structured Illumination Microscopy), SPDM (Spectral Precision Distance Microscopy), Photo - Activated Localization Microscopy (PALM), Fluorescence Photoactivation Localization Microscopy (FPALM), LIMON (3D Light Microscopical Nanosizing Microscopy), Super - resolution optical fluctuation imaging (SOFI), or the like. See, e.g., U.S. Pat. No.7,838,302, issued Nov.23, 2010, entitled “Sub - Diffraction Limit Image Resolution and Other Imaging Techniques, " by Zhuang, et al.; U.S. Pat. No.8,564,792, issued Oct.22, 2013, entitled “Sub - diffraction Limit Image Resolution in Three Dimensions,” by Zhuang, et al.; or Int. Pat. Apl. Pub. No. WO 2013 / 090360, published Jun.20, 2013, entitled “High Resolution Dual - Objective Microscopy,” by Zhuang, et al., each incorporated herein by reference in their entireties.

[0088] In addition, the signaling entity can be inactivated in some cases after detection or imaging. For example, in some embodiments, a first secondary nucleic acid probe containing a signaling entity can be applied to a sample that can recognize a first read sequence, then the signaling entity on the first secondary nucleic acid probe can be inactivated before a second secondary nucleic acid probe is applied to the sample. If multiple signaling entities are used, the same or different techniques can be used to inactivate the signaling entities, and some or all of the multiple signaling entities can be inactivated, e.g., sequentially or simultaneously.

[0089] Inactivation can be caused by removal of the signaling entity (e.g., from the sample, or from the nucleic acid probe, etc.), and / or by chemically altering the signaling entity in some fashion, e.g., by photobleaching the signaling entity, bleaching or chemically altering the structure of the signaling entity, e.g., by reduction, etc.). For instance, in one set of embodiments, a fluorescent signaling entity can be inactivated by chemical or optical techniques such as oxidation, photobleaching, chemically bleaching, stringent washing or 20 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT enzymatic digestion or reaction by exposure to an enzyme, dissociating the signaling entity from other components (e.g., a probe), chemical reaction of the signaling entity (e.g., to a reactant able to alter the structure of the signaling entity) or the like. For instance, bleaching can occur by exposure to oxygen, reducing agents, or the signaling entity could be chemically cleaved from the nucleic acid probe and washed away via fluid flow.

[0090] In some embodiments, various nucleic acid probes (including primary and / or secondary nucleic acid probes) can include one or more signaling entities. If more than one nucleic acid probe is used, the signaling entities can each by the same or different. In certain embodiments, a signaling entity is any entity able to emit light. For instance, in one embodiment, the signaling entity is fluorescent. In other embodiments, the signaling entity can be phosphorescent, radioactive, absorptive, etc. In some cases, the signaling entity is any entity that can be determined within a sample at relatively high resolutions, e.g., at resolutions better than the wavelength of visible light or the diffraction limit. The signaling entity can be, for example, a dye, a small molecule, a peptide or protein, or the like. The signaling entity can be a single molecule in some cases. If multiple secondary nucleic acid probes are used, the nucleic acid probes can comprise the same or different signaling entities.

[0091] Non - limiting examples of signaling entities include fluorescent entities (fluorophores) or phosphorescent entities, for example, cyanine dyes (e.g., Cy2, Cy3, Cy3B, Cy5, Cy5.5, Cy7, etc.), Alexa Fluor dyes, Atto dyes, photoswitchable dyes, photoactivatable dyes, fluorescent dyes, metal nanoparticles, semiconductor nanoparticles or " quantum dots”, fluorescent proteins such as GFP (Green Fluorescent Protein), or photoactivabale fluorescent proteins, such as PAGFP, PSCFP, PSCFP2, Dendra, Dendra2, EosFP, tdEos, mEos2, mEos3, PAmCherry, PAtagRFP, mMaple, mMaple2, and mMaple3. Other suitable signaling entities are known to those of ordinary skill in the art. See, e.g., U.S. Pat. No.7,838,302 or U.S. Pat. Apl. Ser. No.61 / 979,436, each incorporated herein by reference in its entirety.

[0092] In one set of embodiments, the signaling entity can be attached to an oligonucleotide sequence via a bond that can be cleaved to release the signaling entity. In one set of embodiments, a fluorophore can be conjugated to an oligonucleotide via a cleavable bond, such as a photocleavable bond. Non - limiting examples of photocleavable bonds include, but are not limited to, 1- (2 - nitrophenyl) ethyl, 2 - nitrobenzyl, biotin phosphoramidite, acrylic phosphoramidite, diethylaminocoumarin, 1- (4,5 - dimethoxy - 2 - nitrophenyl) ethyl, cyclo - dodecyl (dimethoxy - 2 - nitrophenyl) ethyl, 4 - aminomethyl - 3 - nitrobenzyl, (4 - nitro - 3- (1 - chlorocarbony loxyethyl) phenyl) methyl - S - acetylthioic acid ester, (4 - nitro 3- (1 - thlorocarbonyloxyethyl) phenyl) methyl - 3- (2 - pyridyldi thiopropionic acid) ester, 3- (4,4 ' - 21 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT dimethoxytrityl) -1- (2 nitrophenyl) -propane - 1,3 - diol- [ 2 - cyanoethyl- (N, N diisopropyl)] - phosphoramidite, 1- [ 2 - nitro - 5-6 trifluoroacetylcaproamidomethyl) phenyl] -ethyl- [ 2 - cyano ethyl- (N, N - diisopropyl) l - phosphoramidite, 1- [ 2 - nitro - 5- (6 (4,4 ' - dimethoxytrityloxy) butyramidomethyl) phenyl] -ethyl [ 2 - cyanoethyl- ( N , N - diisopropyl)] - phosphoramidite, 1- [2 nitro - 5-6- (N- (4,4 ' - dimethoxytrityl)) biotinamidocaproamido - methyl) phenyl) -ethyl- [ 2 cyanoethyl- (N, N - diisopropyl)] - phosphoramidite, or similar linkers. In another set of embodiments, the fluorophore can be conjugated to an oligonucleotide via a disulfide bond. The disulfide bond can be cleaved by a variety of reducing agents such as, but not limited to, dithiothreitol, dithioerythritol, beta - mercaptoethanol, sodium borohydride, thioredoxin, glutaredoxin, trypsinogen, hydrazine, diisobutylaluminum hydride, oxalic acid, formic acid, ascorbic acid, phosphorous acid, tin chloride, glutathione, thioglycolate, 2,3 -dimercaptopropanol, 2-mercaptoethylamine, 2 - aminoethanol, tris (2-carboxyethyl) phosphine, bis (2-mercaptoethyl) sulfone, N, N- dimethyl-N, N-bis (mercaptoacetyl) hydrazine, 3-mercaptoproptionate, dimethylformamide, thiopropyl-agarose, tri-n-butylphosphine, cysteine, iron sulfate, sodium sulfite, phosphite, hypophosphite, phosphorothioate, or the like, and / or combinations of any of these. In another embodiment, the fluorophore can be conjugated to an oligonucleotide via one or more phosphorothioate modified nucleotides in which the sulfur modification replaces the bridging and / or non - bridging oxygen. The fluorophore can be cleaved from the oligonucleotide, in certain embodiments, via addition of compounds such as but not limited to iodoethanol, iodine mixed in ethanol, silver nitrate, or mercury chloride. In yet another set of embodiments, the signaling entity can be chemically inactivated through reduction or oxidation. For example, in one embodiment, a chromophore such as Cy5 or Cy7 can be reduced using sodium borohydride to a stable, non - fluorescence state. In still another set of embodiments, a fluorophore can be conjugated to an oligonucleotide via an azo bond, and the azo bond can be cleaved with 2-[(2-N-arylamino) phenylazo] pyridine. In yet another set of embodiments, a fluorophore can be conjugated to an oligonucleotide via a suitable nucleic acid segment that can be cleaved upon suitable exposure to DNAse, e.g., an exodeoxyribonuclease or an endodeoxyribonuclease. Examples include, but are not limited to, deoxyribonuclease I or deoxyribonuclease II. In one set of embodiments, the cleavage can occur via a restriction endonuclease. Non - limiting examples of potentially suitable restriction endonucleases include BamHI, Bsrl, NotI, Xmal, PspAI, Dpnl, Mbol, MnlI, Eco571, Ksp6321, DralII, Ahall, Smal, Mlul, Hpal, Apal, BclI, BstEII, TaqI, EcoRI, SacI, HindII, Haell, Drall, Tsp5091, Sau3AI, Pacl, etc. Over 3000 restriction enzymes have been 22 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT studied in detail, and more than 600 of these are available commercially. In yet another set of embodiments, a fluorophore can be conjugated to biotin, and the oligonucleotide conjugated to avidin or streptavidin. An interaction between biotin and avidin or streptavidin allows the fluorophore to be conjugated to the oligonucleotide, while sufficient exposure to an excess of addition, free biotin could “outcompete” the linkage and thereby cause cleavage to occur. In addition, in another set of embodiments, the probes can be removed using corresponding “toe - hold - probes,” which comprise the same sequence as the probe, as well as an extra number of bases of homology to the encoding probes (e.g., 1-20 extra bases, for example, 5 extra bases). These probes can remove the labeled readout probe through a strand - displacement interaction. Additional detail regarding signaling entities and probes including them, including detail regarding “switchable” signaling entities, among others, is provided in U.S. application No.17 / 374,000, published as US2022 / 0025442 noted above and incorporated herein by reference. Computer-readable medium

[0093] Any of the methods described herein can be performed by a computer program product comprising computer executable logic recorded on a computer readable medium. For example, the computer program can perform some or all of the following functions: (i) align the nucleotide sequences of each pair of nucleic acids, (ii) compute the fraction of nucleotides that match between nucleic acids, (iii) divide each nucleic acid sequence into a series of k- mers of a given length, also known as a k value, (iv) compute the pairwise similarity between all nucleic acids or existing nodes, (v) average the pairwise similarity between each nucleic acid within a node and the selected nucleic acid, (vi) compute the similarity between two nodes by averaging the pairwise similarity between all nucleic acids in the first node and all nucleic acids in the second node.

[0094] The functions or algorithms described herein can be implemented in software or a combination of software and human implemented procedures, for example. The software can consist of computer executable instructions stored on computer readable media such as memory or other type of storage devices. Further, such functions correspond to modules, which are software, hardware, firmware or any combination thereof. Multiple functions can be performed in one or more modules as desired, and the embodiments described are merely examples. The software can be executed on a digital signal processor, ASIC, microprocessor, or other type of processor operating on a computer system, such as a personal computer, 23 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT server or other computer system. In one embodiment, multiple such computer systems are utilized in a distributed network to implement multiple analyses, draw upon information from distributed sources, or facilitate transaction based usage. An object-oriented, service-oriented, or other architecture can be used to implement such functions and communicate between the multiple systems and components.

[0095] For example, the computer can operate in a networked environment using a communication connection to connect to one or more remote computers, such as database servers. The remote computer can include a personal computer (PC), server, router, network PC, a peer device or other common network node, or the like. The communication connection can include a Local Area Network (LAN), a Wide Area Network (WAN) or other networks.

[0096] Datasets of information can be in different forms and from different sources. For example, datasets can be stored and updated in the form of computer-accessible storage. Computer-accessible storage includes random access memory (RAM), read only memory (ROM), erasable programmable read-only memory (EPROM) & electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD ROM), Digital Versatile Disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium capable of storing computer-readable instructions.

[0097] Computer-readable instructions (e.g., for computing a phylogenetic tree for RNA sequences) can be stored on a computer-readable medium and can be executable by a processing unit of the computer. A hard drive, CD-ROM, and RAM are some examples of articles including a non-transitory computer-readable medium. For example, a computer program linked to, or including, the phylogenetic tree programs can be capable of providing a generic technique to perform an access control check for data access and / or for doing an operation on one of the servers in a component object model (COM) based system, or can be included on a CD-ROM and loaded from the CD-ROM to a hard drive. The computer- readable instructions allow computer to provide generic access controls in a COM based computer network system having multiple users and servers.

[0098] Such a hard disk drive, magnetic disk drive, and optical disk drive can couple with a hard disk drive interface, a magnetic disk drive interface, and an optical disk drive interface, respectively. The drives and their associated computer-readable media provide non-volatile storage of computer-readable instructions, data structures, program modules and other data for the computer. It should be appreciated by those skilled in the art that any type of computer-readable media which can store data that is accessible by a computer, such as 24 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT magnetic cassettes, flash memory cards, digital video disks, Bernoulli cartridges, random access memories (RAMs), read only memories (ROMs), redundant arrays of independent disks (e.g., RAID storage devices) and the like, can be used in the exemplary operating environment.

[0099] A plurality of program modules can be stored on the hard disk, magnetic disk, optical disk, ROM, or RAM, including an operating system, one or more application programs, other program modules, and program data. Programming for implementing one or more processes or method described herein can be resident on any one or number of these computer-readable media.

[0100] A user can enter commands and information into computer through input devices such as a keyboard and pointing device. Other input devices (not shown) can include a microphone, touch screen, joystick, game pad, satellite dish, scanner, or the like. These other input devices are often connected to the processing unit through a serial port interface that is coupled to the system bus, but can be connected by other interfaces, such as a parallel port, game port, or a universal serial bus (USB). A monitor or other type of display device can also be connected to the system bus via an interface, such as a video adapter. The monitor can display a graphical user interface for the user, and can include a touchscreen, allowing user interactions to select functions and enter data. In addition to a monitor, computers typically include other peripheral output devices, such as speakers and printers.

[0101] The computer can operate in a networked environment using logical connections to one or more remote computers or servers, such as remote computer. These logical connections are achieved by a communication device coupled to or a part of the computer; the invention is not limited to a particular type of communications device. Such a remote computer can be another computer, a server, a router, a network PC, a client, a peer device or other common network node, and typically includes many or all of the elements described above relative to the computer. The logical connections include a local area network (LAN) and / or a wide area network (WAN). Such networking environments are commonplace in office networks, enterprise-wide computer networks, intranets and the internet, which are all types of networks.

[0102] When used in a LAN-networking environment, the computer can be connected to the LAN through a network interface or adapter, which is one type of communications device. In some embodiments, when used in a WAN-networking environment, the computer typically includes a modem (another type of communications device) or any other type of communications device, e.g., a wireless transceiver, for establishing communications over the 25 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT wide-area network, such as the internet. Such a modem, which can be internal or external, is connected to the system bus via the serial port interface. In a networked environment, program modules depicted relative to the computer can be stored in the remote memory storage device of remote computer, or server. It is appreciated that the network connections described are exemplary and other means of, and communications devices for, establishing a communications link between the computers can be used including hybrid fiber-coax connections, T1-T3 lines, DSL's, OC-3 and / or OC-12, TCP / IP, microwave, wireless application protocol, and any other electronic media through any suitable switches, routers, outlets and power lines, as the same are known and understood by one of ordinary skill in the art. Electronic Apparatus and System

[0103] Example embodiments can therefore be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Example embodiments can be implemented using a computer program product, for example, a computer program tangibly embodied in an information carrier, for example, in a machine- readable medium for execution by, or to control the operation of, data processing apparatus, for example, a programmable processor, a computer, or multiple computers.

[0104] A computer program can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand- alone program or as a module, subroutine, or other unit suitable for use in a computing environment. A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.

[0105] In example embodiments, operations can be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Method operations can also be performed by, and apparatus of example embodiments can be implemented as, special purpose logic circuitry (e.g., a FPGA or an ASIC).

[0106] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In embodiments deploying a programmable computing system, it will be appreciated that both hardware and software architectures merit consideration. Specifically, it will be appreciated that the choice 26 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT of whether to implement certain functionality in permanently configured hardware (e.g., an ASIC), in temporarily configured hardware (e.g., a combination of software and a programmable processor), or a combination of permanently and temporarily configured hardware can be a design choice. Definitions

[0107] The terms “decrease”, “reduced”, “reduction”, or “inhibit” are all used herein to mean a decrease by a statistically significant amount. In some embodiments of any of the aspects, “reduce,” “reduction" or “decrease" or “inhibit” typically means a decrease by at least 10% as compared to a reference level (e.g. the absence of a given treatment or agent) and can include, for example, a decrease by at least about 10%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99% , or more. As used herein, “reduction” or “inhibition” does not encompass a complete inhibition or reduction as compared to a reference level. “Complete inhibition” is a 100% inhibition as compared to a reference level. A decrease can be preferably down to a level accepted as within the range of normal for an individual without a given disorder.

[0108] The terms “increased”, “increase”, “enhance”, or “activate” are all used herein to mean an increase by a statically significant amount. In some embodiments of any of the aspects, the terms “increased”, “increase”, “enhance”, or “activate” can mean an increase of at least 10% as compared to a reference level, for example an increase of at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90% or up to and including a 100% increase or any increase between 10-100% as compared to a reference level, or at least about a 2-fold, or at least about a 3-fold, or at least about a 4-fold, or at least about a 5-fold or at least about a 10-fold increase, or any increase between 2-fold and 10-fold or greater as compared to a reference level. In the context of a marker or symptom, a “increase” is a statistically significant increase in such level.

[0109] As used herein, “contacting" refers to any suitable means for delivering, or exposing, an agent to at least one cell, including, but not limited to a fixed cell. Exemplary delivery methods include, but are not limited to, direct delivery to a sample comprising fixed cells, 27 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT e.g., by overlaying a fixed sample with an agent in solution or suspension, delivery to cell culture medium, perfusion, injection, or other delivery method well known to one skilled in the art. In some embodiments of any of the aspects, contacting comprises physical human activity, e.g., an injection; an act of dispensing, mixing, and / or decanting; and / or manipulation of a delivery device or machine.

[0110] The term “statistically significant" or “significantly" refers to statistical significance and generally means a two-standard deviation (2SD) or greater difference.

[0111] Other than in the operating examples, or where otherwise indicated, all numbers expressing quantities of ingredients or reaction conditions used herein should be understood as modified in all instances by the term “about.” The term “about” when used in connection with percentages can mean ±1%.

[0112] As used herein, the term “comprising” means that other elements can also be present in addition to the defined elements presented. The use of “comprising” indicates inclusion rather than limitation.

[0113] The term "consisting of" refers to compositions, methods, and respective components thereof as described herein, which are exclusive of any element not recited in that description of the embodiment.

[0114] As used herein the term "consisting essentially of" refers to those elements required for a given embodiment. The term permits the presence of additional elements that do not materially affect the basic and novel or functional characteristic(s) of that embodiment of the invention.

[0115] As used herein, the term “specific binding” refers to a physical interaction between two molecules, compounds, cells and / or particles wherein the first entity binds to the second, target entity to the substantial exclusion of one or more non-target third entities which may or may not be in excess relative to the second, target entity. Specific binding generally involves binding with greater affinity for target vs non-target binding. In some embodiments of any of the aspects, specific binding can refer to an affinity for the target entity which is at least 100 times, at least 500 times, at least 1000 times, or at least 10,000 times greater than the affinity for nontarget entities under the conditions of the assay being utilized.

[0116] The singular terms "a," "an," and "the" include plural referents unless context clearly indicates otherwise. Similarly, the word "or" is intended to include "and" unless the context clearly indicates otherwise. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of this disclosure, suitable methods and materials are described below. The abbreviation, "e.g." is derived from the Latin exempli 28 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT gratia, and is used herein to indicate a non-limiting example. Thus, the abbreviation "e.g." is synonymous with the term "for example."

[0117] Groupings of alternative elements or embodiments of the invention disclosed herein are not to be construed as limitations. Each group member can be referred to and claimed individually or in any combination with other members of the group or other elements found herein. One or more members of a group can be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is herein deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.

[0118] Unless otherwise defined herein, scientific and technical terms used in connection with the present application shall have the meanings that are commonly understood by those of ordinary skill in the art to which this disclosure belongs. It should be understood that this invention is not limited to the particular methodology, protocols, and reagents, etc., described herein and as such can vary. The terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which is defined solely by the claims. Definitions of common terms in immunology and molecular biology can be found in The Merck Manual of Diagnosis and Therapy, 20th Edition, published by Merck Sharp & Dohme Corp., 2018 (ISBN 0911910190, 978- 0911910421); Robert S. Porter et al. (eds.), The Encyclopedia of Molecular Cell Biology and Molecular Medicine, published by Blackwell Science Ltd., 1999-2012 (ISBN 9783527600908); and Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 1-56081- 569-8); Immunology by Werner Luttmann, published by Elsevier, 2006; Janeway's Immunobiology, Kenneth Murphy, Allan Mowat, Casey Weaver (eds.), W. W. Norton & Company, 2016 (ISBN 0815345054, 978-0815345053); Lewin's Genes XI, published by Jones & Bartlett Publishers, 2014 (ISBN-1449659055); Michael Richard Green and Joseph Sambrook, Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., USA (2012) (ISBN 1936113414); Davis et al., Basic Methods in Molecular Biology, Elsevier Science Publishing, Inc., New York, USA (2012) (ISBN 044460149X); Laboratory Methods in Enzymology: DNA, Jon Lorsch (ed.) Elsevier, 2013 (ISBN 0124199542); Current Protocols in Molecular Biology (CPMB), Frederick M. Ausubel (ed.), John Wiley and Sons, 2014 (ISBN 047150338X, 9780471503385), Current Protocols in Protein Science (CPPS), John E. Coligan (ed.), John Wiley and Sons, Inc., 2005; and Current Protocols in Immunology (CPI) (John E. Coligan, 29 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT ADA M Kruisbeek, David H Margulies, Ethan M Shevach, Warren Strobe, (eds.) John Wiley and Sons, Inc., 2003 (ISBN 0471142735, 9780471142737), the contents of which are all incorporated by reference herein in their entireties.

[0119] Other terms are defined herein within the description of the various aspects of the invention.

[0120] The following U.S. Patent Application Nos.: 17 / 374,000; 17 / 413,148; 16 / 616,833; 16 / 347,874; 16 / 348,071 are incorporated herein by reference in their entireties.

[0121] All patents and other publications; including literature references, issued patents, published patent applications, and co-pending patent applications; cited throughout this application are expressly incorporated herein by reference for the purpose of describing and disclosing, for example, the methodologies described in such publications that might be used in connection with the technology described herein. These publications are provided solely for their disclosure prior to the filing date of the present application. Nothing in this regard should be construed as an admission that the inventors are not entitled to antedate such disclosure by virtue of prior invention or for any other reason. All statements as to the date or representation as to the contents of these documents is based on the information available to the applicants and does not constitute any admission as to the correctness of the dates or contents of these documents.

[0122] The description of embodiments of the disclosure is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. While specific embodiments of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as those skilled in the relevant art will recognize. For example, while method steps or functions are presented in a given order, alternative embodiments can perform functions in a different order, or functions can be performed substantially concurrently. The teachings of the disclosure provided herein can be applied to other procedures or methods as appropriate. The various embodiments described herein can be combined to provide further embodiments. Aspects of the disclosure can be modified, if necessary, to employ the compositions, functions and concepts of the above references and application to provide yet further embodiments of the disclosure. Moreover, due to biological functional equivalency considerations, some changes can be made in protein structure without affecting the biological or chemical action in kind or amount. These and other changes can be made to the disclosure in light of the detailed description. All such modifications are intended to be included within the scope of the appended claims. 30 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT

[0123] Specific elements of any of the foregoing embodiments can be combined or substituted for elements in other embodiments. Furthermore, while advantages associated with certain embodiments of the disclosure have been described in the context of these embodiments, other embodiments can also exhibit such advantages, and not all embodiments need necessarily exhibit such advantages to fall within the scope of the disclosure.

[0124] Some embodiments of the technology described herein can be defined according to any of the following numbered paragraphs: 1. A method of preparing a nucleic acid probe that discriminates between its RNA target molecule and a closely related non-target RNA molecule in a sample in a fluorescence in situ hybridization (FISH) assay, the method comprising: a) identifying the set of RNAs in the sample that are within a chosen threshold for similarity to the target RNA molecule sequence; b) defining a hierarchy for the set of RNAs identified in (a) that classifies RNAs in the set on the basis of similarity to each other, such that all RNAs in the set are grouped into nodes by similarity; c) identifying a sequence for an in situ hybridization probe that can discriminate in a FISH assay between RNAs in different nodes identified in (b), wherein a probe that can discriminate between RNAs in different nodes will hybridize, under FISH hybridization conditions, to a target sequence in one node, but not to any non-target sequence in any other node; and d) preparing an in situ hybridization (ISH) probe comprising the probe sequence identified in (c). 2. The method of paragraph 1, wherein the step (a) of identifying the set of RNAs in the sample that are within a chosen threshold for similarity to the target RNA molecule sequence comprises: i) aligning target RNA sequence pairwise with each other RNA sequence expressed in the sample, wherein RNAs with target sequence identity greater than the threshold are selected for the set; or ii) dividing each RNA sequence expressed in the sample into a series of all sub- sequences of a given length k (k-mers), wherein the degree of similarity between the target RNA and any non-target RNA expressed in the sample is defined by the fraction of k-mers found in either the target RNA sequence or the non-target RNA sequence that are also found in both sequences, and wherein RNAs with a degree of similarity greater than the threshold are selected for the set. 31 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT 3. The method of paragraph 1, wherein step (b) of defining a hierarchy for the set of RNAs identified in (a) that classifies RNAs in the set (a) on the basis of similarity to each other comprises generating a phylogenetic tree wherein individual RNAs are grouped into different branches and nodes on the tree. 4. The method of paragraph 3, wherein the phylogenetic tree is computed by defining pairwise similarity between individual RNA sequences expressed in the sample, and applying an agglomerative algorithm to systematically group RNAs into branches and nodes based on the RNAs, or nodes comprised of RNAs, that are most similar. 5. The method of paragraph 4, wherein the agglomerative algorithm employs an iterative process comprising: i) computing the pairwise similarity between all RNAs or existing nodes; ii) selecting the pair with the greatest similarity, and grouping these together to form a new node, wherein the pair can represent two RNAs, two existing nodes, or an RNA and a node; and iii) repeating steps (i) and (ii) until all RNAs or nodes have been grouped together. 6. The method of paragraph 5, wherein in step (i), a) the similarity between a given RNA and a node is computed by averaging the pairwise similarity between each RNA within a node and the given RNA; b) the similarity between two nodes is computed by averaging the pairwise similarity between all RNAs in the first node and all RNAs in the second node; and c) a node can consist of multiple other nodes, wherein an RNA is considered to be part of a first node if it is part of a second or third node comprised by the first node, and if the second or third nodes themselves contain nodes, then any RNAs within the second or third nodes or their sub-nodes is considered to be part of the first node. 7. The method of paragraph 4 or paragraph 5, wherein the first iteration of the iterative process uses only RNAs, creating one node by grouping the most similar RNAs, and each subsequent iteration creates another node by grouping two RNAs or the node created in the previous round to another RNA, until all nodes and RNA have been grouped. 8. The method of any one of paragraphs 1-7, wherein step (c) of identifying a sequence for an in situ hybridization probe that can discriminate in a FISH assay between target RNAs in different nodes identified in (b) comprises: a) for each RNA target, consider all possible probe sequences that are of a given length and satisfy a given set of free energy of binding parameters for a probe to its target sequence; 32 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT b) determine the specificity of each possible probe sequence of (a) for the RNA, wherein specificity is determined by considering all possible fragments of length k2for the possible probe sequence and computing the fraction of k-mers within that possible probe sequence that are found in any other RNA sequence expressed in the sample, wherein the probe is considered specific if this fraction is below a given threshold; c) sum the number of unique specific probes determined in (b), wherein: when the number of specific probes determined in (b) is greater than a predetermined threshold, the target RNA is designated as distinguishable from all others expressed in the sample via one or more of such probes; and when the number of specific probes is below the predetermined threshold, the target RNA is considered not distinguishable from others expressed in the sample via one or more of such probes, d) grouping all RNAs designated not distinguishable with the RNA or nodes to which they are assigned based on the hierarchy in paragraph 1(b), and repeating steps (a) – (c), but considering the RNAs grouped into a node as a single object, and e) considering each possible probe within a given RNA, computing its specificity by calculating the fraction of k-mers within that probe found in all other RNA except the RNAs found within the node to which that RNA has been assigned, wherein a probe is considered specific if less than a given fraction of the k-mers within it are found in RNAs outside of that node, and if all RNAs within a node have a number of specific probes greater than a predetermined threshold, the node is considered to be distinguishable from all others in the sample, and if any RNA within the node does not have a number of specific probes greater than the predetermined threshold, the node is considered indistinguishable and grouped with the next nearest RNA or node based on the hierarchy in paragraph 1(b), whereby a set of nodes or RNAs that are designated distinguishable is produced, with a set of ISH nucleic acid probes to each RNA that can discriminate that RNA from all other sequences not found within the node in which it has been grouped. 9. The method of any one of paragraphs 1-8, wherein each ISH probe further comprises a barcode element. 10. The method of paragraph 9, wherein all RNAs within a node designated distinguishable are assigned the same barcode. 11. The method of paragraph 9 or paragraph 10, wherein each barcode element comprises a MERFISH barcode comprised of concatenated readout sub-sequences that can be specifically 33 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT detected in a FISH assay, wherein detection of a given readout sub-sequence in the FISH assay is read as a “1” in a binary barcode. 12. The method of any one of paragraphs 9-11, wherein ISH probes are designed by a) assigning a unique binary barcode to each distinguishable RNA or node; and for each distinguishable RNA or node, b) creating a set of MERFISH probe sequences by concatenating to all or a subset of the ISH probes against that RNA the sequences of the barcode readout elements associated with the bits in which there is a value of “1” in the barcode assigned to the node. 13. A method of detecting a target RNA or set of target RNAs in a fluorescence in situ hybridization assay wherein the target RNA or set of target RNAs is discriminated from one or more closely related non-target RNAs in the same sample, the method comprising hybridizing an ISH probe or set of ISH probes prepared according to any one of paragraphs 1-12 to a sample, and detecting signal from hybridized probe. 14. The method of paragraph 13, wherein detection is preformed via iterative hybridization and detection of members of a set of fluorescently labeled probes specific for barcode elements on the set of ISH probes. 15. The method of paragraph 14, wherein fluorescent labels are removed between iterative rounds of hybridization and detection. 16. The method of any one of paragraphs 1-13, wherein closely related sequences have at least 90% sequence identity. 17. The method of paragraph 2, wherein in (i) the threshold is 90% sequence identity. 18. The method of paragraph 2, wherein in (ii) the threshold degree of similarity is that 90% of the k-mers found in either sequence are found in both sequences. 19. The method of any one of paragraphs 1-13, wherein the closely related sequences are T cell receptor (TCR) variable (V) region or antibody V region coding sequences. 20. The method of paragraph 8, wherein the free energy parameters comprise predicted melting temperature and GC content. 21. The method of paragraph 8, wherein considered probes have 40% to 60% GC content, inclusive. 22. The method of paragraph 8, wherein considered probes have a predicted melting temperature between 55 and 75oC, inclusive. 23. The method of paragraph 8, wherein in step (e), no more than one k-mer is found in RNAs outside that node. 34 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT 24. The method of paragraph 8, wherein considered probe length is 30 nt. 25. The method of paragraph 8, wherein the predetermined threshold number of specific probes in (C) is 10. 26. A computer implemented method of preparing a nucleic acid probe according to the method of any one of paragraphs 1-25. 27. A computer-readable medium having computer-readable signals stored thereon that define instructions which, as a result of being executed in a computer system having a processor and a user interface including a display and an input device, instruct the computer system to perform a method according to paragraph 26. EXAMPLES

[0125] The following provides by non-limiting example a description of the considerations for selecting target sequences and probes that distinguish closely related nucleic acid sequences by in situ hybridization, providing a significant improvement upon the MERFISH technique. The considerations can be advantageously applied to the detection and / or quantitation of different B cell or T cell clones in a given sample, thereby permitting the tracing of B cell or T cell lineages in such sample, but other applications for distinguishably detecting other genes, gene families or cell populations are contemplated. Example 1: Step-by-step procedure for detecting families of RNAs that have different degrees of sequence similarity

[0126] The goal is to develop in situ hybridization (ISH) probes that can fully or partially distinguish individual RNA molecules that belong to families of molecules that share potentially high degrees of sequence similarity. In standard protocols, ISH probes often cannot discriminate between two sequences that are similar because the difference in the free- energy of binding of a probe to two different sequences is often small when the sequence of the targets are similar. For example, if an ISH probe is designed to bind to a 30-nt sequence, it will likely also bind with high probability to any 30-nt sequence that differs from the target sequence by only a single nucleotide variation.

[0127] As a result, one often searches for regions of RNA that are sufficiently dissimilar to regions of any other RNA that might be expressed within the sample of interest, e.g. a cell, in standard ISH probe design approaches. For RNAs that are sufficiently dissimilar in their 35 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT sequence, it is often possible to design sets of ISH probes that target regions of those RNAs and, thus, these ISH probe sets can discriminate these molecules.

[0128] However, when a given RNA has one or more RNAs that could be expressed within the sample to which it is sufficiently similar in sequence, it may not be possible to identify ISH probes that can bind selective to that RNA and not bind to these other, similar, RNAs. In this case, it is often considered that these similar RNAs cannot be targeted with ISH.

[0129] Here a solution is proposed to this problem.

[0130] Step 1: Identify the set of similar or homologous RNAs and define their similarity.

[0131] There are multiple methods for defining the similarity between two RNAs, and any method that can provide a numerical measure of the similarity could be used for these purposes. Formally, two methods have been used.

[0132] In the first, an alignment is performed between the nucleotide sequences of each pair of RNAs and compute the fraction of nucleotides that match between these RNAs. RNAs could be considered too similar to distinguish if they have a have a fraction of nucleotides that exactly match greater than 90%, for example.

[0133] In the second method, each RNA sequence is divided into a series of sub-sequences (often called k-mers) of a given length (this is the k value). For example, each RNA might be divided into all of the 15-nt sub-sequences that can be found within it. The degree of similarity between two RNAs can then be defined by the fraction of k-mers found in either sequence that are also found in both sequences. Again, two RNAs can be considered too similar to distinguish if 90% of the k-mers found in either sequence are found in both sequences.

[0134] Step 2: Organize the set of potentially homologous or similar sequences into a hierarchy that reflects shared similarities between sequences.

[0135] There are multiple ways to define a hierarchy of RNA sequences that share different degrees of similarity to one another. One common implementation of such a hierarchy is a tree, where individual RNAs are grouped into different ‘branches’ or ‘nodes’ based on their respective similarity and then individual branches or nodes themselves are grouped based on a measure of the collective similarity of the sequences within a branch or node to the sequences within a different branch or node.

[0136] To compute a tree for RNA sequences, the pairwise similarity is first defined between individual RNA sequences as described in step 1. Then an agglomerative algorithm is used to systematically group RNAs into branches or nodes based on the pairs of RNAs (or nodes) that are most similar. This iterative algorithm consists of the following steps: 36 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT

[0137] Compute the pairwise similarity between all RNAs or existing nodes:

[0138] The similarity between a given RNA and a node is computed by averaging the pairwise similarity between each RNA within a node and the given RNA.

[0139] The similarity between two nodes is computed by averaging the pairwise similarity between all RNAs in the first node and all RNAs in the second node:

[0140] A node—“A” in this example—can consist of multiple other nodes—“B” and “C” in this example. An RNA is considered to be part of node A if it is part of node B or node C. If node B or C themselves contain nodes, then any RNAs within them or their sub-nodes is considered to be part of node A.

[0141] Select the pair with the greatest similarity and group these together to form a new node.

[0142] This pair could represent two RNAs, two existing nodes, or an RNA and a node.

[0143] Repeat steps 1 and 2 until all RNAs or nodes have been grouped together.

[0144] This protocol will start first with only RNAs and in the first round create one node by grouping the most similar RNAs. Then in the next round it will create a second node by grouping two RNAs or the node it created in the first round to another RNA. In each iterative round of the process above one new node will be created until all RNAs and nodes have been grouped together.

[0145] Step 3: Identify the ISH probe sequences that can discriminate different RNAs or different nodes.

[0146] The task is now to identify sets of ISH sequences that will bind to one set of RNAs but not to another set. To accomplish this task, the following steps are performed.

[0147] For each RNA, consider all possible probe sequences that are of a given length and satisfy properties that are related to the free-energy of binding of a probe to its target sequence, namely the predicted melting temperature and the fraction of the sequence comprised of G or C nucleotides (the GC content).

[0148] The predicted melting temperature can be calculated via a variety of standard methods.

[0149] Only probes of 30-nt in length are considered though other sequence lengths could be considered and it would be possible to consider variable sequence lengths in this process.

[0150] Probes that have a GC content between 40 and 60% are often considered. Probes that have a melting temperature between 55 and 75 C are often considered.

[0151] For each possible probe sequence, determine its specificity to the RNA of interest.

[0152] The specificity can be calculated by breaking down the probe sequence into k-mers 37 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT and then computing the fraction of k-mers within that probe sequence that are found in any other RNA sequence.

[0153] Consider the probe as ‘specific’ if this fraction (‘specificity’) is below a given threshold. For example, if there are 15 k-mers within a given a given probe, it could be considered specific if no more than 1 of these k-mers is found in any other sequence. Count the number of unique ‘specific’ probes that can be found within the RNA under consideration.

[0154] A unique probe can be defined as a probe that does not share an overlapping sequence with another probe to the same RNA of more than a given number of nucleotides. For example, 30-nt probes are considered, but allow them to share as much as 20-nt with another probe to the same RNA. These probes can be considered to overlap, i.e. target a shared portion of RNA sequence, by 20-nt.

[0155] If the number of unique ‘specific’ probes is sufficiently high, consider this RNA as distinguishable from all others and set it aside.10 unique probes are used as the threshold in this example. If the number of unique ‘specific’ probes is not sufficiently high, consider this RNA as not distinguishable. For all RNAs that are judged as not distinguishable, group them with the RNA or nodes to which they are assigned based on the hierarchical tree determined in step 2.

[0156] Once this process has been completed for all RNAs, do one of the following: If all RNAs are judged as distinguishable, then the procedure is done. If some RNAs have been grouped into nodes because they were judged as not distinguishable, repeat the process above but now considering the RNAs grouped into a node as a single object.

[0157] To define the ability to discriminate a node, modify the above process in the following way:

[0158] Consider each possible probe within a given RNA and compute its specificity by calculating the fraction of k-mers within that probe found in all other RNAs, excluding the RNAs found within the node to what that RNA has been assigned. As above, consider a probe specific if less than some fraction of the k-mers within it are found in RNAs outside of the node.

[0159] As above, no more than 1 kmer is typically considered to define a specific probe, although more than 1 kmer could also be considered to define a specific probe. If all RNAs within a node have enough unique, specific probes, then consider that node distinguishable. The same threshold is used as above: 10 probes.

[0160] If any RNA within the node does not have enough unique, specific probes, then 38 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT consider this node as indistinguishable and group it with the next nearest RNA or node as judged by the tree determined in step 2.

[0161] This process will produce a set of nodes or RNAs that are judged as distinguishable and a set of ISH probes to each RNA that can discriminate that RNA from all other sequences not found within the node in which it may have been grouped.

[0162] Step 4: Design ISH probes that can selectively barcode distinguishable RNAs or nodes.

[0163] In this step, the ISH probe sets designed above are leveraged to create MERFISH probes that can be used to assign discriminatory barcodes to individual RNAs. All RNAs within a ‘distinguishable’ node will be assigned the same barcode.

[0164] There are multiple ways to create MERFISH barcodes, but the most common implementation is to create a set of unique binary barcodes that have some error tolerant properties. Error tolerance in binary barcodes arises by not using all possible binary barcodes to encode a given RNA. For example, if one had barcodes of 16-bits in length, these barcodes could be used to distinguish 2^16 possible RNAs. However, if any of those bits were corrupted during measurement, one RNA could be mistaken as another. To provide tolerance for such errors, one uses only a subset of the possible barcodes selected such that a minimum number of bits—termed the Hamming distance—must be corrupted to convert one barcode into another. If the minimum Hamming distance between any two barcodes actually assigned to RNAs is greater than 2, then no single corrupted bit could produce a barcode that is mistaken for a different RNA. Similar approaches can be used to identify corrupted bits so that they can be corrected.

[0165] These binary barcodes, however, need to be translated into nucleic acid sequences. In one common implementation of MERFISH, this process is accomplished by designing a unique 20-nt sequence for each bit in the barcode. Thus, if the barcodes are 16-bits in length, there would be 16 unique 20-nt sequences. These sequences are often termed readout sequences. The nucleic acid sequence that represents a barcode is then defined by the set of readout sequences associated with the bits in which it contains a ‘1’. For example, if an RNA is assigned a barcode of 1001, the nucleic acid sequence associated with the barcode for this RNA would contain readout sequence 1 and 4 but not readout sequences 2 and 3.

[0166] To build ISH probes that target distinguishable nodes and implement this MERFISH barcoding strategy the following is performed:

[0167] Assign a unique binary barcode to each distinguishable RNA or node.

[0168] For each distinguishable RNA or node, create a set of MERFISH probe sequences by 39 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT doing the following:

[0169] For each RNA within a given node, concatenate to all or a subset of the ISH probes against that RNA defined in Step 3 the sequences of the readouts associated with the bits in which there is a value of ‘1’ in the barcode assigned to the node.

[0170] Step 5: Stain a given sample and identify RNAs within it using MERFISH and the probes designed above.

[0171] In this step, the actual measurement is performed to distinguish RNAs or nodes using the probes designed above.

[0172] Fix a sample of interest in which some or all of the RNAs of interest are expressed.

[0173] Standard fixation methods can be used here, including paraformaldehyde or alcohol- based approaches.

[0174] Permeabilize the membrane of the cells within the sample to allow ISH probes to enter.

[0175] Stain or hybridize the sample with the ISH probes designed above in conditions that promote the base-pairing of the complementary regions of these probes to the RNAs they target in the sample.

[0176] Hybridize to the sample a fluorescently labeled probe that is complementary to one of the readout sequences.

[0177] Image the same to determine the molecules within the sample that have ISH probes bound to them that contain this readout sequence.

[0178] Remove this fluorescent signal.

[0179] Repeat steps 4-6 with fluorescently labeled probes complementary to each of the readout sequences used in the MERFISH probes.

[0180] Using the fluorescent pattern observed for each RNA molecule stained in the sample determine the barcode associated with that molecule and, thus, the identity of the distinguishable RNA or node.

[0181] If the barcode is associated not with a single RNA but a distinguishable node that contains multiple RNAs, recognize that the measured RNA could represent any of the RNAs within that node. Example 2

[0182] To test the ability to discriminate B cell clonality, MERFISH probes were designed that can detect the co-expression of VH genes from the heavy chain, and VK / VL genes from the light chain using the strategy described above. VH and VK / VL RNA sequences for 40 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT C57BL / 6 mice were obtained from the IMGT database. All functional alleles from all V gene families were included in the probe design strategy. VH genes from the heavy chain and VK / VL genes from the light chain were treated separately, and two distinct sets of probes were designed for each of these V gene groups. Directly linking heavy and light chain gene expression within a single cell to define B cell clonality is unique to MERFISH.

[0183] All functional alleles from all V gene families were included in the probe design strategy; however, VH and VK / VL genes were treated separately, and two distinct sets of probes were designed for each of these V gene groups.

[0184] To construct a homology tree, the similarity between all pairs of VH sequences were determined by calculating the degree of overlap in k-mers that comprise these sequences. As expected, this analysis revealed high homology between many members of certain VH families. A similarity tree was then created for the VH genes using a standard agglomerative tree construction method and the similarities calculated via k-mer set overlap (FIG.5).

[0185] A public MERFISH probe design pipeline was next used to generate probes for each VH gene, initially treating each possible VH gene as a separate targetable entity. Probes were screened to fall within a predicted melting temperature range, G-C base composition range, and predicted off-target binding to other VH genes. Off-target binding was calculated using a penalty calculated from the fraction of k-mers shared by a probe to given VH and any potential other VH sequence.

[0186] From this analysis, it was found that many VH genes did not have a sufficient number of probes (in this case 10 probes or more were used as the threshold; FIG.5). Any VH gene that did not have a sufficient number of probes was grouped with the nearest VH gene or VH genes on the similarity tree constructed. Grouping of VH genes in this fashion formed nodes, where each node can be thought of as a location on the similarity tree describing either one gene or a set of genes. The same probe design process was repeated as above with one notable difference. When computing the potential off-target penalty associated with a given VH gene other VH genes within that node were not considered, recognizing that the final designed probes might bind to one more VH genes within the node. The number of probes per VH gene were then tabulated again. If any VH gene within a given node did not have a sufficient number of probes, all VH genes within that node were then grouped with the VH genes within the next nearest node, effectively walking up the tree to the next node above the current node.

[0187] Algorithmically, this novel probe design strategy groups highly homologous genes that are adjacent to each other in the similarity tree in the following steps: 1) create probes for 41 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT each gene in the tree.2) Ask whether the number of unique probes designed for a given gene is >10.3) If the probe count is <10, and the next gene in the tree is from the same family, group these two genes together.4) Build a new tree, where the terminal nodes reflect the groupings in step 3.5) Regenerate, de novo, unique probe sets for each node in the newly generated tree.6) Ask whether the number of unique probes designed for a given gene is >10. 7) continuously iterate these steps until ALL nodes have > 10 unique probes that distinguish them from all other nodes. This algorithm was applied for both VH and VK / VL genes with varying stringencies.

[0188] This algorithm was applied to create 3 VH libraries (EYVH, EYVH2, and EYVH3) that each had increasingly strict sequence homology cut offs for V node groupings (60%, 75%, or 90% of all possible k-mers within a probe are not found within any other considered V gene). Each library required a minimum of 10 probes generated per node. EYVH detects 80 nodes, EYVH2 detects 26 nodes, and EYVH3 detects 17 nodes. The same pipeline was then used to design a VK / VL library (EYVK) that detects 68 light chain nodes. Combining each VH library with the VKVL library enables us to track 5440, 1768, or 1156 VH / VL node combinations within B cells, respectively.

[0189] To create a MERFISH-readable encoding probe library from these potential probe regions, each node was assigned a given binary barcode and then concatenated to all possible probes within that node the readout sequences associated with each bit in the assigned binary barcode for which there was a value of ‘1’ (FIG.5 and FIG.6).

[0190] FIG.6 displays a few probe sequences used to detect nodes by EYVH. Note that some nodes contain only a single gene whereas others can contain over dozen (as is the case with VH family 1). Example 3

[0191] To validate the detection capabilities of these VH VK / VL MERFISH libraries in vitro, 192 plasmid constructs were created that each expressed a different fully assembled B cell heavy or light chain from the C57BL / 6 mouse driven by the constitutively active CMV promoter. First, a single plasmid was transiently transfected into HEK293 cells, fixed the cells to coverslips, then processed the samples for MERFISH imaging (FIG.7, left). Every transfected cell fluoresced with the expected optical barcoding pattern for the VH gene used (FIG.7, top right). The average fluorescence was then quantified within the transfected cells for each round of MERFISH imaging to create an intensity trace across all bits for that transfected cell (FIG.7, bottom right). An expected pattern of high and low fluorescence was 42 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT observed that agreed with the barcode assigned to the node in which the single VH was found.

[0192] Next, the in vitro validation efforts were scaled up by performing a parallel transfection with 60 VH genes. HEK293 cells were transfected in individual wells, the contents of each well were trypsinized, and then pooled together before being incubated on and fixed to coverslips (FIG.8, left). These samples were then stained with the VH VK / VL libraries and processed for MERFISH.

[0193] Intensity traces were created for all transfected cells and uncovered unique patterns for each group of transfected cells (FIG.8, top right). To visualize the diversity and resolution of these intensity traces, the dimensionality reduction technique, tSNE, was leveraged. This analysis revealed groups of cells that shared very similar intensity traces, and the inspection of these intensity traces revealed that distinct barcodes for the expected nodes could be clearly distinguished and identified. These distinct groups reflect their VH identity (FIG.8, bottom right; FIG.9).

[0194] To validate VH and VK / VL detection in vivo, a transgenic mouse was used (564Igi; C57BL / 6 background) that harbors a knock-in of a preassembled B cell receptor with a known VH / VK combination (564H / 564L) (FIG.10). Most B cells, but not all, within this mouse model will express the prearranged B cell receptor. Most plasma B cells within the lamina propria of the small intestine expressed the same, and expected, intensity profiles corresponding to the expected combination of VH and VK genes (564H / 564L). Furthermore, several instances of plasma B cells were discovered expressing other VH and VK / VL combinations within the lamina propria, indicated by their unique pixel traces (FIG.11). These pixel traces could also be assigned to specific VH and VK / VL nodes. As only a small fraction of all possible binary barcodes are used in these measurements, this agreement between the measured VH and VK / VL signals and barcodes used for nodes further supports the proper assignment of these nodes.

[0195] The expression level of the BCR can vary dramatically depending on the type of B cell. For example, naïve B cells express much lower levels of BCR than plasma cells. To validate that this approach can discriminate VH nodes for the low expression of levels associated with naïve B cells, the Peyer’s patch of 564Igi mice were prepared and imaged as described above. Instead of observing fluorescent signal filling the entire cell, as expected for the high expression level of plasma cells, individual fluorescent puncta were observed within the naïve B cells in the Peyer’s patch, consistent with the expected low level of expression. The specific fluorescent images associated with the expected barcode elements assigned a ‘1’ 43 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT for the barcode for associated with the node containing 564H showed overlapping fluorescent puncta (FIG.12), indicating the ability of this approach to identify B cell clonality for naïve B cells.

[0196] In addition, the BCR undergoes class switch recombination for the immunoglobulin heavy change, generating different isotypes, e.g., IgA, IgM. To demonstrate that the isotype for individual cells can be determined simultaneously with their BCR clonality, FISH probes were generated that target the various constant regions associated with the heavy chain, and these probes were stained simultaneously with those that identify VH and VL / VK nodes. Measurement of readouts associated with the various constant regions revealed the expected IgA+ plasma cells within the lamina propria of the ileum while measurement of the readouts associated with the barcodes assigned to the VH and VL / VK nodes revealed the clonal identity of each of these IgA+ plasma cells (FIG.13).

[0197] To validate these measurements in vivo and to demonstrate that the identification of VH and VK / VL genes could be performed simultaneously with previous MERFISH protocols targeting individual genes small intestine samples from C57BL / 6 mice were examined (FIG.14). To this end, a library for conventional MERFISH was designed that targeted 591 individual RNA molecules expected to be expressed at various levels within the gut. Individual slices of the mouse ileum were simultaneously stained with the probes targeting VH nodes, VK / VL nodes, all variations of the VH constant region (e.g., IgM, IgE, IgG1, IgA), and these genes. MERFISH was then performed to readout out each of these probe sets sequentially. These measurements revealed a diversity of IgA+ plasma B cells within the lamina propria that each revealed a diverse range of VH and VK / VL genes (FIG. 13, and FIG.14).

[0198] These measurements also simultaneously revealed the expression of the 591 genes within all cells within this slice. With standard cell-segmentation approaches and single-cell visualization approaches (UMAP), it was shown that this technique could distinguish the diversity of cells expected to be found within the ileum in these same measurements (FIG. 14).

[0199] Thus, the VH and VK / VL libraries were used to co-stain Peyer’s Patches and the surrounding ileal epithelium to map not just the BCR diversity throughout the lamina propria, but also the cell type diversity found throughout the tissue (FIG.13 and FIG.14). Example 4 44 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT

[0200] Specific clones of IgA+ plasma cells found within the ileum can be selected by specific bacteria in the gut microbiome. However, it has been observed that in germ-free mice, which lack a gut microbiome, there are specific B cell clones that dominate, and which can be found in multiple different mice. These clones have been termed public clonotypes. To explore the distribution of public clonotypes, C57BL / 6 mice were prepared as germ-free animals. The ileum was harvested from these mice, fixed, permeabilized, and stained with probes that target the 591 genes described above, VH nodes, VL / VK nodes, and all variations of the VH constant region. MERFISH was performed to determine each of these probe sets sequentially. Included in these measurements was a fluorescently labelled antibody against the pan-cell-surface marker, Na-K ATPase. This fluorescence channel in addition to a DAPI stain was also imaged for these samples. These co-stains were used to define the boundaries of cells. Individual IgA+ plasma cells were identified from these images and the specific clonality of these cells was determined from the VH and VK / VL node barcodes. The number of B cell clones that had unique combinations of these VH and VK / VL nodes was determined. The most abundant of B cell clones were all drawn from VH and VK / VL nodes that contain the reported VH and VL / VK sequences that define the most common public clonotypes (FIG.16), and, as expected, most of these clones were found in multiple mice.

[0201] As the addition of the microbiome should drive the selection of B cell clones that are not public clonotypes, these same measurements were repeated in C57BL / 6 mice that contained a specific-pathogen-free (SPF) microbiome. Indeed, the B cell clonal diversity seen in these mice was broader than that observed in germ-free mice, and far more B cell clones were not public clonotypes. Nonetheless, it is noted that several of the abundant B cell clones were still public clonotypes. These measurements indicate that the MERFISH approach described here can faithfully identify B cell clonality and measure changes in the abundance and spatial distribution of B cell clones. Example 5

[0202] BCR-MERFISH can also measure properties of the spatial distribution such as co- occurrence of B cell clones. To demonstrate this capability, BCR-MERFISH was performed as described in the example above against whole mounts of the mouse ileum to determine differential distributions of plasma cell clones along the length of the ileum (FIG.16). Clones were identified that were specifically enriched in either the proximal, medial, or distal regions of the ileum. In addition, clones were identified that were not uniformly distributed throughout these regions but rather were found enriched next to plasma cells of the same 45 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT clonotype or in some cases plasma cells of different clonotypes. The ability to map the spatial distribution of B cell clones both across large tissue areas, to identify local spatial enrichment of clones, and to determine the co-occurrence of specific clones is an advantage of BCR- MERFISH relative to other methods. Example 6

[0203] BCR-MERFISH can also be paired with other forms of BCR characterization, such as BCR sequencing via bulk or single-cell methods. To illustrate this point, paired slices from the mouse ileum were collected and one set of slices was characterized with BCR-MERFISH and the other was characterized with BCR-sequencing from bulk RNA extracted from the slices (FIG.14). Such measurements can be used to validate BCR-MERFISH. For example, in this experiment the abundance of plasma cells using specific VH or VK / VL genes correlated strongly with the abundance of BCR sequences determined via bulk BCR- sequencing that used these VH or VK / VL sequences (FIG.14, top row, right panels). Such correlation cross validates both methods for determining the abundance of B cell clones using these specific V genes. Moreover, BCR-sequencing can also determine the CDR3 sequences, as well as other single-nucleotide-resolution sequence features, associated these specific V gene choices. The combination of BCR-MERFISH and sequencing-based methods can, thus, be used to link elements of the BCR sequence not determined via BCR-MERFISH to the spatial distribution and spatial context in which those B cell clones are found within a tissue sample. 46 4926-8714-3193.3 701039-192190WOPT

Claims

Attorney Docket No.: 701039-192190WOPT CLAIMS 1. A method of preparing a nucleic acid probe that discriminates between its RNA target molecule and a closely related non-target RNA molecule in a sample in a fluorescence in situ hybridization (FISH) assay, the method comprising: a) identifying the set of RNAs in the sample that are within a chosen threshold for similarity to the target RNA molecule sequence; b) defining a hierarchy for the set of RNAs identified in (a) that classifies RNAs in the set on the basis of similarity to each other, such that all RNAs in the set are grouped into nodes by similarity; c) identifying a sequence for an in situ hybridization probe that can discriminate in a FISH assay between RNAs in different nodes identified in (b), wherein a probe that can discriminate between RNAs in different nodes will hybridize, under FISH hybridization conditions, to a target sequence in one node, but not to any non-target sequence in any other node; and d) preparing an in situ hybridization (ISH) probe comprising the probe sequence identified in (c).

2. The method of claim 1, wherein the step (a) of identifying the set of RNAs in the sample that are within a chosen threshold for similarity to the target RNA molecule sequence comprises: i) aligning target RNA sequence pairwise with each other RNA sequence expressed in the sample, wherein RNAs with target sequence identity greater than the threshold are selected for the set; or ii) dividing each RNA sequence expressed in the sample into a series of all sub- sequences of a given length k (k-mers), wherein the degree of similarity between the target RNA and any non-target RNA expressed in the sample is defined by the fraction of k-mers found in either the target RNA sequence or the non-target RNA sequence that are also found in both sequences, and wherein RNAs with a degree of similarity greater than the threshold are selected for the set.

3. The method of claim 1, wherein step (b) of defining a hierarchy for the set of RNAs identified in (a) that classifies RNAs in the set (a) on the basis of similarity to each other comprises generating a phylogenetic tree wherein individual RNAs are grouped into different branches and nodes on the tree. 47 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT 4. The method of claim 3, wherein the phylogenetic tree is computed by defining pairwise similarity between individual RNA sequences expressed in the sample, and applying an agglomerative algorithm to systematically group RNAs into branches and nodes based on the RNAs, or nodes comprised of RNAs, that are most similar.

5. The method of claim 4, wherein the agglomerative algorithm employs an iterative process comprising: i) computing the pairwise similarity between all RNAs or existing nodes; ii) selecting the pair with the greatest similarity, and grouping these together to form a new node, wherein the pair can represent two RNAs, two existing nodes, or an RNA and a node; and iii) repeating steps (i) and (ii) until all RNAs or nodes have been grouped together.

6. The method of claim 5, wherein in step (i), a) the similarity between a given RNA and a node is computed by averaging the pairwise similarity between each RNA within a node and the given RNA; b) the similarity between two nodes is computed by averaging the pairwise similarity between all RNAs in the first node and all RNAs in the second node; and c) a node can consist of multiple other nodes, wherein an RNA is considered to be part of a first node if it is part of a second or third node comprised by the first node, and if the second or third nodes themselves contain nodes, then any RNAs within the second or third nodes or their sub-nodes is considered to be part of the first node.

7. The method of claim 4 or claim 5, wherein the first iteration of the iterative process uses only RNAs, creating one node by grouping the most similar RNAs, and each subsequent iteration creates another node by grouping two RNAs or the node created in the previous round to another RNA, until all nodes and RNA have been grouped.

8. The method of any one of claims 1-7, wherein step (c) of identifying a sequence for an in situ hybridization probe that can discriminate in a FISH assay between target RNAs in different nodes identified in (b) comprises: 48 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT a) for each RNA target, consider all possible probe sequences that are of a given length and satisfy a given set of free energy of binding parameters for a probe to its target sequence; b) determine the specificity of each possible probe sequence of (a) for the RNA, wherein specificity is determined by considering all possible fragments of length k2for the possible probe sequence and computing the fraction of k-mers within that possible probe sequence that are found in any other RNA sequence expressed in the sample, wherein the probe is considered specific if this fraction is below a given threshold; c) sum the number of unique specific probes determined in (b), wherein: when the number of specific probes determined in (b) is greater than a predetermined threshold, the target RNA is designated as distinguishable from all others expressed in the sample via one or more of such probes; and when the number of specific probes is below the predetermined threshold, the target RNA is considered not distinguishable from others expressed in the sample via one or more of such probes, d) grouping all RNAs designated not distinguishable with the RNA or nodes to which they are assigned based on the hierarchy in claim 1(b), and repeating steps (a) – (c), but considering the RNAs grouped into a node as a single object, and e) considering each possible probe within a given RNA, computing its specificity by calculating the fraction of k-mers within that probe found in all other RNA except the RNAs found within the node to which that RNA has been assigned, wherein a probe is considered specific if less than a given fraction of the k-mers within it are found in RNAs outside of that node, and if all RNAs within a node have a number of specific probes greater than a predetermined threshold, the node is considered to be distinguishable from all others in the sample, and if any RNA within the node does not have a number of specific probes greater than the predetermined threshold, the node is considered indistinguishable and grouped with the next nearest RNA or node based on the hierarchy in claim 1(b), whereby a set of nodes or RNAs that are designated distinguishable is produced, with a set of ISH nucleic acid probes to each RNA that can discriminate that RNA from all other sequences not found within the node in which it has been grouped.

9. The method of any one of claims 1-8, wherein each ISH probe further comprises a barcode element. 49 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT 10. The method of claim 9, wherein all RNAs within a node designated distinguishable are assigned the same barcode.

11. The method of claim 9 or claim 10, wherein each barcode element comprises a MERFISH barcode comprised of concatenated readout sub-sequences that can be specifically detected in a FISH assay, wherein detection of a given readout sub-sequence in the FISH assay is read as a “1” in a binary barcode.

12. The method of any one of claims 9-11, wherein ISH probes are designed by a) assigning a unique binary barcode to each distinguishable RNA or node; and for each distinguishable RNA or node, b) creating a set of MERFISH probe sequences by concatenating to all or a subset of the ISH probes against that RNA the sequences of the barcode readout elements associated with the bits in which there is a value of “1” in the barcode assigned to the node.

13. A method of detecting a target RNA or set of target RNAs in a fluorescence in situ hybridization assay wherein the target RNA or set of target RNAs is discriminated from one or more closely related non-target RNAs in the same sample, the method comprising hybridizing an ISH probe or set of ISH probes prepared according to any one of claims 1-12 to a sample, and detecting signal from hybridized probe.

14. The method of claim 13, wherein detection is preformed via iterative hybridization and detection of members of a set of fluorescently labeled probes specific for barcode elements on the set of ISH probes.

15. The method of claim 14, wherein fluorescent labels are removed between iterative rounds of hybridization and detection.

16. The method of any one of claims 1-13, wherein closely related sequences have at least 90% sequence identity.

17. The method of claim 2, wherein in (i) the threshold is 90% sequence identity. 50 4926-8714-3193.3 701039-192190WOPTAttorney Docket No.: 701039-192190WOPT 18. The method of claim 2, wherein in (ii) the threshold degree of similarity is that 90% of the k-mers found in either sequence are found in both sequences.

19. The method of any one of claims 1-13, wherein the closely related sequences are T cell receptor (TCR) variable (V) region or antibody V region coding sequences.

20. The method of claim 8, wherein the free energy parameters comprise predicted melting temperature and GC content.

21. The method of claim 8, wherein considered probes have 40% to 60% GC content, inclusive.

22. The method of claim 8, wherein considered probes have a predicted melting temperature between 55 and 75oC, inclusive.

23. The method of claim 8, wherein in step (e), no more than one k-mer is found in RNAs outside that node.

24. The method of claim 8, wherein considered probe length is 30 nt.

25. The method of claim 8, wherein the predetermined threshold number of specific probes in (C) is 10.

26. A computer implemented method of preparing a nucleic acid probe according to the method of any one of claims 1-25.

27. A computer-readable medium having computer-readable signals stored thereon that define instructions which, as a result of being executed in a computer system having a processor and a user interface including a display and an input device, instruct the computer system to perform a method according to claim 26. 51 4926-8714-3193.3 701039-192190WOPT

Citation Information

Patent Citations

  • Methods and systems for phylogenetic analysis

    US20180223339A1