Method of determining spatial arrangement of RNA species

The method of using dual probe sets and machine learning on colocalized RNA spots addresses false positives in mFISH, improving RNA species localization accuracy and reliability in spatial transcriptomics.

WO2026095867A1PCT designated stage Publication Date: 2026-05-07AGENCY FOR SCI TECH & RES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
AGENCY FOR SCI TECH & RES
Filing Date
2025-10-14
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current decoding analysis in spatial transcriptomics, particularly in multiplexed fluorescence in situ hybridization (mFISH), suffers from false positive detection due to the lack of ground truth information, affecting the sensitivity and accuracy of RNA species localization and identification.

Method used

A method involving two sets of encoding probes and readout probes, with separate hybridization and imaging rounds, followed by colocalization of RNA spots to establish ground truth data, combined with a machine learning model trained on this data to improve decoding accuracy.

Benefits of technology

Enhances the accuracy of RNA species localization by reducing misidentification rates and improving the precision and recall of RNA species detection, thereby enhancing the reliability of spatial transcriptomics analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050669_07052026_PF_FP_ABST
    Figure SG2025050669_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A method of determining a spatial arrangement of a RNA species in a sample, comprising providing a first set of encoding probes and a second set of encoding probes to hybridize with multiple genes of the RNA species, wherein the first set of encoding probes target a subset of the multiple genes, and the second set of encoding probes target each of the multiple genes; performing hybridization rounds of the first set of encoding probes and the second set of encoding probes separately; performing imaging rounds with a first set of readout probes to obtain a first set of RNA spots of the RNA species; performing imaging rounds with a second set of readout probes to obtain a second set of RNA spots of the RNA species; and colocalizing the first set of RNA spots with the second set of RNA spots to obtain a set of colocalized spots.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD OF DETERMINING SPATIAL ARRANGEMENT OF RNA SPECIESCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of priority of Singapore application No. 10202403372Q filed October 30, 2024, the contents of it being hereby incorporated by reference in its entirety for all purposes.TECHNICAL FIELD

[0002] Various embodiments of this disclosure may relate to the field of nucleic acid analysis. More specifically, the present disclosure relates to a method of determining spatial arrangement of RNA species.BACKGROUND

[0003] Spatial transcriptomics technologies reveal gene-expression profiles of cells together with their spatial context in a tissue, promising insight into tissue function and disease etiology. A major subset comprises multiplexed fluorescence in situ hybridization (mFISH) based approaches, including a variety of commercial options that simultaneously measure 100s to 10,000s of individual RNA molecules using oligonucleotide probes tagged with fluorescent dyes. Each RNA species is assigned a different set of read-out probes corresponding to a binary code-word, and read out over multiple imaging cycles. Hence, RNA location and identity must be decoded from stacks of noisy images, a core analysis step preceding all downstream analysis. The effectiveness of the analysis is contingent upon both the sensitivity and accuracy of the hybridization reaction, as well as the subsequent interpretation of the resulting images.

[0004] However, the current decoding analysis poses certain limitations and results in false positive detection, especially when there is a lack of ground truth information embedded within the assay to experimentally measure the accuracy and sensitivity of the decoding algorithm.

[0005] Therefore, there is a need to provide an alternative to improve the analysis of spatial transcriptomics.SUMMARY

[0006] In a first aspect, there is provided a method of determining a spatial arrangement of a RNA species in a sample, comprising providing a first set of encoding probes and a second set of encoding probes to hybridize with multiple genes of the RNA species, wherein the first set of encoding probes target a subset of the multiple genes, and the second set of encoding probes target each of the multiple genes; wherein for each gene in the subset of the multiple genes, the first set of encoding probes targets a set of first positions of the respective gene, and the second set of encoding probes targets a set of second positions of the respective gene, the set of second positions non-overlapping with the set of first positions; performing hybridization rounds of the first set of encoding probes and the second set of encoding probes separately; performing imaging rounds with a first set of readout probes to obtain a first set of RNA spots of the RNA species, wherein the first set of readout probes hybridizable with at least a part of the first set of encoding probes; performing imaging rounds with a second set of readout probes to obtain a second set of RNA spots of the RNA species, wherein the second set of readout probes hybridizable with at least a part of the second set of encoding probes; and colocalizing the first set of RNA spots with the second set of RNA spots to obtain a set of colocalized spots corresponding to a ground truth reference data of the spatial arrangement of the RNA species in the sample.

[0007] In another aspect, there is provided a method of determining a misidentification rate for a MERFISH imaging method, comprising determining a fraction of incorrectly detected negative spots within a second set of imaging spots using the set of colocalized spots as described above to determine the misidentification rate.

[0008] In another aspect, there is provided a gene prediction module for predicting a spatial arrangement of a test RNA species, the gene prediction module comprising a processing unit, and a non-transitory media readable by the processing unit, the media storing instructions that when executed by the processing unit, causes the processing unit to input a test set of imaging spots of the test RNA species into a machine learning model, the machine learning model trained based on the ground truth reference data as disclosed and predicting a spatial arrangement of the test RNA species based on an output of the machine learning model.

[0009] In yet another aspect, there is provided a MERFISH imaging method, comprising determining a misidentification rate for analysing a set of MERFISH imaging spots using the set of colocalized spots obtained from the method as disclosed and using the misidentification rate for the MERFISH imaging method.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In the drawings, like reference characters generally refer to the same parts throughout the different views. The drawings are not necessarily drawn to scale, emphasis instead generally being placed upon illustrating the principles of various embodiments. In the following description, various embodiments of the invention are described with reference to the following drawings.FIG. 1A is a flowchart illustrating a method of determining a spatial arrangement of a RNA species in a sample according to various embodiments of the present disclosure.FIG. IB illustrates a schematic showing Guidestar (GS) probe design and imaging protocol.FIG. 1C shows representative raw images (display range from 0 to 99.9% percentile of the image) from two color channels (Cy5 and Cy7) for GS and MERFISH showing that GS images have higher signal to noise ratio compared to MERFISH.FIG. ID shows example GS images overlaid with the locations of identified GS peaks and MERFISH callouts showing spots called out by MERFISH only, spots detected by GS peakdetection only, and spots detected by both MERFISH and GS.FIG. IE illustrates the percentage of MERFISH spots colocalized with GS spots for each gene (Acly, Gpam, Hnf4u. Ube2z).FIG. 2A illustrates the colocalization of MERFISH callouts with Guidestar (GS) spots was used to designate callouts as true positives, false positives, false negatives and true negatives for calculation of precision, recall and Fl.FIG. 2B illustrates precision, recall and Fl plotted as a function of the target misidentification rate. Target misidentification rate is a parameter in the conventional decoding pipeline, which is the fraction of barcode callouts that correspond to blank controls. Triangle marker denotes maximum Fl score.FIG. 2C illustrates spot features are extracted from true positive, false positive and blank MERFISH callouts of the training data consisting of GS genes to train the random forest classifier. The model is then applied to the rest of the gene callouts to classify the callouts as RNA or artifacts.FIG. 2D depicts GS model and BFT receiver operating characteristic curves (dashed lines: curves for each FOV in test set, solid lines: average across all FOVs). BFT performance at 0.05 misidentification rate is indicated by the crosses for each FOV. Area under the curve (AUC) values are indicated for the average ROC.FIG. 2E depicts graphs showing precision, recall and Fl for the GS model and BFT (at default 0.05 misidentification rate). Center line, median; height of the box, interquartile range (IQR), whiskers, 1.5xIQR. Points represent individual FOVs in the test set.FIG. 2F depicts graphs showing trained GS model and BFT applied to the full gene library. Log correlation to bulk RNA-sequencing fragments per kilobase million (FPKM) values (left) andtotal RNA number of detected RNA (right). Center line, median; height of the box, interquartile range (IQR); whiskers, 1.5 x IQR. Points represent individual FOVs within the dataset.FIG. 3 shows representative raw images from Guidestar imaging rounds. Genes with good image quality and distinctive spots were selected and used for training the random forest classifier, while rejected genes are indicated in red. Numbers next to the gene names indicate FPKM, while the percentages indicate colocalization percentages. Genes were rejected due low FPKM (**) or when distinct spots could not be detected with sufficient quality for colocalization analysis (*).FIG. 4 illustrates fold change in signal intensity (at each distance) over background noise (mean pixel intensity at the edges of a 10x10 ROI from spot center) in cell line data. The steeper curve for Guidestar spots indicates more distinction of signal from background in Guidestar images and better resolved point spread function. Lines indicate the mean while the shaded region indicates the standard deviation.FIG. 5 A illustrates the percentage of MERFISH spots colocalized with Guidestar spots (bars) for each Guidestar gene for AML-12 cell line sample.FIG. 5B illustrates the percentage of MERFISH spots colocalized with Guidestar spots (bars) for each Guidestar gene for liver tissue sample.FIG. 6 illustrates the percentage of Guidestar spots colocalized with MERFISH spots (bars) for each Guidestar gene (left: AML-12 cell line sample; right: liver tissue sample).FIG. 7 illustrates precision, recall and Fl plotted as a function of the target misidentification rate for the liver tissue sample. Target misidentification rate is a parameter in the conventional decoding pipeline, which is the fraction of barcode callouts that correspond to blank controls. Triangle marker denotes maximum Fl score.FIG. 8 illustrates effects of increasing the number of Guidestar genes used on Fl score (5-fold cross validation). Center line, median; height of the box, interquartile range (IQR); whiskers, 1.5 x IQR.FIGs. 9 A and 9B illustrate the effect of sampling (class balancing) methods on Fl score (5-fold cross validation). FIG. 9A: A: All colocalized spots and all MERFISH only spots; B: Augment minority class (MERFISH only spots) with spots decoded as blanks such that positive and negative classes are equal; C: Downsampling of majority class (colocalized spots) such that positive and negative classes are equal; D: Upsampling of minority class such that positive and negative classes are equal. FIG. 9B: A: No sampling; E: Upsampling of minority class (colocalized spots) such that positive and negative classes are equal; F: Downsampling of majority class (MERFISH only spots) such that positive and negative classes are equal. Positive class refers to colocalized spots while negative class refers to MERFISH only spots. Center line, median; height of the box, interquartile range (IQR); whiskers, 1.5 x IQR.FIG. 10 illustrates sequential forward floating feature selection was performed via 5- fold cross validation to maximize the Fl score for the AML- 12 murine liver cell sample. As performance stopped improving beyond 4-5 features, feature selection was stopped at 12 features. The black line shows the mean while the shaded blue area indicates the standard deviation across cross validation sets.FIG. 11 illustrates hyperparameter optimization of random forest classifier to maximize Fl score via 5-fold cross validation. The hyperparameters optimized were the number of estimators in the random forest, criterion for leaf node impurity calculation, the minimum decrease in impurity for a leaf node to be split, the minimum weighted fraction of the sum of total sample weights at a leaf node and the maximum proportion of samples used to train each estimator in the random forest. Center line, median; height of the box, interquartile range (IQR); whiskers,1.5 x IQR.FIGs. 12A to 12C illustrate the Guidestar model optimized on the AML-12 cell line sample can be retrained on a liver tissue dataset to yield improvements over the BFT. FIG. 12A shows that the Guidestar model has reduced precision but increased recall, leading to an overall increase in Fl compared to BFT at 0.05 misidentification rate in the test set. Each point represents a different FOV, which serves as technical replicates, in the test set. FIG.12B shows that the Guidestar model has a similar average receiver-operator-curve (solid lines) to the BFT with slightly higher area-under-the-curve (AUC). BFT performance at 0.05 misidentification rate is indicated by the crosses for each FOV in the test set (dashed lines). AUC values are indicated for the average ROC. FIG. 12C illustrates that when applied to the full library of genes, the Guidestar model filtered counts achieves lower correlation to bulk sequencing but detects 3. lx more RNA than BFT. Center line, median; height of the box, interquartile range (IQR); whiskers, 1.5 * IQR.FIG. 13 illustrates the correlation between bulk RNA sequencing RNA FPKM and BFT or Guidestar model for AML-12 cell line (left) and liver tissue sample (right).DETAILED DESCRIPTION

[0011] The following detailed description refers to the accompanying drawings that show, by way of illustration, specific details and embodiments in which the invention may be practised. These embodiments are described in sufficient detail to enable those skilled in the art to practise the invention. The various embodiments are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments.

[0012] Features that are described in the context of an embodiment may correspondingly be applicable to the same or similar features in the other embodiments. Features that are described in the context of an embodiment may correspondingly be applicable to the other embodiments,even if not explicitly described in these other embodiments. Furthermore, additions and / or combinations and / or alternatives as described for a feature in the context of an embodiment may correspondingly be applicable to the same or similar feature in the other embodiments.

[0013] In the context of various embodiments, the articles “a”, “an” and “the” as used with regard to a feature or element include a reference to one or more of the features or elements.

[0014] Tn the context of various embodiments, the term “about” or “approximately” as applied to a numeric value encompasses the exact value and a reasonable variance, e g. within 10% of the specified value.

[0015] Various aspects of the present disclosure are exemplified in a type of combinatorial multiplexed in situ hybridization referred to as Multiplexed Error-Robust Fluorescence in situ Hybridization (MERFISH) Except where explicitly stated or required, the technology of this disclosure is not limited to this particular implementation.

[0016] In a first aspect, referring to FIG. 1 A, the present disclosure provides a method 100 of determining a spatial arrangement of a RNA species in a sample. The method 100 comprises: in step 110, providing a first set of encoding probes and a second set of encoding probes to hybridize with multiple genes of the RNA species. The first set of encoding probes targets a subset of the multiple genes, and the second set of encoding probes targets each of the multiple genes. For each gene in the subset of the multiple genes, the first set of encoding probes target a set of first positions of the respective gene, and the second set of encoding probes targets a set of second positions of the respective gene. The set of second positions is non-overlapping with the set of first positions. The method 100 further comprises in step 120, performing hybridization rounds of the first set of encoding probes and the second set of encoding probes separately. Further, the method comprises in step 130, performing imaging rounds with a first set of readout probes to obtain a first set of RNA spots of the RNA species, wherein first set of readout probes is hybridizable with at least a part of the first set of encoding probes.Additionally, the method comprises in step 140, performing imaging rounds with a second set of readout probes to obtain a second set of RNA spots of the RNA species, wherein the second set of readout probes is hybridizable with at least a part of the second set of encoding probes. The method further comprises in step 150, colocalizing the first set of RNA spots with the second set of RNA spots to obtain a set of colocalized spots corresponding to a ground truth reference data of the spatial arrangement of the RNA species in the sample.

[0017] The term “probe” as used herein refers to a substrate (a DNA, RNA or nucleic acid analog) that binds specifically or hybridizes to a target sequence (such as target nucleic acid sequence).

[0018] The term “gene” as used herein refers to DNA sequences that code for a product (such as a RNA or a protein)

[0019] The term “sample” as used herein is anything that may contain one or more nucleic acid of interest to be detected, quantified, and / or imaged in accordance with this technology. Such samples are typically referred to in this disclosure as a “cell sample” or “tissue sample”, but this designation should not be considered limiting.

[0020] The Multiplexed Error-Robust Fluorescence in situ Hybridization (MERFISH) referred to in this disclosure may refer to a combination of in situ hybridization (FISH) with multiplexed error-robust barcoding, that hybridizes encoding probes to target nucleic acids (such as RNA species) in a sample of interest (such as cell or tissue sample). Each of the encoding probes may include a barcode sequence complementary to a readout probe. The combination of the encoding probes and the corresponding readout probes may allow sequential hybridization and imaging to identify spatial location of RNA species.

[0021] The term “encoding probes” as used herein refers to probes that hybridize directly to the target nucleic acids inside the sample of interest, such as oligonucleotide probes. Each encoding probe contains targeting sequences which are complementary to the sequence of thetarget nucleic acids. Each encoding probe may carry one or more short “readout sequences” that work or function as barcodes for spatial readout.

[0022] The term “readout probes” as used herein refers to probes that are complementary to the readout sequences, such as fluorescently labelled probes. During the imaging cycle, the readout probes may be applied such that they bind to the corresponding encoding probes, thus illuminating specific spots Tn some instances, the readout probe may be hybridizable with at least a part of the corresponding encoding probe. Different readout probes can be sequentially applied and washed away, allowing multiple rounds of imaging or measurement without permanently altering the sample.

[0023] The length of an encoding probe can be between 50 to 130 nucleotides, or between 50 to 80 nucleotides, or between 80 to 90 nucleotides, or between 90 to 100 nucleotides, or between 100 to 130 nucleotides. The length may be longer depending on the additional readout sequences on the encoding probe.

[0024] The length of a readout probe can be between 25 to 40 nucleotides, or between 25 to 35 nucleotides, or between 30 to 35 nucleotides, or 25 nucleotides, or 26 nucleotides, or 28 nucleotides, or 30 nucleotides, or 32 nucleotides, or 34 nucleotides, or 36 nucleotides, or 38 nucleotides, or 40 nucleotides.

[0025] The length of the targeting sequence can be between 20 to 50 nucleotides, or between 20 to 30 nucleotides, or between 30 to 35 nucleotides, or 35 to 40 nucleotides, or 40 to 50 nucleotides. The targeting sequences are typically about 30 nucleotides in length.

[0026] Tn one example, hybridization is typically conducted for four rounds for Guidestar probes (8 Guidestar genes) and eight rounds for MERFISH probes (16 bits MHD4 MERFISH), whereby each round uses two color channels. The Guidestar probes correspond to the first set of encoding probes, and the MERFISH probes correspond to the second set of encoding probes. The number of rounds depends on the encoding probes, the desired number of genes to bedetected, and the number of color channels used. For instance, if only one color channel is used for Guidestar probes (8 Guidestar genes) and MERFISH probes (16 bits MHD4 MERFISH), twenty four rounds of hybridization would be required. In some instances, if three or more color channels are used for Guidestar probes (8 Guidestar genes) and MERFISH probes (16 bits MHD4 MERFISH), it is possible that eight rounds of hybridization or fewer would be required when Guidestar genes and MERFISH bits are to be imaged on the same hybridization round using different channels.

[0027] In an exemplary embodiment, the number of genes targeted by the first encoding probes may be about 4 to 12 genes. The number of genes targeted by the second encoding probes may be about 100 to 150 genes. In one example, the first set of encoding probes target a subset of the multiple genes, which corresponds to 7 genes out of 133 genes. In another example, the first set of encoding probes target 4 genes out of 133 genes.

[0028] In some embodiments, the proposed method further comprises classifying pixels of the set of colocalized spots as true positive spots; and obtaining the ground truth reference data of the spatial arrangement of the RNA species using the true positive spots.

[0029] The first set of readout probes corresponds to a set of Guidestar readout probes. The second set of readout probes corresponds to a set of MERFISH readout probes. The set of first positions and the set of second positions are interspersed (or interleaved) across a plurality of regions of the respective gene.

[0030] In some embodiments, each of the first set of readout probes has a higher number of fluorophores than each of the second set of readout probes The number of fluorophores in each of the first set of readout probes may be in the range of 1.5 to 5 times more than each of the second set of readout probes. In one example, each of the first set of readout probes (76 fluorophores) contains 2.5 times the number of fluorophores than each of second set of readout probes (30 fluorophores).

[0031] Non-limiting examples of fluorophores include fluorescent entities or phosphorescent entities, for example, cyanine dyes (e.g., Cy3, Cy5 or Cy7), Alexa Fluor dyes (e g. Alexa Fluor 594) or near-infrared (NIR) dyes (e g. IR800 or IR800CW). Other suitable signaling entities are known to the person skilled in the art.

[0032] In some embodiments, the method of the present disclosure further comprises independently performing a respective imaging round for each gene in the subset of the multiple genes, to obtain the first set of RNA spots of the RNA species.

[0033] In a specific example, each of the first set of encoding probes and each of the second set of encoding probes is a DNA probe. In another example, each of the first set of readout probes and each of the second set of encoding probes is a DNA probe.

[0034] In various embodiments, the subset of the multiple genes comprises genes with moderate to high expression levels. In one example, expression level of the genes may be determined by the estimation of bulk measurement based on the RNA copy number. In a specific example, genes are selected based on fragment per kilobase million (FPKM) values from bulk RNA sequencing, which are greater than 0.0015 and less than 260. The median FPKM value is about 3.6. As expression levels may vary in different spatial locations of a sample, bulk measurements may avoid targeting genes that are overly expressed throughout the whole sample.

[0035] In various embodiments, the proposed method further comprises colocalizing the first set of RNA spots with the second set of RNA spots using a maximum euclidean distance threshold In other embodiments, a various predetermined measures or thresholds may be used in colocalizing the first set of RNA spots with the second set of RNA spots.

[0036] In various embodiments, the proposed method further comprises performing imaging rounds of the first set of readout probes to obtain a first set of imaging spots; and identifying local maxima from the first set of imaging spots to obtain the first set of RNA spots.

[0037] In various embodiments, the proposed method further comprises performing imaging rounds of the second set of readout probes to obtain a second set of imaging spots; and downsampling a subset of the second set of imaging spots that do not colocalize with the first set of imaging spots to obtain a set of down-sampled second imaging spots, wherein the set of down- sampled second imaging spots matches the number of spots in the set of colocalized spots. The proposed method further comprises training a machine learning model using the set of colocalized spots as positive examples for the set of down-sampled second imaging spots.

[0038] In another aspect, the present disclosure provides a method of determining a misidentification rate for a MERFISH imaging method, comprising determining a fraction of incorrectly detected negative spots within a second set of imaging spots using the set of colocalized spots as described above to determine the misidentification rate

[0039] In another aspect, the present disclosure provides a gene prediction method, the gene prediction method comprising inputting a test set of imaging spots of a test RNA species into a machine learning model, the machine learning model trained based on the ground truth reference data as disclosed, and predicting a spatial arrangement of the test RNA species based on an output of the machine learning model

[0040] In various embodiments, the test set of imaging spots is obtained by a MERFISH method.

[0041] In some embodiments, the gene prediction method further comprising classifying a selected subset of the test set of imaging spots as true test positive spots using the machine learning model; and predicting a spatial arrangement of the test RNA species based on the true test positive spots.

[0042] In various embodiments, the machine learning model is a classifier-based model selected from one of: a random forest classifier, a decision tree model and a support vector machine.

[0043] In another aspect, the present disclosure provides a gene prediction module for predicting a spatial arrangement of a test RNA species, the gene prediction module comprising a processing unit; and a non-transitory media readable by the processing unit, the media storing instructions that when executed by the processing unit, causes the processing unit to: input a test set of imaging spots of the test RNA species into a machine learning model, the machine learning model trained based on the ground truth reference data as disclosed; and predicting a spatial arrangement of the test RNA species based on an output of the machine learning model.

[0044] In one embodiment, the test set of imaging spots is obtained by a MERFISH method.

[0045] In various embodiments, the processing unit is further configured to: classify a selected subset of the test set of imaging spots as true test positive spots using the machine learning model; and predict a spatial arrangement of the test RNA species based on the true test positive spots.

[0046] In another aspect, the present disclosure provides a MERFISH imaging method, comprising: determining a misidentification rate for analysing a set of MERFISH imaging spots using the set of colocalized spots obtained from the method as disclosed; and using the misidentification rate for the MERFISH imaging method

[0047] In various embodiments, the misidentification rate is in a range of 0.01 to 0.6.

[0048] In the present disclosure, “Guidestar” OR “GS” refers to the technology or platform used to integrate an additional set of probes to label RNA transcripts containing the target sequence.

[0049] This technology or platform may involve three key components. The first component is the design of Guidestar probes (or first set of encoding probes and readout probes) for a specific subset of RNA species which are also being targeted by the multiplexed FISH or in- situ sequencing assays. The second component is the acquisition of Guidestar images and the accurate detection of co-localized Guidestar RNA spots. The third component is the RNAcalling analysis software that uses Guidestar as ground-truth to (1) evaluate both the sensitivity and accuracy of existing RNA decoding pipelines and (2) train a machine-learning model such as a machine-learning based classifier (e g. random forest with adaptive boosting (RF Adaboost), decision tree with gradient boosting (DT Gradboost), or Support Vector Machine (SVM)) to distinguish true positive from false positive RNA calls. The trained classifier can be used to improve RNA calling performance, including on genes not targeted by Guidestar, and may even be generalized to new experiments.

[0050] Guidestar may be viewed as a spike-in control approach that labels a subset of RNA transcripts with an additional set of probes (such as encoding probes). These additional probes may be imaged as part of the MERFISH assay but on separate hybridization cycles, which then serves as a matching ground truth dataset to evaluate detection accuracy for the subset of RNA transcripts. The “guide” probes are integrated with the MERFISH probe-library through the same process and introduced with MERFISH encoding probes in a single hybridization step. This approach allows a user to evaluate if the spots detected in the “guide” images spatially colocalize with the decoded MERFISH spots, allowing for validation of RNA decoding at the single-molecule level.

[0051] Although various aspects of the present disclosure are exemplified in the context of MERFISH, it may be appreciated that the proposed method is not limited to MERFISH, and may also be implemented or adapted into other forms of spatial transcriptomics technology or combinatorial multiplexed FISH.

[0052] To facilitate a better understanding of the present disclosure, the following examples of specific embodiments are given. In no way should the following examples be read to limit or define the entire scope of the disclosure. One skilled in the art will recognize that the examples set out below are not an exhaustive list of the embodiments of this disclosure.EXAMPLES

[0053] 1. Methods

[0054] Probe Design

[0055] 30-nt targeting regions for the Guidestar probes (or first set of encoding probes) andMERFISH probes (or second set of encoding probes) were identified using a previously published algorithm [Goh, J. J. L. et al. Highly specific multiplexed RNA imaging in tissues with split-FISH. Nat.Methods 17, 689-693 (2020)]. Transcript sequences from the GENCODE website (mouse m4) were used as reference. A specificity table was calculated using 15-nt seed and 0.3 specificity cut-off was used for MERFISH probes, while a 0.5 specificity cut-off was used for Guidestar probes. Quartet repeats ('AAAA', 'TTTT', 'CCCC, and 'GGGG’) were excluded from the possible target regions. The Guidestar and MERFISH probes were randomly distributed on the target regions. In one example, the Guidestar probes comprise SEQ ID NO: 1 to 152.

[0056] Table 1 : Guidestar Library Readout Sequences

[0057] Table 2: Multiplexed FISH Library Codebook

[0058] Probe library amplification and preparation

[0059] The probe library (Twist Bioscience) was amplified according to the protocol [Chen, K. H., Boettiger, A. N., Moffitt, J. R., Wang, S. & Zhuang, X. Spatially resolved, highly multiplexed RNA profiling in single cells. Science 348, science. aaa6090- (2015)] as follows. First, the oligonucleotide pool was amplified by limited-cycle PCR using Phusion Hot Start Flex 2X Master Mix (NEB, cat. no. M0536L), using an annealing temperature of 66°C. The T7 promoter sequence was introduced on the reverse primer during PCR. Then, overnight in vitro transcription (IVT) was done using a high-yield 1VT kit (NEB, cat. no. E2050S), before reverse transcription was done on the RNA template (Thermo Fisher, cat. no. EP0753). The RNA was then cleaved by alkaline hydrolysis, generating single-stranded DNA (ssDNA). Finally, the ssDNA were purified using magnetic beads (Beckman Coulter, cat. no. A63882) and elutedwith nuclease-free water. The probes were freeze-dried and stored at -20°C until use. The following primers were used for PCR: 5’-TGGTTCAATCGTATGCCCGT-3’ and 5’- TAATACGACTC ACTATAGGGGTC ACTTAGCCAACGCCGAT-3 ’ .

[0060] Preparation of cell-culture samples

[0061] AML12 (ATCC CRL-2254) cells were cultured in high-glucose DMEM (Hyclone, cat. no. SH30022.01) that was supplemented with 10% FBS (Thermo Fisher, cat. no. 26140079). The cells were grown to -80% confluency on autoclaved 40 mm round coverslips (Warner Instruments, cat. no. 64-1500) placed in 60 mm dishes. The samples were then fixed in 4% vol / vol paraformaldehyde (Electron Microscopy Sciences, cat. no. 15714) in 1 x PBS for 15 min at room temperature. After fixation, the cell samples were quenched with 0.1 M glycine (1 stBASE, cat no. BIO-2085) for 1 min at room temperature before rinsing with 1 x PBS. Fixed samples were stored in -80°C until use. Prior to usage, samples were permeabilized in 70% ethanol overnight at 4°C.

[0062] Coverslip functionalization

[0063] Coverslips were functionalized prior to tissue sectioning. The coverslips (Warner Instruments, cat. no. 64-1500) were cleaned by sonication in IM KOH using an ultrasonic water bath for 20 min, then rinsed with MilliQ water thrice followed by a rinse with 100% methanol. Next, the coverslips were immersed in an amino-silane solution (3% vol / vol (3- aminopropyl)triethoxysilane (Merck cat no. 440140-500ML), 5% vol / vol acetic acid (Sigma, cat. no. 537020) in methanol) for 2 min at room temperature before being rinsed thrice with MilliQ water. The functionalized coverslips were then air-dried overnight and used immediately or stored in a dry environment at room temperature for up to two months.

[0064] Tissue sample preparation

[0065] All animal care and experiments were carried out in accordance with Agency for Science, Technology and Research (A*STAR) Institutional Animal Care and Use Committee(IACUC) guidelines (IACUC #211580). Histology work was performed by the Advanced Molecular Pathology Laboratory, IMCB, A* STAR, Singapore. Briefly, 8 week old C57BL / 6NTac female mice (In Vivos) were euthanized with ketamine, and their livers were collected. The livers were cut into smaller pieces and frozen as soon as possible in Optimal Cutting Temperature compound (Tissue-Tek O.C.T.; VWR, cat. no. 25608-930). The tissue blocks were stored at -80°C. 7 pm sections of the tissue blocks cut with a cryotome onto the functionalized coverslips. After air drying for 5-10 min at room temperature, samples were fixed in a 4% vol / vol paraformaldehyde in l x PBS solution for 15 min. Samples were then rinsed with l x PBS and stored at -80°C for future use. Samples were permeabilized in 70% ethanol overnight at 4°C before use.

[0066] Guidestar and Multiplexed FISH experiments

[0067] After permeabilization, the samples were rehydrated in 2x saline-sodium citrate (SSC, Axil Scientific, cat. no. BUF-3050-20xlL) for 5 min. Samples were then pre-hybridized in a 20% formamide wash buffer, containing 20% deionized formamide (Ambion, cat. no. AM9342, AM9344) and 2x SSC, at 47°C. AML-12 cell samples were pre-hybridized for 30 min, while the mouse liver section samples were pre-hybridized for at least 3 hours. The library probes were diluted to a concentration of 100 pM in 20% hybridization buffer that consists of 20% deionized formamide (vol / vol), 1 mg / ml yeast tRNA (Invitrogen, cat. no. 15401011) and 10% dextran sulfate (Abeam, cat. no. abl46569) (wt / vol) in 2x SSC. The samples were stained with the encoding probes at 47°C overnight in a humidified, air-tight container. The samples were then washed in a 20% formamide wash buffer at 47°C for 15 min, twice. Finally, the samples were washed twice with 2x SSC before being imaged immediately or stored for no longer than 24 h at 4°C in 2x SSC before imaging.

[0068] Guidestar and Multiplexed FISH imaging cycle 1

[0069] All datasets were acquired on a homebuilt imaging system. The setup was based around a Nikon Ti2-E body, an Andor Sona 4.2B-11 sCMOS camera. A Nikon CFI Plan Apo Lambda *60 1.4-n.a. oil-immersion objective was used for all imaging experiments. For illumination, MPB Communications fiber lasers were used for Cy5 (647 nm) and IRDye 800CW (750 nm), respectively: 2RU-VFL-P-1000-647-B1R (1000 mW) and 2RU-VFL-P-500- 750-B1R (500 mW). Samples were secured to the microscope stage using the Bioptechs FCS2 flow chamber, and a custom-built computer controlled fluidics system was used to perform sequential rounds of readout hybridization and imaging. For each round of imaging, readout probe solution was flowed into the chamber and incubated with the sample for 30 mins at room temperature. The readout probe solution consisted of 10 nM of each fluorescently labeled readout probe (Cy5 or IR800) in a 10% hybridization buffer. Unbound readout probes were then removed by washing the sample with a 10% formamide wash buffer, followed by a wash with 2x SSC. Next, the imaging buffer flowed into the chamber before image acquisition. The imaging buffer consisted of 50 mM Tris-HCl pH 8, 2 mM Trolox (Sigma, cat. no. 238813), 10% glucose, 0.5 mg / ml glucose oxidase (Sigma, cat. no. G2133) and 40 gg / ml catalase (Sigma, cat. no. C30-100MG) in 2x SSC. Fluorescent signals were removed by flowing 2x SSC into the chamber and photobleaching the sample by exposure to 647 nm and 750 nm lasers for 20 s. This hybridization and photobleaching cycle were repeated until all the bits were imaged. The readout probes (or first set of readout probes) for Guidestar were imaged first, followed by the MERFISH readouts (or second set of readout probes), to ensure optimal quality of the Guidestar images and avoid potential signal degradation arising from long acquisition times. The total number of hybridization rounds (each round consists of 2 color channels) is the sum of 4 Guidestar rounds (8 Guidestar genes) and 8 MERFISH rounds (16 bits MHD4 MERFISH).

[0070] Data Analysis for RNA decoding

[0071] Image preprocessing and MERFISH analysis

[0072] An analysis pipeline used to identify RNA spots is as follows. Pre-processing steps (image registration, image filtering, normalization) and comparison to codewords were performed as previously described. For the analysis parameters, the minimum magnitude threshold (mean intensity) per spot was set to 0.1, the maximum distance threshold (Euclidean distance to a unit-normalized gene’s codeword) to 0.605 and the minimum spot size (pixel region) to 2 pixels. Bit normalization and image filtering parameters were used as previously described.

[0073] Implementation of adaptive thresholding method

[0074] Adaptive thresholding was further implemented to identify RNA spots at target misidentification rate. Briefly, a 3D histogram of mean intensity, minimum distance to gene’s code-words and number of spots was constructed. A similar 3D histogram was constructed for seven blank barcodes. For each histogram bin, the fraction of blank barcodes over total gene + blank barcodes were calculated as the blank fraction score (BFS). A high BFS suggests a high probability of barcode misidentification for that histogram bin. A threshold was then set to filter spots by their BFS, which is determined by the histogram bin matching each spots’ characteristics This threshold is tuned to achieve the desired misidentification rate which is defined as mean count per blank barcode / mean count per gene barcode. The gross misidentification rate of all barcodes is conventionally set as 0.05, meaning a 5% chance of misidentifying an artifactual spot as a true gene.

[0075] Signal and background brightness comparison

[0076] To quantify the improvement in Guidestar image quality over MERFISH imaging bits, 500 spots were randomly sampled in Guidestar and MERFISH images. The center of each spot was estimated by detecting local maxima and the pixel intensity was plotted as a function of distance to the spot center. The central pixel intensity was taken to be the peak signal while the furthest pixels’ mean intensity within a 10x10 ROI was taken to be the background noise.Signal-to-background ratio was calculated as the mean intensity of pixels at each distance over the background intensity (FIG 4).

[0077] Colocalization analysis

[0078] The peaklocalmax function of the Scikit-image vO.19.3 library22 in Python was used to find Guidestar spots. The spots were detected by finding local maxima in the filtered image. Only peaks above an intensity threshold were called out as Guidestar spots. The threshold for each of the Guidestar genes was manually set. A MERFISH spot is considered to be colocalized if there is a Guidestar spot of the corresponding gene within 4 pixels from the centroid of the MERFISH spot. It may be appreciated that a different number of pixels, such as 2 or 6, may also be used as a threshold value in different applications.

[0079] Training of classifiers using Guidestar

[0080] Quality control and Guidestar gene selection

[0081] Due to differences in probe binding kinetics, some probe sets may have image quality issues due to nonspecific binding or poor probe clearance. Hence, a quality control step was performed on each of the 7 Guidestar genes in the library to ensure that the Guidestar genes used yielded high quality images (FIG. 3). Igfl, 1'igr and Cpsl were removed for the AML-12 cell line sample as these RNAs have low FPKM (the probe-library was designed for liver tissue data). All possible combinations of the remaining 4 Guidestar genes were then tested using 5- fold cross validation and a random forest classifier, and it was found that using all of the 4 high- quality genes: Acly, Gpam, Hnf4a and Ube2z yielded the highest Fl score (FIG. 8). After optimizing gene combinations, the model was retrained using these 4 genes on the full training set and evaluated its performance on the held-out test set.

[0082] Class balancing approach

[0083] As varying numbers of positive training examples (spots colocalized between guidestar and MERFISH) and negative training examples (spots found in MERFISH only) wereobserved, a variety of approaches were tested to even out class imbalances (FIGs. 9A and 9B). It was sought to use as many of the colocalized spots as possible as they provide valuable information for training the classifier. Therefore, if there were more MERFISH-only than colocalized spots, MERFISH-only spots were downsampled randomly to attain a 1 : 1 ratio with colocalized spots within each field of view (FOV). On the other hand, if there were more colocalized than MERFISH-only spots, the colocalized spots were not downsampled. Instead, the negative examples were augmented by adding spots corresponding to blank code-words. To prevent the model from learning excessively from blanks, the inventors did not add more blanks than MERFISH-only spots, such that blanks and MERFISH only spots were at a 1 : 1 ratio.

[0084] Additional features and feature selection

[0085] Conventionally, the mean intensity of pixels within the spot, the hamming distance between the imaged codeword and the closest theoretical barcode, and the number of pixels in the spot (spot size) are considered in calculating the blank fraction score. These 3 features were expanded to include intensity and distance-related characteristics of second and third closest theoretical barcodes (Table 3). Sequential forward feature selection was then performed to obtain the feature set that produces the best Fl score during cross validation out of the 15 features extracted from the data (FIG. 10).

[0086] Table 3: Descriptions of spot features used in Random Forest Classifier

[0087] Machine Learning Model training

[0088] A classifier was constructed for binary classification (0 = not RNA callout, 1 = true RNA callout). The classifier may be a machine learning model, such as Random Forests, Logistic Regression, Decision Trees, Support Vector Machines (SVMs), Naive Bayes, and Artificial Neural Networks (ANNs). All putative MERFISH callouts were then labeled (without adaptive thresholding) as the following: MERFISH callouts that colocalized with Guidestar spots were labeled as 1 (gene spot). MERFISH callouts that did not colocalize with Guidestar spots were labeled as 0 (not gene spot). 80% of the data was used as the training set while the remaining 20% served as the test set. All possible combinations of Guidestar genes (FIG. 8) and sampling methods were tested to reduce class imbalance (FIG. 9). After optimizing for combinations and numbers of genes, it was found that using a smaller subset of genes did not improve accuracy. For this reason, and to maximize the diversity of the training data and improve generalizability, all Guidestar genes that passed initial QC were used. For the cell line dataset, it was found that the best sampling method was to augment the minority negative class with spots decoded as blanks instead of downsampling the majority positive class which contained valuable information on colocalized spots. For the tissue sample dataset, the majoritynegative class was downsampled as there were no sources for positive class augmentation. Sequential forward feature selection was performed using the mlxtend v0.22.0 library in Python (FIG. 10) to optimize the features used for both datasets separately. Finally, hyperparameters of the random forest classifier from the sklearn v 1.1.1 library in Python (FIG. 11) were optimized. The optimized hyperparameter set used was: min_samples_split=0.001, min_impurity_decrease=O .01, max_features=l, max_depth=12, criterion- gini', min_weight_fraction_leaf=0.01, n_estimators=100, max_samples=0.8. All model optimization was conducted via 5-fold cross validation on the training set to maximize the Fl score.

[0089] Model evaluation

[0090] The optimized model was applied to the test set and the precision, recall, and Fl score for each FOV was calculated Precision is defined by the fraction of colocalized MERFISH callouts over all MERFISH callouts that were deemed as gene spots. Recall is defined by the fraction of colocalized MERFISH callouts deemed as gene spots over all guidestar callouts (FIG. 2A). The Fl score is 2 * precision * recall / (precision + recall). A receiver operating characteristic curve (ROC) was generated for each FOV in the test set (FIG. 2D, dashed lines) and the mean area under the curve across FOVs was reported. The true positive rates of individual FOV’s ROC (FIG. 2D, FIG. 12C, dashed lines) was interpolated across false positive rates ranging from 0 to 1, and averaged to obtain the mean ROC curve (FIG. 2D, FIG. 12C, solid lines). The optimized model was applied to all other genes in the library. Pearson’s correlation of log count values between bulk FPKM and gene spot counts for each FOV as well as over the full dataset was calculated. The FPKM values from bulk RNA- seq of mouse liver tissues were downloaded from the ENCODE portal (ENCFF844MJF and ENCFF271DWG) and averaged. FIG. 13 shows the correlation of BFT / Guidestar model callouts with the bulk RNA-seq FPKM.

[0091] 2. Results

[0092] An integrated library was designed that combined a 133-gene MERFISH panel containing genes of interest in liver tissue with additional Guidestar probes targeting 7 of the 133 genes. The subset of genes targeted by Guidestar probes were chosen to have (1) sufficiently long transcripts to accommodate the additional probes and (2) moderate to high expression levels in the tissue of interest as estimated from bulk measurements to ensure sufficient numbers of detected RNA spots This oligonucleotide probe-set was then hybridized with samples from the AML- 12 murine hepatocyte cell line and performed the Guidestar- integrated MERFISH assay protocol (FIG. IB). After performing quality control checks on the 7 Guidestar genes and removing genes with low FPKM on the cell line sample (FIG. 3), 4 Guidestar genes (Acfy, Gpam, Hnf4a and Ube2z) were selected for evaluation.

[0093] As shown in FIG. IB, GS probes are interspersed or interleaved with MERFISH probes on a small subset of genes from the MERFISH library. For each gene, the GS probes may target a set of first positions of the gene and the MERFISH probes may target a set of second positions of the gene. In various embodiments, the set of first positions may be nonoverlapping with the set of second positions. In various embodiments, the set of first positions may be interspersed with the set of second positions Guidestar uses 2.5 times the number of fluorophores compared to single imaging round in MERFISH. Guidestar images are acquired, one gene at a time, before performing the combinatorial MERFISH assy. Guidestar RNA spots are identified by peak finding using local intensity maxima, while MERFISH images are decoded by matching voxel sequences to the codebook. The spots identified in Guidestar are colocalized with those decoded via MERFISH.

[0094] It was observed that the Guidestar images had superior image quality relative to MERFISH images, with RNA spots that could be clearly distinguished from background fluorescence (FIG. 1C). To verify the improvement in Guidestar image quality over MERFISH, the drop off in intensity from the center of each spot to the periphery was quantified (FIG. 4),showing that Guidestar spots had at least 2- fold improved signal-to-background ratio (SNR) at the spot center, relative to spots in the MERFISH imaging rounds. The extent of colocalization between Guidestar spot locations (first set of RNA spots) and MERFISH spot locations (second set of RNA spots) was then assessed, defining colocalization as a maximum euclidean distance of 0.48 um based on percentages of colocalization at different distances. For each of the Guidestar imaging rounds (corresponding to individual chosen genes), RNA spot locations were identified by detecting local maxima [van der Walt, S. et al. scikit-image: Image processing in Python. (2014) doi: 10.48550 / ARXIV.1407.6245], MERFISH images were analyzed (preprocessing, registration and decoding) as previously described to obtain RNA identities and locations [Chen, K. H., Boettiger, A. N., Moffitt, I. R., Wang, S. & Zhuang, X. Spatially resolved, highly multiplexed RNA profiling in single cells. Science 348, science. aaa6090- (2015); Goh, J. J. L. et al. Highly specific multiplexed RNA imaging in tissues with split-FISH. Nat.Methods 17, 689-693 (2020)]. It was found that across the 4 Guidestar genes, 71 - 87% of spots called by MERFISH colocalized with Guidestar (FIGs. ID - IE, FIGs. 5A - 5B) while 27 - 44% of spots detected by Guidestar were also detected by MERFISH (FIG 6) The fraction of spots detected by both Guidestar and MERFISH out of all Guidestar spots was similar to the expected detection efficiency of MERFISH in high- abundance libraries (MERFISH detected 21±4% of RNA spots determined by smFISH previously). As a negative control, colocalization percentages between MERFISH and unmatched guidestar genes (e g. Acly Guidestar spots vs Gpam MERFISH spots) were computed, showing a mean (averaged across all 12 pairs) colocalization of 4.0% (FIGs. 5 A and 5B). Together, these tests suggest that Guidestar can be used as a ground-truth reference to evaluate MERFISH RNA decoding results.

[0095] Guidestar enables parameter optimization in existing MERFISH decoding pipelines

[0096] A key step in MERF1SH decoding is the use of the number of detected negative controls (specifically “blank” codewords) to tune the stringency of RNA callout filtering. ‘Blanks’ are codewords that do not encode any gene in the library but are also at a Hamming distance of 4 away from the other gene-encoding codewords, and hence serve as a negative control. The key hyperparameter for this filtering step is the “misidentification rate”, which measures the fraction of “blank” codewords that would be called out as true spots after filtering. This is often set heuristically at 0.05 (i.e. 5% of the “blanks” would be accepted as true RNA). This method of RNA spot filtering is referred to as the blank fraction threshold (BFT) approach in this present disclosure.

[0097] Since Guidestar uniquely provides ground truth data of the true positive RNA callouts (true positive spots), Guidestar was used to estimate the precision and recall of the BFT approach (FIG. 2A), and varied the misidentification rate to determine the optimal threshold that balances precision and recall, using the Fl score as a summary statistic. It was found that at the conventional 0.05 misidentification rate, BFT has high precision but low recall, suggesting that the BFT approach at default settings is biased toward removing false positives, while filtering out a large proportion of true spots. The analysis in this cell line sample shows that a more optimal trade-off in precision and recall could likely be achieved at a higher misidentification rate of 0.2-0.4 (FIG. 2B), which maximizes the Fl score for the Guidestar genes. These parameter settings yield a substantially higher sensitivity, with 23% - 53% more RNA callouts relative to default settings. This analysis illustrates the value of Guidestar for quantifying both the sensitivity and accuracy of decoding methods, and its utility for optimizing hyperparameters for decoding.

[0098] Guidestar enables the design of binary classifiers for spot filtering

[0099] It was then sought to use the Guidestar-derived ground-truth to train a binary classifier (Random Forest [RF] model) for MERFISH spot filtering (FIG. 2C). The inventorsfirst set aside 5 out of 25 Fields of View (FOVs) as the test set. The test set or “test set of imaging spots” is obtained by a MERFISH method. A quality control check was performed on the 4 Guidestar genes (Acly, Gpam, Hnf4a and Ube2z) and 5-fold cross validation was used to optimize number of genes used (see Methods section Quality control and Guidestar gene selection) (FIG. 8). As the number of positive examples (colocalized) and negative examples (MERFISH only spots) differed, class balancing was also optimized (see Methods section Class balancing approach) with a combination of downsampling and augmentation with ‘blank’ codewords (FIGs. 9A and 9B). The features used for classification from the 3 conventional parameters: mean intensity of pixels within the spot, the hamming distance between the imaged codeword and closest theoretical barcode, and the spot size, were expanded to include an additional 12 intensity, shape, and distance-related features, then used sequential forward feature selection to obtain the reduced feature set (4 features, see Methods section Additional features and feature selection) that produced the best Fl score during cross validation (FIG. 10). Finally, hyperparameter optimization was conducted for the RF classifier (FIG. 11), and selected the model with the best mean Fl score across 5 folds. The model was then retrained with all 4 genes on the full training set and evaluated its performance on the held-out test set.

[0100] On the test set, using different FOVs as technical replicates, the RF model obtained a mean precision of 0.89 and recall of 0.97 while BFT (default settings) had a mean precision of 0.96 and recall of 0.76. The RF classifier had a significantly higher Fl score compared to BFT (0.93 for RF vs 0.85 for BFT; p = 3.68 x 10'5, two-tailed paired student’s t-test, FIG. 2E). To more comprehensively compare the performance of the classifier and BFT, the receiver operating characteristic (ROC) curves for both were computed and found that they had similar area under curve (AUC), with the classifier being higher (0.910 for classifier, 0.904 for BFT; p = 2.16 x 1 O'5, two-tailed paired student’s t-test) (FIG. 2D). This suggests that the performancefor BFT can be made much closer to that of the classifier by tuning the BFT parameters using Guidestar.

[0101] Since Guidestar yielded improved decoding results in the AML- 12 cell line sample, it was reasoned that it would be valuable for spot filtering in tissue samples, where the presence of extracellular matrix and more densely packed cells often result in higher background autofluorescence and nonspecific probe binding, leading to more challenging decoding relative to cell line samples. Using the same procedure described above for cell line samples, the Guidestar integrated probe library was applied to a mouse liver tissue sample. Acly, Gpam, Hnf4a, and Pigr were selected as Guidestars after quality control (FIG. 3).

[0102] Compared to the cell line dataset, there were fewer positive examples relative to blanks or MERFISH-only spots. The percentage ofMERFISH callouts that were colocalized to Guidestar ranged from 14% to 29%. In contrast, 4.3% of unmatched genes colocalized (FIGs. 5A and 5B). As there were more MERFISH-only spots than colocalized spots, the negative examples (MERFISH-only spots) were downsampled to match the number of colocalized spots to avoid class imbalance (FIGs. 9A and 9B). The classifier was then trained (using the optimized hyperparameters determined from the cell line sample) on the liver tissue dataset (training set = 40 FOVs, held out test set = 5 FOVs).

[0103] On the test set, using different FOVs as technical replicates, the RF model obtained a mean precision of 0.56, recall of 0.79, and Fl of 0.64 (FIG. 12A). This was a significant improvement over the default BFT method which, despite a slightly higher mean precision of 0.68, had a much lower recall of 0.36 (46% that of the RF model) and an Fl of 0.46 (p = 1 .38 x 10'3, two-tailed paired student’s t-test). Notably, Guidestar achieved a larger improvement of 40.2% in Fl score for the tissue sample compared to 9.22% on the cell sample. The receiver operator curves were computed again for both methods. Here, the classifier had, on average, a2.51% higher AUC compared to BFT (0.840 vs 0.820; p = 0.217, two-tailed paired student’s t- test) (FIG. 12B).

[0104] Next, the Guidestar-trained classifier’s ability was tested to generalize to unseen RNA species on both cell line and tissue samples. Fhe trained model was applied to the full probe library where the majority of genes were not imaged with corresponding Guidestar probes. For the cell line sample, the RF model resulted in a small but significant decrease in mean FPKM correlation across all FOVs (0.76 for RF model vs 0.82 for BFT, p = 2.26 x 10'22, two-tailed paired student’s t-test), but also yielded on average 71.4% higher spot callouts per FOV than BFT (FIG. 2F). In the tissue sample, it was also observed a decrease in mean FPKM correlation across all FOVs (0.55 for RF model vs 0.61 using BFT, p = 5.81 x 10'5, two-tailed paired student’s t-test). However, the RF model trained with Guidestar called out 3 5-fold more spots on average per FOV compared to BFT (FIG. 12C). This is consistent with the validation results as disclosed above on the guidestar genes, which showed that the RF classifier is less precise but yielded a substantial increase in sensitivity. These results suggest that the model, trained on a subset of genes, is able to generalize to unseen genes.

[0105] 3. Discussion

[0106] Localizing and decoding RNA species from raw images is a critical step for analyzing in-situ high-throughput transcriptomic data. However, methods for systematically validating this analysis step have been lacking. To address this gap, Guidestar, a platform for spike-in ground truth controls in MERFISH assays, has been developed to enable the validation and optimization of existing RNA decoding methodologies. It was demonstrated that Guidestar probes, when applied to a small subset of genes, yield high-quality images with clearly distinguishable RNA spots, and that a significant proportion of these spots colocalize with decoded MERFISH callouts. Importantly, despite imaging only a small subset of genes withGuidestar probes, classifiers trained on this data generalize to the remaining genes in the library, and can be leveraged to for decoding of the full MERFISH panel.

[0107] The primary advantage of Guidestar is its ability to directly validate both the precision and sensitivity of existing MERFISH decoding approaches. The common approach of utilizing the BFT to filter out false spots was evaluated. The sensitivity and precision analyses on a cell line dataset indicate that the default setting for BFT (0.05 misidentification rate) may lead to the removal of a substantial fraction of true spots, with only marginal improvements in accuracy relative to slightly higher settings for this parameter. This underscores the value of Guidestar in assessing the trade-off between precision and recall, allowing users to fine-tune decoding settings, whether prioritizing conservative results with low RNA count or maximizing RNA detection with slightly elevated error rates. Moreover, in the presence of sample-specific image quality fluctuations driven by tissue-specific fluorescent background and varying probe permeability and binding, Guidestar enables the important task of spike-in based parameter optimization. Thus, the inventors believed that Guidestar holds considerable promise for improving the reliability of in-situ RNA imaging across diverse tissue types.

[0108] Furthermore, the present study highlights the limitations of relying solely on the correlation to bulk FPKM values as a metric. The FPKM correlation metric is highly sensitive to precision, since incorrect RNA calls skew the relative abundances of detected genes. However, it does not account for detection sensitivity, since correlation values remain high even if only a small fraction of total detectable RNA molecules are called by the decoder. This may lead to biased decoding favoring precision over recall. To address this limitation, there is a critical need to supplement FPKM correlation with methods that estimate detection sensitivity. Guidestar may address at least partially the limitation as the method that provides a true positive reference from the same tissue slice, enabling accurate estimation of both recall and precision.

[0109] Beyond validation of existing methods, the ground truth reference produced by Guidestar can also be leveraged for developing improved decoding approaches, such as the classifier the inventors’ developed as a proof of concept. The present results demonstrate that the Guidestar-trained classifier significantly improved the Fl score compared to the BFT approach at default parameters, across both cell line and tissue samples. However, the ROC curves, while slightly higher for the classifier, revealed comparable performance between the two methods, underscoring the robustness of current decoding methods and importance of parameter tuning for optimal decoding outcomes. Notably, the trained models generalized to non-Guidestar genes, suggesting the potential of Guidestar for developing and testing more sophisticated spot detection approaches. The classifier used in this study is limited by its reliance on existing decoding pipelines that first identify candidate RNA spots which can then be classified as true RNA spots or false positives. A promising future direction would be to apply more sophisticated models to decode RNA spots directly from raw image stacks, such as with semi-supervised deep-learning, or by adapting deep-learning models used for srnFISH spot detection. In both cases, using Guidestar as ground truth for training or model refinement could greatly enhance accuracy while reducing the reliance on labor-intensive and subjective manual annotation.

[0110] Advantageously, Guidestar is a unique technique that combines spike-in methodology with machine-learning approaches to validate MERF1SH RNA decoding results directly. It is envisioned that Guidestar being utilized as a quality control and validation tool in multiplexed FISH workflows, enhancing the quality of downstream results derived from spatial omics assays.

[0111] Guidestar will also be useful in combinatorial multiplexed FISH (such as MERFISH, veraFISH, seqFISH, etc) and in situ sequencing (such as Xenium, Cartana, etc) assays as both a validation or quality control tool and a performance improvement tool. (1) Guidestar can beused to validate the quality of such assays by providing a direct comparison of RNA counts for a subset of RNA species, as is currently done using smFISH probes, but on adjacent tissue slices Unlike smFISH, Guidestar can also assess the quality of a particular assay by detecting the co-localization fraction between multiplexed FISH and Guidestar spots. (2) Guidestar can be used to improve the accuracy and sensitivity of current RNA decoding pipelines by providing additional ground-truth information on the features of True Positive callouts vs False Positive callouts, which are used to train a variety of machine learning classifiers. Alternatively, it could enable development of cheaper assays (using fewer probes or denser coding schemes) that maintain the output quality of current assays.

[0112] Although embodiments of the invention have been shown and described, the invention is not limited to the described embodiments. Instead, it would be appreciated by those skilled in the art that various modifications and variations can be made to the embodiments of the invention without departing from the scope of the invention, the scoop of which is set forth in the following claims.

Claims

CLAIMS1. A method of determining a spatial arrangement of a RNA species in a sample, comprising: providing a first set of encoding probes and a second set of encoding probes to hybridize with multiple genes of the RNA species, wherein the first set of encoding probes target a subset of the multiple genes, and the second set of encoding probes target each of the multiple genes; wherein for each gene in the subset of the multiple genes, the first set of encoding probes targets a set of first positions of the respective gene, and the second set of encoding probes targets a set of second positions of the respective gene, the set of second positions non-overlapping with the set of first positions; performing hybridization rounds of the first set of encoding probes and the second set of encoding probes separately; performing imaging rounds with a first set of readout probes to obtain a first set of RNA spots of the RNA species, wherein the first set of readout probes hybridizable with at least a part of the first set of encoding probes; performing imaging rounds with a second set of readout probes to obtain a second set of RNA spots of the RNA species, wherein the second set of readout probes hybridizable with at least a part of the second set of encoding probes; and colocalizing the first set of RNA spots with the second set of RNA spots to obtain a set of colocalized spots corresponding to a ground truth reference data of the spatial arrangement of the RNA species in the sample.

2. The method according to claim 1, further comprising: classifying pixels of the set of colocalized spots as true positive spots; and obtaining the ground truth reference data of the spatial arrangement of the RNA species using the true positive spots.

3. The method according to any one of claims 1 to 2, wherein the second set of encoding probes corresponds to a set of MERFISH probes.

4. The method according to any one of claims 1 to 3, wherein the set of first positions and the set of second positions are interspersed across a plurality of regions of the respective gene.

5. The method according to any one of claims 1 to 4, wherein each of the first set of readout probes has a higher number of fluorophores than each of the second set of readout probes.

6. The method according to any one of claims 1 to 5, further comprising: independently performing a respective imaging round for each gene in the subset of the multiple genes, to obtain the first set of RNA spots of the RNA species.

7. The method according to any one of claims 1 to 6, wherein each of the first set of encoding probes and each of the second set of encoding probes is a DNA probe.

8. The method according to any one of claims 1 to 7, wherein the subset of the multiple genes comprises genes with moderate to high expression levels.

9. The method according to any one of claims 1 to 8, further comprising: colocalizing the first set of RNA spots with the second set of RNA spots using a maximum euclidean distance threshold.

10. The method according to any one of claims 1 to 9, further comprising: performing imaging rounds with the first set of readout probes to obtain a first set of imaging spots; and identifying local maxima from the first set of imaging spots to obtain the first set of RNA spots.1 1 . The method according to any one of claims 1 to 10, further comprising: performing imaging rounds with the second set of readout probes to obtain a second set of imaging spots; and down-sampling a subset of the second set of imaging spots that do not colocalize with the first set of imaging spots to obtain a set of down-sampled second imaging spots, wherein the set of down-sampled second imaging spots matches the number of spots in the set of colocalized spots.

12. The method according to claim 11, further comprising: training a machine learning model using the set of colocalized spots as positive examples for the set of down- sampled second imaging spots.

13. The method according to claim 12, wherein the machine learning model is a classifierbased model.

14. The method according to claim 13, wherein the classifier-based model is one of: a random forest classifier, a decision tree model and a support vector machine.

15. A method of determining a misidentification rate for a MERFISH imaging method, comprising: determining a fraction of incorrectly detected negative spots within a second set of imaging spots using the set of colocalized spots according to any one of claims 1 to 14 to determine the misidentification rate.

16. The method according to claim 15, wherein the misidentification rate is in a range of 0.01 to 0.6.

17. A gene prediction method, the gene prediction method comprising: inputting a test set of imaging spots of a test RNA species into a machine learning model, the machine learning model trained based on the ground truth reference data according to any one of claims 1 to 14; and predicting a spatial arrangement of the test RNA species based on an output of the machine learning model18. The gene prediction method according to claim 17, wherein the test set of imaging spots is obtained by a MERFTSH method19. The gene prediction method according to any one of claims 17 to 18, further comprising: classifying a selected subset of the test set of imaging spots as true test positive spots using the machine learning model; andpredicting a spatial arrangement of the test RNA species based on the true test positive spots.

20. The gene prediction method according to any one of claims 17 to 19, wherein the machine learning model is a classifier-based model selected from one of: a random forest classifier, a decision tree model and a support vector machine21. A gene prediction module for predicting a spatial arrangement of a test RNA species, the gene prediction module comprising: a processing unit; and a non-transitory media readable by the processing unit, the media storing instructions that when executed by the processing unit, causes the processing unit to: input a test set of imaging spots of the test RNA species into a machine learning model, the machine learning model trained based on the ground truth reference data according to any one of claims 1 to 14; and predicting a spatial arrangement of the test RNA species based on an output of the machine learning model.

22. The gene prediction module according to claim 21, wherein the test set of imaging spots is obtained by a MERFISH method.

23. The gene prediction module according to any one of claims 21 to 22, wherein the processing unit is further configured to: classify a selected subset of the test set of imaging spots as true test positive spots using the machine learning model; andpredict a spatial arrangement of the test RNA species based on the true test positive spots.

24. The gene prediction module according to any one of claims 21 to 23, wherein the machine learning model is a classifier-based model selected from one of: a random forest classifier, a decision tree model and a support vector machine25. A MERFISH imaging method, comprising: determining a misidentification rate for analysing a set of MERFISH imaging spots using the set of colocalized spots obtained from the method according to any one of claims 1 to 14; and using the misidentification rate for the MERFISH imaging method.

26. The method according to claim 25, wherein the misidentification rate is in a range of0.01 to 0.6.