Merfish-plus: a versatile spatial genomics technology for high-throughput high-resolution tissue imaging

MERFISH+ addresses limitations in hybridization-based spatial transcriptomics by using acrydite-modified probes and a new microscope-microfluidics system for high-throughput 3D imaging, achieving enhanced stability and scalability in imaging large tissue volumes with improved resolution and gene library flexibility.

WO2026155815A1PCT designated stage Publication Date: 2026-07-23RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
RGT UNIV OF CALIFORNIA
Filing Date
2025-11-18
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Current hybridization-based spatial transcriptomics methods face limitations such as a narrow target range, low throughput, complex multistep staining protocols, and limited stability of hybridized probes, which restrict high-throughput single-molecule imaging across large tissue volumes, especially in human samples.

Method used

The MERFISH+ protocol incorporates acrydite-modified probes anchored to a biogel film, enabling extended imaging sessions, flexible barcoding for larger gene libraries, and a new microscope-microfluidics system for high-throughput 3D imaging, allowing imaging of up to 1800 genes and 10-fold area increase.

Benefits of technology

This approach enhances stability, flexibility, and scalability, enabling 3D reconstruction of complex tissues like the human fetal heart with sub-100 pm resolution, revealing intricate structures and cell type compositions, and facilitates simultaneous imaging of spatial transcriptomes, nascent RNA, and histone/cell-type markers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025056012_23072026_PF_FP_ABST
    Figure US2025056012_23072026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure provides a high throughput, barcode-expandable, and flexible fluorescence in situ hybridization (FISH) protocol with improvements in the stability of samples, allowing for extended imaging sessions lasting over three months. The protocol further allows concurrent detection of thousands of target molecules including RNA, DNA and protein in two and three dimensional spatial confirmations in cells and tissues.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Dkt. No.: 114198-3660MERFISH-PLUS: A VERSATILE SPATIAL GENOMICS TECHNOLOGY FOR HIGH-THROUGHPUT HIGH-RESOLUTION TISSUE IMAGING CROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 746,913 filed January 18, 2025, and U.S. Provisional Application No. 63 / 904,501 filed October 23, 2025, which are incorporated herein by reference in their entirety.STATEMENT OF GOVERNMENT SUPPORT

[0002] This invention was made with government support under OD031878 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND

[0003] Hybridization based spatial transcriptomics and genomics technologies are updating the understanding of the cellular and subcellular organization of complex tissues across many fields from plant biology to cancer biology and neuroscience. Despite their high sensitivity, nanoscale spatial resolution, and multimodal imaging capabilities for RNA, DNA, and protein targets, current methods face significant limitations. These include a narrow target range (typically up to 500 genes), low throughput (limited to one or two tissue sections per experiment), and a complex multistep staining protocol for multimodal imaging. These challenges stem from the limited stability of hybridized probes during sequential hybridization cycles and the demanding engineering requirements for high-throughput single-molecule imaging across large tissue volumes, which can be centimeter-scale large tissue area spanning the intact three-dimensional organs.SUMMARY OF THE DISCLOSURE

[0004] Applicant presents herein MERFISH-PLUS (also identified herein as MERFISH+™ herein), a high throughput, barcode-expandable, and flexible fluorescence in situ hybridization (FISH) protocol designed to overcome these limitations. In some embodiments, the method is coupled with high throughput microscopy and fluidics. MERFISH+™ introduces significant improvements in the stability of samples, allowing for extended imaging sessions lasting over three months. The increased stability allowed for a barcoding strategy facilitating the selection of any subset of genes for imaging or expanding to larger gene libraries (>1000 genes) in aAtty. Dkt. No.: 114198-3660gradual way. In some embodiments, the larger gene libraries comprise at least 1800 genes in a modular way. MERFISH+™ was applied to human fetal heart tissue where imaging 1000 genes revealed new cell states and new genes with varying expression in space. In some embodiments, 1800 genes are imaged. MERFISH+™ high-throughput capabilities, with 10 times larger imaginable area than previous methods, allowed for the 3D reconstruction of entire human fetal hearts at sub- 100 pm resolution in the z-dimension. In some embodiments, the resolution in the z-dimension is 100 pm. In some embodiments, the resolution in the z-dimension is 300 nm. These 3D images revealed the intricate structures and cell type composition of the developing heart, including the aorta, valves, vasculature structures, and boundary structures across all four chambers. Furthermore, MERFISH+ enables the simultaneous imaging of spatial transcriptomes, nascent RNA, single chromatin traces, and histone / cell-type markers from a single tissue slice. The technology represents a large-format platform for spatial multi-omics that enables high resolution mapping of gene expression at subcellular resolution and the characterization of cellular organization within 3D organs. This allows for the visualization of genome-scale dynamic chromatin structure and their associated epigenomic protein marks across more than 10 different cell types. MERFISH+™ represents a significant advance in multiplexed spatial genomics technology, offering enhanced stability, flexibility, and scalability. Its broad applicability and cost-effectiveness promise to accelerate spatial genomics research across diverse biological fields, providing unprecedented insights into complex tissue structures and gene regulation dynamics.

[0005] In another aspect, the chemistry applied to MERFISH method as described herein allows the use of microscopy and microfluidics for 3D whole organ reconstruction. The methods described herein provide significant improvements in the stability of target tissue samples, enabling extended imaging sessions lasting over three months. The disclosed barcoding scheme facilitates flexible encoding for imaging various cell types and modalities within a single sample.

[0006] Single-cell RNA sequencing (scRNA-seq) has revolutionized the understanding of cellular diversity and gene expression dynamics within complex tissues with a wide range of applications across many branches of biology ranging from development to microbial communities. However, a significant limitation of scRNA-seq is that the process of dissociating tissues into single cells disrupts their spatial context. Spatial transcriptomicsAtty. Dkt. No.: 114198-3660technologies have made significant strides in addressing these limitations. By enabling the mapping of gene expression and cell types directly within intact tissues, spatial transcriptomics preserves the crucial spatial information that is lost in scRNA-seq.

[0007] Among the various spatial transcriptomics methods, hybridization-based approaches stand out for their high sensitivity and precision. These methods, including Multiplex Error-Robust Fluorescence in-situ Hybridization (MERFISH) and Seq-FISH (Chen et al. 2015; Xia et al. 2019) utilize fluorescently labeled oligonucleotide probes that bind to specific RNA species, imaged over multiple hybridization cycles. Each target RNA species is given a unique barcode of zeros and ones which is read out by a high-resolution fluorescence microscope. If fluorescence signal is present at a single molecule location in the sample during a cycle of imaging, a 1 is recorded for the corresponding bit position; if no fluorescence is present, a 0 is recorded. In this way, the multiple cycle of imaging come together to generate the barcodes of the RNA species, allowing for the spatial localization of many genes. These techniques can detect up to 80% of single mRNA molecules with tens of nanometer precision within cells. This high sensitivity and resolution are essential for understanding the cellular and subcellular spatial organization and interaction of cells within their native tissue environments.

[0008] Hybridization-based spatial transcriptomics are technologies that still face several critical limitations. First, most studies focus on a limited subset of genes, typically a few hundred. As a result, the full complexity of gene expression within tissues may not be captured, potentially overlooking important genes and pathways involved in biological processes and diseases (Chen et al. 2015; Moffitt, Hao, Wang, et al. 2016; Moffitt et al. 2018; Xia et al. 2019). Secondly, most studies imaged gene expression in two-dimensional (2D) tissue sections. While this provides valuable spatial information within the plane of the section, it does not capture the three-dimensional (3D) organization of tissues (Moffitt, Hao, Wang, et al. 2016; Moffitt et al. 2018; M. Zhang et al. 2023; Farah et al. 2024; Lohoff et al. 2022; Liu et al. 2020; Eng et al. 2019). Many biological processes and cellular interactions occur in three dimensions, especially during a variety of developmental stages in human organogenesis. Therefore, the choice of sectioning plane can significantly influence the observed gene expression patterns, potentially leading to biased or incomplete interpretations.

[0009] Thirdly, spatial transcriptomics techniques can require expensive equipment, reagents, making them difficult to access and perform reliably by many research laboratories. The highAtty. Dkt. No.: 114198-3660cost and limited reliability of each spatial transcriptomics experiment limits large-scale projects involving for instance whole organ reconstructions in human samples. Human intact organs are often hard to obtain compared to animal models bred and sacrificed in controlled laboratory settings. In some embodiments, the intact organs are not presenting with the disease or otherwise disease-free. Factors such as post-mortem interval, the condition of the tissue at the time of collection, and handling procedures can affect RNA integrity and overall sample quality, potentially impacting the results of spatial transcriptomics assays.

[0010] Recently, Applicant performed MERFISH during human embryonic heart development (Farah et al. 2024) revealing the spatial organization of single cells into cellular communities that form distinct cardiac structures. However, this data was limited to only 4 human donors with 3 to 42D tissue sections for each donor and covered only 238 genes due to the following challenges: 1) RNA integrity degraded over prolonged imaging periods (>2 days) limiting both the tissue area to only -2 fetal heart sections and number of genes imaged per experiment to -200 genes; 2) the genes selected for imaging could not be easily expanded due to the inflexible design of the original MERFISH library of probes ordered; 3) any error with the equipment (i.e. loss of focus during imaging, valve failure in the fluidics system etc.) or reagents such as imaging buffer or TCEP-based cleavage buffer oxidizing or changing acidity caused the failure of the entire experiment that cannot be repeated or rescued.

[0011] To overcome these limitations, particularly for valuable human tissues, Applicant extended the MERFISH technology with a method called MERFISH+™. This approach incorporates an acrydite group into the encoding / primary probes, covalently anchoring them to a thin (20 pm) biogel film covering the tissues. In some embodiments, the thickness of the biogel film is -20 pm -200 pm. In some embodiments, the biogel film embeds the tissues. These anchored probes are then read out and imaged over an extended period, spanning three or more months without obvious signal decay. In some embodiments, the extended period spanning one week to twelve months. This enhanced stability further enabled 1) a new flexible bar-coding scheme allowing for imaging more than 1000 or 1800 genes with a similar quality to the 300-scale; 2) a 10-fold increase in imaging area allowing 20-30 sections per imaging experiment compared to MERFISH technology, and 3) stable multi-modal imaging of RNA, DNA, and proteins within the same tissue sample. (FIG.l). In some embodiments, the sections are consecutive.Atty. Dkt. No.: 114198-3660[00121 The chemistry applied to the MERFISH probes design as described herein can facilitate developing a new microscope-microfluidics system. In some embodiments, the system is built on top of a flexible Applied Scientific Instrumentation microscope body equipped with higher-power lasers from Lumencor (up to -1-2.5W per laser line). In some embodiments, the microscope body is coupled with 0.4mm square multimode fiber. In some embodiments, the microscope body is coupled with a larger-format camera (20.8 mm sensor) (Teledyne-Kinetix). In some embodiments, the microscope body is coupled with an upgraded microfluidics chamber (Bioptechs). In some embodiments, the system achieves a 10-fold increase in imaging area compared to previous and commercially available MERFISH imaging platforms. In some embodiments, the increasement is at least 2-fold, at least 3-fold, at least 5-fold, at least 7-fold, at least 9-fold, at least 10-fold, at least 12-fold, at least 15-fold, at least 18-fold, or at least 20-fold. Table 7 provides a complete list of parts to assemble this high-throughput microscope-microfluidics system. The microscope set-up described herein can enable imaging large-format tissues sections leading to 3D organ reconstruction.

[0013] An aspect of the disclosure is directed to a method to provide imaged codewords corresponding to distinct target molecules in a spatial organization in a sample comprising:a) contacting a sample comprising a plurality of target molecules and one or more primary nucleic acid probe pools, wherein the primary nucleic acid probes in each of the pools comprises an acrydite group and is bound to a biogel film covering a surface of the sample, andwherein each pool of primary nucleic acid probes is selected to specifically hybridize to a distinct target molecule in the sample, andwherein each primary nucleic acid probe comprises a target sequence and one or more read sequences, andwherein each primary nucleic acid probe produces a bound primary nucleic acid probe upon contact with its target molecule, andwherein each pool of primary nucleic acid probes encodes an N-bit codeword with a Hamming weight of at least 2 that was assigned to each distinct target molecule,wherein each assigned N-bit codeword is a valid codeword with a Hamming distance equal to or greater than 2 between valid codewords and wherein each of the one orAtty. Dkt. No.: 114198-3660more read sequences correspond to a bit value of 1 for the codewords assigned to the target molecules,b) contacting the one or more plurality of primary nucleic acid probes of the pools with a plurality of fluorescently labeled readout probes that can specifically bind to the one or more read sequences of the primary nucleic acid probes to generate detectable signal at the target molecule location,c) imaging the readout probes bound to the primary nucleic acid probes; and d) repeating steps b) and c) in or more sequential binding and imaging rounds until all bit positions of the N-bit codeword have been imaged providing imaged codewords corresponding to each distinct molecule in a spatial organization.

[0014] An aspect of the disclosure is directed to an improved method for imaging a target molecule spatial organization in a sample comprising, or alternatively consisting essentially of, or consisting of, or yet further consisting of:a) contacting a sample comprising, or alternatively consisting essentially of, or consisting of, or yet further consisting of a plurality of target molecules and one or more primary nucleic acid probe pools,wherein each primary nucleic acid probe comprises, or alternatively consists essentially of, or yet further consists of an acrydite group and is bound to a biogel film covering a surface of the sample, andwherein each pool of primary nucleic acid probes is selected to specifically hybridize to a distinct target molecule in the sample, and wherein each primary nucleic acid probe comprises, or alternatively consists essentially of, or yet further consists of a target sequence and one or more read sequences, andwherein each primary nucleic acid probe produces a bound primary nucleic acid probe upon contact with its target molecule, andwherein each pool of primary nucleic acid probes encodes an N-bit codeword with a Hamming weight of at least 2 that was assigned to each distinct target molecule,wherein each assigned N-bit codeword is a valid codeword with a Hamming distance equal to or greater than 2 between valid codewords and wherein each of the read sequences correspond to a bit value of 1 for the codewords assigned to the target molecules,Atty. Dkt. No.: 114198-3660b) contacting the one or more plurality of primary nucleic acid probe pools with a plurality of fluorescently labeled readout probes to specifically bind to read sequences of the primary nucleic acid probes,c) imaging the readout probes bound to the primary nucleic acid probes; and d) repeating steps b) and c) in or more sequential binding and imaging rounds until all N positions in the N-bit codeword have been imaged providing an imaged codeword corresponding to each distinct molecule in a spatial organization.

[0015] In some embodiments, the biogel film embeds the sample.

[0016] In some embodiments, step b) comprises removal of signal. In some embodiments, steps c) comprises removal of signal.

[0017] In some embodiments, the target molecule is selected from a nucleic acid or a polypeptide.

[0018] In some embodiments, the nucleic acid is a DNA or an RNA molecule.

[0019] In some embodiments, steps c) and d) are repeated over an extended time period of at least 1 month, or at least 2 months, or at least 3 months.

[0020] In some embodiments, the acrydite group is derived from a compound of formula (HO)2P(O)(CH2)xNHC(O)C(CH2)CH3, wherein x is 1-10, and integers therebetween. In some embodiments, x is from 2 to 8, from 3-9, from 4-7, from 5-8, or from 6-8.

[0021] In some embodiments, x is 7.

[0022] In some embodiments, the acrydite group is attached to a 5’ end of each of the plurality of primary nucleic acid probes.

[0023] In some embodiments, the assigned N-bit codewords have a Hamming distance equal to or greater than 4 between each of the valid codewords.

[0024] In some embodiments, the N-bit codeword has a Hamming weight of 4.

[0025] In some embodiments, the imaged codeword is matched to a valid codeword assigned to a distinct RNA species.Atty. Dkt. No.: 114198-3660[00261 In some embodiments, the Hamming distance is 4 or greater and the imaged codewords are matched to valid codewords or discarded.

[0027] In some embodiments, the N-bit codeword has a Hamming weight of 4 and the Hamming distance is 4 or greater and the imaged codewords are matched to valid codewords or discarded.

[0028] In some embodiments, the target molecule is an RNA transcript.

[0029] In some embodiments, the method comprises, or alternatively consists essentially of, or yet further consists of determining the spatial organization of the transcriptome from a single cell.

[0030] In some embodiments, the method determines the spatial organization of the proteome from a single cell.

[0031] In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of at least 10 different primary nucleic acid probes.

[0032] In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of at least four distinct read sequences. In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of four distinct read sequences. In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of at least eight distinct read sequences. The organization of the read sequence provided herein can allow for flexible imaging of molecules either in a single-molecule sequential mode or MERFISH combinatorial mode. The flexible imaging may be an important improvement over the traditional technology because single-molecule sequential imaging serves as a quality control matrix for evaluating the performance of MERFISH+.

[0033] In some embodiments, the primary nucleic acid probes comprise, or alternatively consist essentially of, or yet further consist of a target sequence, with an average length of between about 10 and 200 nucleotides, that hybridize the distinct target molecule. In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of at least about 10 different target sequences.Atty. Dkt. No.: 114198-3660[00341 In some embodiments, after each hybridization and imaging round the fluorescent readout probe is quenched to inactivate, wherein after each hybridization and imaging round the fluorescent read out probe is inactivated by chemically or enzymatically cleaving the fluorescent label from the readout probe.

[0035] In some embodiments, the N-bit binary code comprises, or alternatively consists essentially of, or yet further consists of at least a 16-bit code.

[0036] In some embodiments, the plurality of fluorescent readout probes comprises, or alternatively consists essentially of, or yet further consists of at least one distinct fluorescent label. In some embodiments, the plurality of fluorescent readout probes comprises, or alternatively consists essentially of, or yet further consists of at least two distinct fluorescent labels. In some embodiments, the plurality of fluorescent readout probes comprises, or alternatively consist essentially of, or yet further consist of at least three distinct fluorescent labels. In some embodiments, the plurality of fluorescent readout probes comprises, or alternatively consist essentially of, or yet further consist of at least three distinct fluorescent labels.

[0037] In some embodiments, the spatial organization of the target molecule is imaged in 2 dimensions. In some embodiments, the spatial organization of the target molecule is imaged in 3 dimensions.

[0038] In some embodiments, the method comprises, or alternatively consists essentially of, or yet further consists of determining abundance of the target molecule.BRIEF DESCRIPTION OF THE DRAWINGS

[0039] FIGS. 1A - IB: General outline of improvements and outline of improvements. (FIG.1A) Workflow before Applicant’s invention. (FIG. IB) Workflow after Applicant’s improvements-improved stability from acrydite modification to probes, increased adaptability from new codebook, larger scale due to re-engineered microfluidics, and the increased flexibility of multimodality. Applicant has three series of experiments from a total of fifteen experiments, each series consisting of an acrydite probe experiment and a regular probe experiment, as well as multiple reimages of the acrydite probe experiment (FIG. 8, Table 1). The experiment protocol was to image one section of the heart with regular probes on one imaging system (Kiwi), while simultaneously imaging an adjacent heart section with acryditeAtty. Dkt. No.: 114198-3660modified probes on a second system (Lemon). This would disclose the difference between acrydite and regular oligonucleotide probes with the least amount of extrinsic error. Next, the acrydite sample would be stripped with 100% formamide, and then reimaged on Lemon to test how efficient the acrydite probes are to represent the RNA targets themselves. It is expected that this second imaging of the acrydite sample will have around 60-70% detection efficiency as compared to the first image of the acrydite sample.

[0040] FIGS. 2A - 2F: (FIG. 2A), Left, Original MERFISH encoding probe binds only to mRNA and mRNA’s polyA tail binds the polyacrylamide gel matrix; however, as mRNA degrades over time, probes become unstable and MERFISH imaging can no longer be performed. Right, Acrydite-conjugation of MERFISH probes allows probes to also stably bind directly to the polyacrylamide gel matrix so as mRNA degrades over time, the probes remain stable and MERFISH imaging can continue to be conducted. (FIG. 2B), After imaging 53 hybridization rounds, brightness of MYH7 plateaus at 55% of hybridization round l’s brightness. (FIG.2C), Identified MERFISH cells spatially mapped for experiments performed with original MERFISH probes (left), stable MERFISH using acrydite-conjugated probes (middle), and reimaging of stable MERFISH (right). (FIG.2D), UMAPs for these experiments colored by identified cell classes. (FIG. 2E), Pearson correlation between log scale transcript counts per gene for stable MERFISH and original MERFISH experiments is high (0.959; top), and for reimaging of stable MERFISH and original MERFISH experiments is also high (0.937; bottom). (FIG. 2F), Mean expression of marker genes across cell types have similar patterns across all three experiments.

[0041] FIGS. 3A - 3H: Flexibility - changing codebook, on-bits (3 or 4) (FIG. 3A), Left, original MERFISH probe schematic with example single round of MERFISH shows that each read sequence targets multiple genes. Right, MERFISH+ probe schematic shows that each read sequence is unique to each gene, so MERFISH+ probes can be used for smFISH or MERFISH. (FIG. 3B), Spatial heatmap of exemplar gene IRX3 from smFISH using MERFISH+ probes. (FIG. 3C), All 15 genes’ smFISH+ spatial plots. (FIG. 3D), Sample MERFISH codebook. (FIG. 3E), Implementation of codebook with MERFISH+ protocol using custom pipetting robot to pipette MERFISH+ hybridization tubes with adaptors corresponding to codebook. (FIG. 3F), Hybridization tubes are used to hybridize sample which is then imaged. (FIG. 3G), The same sample is imaged with two different codebooks.Atty. Dkt. No.: 114198-3660(FIG. 3H), Pearson correlation between mean transcripts per cell for two different codebooks shown in g is high (0.815).

[0042] FIGS. 4A - 4H: Gene expansion to over 2000 is easy. (FIG. 4A): UMAP and spatial plot of MERFISH+ experiment with the 2000 gene panel. (FIG. 4B): Correlation matrix between 2000-gene MERFISH+ experiment cell-type clusters and sequencing data clusters, c, Spatial expression pattern of vCM-LV-papillary cluster. (FIG. 4D): Volcano plot of X: log2 of fold change; Y: loglO of p value. (FIG. 4E): Spatial expression pattern for CNTNAP5, HCN4 (FIG. 4F): Spatial expression pattern for COL11A1, IGFBP4. (FIG. 4G): Gene expression correlation along apex-base axis. (FIG. 4H): Gene expression correlation along left-right axis.

[0043] FIGS. 5A - 51: MERFISH+™ scales up imaging throughput, enabling 3D reconstruction of larger samples. (FIG. 5A). Schematic of MERFISH+ hardware. 4-fold increase in throughput over older microscopy systems capable of imaging twice the area of the prior model of cameras. (FIG. 5B). New fluidics chamber with 6 cm x 4 cm=24 cm2area, allowing for 13.5 cm2imaging area (left), compared to previous 12.56 cm2area with 1.25 cm2imaging area (right). (FIG. 5C). Human heart donor section map. The atrial and ventricular halves were sliced into 21 and 32 sections, respectively, with a mean thickness of 16 um, spanning a total length of 7 mm. (FIG. 5D). MERFISH+™ spatial maps for all heart sections (left) and corresponding UMAP for entire heart (top-right) with cell-type labels (bottom-right).(FIG. 5E), the Umap of the 3D MERFISH detected cell types; (FIG. 5F), map the published MERFISH data on the 3D MERFISH umap data. (FIG. 5G), table / Pie Chart to show the cells number of each cluster; (FIG. 5H), the diagram of 3D construction, (FIG. 51), show the heart sagittal section.

[0044] FIGS. 6A - 6H: 3D reconstruction of the human fetal heart, illustrating the spatial distribution of various cell types throughout the entire heart. (FIG. 6A) Top: Diagram of valve system of human developing heart consists of tricuspid valve (TV), bicuspid (mitral valve MV), aortic valve (AV) and pulmonary valve (PV). Bottom left: 3D spatial map of valve system; right: separate structure of TV, PV, MV and AV. (FIG. 6B) Spatial map of the valve system on slices 10 to 20 along the apex-base axis; (FIG. 6C) Marker genes for sub-cell types of VIC.(FIG. 6D) Neighborhood analyses of 4 sub-structures of VIC; (FIG. 6E) Left: Diagram of coronary arteries system. Right: spatial map of AO, PA, left anterior artery and right coronaryAtty. Dkt. No.: 114198-3660artery. (FIG. 6F) dot plot of marker genes of arteries. (FIG. 6G) Left: spatial map of left anterior artery, Right: neighborhood analyses of left anterior artery; (FIG. 6H) gradient expression quantification of marker genes in the arteries.

[0045] FIGS. 7A - 7K: Flexibility of multimodality: Differential transcription and 3D structure of the MYH6 / MYH7 locus in developing human hearts in vivo. (FIG. 7A) MYH6 / 7 transcripts / cell. (FIG. 7B). RNA MERFISH of 238 genes across a section of developing human hearts (12pcw) spatially identified 27 cell types (consistent with Farah, et al., Nature, 2024). (FIG. 7C). Number of mRNA transcripts per cell for MYH6 and MYH 7 for the same section in FIG. 7B. (FIG. 7D). Left panel: Image of nascent RNA for MYH 6 (red), MYH7 (green) across the entire heart section. Right panel: Two representative magnified images from the atria (#1) and the ventricle (#2). Nuclei labeled with DAPI in blue. (FIG. 7E).Quantification of nascent transcription rate across specific cardiomyocyte cell types defined in FIG. 7B. Nascent MYH6 (red), nascent MYH7 (green), both (orange). (FIG. 7F). Schematic and example images of chromatin tracing at 2-kb resolution for the 84-kb MYH6 / MYH7 locus imaged across 42 cycles in single cells. (FIG. 7G). Representative single chromatin traces of the MYH6 / MYH7 locus for chromosomes with nascent transcription of only MYH6 (marked as MYH6+, MYH7- only MYH7 (MYH7+, MYH 6-) or with no nascent transcription of either MYH6 or MYH7. (MYH7-, MYH6-) The loci spanning MYH6 mA MYH 7 are marked in red and green, respectively. The MYH 6 and MYH7 promoters are marked in orange and cyan respectively. Loci containing cis-regulatory regions CRE1 and CRE2 are marked in yellow.(FIG. 7H) Box plots quantifying the compaction of MYH6 (red) and MYH7 (green) gene loci based on the average interdistance among the corresponding loci of the following three classes of chromosomes: MYH6+MYH7-, MYH7+MYH6- and MYH7-MYH6- no transcription. The median, first and third quartile are marked. (FIG. 71). Median distance matrix across the MYH6 / MYH7 locus for the three classes of chromosomes in FIG. 7H. Annotated are MYH6 and MYH7 genes in red and green respectively and CRE1, CRE2 in yellow. (FIG. 7 J). The difference of median distance matrices for the chromosomes with MYH6 or MYH7 transcription only versus no nascent transcription. (FIG. 7K). Box plots quantifying the contact frequency c£MYH6 (red) and MYH7 (green) promoters with CRE1 and CRE2 loci in chromosomes with and without nascent transcription of MYH6 or MYH7.Atty. Dkt. No.: 114198-3660

[0046] FIGS. 8A – 8F: (FIG. 8A), Pearson correlation of mean transcripts per cell between original MERFISH and smFISH experiment (FIG. 3B, FIG. 3C) shows high correlation (0.789). (FIG. 8B), Comparison of spatial plots of five marker genes for original MERFISH and smFISH experiment shows similar spatial expression patterns. (FIG. 8C), (FIG. 8D), Correlation matrix between cell types of MERFISH experiments using MERFISH+™ probes and original MERFISH. (FIG. 8E), (FIG. 8F), Pearson correlation of average counts per cell between MERFISH experiment using MERFISH+ probes and original MERFISH (0.576, 0.535).

[0047] FIGS.9A - 9B: (FIG.9A), Pearson correlation between MERFISH experiments using MERFISH+™ probes and original MERFISH (0.58, 0.54, 0.64). (FIG. 9B), Pearson correlation between MERFISH experiments using MERFISH+™ probes (0.77, 0.64, 0.81).

[0048] FIGS. 10A – 10D: (FIG. 10A), Dot plot showing gene expression in the three BEC subtypes (BEC-Venous, BEC, and BEC-Arterial). (FIG. 10B), Spatial expression pattern for the BEC subtypes. (FIG. 10C), Spatial expression pattern for the EPDC subtypes (aEPDC-LA, aEPDC-RA, and EPDC). (FIG. 10D), Spatial expression pattern for AVN-, AVC-, and IFT-like clusters.

[0049] FIGS. 11A – 11E: (FIG. 11A), Spatial plot of all 44 cell types from 53 sections of 12-PCW human heart. (FIG. 11B), Plot of (FIG. 11C), Correlation matrix between 3D heart and published MERFISH clusters. (FIG. 11D), Sub-clustering of ventricular cardiomyocytes.(FIG. 11E), Development of ventricular wall shown in consecutive sections of heart.

[0050] FIGS. 12A - 12B: (FIG. 12A), Pearson correlation between 3D heart and published MERFISH data and scRNA-seq data. (FIG. 12B), Cell type correlation matrix with published MERFISH data (left) and bar chart (right).

[0051] FIGS. 13A - 13F: MERFISH+ utilizes acrydite probes for increased stability. FIG.13 A: Schematic comparing original and acrydite-conjugated MERFISH probe designs. Left: Original encoding probes (green) bind to mRNA (purple), which anchors to the polyacrylamide gel via its poly A tail. As mRNA degrades over time, probe binding becomes unstable, limiting reimaging. Right: Acrydite-conjugated probes covalently bind directly to the gel matrix, enabling probe retention and continued imaging even after mRNA degradation. (Created with BioRender. Eschbach, J. (2025) https: / / BioRender.com / 9xr4cd9). FIG. 13B: Signal stabilityAtty. Dkt. No.: 114198-3660of MYH7 across 54 rounds of hybridization. Fluorescence intensity plateaus at -55% of the initial signal observed in round 1. FIG. 13C: Spatial maps of identified MERFISH cells from three conditions: original probe design (left), stable MERFISH with acrydite-conjugated probes (middle), and reimaging of the stable MERFISH sample (right). FIG. 13D: UMAP representations of cells from each experiment, colored by transcriptionally defined cell classes.FIG. 13E: Pearson correlations of log-transformed gene counts per cell between original and stable MERFISH (top, R = 0.959), and between original MERFISH and reimaging of the stable sample (bottom, R = 0.937), showing high transcriptomic concordance. FIG. 13F: Marker gene expression patterns across cell types remain consistent across all three experimental conditions.

[0052] FIG. 14A – 14H: MERFISH gene expansion to -2000 genes. FIG. 14A: UMAP and spatial visualization of cell clusters identified in a 12 p.c.w. human heart using the MERFISH+ methodology with an 1,863-gene panel. Newly identified cell populations and subclusters are highlighted by boxes, including a previously uncharacterized cardiomyocyte subtype (vCM-LV-Papillary), three subclusters of blood endothelial cells (BEC I, II, III), and two distinct epicardial populations (aEpicardial and vEpicardial). FIG. 14B: Correlation matrix between 1,863-gene MERFISH+ experiment cell-type clusters and sequencing data clusters. FIG. 14C:Left: Spatial expression pattern of VIC (magenta), vCM-RV-AV (teal), and vCM-LV-Papillary clusters (yellow); Right: spatial expression pattern of DLGAP1, ACTA1. FIG. 14D: Volcano plot of differential gene expression between vCM-LV-Papillary and vCM-RV-AV cells. Positive fold change indicates higher expression in vCM-LV-Papillary. X: log2 fold change; Y: –log10 p-value. FIG. 14E: Correlation of gene expression along the left and right axis of the atrial CMs. FIG. 14F: Spatial expression pattern for PITX2 (canonical LA marker), AKAP6, and KIF26B. FIG. 14G: Correlation of gene expression along the left-right axis of the ventricular CMs. FIG. 14H: Spatial expression pattern for HAND1 (canonical LV marker), COL11A1, and WNT9A.

[0053] FIG. 15A – 15J: Flexibility of multimodality imaging. FIG. 15A: Multimodal MERFISH-Plus imaging of RNA MERFISH, nascent RNA, chromatin tracing and epigenetic protein marks. FIG. 15B: imaging FXYD loci on the chromosome 19 with -50 nm spatial resolution and 10 kb genomic resolution. FIG. 15C: MERFISH single-cell expression was clustered into 27 clusters covering the major cell types and multiple cell types. FIG. 15D:Atty. Dkt. No.: 114198-3660zoom-in images of mature RNA (left), nascent RNA (middle) and reconstruction of singlemolecule chromatin conformation for two chromosome 19 homologues in a single nucleus (right). FIG. 15E: Quantification of gene expression by each cell type of the genes present in the FXYD loci imaged with chromatin tracing. FIG. 15F: Scheme for sequential multiplex DNA imaging of 51 genomic regions on chromosome 19, FIG. 15G and FIG. 15H:Reconstruction of chromatin conformation for endothelial cells (FIG. 15G left, spatial map of endothelial cells, FIG. 15G right, heat map of average distance matrix) and vCM (FIG. 15H). FIGS. 15I-15J: antibody staining on the same section with epigenomics markers (left): H3K27acetal (FIG. 15I), H3K27me3 (FIG. 15J), quantification by cell types (right top), and representative images (right bottom).

[0054] FIGS. 16A – 16L: MERFISH+ scales up imaging throughput, enabling 3D reconstruction of a hologram human heart using a 238 gene panel. FIG. 16A: Schematic of MERFISH+ hardware. 10-fold increase in throughput over older microscopy systems. FIG.16B: New fluidics chamber allows for 13.5 cm2imaging area (left), compared to the previous 1.25 cm2imaging area (right). FIG. 16C) Human heart donor section map. The atrial and ventricular halves were sliced into 21 and 32 sections, respectively, with a mean thickness of 16 μm, spanning a total length of 7 mm. (Created in BioRender. Eschbach, J. (2025) https: / / BioRender.com / f3a25hr.) FIGS. 16D - 16E: MERFISH+ spatial maps (FIG. 16D) for all heart sections and corresponding UMAP (FIG. 16E) for the entire heart with 12 major celltype labels. FIG. 16F: MERFISH+ spatial maps for all heart sections with fine annotated cell types. FIG. 16G: The 3D-reconstructed whole heart consists of 53 sections with 112 pm intervals between each section. The cell colors represent the 34 cell types as shown in panel (FIG. 16F). G1-G4 display coronal (G1, G2) and sagittal (G3, G4) cut views of the bisected heart, demonstrating that the reconstruction enables visualization from any angle to explore different regions of the heart. FIG. 16H: Diagrams illustrate the measurement of ventricular wall thickness, where colors denote different ventricular cardiomyocytes layers, highlighting trabecular (1), hybrid (2), compact(3), and free wall region (4) thickness. Legend is in the bottom of panel (J). FIG. 161: Identification of the free wall locations in the right and left ventricles ventral and dorsal view were shown. FIG. 16J; Representative ventricular sections showing the epicardial (red), endocardial (blue), hybrid (purple), and Purkinje (green) surface contours, corresponding to the four layers defined in panel (H). FIG. 16K; Surface heatmapAtty. Dkt. No.: 114198-3660of ventricular wall thickness, where red corresponds to thicker regions and white to thinner regions. FIG. 16K left shows the right view, and FIG. 16K right shows the left view. FIG.16L: Quantitative comparison of wall thickness among the left and right free wall regions; the trabecular, hybrid and compact layers of the left ventricle; assessed by a two-sided Student’s t-test, where *** indicates a p-value < le-3.

[0055] FIGS. 17A–17J: 3D characterization of the human heart valve system. FIG. 17A:Heatmap showing cell-type co-existence scores across all 34 cell types. The tick labels on the left indicate cell types, while a dendrogram constructed from co-existence scores is displayed on the top. Spatial neighborhoods are annotated, and the scale bar in the bottom right indicates the strength of co-existence score. FIG. 17B: Co-existence network graph illustrates positive associations among cell types (represented by different nodes in the graph), where the node size reflects the number of cells for that particular cell type and edge width indicates strength of association. Nodes are grouped into different spatial neighborhoods, encapsulated in unique shapes each colored with a distinct color. FIG. 17C: Visualization of 3D spatial patterns of selected neighborhoods, i.e., VIC and Epicardial neighborhoods, defined in panel A, with different colors representing different cell types. FIG. 17D: Schematic of the valve system of the human developing heart, consisting of the tricuspid valve (TV), bicuspid (mitral valve, MV), aortic valve (AV), and pulmonary valve (PV). Valves are color-coded and referenced throughout this figure. FIG. 17E: Same as in panel B, but for marker genes in VICs from four different heart valves. FIG. 17F: (F) 2D spatial pattern of sections 12 and 17 highlighting the structure of AV / PV, and TV / MV, respectively. The grayscale background represents the overall heart tissue. FIG. 17G: 3D spatial pattern of the valve system from ventral, top, and right views. FIG. 17H: 3D expression maps for four genes (ASPN, VCAN, POSTN, and PENK) in the heart. The color scale indicates relative expression levels, with higher expression in deeper red. White dashed outlines highlight the valve structures. FIG. 171: UMAP scatter plot of single cells, colored with valve structure-specific genes (PENK, FBLN5, TBX5, and HEY2). Shaded represents different valves. FIG. 17J: Cellular composition of valve and cushion structures. For each valve structure, the stacked bar chart (left) illustrates the cell-type composition of neighboring cells in the selected valve region, with each color representing a distinct cell type (reference numbers are used to link to the full cell type names, provided in the legend below). The middle bar chart indicates the high-resolution composition of theAtty. Dkt. No.: 114198-3660highlighted portion from the left chart. 2D spatial distributions of representative cell types enriched in the valve structure are shown on the right, where colored dots indicate cells from a particular enriched cell type.

[0056] FIGS. 18A-18J: 3D spatial analysis of the human heart vasculature. FIG.18A: Schematic of the human developing heart vasculature system formed by vascular smooth muscle cells (VSMCs), consisting of the aorta (AO), pulmonary artery (PA), left coronary artery (LCA), left anterior descending artery (LAD), left circumflex artery (LCX), and right coronary artery (RCA). FIG. 18B: Reconstructed 3D spatial map of the major arteries (see Methods for details). Each color represents a different artery. FIG. 18C: Dot plot illustrating the expression profiles of selected marker genes in AO, PA, and other arteries. FIG. 18D:Three bar plots summarizing morphometric measurements of coronary vessels: branch number (left), total vessel length (middle, in pm), and maximum vessel length (right, in pm). FIG.18E: 3D spatial distribution of three representative coronary arteries (LAD, LCX, and RCA).FIG. 18F: Stacked area plot illustrating neighborhood cell type composition’s dynamic changes along three vessels. Each colored region represents a distinct cell type (reference numbers are used to link to for each the full cell type names, are provided in the legend below).FIG. 18G: Heatmap of cell type enrichment along the distance to the vessel root. Positive correlations are shown in red and negative correlations in blue. FIG. 18H: Gene expression heatmaps along the LAD, RCA, and LCX from root (left on the heatmap) to distal tip (right on the heatmap). Each row represents a different gene, and the color scale indicates relative expression (blue for lower expression and red for higher). FIG. 181: 3D spatial distributions of 12 representative cell types enriched in the neighborhood of the selected vessels, where red dots indicate cells from the enriched cell type. FIG. 18J: 3D expression maps for 12 representative genes in the selected vessels, and differently shaded for lower expression and higher expression.

[0057] FIGS. 19A-19K: Computational integration of spatial transcriptomic data, spatially resolved histone modification maps within a unified 3D reference system. FIG. 19A: Left panel: three different data modalities to Spateo-VI; Middle panel: Schematic of the computational workflow integrating spatial coordinates, gene expression, and chromatin features using Spateo-VI. Right panel: Application of the model: top, projection of 2D multimodal data (1,863-gene expression and histone mark landscapes) into 3D heart space; bottom,Atty. Dkt. No.: 114198-3660inference of spatially enriched cell-cell interactions using COMMON-COT. FIG. 19B:UMAP of the model’s latent space showing separation of transcriptionally distinct cell types across 3.4 million single cells, colored by cell types and batch. FIG. 19C: Relationship between imputation accuracy and spatial specificity. The y-axis represents imputation accuracy, measured by the correlation between true gene expression and imputed gene expression. The x-axis represents spatial specificity, measured by Moran’s I. The left plot shows the true 238 genes measured in 3D heart, and the right plot shows the true 1,863 genes measured in 2D heart. FIG. 19D: Cross 3D-2D’s spatial specificity of imputation results in cell type level. FIG. 19E: Spatial expression maps of representative imputed genes included HEY2, IRX5 and VSTM2L. Top panel: the imputed expression in 3D heart. Bottom panel: the measured expression in 2D heart. FIG. 19F: Scatterplot of and correlation between the mean expression of three characteristic genes profiled with the MERFISH+ 1,863-gene panel experiment and model-imputed data across different cell types across the 3D heart. FIG. 19G:Cross-sections of the left circumflex artery (LCX), left anterior descending artery (LAD), and right coronary artery (RCA) showing cell types and imputed gene expression of two ligandreceptor pairs SEMA3E-PLXND1 and VEGFA-FLT1. FIG. 19H: Schematic illustrating spatial gradient analysis radially from vessel center to periphery. Multiple sections were selected from each vessel (top) and then projected to create an aggregate representation of the cross-section (bottom). FIG. 191: Vector field (represented by arrows) showing the direction and magnitude of SEMA3E-PLXND1 and VEGFA-FLT1 signaling flow in the LCX, LAD, and RCA. FIG. 19J: Imputed 3D spatial chromatin modification landscapes. Left and middle: Imputed H3K4me3 and H3K27me3 histone mark distributions in 3D are shown in the top while the profiled 2D histone mark distribution shown in the bottom. Right: same as in Panel F but for the H3K4me3 or H3K27me3 respectively (r = 0.742 and r = 0.572). FIG. 19K: Web portals for data access and exploration: Spateo-viewer (left) for 3D visualization and spatial analysis and MERFISHEYE (right) for 3D visualization.

[0058] FIG. 20: MERFISH+ workflow with general outline of improvements. MERFISH+ probes are synthesized with a 5’ acrydite modification on the original oligonucleotide probes. Bar-coded readout oligos are prepared using a custom automated liquid handler with a syringe pump (inset). Errors at this step are corrected with a 100% formamide wash, enabling repreparation. Imaging is conducted on larger slides facilitated by the updated fluidics system,Atty. Dkt. No.: 114198-3660followed by complete probe stripping with a 100% formamide wash. A revised encoding system uses identical sequences for each readout probe, simplifying codebook adjustments. Raw images were processed by performing a spot-based fitting method, followed by decoding. For the case of 3D imaging analyses, machine-learning based alignment software Spateo was used. Data visualization was performed using MERFISHEYES, yielding the final MERFISH results. (2 Created in BioRender. Eschbach, J. (2025) https: / / BioRender.com / yzdxjhc)

[0059] FIGS. 21A–21H: MERFISH+ probes allow for flexible codebook. FIG. 21A: Left: Schematic of the original MERFISH probes, where each read sequence targets multiple genes, demonstrated with a single round of MERFISH. Right: MERFISH+ probe schematic, where each read sequence is unique to a single gene, enabling use in both smFISH and MERFISH. (Created in BioRender. Eschbach, J. (2025) https: / / BioRender.com / iwamj94.) FIG. 21B:Spatial heatmap of the exemplary gene IRX3 visualized with smFISH using MERFISH+ probes. FIG. 21C: Spatial plots of all 15 genes analyzed with smFISH-Plus. FIG. 21D:Sample MERFISH codebook. FIG. 21E: Workflow for implementing the codebook in the MERFISH+ protocol, utilizing a custom pipetting robot to dispense hybridization tubes with adaptors corresponding to the codebook. (Created in BioRender. Eschbach, J. (2025) https: / / BioRender.com / lal231r, https: / / BioRender.com / qlqozk7.) FIG. 21F: Hybridization tubes are used to prepare the sample, which is subsequently imaged. FIG. 21G: The same sample is imaged using two distinct codebooks. (Created in BioRender. Eschbach, J. (2025) https: / / BioRender.com / c8pomhq.) FIG. 21H: Pearson correlation of mean transcripts per cell between the two codebooks demonstrates high correlation (0.815).

[0060] FIGS. 22A: Pearson correlation of mean transcripts per cell between original MERFISH and smFISH experiment shows high correlation (0.789). FIGS. 22B: Comparison of spatial plots of five marker genes for original MERFISH and smFISH experiment shows similar spatial expression patterns. FIGS. 22C: Correlation of transcripts detected by MERFISH-PLUS (MER1) and MERFISH. FIGS.22D: Correlation of transcripts detected by MERFISH-PLUS (MER2) and MERFISH.

[0061] FIG. 23A: three subclusters of blood endothelial cells (BEC I, II, III) identified by MERFISH+. FIG.23B: level of a few markers expressed in BEC I, II, III. FIG.23C: SOX17 and HEYL expression identified by MERFISH+. FIG. 23D: two distinct epicardialAtty. Dkt. No.: 114198-3660populations (aEpicardial and vEpicardial) identified by MERFISH+. FIG. 23E: F13A1 and CEMIP expression identified by MERFISH+.

[0062] FIGS. 24A-24C: Flexibility of multimodality imaging. FIG. 24A: Ensemble chromatin structure of vCM (left heat map) highly correlates with that of HiC (middle contact frequency) from cultured differentiated cardiomyocytes. Pearson correlation (right). FIG.24B: significantly high correlation between antibody staining from 15 p.c.w. and 12 p.c.w. heart. Pol2ser2, Pol2ser5, H3K4me3, H3K9me3 andH3K27me3. Left: spatial map of z-score; middle correlation in cell types; right: Pearson correlation. FIG.24C: To validate the accuracy of the protein-staining in MERFISH+, Applicant performed immunofluorescence staining with the same antibodies using adjacent tissue slices from the same heart sample. Quantification of such measurements showed good correlation between the two sets of data.

[0063] FIG. 25A: Spatial plot of all 45 cell types from 53 sections of a 12p.c.w. human heart.FIG. 25B: Number of cells of each cell type across the heart. FIG.25C: Correlation of average transcripts per cell for each gene with previously published snRNA-seq (left) and MERFISH (right) data. FIG. 25D: Heatmap showing pairwise cell type to cell type correlation between this MERFISH data and the previously published MERFISH data.

[0064] FIGS. 26A-26B: Cardiomyocyte sub-populations in the ventricles. FIG. 26A: Pie chart showing the proportions of major cell types in the heart (top) and the proportions of ventricular cardiomyocyte sub-types (bottom). FIG. 26B: Two sections within the ventricles showing the layered composition of the left ventricle. FIG. 26C: All 13 ventricular cardiomyocyte cell types displayed in a 2D view across the heart.

[0065] FIGS. 27A-27B: Three dimensional reconstruction of the molecular hologram of the human heart. FIG. 27A: Serial reconstruction of the human heart from 53 consecutive tissue sections with 112 pm intervals. FIG. 27B: MERFISH+ spatial maps for the 3D human heart with fine annotated cell types. The leftmost panel shows the full 3D heart, while the remaining panels display spatial distributions of individual cell types.

[0066] FIG. 28A: Thickness of the ventricular wall measured from the reconstructed 3D heart, displayed across four views. The LV, RV, LV-anterior / dorsal, and RV-apex regions are highlighted with dotted circles, where red surface indicates thicker areas and white indicates thinner areas. FIG. 28B: Quantitative comparison of thickness in the LV and RV free wallAtty. Dkt. No.: 114198-3660region, LV-anterior / dorsal, LV hybrid and its anterior / dorsal regions, where *** indicates a p-value < le-3. FIG. 28C: Correlation analysis of LV hybrid layer thickness with LV compact and LV trabecular layer thickness. Pearson correlation coefficients (r) and p-values are indicated.

[0067] FIG. 29A: Schematic of the neighborhood definition and composition analysis. Left: A cell’s neighborhood comprises neighboring cells within its own tissue section and those in adjacent sections. Right: For each cell, the relative abundance of each cell type within its neighborhood is tabulated to form an N cells × M cell-types matrix for downstream coexistence analysis. FIG. 29B: Visualization of 3D spatial patterns of selected neighborhoods, i.e., Atria and Trabecular neighborhoods, with different colors representing different cell types.FIG. 29C: Illustration of the workflow used to extract ventricle surface cells from the 3D dataset. FIG. 29D: Pie chart showing the cell-type composition on the ventricle surface; only the top 15 most abundant cell types are displayed. FIG. 29E: 3D neighborhood complexity visualization of the ventricle surface region. The colormap is displayed on the right. FIG.29F: Serial 2D coronal sections (#10-21) through the human heart highlighted with the recognized VICs of the tricuspid (TV, yellow), pulmonary (PV, cyan), mitral (MV, blue) and aortic (AV, magenta) valves. FIG. 29G: 3D reconstruction of each VIC, shown within cubic bounding boxes. Color conventions are as in (A). FIG. 29H: Cellular composition of the neighborhood of the valve and cushion structures. For each valve structure, the stacked bar chart (left) illustrates the cell type enrichment of neighboring cells in the selected valve region within 25 um radii, where each color represents a distinct cell type. The middle stacked bar chart indicates the high-resolution composition of the highlighted portion from the left chart. The right stacked area plot shows how the cell type enrichment changes as the neighborhood radius increases (reference numbers for each cell type are provided in the legend below).DETAILED DESCRIPTION

[0068] Definitions

[0069] As it would be understood, the section or subsection headings as used herein is for organizational purposes only and are not to be construed as limiting and / or separating the subject matter described.Atty. Dkt. No.: 114198-3660

[0070] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, the preferred methods, devices, and materials are now described. All technical and patent publications cited herein are incorporated herein by reference in their entirety. Nothing herein is to be construed as an admission that the disclosure is not entitled to antedate such disclosure by virtue of prior disclosure.

[0071] The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of tissue culture, immunology, molecular biology, microbiology, cell biology and recombinant DNA, which are within the skill of the art. See, e.g., Sambrook and Russell eds. (2001) Molecular Cloning: A Laboratory Manual, 3rd edition; the series Ausubel et al. eds. (2007) Current Protocols in Molecular Biology; the series Methods in Enzymology (Academic Press, Inc., N. Y.); MacPherson et al. (1991) PCR 1: A Practical Approach (IRL Press at Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach; Harlow and Lane eds. (1999) Antibodies, A Laboratory Manual; Freshney (2005) Culture of Animal Cells: A Manual of Basic Technique, 5th edition; Gait ed. (1984) Oligonucleotide Synthesis; U. S. Patent No. 4,683,195; Hames and Higgins eds. (1984) Nucleic Acid Hybridization; Anderson (1999) Nucleic Acid Hybridization; Hames and Higgins eds. (1984) Transcription and Translation; Immobilized Cells and Enzymes (IRL Press (1986)); Perbal (1984) A Practical Guide to Molecular Cloning; Miller and Calos eds. (1987) Gene Transfer Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Makrides ed. (2003) Gene Transfer and Expression in Mammalian Cells; Mayer and Walker eds. (1987) Immunochemical Methods in Cell and Molecular Biology (Academic Press, London); Herzenberg et al. eds (1996) Weir’s Handbook of Experimental Immunology; Manipulating the Mouse Embryo: A Laboratory Manual, 3rd edition (Cold Spring Harbor Laboratory Press (2002)); Sohail (ed.) (2004) Gene Silencing by RNA Interference: Technology and Application (CRC Press).

[0072] As used in the specification and claims, the singular form “a,” “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a cell” includes a plurality of cells, including mixtures thereof.Atty. Dkt. No.: 114198-3660

[0073] As used herein, the term “comprising” is intended to mean that the compounds, agents, compositions and methods include the recited elements, but not exclude others. “Consisting essentially of’ when used to define compounds, agents, compositions and methods, shall mean excluding other elements of any essential significance to the combination. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants, e.g., from the isolation and purification method and pharmaceutically acceptable carriers, preservatives, and the like. “Consisting of’ shall mean excluding more than trace elements of other ingredients. Embodiments defined by each of these transition terms are within the scope of this technology.

[0074] All numerical designations, e.g., pH, temperature, time, concentration, and molecular weight, including ranges, are approximations which are varied (+) or (-) by increments of 1, 5, or 10%. It is to be understood, although not always explicitly stated that all numerical designations are preceded by the term “about.” It also is to be understood, although not always explicitly stated, that the reagents described herein are merely exemplary and that equivalents of such are known in the art.

[0075] The term “about,” as used herein when referring to a measurable value such as an amount or concentration and the like, is meant to encompass variations of 20%, 10%, 5%, 1 %, 0.5%, or even 0.1 % of the specified amount.

[0076] As used herein, comparative terms as used herein, such as high, low, increase, decrease, reduce, or any grammatical variation thereof, can refer to certain variation from the reference. In some embodiments, such variation can refer to about 10%, or about 20%, or about 30%, or about 40%, or about 50%, or about 60%, or about 70%, or about 80%, or about 90%, or about 1 fold, or about 2 folds, or about 3 folds, or about 4 folds, or about 5 folds, or about 6 folds, or about 7 folds, or about 8 folds, or about 9 folds, or about 10 folds, or about 20 folds, or about 30 folds, or about 40 folds, or about 50 folds, or about 60 folds, or about 70 folds, or about 80 folds, or about 90 folds, or about 100 folds or more higher than the reference. In some embodiments, such variation can refer to about 1%, or about 2%, or about 3%, or about 4%, or about 5%, or about 6%, or about 7%, or about 8%, or about 0%, or about 10%, or about 20%, or about 30%, or about 40%, or about 50%, or about 60%, or about 70%, or about 75%, or about 80%, or about 85%, or about 90%, or about 95%, or about 96%, or about 97%, or about 98%, or about 99% of the reference.Atty. Dkt. No.: 114198-3660

[0077] As will be understood by one skilled in the art, for any and all purposes, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Furthermore, as will be understood by one skilled in the art, a range includes each individual member.

[0078] “Optional” or “optionally” means that the subsequently described circumstance may or may not occur, so that the description includes instances where the circumstance occurs and instances where it does not.

[0079] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).

[0080] “Substantially” or “essentially” means nearly totally or completely, for instance, 95% or greater of some given quantity. In some embodiments, “substantially” or “essentially” means 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0081] The terms or “acceptable,” “effective,” or “sufficient” when used to describe the selection of any components, ranges, dose forms, etc., disclosed herein intend that said component, range, dose form, etc., is suitable for the disclosed purpose.

[0082] The term “contacting” means direct or indirect binding or interaction between two or more. A particular example of direct interaction is binding. A particular example of an indirect interaction is where one entity acts upon an intermediary molecule, which in turn acts upon the second referenced entity. Contacting as used herein includes in solution, in solid phase, in vitro, ex vivo, in a cell and in vivo. Contacting in vivo can be referred to as administering, or administration.[00831 As used herein, the phrase “Hamming weight” refers to the number of symbols that are different from the zero-symbol in a string. In binary code, the phrase “Hamming weight” represents the number of l’s in a codeword.

[0084] As used herein, the phrase “Hamming distance” refers to the number of bit positions in which the two codewords are different. In other words, Hamming distance measures the minimum number of substitutions required to change one codeword into the other, orAtty. Dkt. No.: 114198-3660equivalently, the minimum number of errors that could have transformed one codeword into the other.

[0085] As used herein, a “valid codeword” is a codeword assigned to a target molecule. In some embodiments, a codeword is an N-bit codeword (e.g., 8-bit, 16-bit, 32-bit or more). In some embodiments, any two valid codewords have a Hamming distance equal or greater than 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10) between them. An unassigned codeword is an invalid codeword and is assigned to a valid codeword after error correction, e.g., as described in US Patent No. 11,959,075, which is incorporated herein in its entirety.

[0086] As used herein, the term “spatial” refers to relating to, describing, or involving the physical position, arrangement, or organization of a target molecule, cell, or cellular component within a sample. In certain embodiments, “spatial” encompasses the location of such molecules or structures in one-, two-, or three-dimensional coordinate systems, including their relationships or proximities to other molecules, cells, or anatomical features. Spatial information may be expressed relative to a fixed reference within a tissue or organ, or in absolute coordinates derived from imaging data. In the context of spatial transcriptomics, genomics, or proteomics, “spatial” indicates that the measured molecular data retains positional context within intact tissue architecture, thereby enabling analysis of molecular organization, cell-cell interactions, gradients, boundaries, or structures without dissociating the sample.10087] As used herein, “spatial organization” refers to the arrangement and relative positioning of molecules, cells, tissues, or anatomical structures within a biological sample, such that their location and orientation can be determined and correlated with morphological or functional characteristics. Spatial organization may be defined at multiple scales, including subcellular, cellular, tissue, and organ-level organization, and may describe patterns, distributions, gradients, or repetitive structures discernible through imaging or spatial mapping methods. In certain embodiments, spatial organization includes both absolute and relative coordinates, relationships to designated landmarks, or membership within spatial neighborhoods.

[0088] As used herein, “spatial resolution” refers to the smallest physical distance between two points within a sample that can be distinguished as separate in an imaging or mapping modality. Spatial resolution can be expressed in terms of nanometers, micrometers, or other relevant units, depending on the scale of observation. Higher spatial resolution enables greater precisionAtty. Dkt. No.: 114198-3660in determining molecular or cellular positions and in resolving fine structural details, whereas lower spatial resolution permits only coarse positional assignment. In some embodiments, spatial resolution is defined separately in the x, y, and z dimensions of two- or three-dimensional datasets.

[0089] As used herein, “spatial mapping” refers to the process of assigning measured molecular features (such as gene transcripts, DNA loci, or proteins) to defined physical coordinates within a sample, thereby generating a positional map of those features in one-, two-, or three-dimensional space. Spatial mapping may utilize microscopy images, coordinate-based registration across imaging rounds, and / or computational alignment to reconstruct molecular distributions with respect to tissue architecture. In certain embodiments, spatial mapping further includes alignment of multi-modal datasets (e.g., RNA, DNA, protein) into a unified coordinate framework for integrated spatial analysis.

[0090] As used herein, “biogel film” refers to a continuous or semi-continuous layer of polymeric hydrogel material formed on, or covering, a surface of a biological sample. In certain embodiments, the biogel film embeds the sample at least partially or fully within its matrix. The biogel film provides mechanical stabilization, chemical support, and / or a covalent anchoring medium for hybridization probes, antibodies, or other detection reagents, thereby preserving positional information for target molecules during repeated imaging cycles. In some embodiments, the biogel film comprises a polyacrylamide gel, optionally prepared from acrylamide and bis-acrylamide monomers, and is polymerized in contact with acrydite-modified oligonucleotide probes such that the probes are covalently incorporated into the gel matrix. The thickness of the biogel film is sufficient to maintain probe anchoring and sample integrity under stringent hybridization and stripping conditions and, in certain embodiments, is from about 20 μm to about 200 μm, including, for example, about 20 μm, about 50 μm, about 100 μm, about 150 μm, or about 200 μm. Thickness values may vary according to the sample type, imaging modality, or desired mechanical properties, provided that the film maintains continuous coverage of the sample region to be imaged.

[0091] As used herein, “microscope-microfluidics system” refers to an integrated apparatus comprising an optical microscopy platform operatively coupled with a fluid handling system configured to deliver, exchange, and remove reagents in a controlled manner during imaging of biological samples. In certain embodiments, the microscopy platform includes a microscopeAtty. Dkt. No.: 114198-3660body, one or more objectives, an illumination source (e.g., high-power laser or solid-state light engine), optical filters, and a scientific imaging sensor (e.g., sCMOS or large-format camera). The microfluidics subsystem may comprise one or more fluidic chambers, valves, pumps, tubing, and reagent reservoirs capable of sequentially introducing hybridization buffers, fluorescent readout probes, wash solutions, and stripping reagents, while maintaining precise alignment of the sample within the optical path.

[0092] As used herein, “microscope system” refers to an optical imaging apparatus comprising a microscope body with objective lenses, illumination source, imaging sensor, and stage controls, configured to acquire images of a biological sample. In some embodiments, the microscope system is adapted for fluorescence or other imaging modalities and may be integrated with components for high-throughput or large-format imaging.

[0093] As used herein, “quenched” refers to a reduction, elimination, or inactivation of detectable signal from a fluorescent label or other signal-generating moiety on a probe, such that the probe no longer produces a measurable emission under the imaging conditions employed. In certain embodiments, quenching may be achieved by chemical, enzymatic, photochemical, or physical means. Chemical quenching can include cleavage or modification of the fluorophore (e.g., via treatment with formamide, TCEP, or other denaturants) to disrupt its light-emitting properties. Enzymatic quenching can involve removal or alteration of the fluorescent label through enzyme-mediated reactions. In some embodiments, quenching is performed after each hybridization and imaging round to remove residual fluorescence from fluorescently labeled readout probes, thereby preventing carryover signal into subsequent rounds. Quenching may be reversible or irreversible and can occur by direct destruction of the fluorophore or by altering its microenvironment (e.g., bringing the fluorophore into close proximity with a quencher molecule) such that emitted light is substantially reduced or eliminated.

[0094] As used herein, the term “distinct read sequences” refers to oligonucleotide sequences that differ in nucleotide composition or sequence order from one another, the differences being sufficient to confer specificity such that each read sequence hybridizes preferentially to a matching fluorescently labeled readout probe under the hybridization conditions employed, with minimal or no cross-hybridization to other read sequences in the same probe pool. In certain embodiments, distinct read sequences are unique within the probe pool and correspondAtty. Dkt. No.: 114198-3660to different bit positions in an N-bit codeword, enabling independent detection of the associated target molecules during sequential imaging rounds. Distinct read sequences may be appended to the target sequence of the primary nucleic acid probe at a defined location (e.g., 5’ or 3’ end) and typically range from about 8 nucleotides to about 30 nucleotides in length.

[0095] As used herein, the term “proteome” refers to the entire set of proteins, including all isoforms and post-translationally modified variants, produced or present in a cell, tissue, organ, or organism at a given time under specific conditions. In some embodiments, “proteome” refers to the proteins detectable in the sample or region of interest using the disclosed methods, including structural proteins, enzymes, receptors, signaling molecules, and transcriptional regulators. The spatial organization of the proteome includes subcellular localization patterns, relative abundances, and proximity relationships between proteins in intact biological specimens.

[0096] As used herein, “probe pool” refers to a defined collection of two or more nucleic acid probes that share a common design criterion, such as specificity to the same target molecule or to different molecules sharing a functional or experimental grouping, and / or participation in the same codeword assignment. Each probe pool may comprise multiple distinct target sequences and / or multiple distinct read sequences, and may be used together in one hybridization cycle or distributed across cycles according to a predetermined codebook.

[0097] As used herein, “species” in reference to a nucleic acid or protein target denotes a molecule distinguishable from others by virtue of its unique primary sequence, hybridization target region, or epitope. For nucleic acids, “species” includes transcripts or strands with a distinct contiguous sequence of nucleotides; in the case of mRNA, it may include alternatively spliced variants if such variants are specifically targeted by the probe design. For proteins, “species” includes proteins or polypeptides with distinct amino acid sequences or post-translational modifications that can be distinguished by selective binding agents such as antibodies.

[0098] As used herein, “read sequence” refers to a defined oligonucleotide sequence incorporated into a primary nucleic acid probe, intended to hybridize to a complementary “readout probe” during a particular imaging round. “Readout sequence” refers to the same oligonucleotide sequence, but in the context of being detected via a fluorescently labeled probeAtty. Dkt. No.: 114198-3660during the imaging process. The terms may be used interchangeably in describing the nucleotide region within the primary probe that mediates optical detection through hybridization with a fluorophore-conjugated oligo.

[0099] As used herein, “distinct target molecule” refers to a target nucleic acid or protein within the sample that is unique in sequence, epitope, or identifiable characteristic such that it can be specifically recognized and discriminated from other molecules present by virtue of the binding specificity of the primary probe or antibody. In the case of nucleic acids, distinctness may be determined by at least an 8–12 nucleotide unique subsequence accessible to hybridization; in proteins, by a unique epitope recognized by a capture or detection antibody.Modes for Carrying Out the Disclosure

[0100] The present invention generally relates to improved methods for imaging or determining the abundance of nucleic acids or proteins, for instance, within cells, using probes of improved stability. In some embodiments, the transcriptome of a cell may be determined. In some embodiments, the proteome of a cell may be determined. Certain embodiments are directed to determining nucleic acids, such as mRNA, within cells at a single cell resolution. In some embodiments, a plurality of nucleic acid probes may be applied to a sample, and their binding within the sample determined, e.g., using fluorescence, to determine locations of the nucleic acid probes within the sample. In some embodiments, codewords (e.g., N-bit codewords) may be based on the binding of the plurality of nucleic acid probes, and in some cases, the codewords may define an error-correcting code to reduce or prevent misidentification of the nucleic acids. In certain cases, a relatively large number of different targets may be identified using a relatively small number of labels, e.g., by using various combinatorial approaches.

[0101] In some embodiments, primary probes (also called “encoding probes” or “primary nucleic acid probes”) and secondary probes (also called “readout probes”) are used, where the primary probes encode “codewords” (also called “N-bit codewords”) and bind to target molecules (e.g., nucleic acids) in the sample, and the secondary probes are used to read out the codewords from the primary probes. In some embodiments, a plurality of different primary probes containing codewords are divided into as many separate pools as there are positions inAtty. Dkt. No.: 114198-3660the codewords, such that each primary probe pool corresponds to a certain value in a certain position of the codewords (e.g., a “one” in the first position as in “1001”).

[0102] An aspect of the disclosure is directed to an improved method for imaging a target molecule spatial organization in a sample comprising, or alternatively consisting essentially of, or consisting of, or yet further consisting of:a) contacting a sample comprising, or alternatively consisting essentially of, or consisting of, or yet further consisting of a plurality of target molecules and one or more primary nucleic acid probe pools,wherein each primary nucleic acid probe comprises, or alternatively consists essentially of, or yet further consists of an acrydite group and is bound to a biogel film covering a surface of the sample, andwherein each pool of primary nucleic acid probes is selected to specifically hybridize to a distinct target molecule in the sample, andwherein each primary nucleic acid probe comprises, or alternatively consists essentially of, or yet further consists of a target sequence and one or more read sequences, and wherein each primary nucleic acid probe produces a bound primary nucleic acid probe upon contact with its target molecule, andwherein each pool of primary nucleic acid probes encode an N-bit codeword with a Hamming weight of at least 2 that was assigned to each distinct target molecule, wherein each assigned N-bit codeword is a valid codeword with a Hamming distance equal to or greater than 2 between valid codewords and wherein each of the read sequences correspond to a bit value of 1 for the codewords assigned to the target molecules,b) contacting the one or more plurality of primary nucleic acid probe pools with a plurality of fluorescently labeled readout probes to specifically bind to read sequences of the primary nucleic acid probes,c) imaging the readout probes bound to the primary nucleic acid probes; andAtty. Dkt. No.: 114198-3660d) repeating steps b) and c) in or more sequential binding and imaging rounds until all N positions in the N-bit codeword have been imaged providing an imaged codeword corresponding to each distinct molecule in a spatial organization.

[0103] In some embodiments, the biogel film embeds the sample. In some embodiments, the binding of the primary nucleic acid probes to the biogel film stabilizes the primary nucleic acid probes and allows imaging to be done for an extended time after probe binding, e.g., after a period of at least 1 month, or at least 2 months, or at least three months (i.e., after the target molecule (e.g., mRNA, protein or DNA) in the sample has likely degraded).

[0104] In some embodiments, each primary nucleic acid probe comprises, or alternatively consists essentially of, or yet further consists of at least two (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or 16) read sequences unique to the target molecule, wherein each read sequence is different from each other.

[0105] In some embodiments, the biogel film comprises, or alternatively consists essentially of, or yet further consists of a polyacrylamide gel matrix, In some embodiments, each primary nucleic acid probe binds to the polyacrylamide gel matrix.

[0106] In some embodiments, the target molecule is selected from a nucleic acid or a polypeptide. In some embodiments, the nucleic acid is a DNA or an RNA molecule. In some embodiments, when the target molecule is a polypeptide, the read out of the polypeptide is conducted through an specific primary antibody, followed by a secondary fluorescence-labeled antibody

[0107] In some embodiments, steps c) and d) are repeated over an extended time period of at least 1 month, at least 2 months, or at least 3 months.

[0108] In some embodiments, the acrydite group is derived from a compound of formula (HO)2P(O)(CH2)xNHC(O)C(CH2)CH3, wherein x is 1-10 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10). In some embodiments, x is 7. In some embodiments, the acrydite group is attached to a 5’ end of each of the plurality of primary nucleic acid probes.

[0109] In some embodiments, the N-bit codeword comprises, or alternatively consists essentially of, or yet further consists of between 8 and 64 bits (e.g., 8-bit, 9-bit, 10-bit, 11 -bit, 12-bit, 13-bit, 14-bit, 15-bit, 16-bit, 20-bit, 32-bit, 48-bit, 64-bit, or any value therebetween).Atty. Dkt. No.: 114198-3660

[0110] In some embodiments, the assigned N-bit codewords have a Hamming distance equal to or greater than 4 between each of the valid codewords (e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 20, 32, 48, 64, or any value therebetween).

[0111] In some embodiments, a constant Hamming weight (i.e., the number of “ 1” bits in each barcode) is used. In some embodiments, using a constant Hamming weight avoids potential bias in the measurement of different barcodes because of a differential rate of “1” to “0” and “0” to “1” errors. In some embodiments, each N-bit codeword has a Hamming weight of at least 4 (e.g., at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 20, 32, 48, 64, or any value therebetween).

[0112] In some embodiments, the imaged codeword is matched to a valid codeword assigned to a distinct target molecule (e.g., a distinct RNA species, a distinct DNA sequence or a distinct protein).

[0113] In some embodiments, the Hamming distance is 4 or greater and the imaged codewords are matched to valid codewords or discarded. In some embodiments, the N-bit codeword has a Hamming weight of 4 and the Hamming distance is 4 or greater and the imaged codewords are matched to valid codewords or discarded. As used herein, a “Hamming distance of 4” means that at least four errors need to accumulate to convert one codeword into another codeword. Hamming distance of 4 allows the detection of errors up to any two bits, and the correction of errors to any single bit.

[0114] In some embodiments, the target molecule is an RNA molecule. In some embodiments, the target molecule is an mRNA transcript. In some embodiments, the target molecule is a noncoding RNA transcript (e.g., IncRNA, miRNA) or a intronic RNA.

[0115] In some embodiments, the target molecule is a DNA molecule. In some embodiments, the target molecule is packaged in chromatin.

[0116] In some embodiments, the method comprises, or alternatively consists essentially of, or yet further consists of determining the spatial organization of the transcriptome from a single cell.

[0117] In some embodiments, the target molecule is a polypeptide. In some embodiments, the target molecule is an enzyme.Atty. Dkt. No.: 114198-3660

[0118] In some embodiments, the method determines the spatial organization of the proteome from a single cell.

[0119] In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of at least 10 (e.g., at least 10, 15, 20, 25 30, 35, 4045, 50 or any value therebetween) different primary nucleic acid probes.

[0120] In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of at least four (e.g., at least 4, 8, 16, 32, or any value therebetween) distinct read sequences. In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of four distinct read sequences. In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of at least eight (e.g., at least 8, 16, 32, or any value therebetween) distinct read sequences.

[0121] In some embodiments, the primary nucleic acid probes comprise, or alternatively consist essentially of, or yet further consist of a target sequence, with an average length of between 10 and 200 nucleotides (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170,175, 180, 185, 190, 195, 200, or any value therebetween), that hybridize the distinct target molecule.

[0122] In some embodiments, each primary nucleic acid probe pool comprises, or alternatively consists essentially of, or yet further consists of at least 10 (e.g., at least 10, 15, 20, 25 30, 35, 4045, 50 or any value therebetween) different target sequences.

[0123] In some embodiments, after each hybridization and imaging round the fluorescent readout probe is quenched to inactivate, wherein after each hybridization and imaging round the fluorescent read out probe is inactivated by chemically or enzymatically cleaving the fluorescent label from the readout probe.

[0124] In some embodiments, the N-bit binary code comprises, or alternatively consists essentially of, or yet further consists of at least a 8-bit code, at least a 16-bit code, at least a 32-bit code, or at least a 64-bit code (e.g., 8-bit, 9-bit, 10-bit, 11-bit, 12-bit, 13-bit, 14-bit, 15-bit, 16-bit, 20-bit, 32-bit, 48-bit, 64-bit, or any value therebetween).Atty. Dkt. No.: 114198-3660

[0125] In some embodiments, the plurality of fluorescent readout probes comprises, or alternatively consist essentially of, or yet further consist of at least two (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9 or 10) distinct fluorescent labels. In some embodiments, the plurality of fluorescent readout probes comprises, or alternatively consists essentially of, or yet further consists of at least three (e.g., at least 3, 4, 5, 6, 7, 8, 9 or 10) distinct fluorescent labels.

[0126] In some embodiments, the spatial organization of the target molecule is imaged in 2 dimensions. In some embodiments, the spatial organization of the target molecule is imaged in 3 dimensions.

[0127] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of determining abundance of the target molecule (e.g., the abundance of an RNA transcript, DNA or protein).

[0128] The following examples are included to demonstrate some embodiments of the disclosure. However, those of skill in the art should, considering the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the spirit and scope of the invention.

[0129] In some aspects, the disclosed methods enable multiplexed multimodal imaging in which RNA transcripts, genomic DNA loci (e.g., chromatin tracing targets), and protein-based epigenetic markers are measured within a single tissue section or intact sample. This multimodal capability may be facilitated by the covalent anchoring of encoding probes to a biogel film via acrydite chemistry, which maintains probe stability after repeated washes and imaging cycles. In some embodiments, the sample comprises acrydite-modified nucleic acid probes that hybridize to target RNA or genomic DNA sequences, followed sequentially by fluorescent readout probe hybridization for RNA and DNA targets, and subsequent immunofluorescence staining using primary and secondary antibodies for protein detection, including histone modification marks, transcriptional machinery components, or other epigenetic factors. Sequential hybridization, imaging, and signal-stripping steps are performed without loss of probe anchoring or sample integrity, allowing multi-month imaging sessions to acquire multiple data modalities in precise spatial register.Atty. Dkt. No.: 114198-3660

[0130] In some embodiments, a large-format spatial genomics dataset is acquired and reconstructed into a three-dimensional (3D) whole-organ atlas. For example, a developing human heart is serially sectioned into dozens of tissue slices (e.g., 53 sections at approximately 112 pm spacing) and imaged using the disclosed high-throughput microscope-microfluidics systems described herein. Each section is processed with MERFISH+ to profile hundreds to thousands of gene targets. The spatial transcriptomics sections can be computationally aligned and assembled into a seamless 3D organ representation using the Spateo pipeline, which registers adjacent sections via transcriptional similarity and spatial constraints. In some embodiments, the method provides a volumetric molecular atlas at single-cell resolution covering the entire organ, with quantitative anatomical measurements and spatial neighborhood analysis performed in 3D.

[0131] In certain embodiments, the disclosed MERFISH+ probe libraries comprise unique read sequences per target molecule, permitting adaptation between single-molecule fluorescence in situ hybridization (smFISH) and multiplexed error-robust fluorescence in situ hybridization (MERFISH) combinatorial imaging modes using the same physical probe set. In some embodiments, the unique read sequences enable flexible target selection, codebook modification, or codebook swapping between experiments without redesign or resynthesis of the primary probe library. In one example, codebooks for combinatorial MERFISH imaging are assembled via automated robotic liquid handling systems, which combine readout probes in predetermined hybridization rounds according to the selected codebook. The same sample can be sequentially imaged with different codebooks to target distinct gene subsets from the same anchored probe library.

[0132] In some embodiments, multimodal 3D spatial-omics datasets (e.g., RNA, DNA, and protein measurements) are integrated using a computer-implemented system comprising a graph attention-based spatial encoder (Spateo- VI). This computational method constructs spatial neighborhood graphs, applies learned attention weights to neighbor relationships, and generates unified latent space embeddings that harmonize 2D and 3D datasets across multiple modalities. The integrated spatial-omics data can be used for imputation of unmeasured gene expression patterns, reconstruction of histone modification landscapes in 3D space, or inference of spatially enriched cell-cell communication networks (e.g., via C0MM0N-3D), enabling comprehensive molecular and structural analysis of intact organs.Atty. Dkt. No.: 114198-3660[01331 The disclosed microscope-microfluidics system can include a large-format fluidics chamber with an imaging area of about 13.5 cm2or greater, enabling MERFISH+ imaging at about 10-fold throughput compared to standard MERFISH platforms. In certain embodiments, the system incorporates high-power laser illumination (about 1-2.5 W per laser line), 0.4 mm square multimode fiber coupling for enhanced beam uniformity, and scientific cameras with sensors up to 20.8 mm diagonal to increase the field of view per frame. This hardware configuration allows imaging of large-format tissue sections for whole-organ mapping in substantially fewer experiments.

[0134] In some embodiments, acrydite-modified primary probes are synthesized from pooled oligonucleotide libraries using a multi-step method. A pool of oligonucleotides containing probe target sequences can be first amplified via polymerase chain reaction (PCR) using a forward primer bearing a 5’ acrydite modification. The amplified DNA can be transcribed in vitro using T7 RNA polymerase, and the resulting RNA is reverse-transcribed into single- stranded DNA using a reverse transcription primer also bearing a 5' acrydite modification. In some embodiments, the RNA template is degraded chemically (e.g., via alkaline hydrolysis), and the final single-stranded DNA probes are purified for use in spatial hybridization experiments. This workflow can covalently incorporate acrydite groups at the desired probe termini, enabling direct anchoring into polyacrylamide biogel films during polymerization.

[0135] Examples:

[0136] Example 1: MERFISH+ for High-throughput High-resolution Tissue Imaging

[0137] Experimental Methods

[0138] Acrydite FISH-probes improve stability of MERFISH experiments

[0139] To overcome the gradual decay of RNA during imaging and hence a gradual decrease of the on-target signal Applicant modified the MERFISH probe synthesis protocol (Moffitt, Hao, Bambah-Mukku, et al. 2016) to incorporate an acrydite group to the 5’ end of the oligonucleotides of all the encoding / primary probes. The acrydite group is chemically similar to acrylamide or bis-acrylamide monomers. Thus, after hybridization, an acrylamide gel can be cast on top of the sample and, during polymerization, the acrydite-conjugated oligos will be incorporated into the gel matrix (Rehman et al. 1999). Such a modification is expected toAtty. Dkt. No.: 114198-3660decouple the probe stability from the decay of the RNA substrate during imaging allowing for greatly expanding the number of hybridization cycles. (FIG. 2A)

[0140] Applicant conducted MERFISH experiments with acrydite modified oligonucleotide probes on 16-um human fetal heart sections (12-post conceptual week). Briefly, a template library of MERFISH probes produced through the microarray synthesis technology (TWIST), was amplified first using PCR and then T7 RNA synthesis followed by a reverse transcription reaction with an acrydite modified primer at the 5’ end. These probes were hybridized to RNA molecules within the tissue, followed by the casting of a ~20-um-200um acrylamide gel onto the sample. In order to quantify the stability of the new probes, Applicant used acrydite probes targeting MYH7 (a gene expressed in ventricular cardiomyocytes) and compared the signal in the first cycle of hybridization with the signal after 50 cycles of hybridization with highly stringent 100%-formamide washes between each cycle. (FIG. 2B). Applicant found that brightness plateaus at about 55% of the first round’s brightness.

[0141] Applicant then resynthesized a 238-gene MERFISH library used previously and performed MERFISH experiments with the acrydite modified probes across 11 hybridization cycles. On the same sample, Applicant then performed a stringent 100%-formamide wash step to completely remove residual readout probes and Applicant reconducted the MERFISH experiment across an additional 11 hybridization cycles. Applicant quantified the number of mRNA molecules of each of the 238 targeted genes in each cell and performed UMAP embedding and cell type definition using leiden clustering of the cells imaged across the two MERFISH experiments. Applicant obtained a near identical cell type definition and spatial distribution across the 2D section (FIG.2C, FIG.2D) with the same cell type markers defining each cell type (FIG.2E) as in the original MERFISH protocol (Moffitt, Hao, Wang, et al. 2016; Farah et al. 2024)). Furthermore, Applicant obtained a very high Pearson correlation (0.95 and 0.93) between the average expression of each gene in both the new MERFISH experiments compared to prior published results ((Moffitt, Hao, Wang, et al. 2016; Farah et al. 2024)) suggesting that the detection efficiency and accuracy was not appreciably affected by the new acrydite modified probes. (FIG.2F). Importantly these results demonstrate the increased stability of the probes and allow for overcoming some of the lack of reliability in the original MERFISH design caused RNA decay, equipment failures or reagent instability. By performingAtty. Dkt. No.: 114198-3660a stringent 100% formamide wash post failure, the MERFISH experiment can be reset and continued thus preserving precious samples.

[0142] MERFISH+ probes allow for a flexible readout

[0143] The stability of the acrydite-conjugated probes combined with the stringent washing conditions, allow for the development of enhanced MERFISH probes (termed MERFISH+ probes) in which the readout strategy of the original MERFISH design was changed for increased flexibility. The original MERFISH probe design (FIG. 3A, top-left) has built-in read sequences targeting many different genes (FIG. 3A, bottom-left). This design does not enable individual genes to be imaged separately and if any mistake were introduced during the design (i.e. the accidental inclusion of a gene with high expression) the MERFISH probes need to be reordered and resynthesized. In contrast, the MERFISH+ probe design added built-in unique read sequences for all the probes targeting each gene (FIG.3A, top-right). This allows MERFISH+ probes for any gene or subset of genes to be readout using either serial smFISH or using MERFISH with an adaptive combinatorial readout strategy (FIG.3A, bottom-middle and bottom-right).

[0144] Applicant demonstrated these capabilities by designing and ordering a library of MERFISH+ probes targeting a comprehensive set of 1799 human genes hybridized to 16-um human developing heart sections. Applicant first selected 15 well established marker genes in the human heart, such as IRX1, PENK, MPZ, TCF21, and GJA5 (Farah et al. 2024) imaged across 5 hybridization cycles using single-molecule FISH readouts binding sequentially to each probe set of each gene (FIG.3A, FIG.3B, FIG.3C). The spatial distribution of the single-cell expression of these genes matched (FIG.3B, FIG.3C, FIG.8A) prior publications and the average expression per cell of each gene matched quantitatively the prior results with a Pearson correlation of 0.8 (FIG.8B) For instance, IRX3 is localized to the right-trabecular ventricular cardiomyocytes (Christoffels et al. 2000; Farah et al. 2024)), HEY2 is localized to the ventricular compact myocardium and interventricular septum and MPZ is present in the neuronal / progenitor cells in the atrium of the heart.

[0145] The MERFISH+™ probes also facilitate designing custom combinatorial readout strategies for faster quantification of the genes of interest in each cell. To demonstrate this, Applicant first selected the 238 genes Applicant previously imaged (Farah et al. 2024) andAtty. Dkt. No.: 114198-3660Applicant built a custom pipetting robot programmed to mix combinations of readout probes for each cycle of hybridization. The pipetting program corresponds to a predefined MERFISH combinatorial codebook such that, for instance, a value of 1 in columns 1, 2, 8 and 16 for gene IRX3 means that the readout probes for IRX3 were pipetted and mixed for hybridization cycles 1,2,8 and 16 (FIG. 3D, FIG. 3E). Samples were imaged across two sets of 16 hybridization cycles corresponding to two different codebooks (Codebookl and Codebook2) (FIG.3F, FIG.3G). Quantification of the average number of mRNA per cell was correlated between Codebookl and Codebook2 (Pearson correlation coefficient of 0.815 FIG. 3H) and agreed with previous experiments ((Farah et al. 2024) and FIG. 8C, FIG. 8D, FIG. 8E). A near identical cell type definition and gene expression per cell type was observed with the two codebooks. These results demonstrated that MERFISH+™ allows for the flexibility of hybridizing a sample with a large library of probes (-2,000 genes) followed by a defined selection and adaptation to the genes of interest.

[0146] MERFISH+ allows for scaling up the number of genes imaged

[0147] The increased stability and the flexibility of the readout design of the MERFISH+ probes allow for a scalable increase in the number of genes imaged to thousands of targets. Applicant continued imaging the heart sample in FIG. 3 across a total of 1,800 genes using 8 modules of 16 hybridization cycles each which included many of the previously imaged genes (FIG. 4A). Applicant defined the cell types within the 2D section of the developing heart and obtained a good correlation of gene expression for matching cell types when comparing with single-cell RNA sequencing (Farah et al. 2024) (an average Pearson correlation of 0.8) (FIG.4B). Interestingly, upon increasing nearly 10-fold the number of genes imaged Applicant largely defined the same cell types as with the reduced library of genes. This is consistent with previous observations in the murine brain by the Allen Institute where by imaging an optimum group of -500 to 1,000 marker genes the cell type definition of MERFISH matched the comprehensive single-cell sequencing cell type definition across the entire murine brain. However, Applicant noticed two cell types, not previously captured in Applicant’s published work (Farah et al. 2024): the Dorsal Mesocardium Progenitor (DMP) cells specific to the atrium (FIG. 4C, middle) and vCM-LV-papillary cells (FIG. 4C, right). DMP cells are spatially located along the dorsal wall of the developing atrium. These cells originate from the dorsal mesocardium and migrate to the atrial region, where they play a crucial role in forming theAtty. Dkt. No.: 114198-3660atrial septum and other atrial structures. DMP cells are positioned near the junction between the venous pole and the atrial myocardium, and have been considered to play a role in developing atrial tissue by promoting structural organization and cellular differentiation. vCM-LV-papillary cells are spatially located within the left ventricle of the human heart, specifically concentrated in the papillary muscle regions. These cells form along the inner ventricular wall and extend toward the papillary muscles, which are anchored to the ventricular myocardium. Positioned in close proximity to the chordae tendineae, vCM-LV-papillary cells contribute to the structural integrity and contractile function of the papillary muscles, which play a crucial role in valve mechanics by preventing mitral valve prolapse during ventricular contraction. Their specialized spatial arrangement supports their function in coordinating with other ventricular cardiomyocytes to maintain effective ventricular contraction and blood flow dynamics in the left ventricle. For each of these specific cell types, Applicant highlighted the gene expression of the most specific maker genes (FIG. 4D, FIG. 4E). By measuring more comprehensively the spatial distribution of genes across the developing heart Applicant asked which genes display the most gradient expression along the major axis (basal-apical axis) and the left-right axis FIG. 4D, FIG. 4E. By focusing on ventricular cardiomyocytes Applicant noticed that HAND1, PRX1 and PITX2 are specific to the left / right markers whereas RGS6 is a potential conduction system regulator; GATA2 is a potential valve regulator. Together, Applicant show that MERFISH+ with the expanded library identified largely consistent cell types validating the approach. It also enabled discovery of new cell types providing a more comprehensive molecular map of the developing heart.16148] MERFISH+ enables scaling up the imaging area by 10-fold10149] The increased stability of the new probe design allowed for modifying the underlying microscopy-microfluidics system used for MERFISH to enable an increase in the area profiled per experiment. Applicant built a new microscope system with a more powerful laser system, a larger format camera and, importantly, a redesigned microfluidics chamber which combined allowed for a 10-fold larger imageable area compared to the previous experiments (Farah et al., 2024) (see Methods for details) (FIG. 5A, FIG. 5B). Using this system, Applicant imaged 238 genes across a total of 53 consecutive coronal sections from a 12-pcw human heart donor across two experiments (21 sections spanning the atrial half of the heart and 32 sections spanning the ventricular half, FIG. 5C). This covered the entire fetal heart with 120 umAtty. Dkt. No.: 114198-3660resolution along the basal-apical axis amounting to 3.1 million cells (after filtering from 3.1 million cells) profiled within only 2 MERFISH experiments. Applicant obtained a good overall correlation between gene expression with prior single-cell sequencing (supplemental Fig. S4a) or MERFISH performed on 2D sections of the matching cell types.

[0150] Applicant first defined all the major cell types across all the 3.1 million cells imaged in the heart based on the transcriptional profiles in each cell (FIG.5D, FIG.5E). Next, Applicant sub clustered the major cell types based on a combination of gene expression and spatial location of each transcriptional subcluster. Applicant identified a total of 44 distinct cell clusters with spatially organized patterns, highlighting the complexity of cell types and their specific arrangements to support functional specialization. For example, beyond the previously described ventricular cardiomyocyte (vCM) clusters exhibiting left and right spatial patterns (Farah et al., 2024) — including compact and trabecular left ventricular cardiomyocytes (LV-CM), right ventricular cardiomyocytes (RV-CM), proliferating vCMs, and vCMs in the His-Purkinje conduction system — Applicant identified additional vCM subclusters along the base-to-apex axis. These subclusters include vCM-Apex and vCM-Base outflow tract (OFT) populations, which are more enriched in the right ventricle, consistent with the progressive maturation of the right ventricular region (FIG. 12A, FIG. 12B).

[0151] Furthermore, Applicant identified a small vCM cluster located at the junction between the cushion / valve and the trabecular cardiomyocytes in the left ventricle (LV), annotated as vCM-LV-papillary. Interestingly, this cluster was observed in only a few LV sections, potentially explaining its absence in the earlier 3D section MERFISH study. Additionally, Applicant observed a higher density of vCM-LV-hybrid cells in the dorsal interventricular septum (IVS) region, where the IVS appears enlarged. This may signify a rapidly expanding IVS cardiomyocyte population, with hybrid vCMs potentially playing a pivotal role in IVS formation. These spatially organized vCM clusters underscore the complexity of ventricular cell organization and may be integral to the processes of ventricular cell assembly and development.

[0152] In contrast to vCMs, which exhibit pronounced left and right spatial patterns, non-cardiomyocytes, such as fibroblasts and endocardial cells, show less distinction between the left and right sides. Instead, these cells demonstrate more significant differences between atrial and ventricular regions. Notably, only ve-endocardium-left is highly enriched in the leftAtty. Dkt. No.: 114198-3660ventricle. Beyond the atrial and ventricular endocardium and the v-Endocardium-proliferating clusters described previously, Applicant identified additional region-specific endocardial clusters, including VEC and great artery endocardium, localized to the cushion and great artery valve regions. For fibroblasts, in addition to a-Fibro and v-Fibro populations, Applicant identified several distinct subpopulations with specific spatial patterns, such as trabecular fibroblasts, OFT-fibroblasts, vessel fibroblasts, sinus venosus fibroblasts, and endocardium-associated fibroblasts. These spatial differences appear linked to their developmental origins. For instance, it has been reported that endocardium-associated and trabecular fibroblasts likely originate from the endocardium, while vessel fibroblasts are derived from epicardium-associated cells. Most of these small subpopulations are present only in a limited number of sections, which may explain why they were overlooked in the Applicant’s previous study that relied on only three sections for MERFISH analysis. This finding emphasizes that analyzing more sections can uncover a greater diversity of cell types, further enhancing the understanding of their spatial organization and functional roles.

[0153] The 53 sections were aligned to each other using a computational approach called Spateo (X. Qiu et al. 2022) thus resulting in the first comprehensive 3D fetal heart reconstruction with single-cell resolution (FIG. 5G). This comprehensive 3D data allows for observing the continuity of the various structures as well as quantifying the physical parameters of each structure in the heart including the volumes occupied by each major community structure previously defined FIG.5H) as well as how these structures change shape in 3D. For example, by analyzing the boundary between the left and right ventricles, Applicant quantified wall thickness along the long axis, from apex to base (FIG. 51). Applicant observed that the dorsal LV is significantly thicker than the anterior LV, and the base is notably thicker than the apex.

[0154] Applicant next focused on two structures that are difficult to capture comprehensively based on single 2D sections: the valve system and the Coronary arteries system. Applicant defined the four major valve structures: the mitral, tricuspid, pulmonary, and aortic valves based on a combination of transcriptional and spatial features (see Methods for details). (FIGS.6A- 6C). These 4 valves spanned a total of 10 slices, i.e. 1,100 pm along the long axis; Within each valve Applicant quantified the cell community in the valve and showed the cell composition in valve and cushion structure (FIG. 6B). Applicant found that the valve cushionAtty. Dkt. No.: 114198-3660cells contain neural cells, indicating that the valves and cushions have been innervated by the nervous system at this stage already. Additionally, the tricuspid, mitral, and aortic valves showed a higher abundance of VCM-LV-AV cells (wait until Yifan defines percentage correctly), consistent with their association with these valve structures. Comparing gene expression across matching cell types, Applicant observed significant differences between VIC cells in the aortic (AV) and pulmonary (PV) valves versus those in the mitral (MV) and tricuspid (TV) valve cushions. These differences may be linked to the distinct developmental origins of the various regions, as vascular interstitial cells (VICs) in the pulmonary vein (PV) and atrioventricular (AV) regions are neural crest-derived, whereas part of the cushion in the mitral valve (MV) and tricuspid valve (TV) originates from the epicardium and endocardium. This is supported by the high expression of UPK3B, an epicardium marker, and Angptl, a secreted factor that regulates endothelial growth, in these cushions (FIG. 6D). In contrast, Sema6d, a marker of neural crest cells (NCC), is more highly expressed in the PV and AV regions.

[0155] The 3D heart reconstruction allowed us to trace the vascular smooth muscle cells (VSMCs expressing MYH11), defining the origin of the left and right coronary arteries from the aorta. These arteries target the left and right sides of the ventricle surface and then split into several branches to cover the entire left and right ventricles, spanning from the apex to the base of the heart (FIG. 6H-6K). Interestingly, one branch enters the interventricular septum (IVS), where the IVS shows an expanded region. Applicant defined the community of cells surrounding the blood vessels and analyzed their cellular composition along the basal-apical axis. Applicant found that EPDC, Neural, pericytes, LEC, BEC, accompany SMCs, consistent with their function in angiogenesis and the connection between blood vessels and nerves.

[0156] Applicant found that EPDC cell types were enriched at the basal blood vessels while vEndocardium, proliferating vCM cells types were enriched at the tip. By integrating the 3D heart MERFISH data with single-cell RNA sequencing data Applicant interrogated more comprehensively the transcriptional changes along the vessels of the surrounding cell types. Applicant noticed for instance that known angiogenesis makers antithetical increased / decreased along the basal apical direction. Taken together, these data demonstrate the capability of MERFISH+ to define the anatomical and molecular complexity within theAtty. Dkt. No.: 114198-3660developing human heart, connecting structural, cellular and transcriptional features underlying cardiac function.

[0157] MERIFSH+ allows for facile multimodal imaging of RNA and DNA

[0158] Understanding gene expression of over 2,000 genes greatly facilitated the identification of cell types and cell states during the development, in comparison with the published 250 gene panel. Multi-modality sequencing provides additional layer of chromatin states not only shedding light on cell states but also understand the underlying gene regulation mechanism. To this end, Applicant selected one gene locus that demonstrates the most dynamic changes over human cardiomyocyte lineage specification, the myosin-heavy-chain 6 and 7 (Myh6 / 7). Applicant first applied chromatin tracing to measure the 3D genome organization of the Myh6 / 7 locus in a 12-pcw human heart section.

[0159] Applicant hybridized the section with ~ 24,000 oligonucleotide probes targeting this entire 84kb locus (FIG. 7A). Using sequential hybridization and imaging, Applicant readout the signal of each one out of 2-kb 42-segments comprising this locus in a sequential manner (FIG. 7A, FIG. 7B). By fitting the signal of each locus in each chromosome in each nucleus and precisely aligning each of the sequential images the 3D configurations that the Myh6-7 locus adopts cell-by-cell is reconstructed (FIG. 7B, FIG. 7C) with ~50 nm spatial resolution and 2 kb genomic resolution. On a single-cell level Applicant noticed that this locus tended to organize within globular domain-like structures (FIG. 7C) reminiscent of TAD-like structures previously reported (22). Upon averaging the inter-distances between each of the 2-kb segments across -20,000 cells Applicant noticed a few structural features appearing (FIG.7D): the ends of Myh6 promoter acted as additional domain boundaries and together with some internal intronic regions within the gene body displayed loop-like behavior, forming long-range interactions with other loci within the larger domain. Applicant next correlated the ensemble chromatin structure with those from published Hi-C (Y. Zhang et al. 2019). Applicant compared this population average distance map with the chromatin contacts for the same locus measured with Hi-C (FIG. 7E). Applicant saw a remarkable agreement between the Hi-C data and the distance data (compare FIG. 6D and 6E), cross-validating the two orthogonal techniques: Hi-C and Chromatin tracing. Applicant asked if the heterogeneity in single-cell structures (FIG. 7C) arises in part because of cell-type specific differences. Crucially, within the same cells Applicant measured the expression of 250 genes via the same MERFISHAtty. Dkt. No.: 114198-3660strategy as in FIG. 5, FIG. 7H. These genes covered major cell type markers, atrium cardiomyocytes, ventricular cardiomyocytes, etc., sub-type markers. Based on the MERFISH single-cell expression Applicant clustered the -24,000 cells imaged into -20 clusters covering the major cell types and multiple cell types that correspond to all the cell types as in FIG. 41.Within each cell type Applicant calculated the median distance matrices and contact matrices of the Myh6 / 7 locus across the different cell types and plotted the corresponding contact matrices. ( FIG.71). Clearly, Applicant identified a loop formed between Myh6 promoter and two downstream cis-elements El and E2 in atrium-cardiomyocytes. In stark contrast, the same El and E2 cis-elements form another loop with Myh7 promoter instead in ventricular cardiomyocytes. The loop formation varies in other cell types, such as aFibro, vFibro. Interestingly, the differences among different cell types highly correlates with those of the physical zoning distances, demonstrating the high resolution of the chromatin tracing and true single-cell resolution in MERFISH+.

[0160] Applicant also imaged the 1.7Mb TAD in the same manner by using sequential hybridization and imaging, Applicant readout the signal of 50kb-segments comprising this locus in a sequential manner. Upon averaging the inter-distances between each of the 50-kb segments across -20,000 cells Applicant noticed a few structural features appearing (FIG.7K): including forming long-range interactions with other loci within the larger domain. However, the ensemble structures show similarity among different cell types.

[0161] To further understanding the promoter-enhancer activity in different cell types, Applicant applied antibody staining on the same section with 6 different epigenomics markers: pol2ser2, pol2ser5, H3K9me3, H3K4me3, H3K27acetal, LaminA (FIG. 7L).

[0162] Discussion

[0163] Technologies in spatially resolved single-cell genomics are transforming many aspects of biological research. Here, MERFISH+ has been developed to provide significantly more stability by using acrydite-modified probes, which anchors the single-strand DNA probes onto the biogel film and eliminates the requirement of a good RNA integrity during prolonged imaging. Due to the near unlimited stability of the system, the sample can be imaged with a variety of codebooks to meet the need of flexibility for adjusting to image the correct gene expression levels or being focused on different cell types.Atty. Dkt. No.: 114198-3660[01641 In addition, the expansion to up to 4000 genes for this case is easy. This improvement is eliminating the limitation of targeted spatial transcriptomics technology without sacrificing the detection efficiency or being limited by optical crowding. To further enhance the flexibility Applicant imaged human tissues, both heart and brain, in a modular fashion. Specifically, Applicant imaged 1826 genes with 6 different modules, which allows each module imaged with a 48-bit codebook. Applicant computational analyzed each module sequentially and also have the freedom of alternating the imaging order of the total 6 modules modifying the codebook as needed.

[0165] The improvement of the imaging hardware including upgrading cameras with larger imaging area, higher power lasers and enhanced fluidics system with a chamber imaging up to 5cm x 3.5cm capacity, specifically developed a new microscope-microfluidics system which built on top a more flexible Applied Scientific Instrumentation microscope body equipped with higher-power lasers from Lumencor (up to -1-2.5W per laser line) coupled with 0.4mm square multimode fiber, a larger-format camera (20.8 mm sensor) (Teledyne-Kinetix), and an upgraded microfluidics chamber (Bioptechs), achieving a 10-fold increase in imaging area compared to previous and commercially available MERFISH imaging platforms. Applicant imaged 32 tissue sections from one single human developmental heart in one imaging session, profiled 2.3 million MERFISH cells at single-cell resolution. This represents 10 times scaling up in comparison with the previously published work (Farah et al., 2024, Huang et al., 2021). When Applicant imaged another 21 sections, all together the 53 slices allowed us to perform 3D reconstruction of the first human heart, to Applicant’s knowledge, based on spatial transcriptomics-computed cell types. Because these tens of tissue slices were imaged within one experiment, realignment and stitching during 3D reconstruction becomes easy. Applicant expands the previous finding of a hybrid cell type displaying transitional property between trabecular and compact layers that are specific to 10-15 pew human heart development. In the context of 3D along the long axis, these hybrid cells form a continued thin layer becoming thicker towards atrium proportionally to the thickness of the left ventricular wall; Applicant was also surprised to locate traces of hybrid cells in the right ventricle.

[0166] The 10 times higher imaging area also significantly expands the potential for application of spatial transcriptomics in constructing tissue atlas in large organs, such as adult human heart, brain. This is also a quantum leap to apply spatial transcriptomics in the clinicAtty. Dkt. No.: 114198-3660samples to examine pathological samples. Human tissues are often harder to obtain in sufficient quantities and from diverse donors compared to mouse tissues, which can be bred and sacrificed in controlled laboratory settings. This limitation can restrict the sample size and diversity in human tissue studies. Human tissue samples can vary widely in quality due to differences in collection, preservation, and storage methods. Factors such as post-mortem interval, the condition of the tissue at the time of collection, and handling procedures can affect RNA integrity and overall sample quality, potentially impacting the results of spatial transcriptomics assays.

[0167] Multimodal: To more completely define a cell type and its associated cell state, integration of gene expression, chromatin structure and proteomic markers within the same one single cell are highly desirable. MERFISH+ measured expression of 2000 genes, 99 introns and 6 antibodies targeting histone marks, non-histone chromatin binding proteins; lastly, the associated chromatin structure of chr. 14 with 2kb-20kb resolution. The 2-kb resolution revealed a fine-detailed structure within 84kb locus and fine difference in chromatin structure in left ventricle wall and right atrium. The 20-kb resolution of 1MB chromatin structure, 10-kb resolution of 500MB.

[0168] Computational algorithm

[0169] The sheer volume of the data from the 3D imaging poses a severe challenge for data storage, active data analysis and sharing. Although Applicant has adopted an open source python format zarr and compress the data by 50% in comparison with the original DAX file (Moffit, science, 2018), construction of the 3D heart (25%) at a z-resolution of 100 um represents 3.5 million MERFISH cells and 140TB storage space, which is prohibitively expensive using either Network Attachment System (NAS) or cloud storage (AWS).

[0170] Comparing TCEP Cleavage with Formamide Stripping

[0171] The acrydite experiments utilize TCEP Cleavage, whereas the reimagings use formamide stripping to remove the probes. Applicant performed one more set of experiments, R123_N3S5Acrycomb and R123_N3S5AcrycombFMReImage, an acrydite imaging and reimaging, respectively. Results from both experiments, using a flat field correction on DAPI images for segmentation (see Computational Methods: Flat Field Corrections).Atty. Dkt. No.: 114198-3660

[0172] TCEP detection efficiency can be affected by incomplete removal of the readout adaptors. However, 100% formamide stripping could cause deformation of the biogel film. Therefore, a robust drift correction computation step is much desired.

[0173] The following examples are provided to illustrate, but not limit the scope of the claims.

[0174] Methods and Materials:

[0175] Tissue samples

[0176] De-identified tissue samples were collected with previous patient consent in strict observance of the legal and institutional ethical regulations.

[0177] Tissue processing

[0178] Tissue samples were dissected in buffer containing 125 mM NaCl, 2.5 mM KC1, 1mM MgCl₂, and 1.25 mM NaH₂PO₄ under a stereotaxic dissection microscope (Leica).

[0179] For single cell dissociation, tissue samples were further cut into small pieces and enzymatically digested by incubating with collagenase type IV (Gibco) and Accutase (ThermoFisher) at 37°C for 60 min. After removing the dissociation media, cells were resuspended in PBS supplemented with 5% FBS and sorted on a Sony SH800 sorter. Samples were diluted to approximately 1,000 cells per pl before processing for scRNA-seq.

[0180] Samples for MERFISH were washed 1 time with ice-cold PBS, then fixed in 4% PFA at 4°C overnight. On the second day, the sample was washed in ice-cold PBS three times, 10 minutes each, and were incubated in 10% and 20% sucrose at 4°C for 4 hours each, and in 30% sucrose overnight, followed by immersion with OCT (Fisher, cat# 23-730-571) and 30% sucrose (1 vol:l vol) for 1 hour. The sample was then embedded in OCT and stored at -80°C until sectioning.

[0181] One 12 pew human heart donor was sectioned into the atrial half and ventricular half. The atrial half was sliced into 21 sections and mounted onto one big coverglass and was fit onto 8 individual coverglasses. The ventricular half was sliced into 32 sections mounted onto one big cover glass and was fit onto 8 individual cover glasses. The total number of sections is 53x 8=424 sections with a mean thickness of 16 um, spanning a total length of 17 mm. 20-30 sections were collected on the salinized cover-glass (Chen et al. 2015) and mounted into the new fluidics system. Subsequently, one coverglass was subjected to rapid fixation, stainedAtty. Dkt. No.: 114198-3660with the same 238 genes (Farah et al. 2024). Image processing and decoding was performed using a modified version of MERlin (doi: 10.5281 / zenodo.3758540), which integrates Cellpose for cell segmentation and includes an alternative algorithm for drift correction that uses DAPI staining rather than fiducial beads and can correct for drift along the z-axis. Applicant then performed a 3D reconstruction of the tissue sections using Spateo (C. Qiu et al. 2024).

[0182] MERFISH gene selection and Probe Library Design and Construction

[0183] Traditional way: To identify the transcriptionally distinct subpopulations identified in the scRNA-seq dataset, Applicant designed a panel of 238 genes to be imaged with combinatorial barcoded imaging (Moffitt Science 2018). Applicant first identified differential gene markers for each of the 75 subpopulations in the scRNA-seq with differential gene expression (DGE) analyses as well as NS-Forest2, combining all markers from the binary gene analysis. The list was then filtered for genes that were either not long enough to construct 60 48 target regions (each 30-nucleotide long) without overlapping or whose expression levels were outside the range of 0.01 to 300 average UMI per cluster, as measured by scRNA-seq. MERFISH assays were then performed with a MERFISH encoding probe set. Briefly, the encoding scheme used a 22-bit MHD4 code to encode the RNAs. In this encoding scheme, each of the 238 possible barcodes required at least four errors to accumulate to be converted into another barcode. This property permitted the detection of errors up to any two bits, and the correction of errors to any single bit. In addition, this encoding scheme used a constant Hamming weight (i.e., the number of “1” bits in each barcode) of 4, to avoid potential bias in the measurement of different barcodes because of a differential rate of “1” to “0” and “0” to “1” errors, as described previously (Moffitt et al., 2016, PNAS, 113 (39): 11046-51). Applicant used 238 of the 7,315 possible barcodes to encode cellular RNAs and chose 10 barcodes were left unassigned to serve as blank controls. The encoding probe set that Applicant used contained 30-48 encoding probes per RNA, with each encoding probe containing three of the four read sequences assigned to each RNA. Encoding probes were designed using a publicly available pipeline (github.com / bil022 / ProbeDesigner). Briefly, transcript sequences were derived from the human reference genome sequences (hg 38) downloaded from ncbi_refseq: hgdownload.soe.ucsc.edu / goldenPath / hg38 / bigZips / genes / . The encoding probes were designed using ProbeDesigner (github.com / bil022 / ProbeDesigner).Atty. Dkt. No.: 114198-3660

[0184] ProbeDesigner was developed using a fully optimized algorithm for both DNA and RNA probes based on the principles of probe design used in various established algorithms [OligoArray; OligoMiner; OligoPaint; ProbeDealer], The principles mainly include the detection of off-targets based on genome-wide 17-mer off-target counts, 30-42 mers RNA / DNA probe sequences, GC content or melting temperature (TM), and the avoidance of repeated regions. ProbeDesigner is implemented with three steps: 1) Build a 17-mer index based on reference genome (DNA) or genome annotation files (RNA). 2) Scan 17-mer counts for selected loci or genome sequences. 3) Filter and rank probe candidate based on pre-defined selection criteria. ProbeDesigner was implemented in C++.

[0185] Probe library design for RNA-MERFISH and chromatin tracing

[0186] Applicant selected 40bp target sequences for DNA or RNA hybridization by considering each contiguous 40bp subsequence of each target of interest (the mRNA of a targeted gene or the genomic locus of interest) and then filtering out off-targets to the rest of the transcriptome / genome including repetitive regions, or too high / low GC content or melting temperature(TM). More specifically, presently disclosed probe design algorithm was implemented with three steps: 1) Build a 17-mer index based on reference genome hsl assembly (DNA) or the hg38 transcriptome (RNA). 2) Quantify 17-mer off-tager counts for each candidate 40bp target sequence. 3) Filter and rank target sequences based on predefined selection criteria as previously described (Huang et al. 2021)(Su et al. 2020) PMID: 32822575.

[0187] DNA probe smTADHg19: chr14:23850001-23935000 (HiC smTAD)Hg19: chr14:23435001-24405000 (HiC bigTAD)

[0188] MERFISH gene selection

[0189] The MERFISH gene selection was performed by first using a BICCN dataset from 18-19 gw brain and using NSForest v2 (ref) with default parameters to identify marker genes for the cell type clusters in this data. This list of genes was supplemented with additional marker genes from the literature (i.e. DCX, GFAP etc.) as well as genes with differential methylation in the snm3C-seq3 data in HPC of mid-gestational human brains. The target sequences for each gene were concatenated with one or two unique read sequences to facilitate MERFISH orAtty. Dkt. No.: 114198-3660smFISH imaging. The final list of encoding probes for RNA imaging used is shown in Table SX1 (Sequences of Encoding probes for RNA MERFISH).

[0190] Design of chromatin probes

[0191] Applicant designed probes for DNA hybridization similarly as those for the RNA MERFISH as described in (Su et al. 2020; Wang et al. 2016). Briefly, Applicant first partitioned chromosome 14 into 50-kb segments and selected for imaging a fifth of these segments uniformly spaced every 250 kb (amounting to 354 target genomic loci using genome reference hsl). After screening against off-target binding, GC content and melting temperature, -150 unique 40 bp target sequences were selected for each 50-kb segment. Applicant concatenated a unique read sequence to the target probes of each segment to facilitate sequential hybridization and imaging of each locus. The final list of encoding probes for DNA imaging is shown in Table SX2 (Sequences of Encoding probes for DNA Imaging).

[0192] MERFISH+ probe 2500-gene library design

[0193] Applicant has designed 60 primary probes (40 nt per read sequence type) for each mRNA target. For each mRNA species, two distinct read sequences have been used to increase the accuracy of RNA molecule efficiency by co-localizing signals, which will be discussed later. mRNA molecules were tethered to the polyacrylamide gel using their polyA tails to minimize their drift during the imaging process. Each primary probe has a 40nt length target sequence designed complementary to the RNA strand. Moreover, each primary probe has three read sequences which Applicant used later for mRNA detection using fluorophore-conjugated oligos. To amplify the library, Applicant also put PCR primers at each end of all primary probes, which enabled us to amplify primary probes to maintain the storage. Applicant used two distinct read sequences per mRNA species and tried to alternate between these two while hybridizing the primary probes to the mRNA molecule. Also, to stabilize and tether the primary probes to the matrix, Applicant used acrydite-conjugated oligos. Due to their molecular structure, acrydite-conjugated oligos could be incorporated into the hydrogel (polyacrylamide), resulting in lower drift between frames and more accuracy in mRNA detection. From a chemical perspective, the double-bond in the molecular structure of the acrydite group is similar to the double bond of acrylamide monomers and reacts with activated double bonds of acrylamide and bisacrylamide (for cross-linking linear polymers ofAtty. Dkt. No.: 114198-3660acrylamide) monomers which results in the incorporation and tethering of acrydite-conjugated oligos in the matrix. By hybridizing the primary probes, both smFISH and MERFISH could be performed to determine the transcriptome of the cells and classify the cell type. For the smFISH (single molecule fluorescence in situ hybridization) experiment, since there are two distinct read sequences per mRNA species, Applicant could use co-localization of signals to increase the detection accuracy. The fluorophore-conjugated oligos in three colors (750nm, 650nm, 560nm) were used for imaging purposes. The currently developing secondary MERFISH probes could be hybridized with the primary ones. Each MERFISH probe is designed in a way that it has four on-bit regions (20nt each), which could be used for detection as shown in FIG.2A.

[0194] Acrydite Probe Synthesis

[0195] The acrydite probes were generated using oligonucleotide pools, following the previously described method (Moffitt et al., 2016). Firstly, Applicant amplified the oligo pools (Twist Biosciences) using a limited cycle qPCR (approximately 15 - 20 cycles) with a concentration of 0.6 pM of each acrydite primer to create templates. These templates were converted into RNAs using the in vitro T7 transcription reaction (New England Biolabs, E2040S). These RNAs were then converted back to single-stranded DNA using reverse transcription with a concentration of 16 pM of each acrydite primer. Subsequently, the DNA oligos were purified using alkaline hydrolysis to remove RNA templates and cleaned with columns (Zymo Research, D4060). The resulting acrydite probes were stored at -20 °C.

[0196] R128: Aery DCSS 50ug;75uL + Aery QZBBL1(CRO9) 50ug;75uL

[0197] R129: NCBB-intron (RNA probes) 1.5ug; 2uL + NCBB-Every-smTAD (DNA probes) 10ug; 25 uL+ NCBB-smTAD2kb 2.5ug; 3 uL

[0198] RNA MERFISH sample preparation

[0199] Fresh frozen hearts were sectioned at -20°C using a Leica CM3050S cryostat. Series coronal sections of 16 pm thickness were performed at -600 pm along the anterior-posterior axis of human hearts to capture the major cardiac structures. Sections were collected on salinized and poly-L-lysine (Millipore, 2913997) treated 40-mm, round #1.5 coverslips (Bioptechs, 0420-0323-2). After drying for 20 min, tissue sections were stored at -80C until use.Atty. Dkt. No.: 114198-3660

[0200] MERFISH measurements of 238 genes with 200 non-targeting blank controls was performed as previously described, using the encoding sequences and published readout probes, subsequently pre-cleared by immersing into 30%(vol / vol) ethanol, 50% (vol / vol) ethanol, 70% (vol / vol) ethanol, and 100% ethanol, air dried for 5 minutes, then, 70% (vol / vol) ethanol, 50% (vol / vol) ethanol, each for 5 minutes. The tissue was then treated with Protease III (ACDBio), at 40 °C, for 30 minutes prior to be washed with PBS for 5 minutes. Then the tissues were preincubated with hybridization wash buffer (40% (vol / vol) formamide in 2x SSC) for ten minutes at room temperature. After preincubation, the coverslip was moved to a fresh 60 mm petri dish and residual hybridization wash buffer was removed with a lab tissue. In the new dish, 50 uL of encoding probe hybridization buffer (2X SSC, 50% (vol / vol) formamide, 10% (wt / vol) dextran sulfate, and a total concentration of 5 pM encoding probes and 1 pM of anchor probe: a 15-nt sequence of alternating dT and thymidine-locked nucleic acid (dT+) with a 5'-acrydite modification (Integrated DNA Technologies). The sample was placed in a humidified 47°C oven for 18 to 24 hours then washed with 40% (vol / vol) formamide in 2X SSC, 0.5% Tween 20 for 30 minutes at room temperature. Samples were post-fixed with 4% (vol / vol) paraformaldehyde in 2X SSC and washed with 2X SSC with murine RNase inhibitor for five minutes. To anchor the RNAs in place, the encoding-probe-hybridized samples were embedded in thin, 4% poly-acrylamide (PA) gels as described (Farah et al, Nature 2024). Briefly, the hybridized samples on coverslips were washed with a de-gassed 4% polyacrylamide solution, consisting of 4% (vol / vol) of 19:1 acrylamide / bis-acrylamide (BIORAD®, 1610144), 60 mM Tris-HCl pH 8 (ThermoFisher®, AM9856), 0.3 M NaCl (ThermoFisher®, AM9759) supplemented with the polymerizing agents ammonium persulfate (Sigma, A3678) and TEMED (Sigma®, T9281) at final concentrations of 0.01% (wt / vol) and 0.05% (vol / vol), respectively, for 10min to equilibrate. Then the sample was reverted onto a glass plate that is cleaned and siliconized and spotted with 300ul of the 4% (vol / vol) of 19:1 acrylamide / bis-acrylamide (BioRad, 1610144), 60 mM Tris-HCl pH 8 (ThermoFisher®, AM9856), 0.3 M NaCl (ThermoFisher®, AM9759) supplemented with the polymerizing agents ammonium persulfate (Sigma, A3678) and TEMED (Sigma, T9281) at final concentrations of 0.03% (wt / vol) and 0.15% (vol / vol), respectively. The gel was then allowed to cast for 5 h at room temperature. The coverslip and the glass plate were then gently separated, and postfixed with 4% PF A, then the PA film was incubated with a digestion bufferAtty. Dkt. No.: 114198-3660consisting of 50 mM Tris-HCl pH 8, 1 mM EDTA, and 0.5% (vol / vol) Triton X-100 in nuclease-free water and 1% (vol / vol) proteinase K (New England Biolabs, P8107S). The sample was digested in this buffer for >36h in a humidified, 37°C incubator and then washed with 2 x SSC three times.

[0201] MERFISH measurements were conducted on a home-built system.

[0202] Multi-modal Chromatin tracing sample preparation

[0203] The sample preparation is similar to that for RNA MERFISH sample preparation with the exception that the sample is treated with 0.5% Triton X-100 in lx PBS for 10 minutes at room temperature instead of ethanol dehydration. Briefly, Frozen hearts were sectioned at -20°C using a Leica® CM3050S cryostat. Series coronal sections of 14 pm thickness. Sections were collected on salinized and poly-L-lysine (Millipore, 2913997) treated 40-mm, round #1.5 coverslips (Bioptechs, 0420-0323-2). After drying for 20 min, Fixed tissue sections were permeabilized with 0.5% Triton X-100 (Sigma-Aldrich, T8787) in lx PBS with RNase inhibitors for 10 min at room temperature and washed once with IX PBS with RNase inhibitors for 3 min. After permeabilization, coverslips were moved to a fresh 60 mm petri dish and treated with 0.1M hydrochloric acid (Thermo Scientific™, 24308) for exactly 5 min at room temperature, and washed three times for 5 min with lx PBS with RNase inhibitors. Tissue sections were preincubated with the hybridization wash buffer (50% (vol / vol) formamide (Ambion, AM9342) in 2x SSC (Corning, 46-020-CM)) for 10 min at room temperature. Prepared a layer of fresh parafilm with a 0.8 cm diameter hole in the center and placed in a fresh 60 mm petri dish. Then 50 uL of encoding probe hybridization buffer (50% (vol / vol) formamide (Ambion, AM9342), 2x SSC (Corning, 46-020-CM), 10% Dextran Sulfate (Millipore, S4030), 1.5 pg of NCBB-intron, 6 pg of NCBB-exon, 10 pg of NCBB-Every-smTAD, 2.5 pg of NCBB-smTAD2kb, 50 pg of QZBB 5 pM encoding probes and 1 pM of anchor probe, a 15-nt sequence of alternating dT and thymidine-locked nucleic acid (dT+) with a 5'-acrydite modification (Integrated DNA Technologies)) was added to the center. Coverslips were wiped away excess hybridization wash buffer with Kimwipes and placed tissue-side-down onto the droplet. Samples were in a humidity chamber at 47 °C for 18 - 24 hours then washed with 40% (vol / vol) formamide in 2X SSC, 0.5% Tween 20 for 30 minutes at room temperature. Samples were post-fixed with 4% (vol / vol) paraformaldehyde in 2X SSC andAtty. Dkt. No.: 114198-3660washed with 2X SSC with murine RNase inhibitor for five minutes, as described for RNA MERFISH sample.

[0204] Sample preparation for multiplexed imaging experiments

[0205] The samples were prepared similarly as previously described for cell culture samples with notable modifications (Su et al. 2020). Briefly, fresh frozen brain tissues were sectioned into coronal sections of 18 pm thickness at -20°C using a Leica CM3050S cryostat. Sections were collected on salinized and poly-L-lysine (Millipore, 2913997) treated 40-mm, round #1.5 coverslips (Bioptechs, 0420-0323-2). The tissue sections were fixed with 4% PFA (Electron Microscopy Sciences, 15710) in lx PBS with RNase inhibitors (New England Biolabs, M0314L) for 10 min at room temperature before being permeabilized with 0.5% Triton X-100 (Sigma-Aldrich, T8787). Then coverslips were treated with 0.1M hydrochloric acid (Thermo Scientific™, 24308) for 5 min at room temperature. Tissue sections were next incubated with pre-hybridization buffer (40% (vol / vol) formamide (Ambion, AM9342) in 2x SSC (Corning, 46-020-CM)) for 10 min. Then 50 uL of encoding probe hybridization buffer (50% (vol / vol) formamide (Ambion, AM9342), 2x SSC, 10% Dextran Sulfate (Millipore, S4030) containing 15 pg of RNA encoding probes for the targeted genes and 150 pg of DNA encoding probes for the targeted chrl4 loci was incubated with the sample first at 90C for 3 minutes, followed by 47 °C for 18 hours. The sample was then washed with 40% (vol / vol) formamide in 2X SSC, 0.5% Tween 20 for 30 minutes before being embedded in thin, 4% polyacrylamide (PA) gels as described (Moffitt et al. 2018).

[0206] Imaging and adaptor / readout hybridization protocol

[0207] MERFISH measurements were conducted on a custom microscope-microfluidics system with the configuration previously described (Huang et al. 2021)(Su et al. 2020). Briefly, the system was built around a high-speed and high-sensitivity camera (the Kinetix model from Teledyne) capable of imaging twice the area of the prior model of Hamamatsu CMOS cameras. A newer line of Nikon 60x oil immersion objective (1.4 NA) was also used for better spherical and chromatic aberration corrections. The different components were synchronized and controlled using a National Instruments Data Acquisition card (NI PCIe-6353) and custom software (Su et al. 2020).Atty. Dkt. No.: 114198-3660

[0208] To enable multi-modal imaging Applicant first sequentially hybridized fluorescent readout probes and then imaged the targeted genomic loci and then the targeted mRNAs. Then Applicant performed a series of antibody stains and imaging. Specifically the following protocol was used in order:1) 52 rounds of hybridization and chromatin tracing imaging, sequentially targeting the 1000 loci using 2-color imaging (Su et al., 2020)2) 16 rounds of hybridization and MERFISH imaging, combinatorially targeting 298 genes using 3 -color imaging3) 3 rounds of sequential staining for 6 different antibodies using 2-color imaging.

[0209] The protocol for each hybridization included the following steps:1. Incubate the sample with adaptor probes for 30 minutes for DNA imaging or 75 minutes for RNA imaging at room temperature.2. Flow wash buffer and incubate for 7 minutes3. Incubate fluorescence readout probes (one for each color) for 30 minutes at room temperature.4. Flow wash buffer and incubate for 7 minutes.5. Flow imaging buffer. The imaging buffer was prepared as described previously (Su et al. 2020) and additionally included 2.5μg / mL of DAPI.10210] Following each hybridization the sample was imaged and then the signal was removed by flowing 100% formamide for 20 minutes and then re-equilibrating to 2XSSC for 10 minutes. RNA-MERFISH measurement of the 298 gene panel was performed with an encoding scheme of 48-bit binary barcode and a Hamming weight of 4 (HW4). Therefore, in each hybridization round, ~75 adaptor probes were pooled together to target a unique subset of the 298 genes. The genes targeted in each round of hybridization are FOXCI, KIF26B, PTPRC, UPK3B, TNFRSF12A, TFPI2, TLL2, PGF, NKD2, SCN5A, FLRT1, C1QTNF3, MYBL2, NOS3, XPO4, HEY1, DAPK2, IRX4, ADAMTS6, LM0D3, TNNT1, IRX1, ADM, RAD51AP1, RND3, ABCC9, FLRT3, ANGPT1, ASCL1, BTG1, CLDN5, CLEC10A, CLSPN, COBLL1, ETV1, FNDC1, FZD1, HAPLN1, IGFBP4, IGFBP5, JAG1, MAF, MCAM, MFAP5, MRC1, MYH11, NOTCH1, NRXN1, NTM, NTRK3, PECAM1, PRSS23, SCN7A, SFRP1, SLC1A3,Atty. Dkt. No.: 114198-3660TCF21, TNC, VWF, PIEZ02, PAM, CD34, LBH, SERPINI1, INA, SH0X2, ARL6IP1, FN1, RABGAP1L, INMT, CAV1, FBLN5, NKX2-5, ADGRL2, CGNL1, TENM3, RBP1, MMP11, AGT, CRABP2, PCNA, GJA1, PITX2, TBX3, 0LFML2B, C0L2A1, COL6A3, SLC26A7, DHRS3, C0L15A1, AGPS, ENTPD2, HLA-DQB1, KIF20A, CPNE3, RASL11B, PTCH2, FRZB, HAND1, CDT1, HCN4, COL9A2, IRX3, PPM1K, RPRM, SLITRK6, PLK1, HEY2, TNFSF14, MSC, APBB1IP, C1QB, CD9, CERCAM, CPE, CXCL12, DES, EDNRA, FGF12, FREM2, HHIP, IL1B, KLF2, KLF4, LAMC3, LYVE1, MOXD1, NAVI, NRP2, PDE1A, PDGFRA, PDGFRB, PDLIM3, RRM2, RYR2, SEMA6D, SMOC1, SULF1, TBX5, TFF3, TOP2A, TRPS1, VSNL1, WT1, CD24, PRSS35, CTSC, ADAMTS8, CKMT2, NR2F1, NSG1, COLECI 1, COTL1, RSPO3, MEF2C, SPON2, COL14A1, PLK2, TENM4, H0XA5, CACNA1C, GYPC, NPR3, ADGRL3, RRAD, DKK3, SBSPON, TECRL, PRPH, EGR1, TBX18, FBLN2, ADGRL1, TMEM40, RMI2, OSR1, TPBG, TMEM176B, KCNH2, UHRF1, OAF, MSX2, IRX2, CTSV, MCM7, INHBA, CDO1, NDP, PPP1R12B, RCAN1, BMP2, SHISA3, COL26A1, APOE, ARHGAP18, ASPN, BRINP3, CA4, CBLN2, CD36, COL5A2, DLK1, ELAVL2, F13A1, GAS1, GJA5, LEPR, LRP1, LUM, MECOM, MKI67, NEFL, NELL2, PROXI, PRRX1, RGS5, SOX9, SYK, TBXAS1, TRPM3, TSHZ2, VCAN, PTTG1, FMOD, DPYSL3, ECE1, BAMBI, ARHGAP29, KPNA2, KCNJ8, CNN1, TSPO, NR4A3, TENM2, PLN, PENK, GAS7, HAND2, FOXS1, MAZ, ALDH1A2, RAMP1, PCDH7, FILIP1L, NR2F2, MPZ, FLRT2, ISL1, CASQ2, NTS, FBN2, and PKP2.

[0211] Immunofluorescence staining

[0212] Antibody imaging was performed immediately after completing the DNA and RNA imaging. The sample was first stained for 4 hours at room temperature using 2 primary antibodies of two different species (mouse and rabbit), washed in 2X SSC for 15 minutes and then stained for 2 hours using 2 secondary antibodies for each target species conjugated with fluorescent dyes. The table below (Table 1) lists all the antibodies used in this study.

[0213] Table 1Antibody Company Catalog number Dilution Reference Mouse Histone H3Cat# C15200146,trimethylated atDiagenode RRID:AB_292765 1:300 (Takei et al. 2021) lysine 90(H3K9me3)Atty. Dkt. No.: 114198-3660Rabbit Anti-RNApolymerase II CTDCat# ab 193468,repeat YSPTSPSAbeam RRID:AB_290555 1:400 (Takei et al. 2021) (phosphor S2)7antibody[EPR18855]Mouse monoclonal[SC-35] to SC35 - Cat# ab11826, (Su et al. 2020)Abeam 1:400Nuclear Speckle RRID: AB_298608MarkerRabbit Histone Cat# 39133,H3K27ac antibody Active Motif RRID: AB_256101 1:400 (Takei et al. 2021) (pAb) 6Santa Cruz Cat# sc-376248,Mouse Lamin A / CBiotechnolog RRID: AB_109915 1:200 (Takei et al. 2021) (E-l)y 36Cell Cat# 2598,NUP98 (C39A3) (Lincoln et al.Signaling RRID: AB_226770 1:200Rabbit mAb 2022)Technology 0Alexa Fluor® 790JacksonAffiniPure DonkeyImmunoRese Cat# 715-655-150 1:1000 (Su et al. 2020) Anti-Mouse IgGarch Labs(H+L)Alexa Fluor® 647JacksonAffiniPure DonkeyImmunoRese Cat# 111-605-144 1:1000 (Su et al. 2020) Anti-Rabbit IgGarch Labs(H+L)

[0214] Image acquisition: Applicant imaged approximately -650 fields of view (FOVs) covering the HPC. After each round of hybridization, Applicant acquired z-stack images of each FOV in 4 colors:750 nm, 647nm, 560 nm and 405 nm. Consecutive z sections were separated by 300 nm and covered 15 um of the sample. Images were acquired at a rate of 20 Hz.

[0215] DATA ANALYSIS

[0216] After generating the gene-barcode matrix file from Cell Ranger, the individual count matrices were merged together and processed using the Seurat v4.0.1 R packageAtty. Dkt. No.: 114198-3660(satijalab.org / seurat / ). Further filtering and clustering analyses of the scRNA-seq cells were performed with the Seurat package, as described in the tutorials (satijalab.org / seurat / ). Cells with at least 1,000 genes detected, and a mitochondrial read percentage of less than 30% were used for downstream processing. Potential doublets were removed using DoubletFinder (github.com / chris-mcginnis-ucsf / DoubletFinder), using an anticipated doublet rate of 5%, which is the expected rate reported by 10X Genomics for the number of cells loaded onto the 10X Controller. For the aggregated dataset, gene expression was normalized for genes expressed per cell and total expression by the NormalizeData function. The top 3,000 variable genes were detected with the FindVariableFeatures function with default parameters. All the genes were subsequently scaled using the ScaleData function, which utilizes a linear regression model to eliminate technical variability due to number of genes detected, replicate, and mitochondrial read percentage. Principal components were calculated using RunPCA and the 50 significant principal components (determined by ElbowPlot) were used for creating the nearest neighbor graph utilizing the FindNeighbors function with k.param = 50. The generated nearest neighbor graph was then used for graph-based, semi-unsupervised clustering (FindClusters, default resolution of 0.8) and uniform manifold approximation and projection (UMAP) to project the cells into two dimensions. Marker genes were identified using a Wilcoxon rank-sum test (FindAllMarkers, default parameters) for pairwise comparisons between cell clusters. Cell identities were assigned to the clusters by cross-referencing their marker genes with known cardiac cell type markers from human studies, in addition to in situ hybridization data from the literature. Typically, one cell cluster would emerge that expressed marker genes representing multiple populations, as well as contained cells with low UMI and gene counts that escaped the first filtering step. These cells were removed from downstream analyses. The clustering approach was then repeated for each partition of cell types (Cardiomyocyte, Mesenchymal, Endothelial, Neural crest-like, Blood) as described above.

[0217] MERFISH processing:1) MERlin: Individual RNA molecules were decoded using MERlin v0.6.1. Images were aligned across hybridization rounds by maximizing phase cross-correlation on the fiducial bead channel to adjust for drift in the position of the stage from round to round. Background was reduced by applying a high-pass filter and decoding was then performed per-pixel. For each pixel, a vector was constructed of the 16-22 brightnessAtty. Dkt. No.: 114198-3660values from each of the 22–16 rounds of imaging. These vectors were then L2 normalized and their Euclidean distances to each of the L2 normalized barcodes from MERFISH codebook was calculated. Pixels were assigned to the gene whose barcode they were closest to, unless the closest distance was greater than 0.512, in which case the pixel was not assigned to a gene. Adjacent pixels assigned to the same gene were combined into a single RNA molecule. Molecules were filtered to remove potential false positives by comparing the mean brightness, pixel size, and distance to the closest barcode of molecules assigned to blank barcodes to those assigned to genes to achieve an estimated misidentification rate of 5%. The exact position of each molecule was calculated as the median position of all pixels consisting of the molecule.2) Cellpose v2.08 was used to perform image segmentation to determine the boundaries of cells and nuclei. The nuclei boundaries were determined by running Cellpose with the ‘nuclei’ model using default parameters on the DAPI stain channel of the prehybridization images. Cytoplasm boundaries were segmented with the ‘cyto’ model and default parameters using the polyT stain channel. RNA molecules identified by MERlin were assigned to cells and nuclei by applying these segmentation masks to the positions of the molecules. Any segmented cells that did not have any barcodes assigned were removed before constructing the cell-by-gene matrix.3) Cellpose v2.0.28 for MERFISH+: was used to perform image segmentation to determine the boundaries of cells and nuclei. The nuclei boundaries were determined by running Cellpose with the ‘nuclei’ model using default parameters on the DAPI stain channel of each hybridization images. RNA molecules identified by MERlin were assigned to cells and nuclei by applying these segmentation masks to the positions of the molecules. Any segmented cells that did not have any barcodes assigned were removed before constructing the cell-by-gene matrix.

[0218] Single Molecule FISH Computational Analysis for heart data

[0219] Fluorescent spots were identified in smFISH images identically to the process used for MERFISH. Spots were then filtered to ensure a high correlation with the point spread function and a minimum brightness threshold, which was calibrated independently for each color channel by visually inspecting the resulting spots at various cutoffs. The filtered spots wereAtty. Dkt. No.: 114198-3660then assigned to the same cell segmentation mask generated for the accompanying MERFISH so that the gene expression measured by smFISH and MERFISH could be associated with the same cells.

[0220] Single Molecule FISH Computational Analysis for heart data

[0221] Fluorescent spots were identified in smFISH images identically to the process used for MERFISH. Spots were then filtered to ensure a high correlation with the point spread function and a minimum brightness threshold, which was calibrated independently for each color channel by visually inspecting the resulting spots at various cutoffs. The filtered spots were then assigned to the same cell segmentation mask generated for the accompanying MERFISH so that the gene expression measured by smFISH and MERFISH could be associated with the same cells.

[0222] Images were flatfield corrected for the two gene channels (750 nm and 635 nm) and the fiducial marker (405 nm) channel. To reduce background noise, for each hybridization round, the images of the preceding hybridization round were reduced in intensity and subtracted to obtain new background- subtracted images. The images were then locally normalized by subtracting a 15x15 blur from each pixel, before undergoing maximum intensity projection into two dimensions. For transcript detection, the OpenCV function adaptive Threshold was used with a block size of 41 pixels, and a subtracted constant ranging from -80 to -70 among the replicate smFISH experiments. This constant was empirically determined by choosing a value that ensured the resulting mask only captured visible fluorescent spots across diverse imaging planes for each gene. Using the regionprops function from Scikit-Image, Applicant filtered out spots on the max with eccentricity value of 0 and cells with low pixel area (<4 pixels) to combat artifactual fluorescence detection. A global threshold was identified for the images of each gene, by observing the value for which features identified as non-specific background by their irregular shape and low intensity, were separated from higher-intensity transcripts. The coordinates of local brightness maxima that remained unattenuated after applying this global threshold were stored. Coordinates lying within the adaptive Threshold mask boundaries were identified and counted as a single identified gene transcript. The images for each of the smFISH imaging rounds were aligned to the respective initial MERFISH hybridization round images to correct for microscopic drift, using the fiducial marker channel. This was done by fitting spots to the fiducial bead markers of both sets of images, then minimizing the medianAtty. Dkt. No.: 114198-3660distances between them. DAPI segmentation masks obtained from the MERFISH imaging were translated using this drift correction so that all identified gene transcript locations could accurately be assigned to the drift-corrected nuclei, allowing reconstruction of a spatial mosaic of the cellular gene expression for each of the sequentially imaged gene targets.

[0223] Segmentation Using Cellpose

[0224] Cellpose is an advanced image segmentation tool widely used in the analysis of Multiplexed Error Robust Fluorescence in situ Hybridization (MERFISH) data (Stringer, 2021). Cellpose was utilized for image segmentation to delineate the boundaries of cells in MERFISH data. For cytoplasm boundaries, the ‘cyto’ model with default parameters was applied to the PolyA and DAPI stain channels. Applicant used the PolyA channel for the original MERFISH probe experiments and acrydite-modified probe experiments (FIG. 2).However, for the reimages and second reimages, the PolyA signal was weak, so Applicant used the DAPI channel instead to run Cellpose. Additionally, MERFISH+ uses only the DAPI channel for cell segmentation with Cellpose. RNA molecules identified by MERlin were then assigned to the segmented cells and nuclei based on these segmentation masks.

[0225] Cellpose was utilized for image segmentation to delineate the boundaries of cells in MERFISH data. For cytoplasm boundaries, the ‘cyto’ model with default parameters was applied to the PolyA and DAPI stain channels. Applicant used the PolyA channel for original acrydite experiments. For the reimages and second reimages, the PolyA signal was weak, so Applicant used the DAPI channel instead to run Cellpose. RNA molecules identified by MERlin were then assigned to the segmented cells and nuclei based on these segmentation masks.

[0226] One important step in analyzing MERFISH data is to segment the cells, or find the boundaries of cells, so that Applicant can assign certain transcripts to certain cells based on their spatial location. In order to do this, Applicant used an algorithm called Cellpose. Normally, Applicant use the PolyA staining images (from hybridization round 0) to run Cellpose. However, for many of the reimages and second reimages, the PolyA signal was very weak, so Applicant used the DAPI staining images instead to run Cellpose. This led to a great improvement in the results for the acrydite reimages and second reimages (FIG. 11). Note thatAtty. Dkt. No.: 114198-3660all previous data for reimages and second reimages was from using the DAPI staining images for Cellpose rather than the PolyA that was used for the rest of the experiments.

[0227] However, while there was an improvement in the quality of the reimages, because of the lasers used, the DAPI staining images often have a gradient (as shown in FIG. HE).

[0228] While DAPI staining improved the quality control metrics such as the number of segmented cells and the transcripts per cell, the results still are not as high quality as the other acrydite and regular probe experiments.

[0229] Flat Field Corrections:

[0230] In order to fix the DAPI segmentation quality further, Applicant performed a flat field correction to the dax images. This process fixes the gradient that is present in the DAPI staining dax images so that Applicant don’t get that gradient in the final spatial images Applicant performed (using PolyA, DAPI, and flat field corrected DAPI images) for both N3S5A and its reimage. There was a clear improvement in the median transcripts per cell count as Applicant improved the segmentation by using the flat field corrected DAPI images.

[0231] Cell clustering analysis of MERFISH

[0232] With the cell-by-gene matrix, Applicant followed a standard procedure as suggested in the Scanpy version 1.89tutorial using python version 3.9 for processing MERFISH data. Count normalization, Principal Component Analysis (PCA), neighborhood graph construction, and Uniform Manifold Approximation and Projection (UMAP) were performed with SCANPY’ s default parameters. Applicant performed Leiden clustering utilizing a resolution of 2. The top 20 differential genes identified by the rank_gene_groups() function were used to annotate each cluster. Applicant further subclustered the ventricular cardiomyocytes (vCM) clusters using Leiden clustering at a resolution of 1. This allowed us to further annotate vCM cells as Compact and Trabecular vCMs for both the left and right ventricles. To investigate the cellular populations in the ventricle specifically, Applicant manually defined the ventricular region, subset the MERFISH dataset to the ventricular cells and performed Leiden clustering at a resolution of 5. Again, the top 20 genes identified by the rank_gene_groups() function were used to annotate each cluster.

[0233] Statistical analysisAtty. Dkt. No.: 114198-3660[02341 Sample sizes were not predetermined utilizing statistical methods. Tissue samples were not randomized, nor were the investigators blinded to the donors. To identify differentially expressed genes between clusters, a Wilcoxon rank sum test was performed, and the resulting P value was corrected using the Bonferroni procedure. For the qPCR results, Applicant used a two-tailed Student T-test using R (r-project.org / ).[0235 j Chromatin Tracing Analysis[0236| A. Localization of Fluorescent Spots[0237| To calculate fluorescent spot localizations for chromatin tracing data, Applicant followed the following computational steps:1) Applicant computed a point spread function (PSF) for the microscope and a median image across all fields of view for each color channel based on the first round of imaging to be used for homogenizing the illumination across the field of view (called flat-field correction).2) To identify fluorescent spots, the images were flat-field corrected, deconvoluted with the custom PSF, and then local maxima were computed on the resulting images. A flatfield correction was done for each color channel separately.

[0238] B. Image registration and selection of chromatin traces

[0239] Imaging registration was performed by aligning the DAPI channel of each image from the same field of view across imaging rounds. First, the local maxima and local minima of the flatfield corrected and deconvolved DAPI signal were calculated. Next, a rigid translation was calculated using a fast Fourier transform to best align the local maxima / minima between imaging rounds.

[0240] Nuclear segmentation was performed on the DAPI signal of the first round of imaging using the Cellpose [nature. com / articles / s41592-022-01663-4] algorithm using the “nuclei” neural network model. Following image registration, chromatin traces were computed from the drift-corrected local maxima of each imaged locus as previously described (Su et al. 2020).

[0241] RNA MERFISH+ 2000-gene Analysis-Kaifu

[0242] The MERFISH decoding followed a similar strategy as the MERlin algorithm but operated on spots identified in the images rather than individual pixels. Briefly, the drift andAtty. Dkt. No.: 114198-3660chromatic aberration corrected local maxima (spots) were grouped into clusters, where each cluster contained all spots from all imaging rounds within a 2-pixel radius of an anchor spot. Clusters were generated for every possible anchor spot. Any cluster containing spots from at least four images was then assigned a gene identity by best matching the MERFISH codebook. Each cluster was ranked by the average brightness and the interdistance between the contained spots. These measures were used to filter the decoded clustered and best separate the more confident spots from the less confident.

[0243] Protein Density Quantification

[0244] The antibody images were flat field corrected, deconvolved and then registered to the chromatin traces using the DAPI signal as described before. For each chromatin trace the fluorescent signal of each antibody was sampled at each genomic locus 3D location in each cell.

[0245] Refinement of the nuclear segmentation

[0246] Refined 3D nuclear segmentation was performed using the 3D Cellpose ‘nuclei’ model based on the LaminA and NUP98 fluorescent stain.

[0247] Table 2 - QC Data, microscope, probe type, and issues for all fifteen experiments. Since this experiment is on one heart sample, Applicant have labeled each slice as S{ slice number} {A}, where A is for acrydite. Applicant tested the new acrydite probes on sections that are directly adjacent to the corresponding regular probe experiments in order to minimize biological differences between the experiments.

[0248] Table 2. QC Data, microscope, probe type, and issues for all fifteen experiments.Expert Probe Microsco FOVs Cells Cells Transc Genes Filtered Issues merit Type Pe per ripts per barcode FOV per cell s percell FOVR107 N2S1 Acrd heart comb Acrydit Lemon 552 133,3 242 152 44 49,790 Some (S1A / S e 78 empty 2) FOVs N2S1 Acrd heart comb StripR Acrydit Lemon 182 38,40 211 7 5 4,212 Seg issue elm e 1Atty. Dkt. No.: 114198-3660Expert Probe Microsco FOVs Cells Cells Transc Genes Filtered Issues merit Type pe per ripts per barcode FOV per cell s percell FOVN2S2 heart com Regula Kiwi 587 134,9 230 209 54 55,616 Some b r 76 empty FOVs N2S1 Acryd heart Kiwi comb Acrydit Kiwi 182 38,40 211 84 36 29,714 Some e 1 empty FOVs R110 N2S14R comb Regula Kiwi 659 80,60 122 28 12 13,879 Some (S11A / eg r 4 empty S14) FOVs N2S11 Aery com Acrydit Lemon 558 69,58 125 74 34 25,371 Some b e 6 empty FOVs N2S11A Kiwi Rel Acrydit Kiwi 558 49,36 88 82 37 34,920 Seg issue cry m e 3Rill N2S22 Acryd com Acrydit Lemon 707 121,6 172 197 60 57,604 None (S22A / b e 14S21)N2S22 Acryd com FMRe Acrydit Lemon 707 124,3 176 112 45 40,192 Many b Im e 46 empty FOVs N2S22A Lemo Acrydit Lemon 707 79,45 112 24 13 9,584 Seg issue cryd nRel e 6m2N2S22 Acryd Kiwi Relm Acrydit Kiwi 707 54,01 76 21 12 12,112 Seg issue, e 7 couple empty FOVs N2S22A Kiwi Rel Acrydit Kiwi 707 99,48 141 76 37 34,749 Seg issue, cryd m2 e 8 couple empty FOVsAtty. Dkt. No.: 114198-3660Experi Probe Microsco FOVs Cells Cells Transc Genes Filtered Issues merit Type Pe per ripts per barcode FOV per cell s per cell FOV R112 N2S21R comb Regula Kiwi 688 113,4 165 53 20 18,331 Many eg r 50 empty FOVs R113 N2S31 Reg com Regula Kiwi 683 111,9 164 184 53 55,662 Some b r 86 empty FOVs R115 N2S25 Reg com Regula Kiwi 726 117,9 162 23 11 8,838 Couple b r 09 emptyFOVs

[0249] For some samples, another round of reimaging was performed, which was expected to be identical to the first reimage. These second reimages were imaged twice, once on each microscope, in order to investigate the reproducibility among different units of the imaging systems of identical hardware configuration such as that of Lemon and Kiwi.

[0250] The three experiment series are S1A / S2 (R107N2SlAcryheartcomb with R107N2S2heartcomb), S11A / S14 (R110N2S1 lAcrycomb with R110N2S14Regcomb), and S22A / S21 (R11 lN2S22Acrydcomb with R112N2S21Regcomb). Each of these series also had subsequent reimages done on the acrydite experiments (FIG. 9). Aside from the three series, Applicant also have two lone regular probe experiments, N2S25R and N2S31R.N2S25R was imaged for the purpose of understanding the impact of imaging the sequential before the combinatorial. Applicant also use 4C15R, a regular probe experiment, as a baseline as it is the best heart experiment to date. 4C15R was imaged on Kiwi at the end of 2021.

[0251] Heart data

[0252] Regular probes - (no acrydite modification) - 1 dataset

[0253] Acrydite Probes (before 100% and after 100%) - 2 datasets but on the same sample

[0254] Table 3Atty. Dkt. No.: 114198-3660Acrydite Probes Regular Probes N2S1 N2S11 N2S22 N2S2 N2S14Reg N2S21Reg Acrd heart Aery comb Acryd comb heart comb comb combcomb10255} Example 2: MERFISH+, a large-scale, multi-omics spatial technology, resolves the transcriptomic holograms of the 3D human developing heart

[0256] Hybridization-based spatial transcriptomics and genomics technologies are rapidly transforming Applicant’s understanding of cellular and subcellular organization in complex tissues. However, existing methods remain constrained in gene throughput, multimodal compatibility, and field of view. Here, Applicant present MERFISH+, an enhanced version of Multiplexed Error-Robust Fluorescence in Situ Hybridization (MERFISH), which integrates chemical probe anchoring within protective hydrogels, redesigned high-throughput microfluidics, and large-area volumetric microscopy. MERFISH+ simultaneously quantifies >1,800 genes and resolves 3D chromatin architecture of a dynamic regulatory domain critical for cardiac development. With a tenfold expansion in imaging area, MERFISH+ enabled the first 3D molecular atlas of human cardiac morphogenesis at subcellular resolution. Spateo firstly reconstructed the 3D map of heart from 53 sections and a total of 3.1 million cells across 34 distinct populations and then quantified volumetric gene expression across key anatomical structures, including cardiac valves and the developing vasculature. Neighborhood analysis uncovered coordinated signaling between the conduction system and valve / cushion cells during morphogenesis. Moreover, 3D mapping revealed region-specific transcriptional programs in vascular smooth muscle cells and their supporting lineages along coronary arteries, illuminating mechanisms of arterial patterning and maturation. Finally, Applicant introduce Spateo- VI, a generative integration framework that jointly harmonizes single-cell and 3D spatial multimodal data to construct a transcriptome-scale epigenomic atlas of the developing human heart. Together, MERFISH+ establishes a versatile, large-format platform for spatial multi-omics at unprecedented scale and resolution, enabling integrative discovery of gene regulatory mechanisms and organ architecture dynamics in intact 3D space.

[0257] In summary, MERFISH+ expands MERFISH to >1,800 genes and whole-organ 3D imaging scale. MERFISH+ combines chemical probe anchoring, microfluidics, and large-area volumetric microscopy. MERFISH+ generates a 3D molecular atlas of human cardiacAtty. Dkt. No.: 114198-3660morphogenesis at subcellular resolution. MERFISH+ reveals coordinated signaling between conduction, valve, and vascular lineages. MERFISH+ introduces Spateo-VI, a generative framework integrating 3D multimodal datasets.

[0258] Single-cell RNA sequencing (scRNA-seq) has significantly advanced Applicant’s understanding of cellular diversity and gene expression dynamics across diverse biological systems, from developmental biology to microbial ecosystems. Despite its transformative impact, scRNA-seq inherently disrupts tissue integrity, eliminating critical spatial context. Spatial transcriptomics methods overcome this limitation by preserving positional information, essential for understanding tissue architecture and cell interactions in their native environments. Hybridization-based spatial transcriptomics techniques, such as Multiplexed Error-Robust Fluorescence In Situ Hybridization (MERFISH), Seq-FISH and others achieve high sensitivity and spatial resolution by repeatedly imaging fluorescently labeled oligonucleotide probes bound to single RNA molecules. Many groups started employing these methods to unravel the gene expression and cell type composition across an array of human tissues including human kidney, human lung, developing brain etc. Recently, Applicant performed MERFISH in the developing human heart, revealing the spatial organization of single cells into cellular communities that form distinct cardiac structures. However, the current methods face critical limitations: First, most prior studies focus only on RNA and on a limited set of genes (typically a few hundred) hence overlooking important genes and processes such as epigenetic regulation of DNA which characterize biological processes and diseases. Secondly, most studies imaged two-dimensional (2D) tissue sections capturing up to hundreds of thousands of cells. While this provides valuable spatial information within the plane of the section, it does not capture the three-dimensional (3D) organization of tissues. Many biological processes and cellular interactions occur in three dimensions, especially during a variety of developmental stages in human organogenesis. Therefore, the choice of sectioning plane can significantly influence the observed gene expression patterns, potentially leading to biased or incomplete interpretations.

[0259] To overcome these limitations, particularly for valuable human tissues, Applicant developed MERFISH+ by extending the MERFISH technology. First Applicant incorporated an acrydite group into the encoding / primary probes, covalently anchoring them to a layer of acrylamide gel covering the tissues. These anchored probes are then read out and imaged overAtty. Dkt. No.: 114198-3660an extended period, spanning three or more months (FIG.20). This enhanced stability enabled within 2D sections: 1) a new flexible bar-coding scheme allowing for imaging >1,800 genes with a similar quality as that of the 300-gene scale and 2) stable multi-modal imaging of RNA, DNA, and proteins within the same sample. By combining the stability of the probes with a new high throughput microscope and microfluidics design Applicant increase 10-fold the imaging area, from 1~2 cm2to ~12 cm2allowing for millions of cells profiled per experiment. Extending Applicant’s previous Spateo framework15, Applicant developed a series of novel computational pipelines to facilitate the 3D MERFISH data analyses, Applicant reconstructed the first human 3D digital heart, revealing the intricate architecture and cellular composition of each cardiac structure, as well as cell-cell interactions in three dimensions. Notably, Applicant’s Spateo- VI method jointly integrated single cell and spatial multi-omics data to allow the creation of the first transcriptome-wide, multi-modal model of 3D human developing heart, opening doors for many downstream analyses, including the analyses of gene expression and cell-cell communication during coronary artery development.

[0260] Results

[0261] Acrydite FISH-probes improve stability of MERFISH experiments

[0262] To overcome the gradual decay of RNA during imaging and hence a gradual decrease of the on-target signal, Applicant modified the MERFISH probe synthesis protocol to incorporate an acrydite group to the 5’ end of the oligonucleotides of all the encoding / primary probes. The molecular structure of the acrydite group is similar to the acrylamide monomers and reacts with activated double bonds of acrylamide and bisacrylamide, resulting in the tethering of acrydite-conjugated oligos into the gel matrix17. Such a modification is expected to decouple the probe stability from the decay of the RNA substrate during imaging, greatly expanding the possible number of hybridization cycles (FIG. 20 and FIG. 13A).

[0263] Applicant conducted MERFISH experiments with acrydite modified oligonucleotide probes on 16-um human developing heart sections (12 post-conceptual weeks, 12 p.c.w.). Briefly, a template library of MERFISH probes produced through the silicon-based synthesis technology (TWIST Bioscience) was amplified, first using limited amplification cycles of PCR and then T7 RNA synthesis followed by a reverse transcription reaction with an acrydite modified primer at the 5’ end. These probes were hybridized to RNA molecules within theAtty. Dkt. No.: 114198-3660tissue, followed by the casting of an acrylamide gel onto the sample. In order to quantify the stability of the new probes, Applicant used acrydite probes targeting MYH7 (a gene expressed in ventricular cardiomyocytes) and compared the signal in the first cycle of hybridization with the signal after 54 cycles of hybridization with highly stringent 100%-formamide washes between each cycle (FIG. 13B). Applicant found that brightness plateaus at about 55% of the first round’s brightness.

[0264] Applicant then resynthesized the same 238-gene MERFISH library used previously and performed MERFISH experiments with the acrydite modified probes which Applicant named MERFISH+. On the same sample, Applicant first performed the MERFISH experiment as previously described using TCEP cleavage for 11 rounds (FIG. 13C, middle). Applicant then performed a stringent 100%-formamide wash step to completely remove residual readout probes prior to reconducting the MERFISH experiment across an additional 11 hybridization cycles, with each round stripped with 100% formamide (FIG. 13C, right). Applicant quantified the number of mRNA molecules of each of the 238 targeted genes in each cell and performed UMAP embedding and cell type definition using Leiden clustering of the cells imaged across the two MERFISH experiments. Applicant obtained a near-identical cell type definition and spatial distribution across the 2D section (FIGS. 13C and 13D) with the same cell type markers defining each cell type (FIG. 13F) as in Applicant’s prior publication with the original MERFISH protocol. Furthermore, Applicant obtained a very high Pearson correlation (0.95 and 0.93) between the average expression of each gene in both of the new MERFISH+ experiments compared to prior published results8,9, suggesting that the detection efficiency and accuracy was not appreciably affected by the new acrydite modified probes (FIG. 13E) These results highlight the increased stability of MERFISH+ probes, which effectively address the probe degradation issues observed in the original MERFISH design. In conventional MERFISH, signal loss can occur due to RNA degradation, equipment malfunctions, or instability of imaging reagents. However, in MERFISH+, a stringent 100% formamide wash allows the experiment to be reset and resumed without compromising sample integrity. This capability preserves precious biological samples and ensures continued data collection even after unexpected disruptions. Together, these improvements make MERFISH+ a more robust and reliable platform for high-resolution spatial transcriptomics, particularly in challenging experimental conditions.Atty. Dkt. No.: 114198-3660

[0265] MERFISH+ allows for scaling up the number of target genes

[0266] The increased stability of the MERFISH+ probes allow for a scalable increase in the number of genes imaged to thousands of targets (FIGS. 21 and 22). Applicant imaged a 12 p.c.w. heart section across a total of 1,863 genes using 8 modules of 16 hybridization cycles (a 48-bit code book per module), each including a portion of the previously imaged genes (FIG.14A) Cell type identities were defined by their gene expression profiles, which showed strong concordance with matching populations identified by single-cell RNA sequencing8(FIG.14B) With this increased number of MERFISH genes, Applicant identified an additional cardiac cell subcluster not described in Applicant’s published work: vCM-LV-papillary cells (FIGS. 14A-boxed and 14C) vCM-LV-papillary cells are localized within the left ventricle, where they form a distinct subpopulation concentrated around the papillary muscle regions. These cells align along the inner ventricular wall and extend toward the trabecular muscles and mitral valve cushion cells, establishing a continuous structural and signaling interface between the contractile myocardium and valvular tissues (FIG. 14C). Functionally these cells contribute, alongside with the papillary muscles, to valve mechanics, preventing mitral valve prolapse during ventricular contraction. The larger gene panel of 1,863 genes allowed us to identify marker genes specific for this population of cells, by comparing their gene expression with other cardiomyocyte cell types (shown in FIG. 14C, left, vCM-RV-AV, and FIG. 14D), which had the most similar gene expression profile to the vCM-LV-papillary cells. Both DLGAP1 and ACTA1 showed a spatial expression pattern that was highly specific to the papillary regions and were more enriched in the papillary cardiomyocytes compared to the closest cell type (FIG. 14C, right).

[0267] Moreover, Applicant were able to identify more specific subtypes of the BECs (FIGS.14A-boxed and 23), which Applicant named BEC I, II, III. BEC I is the most abundant and is broadly distributed along the ventricular wall and IVS, suggesting that it may represent capillary endothelium. The BEC II endothelial subcluster is located at the sinus venosus region or the atrioventricular canal junction between the atrium and ventricle, consistent with a sinus venosus-derived endothelial lineage. In contrast, BEC III cells are sparsely distributed within the ventricular myocardium, forming distinct clusters. These cells show high expression of SOX17, a key regulator of coronary artery development, as well as DLL4, a NOTCH pathway gene also essential for coronary artery formation. Furthermore, Applicant were able toAtty. Dkt. No.: 114198-3660distinguish the epicardium into two populations: aEpicardial and vEpicardial (FIG. 14A-boxed). These two populations are specifically located in the atrial (aEpicardial, F13A1) surface and in the ventricular (vEpicardial, CEMIP) surface (FIGS. 23D and 23E).

[0268] The larger number of genes profiled enabled us to identify genes with asymmetric expression between the left and right sides of the heart (FIGS. 14E-14H). By focusing on atrial and ventricular cardiomyocytes and ordering genes based on their left-right expression gradients (FIGS. 14E and 14G), Applicant confirmed the left-side bias of PITX2, and HAND1, which are known markers of the left atrium (PITX2) (FIGS. 14E and 14F) and left ventricle (HAND1) (FIGS. 14G and 14H). In addition, in atrial cardiomyocytes, Applicant identified genes such as AKAP6, COL2A1, TRPM3, and IPO8 as enriched in the left atrium, while MAP4, CYFIP2, ADM, MAZ, and KIF26B were more highly expressed in the right atrium. Similarly, in ventricular cardiomyocytes, genes such as COL11 Al, PAP1GAP2, KIF26B, and SLC1A3 were highly enriched in the left ventricle, whereas IGFBP4S, COX4I1, SMYD1, EDA, LBH, WNT9A and DES were enriched in the right ventricle (FIGS. 14G and 14H).Notably, several asymmetrically expressed genes, such as COL11A1, are essential for ventricular morphogenesis. Together, these results demonstrate that MERFISH+ enables detailed spatial mapping of additional cell types and states, providing a more comprehensive molecular atlas of the developing heart.

[0269] Facile multimodal imaging of RNA, DNA and proteins.

[0270] Applicant next wondered whether the improved probe stability of MERFISH+ facilitates multi-modal imaging of RNA transcripts, 3D DNA chromatin structure and nuclear accumulation of epigenetic marks (FIG. 15A). Applicant selected and imaged a dynamic genetic locus - the FXYD region - containing 21 genes, including the FXYD family of ion channel modulators, which show high variability of expression across the cell types of the developing heart8(FIG. 15B). Applicant synthesized -24,000 oligonucleotide probes with acrydite anchor, targeting this entire 510 kb FXYD locus (FIGS. 15A and 15B). Using sequential hybridization and imaging, Applicant read out the signal of 51 10-kb segments comprising this locus across all the cells in a 12 p.c.w. human heart section. By fitting and aligning the signal of each segment, the 3D configurations of FXYD locus in each chromosome in each cell is reconstructed (FIG. 15F) with -50 nm spatial resolution and 10-kb genomic resolution. On a single-cell level, Applicant noticed that this locus tended to organize withinAtty. Dkt. No.: 114198-3660globular domain-like structures (FIG. 15D, right) reminiscent of globular, TAD-like structures previously reported. Upon averaging the inter-distances between each of the 10-kb segments across -36,000 cells, Applicant found excellent agreement between the average structural features predicted by imaging and those predicted by an orthogonal sequencing method called HiC in cultured cardiomyocytes (FIG. 24A). To relate the chromatin structure with transcription within the same cells Applicant further imaged the mature RNA expression of 250 marker genes and both the mature and nascent RNA expression of the genes in the FXYD locus. The MERFISH+ anchoring scheme facilitated the imaging of RNA across additional cycles after the DNA measurements with similar quality to the experiments in which only RNA is measured (FIGS. 15C and 24B). The RNA expression was used to define cell types and quantify the expression of the FXYD genes across cell types (FIGS. 15D and 15E). Applicant asked if there are conserved structures of this locus that are differentially formed across different cell types. Applicant found a “loop” (enriched interaction) between the promoter of the FXYD5 and a distal element close to the promoter regions of LSR and USF2 enriched in endothelial cells compared to all other cell types. Intriguingly this long-range chromatin interaction (70 kb) coincides with the increased expression of the FXYD5 gene within endothelial cells. Alternatively, in ventricular cardiomyocytes Applicant saw a potential compartment switch in which the first chromatin domain (chr19:35.00Mb-35.14Mb) had enriched interaction with the last chromatin domain (chr19: 35.23Mb-35.50Mb).[02711 To further test the multimodal capability of the MERFISH imaging, Applicant performed sequential staining and immunofluorescence imaging on the same sample with different epigenomics markers: pol2ser2, pol2ser5, H3K4me3, H3K27ac, and H3K27me3 (FIGS. 15I, 15J and 24C). The total brightness of the antibody signal in each nucleus provided an approximate measure of the accumulation of that mark within each cell allowing for a comparison across cell types. Consistent with Applicant’s measurements in the developing brain, Applicant found neuronal— like cells had higher levels of H3K27me3 and lower levels of H3K27ac relative to other cell types. To validate the accuracy of the protein-staining in MERFISH+, Applicant performed immunofluorescence staining with the same antibodies using adjacent tissue slices from the same heart sample. Quantification of such measurements showed good correlation between the two sets of data (FIG.24C). Overall MERFISH+ allowsAtty. Dkt. No.: 114198-3660for multi-modal measurements connecting the kb-scale DNA structure, RNA expression and epigenetic states of different cell types.

[0272] Scaling up the imaging area by 10-fold enables 3D organ reconstruction

[0273] Complex tissues contain diverse cell types, with many enriched in specific regions. To fully capture the cellular composition of the heart, broader tissue coverage is required, necessitating spatial methods capable of large-area imaging. MERFISH+, with its stable probes and high gene throughput, enables large-scale scanning in a single run. Applicant developed a new microscope system with higher-power lasers, a larger-format camera, and an upgraded microfluidics chamber, achieving a 10-fold increase in imaging area compared to previous setups (FIGS. 16A, 16B and STAR Methods) To ensure comprehensive sampling of the heart and to obtain an unbiased understanding of its cellular composition and spatial organization, Applicant utilized this system to image 238 genes across a total of 53 consecutive transverse sections from a 12 p.c.w. human heart donor across two experiments, with 21 sections spanning the atrial half of the heart and 32 sections spanning the ventricular half (FIG.16C). This imaging scheme covered the entire developing heart at 112 um resolution along the basal-apical axis, profiling 3.1 million high quality cells (filtered from 3.8 million cells) aggregated from two large-scale MERFISH experiments.

[0274] Comprehensive spatial cell atlas of developing human heart

[0275] The transcriptional profiles within each cell were used to first define all the 12 major cell types in the heart: blood vessel endothelial cells (BEC), atrial cardiomyocytes (aCM), nonchambered cardiomyocytes (ncCM), ventricular cardiomyocytes (vCM), endocardial cells (Endo), epicardial cells (Epi), fibroblasts (Fibro), lymphatic endothelial cells (LEC), pericytes (Peri), white blood cells (WBC), vascular smooth muscle cell (VSMC), and neural cells (FIGS.16D and 16E) Each major cell type was isolated, and subtypes were defined progressively with higher resolution until spatial or transcriptional separation was no longer observed between the subtype divisions. This resulted in identifying a total of 34 distinct cell clusters (FIG. 16F and FIG. 25). Upon averaging across all cells or across matching cell types, a reasonably high correlation was found between the gene expression of the MERFISH+ heart data and prior single-cell RNA-sequencing (Pearson correlation coefficient of 0.62; FIG.25A) or prior MERFISH data on 2D heart sections (Pearson correlation coefficient of 0.80, FIG.Atty. Dkt. No.: 114198-366025B), respectively. The cluster annotations also show a high correlation with the previous MERFISH dataset (Pearson correlation coefficient of 0.89, FIG. 25C).

[0276] Importantly, this comprehensive cell atlas provides an unbiased assessment of the spatial distribution and an estimate of the absolute abundance of each cell subtype within the developing human heart (FIG. 25D). The ventricular cardiomyocytes (vCMs) represent the largest cell population, accounting for 51% of total cells, followed by fibroblasts at 16.6%, and atrial cardiomyocytes (aCMs) at 12% (FIG. 25D right).

[0277] The number of cells of each cell type identified within the heart varied by two orders of magnitude ranging from the most abundant compact ventricular cardiomyocytes (vCM-LV-com, 20.45%, or approximately 633,079*8 number of cells within the 12 p.c.w. heart) to the least abundant ncCM and other populations (2%, 61,200*8). The 53 sections were computationally aligned using Spateo into the first comprehensive 3D reconstruction of the developing human heart at single-cell resolution (FIG. 16G). This digital model provides a 3D visualization of each cardiac structure and cell type (FIG. 27), offering spatial context and anatomical precision that enables a more comprehensive understanding of their organization, distribution, and interactions within the cardiac tissue. In addition, it enables quantitative analysis of key anatomical parameters (Methods), such as measuring compact cardiomyocyte wall thickness as the distance between the epicardial-compact layer interface and the endocardium-trabecular vCM layer (FIGS. 16H-16J). This approach avoids the inclusion of free trabecular muscle in ventricular wall thickness measurements, as these structures can be influenced by the contraction and relaxation state of the heart at the time of dissection. For instance, during contraction, free trabeculae may cluster together and be mistakenly interpreted as part of the compact ventricular wall. The analysis reveals that the LV is generally thicker than the RV. Moreover, the wall thickness is not uniform, showing regional variation in both the LV and RV (FIGS. 16K, 16L, 25D and 25E). Qualitative analysis shows a positive correlation between the thickness of vCM-LV-Hybrid and vCM-LV-Compact layers, suggesting that vCM-LV-Hybrid cells may also contribute to left ventricular wall growth (FIGS. 16J and 25F).

[0278] Spatial analysis highlights the intricate cellular heterogeneity of spatial neighborhoods within the 3D heartAtty. Dkt. No.: 114198-3660

[0279] By clustering cell subtypes based jointly on their gene expression similarity and spatial proximity, Applicant identified several hierarchical anatomical neighborhoods, including the atrial neighborhood, compact ventricular neighborhood, trabecular neighborhood, valve interstitial neighborhood, and epicardial neighborhood (FIGS. 17A and 17B). In the compact ventricular neighborhood localized within the compact layer of the ventricle Applicant found enrichment of blood endothelial cells (BECs) and pericytes, indicative of increased blood vessel formation31and a specific subpopulation of fibroblasts (Compact vFibro). This is consistent with the role of BECs, pericytes and fibroblast in supporting ventricular compaction as demonstrated in mouse models in which defect of these cell types resulted in loss of ventricular compaction32,33. Similarly, in the trabecular neighborhood Applicant found an enrichment of specific fibroblast and endocardial subpopulations (vFibro and vEndocardial), consistent with the role of endocardium in regulating trabecular cardiomyocyte differentiation.

[0280] The epicardium, epicardium-derived cells (EPDCs), adventitial fibroblasts (adFibros), vascular smooth muscle cells (VSMCs), along with lymphatic endothelial cells (LECs) and neuronal-like cells, form a distinct neighborhood near the surface of the heart (FIG. 17C).Many of these non-cardiomyocyte cell types are known to migrate during heart development from the surface of the heart to the ventricular wall, playing a crucial role in cardiomyocyte proliferation, maturation, and ventricular wall morphogenesis. Applicant specifically isolated the cells near from the surface of the ventricle and assessed across the 3D heart the complexity of the local cellular composition (FIGS. 28A-28E). Interestingly, this analysis revealed that the right ventricular surface exhibits a more complex cellular composition than the left. This is consistent with Applicant’s prior measurements on 2D slices of the heart and may reflect the differential development of right versus left ventricle, with the right ventricle developing slower and having a more diverse progenitor origin.

[0281] Applicant next focused on the VIC neighbourhood containing cell types comprising or adjacent to the valve structures of the heart(FIG. 17A, C). The valve interstitial cells (VICs), the most dominant cell type in this neighborhood, were spatially located within 4 anatomical regions corresponding to the four valves: the mitral (MV), tricuspid (TV), pulmonary (PV), and aortic (AV) valves(FIG. 17C). To determine whether VIC cells in the four valves exhibit gene expression differences, Applicant further divided the VIC population into four groups based on their respective valve locations (FIGS. 17D, G, and H; FIGS. 28F-28G).Atty. Dkt. No.: 114198-3660Transcriptionally VIC-TV and VIC-MV show higher expression of PENK, RGS5, TBX5, NSG1, and SCN7A while, in contrast, VIC- AV and VIC-PV exhibit higher expression of HEY2, FRAB, and TENM3 (FIGS. 17F, I, J).

[0282] Using the VICs as landmarks Applicant quantified the surrounding cells to identify cell types differentially enriched within the four valves (FIG. 17K, FIG. 28H). Across all valves, Applicant identified two distinct populations of endothelial cells (named VECs and aEndo) arranged in a complementary manner around the leaflet structures. Valve endocardial cells (VECs) were specifically localized to the surface of all valve rings, whereas atrial endocardial cells (aEndo) extended along the inner wall of the chamber lumen containing the valve and cushion. Interestingly, in addition to the ventricular canal cardiomyocyte populations (ncCM-AVC-like, vCM-LV-AV and vCM-RV-AV) located within the VIC neighborhood, Applicant also observed a significant enrichment of vCM-IVS-His cells in regions surrounding the tricuspid, mitral, and aortic valves (FIG. 17K, FIG. 28H). vCM-IVS-His cells are located in close proximity to and potentially structurally integrated with VICs. This spatial association may reflect a coordinated function between vCM-IVS-His and valve / cushion cells. Notably, clinical conditions such as aortic valve calcification and mitral valve prolapse are frequently associated with arrhythmias, supporting a potential link between valve abnormalities and conduction system defects. Overall, these results highlight the intricate cellular neighborhoods that emerge within the 3D architecture of the developing human heart, providing a valuable spatial framework for the cardiac research community.

[0283] 3D cellular composition and expression gradient for the coronary artery morphogenesis

[0284] The coronary arteries are key cardiac structures, whose defects cause major heart diseases yet whose characterization remains challenging using individual 2D sections. In contrast, the 3D heart reconstruction allowed us to trace the vascular smooth muscle cells (VSMCs expressing MYH11), defining the aorta (AO), the pulmonary artery (PA), left coronary artery (LCA), right coronary artery (RCA), and their major coronary artery segments. The LCA clearly branches into the left circumflex artery (LCX) and the left anterior descending (LAD) artery (FIGS. 18A, 18B and 18D). These arteries exhibited variable lengths, trajectories and gene expression in the 3D heart. The AO and PA VSMCs have differential gene expression compared to those of descending coronary arteries marked by elevatedAtty. Dkt. No.: 114198-3660expression of FBLN5, HAND2, SFRP1, OSR1, and CRABP2 (FIG. 18C). These genes are known markers of neural crest cells (CRABP2) and second heart field (HAND2, OSR1) lineages, consistent with their developmental origin from the neural crest cells and second heart field. In contrast, coronary artery VSMCs showed higher expression of TBX18, MCAM, GJA5, CASQ2, and HEY2 (FIG. 18C), reflecting their EPDC origin as TBX18 is the marker gene of epicardium. Starting from the root region of the aorta, the left and right coronary arteries extended across the left and right ventricular surfaces, giving rise to a branching network that spans the entire length of the ventricles from base to apex, highlighting the ability of MERFISH+ to capture their trajectories in 3D (FIGS. 18B and 18E).

[0285] The 3D heart atlas allows for identifying the cellular environment and molecular pathways facilitating arterial morphogenesis. Applicant quantified the cell type composition along the proximal-distal axis within a 25 pm radial neighborhood of VSMCs, focusing on the RCA, LAD, and LCX branches. (FIGS. 18D-18G). Multiple cell types were proximal to arterial VSMCs including epicardial, EPDC, neural cells, LEC, and BEC consistent with their supportive role in angiogenesis. vEndo cell types were preferentially enriched at the proximal end of the blood vessels whereas LEC and epicardium cells were enriched at the distal blood vessels (FIG. 181) This might reflect that different cell types are differentially involved in either the maturation (proximal end) or elongation (distal end) of coronary artery development. The epicardium can differentiate into coronary smooth muscle cells and is located in the distal part of the artery, indicating the generation of more smooth muscle cells at the tip to promote arterial elongation. LEC has been reported to regulate cardiomyocyte proliferation and regeneration; however, its role in coronary artery development has not yet been demonstrated.

[0286] Applicant next set out to quantify gene expressions along the proximal-distal axis of the VSMC arterial branches. Several genes exhibited differential expression along the proximal-distal axis (FIG. 18H). TCF21, a marker of epi cardial -derived cells (EPDCs) that can differentiate into coronary smooth muscle cells, has been linked to coronary artery disease risk in multiple racial and ethnic groups through human genome-wide association studies. Its high expression in the distal artery suggests that EPDCs may be differentiating into smooth muscle cells at these sites, with TCF21 playing a crucial role in this process; moreover, downregulation of TCF21 is required for VSMC maturation. Additionally, CXCL12 is highly expressed in the distal artery. Beyond its established association with coronary artery disease,Atty. Dkt. No.: 114198-3660CXCL12 has recently been identified as a key regulator of left and right coronary vessel branching during cardiac development. Through CXCL12-CXCR4 signaling, it provides critical endothelial guidance cues that direct coronary vessel patterning and ensure complete vascular coverage of the dorsal ventricular surface. Together, these 3D analyses provide a spatially resolved catalog of transcriptional changes in VSMCs and their supporting cell types along the developing coronary arteries, highlighting region-specific molecular programs that govern arterial maturation and elongation.

[0287] Construction of a unified 3D reference human heart atlas

[0288] Applicant then computationally integrated 2D transcriptomic data and the spatially resolved histone modification maps within a unified 3D reference system (FIG. 19A, left).Building upon the variational autoencoder (VAE) framework for joint probabilistic modeling of multi-omic data, Applicant developed a novel integration framework that explicitly model spatial proximity between cells based on the attention of a neighborhood graph to generate transcriptome-wide spatial multi-omics 3D profiles (FIG. 19A middle, Methods). The reconstructed 3D multi-modal molecular hologram of the human developing heart then opens doors to many novel analyses, including the inference of spatially enriched cell-cell interactions (FIG. 19A right) Gene expression profiles from the MERFISH+ 2k gene panel (1,863 genes) and spatial histone modification data, originally acquired from 2D section, were mapped into the 3D heart transcriptome reference. Using this multi-modality integration, Applicant extended COMMOT to implement COMMON-COT (COordinate-aware MultiMOdal Neighbor-based Cell-Cell interaction analysis in 3D space) to infer spatially structured cellcell communication networks with GPU acceleration.

[0289] After integrating these three datasets, Applicant visualized the latent space using UMAP to evaluate model performance, which separated 32 transcriptionally distinct cell types with minimal batch effect (FIG. 19B). Applicant found that imputation accuracy (Spearman rank correlation between predicted gene expression and the input ground-truth expression) correlated strongly with spatial gene properties; genes with higher Moran’s I spatial autocorrelation exhibited better imputation performance (FIG. 19C). Applicant also evaluated the imputation accuracy between 2D heart and 3D heart in cell type level and found that genes’ imputation correlation increased with higher Moran’s I, validating the model’s spatial sensitivity (FIG. 19D). Spatial gene expression maps for representative genes, includingAtty. Dkt. No.: 114198-3660HEY2, IRX5, and VSTM2L, demonstrated high concordance between imputed 3D patterns and experimental measurements from 2D sections (FIG. 19E). Quantitative comparison between model-predicted and experimentally measured cell type-specific gene expression levels across the 2k MERFISH gene panel confirmed high accuracy (FIG. 19F).

[0290] Building on these validated spatial profiles, Applicant investigated spatially resolved cell-cell communication within the vascular niche in the 3D heart based on the above imputed profiles (FIG. 19). Applicant’s analysis of the coronary artery cell composition revealed distinct cell types distributed along the vessel’s longitudinal axis (FIGS. 18F and 181), suggesting diverse roles for these neighboring cells in coronary artery elongation and branching. This spatial gradient in cell distribution appears to be related to coronary artery growth. To investigate this further, Applicant performed 3D cell-cell interaction analysis focusing on VSMCs, examining their roles as both primary signal receivers and senders to characterize their interactions with neighboring cells (FIG. 19). Applicant found that both the CXCL-CXCR and SEMA3-NRP2-PLXNA signaling pathways were significantly enriched in cell-cell interactions involving VSMCs. Specifically, VSMCs acted as signal senders via SEMA3, targeting several cell types located at the surface of the coronary artery — including EPDCs, LECs, aFibro, and vFibro. This pattern suggests that VSMC-derived SEMA3 signals may function as repulsive cues, preventing these surface-associated cells from migrating into the inner layers of the coronary artery.

[0291] In parallel, VSMCs also sent CXCL ligands, which significantly targeted EPDCs. This suggests a potential balance between attractive (CXCL) and repulsive (SEMA3) cues that may regulate the positioning and fate of EPDCs — encouraging their migration to surround the coronary artery and contribute to the formation of supporting structures, such as smooth muscle cells and pericytes(REFER), while restricting their entry into the innermost vessel layers. When VSMCs were analyzed as signal receivers, EPDCs were identified as key senders of SEMA3 ligands, indicating that EPDC-derived SEMA3 may act to limit VSMC migration or expansion (REFER), thereby contributing to spatial organization and structural integrity of the developing coronary artery.

[0292] Along the vessel axis, the signal density received by vascular smooth muscle cells (VSMCs) and BEC exhibited a longitudinal gradient, with ligand-receptor interactions such as CXCL12-CXCR4 and VEGFA-FLT1 increasing from the root to the tip. This pattern isAtty. Dkt. No.: 114198-3660consistent with the reported role of CXCL12-CXCR4 and VEGFA-FLT1 signaling in coronary artery development. Radial analysis of the coronary artery wall, based on aggregate signal density along the inward-outward axis, revealed a sharp decline in VEGFA-FLT1 signaling interaction density with increasing distance from the vessel center (FIG. 19K, top). Peak signaling activity was observed at the inner vessel wall, gradually decreasing toward the outer regions (FIG. 19K, bottom). Notably, VEGFA-FLT1 signaling was enriched on the inner side of the vessel wall compared to the outer layer.

[0293] To further visualize these spatial dynamics in three dimensions, Applicant generated 3D spatial heatmaps of VEGFA ligand and FLT1 receptor distributions across heart sections, along with vector field representations illustrating the directionality and magnitude of VEGFA-FLT1 -mediated signaling flow (FIG. 19L). This spatial pattern is consistent with the anatomical organization of the vessel wall, where the endothelium resides on the inner side and is positioned to receive VEGFA signals from surrounding cells, including vascular smooth muscle cells and other neighboring cell types.

[0294] Applicant next imputed the 3D spatial chromatin modification landscapes. The imputed H3K4me3 and H3K27me3 histone mark distributions in 3D (FIG. 19J top panels) correlate well with that of the profiled 2D histone mark distribution (FIG. 19J bottom), with the correlation of r = 0.742 and r = 0.572b, respectively.

[0295] To facilitate data access and community exploration, Applicant developed two interactive web-based tools: Spateo-viewer, for 3D spatial transcriptomics visualization and analysis, and cell-cell interaction maps (the website is available at: viewer.spateo.aristoteleo.com / ), and MERFISHEYES (the website is available at: zhu.merfisheyes.com / ) for 3D visualization within the 3D reference system (FIG. 19H).

[0296] In summary, Spateo-VI uniformly integrates various 2D, 3D spatial transcriptomics datasets to create the first 3D multi-modal map of human fetal heart development, and offers a powerful framework for dissecting the molecular architecture of cardiac development in the 3D space.

[0297] Discussion

[0298] This study integrates three major methodological and data-driven advances that together establish a next-generation spatial genomics platform: 1) a molecular anchoringAtty. Dkt. No.: 114198-3660strategy, together with new microscopy and microfluidics instrumentation, extends the capabilities of MERFISH into its next high-throughout form, MERFISH+; 2) a resource capturing the first molecularly defined 3D atlas of the developing human heart; and 3) novel computational tools that enable multigene and multimodal imputations from 2D spatial imaging data onto a reconstructed 3D heart model.

[0299] Advancing MERFISH through molecular anchoring and instrumentation

[0300] Applicant introduced a new probe-anchoring protocol in which spatial transcriptomic probes are covalently linked to polyacrylamide hydrogels that are cast directly on top of the tissue sample. This approach markedly enhances sample stability and imaging longevity in MERFISH experiments. Unlike previous methods for increasing probe stability that focused on anchoring the targeted RNA(Wang et al. 2018) or enzymatically converting the targeted RNA into cDNA, here Applicant provide a cost-effective and enzymatic-free solution of anchoring the target hybridization probes rather than their substrate. Applicant demonstrated that this method of anchoring is highly efficient (capturing more than 50% of the initial signal - FIG. 13) and enabled stringent wash conditions between hybridization cycles that allows imaging experiments to be “reset” if needed, a particularly valuable feature when working with irreplaceable human specimens. Moreover, this anchoring chemistry supported multi-month continuous imaging with consistent signal performance, enabling hundreds of hybridization and imaging cycles without degradation. Here, Applicant demonstrated this advantage progressively by first imaging over 1,800 genes in a modular fashion in 2D heart sample and then integrating multi-modal imaging capturing the 50-nm scale of chromatin organization across tens of genomic loci, the nascent RNA and mature RNA across hundreds of genes and the nuclear accumulation of epigenetic marks via iterative immunofluorescence. Finally, with RNA stability and imagining time no longer limiting factors, Applicant re-engineered the microfluidics and microscopy system to enable 10-fold larger imaging area compared to what is commercially available for MERFISH or other imaging based spatial transcriptomics platforms. This complete imaging platform Applicant call MERFISH+ enabled the molecular profiling of more than 1.5 million cells per experiment at a cost of less than a thousand dollars in total reagent cost.

[0301] Building the first 3D molecular atlas of the human heartAtty. Dkt. No.: 114198-3660

[0302] Leveraging these advances, Applicant conducted two large scale MERFISH+ experiments. Empowered with precise 3D molecular alignment Applicant reconstructed a comprehensive three-dimensional map of the developing human heart at whole-organ scale. The resulting dataset captured 3.1 million cells across 34 distinct populations, including undercharacterized cell types localized to specialized cardiac structures such as valves and vascular compartments(FIGS. 17 and 18).

[0303] Using novel variational autoencoders designed specifically to take advantage of both transcriptional and spatial cues Applicant imputed the 1,863 gene measurement and the multimodal epigenetic marks measurements across the entire 3D heart model. This resource provides a precise molecular and anatomical framework for understanding how cellular populations assemble into large scale anatomical structures. Applicant have made this data available and explorable via Applicant’s visualization tools MERFISHEYES and Spateo. The reconstructed 3D atlas revealed molecular correlates of structural heterogeneity. Quantitative analysis of ventricular wall organization identified distinct contributions of trabecular, hybrid, and compact cardiomyocytes to local wall thickness. Similarly, 3D reconstructions of valves and coronary arteries exposed their cellular compositions and layered organization, which are features difficult to resolve from 2D data. Along the proximal-distal axis of coronary vessels, Applicant detected distinct spatial domains of endothelial and mural cell subtypes accompanied by ligand-receptor signaling gradients (e.g., SEMA3E and VEGF), highlighting potential mechanisms guiding vascular development. Additionally, these findings emphasize the need for tools capable of fully exploring coronary artery development in three dimensions, including along the long axis. Expanding the probe panels to include comprehensive well-characterized ligand-receptor pairs will greatly benefit coronary artery vessel research.

[0304] Broader implications and outlook

[0305] Beyond the cardiac focus within this manuscript, Applicant believe that the MERFISH+ platform and its computational frameworks establish a generalizable and scalable approach for large format spatial genomics. The same strategy can be extended to construct multimodal, spatially resolved cell atlases of other complex human organs, including the brain, kidney, and lung, where existing efforts have been largely limited to 2D. By combining rich volumetric datasets with emerging machine learning tools, including generative Al models for 3D tissue synthesis and disease simulation, MERFISH+ provides a foundation for building virtual organsAtty. Dkt. No.: 114198-3660that bridge molecular and anatomical understanding. Applicant anticipate that these integrated advances will accelerate both basic biological discovery and translational modeling across human organ systems.

[0306] Methods

[0307] Materials availability

[0308] Templates and reagents for making these probes can be purchased from commercial sources, as detailed in the Key Resources Table.

[0309] Data and code availability

[0310] Data reported in this paper are available at Dryad: DOI: 10.5061 / dryad.k98sf7mkx

[0311] Analysis code used for MERFISH-Plus probe library design, image decoding, and data quantification is available at: github.com / deprekate / LibraryDesigner, github.com / epigen-UCSD / MERlin, github.com / epigen-UCSD / MERFISH_Plus_Paper, github.com / ari stoteleo / MERFISHVI,github.com / ari stoteleo / Human Heart MERFISH Analy si s_2024

[0312] Table 4. Key resources tableAntibody Company Catalog number Dilution Reference Mouse Histone H3Cat# C15200146,trimethylated atDiagenode RRID:AB_292765 1:300 58 lysine 90(H3K9me3)Rabbit Anti-RNApolymerase II CTDCat# ab 193468,repeat YSPTSPSAbeam RRID:AB_290555 1:400 58 (phospho S2)7antibody[EPR18855]Mouse monoclonal[SC-35] to SC35 - Cat# ab11826,Abeam 1:400 27Nuclear Speckle RRID: AB_298608MarkerRabbit Histone Cat# 39133,H3K27ac antibody Active Motif RRID:AB_256101 1:400 58(pAb) 6Atty. Dkt. No.: 114198-3660Santa Cruz Cat# sc-376248,Mouse Lamin A / CBiotechnolog RRID: AB_109915 1:200 58 (E-l)y 36Cell Cat# 2598,NUP98 (C39A3)Signaling RRID: AB_226770 1:200 59 Rabbit mAbTechnology 0Alexa Fluor® 790JacksonAffiniPure Donkey Cat# 715-655-150ImmunoRese 1:1000 27 Anti-Mouse IgGarch Labs(H+L)Alexa Fluor® 647JacksonAffiniPure DonkeyImmunoRese Cat# 111-605-144 1:1000 27 Anti-Rabbit IgGarch Labs(H+L)

[0313] Table 5.Resource Source LinkSoftware and algorithms tissueSpate o (version: This paper https: / / github.com / 1.1.0) aristoteleo / spateo- releaseCellpose 2.0 Pachitariu, https: / / www.natureM., Stringer,.com / articles / s415C. Cellpose 92-022-01663-42.0: how totrain your ownmodel.

[0314] Experimental model and study participant details

[0315] Tissue samplesAtty. Dkt. No.: 114198-3660

[0316] Heart samples were collected in strict observance of the legal and institutional ethical regulations. The heart samples were collected under a University of California San Diego (UCSD) Human Research Protections Program Committee Institutional Review Board (IRB)-approved protocol (IRB number 172111) by the UCSD Perinatal Biorepository’s Developmental Biology Resource after informed consent was obtained from the donor families. All experiments were performed within the guidelines and regulations set forth by the IRB (IRB number 101021, registered with the Developmental Biology Resource). Ethical requirements for data privacy include that sequence-level data be shared through controlled-access databases.

[0317] Tissue processing

[0318] Tissue samples were dissected in buffer containing 125 mM NaCl, 2.5 mM KC1, 1mM MgCl₂, and 1.25 mM NaH₂PO₄ under a stereotaxic dissection microscope (Leica). Samples for MERFISH were washed 1 time with ice-cold PBS, then fixed in 4% PF A at 4°C overnight. On the second day, the sample was washed in ice-cold PBS three times, 10 minutes each, and were incubated in 10% and 20% sucrose at 4°C for 4 hours each, and in 30% sucrose overnight, followed by immersion with OCT (Fisher, cat# 23-730-571) and 30% sucrose (1 vol:l vol) for 1 hour. The sample was then embedded in OCT and stored at -80°C until sectioning.

[0319] A 12-week-old human heart was systematically cryosectioned in the transverse plane at 16 pm thickness per slice, starting from the apex and progressing toward the atria. Each tissue section was mounted sequentially and serially onto a set of eight large-format glass slides in a round-robin fashion. Each individual slide therefore contained 21 sections, representing every 8th section along the apical-to-atrial axis. This strategy ensured that each of the eight slides captured a spatially even and parallel representation of the entire heart, preserving the full diversity of cell types, anatomical zones, and structural features across all slides. The process continued until atrial tissue was no longer present, at which point the superior portion of the heart was separately sectioned and similarly distributed. For this superior region, the remaining heart tissue was sectioned and again distributed across a second set of eight slides, each containing 32 tissue sections. This maintained the same l-in-8 sampling strategy and ensured uniform representation of all anatomical regions across the second slide set.Atty. Dkt. No.: 114198-3660

[0320] The total number of sections is 53 x 8 = 424 sections with a mean thickness of 16 um, spanning a total length of 7 mm. One glass slide containing either 21 atrial sections or 32 ventricular sections was selected, and the tissue was transferred onto a salinized coverslip. The coverslip was then mounted into the newly developed fluidics system for downstream analysis. Subsequently, one coverglass was subjected to rapid fixation, stained with the same 238 genes. Image processing and decoding was performed using a modified version of MERlin, which integrates Cellpose for cell segmentation and includes an alternative algorithm for drift correction that uses DAPI staining rather than fiducial beads and can correct for drift along the z-axis. Applicant then performed a 3D reconstruction of the tissue sections using Spateo.

[0321] MERFISH gene selection (1863 genes) and probe library design and construction

[0322] Applicant developed custom primary probes using a Python-based workflow available at website: github.com / deprekate / LibraryDesigner. Probe sequences were selected following a previously established method, in which each gene was targeted by a set of 5 to 60 unique 40-mer sequences complementary to its mRNA.

[0323] Probe selection involved:1. Creating a k-mer index (17-mers) from the unspliced human transcriptome to assess sequence uniqueness2. Screening candidate 40-mer regions for off-target potential based on their 17- mer content3. Applying filtering criteria to retain only the most specific and effective sequences

[0324] Applicant then used transcript sequences from the human reference genome (hsl). Genes that were either too short or highly homologous to other targets were excluded, resulting in final panels of 238 and 1863 genes. To each 40-mer, Applicant appended three unique 20-nt read sequences per gene, along with universal PCR priming sites at both ends. All readout and primer sequences were checked against the human transcriptome to minimize crossreactivity.

[0325] The probes were synthesized from custom oligo pools (Twist Biosciences), amplified via limited-cycle PCR, and purified using Zymo DNA oligo cleanup kits (D4003). RNA wasAtty. Dkt. No.: 114198-3660transcribed in vitro using T7 polymerase (NEB), then reverse-transcribed into single- stranded DNA using Maxima reverse transcriptase (Thermo Scientific). RNA templates were removed through alkaline hydrolysis, followed by another round of DNA purification (Zymo Research, D4006). The forward primer used in PCR and reverse transcription carried a 5’-acrydite modification to enable covalent attachment of probes to the acrylamide hydrogels during tissue processing.

[0326] Linker Probes and Codebook Construction

[0327] Each gene-specific linker probe consisted of a 20-nt region complementary to the gene’s read sequence, followed by two tandem amplification binding sequences (each 20nt). Linkers were ordered from IDT in 384-well plates.

[0328] To assign spatial barcodes for MERFISH imaging, Applicant designed a binary codebook using a custom Python script, available at website: github.com / BogdanBintu / . A binary code was randomly assigned to each gene such that each barcode contained exactly four 1 ’ s across the chosen number of imaging rounds N, while maintaining a Hamming distance of four between codes to ensure error correction. Applicant further refined the codebook using a Metropolis-Hastings optimization algorithm62, minimizing the overlap of high-expression genes within the same hybridization round based on single-cell RNA-seq data from fetal heart tissue.

[0329] The final hybridization mixtures were prepared using a Tecan Fluent robotics system, which dispensed gene-specific linkers into 1.5mL Lo-Bind tubes according to their assigned hybridization rounds.[0330[ Amplification and Readout Probes[03311 Amplification probes were designed in two tiers. Level-1 probes included a 20-nt arm to bind to the linker overhang, followed by four binding sites for the Level-2 probes. Level-2 probes were similarly designed to bind Level- 1 probes and carry four additional sites for fluorescent readout probe binding. All amplification oligonucleotides were synthesized by IDT. Readout probes were 20-nt oligos, double-labeled with fluorescent dyes specific to the imaging channels: Cy3 (561nm), Cy5 (638nm), and Alexa Fluor 488 (488nm). These were purchased from Eurofins Scientific with high-performance liquid chromatography (HPLC) purification.Atty. Dkt. No.: 114198-3660[03321 Probe library design for RNA-MERFISH and chromatin tracing[033 | Applicant selected 40bp target sequences for DNA or RNA hybridization by considering each contiguous 40bp subsequence of each target of interest (the mRNA of a targeted gene or the genomic locus of interest) and then filtering out off-targets to the rest of the transcriptome / genome including repetitive regions, or too high / low GC content or melting temperature (TM). More specifically, Applicant’s probe design algorithm was implemented with three steps: 1) Build a 17-mer index based on reference genome hsl assembly (DNA) or the hg38 transcriptome (RNA). 2) Quantify 17-mer off-target counts for each candidate 40bp target sequence. 3) Filter and rank target sequences based on predefined selection criteria as previously described63,27. The primary probe set was constructed for pan-tissue imaging (human brain, heart, lung, and pancreas) and included up to 60 unique 40-mer sequences per gene, selected to hybridize to either mRNA or genomic DNA. Target sequence design was carried out as previously described (27,63, with modifications to ensure broad tissue compatibility. Target design proceeded through three main steps: (1) A 17-mer index was constructed from the human genome assembly hsl for DNA targets and the hg38 transcriptome for RNA targets; (2) Each contiguous 40-mer candidate from the target mRNA or genomic locus was scanned for potential off-targets by quantifying 17-mer matches across the genome or transcriptome; and (3) candidate sequences were filtered and ranked based on predefined criteria, including melting temperature, GC content, repetitive element masking, and minimal off-target potential. Genes that failed to meet design criteria due to sequence redundancy or insufficient unique regions were excluded. Ultimately, 1,826 genes were selected for the final probe set. For each probe, additional sequence handles were appended, including three 20-nt gene-specific read sequences and two 20-nt PCR primer sites. Read sequences and primers (Table S4) were selected to minimize cross-hybridization, using prior design frameworks27and screened against the human transcriptome. Oligonucleotide pools (Twist Biosciences) were PCR-amplified using limited-cycle PCR, incorporating a 5'-acrydite modified forward primer to enable acrylamide gel incorporation. Amplified DNA was used for in vitro transcription using T7 polymerase (NEB), followed by reverse transcription with Maxima RT (Thermo Scientific) and RNA strand hydrolysis to yield single-stranded DNA probes, which were then purified (Zymo Research D4003 and D4006 kits).

[0334] MERFISH gene selectionAtty. Dkt. No.: 114198-3660[03351 The MERFISH gene selection was performed by first using a scRNA-seq dataset from 12-15 p.c.w. heart and using NSForest v2 with default parameters to identify marker genes for the cell type clusters in this data. This list of genes was supplemented with additional marker genes from the literature. The target sequences for each gene were concatenated with one or two unique read sequences to facilitate MERFISH or smFISH imaging.

[0336] Design of chromatin probes

[0337] Applicant designed probes for DNA hybridization similarly as those for the RNA MERFISH as described in. Briefly, Applicant first partitioned FXYD7 locus into 10-kb segments. After screening against off-target binding, GC content and melting temperature, -100 unique 40 bp target sequences were selected for each 10-kb segment. Applicant concatenated a unique read sequence to the target probes of each segment to facilitate sequential hybridization and imaging of each locus.

[0338] Acrydite Probe Synthesis

[0339] The acrydite probes were generated using oligonucleotide pools, following the previously described method. Firstly, Applicant amplified the oligopools (Twist Biosciences) using a limited cycle qPCR (approximately 15 - 20 cycles) with a concentration of 0.6 pM of each acrydite primer to create templates. These templates were converted into RNAs using the in vitro T7 transcription reaction (New England Biolabs, E2040S). These RNAs were then converted back to single-stranded DNA using reverse transcription with a concentration of 16 pM of each acrydite primer. Subsequently, the DNA oligos were purified using alkaline hydrolysis to remove RNA templates and cleaned with columns (Zymo Research, D4060). The resulting acrydite probes were stored at -20 °C.

[0340] RNA MERFISH sample preparation

[0341] Fresh frozen hearts were sectioned at -20°C using a Leica CM3050S cryostat. Series coronal sections of 16 pm thickness were performed at -600 pm along the anterior-posterior axis of human hearts to capture the major cardiac structures. Sections were collected on salinized and poly-L-lysine (Millipore, 2913997) treated 40-mm, round no.1.5 coverslips (Bioptechs, 0420-0323-2). After drying for 20 min, tissue sections were stored at -80°C until use.Atty. Dkt. No.: 114198-3660

[0342] MERFISH measurements of 238 genes with 200 non-targeting blank controls was performed as previously described, using the published encoding sequences and readout probes (Table S3.2), subsequently pre-cleared by immersing into 30% (vol / vol) ethanol, 50% (vol / vol) ethanol, 70% (vol / vol) ethanol, and 100% ethanol, air dried for 5 minutes, then, 70% (vol / vol) ethanol, 50% (vol / vol) ethanol, each for 5 minutes. The tissue was then treated with Protease III (ACDBio) at 40 °C for 15 minutes prior to being washed with lx PBS for 5 minutes at room temperature. Then the tissues were preincubated with hybridization wash buffer (40% (vol / vol) formamide in 2x SSC) for ten minutes at room temperature. After preincubation, the coverslip was moved to a fresh 60 mm petri dish and residual hybridization wash buffer was removed with a Kimwipe lab tissue. In the new dish, 50 pL of encoding probe hybridization buffer (2X SSC, 50% (vol / vol) formamide, 10% (vol / vol) dextran sulfate, and a total concentration of 5 uM encoding probes and 1 pM of anchor probe: a 15-nt sequence of alternating dT and thymidine-locked nucleic acid (dT+) with a 5'-acrydite modification (Integrated DNA Technologies). The sample was placed in a humidified 47°C oven for 18 to 24 hours then washed with 40% (vol / vol) formamide in 2X SSC, 0.5% Tween 20 for 30 minutes at room temperature. Samples were post-fixed with 4% (vol / vol) paraformaldehyde in 2X SSC and washed with 2X SSC with murine RNase inhibitor for five minutes. To anchor the RNAs in place, the encoding-probe-hybridized samples were embedded in thin, 4% poly-acrylamide (PA) gels as described. Briefly, the hybridized samples on coverslips were washed with a degassed 4% poly-acrylamide solution, consisting of 4% (vol / vol) of 19:1 acrylamide / bis-acrylamide (BioRad, 1610144), 60 mM Tris-HCl pH 8 (ThermoFisher, AM9856), 0.3 MNaCl (ThermoFisher, AM9759) supplemented with the polymerizing agents ammonium persulfate (Sigma, A3678) and TEMED (Sigma, T9281) at final concentrations of 0.01% (wt / vol) and 0.05% (vol / vol), respectively, for 10min to equilibrate. Then the sample was reverted onto a glass plate that is cleaned and siliconized and spotted with 300ul of the 4% (vol / vol) of 19:1 acrylamide / bis-acrylamide (BioRad, 1610144), 60 mM Tris-HCl pH 8 (ThermoFisher, AM9856), 0.3 M NaCl (ThermoFisher, AM9759) supplemented with the polymerizing agents ammonium persulfate (Sigma, A3678) and TEMED (Sigma, T9281) at final concentrations of 0.03% (wt / vol) and 0.15% (vol / vol), respectively. The gel was then allowed to cast for 5 h at room temperature. The coverslip and the glass plate were then gently separated, and postfixed with 4% PF A, then the PA film was incubated with a digestion buffer consisting of 50 mMAtty. Dkt. No.: 114198-3660Tris-HCl pH 8, 1 mM EDTA, and 0.5% (vol / vol) Triton X-100 in nuclease-free water and 1% (vol / vol) proteinase K (New England Biolabs, P8107S). The sample was digested in the digestion buffer for ~5 h in a humidified, 37°C incubator and then washed with 2xSSC three times.

[0343] MERFISH measurements were conducted on a home-built system as described.

[0344] Multi-modal Chromatin tracing sample preparation

[0345] The sample preparation is similar to that for RNA MERFISH sample preparation with the exception that the sample is treated with 0.5% Triton X-100 in IX PBS for 10 minutes at room temperature instead of ethanol dehydration. Briefly, Frozen hearts were sectioned at -20°C using a Leica CM3050S cryostat. Series coronal sections of 14 pm thickness. Sections were collected on salinized and poly-L-lysine (Millipore, 2913997) treated 40-mm, round no.1.5 coverslips (Bioptechs, 0420-0323-2). After drying for 20 min, Fixed tissue sections were permeabilized with 0.5% Triton X-100 (Sigma-Aldrich, T8787) in lx PBS with RNase inhibitors for 10 min at room temperature and washed once with IX PBS with RNase inhibitors for 3 min. After permeabilization, coverslips were moved to a fresh 60 mm petri dish and treated with 0.1M hydrochloric acid (Thermo Scientific™, 24308) for exactly 5 min at room temperature, and washed three times for 5 min with lx PBS with RNase inhibitors. Tissue sections were preincubated with the hybridization wash buffer (50% (vol / vol) formamide (Ambion, AM9342) in 2x SSC (Coming, 46-020-CM)) for 10 min at room temperature. Prepared a layer of fresh parafilm with a 0.8 cm diameter hole in the center and placed in a fresh 60 mm petri dish. Then 50 uL of encoding probe hybridization buffer (50% (vol / vol) formamide (Ambion, AM9342), 2x SSC (Corning, 46-020-CM), 10% Dextran Sulfate (Millipore, S4030), 5 pM encoding probes and 1 pM of anchor probe, a 15-nt sequence of alternating dT and thymidine-locked nucleic acid (dT+) with a 5'-acrydite modification (Integrated DNA Technologies)) was added to the center. Coverslips were wiped away excess hybridization wash buffer with Kimwipes and placed tissue-side-down onto the droplet. Samples were in a humidity chamber at 47 °C for 18 - 24 hours. The sample was then washed with 40% (vol / vol) formamide in 2X SSC, 0.5% Tween 20 for 30 minutes and post-fixed with 4% (vol / vol) paraformaldehyde in 2X SSC followed by two washes in 2X SSC with murine RNase inhibitor for 5 min each at room temperature before being embedded in thin, 4% polyacrylamide (PA) gels as described for RNA MERFISH sample.Atty. Dkt. No.: 114198-3660

[0346] Microscopy and liquid handling equipment set up.

[0347] (1) Microscope Setup for Image Acquisition

[0348] Image acquisition was carried out using a custom-built microscope system, as previously described6’27’66. The core of the setup was a Nikon Ti-U microscope body fitted with a Nikon CFI Plan Apo Lambda 60x oil immersion objective (NA 1.4).

[0349] Illumination was provided based on one of two alternative systems.

[0350] Laser-based Illumination

[0351] Single-mode solid-state lasers (405, 560, 647, and 750 nm) were used, with the 560, 647, and 750 nm lasers controlled via an acousto-optic tunable filter (AOTF), and the 405 nm laser controlled via its internal control unit. Excitation and emission paths were separated using a custom dichroic (Chroma, zy405 / 488 / 561 / 647 / 752RP-UFl) and emission filter (Chroma, ZET405 / 488 / 461 / 647–656 / 752 m).

[0352] Light Engine Illumination

[0353] A Lumencor CELESTA light engine (wavelengths: 405, 477, 546, 638, 749 nm) was used with a penta-bandpass dichroic (IDEX, FF421 / 491 / 567 / 659 / 776-Di01) and filter (IDEX, FF01–441 / 511 / 593 / 684 / 817).

[0354] In most experiments, beam uniformity was enhanced using either a refractive beam shaper (Newport Optics, GBS-AR14) or a vibrational optical fiber (Errol, custom Albedo unit). Images were acquired with a scientific CMOS camera (Hamamatsu FLASH4.0 or C13440). Sample positioning was controlled with a motorized XYZ stage (Ludl), and focus was maintained using a custom-built autofocus system67, employing IR lasers (Thorlabs LP980-SF15) and a secondary CMOS camera (Thorlabs, uc480).

[0355] All hardware components were synchronized and operated via a National Instrument DAQ card (NI PCIe-6353) and a custom software.

[0356] (2) Fluidics System

[0357] The fluidics system consisted of a syringe pump (Gilson MINIPLUS 3), a 24-port Vici valve, a flow chamber (Bioptechs 060319-2), and tubing sealed with pressure adhesive (Blu-tack). Each valve was connected to a distinct buffer, with dedicated lines for imaging,Atty. Dkt. No.: 114198-3660stripping, and washing. Buffers were delivered to the sample chamber, and waste was collected downstream in an open-loop setup.

[0358] The system supported hundreds of hybridization rounds, depending on buffer configuration. For experiments exceeding this capacity, buffer replacement was performed by:1. Disconnecting the chamber and flushing the valve lines with 30% formamide and water.2. Introducing new buffers.3. Reconnecting the chamber and resuming imaging.

[0359] (3) Experimental Control Software

[0360] All components were controlled via custom software available at https: / / github.com / ZhuangLab, composed of the following modules:

[0361] 1. Hal: Controls illumination, microscope components, and imaging parameters.

[0362] 2. Steve: Acquires mosaic images and selects imaging regions.

[0363] 3. Kilroy: Controls fluidics and executes buffer exchange sequences.

[0364] 4. Dave: Automates the experiment by integrating Hal and Kilroy operations.

[0365] At experiment start, imaging parameters and buffer sequences were configured in Hal and Kilroy. A DAPI mosaic was acquired using Steve to define imaging regions. The full experiment protocol was loaded into Dave for automated execution. For long experiments, buffer exchanges were manually performed between Dave runs as needed.

[0366] (4) CNC liquid handler construction and configuration.

[0367] Table 6. The CNC liquid handler consists of the following partsVendor Catalogue DescriptionHamilton 63133-01 PSD6, SYRINGE PUMP, STD, ROHSHamilton 64316-04 PSD-4 OEM STARTER KITHamilton 8300-45 SYRINGE, TLL, 200 pl, LONG LIFE, UHMWHamilton 7750-11 KF720 NDL 6 / PK (20 / * / *) Configuration: DEFAULTAtty. Dkt. No.: 114198-3660Hamilton 57252-01 Y VALVE ASSY, CTFE, HV, SD4 / PS D8Hamilton 99893-01 VALVE, HV 3-5-PCTFE

[0368] Imaging and adaptor / readout hybridization protocol

[0369] MERFISH measurements were conducted on a custom microscope-microfluidics system with the configuration previously described. Briefly, the system was built around a high-speed and high-sensitivity camera (the Kinetix model from Teledyne) capable of imaging twice the area of the prior model of Hamamatsu CMOS cameras. A newer line of Nikon 60x oil immersion objective (1.4 NA) was also used for better spherical and chromatic aberration corrections. The different components were synchronized and controlled using a National Instruments Data Acquisition card (NI PCIe-6353) and custom software.

[0370] To enable multi-modal imaging Applicant first sequentially hybridized fluorescent readout probes and then imaged the targeted genomic loci and then the targeted mRNAs. Then Applicant performed a series of antibody stains and imaging. Specifically the following protocol was used in order:1) 52 rounds of hybridization and chromatin tracing imaging, sequentially targeting the 21 loci using 1 -color imaging272) 16 rounds of hybridization and MERFISH imaging, combinatorially targeting 298 genes using 3 -color imaging3) 21 rounds of hybridization and smFISH imaging, targeting 21 exons and 21 introns using 2-color imaging4) 3 rounds of sequential staining for 6 different antibodies using 2-color imaging. The protocol for each hybridization included the following steps:1. Incubate the sample with adaptor probes for 30 minutes for DNA imaging or 75 minutes for RNA imaging at room temperature.2. Flow wash buffer and incubate for 7 minutes3. Incubate fluorescence readout probes (one for each color) for 30 minutes at room temperature.4. Flow wash buffer and incubate for 7 minutes.Atty. Dkt. No.: 114198-36605. Flow imaging buffer. The imaging buffer was prepared as described previously27and additionally included 2.5μg / mL of DAPI.Following each hybridization the sample was imaged and then the signal was removed by flowing 100% formamide for 20 minutes and then re-equilibrating to 2XSSC for 10 minutes.

[0371] RNA-MERFISH measurement of the 298 gene panel was performed with an encoding scheme of 48-bit binary barcode and a Hamming weight of 4 (HW4). Therefore, in each hybridization round, ~75 adaptor probes were pooled together to target a unique subset of the 298 genes.

[0372] Immunofluorescence staining

[0373] Antibody imaging was performed immediately after completing the DNA and RNA imaging. The sample was first stained for 4 hours at room temperature using 2 primary antibodies of two different species (mouse and rabbit), washed in 2XSSC for 15 minutes and then stained for 2 hours using 2 secondary antibodies for each target species conjugated with fluorescent dyes.

[0374] Image acquisition: After each round of hybridization, Applicant acquired z-stack images of each FOV in 4 colors:750 nm, 647nm, 560 nm and 405 nm. Consecutive z sections were separated by 300 nm and covered 15 um of the sample. Images were acquired at a rate of 20 Hz.

[0375] DATA ANALYSIS

[0376] Analysis of 2D heart sections with acrydite-modified MERFISH probes

[0377] MERFISH data imaged using acrydite-modified MERFISH probes were analyzed using MERlin vO.6.1 exactly as previously described8. Cell types were assigned and UMAPs generated using the “ingest” function of the Scanpy tool to map the newly generated data onto the previously published MERFISH data.

[0378] Analysis of 2D heart sections with MERFISH+ probes

[0379] MERFISH data in FIGS. 14 and 15 were processed with custom scripts developed for MERFISH+ data.

[0380] Image Pre-processingAtty. Dkt. No.: 114198-3660[03811 Flat-field correction[03821 Images were first corrected to remove variations in illumination present on the edges and corners of images (flat-field correction) by calculating the per-pixel median brightness across the first round of imaging, separately for each color. Each image was then inversely scaled by this median brightness.

[0383] Deconvolution

[0384] Next, the images were deconvoluted using the Wiener algorithm with a custom point spread function (PSF) calculated for the specific microscope used for imaging to reduce noise and enhance the discreteness of fluorescent spots. Finally, a high-pass filter was used with a blur sigma of 30 to further improve signal-to-noise ratio.

[0385] Localization of Fluorescent Spots

[0386] After pre-processing images, local maxima within a 1 -pixel radius were identified and further filtered to enforce a brightness of at least 3600 and a correlation with the point spread function of 0.25. The remaining local maxima were used for decoding.

[0387] smFISH quantification

[0388] For quantification of transcripts imaged individually as smFISH targets, the fluorescent spots identified in the previous step were further filtered based on a minimum brightness relative to their local background. This threshold was selected separately for each color channel and manually tuned to remove noisy and dim fluorescence while keeping bright puncta.

[0389] MERFISH decoding

[0390] To resolve RNA molecules from the identified fluorescent spots across imaging rounds, spots were first clustered into groups across images within a range of 2 pixels after correcting for drift using the same method described above (MERlin). Due to the MERFISH codebook designed with a hamming weight of 4, clusters with at least 3 spots were kept to allow for single-bit-error correction, while clusters with 1 or 2 spots were discarded. For each cluster, a brightness vector was constructed using the spot brightness in each round, then L2 normalized and compared with the codebook similar to MERlin described above.Atty. Dkt. No.: 114198-3660[03911 To filter noise from the resulting decoded transcripts, a score was calculated for each transcript based on three metrics: 1) the average brightness across the fluorescent spots, 2) the distance of the brightness vector to the nearest spot in the MERFISH codebook, and 3) the average distance of each fluorescent spot to the median of all spots. These metrics were combined into a single score by calculating the combined Fisher p-value against the distribution of each metric across all decoded transcripts. The distribution of scores for decoded transcripts with a blank barcode identity were then compared to the distribution of scores for decoded transcripts with a gene identity, and a score threshold was chosen which separates the peaks of each distribution.

[0392] Cell Segmentation

[0393] Cell boundaries were determined in 3 dimensions (3D) using Cellpose v269. First, for each z-slice image, a 2D segmentation mask was generated using the DAPI channel with Cellpose’s “dapi” model, after flat field correcting and deconvoluting the DAPI image as described in “Image pre-processing”. Cells were then linked between adjacent z-slices to create a 3D segmentation.

[0394] Assigning transcripts to cells

[0395] Transcripts identified by both smFISH and MERFISH were assigned to cells by first adjusting the coordinates of each molecule by the calculated drift between the images used for segmentation and those in which the RNA were imaged. Next, the transcript coordinates were rounded to the nearest integer and assigned to the cell based on the value of the segmentation mask at each transcript’s coordinates.

[0396] Analysis of 3D heart reconstruction

[0397] Applicant firstly preprocessed the 3D heart data using MERlin as described in Farah et al., 20248, and then reconstructed the 3D heart map using Spateo15, with the following minor modifications for each method:

[0398] Localization of RNA molecules

[0399] MERFISH images generated for the 3D heart reconstruction were analyzed using a modified version of MERlin vO.6.1 to decode individual RNA molecules. Briefly, these modifications added support for images in compressed Zarr format and extended the driftAtty. Dkt. No.: 114198-3660correction functionality to allow for the calculation of drift based on DAPI images instead of fiducial beads, as well as correcting drift along the z-axis (perpendicular to the microscope slide). Images were aligned across hybridization rounds by identifying local maxima and minima in the DAPI channel and calculating the offset of these anchor points using fast Fourier transform. A high-pass filter and deconvolution using a Gaussian point spread function were applied to reduce noise and background of images, then a vector was constructed for each pixel representing the brightness of that pixel across all imaging rounds. An L2 normalization was applied to these vectors to standardize their total brightness, and then the Euclidean distance to each of the barcodes in the MERFISH codebook was calculated. For this calculation, each barcode was represented as a similar L2-normalized brightness vector with a value of 0 in the off-bit rounds and 1 in the on-bit rounds pre-normalization. Each pixel was assigned the gene identity of the nearest barcode, provided the distance was less than 0.6, which was chosen to capture distances representing single bit errors. Adjacent pixels with the same gene identity were combined as a single RNA molecule. Scale factors to equalize the brightness across imaging rounds were calculated by decoding a random sample of 50 z-slices and keeping RNA molecules with at least 5 pixels, then calculating the factors necessary to equalize the mean brightness of the pixels in these molecules in each round. This process was repeated iteratively 10 times, with brightness in each iteration normalized using the scale factors calculated in the previous iteration. The RNA molecules were filtered based on 3 parameters: the number of pixels in the molecule, the mean brightness of the pixels, and the minimum distance to the barcode among the pixels. A 3-dimensional histogram was constructed using these parameters, and then for each histogram bin, the ratio of gene barcodes to blank barcodes was calculated. The barcodes from the bins with the lowest gene-to-blank were removed until a misidentification rate of 5% was reached.

[0400] Cell Segmentation

[0401] Cell segmentation was performed using the same method as above using Cellpose v2.

[0402] Cell clustering

[0403] The cells were first clustered using the Leiden algorithm and manually labeled with major cell types based on gene expression. Clusters with the same major cell type label were merged. Next, each major cell type was sub-clustered individually and manually labeled basedAtty. Dkt. No.: 114198-3660on gene expression and the spatial pattern of the cells’ locations. Some clusters were identified as artifacts. Further inspection of these artifact clusters revealed that some were the result of out-of-focus images in specific rounds of imaging, affecting 72,142 (2.3%) cells. Because only a subset of genes was affected by these problematic images, the unaffected genes were used to impute the correct cell type by nearest-neighbor classifier using cells outside the affected regions. Other artifact clusters were identified to consist of low-quality cells arising from segmentation artifacts or low transcript counts and were removed from the dataset.

[0404] Slice alignment and 3D reconstruction of the human heart

[0405] In order to reconstruct the 3D map of the human developing heart, Applicant utilized Spateo 15 Applicant previously developed and extensively demonstrated its favorable performance when compared to PASTE, PASTE2, Moscot, STAlign, etc.. In brief, Spateo leverages the transcriptional similarity between cells in adjacent slices, together with a spatial constraint on the transformation of the slice, to align the 53 slices of the developing human heart. Specifically, Spateo relies on a Bayesian generative model to efficiently and robustly align the adjacent slices, and to output a probabilistic mapping and a rigid or non-rigid transformation. Although Spateo’ s non-rigid alignment mode can eliminate distortions introduced by tissue sectioning, the cardiac cavities are relatively small, resulting in minimal distortion. Additionally, Applicant cannot determine whether the differences between sections are caused by the sectioning process or represent genuine biological variation. Therefore, Applicant opted for Spateo’ s rigid alignment mode. Furthermore, due to partial overlap between certain sections (such as sections 21 and 22), Applicant set the partial robust level parameter of Spateo’ s morpho align function to 25 for these sections to achieve more precise alignment and reconstruction results. For the input gene expression features, Applicant utilized principal component analysis (PCA) features processed by the Scanpy Python package, which provides a denoising effect compared to the original raw counts. Default values are used for all other parameters in Spateo’ s morpho align alignment function. The entire reconstruction process iteratively applied Spateo’ s alignment method along the z-axis, enabling us to precisely reconstruct a 3D human heart containing 3.1 million cells. All computations were performed on a high-performance server running Ubuntu 22.04.5, equipped with Intel Xeon Platinum 8280 processor, 1.4 TB of RAM, and NVIDIA A100 GPUs with 40 GB memory. The complete reconstruction process required approximately 1 hour of processing time.Atty. Dkt. No.: 114198-3660

[0406] 3D Visualization[0407| After the 3D reconstruction of the human heart, Applicant utilized Spateo’s built-in 3D rendering method for visualization, which is built upon PyVista70and provides convenient and realistic rendering capabilities. For 3D visualization of cell types, Applicant used Spateo’s construct _pc function to convert the h5ad data structure into VTK files. For scenarios displaying 3D scalar values, such as gene expression levels in 3D, Applicant additionally employed Spateo’s add model labels function to insert the desired scalar values into the VTK files. To address potential occlusion issues, e.g., when some parts of the data (e.g., cells, regions, or structures) block or obscure other parts from view, in 3D visualization, Applicant rescaled the scalar values to be between 0 to 1 range and use them as opacity values. Finally, Applicant used Spateo’s three_d_plot function for 3D rendering.

[0408] Morphometric measurement

[0409] Different cell types occupy distinct positions, volumes, and exhibit varying densities within the heart. Accurately measuring these values is crucial for characterizing the three-dimensional composition and organization of the heart.

[0410] Density measurement of each cell type

[0411] For different cell types, Applicant calculated their ^-nearest neighbor density to more accurately capture the local density distribution of cells. Specifically, for each cell, Applicant first calculated the spatial distance d to its A th nearest neighbor of the same cell type (where k=30 in this work), then constructed a sphere with radius d. The density was then defined as the number of cells of the same cell type within this sphere divided by the volume of the sphere, where the volume is defined as 4 / 3πr3:

[0412] Volume measurement of each cell type

[0413] For different cell types, Applicant employed the marching cubes algorithm to construct surface meshes from cell point sets, ensuring to tightly enclose the space occupied by each cell type, and subsequently measured the volume of the constructed surface meshes. It’s important to note that cell type annotations in this study are solely based on gene expression, which may result in spatially dispersed patterns for a small portion of cells. To accurately measure the volume occupied by the majority of cells, Applicant needed to eliminate these dispersed cells.Atty. Dkt. No.: 114198-3660Specifically, Applicant first calculated the density of each cell from a particular cell type. Applicant then removed cells with density less than 2% of all cells, which primarily represented the dispersed cells. The volume calculation is then based on remaining cells.

[0414] Ventricular surface analysis

[0415] To extract ventricular surface cells, Applicant first constructed a vCM surface mesh using vCM, Epicardial, Fibro, LEC, BEC, Endocardial, and Pericyte cells through Spateo’s construct surface function and the marching cubes algorithm. Subsequently, Applicant scaled the vCM surface mesh to 90% of its original size, and identified the ventricular surface cells as those located between the original vCM surface mesh and the 90% scaled surface mesh. Next, Applicant analyzed the cellular composition of the ventricular surface cells and the complexity of their neighboring cell types.

[0416] Neighborhood complexity calculation

[0417] Applicant aimed to measure the complexity of cellular composition in the microenvironment surrounding each cell. To achieve this, Applicant first needed to define a cell’s neighborhood. Given that the inter-section gap is 150pm, using a fixed radius of inter / intra-section neighbors would require an extremely large radius to include cells from adjacent sections, thus losing the meaningful definition of the intra-section neighborhood. To reconcile this, Applicant directly used a circle with radius?T in the current 2D section space to define the neighborhood, while for adjacent sections, Applicant first mapped the current cell’s spatial location to the adjacent sections and then defined the neighborhood using a circle with radiusr2. The final neighborhood of the current cell was determined by the union of these defined regions. This definition allows the neighborhood to include cells from other sections while ensuring obtaining meaningful neighbor cells and making the computation more efficient by constraining calculations within each section. In this work, Applicant setr1 to 100 μm andr2 to 30 μm. Next, Applicant calculated the proportion of each cell type within the neighborhood - of each cell, denoting the proportion of cell type m as Pm. To represent the complexity of the surrounding cellular environment, Applicant computed its entropy value as follows:Entropy = − ∑pmlog pm, ∑pm= 1mAtty. Dkt. No.: 114198-3660

[0418] A high entropy indicates a complex cellular composition in the neighborhood while low entropy indicates a more uniform cellular composition.

[0419] Cell-type co-existence analysis

[0420] Applicant aimed to identify pairs of cell types that are more likely to co-exist, thereby potentially uncovering spatial relationships between different cell types. To this end, Applicant first constructed cell neighborhoods according to the definition described in the “Neighborhood complexity calculation” section and calculated the proportional composition of cell types. Given N cells and M cell types, Applicant obtained an NxM matrix, where the z-th row represents the proportional composition of cell types in the neighborhood of the z-th cell. From a column perspective, Applicant obtained the distribution of the proportional composition of each cell type across all cell neighborhoods, described by an Nxl vector. Then, for any two different cell types ml and m2, Applicant calculated their Pearson correlation coefficient between two Nxl vectors, which represents a co-existence score between cell types. A high Pearson correlation coefficient between two cell types indicates that these cell types tend to either co-exist or be absent together in most neighborhoods. To further discover clusters of spatial co-existence patterns of cell types, Applicant constructed a dendrogram structure based on the cell type correlation coefficient matrix. First, Applicant calculated the pairwise similarity between cell types using SciPy’s pdist function with the “cityblock” metric. Then, Applicant constructed the dendrogram structure using ’SciPy’s linkage function with the “single” method. To visualize the co-existence relationships between cell types more intuitively, Applicant further utilized the igraph Python library to visualize the network relationships between cell types. In the network, nodes represent different cell types, and edges are undirected with weights representing co-existence scores. To reduce spurious correlations, Applicant filtered out connections with co-existence scores below 0.07. Applicant manually annotate the dendrogram with four major niche communities (Atrium community, Trabecular community, Epicardium community, and Compact ventricular community) and further divide the second group into three sub-communities (Coronary artery vessel community, VIC community, Neural and glial community).

[0421] Neighborhood cell type enrichment analysisAtty. Dkt. No.: 114198-3660[04221 Leveraging Applicant’s comprehensive 3D spatial information, Applicant conducted an in-depth analysis of cell-type distributions and their spatial relationships within the human heart. Applicant focused on two distinct spatial distribution patterns to characterize the cellular organization: the radial distance distribution from target cell types and the distribution along the anterior-posterior (AP) axis.[04231 Distance-based distribution[04241 For each cell in the human heart tissue, Applicant quantified its spatial relationship to target cell populations by calculating the minimum distance to the nearest target cell. Applicant then systematically analyzed the cell-type compositions within concentric shells at 25pm intervals, ranging from 0-25pm to 200-225pm. To accurately represent the relative enrichment patterns, Applicant normalized these distributions by dividing the observed proportions by the background cell-type frequencies, thereby highlighting specific spatial associations.

[0425] Apex-base (AB) axis distribution

[0426] In parallel, Applicant examined the cellular organization along the AB axis by implementing cell-type-specific distance thresholds. Specifically, Applicant analyzed the cellular composition within a 200pm radius for valvular interstitial cells (VIC) and a 100pm radius for vascular smooth muscle cells (VSMC). The enrichment levels were quantified by normalizing against background cell-type frequencies, revealing distinct patterns of spatial organization along this anatomical axis.10427} Constructing artery vessels network

[0428] Applicant initially identified coronary vessel cells based on MYH11 enrichment. Specifically, the top 3% of cells with the highest MYH11 expression were selected as a rough set of coronary vessel cells. Cells belonging to the aorta (AO) and pulmonary artery (PA) were then excluded according to their spatial distribution. To further refine the selection, Applicant applied spatial density estimation to filter out randomly scattered outliers (performed in three iterations, removing the lowest-density 10%, 3%, and 0.5% of cells, respectively, see section “Density measurement of each cell type"'). This procedure yielded a high-quality set of coronary vessel cells.Atty. Dkt. No.: 114198-3660

[0429] To identify each coronary vessel on the cross-section, Applicant implemented a spatial clustering method using the Density-Based Spatial Clustering of Applications with Noise72(DBSCAN) algorithm from the scikit-learn library. This approach enabled us to identify and segment distinct vessel cross-sections within each tissue slice, which were subsequently designated as leaf nodes in Applicant’s vascular tree structure. DBSCAN parameters were empirically optimized, with the neighborhood radius (eps) set to 50 pm and the minimum cluster size (min samples) set to 5. To ensure robust vessel identification, Applicant excluded noise points (cells assigned with label -1) from the analysis.

[0430] Leveraging the inherent hierarchical, top-down organization of the cardiac surface vasculature, Applicant developed a top-down tree construction algorithm. For each vessel node identified in slice z, Applicant established hierarchical connections by computing a connection score between the current node and each visited node, defined as:Score(N1, N2) = wℒangle(N1, N2) + (1 − w)ℒeuc(N1, N2)where the angular and Euclidean terms are:ℒangle(N1, N2) = 1 − (1 / π) arccos((X(N2) − X(N1)) · (X(N2) − X(P(N2))) / (||X(N2) − X(N1)|| ||X(N2) − X(P(N2))||))ℒeuc(N1, N2) = ||X(N2) − X(N1)||2

[0431] Here tV'i denotes the current node, ^'2 a visited node, ^ (’) the centroid of cells in a node, and the parent of a node. The hyperparameter w balances the contributions of the two terms and was set to 0.5. At each step, the current node was assigned to the visited node with the lowest connection score, after which the current node was added to the visited set. To maintain biological plausibility, node connections within the same section were disallowed, and only connections to nodes in higher layers were permitted. In cases where the top-down relationship is not strictly satisfied (e.g., at the top of the LCX, where the vessel initially ascends before descending), predefined manual connections were introduced to preserve vascular continuity.

[0432] Thickness calculation for heart layer

[0433] To quantify the 3D thickness between different layers, Applicant first reconstructed the surface geometry of each layer using a mesh-based representation. Since the interventricularAtty. Dkt. No.: 114198-3660septum (IVS) region may affect the thickness estimation and is not of Applicant’s primary interest, Applicant subsequently identified and removed this region. Finally, based on the reconstructed meshes, Applicant computed the physical 3D thickness between layers. The detailed steps are described as follows.

[0434] Construction of Layer Surfaces

[0435] Applicant began by defining the cell types contained within each layer to construct a dense, non-hollow point cloud. For instance, the mesh of the epicardial layer was supported by both left and right ventricular cardiomyocytes (vCMs) together with epicardial cells. To obtain a spatially compact and continuous cell distribution, Applicant performed density-based filtering of the discrete cells. Specifically, for cell types with left-right distinctions, Applicant manually separated left and right cells and then applied the density filtering procedure described in Section “Density measurement of each cell type”. Applicant iteratively removed the lowest 1% of low-density cells until the remaining cells occupied a spatially solid region.

[0436] Next, Applicant reconstructed the surface mesh using the st.tdr.construct surface function in Spateo, employing the marching-cubes algorithm with parameters “mc_scale_factor”: 1.05, ‘dist_sample_num’: 1000, and Tevelset’: 0.4. Because the spacing between slices was relatively large, Applicant compressed the z-axis of the point cloud to one-fifth of its original scale during mesh construction, and subsequently rescaled it fivefold to restore the true physical z-axis dimension. This approach effectively prevented the generation of internal voids. Using this procedure, Applicant reconstructed three major layers: epicardial, endocardial, and hybrid layers.

[0437] Identification of the IVS Region

[0438] The IVS region is generally located between the left and right endocardial layers. Thus, Applicant first identified the IVS on the endocardial meshes. For the left endocardial mesh, Applicant computed the normal vector at each vertex and determined whether the normal intersected the right endocardial mesh. If such an intersection existed, the vertex was labeled as belonging to the IVS region. The same procedure was applied symmetrically to the right endocardial mesh.

[0439] After identifying IVS regions on the endocardial surfaces, Applicant further detected IVS regions on the epicardial mesh. Again, Applicant used the normal-intersection methodAtty. Dkt. No.: 114198-3660and classified the intersection patterns into five cases: (1) The normal vector did not intersect any endocardial mesh. (2) The normal vector intersected the right endocardial mesh, and the intersection point was labeled as part of the IVS. (3) The normal vector intersected the left endocardial mesh, and the intersection point was labeled as part of the IVS. (4) The normal vector intersected the right endocardial mesh but was labeled as a non-IVS region. (5)The normal vector intersected the left endocardial mesh but was labeled as a non-IVS region. Applicant defined cases (1)— (3) as IVS regions, while cases (4) and (5) were assigned to the right and left walls, respectively.

[0440] Computation of 3D Physical Thickness

[0441] Based on the reconstructed layer meshes, Applicant calculated the 3D physical thickness within the non-IVS regions on the epicardial surface. For each vertex on the epicardial mesh, Applicant first computed its surface normal vector and then determined the intersection points of this normal with the endocardial, hybrid, and Purkinje layer meshes. The distances between consecutive intersection points along the normal direction were defined as local layer thicknesses.

[0442] Specifically, the distance between the Purkinje layer and the endocardial layer was defined as the trabecular layer thickness, the distance between the endocardial layer and the hybrid layer was defined as the hybrid layer thickness, and the distance between the hybrid layer and the epicardial layer was defined as the compact layer thickness. This pointwise computation provided a spatially resolved map of heart layer thickness, enabling quantitative comparison across different regions of the heart.

[0443] Gene expression gradient analysis

[0444] Based on the hierarchical tree structure of the vasculature, Applicant identified the path from each vascular cell to its corresponding root of the vessel. Applicant can then perform path-dependent differential gene expression analyses to discover gene expression patterns along the vascular hierarchy.

[0445] Vessel branching analysis

[0446] Building upon the Census statistical framework73, Applicant reformulated the identification of differential vascular gene expression as a model comparison problem betweenAtty. Dkt. No.: 114198-3660two distinct negative binomial Generalized Linear Models (GLMs). Specifically, Applicant established two hypotheses, each correspond to a different model:The null model

[0447] NB(Counts)∼smf.glm(Geodist)assumes the gene being tested is not a branch specific gene, representing a uniform expression pattern across vessel branches, whereas the alternative model NB(Counts)∼smf.glm(Geodist)+Branch+smf.glm(Geodist):Branch accommodates branch-specific expression patterns, allowing for differential gene expression across distinct vessel branches, where: represent an interaction term between branch and geodist and NB means negative binomial distribution. “Geodist” indicates the geodesic distance of each cell from the root vessel. Both models employ natural splines (implemented with three degrees of freedom) to characterize the continuous progression of gene expression along geodist coordinates. The null model constrains the expression trajectory to a single smooth curve across all branches, whereas the alternative model allows for branch-specific expression trajectories, fitting separate curves for each vessel branch. Applicant implemented GLM regression using the smf.glm function from the statsmodels Python library, employed the cr function from the patsy Python library for natural regression splines, and utilized the chi2.sf function from the SciPy Python library for likelihood ratio testing.

[0448] Correlation of RNA expression measurements

[0449] All pseudo-bulk correlations shown in this work are Pearson correlations calculated using the log10(x + 1) transformed mean transcript count per gene across cells based on the raw cell by gene tables (UMI count for scRNA-seq or localized transcript count for MERFISH and smFISH). Cell type to cell type correlation heatmaps were generated by first calculating the mean transcript count per gene among the cells of each cell type, then transforming each mean into a z-score using the distribution of means for each gene across cell types. The Pearson correlations of these z-scores between every pair of cell types were used for the heatmap values, and the average of the maximum correlation for each cell type to any other cell type was used to quantify the overall correlation.

[0450] Spateo-VI integrates various 2D and 3D spatial transcriptomics and other modalities

[0451] Spatial Graph ConstructionAtty. Dkt. No.: 114198-3660[0452 | To leverage spatial context, Applicant construct a graph of cell neighborhoods based on each cell’s spatial coordinates. Applicant provide two options for graph construction: K-nearest neighbors (KNN) and Delaunay triangulation. In KNN mode, each cell is connected to its k nearest neighboring cells in space, yielding an adj acency graph that captures local spatial proximity. In Delaunay mode, Applicant perform Delaunay triangulation on the coordinates, connecting cells that share a triangle edge, which tends to capture both nearest and slightly farther neighbors in a planar arrangement. These options are configurable (e.g., spatial_graph_type=“knn” or “delaunay”). For KNN, Applicant use a default of 10 neighbors (or fewer if fewer cells) and ensure the graph is symmetric (if cell i is among cell j’s neighbors, Applicant adds the reciprocal edge). For Delaunay, Applicant triangulate the points (adding a small jitter if points are collinear or ID) and connect all cells that form the vertices of a common Delaunay simplex. The result in both cases is an undirected graph G = (V, E) with nodes as cells and edges representing spatial adjacency. This graph serves as a prior for sharing information between neighboring cells in the model. If the graph happens to be disjoint or have isolated nodes, Applicant issue a warning and can fall back to a fully-connected or KNN graph to ensure all cells are reachable. The spatial graph construction is a preprocessing step that encodes spatial structure for downstream use in the model’s inference process.

[0453] Graph Attention-based Spatial Encoder

[0454] Given the spatial neighbor graph, Applicant’s model employs a Graph Attention Network (GAT) to incorporate spatial information into the cell representations. Starting from each cell’s initial latent representation (described below), Applicant apply a graph attention layer that propagates information across the neighbors of each cell. Specifically, Applicant use a multi-head GATv2 convolution layer, which allows each cell to attend to its neighbors with learned attention weights. Formally, for a cell i with latent vector zt, and neighbors j ∈ N(i) on the spatial graph, the graph attention layer computes a new featureas a weighted aggregate of neighbor features:hi = o’ aij Wzj lj∈N(i)Atty. Dkt. No.: 114198-3660where V is a learnable weight matrix and is the attention coefficient for neighbor j of i. The attention coefficients atj are computed via a softmax that considers z’ s neighbors’ features, enabling the network to assign different importance to different neighbors. Multi-head attention is supported (configurable via attention heads) so that the model can attend to multiple aspects of the neighborhood; the outputs from multiple heads are averaged (Applicant do not concatenate, ensuring the output dimension remains fixed). This GAT-based message passing effectively smooths and enriches each cell’s latent representation with information from its spatial vicinity, while automatically learning the relevance of each neighbor.

[0455] The output of the graph attention layer for cell z is a spatial feature vector hLof dimension nspatial(a hyperparameter). Applicant interpret this as parameters of a latent spatial factor for each cell. In particular, Applicant assume that for each cell z there is an unobserved spatial latent vectorE Rnspanai that follows a Gaussian distribution conditioned on ht. Applicant use two parallel linear layers to produce a meanE Rnspattaianc[ log-variancelogtTt E Rnspattai fromthe GAT output[4]. This yields a variational posterior for the spatial latent:q(st | z£,lV(i)) = N ^sgg^, diag(a^s^,[0456J where g^ = kl^ / i£and = exp(Wahi) + e (with a small e = 10-4added for numerical stability). In practice, stis sampled by the reparameterization trick as st= g^ + &i1O G with£~ A(0, / ). Intuitively, this spatial encoder looks at a cell’s expression-derived latent vector and its neighbors’ latents, and then proposes a distribution for a spatial effect Si for that cell. If no neighbors are present or the graph is empty, the encoder simply yields a default s ~ 0 with unit variance, meaning no spatial adjustment. By using attention, the model can learn to ignore irrelevant neighbors or focus on certain directions (e.g., along a tissue boundary) as needed.[0457| Variational Inference and Objective]0458| Applicant’s model builds upon the single-cell variational inference (scVI) framework, using a variational autoencoder (VAE) to model gene expression data. Each cell z has an underlying global latent variable z£E Rniatent that captures its transcriptional state (accountingAtty. Dkt. No.: 114198-3660for technical effects like batch, if specified). Applicant assumes a standard normal prior p(z ) = A(0, / ). An encoder neural network (a multilayer perceptron with niayershidden layers and nhiddennodes each) takes the observed gene expression vector xt(and optionally other covariates) and outputs the approximate posterior distribution q zt| x^ = N (zi; p^z\ diag^e?^^. This is analogous to standard sc VI: the encoder learns to compress the high-dimensional gene expression of a cell into a lower-dimensional latent representation. In Applicant’s Spatial VI model, if spatial information is used, the inference proceeds in two stages: first inferfrom the cell’s own data, then infer from ztand neighbors as described above. Thus, the overall approximate posterior factorizes as q(zi \Xi) q(si \ zi, N(i) Applicant emphasize thatis only inferred if use_spatial=True and it encodes residual variation structured by spatial locality.[0459| On the generative side, the model reconstructs the observed gene expression from the latent variables. Applicant model raw gene counts (such as UMI counts) with appropriate count distributions commonly used in single-cell analysis[8]. By default Applicant uses a negative binomial (NB) likelihood (or zero-inflated NB if specified) for each gene’s expression, which accounts for over-dispersion and the many zero counts. Specifically, for each gene j in cell z, Applicant’s decoder outputs a parameter pij (the mean of the NB) and gene-specific overdispersion 0j. The mean is formulated to capture library size and latent effects: ptj = it ■ fj(zi, Si), where it is the cell’s total count (library size) or an estimated size factor, and fj(zi, Si) is the decoder function’s output for gene j. In practice (•) is a neural decoder that takes z(- (and potentially S[ or other covariates) and produces gene-wise rates; when no spatial latent is used, f only depends on zL. Each gene count is then modeled as x^ ~ NB(.ij, d ). The negative log-likelihood (reconstruction loss) for gene j in cell z is:-logp(xtj | zt.st)X; j + 0 J. - 1+ XijlogOj - Xijlog^Bj + ptj') + 0jlog(0j +xij — 0jlog0j,Atty. Dkt. No.: 114198-3660which is the standard NB formulation (for zero-inflated NB, an additional zero-inflation probability is learned per gene). The decoder is trained to maximize the likelihood of the observed counts given the latent variables.The overall training objective is the evidence lower bound (ELBO) from variational inference. For a single cell z, the ELBO is:Li = Eq^t,st\xt)[logp(Xi I zt.st)] - KL,(q(zi | x£) || p(z£)) - KL(q(st I z£) || p(s£)), where p(s£) = A(0, / ) is the prior for the spatial latent. In practice, Applicant introduce a weighting factor a> for the spatial KL term to moderate its influence: the objective becomes L£= Eq[logp{xt| z£,s£)] - KL(q(z£|x£) || p(z£)) - m KL(q(s£|z£) || p(s£)). The hyperparameter m corresponds to spatial kl weight in Applicant’s implementation (default a> = 0.01 ), meaning Applicant initially down-weight the spatial latent regularization to encourage the model to learn meaningful spatial factors without dominating the gene expression term. This weighted ELBO approach is akin to a P-VAE strategy where = a> for the spatial part, and it can be tuned. The KL divergences have closed forms since q and p are Gaussians. For example, KL(q(s£) || p(s£)) = ^Y^=itial(k?)2 / i + ^<s’ / i -l°g fkS)~ 1 During training, Applicant average L£over all cells (or sum, equivalently) and maximize it. Applicant use stochastic optimization (minibatches of cells) with Adam, as in standard scVI. Gradients flow through both the encoder for z and the GAT -based encoder for s. The inclusion of the spatial graph and latent does not drastically increase computational complexity; the KNN graph can be precomputed efficiently, and the GAT convolution is applied on latent dimensions (typically much smaller than gene counts). Graph operations are accelerated via PyTorch Geometric.

[0460] After training, the model can output various representations: the global latent embeddings Z = {zi} for all cells (useful for visualization or clustering), as well as spatial latent embeddings S’ = {s£} which capture spatially correlated variation. Applicant can also impute or denoise gene expression, generate samples, or perform differential expression by leveraging the learned generative model, similar to sc VI’ s downstream applications. The learned attention weights in the GAT provide interpretability on which neighbors most influence each cell’s state, potentially highlighting spatial patterns or boundaries.Atty. Dkt. No.: 114198-3660[04611 Multi-Modal Data Integration[04621 A key feature of Applicant’s Spateo-VI model is the ability to jointly model multiple data modalities in a shared latent space. In many experiments, one might have spatial transcriptomics data (e.g., MERFISH with spatial coordinates) as well as single-cell RNA-seq data from dissociated cells of the same tissue, or protein expression data or epigenetics histone modification. Applicant’s model extends the single-modality VAE to handle two gene expression matrices (spatial and non-spatial) and an optional protein expression / Histone modification matrix. Applicant maintains a single latent variable z(- per cell that captures the cell’s state in a way that is common to all modalities. Cells from the spatial dataset and the non-spatial dataset are all embedded in this same latent space, allowing direct integration and comparison. Applicant instantiate separate encoders for each modality’s observed data, but they map to the same type of latent z distribution (with possibly modality-specific encoder networks). Likewise, for decoding, Applicant have modality-specific decoders: one decodes Zj (and Si for spatial cells) to gene counts in the spatial transcriptomics, another decodes z;- for cells j in the non-spatial dataset to gene counts, and optionally a decoder for proteins (which typically uses a negative binomial or Poisson with a mixture to model background, following the strategy of Total VI). The loss function in the joint model is a sum of the reconstruction losses for each modality plus the KL terms. Applicant allow weighting of each modality’s contribution via user-defined modality weights. For example, one can set higher weight to the spatial transcriptomic loss if those counts are more reliable, or balance them equally. The default is equal weighting (1.0 each for spatial and non-spatial)

[0014] . The modality weights effectively appear as multipliers on the log-likelihood terms for each dataset in the ELBO. By tuning these, users can emphasize one modality over another as needed (analogous to multitask learning weighting).

[0463] When protein or histone data are included, Applicant’s model extends the generative part to include protein observations. In such cases, the decoder for proteins incorporates a background noise model per protein / hi stone, similar to how TotalVI models protein counts as a mixture of a specific (foreground) signal and an unspecific background component. In Applicant’s implementation, placeholders exist for learning protein-specific background rate and scale parameters and for computing a “foreground probability” for each protein in each cell. These parameters would be learned by maximizing the protein likelihood (often modeledAtty. Dkt. No.: 114198-3660with a zero-inflated negative binomial for foreground and a separate term for background). Although the current code provides functions for retrieving protein background parameters and normalized protein expression, those follow the approach of Gayoso et al. (2021) in concept. For instance, the model can output an estimated probability that a given protein count is “foreground” signal rather than background noise, and adjust the protein expression accordingly.

[0464] Overall, by integrating multiple modalities, Spateo-VI can learn a unified latent representation that synthesizes information from spatial gene expression, conventional scRNA-seq, and histone markers. This shared latent space facilitates dataset integration (e.g., aligning a dissociated single-cell atlas with spatial data), as well as multimodal downstream analyses. For example, one can cluster cells based on the latent space that reflects both RNA and histone, or transfer labels between the non-spatial and spatial datasets (since they occupy the same latent manifold). Applicant’s approach is analogous to Total Vi’s joint modeling of RNA and proteins, but extended to spatial transcriptomics and epigenetics: Applicant represent all modalities probabilistically and use a variational inference scheme to decompose observed variation into a common biological latent variable z, modality-specific technical effects (e.g., histone background, batch effects), and in the case of spatial data, a spatial effect s. This provides a cohesive solution for multi-modal single-cell data integration, allowing tasks such as dimensionality reduction, data integration across modalities, and differential expression to be performed in a unified framework. By preserving uncertainty through the variational distributions, the model can also propagate measurement uncertainty from all sources when making biological inferences.[04651 Spatial cell-cell interaction analysis

[0466] Furthermore, to investigate spatial cell interactions, Applicant extended COMMOT by developing the COMMOT-COT algorithm, enabling its application in 3D spaces and accelerating the optimal transport process via GPU acceleration. Specifically, since 3D space introduces an additional z-axis compared to 2D space, neighborhoods along the z-axis exhibit significant differences from those along the x / y axes. Therefore, Applicant first applied Spateo to minimize spatial distances between different slices. Subsequently, Applicant employed an optimized optimal transport algorithm to bring adjacent cells on different slices closer together, transforming uniform slices into 3D dynamic slices. Applicant then utilized the CellChatAtty. Dkt. No.: 114198-3660database to compute the spatial communication strength of each signal on every cell. Furthermore, Applicant developed the CellChatViz function to visualize the underlying signal processes.

[0467] MERFISHEYES Visualization Implementation

[0468] The MERFISHEYES website delivers fast and reliable visualization of millions of cells within a web browser. The rendering is optimized through WebGL (Khronos Group. (2009- 2025). WebGL [Computer graphics API], https: / / www.khronos.org / webgl / ) in combination with the Three. is framework (Dirksen, R., Cabello, R., & Contributors. (2010-2025). Three.js [Computer software], GitHub. https: / / threejs.org / ) enabling @responsive hovering interactions, fast performance, and utilization across desktop and mobile devices. The backend is implemented using the Rust Rocket (Rust Project Developers. (2010-2025). The Rust programming language [Programming language], Mozilla Research & Rust Foundation. https: / / www.rust-lang.org / && Rocket Contributors. (2016-2025). Rocket: A web framework for Rust [Computer software], GitHub. https: / / rocket.rs / ) web api with file compression to reduce latency and server load.[0469J Chromatin Tracing Analysis[0470[ A. Localization of Fluorescent Spots[04711 To calculate fluorescent spot localizations for chromatin tracing data, Applicant followed the following computational steps:1. Applicant computed a point spread function (PSF) for Applicant’s microscope and a median image across all fields of view for each color channel based on the first round of imaging to be used for homogenizing the illumination across the field of view (called flat-field correction).2. To identify fluorescent spots, the images were flat-field corrected, deconvoluted with the custom PSF, and then local maxima were computed on the resulting images. A flat-field correction was done for each color channel separately.

[0472] B. Image registration and selection of chromatin traces

[0473] Imaging registration was performed by aligning the DAPI channel of each image from the same field of view across imaging rounds. First, the local maxima and local minima of theAtty. Dkt. No.: 114198-3660flatfield corrected and deconvolved DAPI signal were calculated. Next, a rigid translation was calculated using a fast Fourier transform to best align the local maxima / minima between imaging rounds. FXYD Coordinates: chrl9:35003250-35533250(hg38); FXYD Coordinates: chrl9:37547865-38078198 (hsl)

[0474] Nuclear segmentation was performed on the DAPI signal of the first round of imaging using the Cellpose69algorithm using the “nuclei” neural network model. Following image registration, chromatin traces were computed from the drift-corrected local maxima of each imaged locus as previously described27.

[0475] Protein Density Quantification

[0476] The antibody images were flat field corrected, deconvolved and then registered to the chromatin traces using the DAPI signal as described before. For each chromatin trace, the fluorescent signal of each antibody was sampled at each genomic locus 3D location in each cell.

[0477] Code availability

[0478] Custom code used for analyzing chromatin tracing and MERFISH datasets in this study are available here: https: / / github.com / cfgOO / MERFISH Chromatin Tracing 2024

[0479] Custom code used for 3D reconstruction, visualization, analysis, and imputation in this study are available here: https: / / ithub.com / ari stoteleo / Human Heart MERFISH Analy si s 2024

[0480] Data availability

[0481] Zenodo - single scanpy object - traces, RNA, proteins and volume

[0482] Raw imaging data will be provided upon request. Processed data is available on Zenodo as a scanpy,h5ad file. The main “. X” matrix of the object contains log-normalized counts. The full contents of the scanpy object is described below. For brevity, standard contents added by scanpy, e.g., connectivities and distances added by the sc. pp. neighbors, are not listed.• obso volm - Total pixel volume of the cell based on DAPI segmentation o x um abs, y um abs - Global x and y coordinates of the cell in microns o zc, xc, yc - Pixel coordinates of the cell center relative to the field of viewAtty. Dkt. No.: 114198-3660o leideno regiono batcho scToSpatialo leiden loado confidenceo LIo Ll leideno dpt_pseudotimeo inverted_dpto final annoo final_anno_v2o final_anno_v3o hpcro hiluso fimbriao hpcRGo hilusRGo fimbriaRGo ventricularo ventricularRGo Refined volume - Re-calculated cell volume based on Nup98 antibodies • varo mean - Average expression of the gene across cellso std - Standard deviation of the gene expression across cells• unso X h score shape - Original shape of X h score in obsmo antibody shape - Original shape of each antibody matrix in obsm.• obsmo X fov - The field of view (FOV) identifier each cell was imaged in o X_raw - Raw count matrixAtty. Dkt. No.: 114198-3660o X spatial - The spatial coordinates of the cellso blank - The count of each blank barcode per cello X h score - A csr sparse matrix containing chromatin trace results. The matrix should be reshaped to 50374x4x354x5 representing the number of cells, maximum number of homologs, number of chromatin regions, and the z, x, y coordinates followed by the brightness and score of the fluorescent spot. Missing data, i.e. containing fewer than 4 homologs or missing regions, are filled with Os.o H3K9me3, H3K27Ac, H3K27me3, H3K4me3, PolIISer2P, SC35, Lamin A, Nup98 - Antibody signals localized at each chromatin region. Stored as a csr sparse matrix and can be reshaped to 50374x4x354 similarly to X_h_score

[0483] Table 7: List of parts to assemble high-throughput microscope-microfluidics system described hereinCatalogueVendor Manufacturer number Part name Quantity Stage Option Piezo Top PlateASI ASI PZ-2500 INV 1 ASI ASI S551-2201B Stage, Flat-top RAMM 1 ASI ASI TG8-BASIC Controller Tiger 8 System 1 MIM3- MIM3 w / automated 2 positionASI ASI O2SM25 objective slide 1 ASI ASI RAMM-BASIC RAMM basic 1 ASI ASI TGADEPT Control Tiger Card Piezo 1 MIM-CUBE- CUBE III with D-Cube andASI ASI III-KD C60-RINGs 2 MIM-CUBE- ASI ASI III-K MIM Cube-Ill Kit 1 Tube Lens Nikon 200mm,ASI ASI C60-TUBE-B 30mm diameter aperture 1 ASI ASI C60-TUBE-300 Tube Lens 300mm achromatic 1 ASI ASI C60-TUBE-250 Tube Lens 250mm achromatic 1 ASI ASI TGPLC Control Tiger Card TTL 2-slot 1 RAMM- ASI ASI STILES RAMM STILTS 1 C60-30MM- Mirror Right Angle 30 mmASI ASI RA-MIRROR-S Silver F-Dovetail 1 C60-5060-C- ASI ASI MOUNT C-Mount Adapter Male 5060 1ASI ASI C60-F-MOUNT F-mount C60 Nikon Tube B 3Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity C60-30CRM- ASI ASI 30LM Cage C60 ring 30mm 2 C60-RING- Ring, 38mm Male to SMIASI ASI SM1F Female 2 C60-RING- ASI ASI SM1M C60 Coupling Ring 38mm Male 2 imaging Chamber for 64mm x50mm Slides, includes SBS4,563.25 4,563.25 base plate (Aluminum) with lock andBioptechs Bioptechs 0903-6450-NH perfusion insert (stainless). 1 0903-6450- 64mm x 50mm Microaqueduct Bioptechs Bioptechs 1319-NH slide (5pk) 1 Microaqueduct Slide, TPG (non Bioptechs Bioptechs 130119-5NC coated) T Grooves, Holes (5pk) 1 0903-6450- 64 mm x 50mm Micro Slide Bioptechs Bioptechs 131907 Gasket (5 pk) 1 0903-6450- 64 mm x 50mm Inner Perfusion Bioptechs Bioptechs 091607 Gasket (5 pk) 1 No Heat FCS2 Chamberincludes assembled base with 03060319-2-3- 30mm aperature and standard Bioptecs Bioptecs NH-30 white top 1 FCS2 Zeiss K & Ludl StageBioptecs Bioptecs 060319-2-2611 Adapter 1 Silicone Gasket cut with die1907-449673- #449673-A 0.75mm Thick 5Bioptecs Bioptecs A-750 pack 11907-P47132- Silicone Gasket cut with die # Bioptecs Bioptecs 1000 P47132 1.0mm Thick 5 pack 11907-P47132- Silicone Gasket cut with die # Bioptecs Bioptecs 750 P471320.75mm Thick 5 pack 1 Bioptecs Bioptecs Engineering TBD 1 DAQ Device MultifunctionalDigi-Key Digi-Key 781048-01 I / O PCIE 1 SCB-68A Shielded ConnectorDigi-Key Digi-Key 782536-01 Block 1 SHC68-68-EPM ShieldedDigi-Key Digi-Key 192061-01 Cable, 68- 1 WIRE KIT 24AWG HOOK-UP3050 ROHS3 COMP REACH DigiKey DigiKey HUKIT-40-ND UNAFFECTED 1 AC / DC DESKTOP ADAPTERDigiKey DigiKey EPS513-ND 24V 221W ROHS COMP 1Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity HEATSHRINK KIT Q2F 141 PCCOLOR ROHS3 COMPDigiKey DigiKey Q2F15-KIT-ND REACH UNAFFECTED 1 ARDUINO UNO SMD R3 ATMEGA328 ROHS NADigiKey DigiKey 1050- 1041 -ND REACH UNAFFECTED 1 TEST LEAD BNC TO WIRE LEADS 20” ROHS NADigiKey DigiKey 501-2043-ND REACH UNAFFECTED 10 CONN ADAPT PLUG TO JACK BNC ROHS3 COMPDigiKey DigiKey 501-1134-ND REACH UNAFFECTED 2 CONN RCPT FMALE DIN 8P SOLDER ROHS3 COMPDigiKey DigiKey SC2007-ND REACH UNAFFECTED 1 ARDUINO STACKABLE HEADER KIT - R ROHS3DigiKey DigiKey 1568-1413-ND COMP 123- CONN RCPT HSG 13P0S 2.54 0050579413- MM ROHS3 COMP REACH DigiKey DigiKey ND UNAFFECTED 10 CONN RCPT HSG 2P0S 2.54MM ROHS3 COMP REACH DigiKey DigiKey WM2900-ND UNAFFECTED 10 CONN SOCKET 24-30AWGCRIMP GOLD ROHS3 COMP DigiKey DigiKey WM2513-ND REACH UNAFFECTED 500RES 4.7K OHM 5% 1 / 4W CF14JT4K70C AXIAL ROHS3 COMPDigiKey DigiKey T-ND REACH 10 PSD6, SYRINGE PUMP, STD, Hamilton Hamilton 63133-01 ROHS 1 Hamilton Hamilton 64316-04 PSD-4 OEM STARTER KIT 1 SYRINGE, TLL,5ML, LONG Hamilton Hamilton 8300-45 LIFE, UHMW 3 KF720 NDL 6 / PK (20 / * / *)Hamilton Hamilton 7750-11 Configuration: DEFAULT 6 Y VALVE ASSY, CTFE, HV, Hamilton Hamilton 57252-01 SD4 / PS D8 1 Hamilton Hamilton 99893-01 VALVE, HV 3-5-PCTFE 1 Celesta 5-Channel Light Engine5 independently controlled laser light sources VCGRnIR withLumencor Lumencor 90-10791 output despeckler. ~50mw of 1Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity Violet, -800 mW output forCGRand -2W for NIR at distal endof 0.4mm square multimodefiber.Pentaband dichroic optimizedfor reflection of405 / 477 / 545 / 639 / 748 nm laser outputs of CELESTA andZIVA light engines.Unmounted, 25 mm x 36 mm x Lumencor Lumencor 10-10858 1 mm 1 Lumencor Lumencor 10-10751 Pentaband emission filter 1 Directs TTL signals fromseparate BNC (M) connectorsto SPECTRA or CELESTAlight engine for individualLumencor Lumencor 29-10156 channel triggering. 2 m length. 1 VIS IsoStation, 30 in. x 48 in., 4 Newport in.Corporati Newport VIS3048-PG4- thick PG Breadboard, 1-325on Corporation 325A Isolators 1 Newport Overhead Shelf, VisionCorporati Newport Isostation, 48 in. Long,on Corporation VIS-ATS-48 Electrical Sockets 1 NewportCorporati Newport Keyboard Monitor Mountingon Corporation VIS-LS-KBM Kit, No Hip Guard Required 1 NewportCorporati Newport Air Compressor, Low Noise,on Corporation ACWS 110V 1 NewportCorporati Newport Computer Shelf, Visionon Corporation VIS-SCS-30 Isostation, 30 in. 1 CFI APO LWD 40X WI 1.15NA LAMBDA S CFI60Apochromat Lambda S LWD40x Water Immersion Objective Lens, N. A. 1.15, W. D. 0.59- 0.61mm, F. O. V. 22mm, DIC,Nikon Nikon MRD77410 Correction Collar 0.15-0.19mm 1 CFI60 PLAN APOCHROMAT LAMBDA D 60X Oil CFI60Nikon Nikon MRD71670 Plan Apochromat Lambda D 1Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity 60x Oil Immersion ObjectiveLens, N. A. 1.42, W. D. 0.15mm, F. O. V. 25mm, DIC, SpringLoadedCFI60 PLAN APOCHROMAT LAMBDA D 10X CFI60 Plan Apochromat Lambda D lOx Objective Lens, N. A. 0.45,W. D. 4.0mm, F. O. V. 25mm,Nikon Nikon MRD70170 DIC 1 PLAN APOCHROMAT MRD70040 LAMBDA D 4X 1 CFI Plan Apo Lambda S 25XC MRD73250 Sil 130CC SILICONEIMMERSION OIL 30ccSilicone Immersion OilNikon Nikon MXA22179 (ne=1.406) 130CC NON-FLUORESCING IMMERSION OIL F 30cc Non- Nikon Nikon MXA22168 Fluorescing Immersion Oil F 1 Imaging Chamber for 64mm x50mm Slides, includes SBSbase plate (Aluminum) withlock and perfusion insertOther Bioptecs 0903-6450-NH (stainless). 1 Microaqueduct Slides noOther Bioptecs 130119-5NC coating (5 / pk) 1 0903-6450- Microaqueduct Slides, 64mm xOther Bioptecs 1319-NH 50mm 5 0903-6450- Microaqueduct Slide Gasket,Other Bioptecs 131907 64mm x 50mm 1 0903-6450- Inner Perfusion Gasket, 64mmOther Bioptecs 091607 x 50mm 1 Tube Tfzl Nat 1 / 16 OD x0.04Other IDEX 1517XL IDx 100ft 1 BleedZone Female Luer Lock Connector - 1 / 16” Hose BarbFittings PP PolypropyleneOther BleedZone Hose, 25x Luer Lock Adapter 10 Adapter, Luer; LuerTight;Quick connect; 1 / 4-28 toOther IDEX P-658 Female; with quote 202433341 10Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity Luer Adapter Male Luer x 1 / 4- 28 Female Tefzel Red withOther IDEX P-675 quote 202433341 10 uxcell Silicone Tubingl / 16”(1.5mm) ID x l / 8”(3mm)OD 16ft(5m) Silicone RubberTube Air Hose Water Pipe forOther Uxcell Pump Transfer Clear 1 BleedZone - PP PolypropyleneHose Barb Adapter, 5x HoseBarb Fittings with Luer Lock Connector, 1 / 16” Hose BarbOther BleedZone Size 10 Flngls Rd Delrin 1 / 16in Moq 25 Other IDEX XP-202 with quote 202433341 25 Other Nikon 25x objective?Kinetix 10 cameraKinetix 10MP PCI ExpressCamera USB 3.2 Gen 2 BackSide Illuminated CMOS - 95% QE, 6.5pm x 6.5pm pixelsize, 3200 x 3200 array (29mmField of View), 498Teledyne frames per second, PCI-Express Acton Teledyne O1_KINETIX_ Card Kit Included. 3 YearOptics Acton Optics 10MP PC IE Warranty. 1 PM100D Power Energy Meter Thorlabs Thorlabs PM100D w. Graphics Display 1 Benchtop LD CurrentThorlabs Thorlabs LDC205C Controller, ±500mA HV 1 LD / TEC Mount for Thorlabs Thorlabs Thorlabs LDM9LP FiberPigtailed Laser Diodes 125 x 36mm SP Dichroic Mirror, Thorlabs Thorlabs DMSP900R 50% T / R at 900nm 1 Slim Silicon Power Head, 400- Thorlabs Thorlabs S130C 1100nm, 5mW / 500mW 1 LPS, 980nm, 15mW, 5EPinThorlabs Thorlabs LP980-SF15 FC / PC 11 / 2” Cage Translation StageThorlabs Thorlabs CT1A TTN255233 1 CMOS Scientific 1.6MP MonoThorlabs Thorlabs CS165MU Zelux Camera, USB3.0 1Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity 30mm Cage Mounted Non-PolBS Cube 50:50 0.7-1.1um, 8-32 Thorlabs Thorlabs CCM1-BS014 Tap 21 Inch Elliptical KinematicComer Block TTN051747, 1Inch Elliptical KinematicThorlabs Thorlabs KCB1EC Comer Block 4 Breadboard 12x18x1 / 26276- Thorlabs Thorlabs MB1218 001, Breadboard 12x18x1 / 2 1 Breadboard 12x18x1 / 26276- 001, Breadboard 12x18x1 / 260 Thorlabs Thorlabs MB1218 X 34 X 5 CM @ 5 KG 2 Thorlabs Thorlabs SM1ZA Z-Axis Translator 2 USB Laser Module 520 nm 1 Thorlabs Thorlabs PL201 mW 1 Laser Safety Glasses, LightOrange Lenses, 48% VisibleThorlabs Thorlabs LG3B Light 2 Thorlabs Thorlabs RMS4X 4X Microscope Objective 1 01” Mounted AC254-100-AB, AC254-100- f=100 mm, SMI -ThreadedThorlabs Thorlabs AB-ML Mount 1 01” Mounted AC254-150-AB, AC254-150- f=150 mm, SMI -ThreadedThorlabs Thorlabs AB-ML Mount 1 01” Mounted AC254-200-AB, AC254-200- f=200 mm, SMI -ThreadedThorlabs Thorlabs AB-ML Mount 1 Premium Longpass Filter, Cut- Thorlabs Thorlabs FELH0950 On Wavelength: 950nm 11 / 4-20 bolt kit over 1000 pieces TTN022117, 1 / 4-20 bolt kitThorlabs Thorlabs HW-KIT2 over 1000 pieces 1 #8-32 Set Screw KitThorlabs Thorlabs HW-KIT3 TTN022119 1 AC254-075-B- Mounted AC254-075-B -01.0”, Thorlabs Thorlabs ML efl=75mm 1 AC254-400-B- Mounted AC254-400-B -01.0”, Thorlabs Thorlabs ML efl=400mm 125.4mm Broadband Dielectric Thorlabs Thorlabs BBE1-E02 Elliptical Mirror, -E02 425.4mm Elliptical Mirror, -E03Thorlabs Thorlabs BBE1-E03 Coated 4Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity f=50 mm, 01” AchromaticAC254-050-A- Doublet, SMI-Threaded Mount, Thorlabs Thorlabs ML ARC: 400-700 nm 1 f=75 mm, 01” AchromaticAC254-075-A- Doublet, SMI-Threaded Mount, Thorlabs Thorlabs ML ARC: 400-700 nm 1 00.5” Mounted AC127-019- AC127-019- AB, f=19 mm, SM2-Threaded Thorlabs Thorlabs AB-ML Mount 1 00.5” Mounted AC127-030- AC127-030- AB, f=30 mm, SM2-Threaded Thorlabs Thorlabs AB-ML Mount 11 / 4” -20 Set Screw KitThorlabs Thorlabs HW-KIT4 TTN022121 1 Mounted AC254-125-A-01. O”, AC254-125-A- efl=125mm f=125.0 mm, 01” Thorlabs Thorlabs ML Achromatic Doublet 1 AC254-300-A- Mounted AC254-300-A -01.0”, Thorlabs Thorlabs ML efl=300mm 1 SMI Threaded MetricKinematic 30mm CageCompatible Mount forTTN258140, SMI ThreadedMetric Kinematic 30mm Cage Compatible Mount for Ø1 inThorlabs Thorlabs KC1T / M Optics 2 Thorlabs Thorlabs P14 1.5” dia Mounting Post x 14” 4 0=9.24mm, f=7.50mm,NA=0.30, Mounted H-LAK54 Thorlabs Thorlabs A375TM-A Asphere,-A 1 Kinematic 30mm CageCompatible Mount for Ø1inThorlabs Thorlabs KC1L Optics 2 Thorlabs Thorlabs TC2 20 pcs Wrench Set with Stand 115 Piece Metric Wrench KitThorlabs Thorlabs TC3 / M with Stand 1 Mounted Ø6.50mm, D-ZK3Asphere, f= 18.40mm,Thorlabs Thorlabs C280TMD-B NA=0.15, -B Coa 1 Ø=12.18mm, f=8.00mm,NA=0. 50, Mounted D-LAK6 Thorlabs Thorlabs A240TM-A Asphere, -A 1 Viewing Card for 400-640nm &Thorlabs Thorlabs VRC2 800-1700nm 1Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity 50 MC-5 Lens Tissue in aThorlabs Thorlabs MC-50E Closeable Box 260mm Quick Release CageThorlabs Thorlabs QRC2A Mount 2 MOUNTED 25.0 mmABSORPTIVE ND FILTER, BBAR 650-1050 nm Thorlabs Thorlabs NE10A-B OD: 1.0 1 MOUNTED 25.0 mmABSORPTIVE ND FILTER, BBAR 650-1050 nm Thorlabs Thorlabs NE20A-B OD: 2.0 1 OIMMOIL-F30CC; IMMER.Thorlabs Thorlabs MOIL-30 OIL 30CC, LOW AUTOFL 1 Kit: 4 Wash Bottles and 3Thorlabs Thorlabs B2939 Plastic Dropper Bottles 1 Ø25.4mm Mirror, Broadband - Thorlabs Thorlabs BB1-E03 E03 Coated 3 SMI Slotted Lens Tube w / Dust Cover TTN004878, SMISlotted Lens Tube w / DustThorlabs Thorlabs SM1L30C Cover 1 SM1 Series Quick ReleaseThorlabs Thorlabs QRC1A Cage Mount 3 SMI Slotted lens tubes, 2inThread depth TTN094267, SMI Thorlabs Thorlabs SM1L20C Slotted lens tubes, 2in Thread 14” X 6” DD Mini SeriesBreadboard, 1 / 4-20 and 8-32Tap TTN127332, 4” X 6” DDMini Series Breadboard, 1 / 4 -20 Thorlabs Thorlabs MSB46 and 8-32 Tap 1 Iris Diaphragm used in SMIseries TTN137232, 0077, Iris Thorlabs Thorlabs SM1D12 Diaphragm used in SMI series 4 #8-32 bolt kit over 750 pieces TTN022116, #8-32 bolt kit over Thorlabs Thorlabs HW-KIT1 750 pieces 1 Small Universal Clamping Forkw / 1 / 4-20 captive screw, 5 pack TTN043771, Small Universal Clamping Fork w / 1 / 4- 20Thorlabs Thorlabs CF125C-P5 captive screw, 5 pack 2Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity Black posterboardThorlabs Thorlabs TB5 20”x30”xl / 16” 2 Thorlabs Thorlabs RA90-P5 Right Angle Post Clamp, 5 pack 2 Thorlabs Thorlabs SM1E60 SMI Series Extension Tube, 6” 260.0mm CAGE PLATE, 8.0mm Thorlabs Thorlabs LCP8S THICK 330mm to 60mm Cage Adapter TTN247582, 30mm to 60mmThorlabs Thorlabs LCP33 Cage Adapter 5 Microscope Objective Adapter Thorlabs Thorlabs E09RMS Ext. Tube 1 Mounted Frosted GlassDG10-1500- Alignment Disk 01” w / 01mm Thorlabs Thorlabs Hl-MD Hole 31 / 2” Dia. x 6” Length: Pack of Thorlabs Thorlabs TR6-P5 5 Post 1 Matte Black AluminumThorlabs Thorlabs BKF12 Foil,.002 x 12” x 50’ 1 XE25 Series “T” Nut, pack of Thorlabs Thorlabs XE25T4 10 1 Extension Rod 6 inch, Pack of 4 TTN010250, Extension Rod 6 Thorlabs Thorlabs ER6-P4 inch, Pack of 4 130 mm Cage System Alignment Thorlabs Thorlabs VRC4CPT Plate with IR Disk 2 Thorlabs Thorlabs SM2RR-P5 SM2 Retaining Ring, 5 pack 1 Thorlabs Thorlabs SM1FC Fiber Adapters 1 Thorlabs Thorlabs SM1FC Fiber Adapters 1 Fiber Collimation Adapter forSMI Series MountsThorlabs Thorlabs AD11F TTN100768, 2314 1 SMA Fiber Adapter Plate with External SMI (1.035”- 40)Thorlabs Thorlabs SM1SMA Thread 1 Stackable lens tube for 2”opticusable depth 1”AC127- Thorlabs Thorlabs SM2L10 019-AB-ML 3 Thorlabs Thorlabs P2 1.5” dia. Mounting Post x 2” 4 Thorlabs Thorlabs XE25A90 Angle Bracket 4 Pedestal Style Post Holder, 6 Thorlabs Thorlabs PH6E inch 5 Thorlabs Thorlabs SPW602 Spanner Wrench For SM1RR 11 / 2” Dia. x 4” Length: Pack ofThorlabs Thorlabs TR4-P5 5 Post 1Atty. Dkt. No.: 114198-3660CatalogueVendor Manufacturer number Part name Quantity Extension Rod 4 inch, Pack of 4 TTN010249, Extension Rod 4 Thorlabs Thorlabs ER4-P4 inch, Pack of 4 2 SMI to SM2 Series ThreadThorlabs Thorlabs SM2A6 Adapter 1 Extension Rod 18 inch 19231- Thorlabs Thorlabs ER18 001 8 Thorlabs Th...

Claims

Atty. Dkt. No.: 114198-3660WHAT IS CLAIMED IS:

1. A method to provide imaged codewords corresponding to distinct target molecules in a spatial organization in a sample comprising:a) contacting a sample comprising a plurality of target molecules and one or more primary nucleic acid probe pools, wherein the primary nucleic acid probes in each of the pools comprises an acrydite group and is bound to a biogel film covering a surface of the sample, andwherein each pool of primary nucleic acid probes is selected to specifically hybridize to a distinct target molecule in the sample, andwherein each primary nucleic acid probe comprises a target sequence and one or more read sequences, andwherein each primary nucleic acid probe produces a bound primary nucleic acid probe upon contact with its target molecule, andwherein each pool of primary nucleic acid probes encodes an N-bit codeword with a Hamming weight of at least 2 that was assigned to each distinct target molecule,wherein each assigned N-bit codeword is a valid codeword with a Hamming distance equal to or greater than 2 between valid codewords and wherein each of the one or more read sequences correspond to a bit value of 1 for the codewords assigned to the target molecules,b) contacting the one or more plurality of primary nucleic acid probes of the pools with a plurality of fluorescently labeled readout probes that can specifically bind to the one or more read sequences of the primary nucleic acid probes to generate detectable signal at the target molecule location,c) imaging the readout probes bound to the primary nucleic acid probes; and d) repeating steps b) and c) in or more sequential binding and imaging rounds until all bit positions of the N-bit codeword have been imaged providing imaged codewords corresponding to each distinct molecule in a spatial organization.

2. The method of claim 1, wherein the target molecules are selected from nucleic acids or polypeptides.

3. The method of claim 2, wherein the nucleic acids are DNA or RNA molecules.Atty. Dkt. No.: 114198-36604. The method of any one of claims 1-3, wherein steps c) and d) are repeated over a period of at least 1 month, or at least 2 months, or at least three months.

5. The method of any one of claim 1-4, wherein the acrydite group is derived from a compound of formula (HO)2P(O)(CH2)xNHC(O)C(CH2)CH3, wherein x is 1-10.

6. The method of claim 5, wherein x is selected from 2-9, 3-8, 6-7, or 7.

7. The method of any one of claim 1-7, wherein the acrydite group is attached to a 5’ end of the primary nucleic acid probes.

8. The method of any one of claims 1-7, wherein the assigned N-bit codewords have a Hamming distance equal to or greater than 4 between each of the valid codewords.

9. The method of any one of claims 1-8, wherein each N-bit codeword has a Hamming weight of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10.

10. The method of any one of claims 1-9, wherein the imaged codeword is matched to a valid codeword assigned to a distinct RNA species or a distinct DNA species.

11. The method of any one of claims 1-10, wherein the Hamming distance is 4 or greater and the imaged codewords are matched to valid codewords or discarded.

12. The method of any one of claims 1-11, wherein the N-bit codeword has a Hamming weight of 4 and the Hamming distance is 4 or greater and the imaged codewords are matched to valid codewords or discarded.

13. The method of any one of claims 1-12, wherein the target molecule is an RNA transcript.

14. The method of any one of claims 1-13, wherein the spatial organization of the transcriptome is determined from a single cell.

15. The method of any one of claims 1, 2 or 4-12, wherein the spatial organization of the proteome is determined from a single cell.Atty. Dkt. No.: 114198-366016. The method of any one of claims 1-15, wherein each primary nucleic acid probe pool comprises at least 10 different primary nucleic acid probes, wherein each of the at least 10 different primary nucleic acid probe hybridizes to a distinct target molecule.

17. The method of any one of claims 1-16, wherein each primary nucleic acid probe pool comprises at least four distinct read sequences of the one or more read sequences.

18. The method of any one of claims 1-17, wherein each primary nucleic acid probe pool comprises four distinct read sequences of the one or more read sequences.

19. The method of any one of claims 1-18, wherein each primary nucleic acid probe pool comprises at least eight distinct read sequences of the one or more read sequences.

20. The method of any one of claims 1-19, wherein the target sequence has an average length of between 10 and 200 nucleotides and is configured to hybridize to the distinct target molecule.

21. The method of any one of claims 1-20, wherein each primary nucleic acid probe pool comprises at least 10 different target sequences.

22. The method of any one of claims 1-21, further comprising quenching each fluorescent readout probe to inactivate the readout probe after each imaging round, optionally by chemically or enzymatically cleaving the fluorescent label from the readout probe.

23. The method of any one of claims 1-22, wherein the N-bit binary code comprises at least a 16-bit code, at least a 18-bit code, at least a 20-bit code, at least a 22-bit code, at least a 24-bit code, at least a 26-bit code, at least a 28-bit code, or at least a 30-bit code.

24. The method of any one of claims 1-23, wherein the plurality of fluorescent readout probes are the same or different from each other.

25. The method of any one of claims 1-24, wherein the plurality of fluorescent readout probes comprise at least three distinct fluorescent labels.

26. The method of any one of claims 1-25, wherein the spatial organization of the target molecule is imaged in 2 dimensions.Atty. Dkt. No.: 114198-366027. The method of any one of claims 1-26, wherein the spatial organization of the target molecule is imaged in 3 dimensions.

28. The method of any one of claims 1-27, further comprising determining abundance of the target molecule.

29. The method of any one of claims 1-28, wherein the imaging of step c) is conducted on a microscope-microfluidics system.

30. The method of claim 29, wherein the microscope-microfluidics system comprises a high-power laser.

31. The method of claim 30, wherein the high-power laser operates at about 1 to 2.5W per laser line.

32. The method of any one of claims 29-31, wherein the microscope-microfluidics system comprises a microscope body.

33. The method of claim 32, wherein the microscope body is coupled with a square multimode fiber.

34. The method of claim 33, wherein the square multimode fiber is a 0.4mm square multimode fiber.

35. The method of any one of claims 32-34, wherein the microscope body is coupled with a larger-format camera.

36. The method of claim 35, wherein the large-format camera comprises a 20.8 mm sensor.

37. The method of any one of claims 32-36, wherein the microscope body is coupled with a microfluidics chamber.

38. The method of any one of claims 29-37, wherein the microscope-microfluidics system provides a at least 2-fold, at least 3-fold, at least 5-fold, at least 7-fold, at least 9-fold, at least 10-fold, at least 12-fold, at least 15-fold, at least 18-fold, or at least 20-fold increase inAtty. Dkt. No.: 114198-3660imaging area compared to an otherwise identical system wherein non-acrydite-modified probes are used.

39. The method of claim 38, wherein the imaging area is at least 1 cm2, 2 cm2, 3 cm2, 4 cm2, 5 cm2, 6 cm2, 7 cm2, 8 cm2, 9 cm2, 10 cm2, 12 cm2, 15 cm2, 18 cm2, or 20 cm2.

40. The method of any one of claims 1-39, wherein the sample is an organ from a subject.

41. The method of claim 40, wherein the subject is a mammal.

42. The method of claim 41, wherein the mammal is a mouse or a human.

43. The method of any one of claims 40-42, wherein the organ is selected from a group consisting of heart, lung, trachea, bronchus, mouth, esophagus, stomach, small intestine, large intestine, rectum, liver, gallbladder, pancreas, kidney, ureter, urinary bladder, urethra, brain, spinal cord, pituitary gland, hypothalamus, thyroid gland, parathyroid gland, adrenal gland, ovary, testis, fallopian tube, uterus, vagina, prostate, seminal vesicle, penis, skin, spleen, thymus, lymph node, and tonsil.