Method for identifying cell boundary applied in spatial transcriptomic analysis
By using fluorescent transgenic animal models and fluorescence imaging technology, combined with cell segmentation algorithms and gene expression matrix alignment, the problems of misalignment and low efficiency caused by traditional staining methods have been solved, achieving efficient and accurate cell boundary identification and single-cell transcriptome analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- UNIV OF MACAU
- Filing Date
- 2025-10-10
- Publication Date
- 2026-04-23
AI Technical Summary
Existing spatial transcriptomics techniques struggle to accurately and efficiently identify cell boundaries in high-resolution analyses, especially in high-density cell environments where adjacent cells merge. Furthermore, traditional staining methods can affect mRNA capture efficiency and lead to misalignment.
Using a fluorescent transgenic animal model, fluorescent proteins A and B are used to label cell membranes and specific cells. Fluorescent images are acquired using a fluorescence imaging device, and the images are processed using a cell segmentation algorithm. Combined with the gene expression matrix, alignment and segmentation are performed to achieve single-cell transcriptome analysis.
It can accurately identify cell boundaries without traditional staining, improving the efficiency of cell boundary identification and the accuracy of subsequent algorithm-assisted cell image segmentation, reducing the impact of mRNA capture efficiency, and realizing the visualization and information integration of different tissue structures.
Smart Images

Figure CN2025126784_23042026_PF_FP_ABST
Abstract
Description
A method for identifying cell boundaries in spatial transcriptomics analysis Technical Field
[0001] This invention belongs to the field of spatial transcriptomics technology, specifically relating to a method for identifying cell boundaries in spatial transcriptomics analysis. Background Technology
[0002] Spatial transcriptomics is a technique that preserves spatial information about gene expression. Compared to traditional sequencing techniques or other techniques aimed at understanding the gene expression of the whole cell or parts of a tissue, spatial transcriptomics provides a powerful tool for studying the gene expression of cell populations or single cells, as well as the interaction between cells and their surrounding environment. It has wide applications in fields such as neuroscience, cancer, neuroscience, immunology, and regenerative medicine, and is particularly important for understanding tumor heterogeneity, degenerative neurological diseases, and the distribution of immune cells. However, as the sampling density of spatial transcriptome chips increases, the transcriptional information at each sampling point becomes increasingly sparse, making it impossible to identify cell types. At this point, clues about cell boundaries are needed to align and integrate sampling points, allowing cell types to be inferred through cluster analysis. In this high-resolution spatial transcriptomics, tissue alignment needs to be at the cellular level, and accurately analyzing and identifying cell boundaries is a challenge. Therefore, how to accurately and efficiently analyze and identify cell boundaries has become a key challenge in the development of this technology.
[0003] The specific workflow of spatial transcriptomics includes tissue section sample preparation, mRNA capture and reverse transcription, cDNA synthesis and amplification, high-throughput sequencing, and data analysis and visualization. These steps together constitute the process from gene expression data to the generation of spatial gene expression maps. Existing advanced technology platforms include Visium (10×Genomics) and Slide-seq. Visium is suitable for large-scale experiments and has high resolution, enabling single-cell spatial transcriptome analysis; while Slide-seq uses its unique microbead technology to capture fine and high-resolution gene expression maps.
[0004] In the process of sample processing, in order to obtain imaging information related to cell structure, traditional nuclear staining methods are usually adopted, which involve injecting chemical stains such as DAPI or Hoechst onto tissue sections. However, this action may cause displacement of the tissue sections. In addition, the affinity of chemical stains for nucleic acids may reduce the capture efficiency of mRNA. These problems are particularly prominent in high-precision cell alignment.
[0005] Traditional methods for identifying cell boundaries rely on staining the cell nucleus or cell membrane to roughly pinpoint the location and extent of the nucleus or cell. Then, trained algorithms, such as Cellpose and StarDist, are used to predict cell boundaries and perform cell image segmentation. Cellpose is a general-purpose cell boundary identification algorithm that addresses the problem of inaccurate identification due to the diversity of cell shapes. Its core method involves training a deep neural network to predict cell shape. Its main advantages are flexibility and automation; it can adapt to different microscopic image data and automatically identify cell boundaries, while also providing a user-friendly graphical interface. However, the Cellpose algorithm requires boundary image comparison, and when dealing with high-density cellular environments, it may experience the problem of adjacent cells merging, especially when cells are close together. Furthermore, because Cellpose relies on a large training dataset, manual data labeling and adjustments are still necessary before applying it to novel tissues. StarDist, on the other hand, is another cell boundary identification method that uses star-shaped convex polygons to represent cell boundaries, overcoming the limitations of traditional boundary identification methods. This method accurately depicts the circular shape of cell nuclei in microscope images, flexibly describes cell boundaries, and avoids the problem of adjacent cells merging. It performs well in crowded or high-density cell environments and does not require subsequent shape correction. However, StarDist is primarily designed for circular cell nuclei, and its effectiveness remains limited when dealing with cells with more complex shapes.
[0006] The biggest problems with this method of cell boundary identification and segmentation are the need for pre-training and its lack of general applicability. The algorithm model requires manual annotation for training before execution, making it difficult to obtain sufficient training data for unfamiliar tissues or complex cell compositions. This results in a specific model only being able to analyze specific samples, making it difficult to analyze samples with different tissues or complex cell compositions. Furthermore, because cell membrane staining affects the quality of transcriptome capture, the slides used for staining are adjacent slides to those used for spatial transcriptome analysis. However, there is a "misalignment" problem between adjacent slides and transcriptome slides at the cellular scale. Therefore, when the sampling density of spatial transcriptome reaches the subcellular level, boundary confirmation must be completed on the same slide.
[0007] Currently, boundary identification on the same slice relies solely on nuclear staining methods. For example, SCS (Subcellular Spatial Transcriptomics Cell Boundary Identification) is a method specifically designed to address the problem of identifying cell boundaries in high-resolution spatial transcriptomics. SCS combines spatial transcriptomics gene expression data with nuclear staining images to improve the accuracy of cell boundary identification. SCS mainly consists of three steps: using the Watershed algorithm to identify cell nuclei from the stained image; using the Transformer model to infer whether each point is a cell or background; and combining gene expression signals from spatial transcriptomics analysis to determine cell location, and assigning gene expression points to specific cells based on gradient flow. This method can accurately identify cell boundaries, especially in mouse brain and liver datasets, where SCS outperforms traditional boundary identification methods.
[0008] In general, existing boundary segmentation algorithms, such as Cellpose and StarDist, rely on staining cell membranes in adjacent slices to obtain fluorescence images, and then use algorithms to automatically select and segment cell boundaries. These methods suffer from problems such as staining interference, the need for pre-training, lack of general applicability, and misalignment. SCS, on the other hand, is based on a single slice, overcoming the misalignment problem by using nuclear staining to obtain fluorescence images that are less disruptive to cell tissue state. This is combined with spatial transcriptomics analysis to obtain the spatial distribution of gene expression, which is then used to algorithmically estimate cell location and extent. However, SCS still depends on staining techniques; when staining is ineffective or the gene expression data quality is poor, the cell boundary identification results will be affected. These issues represent bottlenecks that current single-cell, high-resolution, subcellular level spatial transcriptomics analysis techniques cannot overcome.
[0009] Therefore, it is necessary to develop a method for identifying cell boundaries in spatial transcriptomics analysis that does not require staining, has fewer misalignment problems, and is more versatile. Summary of the Invention
[0010] To address the aforementioned technical problems, this invention provides a method for identifying cell boundaries and / or performing single-cell transcriptome analysis in spatial transcriptome analysis, the technical solution of which is as follows:
[0011] This invention provides a method for at least one purpose (a1) or (a2) in spatial transcriptome analysis.
[0012] a1) Identify cell boundaries;
[0013] a2) Perform single-cell transcriptome analysis;
[0014] Includes the following steps:
[0015] A fluorescent transgenic animal model was selected;
[0016] A sample was prepared from the transgenic animal model using a specific preparation method;
[0017] At least one fluorescence image of the sample is acquired using a fluorescence imaging device;
[0018] Gene expression matrix of spatial transcriptome was obtained based on the sample;
[0019] The acquired fluorescence images were aligned with the gene expression matrix; and
[0020] The fluorescence image is processed using a single-cell segmentation algorithm to generate a single-cell segmentation region image.
[0021] Analysis results.
[0022] In some embodiments of the present invention, the fluorescent transgenic animal model is capable of exhibiting at least one fluorescent protein A, and the fluorescent protein A is anchored on the cell membrane of the fluorescent transgenic animal model.
[0023] In some embodiments of the present invention, the fluorescent transgenic animal model is capable of exhibiting at least one additional fluorescent protein B, which is used to label specific cells.
[0024] In some embodiments of the present invention, the fluorescent proteins A and B are independently selected from the following fluorescent proteins: green fluorescent protein, blue / cyan fluorescent protein, yellow fluorescent protein, orange / red fluorescent protein, and far-infrared fluorescent protein; and the fluorescent proteins A and B are different from each other.
[0025] In some embodiments of the present invention, the green fluorescent protein includes one of GFP, EGFP, and Superfolder GFP.
[0026] In some embodiments of the present invention, the blue / cyan fluorescent protein includes one of mTagBFP2 and mTurquoise.
[0027] In some embodiments of the present invention, the yellow fluorescent protein includes one of YPet and Venus.
[0028] In some embodiments of the present invention, the orange / red fluorescent protein includes one of mCherry, tdTomato, and mKate2.
[0029] In some embodiments of the present invention, the far-infrared fluorescent protein includes iRFP.
[0030] In some embodiments of the present invention, the specific cells may be: myeloid macrophages, fibroblasts, epithelial cells (e.g., vascular epithelial cells, lymphatic epithelial cells), stellate cells, or nerve cells. Those skilled in the art can adjust the cell type according to the experimental purpose.
[0031] In some embodiments of the present invention, the fluorescent transgenic animal model may also express multiple fluorescent proteins such as 1, 2, 3, 4, 5, 6, 7 or 8, which can be used to label different specific cell types according to different experimental purposes.
[0032] In some embodiments of the present invention, the fluorescent protein A anchored to the cell membrane is tdTomato.
[0033] In some embodiments of the present invention, the fluorescent protein B used to label specific cells is EGFP.
[0034] In some embodiments of the present invention, the fluorescent transgenic animal model is selected from at least one of mice, rats, rabbits, monkeys, dogs, sheep, fruit flies, nematodes, chickens, African clawed frogs, pigs, and fish (zebrafish).
[0035] In some embodiments of the present invention, the fluorescent transgenic animal model is selected from mice.
[0036] In some embodiments of the present invention, the fluorescent transgenic mouse model includes dual fluorescent mice (mT / mG), LysM-Cre mT / mG mice, and R26R-Brainbow2.1 mice.
[0037] Dual-fluorescent mice achieved dual-color fluorescent labeling of the cell membrane through gene design targeting tdTomato (mT) and EGFP (mG). In the presence of Cre recombinase, the mT cassette was cleaved, resulting in the expression of only the green fluorescence of mG.
[0038] The LysM-Cre mT / mG mouse model utilizes a myeloid-specific Cre recombinase (LysM-Cre) to drive the cell membrane localization of red fluorescent protein (tdTomato) and green fluorescent protein (EGFP). tdTomato marks the cell boundary, while EGFP marks myeloid cells (such as Kupffer cells and cardiac macrophages).
[0039] Rainbow mice (Brainbow technology) are a genetic engineering model that achieves multicolor cell labeling through the random combination and expression of multicolor fluorescent proteins. The direction and location of random recombination of fluorescent protein genes can be controlled by designing different loxP sites.
[0040] In some embodiments of the present invention, visualization of different structures and proteins can be further achieved. When the transgenic animal used is ROSA... mT / mG In mice, myeloid macrophages expressed green fluorescence (EGFP), while other cell membranes expressed red fluorescence (tdTomato). Furthermore, different tissues and structures could be visualized using excitation light of different wavelengths. For example, collagen fibers could be contrasted using second harmonic excitation at 920 nm (460 nm), and pseudo-coloring was performed using blue light in the processing software.
[0041] Furthermore, its myeloid macrophages express green fluorescence (EGEP), while blood vessels and lymphatic vessels express red fluorescence (tdTomato). Under dual-color excitation at 920nm and 1040nm, the photodetector channel corresponding to the green overlay (500-600nm) can detect the EGFP fluorescence emitted by macrophages and the short-wavelength red fluorescence of tdTomato, while the photodetector channel corresponding to the red overlay (600-690nm) can detect the long-wavelength red fluorescence of tdTomato. After the two channels are superimposed, yellow blood vessels and lymphatic vessels and green macrophages are formed, thus presenting the outlines of different cells, tissues and organs.
[0042] In some embodiments of the present invention, the sample preparation method includes: tissue collection, cryo-embedding, sectioning, tissue permeabilization, and obtaining tissue sections.
[0043] In some embodiments of the present invention, fluorescent images can be subjected to color-coded excitation light of different wavelengths to obtain visualization results of proteins, cells or tissues of different colors.
[0044] In some embodiments of the present invention, the cell segmentation algorithm includes at least one of Cellpose, Mesmer, StarDist, ilastik, Bin2Cell, and scBERT.
[0045] In some embodiments of the present invention, the alignment includes coarse alignment and precise alignment.
[0046] In some embodiments of the present invention, the coarse alignment includes rotating and moving the fluorescence image and the gene expression matrix so that the edge contour overlaps with the tissue morphology features.
[0047] In some embodiments of the present invention, the precise alignment includes: aligning a fluorescence image with at least one characteristic region on a gene expression matrix.
[0048] In some embodiments of the present invention, the feature region includes at least one of the following:
[0049] (1) Fluorescent images of the cell membrane serve as characteristic regions of the cell membrane;
[0050] (2) The fluorescent images of labeled cells are used as cellular feature regions in local areas;
[0051] (3) Cellless void regions in the sample are considered as vacancy characteristic regions; and
[0052] (4) Regions without fluorescent labels are considered negative feature regions with no gene expression.
[0053] Cellless cavitary regions in a sample are cellless areas formed by the lumen structures in the tissue. The shape of these lumen structures can be used as a feature region for alignment comparison.
[0054] Regions without fluorescent labels can be heterologously seeded cells in animal models, such as tumor cell models. Since these tumor cells originate from other individuals and do not express fluorescent genes, they lack any gene markers on imaging and can be used as negative feature regions without gene expression for comparative analysis.
[0055] Regions without fluorescent labels can also be extracellular components, such as collagen fibers, extracellular matrix, or other substances. Since they lack any gene markers on imaging, they can be used as negative feature regions lacking gene expression for in-situ comparison.
[0056] In some embodiments of the present invention, the analysis results include steps of transcriptome integration and single-cell annotation.
[0057] In some embodiments of the present invention, at least one of the following methods—cell boundary-guided binning, spatial data convolution, and cluster analysis—is used to optimize the analysis results.
[0058] In some embodiments of the present invention, the resolution of the cell boundary guided bin is 13-15 micrometers, preferably 14 micrometers.
[0059] The beneficial effects of this invention are:
[0060] This invention proposes a method for identifying cell boundaries in spatial transcriptome analysis. Generally, spatial transcriptome analysis involves sample preparation, optical structural image generation, spatial transcriptome information collection, alignment of optical structural images with spatial transcriptome information, and single-cell annotation. The key feature of this invention is the use of a specific fluorescent protein transgenic animal model, coupled with an optical and transcriptome image alignment method tailored to that model. This allows for clear visualization of the structure and outline of individual cells under a fluorescence imaging system, eliminating the need for cell membrane or nuclear staining to obtain fluorescence images. Therefore, it avoids the drawbacks of staining affecting mRNA capture efficiency and sample displacement. Besides significantly improving cell boundary identification efficiency, this method also enhances the accuracy of subsequent algorithms for cell image segmentation and large-scale language model-assisted cell annotation.
[0061] The method of this invention enables the visualization of different tissue structures, which is beneficial for connecting with the interpretation of clinical histopathology. While visualizing boundaries, it simultaneously achieves: 1. Integration of spatial omics information with tissue structure alignment; 2. Connection of molecular and cellular information with the interpretation of clinical histopathology. Attached Figure Description
[0062] The present invention will be further described below with reference to the accompanying drawings and embodiments, wherein:
[0063] Figure 1 is a flowchart of the method of the present invention.
[0064] Figure 2 shows two-photon fluorescence images of the ears of transgenic mice. Excitation wavelengths: 960 nm and 1040 nm; field of view: 310 × 310 μm.
[0065] Figure 3 shows unstained two-photon fluorescence images of cell boundaries and myeloid cell distribution in frozen tissue sections from LysM-Cre mT / mG mice. (A) Frozen tissue sections of liver; (B) heart; (C) spleen. tdTomato (red) and EGFP (green) fluorescence were excited at wavelengths of 1040 nm and 960 nm, respectively. Field of view: 317 × 317 μm.
[0066] Figure 4 shows fluorescence images of fresh and frozen sections of liver tissue.
[0067] Figure 5 shows fluorescence images of fresh and frozen sections of tumor tissue.
[0068] Figure 6 shows: A is an image of the red fluorescence signal of tdTomato in liver tissue obtained before the spatial transcription experiment; B is an image of the green fluorescence signal of GFP in liver tissue obtained before the spatial transcription experiment; C is a distribution map of the number of transcriptomes captured by the spatial transcription chip; and D is an image of the image alignment program.
[0069] Figure 7 is a schematic diagram of merging spatial transcription data into single-cell transcriptome information based on image cell contour information.
[0070] Figures 8A-8H show the cell segmentation and binning analysis results of spatial transcriptome data, where: 8A is a large-scale unstained fluorescent image of liver slices (red: tdTomato; green: EGFP); 8B is the distribution map of total reads in liver slices obtained by spatial transcriptome analysis; 8C is the structural similarity (SSIM) analysis of the final registration between the fluorescent image and the read distribution map; 8D is the segmentation of tdTomato-labeled cell membranes (red) in liver slices using Cellpose, with a field of view of 317×317μm; 8E is the segmentation of EGFP-labeled myeloid cells (green) in liver slices using ImageJ, with a field of view of 317×317μm; 8F is the area distribution of tdTomato-labeled cells segmented by Cellpose; 8G is the average UMI count / gene count per cell under different square binning sizes; 8H is the average UMI count / gene count per cell after cell boundary-guided binning.
[0071] Figures 9A-9F show the data analysis results after convolution processing, where: 9A shows the average UMI count / gene count per cell after convolution and cell boundary guided binning; 9B shows the 20 genes with the highest expression levels in the liver sample after convolution and cell boundary guided binning; 9C shows the expression of Alb, Cyp2f2, and Cyp2c29 genes in liver slices; 9D shows the results of cluster analysis; 9E shows the expression levels of myeloid cell marker genes in each cluster; and 9F shows the spatial distribution of myeloid cell marker genes and EGFP-labeled cells. Detailed Implementation
[0072] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.
[0073] The experimental method is as follows:
[0074] Mouse model and tissue section preparation:
[0075] Transgenic mice LysM-Cre\mTmG(ROSA) mT / mGThe transgenic mice (LysM-Cre\mTmG) were generated by crossing B6.129P2-Lyz2tm1(cre)Ifo / J mice with STOCK Gt(ROSA)26Sortm4(ACTB-tdTomato,-EGFP)Luo / J mice. In these mice, myeloid cells express EGFP, while other cells express tdTomato (see reference: Li, Y., et al., Imaging of macrophage mitochondria dynamics in vivo reveals cellular activation phenotype for diagnosis. Theranostics, 2020.10(7):p.2897-2917). In this embodiment, 6- to 8-week-old transgenic mice (LysM-Cre\mTmG) were used. In this embodiment, various fresh tissue samples, including heart, spleen, and liver, were collected and embedded in optimal cutting temperature compound (OCT). The embedded samples were rapidly frozen in liquid nitrogen using isopentane as the medium. Finally, the tissue samples were sectioned using a cryostat (Leica, CM3050S, Germany) and stored at -80°C before use.
[0076] Multiphoton imaging of tissue sections:
[0077] Before imaging, sections were fixed with 4% paraformaldehyde for 10 minutes at room temperature. The sections were then washed with phosphate-buffered saline (PBS) to remove the paraformaldehyde. The sections were then immersed in PBS and covered with coverslips. Finally, the sections were observed under a Nikon Eclipse Inverted Multiphoton Microscope (A1MP+Eclipse Ti-2E, Nikon Instruments, Japan) using a 40x NA=1.15 water immersion objective. Fluorescence images were acquired at excitation wavelengths of 960 nm and 1040 nm. At 960 nm excitation, EGFP two-photon fluorescence signals were collected in the wavelength range of 506 nm to 593 nm. At 1040 nm excitation, tdTomato fluorescence signals were collected in the wavelength range of 604 nm to 678 nm. To avoid sample damage, the laser power behind the objective was kept below 20 mW.
[0078] Large-scale confocal imaging with a resolution of 19742×15020 pixels and a pixel dwell time of 2.42 μs was acquired within 1 hour using excitation wavelengths of 488 nm and 561 nm. EGFP fluorescence signals in the wavelength range of 500 nm to 550 nm were collected under 488 nm excitation. tdTomato fluorescence signals in the wavelength range of 570 nm to 620 nm were collected under 561 nm excitation. The images consist of 20×15 image patches with 5% overlap between each patch.
[0079] For spatial transcriptomics liver sample processing
[0080] Following microscopic imaging, the same liver slice was used for spatial transcriptomics analysis based on a spatial chip manufactured by Centrillion Technologies (Ding, X., et al., Spatial Transcriptomics Sequencing of Mouse Liver at 2μm Resolution Using a Novel Spatial DNA Chip. BioRxiv, 2024.). The experimental procedures were performed according to the Yosemite product instructions. The procedure was as follows: Sample slices were permeabilized with pepsin at 37°C for 5 minutes. The samples were then inverted onto a chip soaked in hybridization buffer and incubated at 42°C for 4 hours. During this incubation step, messenger RNA (mRNA) diffused to the bottom of the chip, where the poly(A) tail could capture the poly(A) cap of the mRNA transcript. Amplified complementary deoxyribonucleic acid (cDNA) was obtained after reverse transcription, second-strand synthesis, and polymerase chain reaction (PCR). cDNA samples of 300-700 base pairs (bp) were then collected using magnetic bead screening (Beckman Laboratories, B23317). Libraries were constructed using the NEBNext Multiplex Oligonucleotide Kit (New England Biolabs, E7335) through fragmentation, end repair, ligation, and adapter addition, followed by sequencing using the DNBSEQ-T7-PE150 (BGI Genomics, China).
[0081] Spatial transcriptome data preprocessing:
[0082] The raw sequencing data is stored in FASTQ files, which need to be converted to h5ad files for subsequent analysis using Python on a Linux system. First, PostMaster, developed by Zhongyimingke Biotechnology Co., Ltd., is used to extract unique molecular identifiers (UMIs) and coordinates from the raw FASTQ files. Then, reads are mapped to the reference genome GRCm39 using the STAR alignment tool. The resulting files are processed by FeatureCounts to generate a counting matrix. Subsequently, Bamreader, also developed by Zhongyimingke Biotechnology Co., Ltd., is used to create comma-separated value (CSV) files containing gene IDs, UMIs, and spatial coordinates. Finally, the resulting files are converted to h5ad format for downstream analysis.
[0083] Spatial registration between captured unique molecular identifiers (UMIs) and fluorescence imaging:
[0084] A UMI counting matrix containing 22,291 genes and measuring 5000×5000 pixels was projected onto a single image, with expression levels normalized to the range of 0-255. For spatial registration, a graphical user interface (GUI) based on the PyQt5 toolkit was developed to adjust parameters of the image's rigid transformations (including rotation and translation). First, the gene image was fixed and then adjusted to match the screen resolution. Subsequently, a fluorescence image was added, similarly adjusted to match the screen resolution, and rescaled to the same resolution as the gene image. A rescaling factor of 0.155 was determined using the pixel resolution ratio of the fluorescence image (0.31 μm) to the gene image (2 μm). After this, the fluorescence image was aligned with the gene image by rotation and translation, and any fluorescence image portions exceeding the overlap were cropped. The fluorescence intensity of missing fluorescence image portions relative to the gene image was filled with 0. During this process, an affine transformation matrix, which can be defined within the software, was applied, as shown below:
[0085] Where P(x,y) and P(x′,y′) are the original point and the point after the affine transformation, respectively. θ controls the rotation of the image, while t... x ,t y It enables translation in both horizontal and vertical directions.
[0086] Next, an image similarity evaluation method—structural similarity index measurement (SSIM)—is used to optimize the best registration for fluorescence imaging: c1 = (K1L) 2 c2 = (K2L) 2 ;
[0087] Where, μ xμ y These are the mean values of the total gene expression image and the fluorescence image, respectively; These are their variances and covariances, respectively; K1 = 0.01, K2 = 0.03, and L = 255 are constants.
[0088] Cell boundary-guided spatial transcriptional partitioning
[0089] After obtaining the aligned fluorescence images, Cellpose software and its Cyto2 model were used to segment cell boundaries with tdTomato signals. ImageJ software was used to identify myeloid cells with EGFP fluorescence signals, and cells with an area smaller than 60 pixels were filtered out. 2 To match the coordinates of the fluorescence images to the gene expression data, each pixel of the gene expression image was magnified into a square region on the fluorescence image based on their pixel resolution and registration information. Finally, the entire spatial transcriptome was partitioned using the pixel coordinates on the gene chip and the cell center coordinates. Different standard square partition sizes were also used to process the data for comparison.
[0090] Cluster analysis of spatial transcriptomes
[0091] To enhance gene abundance in downstream analysis, a 5×5 pixel square window was applied to the h5ad file for convolution, followed by cell boundary-guided partitioning of the data. Subsequently, the Scanpy package in Python was used to process the data for cluster analysis. In short, the Pearson residual algorithm (with a component count of 15) was used to identify highly variable genes, from which the top 5000 genes were selected. Dimensionality reduction was then performed using Principal Component Analysis (PCA) and Unified Manifold Approximation and Projection (UMAP). Finally, the Leiden algorithm was applied for cluster analysis at a resolution of 0.3.
[0092] Example 1: Fluorescence images of mouse ears
[0093] This example demonstrates the visualization results of various proteins, cells, tissues, and organs based on different fluorescent proteins.
[0094] Transgenic mice with fluorescent cell membranes (ROSA) were produced by crossing Jackson mice with LysM Cre. mT / mG It can reveal complete tissue structures under a multiphoton fluorescence microscope, including blood vessels (yellow), lymphatic vessels (yellow), collagen fibers (blue), myeloid macrophages (green), and the cell membranes of other cells are labeled with red fluorescent protein (tdTomato) to show the outline of the cells.
[0095] In this design, green represents EGFP expressed by myeloid macrophages, while red represents tdTomato anchored on the cell membrane. Blue collagen fibers are generated using a 920nm excitation frequency second harmonic (460nm) contrast filter, and are overlaid in blue in the processing software. Yellow blood vessels and lymphatic vessels are generated under dual-color excitation at 920nm and 1040nm. The photodetector channel corresponding to the green overlay (500-600nm) detects the EGFP fluorescence emitted by macrophages and the short-wavelength red fluorescence of tdTomato, while the photodetector channel corresponding to the red overlay (600-690nm) detects the long-wavelength red fluorescence of tdTomato. The superposition of these two channels creates the yellow blood vessels and lymphatic vessels, and the green macrophages, thus revealing the outlines of different cells, tissues, and organs.
[0096] This method of stimulating color matching enables the visualization of different tissue structures, which is beneficial for connecting the interpretation of clinical histopathology.
[0097] Example 2
[0098] Using two-photon fluorescence imaging, this example obtained unstained images from frozen sections of the liver, heart, and spleen of LysM-Cre mT / mG mice (Fig. 3). tdTomato fluorescence (red) marked clear cell boundaries, while EGFP fluorescence (green) marked myeloid cells. In the liver, hepatocytes were distributed within the lobules, showing prominent tdTomato signals (orange) outlining their membranes, with myeloid cells (green) scattered within them. Longitudinal sections of the heart showed cardiomyocytes with elongated tdTomato-positive membranes (orange), closely packed together, while cardiac macrophages (green) were visible within the cardiomyocytes. In the spleen, spleen cells (orange) were small, with splenic macrophages (green) scattered within them. These results demonstrate that unstained tissue sections from LysM-Cre mT / mG mice provide high-quality imaging of cell boundaries, clearly showing tissue structure and myeloid cell distribution.
[0099] Further observation of the fluorescence images of liver sections (Figure 4) revealed the location of macrophages using EGFP green fluorescence (500-550 nm), clearly showing the pseudopodia characteristic of macrophages. The boundaries of liver cells were visible using tdTomato red fluorescence (570-738 nm). Within the area enclosed by the red fluorescence signal, there were some circular black regions without either red or green fluorescence. Hoechst blue fluorescence (425-475 nm) excited by a 405 nm laser confirmed that these areas were the nuclei. Image comparison showed that the boundaries of fresh samples were more detailed and distinct, and the macrophage morphology was more clearly defined, while the ice-cut samples showed fewer pseudopodia.
[0100] Example 3: Cell Boundary Differentiation in an Exogenous Tumor Model
[0101] A cancer animal model was constructed by collecting Py8119 breast cancer cells. Py8119 was collected to a concentration of 1×10⁻⁶. 7 / mL, injected into transgenic mice (ROSA) mT / mG The tumor was injected into the fat pad (100 μL) in mice, and mice were anesthetized by intraperitoneal injection of avertin (0.25 mg / g). Mice were fed until the tumor volume reached the desired size (approximately 500 mm²). 3 The mice were euthanized with carbon dioxide. The tumor tissue and normal tissue were cryo-embedded and sectioned for imaging (Figure 5). Since tumor cells are foreign cells, they should not normally have clear cell boundaries. However, in the microenvironment, tumor-associated macrophages are ubiquitous, so the black areas become tumor areas. After cryo-sectioning, the tumor boundaries were stained by green fluorescent protein on the macrophages, thus revealing the cell boundaries. This successfully identified cell boundaries without external staining or the use of dyes.
[0102] Example 4: Alignment of Fluorescent Images with Gene Expression Matrix
[0103] Existing techniques for spatial transcriptome analysis typically involve two adjacent tissue sections. One section is stained to obtain fluorescent images of cell structure and boundaries, while the other is used for mRNA capture, reverse transcription, sequencing, localization, and quantification to obtain gene expression matrix information. Finally, the fluorescent images and gene expression matrices are aligned. However, even adjacent sections can be expected to have slight differences in shape and size, especially at high resolution and small scale cell alignment. This invention, however, utilizes a fluorescent protein transgenic animal model. After permeation, one side of the tissue section is used for mRNA capture and hybridization in the spatial transcriptome, while the other side allows the fluorescence imaging system to capture cellular fluorescence images. This enables the acquisition of both fluorescent images and the gene expression matrix information of the spatial transcriptome on a single section.
[0104] Since both the fluorescence image and the gene expression matrix information come from the same sample, the overlap can be significantly improved in the subsequent alignment. This embodiment employs an optical and transcriptional image alignment method in conjunction with a fluorescent transgenic animal model, and develops a program to execute each alignment step, producing suitable images and information for subsequent single-cell annotation. The alignment steps are roughly divided into coarse alignment and precise alignment. In this embodiment, after obtaining the fluorescence images of tissue sections (A and B in Figure 6) and the spatial distribution of transcriptome expression levels (gene expression matrix) obtained from spatial transcription technology (C in Figure 6), the periphery of the sections is rotated and moved so that the images can be superimposed and aligned according to the edge contours of the sections and the morphological features of the tissue (D in Figure 6). This step is coarse alignment. The next step is to perform precise alignment using several feature landmarks, including: the EGFP fluorescence image of macrophages as a cellular landmark of the local area, the cavities formed by the lumen in the section as a feature area without gene expression, or the collagen fiber area as a feature area without gene expression, finally achieving precise alignment between the fluorescence image and the gene expression matrix.
[0105] After obtaining the boundary image of a single-cell plate, the gene expression matrix of that region can be merged (Figure 7).
[0106] Example 5: Cell boundary-guided transcript binning can improve the UMI count and gene count per cell in spatial transcriptomics.
[0107] This embodiment utilizes a high-density spatial barcode oligonucleotide array (Yosemite, Centrillion Technologies) for high-resolution spatial transcriptomics analysis. The array contains 5000 × 5000 capture sites, each 2 × 2 micrometers in size with no gaps, for a total capture area of 10 × 10 millimeters. After large-area two-photon fluorescence imaging, the same liver slice was transferred to the spatial transcriptomics chip. The RNA integrity index (RIN) of fresh tissue slices was higher than 6.5, ensuring high-quality RNA capture. Because tissues are rich in proteins such as collagen, laminin, and proteoglycans, which may form barriers and capture mRNA, these proteins were partially digested using optimized pepsin treatment to release the mRNA within the tissue. Subsequently, the tissue slices were immersed in hybridization buffer to promote mRNA capture. The barcode oligonucleotides on the chip capture mRNA through their 3' poly-A tails. Each oligonucleotide contains a unique molecular identifier (UMI) sequence, supporting high-throughput sequencing and precise labeling of captured mRNA molecules. After reverse transcription, on-chip second-strand synthesis, cDNA amplification, and purification, high-quality cDNA fragments (300-700 bp) were obtained. Finally, a sequencing library (200 Gbps data volume, sequencing saturation exceeding 94%) was constructed, and the x and y coordinates on the chip were decoded using STAR mapping. The UMI value of each gene was obtained through gene counting. The address decoding accuracy exceeded 95%, and the gene alignment rate was approximately 70%. A 5000×5000 expression matrix was stored in h5ad format for subsequent analysis.
[0108] To align gene expression data with imaging data, fluorescence images (Fig. 8A) were compared with total count readouts of the same liver slice (Fig. 8B). Image registration was performed using structural similarity index (SSIM) analysis (Fig. 8C). Subsequently, cell domains enclosed by tdTomato-labeled cell membranes were segmented using Cellpose (Fig. 8D), and EGFP-labeled myeloid cells were segmented using ImageJ (Fig. 8E). After segmentation, the registered x'-y' pixel coordinates (0.31 μm resolution) of each cell domain were obtained and projected onto xy feature sites (2 μm resolution) on a gene chip for spatial transcriptomics analysis.
[0109] Statistical analysis of each 2-micron capture site showed low UMI counts and the number of detected genes (Fig. 8G). This sparsity is a direct result of the high-density spatial sampling of the microarray. To improve the gene detection efficiency of cluster analysis, naive square binning is commonly used to merge transcripts from neighboring sites. The effect of different bin sizes on the spatial transcriptome was tested (Fig. 8G). The average UMI count per cell obtained by the commonly used 8-micron binning (the standard size in recent single-cell resolution studies) was 79.66, and the number of genes was 51.15, both higher than the results of the original 2-micron binning. Binning significantly reduced the number of cells with zero counts or zero genes, and showed a clear statistical peak in the histogram. By calculating the area of the cell domain segmented by Cellpose, the median area was determined to be 187.97 microns. 2 The optimal binning size is 14 micrometers. Under this condition, the average UMI count per cell is 228.79, and the average number of genes is 125.10. When using 16-micrometer binning, the average UMI count and the average number of genes reach 318.64 and 167.44, respectively.
[0110] After employing cell boundary-guided binning, the results showed that the average UMI count and gene count were 234.50 and 121.77, respectively, both higher than those of 8-micron binning (Figure 8H). This indicates that traditional 8-micron binning over-segments cell domains, leading to an underestimation of UMI counts per cell. Conversely, 16-micron binning may overestimate the binning range and introduce mRNA counts from adjacent cell domains. However, cell boundary-guided binning results are close to those of 14-micron binning. This method not only accurately reflects but also enhances the abundance of UMI counts and gene counts per cell. The study demonstrates that stain-free cell boundary-guided binning using LysM-Cre mT / mG mice provides better data quality than naive square binning in spatial transcriptomics analysis.
[0111] Example 6: Spatial data, after convolution and cell boundary-guided binning, can be used for downstream analysis.
[0112] During microarray hybridization, mRNA may diffuse beyond cell boundaries, blurring the gene map of each cell. To address this, this embodiment employs a convolution function to integrate mRNA signals from neighboring regions to enhance gene expression levels, combined with cell boundary-guided binning. Through tissue alignment analysis, a 5×5 pixel square convolution window (approximately 5 μm radius) was initially tested. Results showed a significant increase in the average UMI count and gene number per cell (Figure 9A). Among the top 20 highly expressed genes, many are closely related to the liver microenvironment, including the hepatocyte marker gene Alb (albumin) and Mup22 and Mup3 (highly expressed in mouse liver) involved in chemical communication (Figure 9B). Subsequently, three representative gene markers—Alb (hepatocyte marker), Cyp2f2 (periportal region marker), and Cyp2c29 (pericentral region marker)—were selected to analyze their spatial expression patterns. Results showed that Alb was widely expressed throughout the liver slices, while Cyp2f2 and Cyp2c29 exhibited distinct clustering and regional distribution characteristics (Figure 9C). Therefore, convolution and cell boundary-guided binning not only improve the data quality of downstream analysis, but also accurately present the spatial information of gene expression.
[0113] Subsequently, cluster analysis was performed in this embodiment, identifying a total of 19 cell clusters (Figure 9D). The results showed that EGFP-labeled myeloid cells were significantly enriched in clusters 0 and 1, and highly expressed myeloid cell marker genes (such as Mrc1 and CD74) (Figures 9D and 9E). Notably, the myeloid cell marker gene Adgre1 (encoding the F4 / 80 antigen) did not show significantly high expression in EGFP-labeled cells (Figure 9F). This may be due to the low expression level of this gene in the tissue itself. Furthermore, the spatial distribution of Mrc1 and CD74-labeled myeloid cells was highly consistent with that of EGFP-labeled cells. Overall, these results indicate that the data are of high quality and accuracy, suitable for downstream analysis.
Claims
1. A method for a1), a2) at least one purpose in spatial transcriptome analysis; a1) identifying cell boundaries; a2) performing single-cell transcriptome analysis; comprising the following steps: selecting a fluorescent transgenic animal model; preparing a sample from the transgenic animal model by a preparation method; acquiring at least one fluorescent image of the sample using a fluorescent imaging device; obtaining a gene expression matrix of the spatial transcriptome based on the sample; aligning the acquired fluorescent image with the gene expression matrix; and processing the fluorescent image by a cell segmentation algorithm and generating a single-cell segmented region image; analyzing the results.
2. The method of claim 1, wherein: the fluorescent transgenic animal model is capable of expressing at least one fluorescent protein A, and the fluorescent protein A is anchored on the cell membrane of the fluorescent transgenic animal model; preferably, the fluorescent transgenic animal model is capable of expressing at least one additional fluorescent protein B, which is used to label specific cells.
3. The method of claim 2, wherein: the fluorescent proteins A and B are independently selected from the following fluorescent proteins: green fluorescent protein, blue / cyan fluorescent protein, yellow fluorescent protein, orange / red fluorescent protein, and far-red fluorescent protein; and the fluorescent proteins A and B are different from each other.
4. The method of claim 3, wherein: The fluorescent transgenic animal model is selected from the group consisting of bi-fluorescent mice, preferably ROSA mT / mG mice.
5. The method of claim 1, wherein: the sample preparation method comprises tissue collection, frozen embedding, sectioning, tissue permeation, and obtaining a tissue section.
6. The method of claim 1, wherein: the cell segmentation algorithm comprises at least one of Cellpose, Mesmer, StarDist, ilastik, Bin2Cell, and scBERT.
7. The method of claim 1, wherein: the alignment comprises coarse alignment and fine alignment.
8. The method of claim 1, wherein: the coarse alignment comprises rotating and moving the fluorescent image and the gene expression matrix so that the edge profile overlaps with the tissue topography feature; the fine alignment comprises aligning at least one feature region on the fluorescent image and the gene expression matrix.
9. The method of claim 7, wherein: the feature region comprises at least one of: (1) the fluorescent image of the cell membrane as a cell membrane feature region; (2) the fluorescent image of the labeled cells as a cell feature region in the local region; (3) the hollow region in the sample without cells as a void feature region; and (4) the cell region without fluorescent labeling as a negative feature region without gene expression.
10. The method of claim 1, wherein: the analysis results comprise the steps of transcriptome integration and single-cell annotation; preferably, at least one of cell boundary guided binning, spatial data convolution, and clustering analysis is used to optimize the result analysis.
Citation Information
Patent Citations
Fluorescence in situ hybridization (FISH) image parallel processing and analysis method
CN106296635A
Spatial transcriptome data processing method and device and computer readable storage medium
CN114267414A
High-precision spatial transcriptome analysis method and system based on immunofluorescence
CN115308414A
Cell boundary recognition model training method and cell boundary recognition method
CN118334655A
Colony Detection
US20100074507A1