A spatial omics multi-modal fusion method at single cell level

By integrating spatial transcriptomics, single-cell omics, and histological image data using the STEP method, the limitations of resolution and high cost in existing technologies are addressed, enabling efficient multi-omics information reconstruction at the whole tissue scale and improving the resolution and data utilization efficiency of spatial omics.

CN121011247BActive Publication Date: 2026-02-10HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511543192.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-02-10
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing spatial omics technologies face limitations in multi-omics reconstruction at single-cell resolution and whole-tissue scales, including resolution limitations, insufficient utilization of complementary data types, and high costs, making it difficult to efficiently integrate spatial transcriptomics, single-cell omics, and histological image data.

Method used

We employ a single-cell level spatial omics multimodal fusion method (STEP) that integrates spatial transcriptomics, single-cell omics, and histological image data through probabilistic inference and deep learning. We establish a hybrid computational framework that integrates probabilistic inference and deep learning, construct a spatial cell network, realize the spatial diffusion and identification of cell type information, and complete gene expression and reconstruct multi-omics information.

Benefits of technology

This method efficiently integrates multi-omics data at single-cell resolution, improves spatial resolution, reduces costs, and achieves high-resolution multi-omics information reconstruction at the whole tissue scale. It significantly overcomes the limitations of existing methods, supports seamless integration across platforms and slices, and enhances generalization ability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121011247B_ABST
    Figure CN121011247B_ABST
Patent Text Reader

Abstract

A single-cell level spatial omics multi-modal fusion method, comprising: extracting differential expression genes and nuclear spatial morphological features from spatial transcriptome data, single-cell sequencing data and histological images, and realizing field adaptation between different platforms by using conditional variational autoencoder. Based on a probability inference model, spatial transcriptome expression, single-cell omics and morphological features are fused to jointly infer the type and gene expression level of each cell. A spatial cell network is constructed by a graph attention mechanism to realize the spatial diffusion and recognition of cell types in the whole slice range. Combined with a multi-omics enhancement module, the unmeasured gene and protein expression are completed based on expression similarity, and the prediction consistency is improved through spatial correction. The method realizes high-resolution reconstruction of single-cell multi-omics information in three-dimensional space, improves the information coverage and spatial resolution of spatial omics data, and provides an efficient and low-cost solution for spatial biology and precision medicine research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of machine learning technology with spatial omics, bioinformatics and computational biology, and in particular to a multimodal fusion method for spatial omics at the single-cell level. Background Technology

[0002] Space omics technology has greatly expanded our understanding of the structure and function of biological systems, enabling researchers to precisely map the spatial distribution of molecules such as gene expression, proteins, and metabolites in the in-situ environment of tissues. Compared with traditional methods, space omics preserves crucial spatial information, providing new means to reveal cellular tissue structure and intercellular interactions. Achieving whole-tissue maps with single-cell resolution, multi-omics coverage, and multi-slice coverage is an important goal of current space biology, but it still faces multiple challenges in practice, including resolution, cost, and technical difficulty.

[0003] Existing spatial transcriptomics technologies are divided into two categories: sequencing-based and imaging-based. While sequencing-based technologies can achieve full transcriptome coverage, their spatial resolution is limited, making it difficult to accurately resolve single-cell heterogeneity. Imaging-based technologies offer single-cell resolution but are limited by the number of genes that can be detected. Single-cell sequencing can obtain cellular-level molecular information but completely loses spatial location information. In recent years, spatial multi-omics technologies have made continuous progress, but their high cost and complexity have limited their widespread adoption. Computational methods that integrate spatial transcriptomics, single-cell data, and histological images provide a more flexible and low-cost implementation path for spatial multi-omics.

[0004] Currently, most methods focus on pairwise integration of the three modalities mentioned above. For example, the SpaGE method proposed by Tamim Abdelaal et al. utilizes single-cell sequencing data for gene expression completion, and the CARD method proposed by Ying Ma et al. achieves spatial cell type resolution or spatial location prediction. However, these methods have not fully utilized the potential of trimodal synergy. Even recent advances such as the SpatialScope method proposed by Xiaomeng Wan et al. and the STIE method proposed by Shijia Zhu et al., while attempting trimodal integration, often neglect histological images, an important source of morphological information, making it difficult to achieve high-resolution multi-omics reconstruction at the whole-tissue scale. In addition, the acquisition of existing spatial multi-omics data remains expensive, while the widespread availability of histological images provides new opportunities for low-cost inference of molecular information. Currently, there is an urgent need for a novel computational method that can efficiently integrate spatial transcriptomics, single-cell omics, histological images, and multi-omics data to achieve joint spatial multi-omics analysis at single-cell resolution and the whole-tissue scale, thereby promoting the further development of space biology and precision medicine.

[0005] It should be noted that the information disclosed in the background section above is only for understanding the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The main objective of this invention is to overcome the deficiencies in the aforementioned background technology and provide a spatial omics multimodal fusion method at the single-cell level, which can obtain spatial multi-omics integrated information at the whole tissue scale and single-cell resolution at low cost and high efficiency.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A method for multimodal fusion of spatial omics at the single-cell level includes the following steps:

[0009] S1. Preprocess the input spatial transcriptome data, single-cell sequencing data, and histological images;

[0010] S2. Perform cell segmentation on the histological image and extract the spatial location and morphological features of each cell;

[0011] S4. Perform domain-adaptive processing on spatial transcriptome data and single-cell data to eliminate inter-platform differences;

[0012] S5. Based on the spatial relationship between spatial points and cell locations, establish the correspondence between spatial points and cells;

[0013] S6. Construct a probabilistic inference model that integrates spatial transcriptome expression, single-cell data, and morphological features to jointly infer the type and gene expression level of each cell;

[0014] S7. Introduce a prior distribution to regularize and spatially constrain the model parameters;

[0015] S8. Construct a spatial cell network based on graph neural networks to realize the spatial diffusion and recognition of cell type information;

[0016] S9. Based on the expression similarity of single-cell data, complete the gene expression and protein expression of spatial cells;

[0017] S11. Integrate data from different sources and adjust the model;

[0018] S12. Achieve reconstruction and visualization of single-cell multi-omics information at a three-dimensional spatial scale.

[0019] The present invention has the following beneficial effects:

[0020] To overcome the limitations of existing spatial omics multimodal data integration methods, such as limited spatial resolution, insufficient utilization of data type complementarity, and high cost, this invention provides a single-cell-level spatial omics multimodal fusion method. This method can efficiently integrate and infer spatial transcriptomics, single-cell omics, histological images, and multi-omics data at single-cell resolution, achieving high-resolution multi-omics information reconstruction at the entire tissue scale. The method establishes a hybrid computational framework (STEP) that integrates probabilistic inference and deep learning. This method can flexibly integrate multiple types of data, including spatial transcriptomics, single-cell omics, and histological images. Through a specially designed information integration strategy, it fully leverages the complementary advantages of the three modalities, achieving efficient reconstruction of spatial multi-omics maps at both single-cell resolution and the entire tissue scale, significantly overcoming the limitations of existing methods.

[0021] Specifically, this invention utilizes the spatial relationship between spatial points and single cells, combined with cell morphology and gene expression characteristics, to construct a joint probabilistic inference model, effectively bridging the information gap between spatial detection points and single cells. Regarding platform compatibility, this invention supports both sequencing-based and imaging-based spatial transcriptomics, enabling spatial resolution improvements from the spatial point level to the single-cell level in sequencing-based spatial transcriptomics, and expanding the number of detected genes from hundreds to the entire transcriptome level in imaging-based spatial transcriptomics. By introducing domain-adaptive adaptation and spatial prior constraints, seamless data integration across different platforms, slices, and batches is achieved, enhancing the generalization ability and stability of multimodal fusion. Simultaneously, the spatial cell network constructed based on a graph attention network enables efficient propagation and identification of cell types and multi-omics features across the entire tissue.

[0022] Experimental results on large-scale public datasets (such as Xenium, CITE-seq, and IMC) demonstrate that this invention outperforms existing mainstream methods in tasks such as multi-omics information reconstruction, cell type identification, and 3D reconstruction at single-cell resolution, particularly in multimodal integration, spatial resolution enhancement, and 3D spatial reconstruction. Furthermore, this invention fully leverages the widespread availability of histological images, enabling high-resolution, multi-omics spatial information inference and 3D tissue reconstruction at low cost, even with only a small number of spatial omics slices, significantly reducing experimental costs and technical barriers. Overall, this invention provides an efficient, universal, and scalable computational solution for large-scale spatial multi-omics research and precision medicine applications at single-cell resolution.

[0023] Other beneficial effects of the embodiments of the present invention will be further described below. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the overall algorithm architecture of the STEP method for multimodal spatial omics integration and enhancement, as described in an embodiment of the present invention.

[0025] Figure 2 The STEP method, an embodiment of the present invention, demonstrates how cell types can be identified at single-cell resolution using nuclear morphology.

[0026] Figure 3 This invention demonstrates how STEP, an embodiment of the STEP method, predicts gene expression at single-cell resolution using cell nuclear morphology.

[0027] Figure 4 This invention demonstrates how STEP enables cost-effective 3D cell type identification at single-cell resolution.

[0028] Figure 5 This invention demonstrates how STEP enables cost-effective and efficient rendering of three-dimensional multi-omics maps at single-cell resolution.

[0029] Figure 6 This is the overall flowchart of the spatial omics multimodal fusion method at the single-cell level of the present invention. Detailed Implementation

[0030] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.

[0031] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of the present invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0032] Spatial omics technology aims to obtain spatial distribution information of multi-omics molecules such as gene expression, proteins, and metabolites in the in situ environment of tissues, enabling in-depth analysis of cellular spatial structure, function, and interactions. Existing spatial omics data integration methods fail to fully utilize the complementarity of multimodal data, making it difficult to balance requirements such as spatial resolution, detection throughput, and data acquisition cost. To overcome these problems, this invention proposes a single-cell resolution spatial omics multimodal fusion method (STEP). By fusing probabilistic inference and deep learning, it flexibly integrates three types of data: spatial transcriptomics, single-cell omics, and histological images, achieving high-resolution, full-tissue-scale multi-omics information reconstruction. The innovative design of STEP mainly includes: 1. Based on histological image segmentation and spatial point localization, combined with cell nuclear morphology features and spatial coordinates, constructing a correspondence between spatial detection points and single cells to accurately infer the spatial location of single cells. 2. Utilizing a probabilistic inference model to jointly model the transcriptomic expression of spatial detection points and single-cell expression, bridging the information gap between spatial point resolution and cell resolution. 3. Based on deep graph neural networks, spatial neighborhood, expression features, and morphological features are integrated to achieve global inference of cell type and multi-omics features. 4. Domain adaptation and spatial prior constraints are employed to improve data integration capabilities across slices, platforms, and batches.

[0033] See Figure 6 This invention provides a method for multimodal fusion of spatial omics at the single-cell level, comprising the following steps:

[0034] Step S1. Preprocess the input spatial transcriptome data, single-cell sequencing data, and histological images.

[0035] In some embodiments, in step S1, the spatial transcriptome data and single-cell sequencing data are normalized to convert expression levels into a uniform counting matrix; color normalization and noise suppression are performed on the histological images to eliminate technical variations and batch effects among multi-source data.

[0036] Step S2. Perform cell segmentation on the histological image and extract the spatial location and morphological features of each cell.

[0037] Step S4. Perform domain-adaptive processing on spatial transcriptome data and single-cell data to eliminate inter-platform differences.

[0038] In some embodiments, in step S4, spatial transcriptome data and single-cell data are mapped to a unified latent feature space by a conditional variational autoencoder, and domain adaptation is performed by combining batch conditional information to eliminate differences in expression distribution between platforms.

[0039] Step S5. Based on the spatial relationship between spatial points and cell locations, establish the correspondence between spatial points and cells.

[0040] In some embodiments, in step S5, based on the spatial coordinates of the cell nucleus and the location of the spatial point, each spatial point is matched with the set of cells it contains through distance threshold and area overlap analysis, and the area ratio of each cell within the spatial point is calculated as the weight basis for the cell expression contribution in the subsequent probabilistic inference model.

[0041] Step S6. Construct a probabilistic inference model that integrates spatial transcriptome expression, single-cell data, and morphological features to jointly infer the type and gene expression level of each cell.

[0042] In some embodiments, in step S6, the probabilistic inference model models the observed expression level of a spatial point as a weighted superposition of the expression levels of multiple cells, wherein the expression level of each cell is jointly affected by its cell type-specific expression profile and morphological features, and the morphological features are weighted by learnable parameters to adjust the gene expression level, thereby achieving joint inference of cell type, gene expression and morphological features.

[0043] Step S7. Introduce a prior distribution to regularize and spatially constrain the model parameters.

[0044] In some embodiments, in step S7, the relationship between morphological features and gene expression is sparsely modeled by introducing a spike-and-slab prior, and features with significant influence are automatically screened; and the type distribution of adjacent cells is smoothed by using a Potts spatial prior, thereby improving the spatial continuity of type distribution based on the morphological similarity and spatial proximity between cells.

[0045] Step S8. Construct a spatial cell network based on a graph neural network to realize the spatial diffusion and recognition of cell type information.

[0046] In some embodiments, in step S8, the graph neural network is a graph attention network based on a multi-head attention mechanism. It constructs an undirected adjacency graph using cell spatial coordinates and morphological features, with cells as nodes and physical proximity relationships as edges. It dynamically aggregates neighbor information through a multi-layer network, assigns attention weights to different neighbors, and realizes the spatial propagation and recognition of cell type information across the entire slice.

[0047] Step S9. Based on the expression similarity of single-cell data, complete the gene expression and protein expression of spatial cells.

[0048] In some embodiments, in step S9, based on single-cell reference data, the k-nearest neighbor method is used to screen similar neighbors for spatial cells according to expression similarity, retaining only neighbors with similarity greater than zero, and weighting the expression of neighbor cells by similarity weight to complete the expression of untested genes and proteins in spatial cells.

[0049] Step S11. Integrate data from different sources and adjust the model.

[0050] In some embodiments, in step S11, feature normalization and domain adaptive adjustment are performed on tissue slices from different sources, data from different sample processing methods, and data from different spatial omics platforms to achieve cross-platform generalization of the model and efficient prediction of external data.

[0051] Step S12. Reconstruct and visualize single-cell multi-omics information in three-dimensional space.

[0052] In step S12, the fused, completed, and spatially corrected multimodal data are registered and reconstructed in three-dimensional space to output a high-resolution single-cell multi-omics three-dimensional spatial map, supporting spatial heterogeneity analysis and precision medicine applications.

[0053] In some embodiments, the method further includes the following steps before step S4:

[0054] Step S3. Perform differential expression analysis on single-cell sequencing data to screen information gene sets, which are used in the domain adaptive processing in step S4 and the probabilistic inference model in step S6 to focus on gene features with high cell type discrimination, thereby improving the efficiency and specificity of subsequent modeling.

[0055] In some embodiments, the method further includes the following steps before step S11:

[0056] Step S10. Spatial correction is performed on the completion results of step S9 using a spatial Gaussian kernel. Specifically, the expression results after completion are corrected by weighted averaging of the expression values ​​at neighboring points using a Gaussian kernel function based on spatial coordinates, improving the spatial continuity and biological plausibility of the expression distribution. The correction results from step S10 are used for model integration in step S11 and three-dimensional reconstruction in step S12 to further improve the spatial continuity of gene expression prediction and the biological plausibility of tissue structure.

[0057] In a preferred embodiment, a spatial omics multimodal fusion method at the single-cell level includes the following steps:

[0058] S1. Preprocess and standardize the input spatial transcriptome data, single-cell sequencing data, and histological images;

[0059] S2. Perform cell nucleus segmentation on the histological images and extract the spatial location and morphological features of each cell;

[0060] S3. Perform differential expression analysis on single-cell sequencing data to screen for informative gene sets;

[0061] S4. Utilize conditional variational autoencoders to perform domain-adaptive processing on spatial transcriptome data and single-cell data, eliminating differences between sequencing platforms;

[0062] S5. Based on the spatial relationship between spatial points and cell nucleus locations, calculate the area percentage of each cell within a spatial point and establish the correspondence between spatial points and cells;

[0063] S6. Construct a probabilistic inference model based on Poisson distribution, integrating spatial transcriptome expression, single-cell data and morphological features to infer the type and gene expression level of each cell;

[0064] S7. Introduce spike-and-slab priors and Potts spatial priors to regularize the model parameters and impose spatial continuity constraints.

[0065] S8. Construct a spatial cell network based on graph attention network to achieve spatial diffusion of cell type information and whole-slice recognition;

[0066] S9. Based on the expression similarity of single-cell data, multi-omics enhancement is performed on spatial cells to complete the expression of untested genes and proteins.

[0067] S10. Spatial correction is performed on the completion results using spatial Gaussian kernel to improve the consistency and biological rationality of the predictions;

[0068] S11. Normalize and adaptively adjust the model for data from different slices, processing methods, and platforms to achieve generalized integration;

[0069] S12. After fusing multimodal data, high-resolution reconstruction and visualization of single-cell multi-omics information at a three-dimensional spatial scale are achieved.

[0070] In some embodiments, step S4 specifically includes: applying a Conditional Variational Autoencoder (CVAE) to both spatial transcriptome data and single-cell sequencing data, mapping the two types of data to a unified latent feature space. Specifically, the spatial transcriptome expression matrix and the single-cell expression matrix are first encoded as latent variables, where the latent representation of each sample not only reflects its native expression characteristics but also incorporates conditional information such as batch origin. By maximizing the reconstruction probability and minimizing the KL divergence loss, the model can capture common features between different data sources and effectively eliminate potential sequencing differences such as sequencing platforms. Finally, the expression data mapped to the same feature space provides a unified basis for subsequent probabilistic inference and integration analysis.

[0071] In some embodiments, step S6 specifically includes: constructing a probabilistic inference model that integrates multi-omics expression, cell nuclear morphology, and other multimodal features based on the relationship between spatial points and single cells. Based on the Poisson distribution, it is assumed that the gene expression level of a spatial point is jointly contributed by the expression of the cells it covers, with the contribution ratio of each cell weighted according to its area proportion, morphological characteristics, and other factors. The model models the observed expression of a spatial point as a weighted superposition of the expression of multiple cells, and uses the single-cell expression profile and morphological features as priors. Maximum likelihood estimation or Bayesian inference is used to jointly infer the type and specific gene expression level of each cell. This method can achieve spatial expression reconstruction at single-cell resolution.

[0072] In some embodiments, step S7 specifically includes: the model introduces a spike-and-slab prior to achieve automatic screening and sparse modeling of the relationship between morphological features and gene expression, automatically identifying and retaining a few key features that significantly affect gene expression, effectively improving the simplicity and prediction accuracy of the model. The model introduces a Potts spatial prior to smooth the distribution of spatially adjacent cell types. Based on spatial biology principles, it is assumed that spatially adjacent and morphologically similar cells are more likely to belong to the same type. This method enables the model to both automatically screen key features and ensure the reasonable spatial continuity of cell types, thereby improving the spatial continuity and biological rationality of the prediction results and effectively suppressing noise and abnormal classifications.

[0073] In some embodiments, step S8 specifically includes: constructing a spatial cell adjacency graph with each cell as a node, where node attributes are morphological features, and edge establishment is based on spatial distance using a k-nearest neighbor strategy. A spatial cell network based on a graph attention network (GAT) is designed to process graph-structured data. It dynamically aggregates neighborhood information through multi-layer networks, assigning different weights to different neighbors, without requiring a pre-defined graph structure, making it suitable for semi-supervised learning scenarios. In this invention, the graph attention mechanism is used to propagate the type information of identified cells from within a spatial point to unidentified cells outside that point, achieving efficient diffusion of spatial information and morphological features, and single-cell type identification across the entire slice.

[0074] In some embodiments, step S9 specifically includes: screening all dissociated single cells centered on the spatial cell. Cells with the most similar expression patterns were identified as nearest neighbors, and cosine distance was used to measure expression similarity, retaining only cells with positive cosine similarity. Subsequently, a weighted average of the expression of nearest neighbor cells was applied based on similarity weights to complete the expression of all target genes in spatial cells. This strategy effectively enhances the accuracy of spatial gene expression prediction and improves the biological interpretability of spatial multi-omics data.

[0075] The Adaptive Multimodal Spatial Omics Integration and Enhancement Method (STEP) of this invention addresses the shortcomings of existing spatial omics methods in terms of resolution, omics coverage, and multimodal fusion by incorporating innovative technologies such as probabilistic inference and deep learning. It offers an efficient, flexible, and highly adaptable overall solution. This method fully integrates multimodal data such as spatial transcriptomics, single-cell sequencing, and histological images. It utilizes a Conditional Variational Autoencoder (CVAE) to eliminate plateau effects and accurately depicts tissue structures through cell nuclear segmentation and morphological feature extraction. In the cell identification and gene expression prediction stages, STEP innovatively introduces a spike-and-slab prior to achieve sparse modeling of morphological features and combines it with a Potts spatial prior to achieve smooth constraints on spatially adjacent cell types, thereby improving the spatial continuity of cell identification and the accuracy of gene expression prediction. Simultaneously, STEP employs a Graph Attention Network (GAT) for deep modeling of cell spatial relationships, dynamically integrating spatial and morphological information through a multi-head self-attention mechanism to achieve accurate cell type identification and diffusion across the entire tissue scale. In terms of multi-omics enhancement, STEP employs weighted regression based on expression similarity and spatial correction strategies to accurately complete the expression of undetected genes and proteins in spatial cells, ensuring both expression consistency and spatial coherence. Experimental results show that STEP demonstrates excellent generalization ability and accuracy on multiple mainstream spatial omics platforms and multimodal datasets, enabling the low-cost and efficient reconstruction of high-resolution, whole-transcriptome, multi-omics 3D spatial cell atlases. Overall, STEP, with its multimodal collaborative integration, dynamic spatial reasoning, and efficient multi-omics enhancement capabilities, significantly improves the resolution, coverage, and biological interpretability of spatial omics data, providing a powerful, economical, and practical technical solution for space biology and precision medicine research.

[0076] The following further describes specific embodiments of the present invention, algorithm examples, and experimental verification.

[0077] STEP is a hybrid computational framework combining probabilistic inference and deep learning. Its core design integrates cell morphology information with spatial omics data, enabling cross-scale resolution from point data to single-cell data. This platform-independent framework is applicable to both sequencing-based and imaging-based spatial transcriptomics data, and supports integration with single-cell omics, multi-omics, and histology. In sequencing-based spatial transcriptomics, STEP can improve resolution from point data to single-cell data through precise cell identification; in imaging-based spatial transcriptomics, it can expand gene expression coverage from hundreds of genes to the entire transcriptome.

[0078] The STEP method, based on multimodal spatial omics integration, firstly, in the cell identification stage, integrates spatial transcriptomics, single-cell data, and histological image multimodal data through probabilistic modeling to achieve single-cell type identification and gene expression prediction. Secondly, it utilizes graph attention networks to fully integrate intercellular spatial and morphological information, completing the spatial diffusion of cell types and expression inference across the entire slice. Finally, in the multi-omics enhancement module, it combines expression similarity and spatial neighborhood features to perform high-precision completion of undetected gene and protein expressions, improving the accuracy and spatial consistency of multi-omics prediction. The overall STEP process is as follows: Figure 1 As shown below. A detailed introduction to each core module will follow.

[0079] Data preprocessing: First, the QuPath plugin, integrating the StarDist model, was used to segment H&E-stained tissue sections, automatically extracting the spatial location and morphological features of each cell. To improve processing efficiency, large images were divided into 2000×2000 pixel blocks with a 100-pixel overlap to prevent cell fragmentation. Within the overlap area, duplicate cells were removed based on cell size and location, retaining only the largest cells for subsequent analysis.

[0080] Secondly, to eliminate batch effects between single-cell and spatial transcriptome data, a conditional variational autoencoder (CVAE) is used to achieve neighborhood alignment. First, the average number of cells per spatial point is determined based on the nuclear segmentation results. Single-cell expression is then randomly selected and weighted to generate pseudo-spatial point data. The CVAE model, through its encoder and decoder structure, maps data from different sources to a unified latent space, achieving consistency in the distribution of the two types of data and correction of expression levels.

[0081] Cell identification: The STEP method integrates spatial transcriptome data and single-cell data as input. For single-cell data, the following is first selected... Genes specifically expressed in different cell types, constructing Cell type-specific expression matrix ,in Indicates cell type Zhonggen The average expression level. For spatial transcriptome data, the average expression level of each spatial point. genes The expression value after CVAE transformation is denoted as Furthermore, the definition for Morphological feature matrix, Represents cells In features The morphological features are standardized quantitative values, including cell shape, color indices, and other quantifiable and interpretable attributes. Based on this, a Poisson distribution-based inference model is constructed to identify single-cell types at spatial points.

[0082] (1)

[0083] in Representing a spatial point The total number of unique molecular identifiers (UMIs) in the database. Indicates gene At a point in space The relative expression level, Single cells obtained through cell segmentation At point The area ratio in For cells In genes The relative expression value on [the surface]. This invention assumes a single cell [cell]. For spatial points The contribution of expression is proportional to its area proportion at that point. It is Gaussian noise. For spatial points The specific offset is used to capture genetic differences between different spatial points.

[0084] In cell identification, many methods (such as RCTD and SpatialScope) assume that all cells of the same type have the same expression characteristics, i.e. However, in practice, even among cells of the same type, their expression is often affected by morphological differences. Therefore, this invention introduces... to replace Modeling:

[0085] (2)

[0086] in This is a one-hot matrix, representing cells. The type, if it belongs to type ,but It is 1 if it is true, otherwise it is 0. These are learnable parameters used to represent morphological features. On genes The effect of expression logarithm. Considering that typically only a few morphological features significantly affect specific gene expression, and that some features may be highly correlated, to avoid overfitting and multicollinearity, the following considerations are taken: Introducing spike-and-slab sparse priors:

[0087] (3)

[0088] when hour, Follows a normal distribution, representing characteristics On genes The expression has an influence; conversely, and If so, it means that the two are unrelated.

[0089] Many spatial transcriptome deconstruction methods (such as CARD and SpatialScope) have found that spatially adjacent cells tend to belong to the same type. Furthermore, in some nuclear segmentation tasks, cell type can even be determined solely based on nuclear morphological characteristics. Therefore, this invention introduces the Potts model as a prior hypothesis: the more morphologically similar and geographically closer two cells are, the higher the probability that they belong to the same type. Type Its prior probability is:

[0090] (4)

[0091] in The normalization constant is The smoothing parameter controls the strength of spatial consistency. For cells The neighborhood group, For indicator functions, For cells and Cosine similarity in morphological features:

[0092] (5)

[0093] Spatial diffusion: This module utilizes graph structures to model the spatial relationships between cells. It can automatically learn the weight allocation between neighboring cells without prior graph structure information, making it particularly suitable for semi-supervised learning tasks. In this way, identified cell types can be diffused from cells within a spatial point to unidentified cells outside that spatial point, achieving cell identification across the entire slice.

[0094] Specifically, an undirected graph is first constructed based on the spatial coordinates of the cells, where each node represents a cell. If two cells are physically adjacent, an edge is established between them. This graph is... adjacency matrix It means that if the cell It is a cell of One of the neighbors, Otherwise, it is 0. Subsequently, the morphological feature matrix of the cells is used. and adjacency matrix Construct a Graph Attention Network (GAT) model with a multi-head attention mechanism, denoted as... :

[0095] (6)

[0096] in The index of the graph attention layer is the initial input. That is, using morphological features as the initial node representation. The output of each layer... Composed of the feature vectors of all cells, denoted as Attention weights The result is obtained after softmax normalization, and the specific formula is as follows:

[0097] (7)

[0098] in, For the first The weight matrix of the layer, Indicates feature splicing, For cells The set of neighbors. Based on this attention mechanism, the feature update rule for each layer can be expressed as:

[0099] (8)

[0100] The GAT model in STEP consists of four stacked layers, starting from the initial input. spread to The final layer is used to output the predicted type for each cell.

[0101] Gene enhancement: A gene enhancement module is introduced to complete the expression of undetected genes and proteins. In the gene and protein enhancement section, the STEP framework can predict gene and protein expression at the single-cell level; however, because only informative genes are selected in the preprocessing stage, and only a very small number of genes are included in imaging-based spatial omics, The matrix contains only a subset of spatial molecules. Before cell type inference, a conditional variational autoencoder (CVAE) model is first used to eliminate the differences between the spatial transcriptome and the single-cell data platform, making... The distribution is consistent with the single-cell data, providing a reliable basis for subsequent completion. The original single-cell data contains... Based on the data of individual genes and proteins, STEP predicts the expression of all spatial molecules.

[0102] In the gene mapping stage, for any cell in space Using cosine similarity, from Searching for its in dissociated cells in a single cell The nearest neighbor cells Then calculate the cells. Its nearest neighbor cell Weights between Only neighbors with a cosine similarity greater than 0 are retained, and normalization is performed in the following way:

[0103] (9)

[0104] Space cells Untested genes or proteins Expression prediction value It is obtained by weighted average of its nearest neighbor dissociated cells:

[0105] (10)

[0106] in For cells in single-cell data In genes or proteins The amount of expression on.

[0107] In the spatial mapping stage, considering that spatially neighboring cells usually have similar types and expression patterns, spatial correction is performed on the above prediction results. First, the cells are identified... Adjacent spatial points Construct a Gaussian kernel function based on its coordinates:

[0108] (11)

[0109] in The square of the Euclidean distance between the coordinates. The scale parameter, used to represent the diffusion range, is fixed at 0.1. Subsequently, a Gaussian kernel function is used to represent the original space. We perform weighting to obtain the correction factor:

[0110] (12)

[0111] Finally, spatial correction factor As a scaling factor, it is introduced into the predicted expression results to complete the correction:

[0112] (13)

[0113] experiment:

[0114] Dataset and evaluation metrics:

[0115] Dataset: To evaluate the performance of STEP in cell recognition and gene enhancement, a benchmark experiment was conducted based on the human pancreatic ductal adenocarcinoma (PDAC) dataset. The PDAC dataset has been extensively studied for its relevance to understanding cellular heterogeneity and the tumor microenvironment in pancreatic cancer. This dataset contains two slices: PDAC-A and PDAC-B. Slice PDAC-A corresponds to patient ID GSM3036911 and contains 19,738 genes and 428 spatial points; slice PDAC-B corresponds to patient ID GSM3405534 and contains 19,738 genes and 224 spatial points. In addition, two single-cell sequencing datasets from the same study were used as references: the PDAC-A-inDrop dataset from patient GSM3036911, containing 19,738 genes and 1,926 cells classified into 20 cell types in the original study; and the PDAC-B-inDrop dataset from patient GSM3405534, containing 19,738 genes and 1,733 cells classified into 13 cell types. These datasets cover rich information from the same tissue at both the spatial and single-cell levels, providing a comprehensive resource for evaluating the present invention.

[0116] To demonstrate that STEP can reconstruct high-resolution 3D multi-omics brain cell atlases using minimal low-resolution spatial transcriptome slices, this method integrates four datasets: spatial transcriptomics, histological images, single-cell sequencing, and inCITE-seq. Two key datasets were used in the experiments: an adult mouse brain molecular atlas and single-cell sequencing. The former, from the GEO database (GSE147747), contains 35 of 75 coronal slices, revealing the subregional divisions of the brain. The latter, from the EMBL-EBI database (E-MTAB-11115), contains cell nuclei from multiple regions of the mouse brain, providing detailed cell type references for this invention.

[0117] Evaluation indicators:

[0118] Accuracy (ACC): Accuracy is a direct metric that represents the proportion of correct predictions out of the total number of predictions. For multi-class classification problems, the formula is as follows:

[0119] (14)

[0120] Accuracy measures the overall correctness of a model's predictions. A higher ACC value indicates higher model accuracy.

[0121] Adjusted Rand Index (ARI): The Adjusted Rand Index (ARI) is a measure of similarity between two clustering results, particularly useful for comparing clustering results with known reference clusters. This index improves the Rand Index (RI) by correcting for randomness, and its score ranges from -1 to 1. Where 1 represents perfect consistency (clustering results are completely identical), 0 represents random assignment (no significant consistency), and negative values ​​indicate consistency below the level of randomness. Its calculation formula is as follows:

[0122] (15)

[0123] in, This indicates that they both belong to the same cluster in the first clustering. and clusters in the second cluster Number of points Indicates the cluster in the first clustering The number of points contained, and This indicates that the cluster in the second clustering is... The number of points contained.

[0124] F1-Score. The macro-averaging method calculates the F1-score for each category separately and then averages them. This method is more sensitive to the performance of individual categories and is particularly suitable for situations where all categories are equally important. Its definition is as follows:

[0125] (16)

[0126] in, Indicates the total number of categories. Indicates the first The F1 score for each category.

[0127] Normalized Mutual Information (NMI): NMI is a metric for measuring the similarity between two cluster assignments. It measures similarity by quantifying the amount of information shared between them and eliminates the influence of differences in cluster size through normalization. Its definition is as follows:

[0128] (17)

[0129] in, and Representing cluster assignment (Predictive clustering) and The entropy of (true clustering) is used to measure the uncertainty (or information content) of each cluster. for and Mutual information between them, used to characterize the amount of information shared between them, takes the form of:

[0130] (18)

[0131] in, This indicates that a data point simultaneously belongs to medium cluster and medium cluster The joint probability, and These respectively indicate that the data points belong to medium cluster and medium cluster The marginal probability. The value of NMI ranges from 0 to 1, and the larger the value, the higher the similarity between the clustering results.

[0132] Purity: Purity is calculated by assigning each cluster to the most frequently occurring class of data points within it. Purity is defined as the proportion of correctly assigned data points out of the total number of data points.

[0133] (19)

[0134] in, Indicates the total number of data points. Indicates the number of clusters. Indicates the first The set of points in a cluster Indicates the first The set of points in each category. The Purity value ranges from 0 to 1, with higher values ​​indicating better clustering results. However, it's important to note that Purity tends to favor larger clusters and does not penalize a large number of small clusters.

[0135] Spatial Accuracy (sACC): The sACC metric calculates accuracy by simultaneously considering both the exact match between the predicted label and the true label, and the correct label within a specified neighborhood.

[0136] (20)

[0137] in, Represents the true label of the data point. Indicates the predicted label, Indicates passage Points calculated by the nearest neighbor method The set of neighbor indexes. The value of sACC ranges from 0 to 1, with a value of 1 indicating perfect classification accuracy, meaning that all points are either correctly classified directly or correctly within their neighborhood.

[0138] Pearson correlation (PCC): PCC measures the linear correlation between predicted and actual results, and is defined as follows:

[0139] (twenty one)

[0140] in, Indicating the first in the true result Gene expression per spot / cell Indicating the first in the true result Average gene expression per spot / cell Indicating the first in the true result Standard deviation of gene expression per spot / cell; , and These represent the corresponding values ​​in the prediction results. For a given gene, a higher PCC value indicates a more accurate expression imputation result for that gene.

[0141] Results analysis:

[0142] To evaluate STEP's performance in cell identification and gene imputation, benchmark tests were first conducted using the human pancreatic ductal adenocarcinoma (PDAC) dataset. STEP was compared with six state-of-the-art methods, including CARD, SpaDecon, CellDART, Cell2location, STIE, and SpatialScope. SpaDecon and CellDART are deep learning methods, CARD and Cell2location are based on decomposition and probabilistic modeling, while STIE and SpatialScope are capable of inferring cell type at the single-cell level.

[0143] In PDAC tissues, studies have shown that certain cell types are primarily enriched in specific regions; for example, acinar cells are mainly distributed in the pancreatic region, while cancer cells are concentrated in the cancerous area. Based on this phenomenon, various assessment indicators were designed to compare different methods. Using a label-free assessment system, STEP showed at least a 15.2% improvement over dissociation methods in indicators such as ARI and F1-Score, as shown in Table 1.

[0144] Table 1

[0145]

[0146] Compared to the two methods capable of single-cell recognition, the F1-score was improved by 34.0%, ARI by 35.0%, ACC by 24.0%, and NMI by 33.8%, as shown in Table 2:

[0147] Table 2

[0148]

[0149] While the Purity metric is biased against methods that cover the entire slide, STEP still maintains its lead. It's important to note that the Purity metric is easily affected by variations in cell number, but STEP still outperforms the control method in this regard.

[0150] In qualitative analysis, such as Figure 2 STEP provides a clearer spatial distribution prediction for acinar cells and accurately characterizes the distribution of rare ductal hyperoxygenated cells and cancer clonal subregions in cancerous areas, phenomena that other methods often miss. It also exhibits strong robustness and is less susceptible to experimental noise. Morphological analysis shows that cancer cells are significantly different from other cells in terms of nuclear size, chromatin distribution, and nucleocell ratio, while the characteristics of rare ductal hyperoxygenated cells are similar to those of cancer cells, consistent with biological expectations.

[0151] In gene expression prediction, such as Figure 3 As shown, STEP provides fine-grained single-cell expression results, outperforming BayesSpace, SpatialScope, iStar, and XFuse. It enables full-slice-area spatial marker mapping, revealing heterogeneity between tumor and stromal regions, as well as between different cancer subregions. Further quantitative analysis indicates that specific marker genes are correctly enriched in the expected regions, validating STEP's accuracy and resolution-enhancing capabilities in capturing spatial gene expression heterogeneity.

[0152] To achieve three-dimensional multi-omics reconstruction using limited slides, a four-modal integration of spatial transcriptomics, histological images, single-cell sequencing data, and inCITE-seq was conducted in mouse brains. Based on 35 coronal slides along the anterior-posterior axis and 59 types of cell reference data, STEP was able to reconstruct high-resolution three-dimensional gene expression maps using only 6 ST slides. The predictions reconstructed by STEP under low-cost conditions were highly consistent with those using all 35 slides, and the spatial accuracy (sACC) remained at 87.7% with 6 ( / 35) slides. Accuracy further improved with increasing slide number. Figure 4 As shown, the visualization results demonstrate that STEP can accurately locate subcortical structures such as the hippocampus and striatum. In the hippocampus, the original ST data, limited by a resolution of approximately 100 µm, struggles to resolve fine structures such as CA and DG layers. In contrast, the three-dimensional gene expression maps reconstructed by STEP at single-cell resolution exhibit clear subregion boundaries and show a high degree of agreement with Allen Brain Atlas. Figure 5 As shown. Further combining inCITE-seq reference data, STEP achieved spatial multi-omics mapping of genes and proteins, such as... Figure 5As shown, the three-dimensional distributions of NeuN and c-Fos are consistent with existing studies, revealing neuronal distribution and the dynamics of neural circuit activity. These results demonstrate that STEP not only overcomes the resolution limitations of traditional ST methods to achieve low-cost three-dimensional reconstruction, but also exhibits strong potential in multi-omics integration and expansion.

[0153] This invention also provides a storage medium for storing a computer program, which, when executed, performs at least the methods described above.

[0154] This invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein the processor executes the computer program by performing at least the method described above.

[0155] This invention also provides a processor that executes a computer program, at least performing the methods described above.

[0156] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc or CD-ROM; magnetic surface memory can be disk storage or magnetic tape storage. The storage media described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0157] In the several embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0158] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0159] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0160] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0161] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0162] The methods disclosed in the several method embodiments provided by this invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0163] The features disclosed in the several product embodiments provided by this invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0164] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0165] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various equivalent substitutions or obvious modifications can be made without departing from the concept of the present invention, and all such modifications, achieving the same performance or application, should be considered within the scope of protection of the present invention.

Claims

1. A method for multimodal fusion of spatial omics at the single-cell level, characterized in that, Includes the following steps: S1. Preprocess the input spatial transcriptome data, single-cell sequencing data, and histological images; S2. Perform cell segmentation on the histological image and extract the spatial location and morphological features of each cell; S4. Perform domain-adaptive processing on spatial transcriptome data and single-cell data to eliminate inter-platform differences; S5. Based on the spatial relationship between spatial points and cell locations, establish the correspondence between spatial points and cells; among them, based on the spatial coordinates of the cell nucleus and the location of the spatial point, calculate the area ratio of each cell within the spatial point, which serves as the weight basis for the cell expression contribution in the subsequent probabilistic inference model; S6. Construct a probabilistic inference model that integrates spatial transcriptome expression, single-cell data, and morphological features to jointly infer the type and gene expression level of each cell; wherein, the probabilistic inference model models the observed expression level at a spatial point as a weighted superposition of the expression levels of multiple cells, wherein the expression level of each cell is jointly affected by its cell type-specific expression profile and morphological features, and the morphological features are weighted and adjusted for gene expression level through learnable parameters, thereby achieving joint inference of cell type, gene expression, and morphological features; S7. Introduce a prior distribution to regularize and spatially constrain the model parameters; S8. Construct a spatial cell network based on a graph neural network to realize the spatial diffusion and recognition of cell type information; wherein, the graph neural network is a graph attention network based on a multi-head attention mechanism, which constructs an undirected adjacency graph through cell spatial coordinates and morphological features, with cells as nodes and physical proximity relationships as edges, dynamically aggregates neighbor information through a multi-layer network, assigns attention weights to different neighbors, and realizes the spatial propagation and recognition of cell type information within the entire slice range; S9. Based on the expression similarity of single-cell data, gene and protein expression of spatial cells are completed; among them, based on single-cell reference data, similar neighbors of spatial cells are screened according to expression similarity, only neighbors with similarity greater than zero are retained, and the expression of neighbor cells is weighted by similarity weight to complete the expression of untested genes and proteins in spatial cells. S11. Integrate data from different sources and adjust the model; S12. Achieve reconstruction and visualization of single-cell multi-omics information at a three-dimensional spatial scale.

2. The method for multimodal fusion of spatial omics at the single-cell level as described in claim 1, characterized in that, In step S1, the spatial transcriptome data and single-cell sequencing data are normalized to convert expression levels into a uniform counting matrix; color normalization and noise suppression are performed on histological images to eliminate technical variations and batch effects among multi-source data.

3. The method for multimodal fusion of spatial omics at the single-cell level as described in claim 1, characterized in that, In step S4, spatial transcriptome data and single-cell data are mapped to a unified latent feature space by a conditional variational autoencoder, and domain adaptation is performed by combining batch conditional information to eliminate differences in expression distribution between platforms.

4. The method for multimodal fusion of spatial omics at the single-cell level as described in claim 1, characterized in that, In step S5, each spatial point is matched with the set of cells it contains by means of distance threshold and area overlap analysis.

5. The method for multimodal fusion of spatial omics at the single-cell level as described in claim 1, characterized in that, In step S7, the relationship between morphological features and gene expression is sparsely modeled by introducing a spike-and-slab prior, and features with significant influence are automatically screened. Furthermore, the type distribution of adjacent cells is smoothed by using a Potts spatial prior, and the spatial continuity of type distribution is improved based on the morphological similarity and spatial proximity between cells.

6. The method for multimodal fusion of spatial omics at the single-cell level as described in claim 1, characterized in that, In step S9, the k-nearest neighbor method is used to screen similar neighbors for spatial cells based on expression similarity.

7. The method for multimodal fusion of spatial omics at the single-cell level as described in claim 1, characterized in that, In step S11, feature normalization and domain adaptive adjustment are performed on tissue slices from different sources, data from different sample processing methods, and data from different spatial omics platforms to achieve cross-platform generalization of the model and efficient prediction of external data. In step S12, the fused, completed, and spatially corrected multimodal data are registered and reconstructed in three-dimensional space to output a high-resolution single-cell multi-omics three-dimensional spatial map, supporting spatial heterogeneity analysis and precision medicine applications.

8. The method for multimodal fusion of spatial omics at the single-cell level as described in any one of claims 1 to 7, characterized in that, It also includes one or both of the following steps in sequence: S3. Perform differential expression analysis on single-cell sequencing data to screen information gene sets, which are then used in the domain adaptive processing in step S4 and the probabilistic inference model in step S6 to focus on gene features with high cell type discrimination, thereby improving the efficiency and specificity of subsequent modeling. S10. Spatial correction is performed on the completion result of step S9 using a spatial Gaussian kernel. Specifically, the expression result after completion is corrected by weighted averaging of the expression values ​​of neighboring points based on the spatial coordinates using a Gaussian kernel function, thereby improving the spatial continuity and biological rationality of the expression distribution. The correction results are used for model integration in step S11 and three-dimensional reconstruction in step S12 to improve the spatial continuity of gene expression prediction and the biological rationality of tissue structure.

Citation Information

Patent Citations

  • Method for identifying cross-modal features from spatially resolved datasets

    CN118176527A

  • Space omics data completion method and system based on variational graph auto-encoder

    CN120783869A