Single cell and space transcriptome data integrated multi-omics analysis system

By employing a strategy of dual feature anchoring and probabilistic graphical model alignment, the problems of insufficient mapping accuracy and fragmented analysis workflow in the integration of single-cell and spatial transcriptome data were solved, achieving high-confidence spatial mapping of cell states and integrated analysis, thereby enhancing biological discovery capabilities.

CN121983136APending Publication Date: 2026-05-05SHANXI MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANXI MEDICAL UNIV
Filing Date
2026-01-21
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies, when integrating single-cell and spatial transcriptome data, suffer from insufficient mapping accuracy, fragmented analysis processes, reliance on experience for feature selection, and failure to fully utilize the expression patterns and topological structures of spatial data, resulting in information loss and inconsistent conclusions.

Method used

Employing a strategy of dual feature anchoring and probabilistic graphical model alignment, this approach achieves high-confidence cell state space mapping and integrated multi-omics analysis through input processing, data integration and alignment, joint analysis, and visualization output modules. It combines differentially expressed genes with covariation relationships to construct a shared feature set and utilizes a probabilistic graphical model for spatial alignment.

Benefits of technology

It enables precise localization of cell types and quantification of their spatial distribution, quantifies intercellular communication events, reveals local microenvironment regulatory mechanisms, enhances biological discovery capabilities, and avoids information loss and contradictory conclusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983136A_ABST
    Figure CN121983136A_ABST
Patent Text Reader

Abstract

The invention provides a single cell and spatial transcriptome data integrated multi-omics analysis system, which relates to the technical field of bioinformatics and computational biology, and comprises an input processing module, a data integration and alignment module, a conjoint analysis module and a visual output module, according to the method, a shared feature set with clear biological significance is constructed by introducing a dual feature anchoring strategy and combining a co-variant relationship between a differential expression gene and a conservative gene, a high-confidence constraint basis is provided for subsequent integration, spatial alignment is further performed by adopting a probability graph model, accurate positioning of cell types is realized, and the accuracy of the cell types is improved. And the uncertainty of spatial distribution is quantified, so that the mapping result breaks through the limitation of traditional hard distribution, the biological reality that cell states are mixed in a tissue microenvironment is better fitted, and the whole process height of data correction, feature anchoring and probability mapping to spatial communication and co-expression network analysis is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and computational biology, and in particular to a multi-omics analysis system that integrates single-cell and spatial transcriptome data. Background Technology

[0002] With the rapid development of single-cell sequencing and spatial transcriptomics technologies, life science research has entered a new stage capable of simultaneously analyzing cellular heterogeneity and its spatial organizational patterns. Single-cell transcriptomics provides gene expression profiles at the resolution of a single cell, but loses information about the spatial location of cells within the original tissue. Spatial transcriptomics preserves the spatial coordinates of gene expression, but its resolution is typically a "spot" containing multiple cells, and the cell type composition is unknown.

[0003] Currently, the mainstream methods for integrating these two approaches fall into two main categories: one is pairwise label transfer methods, which "map" or "deconvolve" the cell type labels annotated in single-cell data onto spatial blobs; the other is low-dimensional spatial alignment methods based on co-embedding. However, these methods have the following limitations:

[0004] (1) Insufficient mapping accuracy: Most methods perform "hard" allocation or simple linear decomposition, failing to make full use of the spatial data's own expression pattern and topological structure as constraints, and ignoring the spatial continuity of cell states.

[0005] (2) Fragmented analysis process: Integration and downstream analysis (such as cell communication and co-expression module identification) are usually separate steps, which do not form a closed loop and may lead to information loss and inconsistent conclusions.

[0006] (3) Feature selection depends on experience: The gene feature (anchor) selection strategy relied upon in the integration process is relatively simple, usually based only on hypervariable genes or marker genes, and fails to take advantage of the stronger constraint of the consistency of gene relationships in spatial and single-cell dimensions. Therefore, this invention proposes a multi-omics analysis system for integrating single-cell and spatial transcriptome data to solve the problems existing in the prior art. Summary of the Invention

[0007] To address the aforementioned problems, the present invention aims to propose a multi-omics analysis system that integrates single-cell and spatial transcriptome data. This invention achieves high-confidence spatial mapping of cell states by introducing a strategy of dual feature anchoring and probabilistic graphical model alignment, and performs integrated spatial multi-omics joint analysis on this basis, ultimately generating interactively verifiable biological insights, thereby solving the problems in the prior art.

[0008] To achieve the objectives of this invention, the invention is implemented through the following technical solution: a multi-omics analysis system integrating single-cell and spatial transcriptome data, comprising:

[0009] The input processing module receives single-cell transcriptome data and spatial transcriptome data, performs standardization and cross-platform batch effect correction on them respectively, and outputs the corrected single-cell gene expression data and spatial gene expression data.

[0010] The data integration and alignment module is used to receive single-cell gene expression data and spatial gene expression data. Then, based on the cell type information inferred from the single-cell data, it uses a probabilistic graphical model to associate the cell type information with the spatial coordinates of the spatial gene expression data and outputs a spatial cell type composition probability matrix.

[0011] The joint analysis module receives the spatial cell type composition probability matrix, then performs spatial resolution intercellular communication inference and multi-omics co-expression network analysis, and outputs the analysis results.

[0012] The visualization output module receives the analysis results and transforms them into interactive spatial visualization charts and structured analysis reports.

[0013] A further improvement is that the input processing module includes a cross-platform correction unit, which is used to identify the housekeeping gene set common to single-cell and spatial transcriptome data, and apply a batch effect correction algorithm based on a linear model to make the gene expression profiles of the two sets of data comparable.

[0014] A further improvement is that the data integration and alignment module includes a dual feature anchoring unit, which is used to perform a first anchoring analysis and a second anchoring analysis to obtain a first feature set and a second feature set respectively. Then, the first feature set and the second feature set are combined to form a shared feature set. The first anchoring analysis is based on single-cell gene expression data to identify cell type-specific differentially expressed genes and on spatial gene expression data to identify spatially differentially expressed genes. The intersection of the two is taken to form the first feature set.

[0015] The second anchoring analysis involves screening gene pairs with stable cross-cell type correlations from single-cell gene expression data and verifying their consistency in correlations in spatial gene expression data. Consistent gene pairs are retained to form the second feature set.

[0016] A further improvement is that the data integration and alignment module also includes a probabilistic graphical model alignment unit, which uses the expression data of genes in the shared feature set as the observation input, and the cell type label and its expression characteristics inferred from the single cell data as the hidden state prior to construct a conditional random field model with spatial points as nodes. The model infers and calculates the probability that each spatial point belongs to each cell type, thereby generating the spatial cell type composition probability matrix.

[0017] A further improvement is that the joint analysis module includes a spatial cell communication inference unit, which is used to quantify the ligand-receptor interaction strength between different cell types in a local spatial region by assigning cell type origin weights to the ligand and receptor expression levels of each spatial point based on the probability matrix of spatial cell types and the proximity relationship network of spatial points.

[0018] Further improvements are made in that: the joint analysis module includes a multi-omics co-expression analysis unit, which is used to merge single-cell gene expression data and spatial gene expression data based on a shared feature set, perform weighted gene co-expression network analysis on the merged data, identify gene modules conserved in spatial and cell type dimensions, and finally associate gene modules with their spatial activity distribution and biological functions.

[0019] A further improvement is that the visualization output module outputs an interactive spatial distribution map, which can display at least one of the following information layers in a linked manner: spatial cell type composition probability, spatial ligand-receptor interaction strength and direction, and the active hotspot region of the spatial co-expression module.

[0020] A further improvement is that the visualization output module includes an automated report generation unit, which integrates spatially differentially expressed genes, significant cell interaction pairs, co-expression modules and their key functional annotation results to automatically generate a structured analysis report.

[0021] The beneficial effects of this invention are as follows: By introducing a dual feature anchoring strategy and combining differentially expressed genes with conserved gene covariation relationships, this invention constructs a biologically meaningful shared feature set, providing a high-confidence constraint basis for subsequent integration. Furthermore, by employing a probabilistic graphical model for spatial alignment, it not only achieves precise localization of cell types but also quantifies the uncertainty of their spatial distribution. This allows the mapping results to break through the limitations of traditional hard allocation and better reflect the biological reality of mixed cell states in the tissue microenvironment. It ensures that the entire process of data correction, feature anchoring, probabilistic mapping to spatial communication and co-expression network analysis is highly cohesive and logically consistent, eliminating information loss or contradictory conclusions caused by process fragmentation.

[0022] This invention significantly enhances the biological discovery capabilities of the present invention, enabling the identification of novel cell subpopulations or functional units with specific spatial aggregation patterns through in-depth analysis of spatial probability matrices. Furthermore, by utilizing cell type probability weights to finely correct local expression levels, spatially constrained intercellular communication events can be more accurately quantified, revealing local microenvironment regulatory mechanisms that are difficult to capture using traditional methods. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the system structure of the present invention.

[0024] Figure 2 This is a schematic diagram of the workflow of the present invention. Detailed Implementation

[0025] To enhance understanding of the present invention, the present invention will be further described in detail below with reference to embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the scope of protection of the present invention.

[0026] Example 1

[0027] according to Figure 1 As shown, this embodiment proposes a multi-omics analysis system integrating single-cell and spatial transcriptome data, including:

[0028] The input processing module receives single-cell transcriptome data and spatial transcriptome data, performs standardization and cross-platform batch effect correction on them respectively, and outputs corrected single-cell gene expression data and spatial gene expression data. It includes a cross-platform correction unit to identify the housekeeping gene set common to the single-cell and spatial transcriptome data, and applies a batch effect correction algorithm based on a linear model to make the gene expression profiles of the two sets of data comparable.

[0029] Specifically, the input quantity module receives two types of input data:

[0030] The first type is single-cell transcriptome data, which is a gene expression matrix that has undergone preliminary processing. The rows of the matrix represent cells, the columns represent genes, and it contains cell type label information obtained through clustering and annotation analysis based on the matrix.

[0031] The second type is spatial transcriptome data, which includes gene expression matrices and their corresponding two-dimensional spatial coordinates. The rows of the matrix represent spatial locations, and the columns represent genes.

[0032] Therefore, independent quality control and standardization were performed on the two sets of data. For single-cell data, the Scanpy workflow was used, specifically including filtering out low-quality cells with an excessively high proportion of mitochondrial genes or an insufficient total number of detected genes, followed by library size standardization (e.g., per million count conversion) and logarithmic transformation. For spatial transcriptome data, the Squidpy workflow was used to filter out spatial sites with insufficient total molecular counts or insufficient number of detected genes, followed by similar library size standardization.

[0033] Next, the cross-platform calibration unit begins its work. This unit identifies the set of housekeeping genes that are stably expressed in both datasets. Subsequently, the ComBat algorithm based on a linear model is applied, using the "sequencing technology platform" as the batch variable and the identified housekeeping gene expression values ​​as a benchmark, to calibrate the two complete gene expression matrices of the single-cell and spatial transcriptomes. This eliminates systematic differences in expression levels between platforms and outputs two expression matrices at a comparable scale, namely, the calibrated single-cell gene expression data and spatial gene expression data.

[0034] The data integration and alignment module receives single-cell gene expression data and spatial gene expression data. Based on the cell type information inferred from the single-cell data, it uses a probabilistic graphical model to associate this cell type information with the spatial coordinates of the spatial gene expression data, outputting a spatial cell type composition probability matrix. It includes a dual feature anchoring unit and a probabilistic graphical model alignment unit, specifically:

[0035] The dual-feature anchoring unit performs a first anchoring analysis and a second anchoring analysis, yielding a first feature set and a second feature set, respectively. These two feature sets are then combined to form a shared feature set. For the first anchoring analysis, in single-cell data, for each annotated cell type, statistical tests (such as the Wilcoxon rank-sum test) are used to identify marker genes with specific high expression. In spatial data (referring to spatial gene expression data), spatial differential expression analysis is used to identify genes exhibiting significant expression patterns on spatial coordinates. The intersection of these two gene lists constitutes the first feature set. The genes in this set are both key features defining cell types and exhibit spatial expression patterns in tissues.

[0036] For the second anchoring analysis, in single-cell data, the expression correlation of all gene pairs within each defined cell type is calculated. Gene pairs that maintain stable correlations across multiple cell types (e.g., both are strongly positively correlated or both are strongly negatively correlated) are screened. In spatial data, the expression correlation of these candidate gene pairs is calculated across all spatial locations. Gene pairs whose correlation direction and strength in spatial data are consistent with those in single-cell data are retained, forming the second feature set. This set represents conserved gene regulation or co-expression relationships across different biological contexts (single-cell type and spatial tissue structure).

[0037] Finally, the first feature set and the second feature set are merged to form a shared feature set for subsequent model alignment. This dual strategy combines information from two dimensions: gene expression levels and inter-gene relationships, providing stronger and more robust constraints for integration.

[0038] The probabilistic graphical model alignment unit is used to construct a conditional random field model with spatial points as nodes. This model uses gene expression data from a shared feature set as observation input and cell type labels and their expression characteristics inferred from single-cell data as hidden state priors. The model infers and calculates the probability that each spatial point belongs to a particular cell type, thereby generating the spatial cell type composition probability matrix. Specifically, it includes the following steps:

[0039] Step S1: Construct a conditional random field model with spatial locations as nodes. The observation data for each node is the expression vector of all genes in the shared feature set at that location. The goal of model learning is to infer a hidden state variable for each node, i.e., the probability distribution of that spatial location belonging to each cell type;

[0040] Step S2: The model uses the typical expression patterns of each cell type on the shared feature set learned from single-cell data as prior knowledge. At the same time, the model defines two constraints: first, the observed expression spectrum of a single spatial point should match the probability distribution of the cell type it is assigned; second, spatially adjacent points tend to have similar cell type compositions, thereby introducing spatial smoothness.

[0041] Step S3: The above model is solved using a standard inference algorithm for probabilistic graphical models (such as the belief propagation algorithm). The algorithm iteratively integrates the observation data (gene expression) with prior constraints (cell type features and spatial smoothness) and finally calculates the most likely cell type composition probability for each spatial point. The output is a probability matrix, where each row corresponds to a spatial point, each column corresponds to a cell type, and the values ​​of the matrix elements represent the composition proportion.

[0042] The joint analysis module receives the spatial cell type composition probability matrix, then performs spatial resolution intercellular communication inference and multi-omics co-expression network analysis, outputting the analysis results. It includes a spatial cell communication inference unit and a multi-omics co-expression analysis unit, specifically:

[0043] The spatial cell communication inference unit integrates a known database of ligand-receptor interactions. For each pair of spatially adjacent sites in a tissue, the unit analyzes the potential cell communication between them. Specifically, for a pair of adjacent sites A and B, the unit considers the possible interactions between ligands expressed by all cell types in site A (weighted by their composition probabilities) and receptors expressed by all cell types in site B (weighted by their composition probabilities). Finally, by using cell type composition probabilities as weights, the expression levels of ligands and receptors are weighted and summed to quantify the strength of a specific ligand-receptor interaction signal from a cell type in site A to a cell type in site B. By traversing all adjacent site pairs and all possible cell type combinations and ligand-receptor pairs, a spatially resolved intercellular communication network is constructed, and cell interaction events active in specific tissue regions are identified.

[0044] The multi-omics co-expression analysis unit merges single-cell data and spatial data along the gene dimension of a shared feature set to create a fused expression matrix. A weighted gene co-expression network analysis method is then applied to this fused matrix. This method constructs a network based on the expression correlations between genes and uses hierarchical clustering to divide genes into different co-expression modules. Each module contains a set of genes with highly coordinated expression patterns. Subsequently, the biological characteristics of these gene modules are analyzed. Specifically, gene function enrichment analysis is first used to interpret the biological processes or pathways that each module may represent. Then, a "module activity score" is calculated for each module at each point in the spatial data, thereby revealing the spatial activity pattern of the co-expression program in the tissue, such as whether a specific module is specifically active at the tumor margin, around blood vessels, or in specific functional regions.

[0045] The visualization output module is used to receive the analysis results and convert them into interactive spatial visualization charts and structured analysis reports. It includes an automated report generation unit, which integrates key results of spatially differentially expressed genes, significant cell interaction pairs, co-expressed modules and their functional annotations, and automatically generates a structured analysis report. That is, the automated report generation unit generates an HTML report in parallel, which summarizes: (1) a list of the top 10 spatially differentially expressed genes; (2) the top 5 cell interaction pairs in terms of intensity and their possible functions; (3) the gene ontology enrichment analysis results of the 5 co-expressed modules; and (4) a summary of key findings, such as: "The analysis found that microglia and neurons have strong spatial interactions in the V layer of the cortex through the CX3CL1-CX3CR1 signal axis, and the region also highly expresses the ME4 (inflammation-related) module."

[0046] Specifically, the visualization output module automatically generates an interactive HTML page to display an interactive spatial distribution map, which displays at least one of the following information layers in conjunction with the map:

[0047] Spatial cell type composition probability: This is a spatial composition heatmap. Users can slide the slider to view the probability distribution heatmap of different cell types in the tissue space.

[0048] Spatial ligand-receptor interaction strength and direction: This refers to the interaction network and spatial mapping diagram. One side of the page displays the interaction network diagram between cell types (node ​​size represents the proportion of cell types, and edge thickness represents the interaction strength). Clicking on any edge (e.g., "microglia -> neurons") will highlight the spatial region with the highest interaction strength on the other side of the spatial map (e.g., a cluster of spots near the lesion area).

[0049] Activity hotspots of spatial co-expression modules: i.e. co-expression module activity map. Users can select different gene modules (ME1-ME5), and the spatial map will show the spatial distribution of the module's activity using a color gradient.

[0050] Example 2

[0051] according to Figure 2 As shown, this embodiment takes the analysis of a publicly available mouse prefrontal cortex spatial transcriptome dataset and its matching single-cell RNA-seq dataset as an example. The workflow is as follows:

[0052] Step 1: Data Input and Preprocessing

[0053] Users upload single-cell data (in h5ad format, including the X matrix and cell_type annotation column) and spatial data (in h5ad format, including the X matrix and spatial coordinate information) via the system interface. The input processing module then calls the Scanpy library to perform basic quality control on the single-cell data: filtering cells with fewer than 200 expressed genes, filtering genes expressed in fewer than 3 cells, performing library size normalization (sc.pp.normalize_total) and logarithmic transformation (sc.pp.log1p). Next, the Squidpy library is called to perform quality control on the spatial data: filtering spots with excessively low total counts or abnormal gene counts. Finally, the cross-platform correction unit is invoked, and the system automatically selects a set of approximately 100 housekeeping genes from the HKG gene library. Using the spatial data as a reference, the ComBat algorithm implemented in scikit-learn is used to perform batch effect correction on the expression profiles of the single-cell data, making the expression distribution of housekeeping genes in the two sets of data more consistent.

[0054] Step 2: Data Integration and Alignment

[0055] Two anchoring analyses are performed using dual-feature anchoring elements, wherein:

[0056] First anchoring analysis: In single-cell data, the `sc.tl.rank_genes_groups` function was used to group genes by `cell_type`, identifying the top 50 specific marker genes for each cell type (logFC > 1, p_val_adj < 0.01). In spatial data, the SpatialDE algorithm was run, identifying approximately 300 genes with significant spatial expression patterns (qval < 0.05). The intersection of these two groups of genes was calculated, yielding a first feature set containing approximately 80 genes.

[0057] The second anchoring analysis: In the single-cell data, for all gene pairs, their expression correlation (Spearman correlation coefficient) was calculated within each cell type. Gene pairs with consistent correlation directions (both positive or both negative) and a mean absolute value greater than 0.4 were selected. In the spatial data, the global correlation of these candidate gene pairs was calculated across all spatial blobs. Gene pairs with correlation directions consistent with the average direction of single-cell correlation in the spatial data were selected, ultimately yielding approximately 200 gene pairs, constituting the second feature set;

[0058] Finally, the genes in the first feature set are merged with all the unique genes in the second feature set to form the final shared feature set (approximately 250 genes).

[0059] Then, the probabilistic graphical model alignment unit begins its work, using the expression vectors of genes in the shared feature set across all spatial spots as the model's observed features, and the cell type labels of all cells in the single-cell data and their average expression spectra on the shared feature set as the prior distribution of the hidden states (i.e., the "expression template" for each cell type). A conditional random field model is then constructed, in which each spatial spot is a node, its observed value being the shared gene expression spectrum at that point, and edges are constructed based on the K nearest neighbors (K=6) of spatial coordinates. The potential function combines observation-state matching (based on a Gaussian distribution) and neighborhood state consistency (based on homogenization potential). Finally, a cyclic belief propagation algorithm is used for model inference, outputting an N_spots × N_celltypes probability matrix. For example, a spot might be assigned a probability composition of "oligodendrocyte: 0.85, neuron: 0.15". Simultaneously, a spatial proximity network based on K nearest neighbors is output.

[0060] Step 3: Joint Analysis

[0061] The obtained probability matrix and the neighboring network are input into the spatial cell communication inference unit. Then, a list of ligand-receptor pairs is loaded from the CellPhoneDB public database. For each pair of adjacent spots (A and B), the interaction potential from cell type i in spot A to cell type j in spot B is calculated. Specifically:

[0062] The "effective expression level" of ligand L in spot A = the original expression level of L in A × the probability of cell type i in spot A.

[0063] The "effective expression level" of receptor R in spot B = the original expression level of R in B × the probability of cell type j in spot B.

[0064] The contribution score of the interaction (i, j, L, R) between the adjacent points = effective expression level (L)_A × effective expression level (R)_B.

[0065] The contribution scores of all adjacent spots are summed and normalized to obtain a global, spatially weighted inter-cell type interaction strength matrix.

[0066] The corrected single-cell data matrix and spatial data matrix were then extracted, retaining only the gene columns with shared feature sets. A multi-omics co-expression analysis unit merged the two matrices row-wise (cell / spot) to obtain a (N_cells + N_spots) × N_shared_genes joint expression matrix. The WGCNA algorithm was run on this joint matrix with a soft threshold power of 6 and a minimum module gene count of 30, identifying five significant co-expressed gene modules (labeled ME1 to ME5). The activity score of the eigengene for each module was calculated at each spatial spot. Module ME2 (enriched in the "synaptic transmission" pathway) showed significantly increased activity in specific lamellar structures of the cortex; module ME4 (enriched in the "inflammatory response" pathway) was enriched in the perivascular region.

[0067] Step 4: Visualization and Report Generation

[0068] Interactive HTML pages are automatically generated by the visualization output module.

[0069] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the present invention without departing from its framework and scope of application, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-omics analysis system integrating single-cell and spatial transcriptome data, characterized in that: include: The input processing module receives single-cell transcriptome data and spatial transcriptome data, performs standardization and cross-platform batch effect correction on them respectively, and outputs the corrected single-cell gene expression data and spatial gene expression data. The data integration and alignment module is used to receive single-cell gene expression data and spatial gene expression data. Then, based on the cell type information inferred from the single-cell data, it uses a probabilistic graphical model to associate the cell type information with the spatial coordinates of the spatial gene expression data and outputs a spatial cell type composition probability matrix. The joint analysis module receives the spatial cell type composition probability matrix, then performs spatial resolution intercellular communication inference and multi-omics co-expression network analysis, and outputs the analysis results. The visualization output module receives the analysis results and transforms them into interactive spatial visualization charts and structured analysis reports.

2. The multi-omics analysis system integrating single-cell and spatial transcriptome data according to claim 1, characterized in that: The input processing module includes a cross-platform correction unit, which is used to identify the housekeeping gene set common to single-cell and spatial transcriptome data, and apply a batch effect correction algorithm based on a linear model to make the gene expression profiles of the two sets of data comparable.

3. The multi-omics analysis system integrating single-cell and spatial transcriptome data according to claim 1, characterized in that: The data integration and alignment module includes a dual feature anchoring unit, which is used to perform a first anchoring analysis and a second anchoring analysis to obtain a first feature set and a second feature set, respectively. Then, the first feature set and the second feature set are combined to form a shared feature set. The first anchoring analysis is based on single-cell gene expression data to identify cell type-specific differentially expressed genes and on spatial gene expression data to identify spatially differentially expressed genes. The intersection of the two is taken to form the first feature set. The second anchoring analysis involves screening gene pairs with stable cross-cell type correlations from single-cell gene expression data and verifying their consistency in correlations in spatial gene expression data. Consistent gene pairs are retained to form the second feature set.

4. The multi-omics analysis system integrating single-cell and spatial transcriptome data according to claim 3, characterized in that: The data integration and alignment module also includes a probabilistic graphical model alignment unit, which uses the expression data of genes in the shared feature set as the observation input, and the cell type labels and their expression features inferred from single-cell data as the hidden state priors to construct a conditional random field model with spatial points as nodes. The model infers and calculates the probability that each spatial point belongs to each cell type, thereby generating the spatial cell type composition probability matrix.

5. The multi-omics analysis system integrating single-cell and spatial transcriptome data according to claim 1, characterized in that: The joint analysis module includes a spatial cell communication inference unit, which is used to quantify the ligand-receptor interaction strength between different cell types in local spatial regions by assigning cell type origin weights to the ligand and receptor expression levels of each spatial location based on the spatial cell type composition probability matrix and the proximity relationship network of spatial locations.

6. The multi-omics analysis system integrating single-cell and spatial transcriptome data according to claim 4, characterized in that: The joint analysis module includes a multi-omics co-expression analysis unit, which is used to merge single-cell gene expression data and spatial gene expression data based on a shared feature set, perform weighted gene co-expression network analysis on the merged data, identify gene modules conserved in spatial and cell type dimensions, and finally associate gene modules with their spatial activity distribution and biological functions.

7. The multi-omics analysis system integrating single-cell and spatial transcriptome data according to claim 1, characterized in that: The visualization output module outputs an interactive spatial distribution map, which can display at least one of the following information layers in a linked manner: spatial cell type composition probability, spatial ligand-receptor interaction strength and direction, and active hotspot regions of spatial co-expression modules.

8. The multi-omics analysis system integrating single-cell and spatial transcriptome data according to claim 1, characterized in that: The visualization output module includes an automated report generation unit, which integrates spatially differentially expressed genes, significant cell interaction pairs, co-expression modules, and key results of their functional annotations to automatically generate structured analysis reports.