Spatial multi-omics integration method and device, equipment and storage medium
By using gradient coordination and adaptive eigenvalue decomposition to dynamically adjust omics contributions and combining internal and external omics learning, the problem of omics imbalance in spatial multi-omics data integration is solved, achieving more accurate cross-omics alignment and biological signal parsing, and improving the integration effect.
Patent Information
- Application Number
- CN202511164113.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-01-09
AI Technical Summary
Existing spatial multi-omics data integration methods are difficult to effectively integrate data from different omics levels. Especially after introducing spatial information, they cannot make full use of the complementary information between different omics levels, resulting in significant limitations in the analysis of multi-omics interaction relationships. Furthermore, omics imbalance is prone to occur during the integration process, leading to the loss of key biological information.
A gradient coordination and adaptive feature decomposition method is adopted. By constructing an adjacency graph and a multi-head attention mechanism, the contributions of different omics are dynamically adjusted. By combining intra-omics private learning and inter-omics shared learning, the collaborative optimization of cross-omics learning is achieved. A two-stream architecture is adopted to retain omics-specific features while learning shared representations.
It achieves efficient integration of spatial multi-omics data, improves the accuracy of cross-omics alignment and the ability to resolve biological signals, significantly outperforming existing methods, and can more comprehensively reveal the complex characteristics of cells and the multi-level biological mechanisms behind them.
Smart Images

Figure CN121306265A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of space multi-omics technology, and in particular to a space multi-omics integration method, apparatus, device and storage medium. Background Technology
[0002] In recent years, spatial multi-omics technology has developed rapidly, enabling researchers to simultaneously acquire multi-level omics data and their spatial location information. This technological breakthrough overcomes the limitations of traditional sequencing methods, which rely on aligning sequencing reads to reconstruct transcript structure and measure gene expression levels, but often lack spatial information. In contrast, spatial multi-omics integrates spatial location with molecular feature analysis, enabling spatial mapping of molecular features at the tissue level. It encompasses not only spatial transcriptomics but also extends to spatial epigenomics and spatial proteomics. These technologies can simultaneously resolve genomic structure, molecular state, and cellular spatial distribution, thereby revealing the complex biological relationships between gene regulatory mechanisms and the tissue microenvironment. Among them, spatially resolved transcriptomics (SRT) was selected as the "Method of the Year" by *Nature* magazine in 2020 for its ability to resolve the spatial distribution of transcripts while preserving tissue structural information. As an important branch of this field, the application of spatial transcriptomics has promoted the study of the spatial characteristics of cellular function within tissues. Currently, spatial multi-omics has received widespread attention and was again listed by *Nature* as one of the seven key technologies in 2022. This method not only integrates multi-omics data such as transcriptomics, proteomics and chromatin accessibility, but also captures the spatial heterogeneity of cell function and microenvironment-dependent regulatory mechanisms at single-cell or even subcellular resolution. In recent years, a number of commercial imaging spatial omics platforms with single-cell resolution have emerged, including NanoString CosMx based on smFISH, Vizgen MASCOPE based on MERFISH
[16] and 10x GenomicsXechium based on ISS, which significantly enhance the ability of high-resolution spatial analysis. The core goal of spatial multi-omics is to integrate heterogeneous omics data and preserve tissue spatial structure, thereby accurately characterizing the spatial identity of cells and deeply analyzing gene regulatory networks, cell differentiation trajectories and intercellular interactions. It is worth emphasizing that the introduction of spatial information enables researchers to identify microenvironment-dependent molecular mechanisms and explore biological characteristics at different tissue levels. However, due to the significant differences in data characteristics and distribution at different omics levels, how to efficiently and accurately integrate spatial multi-omics data is still an important research problem that needs to be solved.
[0003] Currently, the integration of spatial multi-omics data still faces many challenges, mainly due to significant differences in the number of features among different omics (e.g., the mismatch between protein and transcript measurements) and their different statistical distributions. This integration problem becomes even more complex when spatial information is combined with the features of each omics. Existing multimodal integration methods can be broadly classified into three categories: non-spatial multi-omics data integration, spatial single-omics data integration, and spatial multi-omics data integration. Non-spatial multi-omics integration methods include Seurat WNN, MOFA+, totalVI, and MultiVI. Among them, totalVI is specifically designed for RNA and proteomics in CITE-seq data. Based on the variational autoencoder (VAE) framework, it maps RNA and protein data to a shared low-dimensional latent space, effectively correcting technical biases such as batch effects and protein background noise, thereby achieving unbiased cross-dataset integration and filling in missing protein measurements. Seurat WNN extends support for multi-omics data through the weighted nearest neighbor (WNN) algorithm, which dynamically adjusts the weights of each omics based on data similarity. MOFA+, based on a factor analysis framework, provides a flexible approach for cross-omics integration. MultiVI combines a modular encoder-decoder architecture with an adversarial training mechanism, enabling joint modeling of RNA, ATAC, and protein data, thus expanding the range of omics that can be integrated. While these methods demonstrate good performance in integrating multi-omics data, they do not incorporate spatial information, making it difficult to accurately resolve spatially specific biological structures and limiting their ability to capture spatial heterogeneity and identify spatial domains.
[0004] To address these challenges, spatial single-omics data integration tools have emerged, such as STAGATE, GAAEST, SpaGCN, GraphST, and iSpatial. These methods combine spatial information with single-omics modalities, demonstrating unique advantages in spatial organizational structure analysis. STAGATE utilizes graph attention autoencoders to extract spatially dependent features, while GAAEST optimizes spatial embeddings through a three-level graph attention contrastive learning framework that integrates local location information, global features, and contextual features. GraphST employs a cross-graph self-supervised contrastive learning strategy to integrate spatial transcriptome data, while SpaGCN analyzes spatial transcriptome data based on graph convolutional networks (GCN) and supports joint analysis with histological images. Furthermore, iSpatial combines single-cell transcriptome (scRNA-seq) data with spatial transcriptome (ST) data, using dimensionality reduction, weighted k-nearest neighbor (KNN) algorithms, and the Harmony algorithm to infer genome-wide single-cell spatial expression patterns. These methods significantly improve the ability to resolve spatial structures, providing strong technical support for spatial analysis of single-omics data.
[0005] However, because these tools rely on data from only a single omics level and fail to fully utilize complementary information between different omics levels, they struggle to comprehensively reveal the complex characteristics of cells and the multi-layered biological mechanisms behind them. Consequently, they have significant limitations in elucidating multi-omics interactions.
[0006] With the emergence of more and more spatial multi-omics technologies with single-cell resolution and the ability to simultaneously resolve multiple molecular components, vertical integration strategies based on existing non-spatial single-cell methods are also rapidly developing. By leveraging spatial information that can link cell states with their microenvironment and macro-organic environment (e.g., through graph neural networks), researchers hope to obtain more detailed multi-omics cell state representations. Current state-of-the-art spatial multi-omics integration methods have made significant progress. For example, SpatialGlue uses a dual attention mechanism to strengthen the association between spatial features and multi-omics features, while PRAGA achieves deep fusion of spatial information and feature semantics by constructing dynamic graphs to mine latent semantic relationships. However, existing research mainly focuses on how to efficiently integrate multi-omics data, often neglecting the potential conflicts between multi-omics and single-omics learning objectives during model optimization. These conflicts can lead to inconsistent convergence progress among different omics, thus limiting the model's ability to understand and interpret multi-omics data. Despite significant technological advancements, there is still room for further improvement. The root cause of this problem lies in the phenomenon of "omics imbalance," where a dominant omics maskes the signals of other omics during training, suppressing their contributions and hindering the full utilization of multi-omics information. When different omics converge at inconsistent speeds under the same optimization objective, it leads to an imbalance in learning efficiency. Furthermore, during the integration process, deep multi-omics models often prioritize learning shared representations among omics while neglecting omics-specific features. This can result in the loss of crucial biological information and weaken the model's ability to capture biological signals. Summary of the Invention
[0007] This application provides a spatial multi-omics integration method, apparatus, device, and storage medium, aiming to achieve collaborative optimization of cross-omics learning through gradient coordination and adaptive feature decomposition. It introduces a novel gradient balancing mechanism that dynamically adjusts the contributions of different omics during backpropagation, and resolves gradient conflicts through a task-specific prioritization mechanism without manually setting weights. Simultaneously, a two-stream architecture is employed to preserve omics-specific features while learning shared representations.
[0008] Firstly, this application provides a spatial multi-omics integration method, including:
[0009] Obtain a dual-omics dataset; wherein the dual-omics dataset includes first omics data and second omics data;
[0010] Based on the aforementioned dual-omics dataset, an adjacency graph is constructed; wherein, the adjacency graph includes a spatial graph and a feature graph;
[0011] An intra-omics integration module is constructed. In response to the input spatial graph and feature graph, the intra-omics integration module performs graph convolution operation to obtain spatial latent representation and feature latent representation respectively. The spatial latent representation and feature latent representation are concatenated on the feature dimension to obtain concatenated features. The concatenated features are input into a multilayer perceptron to obtain intra-omics fusion representation.
[0012] An inter-omics integration module is constructed. The inter-omics integration module responds to the input first intra-omics fusion representation and second intra-omics fusion representation. The first intra-omics fusion representation and the second intra-omics fusion representation are output through the intra-omics integration module. A multi-head attention mechanism is applied to the first intra-omics fusion representation and the second intra-omics fusion representation to obtain the attention weights of each intra-omics fusion representation. The attention weights and the corresponding intra-omics fusion representations are weighted and summed to obtain the multi-omics latent representation.
[0013] The intra-omics integration module and the inter-omics integration module are combined to form a spatial multi-omics integration model. A dual learning strategy is used to train the spatial multi-omics integration model, and spatial multi-omics integration is achieved based on the trained spatial multi-omics integration model. The dual learning strategy includes intra-omics private learning and inter-omics shared learning. The intra-omics private learning is used to guide the maintenance of the mapping relationship between the potential representation of multi-omics and the original normalized representation space, and self-supervised contrastive learning is introduced to enhance the model's ability to identify omics-specific features. The goal of the inter-omics shared learning is to impose constraints between the representation of one omics and the corresponding representation generated by another omics through the decoder-encoder path, thereby ensuring that data from different omics remain aligned in the shared space.
[0014] In one possible design, the intra-omics integration module, in response to the input spatial graph and feature graph, performs graph convolution operations to obtain the computation process of the spatial latent representation and the feature latent representation, as follows:
[0015]
[0016] In the formula, and These are the normalized adjacency matrices of the spatial graph and the feature graph, respectively. and Let be the trainable weight matrix of the spatial graph and the feature map. and This is the bias vector between the spatial map and the feature map. and Let represent the spatial latent representation and feature latent representation learned from the spatial map and feature map, respectively, and σ represent the activation function;
[0017] The spatial latent representation and the feature latent representation are concatenated along the feature dimension using the following formula to obtain the concatenated feature:
[0018]
[0019] In the formula, x represents the concatenation feature, concat represents the concatenation operation, R represents the set of real numbers, N represents the number of spatial points, and d latent Indicates feature dimension;
[0020] The stitched features are input into a multilayer perceptron, and the intra-omics fusion representation is calculated using the following formula:
[0021] x (i) =Dropout(σ(W) i x i-1 +b i ))
[0022]
[0023] In the formula, x (i) x represents the output of the i-th layer neuron in a multilayer perceptron. i-1 W represents the output of the (i-1)th layer neuron in a multilayer perceptron. i and b i x represents the weights and biases of neurons in the i-th layer of a multilayer perceptron, and Dropout represents randomly assigning neurons to their positions. (L) x represents the output after processing by an L-layer multilayer perceptron. (0) =x is the initial input of the multilayer perceptron, d hidden H represents the dimension of latent features. final This represents the fusion representation within omics.
[0024] In one possible design, a multi-head attention mechanism is applied to the first and second intra-omics fusion representations to obtain attention weights for each intra-omics fusion representation. The attention weights and their corresponding intra-omics fusion representations are then weighted and summed to obtain a multi-omics latent representation, including:
[0025] Based on the first and second sets of intra-learning fusion representations, the query vector, key vector, and value vector are determined using the following formulas:
[0026]
[0027] In the formula, Q, K, and V represent the query vector, key vector, and value vector, respectively, and W... Q WK and W V Let represent the trainable linear mapping matrices for the query vector, key vector, and value vector, respectively. and These are the first group of intra-learning fusion representations and the second group of intra-learning fusion representations, respectively.
[0028] Based on the query vector, key vector, and value vector, the representation of omics-specific features is determined using the following formula:
[0029]
[0030] The attention representation output by the nth attention head, where Attention is the attention function, Q. n ,K n V n Let be the query vector, key vector, and value vector of the nth attention head, respectively, and T be the matrix transpose. h_size Let H be the dimension of the attention head, Softmax be the normalized exponential function, and H be the dimension of the attention head. mean For omics-specific feature representation, head_num is the number of attention heads, and n is the number of the nth attention head currently being computed;
[0031] Based on the omics-specific feature representation, the attention score is determined using the following formula:
[0032]
[0033] In the formula, v represents the attention score of omics m for sample i. m W m and b m These are the learnable parameters related to omics m. This represents the omics-specific features of sample i;
[0034] Based on the attention score, the attention weight is determined using the following formula:
[0035]
[0036] In the formula, Let m be the attention weight of omics m for sample i, exp be an exponential function with the natural constant as the base, M be the number of omics, and m be the index of the omics.
[0037] The multi-omics latent representation is obtained by weighted summation of the attention weights and the corresponding intra-omics fusion representations using the following formula:
[0038]
[0039] In the formula, z i The multi-omics latent representation of sample i. This represents the intra-omics fusion representation of omics m.
[0040] In one possible design, after constructing the adjacency graph, the method further includes: performing data augmentation on the adjacency graph to obtain a perturbed graph; wherein the data augmentation includes: performing random permutations on the gene expression matrix while maintaining the topological structure of the adjacency graph to obtain a perturbed gene expression matrix, constructing a perturbed graph based on the perturbed gene expression matrix, and introducing a symmetric contrastive loss function into the perturbed graph to optimize spatial representation consistency when constructing the perturbed graph, wherein the contrastive loss function is expressed as:
[0041]
[0042] In the formula, To compare the losses, For the representation of position i in the perturbation graph, g ' i Let i be the local environment vector at position i in the perturbation map. Let E represent position i in the adjacency graph. (X,A) To represent the expectation under the original graph (adjacency graph) distribution, E (X',A') To represent the expectation under the perturbation graph distribution, Φ is a similarity function that measures the degree of similarity between two vectors.
[0043] In one possible design, the intra-omics private learning includes:
[0044] A decoder is constructed to generate reconstructed features based on the intra-omics fusion representation, wherein the reconstructed representation of the decoder at layer l is calculated by the following formula:
[0045]
[0046] In the formula, The reconstructed features are the output of the decoder at layer l. Let be the normalized adjacency matrix of the spatial graph, and σ be the activation function. and Z is a trainable weight matrix and bias vector. l-1 This is the output of the decoder at layer l-1;
[0047] Based on the reconstructed features, a reconstruction loss function is constructed, which is expressed as follows:
[0048]
[0049] In the formula, Lrecon The reconstruction loss is defined as follows: i is the index of the sample, N is the number of samples, and x... i Let i be the intra-omics fusion representation of sample i. The reconstructed features of sample i;
[0050] Contrastive learning based on adjacency graphs and perturbation graphs includes: pairing a representation of a location in the adjacency graph with its corresponding local environment vector to form a positive sample; pairing a representation of a location in the perturbation graph with its corresponding local environment vector to form a negative sample; and performing contrastive learning using the positive and negative samples according to a predefined contrastive learning loss function; wherein the contrastive learning loss function is expressed as:
[0051]
[0052] In the formula, L scl To compare learning loss, For the representation of position i in the adjacency graph, g i Let i be the local environment vector at position i in the adjacency graph;
[0053] An intra-omics loss function is constructed based on the reconstruction loss function, the contrastive loss function, and the contrastive learning loss function, and intra-omics private learning is performed according to the intra-omics loss function; wherein, the intra-omics loss function is expressed as:
[0054]
[0055] In the formula, L omics The loss is omics-based, and λ1 and λ2 are hyperparameters.
[0056] In one possible design, the inter-omics shared learning includes:
[0057] Obtain the omics latent representation extracted from one omics and the cross-omics representation generated through the decoder-encoder path of another omics, and optimize the mapping relationship between the first omics latent representation and the second omics latent representation using the following formula:
[0058]
[0059] In the formula, and Let be the omics potential representations of the i-th sample in omics 1 and omics 2, respectively. and W represents the cross-omics representations generated via omics 2 and omics 1 paths, respectively. d and W e These are the trainable parameters for the decoder and encoder, respectively. This refers to the l-th layer cross-omics representation generated through the omics 2 to omics 1 path, i.e., reconstructing omics 1 using information from omics 2. The representation input of omics 1 at layer l-1 is the input feature of the current graph neural network. and These are the trainable weights and bias parameters of the decoder at layer l-1. and Let L be the trainable weights and bias parameters of the encoder at layer l-1. corr Let γ1 and γ2 be the cross-omics alignment loss function, where γ1 and γ2 are the weighting coefficients controlling the alignment loss between the two omics in the total loss. and Let be the squared norm between the representation of the i-th sample and its cross-modal reconstructed representation.
[0060] Establish a shared learning contrastive loss function, expressed as:
[0061]
[0062] In the formula, L bt To share the learning contrast loss, Representing the latent representation of two omics modes and The cross-correlation matrix between them, where N is the number of samples, d is the feature dimension, and C ii As diagonal elements, C ii For off-diagonal elements, γ3 is a hyperparameter used to balance diagonal and off-diagonal terms;
[0063] exchange and Given the position, calculate the mapping relationship in both directions, and take the average value to obtain the final optimization objective:
[0064]
[0065] In the formula, β is a weighting parameter that controls the strength of omics alignment, and L multi_omics This represents a loss between omics studies.
[0066] In one possible design, the dual learning strategy further includes degree-omics balanced learning, which includes:
[0067] Cosine similarity is determined using the following formula:
[0068]
[0069] In the formula, The gradient of the loss function specific to each omics. is the gradient of the joint loss between omics, and cosβ is the cosine similarity;
[0070] When cosβ≥0 and These are the weights for inter-omics loss and intra-omics loss, respectively.
[0071] When cosβ < 0, it is determined by the following formula. and
[0072]
[0073] Based on the determined weights of intra-omics and inter-omics losses, the total loss function is established using the following formula:
[0074]
[0075] In the formula, L total Let represent the total loss, n be the number of training samples, and i be the index of the sample.
[0076] The spatial multi-omics integration model is trained based on the total loss function.
[0077] Secondly, this application provides a space multi-omics integration device, comprising:
[0078] The data acquisition module is configured to acquire a dual-omics dataset; wherein the dual-omics dataset includes first omics data and second omics data;
[0079] The graph feature construction module is configured to construct an adjacency graph based on the dual-omics dataset; wherein the adjacency graph includes a spatial graph and a feature graph;
[0080] An intra-omics integration module is configured to construct an intra-omics integration module. The intra-omics integration module responds to the input spatial graph and feature graph by performing graph convolution operations to obtain spatial latent representation and feature latent representation, concatenates the spatial latent representation and feature latent representation on the feature dimension to obtain concatenated features, and inputs the concatenated features into a multilayer perceptron to obtain an intra-omics fusion representation.
[0081] An inter-omics integration module is configured to construct an inter-omics integration module that responds to input first intra-omics fusion representation and second intra-omics fusion representation. The first intra-omics fusion representation and the second intra-omics fusion representation are output by the inter-omics integration module. A multi-head attention mechanism is applied to the first intra-omics fusion representation and the second intra-omics fusion representation to obtain the attention weights of each intra-omics fusion representation. The attention weights and the corresponding intra-omics fusion representations are weighted and summed to obtain a multi-omics latent representation.
[0082] The learning and training module is configured to combine the intra-omics integration module and the inter-omics integration module to form a spatial multi-omics integration model. A dual learning strategy is employed to train the spatial multi-omics integration model, and spatial multi-omics integration is achieved based on the trained spatial multi-omics integration model. The dual learning strategy includes intra-omics private learning and inter-omics shared learning. Intra-omics private learning guides the maintenance of a mapping relationship between the multi-omics latent representation and the original normalized representation space, and introduces self-supervised contrastive learning to enhance the model's ability to identify omics-specific features. The goal of inter-omics shared learning is to impose constraints between the representation of one omics and the corresponding representation generated by another omics through a decoder-encoder path, thereby ensuring that data from different omics remain aligned in the shared space.
[0083] Thirdly, embodiments of this application provide an electronic device, including: at least one processor and a memory; the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, causing the at least one processor to perform the spatial multi-omics integration method as described in the first aspect and various possible designs of the first aspect.
[0084] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the spatial multi-omics integration method described in the first aspect and various possible designs of the first aspect.
[0085] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the spatial multi-omics integration method as described in the first aspect and various possible designs of the first aspect.
[0086] The beneficial effects of this application are as follows:
[0087] This application proposes SpaBalance, a balanced deep learning model for efficient integration of spatial multi-omics. SpaBalance achieves accurate cross-omics alignment by dynamically adjusting the learning weights of each omics, thus fully integrating spatial information with multi-omics features. Specifically, SpaBalance introduces a gradient conflict coordination mechanism, which dynamically adjusts the gradient direction and magnitude of the learning objectives of multi-omics and single-omics learning, effectively mitigating the optimization imbalance caused by inter-omics conflicts, thereby improving the integration effect of spatial multi-omics data.
[0088] Furthermore, SpaBalance employs a dual learning strategy combining inter-omics shared learning and intra-omics specific learning. Inter-omics shared learning enhances cross-omics complementarity and strengthens integration by improving consistency among different omics within a unified feature space; while intra-omics specific learning focuses on preserving the unique information within each omics, preventing the loss of key omics-specific features during integration, thus ensuring more accurate biosignal analysis. Extensive qualitative and quantitative experimental results demonstrate that SpaBalance significantly outperforms state-of-the-art methods in integrating spatial multi-omics information into structured, interpretable representations, validating its effectiveness and broad applicability. Further, to verify the practical effect of SpaBalance in balancing multi-omics learning, this application designed and implemented ablation experiments. The results show that the dynamic weight adjustment mechanism plays a crucial role in improving model performance, thus confirming SpaBalance's significant contribution to spatial multi-omics data integration. Attached Figure Description
[0089] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0090] Figure 1 A schematic diagram of a deep integration model for spatial multi-omics data analysis provided in this application embodiment;
[0091] Figure 2 A flowchart illustrating a spatial multi-omics integration method provided in this application embodiment;
[0092] Figure 3 The experimental results of SpaBalance integrating three omics datasets are shown in the figure provided in this embodiment of the invention.
[0093] Figure 4 This is a schematic diagram illustrating how SpaBalance accurately identifies spatial structural domains in human lymph node A1, as provided in an embodiment of the present invention.
[0094] Figure 5 Supplementary results of the ablation experiment on human lymph node A1 using SpaBalance provided in the embodiments of the present invention;
[0095] Figure 6 An integrated result diagram of human lymph node sample D1 provided in an embodiment of the present invention;
[0096] Figure 7 A schematic diagram of SpaBalance's higher resolution analysis of spatial epigenome and transcriptome samples of the mouse brain, provided as an embodiment of the present invention;
[0097] Figure 8This is a schematic diagram of SpaBalance's analysis of more complex mouse brain spatial epigenome-transcriptome samples, provided in an embodiment of the present invention.
[0098] Figure 9 A schematic diagram of SpaBalance integrating mouse thymus (based on RNA and protein data obtained from Stereo-CITE-seq) provided in an embodiment of the present invention;
[0099] Figure 10 This is a diagram showing the integrated results of the mouse thymus 1 and mouse thymus 2 datasets provided in an embodiment of the present invention.
[0100] Figure 11 A schematic diagram of a mouse spleen integrated with SpaBalance (RNA and protein data obtained based on SPOTS technology) provided in an embodiment of the present invention;
[0101] Figure 12 This is a diagram showing the integrated results of the mouse spleen 1 and mouse spleen 2 datasets provided in an embodiment of the present invention.
[0102] Figure 13 This is a schematic diagram of the structure of a spatial multi-omics integration device provided in an embodiment of this application.
[0103] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0104] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0105] The collection, storage, use, processing, transmission, provision, and disclosure of information such as financial data, user data, or medical image data involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0106] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0107] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0108] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0109] Example 1: A Spatial Multi-Omics Integration Method (SpaBalance)
[0110] Recent breakthroughs in spatially resolved multi-omics technologies have enabled researchers to simultaneously resolve multiple molecular levels at the tissue level, gaining an unprecedented understanding of their synergistic roles in development and disease. Despite these significant advances, the integration and analysis of multi-omics data remains a major challenge due to inherent biological and technical differences between experiments, which often lead to gradient conflicts during joint learning. Gradient conflicts arise from competition or contradictions among different omics on optimization paths, thus limiting the improvement of integration performance. To address this issue, this application provides a spatial multi-omics integration method. This method can accurately predict cell type abundance from spatial histopathological images. In this application, this spatial multi-omics integration method is named SpaBalance for ease of subsequent description. SpaBalance introduces a novel gradient balancing mechanism that dynamically adjusts the contributions of different omics during backpropagation, resolving gradient conflicts through a task-specific prioritization mechanism without manually setting weights. Simultaneously, SpaBalance employs a two-stream architecture, preserving omics-specific features while learning shared representations. This embodiment extensively evaluates SpaBalance on various spatial omics datasets, including epigenome-transcriptome and proteome-transcriptome paired data from human tumors and brain tissues. The results demonstrate SpaBalance's outstanding ability to reveal complex spatial domains and discover previously hidden multi-omics regulatory hubs. Comparisons with existing mainstream methods show that SpaBalance achieves significant improvements in both clustering accuracy and biological interpretability. Furthermore, SpaBalance's flexible design supports the integration of three or more omics datasets and achieves near-linear computational scalability. By bridging the critical gap between multi-omics data integration and biological discovery, SpaBalance significantly advances the study of tissue structure and lays a solid foundation for future research in space systems biology.
[0111] like Figure 1The diagram shows a deep integration model for spatial multi-omics data analysis provided in this application embodiment. Here, 'a' represents the SpaBalance model architecture. To capture the interactions between different omics, the SpaBalance learning framework is divided into private learning modules within omics and shared learning modules between omics. The private learning modules within omics extract omics-specific representations through graph structure feature integration. To ensure the importance of each omics is preserved, SpaBalance introduces a multi-head attention mechanism in the shared learning between omics, effectively integrating the omics-specific representations to generate the final unified representation. Furthermore, to avoid gradient dominance in one omics during training, SpaBalance achieves dynamic balance in the learning process by coordinating gradient conflicts, thereby improving the accuracy of the fused representation. 'b' represents the integration of spatial and feature information in SpaBalance. To construct a spatial adjacency graph, SpaBalance first applies the k-nearest neighbor (KNN) algorithm based on the spatial coordinates and normalized representation features of each omics. For each omics, a graph neural network (GNN) encoder learns two graph structure feature representations through iterative aggregation of adjacency node information. Subsequently, the two graph feature representations of this omics are integrated through a multilayer perceptron (MLP) module to generate a unified single-omics representation. c represents a downstream application of SpaBalance. Through cluster analysis of the final integrated representation, visualization of intra-cluster compactness and inter-cluster separation can be achieved, spatial cellular domains can be identified, and differential expression analysis can be performed between different spatial regions.
[0112] SpaBalance is a balanced learning integration framework for spatial multi-omics, aiming to achieve a balance in cross-omics learning during data integration while preserving the specific characteristics of each omics. To achieve this goal, SpaBalance employs a dual strategy: within each omics, it combines graph neural networks (GNNs) and multilayer perceptrons (MLPs) to fuse spatial information with observational features. Figure 1 (b) In inter-omics learning, a multi-head attention mechanism is employed to promote the interaction and effective integration of cross-omics information. During the integration process, SpaBalance coordinates gradient conflicts to prevent any one omics from dominating the optimization direction during training and alleviates the optimization imbalance caused by the dynamic differences in learning between different omics, thereby achieving a balance in multi-omics learning. To preserve omics-specific information, SpaBalance implements a dual learning strategy consisting of shared learning between omics and specific learning within omics. Shared learning enhances the consistency and interaction capabilities of features between omics, while specific learning effectively preserves the unique features of each omics. This strategy enables SpaBalance to maintain the uniqueness of each omics while integrating cross-omics information, ultimately achieving a more comprehensive and biologically meaningful spatial multi-omics integration. Figure 1(a). The integrated feature representation can further support downstream analysis tasks, such as cluster analysis, cell type identification, and differential gene expression analysis. Figure 1 (c).
[0113] Unlike traditional multi-omics methods, spatial multi-omics data carries both molecular features and spatial coordinates within each omics. Due to the inherent spatial heterogeneity of cells, encoding only omics features may not fully capture spatial information. To address this issue, SpaBalance constructs a spatial adjacency graph and embeds preprocessed feature count data into a shared low-dimensional space. This method not only preserves the unique semantic information of each omics but also accurately characterizes the spatial heterogeneity of cells. To adaptively learn the spatial and feature relationships within omics and ensure robust information integration, SpaBalance employs a strategy combining Multilayer Perceptron (MLP) and Dropout during intra-omics integration. MLP fuses information from different levels layer by layer, while Dropout effectively reduces the risk of overfitting and improves model generalization. Regarding inter-omics integration, simple weighted summation may ignore complex relationships and feature differences between omics, leading to decreased integration performance. To overcome this limitation, SpaBalance introduces a multi-head attention mechanism, dynamically assigning different weights to each omics, enabling the model to accurately capture inter-omics relationships and fine-grained interactions. Multi-head attention mechanisms support the parallel learning of correlations between omics in multiple subspaces, thereby obtaining more accurate cross-omics representations and improving the ability to classify spatial domains or cell types.
[0114] However, during model training, uneven convergence speeds among different omics can lead to conflicts between multi-omics and single-omics gradients. These conflicts can disrupt the optimization path, resulting in underlearning of some omics or even misleading the single-omics encoder. To mitigate this issue, SpaBalance calculates the gradient directions of the multi-omics and single-omics learning objectives and dynamically adjusts the loss weights between the single-omics and multi-omics objectives, thereby ensuring a balance in learning contributions among omics during the optimization process and effectively reducing gradient conflicts. In addition to improving the accuracy of shared omics representations, SpaBalance also focuses on preserving key omics-specific information. To this end, this application proposes a dual learning strategy, including shared learning between omics and omics-specific learning within omics. Shared learning aims to enhance cross-omics feature consistency and improve the accuracy of shared representations; while omics-specific learning focuses on preserving omics-specific features, ensuring that key omics-specific information is fully preserved during the integration process.
[0115] Ultimately, the integrated multi-omics representations generated by SpaBalance can be used for biologically relevant spatial domain identification and subsequent clustering analysis, thereby uncovering biological features and potential patterns in spatial multi-omics data. To evaluate the accuracy of the SpaBalance model, this embodiment conducted comparative experiments with other tools on the same dataset. Furthermore, this embodiment also carried out a series of ablation experiments to verify the effectiveness of the multi-omics balanced learning strategy in SpaBalance.
[0116] SpaBalance is a novel spatial multi-omics integration framework based on graph neural networks, designed to address the challenges posed by heterogeneous molecular omics and complex organizational structures. This method achieves coordinated integration between spatial topology and molecular specificity through a multi-level feature interaction mechanism and a dynamic equilibrium learning strategy. The framework maps these heterogeneous features into a unified latent space, thereby preserving spatial co-location information and functional complementarity. The overall architecture of SpaBalance consists of five core modules: (1) a multi-omics data augmentation module; (2) an intra-omics spatial and feature integration module; (3) an inter-omics multi-head attention integration module; (4) an intra-omics and inter-omics dual learning module; and (5) a multi-omics equilibrium learning module. These modules work together to enable SpaBalance to achieve high-fidelity integration of spatial multi-omics data and support downstream tasks such as spatial structural domain identification, functional region annotation, and multimodal biomarker discovery. Detailed information about each module will be provided below.
[0117] Specifically, such as Figure 2 The diagram shows a flowchart of a spatial multi-omics integration method. This spatial multi-omics integration method includes the following steps S10-S50.
[0118] S10: Obtain the dual-omics dataset; wherein, the dual-omics dataset includes the first omics data and the second omics data.
[0119] In this embodiment, to accommodate the dual-omics dataset, the input features (first omics data and second omics data) are defined as follows: and Where N represents the number of spatial locations, and d1 and d2 represent the feature dimensions of each omics, such as transcriptome and epigenome or proteome data.
[0120] S20: Construct an adjacency graph based on a dual-omics dataset; the adjacency graph includes a spatial graph and a feature graph.
[0121] In complex tissue samples, the spatial distribution of cells is often non-uniform. Typically, spatially adjacent points are more likely to have similar cell types or states, thus spatial information can be used to capture relationships between local cells. However, relying solely on spatial location may overlook some important biological characteristics, as points with the same cell type or state may be spatially distant and cannot be directly captured through spatial neighborhood relationships. Therefore, in addition to constructing a spatial graph G to reflect the spatial adjacency relationships of cells, [further research is needed]. s =(V s E s In addition, this embodiment also requires constructing a feature map G in the latent space. f =(V f E f This method simulates the similarity between cellular features, thereby compensating for the limitations of spatial information and capturing the relationships between cells more comprehensively.
[0122] Spatial diagram G s =(V s E s A spatial graph is used to represent the spatial adjacency relationships between cells in a tissue. Each point is considered a cell in space, where the set of nodes V is... s Let E represent all cells and the set of edges. s This represents the adjacency relationship between cells. The adjacency relationship is defined by calculating the Euclidean distance between points. Specifically, for a given cell i∈V... s In this embodiment, the nearest r neighbors in space are selected as its adjacent cells. When the distance between two cells is within the range of nearest neighbor r = 3, the adjacency matrix A is... s ∈R N×N The corresponding position is set to 1; otherwise, it is set to 0.
[0123] Feature map G f =(V f E f Feature maps are used to capture the similarity relationships between cells in the feature space. The node set V f Let E represent all points in the cell feature space and the set of edges. f This represents the similarity relationships between cells in the feature space. First, principal component analysis (PCA) is performed on the data for each omics to extract key features; then, the k-nearest neighbor (KNN) algorithm is applied in the PCA embedding space. For each cell i∈V f Cell j is selected as its k=20 nearest neighbors in the feature space. If cell j is one of the k nearest neighbors of cell i in the feature space, then the adjacency matrix A of the feature graph is... f ∈R N×N The corresponding position is set to 1, otherwise it is set to 0. This is achieved by combining the spatial graph G.s With feature map G f This allows for a more comprehensive capture of the multidimensional relationships between cells in terms of space and features.
[0124] In some embodiments, a data augmentation step is included after constructing the adjacency graph. Data augmentation plays a crucial role in self-supervised contrastive learning, enhancing model robustness, preventing overfitting to specific spatial graph structures or gene expression features, and helping the model learn more general and stable cell relationship representations. After constructing the adjacency graph, this embodiment generates positive and negative samples through data augmentation to assist the contrastive learning task. Specifically, given an adjacency graph G = (V, E) and a gene expression matrix X ∈ R... N×F (Where N represents the number of spatial points and F represents the dimension of gene features), in this embodiment, while maintaining the adjacency graph topology, a random permutation operation is performed on the gene expression matrix XXX to obtain the perturbed gene expression matrix X'={x'1,x'2,..,x' N}∈R N×F And construct a new perturbation graph G' = (V', E'). The node set and edge set remain the same as the original. Figure 1 Since V' = V and E' = E, the adjacency matrix is regularized. It is the same as the original figure, except that the gene expression matrix is replaced with X' after random permutation.
[0125] This data augmentation method generates perturbation data with a structure similar to the original graph while disrupting the feature information of spatial points. This helps the model distinguish between real biological patterns and random noise during comparative learning, thereby improving the model's ability to generalize and understand complex spatial relationships.
[0126] S30: Construct an intra-omics integration module. In response to the input spatial graph and feature graph, the intra-omics integration module performs graph convolution operations to obtain spatial latent representation and feature latent representation. The spatial latent representation and feature latent representation are concatenated along the feature dimension to obtain concatenated features. The concatenated features are input into a multilayer perceptron to obtain intra-omics fusion representation.
[0127] The intra-omics integration module is used to achieve intra-omics integration. In some embodiments, step S30 can be implemented through the following steps S301-S302.
[0128] S301: Spatial and Feature Representation Encoding Based on Graph Neural Networks.
[0129] This embodiment uses a Graph Neural Network (GNN) encoder to learn the representation of cell points, capturing important features in gene expression profiles and spatial location information. The encoder takes an adjacency graph G and a normalized gene expression matrix X as input, and uses a Graph Convolutional Network (GCN) to learn a latent representation Z for each cell point i. i This process is achieved by iteratively aggregating the representations of neighboring nodes to update the potential representation at each layer.
[0130] The encoder in layer l is represented as follows:
[0131]
[0132] in, Let D be the normalized adjacency matrix, A be the degree matrix, and W be the adjacency matrix of the graph. l-1 and b l-1 Z represents the trainable weight matrix and bias vector, respectively, σ(·) represents a non-linear activation function (such as ReLU), and Z... 0 This represents the input gene expression matrix. The output Z of each layer... l Represents the potential representation of a node.
[0133] In order to capture both spatial information and feature information simultaneously, this embodiment focuses on the spatial adjacency graph G. s Similarity graph G f An independent encoding process was designed. The spatial graph reflects the spatial adjacency of cells in the tissue, while the feature graph characterizes the similarity relationships of cells in the phenotypic feature space. Through this design, the encoder can aggregate information from adjacent nodes through graph convolution operations, thereby effectively extracting different local patterns and dependencies.
[0134] For each type of omics, the encoder is located in the spatial adjacency graph G. s Similarity graph G f The graph convolutional network (GCN) operation is performed on the above layer. The representation of the l-th layer is calculated as follows:
[0135] Spatial diagram representation:
[0136]
[0137] Feature map representation:
[0138]
[0139] in, and These are the normalized adjacency matrices of the spatial graph and the feature graph, respectively. and For a trainable weight matrix, and For bias vectors, and These represent the latent representations learned from the spatial map and the feature map, respectively.
[0140] Finally, the encoder output is: Local environmental features representing the spatial location of cells can capture physical adjacency relationships and microenvironment information; and These two latent representations, representing feature-based cell relationships, are used to reveal molecular similarities between non-neighboring spatial points. They complement each other, encoding different semantic and topological information respectively, jointly characterizing the complex data features of a single genome (e.g., RNA, ADT, or ATAC). To construct a comprehensive representation for each spatial point, this embodiment concatenates these two latent representations along the feature dimension, i.e.:
[0141]
[0142] S302: Feature Transformation and Integration.
[0143] A multi-layer perceptron (MLP) with dropout mechanism is used to transform and fuse the concatenated representations. This MLP contains multiple linear layers, non-linear activation functions, dropout layers, and optional layer normalization (LayerNorm) layers, aiming to map high-dimensional features to a more compact and expressive latent space, generating the final intra-omics representation.
[0144] For an input representation x, each layer in the MLP is computed as follows:
[0145] x (i) =Dropout(σ(W) i x i-1 +b i ))
[0146] Where x (i) W represents the output of the (i-th)th layer. i and b i These are learnable parameters (weights and biases). Dropout is a regularization technique that randomly sets neurons to zero to prevent overfitting. (0) =x is the initial input of the MLP.
[0147] After L-layer MLP processing, the final intra-omics fusion representation can be obtained:
[0148]
[0149] Where d hidden This represents the dimension of the hidden features. This representation integrates multi-level semantic information from the spatial map and the feature map, which can comprehensively characterize the feature properties of each site and can be used as input for subsequent tasks such as classification, regression or generative modeling.
[0150] S40: Construct an inter-omics integration module. The inter-omics integration module responds to the input first intra-omics fusion representation and second intra-omics fusion representation. The first intra-omics fusion representation and the second intra-omics fusion representation are output by the intra-omics integration module. A multi-head attention mechanism is used on the first intra-omics fusion representation and the second intra-omics fusion representation to obtain the attention weights of each intra-omics fusion representation. The attention weights and the corresponding intra-omics fusion representations are weighted and summed to obtain the multi-omics latent representation.
[0151] Each omics approach can only capture a portion of the features of complex biological samples; therefore, effective multi-omics data integration is necessary to comprehensively and accurately describe the biological information of the samples. Since different omics approaches exhibit both complementarity and differences, assigning different weights based on each omics' contribution to the overall representation helps achieve more rational integration. To this end, this embodiment employs a multi-head cross-omics attention mechanism to adaptively fuse feature representations from different omics. The multi-head mechanism enables the model to learn diverse inter-omics associations across multiple subspaces, capturing complex cross-omics dependencies and thus enhancing the model's understanding of multi-omics information. The cross-omics attention mechanism supports interaction modeling between omics and can dynamically assess the relative importance of each omics, thereby adaptively assigning weights based on the distribution of omics contributions.
[0152] In some embodiments, step S40 can be implemented by the following steps S401-S402.
[0153] S401: Omics-specific weight learning.
[0154] For a unified potential representation from different omics and This embodiment first performs a linear transformation to extract features and generate Query(Q), Key(K), and Value(V) representations for modeling inter-omics interactions:
[0155]
[0156] Among them, W Q W K and W V It is a trainable linear mapping matrix, where Q, K, and V all have dimensions of 1.
[0157] Next, Q, K, V will be reshaped into the format required for a multi-head attention mechanism, with each attention head having the following dimensions. After transformation, we get: For each cell i and each omics m, this embodiment performs average aggregation on the outputs of all attention heads to obtain omics-specific feature representations:
[0158]
[0159] Subsequently, an omics-specific linear transformation layer is used to generate the attention score for each omics on sample i.
[0160]
[0161] Among them, v m W m and b m It is a learnable parameter related to omics m. This indicates the importance of omics m to sample i.
[0162] To give the attention scores between different omics a probabilistic meaning, this embodiment uses the Softmax function to normalize them, obtaining the weights of omics m with respect to sample i.
[0163] Where M represents the total number of omics, The normalized attention coefficient of omics m for sample i represents its contribution to the final representation.
[0164] S402: Generation of multi-omics latent representations.
[0165] After calculating the weight of each omics, this embodiment performs a weighted summation of the latent representations of each omics to generate the final multi-omics latent representation Z:
[0166]
[0167] Among them, z i ∈R d_hidden This represents the multi-omics ensemble representation of sample i. The final multi-omics representation matrix is obtained. It can be used for a variety of downstream analysis tasks, including cell clustering, data visualization, and differential gene expression analysis.
[0168] S50: Combine the intra-omics integration module and the inter-omics integration module to form a spatial multi-omics integration model. Use a dual learning strategy to train the spatial multi-omics integration model and realize spatial multi-omics integration based on the trained spatial multi-omics integration model. The dual learning strategy includes intra-omics private learning and inter-omics shared learning. Intra-omics private learning is used to guide the maintenance of the mapping relationship between the potential representation of multi-omics and the original normalized representation space, and introduces self-supervised contrastive learning to enhance the model's ability to identify omics-specific features. The goal of inter-omics shared learning is to impose constraints between the representation of one omics and the corresponding representation generated by another omics through the decoder-encoder path, thereby ensuring that data from different omics remain aligned in the shared space.
[0169] This model employs a dual learning strategy, aiming to achieve accurate alignment of shared information across omics while preserving the specific features of each omics' data. This dual learning strategy comprises two parts: private learning within an omics and shared learning between omics.
[0170] In some embodiments, step S50 can be implemented by the following steps S501-S502.
[0171] S501: Private learning within omics.
[0172] Intra-omics private learning aims to ensure that the fused multi-omics representation Z can share expression patterns across different omics while preserving information unique to each omics. To achieve this goal, this embodiment guides Z to maintain a mapping relationship with the original normalized expression space and introduces self-supervised contrastive learning (SCL) to enhance the model's ability to identify omics-specific features.
[0173] To further strengthen this mapping relationship, this embodiment designs an independent decoder for each omics, enabling the fused multi-omics representation Z to effectively reconstruct the original expression spectrum. Specifically, the model learns the mapping relationship from Z to the original feature space, ensuring that while fusing multi-omics information, it can still accurately capture the expression features of each omics. To achieve this goal, this embodiment introduces a learning process that minimizes the difference between the reconstructed features and the original features, thereby constraining the optimization of Z so that it can both retain the overall information integration capability and take into account the local omics specificity. The reconstructed expression at layer l (l∈{1,2,…,L-1,L}) is calculated using the following formula:
[0174]
[0175] in, and Let be the trainable weight matrix and bias vector, and σ be the activation function. To optimize the learning process, this embodiment defines the following reconstruction loss function to minimize the difference between the reconstructed features and the original features:
[0176]
[0177] Where, x i Represents the original features of sample i. This represents the corresponding reconstructed feature, where N is the total number of samples.
[0178] This optimization strategy ensures that while learning omics share information, the fused representation Z retains the unique expression patterns of each omics, providing more accurate feature representations for subsequent spatial multi-omics integration tasks. To enhance the spatial representation H... s To enhance information richness and discriminative ability, self-supervised contrastive learning (SCL) is introduced to ensure the model can capture local contextual information about spatial locations. Specifically, a graph neural network (GNN) encoder is used to extract spatial location information from the original graph G and the perturbed graph G', respectively, to generate a spatial representation matrix H. s ∈R N×d and H' s ∈R N×d Both the original and perturbation maps originate from the same ensemble and construct different data perspectives for representation learning. Because this constraint is applied within each ensemble, it ensures that spatial relationships within the same ensemble are more compact and consistent.
[0179] During the contrastive learning process, for each spatial location i, the local environment vector g is calculated by aggregating the average representations of its neighboring points. i This is to capture information about its microenvironment. In contrastive learning, it is defined as follows:
[0180] Positive sample: Representation of position i in the original graph Pair it with its local environment vector i;
[0181] Negative sample: Representation of position i in the perturbation graph The local environment vector g at position i in the original graph i pair.
[0182] This model maximizes the mutual information of positive sample pairs and minimizes the mutual information of negative sample pairs, making the representations of adjacent spatial locations more similar while increasing the discriminative power between non-adjacent locations. This effectively captures the local structure of spatial data and improves the consistency and discriminative power of spatial representations within omics. The loss function for self-supervised contrastive learning (SCL) is defined as follows:
[0183]
[0184] Where X and A represent the values used to calculate H, respectively. s The normalized gene expression matrix and normalized adjacency matrix are given, and X' and A' correspond to H' generated from the perturbation graph. The function Φ(·) is a discriminator used to determine whether a sample pair is a positive match and outputs the probability that it is a positive sample pair.
[0185] To enhance the stability of the model, this embodiment introduces a symmetric contrastive loss function to the perturbation graph G':
[0186]
[0187] This symmetric loss ensures that the model learns spatial information not only from the original graph but also from the perturbed graph, thereby enhancing the robustness of the spatial representation. This contrastive learning strategy promotes closer representations of adjacent spatial locations while preserving the differences in representations of non-adjacent locations, effectively capturing the local structure and biological characteristics of spatial data.
[0188] The final intra-omics loss function combines the reconstruction loss with the omics-specific contrastive learning loss, and is defined as follows:
[0189]
[0190] λ1 and λ2 are hyperparameters used to balance the contributions of feature reconstruction and expression discriminability, ensuring that the model achieves a good balance between information fusion and expression preservation.
[0191] S502: Inter-omics shared learning.
[0192] To ensure alignment consistency among different omics within a shared latent space, this embodiment designs an inter-omics learning mechanism, including consistency alignment learning and BarlowTwins contrastive learning, to improve the quality of feature representations in the shared space. Although models can learn omics-specific representations through intra-omics optimization, applying a learning strategy to each omics individually does not guarantee that their representation manifolds remain consistent in the shared space. Due to differences in measurement techniques and biological characteristics, direct integration may lead to misalignment between feature distributions of different omics, thus affecting the effectiveness of data integration. To address this issue, this embodiment introduces an inter-omics consistency learning strategy. Its core objective is to impose constraints between the representation of one omics and the corresponding representation generated by another omics through a decoder-encoder path, thereby ensuring that data from different omics remain aligned in the shared space. Specifically, let Y be the latent representation extracted from one omics, and Y′ be the corresponding representation generated through the decoder-encoder path of another omics. By optimizing this mapping relationship, the consistency between the original representation and the cross-omics generated representation is enhanced, thereby improving the semantic alignment capability between omics. This process can be described as follows:
[0193]
[0194] in, and Let represent the representation of the i-th sample in omics 1 and omics 2, respectively. and These are cross-omics representations generated via the paths of Omics 2 and Omics 1, respectively. W d and W e These are the trainable parameters for the decoder and encoder, respectively. By optimizing the above process, the mapping relationship between different omics in the shared space can be kept stable, thereby improving the reliability of multi-omics data integration.
[0195] To further enhance alignment in the shared omics space, this embodiment employs the BarlowTwins contrastive learning method. This method promotes consistency in the latent representations of different omics while reducing information redundancy, thereby ensuring the effectiveness of cross-omics integration. Let the latent representations of the two omics groups be... and Its contrastive loss function is defined as follows:
[0196]
[0197] in, Let represent the cross-correlation matrix between the latent representations of two omics modalities, where N is the number of samples and d is the feature dimension. ii (Diagonal elements) are constrained to approach 1 to ensure consistency of corresponding features; while C iiOff-diagonal elements are constrained to approach 0 to reduce redundant information between modes. γ3 is a hyperparameter used to balance diagonal and off-diagonal terms.
[0198] To further enhance the alignment effect, this embodiment introduces a two-way comparison strategy, namely, swapping. and Given the position, calculate the mapping relationship in both directions and take their average. The final optimization objective is:
[0199]
[0200] Here, β is a weighted parameter that controls the strength of omics alignment. In the early stages of training, cross-omics relationships are relatively simple, and a fixed β can be used. As training progresses, the relationships between omics gradually become more complex, and β can be dynamically adjusted to finely control the alignment process. This design ensures that the model focuses primarily on learning omics-specific features in the early stages, while gradually enhancing cross-omics fusion in the later stages, ultimately improving the accuracy of multi-omics data integration.
[0201] In some embodiments, the dual learning strategy in S50 also includes multi-omics balancing learning. In multi-omics learning, different omics (such as gene expression and phenotypic features) each have independent optimization objectives, which may be inherently different, leading to gradient conflicts. Traditional methods typically use a fixed-weight approach to weighted summation of intra-omics and inter-omics losses. However, this fixed-weight strategy may cause the gradient of one omics to dominate the model's parameter updates during training, thereby inhibiting the effective learning of other omics. Furthermore, the gradient directions of intra-omics and inter-omics objectives may contradict each other, causing model parameter updates to deviate from the optimal solution and affecting the model's ability to learn omics-specific features and cross-omics joint representations. To address these issues, this embodiment proposes a dynamic balancing strategy for multi-omics learning. This strategy can adjust the contributions of different loss terms in real time, ensuring that the model enhances the alignment ability between different omics in the shared space while learning omics features.
[0202] Multi-omics balanced learning includes the steps of calculating gradients within and between omics and calculating the total loss function.
[0203] When calculating the gradients within and between omics, we first calculate the gradient of the omics-specific loss function, denoted as . (in and Let represent the gradients of omics 1 and omics 2, respectively. Simultaneously, the gradient of the joint loss between omics is calculated, denoted as . To assess the consistency between intra- and inter-omics gradient directions, cosine similarity is introduced as a metric, defined as:
[0204]
[0205] When cosβ≥0 (i.e., the gradient directions are consistent), this embodiment assigns equal weights to the intra-omics loss and the inter-omics loss:
[0206]
[0207] When cosβ < 0 (i.e., gradient direction conflict), to alleviate this conflict, this embodiment dynamically adjusts the weights of each loss by minimizing the combined gradient norm. At this point, the optimal weighting coefficients for the intra-omics loss and the inter-omics loss are respectively:
[0208]
[0209] This strategy ensures that the optimization direction remains coordinated across different tasks during multi-omics training, enhancing the model's ability to simultaneously learn omics-specific expressions and shared features among omics.
[0210] When calculating the total loss function, dynamically calculated weights are used. The final total loss function is jointly determined by the intra-omics loss and the inter-omics loss, specifically expressed as follows:
[0211]
[0212] Among them, L multi_omics L represents the joint loss between omics (such as contrast loss, correlation loss). omics This represents the loss specific to each omics (such as reconstruction loss, classification loss). and These are the dynamically adjusted weights for the inter-omics and intra-omics losses of sample i, respectively.
[0213] By dynamically adjusting the weights between internal and external omics losses based on gradient similarity, this method effectively coordinates the learning processes of different omics, alleviates gradient conflicts, and ensures more stable optimization of shared parameters. This strategy enhances the ability of feature fusion in multi-omics spaces, strengthens cross-omics consistency while maintaining omics-specific learning, and thus significantly improves the overall performance of the model.
[0214] During training on all datasets, this embodiment uses a learning rate of 0.0001, β of 0.5, λ2, γ1, γ2, and γ4 of 1, and γ3 of 0.005. To accommodate differences in feature distribution across datasets, this embodiment assigns a specific loss weight factor λ1 to each dataset based on experience. Specifically, λ1 is set to 120 for the SpOTS mouse spleen dataset; 30 for the 10x Genomics Visium human lymph node dataset; 60 for the Stereo-CITE-seq mouse thymus dataset; and 50 for the spatial epigenomics-transcriptomics combined mouse brain tissue dataset. These weight factors are used to adjust the contribution ratio of each loss term to the total loss function, ensuring that the model can effectively capture multi-omics feature information on different datasets.
[0215] Example 2: Performance of SpaBalance on a Tri-omics Integrated Dataset
[0216] To verify SpaBalance's integration capabilities on a three-omics spatial dataset, this embodiment conducted performance tests on a simulated dataset containing three omics. This embodiment initially selected seven representative comparison methods—STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, and PRAGA—covering spatial single-omics integration tools, non-spatial multi-omics integration tools, and spatial multi-omics integration tools. This comprehensive benchmarking framework provides sufficient evidence for evaluating SpaBalance's performance and applicability in cross-omics data integration. The dataset contains real-world annotation information, providing a reliable reference for model evaluation. To comprehensively and objectively evaluate the integration effects of each method, this embodiment employed various supervised evaluation metrics, including Normalized Mutual Information (NMI), Adjusted Mutual Information (AMI), Fowlkes–Mallows Index (FMI), Adjusted Rand Index (ARI), V-measure, and Completeness. In addition, this embodiment also uses the unsupervised Jaccard similarity index to quantitatively analyze the intra-cluster compactness and inter-cluster separation of the clustering results.
[0217] The results are as follows Figure 3 As shown, Figure 3The figures below show the experimental results of SpaBalance integrating three omics datasets provided in this application embodiment. Figure a shows the spatial distribution of the ground truth labels. Figure b shows the spatial distribution of the three omics datasets, displayed from left to right: the original data for each omics, and the clustering results from single-cell and spatial multi-omics integration methods—including STAGATE, SeuratWNN, TotalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance. The labels shown correspond to the SpaBalance clustering results; due to the lack of uniformity in clustering color codes across different methods, the colors may not represent the same tissue structure. Figure c shows the clustering visualization results of STAGATE, SeuratWNN, TotalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance in the integrated embedded space. Figure d shows the quantitative evaluation results of STAGATE, SeuratWNN, TotalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance. e is a heatmap of intra-cluster compactness and inter-cluster separability calculated by SpaBalance based on unsupervised Jaccard similarity on a simulated tri-omics dataset.
[0218] Single-omics visualization results ( Figure 3 b) shows that RNA is primarily associated with factors 1, 2, and 3, while ADT and ATAC provide supplementary information for factors 1, 2, 3, and 4. Regarding the recovery of spatial factors ( Figure 3 In section b), both SpaBalance and SpatialGlue clearly captured all four spatial factors. However, compared to SpatialGlue, SpaBalance showed a higher degree of matching between spatial factors and true labels, demonstrating superior factor identification accuracy. Among other methods, PRAGA recovered factors 1, 2, and 4 relatively accurately, but its identification of factor 3 was still insufficient; totalVI performed slightly worse than PRAGA, mainly recovering factors 2 and 3. SeuratWNN, MultiVI, and MOFA+ could identify some factors, but their high noise levels affected the accurate identification of factors; STAGATE performed the weakest in spatial factor recovery, failing to accurately reconstruct the true distribution of spatial factors. Overall, SpaBalance demonstrated superior spatial factor recovery capabilities in multi-omics data integration, providing a more precise analytical tool for spatial omics research.
[0219] Clustering visualization results ( Figure 3c) shows that SpaBalance performs best overall, with compact and clear clustering results, distinct and complete category boundaries, and accurate preservation of spatial structure and omics specificity, demonstrating excellent multi-omics integration capabilities and spatial consistency. SeuratWNN and PRAGA clustering results are relatively sparse, with significant overlap between categories and poor separation, failing to accurately capture complex spatial relationships in multi-omics data. SpatialGlue clustering results show some separation between categories, but samples within categories are scattered, with numerous outliers, and spatial consistency is inferior to SpaBalance. STAGATE can capture local spatial topology, but it has limitations in integrating multi-omics information, exhibiting abnormal cluster shapes, making it more suitable for single-omics spatial data analysis than multi-omics integration tasks. MOFA+, MultiVI, and totalVI clustering results are even more scattered, with blurred category boundaries and significant overlap, failing to effectively integrate multi-omics and spatial information, and their overall integration performance is significantly lower than frameworks specifically designed for spatial multi-omics tasks.
[0220] In quantitative assessment ( Figure 3 In comparison with methods such as SeuratWNN, MOFA+, MultiVI, TotalVI, STAGATE, PRAGA, and SpatialGlue, SpatialGlue performed best, with scores of 0.965 for NMI, 0.965 for AMI, 0.979 for FMI, 0.973 for ARI, 0.965 for V-measure, and 0.965 for Completeness(comp). Compared to SpatialGlue, SpaBalance improved by 3.5%, 3.498%, 2.1%, 2.7%, 3.499%, and 3.5% in the above metrics, respectively. These results indicate that the clustering distribution obtained by SpaBalance is highly consistent with the true distribution and characteristics of the simulated data. Furthermore, simulated data testing further validated the scalability of SpaBalance, demonstrating its suitability for integrating three or more modalities in multi-omics tasks.
[0221] Example 3: Performance of SpaBalance on the Human Lymph Node Dataset
[0222] To evaluate the effectiveness of SpaBalance in integrating spatial transcriptomics and proteomics data and to verify its contribution to multi-omics balance learning mechanisms, this embodiment conducted systematic benchmarking and ablation studies on a real human lymph node dataset. This dataset, generated by the 10x Genomics Visium platform, originates from region A1 and contains both RNA and proteomics data, along with authentic biological annotations, providing a reliable reference for model evaluation.
[0223] The results are as follows Figure 4 As shown, Figure 4 This diagram illustrates how SpaBalance accurately identifies spatial domains in human lymph node A1, as provided in this embodiment. In the diagram, a represents a manually annotated human lymph node sample A1. b represents the clustering space diagram of single-cell RNAomics and proteomics in sample A1. c represents the clustering space diagram of various spatial multi-omics integration methods in sample A1, including STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, SpaBalance_withoutMBL (i.e., SpaBalance without the multi-omics balancing learning module), and the complete SpaBalance model. It should be noted that the clustering colors of different methods do not necessarily correspond to the same tissue structure. d represents the clustering visualization results of STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, SpaBalance_withoutMBL, and SpaBalance in the final representation space. e represents the quantitative evaluation results of the above methods. f is a heatmap of intra-cluster compactness and inter-cluster separability calculated by SpaBalance based on unsupervised Jaccard similarity on lymph node sample A1.
[0224] As shown in the visualization results of single-omics features ( Figure 4 (b) Both omics approaches were able to identify follicular regions and parts of the cortex, but RNA omics showed better resolution, enabling more detailed identification of myelin cord structures. In the comparison of integration and embedding results ( Figure 4(c) While STAGATE exhibits some clustering ability, its clustering clarity and accuracy are relatively limited. In contrast, SeuratWNN, MOFA+, MultiVI, PRAGA, SpatialGlue, and SpaBalance can all accurately identify follicular regions, while totalVI fails to effectively identify this region. Regarding cortical region identification, MOFA+ and PRAGA tend to overestimate, while SeuratWNN, MultiVI, totalVI, and SpatialGlue generally underestimate their boundary range. Notably, SeuratWNN has low overall clustering clarity, while SpaBalance demonstrates higher accuracy in cortical region boundary delineation, showing a clear advantage. Furthermore, PRAGA, SpatialGlue, and SpaBalance can all accurately identify pericapsular adipose tissue, with SpaBalance showing the best identification performance. Comparatively, PRAGA fails to accurately identify the medullary cord region, while totalVI cannot effectively identify medullary sinuses. Overall, SpaBalance outperforms all other methods in terms of accuracy in multi-region identification, especially in the identification of cortical areas and pericapsular adipose tissue, demonstrating excellent integration capabilities.
[0225] Comparison of clustering results from eight multi-omics data fusion methods ( Figure 4 The methods (d) include SeuratWNN, MOFA+, MultiVI, PRAGA, SpaBalance, STAGATE, totalVI, and SpatialGlue. It can be seen that SpaBalance exhibits optimal performance by flexibly integrating different omics while maintaining strong class separation. Unlike other methods, SpaBalance achieves balanced integration of multi-omics data, preserving detailed information between omics while avoiding excessive cluster overlap, thus achieving the best integration effect. In comparison, while MOFA+ and MultiVI have relatively clear cluster boundaries, they lack sufficient flexibility in multi-omics integration; and PRAGA clustering results have more overlap, making it less suitable for tasks requiring high clarity of cluster boundaries. Furthermore, this embodiment also uses the unsupervised Jaccard similarity index to quantitatively analyze the intra-cluster compactness and inter-cluster separation of the clustering results. Figure 4(f). The results showed good cohesion within each cluster, indicating high consistency in sample grouping within the same cluster. Regarding inter-cluster separation, the separation scores for cluster 1 and cluster 3, and cluster 3 and cluster 6 were relatively low, possibly due to the spatial proximity of cells in these regions. However, the separation between the remaining cluster pairs was high, further validating the spatial discriminativeness and biological rationale of the clustering results.
[0226] Using hematoxylin-eosin (H&E) staining as a true reference ( Figure 4 (a) In this embodiment, eight integration tools were quantitatively evaluated. Compared with the quantitative results of SeuratWNN, MOFA+, MultiVI, totalVI, STAGATE, PRAGA, and SpatialGlue, PRAGA performed best, with scores of: NMI 0.384, AMI 0.386, FMI 0.426, ARI 0.275, V-measure 0.347, and Completeness (comp) 0.394. In contrast, SpaBalance outperformed PRAGA on all indicators, improving by 2.5%, 2.1%, 3.9%, 4.1%, 3.5%, and 1.7%, respectively, further demonstrating its advantage in multi-omics data integration.
[0227] To evaluate the role of the multi-omics balanced learning mechanism, an ablation experiment was conducted in this embodiment, and the results are as follows: Figure 5 As shown. Figure 5 Supplementary results figures for the ablation experiment of human lymph node A1 using SpaBalance provided in the embodiments of this application. Figure a shows the trend of the loss function of RNA and ADT omics during SpaBalance training. Figure b shows the trend of the loss function of RNA and ADT omics during SpaBalance training without using multi-omics balanced learning. Figure c shows the trend of the accuracy of RNA, ADT, and their fusion representations during SpaBalance training. Figure d shows the trend of the accuracy of RNA, ADT, and their fusion representations during SpaBalance training without using multi-omics balanced learning.
[0228] The results show that, in the absence of a balanced learning rate among omics, the convergence of multi-omics data is uneven during training, leading to significant fluctuations in the loss of each omics. Figure 5 (b) ultimately led to a decrease in model accuracy of 2.4%, 3.4%, 5.4%, 3.7%, 2.3%, and 2.6% across multiple evaluation metrics. Further analysis of the accuracy trends for single-omics and integrated representations revealed that the overall accuracy declined when a multi-omics balancing learning mechanism was not introduced. Figure 5c); and after adding this mechanism, the accuracy rate showed a significant upward trend. Figure 5 (d). The above results strongly demonstrate that the multi-omics balancing learning mechanism can effectively improve integration accuracy, achieve optimized balance across omics, and improve the integration effect of multi-omics data.
[0229] To ensure the results are not biased by specific tissue sections, this embodiment further conducted experiments on a real human lymph node dataset derived from the D1 region. Figure 6 The image shows the integrated result of human lymph node sample D1 provided in this embodiment of the application. Wherein a is the manually annotated human lymph node sample D1. b is a spatial clustering visualization of RNA and proteomics for sample D1, and the clustering results on sample D1 using spatial multi-omics integration methods (STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance). Note: Clustering colors in different methods may not correspond to the same tissue structure. c is the final clustering visualization result generated by STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance. d is a heatmap of intra-cluster compactness and inter-cluster separability on lymph node sample D1 using SpaBalance, calculated based on unsupervised Jaccard similarity. Figure 6 It can be seen that SpatialGlue, PRAGA, and SpaBalance achieve clearer spatial representations compared to other methods. Notably, PRAGA and SpaBalance outperform SpatialGlue in terms of localization accuracy in cortical regions. Furthermore, compared to SpatialGlue and PRAGA, SpaBalance performs better in accurately identifying B-cell follicular regions and exclusion regions, highlighting its stronger spatial recognition capability.
[0230] Example 4: Performance of SpaBalance on Coronal Sections of Mouse Brain Tissue
[0231] To further validate the performance of SpaBalance and its ability to preserve key omics-specific information during integration, this example was tested on a coronal slice dataset of mouse brains from day 22 (P22). This dataset, generated by spatial ATAC-RNA sequencing technology, contains 9,215 cells and covers both mRNA expression data and open chromatin regions. Due to the high complexity of this dataset and the difficulty of integration, this experiment provides a more challenging testing scenario for evaluating the robustness and applicability of SpaBalance in cross-omics data integration.
[0232] like Figure 7 The diagram illustrates how SpaBalance, provided in this application, resolves spatial epigenomic and transcriptomic samples of the mouse brain at higher resolution. Figure a shows a reference diagram with annotations of Allen's mouse brain atlas on coronal sections of the mouse brain. Figure b shows unipolar clustering (left) and spatial clustering results (right) based on RNA-seq and ATAC-seq data using single-cell and spatial multi-omics integration methods (MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance). The annotations correspond to the SpaBalance results; clustering colors in different methods do not necessarily represent the same structures. The abbreviations in the legend are as follows: ctx represents the cerebral cortex, cp represents the caudate putamen, vl represents the lateral ventricle, lpo represents the lateral preoptic area, aca represents the anterior cingulate cortex, ls represents the lateral septal nucleus, aco represents the olfactory commissure, acb represents the nucleus accumbens, and cc represents the corpus callosum. Figure c shows the clustering visualization results of MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance in the final spatial representation. d shows the heatmap of intra-cluster compactness and inter-cluster separability of SpaBalance based on unsupervised Jaccard similarity calculation. e shows the quantitative analysis results of MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance under the ATAC tag. f shows the quantitative analysis results of MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance under the RNA tag.
[0233] This embodiment uses the Allen brain atlas as a reference. Figure 7In section a), the following brain regions were labeled: cerebral cortex (ctx), caudate putamen (cp), lateral ventricle (vl), lateral preoptic area (lpo), anterior cingulate area (aca), lateral septal nucleus (ls), olfactory commissure (aco), nucleus accumbens (acb), and corpus callosum (ccg).
[0234] In benchmark tests, since totalVI is only applicable to CITE-seq data and STAGATE experienced memory overflow issues under the same experimental conditions, this embodiment selected MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance for comparison. This embodiment first analyzes the visualization results of single-omics data. Figure 7 (b) The study found that different omics approaches exhibited varying degrees of accuracy in identifying different brain regions. Results showed that RNA omics data clearly identified the corpus callosum (CCG) and the olfactory commissure (ACO), and captured the caudate putamen (CP) relatively well; while ATAC omics could roughly reflect the spatial structure of the lateral ventricle (VL) and the lateral preoptic area (LPO). In the comparison of the performance of different methods ( Figure 3 (b) It was found that all methods could effectively identify the corpus callosum (CCG) and the anterior visual area (LPO), but MOFA+ had lower clarity in identifying the corpus callosum. Except for PRAGA, all other tools could identify the olfactory association area (ACO). Overall, MultiVI and MOFA+ have certain limitations in spatial clarity and detail representation. In contrast, SpaBalance and SpatialGlue performed exceptionally well, accurately identifying all key anatomical regions, and their performance was superior to non-spatial multi-omics integration tools after integrating spatial information. Notably, SpaBalance provided more complete identification of the caudate putamen (CP) region. Although PRAGA could identify all regions except the olfactory association area (ACO), its clustering diagram ( Figure 7c) shows a high degree of overlap between omics. This embodiment further employs the unsupervised Jaccard similarity index to quantify the intra-cluster compactness and inter-cluster separation of the clustering results. Figure 7 d), the results show that the performance within and between clusters is generally good, demonstrating good grouping consistency and separability in spatial structure capture.
[0235] Under the ATAC_clusters tag ( Figure 7 Compared to MOFA+, MultiVI, TotalVI, PRAGA, and SpatialGlue, PRAGA performed best across multiple metrics, scoring 0.336, 0.376, 0.289, 0.331, and 0.329 for NMI, FMI, ARI, V-measure, and Completeness (comp), respectively. SpatialGlue achieved the highest score for AMI at 0.326. In contrast, SpaBalance improved by 1.3%, 2%, 2%, 1.6%, 1.8%, and 2.1% across NMI, AMI, FMI, ARI, V-measure, and comp, respectively, ranking first overall. Under the RNA_clusters tag ( Figure 4 f) Among MOFA+, MultiVI, TotalVI, and PRAGA, SpatialGlue performed best, with the following metrics: NMI 0.382, AMI 0.380, FMI 0.353, ARI 0.254, V-measure 0.382, and comp 0.331. Compared to SpatialGlue, SpaBalance improved NMI by 1.4%, AMI by 0.6%, FMI by 0.8%, ARI by 1.4%, V-measure by 0.9%, and comp by 0.6%.
[0236] This embodiment further extends the analysis to another dataset from the P22 mouse brain, which is also a coronal section, similar to the aforementioned mouse brain data, but includes two omics approaches: RNA-seq and CUT&Tag. The CUT&Tag data is primarily used to capture H3K27ac histone modification information. This modification is an acetylation modification at lysine 27 of histone H3, an important epigenetic marker closely related to active enhancer regions and gene transcriptional activation. Since this dataset also lacks real-world labeling, this embodiment again uses the Allen brain atlas as a spatial reference to label and compare anatomical regions. Figure 8The diagram illustrates the SpaBalance analysis of more complex mouse brain spatial epigenomic-transcriptomic samples provided in this embodiment of the application. Figure a shows the single-omics clustering results (left) and spatial clustering results (right) of single-cell and spatial multi-omics integration methods based on RNA-seq and CUT&Tag-seq (H3K27ac) data, covering methods such as MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance. The annotation labels in the figure correspond to the SpaBalance clustering results; clustering colors may not represent the same tissue structures in different methods. Figure b is a reference annotation diagram of the Allen mouse brain atlas for coronal sections of the mouse brain. Figure c shows the clustering visualization results of MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance in the final representation space. Figure d is a heatmap of differentially expressed genes (DEG) in each cluster unit. Figure e is a heatmap of SpaBalance intra-cluster compactness and inter-cluster separability based on unsupervised Jaccard similarity calculation.
[0237] From single-origin images ( Figure 8 As shown in a), RNA data can identify the approximate regions of the corpus callosum (CCG) and olfactory association area (ACO), while H3K27ac data reflects the lateral ventricle (VL) region. In this dataset, all six tools can identify the CCG, ACO, and VL regions, but PRAGA, MOFA+, and MultiVI have relatively low overall image clarity. SpaBalance successfully identified the caudate putamen (CP, regions 6 and 9), nucleus accumbens (ACB, region 13), and lateral septal nucleus (LS, region 1). Although SpatialGlue can also identify these regions, it has difficulty distinguishing between the ACB and CP regions and has relatively high overall image noise.
[0238] Based on the SpaBalance clustering results ( Figure 8 e) Unsupervised Jaccard similarity index analysis was performed. This embodiment observed significant inter-cluster differences even among clusters classified into the same tissue region. This phenomenon suggests that although these clusters are spatially adjacent, they may have underlying molecular-level differences. Further analysis of differentially expressed genes (DEGs) within each cluster was conducted. Figure 8(d) The study found that multiple subclusters within the cerebral cortex (ctx) region—including clusters 2, 5, 7, 9, 10, 12, 13, and 15—exhibited significant differences in gene expression profiles. This finding reveals a high degree of spatial heterogeneity in the cortical region, indicating that it is not a homogeneous region at the transcriptional level, but rather composed of multiple functional partitions with distinct molecular characteristics. This spatial heterogeneity provides important molecular-level evidence for understanding the specific functions of cortical regions and lays the foundation for further research into the regulatory mechanisms behind cortical functional differentiation.
[0239] Notably, clusters 4, 8, and 14 in the caudate putamen (cp) region exhibit a clear differentiation between two highly expressed gene modules. Module 1 includes Scn4b, Syndig1, Lrrc10b, Nexn, and Dach1. Scn4b encodes sodium channel auxiliary subunit 4, which is widely involved in the regulation of neurophysiological signals, while Syndig1 regulates synaptic AMPA receptor content and promotes the maturation of excitatory synapses. The high expression of this module suggests that the cp region may be enriched with excitatory neural networks involved in the regulation of electrical activity. Module 2 includes Rgs9, D830015G02Rik, Drd1a, Pde1b, Pde7b, and Rarb. Rgs9 encodes a GTPase activator protein that accelerates the inactivation of G protein signaling pathways; Drd1a is a classic dopamine D1 receptor involved in neurotransmitter signal transduction, synaptic structure regulation, and central nervous system development; Pde7b is a phosphodiesterase that specifically degrades cAMP and is significantly highly expressed in brain tissue. The significant expression of this module suggests that the cp region may play a key role in the regulation of dopamine-related signals, indicating that this region is enriched with dopaminergic neurons involved in motor control and reward pathways.
[0240] In the corpus callosum / olfactory association region (ccg / aco, cluster 11), a gene module consisting of Cldn11, Mal, Mobp, Gsn, Mog, Adamts4, Pllp, Cryab, and Enpp6 was identified. Most of these genes are closely related to myelination and glial cell function. For example, Mobp is a marker gene for oligodendrocytes and is associated with multiple sclerosis; Enpp6 is a lipid hydrolase involved in glial cell metabolism. High expression of this module suggests that the ccg / aco region may be enriched with mature oligodendrocytes, further supporting their crucial role in white matter conduction pathways. In the lateral septal nucleus (ls, cluster 3) region, high expression of Zic1, Trpc4, and Dgkg was observed in this embodiment. Zic1 is a classic transcription factor involved in nervous system development, playing a crucial role, particularly in forebrain formation; Trpc4 belongs to the TRP calcium channel family and participates in the regulation of neuronal excitability; Dgkg is involved in phospholipid signaling pathways and plays an important role in early embryonic neural tube development. These genes suggest that the ls region may actively participate in neural development and the regulation of neuronal excitability. Furthermore, Top2a showed significantly high expression in the lateral ventricle (vl, cluster 16) region, suggesting it may be a potential marker gene for this region. Top2a encodes DNA topoisomerase II, a key molecule indispensable for cell proliferation and chromatin remodeling; its high expression suggests that the vl region may be rich in highly proliferative cells, indicating that this region may aggregate neural stem cells or progenitor cells.
[0241] Example 5: SpaBalance integrates mouse thymus and spleen datasets
[0242] This embodiment applies SpaBalance to a mouse thymus tissue sample dataset. This dataset, obtained through Stereo-CITE-seq technology, combines mRNA and proteomics information, providing a multi-dimensional analytical perspective for thymus tissue. The thymus is a crucial immune organ for T cell development and differentiation, consisting of two lobes, each divided into a central medullary region and a peripheral cortical region. The cortical region primarily handles positive selection of immature T cells, while the medullary region participates in negative selection to eliminate T cells that respond to their own antigens.
[0243] like Figure 9The figure shows a schematic diagram of SpaBalance integrating mouse thymus (based on RNA and protein data obtained from Stereo-CITE-seq) according to an embodiment of this application. Figure a shows a spatial map of RNA and protein data obtained from Stereo-CITE-seq in mouse thymus tissue, including single-omics clustering results (left) and a comparison with clustering results from multiple single-cell and spatial multi-omics integration methods (right)—including STAGATE, SeuratWNN, TotalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance. The annotation labels in the figure correspond to the clustering results of SpaBalance; the clustering colors in different methods may not represent the same tissue structure. Figure b shows the visualization results of the integrated embedding spatial clustering of STAGATE, SeuratWNN, TotalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance in mouse thymus. Figure c shows a heatmap of intra-cluster compactness and inter-cluster separability calculated by SpaBalance based on unsupervised Jaccard similarity in mouse spleen samples.
[0244] Visual analysis of RNA and proteomics data ( Figure 9 (a) Clearly, the outer cortex and connective tissue capsule (containing fibroblasts, erythrocytes, and myeloid cells) exhibit distinct spatial distribution characteristics in both omics studies, while RNA omics data further reveals the medullary region. This example compares SeuratWNN, TotalVI, MultiVI, MOFA+, STAGATE, SpatialGlue, PRAGA, and SpaBalance, revealing significant differences in their ability to resolve the spatial structure of the thymus. MultiVI and STAGATE failed to accurately identify the bilobed structure of the thymus, especially MultiVI, which performed poorly in distinguishing the outer cortex from the medulla. SeuratWNN could identify some connective tissue capsule regions, while TotalVI could capture the outer cortex, medulla, and cortico-medullary junction, but both still lacked sufficient spatial resolution. MOFA+, PRAGA, and SpatialGlue performed well in separating the medulla and cortex, and were able to identify the bilobed structure of the thymus. However, MOFA+ and SpatialGlue were not clear enough in distinguishing between inner cortical region 1 and middle cortical region 2, while PRAGA showed regional overlap between inner cortical region 1 and the cortico-medulloeal junction. In contrast, SpaBalance performed better in spatial structure division, accurately distinguishing between the inner cortex and central medulla, with clear region boundaries, and effectively avoiding cell type mixing problems caused by over-clustering. Its clustering results ( Figure 9(b) exhibits high stability and precision, capable of capturing local details in complex spatial structures while maintaining the independence of cell populations. Unsupervised Jaccard similarity heatmap ( Figure 9 c) further demonstrates that the model performs well in terms of both intra-cluster compactness and inter-cluster separability.
[0245] This embodiment also conducted experiments on two other mouse thymus datasets generated using Stereo-CITE-seq technology.
[0246] like Figure 10 The diagram shows the integration results of the mouse thymus 1 and mouse thymus 2 datasets provided in this embodiment of the application. Figure a shows the spatial clustering visualization of the RNA and proteomics of mouse thymus 1, along with the clustering results generated by spatial multi-omics integration methods (STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance). Note: Clustering colors in different methods may not correspond to the same tissue structures. Figure b shows the final representation clustering visualization of STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance on the mouse thymus 1 dataset. Figure c shows a heatmap of intra-cluster compactness and inter-cluster separability of SpaBalance on the mouse thymus 1 dataset, calculated based on unsupervised Jaccard similarity. Figure d shows the spatial clustering visualization of the RNA and proteomics of mouse thymus 2, along with the clustering results of the same integration methods described above. Figure e shows the final representation clustering visualization of all integration methods on the mouse thymus 2 dataset. f, Heatmap of intra-cluster compactness and inter-cluster separability of SpaBalance on the mouse thymus 2 dataset, based on unsupervised Jaccard similarity calculation.
[0247] Figure 10 The results showed that SpatialGlue, PRAGA, and SpaBalance demonstrated higher accuracy in spatial domain segmentation compared to non-spatial multi-omics integration methods. However, further analysis of RNA and ADT single-omics features revealed that SpatialGlue and PRAGA neglected key information derived from RNA in some regions, leading to a decrease in the accuracy of region segmentation. In contrast, SpaBalance performed exceptionally well in the analysis of thymus data, accurately capturing key anatomical structures and finely distinguishing complex cell types, fully demonstrating its excellent spatial resolution and multi-omics integration performance.
[0248] In another experiment, this embodiment benchmarked SpaBalance on a spatial analysis dataset of mouse spleens. This dataset was acquired using the SPOTS technique, which simultaneously captures transcriptomic and extracellular protein data. The spleen is an important organ in the lymphatic system, whose main functions include promoting B cell maturation in the germinal centers of B cell follicles. These structures are complex and contain various types of immune cells.
[0249] like Figure 11 The diagram shows a schematic of SpaBalance integrating mouse spleen data (RNA and protein data obtained based on SPOTS technology) according to an embodiment of this application. a) is a visualization of the integration embedding spatial clustering results of STAGATE, SeuratWNN, TotalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance on mouse thymus samples. b) is a visualization of the integration embedding spatial clustering results of STAGATE, SeuratWNN, TotalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance on mouse spleen samples. c) is a heatmap of differentially expressed antibody-derived tags (ADTs) in each cluster. d) is a heatmap of intra-cluster compactness and inter-cluster separability calculated by SpaBalance based on unsupervised Jaccard similarity on mouse spleen samples.
[0250] From the visualization results ( Figure 11 (a) SpaBalance can identify clearer spatial structural domains. Analysis of clustering results ( Figure 11 (b) MOFA+ showed poor clustering performance, with low separation between cell types and significant overlap. STAGATE's clustering results were chaotic, particularly in cell type differentiation, with significant overlap between MZMΦ and B cells. Although SeuratWNN and MultiVI improved clustering performance, significant overlap remained, especially in distinguishing RpMΦ, B cells, and T cells. TotalVI performed well in separating certain cell types, but its overall clustering performance was still unsatisfactory. PRAGA's clustering results were even more complex, with low separation between some cell populations (such as MZMΦ and B cells), resulting in weak overall clustering performance. SpatialGlue could distinguish some cell types relatively well, but its clustering results were still not as clear as SpaBalance, with some cell populations still overlapping. In contrast, SpaBalance could clearly distinguish different cell types (such as MZMΦ, MMMΦ, RpMΦ, B cells, and T cells), with high separation between clusters and minimal overlap, demonstrating superior cell type resolution capabilities.
[0251] When observing ADT biomarker heatmaps ( Figure 11 As shown in (c), different cell types exhibit specific gene expression patterns. These markers not only reveal the characteristics of cell types but also reflect their functions in the immune system. While there is some overlap in the highly expressed genes between marginal zone macrophages (MZMΦ) and red pulp macrophages (RpMΦ), significant differences remain in key markers. For example, CD68 and CD163 are expressed at higher levels in both MZMΦ and RpMΦ compared to marginal sinus macrophages (MMMΦ). CD163 is a highly specific marker for M2 tumor-associated macrophages, primarily expressed on the surface of monocytes and macrophages; while CD68 is closely related to the phagocytic activity of macrophages. These findings suggest that CD68 and CD163 could serve as potential molecular markers for differentiating MMMΦ. MMMΦ is characterized by high expression of CD19, CD169 (Siglec-1), and B220 (PTPRC, CD45R). CD19 and B220 are key markers involved in B cell maturation and activation, while CD169 plays an important role in phagocytosis of apoptotic cells and antigen presentation. These characteristics suggest that MMMΦ may be involved in the regulation of B cell-mediated immune responses. T cells significantly express CD3, CD4, and CD8, corresponding to the T cell receptor complex (CD3), helper T cells (CD4), and cytotoxic T cells (CD8), respectively, indicating their important roles in immune regulation, antigen recognition, and cytotoxic killing. In contrast, B cells did not show significant gene overexpression, possibly due to expression differences between subsets or low expression levels at rest. However, some B cell-related markers (such as CD19 and B220) can still be observed in specific cell populations, suggesting their role in humoral immunity.
[0252] To further verify the performance of SpaBalance, supplementary experiments were conducted on two additional mouse spleen sections in this embodiment. Figure 12The figures show the integration results of mouse spleen 1 and mouse spleen 2 datasets provided in this embodiment of the application. a) is a spatial clustering visualization of RNA and proteomics for mouse spleen 1, along with the clustering results generated by spatial multi-omics integration methods (STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance). Note: Clustering colors in different methods may not correspond to the same tissue structures. b) is the final representation clustering visualization of STAGATE, SeuratWNN, totalVI, MultiVI, MOFA+, SpatialGlue, PRAGA, and SpaBalance on the mouse spleen 1 dataset. c) is a heatmap of intra-cluster compactness and inter-cluster separability of SpaBalance on the mouse spleen 1 dataset, calculated based on unsupervised Jaccard similarity. d) is a spatial clustering visualization of RNA and proteomics for mouse spleen 2, along with the clustering results obtained using the same integration methods described above. e) is the final representation clustering visualization of all integration methods on the mouse spleen 2 dataset. f is a heatmap of intra-cluster compactness and inter-cluster separability of SpaBalance on the mouse spleen 2 dataset, based on unsupervised Jaccard similarity calculation.
[0253] Depend on Figure 12 It is evident that SpaBalance is significantly superior to other methods in capturing spatial structure, and its clustering results also demonstrate higher accuracy and resolution, highlighting its comprehensive advantages in multi-omics spatial integration.
[0254] These different cell types and their marker expression patterns not only demonstrate the diversity of cells in the immune system but also reflect the functions of specific genes within cells. For example, through the expression of differentially expressed markers, this application can more clearly understand the roles of various cell types in immune responses, inflammatory responses, and tissue repair. The expression intensity of each marker reflects the activity of the gene within the cell, further revealing the functional localization and contribution of cells in specific immune environments. Analysis of these markers allows for a deeper exploration of the specific manifestations and mechanisms of genes in regulating cellular function.
[0255] As illustrated in Examples 1 to 5 above, SpaBalance is an innovative spatial multi-omics balanced learning framework that effectively integrates data from different omics through a multi-omics balanced learning mechanism and a dual learning strategy. While achieving learning balance between omics, it preserves the specific information of each omics, thereby improving the analytical accuracy and biological interpretability of spatial multi-omics data. Leveraging its multi-omics balanced learning mechanism, SpaBalance adaptively learns and weighs the importance of different omics modalities, accurately capturing the dependencies between spatial location and omics features within each omics, giving it stronger expressive power and biological interpretability in spatial multi-omics analysis, providing a more comprehensive and accurate solution for multi-omics integration analysis. Its dual learning strategy aims to improve the learning accuracy of shared representations across omics while ensuring the preservation of omics-specific information. Specifically, this strategy includes shared learning between omics and specific learning within omics: shared learning aims to minimize the feature distance between omics, enhance inter-omics collaborative expression, and improve the consistency of cross-omics information; specific learning focuses on preserving the unique information of each omics, avoiding the loss of key omics-specific features during cross-omics integration. This strategy effectively prevents excessive fusion of information between omics, maintains the accuracy of multi-omics integration, and significantly improves the model's performance in multi-omics analysis. Furthermore, SpaBalance boasts a flexible architecture, demonstrating excellent scalability and adaptability. Its multi-level design divides the integration process into two stages: intra-omics integration and inter-omics integration. Intra-omics integration employs a lightweight network structure to enhance single-omics feature learning and spatial information embedding, while leveraging data augmentation techniques to improve model robustness and spatial feature capture capabilities. Inter-omics integration introduces a multi-head attention mechanism to effectively capture the complex interactions between different omics and spatial locations, further enhancing the model's ability to model local and global dependencies. This modular design not only supports seamless expansion of multi-omics data but also ensures compatibility with various spatial omics technology platforms. Compared to existing methods, SpaBalance, through its multi-omics balanced learning algorithm and modular architecture, significantly enhances the scalability and applicability of the model while improving the integration effect of spatial multi-omics data.
[0256] In integration experiments on human lymph node datasets, SpaBalance demonstrated superior performance compared to other tools, further validating the crucial role of multi-omics balancing learning mechanisms in optimizing learning weights among different omics and improving the integration effect of multi-omics data. On the largest test dataset (containing 9752 spatial points of mouse brain spatial epigenome-transcriptome data), the model ran on a server equipped with an NVIDIA RTX A6000 GPU. Results showed that SpaBalance could resolve more detailed cortical hierarchical structures compared to other methods. Furthermore, in quantitative analysis using RNA and ATAC omics-specific tags, SpaBalance's dual learning strategy effectively preserved key omics features of RNA and ATAC. Notably, in terms of computational efficiency, MOFA+, MultiVI, TotalVI, and PRAGA all required several hours to run, with MOFA+ and PRAGA taking the longest, reaching several hours; while STAGATE could not run due to excessive memory consumption. In contrast, SpaBalance and SpatialGlue can complete integration in just a few minutes, further validating SpaBalance's efficiency. Furthermore, in systematic benchmark tests using a three-omics simulation dataset, SpaBalance significantly outperformed six state-of-the-art single-omics and non-spatial methods in multi-omics data integration and spatial feature parsing tasks, fully validating its accuracy and effectiveness in spatial multi-omics integration tasks. Direct comparisons with existing spatial multi-omics methods (such as SpatialGlue and PRAGA) further demonstrate that SpaBalance can capture more complex histological structural features, significantly improving the analytical capabilities for complex spatial structures and providing a more accurate solution for multi-omics spatial omics research.
[0257] Example 6:
[0258] Figure 13 This is a schematic diagram of the spatial multi-omics integration device provided in an embodiment of this application. This application also provides a spatial multi-omics integration device. Figure 13 As shown, this space-based multi-omics integration device includes:
[0259] The data acquisition module 1301 is configured to acquire a dual-omics dataset; wherein the dual-omics dataset includes first omics data and second omics data;
[0260] The graph feature construction module 1302 is configured to construct an adjacency graph based on the dual-omics dataset; wherein the adjacency graph includes a spatial graph and a feature graph;
[0261] The intra-omics integration module 1303 is configured to construct an intra-omics integration module. The intra-omics integration module responds to the input spatial graph and feature graph by performing graph convolution operations to obtain spatial latent representation and feature latent representation, concatenates the spatial latent representation and feature latent representation on the feature dimension to obtain concatenated features, and inputs the concatenated features into a multilayer perceptron to obtain an intra-omics fusion representation.
[0262] Inter-omics integration module 1304 is configured to construct an inter-omics integration module. The inter-omics integration module responds to the input first intra-omics fusion representation and second intra-omics fusion representation. The first intra-omics fusion representation and the second intra-omics fusion representation are output through the intra-omics integration module. A multi-head attention mechanism is applied to the first intra-omics fusion representation and the second intra-omics fusion representation to obtain the attention weights of each intra-omics fusion representation. The attention weights and the corresponding intra-omics fusion representations are weighted and summed to obtain a multi-omics latent representation.
[0263] The learning and training module 1305 is configured to combine the intra-omics integration module and the inter-omics integration module to form a spatial multi-omics integration model. A dual learning strategy is employed to train the spatial multi-omics integration model, and spatial multi-omics integration is achieved based on the trained spatial multi-omics integration model. The dual learning strategy includes intra-omics private learning and inter-omics shared learning. Intra-omics private learning guides the maintenance of a mapping relationship between the multi-omics latent representation and the original normalized representation space, and introduces self-supervised contrastive learning to enhance the model's ability to identify omics-specific features. The goal of inter-omics shared learning is to impose constraints between the representation of one omics and the corresponding representation generated by another omics through a decoder-encoder path, thereby ensuring that data from different omics remain aligned in the shared space.
[0264] The spatial multi-omics integration device provided in this application embodiment can be used to execute the technical solution of the spatial multi-omics integration method in the above embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0265] This application provides an electronic device. The electronic device may include a processor and a memory, wherein the processor and the memory can communicate; exemplarily, the processor and the memory communicate via a communication bus.
[0266] The processor executes computer execution instructions stored in memory, causing the processor to perform the schemes in the above embodiments. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0267] Communication buses can be Peripheral Component Interconnect (PCI) buses or Extended Industry Standard Architecture (EISA) buses, etc. System buses can be divided into address buses, data buses, control buses, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. Transceivers are used to enable communication between database access devices and other computers (e.g., clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) and may also include non-volatile memory.
[0268] The electronic device provided in this application embodiment can be the terminal device described in the above embodiments.
[0269] This application also provides a computer-readable storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the technical solution of the spatial multi-omics integration method described above.
[0270] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, it can implement the technical solution of the spatial multi-omics integration method in the above embodiments.
[0271] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0272] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.
[0273] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit composed of the above modules can be implemented in hardware or in the form of hardware plus software functional units.
[0274] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0275] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0276] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0277] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, the bus in this application is not limited to a single bus or a single type of bus.
[0278] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.
[0279] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic control unit or main control device.
[0280] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0281] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A spatial multi-omics integration method, characterized in that, include: Obtain a dual-omics dataset; wherein the dual-omics dataset includes first omics data and second omics data; Based on the aforementioned dual-omics dataset, an adjacency graph is constructed; wherein, the adjacency graph includes a spatial graph and a feature graph; An intra-omics integration module is constructed. In response to the input spatial graph and feature graph, the intra-omics integration module performs graph convolution operation to obtain spatial latent representation and feature latent representation respectively. The spatial latent representation and feature latent representation are concatenated on the feature dimension to obtain concatenated features. The concatenated features are input into a multilayer perceptron to obtain intra-omics fusion representation. An inter-omics integration module is constructed. The inter-omics integration module responds to the input first intra-omics fusion representation and second intra-omics fusion representation. The first intra-omics fusion representation and the second intra-omics fusion representation are output through the intra-omics integration module. A multi-head attention mechanism is applied to the first intra-omics fusion representation and the second intra-omics fusion representation to obtain the attention weights of each intra-omics fusion representation. The attention weights and the corresponding intra-omics fusion representations are weighted and summed to obtain the multi-omics latent representation. The intra-omics integration module and the inter-omics integration module are combined to form a spatial multi-omics integration model. A dual learning strategy is used to train the spatial multi-omics integration model, and spatial multi-omics integration is achieved based on the trained spatial multi-omics integration model. The dual learning strategy includes intra-omics private learning and inter-omics shared learning. The intra-omics private learning is used to guide the maintenance of the mapping relationship between the potential representation of multi-omics and the original normalized representation space, and self-supervised contrastive learning is introduced to enhance the model's ability to identify omics-specific features. The goal of the inter-omics shared learning is to impose constraints between the representation of one omics and the corresponding representation generated by another omics through the decoder-encoder path, thereby ensuring that data from different omics remain aligned in the shared space.
2. The method according to claim 1, characterized in that, The intra-omics integration module responds to the input spatial map and feature map by performing graph convolution operations to obtain the spatial latent representation and feature latent representation. The computation process is as follows: In the formula, and These are the normalized adjacency matrices of the spatial graph and the feature graph, respectively. and Let be the trainable weight matrix of the spatial graph and the feature map. and This is the bias vector between the spatial map and the feature map. and Let represent the spatial latent representation and feature latent representation learned from the spatial map and feature map, respectively, and σ represent the activation function; The spatial latent representation and the feature latent representation are concatenated along the feature dimension using the following formula to obtain the concatenated feature: In the formula, x represents the concatenation feature, concat represents the concatenation operation, R represents the set of real numbers, N represents the number of spatial points, and d latent Indicates feature dimension; The stitched features are input into a multilayer perceptron, and the intra-omics fusion representation is calculated using the following formula: x (i) =Dropout(σ(W i x i-1 +b i )) In the formula, x (i) x represents the output of the i-th layer neuron in a multilayer perceptron. i-1 W represents the output of the (i-1)th layer neuron in a multilayer perceptron. i and b i x represents the weights and biases of neurons in the i-th layer of a multilayer perceptron, and Dropout represents randomly assigning neurons to their positions. (L) x represents the output after processing by an L-layer multilayer perceptron. (0) =x is the initial input of the multilayer perceptron, d hidden H represents the dimension of latent features. final This represents the fusion representation within omics.
3. The method according to claim 1, characterized in that, A multi-head attention mechanism is applied to the first and second intra-omics fusion representations to obtain attention weights for each intra-omics fusion representation. The attention weights and their corresponding intra-omics fusion representations are then weighted and summed to obtain a multi-omics latent representation, including: Based on the first and second sets of intra-learning fusion representations, the query vector, key vector, and value vector are determined using the following formulas: In the formula, Q, K, and V represent the query vector, key vector, and value vector, respectively, and W... Q W K and W V Let represent the trainable linear mapping matrices for the query vector, key vector, and value vector, respectively. and These are the first group of intra-learning fusion representations and the second group of intra-learning fusion representations, respectively. Based on the query vector, key vector, and value vector, the representation of omics-specific features is determined using the following formula: In the formula, The attention representation output by the nth attention head, where Attention is the attention function, Q. n ,K n V n Let be the query vector, key vector, and value vector of the nth attention head, respectively, and T be the matrix transpose. h_size Let H be the dimension of the attention head, Softmax be the normalized exponential function, and H be the dimension of the attention head. mean For omics-specific feature representation, head_num is the number of attention heads, and n is the number of the nth attention head currently being computed; Based on the omics-specific feature representation, the attention score is determined using the following formula: In the formula, v represents the attention score of omics m for sample i. m W m and b m These are the learnable parameters related to omics m. This represents the omics-specific features of sample i; Based on the attention score, the attention weight is determined using the following formula: In the formula, Let m be the attention weight of omics m for sample i, exp be an exponential function with the natural constant as the base, M be the number of omics, and m be the index of the omics. The multi-omics latent representation is obtained by weighted summation of the attention weights and the corresponding intra-omics fusion representations using the following formula: In the formula, z i The multi-omics latent representation of sample i. This represents the intra-omics fusion representation of omics m.
4. The method according to any one of claims 1 to 3, characterized in that, After constructing the adjacency graph, the method further includes: performing data augmentation on the adjacency graph to obtain a perturbed graph; wherein the data augmentation includes: randomly permuting the gene expression matrix while keeping the topology of the adjacency graph unchanged to obtain a perturbed gene expression matrix, constructing a perturbed graph based on the perturbed gene expression matrix, and introducing a symmetric contrastive loss function into the perturbed graph to optimize spatial representation consistency when constructing the perturbed graph, wherein the contrastive loss function is expressed as: In the formula, To compare the losses, For the representation of position i in the perturbation graph, g' i Let i be the local environment vector at position i in the perturbation map. Let E represent position i in the adjacency graph. (X,A) To represent the expectation under the original graph (adjacency graph) distribution, E (X',A') To represent the expectation under the perturbation graph distribution, Φ is a similarity function that measures the degree of similarity between two vectors.
5. The method according to claim 4, characterized in that, The intra-omics private learning includes: A decoder is constructed to generate reconstructed features based on the intra-omics fusion representation, wherein the reconstructed representation of the decoder at layer l is calculated by the following formula: In the formula, The reconstructed features are the output of the decoder at layer l. Let be the normalized adjacency matrix of the spatial graph, and σ be the activation function. and Z is a trainable weight matrix and bias vector. l-1 This is the output of the decoder at layer l-1; Based on the reconstructed features, a reconstruction loss function is constructed, which is expressed as follows: In the formula, L recon The reconstruction loss is defined as follows: i is the index of the sample, N is the number of samples, and x... i Let i be the intra-omics fusion representation of sample i. The reconstructed features of sample i; Contrastive learning based on adjacency graphs and perturbation graphs includes: pairing a representation of a location in the adjacency graph with its corresponding local environment vector to form a positive sample; pairing a representation of a location in the perturbation graph with its corresponding local environment vector to form a negative sample; and performing contrastive learning using the positive and negative samples according to a predefined contrastive learning loss function; wherein the contrastive learning loss function is expressed as: In the formula, L scl To compare learning loss, For the representation of position i in the adjacency graph, g i Let i be the local environment vector at position i in the adjacency graph; An intra-omics loss function is constructed based on the reconstruction loss function, the contrastive loss function, and the contrastive learning loss function, and intra-omics private learning is performed according to the intra-omics loss function; wherein, the intra-omics loss function is expressed as: In the formula, L omics The loss is omics-based, and λ1 and λ2 are hyperparameters.
6. The method according to claim 5, characterized in that, The inter-omics shared learning includes: Obtain the omics latent representation extracted from one omics and the cross-omics representation generated through the decoder-encoder path of another omics, and optimize the mapping relationship between the first omics latent representation and the second omics latent representation using the following formula: In the formula, and Let be the omics potential representations of the i-th sample in omics 1 and omics 2, respectively. and W represents the cross-omics representations generated via omics 2 and omics 1 paths, respectively. d and W e These are the trainable parameters for the decoder and encoder, respectively. This refers to the l-th layer cross-omics representation generated through the omics 2 to omics 1 path, i.e., reconstructing omics 1 using information from omics 2. The representation input of omics 1 at layer l-1 is the input feature of the current graph neural network. and These are the trainable weights and bias parameters of the decoder at layer l-1. and Let L be the trainable weights and bias parameters of the encoder at layer l-1. corr Let γ1 and γ2 be the cross-omics alignment loss function, where γ1 and γ2 are the weighting coefficients controlling the alignment loss between the two omics in the total loss. and Let be the squared norm between the representation of the i-th sample and its cross-modal reconstructed representation; Establish a shared learning contrastive loss function, expressed as: In the formula, L bt To share the learning contrast loss, Representing the latent representation of two omics modes and The cross-correlation matrix between them, where N is the number of samples, d is the feature dimension, and C ii As diagonal elements, C ii For off-diagonal elements, γ3 is a hyperparameter used to balance diagonal and off-diagonal terms; exchange and Given the position, calculate the mapping relationship in both directions, and take the average value to obtain the final optimization objective: In the formula, β is a weighting parameter that controls the strength of omics alignment, and L multi_omics This represents a loss between omics studies.
7. The method according to claim 6, characterized in that, The dual learning strategy also includes degree-omics balanced learning, which includes: Cosine similarity is determined using the following formula: In the formula, The gradient of the loss function specific to each omics. is the gradient of the joint loss between omics, and cosβ is the cosine similarity; When cosβ≥0 and These are the weights for inter-omics loss and intra-omics loss, respectively. When cosβ < 0, it is determined by the following formula. and Based on the determined weights of intra-omics and inter-omics losses, the total loss function is established using the following formula: In the formula, L total Let represent the total loss, n be the number of training samples, and i be the index of the sample. The spatial multi-omics integration model is trained based on the total loss function.
8. A space-based multi-omics integration device, characterized in that, include: The data acquisition module is configured to acquire a dual-omics dataset; wherein the dual-omics dataset includes first omics data and second omics data; The graph feature construction module is configured to construct an adjacency graph based on the dual-omics dataset; wherein the adjacency graph includes a spatial graph and a feature graph; An intra-omics integration module is configured to construct an intra-omics integration module. The intra-omics integration module responds to the input spatial graph and feature graph by performing graph convolution operations to obtain spatial latent representation and feature latent representation, concatenates the spatial latent representation and feature latent representation on the feature dimension to obtain concatenated features, and inputs the concatenated features into a multilayer perceptron to obtain an intra-omics fusion representation. An inter-omics integration module is configured to construct an inter-omics integration module that responds to input first intra-omics fusion representation and second intra-omics fusion representation. The first intra-omics fusion representation and the second intra-omics fusion representation are output by the inter-omics integration module. A multi-head attention mechanism is applied to the first intra-omics fusion representation and the second intra-omics fusion representation to obtain the attention weights of each intra-omics fusion representation. The attention weights and the corresponding intra-omics fusion representations are weighted and summed to obtain a multi-omics latent representation. The learning and training module is configured to combine the intra-omics integration module and the inter-omics integration module to form a spatial multi-omics integration model. A dual learning strategy is employed to train the spatial multi-omics integration model, and spatial multi-omics integration is achieved based on the trained spatial multi-omics integration model. The dual learning strategy includes intra-omics private learning and inter-omics shared learning. Intra-omics private learning guides the maintenance of a mapping relationship between the multi-omics latent representation and the original normalized representation space, and introduces self-supervised contrastive learning to enhance the model's ability to identify omics-specific features. The goal of inter-omics shared learning is to impose constraints between the representation of one omics and the corresponding representation generated by another omics through a decoder-encoder path, thereby ensuring that data from different omics remain aligned in the shared space.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.