Spatial domain identification method based on spatial multi-omics data
By combining the deep neural network model of autoencoder and multi-scale adaptive graph convolution network, integrating multi-omics data, the problems of insufficient utilization and inaccurate identification of spatial multi-omics data in the prior art are solved, and accurate identification of spatial domains and clear separation of adjacent organizational structures are achieved.
Patent Information
- Application Number
- CN202510063037.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The prior art is difficult to make full use of spatial multiomics data, especially when identifying discrete spatial domains and adjacent organizational structures, there are problems of inaccurate identification and neglect of information.
A deep neural network model combining autoencoder and multi-scale adaptive graph convolution network is adopted to integrate multiple omics data such as RNA-seq and ATAC-seq. It extracts attribute features and structural features through feature extraction and fusion modules, and performs multi-step fusion, and finally recognizes the spatial domain through K-means clustering.
Accurate identification of the same spatial domain of discrete distribution and clear separation of adjacent spatial domains are achieved, which improves the performance and accuracy of spatial domain identification, can process multi-slice data and eliminate batch effects.
Smart Images

Figure CN120089200A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for spatial domain recognition based on spatial omics data, belonging to the technical field of bioinformatics. Background Art
[0002] The advent of single-cell RNA sequencing (scRNA-seq) technology has enabled biologists to obtain the expression profile data of the entire genome at the single-cell level, and thus effectively analyze heterogeneous cell populations in complex samples. The rapid development of single-cell technology provides a unique perspective for understanding cell heterogeneity and dynamic changes in complex biological systems. In recent years, the emergence of emerging multi-omics sequencing technologies such as CITE-seq, SNARE-seq, and TEA-seq has enabled researchers to simultaneously capture multiple omics data such as transcriptome and epigenome at the single-cell level, and thus can dissect cell heterogeneity from multiple perspectives. However, in the tissues of multicellular organisms, cells usually form specific structures or tissue regions in an intertwined manner according to their specific spatial positions and functional similarities in a given tissue, and these regions contribute to the realization of specific biological functions. For example, hepatocytes in the liver can form hepatic lobules, and cardiomyocytes in the heart are arranged into myocardial layers. Linking gene expression information with its spatial distribution is crucial for understanding the related functions of tissues. Currently, spatial transcriptome sequencing technologies are mainly divided into in-situ sequencing-based technologies represented by 10×Visiu, Slide-seq, and Stereo-seq, and imaging-based technologies represented by MERFISH and seqFISH+. Using these spatial transcriptome sequencing technologies, not only can the transcriptome data of spots be obtained, but also spatial position information can be obtained, making it possible to decipher tissue functions and cell compositions. However, these technologies can only capture one type of omics information, namely RNA-seq, which results in an inherent defect of information deficiency. Therefore, spatial transcriptome technology is being extended to spatial multi-omics, enabling biologists to simultaneously capture the expression patterns of different omics data on a single tissue section. Recently, some spatial multi-omics sequencing methods that combine multi-omics sequencing technologies have been developed. For example, DBiT-seq and spatial-CITE-se can simultaneously obtain the gene expression profiles, epiprotein data, and spatial position information of spots, and spatialATAC–RNA-seq and CUT&Tag-RNA-seq can capture gene expression profiles, chromatin accessibility, and spatial position information. These additional omics information provides multiple perspectives for in-depth understanding and analysis of tissue heterogeneity and the functions of different tissue structures.
[0003] Identifying the spatial domain is one of the key steps in analyzing spatial transcriptome data and is the basis for understanding tissue heterogeneity. Over the years, methods for analyzing spatial transcriptome data have been continuously enriched and developed. However, the information contained in single-omics data is limited. Especially in the case of complex tissue structures, the spatial domain identification methods developed for transcriptome data often fail to achieve ideal results. The emergence of spatial multi-omics sequencing technology has broken the limitation of the traditional spatial transcriptome sequencing technology that only captures single-omics data, resulting in information scarcity. Therefore, the spatial domain partitioning method is also shifting from applying single spatial transcriptome data to the integration of spatial multi-omics data.
[0004] In the analysis of spatial multi-omics data, the information provided by different omics data is complementary and can characterize the tissue structure from multiple perspectives. By integrating spatial multi-omics data, the inherent limitations of single-omics data can be overcome, thus achieving more accurate spatial domain identification. Although the recently developed SpatialGlue and SSGATE methods can use spatial multi-omics data to decipher spatial domains, their performance still needs to be improved. Developing a more reliable method for identifying spatial domains using spatial multi-omics data is still an urgent problem to be solved. In addition, the present invention notes that when these methods identify the same spatially discrete domain or a spatially fine but closely connected domain with other tissue layers in the tissue, due to overemphasizing the importance of spatial information and ignoring the attribute information of the spot itself, these tissue structures are divided into different spatial domains and it is difficult to separate adjacent tissue structures. Summary of the Invention
[0005] In order to solve the problem that it is difficult to make full use of spatial multi-omics data and difficult to identify discrete spatial domains in the prior art, the present invention further proposes a method for identifying spatial domains based on spatial multi-omics data.
[0006] The technical solution adopted by the present invention to solve the above problems is as follows: The present invention includes the following steps:
[0007] Step 1: Obtain a comprehensive spatial multi-omics dataset and perform preprocessing;
[0008] Step 2: Construct a deep neural network model;
[0009] Step 3: Train the deep neural network model based on the preprocessed comprehensive spatial multi-omics dataset;
[0010] Step 4: Input the data to be tested into the trained deep neural network model for spatial domain identification.
[0011] Preferably, the spatial multi-omics comprehensive dataset in step 1 includes a multi-omics dataset composed of RNA-seq data and ATAC-seq data and a multi-omics dataset composed of RNA-seq data and ADT-seq data.
[0012] Preferably, the preprocessing of the spatial multi-omics comprehensive dataset in step 1 specifically includes:
[0013] Step 1.1: Filter out the genes expressed in less than 10 spots in the RNA-seq data, perform logarithmic transformation and normalization on the filtered gene expression counts using the SCANPY package, and select the top 3000 highly variable genes in the normalized RNA-seq data for dimensionality reduction. The dimensionality reduction dimension setting is consistent with the feature dimension of the proteomics. Among them, the SCANPY package is a single-cell analysis toolbox based on python to complete the preprocessing of the RNA-seq data;
[0014] Step 1.2: Perform centered log-ratio normalization on the original protein expression counts in the ADT-seq data and the ATAC-seq data, and perform dimensionality reduction on the normalized ADT-seq data and ATAC-seq data to complete the preprocessing of the ADT-seq data and the ATAC-seq data.
[0015] Preferably, the deep neural network model constructed in step 2 includes:
[0016] A feature extraction module and a feature fusion module;
[0017] The feature extraction module includes an autoencoder and a multi-scale adaptive graph convolution module;
[0018] The autoencoder is used to extract the attribute features of different omics data;
[0019] The multi-scale adaptive graph convolution module includes three MACM sub-modules, which are used to extract the structural features of different omics data containing different-order neighbor information;
[0020] The feature fusion module is used to fuse the attribute features extracted by the autoencoder and the structural features extracted by the multi-scale adaptive graph convolution module.
[0021] Preferably, step 3 specifically includes:
[0022] Step 3.1: Select the first n - 1 principal components of the preprocessed ADT-seq data and ATAC-seq data and the preprocessed RNA-seq data as the input data of the autoencoder;
[0023] Step 3.2: Use the k-nearest neighbor algorithm to convert the spatial location information of the input data spot into a neighborhood graph G=(V, E). For a specific point i in V of the neighborhood graph, calculate the Euclidean distance between point i and all other points according to the spatial location information, and select k nearest points as the neighbors of point i;
[0024] Step 3.3: Use an autoencoder to extract the attribute features of the input data;
[0025] Step 3.4: Use a multi-scale adaptive graph convolution module to extract the structural features of the input data;
[0026] Step 3.5: Use a feature fusion module to adaptively fuse the attribute features and structural features of the extracted corresponding omics data to obtain the multi-source fusion features of the corresponding omics data;
[0027] Step 3.6: Repeat Step 3.5 to obtain the multi-source fusion features of different omics data. Concatenate the multi-source fusion features of different omics, and combine the common constraints of multiple losses to obtain the fusion feature Z containing the attribute information and structural information of multiple omics data;
[0028] Step 3.7: Use the K-means method to perform clustering analysis on the fusion feature Z to identify the spatial domains of different tissues and complete the training of the deep neural network model.
[0029] Preferably, Step 3.3 specifically includes:
[0030] Step 3.3.1: Based on the encoder in the autoencoder, extract the gene expression attribute features of different sopt in the input data;
[0031] Step 3.3.2: Based on the decoder in the autoencoder, reconstruct the gene expression attribute features of different sopt to obtain the attribute features of the input data;
[0032] Step 3.3.3: Calculate the mean square error between the reconstructed data and the original data of the input data and use it as the reconstruction loss L R , if the reconstruction loss L R is greater than the first preset value, repeat Step 3.3.2 until the reconstruction loss L R is less than the first preset value;
[0033] The expression formula for the gene expression attribute features of RNA-seq data is:
[0034]
[0035] In formula (1), is the encoder of RNA-seq data;
[0036] The gene expression attribute feature expression of ADT-seq data is:
[0037]
[0038] In formula (1), is the encoder of ADT-seq data;
[0039] The expression for reconstructing the gene expression profile of RNA-seq data is:
[0040]
[0041] In formula (3), are different layers of the decoder for RNA-seq data;
[0042] The expression for reconstructing the gene expression profile of ADT-seq data is:
[0043]
[0044] In formula (4), are different layers of the decoder for ADT-seq data;
[0045] The reconstruction loss L R has the following expression:
[0046]
[0047] In formula (5), n is the number of spots.
[0048] Preferably, step 3.4 specifically includes:
[0049] Step 3.4.1: Obtain an omics feature matrix of all spots based on the structural features of the input data;
[0050] Step 3.4.2: Calculate a feature encoding containing k-order neighbor information based on the omics feature matrix of all spots, and merge the feature encodings of k-order neighbor information through the MACM sub-module to obtain a multi-scale feature encoding H containing different neighbor information;
[0051] Step 3.4.3: Perform average pooling on the feature encodings of each scale and splice them together, and use the MLP module to adaptively learn the importance of the feature encodings of different scales and calculate an importance scoring matrix;
[0052] Step 3.4.4: Use the importance scoring matrix to fuse the multi-scale feature encoding H containing different-order neighbor information to obtain a final feature representation h containing multi-scale structural information;
[0053] The expression for the feature encoding of the k-th order neighbor information is:
[0054] h k = A k Tanh(XW (k) ) (6);
[0055] In formula (6), h k is the feature encoding containing the k-th order neighbor information, A is the adjacency matrix, Tanh is the activation function, X is the omics feature matrix, and W (k) is the learnable parameter of the k-th order neighbor information;
[0056] The expression for the multi-scale feature encoding H of different neighbor information is:
[0057] H = [h 1 ; h 2 ; h 3 ; …; h k (7);
[0058] The expression for average pooling and concatenation of the feature encoding is:
[0059] h m = [f mean (h 1 , dim = 0), f mean (h 2 , dim = 0), …, f mean (h k , dim = 0)] (8);
[0060] In formula (8), f mean is the average pooling operation on the feature matrix, and h m is the concatenated feature encoding of different scales;
[0061] The expression for the importance scoring matrix is:
[0062] W = softmax(f m (h m )) (9);
[0063] In formula (9), f m is to perform importance adaptive learning on h m through the MLP module;
[0064] The calculation formula for the feature representation h of the multi-scale structural information is:
[0065] h = h 1 W (0) + h 2 W (1) + h 3 W(3) +…+h k W (k) (10);
[0066] In formula (10), h is the feature representation that integrates multi-scale structural information, and W (k) is the importance of the feature matrix containing the k-th order neighborhood information.
[0067] Preferably, step 3.5 specifically includes:
[0068] Step 3.5.1: Perform multi-step fusion on the attribute features extracted by the autoencoder and the multi-scale structural features of the corresponding omics data extracted by the multi-scale adaptive graph convolution module to obtain the multi-source fusion feature H of the corresponding omics data RNA ;
[0069] Step 3.5.2: Repeat step 3.5.2 to obtain the multi-source fusion feature H of the omics data ADT ;
[0070] Step 3.5.3: Use the graph decoder to perform node embedding reconstruction on the multi-source fusion features of different omics, where the structure of the graph decoder is symmetrically distributed with that of the autoencoder;
[0071] Step 3.5.4: Calculate the graph feature reconstruction loss L G , if the graph feature reconstruction loss is greater than the second preset value, then repeat steps 3.5.1 - 3.5.3 until the graph feature reconstruction loss is less than the second preset value, and output the reconstructed multi-source fusion features of different omics;
[0072] The multi-source fusion feature H of the omics data RNA has the following calculation formula:
[0073]
[0074] In formulas (11)-(13), a is a hyperparameter, n is the number of layers of the graph encoder, A is the adjacency matrix constructed according to the spatial position information, and H RNA is the multi-source feature fusion representation of the attribute features and multi-scale structural features of the RNA omics data, are the attribute features extracted from different layers of the autoencoder;
[0075] The expression for node embedding reconstruction is:
[0076]
[0077]
[0078] In formulas (14) and (15), Z RNA and Z ADTFor the reconstructed omics embedding features, including neighbor information and related topological information, and are the omics embedding decoders for different omics data;
[0079] The graph feature reconstruction loss L G The calculation formula is:
[0080]
[0081] Preferably, step 3.6 specifically includes:
[0082] Step 3.6.1: Concatenate the reconstructed multi-source fusion features of different omics to obtain the feature representation Z;
[0083] Step 3.6.2: Introduce the constraint based on the graph structure, reconstruct the graph structure based on the feature representation Z and calculate the binary cross-entropy loss as the structure reconstruction loss L S And calculate the contrastive learning loss L con ;
[0084] Step 3.6.3: Based on the structure reconstruction loss L S , the contrastive learning loss L con , the reconstruction loss L R and the graph feature reconstruction loss L G Construct the total loss L, and train the deep neural network model based on the total loss L;
[0085] The calculation formula of the feature representation Z is:
[0086] Z = [H RNA , H ADT (17);
[0087] The calculation formula of the structure reconstruction loss L S is:
[0088] A' = sigmoid(ZZ T ) (18);
[0089]
[0090] In formulas (18) and (19), A' is the reconstructed graph decoded from the fusion feature Z;
[0091] The calculation formula of the contrastive learning loss L con is:
[0092]
[0093] In formula (20), N i is the neighbor set i of the spot, cos(Zi , Z j ) is the cosine similarity between spot i and j, and τ is a hyperparameter;
[0094] The calculation formula of the total loss L is:
[0095] L = L R + L G + L con + L S (21).
[0096] The beneficial effects of the present invention are:
[0097] 1. The present invention proposes a spatial domain recognition method spaMGCN that combines an autoencoder and a multi-scale adaptive graph convolutional network. spaMGCN integrates transcriptomic data and epigenomic data, combines the attribute information contained in different omics data of spots and the structural information considering the attribute information of adjacent spots, and realizes accurate recognition of discrete distributed same spatial domains and clear separation of adjacent spatial domains.
[0098] 2. On various types of spatial multi-omics datasets, the spaMGCN model (deep neural network model) of the present invention can achieve excellent spatial domain recognition performance. By comparing with other latest spatial domain recognition methods, the spaMGCN model shows better performance on multiple evaluation indicators;
[0099] 3. The spaMGCN model constructed by the present invention can also perform parallel processing on multiple slice datasets and eliminate batch effects in spatial transcriptomic data at the same time. Description of the Drawings
[0100] Figure 1 is a flowchart of a method for spatial domain recognition based on spatial multi-omics data provided by the present invention;
[0101] Figure 2 is a schematic structural diagram of the deep neural network model provided by the present invention;
[0102] Figure 3 is a comparison diagram of the spatial domain recognition performance of 4 datasets provided by the present invention under different methods, Figure 3 in which, Figure 3 a is the clustering performance evaluation index of the spaMGCN method and other methods on the S1 dataset, Figure 3 b is the clustering performance evaluation index of the spaMGCN method and other methods on the S3 dataset, Figure 3 c is a comparison schematic diagram of the spatial domain recognition performance of the spaMGCN method and other methods on the Mouse Thymu dataset, Figure 3d is a schematic diagram comparing the spatial domain recognition performance of the spaMGCN method with other methods on the Mouse Spleen dataset;
[0103] Figure 4 This is a visualization graph of the comparison results of spatial domain prediction between the present invention and other methods on the human lymph node S3 dataset;
[0104] Figure 5 This is a comparison graph of the discrete spatial domain recognition performance of different spatial multi-omics spatial domain recognition methods provided by the present invention. Figure 5 In Figure 5 a is a schematic diagram of the F1 score of T cell Zone recognition between the present invention and other methods on the Mouse Spleed dataset. Figure 5 b is a schematic diagram of the F1 score of follicle recognition between the present invention and other methods on the human lymph node S3 dataset. Figure 5 c is a schematic diagram of the true distribution of T cell Zone. Figure 5 d, Figure 5 e and Figure 5 f are schematic diagrams of the prediction situations of three methods;
[0105] Figure 6 This is a schematic diagram of the recognition performance of capsule by different methods provided by the present invention. Figure 6 In Figure 6 a and Figure 6 b are schematic diagrams comparing the spatial domains recognized by the spaMGCN method and the SpatialGlue method. Figure 6 c is a schematic diagram of the capsule distribution situation recognized by spaMGCN and the true distribution situation;
[0106] Figure 7 This is a graph of the transcriptome data distribution before and after feature reconstruction provided by the present invention. Figure 7 In Figure 7 a is a schematic diagram of the data distribution of slice 3, slice 1, and slice 2. Figure 7 c is the data distribution situation of the three slices after reconstruction by spaMGCN. Figure 7 b and Figure 7 d are schematic diagrams of the spot distribution in the same spatial domain after data reconstruction;
[0107] Figure 8 This is a performance comparison graph of different methods on the Human placenta dataset provided by the present invention. Figure 8 In Figure 8 a, Figure 8 b and Figure 8 c are schematic diagrams of the cell clustering performance of different methods.Figure 8 d, Figure 8 e, Figure 8 f, Figure 8 g, and Figure 8 h are schematic diagrams of the feature spaces of different methods visualized using UMAP;
[0108] Figure 9 This is a differential analysis diagram of the spatial domain divided by spaMGCN provided by the present invention. Figure 9 In Figure 9 a is a bubble chart of the differentially expressed genes plotted; Figure 9 b is a bubble chart of the surface proteins plotted; Figure 9 c is a schematic diagram of the distribution of differentially expressed genes strongly correlated with the true spatial domain in tissues; Figure 9 d is a visualization diagram of spatially differential proteins with strong correlation. Detailed implementation manners
[0109] Detailed implementation manner 1: In combination with Figure 1 and Figure 2 this implementation manner is described. In this implementation manner, the spaMGCN method is a method for spatial domain identification based on spatial omics data proposed. As Figure 1 shown, the steps of a method for spatial domain identification based on spatial omics data described in this implementation manner include:
[0110] S1: Obtain a comprehensive spatial omics dataset and perform preprocessing;
[0111] The comprehensive spatial omics dataset includes a multi-omics dataset composed of RNA-seq data and ATAC-seq data and a multi-omics dataset composed of RNA-seq data and ADT-seq data;
[0112] In this embodiment, five sets of multi-omics datasets were obtained from the authors of the papers on SpatialGlue and SpaMosaic. These datasets were measured by three spatial multi-omics sequencing technologies, namely 10x Genomics Visium, Stereo-CITE-seq, and SPOTS, and could simultaneously measure protein markers, the whole transcriptome, and spatial location information. The Mouse Thymus dataset was measured by the Stereo-CITE-seq sequencing technology, the Mouse Spleen dataset was measured by the SPOTS technology, and the human lymphnode dataset contained data from three slices, all of which were measured by the 10x Genomics Visium technology. In addition, in this embodiment, a human placenta spatial multi-omics dataset human_placenta with single-cell resolution was obtained from the Broad Institute Single Cell Portal SCP2601. This dataset provided gene expression data, chromatin accessibility data, and spatial location information of cells.
[0113] The preprocessing specifically includes:
[0114] S101: Filter out the genes expressed in less than 10 spots in the RNA-seq data, perform logarithmic transformation and normalization on the filtered gene expression counts using the SCANPY package, and select the top 3000 highly variable genes in the normalized RNA-seq data for dimensionality reduction. The dimensionality reduction dimension setting is consistent with the feature dimension of the proteomics. Among them, the SCANPY package is a single-cell analysis toolbox based on python, and the preprocessing of the RNA-seq data is completed.
[0115] S102: Perform centered log-ratio normalization on the original protein expression counts in the ADT-seq data and the ATAC-seq data, and perform dimensionality reduction on the normalized ADT-seq data and ATAC-seq data to complete the preprocessing of the ADT-seq data and the ATAC-seq data.
[0116] S2: Build a deep neural network model;
[0117] The spaMGCN model constructed in this embodiment is a deep neural network model as Figure 2 shown, including a feature extraction module and a feature fusion module;
[0118] The feature extraction module includes an autoencoder and a multi-scale adaptive graph convolutional module;
[0119] The autoencoder is used to extract the attribute features of different omics data;
[0120] The multi-scale adaptive graph convolution module includes three MACM sub-modules, which are used to extract the structural features of different omics data containing different orders of neighbor information;
[0121] The feature fusion module is used to fuse the attribute features extracted by the autoencoder and the structural features extracted by the multi-scale adaptive graph convolution module.
[0122] S3: Train the deep neural network model based on the preprocessed spatial multi-omics comprehensive dataset;
[0123] S301: Construct a neighborhood graph. In this embodiment, the k-nearest neighbor algorithm is used to convert the spatial position information of spots into a neighborhood graph G=(V, E). Where V represents the set of points, E represents the edge set connecting the points, and k is the number of neighbors. Specifically, for a specific point i belonging to V, the Euclidean distance between this point and all other points is calculated according to the spatial position information, and k nearest points are selected as its neighbors;
[0124] S302: Use the autoencoder to extract the attribute features of different omics data. The extraction process of the gene expression attribute features of different spots is simply expressed as follows:
[0125]
[0126] In formula (1), is the encoder of RNA-seq data;
[0127] The extraction of the attribute features of other omics data is similar. Below, taking the surface protein as an example, it is expressed as follows:
[0128]
[0129] In formula (1), is the encoder of ADT-seq data, which is used to capture the multiple omics attribute features of different spots; the decoder is an inverse process of the encoder and is used to reconstruct different omics data. The reconstruction process of the gene expression profile is expressed as follows:
[0130]
[0131] In formula (3), are the different layers of the decoder of RNA-seq data;
[0132] The reconstruction process of other omics data is:
[0133]
[0134] In formula (4), are the different layers of the decoder of ADT-seq data;
[0135] In this embodiment, the mean square error between the reconstructed data of different omics and the original data is calculated as the reconstruction loss L R , to ensure that the features extracted by the autoencoder are consistent with the original features, which is expressed as follows:
[0136]
[0137] In formula (5), n is the number of spots. The extraction of the spot's own attribute information helps to group together the spatially discrete same spatial domain.
[0138] S303: Use the multi-scale adaptive convolution module to extract structural features and perform multi-step fusion of attribute features and structural features. Inspired by the Multi-scale graph clustering network, this embodiment uses the MACM sub-module it proposed to capture the structural information of the graph and adaptively fuse multiple feature matrices containing different-order neighborhood information to obtain more complex and comprehensive structural features; the MACM sub-module conveys multi-order neighbor information by constructing multiple high-order adjacency matrices. This method overcomes the limitation of direct adjacency and combines multi-order neighbor information, thus making the structural representation of spots more comprehensive and rich; the specific implementation is as follows:
[0139] S30301: X ∈ R N×d is the omics feature matrix of all spots on an tissue section, where N is the number of spots and d is the feature dimension. The feature encoding containing the k-th order neighbor information is expressed as follows:
[0140] h k = A k Tanh(XW (k) ) (6);
[0141] In formula (6), h k is the feature encoding containing the k-th order neighbor information, A is the adjacency matrix, Tanh is the activation function, X is the omics feature matrix, and W (k) is the learnable parameter of the k-th order neighbor information; in this way, this embodiment can obtain a set of multi-scale feature encodings H containing different neighbor information, which is expressed as follows:
[0142] H = [h 1 ; h 2 ; h 3 ; …; h k (7);
[0143] S30302: Perform average pooling on the feature encodings of each scale and splice them together, which is expressed as follows:
[0144] h m = [f mean (h 1 , dim = 0), f mean (h 2 , dim = 0), …, f mean (h k , dim = 0)] (8);
[0145] In formula (8), f mean performs average pooling operation on the feature matrix, and h m is the feature encoding of different scales after splicing;
[0146] S30303: Use an MLP module to adaptively learn the importance of feature encodings of different scales. The importance scoring matrix is expressed as follows:
[0147] W = softmax(f m (h m )) (9);
[0148] In formula (9), f m performs importance adaptive learning on h m through the MLP module;
[0149] S30304: Use the importance scoring matrix to fuse multi-scale feature encodings containing different orders of neighbor information to obtain the final feature representation containing multi-scale structural information, specifically as follows:
[0150] h = h 1 W (0) + h 2 W (1) + h 3 W (3) + … + h k W (k) (10);
[0151] In formula (10), h is the feature representation fused with multi-scale structural information, and W (k) is the importance of the feature matrix containing the k-th order neighborhood information.
[0152] S30305: The structure feature extraction module consists of three layers of MACM and acts as a graph encoder to extract structure features. During the process of extracting structure features, in this embodiment, the attribute features extracted from different layers of the autoencoder are fused with the multi-scale structure features extracted by the MACM module in multiple steps, so as to ensure the full fusion of attribute information and structure information, and avoid only considering its spatial structure information and ignoring the attribute information of the spot itself during the process of extracting the features of the central node. It is expressed as follows:
[0153]
[0154] In formulas (11)-(13), a is a hyperparameter, n is the number of layers of the graph encoder, A is the adjacency matrix constructed according to the spatial position information, and H RNA is the multi-source feature fusion representation of the attribute features and multi-scale structural features of RNA omics data, is the attribute features extracted from different layers of the autoencoder; the extraction of multi-scale structural information can take into account the fact that in actual situations, multicellular biological tissues often aggregate together during the spatial domain division process, enabling spaMGCN to not lose the recognition accuracy of continuous and aggregated spatial domains. Through the multi-step fusion method, spaMGCN comprehensively considers the attribute information and structural information of different omics data itself, enabling the extracted low-dimensional representation to be applied to the recognition of more similar spatial domains.
[0155] S304: Use the graph decoder to reconstruct the node embeddings of the multi-source fusion features of different omics. Among them, the structure of the decoder is symmetric with that of the encoder, and this symmetry simplifies the learning process. The node embedding reconstruction is expressed as follows:
[0156]
[0157] In formulas (14) and (15), Z RNA and Z ADT are the reconstructed omics embedding features, including neighbor information and related topological information, and are the omics embedding decoders of different omics data;
[0158] S305: Calculate the graph feature reconstruction loss of different omics data, and the process is as follows:
[0159]
[0160] The graph feature reconstruction loss calculates the reconstruction loss of the omics embedding containing neighbor information, and the omics embedding reconstruction is crucial for a deeper understanding of the intrinsic attributes of the spot and its spatial information.
[0161] S306: Multi-omics data fusion: In order to further fuse the low-dimensional representations of different omics, in this embodiment, the multi-source fusion features of different omics are first concatenated, and the expression is as follows:
[0162] Z = [H RNA , H ADT (17);
[0163] The fusion of multiple omics features enables the model to analyze the spatial domain from multiple perspectives. In addition, inspired by Deep Graph Structural Infomax, this embodiment introduces graph-structure-based constraints to facilitate the learning of point representations. In this embodiment, the graph structure is reconstructed based on the feature representation Z, and the binary cross-entropy loss is calculated as the structure reconstruction loss L S , which is expressed as follows:
[0164] A’ = sigmoid(ZZ T ) (18);
[0165]
[0166] In formulas (18) and (19), A’ is the reconstructed graph decoded from the fused feature Z;
[0167] To make the feature representation Z of spots more informative and discriminative, this embodiment calculates a contrastive learning loss similar to InfoNCE to make spots adjacent in physical space closer to each other in feature space and spots not adjacent farther away in feature space, which is expressed as follows:
[0168]
[0169] In formula (20), N i is the neighbor set i of spot, cos(Z i , Z j ) is the cosine similarity between spot i and j, and τ is a hyperparameter; for a given spot, its adjacent spots are defined as positive pairs, while its non-adjacent spots are defined as negative pairs. The final overall loss is defined as:
[0170] L = L R + L G + L con + L S (21).
[0171] S307: After obtaining the spatial multi-omics fusion data, use the K-means method to perform clustering analysis on the fused feature Z to identify the spatial domains of different tissues and complete the training of the deep neural network model.
[0172] S4: Input the data to be tested into the trained deep neural network model for spatial domain identification.
[0173] Specific Embodiment 2: To verify the technical effects in Specific Embodiment 1, the following experiment is designed in this embodiment:
[0174] In this embodiment, experiments were first conducted on three sets of human lymph node spatial multi-omics datasets measured by 10x Genomics Visium technology, and the spaMGCN method was compared with six other algorithms. Figure 3 a, Figure 3 b shows the clustering performance evaluation metrics of these methods on the S1 and S3 datasets, including AMI, NMI, and ARI; the results show that spaMGCN is superior to other methods in all three clustering metrics. In addition, during the experiment, this embodiment found that most methods for identifying spatial domains using spatial multi-omics data have better performance than conventional methods using single-omics data to identify spatial domains. This result further verifies that using spatial multi-omics data to identify spatial domains is effective and necessary, which is attributed to the fact that these multi-omics data provide richer tissue heterogeneity information compared to single-omics data and can explain different tissue structures from different perspectives. Therefore, integrating multi-omics data for analysis can help this embodiment more accurately identify spatial domains.
[0175] To evaluate the generalization ability of the spaMGCN model, this embodiment conducted experiments on two other spatial multi-omics datasets of Mouse Thymu and Mouse Spleen measured by Stereo-CITE-seq and SPOTS technologies. Figure 3 c and 3d respectively show the spatial domain identification performance of different methods on the Mouse Thymu and Mouse Spleen datasets. Table 1-3 shows the data type capabilities of different methods on all datasets. spaMGCN also achieved the best spatial domain identification performance on these two datasets. Overall, spaMGCN can achieve good performance when facing spatial multi-omics data measured by different sequencing technologies, demonstrating its strong generalization ability. At the same time, during the experimental setup process, this embodiment applied a classic single-cell multi-omics clustering method, scDMC, to identify spatial domains, and its performance on all datasets was inferior to that of methods using spatial multi-omics data to identify spatial domains. This is attributed to the fact that it does not utilize the key information of spatial location for identifying spatial domains, resulting in the fragmentation of most spatially continuous spatial domains and the emergence of a series of isolated points. However, compared with other methods using single-omics data to identify spatial domains, its performance varies on different datasets. This embodiment speculates that this is because for tissues with continuous spatial domains, spatial location information is often more important; when the spatial domains are relatively discrete, multiple omics features can often provide richer node attribute information to help identify those discontinuous but spatially related spatial domains belonging to the same tissue structure.
[0176] To further intuitively demonstrate the spatial domain recognition effect of spaMGCN, the clustering results of spaMGCN and four other spatial domain recognition methods on human lymph node S3 are visualized as follows Figure 4 As shown, generally speaking, the prediction results of spaMGCN are roughly consistent with the actual situation of the spatial domain. The important organizational structure of the medulla is also clearly predicted and separated from the cortex and the capsule, which is highly consistent with the fact that human lymph nodes can be roughly divided into the outer cortex area and the inner medulla area.
[0177] Table 1
[0178] Mouse_Thymus Mouse_Spleen S1 S2 S3 spaMGCN 0.4626 0.1477 0.2824 0.244 0.2926 SpatialGlue 0.45296 0.0974 0.2385 0.2273 0.2314 SSGATE 0.29258 0.07242 0.15584 0.13098 0.18134 GraphST 0.16294 0.0128 0.11562 0.07936 0.13338 GAAEST 0.0656 0.0084 0.06104 0.0197 0.08188 scMDC 0.2673 0.0373 0.094 0.0834 0.0777 SpaGIC 0.1372 0 0.0967 0.0921 0.1138
[0179] Table 2
[0180]
[0181]
[0182] Table 3
[0183] Mouse_Thymus Mouse_Spleen S1 S2 S3 spaMGCN 0.5012 0.2064 0.4209 0.3779 0.4347 SpatialGlue 0.48472 0.20422 0.379 0.3654 0.3839 SSGATE 0.3914 0.13582 0.30096 0.255 0.32834 GraphST 0.27668 0.03896 0.23712 0.18262 0.25344 GAAEST 0.1060 0.0240 0.10952 0.03398 0.11706 scMDC 0.3716 0.0666 0.16 0.1413 0.1312 SpaGIC 0.224 0.0398 0.2064 0.2088 0.2302
[0184] To prove that spaMGCN can more accurately identify the same spatial domain discretely distributed in tissues. Figure 5 The performance of different methods in identifying discrete spatial domains is shown.
[0185] In multicellular biological tissues, the organizational structure is mostly layered and continuous in most cases. However, there are also some cell clusters that belong to the same organizational structure and perform the same function but are discretely distributed throughout the tissue. For example, T cell zones play a key role in the kidney and are important aggregation regions of immune cells in the spleen, responsible for monitoring and responding to pathogens and foreign substances in the system, and are often in a discretely distributed state in the kidney tissue; in lymph nodes, follicles play a crucial role in the immune system. They are mainly composed of B cells and are the main origin of B cells. Follicles play a role in filtering and capturing antigens (such as bacteria and viruses) in lymph nodes, ensuring that these antigens can be recognized by the immune system and trigger an appropriate immune response, and they are also discretely distributed in lymphoid tissue. These discrete organizational structures also play important roles, and it is crucial to clearly identify such organizational structures. To verify the effectiveness of spaMGCN in re-identifying discrete organizational structures, this embodiment analyzes two datasets, Mouse Spleen and human lymphnode S3. This embodiment uses different spatial domain recognition methods to analyze these two datasets, and then uses the Hungarian algorithm to match the prediction results with the true labels and calculates the F1 scores for T cell Zone recognition on the MouseSpleed dataset and follicle recognition on the human lymph node S3 dataset for different methods, as Figure 5 a and Figure 5 b shown. To more intuitively display the performance of spaMGCN in recognizing discrete spatial domains, this embodiment visualizes the recognition status of spaMGCN, SpatialGlue, and SSGATE, three methods that use spatial omics data for the discrete spatial domain of T cellZone. Figure 5 c shows the true distribution of the T cell Zone, Figure 5 d, 5e, 5f show the prediction situations of the three methods. Observing the red box part, it can be seen that the prediction situation of spaMGCN is the most consistent with the true situation. This is attributed to the fact that during the model training process, this embodiment uses an autoencoder to extract the attribute information of spots, and at the same time, uses MACM to extract the multi-order neighbor information of spots, so that the boundaries are clear and discrete regions can be classified into the same category. SSTAGE is a two-path graph neural network model. Due to overemphasizing the importance of spatial information and ignoring the attribute information of spots themselves, the boundaries are not clear.
[0186] In addition, to further verify the ability of spaMGCN to separate adjacent tissue structures, experiments were conducted on the lymph node dataset in this embodiment. Lymph nodes are an important part of the lymphatic system and play an important role in the human immune system. Generally, they are composed of a capsule, cortex, medulla, and some other tiny tissues. Since the capsule is adjacent to the cortex, Pericapsular adipose tissue often classifies these two tissues into the same category when performing spatial domain partitioning methods, making it difficult to separate the tissue structures. To demonstrate the recognition performance of spaMGCN for such adjacent tissues in lymphatic tissues, this embodiment compares the spatial domains identified by the currently optimal SpatialGlue method, as Figure 6 shown; observing Figure 6 Figures 6a and 6b, it can be found that spaMGCN clearly identifies the capsule tissue and does not mix with the surrounding intra-lymphatic fat and cortex; in the spatial domain identified by SpatialGlue, the capsule is included in the adjacent and larger cortex and Pericapsular adipose tissue structures. To further highlight the recognition performance of spaMGCN for this important tissue structure of the capsule, this embodiment specifically visualizes the distribution of the capsule identified by spaMGCN and the actual distribution in Figure 6 Figure 6c. It can be seen that the distribution identified by spaMGCN is consistent with the actual distribution, and it can well solve the problem of adjacent tissues being classified into the same spatial domain.
[0187] In the actual analysis process, there are some datasets containing multiple slice data. Being able to process multiple slice data in parallel will greatly simplify the analysis process. For this reason, this embodiment combines the adjacency matrices of three slice datasets of human lymph nodes. Similarly, this embodiment performs parallel analysis on the dataset containing three slice data in a five-fold cross-validation manner with 5 random seeds, and its performance is shown in Table 4.
[0188] Table 4
[0189] Dataset ARI AMI NMI humanlymphnodeS1* 0.2627 0.4201 0.4384 humanlymphnodeS2* 0.2465 0.3272 0.3486 humanlymphnodeS3* 0.3272 0.4238 0.4383 humanlymphnodeS1 0.2824 0.4025 0.4209 humanlymphnodeS2 0.2440 0.3570 0.3779 humanlymphnodeS3 0.2926 0.4162 0.4347
[0190] Observing Table 4, it can be seen that during the parallel analysis of multiple slices, its performance is comparable to the performance of individual analysis, indicating that spaMGCN can perform parallel analysis on multi-slice datasets while maintaining comparable spatial domain recognition performance. In addition, during the experiment, this embodiment used UMAP to visualize the data distribution before and after the reconstruction of the multi-slice dataset, as Figure 7 shown.
[0191] Observation Figure 7 It can be seen that the data distribution of slice 3 deviates significantly from that of slice 1 and slice 2 because of the batch effect; Figure 7 c shows the data distribution of the three slices after reconstruction by spaMGCN. It can be seen that the data of the three slices are mixed together and the batch effect is successfully removed. In addition, by comparing Figure 7 b and 7d, it can be clearly seen that the spots in the same spatial domain are clustered together after data reconstruction, which to a certain extent proves that spaMGCN can impute transcriptomic data and make the reconstructed data more in line with the real biological distribution. These findings further prove that spaMGCN also has certain competitiveness in removing data noise.
[0192] To verify the applicability of spaMGCN to high-resolution spatial omics data, this embodiment obtained high-resolution spatial omics data of human placenta in recent research. spaMGCN was compared with four other methods. Figure 8 a-c show the cell clustering performance of different methods. By comparison, it can be seen that the two spatial omics methods have the worst cell clustering performance on the high-resolution spatial omics dataset. SSGATE_F and scMDC are two single-cell omics clustering methods, and their performance is better than that of the two spatial omics clustering methods, SpatialGlue and SSGATE_S, and they are more suitable for this kind of high-resolution spatial omics data. The spatial omics clustering method spaMGCN proposed in this embodiment has achieved the best performance because spaMGCN can use MACM to obtain multi-scale structural information. This method enables spaMGCN to comprehensively consider the multi-order neighbor information of the graph and at the same time use the autoencoder to retain the attribute information of different spots, making it applicable to both low-resolution spatial omics data and high-resolution datasets. To visually show the distribution of different cell types in the feature space of different methods, this embodiment uses umap to visualize the feature space of different methods, as shown in Figure 8 d-h. From the umap graph, it can be seen that the separation of different types of cells in the feature space of spaMGCN is better, and the cell clusters of SSGATE_F are also well separated, but there is the same situation of cell cluster splitting. In the feature spaces of SpatialGlue and scMDC, the same type of cells do not cluster well. Due to the over-smoothing of graph convolution in SSGATE_S, the features of different types of cells have no distinguishability and are difficult to separate.
[0193] Marker genes can be used to annotate specific cell types. By studying their expression patterns, this embodiment can clarify cell heterogeneity and potential gene regulatory mechanisms. There are also specific gene markers for different spatial domains, and the spatial patterns of different spatial domains are different. To further prove the accuracy of the spatial domain division of the spaMGCN method, this embodiment uses the non-parametric Wilcoxon rank-sum test to screen the top three differential features (genes, epigenetic proteins, chromatin open regions) of each predicted spatial domain and verify their association with the true tissue structure through previous research papers. In addition, this embodiment plots bubble charts of differential genes and epigenetic proteins, as shown in Figure 9 Figures a and 9b. Some of the predicted domain-specific features are significantly different in different spatial domains and can identify specific spatial domains to a certain extent. They may be potential spatial domain-specific genes and epigenetic proteins. This embodiment also visualizes the distribution of some differential genes strongly related to the true spatial domains in the tissue. For example: Figure 9 In Figure c, CCL21, TRBC2, and TRAC clearly mark the Cortex region. Previous studies have shown that CCL21 is mainly expressed in lymphatic endothelial cells, which is consistent with the true annotated region. The spatial distribution regions of CLEC4G and CCL14 are consistent with the medulla region. The CLEC4G gene encodes a member of the glycan-binding receptor and C-type lectin family, and the latter plays a role in the immune response, which is highly consistent with the function of the medulla in providing lymphocytes; the spatial distribution patterns of FABP4, GPX3, GSN, and ADIRF are consistent with the Pericapsular Adipose tissue. FABP4 is mainly expressed in adipocyte-encoding cells, and ADIR plays a role in adipocyte development. In the early stage of pre-differentiation of adipocytes, it promotes adipogenesis and stimulates major adipogenic factors, which is very consistent with the fact that Pericapsular Adipose tissue is mainly composed of adipocytes; the distribution of CXCL13 is consistent with the follice region. It is an exhaustion marker of T cells and can be used for the homing and localization of lymphocytes within the follicles of secondary lymphoid tissues. By analyzing the spatial domains divided by spaMGCN, this embodiment can identify some marker genes that are strongly associated with the functions of spatial domains or have been proven, which further proves the effectiveness of spaMGCN in identifying spatial domains consistent with biological facts. In addition, these domain-specific genes can help biologists analyze the cell composition of different tissue structures and provide strong support for further understanding tissue heterogeneity.
[0194] At the same time, this embodiment also performs differential analysis on the epigenetic protein data in the human lymph node S3 dataset. Figure 9b shows the spatially differential proteins obtained by spatial domain analysis identified by spaMGCN. It can be seen that some surface proteins have strong associations with certain regions. In this embodiment, these spatially differential proteins with strong associations are further visualized, such as Figure 9 shown in d. HLA-DRA is mainly expressed in B lymphocytes, which is consistent with the fact that the Cortex region is mainly composed of B cells, indirectly reflecting that the spatial domain division of spaMGCN has a biological basis. At the same time, in Figure 9 d, the spatial distribution of HLA-DRA and CD3E proteins is consistent with the Cortex region and can better identify the Cortex relative to the CCL21 gene. This also indirectly reflects the necessity of dividing the spatial domain by integrating multiple omics data, and explains why the spatial domain boundaries divided by using multiple omics data are clearer and more consistent with biological facts.
[0195] During the differential analysis process, by performing differential analysis on the spatial domains divided by spaMGCN, this embodiment finds that the distribution and related functions of the spatial differential features found in some spatial domains have strong associations with the real spatial domains, further proving that the spatial domains divided by spaMGCN have biological significance and are consistent with biological facts. In addition, for some genes and surface proteins with spatial patterns, although there is no previous research proving their associations with certain tissue structures, their spatial distribution patterns are highly consistent with real biological tissues. These may be potential spatial domain marker features, indicating that the spaMGCN model has potential value and application prospects in identifying spatial domain specific features and is expected to provide guidance for biologists.
[0196] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the technical solution content of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement of the above embodiments still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for spatial domain identification based on spatial multi-omics data, characterized in that: The method for performing spatial domain identification based on spatial multi-omics data comprises the following steps: Step 1: Obtain spatial multi-omics comprehensive dataset and preprocess it; Step 2: Build a deep neural network model; Step 3: Train the deep neural network model based on the preprocessed spatial multi-omics comprehensive dataset; Step 4: Input the test data into the trained deep neural network model for spatial domain recognition.
2. The method for spatial domain identification based on spatial multi-omics data according to claim 1, characterized in that: The spatial multi-omics comprehensive dataset in step 1 includes a multi-omics dataset consisting of RNA-seq data and ATAC-seq data and a multi-omics dataset consisting of RNA-seq data and ADT-seq data.
3. The method for spatial domain identification based on spatial multi-omics data according to claim 1, characterized in that: The preprocessing of the spatial multi-omics comprehensive dataset in step 1 specifically includes: Step 1.1: Filter out genes expressed in less than 10 spots in the RNA-seq data, use the SCANPY package to logarithmically transform and standardize the filtered gene expression counts, select the top 3000 highly variable genes in the standardized RNA-seq data for dimensionality reduction, and the dimension setting of dimensionality reduction is consistent with the characteristic dimension of epiproteomics. Among them, the SCANPY package is a single-cell analysis toolkit based on Python, which completes the preprocessing of RNA-seq data; Step 1.2: The original protein expression counts in the ADT-seq data and ATAC-seq data were normalized by centered logarithmic ratio, and the normalized ADT-seq data and ATAC-seq data were reduced in dimension to complete the preprocessing of the ADT-seq data and ATAC-seq data.
4. The method for spatial domain identification based on spatial multi-omics data according to claim 1, characterized in that: The deep neural network model constructed in step 2 includes: Feature extraction module and feature fusion module; The feature extraction module includes an autoencoder and a multi-scale adaptive graph convolution module; The autoencoder is used to extract attribute features of different omics data; The multi-scale adaptive graph convolution module includes three MACM sub-modules for extracting structural features of different omics data containing different order neighbor information; The feature fusion module is used to perform feature fusion on the attribute features extracted by the autoencoder and the structural features extracted by the multi-scale adaptive graph convolution module.
5. The method for spatial domain identification based on spatial multi-omics data according to claim 1, characterized in that: Step 3 specifically includes: Step 3.1: Select the preprocessed ADT-seq data, the first n-1 principal components of the ATAC-seq data, and the preprocessed RNA-seq data as the input data of the autoencoder; Step 3.2: Use the k-nearest neighbor algorithm to convert the spatial location information of the input data spot into a neighborhood graph G = (V, E). For a specific point i in the neighborhood graph V, calculate the Euclidean distance between point i and all other points based on the spatial location information, and select the k nearest points as neighbors of point i; Step 3.3: Extract attribute features of input data using the autoencoder; Step 3.4: Extracting structural features of input data using the multi-scale adaptive graph convolution module; Step 3.5: Adaptively fuse the attribute features and structural features of the extracted corresponding omics data using the feature fusion module to obtain multi-source fusion features of the corresponding omics data; Step 3.6: Repeat step 3.5 to obtain multi-source fusion features of different omics data, concatenate the multi-source fusion features of different omics, and combine the common restrictions of multiple losses to obtain the fusion feature Z containing attribute information and structural information of multiple omics data; Step 3.7: Use the K-means method to perform cluster analysis on the fused feature Z, identify the spatial domains of different tissues, and complete the training of the deep neural network model.
6. The method for spatial domain identification based on spatial multi-omics data according to claim 5, characterized in that: Step 3.3 specifically includes: Step 3.3.1: Extract the gene expression attribute features of different sopt in the input data based on the encoder in the autoencoder; Step 3.3.2: Reconstruct the gene expression attribute features of different sopts based on the decoder in the autoencoder to obtain the attribute features of the input data; Step 3.3.3: Calculate the mean square error between the reconstructed data and the original data of the input data and use it as the reconstruction loss L R , if the reconstruction loss L R If it is greater than the first preset value, repeat step 3.3.2 until the reconstruction loss L R is less than a first preset value; The gene expression attribute feature expression of RNA-seq data is: In formula (1), It is an encoder for RNA-seq data, used to capture multiple omics attribute features of different spots; The gene expression attribute feature expression of ADT-seq data is: In formula (1), It is an encoder for ADT-seq data, used to capture multiple omics attribute features of different spots; The expression for reconstructing the gene expression profile of RNA-seq data is: In formula (3), Different layers of the decoder for RNA-seq data; The expression for reconstructing the gene expression profile of ADT-seq data is: In formula (4), Different layers of the decoder for ADT-seq data; Reconstruction loss L R The expression is: In formula (5), n is the number of spots.
7. The method for spatial domain identification based on spatial multi-omics data according to claim 5, characterized in that: Step 3.4 specifically includes: Step 3.4.1: Obtain a set of learning feature matrices for all spots based on the structural features of the input data; Step 3.4.2: Based on a set of learning feature matrices of all spots, a feature code containing k-order neighbor information is calculated, and the feature codes of k-order neighbor information are merged through the MACM submodule to obtain a set of multi-scale feature codes H containing different neighbor information; Step 3.4.3: Average pool the feature codes of each scale and concatenate them together, use the MLP module to adaptively learn the importance of feature codes of different scales and calculate the importance score matrix; Step 3.4.4: Use the importance score matrix to fuse the multi-scale feature encoding H containing different order neighbor information to obtain the final feature representation h containing multi-scale structural information; The expression of feature encoding of k-order neighbor information is: h k =A k Fishy(XW) (k) ) (6); In formula (6), h k is the feature encoding containing the k-th order neighbor information, A is the adjacency matrix, Tanh is the activation function, X is the omics feature matrix, and W (k) is the learnable parameter of the k-th order neighbor information; The expression of multi-scale feature encoding H of different neighbor information is: H=[h1;h2;h3;…;h k ] (7); The expression for average pooling and concatenation of feature codes is: h m =[f mean (h1,dim=0),f mean (h2,dim=0),…,f mean (h k ,dim=0)] (8); In formula (8), f mean To perform average pooling on the feature matrix, h m Encode features of different scales after splicing; The expression of the importance score matrix is: W=softmax(f m (h m )) (9); In formula (9), f m To use the MLP module to m Conduct importance-adaptive learning; The calculation formula of the feature representation h of multi-scale structural information is: h=h1W (0) +h2W (1) +h3W (3) +…+h k W (k) (10); In formula (10), h is the feature representation that integrates multi-scale structural information, W (k) is the importance of the feature matrix containing the k-th order neighborhood information.
8. The method for spatial domain identification based on spatial multi-omics data according to claim 5, characterized in that: Step 3.5 specifically includes: Step 3.5.1: The attribute features extracted by the autoencoder and the multi-scale structural features of the corresponding omics data extracted by the multi-scale adaptive graph convolution module are multi-step fused to obtain the multi-source fusion features H of the corresponding omics data. RNA ; Step 3.5.2: Repeat step 3.5.2 to obtain the multi-source fusion feature H of omics data ADT ; Step 3.5.3: Use the graph decoder to reconstruct node embedding of multi-source fusion features of different omics, where the structure of the graph decoder and the autoencoder are symmetrically distributed; Step 3.5.4: Calculate the graph feature reconstruction loss L for different omics data G , if the graph feature reconstruction loss is greater than the second preset value, repeat steps 3.5.1-3.5.3 until the graph feature reconstruction loss is less than the second preset value, and output the reconstructed multi-source fusion features of different omics; Multi-source fusion features of omics data RNA The calculation formula is: In formulas (11)-(13), a is a hyperparameter, n is the number of layers of the graph encoder, A is the adjacency matrix constructed based on spatial position information, and H RNA It is a multi-source feature fusion representation of the attribute features and multi-scale structural features of RNA-omics data. Attribute features extracted from different layers of the autoencoder; The expression for node embedding reconstruction is: In formulas (14) and (15), Z RNA and Z ADT To reconstruct the omics embedding features, including neighbor information and related topological information, and It is an omics embedding decoder for different omics data; Graph feature reconstruction loss L G The calculation formula is:
9. The method for spatial domain identification based on spatial multi-omics data according to claim 5, characterized in that: Step 3.6 specifically includes: Step 3.6.1: Concatenate the reconstructed multi-source fusion features of different omics to obtain the feature representation Z; Step 3.6.2: Introduce constraints based on graph structure, reconstruct the graph structure based on feature representation Z and calculate binary cross entropy loss as the structure reconstruction loss L S And calculate the contrastive learning loss L con ; Step 3.6.3: Structure-based reconstruction loss L S , contrastive learning loss L con , reconstruction loss L R And the image feature reconstruction loss L G Construct a total loss L, and train the deep neural network model based on the total loss L; The calculation formula of feature representation Z is: Z=[H RNA ,H ADT ] (17); Structural reconstruction loss L S The calculation formula is: A’=sigmoid(ZZ T ) (18); In formulas (18) and (19), A' is the reconstructed image obtained by decoding the fused feature Z; Contrastive learning loss L con The calculation formula is: In formula (20), N i is the neighbor set of spot i, cos(Z i ,Z j ) is the cosine similarity between spot i and j, and τ is a hyperparameter; The calculation formula of total loss L is: L=L R +L G +L con +L S (21)
Citation Information
Patent Citations
Spatial domain identification method based on spatial transcriptomics data feature extraction
CN116189785A
Spatial domain identification method integrating spatial transcriptome multi-modal information
CN118016149A
Heterogeneous graph-based spatial domain identification method and system for spatial multi-omics technology
CN118800328A
Single cell spatial metabolomics for multiplexed chemical analysis
US20240069031A1