Spatial domain identification method based on artificial intelligence

By constructing a cross-modal graph convolutional network model and combining gene expression and histological image data, the problems of insufficient accuracy and robustness of existing spatial domain recognition methods are solved, and more efficient spatial domain recognition and clustering are achieved.

CN120852826AActive Publication Date: 2025-10-28QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Patent Information

Application Number
CN202511351484.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-10-28
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing spatial domain identification methods are insufficient in terms of accuracy and robustness. In particular, traditional algorithms ignore the relationship between adjacent spatial domains, resulting in discontinuity in identification and failure to effectively utilize spatial coordinate information.

Method used

A cross-modal graph convolutional network model was constructed, combining gene expression and histological image data. A spatial clustering graph was generated through multi-view graph convolutional layers, self-attention mechanism and cross-modal joint embedding learning. Feature fusion was performed using spatial adjacency matrix and image features, and the feature vector was reconstructed through ZINB decoder. Clustering was performed using KMeans algorithm.

Benefits of technology

The flexibility and accuracy of the model are improved, and it can adaptively find the most appropriate number of neighbors and radius values, reduce noise dependence, enhance the generalization ability of the model, and ensure the consistency of spatial structure and clustering accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852826A_ABST
    Figure CN120852826A_ABST
Patent Text Reader

Abstract

The invention provides a spatial domain identification method based on artificial intelligence, which belongs to the field of spatial domain identification, combines a traditional graph convolutional network with multiple modals, learns potential representations of data of different modals, and pre-processes and standardizes gene expression data and histological image data to obtain gene features and image features. The method comprises the steps of generating a spatial neighborhood graph and a spatial adjacency matrix based on spatial coordinate information, constructing a cross-modal graph convolution network model, reconstructing feature vectors and performing clustering after multi-view graph convolution, a self-attention mechanism with pruning operation and cross-modal joint embedded learning module learning, generating a spatial clustering graph, and evaluating a clustering effect by using an evaluation index. According to the method provided by the invention, different representations can be ensured to have consistency in space, and the robustness of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spatial domain recognition, and more specifically to a spatial domain recognition method based on artificial intelligence. Background Technology

[0002] In bioinformatics, spatial domain identification technology combines single-cell transcriptomics, tissue section analysis, and biological information, enabling researchers to accurately locate and quantify RNA expression in tissue sections. This technology reveals cell type diversity, heterogeneity in gene expression patterns, and intercellular interactions.

[0003] Traditional molecular biology research typically relies on complex and resource-intensive experiments to determine cellular interactions within specific spatial domains. Utilizing data mining techniques to reconstruct the spatial domains of transcriptomics and pinpoint cellular locations effectively overcomes the limitations of these methods, providing a new perspective for understanding the biological functions of complex tissues.

[0004] With the development of high-throughput sequencing technology, spatial transcriptomics (ST) has become an important tool for revealing intercellular interactions and tissue structure. Based on spatial transcriptomics data, researchers have proposed a variety of non-spatial clustering and spatial clustering methods to identify and analyze cell populations and their gene expression patterns in spatial domains.

[0005] Several computational methods have been proposed for spatial domain identification based on spatial transcriptome data.

[0006] Traditional techniques, such as K-means, Louvain, and Leiden algorithms, are typically used in Scanpy or Seurat packages to form integrated analysis workflows. In addition, graph neural networks (GNNs) are employed to identify spatial domains. Although numerous algorithms have been developed to identify spatial domains in spatial transcriptome data, their accuracy and robustness generally require further improvement.

[0007] Traditional non-spatial algorithms, such as Seurat, perform linear dimensionality reduction clustering based solely on expression data. While they can identify spatial domains, they often suffer from discontinuities because they ignore the relationships between spatially adjacent domains. Probabilistic non-spatial clustering algorithms, such as Giotto, employ Hidden Markov Random Fields (HMRFs) to identify spatial domains by comparing gene expression patterns of neighboring cells. Another approach, BayesSpace, is based on Bayesian statistics and promotes the grouping of neighboring cells into the same group through predefined spatial priors. However, these methods fail to effectively utilize available spatial coordinate information. Summary of the Invention

[0008] The purpose of this invention is to provide a spatial domain recognition method based on artificial intelligence, which can enhance the flexibility and generalization ability of the model.

[0009] To achieve the above objectives, the present invention employs the following technical solutions.

[0010] A spatial domain recognition method based on artificial intelligence includes the following steps: Acquire spatial information, gene expression, and histological images from spatial transcriptome data; Gene expression data and histological image data are preprocessed and standardized to obtain gene features and image features; A spatial neighborhood graph is generated based on spatial coordinate information and then transformed into a spatial adjacency matrix. A cross-modal graph convolutional network model is constructed based on gene features, image features, and spatial adjacency matrix. The network includes multi-view graph convolutional layers, a self-attention mechanism with pruning operations, and a cross-modal joint embedding learning module. The decoder is used to reconstruct the feature vectors learned from the cross-modal joint embedding; Cluster the reconstructed feature vectors, assign cluster labels to each node, and generate a spatial clustering graph. Use evaluation metrics to assess clustering effectiveness.

[0011] Furthermore, the pretreatment of the original gene expression includes the following steps: Gene screening preprocessing was performed on the raw gene expression data to remove genes whose expression coverage was lower than the preset cell number standard; Highly variable genes were screened using the Seurat V3 method; Each cell is normalized to eliminate differences in sequencing depth; The gene data is transformed using a logarithmic transformation to obtain the transformed gene feature matrix. , where P is the number of highly variable genes and N is the total number of points; For histological image data, a pre-trained Vision Transformer model is used to extract the image features corresponding to each pixel, generating an image feature matrix. , where M is the feature dimension of the model output; The standardization process for raw gene expression includes the following steps: The StandardScaler method is used to transform the features into a distribution with a mean of 0 and a standard deviation of 1, resulting in standardized features: definition , ( ) represent the i-th gene feature and the image feature, respectively; The standardized features are , in and These are the mean and standard deviation of the feature, respectively. The standardized gene expression data and histological image data are respectively... and .

[0012] Furthermore, KNN and Radius are used to generate spatial neighborhood graphs and transform them into spatial adjacency matrices. The two different spatial neighborhood graphs are... Where n is the graph number, n=1 indicates the KNN method, n=2 indicates the Radius method, V represents the set of points, and R represents the set of real numbers. Let the set of edges connecting the points in the nth graph be the set of edges connecting the points in the nth graph. The adjacency matrix is ​​represented as .

[0013] Furthermore, generating the spatial adjacency matrix using the KNN method includes the following steps: Calculate the Euclidean distances between all nodes to generate the distance matrix: in, , Let i and j represent the i-th and j-th nodes in V, respectively. By adjusting the threshold t, a suitable number of neighbors is found, so that the number of neighbors for each node is close to a specified value k. ,but Otherwise, it is 0; Set the diagonal to 0 to remove self-loops; Based on the number of neighbors The median is used to determine the optimal value of k. If the specified k is different from the optimal value of k, the adjacency matrix is ​​adjusted. Construct positive and negative adjacency matrix , And generate a sparse adjacency matrix, where It is the element in the i-th row and j-th column of matrix A; Generating a spatial adjacency matrix using the Radius method includes the following steps: Extract spatial coordinates from spatial transcriptome data and initialize the adjacency matrix: ; By using nearest neighbor search for different radius values, the average number of neighbors for each cell is calculated to find the optimal radius value; An adjacency matrix is ​​constructed based on the optimal radius to represent the connections between cells, and a positive and negative graph representation is created, with the positive graph representing the adjacency relationships between cells. Negative graphs represent non-adjacent relationships. And generate a sparse adjacency matrix, where It is the element in the i-th row and j-th column of matrix A.

[0014] Furthermore, the construction of a cross-modal graph convolutional network model includes the following steps: Multi-view graph convolutional layer: Combines two spatial neighborhood graphs with gene expression, PCA dimensionality reduction is used to generate spatial embeddings, and a soft attention mechanism is used to fuse the features of the generated embeddings; Self-attention mechanism with pruning operation: The feature fusion embedding is linearly transformed, attention weights are calculated, and low-weight connections are pruned through a preset threshold to obtain pruned features; Cross-modal joint embedding learning module: A projection head is constructed through a multilayer perceptron, which performs linear transformation and GELU activation on the pruned features, introduces residual connections, fuses the generated latent representations, and generates feature vectors learned through cross-modal joint embedding through fully connected layers.

[0015] Furthermore, the multi-view graph convolutional layer steps include: Graph convolution is performed between two spatial neighborhood graphs and gene expression: ; Perform PCA dimensionality reduction: ,in It is the principal component matrix obtained by PCA; Perform two-layer graph convolution: Two spatial embeddings are obtained and ,in yes The corresponding diagonal matrix, , These are the weight matrices of the two graph convolution layers. , These are the bias terms of the nth spatial adjacency matrix of the two-layer graph convolution; Perform graph convolution between the spatial neighborhood map and image features, i.e. Then perform PCA operation: ,in The principal component matrix is ​​obtained from PCA, resulting in two spatial embeddings. and ; Feature fusion is performed on the generated embeddings using a soft attention mechanism: , ,in and , and All of them are trainable weight parameters; The result and The input is fed into the self-attention mechanism.

[0016] Furthermore, the self-attention mechanism with pruning operations includes the following steps: Perform a linear transformation on the features: , ,in and It is a learnable weight matrix; Calculate attention weights: , ,in and Each feature vector and The dimension; We perform weighted average to obtain , And prune according to a preset threshold; attention weights below the threshold will be set to 0. The result and Input is fed into cross-modal joint embedding learning for joint learning.

[0017] Furthermore, the cross-modal joint embedding learning module includes the following steps: Perform a first-level linear transformation on each modal data point and calculate the projection head: , ,in and It is a learnable bias term; Applying the GELU activation function to the projection result introduces nonlinearity: ; right and Perform the second level of linear transformation: , ,in and These are learnable parameters, obtained by introducing residual connections: , ; The generated latent representations are then fused with features. and Generate by splicing together Then, through another fully connected layer: ,Will Mapping to the output dimension generates cross-modal joint embedding learned feature vectors. The data is then input into the ZINB decoder for reconstruction.

[0018] Furthermore, the ZINB decoder is used to process the feature vector. Reconstruction: Given ST data, the gene expression data x follows a ZINB distribution: ,in These are the parameter matrices for the zero-inflation parameter, the mean, the discretized output of the decoder, and the bias vector, respectively. Minimize the ZINB loss: .

[0019] Furthermore, the KMeans algorithm is used to process the trained feature vectors. Perform cluster analysis: Create a KMeans clustering model and use the fit method on the input data. The training is performed on the CPU, the data is moved off the CPU, the data is cleaned, the data is divided into k clusters, and a cluster label is assigned to each point to generate a spatial clustering graph. The clustering effect is evaluated using the following formula: in This represents the number of samples among N cell points whose clustering results match the true labels. This represents the number of samples among N cell points whose clustering results are inconsistent with their true labels. It is the expected random consistency number. For the accuracy of clustering, The higher the value, the more accurate the clustering.

[0020] The advantages of this invention are: A cross-modal graph convolutional network is constructed. By fusing data from multiple modalities in a graph structure, graph convolutional layers (GCNs) are used to extract node features for each modality. Information sharing and fusion between modalities are achieved through cross-modal information transfer. The graph convolutional layer of each modality updates node features through aggregation operations on information from the local neighborhood, thereby capturing semantic relationships between different modalities.

[0021] It can adaptively find the most suitable number of neighbors k, avoiding the overfitting or underfitting problems that may be caused by fixing the k value, and ensuring that the connectivity of the graph structure is neither too sparse nor too dense, thereby improving the flexibility and accuracy of the model.

[0022] By iterating through different radius values, calculating the number of neighbors for each node at that radius, and comparing this number with the expected value of the radius value, the best-matching radius value is selected. This ensures that the neighborhood structure of each node is most reasonable, avoiding a sparse or overly dense graph structure due to an excessively large or small radius.

[0023] The projector uses a linear transformation to embed and map the input of each modality into a unified feature space, enabling features from different modalities to understand each other. Feature fusion then integrates these features, allowing the model to obtain richer information as a whole.

[0024] The process of combining the projection head with feature fusion aims to achieve spatial consistency in the latent representations across different modalities. This method ensures that features from different modalities are processed in a unified feature space, thereby reducing intermodal differences.

[0025] By incorporating spatial regularization constraints during training, it is ensured that nodes in the graph follow their spatial neighborhood relationships in the latent space. Specifically, spatially adjacent points should have small distances in the latent space, while non-adjacent points should have large distances. This regularization mechanism helps the model better maintain the consistency of the spatial structure during learning, avoiding mismatches between the node distribution in the latent space and the actual spatial structure.

[0026] A self-attention mechanism with pruning is employed, introducing dynamic pruning into attention calculation to adjust the attention weights between nodes. By pruning unimportant connections, computational burden is reduced and efficiency is improved. During training, the pruning threshold can be dynamically adjusted, thereby reducing the model's dependence on noisy or irrelevant connections, enhancing the model's generalization ability, and avoiding overfitting. Attached Figure Description

[0027] Figure 1 This is a flowchart of the spatial domain recognition method based on artificial intelligence of the present invention; Figure 2 This is a flowchart of the preprocessing process for gene expression data and histological image data in this invention; Figure 3 This is a flowchart illustrating the construction process of the cross-modal graph convolutional network model of the present invention. Detailed Implementation

[0028] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.

[0029] This embodiment discloses an artificial intelligence-based spatial domain recognition method, and evaluates its performance on five ST datasets: (1) human dorsolateral prefrontal cortex cells (DLPFC); (2) human breast cancer cells; (3) lung cancer dataset; (4) human pancreatic ductal adenocarcinoma (PDAC); and (5) mouse brain dataset. For detailed procedures, please refer to [reference needed]. Figure 1 The following will explain each step in detail.

[0030] S1. Use spatial transcriptome data as input to the model, including gene expression, spatial information, and histological images, and preprocess it. Please refer to [link / reference]. Figure 2 The flowchart.

[0031] Gene screening preprocessing was performed on the raw gene expression data to remove genes expressed in fewer than 3 cells. 3000 hypervariable genes were selected using the Seurat V3 method. Normalization was performed on each cell to eliminate sequencing depth differences. Finally, log transformation was performed on the data to obtain the gene feature matrix. , where P is the number of highly variable genes and N is the total number of points.

[0032] For histological image data, a pre-trained Vision Transformer model is used to extract the image features corresponding to each pixel, generating an image feature matrix. , where M is the feature dimension of the model output.

[0033] To ensure that features from different modalities are on the same scale, the StandardScaler method is used to transform the features into a distribution with a mean of 0 and a standard deviation of 1, resulting in standardized features: definition , ( ) represent the i-th gene feature and the image feature, respectively; The standardized features are , in and These are the mean and standard deviation of the feature, respectively. The standardized gene expression data and histological image data are respectively... and .

[0034] S2. For spatial information, we use different similarity metrics (i.e., KNN and Radius) to generate two different spatial neighborhood maps. n is the graph number, n=1 indicates the KNN method, n=2 indicates the Radius method, V represents the set of points, and R represents the set of real numbers. Let the set of edges connecting the points in the nth graph be the set of edges connecting the points in the nth graph. The adjacency matrix is ​​represented as .

[0035] Generating a spatial adjacency matrix using the KNN method includes the following steps: Calculate the Euclidean distances between all nodes to generate the distance matrix: in, , Let i and j represent the i-th and j-th nodes in V, respectively.

[0036] By adjusting the threshold t, a suitable number of neighbors is found, so that the number of neighbors for each node is close to a specified value k. ,but Otherwise, set to 0. Set the diagonal to 0 to remove self-loops; based on the number of neighbors. The median is used to determine the optimal value of k. If the specified k differs from the optimal k value, the adjacency matrix is ​​adjusted. Positive and negative adjacency matrices are constructed. , This generates a sparse adjacency matrix, which facilitates regularization constraints in the space. It is the element in the i-th row and j-th column of matrix A.

[0037] Generating a spatial adjacency matrix using the Radius method includes the following steps: Extract spatial coordinates from spatial transcriptome data and initialize the adjacency matrix: .

[0038] Nearest neighbor search is used to find the optimal radius value by calculating the average number of neighbors for each cell across different radius values. An adjacency matrix is ​​constructed based on the optimal radius to represent the connections between cells, and a positive and negative graph representation is created, with the positive graph representing the adjacency relationships between cells. Negative graphs represent non-adjacent relationships. This generates a sparse adjacency matrix, which facilitates regularization constraints in the space. It is the element in the i-th row and j-th column of matrix A.

[0039] S3. Construct a cross-modal graph convolutional network model. Please refer to [link / reference]. Figure 3 The model includes a multi-view graph convolutional network, cross-modal joint embedding learning, and a self-attention mechanism with pruning operations.

[0040] Multi-view graph convolutional layer: First, it processes the spatial adjacency matrices of different views. Graph convolution follows these propagation rules: ,in It is the weight matrix of the l-th layer. It is the ReLU activation function.

[0041] Graph convolution is performed between two spatial neighborhood graphs and gene expression: To reduce computational complexity while preserving important feature information, a PCA operation is incorporated into the convolution process: ,in This is the principal component matrix calculated by PCA. Then, two layers of graph convolution are performed: Two spatial embeddings are obtained and ,in yes The corresponding diagonal matrix, , These are the weight matrices of the two graph convolution layers. , These are the bias terms of the nth spatial adjacency matrix of the two-layer graph convolution.

[0042] Perform graph convolution between the spatial neighborhood map and image features, i.e. Then perform PCA operation: ,in The principal component matrix is ​​obtained from PCA calculation, and similarly, the two spatial embeddings are obtained. and The generated embeddings are fused using a soft attention mechanism. , ,in and , and All are trainable weight parameters: , , , Similarly, we can obtain The result will be and The input is fed into the self-attention mechanism.

[0043] Secondly, a self-attention mechanism with pruning is used to adaptively fuse features of each point in each modality of data.

[0044] Perform a linear transformation on the features: , ,in and It is a learnable weight matrix. Calculate the attention weights: , ,in and Each feature vector and The dimensions are then weighted to obtain the final result. , To reduce computation, pruning is used to remove low-weight connections, and attention weights below a threshold are set to 0. The resulting... and Input is fed into cross-modal joint embedding learning for joint learning.

[0045] Finally, in order to learn data from different modalities, a multilayer perceptron (MLP) is used for feature learning to ensure that the model captures similar structures when processing data from different modalities, thereby improving the model's generalization ability.

[0046] Perform a first-level linear transformation on each modal data point and calculate the projection head: , ,in and It is a learnable bias term. Applying the GELU activation function to the projection result introduces nonlinearity: ; right and Perform the second level of linear transformation: , ,in and These are learnable parameters. To facilitate the flow of information, residual connections are introduced, resulting in: , The generated latent representations are then fused with features. and Generate by splicing together Then through a fully connected layer ,Will Mapping to output dimension generation Jointly embed the learned vectors across modalities The data is input into the ZINB decoder for reconstruction.

[0047] S4. To address the high sparsity and discreteness of ST data, a ZINB decoder is used to process the feature vectors. Reconstruction will be carried out.

[0048] Given gene expression data x from ST data, assuming it follows a ZINB distribution, ZINB is defined as follows: ,in These are the parameter matrices for the zero-inflation parameter, the mean, the discretization of the decoder output, and the bias vector, respectively. This is achieved by minimizing the ZINB loss, i.e., eigenvectors It contains more biological information.

[0049] S5. Use the KMeans algorithm to process the trained feature vectors. Perform cluster analysis. Create a KMeans clustering model, where k specifies the number of clusters, and a random seed (random_state) ensures the reproducibility of the clustering process. Use the `fit` method on the input data. The training is performed on the CPU, the data is moved off the CPU, the data is cleaned, the data is divided into k clusters, and a cluster label is assigned to each point to generate a spatial clustering graph.

[0050] S6. Use evaluation metrics to assess clustering performance: in This represents the number of samples among N cell points whose clustering results match the true labels. This represents the number of samples among N cell points whose clustering results are inconsistent with their true labels. It is the expected random consistency number. For the accuracy of clustering, The higher the value, the more accurate the clustering, and the generated spatial clustering map can reveal the complex biological characteristics of cells in important spatial locations.

[0051] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A spatial domain recognition method based on artificial intelligence, characterized in that, Including the following steps: Acquire spatial information, gene expression, and histological images from spatial transcriptome data; Gene expression data and histological image data are preprocessed and standardized to obtain gene features and image features; A spatial neighborhood graph is generated based on spatial coordinate information and then transformed into a spatial adjacency matrix. A cross-modal graph convolutional network model is constructed based on gene features, image features, and spatial adjacency matrix. The network includes multi-view graph convolutional layers, a self-attention mechanism with pruning operations, and a cross-modal joint embedding learning module. The decoder is used to reconstruct the feature vectors learned from the cross-modal joint embedding; Cluster the reconstructed feature vectors, assign cluster labels to each node, and generate a spatial clustering graph. Use evaluation metrics to assess clustering effectiveness.

2. The spatial domain recognition method based on artificial intelligence according to claim 1, characterized in that, Pretreatment of raw gene expression includes the following steps: Gene screening preprocessing was performed on the raw gene expression data to remove genes whose expression coverage was lower than the preset cell number standard; Highly variable genes were screened using the Seurat V3 method; Each cell is normalized to eliminate differences in sequencing depth; The gene data is transformed using a logarithmic transformation to obtain the transformed gene feature matrix. , where P is the number of highly variable genes and N is the total number of points; For histological image data, a pre-trained Vision Transformer model is used to extract the image features corresponding to each pixel, generating an image feature matrix. , where M is the feature dimension of the model output; The standardization process for raw gene expression includes the following steps: The StandardScaler method is used to transform the features into a distribution with a mean of 0 and a standard deviation of 1, resulting in standardized features: definition , ( ) represent the i-th gene feature and the image feature, respectively; The standardized features are , in and These are the mean and standard deviation of the feature, respectively. The standardized gene expression data and histological image data are respectively... and .

3. The spatial domain recognition method based on artificial intelligence according to claim 2, characterized in that, Spatial neighborhood graphs are generated using KNN and Radius and then transformed into spatial adjacency matrices. Two different types of spatial neighborhood graphs are shown below. Where n is the graph number, n=1 indicates the KNN method, n=2 indicates the Radius method, V represents the set of points, and R represents the set of real numbers. Let the set of edges connecting the points in the nth graph be the set of edges connecting the points in the nth graph. The adjacency matrix is ​​represented as .

4. The spatial domain recognition method based on artificial intelligence according to claim 3, characterized in that, Generating a spatial adjacency matrix using the KNN method includes the following steps: Calculate the Euclidean distances between all nodes to generate the distance matrix: in, , Let i and j represent the i-th and j-th nodes in V, respectively. By adjusting the threshold t, a suitable number of neighbors is found, so that the number of neighbors for each node is close to a specified value k. ,but Otherwise, it is 0; Set the diagonal to 0 to remove self-loops; Based on the number of neighbors The median is used to determine the optimal value of k. If the specified k is different from the optimal value of k, the adjacency matrix is ​​adjusted. Construct positive and negative adjacency matrix , And generate a sparse adjacency matrix, where It is the element in the i-th row and j-th column of matrix A; Generating a spatial adjacency matrix using the Radius method includes the following steps: Extract spatial coordinates from spatial transcriptome data and initialize the adjacency matrix: ; By using nearest neighbor search for different radius values, the average number of neighbors for each cell is calculated to find the optimal radius value; An adjacency matrix is ​​constructed based on the optimal radius to represent the connections between cells, and a positive and negative graph representation is created, with the positive graph representing the adjacency relationships between cells. Negative graphs represent non-adjacent relationships. And generate a sparse adjacency matrix, where It is the element in the i-th row and j-th column of matrix A.

5. The spatial domain recognition method based on artificial intelligence according to claim 3, characterized in that, The construction of the cross-modal graph convolutional network model includes the following steps: Multi-view graph convolutional layer: Combines two spatial neighborhood graphs with gene expression, PCA dimensionality reduction is used to generate spatial embeddings, and a soft attention mechanism is used to fuse the features of the generated embeddings; Self-attention mechanism with pruning operation: The feature fusion embedding is linearly transformed, attention weights are calculated, and low-weight connections are pruned through a preset threshold to obtain pruned features; Cross-modal joint embedding learning module: A projection head is constructed through a multilayer perceptron, which performs linear transformation and GELU activation on the pruned features, introduces residual connections, fuses the generated latent representations, and generates feature vectors learned through cross-modal joint embedding through fully connected layers.

6. The spatial domain recognition method based on artificial intelligence according to claim 5, characterized in that, The multi-view graph convolutional layer steps include: Graph convolution is performed between two spatial neighborhood graphs and gene expression: ; Perform PCA dimensionality reduction: ,in It is the principal component matrix obtained by PCA; Perform two-layer graph convolution: Two spatial embeddings are obtained and ,in yes The corresponding diagonal matrix, , These are the weight matrices of the two-layer graph convolution. , These are the bias terms of the nth spatial adjacency matrix of the two-layer graph convolution; Perform graph convolution between the spatial neighborhood map and image features, i.e. Then perform PCA operation: ,in The principal component matrix is ​​obtained from PCA, resulting in two spatial embeddings. and ; Feature fusion is performed on the generated embeddings using a soft attention mechanism: , ,in and , and All of them are trainable weight parameters; The result and The input is fed into the self-attention mechanism.

7. The spatial domain recognition method based on artificial intelligence according to claim 6, characterized in that, The self-attention mechanism with pruning operation includes the following steps: Perform a linear transformation on the features: , ,in and It is a learnable weight matrix; Calculate attention weights: , ,in and Each feature vector and The dimension; We perform weighted summation to obtain , And prune according to a preset threshold; attention weights below the threshold will be set to 0. The result and Input is fed into cross-modal joint embedding learning for joint learning.

8. The spatial domain recognition method based on artificial intelligence according to claim 7, characterized in that, The cross-modal joint embedding learning module includes the following steps: Perform a first-level linear transformation on each modal data point and calculate the projection head: , ,in and It is a learnable bias term; Applying the GELU activation function to the projection result introduces nonlinearity: ; right and Perform the second level of linear transformation: , ,in and These are learnable parameters, obtained by introducing residual connections: , ; The generated latent representations are then fused with features. and Generate by splicing together Then, through another fully connected layer: ,Will Mapping to the output dimension generates cross-modal joint embedding learned feature vectors. The data is then input into the ZINB decoder for reconstruction.

9. The spatial domain recognition method based on artificial intelligence according to claim 8, characterized in that, Use the ZINB decoder on the feature vector Reconstruction: Given ST data, the gene expression data x follows a ZINB distribution: ,in These are the parameter matrices for the zero-inflation parameter, the mean, the discretized output of the decoder, and the bias vector, respectively. Minimize the ZINB loss: .

10. The spatial domain recognition method based on artificial intelligence according to claim 8, characterized in that, The KMeans algorithm is used to process the trained feature vectors. Perform cluster analysis: Create a KMeans clustering model and use the fit method on the input data. The training is performed on the CPU, the data is moved off the CPU, the data is cleaned, the data is divided into k clusters, and a cluster label is assigned to each point to generate a spatial clustering graph. The clustering effect is evaluated using the following formula: ; in This represents the number of samples among N cell points whose clustering results match the true labels. This represents the number of samples among N cell points whose clustering results are inconsistent with their true labels. It is the expected random consistency number. For the accuracy of clustering, The higher the value, the more accurate the clustering.

Citation Information

Patent Citations

  • Spatial domain identification method integrating spatial transcriptome multi-modal information

    CN118016149A

  • Spatial domain identification method based on multi-view weighted fusion GCN network

    CN120148635A

  • Spatial transcriptome data clustering method based on progressive learning and multi-modal fusion

    CN120256988A

  • Deep clustering method and system based on cross-modal fusion

    WO2022166361A1

Cited By

  • Spatial transcriptome region identification method driven by multi-modal graph fusion

    CN121354671A