Space transcription data clustering method based on dual multi-scale graph learning

By using the multi-scale mask graph autoencoder and dual representation learning mechanism of graph attention networks in spatial transcriptome data clustering, the problem that existing methods cannot fully utilize spatial coordinate information and can only extract single-scale representations is solved, and more accurate and robust spatial domain annotation is achieved.

CN120126580APending Publication Date: 2025-06-10NANTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510175860.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing spatial transcriptome data clustering methods cannot fully utilize spatial coordinate information and can only extract potential representations of a single scale, which cannot ensure that the representation contains sufficient discriminant information, resulting in inaccurate spatial domain division.

Method used

A multi-scale masked graph autoencoder based on graph attention network is adopted to extract potential multi-scale embeddings through multi-scale processing, and a scaling cosine loss function and dual representation learning mechanism are introduced to improve the accuracy and robustness of spatial domain annotation.

Benefits of technology

By exploring multi-scale information in spatial transcription data, the expressiveness and robustness of the model are enhanced, and the accuracy of clustering and the effect of spatial domain annotation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126580A_ABST
    Figure CN120126580A_ABST
Patent Text Reader

Abstract

The invention provides a spatial transcription data clustering method based on dual multi-scale graph learning, and solves the technical problems of limitation of single-scale feature representation and insufficient spatial information extraction. According to the technical scheme, firstly, graph data are established, and a certain proportion of node features are randomly masked in the training stage; secondly, constructing a multi-scale mask graph automatic encoder, extracting multi-scale potential embedding through the encoder, and generating reconstruction features by using a decoder; then, a scaling cosine loss function is introduced, an Adam optimizer is adopted to update network parameters, and a multi-scale mask graph automatic encoder model is stored; and finally, extracting multi-scale embedding of spatial transcription data in a test stage, and realizing spatial domain annotation through multi-scale clustering. The method has the advantages that higher robustness and adaptability are shown in the aspects of spatial information extraction and clustering precision, and the annotation effect and clustering accuracy of spatial transcription data are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of spatial transcriptomics, and particularly to a spatial transcriptome data clustering method based on dual multi-scale graph learning. Background Art

[0002] In recent years, spatial transcriptomics technology has become a cutting-edge tool for understanding cell dynamics and its in-situ microenvironment. Different from single-cell RNA sequencing, spatial transcriptomics not only captures gene expression information but also includes spatial location information, enabling in-depth understanding of molecular communication and tissue structure.

[0003] Using clustering methods to process and annotate spatial transcript data has become a popular research direction in recent years. In addition to the classical K-means, Louvain, and Leiden methods for spatial domain partitioning, some researchers have also developed some clustering modeling methods based on distance calculation or probability estimation. In the paper titled "Spatial transcriptomics at subspot resolution with BayesSpace", the partitioning of the spatial domain was achieved based on Bayesian statistical methods and prior knowledge of the spatial domain. In the paper titled "SC-MEB: spatial clustering with hidden Markov random field using empirical Bayes", a hidden Markov random field was introduced to explore spatial dependence and the association between neighboring cells, and at the same time, the model parameters were estimated by the empirical Bayes method.

[0004] All of the above spatial domain annotation methods are based on traditional machine learning methods, but they ignore most of the valuable spatial coordinate information. To solve this problem, several deep learning-based graph clustering methods have been proposed in recent years and have attracted wide attention. In the paper titled "SpaGCN: Integrating gene expression, spatial location and histology to identify spatial domains and spatially variable genes by graph convolutional network", an undirected weighted graph was introduced to represent the dependence relationship of spatial data, and hidden embeddings were extracted by a graph convolutional network, and finally, the spatial domain partitioning was achieved through an iterative clustering method.

[0005] Although these deep learning-based graph clustering methods have shown effectiveness, they still face significant challenges. First, understanding the complexity of organisms in the field of life science is a major challenge because this complexity stems from its hierarchical structure and multidimensional interactions. Researchers are increasingly aware that relying solely on single-level observations and analyses often fails to comprehensively understand the functions and complex mechanisms of organisms. At the same time, some studies have shown that exploring multi-scale information can more comprehensively capture the data distribution, thereby enhancing the robustness of the model. However, existing spatial domain annotation methods are limited to extracting single-scale latent representations, which often cannot ensure that the representations contain sufficient discriminative information to accurately perform spatial domain partitioning. Therefore, how to design a novel graph neural network to explore multi-scale information in spatial transcriptomic data has become a research question of great significance. Second, current methods directly apply classical clustering methods to the learned representations for spatial domain annotation, which may fail to fully exploit the information embedded in the latent representations. Therefore, designing a clustering method suitable for spatial transcriptome representation extraction is of great significance for further improving clustering performance and achieving more accurate spatial domain annotation. Summary of the Invention

[0006] The object of the present invention is to provide a spatial transcript data clustering method based on dual multi-scale graph learning, mainly solving the technical problems of being unable to fully utilize spatial coordinate information and only being able to extract single-scale latent representations. This method aims to explore multi-scale information in spatial transcript data by designing a novel graph neural network structure and improve the accuracy and robustness of spatial domain annotation through an improved clustering method.

[0007] To achieve the above object of the invention, the technical solutions adopted by the present invention are specifically as follows: A spatial transcript data clustering method based on dual multi-scale graph learning, comprising the following steps:

[0008] S1: Based on spatial location information and gene expression data, establish graph data where is the node set, A ∈ R N×N is the adjacency matrix, X ∈ R d×N is the node unit matrix, N is the number of nodes, and d is the dimension; and randomly mask a certain proportion of the nodes features during the training phase to obtain graph data where is the masked feature;

[0009] S2: Construct a multi-scale masked graph autoencoder based on the graph attention network, perform multi-scale processing on the input masked graph data through the encoder to extract latent multi-scale embeddings, and use the decoder to re-mask the latent embeddings obtained by the encoder and input them into the decoding process, and generate accurate reconstruction features through the multi-scale decoder;

[0010] S3: Introduce a scaled cosine loss function to calculate the error between the reconstructed node features and the original node features

[0011] S4: According to the loss function obtained in step S3, use the Adam optimizer to update the network parameters and save the multi-scale mask graph autoencoder model;

[0012] S5: In the test phase, use the trained encoder to extract the multi-scale embeddings of the spatial transcriptomics data, mine the scale-common information and scale-specific information from the multi-scale embeddings through the dual representation learning method, and input them into the multi-scale clustering method to achieve spatial domain annotation.

[0013] Furthermore, the step S2 includes the following steps:

[0014] S21: The present invention uses a graph attention network as the basic model to extract highly discriminative spatial transcriptome features with different scales; the multi-scale encoder consists of two parts: the first part is the common encoder where f E (*) represents that the common encoder extracts information at different scales and provides the initial representation learning of spatial transcript data. The second part is the scale-specific encoder where f E_m (*) represents that the m-th encoder extracts information at different scales, and M is the number of scales; calculate the attention coefficient of each node in the graph:

[0015]

[0016] where, a i,j is the attention coefficient obtained by normalizing the similarity score, e i,j are all adjacent nodes, V i is the set of adjacent nodes obtained from the adjacency matrix A, LeakyReLU(*) is the activation function; e i,j is the similarity coefficient between nodes i and j, W ∈ R d′×d is the common mapping matrix, d ′ is the dimension of the mapping space, is the feature vector of the i-th node, || is the concatenation operation, set g(*) as a single-layer feedforward network, and θ is the learnable parameter;

[0017] S22: Similar to Equation (2), calculate K coefficients to obtain the multi-head attention coefficients; after obtaining the multi-head attention coefficients, use the multi-head attention mechanism to update the features of each node:

[0018]

[0019] Among them, and W k are the attention coefficient and the common mapping matrix of the k-th head respectively, and σ(*) is the PreLU activation function that enhances the flexibility of the network. is the updated feature vector of the i-th node with a common encoder;

[0020] S23: represents the m-th extracted embedding. Similar to f E (*), it is as follows:

[0021]

[0022] Among them, is the mapping matrix of the k-th head at the m-th scale, is the attention coefficient of the k-th head at the m-th scale, and is the mapped feature vector of the i-th node's m-th encoder, and d ′ m is the feature dimension;

[0023] S24: After obtaining the multi-scale embeddings compressed by the encoder, apply another set of masks to replace the node indices in the previous masks, that is Similar to (5), is defined as follows:

[0024]

[0025] Consistent with the encoder, use a single-layer graph attention network as the decoder for each scale; this method allows the model to recover the features of nodes based on a set of nodes rather than relying only on the nodes themselves, thus supporting the encoder to learn highly discriminative embeddings, that is where f D_m (*) represents the m-th encoder; the specific situation is as follows:

[0026]

[0027] Among them, is the attention coefficient of the k-th head at the m-th scale decoder, is the mapping matrix of the k-th head at the m-th scale decoder, is the feature vector reconstructed by the decoder at different scales.

[0028] Furthermore, the step S3 includes the following steps:

[0029] S31: Introduce the scaled cosine error as the loss function and calculate the error between the reconstructed node features and the original node features:

[0030]

[0031] Among them, X i and are the original feature and the reconstructed feature, γ≥1 is a scale factor, used as a hyperparameter, and the L2 norm in the scaled cosine error maps the vector onto the unit hypersphere.

[0032] Furthermore, the step S5 includes the following steps:

[0033] S51: Based on the multi-scale masked graph autoencoder, obtain the multi-scale hidden embedding of the spatial transcriptome data; denote H m =(H m ), m = 1, 2, …, M. Based on the matrix factorization technique, use a dual representation learning mechanism to simultaneously mine common information and scale-specific information from the multi-scale data: T

[0034]

[0035] Among them, is the common representation between views, is the specific representation of the m-th view, d c and d s are the feature dimensions of the common and specific representations, and are the mapping matrices of the m-th scale, and β is the regularization parameter; in addition, a regularization term is introduced to make the learned representation more robust;

[0036] S52: After extracting the common and specific representations between different scales, adopt a unified clustering framework that introduces the Shannon entropy mechanism and orthogonal constraints; set the common representation as the M+1 scale, then the objective function is as follows:

[0037]

[0038] Among them, the fourth and fifth terms cluster the common representation and the specific representation respectively, and are the clustering centers of the common representation and the specific representation respectively, C is the number of clusters; U∈R C×N is the clustering indicator matrix shared by the two types of representations. When the j-th instance is clustered into the i-th class, U i,j =1, otherwise, U i,j =0; α m is the weight of different representations, I∈R C×C is the identity matrix, and β, λ, δ≥0 are parameters.

[0039] Compared with the prior art, the beneficial effects of the present invention are:​

[0040] (1) Enhanced multi-scale information extraction: The present invention processes spatial transcriptomic data by introducing a multi-scale masked graph autoencoder based on Graph Attention Networks (GAT) to extract features from different scales. This method can fully exploit different levels of spatial information, overcome the limitation of traditional methods that only rely on a single scale, and improve the expressiveness and robustness of the model.

[0041] (2) Improved robustness: The present invention adopts a self-supervised feature masking mechanism to randomly replace node features to simulate missing data, thereby enhancing the model's knowledge compression and extraction capabilities, and enabling it to maintain stable performance in complex data distributions.

[0042] (3) Improved clustering accuracy: For the clustering of spatial transcriptomic data, the present invention designs a novel multi-scale clustering method that performs dual representation learning through matrix factorization, which can effectively fuse common and specific information between scales, thereby enhancing the clustering accuracy.

[0043] (4) Adaptive adjustment of scale weights: By introducing the Shannon entropy mechanism, the present invention can dynamically adjust the importance of representations at different scales, optimize the spatial domain annotation effect, and enable the model to automatically adapt according to the characteristics of the data without relying on fixed scale weights. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention.

[0045] Figure 1 It is a schematic flow chart of the method for clustering spatial transcriptomic data based on dual multi-scale graph learning of the present invention;

[0046] Figure 2 It is an overall block diagram of the method for clustering spatial transcriptomic data based on dual multi-scale graph learning of the present invention;

[0047] Figure 3 It is a schematic diagram of the comparison results on DLPFC in Embodiment 2 of the present invention. Among them, (A) is a schematic diagram of the ARI metric clustering results of all methods; (B) is a schematic diagram of the visualization results of the present invention on slice 151672;

[0048] Figure 4 It is a schematic diagram of the clustering results of all methods based on the NMI and Purity metrics in Embodiment 2 of the present invention on DLPFC;

[0049] Figure 5Schematic diagram of the UMAP graph and PAGA trajectory inference results of Example 2 of the present invention on the DLPFC, where (A) is the visualization result of the UMAP graph of the proposed m2ST on slice 151672; (B) is the visualization result of the UMAP graph of the proposed m2ST on slice 151672;

[0050] Figure 6 Schematic diagram of the results of the mouse hippocampus and cerebellum datasets of Example 3 of the present invention under the silhouette coefficient and the Davies-Bouldin index, where (A) is the clustering result of six methods on the mouse hippocampus dataset; (B) is the clustering result of six methods on the mouse cerebellum dataset;

[0051] Figure 7 Visualization results of Example 3 of the present invention, where (A) is a schematic diagram of the visualization results of the present invention on the mouse hippocampus dataset; (B) is a schematic diagram of the layered structures CA 1, CA 3, and DG of the mouse hippocampus; (C) is a schematic diagram of the visualization results of the present invention on the mouse cerebellum dataset. Detailed implementation manners

[0052] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0053] Example 1:

[0054] See Figures 1 to 2 , the technical solution provided in this embodiment is a spatial transcript data clustering method based on dual multi-scale graph learning, including the following steps:

[0055] S1: Based on spatial position information and gene expression data, a graph data is established where is the node set, A ∈ R N×N is the adjacency matrix, X ∈ R d×N is the node identity matrix, N is the number of nodes, and d is the dimension; and 0.5 of the nodes are randomly masked during the training phase of the features to obtain the graph data where is the masked feature;

[0056] S2: Based on the graph attention network, a multi-scale masked graph autoencoder is constructed. The input masked graph data is multi-scale processed by the encoder to extract the potential multi-scale embedding, and the potential embedding obtained by the encoder is re-masked by the decoder and passed into the decoding process, and the accurate reconstruction features are generated by the multi-scale decoder;

[0057] S3: Introduce a scaled cosine loss function to calculate the error between the reconstructed node features and the original node features

[0058] S4: According to the loss function obtained in step S3, use the Adam optimizer to update the network parameters and save the multi-scale mask graph autoencoder model;

[0059] S5: In the test phase, use the trained encoder to extract the multi-scale embeddings of the spatial transcriptomics data, mine the scale common information and scale-specific information from the multi-scale embeddings through the dual representation learning method, and input them into the multi-scale clustering method to achieve spatial domain annotation.

[0060] Specifically, the specific steps of step S2 are as follows:

[0061] S21: The present invention uses a graph attention network as the basic model to extract highly discriminative spatial transcriptome features with different scales; the multi-scale encoder consists of two parts: the first part is the common encoder where f E (*) represents that the common encoder extracts information at different scales and provides initial representation learning for spatial transcript data. The second part is the scale-specific encoder where f E_m (*) represents that the m-th encoder extracts information at different scales, and M = 2 is the number of scales; calculate the attention coefficient of each node in the graph:

[0062]

[0063] where, a i,j is the attention coefficient obtained by normalizing the similarity score, e i,j is all adjacent nodes, V i is the set of adjacent nodes obtained from the adjacency matrix A, LeakyReLU(*) is the activation function; e i,j is the similarity coefficient between nodes i and j, W ∈ R d′×d is the common mapping matrix, d ′ is the dimension of the mapping space, is the feature vector of the i-th node, || is the concatenation operation, set g(*) as a single-layer feedforward network, and θ is the learnable parameter;

[0064] S22: Similar to equation (2), calculate K coefficients to obtain the multi-head attention coefficients; after obtaining the multi-head attention coefficients, use the multi-head attention mechanism to update the features of each node:

[0065]

[0066] where, and W k are the attention coefficient and the common mapping matrix of the k-th head respectively, and σ(*) is the PreLU activation function that enhances the flexibility of the network, is the updated feature vector of the i-th node with a common encoder;

[0067] S23: represents the m-th extracted embedding, similar to f E (*), as follows:

[0068]

[0069] where, is the mapping matrix of the k-th head at the m-th scale, is the attention coefficient of the k-th head at the m-th scale, and is the mapped feature vector of the m-th encoder of the i-th node, d ′ m is the feature dimension;

[0070] S24: After obtaining the multi-scale embeddings compressed by the encoder, apply another set of masks to replace the node indices in the previous masks, that is Similar to (5), is defined as follows:

[0071]

[0072] Consistent with the encoder, use a single-layer graph attention network as the decoder for each scale; this method allows the model to recover the features of nodes based on a set of nodes rather than relying only on the nodes themselves, thus supporting the encoder to learn highly discriminative embeddings, that is where f D_m (*) represents the m-th encoder; the specific situation is as follows:

[0073]

[0074] where, is the attention coefficient of the k-th head at the m-th scale decoder, is the mapping matrix of the k-th head at the m-th scale decoder, is the feature vector reconstructed by the decoder at different scales.

[0075] Specifically, the specific steps of the step S3 are as follows:

[0076] S31: Introduce the scaled cosine error as the loss function and calculate the error between the reconstructed node features and the original node features:

[0077]

[0078] Among them, X i and are the original feature and the reconstructed feature, γ≥1 is the scale factor, used as a hyperparameter, and the L2 norm in the scaled cosine error maps the vector to the unit hypersphere.

[0079] Specifically, the specific steps of the step S5 are as follows:

[0080] S51: Based on the multi-scale masked graph autoencoder, obtain the multi-scale hidden embedding of the spatial transcriptome data; denote H m =(H m ) T , m = 1, 2, …, M = 2. Based on the matrix factorization technology, use a dual representation learning mechanism to simultaneously mine the common information and scale-specific information from the multi-scale data:

[0081]

[0082] Among them, is the common representation between views, is the specific representation of the m-th view, d c and d s are the feature dimensions of the common and specific representations, and are the mapping matrices of the m-th scale, and β is the regularization parameter; in addition, a regularization term is introduced to make the learned representation more robust;

[0083] S52: After extracting the common and specific representations between different scales, adopt a unified clustering framework that introduces the Shannon entropy mechanism and orthogonal constraints; set the common representation to the M + 1 scale, then the objective function is as follows:

[0084]

[0085] Among them, the fourth and fifth terms cluster the common representation and the specific representation respectively, and are the clustering centers of the common representation and the specific representation respectively, C is the number of clusters; U ∈ R C×N is the clustering indicator matrix shared by the two types of representations. When the j-th instance is clustered into the i-th class, U i,j = 1, otherwise, U i,j = 0; α m is the weight of different representations, I ∈ R C×C is the identity matrix, and the parameter values of β, λ, δ are in [1e - 5, 1e - 4, …, 1e5].

[0086] Example 2:

[0087] Referring to Embodiment 1, this embodiment will use the parameters calculated in Embodiment 1 to compare with other algorithms on the dorsolateral prefrontal cortex (DLPFC) dataset to prove the superiority of the present invention. The final results show that the present invention is preferred compared to other algorithms.

[0088] 1. Comparative Algorithms and Metrics

[0089] Referring to relevant research, the following seven comparative algorithms are selected in this embodiment:

[0090] 1) SpaGCN: A graph convolutional network method that combines gene expression, spatial location, and histology is proposed. It aggregates gene expression of adjacent points, identifies spatial domains, and performs differential expression analysis to effectively detect genes with rich spatial expression patterns. (Paper title: "SpaGCN: Integrating gene expression, spatial location and histology to identify spatial domains and spatially variable genes by graph convolutional network")

[0091] 2) DeepST: An accurate and general deep learning framework is proposed. By identifying spatial domains, it can refine the analysis of tissue spatial structure and improve the effect of spatial transcriptome data analysis. (Paper title: "DeepST: identifying spatial domains in spatial transcriptomics by deep learning")

[0092] 3) BayesSpace: A fully Bayesian statistical method is proposed. By using spatial neighborhood information, it enhances the resolution of spatial transcriptome data and performs clustering analysis, improving the identification of tissue transcriptional profiles. (Paper title: "Spatial transcriptomics at subspot resolution with BayesSpace")

[0093] 4) Seruat: A weighted nearest neighbor analysis method is proposed. By integrating multimodal data, it learns the relative importance of each data in cells and constructs a multimodal reference map of the circulatory immune system. (Paper title: "Integrated analysis of multimodal single-cell data")

[0094] 5) STAGATE: By integrating spatial information and gene expression, it uses graph attention mechanism to accurately identify spatial domains, improve the accuracy of spatial domain identification, denoise and retain spatial expression patterns. (Paper title: "Deciphering spatial domains from spatially resolved transcriptomics with an adaptive graph attention auto-encoder")

[0095] 6) CCST: Proposed a cell clustering method for spatial transcriptomic data based on graph convolutional network, which effectively improves the accuracy of cell clustering by using spatial information and gene expression, and discovers the interaction between cell subtypes and their microenvironments. (Paper title: "Cell clustering for spatial transcriptomics data with graph neural networks")

[0096] 7) stAA: Proposed an adversarial variational graph auto-encoder, which accurately identifies the boundaries of spatial domains by combining gene expression and spatial information, and uses graph neural network and Wasserstein distance to improve the accuracy of spatial clustering. (Paper title: "stAA: adversarial graph autoencoder for spatial clustering task of spatially resolved transcriptomics")

[0097] To verify the effectiveness of the present invention, NMI, ARI and Purity are used as evaluation metrics, where higher values indicate better performance.

[0098] 2. Comparison results

[0099] The DLPFC dataset consists of 12 slices, and a comprehensive comparison was made on all 12 slices. Figure 3 Shows the results of all methods for the ARI metric, and the results of all methods for the NMI and Purity metrics are as Figure 4 shown. It can be seen from the results that the present invention is always superior to other methods in all three metrics, and has a particularly significant advantage in the ARI metric. Compared with the graph auto-encoder-based methods (SpaGCN, stAA, CCST, DeepST), the performance of the present invention is the best, which indicates that exploring multi-scale information simultaneously is effective. At the same time, Figure 3 (B) Shows the visualization results of the present invention on 151672 slices. It can be seen from the figure that the clustering of the present invention is depicted better, and it can accurately identify the spatial domain structure of cells.

[0100] Finally, Figure 5 (A) and Figure 5 (B) are the UMAP plot and PAGA trajectory inference results of the present invention, respectively. The UMAP results clearly show different regions of each layer, indicating that the method of the present invention can effectively distinguish each layer domain. In addition, the PAGA plot reveals a linear trajectory from the WM to the third layer, further demonstrating that the developmental trajectory inferred by the present invention is highly consistent with the spatial topology of the slice.

[0101] Example 3:

[0102] Referring to Example 1, this example will use the parameters calculated in Example 1 to compare with the five comparative algorithms in Example 2 on the unlabeled mouse hippocampus and mouse cerebellum datasets to prove the superiority of the present invention. The final results show that the present invention is preferred compared to other algorithms.

[0103] 1. Comparative algorithms and metrics

[0104] Referring to Example 2, this example will select the five comparative algorithms in Example 2: SpaGCN, DeepST, STAGATE, CCST, and stAA. To verify the effectiveness of the present invention, NMI, ARI, and Purity are used as evaluation metrics, where higher values indicate better performance. In addition, for unlabeled datasets, the Silhouette Coefficient (SC) and Davies-Bouldin index (DB) are used as evaluation metrics. Specifically, a higher SC value indicates better clustering performance, while a lower DB value indicates better performance.

[0105] 2. Comparative results

[0106] First, the five methods were compared on the mouse hippocampus dataset. Figure 6 (A) shows the performance of the six algorithms in terms of the SC and DB metrics. From the results, the proposed method is superior to the other five algorithms in terms of the SC metric, especially exceeding SpaGCN by nearly 5%. At the same time, in terms of the DB metric, the value of the present invention is significantly lower than that of other methods (the lower the value, the better the performance), further verifying the effectiveness of the method on unlabeled datasets. Figure 7 (A) shows the visualization results of the present invention, indicating that it has the most obvious spatial depiction effect. In addition, as Figure 7 shown in the first figure of (B), the mouse hippocampus mainly consists of three regions: CA1, CA3, and dentate gyrus (DG), indicating that the present invention can accurately identify these three regions.

[0107] Figure 6(B) shows the clustering results of six methods on the mouse cerebellum dataset. Among these methods, the SC value of the present invention is significantly higher than that of other methods, while the DB value is the lowest, indicating its superior performance in spatial domain segmentation. Figure 7 (C) further shows the visualization results of the present invention, demonstrating its ability to effectively segment and identify complex spatial domains.

[0108] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A spatial transcription data clustering method based on dual multi-scale graph learning, characterized in that: The following steps are involved: S1: Building graph data based on spatial location information and gene expression data in is a set of nodes, A∈R N×N is the adjacency matrix, X∈R d×N is the node identity matrix, N is the number of nodes, and d is the dimension; And randomly mask a certain proportion of nodes during the training phase The features of in is a masking feature; S2: Based on the graph attention network, a multi-scale mask graph autoencoder is constructed. The encoder processes the input mask graph data at multiple scales to extract the potential multi-scale embedding. The decoder is used to re-mask the potential embedding obtained by the encoder and pass it into the decoding process. The multi-scale decoder generates accurate reconstruction features. S3: Introduce the scaled cosine loss function to calculate the error between the reconstructed node features and the original node features S4: According to the loss function obtained in step S3, the Adam optimizer is used to update the network parameters and save the multi-scale mask image autoencoder model; S5: In the testing phase, the trained encoder is used to extract multi-scale embeddings of spatial transcriptomics data. The scale-common information and scale-specific information are mined from the multi-scale embeddings through a dual representation learning method and input into a multi-scale clustering method to achieve spatial domain annotation.

2. The spatial transcription data clustering method based on dual multi-scale graph learning according to claim 1, characterized in that: The step S2 comprises the following steps: S21: Graph Attention Network is used as the basic model to extract high-resolution spatial transcriptome features at different scales. The multi-scale encoder consists of two parts: the first part is the common encoder where f E (*) indicates that the common encoder extracts information of different scales and provides initial representation learning for spatial transcription data. The second part is the scale-specific encoder. where f E_m (*) indicates that the mth encoder extracts information of different scales, M is the number of scales; the attention coefficient of each node in the graph is calculated: Among them, a i,j is the attention coefficient obtained by normalizing the similarity score, e i,j are all adjacent nodes, V i is the set of adjacent nodes obtained from the adjacency matrix A, LeakyReLU(*) is the activation function; e i,j is the similarity coefficient between nodes i and j, W∈R d′×d is the common mapping matrix, d ′ is the dimension of the mapping space, is the feature vector of the i-th node, || is the concatenation operation, g(*) is set to a single-layer feedforward network, and θ is a learnable parameter; S22: Similar to formula (2), calculate K coefficients Get the multi-head attention coefficient; after obtaining the multi-head attention coefficient, use the multi-head attention mechanism to update the features of each node: in, and W k are the attention coefficient and common mapping matrix of the kth head, σ(*) is the PreLU activation function that enhances the flexibility of the network, is the updated feature vector of the ith node with a common encoder; S23: represents the extracted mth embedding, and f E (*)Similar to the following: in, is the mapping matrix of the kth head at the mth scale, is the attention coefficient of the kth head at the mth scale, and is the mapping feature vector of the mth encoder of the i-th node, d ′ m is the characteristic dimension; S24: After obtaining the multi-scale embedding compressed by the encoder, another set of masks is applied to replace the node indices in the previous masks, i.e. Similar to (5), The definition is as follows: Consistent with the multi-scale encoder, a single-layer graph attention network is used as the decoder for each scale; this method allows the model to recover the features of a node based on a set of nodes, rather than relying solely on the node itself, thereby supporting the encoder to learn highly discriminative embeddings, i.e. where f D_m (*) indicates the mth encoder; the details are as follows: in, is the attention coefficient of the kth head at the mth scale decoder, is the mapping matrix of the kth head at the mth scale decoder, is the feature vector reconstructed by the decoder at different scales.

3. The spatial transcription data clustering method based on dual multi-scale graph learning according to claim 1 is characterized in that: The step S3 comprises the following steps: S31: Introduce scaled cosine error as the loss function to calculate the error between the reconstructed node features and the original node features: Among them, X i and are the original features and the reconstructed features, γ ≥ 1 is the scaling factor used as a hyperparameter to scale the L2 norm in the cosine error to map the vector onto the unit hypersphere.

4. The spatial transcription data clustering method based on dual multi-scale graph learning according to claim 1, characterized in that: The step S5 comprises the following steps: S51: Based on the multi-scale mask map autoencoder, we obtain the multi-scale hidden embedding of spatial transcriptome data; m =(H m ) T ,m=1,2,…,M, based on matrix decomposition technology, a dual representation learning mechanism is used to simultaneously mine common information and scale-specific information from multi-scale data: in, is a common representation between views, is the specific representation of the mth view, d c and d s is the characteristic dimension of common and specific representations, and is the mapping matrix of the mth scale, and β is the regularization parameter; in addition, the introduction of the regularization term makes the learned representation more robust; S52: After extracting the common and specific representations between different scales, a unified clustering framework is adopted that introduces the Shannon entropy mechanism and orthogonal constraints; the common representation is set to the M+1 scale, and the objective function is as follows: Among them, the fourth and fifth items cluster the common representation and specific representation respectively. and are the cluster centers of the public representation and the specific representation respectively, C is the number of clusters; U∈R C×N is a cluster indicator matrix shared by two types of representations. When the jth instance is clustered into the i-th class, U i,j =1, otherwise, U i,j =0; α m is the weight of different representations, I∈R C×C is the identity matrix, and β,λ,δ≥0 are parameters.

Citation Information

Cited By

  • Spatial transcriptome data spatial domain identification method based on multi-space self-supervised contrast learning

    CN120954500A