Multi-modal spatial domain identification method based on clustering-guided gradient contrast learning

The multimodal spatial domain identification method based on clustering-guided gradient contrastive learning solves the problem that the existing model focuses on individual differences while ignoring the semantic information of the same spatial domain, and achieves more efficient multimodal information integration and improved spatial domain identification accuracy.

CN120690281AActive Publication Date: 2025-09-23TIANJIN UNIV
3 Cites 0 Cited by

Patent Information

Application Number
CN202510793466.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-23
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Existing multimodal spatial domain identification methods rely on the supervision information provided by original contrastive learning in self-supervised learning, which causes the model to focus too much on individual differences and ignore the semantic information in the same spatial domain, and lack the adaptive ability to handle the differences between different samples.

Method used

The clustering-guided gradient contrastive learning method is adopted to construct a multimodal spatial domain identification model, including a Graph construction module, a GraphMAE training module, a multimodal representation fusion module and a clustering-guided gradient contrastive learning module. The K-nearest neighbor relationship is used to construct a graph, the graph masker and decoder are trained in a self-supervised manner, a modal attention fusion module is designed, and negative samples are selected through clustering guidance to enhance semantic consistency and difference.

Benefits of technology

It improves the quality and robustness of spatial domain identification, can better integrate multimodal information, and improves the accuracy and consistency of spatial domain identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690281A_ABST
    Figure CN120690281A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal spatial domain identification method based on clustering guide gradient contrast learning. The method comprises the following steps: step 1, constructing a data set; step 2, constructing a multi-modal spatial domain identification model; and step 3, training the multi-modal spatial domain identification model by using a data set to obtain multi-modal representation, and clustering by using an unsupervised clustering method mcluster to obtain a spatial domain identification result. According to the method, the distance between the samples and different cluster centers is adjusted, so that the semantic consistency between the samples of the same cluster and the difference between the samples of different clusters are increased; and an inter-modal attention fusion module is designed, the weight is adjusted according to the information amount of different modal data, adaptive fusion of multi-modal features is realized, the quality of spatial domain identification is improved, and good robustness is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular to a multimodal spatial domain identification method based on clustering-guided gradient contrastive learning. Background Art

[0002] The cells in organisms form complex ecosystems in the form of tissues. Different tissues have different cellular compositions and biological functions. Spatial domain identification is committed to depicting the spatial structure of tissues, helping people better understand the cell coordination mode, tissue structure and functional changes of organisms during growth, development, disease, aging, etc., which is of great significance.

[0003] Spatial transcriptome sequencing can measure gene expression at different locations while preserving spatial location information, facilitating spatial domain identification. Existing spatial transcriptome sequencing technologies fall into two main categories. The first category involves image-based sequencing technologies, such as STARmap, seqFISH, and MERFISH. These technologies can achieve subcellular resolution but only measure a subset of genes. The second category involves next-generation sequencing technologies, such as 10×Visium, SLIDE-seq V2, and Stereo-seq. These technologies can measure genome-wide gene expression but typically only achieve multi-cellular resolution. Spatial transcriptome data are often paired with pathology images, where the spatial transcriptome data reflects gene expression at different locations, while the pathology images reflect morphological information of cells and tissues.

[0004] Early spatial domain identification is represented by K-Means, Louvain's method, Seurat, etc. These methods do not use the spatial information in the data, and the identified spatial domains are usually discontinuous. Subsequent methods use graph neural networks to model spatial information at different positions, effectively solving the problem of discontinuous spatial domains. Depending on the learning method, these methods can be divided into unsupervised learning methods and self-supervised learning methods. Unsupervised learning methods are represented by Gitto, SpaGCN, stLearn, BayesSpace, STAGATE and CCST. These methods have no supervisory signals, and the accuracy of spatial domain identification is usually limited. Among the self-supervised learning methods, SpaceFlow maximizes the mutual information between a single spot representation and the average representation of all spots; ConST maximizes the mutual information between different spots at multiple levels such as local-local, local-global, and local-context;

[0005] GraphST uses contrastive learning to increase the similarity between a spot and its neighbors and reduce the similarity between a spot and its non-neighbors. By introducing self-supervisory information, these methods are able to extract more representative representations and generally have better performance.

[0006] Based on the different data used, existing methods can also be divided into single-modal methods and multi-modal methods. Single-modal methods such as SpaceFlow, GraphST, and SEDR only use spatial transcriptome data, ignoring the complementary information of other modalities, which limits the performance of spatial domain identification. Among the multi-modal methods, SpaGCN converts the mean and variance of the pathological image blocks corresponding to the spots into position coordinates, and constructs a weighted graph based on the original position information in the spatial transcriptome data; stLearn assumes that the gene expression of spots with similar morphology is more similar, and proposes to use pathological image features to correct the gene expression data; conST realizes the fusion of pathological and spatial transcriptome representations through concatenation; stMDA first aligns the pathological and spatial transcriptome representations at the local and global levels, and then uses the attention mechanism to fuse the modalities to obtain a common representation; STAIG calculates the distance between different spots through pathological image representation, and further converts the distance into a mask probability in replica generation. These methods integrate the complementary information in pathological images and spatial transcriptome data, and generally have better performance.

[0007] Although significant progress has been made, existing methods still have significant limitations. First, in terms of learning methods, although existing methods have introduced self-supervised learning techniques, they often rely on primitive contrastive learning methods to provide supervisory information, in which positive samples are usually different copies of the same sample, and different samples serve as negative samples for each other. However, in the task of spatial domain identification, the semantic information of samples in the same spatial domain is very similar. Using these samples as negative samples will cause the model to focus too much on the individual differences between different samples and ignore the common semantic information of samples in the same spatial domain. Secondly, pathological images and spatial transcriptomes reflect different information and have different values ​​in different samples. However, most existing multimodal methods lack adaptive capabilities and find it difficult to model the differences between different samples.

[0008] Glossary:

[0009] Spot: Data points in the spatial transcriptome.

[0010] RetCCL: A comparative clustering model. Summary of the Invention

[0011] To solve the above problems, the present invention discloses a multimodal spatial domain identification method based on clustering-guided gradient contrastive learning.

[0012] To achieve the above object, the technical solution of the present invention is:

[0013] A multimodal spatial domain identification method based on clustering-guided gradient contrastive learning includes the following steps:

[0014] Step 1: constructing a data set, wherein the data set includes spatial transcriptome data and corresponding pathological data;

[0015] Step 2: Construct a multimodal spatial domain identification model; the multimodal spatial domain identification model includes a Graph construction module, a GraphMAE training module, a multimodal representation fusion module, and a cluster-guided gradient contrast learning module in accordance with the data processing direction;

[0016] The Graph building module is used to construct the spatial transcriptome graph G st and pathological figure G hist ;

[0017] The GraphMAE training module is used to train the spatial transcriptome graph G according to the input st and pathological figure G hist Obtaining pathological latent space map Reconstructed pathological map, spatial transcriptome latent space map and reconstructed spatial transcriptome maps The multimodal representation fusion module is used to map the input pathological latent space and spatial transcriptome latent space graphs Obtain multimodal representation F m ;

[0018] Clustering-guided gradient contrastive learning module for multimodal representation F m Perform semantic information enhancement to obtain the updated multimodal representation F m' ;

[0019] Step 3: Use the dataset to train the multimodal spatial domain identification model to obtain the updated multimodal representation F m' Then use the unsupervised clustering method mclust to update the multimodal representation F m' Clustering is used to obtain spatial domain identification results.

[0020] For further improvement, the data processing process of the Graph building module is as follows:

[0021] According to the positional correspondence between the pathological image and the spatial transcriptome, the pathological image is cropped into pathological image blocks; then the pathological image representation of the pathological image blocks is extracted through the pre-training model, and the highly variable genes in the spatial transcriptome data are selected; finally, according to the spatial position relationship, the K nearest neighbor method is used to construct the pathological map G hist and spatial transcriptome maps G stPathological Figure G hist and spatial transcriptome maps G st Each node in the graph corresponds to a spot, and the edge represents the K-nearest neighbor relationship. The node representation in the pathology graph is the pathology image representation extracted by the pre-training model, and the node representation in the spatial transcriptome graph is the selected highly variable genes.

[0022] As a further improvement, the pre-training model is RetCCL, and the K value of the K-nearest neighbor method is selected as 3.

[0023] Further improvement, the GraphMAE training module includes a pathology training module and a spatial transcriptome graph training module; the pathology training module and the spatial transcriptome graph training module have the same structure, both including a masker, an encoder and a decoder; the masker is used to randomly mask the input graph to obtain a mask graph, and then the encoder characterizes and extracts the mask graph to obtain a latent space graph, and the decoder decodes the latent space graph to obtain a reconstructed graph; the latent space graph includes a pathology latent space graph and spatial transcriptome latent space graphs The reconstructed map includes a reconstructed pathology map and a reconstructed spatial transcriptome map

[0024] The GraphMAE training module performs self-supervised training, and the loss function L of self-supervised training SCE To scale the cosine error:

[0025]

[0026] Among them, V represents the set of masked nodes, x i represents the i-th node in the input latent space graph, z i represents the i-th node in the reconstructed graph, γ is a hyperparameter; T represents the matrix transpose, v i represents the i-th masked node; L SCE is the GraphMAE loss; GraphMAE loss includes For pathological GraphMAE loss and GraphMAE loss for the spatial transcriptome.

[0027] As a further improvement, both the encoder and decoder use a one-layer graph attention network, the feature dimension of the latent space is set to 128, and the value of γ is 3.

[0028] As a further improvement, the data processing method of the multimodal representation fusion module is as follows:

[0029] Input pathological latent space map and spatial transcriptome latent space graphs The corresponding node feature sets are and After passing through two mapping networks respectively, the representation dimension is reduced to 1 dimension, namely:

[0030]

[0031] Among them, f hist and f st are all vectors of length N, where N is the number of nodes in the graph; M hist () represents a fully connected network with one pathological layer, M st () represents a fully connected network of the spatial transcriptome, and the variance is used to measure the information content of the two modalities, namely:

[0032] v hist =Var(f hist ) (4)

[0033] v st =Var(f st ) (5)

[0034] Var() means calculating variance, v hist represents the variance of pathological representation after dimensionality reduction, v st represents the variance of the spatial transcriptome representation after dimensionality reduction;

[0035] And after Softmax normalization, the modal weight is obtained:

[0036] [w hist ,w st ]=softmax([v hist ,v st ]) (6)

[0037] Among them, w hist represents the weight of the pathological modality, w st represents the weight of the spatial transcriptome modality, and softmax() represents the normalized exponential function;

[0038] Finally, multimodal weighting is performed to obtain the multimodal representation F m :

[0039]

[0040] w hist represents the pathological modality weight, w st represents the spatial transcriptome modality weight.

[0041] 7. The multimodal spatial domain identification method based on cluster-guided gradient contrastive learning according to claim 1, wherein the data processing flow of the cluster-guided gradient contrastive learning module is as follows:

[0042] For the multimodal representation Fm , first use K-Means for clustering, the cluster number to which the i-th sample belongs is τ i , and calculate the cluster center of each cluster:

[0043]

[0044] Among them, N n is the number of samples in cluster n, is the jth sample in cluster n; C n is the cluster center of cluster n; for the i-th multimodal representation Randomly select a positive sample from the neighbors of the same cluster, and the corresponding representation is The cluster centers of all clusters are regarded as negative samples, and the corresponding representation is C n , the set of all negative samples is S neg ; In order to increase the semantic consistency of samples within a cluster and the semantic difference of samples between clusters, different negative samples are distinguished by different weights. For the cluster center of cluster n:

[0045]

[0046] Among them, β n is the weight corresponding to the cluster center of cluster n, ε is a hyperparameter with a value between 0 and 1;

[0047] The loss function L of the clustering-guided gradient contrastive learning module Con for:

[0048]

[0049] , exp() represents the natural exponential function, sim() is the similarity function; C neg Represents the set of negative samples.

[0050] As a further improvement, in step 3, the GraphMAE training module is first pre-trained, and then the multimodal spatial domain identification model is trained as a whole;

[0051] The loss function of pre-training is:

[0052]

[0053] α hist is the pathological GraphMAE loss weight, α st is the spatial transcriptome GraphMAE loss weight, For pathological GraphMAE loss, GraphMAE loss for the spatial transcriptome;

[0054] The loss function L of the overall training is:

[0055]

[0056] α Con is the contrastive learning loss weight;

[0057] When the loss function L of the overall training is minimized, the updated multimodal representation F is obtained m' Then the unsupervised clustering method mclust is used to categorize the updated multimodal representation F m' Clustering is used to obtain spatial domain identification results.

[0058] Advantages of the present invention:

[0059] The present invention increases the semantic consistency between samples in the same cluster and the difference between samples in different clusters by adjusting the distance between samples and different cluster centers; and designs an inter-modal attention fusion module to adjust its weight according to the amount of information of different modal data, thereby realizing the adaptive fusion of multimodal features, improving the quality of spatial domain identification, and having good robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 Schematic diagram of the method of the present invention, including: (a) pathological and spatial transcriptome graph construction; (b) clustering-guided contrastive learning; (c) modality attention fusion module.

[0061] Figure 2 The performance comparison between the single-modality and multi-modality methods of the present invention is shown in FIG.

[0062] Figure 3 These are the experimental results under different parameters.

[0063] Figure 4 GraphST is a performance comparison chart of the method of the present invention and the method.

[0064] Figure 5 The results for samples 151508, 151671, and 151673 are visualized. (a) is the ground truth for sample 151508, and (b) is the identification result for sample 151508; (c) is the ground truth for sample 151671, and (d) is the identification result for sample 151671; (e) is the ground truth for sample 151673, and (f) is the identification result for sample 151673. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0066] 1. The method of the present invention is as follows:

[0067] 1.1 Method Overview

[0068] The principle of the proposed multimodal spatial domain identification method based on clustering guided contrastive learning is as follows Figure 1 As shown in the figure, the method consists of four steps: (1) construction of pathological and spatial transcriptome graphs; (2) training of pathological and spatial transcriptome encoder-decoders based on GraphMAE; (3) multimodal representation fusion; and (4) clustering-guided gradient contrastive learning.

[0069] 1.2Graph Construction

[0070] In order to better utilize spatial information, this method intends to use graph neural networks. Therefore, graph structure data is first constructed for pathology and spatial transcriptome data. First, the pathology image is cropped into small image blocks according to the positional correspondence between the pathology image and the spatial transcriptome. Then, the pre-trained model is used to extract the representation of the pathology image blocks and select the highly variable genes in the spatial transcriptome data. Finally, the K-nearest neighbor method is used to construct the pathology graph G according to the spatial position relationship. hist and spatial transcriptome maps G st Each node in the pathology map and the spatial transcriptome map corresponds to a spot, and the edges represent the K-nearest neighbor relationship. The node representation in the pathology map is the pathology image representation extracted by the pre-training model, and the node representation in the spatial transcriptome map is the selected highly variable genes.

[0071] In the specific implementation process, in order to cover the complete area of ​​a spot, the size of the pathological image block is set to 100×100μm 2 The pathological image pre-training model selected is RetCCL, K in K nearest neighbors is selected as 3, and 3000 highly variable genes are selected using seurat_v3.

[0072] 1.3 GraphMAE Training

[0073] To obtain a good initial representation of both pathology images and spatial transcriptome data, this method first uses a Graph Masked AutoEncoder (GraphMAE) for self-supervised training on both modalities. The model structure is identical for both modalities, and the following example uses spatial transcriptome data as an example.

[0074] The spatial transcriptome GraphMAE consists of three modules: masker, encoder, and decoder. st , the masker first randomly masks the nodes in the graph and obtains Afterwards, the encoder E st right Perform representation extraction to obtain the latent space graph Finally, the decoder D st According to the latent space graph Decode and obtain the reconstructed spatial transcriptome map

[0075] The loss function of GraphMAE is scaled cosine error (SCE), which is calculated as follows:

[0076]

[0077] Among them, V represents the set of masked nodes, x i Represents the input spatial transcriptome graph G st The i-th node in z i Representation of the reconstructed spatial transcriptome map The i-th node in , γ is a hyperparameter.

[0078] In the specific implementation process, the masking rate in the masker is set to 0.6, that is, 60% of the nodes are randomly selected for masking; both the encoder and decoder use a one-layer Graph Attentional Network (GAT), and the feature dimension of the latent space is set to 128; γ in the loss function is set to 3.

[0079] 1.4 Multimodal Representation Fusion

[0080] In different samples, the value of pathology and spatial transcriptome for spatial domain identification is different. In order to better model the differences between samples, this method uses the modality attention fusion module to adaptively adjust the weights of different modalities. The structure of the modality attention fusion module is as follows: Figure 1 (c) shows the input pathological latent space map. and spatial transcriptome latent space graphs The corresponding node feature sets are and After passing through two mapping networks respectively, the representation dimension is reduced to 1 dimension, namely:

[0081]

[0082] Among them, f hist and f st are both vectors of length N, where N is the number of nodes in the graph. Then, the variance is used to measure the information content of the two modes, namely:

[0083] v hist =Var(f hist ) (4)

[0084] v st =Var(f st ) (5)

[0085] And after Softmax normalization, the modal weight is obtained:

[0086] [w hist ,w st ]=softmax([v hist ,v st ]) (6)

[0087] Finally, multimodal weighting is performed:

[0088]

[0089] Through the above methods, this method can adaptively adjust the modal weights according to the amount of information in different modalities. hist () and M st () are all one-layer fully connected networks.

[0090] 1.5 Clustering-guided Gradient Contrastive Learning

[0091] In conventional contrastive learning, positive samples are often different copies of the same sample, while different samples are set as negative samples. However, in the task of spatial domain identification, samples belonging to the same spatial domain have very similar semantic information, causing them to become negative samples, affecting model performance. To address this problem, this method proposes cluster-guided contrastive learning, which reduces the impact of negative samples by selecting reliable negative samples.

[0092] For the obtained multimodal representation F m First, use K-Means to cluster it, and the cluster number to which the i-th sample belongs is τ i , and calculate the cluster center of each cluster:

[0093]

[0094] Among them, N n is the number of samples in cluster n, is the jth sample in cluster n. For each sample representation We randomly select one of its neighbors in the same cluster as a positive sample, and the corresponding representation is The cluster centers of all clusters are regarded as negative samples, and the corresponding representation is C n , the set of all negative samples is S neg In order to increase the semantic consistency of samples within a cluster and the semantic difference of samples between clusters, we distinguish different negative samples by different weights. Specifically, for sample i:

[0095]

[0096] Among them, β n is the weight corresponding to the cluster center of cluster n, and ε is a hyperparameter with a value between 0 and 1.

[0097] The loss function of cluster-guided contrastive learning is:

[0098]

[0099] In the specific implementation process, the number of K-Means clusters is specified by prior knowledge.

[0100] 1.6 Training Process

[0101] The model training is divided into two stages. In the first 500 epochs, only the encoder and decoder of the pathological and spatial transcriptomes are trained using GraphMAE, and the corresponding training loss is:

[0102]

[0103] in, and are the reconstruction losses of pathological and spatial transcriptomes, respectively, and α hist and α st is the corresponding weight. In this method, the value is 1. From the 500th epoch to the 1200th epoch, the entire model is trained. The loss function at this time is:

[0104]

[0105] Among them, L Con is the contrastive learning loss, α Con For contrastive learning loss weight, the value is taken as 0.05 in this method.

[0106] 2. Performance comparison evaluation:

[0107] 2.1 Data and Evaluation Metrics

[0108] To evaluate the performance of our method, we conducted experiments on the human dorsolateral prefrontal cortex (DLPFC) dataset. The DLPFC dataset, acquired using 10× Visium technology, contains 12 slices, each containing between 3,460 and 4,789 spots, capturing a total of 33,538 genes. Each slice in the DLPFC dataset is divided into five to seven spatial domains, corresponding to multiple cortical layers and one white matter layer.

[0109] We use the Adjusted Rand Index (ARI) to evaluate the performance of this method. The ARI takes into account the possibility that two clusters are consistent by chance. The ARI ranges from 0 to 1, where the closer the ARI value is to 1, the more similar the clusters are. It is calculated as follows:

[0110]

[0111] Among them, TP is the number of spots that belong to the same cluster in both the true label and the predicted result, TN is the number of spot pairs that belong to different clusters in both the true label and the predicted result, FN is the number of spot pairs that belong to the same cluster in the true result but different clusters in the predicted result, and FP is the number of spot pairs that belong to different clusters in the true result but the same cluster in the predicted result. E is the expected value of this indicator, that is, the value under completely random clustering, and is calculated as follows:

[0112]

[0113] 2.2 Comparison of single-modal and multi-modal performance

[0114] We first compared the effects of spatial transcriptome single modality and spatial transcriptome plus pathology multimodality. Figure 2 In most samples, the multimodal method outperforms the single-modal method. Specifically, the average ARI of the single-modal method is 0.602, while the average ARI of the multimodal method is 0.622. This shows that integrating pathological and spatial transcriptomic data can improve the quality of spatial domain identification, and our proposed method can effectively integrate multimodal information.

[0115] 2.3 Results of different parameters

[0116] In the proposed method, the weight of the negative samples in the same cluster is a key parameter. Therefore, we compared the performance of different weights. The results are shown in the figure. Figure 3As shown in Figure 2, the proposed method can achieve good results under different weights. Specifically, when the weights are 0.5, 0.7, and 0.9, the corresponding average ARIs are 0.613, 0.622, and 0.608, respectively. This shows that the proposed method is robust to the weights of negative samples in the same cluster.

[0117] 3.4 Comparison of different methods

[0118] We further compared the proposed method with the state-of-the-art method (GraphST), and the results are shown in Figure 2. Figure 4 As shown in the figure, the proposed method outperforms GraphST on most samples. Specifically, the average ARI of GraphST is 0.558, while that of the proposed method is 0.622, which fully demonstrates the excellent performance of the proposed method.

[0119] 3.5 Visualization Results

[0120] In order to intuitively evaluate the spatial domain identification effect of the proposed method, we visualized the results of three samples 151508, 151671 and 151673. The results are as follows: Figure 5 As shown in Figure 2, on all three samples, the proposed method can accurately achieve spatial domain identification and has good consistency with the true labels, which intuitively demonstrates the performance of the proposed method.

[0121] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and the embodiments. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and shown here.

Claims

1. A multimodal spatial domain identification method based on clustering-guided gradient contrastive learning, characterized in that: The steps include: Step 1: constructing a data set, wherein the data set includes spatial transcriptome data and corresponding pathological data; Step 2: Construct a multimodal spatial domain identification model; the multimodal spatial domain identification model includes a Graph construction module, a GraphMAE training module, a multimodal representation fusion module, and a cluster-guided gradient contrast learning module in accordance with the data processing direction; The Graph building module is used to construct the spatial transcriptome graph G st and pathological figure G hist ; The GraphMAE training module is used to train the spatial transcriptome graph G according to the input st and pathological figure G hist Obtaining pathological latent space map Reconstructed pathological map, spatial transcriptome latent space map and reconstructed spatial transcriptome maps The multimodal representation fusion module is used to map the input pathological latent space and spatial transcriptome latent space graphs Obtain multimodal representation F m ; Clustering-guided gradient contrastive learning module for multimodal representation F m Perform semantic information enhancement to obtain the updated multimodal representation F m' ; Step 3: Use the dataset to train the multimodal spatial domain identification model to obtain the updated multimodal representation F m' Then use the unsupervised clustering method mclust to update the multimodal representation F m' Clustering is used to obtain spatial domain identification results.

2. The multimodal spatial domain identification method based on cluster-guided gradient contrastive learning according to claim 1, characterized in that: The data processing process of the Graph building module is as follows: According to the positional correspondence between the pathological image and the spatial transcriptome, the pathological image is cropped into pathological image blocks; then the pathological image representation of the pathological image blocks is extracted through the pre-training model, and the highly variable genes in the spatial transcriptome data are selected; finally, according to the spatial position relationship, the K nearest neighbor method is used to construct the pathological map G hist and spatial transcriptome maps G st Pathological Figure G hist and spatial transcriptome maps G st Each node in the graph corresponds to a spot, and the edge represents the K-nearest neighbor relationship. The node representation in the pathology graph is the pathology image representation extracted by the pre-training model, and the node representation in the spatial transcriptome graph is the selected highly variable genes.

3. The multimodal spatial domain identification method based on cluster-guided gradient contrastive learning according to claim 2, characterized in that: The pre-trained model is RetCCL, and the K value of the K-nearest neighbor method is selected as 3.

4. The multimodal spatial domain identification method based on cluster-guided gradient contrastive learning according to claim 1, characterized in that: The GraphMAE training module includes a pathology training module and a spatial transcriptome graph training module; the pathology training module and the spatial transcriptome graph training module have the same structure, both including a masker, an encoder and a decoder; the masker is used to randomly mask the input graph to obtain a mask graph, and then the encoder characterizes and extracts the mask graph to obtain a latent space graph, and the decoder decodes the latent space graph to obtain a reconstructed graph; the latent space graph includes a pathology latent space graph and spatial transcriptome latent space graphs The reconstructed map includes a reconstructed pathology map and a reconstructed spatial transcriptome map The GraphMAE training module performs self-supervised training, and the loss function L of self-supervised training SCE To scale the cosine error: Among them, V represents the set of masked nodes, x i represents the i-th node in the input latent space graph, z i represents the i-th node in the reconstructed graph, γ is a hyperparameter; T represents the matrix transpose, v i represents the i-th masked node; L SCE is the GraphMAE loss; GraphMAE loss includes For pathological GraphMAE loss and GraphMAE loss for the spatial transcriptome.

5. The multimodal spatial domain identification method based on cluster-guided gradient contrastive learning according to claim 4, characterized in that: Both the encoder and decoder use a one-layer graph attention network, the feature dimension of the latent space is set to 128, and the value of γ is 3.

6. The multimodal spatial domain identification method based on cluster-guided gradient contrastive learning according to claim 1, characterized in that: The data processing method of the multimodal representation fusion module is as follows: Input pathological latent space map and spatial transcriptome latent space graphs The corresponding node feature sets are and After passing through two mapping networks respectively, the representation dimension is reduced to 1 dimension, namely: Among them, f hist and f st are all vectors of length N, where N is the number of nodes in the graph; M hist () represents a fully connected network with one pathological layer, M st () represents a fully connected network of the spatial transcriptome, and the variance is used to measure the information content of the two modalities, namely: v hist =There is(f hist ) (4) v st =There is(f st ) (5) Var() means calculating variance, v hist represents the variance of pathological representation after dimensionality reduction, v st represents the variance of the spatial transcriptome representation after dimensionality reduction; And after Softmax normalization, the modal weight is obtained: [w hist ,w st ]=softmax([v hist ,v st ]) (6) Among them, w hist represents the weight of the pathological modality, w st represents the weight of the spatial transcriptome modality, and softmax() represents the normalized exponential function; Finally, multimodal weighting is performed to obtain the multimodal representation F m : w hist represents the pathological modality weight, w st represents the spatial transcriptome modality weight.

7. The multimodal spatial domain identification method based on cluster-guided gradient contrastive learning according to claim 1, characterized in that: The data processing flow of the cluster-guided gradient contrastive learning module is as follows: For the multimodal representation F m , first use K-Means for clustering, the cluster number to which the i-th sample belongs is τ i , and calculate the cluster center of each cluster: Among them, N n is the number of samples in cluster n, is the jth sample in cluster n; C n is the cluster center of cluster n; for the i-th multimodal representation Randomly select a positive sample from the neighbors of the same cluster, and the corresponding representation is The cluster centers of all clusters are regarded as negative samples, and the corresponding representation is C n , the set of all negative samples is S neg ; In order to increase the semantic consistency of samples within a cluster and the semantic difference of samples between clusters, different negative samples are distinguished by different weights. For the cluster center of cluster n: Among them, β n is the weight corresponding to the cluster center of cluster n, ε is a hyperparameter with a value between 0 and 1; The loss function L of the clustering-guided gradient contrastive learning module Con for: , exp() represents the natural exponential function, sim() is the similarity function; C neg Represents the set of negative samples.

8. The multimodal spatial domain identification method based on cluster-guided gradient contrastive learning according to claim 1, characterized in that: In the step 3, the GraphMAE training module is first pre-trained, and then the multimodal spatial domain identification model is trained as a whole; The loss function of pre-training is: α hist is the pathological GraphMAE loss weight, α st is the spatial transcriptome GraphMAE loss weight, For pathological GraphMAE loss, GraphMAE loss for the spatial transcriptome; The loss function L of the overall training is: α Con is the contrastive learning loss weight; When the loss function L of the overall training is minimized, the updated multimodal representation F is obtained m' Then the unsupervised clustering method mclust is used to categorize the updated multimodal representation F m' Clustering is used to obtain spatial domain identification results.

Citation Information

Patent Citations

  • Spatial transcriptome spot region clustering method fusing image gene data

    CN116312782A

  • Spatial transcriptomics cell clustering method based on multi-scale contrast learning

    CN120015132A

  • Inferring super-resolution tissue architecture by integrating spatial transcriptomics with histology

    US20250014681A1