Spatial transcriptome data clustering method based on progressive learning and multi-modal fusion

Through the method of progressive learning and multimodal fusion, the graph convolution network and cross-attention mechanism are used to solve the problem of poor clustering effect in spatial transcriptomic clustering analysis, achieving more efficient fusion of gene expression and organizational structure information, and improving cluster accuracy and robustness.

CN120256988AActive Publication Date: 2025-07-04YUNNAN UNIV

Patent Information

Application Number
CN202510747820.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing spatial transcriptomic clustering analysis methods have insufficient clustering effect, and cannot effectively utilize the spatial background information provided by histological images, and are limited by data sparseness and noise patterns, resulting in poor clustering effect.

Method used

Using a method based on progressive learning and multimodal fusion, high-order structural information of histological images and gene expression data is extracted through graph convolutional networks, and combined with complementary masks and cross-attention mechanisms, neural network models are constructed for data fusion and cluster analysis.

Benefits of technology

It improves the accuracy and robustness of clustering, can better understand the relationship between gene expression and tissue structure, improves clustering performance and model generalization ability, and reduces noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256988A_ABST
    Figure CN120256988A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of spatial transcriptome data clustering analysis, and discloses a spatial transcriptome data clustering method based on progressive learning and multi-modal fusion. According to the method, a high-order structure relationship among sample points is captured by using a multilayer graph convolutional network, and various noises of gene expression data are removed to the greatest extent through comparative learning and a complementary mask mechanism; meanwhile, gene expression data, spatial position information and histological image information are ingeniously fused to the maximum degree through two times of cross attention, and then fusion feature latent representation Latent is obtained. Compared with a mainstream spatial transcriptomics clustering analysis method, the method has the advantages that clustering based on the fusion feature latent representation Latent shows higher capability and excellent clustering performance, and has remarkable robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of spatial transcriptome data clustering methods, and particularly to a spatial transcriptome data clustering method based on progressive learning and multimodal fusion. Background Art

[0002] Spatial Transcriptomics (ST) is one of the important breakthroughs in the field of bioinformatics in recent years. Different from traditional single-cell RNA sequencing (scRNA-seq), spatial transcriptome technology not only provides gene expression information but also retains the spatial positional relationship of cells in tissues, which plays a key role in revealing cell-cell interactions, tissue microenvironment structure, and spatial heterogeneity of diseases, and has been widely applied in many fields such as cancer research, neuroscience, and developmental biology.

[0003] Identifying shared and specific spatial domains (clusters with similar spatial expression patterns) is one of the important tasks for the comprehensive analysis of ST datasets. The latest progress in spatial-resolved transcriptomics has made it possible to comprehensively measure gene expression patterns while retaining the spatial context of the tissue microenvironment. However, due to the limitations of sequencing technology, spatial transcriptome data has a high degree of sparsity and complex noise patterns, and traditional clustering methods such as kmeans and hierarchical clustering rely on similarity metrics and cannot well meet the requirements of single-cell data clustering.

[0004] To better identify the spatial domains of spatial transcriptome data, deep clustering methods such as SpaGCN, GraphST, SEDR, and STAGATE have been proposed and used for tissue spatial domain clustering tasks, and further utilize the spatial position information of spatial transcriptome data to cluster spatial transcriptome data. However, the histological images of spatial transcriptomics provide an intuitive, anatomical spatial background for gene expression data, and this combination can help the model more comprehensively understand the relationship between gene expression and tissue structure, but these methods all ignore this important modal data.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0006] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a spatial transcriptome data clustering method based on progressive learning and multimodal fusion, aiming to solve the problem of poor clustering effect in existing spatial transcriptomics clustering analysis methods.

[0007] The first aspect of the present invention provides a method for clustering spatial transcriptome data based on progressive learning and multimodal fusion, including: obtaining histological images and their corresponding spatial transcriptome data; cutting the histological images into blocks according to the size of sample points, and then extracting histological image feature information from each cut small block through the medical image large model UNI ; calculating the spatial Euclidean distance between each pair of sample points according to the spatial position information of each sample point, selecting the k nearest sample points of each sample point as neighbors, and constructing a spatial adjacency matrix A; screening and regularizing the gene expression data of sample points in the spatial transcriptome data, performing principal component analysis on the gene expression matrix of the top 3000 highly variable genes, and extracting the gene expression matrix of the top 200 principal components as gene expression information X; randomly masking each row of the gene expression information X according to a predetermined ratio to obtain a masked gene expression matrix , and then performing complementary masking on the gene expression information X to obtain a complementary masked gene expression matrix ; constructing a neural network model, the neural network model includes a first encoder, a second encoder, a third encoder, a decoder, a first cross-attention fusion module, a second cross-attention fusion module, a progressive learning module, and a contrast learning module; using the spatial adjacency matrix A, histological image feature information , masked gene expression matrix , and complementary masked gene expression matrix as input data, using the reconstruction loss and the contrast loss as the total loss function, training the neural network model to obtain a trained neural network model; inputting the histological image feature information , spatial adjacency matrix A, and gene expression information X obtained after preprocessing the histological image and the corresponding spatial transcriptome data to be analyzed into the trained neural network model, and outputting a fused feature latent representation; performing clustering analysis on the fused feature latent representation and outputting a clustering result.

[0008] Optionally, in the first implementation manner of the first aspect of the present invention, using the spatial adjacency matrix, histological image feature information, masked gene expression matrix, and complementary masked gene expression matrix as input data, and using the reconstruction loss and contrast loss as the total loss function, training the neural network model to obtain a trained neural network model, including the steps of: inputting the histological image feature information and the spatial adjacency matrix into the first encoder to output a latent representation of the image feature information; inputting the masked gene expression matrix and the spatial adjacency matrix into the second encoder, and after the latent representation of the masked gene output is remasked, inputting it into the third encoder to obtain a re-masked gene latent representation; inputting the re-masked gene latent representation into the progressive learning module to output an average re-masked gene latent representation; inputting the average re-masked gene latent representation and the latent representation of the image feature information into the first cross-attention fusion module to output a first fusion latent representation; inputting the first fusion latent representation and the re-masked gene latent representation into the second cross-attention fusion module to output a fusion feature latent representation; inputting the fusion feature latent representation into the decoder to output a reconstructed feature representation; inputting the complementary masked gene expression matrix and the spatial adjacency matrix into the second encoder to output a complementary masked gene latent representation; inputting the complementary masked gene latent representation and the re-masked gene latent representation into the contrast learning module to output the cosine similarity of the positive sample pair and the negative sample pair, and constructing a contrast loss based on the cosine similarity of the positive sample pair and the negative sample pair ; constructing a reconstruction loss based on the reconstructed feature representation ; based on the reconstruction loss function and the contrast loss function constructing a total loss function training the neural network model to obtain a trained neural network model, where and are hyperparameters that control the proportions of the reconstruction loss and the contrast loss in the total loss

[0009] Optionally, in the second implementation manner of the first aspect of the present invention, inputting the histological image feature information and the spatial adjacency matrix into the first encoder to output a latent representation of the image feature information, including the steps of: the first encoder includes several layers of graph convolutional layers, and performing a convolutional operation on the input histological image feature information and spatial adjacency matrix through the several layers of graph convolutional layers to obtain a latent representation of the image feature information, where the convolutional operation expression is: , where is the feature representation of all sample points in the current layer, is the feature representation after all sample points have undergone a convolutional operation once, , , I is the identity matrix, is the degree matrix of, is the activation function, is the trainable parameter matrix of the current layer.

[0010] Optionally, in the third implementation manner of the first aspect of the present invention, the masked gene expression matrix and the spatial adjacency matrix are input into the second encoder, and the output masked gene latent representation is input into the third encoder after re-masking processing to obtain a re-masked gene latent representation, including the steps: The second encoder includes several layers of graph convolutional layers, and the input masked gene expression matrix and spatial adjacency matrix are convolved through the several layers of graph convolutional layers to obtain a masked gene latent representation; the masked gene latent representation is re-masked, and the masked rows of the masked gene latent representation are the same as the masked rows in the masked gene table matrix, to obtain a re-masked gene latent representation; the third encoder includes several layers of graph convolutional layers, and the input re-masked gene latent representation is convolved through the several layers of graph convolutional layers to obtain a re-masked gene latent representation.

[0011] Optionally, in the fourth implementation manner of the first aspect of the present invention, the re-masked gene latent representation is input into the progressive learning module, and an average re-masked gene latent representation is output, including the steps: The progressive learning module randomly selects a certain percentage of neighbor points for each sample point according to the increase in the number of training rounds and based on the spatial adjacency matrix and averages the re-masked gene latent representations of the neighbor points to obtain an average re-masked gene latent representation, where the selection of the percentage is set according to the following formula: , where, is the number of rounds required for training the neural network model, is the current round of training of the neural network model.

[0012] Optionally, in the fifth implementation manner of the first aspect of the present invention, the average re-masked gene latent representation and the image feature information latent representation are input into the first cross-attention fusion module, and a primary fusion latent representation is output, including the steps: The first cross-attention fusion module is used to capture the correlation between different feature spaces, and its calculation formula is as follows: , where Q, K, and V are the query matrix, key matrix, and value matrix respectively, is a weight matrix, represents the matrix concatenation operation; the average re-masked gene latent representation and the image feature information latent representation are input into the first cross-attention fusion module, and a primary fusion latent representation is output, where the image feature information latent representation serves as the data source of the query matrix Q, and the average re-masked gene latent representation serves as the data source of the key matrix K and the value matrix V.

[0013] Optionally, in a sixth implementation of the first aspect of the present invention, the complementary masked gene latent representation and the heavily masked gene latent representation are input into the contrastive learning module, the cosine similarity of the positive sample pair and the negative sample pair is output, and a contrast loss is constructed based on the cosine similarity of the positive sample pair and the negative sample pair. , comprising the steps of: selecting corresponding sample rows in the complementary masked gene latent representation and the heavy masked gene latent representation to construct positive sample pairs according to the masked sample rows in the heavy masked gene latent representation, and randomly extracting sample rows that are inconsistent with the positive sample pairs from the complementary masked gene latent representation and the heavy masked gene latent representation as negative sample pairs; according to the formula Calculate the cosine similarity of positive sample pairs and negative sample pairs, where (A, B) is a sample pair, is the modulus symbol; by pulling the cosine similarity of the positive sample closer , push the negative sample to cosine similarity The contrast loss is constructed in the following way .

[0014] The second aspect of the present invention provides a spatial transcriptome data clustering device based on progressive learning and multimodal fusion, comprising: a data acquisition module for acquiring a histological image and its corresponding spatial transcriptome data; an image processing module for cutting the histological image into blocks according to the size of the sample points, and then extracting histological image feature information from each small block through a medical image large model UNI ; an adjacency matrix construction module, used to calculate the spatial Euclidean distance between each sample point according to the spatial position information of each sample point, select the k nearest sample points of each sample point as neighbors, and construct a spatial adjacency matrix A; a gene expression data processing module, used to screen and regularize the gene expression data of the sample points in the spatial transcriptome data, select the gene expression matrix of the top 3000 highly variable genes for principal component analysis, and extract the top 200 principal component gene expression matrices as gene expression information X; a mask module, used to randomly mask each row of the gene expression information X according to a predetermined ratio to obtain a masked gene expression matrix , and then perform complementary masking on the gene expression information X to obtain the complementary masked gene expression matrix ; A model building module, used to build a neural network model, the neural network model includes a first encoder, a second encoder, a third encoder, a decoder, a first cross attention fusion module, a second cross attention fusion module, a progressive learning module and a contrast learning module; a model training module, used to use the spatial adjacency matrix A, the histological image feature information , masked gene expression matrix and complementary masked gene expression matrix As input data, using the reconstruction loss and the contrastive loss as the total loss function, the neural network model is trained to obtain a trained neural network model; a feature output module for obtaining histological image feature information after preprocessing the histological image to be analyzed and the corresponding spatial transcriptome data , the spatial adjacency matrix A and the gene expression information X are input into the trained neural network model to output a fused feature latent representation; an aggregation module for performing clustering analysis on the fused feature latent representation and outputting a clustering result.

[0015] The third aspect of the present invention provides a spatial transcriptome data clustering device based on progressive learning and multi-modal fusion, including: a memory and at least one processor, wherein computer-readable instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor calls the computer-readable instructions in the memory to enable the spatial transcriptome data clustering device based on progressive learning and multi-modal fusion to execute each step of the spatial transcriptome data clustering method as described above.

[0016] The fourth aspect of the present invention provides a computer-readable storage medium, in which computer-readable instructions are stored, and when it runs on a computer, it enables the computer to execute each step of the spatial transcriptome data clustering method as described above.

[0017] Beneficial effects: The present invention provides a spatial transcriptome data clustering method based on progressive learning and multi-modal fusion. First, the present invention introduces a graph convolutional network as an encoder and a decoder, enabling the model to fully extract the information of adjacent nodes to capture higher-order structural information between sample points; the gene expression information and the image feature information are respectively fused with the structural information (spatial adjacency matrix), allowing the information of the two modalities to first complete the extraction of their own spatial information. Second, when the gene expression information is encoded by the GCN, complementary masked data is used. The complementary data makes the most of the samples and provides different perspectives for the model to learn its data structure, making it easier for the model to capture the key patterns of the data and facilitating the elimination of possible technical noise (such as sequencing errors) or biological noise. Moreover, due to the data missing caused by the mask, the complementary mask can ensure that the model still maintains spatial consistency during the training process, so that the learned fused feature latent representations will not be damaged by the randomness of the missing values and promotes the model to utilize more spatial information; at the same time, the model learns to still capture the key patterns of the data in the case of partial information missing, thereby improving the generalization ability and robustness. Third, the present invention designs a progressive learning module, which is designed based on the idea from local to global. This module will randomly provide adjacent node information from less to more for the next step of modality fusion as the number of training rounds increases. Since there are large modality differences between multi-modal features, if all features are jointly modeled in one stage, the model may be difficult to converge or easily fall into a local optimum. The way of randomly selecting from less to more can first perform a preliminary alignment or pre-training on each part of the representation, and then gradually "unfreeze" or "combine" more elements, so as to obtain a relatively stable and more controllable update at each stage. At the same time, it also avoids the model "thinking too carefully about the problem at first and instead ignoring the global structure". Finally, the present invention also designs a dual cross-attention modality fusion mechanism. The first cross-attention fusion module enables the image features to perceive the overall distribution or global semantics of the gene expression at the "preliminary" stage, so as to "highlight" the part of the image features that is highly correlated with the overall gene expression, which is equivalent to a coarse-grained or "global" cross-modal fusion. The model first learns the correspondence between the overall image and the global information of the gene expression, allowing the image representation to first incorporate a macroscopic understanding of the gene expression; the second cross-attention fusion module then enables the model to "refine" the alignment of the image and the specific dimensions or pathways of the gene expression, so as to capture more discriminative or more interpretable cross-modal patterns, which is equivalent to a fine-grained cross-modal fusion; on the premise that the model has obtained the global context in the first step, it can better grasp the correspondence between each pathway / gene dimension of the gene expression and the local features of the image.The latent representation of the fused features obtained based on the model shows stronger ability and excellent clustering performance in cell clustering after clustering analysis, and has remarkable robustness. Description of the Drawings

[0018] Figure 1 It is a flowchart of the spatial transcriptome data clustering method based on progressive learning and multi-modal fusion provided by the embodiment of the present invention.

[0019] Figure 2 It is a schematic diagram of the principle of the spatial transcriptome data clustering method based on progressive learning and multi-modal fusion of the present invention.

[0020] Figure 3 It is a graph of the ARI and NMI clustering evaluation results of the DLPFC (human dorsolateral prefrontal cortex) multi-slice dataset under different methods. Among them, A is the manual annotation classification of the DLPFC 151675 slices; B is the box plot of the clustering ARI and NMI evaluation results of all slices of the DLPFC multi-slice dataset under different methods (the method of the present invention, CCST, Deep ST, DiffusionST, GraphST, SEDR, SpaGCN, STAGATE); C is the visualization display of the spatial domain clustering results of the 151675 slices of the DLPFC for each method.

[0021] Figure 4 It is the ARI and NMI clustering evaluation results of the single-slice dataset under different methods. Among them, Figure 4 on the left is the ARI and NMI clustering evaluation result graph of the BRCA (human breast cancer) single-slice dataset under different methods, Figure 4 and on the right is the ARI and NMI clustering evaluation result graph of the MBA (mouse brain) single-slice dataset under different methods.

[0022] Figure 5 It is a schematic diagram of the structure of the spatial transcriptome data clustering device based on progressive learning and multi-modal fusion provided by the embodiment of the present invention.

[0023] Figure 6 It is a schematic diagram of the structure of the spatial transcriptome data clustering device based on progressive learning and multi-modal fusion provided by the embodiment of the present invention. Detailed Embodiments

[0024] An embodiment of the present invention provides a method for clustering spatial transcriptome data based on progressive learning and multimodal fusion. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the term "including" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] Please refer to Figure 1 , Figure 1 which is a flowchart of a preferred embodiment of a method for clustering spatial transcriptome data based on progressive learning and multimodal fusion provided by the present invention. As shown in the figure, it includes the steps: S10. Obtain histological images and their corresponding spatial transcriptome data; Specifically, histological images are usually obtained through microscopic imaging techniques (such as H&E staining, fluorescence staining, etc.), which show the morphological characteristics of cells and structures in tissue sections. Its function is to provide a "map" for spatial localization, which is used for the subsequent spatial analysis of gene expression data; the spatial transcriptome data is obtained through high-throughput sequencing technology, which records the gene expression profiles of each sample point (spot) in the tissue section, and the position information (such as coordinates) of each sample point corresponds one-to-one with the physical position in the histological image. As an example, a tissue section is placed on a Visium chip, and the surface of the chip is covered with sample points containing oligonucleotide probes. The oligonucleotide probes capture the mRNA released from the section and assign a unique spatial barcode to each sample point. By sequencing, the gene expression profile of each sample point is obtained, and the expression data is mapped to the corresponding spatial coordinates using the barcode of the probe, and the spatial transcriptome data can be obtained, which correspondingly includes a gene expression matrix (each row represents a gene, and each column represents a sample point); spatial coordinate data (x / y coordinates of each sample point).

[0026] S20. Cut the histological image according to the size of the sample points, and then extract the histological image feature information from each cut small block through the medical image large model UNI ; Specifically, the histological image is centered according to the spatial position information corresponding to its sample points and cut according to the size of the sample points, and then the image feature information is extracted from each cut small block through the medical image large model UNI as For example, according to the coordinates of the spatial transcriptome data (such as the sample point positions in Visium), centered on each sample point, image patches are cropped according to the actual size (such as a diameter of 55 μm). In this embodiment, the medical image large model UNI is a general pre-trained model designed for medical images. UNI incorporates prior knowledge of medical images (such as tissue staining patterns, cell arrangement rules) during pre-training, can better distinguish pathological features (such as tumor regions, inflammatory responses), and UNI has stronger robustness to common noises (such as uneven staining, section folding) and artifacts (such as bubbles, knife marks) in medical images; even if the cut blocks come from different laboratories or staining protocols (such as H&E vs. IHC), the features extracted by UNI still remain consistent, reducing data batch effects.

[0027] S30. Calculate the spatial Euclidean distance between each pair according to the spatial position information of each sample point, select the k nearest sample points of each sample point as neighbors, and construct a spatial adjacency matrix A; Specifically, on the spatial transcriptome section, calculate the spatial Euclidean distance between sample points according to the spatial position information of each sample point, and use the K-nearest neighbor algorithm to ensure that each sample point has k (for example, k is 6 - 12) nearest sample points as neighbors, that is, the spatial adjacency matrix is obtained. There is a spatial adjacency matrix A on the nth spatial transcriptome section. n , if the jth sample point is a neighbor of the ith sample point, then A nij = 1, otherwise A nij = 0, that is .

[0028] S40. Screen and regularize the gene expression data of the sample points in the spatial transcriptome data, perform principal component analysis on the gene expression matrix of the top 3000 highly variable genes, and extract the first 200 principal component gene expression matrices as gene expression information X; Specifically, first screen out the sample points without expressed genes and the genes without expression in the spatial transcriptome data, then use the functions provided by the SCANPY library to regularize the screened gene expression data, and then select the gene expression matrix of the top 3000 highly variable genes , and finally perform principal component analysis (PCA) on this gene expression matrix , and extract the first 200 principal components (PCs) gene expression matrices as the basic input of the model. .

[0029] S50. Randomly mask each row of the gene expression information X according to a predetermined ratio to obtain a masked gene expression matrix. , and then perform complementary masking on the gene expression information X to obtain a complementary masked gene expression matrix ; In this embodiment, taking the gene expression information X including 5 sample points, namely sample point 1, sample point 2, sample point 3, sample point 4, and sample point 5 as an example, as Figure 2 shown, first mask the gene expressions corresponding to sample point 2 and sample point 4 in the gene expression information X, that is, obtain a masked gene expression matrix ; then perform complementary masking on the gene expressions corresponding to sample point 1, sample point 3, and sample point 5 in the gene expression information X, that is, obtain a complementary masked gene expression matrix , and at this time .

[0030] In this embodiment, when the gene expression information is encoded by GCN, complementary masked data is used. The complementary masked data makes the most of the samples and provides different perspectives for the model to learn its data structure, making it easier for the model to capture the key patterns of the data and facilitating the elimination of possible technical noise (such as sequencing errors) or biological noise. Moreover, due to the data missing generated by masking, complementary masking can ensure that the model still maintains spatial consistency during the training process, so that the learned latent representations will not be damaged by the randomness of the missing values and promotes the model to utilize more spatial information. At the same time, the model must learn to capture the key patterns of the data even when part of the information is missing, thereby improving the generalization ability and robustness.

[0031] S60. Construct a neural network model, where the neural network model includes a first encoder, a second encoder, a third encoder, a decoder, a first cross-attention fusion module, a second cross-attention fusion module, a progressive learning module, and a contrast learning module; S70. Use the spatial adjacency matrix A, histological image feature information , masked gene expression matrix , and complementary masked gene expression matrix as input data, and use the reconstruction loss and the contrast loss as the total loss function to train the neural network model to obtain a trained neural network model; Specifically, as Figure 2 shown, the first encoder includes several layers of graph convolutional layers. During the training process, input the histological image feature information and the spatial adjacency matrix A into the first encoder, and perform convolutional operations on the input data through the several layers of graph convolutional layers to obtain a latent representation of the image feature information, where the convolutional operation expression is: , where is the feature representation of all sample points in the current layer, is the feature representation indicating all sample points after one convolution operation, , , where I is the identity matrix, is the degree matrix of, is the activation function, is the trainable parameter matrix of the current layer.

[0032] As Figure 2 shown, the second encoder and the first encoder also include several layers of graph convolutional layers. During training, the masked gene expression matrix and the spatial adjacency matrix A are input into the second encoder. Through the several layers of graph convolutional layers, convolution operations are performed on the input masked gene expression matrix and spatial adjacency matrix to obtain the masked gene latent representation; then, the masked gene latent representation is remasked. The rows of the masked gene latent representation that are masked are the same as the rows of the masked gene table matrix, obtaining the remasked gene latent representation; the third encoder includes several layers of graph convolutional layers. The remasked gene latent representation is input into the third encoder, and through the several layers of graph convolutional layers, convolution operations are performed on the input remasked gene latent representation again to obtain the re-remasked gene latent representation. In this embodiment, by remasking the masked gene expression matrix, the robustness of the model can be further improved. The obtained re-remasked gene latent representation has multiple functions: it can be used as the positive sample of the contrast learning module, can also be used as the data source of the progressive learning module, and can also be used as the input for cross-fusion with the subsequent one-time fusion latent representation.

[0033] As Figure 2 stated, the progressive learning module is a dynamic information extraction module. Its idea comes from curriculum learning (CL). CL is a machine learning method inspired by the human learning process. Its core idea is that the model gradually learns from simple to complex, from easy to difficult samples or tasks, rather than being directly exposed to all data. This strategy can improve the convergence speed, generalization ability of the model, and avoid falling into local optima. During training, the re-remasked gene latent representation is input into the progressive learning module, and the average re-remasked gene latent representation can be output. Specifically, the progressive learning module will, according to the increase in the number of training epochs, randomly select a corresponding certain percentage (increasing with the epoch) of neighbor points for each sample point according to the spatial adjacency matrix and take the average of the re-remasked gene latent representations of these neighbor points as the condition for fusing with the image feature information latent representation; the selection of the percentage is set according to the following formula: , where, is the number of rounds required for training the neural network model. is the current round of training of the neural network model. This design is to enable the model to extract as much information as possible from each neighbor point. Through a progressive ratio, the model gradually adapts to the increasing information; and through random selection, the model does not only rely on the information of a specific few neighbor points, thereby enhancing the robustness of the model. Finally, a 10% margin is provided to allow the model to adapt to the full amount of data input; after training is completed, it is assumed that the model has adapted to the full amount of data input, will be continuously set to 1 and will not be changed again.

[0034] This embodiment designs a progressive learning module, which is a design based on the idea of from local to global. This progressive learning module will randomly provide adjacent node information from less to more for the next modal fusion as the number of training rounds increases. Since there are large modal differences between multi-modal features, if all features are jointly modeled in one stage, the model may be difficult to converge or easily fall into local optimality. And in this embodiment, through random selection and the way from less to more, it is possible to first perform preliminary alignment or pre-training on the representations of each part, and then gradually "thaw" or "combine" more elements, so as to obtain relatively stable and more controllable updates at each stage. At the same time, it also avoids the model "thinking too carefully about the problem first and instead ignoring the global structure".

[0035] The latent representations after separately encoding the histological image feature information and gene expression data are both independent. How to fuse the two parts of data is a very important issue. The fusion of multi-modal data has always been a difficult task. Simply adding the two outputs to obtain a fused representation is a simple approach, but this approach may lose some important expression patterns or structural information, and may even cause a large number of irrelevant dimensions or mismatched noise features to be mixed together. Thanks to the wide application of Transformer, so in this embodiment, the data of the two modalities are fused through a cross-attention mechanism. The cross-attention will automatically assign higher weights to the parts with higher correlation during the matching process, thereby playing a screening and weighting effect, weakening the noise, greatly improving the fusion quality, and obtaining a richer representation.

[0036] The attention mechanism, simply put, maps a query and a set of key-value pairs to obtain an output, where the query, key, value, and output are all vectors. Specifically, by calculating the dot product of the query and the key and applying the softmax function to obtain the relevant weights of the values, the final result is obtained by multiplying these weights by the values. To simplify the calculation, in the actual operation process, a set of key-value pairs are calculated simultaneously each time, converting the vector calculation into a matrix operation. The specific calculation method of the attention mechanism is as follows: , where are the query (Query) matrix, key (Key) matrix, and value (Value) matrix respectively.

[0037] On this basis, the present invention further introduces the multi-head attention mechanism, which can simultaneously focus on the information in each subspace at different positions of the data, so that the obtained attention fusion representation contains richer and more robust information. Specifically, the present invention will perform M times of different attention operations respectively, and then execute the attention function in parallel, connect them and project again to obtain the final result. The operation of the th head in each attention calculation is expressed as follows: , where is a transformation matrix, which is responsible for projecting the corresponding three matrices respectively; the calculation of each multi-head attention (MultiHead) is as follows: , where is a weight matrix, represents the matrix concatenation operation.

[0038] Cross-Attention is a special attention mechanism widely used in tasks such as multi-modal learning, sequence alignment, and data fusion. This mechanism allows the query (Query, Q) of one data source to focus on the key (Key, K) and value (Value, V) of another data source, thereby establishing an interaction relationship between the two feature spaces. The main difference between cross-attention and self-attention is that the Q, K, and V of self-attention often come from the same data source, while the Q, K, and V of cross-attention come from different data sources. The first cross-attention fusion module and the second cross-attention module of the present invention both use the cross-attention mechanism to effectively capture the correlation between different modalities or different feature spaces.

[0039] Such as Figure 2As shown, the first cross-attention fusion module is used to capture the correlation between different feature spaces. During training, the average re-masked gene latent representation and the image feature information latent representation are input into the first cross-attention fusion module, and a first-stage fusion latent representation is output. Among them, the image feature information latent representation serves as the data source for the query matrix Q, and the average re-masked gene latent representation serves as the data sources for the key matrix K and the value matrix V.

[0040] Then, the first-stage fusion latent representation and the re-masked gene latent representation are input into the second cross-attention fusion module, and a fused feature latent representation Latent is output. Among them, the re-masked gene latent representation serves as the data source for the query matrix Q in the second cross-attention fusion module, and the first-stage fusion latent representation serves as the data sources for the key matrix K and the value matrix V in the second cross-attention fusion module. The fused feature latent representation Latent output by the second cross-attention has important uses. First, during the training stage, it will be used as the decoder input for reconstructing the original gene expression. When the training stage is completed, the output fused feature latent representation Latent can be used for downstream task analysis such as spatial domain recognition and visualization clustering.

[0041] This embodiment designs a dual cross-attention modality fusion mechanism. The first cross-attention fusion module enables image features to perceive the overall distribution or global semantics of gene expression at the "preliminary" stage, thus "highlighting" the parts in the image features that are highly correlated with the overall gene expression. This is equivalent to a coarse-grained or "global" cross-modal fusion. The model first learns the correspondence between the overall image and the global information of gene expression, allowing the image representation to first incorporate a macroscopic understanding of gene expression. The second cross-attention fusion module then enables the model to "refinely" align the image with the specific dimensions or pathways of gene expression, thereby capturing more discriminative or interpretable cross-modal patterns, which is equivalent to a fine-grained cross-modal fusion. On the premise that the model has obtained the global context in the first step, it can better grasp the correspondence between each pathway / gene dimension of gene expression and the local features of the image.

[0042] As Figure 2 shown, during the training process, the fused feature latent representation is input into the decoder, and a reconstructed feature representation is output, and a reconstruction loss is constructed based on the reconstructed feature representation. Specifically, the decoder also consists of a multi-layer graph convolutional layer (GCN) as the backbone module, with the fused feature latent representation as the input of this component, aiming to reconstruct the original gene expression matrix and obtain the reconstructed feature representation. The process is as follows: , where there is ; when the reconstructed feature representation Z is obtained, it will be used to calculate the reconstruction loss 。To reconstruct the masked features from the given partially observable input features, this embodiment uses the Scaled Cosine Error (SCE) as the objective function. The normalized cosine error enhances the stability of the embedded representation learning. Under the predefined scale factor γ, the similarity between the reconstruction prediction and the original input is calculated only on the masked genes, and the reconstruction loss has the following mathematical formula: , where is each sample point in the masked gene expression matrix (sample point), is the corresponding sample point in the reconstructed feature representation to , γ is fixed at 2 throughout the experiment to reduce the weight of the contributions from simple samples during the training process, and |V| represents the number of nodes in the masked set, refers to transpose of

[0043] As Figure 2 shown, during the training process, the complementary masked gene expression matrix and the spatial adjacency matrix are input into the second encoder to output the complementary masked gene latent representation; the complementary masked gene latent representation and the re-masked gene latent representation are input into the contrastive learning module to output the cosine similarity of the positive sample pairs and negative sample pairs, and a contrastive loss is constructed based on the cosine similarity of the positive sample pairs and negative sample pairs. Specifically, contrastive learning is a self-supervised learning method that learns the similarity or difference between samples by constructing positive samples (positive pairs) and negative samples (negative pairs). The goal is to make similar data points close in the latent space and dissimilar data points far away. In this embodiment, there are two data input sources for the contrastive learning module, namely the complementary masked gene latent representation and the re-masked gene latent representation. According to the masked sample rows in the re-masked gene latent representation, corresponding sample rows are selected in the complementary masked gene latent representation and the re-masked gene latent representation to construct positive sample pairs, and at the same time, sample rows inconsistent with the positive sample pairs are randomly drawn from the complementary masked gene latent representation and the re-masked gene latent representation as negative sample pairs. As an example, as Figure 2As shown, the masked sample behavior in the heavy-masked gene latent representation is sample point 2 and sample point 4, so the corresponding sample point 2 and sample point 4 are selected from the complementary masked gene latent representation and the heavy-masked gene latent representation as positive sample pairs, and sample rows inconsistent with the positive sample pairs are randomly extracted from the complementary masked gene latent representation and the heavy-masked gene latent representation as negative sample pairs. For example, sample point 1 in the complementary masked gene latent representation can be combined with any sample point in the heavy-masked gene latent representation as a negative sample pair, and the complementary masked gene latent representation can be combined with any sample point in the heavy-masked gene latent representation as a negative sample pair. Sample point 2 in the latent representation can be combined with any sample point in the heavy-masked gene latent representation except sample point 2 as a negative sample pair, sample point 3 in the complementary masked gene latent representation can be combined with any sample point in the heavy-masked gene latent representation as a negative sample pair, sample point 4 in the complementary masked gene latent representation can be combined with any sample point in the heavy-masked gene latent representation except sample point 4 as a negative sample pair, and sample point 5 in the complementary masked gene latent representation can be combined with any sample point in the heavy-masked gene latent representation as a negative sample pair.

[0044] Furthermore, according to the formula Calculate the cosine similarity of positive sample pairs and negative sample pairs, where (A, B) is the modulus symbol; by pulling the cosine similarity of the positive sample closer , push the negative sample to cosine similarity The contrastive loss is constructed in this way.

[0045] The loss function plays a vital role in the training process of deep learning models. It is used to measure the gap between the model's prediction results and the true value and guide the optimization of model parameters. For the special design of this model, the total loss function consists of two parts: reconstruction loss and And Contrastive Loss , to ensure that the model can effectively interpolate missing spatial transcriptome data and learn reasonable spatial dependencies. In order to allow the model to learn spatial relationships, this embodiment introduces contrast loss so that points with similar spatial positions or expression characteristics are close in the latent space, while points with different expression patterns are far away. The contrast loss is also divided into two parts. The first part is the cosine distance between positive sample pairs, and the second part is the cosine distance between negative sample pairs. By shortening the cosine distance between positive sample pairs , push away the cosine distance of negative samples In this way, it is ensured that while the model learns the spatial structure, it can effectively distinguish different gene expression patterns; and by providing complementary masked data sources, local noise interference is reduced, and the self-supervised learning ability and model generalization ability are improved.

[0046] Since the core task during the training of this model is to impute missing gene expression values, a loss function that can measure the error between the predicted value and the true value is required. In this embodiment, the Scaled Cosine Error (SCE) is used to optimize the imputation result. The SCE loss measures the direction consistency through cosine similarity, combines an exponential weighting mechanism to dynamically adjust the sample weights, is robust to noise and amplitude changes, and pays more attention to semantic alignment rather than numerical exact matching, which is very suitable for scenarios such as self-supervised learning. Thus, in this embodiment, there is a total loss function , where and are hyperparameters that control the proportions of the reconstruction loss and the contrast loss in the total loss.

[0047] S80. The histological image features information obtained after preprocessing the histological image to be analyzed and the corresponding spatial transcriptome data , the spatial adjacency matrix A, and the gene expression information X are input into the trained neural network model, and a fused feature latent representation is output; After the neural network model is trained, negative samples will no longer be used, and there is no longer a need to mask the gene expression information X. Only by following the same steps as in the previous training steps, the histological image to be analyzed and the corresponding spatial transcriptome data are preprocessed to obtain the corresponding histological image features information , the spatial adjacency matrix A, and the gene expression information X; then the histological image features information , the spatial adjacency matrix A are input into the first encoder of the trained neural network model, and the gene expression information X and the spatial adjacency matrix A are input into the second encoder of the trained neural network model. After being processed by each module in the trained neural network model in sequence, finally, after the feature fusion of the second cross-attention fusion module, a fused feature latent representation Latent is output.

[0048] S90. Perform clustering analysis on the fused feature latent representation and output the clustering result.

[0049] The present invention utilizes the large medical image model UNI that has emerged in recent years to extract image feature information, and extracts the features of histological image information to the greatest extent; and uses a multi-layer graph convolutional network (Graph Convolutional Networks, GCN) to capture the high-order structural relationships between each sample point (sample point), and then through contrast learning and complementary mask mechanism, various noises in the gene expression data itself are removed to the greatest extent; at the same time, the histological image feature information and gene expression information are cleverly fused through two cross-attention mechanisms. Among them, the first cross-attention dynamically fuses the neighboring gene latent representations of the sample points as conditions with the image feature latent representations, and the second time fuses the fusion result of the first time as conditions with the gene expression latent representations. Thus, the gene expression data, spatial position information and histological image information are fused to the greatest extent, that is, the fused feature latent representation Latent is obtained. In order to more accurately analyze the potential structure of spatial transcriptome data, this embodiment can adopt the mclust clustering method to cluster the fused feature latent representation Latent learned by the trained neural network model. mclust is a clustering method based on the Gaussian Mixture Model (GMM), which can automatically select the optimal number of clusters in the high-dimensional latent space and provide the probability distribution structure of the data.

[0050] In the deep learning environment based on PyTorch, the clustering performance differences between the method of the present invention (OURS) and other mainstream methods (CCST, Deep ST, DiffusionST, GraphST, SEDR, SpaGCN, STAGATE) are compared through specific experiments below. This embodiment uses different datasets of multi-slice and single-slice to evaluate the generalization ability and robustness of the algorithm. The experimental datasets are shown in Table 1, and each group of data is evaluated in detail through clustering performance evaluation indicators and chart analysis.

[0051] Table 1 Statistical information of spatial transcript data

[0052] When evaluating unsupervised clustering algorithms, NMI (Normalized Mutual Information) and ARI (Adjusted Rand Index) are two commonly used indicators for measuring the similarity degree between the clustering result and the true label. Their definitions and calculation formulas will be introduced in detail below: 1. ARI (Adjusted Rand Index) is an indicator used to compare the similarity between two clustering results, which considers the consistency between pairwise samples in the clustering results. Its calculation formula is shown as follows: , where, represents the number of sample pairs that belong to the same cluster in the th clustering result, represents the number of sample pairs that belong to the same class in the th true label, represents the number of sample pairs that are assigned to the th clustering result and the th true label, represents the number of combinations of choosing elements from elements, is the total number of samples. The value of this metric ranges between 0 and 1, and the higher the value, the higher the similarity between the clustering result and the true label.

[0053] 2. NMI (Normalized Mutual Information) is a metric used to evaluate the quality of clustering. It is based on the concept of information theory and measures the normalized value of the mutual information between the clustering result and the true label. Its calculation formula is as follows: , where, represents the set of clusters in the clustering result, represents the number of clusters, represents the number of classes in the true label, represents the number of samples that belong to both the th cluster in the clustering and the th true label, represents the number of samples in the th cluster in the clustering, represents the number of samples in the th true label, represents the total number of samples.

[0054] Both of the above metrics are values between 0 and 1, and the higher the value, the higher the similarity between the clustering result and the true label.

[0055] Figure 3 The charts in Figure 3 show the ARI and NMI clustering evaluation results of the DLPFC (human dorsolateral prefrontal cortex) multi-slice dataset under different methods, and the clustering results of its 151675 slices are visually presented in Figure 3In it, A is the manual annotation classification of the DLPFC 151675 slices; B is the box plot of the clustering ARI and NMI evaluation results of all slices of the DLPFC multi-slice dataset under different methods (the method of the present invention, CCST, Deep ST, DiffusionST, GraphST, SEDR, SpaGCN, STAGATE); C is the visual display of the spatial domain clustering results of the 151675 slices of DLPFC for each method.

[0056] Figure 4 are the ARI and NMI clustering evaluation results of the single-slice dataset under different methods. Among them, Figure 4 on the left in it is the ARI and NMI clustering evaluation result diagram of the BRCA (human breast cancer) single-slice dataset under different methods, Figure 4 and on the right in it is the ARI and NMI clustering evaluation result diagram of the MBA (mouse brain) single-slice dataset under different methods.

[0057] From Figure 3 and Figure 4 the results, it can be seen that compared with the mainstream spatial transcriptomics clustering analysis methods, the method of the present invention shows stronger ability and excellent performance in cell clustering, and has remarkable robustness. Specifically, in order to quantify the overall performance of different methods in the spatial clustering task, we calculated the ARI and NMI of all 12 slices of the DLPFC dataset and visualized the results as Figure 3 the box plot of B in it. Compared with the seven state-of-the-art methods, the present invention achieved higher averages and medians on both of these two metrics, and at the same time showed lower variances, which indicates its excellent consistency and hierarchical resolution among slices.

[0058] Figure 3 C in it intuitively compares the spatial domain recognition results of different methods on 151675 slices. The clustering boundaries obtained by the present invention are closely aligned with the manually annotated layers, the transition between layers is clear, and the noise interference is minimal. In contrast, some baseline methods (such as SEDR, SpaGCN) show "mixed" or "cross-domain" clustering in the boundary region, resulting in significantly lower accuracy than the present invention.

[0059] To further demonstrate the robustness of the present invention, we also compared the present invention with the baseline methods on more datasets BRCA ( Figure 4 left side), MBA ( Figure 4 right side), and the present invention still achieved the best performance on ARI and NMI, demonstrating remarkable robustness.

[0060] The above described the method for clustering spatial transcriptome data based on progressive learning and multimodal fusion in the embodiments of the present invention. Next, the device for clustering spatial transcriptome data based on progressive learning and multimodal fusion in the embodiments of the present invention will be described. Please refer to Figure 5 One embodiment of the device for clustering spatial transcriptome data based on progressive learning and multimodal fusion in the embodiments of the present invention includes: A data acquisition module 10, configured to acquire histological images and their corresponding spatial transcriptome data; An image processing module 20, configured to cut the histological image according to the size of sample points, and then extract histological image feature information from each cut small block through a medical image large model UNI ; An adjacency matrix construction module 30, configured to calculate the spatial Euclidean distance between each pair according to the spatial position information of each sample point, select the k nearest sample points of each sample point as neighbors, and construct a spatial adjacency matrix A; A gene expression data processing module 40, configured to screen and regularize the gene expression data of sample points in the spatial transcriptome data, perform principal component analysis on the gene expression matrix of the top 3000 highly variable genes, and extract the first 200 principal component gene expression matrices as gene expression information X; A masking module 50, configured to randomly mask each row of the gene expression information X according to a predetermined ratio to obtain a masked gene expression matrix , and then perform complementary masking on the gene expression information X to obtain a complementary masked gene expression matrix ; A model construction module 60, configured to construct a neural network model, where the neural network model includes a first encoder, a second encoder, a third encoder, a decoder, a first cross-attention fusion module, a second cross-attention fusion module, a progressive learning module, and a contrast learning module; A model training module 70, configured to use the spatial adjacency matrix A, histological image feature information , masked gene expression matrix , and complementary masked gene expression matrix as input data, and use the reconstruction loss and the contrast loss as the total loss function to train the neural network model to obtain a trained neural network model; A feature output module 80, configured to input the histological image feature information obtained after preprocessing the histological image and the corresponding spatial transcriptome data to be analyzed, the spatial adjacency matrix A, and the gene expression information X into the trained neural network model, and output a fused feature latent representation; An aggregation module 90 is configured to perform clustering analysis on the fused feature latent representation and output a clustering result.

[0061] above Figure 5 The spatial transcriptome data clustering device based on progressive learning and multi-modal fusion in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the spatial transcriptome data clustering device based on progressive learning and multi-modal fusion in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0062] Figure 6 FIG. is a schematic structural diagram of a spatial transcriptome data clustering device based on progressive learning and multi-modal fusion provided by an embodiment of the present invention. The spatial transcriptome data clustering device 100 based on progressive learning and multi-modal fusion may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 11 (for example, one or more processors) and a memory 12, and one or more storage media 13 for storing application programs 133 or data 132 (for example, one or more mass storage devices). Among them, the memory 12 and the storage media 13 may be transient storage or persistent storage. The program stored in the storage media 13 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the spatial transcriptome data clustering device 100 based on progressive learning and multi-modal fusion. Further, the processor 11 may be configured to communicate with the storage media 13 and execute a series of instruction operations in the storage media 13 on the spatial transcriptome data clustering device 100.

[0063] The spatial transcriptome data clustering device 100 based on progressive learning and multi-modal fusion may further include one or more power supplies 14, one or more wired or wireless network interfaces 15, one or more input / output interfaces 16, and / or one or more operating systems 131, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 6 The shown device structure does not limit the spatial transcriptome data clustering device 100 based on progressive learning and multi-modal fusion, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0064] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the method for clustering spatial transcriptome data based on progressive learning and multimodal fusion.

[0065] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, or unit can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0066] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0067] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for clustering spatial transcriptome data based on progressive learning and multimodal fusion, characterized in that, Including steps: Obtain histological images and their corresponding spatial transcriptome data; Cut the histological image into pieces according to the size of the sample points, and then extract the histological image feature information of each cut piece through the medical image large model UNI ; Calculate the spatial Euclidean distance between each pair of sample points according to the spatial position information of each sample point, select the k nearest sample points of each sample point as neighbors, and construct a spatial adjacency matrix A; Screen and regularize the gene expression data of the sample points in the spatial transcriptome data, perform principal component analysis on the gene expression matrix of the top 3000 highly variable genes, and extract the top 200 principal component gene expression matrices as gene expression information X; Randomly mask each row of the gene expression information X according to a predetermined ratio to obtain a masked gene expression matrix , and then perform complementary masking on the gene expression information X to obtain a complementary masked gene expression matrix ; Construct a neural network model, which includes a first encoder, a second encoder, a third encoder, a decoder, a first cross-attention fusion module, a second cross-attention fusion module, a progressive learning module, and a contrast learning module; Using the spatial adjacency matrix A, histological image feature information , masked gene expression matrix and complementary masked gene expression matrix as input data, and using the reconstruction loss and contrast loss as the total loss function, training the neural network model to obtain a trained neural network model; The histological image feature information obtained after preprocessing the histological image to be analyzed and the corresponding spatial transcriptome data , the spatial adjacency matrix A, and the gene expression information X are input into the trained neural network model, and a fused feature latent representation is output; Perform clustering analysis on the fused feature latent representation and output the clustering result.

2. The method for clustering spatial transcriptome data based on progressive learning and multimodal fusion according to claim 1, wherein Using the spatial adjacency matrix, histological image feature information, masked gene expression matrix, and complementary masked gene expression matrix as input data, and using the reconstruction loss and contrast loss as the total loss function, train the neural network model to obtain a trained neural network model, including the steps: Input the histological image feature information and the spatial adjacency matrix into the first encoder, and output the latent representation of the image feature information; Input the masked gene expression matrix and the spatial adjacency matrix into the second encoder, and the output masked gene latent representation is input into the third encoder after remasking to obtain a remasked gene latent representation; Input the remasked gene latent representation into the progressive learning module, and output the average remasked gene latent representation; Input the average remasked gene latent representation and the latent representation of the image feature information into the first cross-attention fusion module, and output the first fusion latent representation; Input the first fusion latent representation and the remasked gene latent representation into the second cross-attention fusion module, and output the fused feature latent representation; Input the fused feature latent representation into the decoder, and output the reconstructed feature representation; Input the complementary masked gene expression matrix and the spatial adjacency matrix into the second encoder, and output the complementary masked gene latent representation; Input the complementary masked gene latent representation and the re-masked gene latent representation into the contrastive learning module, output the cosine similarities of the positive and negative sample pairs, and construct a contrastive loss based on the cosine similarities of the positive and negative sample pairs ; Construct a reconstruction loss based on the reconstructed feature representation ; Based on the reconstruction loss function and the contrastive loss function construct the total loss function Train the neural network model to obtain a trained neural network model, where and are hyperparameters that control the proportions of the reconstruction loss and the contrastive loss in the total loss.

3. The spatial transcriptome data clustering method based on progressive learning and multimodal fusion according to claim 2, wherein Input the histological image feature information and the spatial adjacency matrix into the first encoder, and output the latent representation of the image feature information, including the steps: The first encoder includes several layers of graph convolutional layers. Through the several layers of graph convolutional layers, convolutional operations are performed on the input histological image feature information and the spatial adjacency matrix to obtain a latent representation of the image feature information. The convolutional operation expression is: , where is the feature representation of all sample points in the current layer, is the feature representation after all sample points have undergone a single convolutional operation, , , I is the identity matrix, is 's degree matrix, is the activation function, is the trainable parameter matrix of the current layer.

4. The method for clustering spatial transcriptome data based on progressive learning and multimodal fusion according to claim 3, wherein, Input the masked gene expression matrix and the spatial adjacency matrix into the second encoder, and the output masked gene latent representation is input into the third encoder after remasking to obtain a remasked gene latent representation, including the steps: The second encoder includes several layers of graph convolutional layers, and performs convolutional operations on the input masked gene expression matrix and spatial adjacency matrix through the several layers of graph convolutional layers to obtain the masked gene latent representation; Perform remasking on the masked gene latent representation, and the rows masked in the masked gene latent representation are the same as the rows masked in the masked gene table matrix, to obtain the remasked gene latent representation; The third encoder includes several layers of graph convolutional layers, and performs convolutional operations on the input remasked gene latent representation through the several layers of graph convolutional layers to obtain the remasked gene latent representation.

5. The spatial transcriptome data clustering method based on progressive learning and multimodal fusion according to claim 4, wherein Input the latent representation of the re-masked genes into the progressive learning module to output the average latent representation of the re-masked genes, including the steps of: The progressive learning module, according to the increase in the number of training rounds, based on the spatial adjacency matrix randomly selects a certain percentage of neighbor points for each sample point, and averages the re-masked gene latent representations of the neighbor points to obtain an average re-masked gene latent representation, where the selection of the percentage is set according to the following formula: , where is the number of rounds required for the training of the neural network model, is the round to which the neural network model has been trained currently.

6. The spatial transcriptome data clustering method based on progressive learning and multimodal fusion according to claim 5, wherein Input the average latent representation of the re-masked genes and the latent representation of the image feature information into the first cross-attention fusion module to output the first fusion latent representation, including the steps of: The first cross-attention fusion module is used to capture the correlation between different feature spaces, and its calculation formula is as follows: , where Q, K, and V are the query matrix, the key matrix, and the value matrix respectively, is a weight matrix, represents the matrix concatenation operation; Input the average latent representation of the re-masked genes and the latent representation of the image feature information into the first cross-attention fusion module to output the first fusion latent representation, where the latent representation of the image feature information serves as the data source for the query matrix Q, and the average latent representation of the re-masked genes serves as the data sources for the key matrix K and the value matrix V.

7. The spatial transcriptome data clustering method based on progressive learning and multimodal fusion according to claim 6, wherein Input the complementary masked gene latent representation and the re-masked gene latent representation into the contrastive learning module to output the cosine similarities of the positive sample pairs and the negative sample pairs, and construct a contrastive loss based on the cosine similarities of the positive sample pairs and the negative sample pairs , including the steps of: According to the masked sample rows in the latent representation of the re-masked genes, select the corresponding sample rows in the latent representation of the complementary masked genes and the latent representation of the re-masked genes to construct positive sample pairs, and at the same time randomly extract sample rows inconsistent with the positive sample pairs from the latent representation of the complementary masked genes and the latent representation of the re-masked genes as negative sample pairs; According to the formula calculate the cosine similarity of positive and negative sample pairs, where (A, B) is a sample pair, is the modulus symbol; By pulling the cosine similarity of the positive samples closer , push the negative sample to cosine similarity The contrastive loss is constructed in this way.

8. A spatial transcriptome data clustering device based on progressive learning and multimodal fusion, characterized in that, Include: A data acquisition module for acquiring histological images and their corresponding spatial transcriptome data; An image processing module, which is used to cut a histological image according to the size of sample points, and then extract histological image feature information of each cut small block through a medical image large model UNI ; An adjacency matrix construction module for calculating the spatial Euclidean distance between each pair of sample points according to the spatial position information of each sample point, selecting the k nearest neighbor sample points of each sample point as neighbors, and constructing a spatial adjacency matrix A; A gene expression data processing module for screening and regularizing the gene expression data of the sample points in the spatial transcriptome data, performing principal component analysis on the gene expression matrix of the top 3000 highly variable genes, and extracting the top 200 principal component gene expression matrices as gene expression information X; A masking module, configured to randomly mask each row of the gene expression information X according to a predetermined ratio to obtain a masked gene expression matrix , and then perform complementary masking on the gene expression information X to obtain a complementary masked gene expression matrix ; A model construction module for constructing a neural network model, where the neural network model includes a first encoder, a second encoder, a third encoder, a decoder, a first cross-attention fusion module, a second cross-attention fusion module, a progressive learning module, and a contrast learning module; A model training module for using the spatial adjacency matrix A, histological image feature information , masked gene expression matrix and complementary masked gene expression matrix as input data, using reconstruction loss and contrast loss as the total loss function to train the neural network model, and obtaining a trained neural network model; A feature output module, which is used to input the histological image features information obtained after preprocessing the histological image to be analyzed and the corresponding spatial transcriptome data , the spatial adjacency matrix A, and the gene expression information X into the trained neural network model, and output the fused feature latent representation; An aggregation module for performing clustering analysis on the fusion feature latent representation and outputting a clustering result.

9. A spatial transcriptome data clustering device based on progressive learning and multimodal fusion, characterized in that, Include a memory and at least one processor, and computer-readable instructions are stored in the memory; The at least one processor calls the computer-readable instructions in the memory to execute each step of the spatial transcriptome data clustering method according to any one of claims 1-7.

10. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by the processor, each step of the spatial transcriptome data clustering method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Spatial domain identification method in spatial transcriptomics based on deep graph learning

    CN117708628A

  • Global modeling and semantic integration combined generalized small sample remote sensing segmentation algorithm

    CN118537745A

  • Spatial domain identification and batch effect removal method for multi-slice spatial transcriptomics data set

    CN118899041A

  • Spatial transcriptomics cell type deconvolution method based on generative adversarial network and comparative learning framework

    CN119724362A

  • Spatially informed clustering, integration, and deconvolution of spatial transcriptomics with graphst

    WO2024155235A1

Cited By

  • Spatial domain identification method based on artificial intelligence

    CN120852826A

  • Empty transgene expression filling method based on conditional variation auto-encoder

    CN120913643A

  • Spatial transcriptome data recognition method and system based on kan distance perception space graph

    CN122619128A

  • Spatial transcriptome data recognition method and system based on kan distance perception space graph

    CN122619128B