Spatial transcriptomics cell clustering method based on multi-scale contrast learning

By adopting multi-scale contrast learning and ContraNorm layer method in spatial transcriptomics, the problem of discontinuous clustering results in the prior art is solved, and more accurate and continuous cell clustering is achieved.

CN120015132AActive Publication Date: 2025-05-16ANHUI UNIV

Patent Information

Application Number
CN202411845699.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-16
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The existing clustering methods are difficult to effectively utilize spatial information, resulting in discontinuous clustering results in tissue sections and are difficult to adapt to cell or subcellular discrimination.

Method used

Using a multi-scale contrast learning method, cell gene expression information is further learned through data preprocessing, graph construction, data augmentation, GCN extraction of cell gene expression information, ContraNorm layer enrichment node information, and multi-scale graph comparison learning.

Benefits of technology

The problems of high noise and zero expansion in ST data are effectively solved. Through the combination of multi-scale comparison learning and the ContraNorm layer, more complete and accurate cellular gene expression information are extracted, improving the continuity and accuracy of clustering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015132A_ABST
    Figure CN120015132A_ABST
Patent Text Reader

Abstract

The invention discloses a spatial transcriptomics cell clustering method based on multi-scale contrast learning. The method comprises the following steps: S1, carrying out data preprocessing; s2, performing graph construction by using the processed data; s3, performing data enhancement; s4, extracting cell gene expression information by using GCN; s5, enriching node information by using a ContraNorm layer; and S6, further learning information by using multi-scale image comparison learning, and performing biological analysis by using final information obtained by learning. According to the method, computer-aided cell type analysis is utilized, a large amount of high-quality data is not needed, only gene expression data and cell space position information data are utilized, which cells belong to the same category can be predicted, and more effective information is provided to help researchers to identify the onset of cancer and the development of diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cell clustering, and in particular to a spatial transcriptomics cell clustering method based on multi-scale contrast learning. Background Art

[0002] Recent technological advances in spatial transcriptomics (ST) have made it possible to analyze gene expression with location information in tissues. ST refers to transcriptomic technologies that can preserve the spatial information and gene expression profiles of tissue samples. Spatial transcriptomics (ST) has made significant progress in the past few years. According to the different data generation methods, ST technologies can be divided into NGS-based (next-generation sequencing) and image-based methods. NGS-based ST technologies obtain data with spatial resolution by attaching spatial barcodes with fixed positions to tissue sections. Therefore, each spot captured by NGS-based ST datasets usually contains multiple cells. Many NGS-based ST methods have been developed, including 10XGenomics' Visium, NanoString's GeoMx, Slide-Seq, Slide-SeqV2, Stereo-Seq and other methods. These methods obtain RNA transcripts through in situ sequencing or in situ hybridization, and retain the spatial information of cells through images of stained tissue samples. Image-based ST technologies, such as STARMap, merFISH and seqFISH+, can usually achieve single-cell or subcellular resolution. Typically, a ST dataset consists of a gene expression matrix, where each row represents a gene and each column represents a spot or cell, and a spatial position matrix, which records the spatial coordinates of the spot or cell.

[0003] Cluster analysis is an indispensable step in transcriptome data analysis. Clustering is very important for annotating cell types, understanding tissue architecture, identifying co-expressed gene modules, and many downstream analyses such as trajectory inference and intercellular communication. Most existing clustering methods cannot effectively use available spatial information. These non-spatial methods can be roughly divided into two categories. The first category uses traditional clustering methods such as K-means and Louvain algorithms. Due to the different resolutions of spatial transcriptome technologies, these methods are limited to a small number of sposts, and the clustering results may be discontinuous in tissue sections. The second category is to deconvolute spots using cell type features defined by single-cell RNA sequences, but they are not suitable for resolution at the cellular or subcellular level. Summary of the invention

[0004] In order to solve the existing problems, the present invention provides a method for spatial transcriptomics cell clustering based on multi-scale contrastive learning. The specific scheme is as follows:

[0005] A method for spatial transcriptomics cell clustering based on multi-scale contrastive learning, comprising the following steps:

[0006] S1, data preprocessing;

[0007] S2, graph construction using processed data;

[0008] S3, perform data enhancement;

[0009] S4, extracting cell gene expression information using GCN;

[0010] S5, enrich node information using ContraNorm layer;

[0011] S6, use multi-scale graph contrast learning to further learn information, and use the final information learned for biological analysis.

[0012] Preferably, step S1 uses the SCANPY software package to log-transform the raw gene expression information and normalize the library size, and selects the top 3000 highly variable genes as input to the model.

[0013] Preferably, step S2 converts the spatial location information into an undirected graph structure G=(V,A) with a preset number of neighbors k, where is a collection of gene expression information. is the adjacency matrix of N points; if point i and point j are neighbors here, let A ij =1, otherwise 0. Therefore, for a given point, its neighbor nodes are defined by calculating the Euclidean distance of the spatial position information. Finally, the present invention selects the nearest k nodes from the nodes as neighbors of the given node.

[0014] Preferably, two methods can be selected for data enhancement in step S3. One method is to perform data augmentation on a given graph structure G = (V, A), by randomly disrupting the gene expression vectors of all points to construct a damaged graph structure, and keeping the original adjacency matrix unchanged; the second method introduces singular value decomposition SVD for data enhancement, which generates a new graph structure by considering maintaining the global collaborative signal. The formula is as follows:

[0015]

[0016] The input matrix is ​​expressed by a low-rank matrix approximation, and then the SVD is performed on the matrix; where q is the rank required for the decomposition matrix, in the formula It's U q , S q , V qis an approximate expression of . U is an orthogonal matrix, called a left singular matrix. S is a diagonal matrix, and the elements on the diagonal are called singular values. V is an orthogonal matrix, called a right singular matrix.

[0017] Preferably, step S4 learns the gene embedding of the point through an encoder composed of a graph convolutional network, the input source is the adjacency matrix G and the gene expression information X of each spot; the gene expression information and the spatial position information are integrated through GCN. Specifically, the representation formula of the lth layer in the encoder is as follows:

[0018]

[0019] in, represents the normalized adjacency matrix, D represents a diagonal matrix, and the diagonal elements are W e is the trainable weight matrix of the l-1th layer in the graph convolutional network, represents the output of layer l and Set to the original input gene expression information X; σ(·) is the nonlinear activation function, Z s is the final output of the encoder, and z i represents the gene expression vector of the i-th row;

[0020] Then the potential representation Z s Input into the decoder, and finally obtain the reconstructed gene expression information through decoding operation. The specific formula is as follows

[0021]

[0022] in, represents the reconstructed gene expression information of the lth layer, Set to the final output Z of the encoder s , W d is the trainable weight matrix of the l-1th layer in the decoder's graph convolutional network, and the final output H s Represents the reconstruction of gene expression information; in order to make gene expression information play an important role, the reconstruction loss is used to train the model as follows:

[0023]

[0024] where x i represents the original gene expression vector, h i represents the reconstructed gene expression vector, and the entire loss function is to calculate the Euclidean distance between the original expression and the reconstructed expression.

[0025] Preferably, in step S5, in order to use the ContraNorm layer to enrich the information of the node, the model has better performance in identifying the cell domain, and the formula is as follows:

[0026]

[0027] Among them, Z s represents the low-dimensional embedding expression after the encoder processing, and s represents the step size of the gradient descent. It is similar to the self-attention matrix in Transformers, which enriches node information by comparing the similarity between a specified point and other points.

[0028] Preferably, the multi-scale contrastive learning in step S6 is divided into two types; the first type is local-to-local contrast, and the graph structure regenerated by SVD composition is used to rewrite the information transmission rule, that is, the original information passing through the GNN encoder is rewritten, and the rewritten embedding is expressed as follows:

[0029]

[0030] where Z′ b Represents the low-dimensional embedding generated from the newly constructed graph structure view. The extended view is created through global collaboration, which can enhance the representation of the main view. In the local-to-local comparison, the model uses InfoNCE loss, and the contrast loss is defined as:

[0031]

[0032] Among them, z i,l represents the potential embedding of the main view, g i,l represents the latent embedding obtained by SVD enhanced view, s(·) and Represent cosine similarity and temperature respectively, thereby learning the global semantic information brought by SVD;

[0033] The second is the comparison between local and context. The original graph structure and the graph structure after randomly shuffling the gene expression vectors are used as the input of the model. Through the learning of the encoder, two different gene expression information can be obtained, namely Z b and Z″ b , and the average expression of the first-order neighbors of a given point is recorded as g i , which represents the neighbor environment of the point. According to the characteristics of spatial transcriptome data, the gene expression and cell type of a point are more likely to be the same as those of its neighbors. Therefore, a method of comparative learning based on the average expression of first-order neighbors is defined. For a specified point i, the gene expression z of the point is set i and the average expression of the first-order neighbors g i form a positive pair, and the gene expression from this point is z i The average expression g′ of the first-order neighbors of the gene expression vector in the randomly shuffled graph iConstruct negative pairs; learn the feature information of nodes by maximizing the mutual information between positive pairs and minimizing the mutual information between negative pairs; use the binary cross entropy BCE loss function in the comparison between local and context, and the contrast loss is defined as:

[0034]

[0035] where Φ(·) is a discriminator A dual neural network that distinguishes positive and negative pairs; Φ(z i ,g i ) indicates that the pair (z i ,g i ) and define a loss similar to the original topology to improve the stability of the model:

[0036]

[0037] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. After the computer program is run, any of the above methods is executed.

[0038] The present invention also discloses a computer system, including a processor and a storage medium, wherein a computer program is stored on the storage medium, and the processor reads and runs the computer program from the storage medium to execute any of the methods described above.

[0039] The beneficial effects of the present invention are:

[0040] (1) In order to solve the problems of high noise and zero expansion in ST data, the present invention introduces singular value decomposition (SVD) in contrastive learning for data enhancement, which generates a new graph structure by considering maintaining global collaborative signals. At the same time, the traditional data enhancement method is retained, and important information can be more comprehensively retained through two different data enhancement methods.

[0041] (2) The graph neural network is combined with sub-supervised contrastive learning to extract the potential embedded gene expression information, and a multi-scale contrast is applied in contrastive learning, namely point-to-point contrast at the same scale and point-to-surround contrast across scales. This makes contrastive learning more comprehensive and can learn more complete information.

[0042] (3) In order to alleviate the over-smoothing problem of graph convolutional neural networks, the present invention introduces a ContraNrom layer to enrich the gene expression information between nodes and provide better gene expression for subsequent work. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0044] Figure 1 The workflow of MCST;

[0045] Figure 2 A is the manually annotated layer structure of 151674 slices in the DLPFC dataset;

[0046] Figure 2 B is a box plot of the adjusted Rand index (ARI) values ​​of 6 different methods on 12 DLPFC slices;

[0047] Figure 2 C is a box plot of the normalized mutual information (NMI) values ​​of 6 different methods on 12 DLPFC slices;

[0048] Figure 2 D is the identification of the spatial domain of 151674 slices by MCST, GraphST, STAGATE, SpaceFlow, SpaGCN and SCANPY methods;

[0049] Figure 3 A is the UMAP visualization and PAGA graph generated by MCST, GraphST, STAGATE, SpaceFlow, SpaGCN and SCANPY methods on 151674 slices;

[0050] Figure 3 B is a box plot of the Moran I values ​​of SVG detected by three different methods;

[0051] Figure 3 C is the spatial expression distribution map of SVG detected by MCST in 151674 slices;

[0052] Figure 4 A HBRC dataset was manually labeled according to pathological features;

[0053] Figure 4 B is the ARI and NMI bar graph of MCST and other comparison methods;

[0054] Figure 4 C is the identification of structural domains on the HBRC dataset using MCST, GraphST, STAGATE, SpaGCN, SpaceFlow, and Seurat methods;

[0055] Figure 4 D is the box plot of Moran's I values ​​of SVG detected by three different methods, and the spatial expression patterns of SVGs detected by MCST;

[0056] Figure 5 A is the MBA dataset manually labeled according to pathological features;

[0057] Figure 5 B is the identification of structural domains on the MBA dataset using MCST, GraphST, STAGATE, SpaGCN, SpaceFlow and Seurat methods; Figure 5 C is the ARI and NMI bar graph of MCST and other comparison methods; Figure 6 A is the comparison results of ARI and NMI of MCST with W / O SVD, W / O context, and W / O ContraNorm on the DLPFC dataset; Figure 6 B shows the comparison results of MCST with W / OSCD, W / O context, and W / O ContraNorm in terms of ARI and NMI on the HBC and MBA datasets. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] The present invention uses the KNN algorithm to construct a spatial neighbor network, then uses the GCN algorithm to obtain cell gene expression information, uses multi-scale graph contrast learning to further learn the characteristic representation of cell genes, and finally performs biological analysis through the final embedded information obtained. Figure 1 The overall workflow is shown.

[0060] The present invention utilizes computer-assisted cell type analysis, which does not require a large amount of high-quality data. It only requires gene expression data and cell spatial location information data to predict which cells belong to the same category, providing more effective information to help researchers identify the onset of cancer and the development of the disease.

[0061] A method for spatial transcriptomics cell clustering based on multi-scale contrastive learning, comprising the following steps:

[0062] S1, data preprocessing; for spatial clustering, the present invention uses gene expression information and spatial location information. The original gene expression information is logarithmically transformed and the library size is normalized using the SCANPY software package. Finally, the first 3000 highly variable genes are selected as the input of the model.

[0063] Specifically, the present invention uses three public ST data sets as benchmark data sets to verify the performance of the present invention. The first data set is the LIBD human dorsolateral prefrontal cortex (DLPFC) data set, obtained by 10x Visium. The number of spots in each slice ranges from 3460 to 4789, and Maynard et al. manually annotated it as the DLPFC layer and white matter (WM). Each slice has five or seven manually annotated areas. The second data set is the 10xVisium data set of human breast cancer. Fu et al. annotated it into 20 areas and four main morphological types: ductal carcinoma in situ / lobular carcinoma in situ (DCIS / LCIS), invasive ductal carcinoma (IDC), healthy area and low malignant tumor margin. The third data set is mouse brain tissue. This data set has two parts, and the present invention selects the front part. The selected part contains 2695 points, in which 32285 genes are captured, and 52 manually annotated areas of the Allen brain atlas are used for reference.

[0064] S2, using the processed data to construct a graph. The advantage of spatial transcriptomics compared to other data is that it has the spatial location information of points. In order to combine the information of similar points, the present invention converts the spatial information into an undirected graph structure G = (V, A) with a preset number of neighbors k, where is a collection of gene expression information. is the adjacency matrix of N points. If point i and point j are neighbors here, let A ij =1, otherwise 0. Therefore, for a given point, its neighbor nodes are defined by calculating the Euclidean distance of spatial location information. Finally, the nearest k nodes are selected from the nodes as neighbors of the given node.

[0065] S3, perform data enhancement.

[0066] The present invention uses two different data augmentation methods.

[0067] The first method is to perform data augmentation on a given graph structure G = (V, A) by randomly shuffling the gene expression vectors of all points to construct a damaged graph structure while keeping the original adjacency matrix unchanged.

[0068] The second method introduces singular value decomposition (SVD) for data enhancement, which generates a new graph structure by considering maintaining global collaborative signals. Since using SVD on large matrices will occupy a relatively large amount of memory and consume computing resources, the random SVD algorithm is used here, and the formula is as follows: U is an orthogonal matrix, called a left singular matrix. S is a diagonal matrix, and the elements on the diagonal are called singular values. V is an orthogonal matrix, called a right singular matrix.

[0069] It approximates the input matrix through a low-rank matrix, and then performs SVD on this smaller matrix. Where q is the rank required for the decomposition matrix. It's U q , S q , V q An approximate expression of .

[0070] S4, using GCN to extract cell gene expression information, that is, using GCN to obtain cell feature vectors for analysis.

[0071] Specifically, the present invention designs an encoder composed of a graph convolutional network to learn the gene embedding of points, where the input source is the adjacency matrix G and the gene expression information X of each spot. GCN is a powerful graph neural network that can directly process graph-structured data and utilize graph structure information. It can aggregate the information of neighbor nodes through iteration and capture the dependencies between nodes to generate low-dimensional embedding information. Therefore, the present invention integrates gene expression information and spatial location information through GCN. Specifically, the representation of the lth layer in the encoder is shown in the following formula:

[0072]

[0073] in, represents the normalized adjacency matrix, where D represents a diagonal matrix with diagonal elements W e is the trainable weight matrix of the l-1th layer in the graph convolutional network, represents the output of the lth layer and Set to the original input gene expression information X. σ(·) is a nonlinear activation function, such as RELU (Rectified Linear Unit). So Z is finally s is defined as the final output of the encoder, and z i Represents the gene expression vector of the i-th row.

[0074] Then the potential representation Z sInput into the decoder, and finally obtain the reconstructed gene expression information through decoding operation. Specifically, it is shown in the following formula:

[0075]

[0076] in, represents the reconstructed gene expression information of the lth layer, Set to the final output Z of the encoder s .W d is the trainable weight matrix of the l-1th layer in the decoder's graph convolutional network. The final output H s represents our reconstructed gene expression information. In order to make the gene expression information play an important role, we use reconstruction loss to train the model as follows:

[0077]

[0078] where x i represents the original gene expression vector, h i represents the reconstructed gene expression vector. The entire loss function is to calculate the Euclidean distance between the original expression and the reconstructed expression.

[0079] S5, enrich node information using ContraNorm layer.

[0080] Specifically, after the gene expression features are processed by the graph convolutional network, the expression vectors between each point may gradually become similar. In order to make the model have better performance in identifying cell domains, the present invention introduces a ContraNorm layer to enrich the node information and solve the above problem.

[0081]

[0082] Among them, Z s represents the low-dimensional embedding expression after being processed by the encoder, and s represents the step size of the gradient descent. It is similar to the self-attention matrix in Transformers, so it mainly enriches node information by comparing the similarity between the specified point and other points.

[0083] S6, use multi-scale graph contrast learning to further learn information, and use the final information learned for biological analysis.

[0084] Specifically, the present invention uses a self-supervised multi-scale contrastive learning strategy, which can more accurately capture the local surrounding information and global information of the spots. Specifically, multi-scale contrastive learning is mainly divided into two types. The first is local-to-local contrast, which mainly rewrites the information transmission rules through the graph structure regenerated by the above-mentioned SVD composition, that is, rewrites the original information passed through the GNN encoder. The specific rewritten embedding representation is as follows:

[0085]

[0086] where Z′ b represents a low-dimensional embedding generated from a newly constructed graph structure view. The present invention does not require the calculation of a large dense matrix Because the entire dense matrix takes up a lot of space to store. Therefore, the model can store low-dimensional To calculate and Thereby improving the efficiency of the model. In the present invention, the extended view is created through a global collaborative relationship, which can enhance the representation of the main view. In the local-to-local comparison, the model uses the InfoNCE loss. The contrast loss can be defined as:

[0087]

[0088] Among them, z i,l represents the potential embedding of the main view, g i,l represents the latent embedding obtained by SVD enhanced view. s(·) and Represent cosine similarity and temperature respectively. Through this comparison, we can learn the global semantic information brought by SVD.

[0089] The second is the comparison between local and context. The present invention uses the original graph structure and the graph structure after randomly shuffling the gene expression vectors as the input of the model. Through the learning of the encoder, two different gene expression information can be obtained, namely Z b and Z″ b . And the average expression of the first-order neighbors of a given point is denoted as g i , which represents the neighbor environment of the point. It can be concluded from the characteristics of spatial transcriptome data that the gene expression and cell type of a point are more likely to be the same as those of its neighbors. Therefore, the present invention defines a method for comparative learning based on the average expression of first-order neighbors. For a specified point i, set the gene expression z of the point i and the average expression of the first-order neighbors g i form a positive pair, and the gene expression from this point is z i The average expression g′ of the first-order neighbors of the gene expression vector in the randomly shuffled graphi Construct negative pairs. By maximizing the mutual information between positive pairs and minimizing the mutual information between negative pairs, the feature information of the node is learned. In the comparison between local and context, the binary cross entropy (BCE) loss function is used. The contrast loss can be defined as:

[0090]

[0091] where Φ(·) is a discriminator It distinguishes between positive and negative pairs of dual neural networks. Φ(z i ,g i ) indicates that the pair (z i ,g i ). Since the gene vectors are randomly shuffled and the adjacency matrix is ​​kept unchanged during data augmentation, the original graph structure and the damaged graph structure have the same topological structure. Therefore, the present invention defines a loss similar to the original topology to improve the stability of the model:

[0092]

[0093] The present invention uses reconstruction loss and contrast loss as the total loss function.

[0094] Specifically, the loss function of the model mainly includes two parts, one is the reconstruction loss of the encoder, and the other is the contrast loss of multi-scale contrastive learning. The model of the present invention is trained by minimizing the loss functions of these two parts. In short, the total loss function is as follows:

[0095]

[0096] Among them, λ1, λ2 and λ3 are weight factors that weigh the impact of reconstruction loss and contrast loss. According to the experimental results, these three parameters are set to 10, 1 and 0.5 by default.

[0097] In addition, the cell clustering performance of the present invention can be evaluated using the following formula:

[0098]

[0099] Among them, a represents the number of point pairs that belong to the same cluster in both the real and experimental cases, b represents the number of point pairs that belong to the same cluster in the real case but not in the experimental case, c represents the number of point pairs that do not belong to the same cluster in the real case but belong to the same cluster in the experimental case, and d represents the number of point pairs that do not belong to the same cluster in both the real and experimental cases. The value range of ARI is [-1, 1]. The larger the value, the more consistent it is with the real result, that is, the better the clustering effect.

[0100]

[0101] Where X is the experimental result and Y is the actual result.

[0102] About the experimental results of the present invention:

[0103] The present invention uses 151674 slices in the DLPFC dataset to illustrate the clustering effect of the algorithm. Figure 2 (A) The manually annotated layer structure of 151,674 slices in the DLPFC dataset. Figure 2 As shown in (D), it can be clearly seen in the clustering effect diagram that Seurat only recognizes WM and Layer_1, and the points of the remaining layers are mixed together without any distinguishing effect, and there is no obvious boundary between the layers. It can be seen from the result diagram that SpaGCN divides the WM domain into two domains. Although it has been improved in terms of the point mixing of the remaining layers, there are still problems, and the remaining layers are severely modularized, which is inconsistent with the actual cell domain. For SpaceFlow, its effect is slightly improved over the previous two methods, and there will be no mixing between different cell domains. However, SpaceFlow has the same problem as SpaGCN, and the remaining layers are severely modularized, which is inconsistent with the actual cell domain. STAGATE and GraphST have made obvious improvements, and the distribution of cell domains is closer to the manually annotated tissue layers, but they both divide Layer6 into two regions, which is a wrong division. MCST does not have the same error, and clearly divides the layers, which is closer to the manually annotated tissue layers. When we focus on the performance of all slices of the entire DLPFC, this method is superior to other methods. In the results shown, MCST achieved the highest values ​​in both ARI and NMI. As shown Figure 2 As shown in (B), the median ARI of MCST reached 0.62, while the medians of the other five comparison methods were all below 0.6. Figure 2 As shown in (C), compared with other methods, MCST achieves excellent results in NMI indicator. The median NMI of MCST exceeds 0.7, while other methods fail to reach this level. This shows that MCST performs better than other comparison methods in clustering tasks.

[0104] The generated embeddings are visualized using UMAP, and the spatial trajectories are presented on the UMAP graph. The 151674 slices are used as the display result, as shown in the figure below: Figure 3 (A) As shown. For the non-spatial method Seurat, Figure 3As shown in (A), except for the WM layer, the spots of the remaining cortical layers are concentrated together and lack obvious separation. Neither SpaceFlow nor SpaGCN can clearly distinguish the spots of Layers_1 and Layer_3. GraphST has improved the results of the above models, but the effect of distinguishing Layers_1 and Layer_3 is still not good. STAGATE has suboptimal UMAP visualization results, but Layers_1 and Layer_3 are still a little mixed. In contrast, MCST can clearly distinguish each tissue layer and accurately reflect the developmental order of the cortex. The above conclusions can also be proved in the trajectory diagram. The PAGA diagram was obtained through analysis, in which the PAGA diagram shows an almost linear development trajectory from layer 1 to layer 6 and the similarity between adjacent layers.

[0105] In order to improve the credibility of the identified spatial domains, the present invention will use the same process as SpaGCN to detect SVGs and compare the results with the SVGs determined by SpaGCN. A total of 111 SVGs were detected on 151674 slices, which were scattered in different domains. These SVGs include 18 SVGs in domain 1, 35 SVGs in domain 3, 50 SVGs in domain 4, 1 SVG in domain 5, and 6 SVGs in domain 6. Figure 3 As shown in (B), in contrast, SpaGCN detected 72 SVGs with lower Moran’s I (which is used to evaluate the degree of spatial aggregation of a variable) values ​​than the obtained values, i.e., the median Moran’s I value of SpaGCN was 0.39, while the median Moran’s I value of our model was 0.51. SpatialDE

[36] detected 3221 SVGs, however, the median Moran’s value of SVGs obtained by SpatialDE was 0.28, which is much lower than that of our method. Figure 3 (C) The spatial expression distribution of SVG detected by MCST in 151,674 slices.

[0106] Next, the present invention analyzed the human breast cancer dataset from the 10x Visium platform and demonstrated the performance of MCST by comparing the results of other methods. Figure 4 (A) The HBRC dataset is manually labeled according to pathological features. For HBRC, there are four main morphological types: ductal carcinoma in situ / lobular carcinoma in situ (DCIS / LCIS), healthy tissue (healthy), invasive ductal carcinoma (IDC), and low-malignant tumor surrounding areas (tumor margins). Figure 4As shown in (C), the results show that the spatial domains identified by MCST are more consistent with the results of manual annotation, and can correctly identify many domains such as Healthy_1 and IDC_4, and each domain presents a smoother segmentation. The results of STAGATE in identifying the edges of spatial domains show that the spatial domains are mixed with each other, while the results of SpaceFlow in identifying the edges of spatial domains are better than STAGATE, but it does not identify the spatial domains more accurately. Seurat identifies more distinguishable structural domains, but there are still outliers and rough boundaries between the identified structural domains. SpaGCN and GraphST obtained suboptimal results and are still insufficient in correctly identifying the categories of spatial domains. Compared with these methods, MCST can better identify the Healthy_1, IDC_4, and DCIS / LCIS_4 domains.

[0107] In addition, if Figure 4 As shown in (B), MCST obtained the highest ARI (0.61) and NMI (0.69). It can be clearly seen from the result graph that the ARI of all other methods is lower than 0.6, and MCST also obtained the highest NMI. For the identified spatial domain, the difference between the method of the present invention and SpaGCN and SpatialDE in detecting information SVG was further compared. MCST detected 439 SVGs, SpaGCN and SpatialDE identified 358 SVGs and 8569 SVGs, respectively, but the method of the present invention obtained the highest median Moran's I value among the three methods. In addition, the detected SVGs showed obvious spatial expression patterns, such as Figure 4 As shown in (D), genes identified in a given domain also clearly exhibited spatial expression patterns.

[0108] Finally, the present invention applies MCST to the mouse brain anterior dataset. Figure 5 As shown in (A), the manual labeling of pathological features shows that this dataset has 52 clusters. The dataset is processed by the model and spatially clustered. From the clustering results, it can be seen that Figure 5As shown in (B), Seuart obtained the worst result. It not only showed a fuzzy display at the boundary of the recognition domain, but also had a large difference between the recognition of the spatial domain and the real label. For example, it almost did not classify the CPa area, and there were many areas mixed together. STAGATE and SpaceFlow have improved their feasibility. They showed a clearer display at the boundary of the recognition domain, but there were many misjudgments in the recognition of the spatial domain. For example, the CPa area is the entire area, but they are divided into many small areas. SpaGCN and GraphST obtained suboptimal clustering results, but they all had the same problem. In the upper right corner of the slice, they divided this area into long strips, but in the manually labeled images, they have a dividing line in the middle and are not long strips. In the recognition area of ​​MCST, the dividing line can be observed, and the CPa area is accurately distinguished. It also shows the best results in distinguishing other domains.

[0109] In addition, if Figure 5 As shown in (C), MCST achieved the highest ARI (0.51) and NMI (0.72). It can be clearly seen from the result graph that the ARI of all other methods is lower than 0.5, and MCST also achieved the highest NMI.

[0110] In order to demonstrate the effectiveness of multi-scale contrastive learning and ContraNorm in the model, the present invention conducted ablation experiments. In these experiments, we systematically deleted the local-to-local contrast, local-to-surrounding contrast, and ContraNorm layers. The designed ablation experiments were conducted on the DLPFC dataset, HBC dataset, and MBA dataset. Figure 6 As shown in (A) and (B), the ARI and NMI values ​​obtained by MCST-ablating SVD, MCST-ablating context, and MCST-ablating ContraNorm are much lower than those of MCST. The experimental results clearly verify the advantages of the multi-scale learning strategy and ContraNorm layer in MCST.

[0111] In order to solve the problems of high noise and zero expansion in ST data, the present invention introduces singular value decomposition (SVD) in contrastive learning for data enhancement, which generates a new graph structure by considering maintaining global collaborative signals. At the same time, the traditional data enhancement method is retained, and important information can be more comprehensively retained through two different data enhancement methods. At the same time, the present invention combines graph neural networks with sub-supervised contrastive learning to extract potential embedded gene expression information, and applies a multi-scale contrast in contrastive learning, which is a point-to-point comparison of the same scale and a cross-scale point-to-surrounding comparison. The perspective of contrastive learning is more comprehensive, and more complete information can be learned. In addition, in order to alleviate the problem of over-smoothing of graph convolutional neural networks, the present invention introduces a ContraNrom layer to enrich the gene expression information between nodes, providing better gene expression for subsequent work.

[0112] The present invention also discloses a computer-readable storage medium and a computer system, wherein a computer-readable storage medium stores a computer program, and after the computer program is run, the method described in any one of the above is executed. A computer system includes a processor and a storage medium, wherein the storage medium stores a computer program, and the processor reads and runs the computer program from the storage medium to execute any one of the above methods.

[0113] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed in the present invention can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, various illustrative components, boxes, modules, circuits, and steps are generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. The technician can implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as resulting in a departure from the scope of the present invention.

[0114] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or executed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in cooperation with a DSP core, or any other such configuration.

[0115] The steps of the method or algorithm described in conjunction with the embodiments disclosed in the present invention may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read and write information from / to the storage medium. In an alternative, a storage medium may be integrated into a processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and the storage medium may reside in a user terminal as discrete components.

[0116] In one or more exemplary embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented as a computer program product in software, each function may be stored on or transmitted by a computer-readable medium as one or more instructions or codes. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. Storage media may be any available medium that can be accessed by a computer. As an example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, a server, or other remote source using a coaxial cable, a fiber optic cable, a twisted pair, a digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of the medium. Disks and discs as used in the present invention include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, wherein disks often reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0117] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the universal principles defined in the disclosure may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described in the disclosure, but should be granted the widest scope consistent with the principles and novel features disclosed in the disclosure.

[0118] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for spatial transcriptomics cell clustering based on multi-scale contrastive learning, characterized in that: The following steps are involved: S1, data preprocessing; S2, graph construction using processed data; S3, perform data enhancement; S4, extracting cell gene expression information using GCN; S5, enrich node information using ContraNorm layer; S6, use multi-scale graph contrast learning to further learn information, and use the final information learned for biological analysis.

2. The method according to claim 1, characterized in that: Step S1 uses the SCANPY software package to log-transform the raw gene expression information and normalize the library size, and selects the top 3000 highly variable genes as the input of the model.

3. The method according to claim 1, characterized in that: Step S2 converts the spatial location information into an undirected graph structure G = (V, A) with a preset number of neighbors k, where is a collection of gene expression information. is the adjacency matrix of N points; If point i and point j are neighbors here, that is, let A ij =1, otherwise 0; select the nearest k nodes from the node as the neighbors of the given node.

4. The method according to claim 1, characterized in that: There are two ways to enhance the data in step S3. One is to augment a given graph structure G = (V, A) by randomly disrupting the gene expression vectors of all points to construct a damaged graph structure, while keeping the original adjacency matrix unchanged. The second method introduces singular value decomposition (SVD) for data enhancement, which generates a new graph structure by considering maintaining the global collaborative signal. The formula is as follows: Approximate the input matrix through a low-rank matrix, and then perform SVD on this matrix; Where q is the rank required to decompose the matrix. It's U q , S q , V q An approximate expression of , U is an orthogonal matrix, called a left singular matrix. S is a diagonal matrix, the elements on the diagonal are called singular values, and V is an orthogonal matrix, called a right singular matrix.

5. The method according to claim 1, characterized in that: Step S4 learns the gene embedding of the point through the encoder composed of graph convolutional network. The input source is the adjacency matrix G and the gene expression information X of each spot. The gene expression information and spatial position information are integrated through GCN. Specifically, the representation formula of the lth layer in the encoder is as follows in, represents the normalized adjacency matrix, D represents a diagonal matrix, and the diagonal elements are W e is the trainable weight matrix of the l-1th layer in the graph convolutional network, represents the output of layer l and Set to the original input gene expression information X; σ(·) is the nonlinear activation function, Z s is the final output of the encoder, and z i represents the gene expression vector of the i-th row; Then the potential representation Z s Input into the decoder, and finally obtain the reconstructed gene expression information through decoding operation. The specific formula is as follows in, represents the reconstructed gene expression information of the lth layer, Set to the final output Z of the encoder s , W d is the trainable weight matrix of the l-1th layer in the decoder's graph convolutional network, and the final output H s Represents the reconstruction of gene expression information; in order to make gene expression information play an important role, the reconstruction loss is used to train the model as follows: where x i represents the original gene expression vector, h i represents the reconstructed gene expression vector, and the entire loss function is to calculate the Euclidean distance between the original expression and the reconstructed expression.

6. The method according to claim 1, characterized in that: Step S5: In order to use the ContraNorm layer to enrich the node information, the model has better performance in identifying cell domains. The formula is as follows: Among them, Z s represents the low-dimensional embedding expression after the encoder processing, and s represents the step size of the gradient descent. It is similar to the self-attention matrix in Transformers, which enriches node information by comparing the similarity between a specified point and other points.

7. The method according to claim 4, characterized in that: There are two types of multi-scale contrastive learning in step S6. The first is local-to-local contrast, which rewrites the information transfer rules through the graph structure regenerated by SVD composition, that is, rewrites the original information passing through the GNN encoder. The rewritten embedding is expressed as follows: where Z′ b Represents the low-dimensional embedding generated from the newly constructed graph structure view. The extended view is created through global collaboration, which can enhance the representation of the main view. In the local-to-local comparison, the model uses InfoNCE loss, and the contrast loss is defined as: Among them, z i,l represents the potential embedding of the main view, g i,l represents the latent embedding obtained by SVD enhanced view, s(·) and Represent cosine similarity and temperature respectively, thereby learning the global semantic information brought by SVD; The second is the comparison between local and context. The original graph structure and the graph structure after randomly shuffling the gene expression vectors are used as the input of the model. Through the learning of the encoder, two different gene expression information can be obtained, namely Z b and Z″ b , and the average expression of the first-order neighbors of a given point is recorded as g i , which represents the neighbor environment of the point. According to the characteristics of spatial transcriptome data, the gene expression and cell type of a point are more likely to be the same as those of its neighbors. Therefore, a method of comparative learning based on the average expression of first-order neighbors is defined. For a specified point i, the gene expression z of the point is set i and the average expression of the first-order neighbors g i form a positive pair, and the gene expression from this point is z i The average expression g′ of the first-order neighbors of the gene expression vector in the randomly shuffled graph i Construct negative pairs; learn the feature information of nodes by maximizing the mutual information between positive pairs and minimizing the mutual information between negative pairs; use the binary cross entropy BCE loss function in the comparison between local and context, and the contrast loss is defined as: where Φ(·) is a discriminator A dual neural network that distinguishes positive and negative pairs; Φ(z i ,g i ) indicates that the pair (z i ,g i ) and define a loss similar to the original topology to improve the stability of the model:

8. A computer-readable storage medium, characterized in that: A computer program is stored on the medium, and after the computer program is run, the method according to any one of claims 1 to 7 is executed.

9. A computer system, characterized in that: The method comprises a processor and a storage medium, wherein a computer program is stored in the storage medium, and the processor reads and runs the computer program from the storage medium to execute the method as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Spatial domain identification method based on spatial transcriptomics data feature extraction

    CN116189785A

  • Spatial transcriptome data clustering method

    CN117253550A

  • Space transcriptome data processing method and system based on hypergraph

    CN117457081A

  • MiRNA and disease association prediction method and device, electronic equipment and storage medium

    CN118314948A

  • Missing value interpolation method for single cell sequencing data

    CN118335191A

Cited By

  • Multi-modal spatial domain identification method based on clustering-guided gradient contrast learning

    CN120690281A

  • A multi-modal spatial domain identification method based on clustering guided gradient contrast learning

    CN120690281B

  • Spatial domain identification method based on data interpolation and cell type deconvolution

    CN120748500A