A method for analyzing the substructure of biological tissues in spatial transcriptomics integrating single-cell transcriptomics
The spatial transcriptome data were encoded and clustered through the STAGATE framework and the Louvain algorithm, and the single-cell RNA sequencing data was classified in combination with the XGBoost classification model. Finally, through the hypergraph segmentation integration results, the problem of difficulty in combining single-cell and spatial transcriptome data in the existing technology was solved, significantly improving the clustering accuracy and classification effect.
Patent Information
- Application Number
- CN202210944249.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-05
AI Technical Summary
When existing spatial transcriptome technology analyzes the substructure of biological tissues, it is difficult to effectively bind single-cell transcriptome data, resulting in poor clustering accuracy and single-cell data classification.
The STAGATE framework is used to encode the spatial transcriptome data, and cluster it with the Louvain algorithm. At the same time, the XGBoost classification model is used to classify single-cell RNA sequencing data. Finally, the two results are integrated through hypergraph segmentation to improve clustering accuracy and classification effect.
The clustering accuracy of spatial transcription data and the classification effect of single-cell data were significantly improved, and more accurate biological tissue substructure analysis results were obtained through integrated analysis.
Smart Images

Figure CN115359845B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics technology, and more particularly, to a method, system and computer-readable storage medium for analyzing the substructure of a spatial transcriptome biological tissue by integrating a single-cell transcriptome. Background Art
[0002] With the rapid development of bioinformatics technology, especially the research on transcriptomics and genetics has changed people's understanding of cancer. The progress of single-cell RNA sequencing (scRNA-seq) technology enables researchers to better understand the internal structure of the cell composition of tumors. By using scRNA-seq technology to study and analyze tumor-related cells, and dividing the cell types into more refined cell subpopulations according to the molecular profiles of the cells. In scRNA-seq technology, clustering analysis technology is extremely crucial. The existing gene expression-based methods mainly use indicators such as pearson correlation coefficient and spearman correlation coefficient for analysis. The cell subpopulations form a complex ecosystem, and their interactions will affect tumor progression and treatment outcomes, but the way of interaction between tumor-related cell subpopulations has not been thoroughly studied. The defect of scRNA-seq is that it loses the tissue spatial background (i.e., cell environment) when processing tissue samples, while spatial transcriptome sequencing can simultaneously obtain the spatial position information and gene expression data of cells, which is more suitable for studying cell interactions and spatial gene expression in tumor stroma.
[0003] Currently, there are mainly two types of spatial transcriptome technologies: methods based on NGS technology and imaging-based methods (including ISS-based and ISH-based).
[0004] NGS-based methods: In 2016, spatial transcriptomics (ST) technology was proposed to obtain spatially resolved whole-transcriptome information. By the end of 2018, ST technology was further developed into 10x Visium. The 10x Visium assay has improvements in both resolution and running time. Slide-seq uses randomly placed barcode (a type of encoding for differentiation) beads on a glass slide to capture mRNA. Shortly after the publication of the Slide-seq method, another technology using smaller barcode beads emerged - high-definition spatial transcriptomics (HDST). DBiT-seq can perform spatial group sequencing using deterministic barcodes in tissues. This method is based on a microfluidics approach to deliver barcodes to the surface of tissue slides to achieve a resolution of 10 μm pixel size. Stereo-seq uses randomly placed barcode DNA nanospheres deposited in an array pattern to achieve nanoscale resolution. Seq-scope has achieved subcellular-resolution spatial barcodes that can be used to visualize nuclear and cytoplasmic transcription. The Nanostring GeoMX DSP technology captures data in individual circular regions of interest (ROIs). It irradiates the ROIs with ultraviolet light to release photocleavable gene tags for sequencing quantification. In all NGS-based methods, spatial barcode RNAs are collected and sequenced, where the basic unit of sequencing data is reads (short sequencing fragments). The barcode of each sequencing short fragment (reads) is used to map the spatial location, while the rest of the sequencing reads are mapped to the genome to identify the transcription source, jointly generating a gene expression matrix.
[0005] ISH (in situ hybridization)- and ISS (in situ sequencing)-based methods:
[0006] Both of the above two types of methods generate gene expression matrices through image processing. ISH-based methods are based on ISH technology and detect target sequences through complementary fluorescent probe hybridization. smFISH uses multiple short oligonucleotide probes to target different regions of the same mRNA transcript. Although smFISH has high sensitivity and subcellular spatial resolution, due to the inherent limitation of spectral overlap in standard microscopes, it can only target a few genes at a time. seqFISH is a multiplexed smFISH method that detects a single transcript multiple times through successive rounds of hybridization, imaging, and probe stripping, but it is both expensive and time-consuming. To make up for the substantial time consumption of seqFISH, the MERFISH technology was published in 2015. This technology can identify the copy number and spatial localization of thousands of RNAs in individual cells. It uses techniques such as combinatorial tagging and sequential imaging to increase the detection throughput and uses binary barcodes to counteract single-molecule labeling and detection errors.
[0007] The ISS-based method directly reads the sequences of transcripts in tissues. BaristaSeq is a nicking-fill padlock-based method with its read length increased to 15 bases. STARmap uses barcode padlock probes, hybridizes with targets, and by adding a second primer targeting the site next to the padlock probe, avoids the reverse transcription (RT) step. This method avoids the efficiency barrier of cDNA conversion and reduces noise by adding a second hybridization step. The methods mentioned above are all based on prior knowledge of the targets, while FISSEQ is a non-targeted method that captures all kinds of RNAs, but non-targeted amplification leads to optical crowding and reduced sensitivity.
[0008] To improve the accuracy of spatial data, in the absence of breakthroughs in spatial transcriptomics technology, integrating multi-level and multi-dimensional data is a feasible approach. The computational integration of two or more data modalities can better characterize the spatial cell type composition and local cell states in tissues. For example, integrating scRNA-seq data with spatial transcriptome data for clustering analysis can obtain more accurate classification results. Summary of the Invention
[0009] The present invention provides a method, system, and computer-readable storage medium for analyzing the substructure of a spatial transcriptome of a biological tissue by integrating a single-cell transcriptome, which improves the clustering accuracy of spatial transcript data and the classification effect of single-cell data.
[0010] The primary object of the present invention is to solve the above technical problems, and the technical solution of the present invention is as follows:
[0011] The first aspect of the present invention provides a method for analyzing the substructure of a spatial transcriptome of a biological tissue by integrating a single-cell transcriptome, including the following steps:
[0012] S1. Obtain publicly available spatial transcriptome data and perform preprocessing;
[0013] S2. Encode the preprocessed spatial transcriptome data using the STAGATE framework, and use the Louvain algorithm to cluster the encoded results to obtain the clustering results of the spatial transcriptome data;
[0014] S3. Obtain publicly available single-cell RNA sequencing data and perform preprocessing, and divide the preprocessed single-cell RNA sequencing data into a training data set and a test data set;
[0015] S4. Use the training data set and the test data set to train the XGBoost classification model to classify the single-cell sequencing data set homologous to the spatial transcriptome data to obtain the single-cell classification results;
[0016] S5. Integrate the clustering results of spatial transcriptome data and the single-cell classification results using hypergraph partitioning.
[0017] Furthermore, the preprocessing of the publicly available spatial transcriptome data in step S1 includes: data normalization and data format adjustment.
[0018] Furthermore, the STAGATE framework includes: a spatial neighbor network SNN and a graph attention autoencoder. Among them, the spatial neighbor network is used for... The graph attention encoder is used to learn a low-dimensional latent vector embedding with spatial information and gene expression.
[0019] Furthermore, the specific process of constructing the spatial neighbor network SNN is as follows:
[0020] Convert the spatial information into an undirected neighbor network according to a predefined radius r. Define A as the adjacency matrix of SNN. When and only when the Euclidean distance between node i and node j is less than r, A ij = 1, A ij represents the element in the i-th row and j-th column of the adjacency matrix A; for spatial transcriptome data of other different technologies, r is selected according to the specific resolution of the data. With each node as the center and r as the radius, an average of 6 - 15 neighbor nodes are included; finally, a self-loop is added to each node.
[0021] Furthermore, the graph attention autoencoder includes: an encoder, a decoder, and a graph attention layer. The graph attention layer is embedded in the encoder and the decoder;
[0022] Among them, the encoder takes the normalized gene expression of the node as input and generates a node vector spotembedding by aggregating the information of the neighbors of the node. The graph attention layer in the encoder has a total of L - 1 layers (k ∈ {1, 2,..., L - 1});
[0023] x i is the normalized expression of node i, L is the number of layers of the encoder, is the node vector spotembedding output by the k-th layer of the encoder, S i is the set of neighbors of node s, W k is a trainable weight matrix;
[0024] Taking the expression profile of the node as the initial node vector spotembedding, then there is:
[0025]
[0026] where is the edge weight between node i and node j in the output of the k-th graph attention layer;
[0027] The edge weight from node i to its neighbor node j where and are trainable weight vectors, and Sigmoid represents the sigmoid activation function;
[0028] To make the spatial similarity weights comparable, they are normalized by the softmax function: That is, the edge weight between node i and node j in the output of the k-th graph attention layer;
[0029] The L-th layer of the encoder does not adopt the attention mechanism, and the output is That is, the final output node vector spotembedding;
[0030] The decoder reconstructs the embedding of node i at the (k - 1)-th layer at the penultimate k-th layer: The output of node i at the last layer of the decoder
[0031] where
[0032] The loss function is
[0033] Furthermore, when training the XGBoost classification model, the parameter settings are as follows: set the pre-tuned parameter learning rate eta = 0.7, the number of iterations nround = 20, the minimum loss function decrease value gamma = 0.001 required for node splitting, the maximum depth of the tree max_depth = 5, and the sum of the minimum sample weights min_child_weight = 10.
[0034] Furthermore, the specific steps for integrating the clustering results of spatial transcriptome data and the single-cell classification results using hypergraph partitioning are as follows:
[0035] First, construct a hypergraph G, where V represents the set of all base clustering results (clusters), C i is one of the clusters, and E represents the set of hyperedges e i constructed based on V. The number of points simultaneously connected by each hyperedge is greater than or equal to 2. Each hyperedge contains multiple nodes, and the nodes included in different hyperedges can be repeated. The weight
[0036] After constructing the hypergraph, use the MCLA algorithm to partition the graph G into k balanced meta-cluster classes Each meta-cluster class is represented by a characterization example and an m-dimensional indicator vector representing the degree of association between meta-cluster classes As shown, each example is then assigned to the meta-cluster class most relevant to it to obtain the integrated clustering cluster λ, which is the optimized final clustering result.
[0037] The second aspect of the present invention provides a system for analyzing substructures of spatial transcriptome biological tissues integrating single-cell transcriptomes, the system comprising: a memory, a processor, wherein the memory includes a program for a method for analyzing substructures of spatial transcriptome biological tissues integrating single-cell transcriptomes, and when the program for the method for analyzing substructures of spatial transcriptome biological tissues integrating single-cell transcriptomes is executed by the processor, the following steps are implemented:
[0038] S1. Obtain publicly available spatial transcriptome data and perform preprocessing;
[0039] S2. Encode the preprocessed spatial transcriptome data using the STAGATE framework, and perform clustering on the encoded result using the Louvain algorithm to obtain the clustering result of the spatial transcriptome data;
[0040] S3. Obtain publicly available single-cell RNA sequencing data and perform preprocessing, and divide the preprocessed single-cell RNA sequencing data into a training data set and a test data set;
[0041] S4. Use the training data set and the test data set to train the XGBoost classification model, and classify the single-cell sequencing data set homologous to the spatial transcriptome data to obtain the single-cell classification result;
[0042] S5. Integrate the clustering result of the spatial transcriptome data and the single-cell classification result using hypergraph partitioning.
[0043] The third aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a program for a method for analyzing substructures of spatial transcriptome biological tissues integrating single-cell transcriptomes, and when the program for the method for analyzing substructures of spatial transcriptome biological tissues integrating single-cell transcriptomes is executed by a processor, the steps of the method for analyzing substructures of spatial transcriptome biological tissues integrating single-cell transcriptomes are implemented.
[0044] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0045] The present invention uses the STAGATE transcriptome data for dimensionality reduction, analysis and clustering, uses XGBoost to cluster single-cell transcript data, improves the clustering accuracy of spatial transcript data and the classification effect of single-cell data, and at the same time uses hypergraph partitioning to integrate the two clustering results to obtain a clustering result with higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a flowchart of a method for analyzing the substructure of biological tissues in spatial transcriptomics integrated with single-cell transcriptomics according to the present invention.
[0047] Figure 2 This is a block diagram of a system for analyzing the substructure of biological tissues in spatial transcriptomics integrated with single-cell transcriptomics according to the present invention. Detailed implementation manners
[0048] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0049] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0050] Embodiment 1
[0051] As Figure 1 shown, the first aspect of the present invention provides a method for analyzing the substructure of biological tissues in spatial transcriptomics integrated with single-cell transcriptomics, including the following steps:
[0052] S1. Obtain publicly available spatial transcriptome data and perform preprocessing;
[0053] It should be noted that the preprocessing of the transcriptome data includes: normalization of the data and adjustment of the data format. Normalize the transcriptome data (such as screening highly differential genes, etc.) and convert the data format into a format that conforms to the input data of the algorithm.
[0054] S2. Encode the preprocessed spatial transcriptome data using the STAGATE framework, and use the Louvain algorithm to cluster the encoding results to obtain the clustering results of the spatial transcriptome data;
[0055] It should be noted that the STAGATE framework includes: a Spatial Neighbor Network (SNN) and a Graph Attention Autoencoder. Among them, the Spatial Neighbor Network is used for... The Graph Attention Encoder is used to learn a low-dimensional latent embedding with spatial information and gene expression. STAGATE first constructs a Spatial Neighbor Network (SNN) based on the relative spatial positions of nodes, and then learns a low-dimensional latent embedding (representing an abstract vector of an object) with spatial information and gene expression through the Graph Attention Autoencoder. The normalized expression of each node is first converted into a d-dimensional latent embedding by the encoder, and then inverted back to the reconstructed expression profile through the decoder. Among them, the attention mechanism is adopted in the middle layer of the encoder and decoder, which can adaptively learn the edge weights of the SNN (i.e., the similarity between adjacent nodes), and update the expression of a node with the SNN by aggregating the information of the neighbors of that node.
[0056] In a specific embodiment, the specific process of constructing the Spatial Neighbor Network (SNN) is as follows:
[0057] Convert the spatial information into an undirected neighbor network according to a predefined radius r. Define A as the adjacency matrix of the SNN. When and only when the Euclidean distance between node i and node j is less than r, A ij = 1, and A ij represents the element in the i-th row and j-th column of the adjacency matrix A. For example, for 10xVisium data, we set the radius r of the SNN network to a value that can include the six nearest nodes of each node. For spatial transcriptome data of other different technologies, r is selected according to the specific resolution of the data. With each node as the center and r as the radius, an average of 6 - 15 neighbor nodes are included; finally, a self-loop is added to each node.
[0058] The Graph Attention Autoencoder includes: an encoder, a decoder, and a graph attention layer. The graph attention layer is embedded in the encoder and the decoder;
[0059] Among them, the encoder takes the normalized gene expression of the node as input and generates a node vector spotembedding by aggregating the information of the neighbors of that node. The graph attention layer in the encoder has a total of L - 1 layers (k ∈ {1, 2,..., L - 1});
[0060] x i is the normalized expression of node i, L is the number of layers of the encoder, is the node vector spotembedding output by the k-th layer of the encoder, S i is the set of neighbors of node s, and W k is a trainable weight matrix;
[0061] Taking the expression profile of the node as the initial node vector spotembedding, we have:
[0062]
[0063] where is the edge weight between node i and node j in the output of the k-th graph attention layer;
[0064] The edge weight from node i to its neighbor node j where and are trainable weight vectors, and Sigmoid represents the sigmoid activation function;
[0065] To make the spatial similarity weights comparable, they are normalized by the softmax function: That is, the edge weight between node i and node j in the output of the k-th graph attention layer;
[0066] The L-th layer of the encoder does not adopt the attention mechanism, and the output is That is, the final output node vector spotembedding;
[0067] The decoder reconstructs the embedding of node i at the (k - 1)-th layer at the k-th layer from the bottom: The output of node i at the last layer of the decoder
[0068] The formula of the decoder is similar to that of the encoder, where
[0069] The loss function is
[0070] It should be noted that the present invention uses the Louvain algorithm to cluster the comparison results of the encoder output (i.e., the node vector spotembedding) to obtain the clustering results of the spatial transcriptome data. The resolution of the Louvain algorithm can be manually selected to adapt to spatial transcriptome data with different resolutions.
[0071] S3. Obtain the publicly available single-cell RNA sequencing data and perform preprocessing, and divide the preprocessed single-cell RNA sequencing data into a training data set and a test data set;
[0072] It should be noted that the preprocessing includes operations such as normalizing the data and adjusting the data format, and dividing the preprocessed single-cell RNA sequencing data into a training data set and a test data set.
[0073] S4. Use the training dataset and the test dataset to train the XGBoost classification model, classify the single-cell sequencing dataset homologous to the spatial transcriptome data, and obtain the single-cell classification results;
[0074] It should be noted that when training the XGBoost classification model, the parameter settings are as follows: set the pre-adjusted parameter learning rate eta = 0.7, the number of iterations nround = 20, the minimum loss function decrease value gamma required for node splitting = 0.001, the maximum depth of the tree max_depth = 5, and the sum of the minimum sample weights min_child_weight = 10. Among them, if you are not satisfied with the classification accuracy after training, you can adjust the parameters accordingly on this basis. Finally, use the trained model to classify the single-cell sequencing dataset homologous (the same sample) to the spatial transcriptome data to obtain the single-cell classification results.
[0075] S5. Use hypergraph partitioning to integrate the clustering results of spatial transcriptome data and the single-cell classification results.
[0076] It should be noted that the specific steps of using hypergraph partitioning to integrate the clustering results of spatial transcriptome data and the single-cell classification results are as follows:
[0077] First, construct a hypergraph G, where V represents the set of all base clustering results (clusters), C i is one of the clusters, and E represents the set of hyperedges e i constructed based on V. The number of points connected by each hyperedge is greater than or equal to 2. Each hyperedge contains multiple nodes, and the nodes contained in different hyperedges can be repeated. The weight
[0078] After constructing the hypergraph, use the MCLA algorithm to partition the graph G into k balanced meta-cluster classes Each meta-cluster class is represented by a characterization example and an m-dimensional indicator vector representing the degree of association between meta-cluster classes Next, assign each example to the meta-cluster class most relevant to it to obtain the integrated clustering cluster λ, which is the optimized final clustering result.
[0079] As Figure 2 shown, the second aspect of the present invention provides a spatial transcriptome biological tissue substructure analysis system integrating single-cell transcriptome. The system includes: a memory and a processor. The memory includes a program for a method for analyzing the substructure of a spatial transcriptome biological tissue integrating single-cell transcriptome. When the program for the method for analyzing the substructure of a spatial transcriptome biological tissue integrating single-cell transcriptome is executed by the processor, the following steps are implemented:
[0080] S1. Obtain the publicly available spatial transcriptome data and perform preprocessing;
[0081] S2. Encode the preprocessed spatial transcriptome data using the STAGATE framework, and use the Louvain algorithm to cluster the encoding results to obtain the spatial transcriptome data clustering results;
[0082] S3. Obtain the publicly available single-cell RNA sequencing data and perform preprocessing, and divide the preprocessed single-cell RNA sequencing data into a training dataset and a test dataset;
[0083] S4. Use the training dataset and the test dataset to train the XGBoost classification model to classify the single-cell sequencing dataset homologous to the spatial transcriptome data to obtain the single-cell classification results;
[0084] S5. Integrate the spatial transcriptome data clustering results and the single-cell classification results using hypergraph partitioning.
[0085] The third aspect of the present invention provides a computer-readable storage medium, which includes a program for the method for analyzing the substructure of a spatial transcriptome biological tissue integrating single-cell transcriptomes. When the program for the method for analyzing the substructure of a spatial transcriptome biological tissue integrating single-cell transcriptomes is executed by a processor, the steps of the method for analyzing the substructure of a spatial transcriptome biological tissue integrating single-cell transcriptomes are implemented.
[0086] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for analyzing the substructure of a spatial transcriptome of biological tissues integrating single-cell transcriptomes, characterized in that, It includes the following steps: S1. Obtain the publicly available spatial transcriptome data and perform preprocessing; S2. Encode the preprocessed spatial transcriptome data using the STAGATE framework, and use the Louvain algorithm to cluster the encoding results to obtain the clustering results of the spatial transcriptome data; The STAGATE framework includes: a spatial neighbor network SNN and a graph attention autoencoder. Among them, the graph attention autoencoder is used to learn a low-dimensional latent vector with spatial information and gene expression. The specific process of constructing the spatial neighbor network SNN is as follows: Convert the spatial information into an undirected neighbor network according to the predefined radius r. Define A as the adjacency matrix of the SNN. A ij = 1, where A ij represents the element in the i-th row and j-th column of the adjacency matrix A. For spatial transcriptomic data of other different technologies, r is selected according to the specific resolution of the data. With each node as the center and r as the radius, an average of 6 - 15 neighbor nodes are included. Finally, a self-loop is added to each node; S3. Obtain the publicly available single-cell RNA sequencing data and perform preprocessing, and divide the preprocessed single-cell RNA sequencing data into a training dataset and a test dataset; S4. Use the training dataset and the test dataset to train the XGBoost classification model to classify the single-cell sequencing dataset homologous to the spatial transcriptome data to obtain the single-cell classification results; S5. Use hypergraph partitioning to integrate the clustering results of the spatial transcriptome data and the single-cell classification results. The specific steps are as follows: First, construct a hypergraph G, where V represents the set of result clusters of all base clusters, and C i is one of the clusters, and E represents the set of hyperedges e i constructed based on V. The number of points connected by a hyperedge is greater than or equal to 2. Each hyperedge contains multiple nodes, and the nodes contained in different hyperedges can be repeated. The weight After constructing the hypergraph, the MCLA algorithm is used to partition the graph G into k balanced meta-cluster classes Each meta-cluster class is represented by an m-dimensional indicator vector that characterizes the example and the degree of association between meta-cluster classes represented by. Next, each example is assigned to the meta-cluster class that is most relevant to it, obtaining the integrated clustering cluster λ, which is the optimized final clustering result 2. A method for analyzing the substructure of a spatial transcriptome biological tissue integrating single-cell transcriptomes according to claim 1, characterized in that, The preprocessing of the publicly available spatial transcriptome data in step S1 includes: data normalization and data format adjustment.
3. A method for analyzing the substructure of a spatial transcriptome biological tissue integrating single-cell transcriptomes according to claim 1, characterized in that, The graph attention autoencoder includes: an encoder, a decoder, and a graph attention layer, and the graph attention layer is embedded in the encoder and the decoder; Among them, the encoder takes the normalized gene expression of the node as input and generates a node vector by aggregating the information of the neighbors of the node. The graph attention layer in the encoder has a total of L - 1 layers; x i is the normalized expression of node i, and L is the number of layers of the encoder. is the node embedding output by the k-th layer of the encoder, and S i is the set of neighbors of node s, and W k is a trainable weight matrix; Taking the expression profile of the node as the initial node vector, then there is: where is the edge weight between node i and node j in the output of the k-th graph attention layer, where k ∈ {1, 2, ..., L - 1}; The edge weight from node i to its neighbor node j where and are trainable weight vectors, and Sigmoid represents the sigmoid activation function; To make the spatial similarity weights comparable, they are normalized by the softmax function: That is, the edge weight between node i and node j in the output of the k-th graph attention layer; The L-th layer of the encoder does not adopt the attention mechanism, and the output is That is, the node vector of the final output; The decoder reconstructs the vector of node i at the (k-1)-th layer at the penultimate k-th layer: The output of node i at the last layer of the decoder Among them The loss function is 4. A method for analyzing the substructure of a spatial transcriptome biological tissue integrating single-cell transcriptomes according to claim 1, characterized in that When training the XGBoost classification model, the parameter settings are as follows: set the pre-tuned parameter learning rate eta = 0.7, the number of iterations nround = 20, the minimum loss function decrease value gamma required for node splitting = 0.001, the maximum depth of the tree max_depth = 5, and the sum of the minimum sample weights min_child_weight = 10.
5. A spatial transcriptome biological tissue substructure analysis system integrating single-cell transcriptomes, characterized in that, The system includes: a memory and a processor. The memory includes a program for a method for analyzing the substructure of a spatial transcriptome biological tissue that fuses a single-cell transcriptome. When the program for the method for analyzing the substructure of a spatial transcriptome biological tissue that fuses a single-cell transcriptome is executed by the processor, the following steps are implemented: S1. Obtain the publicly available spatial transcriptome data and perform preprocessing; S2. Encode the preprocessed spatial transcriptome data using the STAGATE framework, and use the Louvain algorithm to cluster the encoding results to obtain the clustering results of the spatial transcriptome data; The STAGATE framework includes: a spatial neighbor network SNN and a graph attention autoencoder. Among them, the graph attention autoencoder is used to learn a low-dimensional latent vector with spatial information and gene expression. The specific process of constructing the spatial neighbor network SNN is as follows: Convert the spatial information into an undirected neighbor network according to a predefined radius r. Define A as the adjacency matrix of the SNN. When and only when the Euclidean distance between node i and node j is less than r, A ij = 1, A ij represents the element in the i-th row and j-th column of the adjacency matrix A; for spatial transcriptome data of other different technologies, r is selected according to the specific resolution of the data. With each node as the center and r as the radius, on average, 6 - 15 neighbor nodes are included; finally, a self-loop is added to each node; S3. Obtain the publicly available single-cell RNA sequencing data and perform preprocessing, and divide the preprocessed single-cell RNA sequencing data into a training dataset and a test dataset; S4. Train the XGBoost classification model using the training dataset and the test dataset to classify the single-cell sequencing dataset homologous to the spatial transcriptome data, and obtain the single-cell classification results; S5. Integrate the spatial transcriptome data clustering results and the single-cell classification results using hypergraph partitioning. The specific steps are as follows: First, construct a hypergraph G, where V represents the set of result clusters of all base clusters, and C i is one of the clusters, and E represents the set of hyperedges e i constructed based on V. The number of points connected by each hyperedge is greater than or equal to 2. Each hyperedge contains multiple nodes, and the nodes contained in different hyperedges can be repeated. The weight After constructing the hypergraph, the MCLA algorithm is used to partition the graph G into k balanced meta-cluster classes Each meta-cluster class is represented by an m-dimensional indicator vector that characterizes the example and the degree of association between meta-cluster classes represented by. Next, each example is assigned to the meta-cluster class that is most relevant to it, resulting in the integrated clustering cluster λ, which is the optimized final clustering result 6. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a program for the method of analyzing the substructure of a spatial transcriptome biological tissue by integrating single-cell transcriptomes. When the program for the method of analyzing the substructure of a spatial transcriptome biological tissue by integrating single-cell transcriptomes is executed by a processor, the steps of a method of analyzing the substructure of a spatial transcriptome biological tissue by integrating single-cell transcriptomes as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Spatial transcriptome cell clustering and analyzing method
CN114091603A