Method for dividing organization structure of space transcriptome data
Through the autoencoder architecture and hypergraph attention layer processing spatial transcriptome data, the problems of discontinuity and misjudgment of organizational structure division in the prior art are solved, and high-precision and biologically interpretable organizational structure division are achieved.
Patent Information
- Application Number
- CN202510427044.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-11
AI Technical Summary
In the analysis of spatial transcriptome data, the existing technology has problems of ignoring the discontinuity of tissue structure division caused by high-order tissue structure, gene expression heterogeneity and structural boundary ambiguity, as well as misjudgment problems caused by complex associations between multicellular programs, making it difficult to achieve high-precision tissue structure division.
The autoencoder architecture is adopted to explicitly characterize the complex higher-order relationships between spatially adjacent multicellular units through hypergraph structures. Combined with the gene expression-guided hyper-edge decomposition algorithm and autoencoder architecture, the low-dimensional characterization of spatial topology and gene expression is optimized, and data processing is used using the k nearest neighbor algorithm and hypergraph attention layer to generate low-dimensional characterization and cluster analysis.
It significantly improves the accuracy and continuity of organizational structure division, eliminates pseudo-neighborhood noise interference, improves biological interpretability and semantic consistency of cross-resolution data, and supports high-precision organizational structure division.
Smart Images

Figure CN120299524A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and bioinformatics, and in particular to a method for partitioning the organizational structure of spatial transcriptome data. Background Art
[0002] In biological systems, the precise regulation of life activities highly depends on the three-dimensional microenvironment in which cells are located. The spatial localization of cells and their inherent molecular characteristics together constitute the dual determinants of functional regulation, directly affecting cell differentiation fate, metabolic status, and signal response characteristics. Although single-cell omics technologies can achieve cell-level analysis of molecular characteristics, the inevitable single-cell dissociation operation during sample preparation will lead to the loss of key spatial topological information. The emerging spatial transcriptome technology realizes the determination of the whole-genome expression profile on the premise of retaining the integrity of the tissue spatial architecture through in situ molecular capture technology, providing a revolutionary technical means for revealing the spatial regulation mechanism of complex biological systems.
[0003] Among them, the partitioning of the organizational structure, as the core task of spatial transcriptome data analysis, directly determines the effectiveness of downstream analyses such as gene expression denoising, cell communication network reconstruction, and spatio-temporal developmental trajectory inference. Currently, there are mainly two types of computational methods in the field of organizational structure partitioning: 1) Traditional clustering methods: directly transplant single-cell analysis methods, completely ignoring spatial neighborhood information, resulting in discontinuous fragmentation of organizational structure partitioning; 2) Space-enhanced methods: improve the clustering effect through the linear combination of spatial coordinates and gene expression, but still have the following technical bottlenecks: only model the binary neighborhood relationship between cells, and fail to characterize the higher-order functional modules formed by multicellular cooperation; the resolution of organizational structure boundary determination is insufficient, making it difficult to meet the requirements of high-precision spatial transcriptome data structure analysis. These technical defects severely restrict the precise analysis of spatial functional units in complex biological tissues and urgently require breakthrough progress through method innovation. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a method for partitioning the organizational structure of spatial transcriptome data, which can effectively solve the problem of discontinuous organizational structure partitioning in spatial transcriptome data analysis caused by ignoring higher-order organizational structures, gene expression heterogeneity, and fuzzy structure boundaries, as well as the misjudgment problem caused by the complex association between multicellular programs, and significantly improve the accuracy of organizational structure partitioning.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: A method for partitioning the organizational structure of spatial transcriptome data, comprising the following steps:
[0006] S1: Perform tissue region spatial locus filtering, screening of highly variable genes, and gene expression normalization on the original spatial transcriptome data to obtain a gene expression matrix containing the gene expression characteristics of each spatial locus and a spatial coordinate matrix of the spatial position relationship. The two together constitute a standardized data set;
[0007] S2: Use the standardized data set to construct hyperedges according to the spatial coordinate matrix through the k-nearest neighbor algorithm; dynamically adjust the hyperedge connection weights of the obtained hyperedges through gene expression similarity; design a gene expression-guided hyperedge decomposition algorithm to eliminate cross-domain noise: if all the loci within the hyperedge belong to the same pre-clustering cluster, retain them; if there are cross-cluster loci, then select to eliminate outliers or split the hyperedges according to the quantity to ensure the gene expression homogeneity of the loci within each hyperedge, and finally generate a hypergraph structure that fuses spatial proximity and gene expression consistency;
[0008] S3: Input the obtained hypergraph structure and gene expression matrix into an autoencoder architecture containing an encoder and a decoder; the encoder gradually reduces the dimension through three layers of hypergraph attention layers, and uses a gene expression correlation attention mechanism to map high-dimensional features to a low-dimensional latent space; the decoder uses three layers of hypergraph attention layers symmetric to the encoder to reconstruct the gene expression matrix; drive the architecture optimization by minimizing the mean square error between the reconstructed gene expression matrix and the input gene expression matrix, and finally output a low-dimensional representation;
[0009] S4: Cluster the obtained low-dimensional representation and divide the tissue structure according to the clustering results.
[0010] Furthermore, in step S1, the following processing flow is performed: First, perform spatial locus filtering outside the tissue region to eliminate spatial loci located outside the main tissue region to focus on biologically functional regions; then perform gene expression feature screening to extract highly variable genes as the key feature set; perform logarithmic transformation on the expression values of the selected key feature set to achieve data normalization and obtain a standardized highly variable gene expression matrix; finally, use the standardized highly variable gene expression matrix as the gene expression matrix of the standardized data set. Through the above processing flow, complex biological associations between spatial loci can be effectively captured, thereby improving the division accuracy and biological interpretability of tissue structures.
[0011] Further, in step S2, based on the spatial coordinate matrix, the k-nearest neighbor algorithm is first used to generate initial hyperedges according to the spatial coordinates to capture the local associations of spatially adjacent sites; a gene expression-guided hyperedge decomposition algorithm is designed to eliminate cross-domain noise: if all sites within a hyperedge belong to the same pre-cluster, the original hyperedge is retained; if there is only one cross-cluster site, the cross-cluster site is removed to maintain the homogeneity of gene expression within the hyperedge; if there are multiple cross-cluster sites, the original hyperedge is decomposed into two new hyperedges to ensure that each new hyperedge contains only sites within the same cluster; through the above hyperedge decomposition algorithm, the noise connections that are spatially adjacent but biologically inconsistent are eliminated, enabling the hypergraph structure to encode both spatial proximity and gene expression similarity simultaneously, thereby improving the accuracy of tissue structure division.
[0012] Further, step S3 includes the following steps:
[0013] 1) Denote the hypergraph structure generated in step S2 as G=(X, H), where the feature matrix X represents the gene expression matrix of spatial sites, and the incidence matrix H defines the connection relationship between vertices and hyperedges. The element h at the i-th row and j-th column of H ij is defined as:
[0014]
[0015] In the formula, x i represents the i-th vertex, e j represents the j-th hyperedge, represents the set of vertices connected to the hyperedge e j and ∈ represents the membership symbol;
[0016] 2) Use the encoder Encoder(·) for encoding: The encoder takes the gene expression matrix X0 output in step S1 as the initial layer and generates a low-dimensional representation of the spatial sites by aggregating the information of adjacent nodes within the same hyperedge. To adaptively learn the weights of hyperedges, three layers of hypergraph attention layers are used;
[0017] 3) Hypergraph attention layer: The information transfer mechanism on the hypergraph is more complex than that on the ordinary graph, including two key stages: the hyperedge aggregation stage based on the attention mechanism and the vertex aggregation stage based on the attention mechanism;
[0018] Hyperedge aggregation stage based on the attention mechanism: In this stage, the features of all connected vertices within the same hyperedge will be aggregated to generate the corresponding hyperedge embedding representation; among them, the hyperedge feature E t-1 at the (t - 1)-th layer is calculated as:
[0019] E t-1 = H T X t-1
[0020] In the formula, Xt-1 Denote the vertex feature matrix of layer t - 1, H T Denote the transpose of H; for each pair of hyperedges (e i , e j ), where Denote the set of hyperedges adjacent to hyperedge e i . By introducing the attention mechanism, calculate the attention coefficient a i between the i-th hyperedge e j and the j-th hyperedge e ij :
[0021]
[0022] In the formula, Denote the current layer feature vector of hyperedge e i , Denote the current layer feature vector of hyperedge e j , Denote the current layer feature vector of the hyperedge e i adjacent to hyperedge e k , W a Denote the learnable linear transformation weight matrix related to the attention coefficient a ij between hyperedges, Denote the attention parameter vector related to the attention coefficient a ij between hyperedges, ∑ denotes summation, LeakyReLU denotes the leaky rectified linear unit function, ‖ denotes the vector concatenation operation, exp denotes the natural exponential function; the hyperedge feature matrix E t of the output of layer t obtained by attention hyperedge aggregation can be calculated as follows:
[0023] E t = AWH T X t-1
[0024] In the formula, A denotes the attention weight matrix between hyperedges, H denotes the incidence matrix of the hypergraph, X t-1 denotes the vertex feature matrix of layer t - 1; in the vertex aggregation stage based on the attention mechanism, the vertex embedding is generated by integrating the information of connected hyperedges, and the vertex feature matrix X t of layer t obtained by hyperedge aggregation can be calculated by the following formula:
[0025] X t = HE t
[0026] In the vertex aggregation stage based on the attention mechanism: First, based on the attention mechanism, calculate the i-th hyperedge e i and the j-th vertex x jThe attention coefficient b between ij :
[0027]
[0028] In the formula, represents the set of vertices connected to the hyperedge e i , x k represents the k-th vertex belonging to , represents the current layer feature vector of the hyperedge e i , represents the current layer feature vector of the vertex x j , represents the current layer feature vector of the vertex x k , represents the attention parameter vector related to the attention coefficient b between the hyperedge and the vertex; Next, the output low-dimensional representation X of the vertex is calculated as follows ij : e :
[0029] X e = σ(BHAWH T X t-1 )
[0030] In the formula, e in X e indicates that the obtained low-dimensional representation of the vertex is a low-dimensional embedding, σ represents the activation function, B represents the attention weight matrix of the hyperedge-vertex composed of b ij , H T represents the transpose of the hypergraph incidence matrix, A represents the attention weight matrix of the hyperedge-hyperedge composed of a ij , X t-1 represents the vertex feature matrix of the t-1 layer;
[0031] 4) Use the decoder Decoder(·) for decoding: Use the low-dimensional representation X of the vertex of the spatial site obtained in step 3) e as the input, and reconstruct the site features through a symmetric hypergraph attention network architecture. The specific implementation is as follows: Construct a decoder structure using a hypergraph attention network symmetric to the encoder. After being processed by three layers of hypergraph attention networks, the output of the last layer is used as the reconstructed vertex feature; The output X' t of the t-th layer in the decoder is obtained through the following transformation:
[0032]
[0033] In the formula, σ represents the activation function, X' t represents the vertex feature matrix reconstructed by the t-th layer of the decoder, H represents the hypergraph incidence matrix, X' t-1represents the vertex feature matrix after reconstruction of the (t - 1)-th layer of the decoder, W k-t+1 represents the learnable linear transformation weight matrix of the (k - t + 1)-th layer of the decoder, where k represents the total number of layers of the decoder, A k-t+1 represents the hyperedge-hyperedge attention weight matrix of the (k - t + 1)-th layer of the decoder, B k-t+1 represents the hyperedge-vertex attention weight matrix of the (k - t + 1)-th layer of the decoder;
[0034] 5) Training using a loss function: Optimize the autoencoder architecture by minimizing the following loss function:
[0035]
[0036] wherein, represents the summation operation over all N vertices in the dataset, X in represents the gene expression matrix of the input of the encoder, X out represents the vertex feature matrix finally reconstructed by the decoder, ||·||2 represents the L2 norm of the vector, where:
[0037] X in = X0
[0038] X out = X′ k
[0039] wherein, X′ k represents the vertex feature matrix after reconstruction of the k-th layer of the decoder.
[0040] Furthermore, in step S4, using the vertex low-dimensional representation X e output in step S3 as the input, perform clustering analysis using the mclust algorithm based on the Gaussian mixture model to obtain the clustering labels of each spatial site; subsequently, align the clustering labels of each spatial site with the prior labels of known tissue structures: First, calculate the average gene expression matrix of each clustering label and the prior labels of known tissue structures, construct a cost matrix through negative similarity measurement, where the higher the similarity of the expression matrix, the lower the matching cost, and the greater the difference, the higher the matching cost; subsequently, based on this cost matrix, use the entropy-regularized Sinkhorn algorithm to solve the minimum transport strategy, and assign the clustering labels to the most matching known tissue structure categories; finally, divide each spatial site into the corresponding tissue structure category according to its clustering label and label matching result, and output a biologically interpretable and spatially topologically coherent tissue structure division result.
[0041] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0042] 1. Explicitly represent the complex high-order relationships between spatially adjacent multi-cellular units using a hypergraph structure, breaking through the limitation of traditional methods that only model binary proximity relationships, and accurately partitioning the organizational structure.
[0043] 2. Through a gene-expression-guided hyperedge decomposition algorithm, adaptively correct the consistency between spatial neighborhood information and molecular features, and solve the problem of pseudo-neighborhood interference in the boundary regions of the organizational structure.
[0044] 3. Design an autoencoder architecture to jointly optimize the low-dimensional representations of spatial topology and gene expression, and support the continuous partitioning of the organizational structure across different resolutions.
[0045] In summary, through the autoencoder architecture, the present invention realizes the efficient integration of spatial topological relationships and molecular functional features, significantly improving the spatial continuity of the continuous partitioning of the organizational structure; its innovative architecture eliminates the interference of pseudo-neighborhood noise while effectively bridging the semantic gap of cross-resolution data, synchronously enhancing the accuracy and interpretability of the partitioning of complex organizational structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic diagram of the relationship between the various steps of the present invention.
[0047] Figure 2 It is a schematic diagram of the data processing of the present invention.
[0048] Figure 3 It is a schematic diagram of the construction of the hypergraph structure of the present invention.
[0049] Figure 4 It is a schematic diagram of the autoencoder architecture and the partitioning of the organizational structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The present invention will be further described in detail below in conjunction with the embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0051] As Figure 1 shown, this embodiment discloses a method for partitioning the organizational structure of spatial transcriptome data, and the specific situation is as follows:
[0052] S1: As Figure 2 shown, perform tissue region spatial site filtering, screening of highly variable genes, and gene expression normalization processing on the original spatial transcriptome data to obtain a gene expression matrix containing the gene expression characteristics of each spatial site and a spatial coordinate matrix of spatial position relationships, and the two together constitute a standardized data set.
[0053] Execute the following processing flow: First, perform spatial locus filtering outside the tissue region to eliminate spatial loci outside the main tissue region to focus on biologically functional regions; then perform gene expression feature screening to extract highly variable genes as the key feature set; perform logarithmic transformation on the expression values of the selected key feature set to achieve data normalization and obtain the normalized highly variable gene expression matrix; finally, use the normalized highly variable gene expression matrix as the gene expression matrix of the normalized data set. Through the above processing flow, complex biological associations between spatial loci can be effectively captured, thereby improving the division accuracy and biological interpretability of tissue structures.
[0054] S2: As Figure 3 shown, use the normalized data set to construct hyperedges according to spatial positions through the k-nearest neighbor algorithm; dynamically adjust the hyperedge connection weights of the obtained hyperedges through gene expression similarity; design a gene expression-guided hyperedge decomposition algorithm to eliminate cross-domain noise: if all loci within a hyperedge belong to the same pre-cluster, retain it; if there are cross-cluster loci, then select to eliminate cross-cluster loci or split the hyperedge according to the quantity to ensure the gene expression homogeneity of loci within each hyperedge. As Figure 3 shown, hyperedge e1 is not decomposed because the pre-clustering of each vertex in it belongs to one cluster; hyperedge e2 is decomposed into hyperedge e′2 and hyperedge e4 because the vertices in it are pre-clustered into two clusters; hyperedge e3 is decomposed into hyperedge e′3 and hyperedge e5 because the vertices in it are pre-clustered into two clusters. Through the above hyperedge decomposition algorithm, noise connections that are spatially adjacent but biologically inconsistent are eliminated, enabling the hypergraph structure to simultaneously encode spatial proximity and gene expression similarity, thereby improving the accuracy of tissue structure division. Finally, a hypergraph structure that fuses spatial proximity and gene expression consistency is generated.
[0055] S3: As Figure 4 shown, input the obtained hypergraph structure and gene expression matrix into an autoencoder architecture containing an encoder and a decoder; the encoder gradually reduces the dimension through three layers of hypergraph attention layers, and uses the gene expression correlation attention mechanism to map high-dimensional features to a low-dimensional latent space; the decoder uses three layers of hypergraph attention layers symmetric to the encoder to reconstruct the gene expression matrix; optimize the architecture by minimizing the mean square error between the reconstructed gene expression matrix and the input gene expression matrix, and finally output a low-dimensional representation. It includes the following steps:
[0056] 1) Denote the hypergraph structure generated in step S2 as G=(X,H), where the feature matrix X represents the gene expression matrix of spatial loci, and the incidence matrix H defines the connection relationship between vertices and hyperedges. The element h at the i-th row and j-th column of H ij is defined as:
[0057]
[0058] In the formula, xi represents the \(i\)-th vertex, \(e\) j represents the \(j\)-th hyperedge, represents the set of vertices connected to the hyperedge \(e\) j and \(\in\) represents the membership symbol;
[0059] 2) Encode using the Encoder(·): The encoder takes the gene expression matrix \(X_0\) output in step S1 as the initial layer and generates a low-dimensional representation of the spatial positions by aggregating the information of adjacent nodes within the same hyperedge. To adaptively learn the weights of the hyperedges, three layers of hypergraph attention layers are used;
[0060] 3) Hypergraph attention layer: The information transfer mechanism on the hypergraph is more complex than that on the ordinary graph and includes two key stages: the hyperedge aggregation stage based on the attention mechanism and the vertex aggregation stage based on the attention mechanism;
[0061] Hyperedge aggregation stage based on the attention mechanism: In this stage, the features of all connected vertices within the same hyperedge will be aggregated to generate the corresponding hyperedge embedding representation; among them, the hyperedge feature \(E\) of the \((t - 1)\)-th layer t-1 is calculated by the formula:
[0062] \(E\) t-1 = \(H\) T \(X\) t-1
[0063] where \(X\) t-1 represents the vertex feature matrix of the \((t - 1)\)-th layer, and \(H\) T represents the transpose of \(H\); for each pair of hyperedges \((e\) i , \(e\) j ), where represents the set of hyperedges adjacent to the hyperedge \(e\) i , by introducing the attention mechanism, the attention coefficient \(a\) i between the \(i\)-th hyperedge \(e\) j and the \(j\)-th hyperedge \(e\) ij is calculated as follows:
[0064]
[0065] where represents the current layer feature vector of the hyperedge \(e\) i , represents the current layer feature vector of the hyperedge \(e\) j , represents the current layer feature vector of the hyperedge \(e\) i adjacent to the hyperedge \(e\) k , and \(W\) a represents the learnable linear transformation weight matrix related to the attention coefficient \(a\) ij between hyperedges, Denote the attention coefficient a between hyperedges ij The related attention parameter vector, ∑ denotes summation, LeakyReLU denotes the leaky rectified linear unit function, ‖ denotes the vector concatenation operation, and exp denotes the natural exponential function; The hyperedge feature matrix E of the output of the t-th layer obtained by attention hyperedge aggregation t Can be calculated as follows
[0066] E t = AWH T X t-1
[0067] In the formula, A represents the attention weight matrix between hyperedges, H represents the incidence matrix of the hypergraph, and X t-1 Represents the vertex feature matrix of the (t - 1)-th layer; In the vertex aggregation stage based on the attention mechanism, the vertex embedding is generated by integrating the information of connected hyperedges. The vertex feature matrix X of the t-th layer obtained by hyperedge aggregation t Can be calculated by the following formula
[0068] X t = HE t
[0069] Vertex aggregation stage based on the attention mechanism: First, based on the attention mechanism, calculate the attention coefficient b i Between the i-th hyperedge e j And the j-th vertex x ij :
[0070]
[0071] In the formula Denotes the set of vertices connected to the hyperedge e i , x k Represents the k-th vertex belonging to , Denotes the current layer feature vector of the hyperedge e i , Denotes the current layer feature vector of the vertex x j , Denotes the current layer feature vector of the vertex x k , Denotes the attention parameter vector related to the attention coefficient b ij Between hyperedges and vertices; Next, the output vertex low-dimensional representation X is calculated as follows e :
[0072] X e = σ(BHAWH T X t-1 )
[0073] In the formula, X e where e in e indicates that the obtained vertex low-dimensional representation is a low-dimensional embedding, σ represents the activation function, B represents the hyperedge-vertex attention weight matrix composed of b ij and H T represents the transpose of the hypergraph incidence matrix, A represents the hyperedge-hyperedge attention weight matrix composed of a ij and X t-1 represents the vertex feature matrix at layer t - 1;
[0074] 4) Decode using the decoder Decoder(·): Use the vertex low-dimensional representation X e of the spatial sites obtained in step 3) as the input, and reconstruct the site features through a symmetric hypergraph attention network architecture. The specific implementation is as follows: Construct a decoder structure using a hypergraph attention network symmetric to the encoder. After being processed by three layers of hypergraph attention networks, the output of the last layer is used as the reconstructed vertex features; the output X t ' of the t-th layer in the decoder is obtained through the following transformation:
[0075]
[0076] In the formula, σ represents the activation function, X t ' represents the vertex feature matrix reconstructed by the t-th layer of the decoder, H represents the hypergraph incidence matrix, X' t-1 represents the vertex feature matrix reconstructed by the (t - 1)-th layer of the decoder, W k-t+1 represents the learnable linear transformation weight matrix of the (k - t + 1)-th layer of the decoder, where k represents the total number of layers of the decoder, A k-t+1 represents the hyperedge-hyperedge attention weight matrix of the (k - t + 1)-th layer of the decoder, and B k-t+1 represents the hyperedge-vertex attention weight matrix of the (k - t + 1)-th layer of the decoder;
[0077] 5) Train using the loss function: Optimize the autoencoder architecture by minimizing the following loss function:
[0078]
[0079] In the formula, represents the summation operation over all N vertices in the dataset, X in represents the gene expression matrix of the input of the encoder, X out represents the vertex feature matrix finally reconstructed by the decoder, ||·||2 represents the L2 norm of the vector, where:
[0080] X in = X0
[0081] Xout = X' k
[0082] where X' k represents the vertex feature matrix after reconstruction of the k-th layer of the decoder.
[0083] S4: As Figure 4 shown, perform mclust clustering on the obtained low-dimensional representation, and divide the organizational structure according to the clustering results, specifically as follows:
[0084] Taking the vertex low-dimensional representation X e output in step S3 as the input, perform clustering analysis using the mclust algorithm based on the Gaussian mixture model to obtain the clustering labels of each spatial site; subsequently, align the clustering labels of each spatial site with the prior labels of the known organizational structure: first calculate the average gene expression matrix of each clustering label and the prior labels of the known organizational structure, and construct a cost matrix through negative similarity measurement, where the higher the similarity of the expression matrix, the lower the matching cost, and the greater the difference, the higher the matching cost; subsequently, based on this cost matrix, use the entropy-regularized Sinkhorn algorithm to solve the minimum transport strategy, and assign the clustering labels to the most matching known organizational structure categories; finally, divide each spatial site into the corresponding organizational structure category according to its clustering label and label matching result, and output a biologically interpretable and spatially topologically coherent organizational structure division result.
[0085] In summary, the method of the present invention effectively solves the problem of discontinuous domain division in spatial transcriptome data analysis caused by ignoring high-order organizational structure, gene expression heterogeneity, and spatial boundary ambiguity, as well as the problem of misjudgment of functional modules caused by complex associations between multicellular programs, significantly improves the accuracy of spatial domain recognition, the sensitivity of multi-omics integration, and the ability to analyze the spatial evolution path of cross-platform data, has practical application value, and is worthy of promotion.
[0086] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A method for dividing the organizational structure of spatial transcriptome data, characterized in that It includes the following steps: S1: Perform tissue region spatial locus filtering, screening of highly variable genes, and gene expression normalization on the original spatial transcriptome data to obtain a gene expression matrix containing the gene expression characteristics of each spatial locus and a spatial coordinate matrix of the spatial position relationship. The two together constitute a standardized data set; S2: Use the standardized data set to construct hyperedges according to the spatial coordinate matrix through the k-nearest neighbor algorithm; dynamically adjust the hyperedge connection weights of the obtained hyperedges through gene expression similarity; design a gene expression-guided hyperedge decomposition algorithm to eliminate cross-domain noise: if all loci within the hyperedge belong to the same pre-clustering cluster, then retain it; If there are cross-cluster loci, then select to eliminate outliers or split the hyperedge according to the quantity to ensure the gene expression homogeneity of the loci within each hyperedge, and finally generate a hypergraph structure that fuses spatial proximity and gene expression consistency; S3: Input the obtained hypergraph structure and gene expression matrix into an autoencoder architecture including an encoder and a decoder; the encoder gradually reduces the dimension through three layers of hypergraph attention layers, and uses the gene expression correlation attention mechanism to map high-dimensional features to a low-dimensional latent space; the decoder uses three layers of hypergraph attention layers symmetric to the encoder to reconstruct the gene expression matrix; optimize the architecture by minimizing the mean square error between the reconstructed gene expression matrix and the input gene expression matrix, and finally output a low-dimensional representation; S4: Cluster the obtained low-dimensional representation and divide the tissue structure according to the clustering results.
2. The method for dividing the organizational structure of spatial transcriptome data according to claim 1, characterized in that, In step S1, the following processing flow is executed: First, perform spatial locus filtering outside the tissue region to eliminate spatial loci outside the main tissue region to focus on biologically functional regions; Subsequently, perform gene expression feature screening to extract highly variable genes as the key feature set; Perform logarithmic transformation on the expression values of the selected key feature set to achieve data normalization, and obtain a standardized highly variable gene expression matrix; finally, use the standardized highly variable gene expression matrix as the gene expression matrix of the standardized data set; through the above processing flow, complex biological associations between spatial loci can be effectively captured, thereby improving the division accuracy and biological interpretability of the tissue structure.
3. The organizational structure division method for spatial transcriptome data according to claim 1, wherein In step S2, based on the spatial coordinate matrix, first use the k-nearest neighbor algorithm to generate initial hyperedges according to the spatial coordinates to capture the local associations of spatially adjacent loci; design a gene expression-guided hyperedge decomposition algorithm to eliminate cross-domain noise: if all loci within the hyperedge belong to the same pre-clustering cluster, then retain the original hyperedge; If there is only one cross-cluster locus, then eliminate the cross-cluster locus to maintain the gene expression homogeneity within the hyperedge; if there are multiple cross-cluster loci, then decompose the original hyperedge into two new hyperedges to ensure that each new hyperedge only contains loci within the same cluster; through the above hyperedge decomposition algorithm, eliminate noise connections that are spatially adjacent but biologically inconsistent, so that the hypergraph structure encodes both spatial proximity and gene expression similarity, thereby improving the accuracy of tissue structure division.
4. A method for dividing the organizational structure of spatial transcriptome data according to claim 1, characterized in that, The said step S3 includes the following steps: 1) Denote the hypergraph structure generated in step S2 as G = (X, H), where the feature matrix X represents the gene expression matrix of spatial sites, the incidence matrix H defines the connection relationship between vertices and hyperedges, and the element h in the i-th row and j-th column of H ij is defined as: where x i represents the i-th vertex, and e j represents the j-th hyperedge, represents the set of vertices connected to the hyperedge e j and ∈ represents the membership symbol; 2) Encode using the encoder Encoder(·): The encoder takes the gene expression matrix X0 output in step S1 as the initial layer and generates a low-dimensional representation of the spatial positions by aggregating the information of adjacent nodes within the same hyperedge. To adaptively learn the weights of the hyperedges, three layers of hypergraph attention layers are used; 3) Hypergraph attention layer: The information transfer mechanism on the hypergraph is more complex than that on the ordinary graph and includes two key stages: the hyperedge aggregation stage based on the attention mechanism and the vertex aggregation stage based on the attention mechanism; Hyperedge Aggregation Phase Based on Attention Mechanism: In this phase, the features of all connected vertices within the same hyperedge will be aggregated to generate the corresponding hyperedge embedding representation; among them, the hyperedge feature E t-1 of the (t-1)-th layer is calculated as follows: E t-1 = H T X t-1 where X t-1 represents the vertex feature matrix of layer t-1, H T represents the transpose of H; for each pair of hyperedges (e i , e j ), where represents the set of hyperedges adjacent to hyperedge e i , the attention coefficient a i between the i-th hyperedge e j and the j-th hyperedge e ij is calculated by introducing an attention mechanism: In the formula, represents the current layer feature vector of the hyperedge e i , represents the current layer feature vector of the hyperedge e j , represents the current layer feature vector of the hyperedge e i adjacent to the hyperedge e k , W a represents the learnable linear transformation weight matrix related to the attention coefficient a ij between hyperedges, represents the attention parameter vector related to the attention coefficient a ij between hyperedges, ∑ represents summation, LeakyReLU represents the leaky rectified linear unit function, ‖ represents the vector concatenation operation, exp represents the natural exponential function; the hyperedge feature matrix E t of the output of the t-th layer obtained by attention hyperedge aggregation can be calculated as follows: E t = AWH T X t-1 where A represents the attention weight matrix between hyperedges, H represents the incidence matrix of the hypergraph, and X t-1 represents the vertex feature matrix of the (t-1)-th layer; the vertex aggregation stage based on the attention mechanism generates vertex embeddings by integrating the information of connected hyperedges, and the vertex feature matrix X t of the t-th layer obtained by hyperedge aggregation can be calculated by the following formula: X t = HE t Vertex aggregation stage based on the attention mechanism: First, based on the attention mechanism, calculate the attention coefficient b i between the i-th hyperedge e j and the j-th vertex x ij : Where, V ei represents the set of vertices connected to the hyperedge e i , x k represents the k-th vertex belonging to V ei , represents the current layer feature vector of the hyperedge e i ; represents the current layer feature vector of the vertex x j ; represents the current layer feature vector of the vertex x k ; represents the attention parameter vector related to the attention coefficient b ij between the hyperedge and the vertex; Next, the output low-dimensional representation X e of the vertex is calculated in the following way: X e = σ(BHAWH T X t-1 ) where X e in which e indicates that the obtained vertex low-dimensional representation is a low-dimensional embedding, σ represents the activation function, and B represents the hyperedge-vertex attention weight matrix composed of b ij ; H T represents the transpose of the hypergraph incidence matrix, A represents the hyperedge-hyperedge attention weight matrix composed of a ij ; and X t-1 represents the vertex feature matrix at layer t-1; 4) Decode using the decoder Decoder(·): Use the low-dimensional representation X of the vertices of the spatial sites obtained in step 3) e as the input, and reconstruct the site features through a symmetric hypergraph attention network architecture. The specific implementation is as follows: Construct a decoder structure using a hypergraph attention network symmetric to the encoder. After being processed by three layers of hypergraph attention networks, the output of the last layer is used as the reconstructed vertex features; the output X t ' of the t-th layer in the decoder is obtained through the following transformation: wherein, σ represents an activation function, and X t ' represents the vertex feature matrix after reconstruction of the t-th layer of the decoder, H represents the hypergraph incidence matrix, and X' t-1 represents the vertex feature matrix after reconstruction of the (t-1)-th layer of the decoder, and W k-t+1 represents the learnable linear transformation weight matrix of the (k-t+1)-th layer of the decoder, where k represents the total number of layers of the decoder, and A k-t+1 represents the hyperedge-hyperedge attention weight matrix of the (k-t+1)-th layer of the decoder, and B k-t+1 represents the hyperedge-vertex attention weight matrix of the (k-t+1)-th layer of the decoder; 5) Train using the loss function: Optimize the autoencoder architecture by minimizing the following loss function: In the formula, denotes the summation operation over all N vertices in the dataset, X in represents the gene expression matrix of the input to the encoder, X out represents the vertex feature matrix after the final reconstruction by the decoder, ||·||2 denotes the L2 norm of a vector, where: X in = X0 X out = X' k where, X′ k represents the vertex feature matrix after reconstruction of the k-th layer of the decoder.
5. A method for dividing the organizational structure of spatial transcriptome data according to claim 1, characterized in that, In step S4, using the vertex low-dimensional representation X output in step S3 e as the input, perform clustering analysis using the mclust algorithm based on the Gaussian mixture model to obtain the clustering labels of each spatial site; Subsequently, align the clustering labels of each spatial site with the prior labels of the known tissue structures: First, calculate the average gene expression matrices of each clustering label and the prior labels of the known tissue structures, and construct a cost matrix through negative similarity measurement, where the higher the similarity of the expression matrices, the lower the matching cost, and the greater the difference, the higher the matching cost; Subsequently, based on this cost matrix, use the entropy-regularized Sinkhorn algorithm to solve the minimum transport strategy and assign the clustering labels to the most matching known tissue structure categories; Finally, divide each spatial site into the corresponding tissue structure category according to its clustering label and label matching result, and output a biologically interpretable and spatially topologically coherent tissue structure division result.
Citation Information
Cited By
In-vitro tissue space transcriptome data analysis method based on graph neural network
CN121838890A