Histological image gene expression prediction method based on hypergraph neural network
Through the hypergraph neural network combined with node feature embedding and self-attention mechanism, the problem of ignoring the relationship between the macroscopic structure of tissues and the microscopic composition of cells in the existing technology is solved, and a higher-precision gene expression prediction is achieved.
Patent Information
- Application Number
- CN202510634875.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-22
AI Technical Summary
When integrating spatial transcriptome data, the prior art ignores the relationship between the macrostructure of tissues and the microscopic composition of cells in histopathological images, and lacks the utilization of image-based features, and lacks sufficient mining of complex relationships between nodes in tissue images, resulting in insufficient prediction accuracy.
The hypergraph neural network is combined with the node feature embedding method, and by introducing spatial point feature embedding into the hypergraph structure, the complex relationship learning between histological images and gene expression is strengthened, and the self-attention mechanism is used to optimize feature interactions to improve model training efficiency and prediction accuracy.
It significantly improves the accuracy of gene expression prediction, enhances the characterization ability of tissue microenvironment, and improves the feature expression ability, especially in long-term dependence modeling.
Smart Images

Figure CN120526421A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of gene expression prediction in bioinformatics, and in particular relates to a histological image gene expression prediction method based on a hypergraph neural network. Background Art
[0002] Over the past decade, single-cell transcriptomics (scRNA-seq) has gained significant traction in biomedical research, particularly in developmental biology, cancer, immunology, and neuroscience. However, most commercial scRNA-seq protocols require intact, viable cells recovered from tissues, which not only makes many cell types difficult to study but also destroys the spatial context used for analyzing cell identity and function. In recent years, the development of spatial transcriptomics has enabled researchers to analyze the spatial distribution of gene expression within tissue microenvironments, driving advances in high-dimensional transcriptional assessment techniques.
[0003] Spatially resolved transcriptomics (ST) is booming as an extension of single-cell RNA sequencing (scRNA-seq) technologies such as Slideseq, XYZeq, and 10x Visium. ST technology can analyze gene expression of the entire transcriptome at near single-cell resolution while retaining spatial location information, and usually has a matching hematoxylin and eosin (H&E) stained histological image (WSI) of the entire slide. Depending on the ST technology, the sequencing expression of each point contains two to dozens of cells. These powerful ST technologies have changed researchers' understanding of human heart development, amyotrophic lateral sclerosis, and Alzheimer's disease. However, integrating the unique properties of spatial transcriptome data to reveal spatially variable features and cell identity remains challenging.
[0004] Despite rapid advancements in ST technology, it has yet to be applied in large-scale studies due to its high cost. In contrast, WSI is cheaper and more readily available, typically obtained in hospitals. Recent research by Schmauch et al. demonstrates that WSI can be used to predict the expression of large numbers of genes. Their developed HE2RNA method robustly captures subtle structures in WSI, revealing tumor regions that are suggestive of the development of specific cancer types. Similarly, WSI can also be used to predict spatial gene expression.
[0005] Currently, some methods for predicting spatial gene expression have emerged, such as ST-Net and HisToGene, and have achieved certain results. However, these existing methods still have two major limitations: (i) they ignore the relationship between tissue macrostructure and cellular microstructure in histopathological images; (ii) they insufficiently utilize image-based features and lack sufficient exploration of the complex correlation relationships between nodes in tissue images. Summary of the Invention
[0006] This invention aims to address the issues raised in the background art by proposing a gene expression prediction method that integrates a hypergraph neural network with node feature embedding, aiming to expand upon existing models. By incorporating spatial point feature embedding into the hypergraph structure, this method enhances learning of the complex relationship between histological images and gene expression, thereby improving model training efficiency and prediction accuracy.
[0007] To achieve the purpose of the present invention, the present invention discloses a method for predicting gene expression in histological images based on a hypergraph neural network, comprising the following steps:
[0008] Step 1: Spatial point feature embedding: For a given histological image, first use the pre-trained large model to extract the node feature tensor, then use the cell classification large model to obtain the cell type information within the node and encode it into a tensor, and finally aggregate to obtain the hybrid feature;
[0009] Step 2: Construct hyperedges to model node associations: Construct multi-scale hyperedges based on node spatial distance and feature distance, and perform hyperedge convolution after merging.
[0010] Step 3: Optimize features based on the self-attention mechanism. Apply the self-attention mechanism to the feature tensor after hyperedge convolution to dynamically optimize the feature interaction between nodes and obtain the prediction results.
[0011] Furthermore, in step 1, the pre-trained UNI large model is used to extract the node feature tensor f in the image, and the cell type information is obtained and encoded through the HoverNet cell classification large model, and the mixed features of the feature dimension D are aggregated.
[0012] Furthermore, in step 2, a hyperedge is constructed based on the node spatial distance using the nearest neighbor node method. Based on the feature tensor f extracted in step 1, a hyperedge is constructed using the Euclidean distance method. The nearest N and M nodes are selected respectively, and the hyperedges are merged to form a hypergraph G. After the hyperedge convolution, the new feature F2 is obtained.
[0013] Furthermore, in step 3, the feature F2 obtained in step 2 is dynamically optimized through the self-attention mechanism to obtain the prediction result P.
[0014] Furthermore, step 1 is as follows:
[0015] Given a histological image and the coordinates of the nodes in the image, the node feature tensor f is extracted using the UNI pre-trained large model:
[0016] f=UNI(x in ) (1)
[0017] Obtain cell type information within the node through the cell classification model HoverNet:
[0018] h1,h2,h3,h4,h5,h6=HoverNet(x in ) (2)
[0019] Among them, h 1-6 Corresponding to the number of six cell types in the histological image; then, the cell type information is encoded into the required tensor:
[0020]
[0021] Finally, the two tensors are fused to obtain the mixed features:
[0022] F=Concat(x,y,dim=-1) (4)
[0023] Among them, x and y are the results of the dimension transformation of two tensors through the linear layer, and the dimension D of the mixed feature is a pre-set hyperparameter.
[0024] Furthermore, step 2 is as follows:
[0025] First, the nearest neighbor node method is used to construct hyperedges based on the spatial distance between nodes:
[0026]
[0027] Among them, node v i The coordinates of (x i ,y i ), node v j The coordinates of (x j ,y j ); Then, based on the features extracted in step 1, the Euclidean distance method is used to construct the hyperedge:
[0028]
[0029] The number of nearest nodes selected by the two methods are N and M, respectively, which are pre-set hyperparameters;
[0030] F2=HypergraphConv(F,G) (7)
[0031] After merging the hyperedges, a hypergraph G is formed, and the hyperedges are convolved to obtain the new feature F2.
[0032] Furthermore, step 3 is as follows:
[0033] First, the spatial position of each point is encoded as an embedding, which is then fused with the feature F2 obtained in step 2:
[0034] e m =F2+e x +e y(8)
[0035] Among them, e x , e y are the two-dimensional coordinates of each point after being encoded by the torch.nn.Embedding function; then, the fused features e m The multi-head attention layer in the input Transformer module learns the spatial dependencies of each light spot / image patch. The multi-head attention layer consists of multiple attention heads in a linear combination, as follows:
[0036] MultiHead(Q,K,V)=[head1,...,head c ]×W0 (9)
[0037] Among them, W0 is the weight matrix of the aggregated attention head, c is the number of heads; Q, K and V represent Query, Key and Value respectively.
[0038] Furthermore, the attention mechanism is defined as follows:
[0039] head i =Attention(QW i Q ,KW i K ,VW i V ) (10)
[0040]
[0041] Among them, W i Q ,W i K and W i V is the weight matrix, It is called Attention Map, and its shape is N×N; V is the value of the self-attention mechanism, where V=K=Q; the attention weights contributed by other points are stored in each column of the Attention Map, and the final output is the prediction result P.
[0042] Compared with the existing technology, the significant progress of the present invention lies in: 1) Improved prediction accuracy: By embedding node features to integrate the macroscopic structure of the tissue and the microscopic composition information of the cells, the ability to characterize the tissue microenvironment is enhanced. At the same time, the hypergraph neural network can model the complex multi-scale associations between nodes, improving the ability to express features. Experimental results show that the accuracy of the method of the present invention in gene expression prediction tasks is significantly higher than that of existing methods; 2) Self-attention mechanism optimization: The present invention captures global dependencies by introducing a self-attention mechanism, dynamically optimizes the feature interactions of key nodes, and compensates for the shortcomings of the hypergraph module in long-range dependency modeling, thereby generating more accurate gene expression prediction results.
[0043] In order to more clearly illustrate the functional characteristics and structural parameters of the present invention, further description is given below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0045] Figure 1 This is a schematic diagram of a model based on a hypergraph neural network; Figure 2 It is a schematic diagram of the PCC comparison results of each method; Figure 3 This is a schematic diagram of the RMSE comparison results of each method; Figure 4 This is a diagram showing the visualization results of the top five genes with -log10 (P value) in HER2+; Figure 5 Schematic diagram of the clustering results of gene expression predicted in 6 tissue sections in HER2+. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0047] like Figure 1As shown in the figure, a hypergraph neural network-based method for predicting gene expression in histological images includes the following steps: First, for a given histological image, a pre-trained large model is used to extract node feature tensors. Cell type information within the nodes is obtained through a cell classification large model, encoded into tensors, and finally aggregated to obtain mixed features. Next, multi-scale hyperedges are constructed, and corresponding hyperedges are constructed based on node spatial distance and feature distance. The resulting hyperedge groups are combined and input into hyperedge convolution along with the mixed features. Finally, a self-attention mechanism is used to dynamically optimize the feature interactions between nodes in the feature tensor after hyperedge convolution, thereby obtaining the final gene expression prediction results.
[0048] The input of the model of the present invention is a histological image, which contains the spatial structure and cell composition information of the tissue, providing basic features for gene expression prediction. At the same time, the node position coordinates are used to further characterize the relative relationship between different nodes in space, assisting in modeling spatial dependencies and local feature interactions. By comprehensively utilizing histological image information and combining node position coordinates for spatial constraints, the present invention can more accurately infer the complex associations between nodes, thereby improving the accuracy of gene expression prediction. The specific implementation method is as follows:
[0049] Step 1: Spatial point feature embedding
[0050] Given a histological image and the coordinates of the nodes in the image, the node feature tensor f is extracted using the UNI pre-trained large model:
[0051] f=UNI(x in ) (1)
[0052] Obtain cell type information within the node through the cell classification model HoverNet:
[0053] h1,h2,h3,h4,h5,h6=HoverNet(x in ) (2)
[0054] Among them, h 1-6 Corresponding to the number of six cell types in the histological image. Next, the cell type information is encoded into the required tensor:
[0055]
[0056] Finally, the two tensors are fused to obtain the mixed features:
[0057] F=Concat(x,y,dim=-1) (4)
[0058] Among them, x and y are the results of the dimension transformation of two tensors through the linear layer, and the dimension D of the mixed feature is a pre-set hyperparameter.
[0059] Step 2: Construct hyperedge modeling node association
[0060] First, the nearest neighbor node method is used to construct hyperedges based on the spatial distance between nodes:
[0061]
[0062] Among them, node v i The coordinates of (x i ,y i ), node v j The coordinates of (x j ,y j ). Then, based on the features extracted in step 1, the Euclidean distance method is used to construct hyperedges:
[0063]
[0064] The number of nearest nodes selected by the two methods is N and M, respectively, which are pre-set hyperparameters. After merging the hyperedges, a hypergraph G is formed, and the hyperedges are convolved to obtain a new feature F2:
[0065] F2=HypergraphConv(F,G) (7)
[0066] Step 3: Dynamically optimize features based on self-attention mechanism
[0067] First, the spatial position of each point is encoded as an embedding, which is then fused with the feature F2 obtained in step 2:
[0068] e m =F2+e x +e y (8)
[0069] Among them, e x , e y are the two-dimensional coordinates of each point after being encoded by the torch.nn.Embedding function. m The multi-head attention layer in the Transformer module is input to learn the spatial dependencies of each light spot / image patch. The multi-head attention layer consists of multiple attention heads in a linear combination as follows:
[0070] MultiHead(Q,K,V)=[head1,...,head c ]×W0 (9)
[0071] Where W0 is the weight matrix of the aggregated attention head, and c is the number of heads. Q, K, and V represent Query, Key, and Value, respectively. The attention mechanism is defined as follows:
[0072] head i =Attention(QW i Q ,KW i K ,VW i V ) (10)
[0073]
[0074] Among them, W i Q ,W i K and W i V is the weight matrix, This is called an Attention Map, and its shape is N×N. V is the value of the self-attention mechanism, where V = K = Q. The attention weights contributed by other points are stored in each column of the Attention Map, and the final prediction result P is output.
[0075] Example
[0076] Our method was validated using spatial transcriptome (ST) datasets from two tumor types: HER2-positive breast cancer (HER2+) and cutaneous squamous cell carcinoma (cSCC). Both datasets provide histological images, including H&E-stained whole-slide images (WSIs), spatially resolved gene expression matrices, and corresponding spatial coordinate maps. The HER2+ dataset contains 32 tissue sections (from seven patients, each containing at least 180 probes) obtained after quality screening from an initial set of 36 sections (from eight patients). The cSCC dataset contains 12 tissue sections (from four patients, each providing three sections) analyzed using the 10x Visium platform.
[0077] To evaluate model performance, we conducted a comprehensive evaluation of all models on the HER2+ dataset (32 tissue sections) and the cSCC dataset (12 tissue sections) using leave-one-out cross-validation. Figure 2 As shown, Figure 2 The results of the Pearson correlation coefficient (PCC) comparison of the various methods are presented. Specifically, on the HER2+ dataset, the mean and median PCC of the HRRGE model were approximately 2% higher than those of the second-best method, THItoGene. Similarly, on the cSCC dataset, HRRGE's PCC was approximately 3% higher than that of THItoGene. Figure 2Most comparison models performed poorly in sections A2–A6 and F1–F3, with PCC values consistently below 0.07. In contrast, HRRGE improved performance by approximately 2% in these sections. Notably, HRRGE achieved the most significant improvement over other models in sections C1–C6, with a PCC increase of approximately 3%. These results demonstrate that HRRGE outperforms existing methods in predicting gene expression from histological images.
[0078] To comprehensively evaluate the performance of HRRGE in gene expression prediction, we also compared the root mean square error (RMSE) of each model. RMSE measures the difference between the predicted value and the actual gene expression value by calculating the square root of the average square error between the predicted value and the true value. Therefore, the lower the RMSE, the better the model performance. Figure 3 As shown, Figure 3 The comparison results of the RMSE indicators of various models are shown. Obviously, HRRGE has the lowest RMSE, indicating that our proposed HRRGE model performs best in the gene expression prediction task.
[0079] To further evaluate the predicted gene expression, we ranked the genes based on the highest mean -log10 (P value) across all tissue sections in each dataset. The P value for each tissue section was calculated based on the correlation between the predicted and actual gene expression. In the HER2+ dataset, we visualized the top five genes: FN1, GANS, SCD, MYL12B, and FASN. The visualization results showed that the proposed HRRGE achieved the highest PCC for each gene (see Figure 4 ). In addition, these five highly ranked predictive genes have all been shown to be associated with breast cancer.
[0080] To evaluate the performance of spatial region detection in complete histological images, we applied K-means clustering to the predicted gene expression. It is worth noting that only the HER2+ dataset contains 6 tissue sections (B1, C1, D1, E1, F1 and G2) annotated by pathologists, so a direct comparative analysis can be performed. Figure 5 As shown, HRRGE achieved the highest average adjusted Rand Index (ARI) in these slices, about 5% higher than the second-ranked method THItoGene. Specifically, HRRGE achieved the highest ARI in slices B1, C1, D1, and G2, especially in slice B1, where its ARI increased by about 10% compared to Hist2ST. These results highlight the excellent accuracy of HRRGE in spatial region detection tasks.
[0081] In summary, the present invention uses a spatial point feature embedding method to efficiently represent node features in histological images and incorporates cell type information to enrich feature expression. Next, multi-scale hyperedges are constructed based on spatial distance and feature similarity, and complex relationships between nodes are modeled through hyperedge convolution. Finally, self-attention is used to dynamically optimize the features after hyperedge convolution, enhancing the ability to model long-range dependencies, thereby improving the accuracy of gene expression prediction and generating precise prediction results.
[0082] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0083] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for predicting gene expression in histological images based on a hypergraph neural network, characterized in that: The following steps are involved: Step 1: Spatial point feature embedding: For a given histological image, first use the pre-trained large model to extract the node feature tensor, then use the cell classification large model to obtain the cell type information within the node and encode it into a tensor, and finally aggregate to obtain the hybrid feature; Step 2: Construct hyperedges to model node associations: Construct multi-scale hyperedges based on node spatial distance and feature distance, and perform hyperedge convolution after merging. Step 3: Optimize features based on the self-attention mechanism. Apply the self-attention mechanism to the feature tensor after hyperedge convolution to dynamically optimize the feature interaction between nodes and obtain the prediction results.
2. The method for predicting gene expression in histological images based on a hypergraph neural network according to claim 1, characterized in that: In step 1, the pre-trained UNI large model is used to extract the node feature tensor f in the image, and the cell type information is obtained and encoded through the HoverNet cell classification large model, and the mixed features of the feature dimension D are aggregated.
3. The method for predicting gene expression in histological images based on a hypergraph neural network according to claim 1, characterized in that: In step 2, a hyperedge is constructed based on the node spatial distance using the nearest neighbor node method. Based on the feature tensor f extracted in step 1, a hyperedge is constructed using the Euclidean distance method. The nearest N and M nodes are selected respectively, and the hyperedges are merged to form a hypergraph G. After the hyperedge convolution, the new feature F2 is obtained.
4. The method for predicting gene expression in histological images based on a hypergraph neural network according to claim 1, characterized in that: In step 3, the feature F2 obtained in step 2 is dynamically optimized through the self-attention mechanism to obtain the prediction result P.
5. The method for predicting gene expression in histological images based on a hypergraph neural network according to claim 2, characterized in that: Step 1 is as follows: Given a histological image and the coordinates of the nodes in the image, the node feature tensor f is extracted using the UNI pre-trained large model: f=UNI(x in ) (1) Obtain cell type information within the node through the cell classification model HoverNet: h1,h2,h3,h4,h5,h6=HoverNet(x in ) (2) Among them, h 1-6 Corresponding to the number of six cell types in the histological image; then, the cell type information is encoded into the required tensor: Finally, the two tensors are fused to obtain the mixed features: F=Concat(x,y,dim=-1) (4) Among them, x and y are the results of the dimension transformation of two tensors through the linear layer, and the dimension D of the mixed feature is a pre-set hyperparameter.
6. The method for predicting gene expression in histological images based on a hypergraph neural network according to claim 3, characterized in that: Step 2 is as follows: First, the nearest neighbor node method is used to construct hyperedges based on the spatial distance between nodes: Among them, node v i The coordinates of (x i ,y i ), node v j The coordinates of (x j ,y j ); Then, based on the features extracted in step 1, the Euclidean distance method is used to construct the hyperedge: The number of nearest nodes selected by the two methods are N and M, respectively, which are pre-set hyperparameters; F2=HypergraphConv(F,G) (7) After merging the hyperedges, a hypergraph G is formed, and the hyperedges are convolved to obtain the new feature F2.
7. The method for predicting gene expression in histological images based on a hypergraph neural network according to claim 4, characterized in that: Step 3 is as follows: First, the spatial position of each point is encoded as an embedding, which is then fused with the feature F2 obtained in step 2: And m =F2+e x +e y (8) Among them, e x , e y are the two-dimensional coordinates of each point after being encoded by the torch.nn.Embedding function; then, the fused features e m The multi-head attention layer in the input Transformer module learns the spatial dependencies of each light spot / image patch. The multi-head attention layer consists of multiple attention heads in a linear combination, as follows: MultiHead(Q,K,V)=[head1,...,head c ]×W0 (9) Among them, W0 is the weight matrix of the aggregated attention head, c is the number of heads; Q, K and V represent Query, Key and Value respectively.
8. The method for predicting gene expression in histological images based on a hypergraph neural network according to claim 7, characterized in that: The attention mechanism is defined as follows: head i =Attention(QW i Q ,KW i K ,VW i V ) (10) Among them, W i Q ,W i K and W i V is the weight matrix, It is called Attention Map, and its shape is N×N; V is the value of the self-attention mechanism, where V=K=Q; the attention weights contributed by other points are stored in each column of the Attention Map, and the final output is the prediction result P.
Citation Information
Cited By
Space transcriptome gene expression prediction method and system, terminal and storage medium
CN121354663A