PPI prediction method based on dynamic hypergraph
By dynamically adjusting the hypergraph structure and the number of super edges, and using attention mechanism and feature splicing to optimize the super edge features, the problem of unadjustable super edges in the existing methods is solved, the accuracy and efficiency of PPI prediction is improved, and the representation ability and information flow of the protein network are enhanced.
Patent Information
- Application Number
- CN202510400476.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-08-01
AI Technical Summary
The existing graph neural network methods cannot dynamically adjust the number and structure of super edges, resulting in the inability to fully capture the complex dynamic relationship between proteins, affecting the accuracy and efficiency of PPI prediction.
By dynamically adjusting the hypergraph structure and the number of hyperedges, the attention mechanism is used to calculate the correlation between the node and the hyperedge, the hyperedge features are updated by combining multi-layer perceptrons and feature stitching, the node features are optimized using Fourier transform and regularization terms, and the superedge number and structure are adaptively adjusted.
It improves the accuracy and efficiency of PPI prediction, avoids information loss, enhances the representation ability of the protein network, and ensures information mobility and robustness.
Smart Images

Figure CN120412701A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of PPI prediction, and particularly to a PPI prediction method based on a dynamic hypergraph. Background Art
[0002] Protein-protein interaction (PPI) is a key link in cell functions and biological processes, involving various biological activities such as signal transduction, cell cycle, and metabolic pathways. Therefore, understanding, identifying, and predicting PPI are of great significance in medical, pharmaceutical, and genetic research. However, PPI detection methods in the laboratory, such as yeast two-hybrid screening and mass spectrometry protein complex identification, are usually expensive and time-consuming, which limits the development of large-scale PPI prediction. Therefore, there is an urgent need to develop computational methods, especially deep learning techniques, to effectively predict unknown PPIs. In recent years, with the rapid development of artificial intelligence technology, methods based on graph neural networks (GNNs) have gradually become the mainstream in PPI prediction. Traditional PPI prediction methods mainly rely on machine learning-based feature engineering, and the performance of these methods is often limited when dealing with large-scale data. Graph neural networks (GNNs) can capture the complex relationships between proteins through deep learning of protein networks and provide more accurate prediction results than traditional methods.
[0003] However, most of the existing graph neural network methods are based on static graph structures and cannot dynamically adjust the connection relationships in the graph, thus unable to fully capture the complex dynamic relationships between proteins. To solve this problem, dynamic hypergraph neural networks (DHGNNs) have been proposed, aiming to capture more hidden relationships by adaptively optimizing the hypergraph structure. In a hypergraph, a hyperedge connects not just a pair of nodes but a group of nodes, which can better represent the higher-order relationships between proteins. Although DHGNNs have advantages in capturing implicit relationships, most of the existing DHGNNs cannot dynamically adjust the number of hyperedges, resulting in the inability to fully explore the potential of the underlying hypergraph structure. In other words, the existing methods cannot flexibly adjust the number and structure of hyperedges according to the needs of actual data, thus affecting the performance of the model. Summary of the Invention
[0004] The purpose of the present invention is to provide a PPI prediction method based on a dynamic hypergraph to solve the problems of fixed hypergraph structure and non-adjustable number of hyperedges in existing PPI prediction methods. By dynamically adjusting the hypergraph structure and the number of hyperedges, it is possible to more flexibly capture the complex interaction relationships between proteins and improve the accuracy and efficiency of PPI prediction.
[0005] The technical solution of the present invention is as follows:
[0006] A PPI prediction method based on a dynamic hypergraph, comprising the following steps:
[0007] Hyperedge Sampling: According to the protein interaction network G = (V, E), sample hyperedges using the learned hyperedge feature distribution P(X e ∣X v ) to obtain the initial hyperedge feature matrix X e ;
[0008] Hyperedge Feature Update: Map node features to the hyperedge feature space through the mapping function f():
[0009]
[0010] Adopt the attention mechanism to calculate the correlation between each node and the hyperedge, forming the hyperedge attention matrix A e , and the calculation formula for the attention coefficient is:
[0011]
[0012] where A e ∈R m×n is the attention matrix of the hyperedge, and α ei,vj is the attention coefficient between the hyperedge e i and the node v j ;
[0013] Aggregate the features of the top K n node features, and update the hyperedge features through a multi-layer perceptron and feature concatenation. Finally, optimize the model through backpropagation;
[0014] Hypergraph Construction: Use the attention mechanism to calculate the attention coefficients between nodes and hyperedges, and assign each node to the top K e hyperedges with the highest attention coefficients;
[0015] Hypergraph Convolution: Combine the Fourier transform and regularization terms to update node features through the weight matrix, add self-loop node features to enhance node information representation, and update node features through trainable parameters.
[0016] ?Furthermore, the optimization of the model through backpropagation is as follows:
[0017]
[0018] where MLP(·) represents the multi-layer perceptron, and concate(·) represents feature concatenation, represents that for each row of A e , retain the top k n values, and set the rest to zero.
[0019] Furthermore, the hypergraph construction specifically includes the following steps:
[0020] Use the hyperedge query q(X e ) to match the keywords of the nodes Different. First, use the node query to match the keywords of the hyperedges
[0021]
[0022] Put each node into the top k e hyperedges with the highest attention coefficients:
[0023]
[0024] Here means that for each row of A v , keep the top k e values and set the rest to zero.
[0025] Furthermore, the hypergraph construction is constrained by the supervised constraint and the supervised constraint function. The loss function in the supervised constraint and the supervised constraint function is:
[0026]
[0027] where C represents the set of node categories in the dataset, c ∈ C represents a category, and d(·) is the distance function, and d(h i , h j ) represents the minimum distance between h i and h j .
[0028] Furthermore, the hypergraph construction adjusts the number of hyperedges according to the saturation score S H of the hypergraph and optimizes the hypergraph structure; the saturation score S H of the hypergraph is:
[0029]
[0030] where E empty ={e∣e∈E, 「e = 0} is the set of empty hyperedges in the hypergraph, SH is the saturation score of the hypergraph H, that is, the proportion of non-empty hyperedges in all hyperedges;
[0031] The hyperparameters β ∈ [0, 1] and γ ∈ [0, 1] represent the lower and upper bounds of the saturation score respectively. After each iteration, adjust the sample number m according to the saturation score:
[0032]
[0033] Furthermore, the hyperedge sampling specifically includes the following steps:
[0034] According to the protein interaction network G=(V, E), learn the hyperedge feature distribution P(Xe|Xv) to generate hyperedge features;
[0035] Based on the learned hyperedge feature distribution, sample the hyperedges to obtain the initial hyperedge feature matrix X e ;
[0036] Use a trainable Gaussian distribution for hyperedge feature sampling, and ensure the trainability of the distribution through the reparameterization trick to improve the flexibility and generalization ability of the model.
[0037] Furthermore, the initial hyperedge feature matrix X e is obtained through the following steps:
[0038] The feature dimension of each hyperedge is independent and follows a trainable Gaussian distribution (X e ) i,j ~N(μ j , diag(σ j ))), where μ ∈ R 1×de and σ ∈ R 1×de ;
[0039] Randomly initialize the distribution and obtain the initial hyperedge feature matrix X through sampling m times e ∈ R m×de ;
[0040] Use the reparameterization trick to make μ and σ trainable, that is, sample Q ∈ R from N(0, 1) m×de , and then use the following formula to obtain X e :
[0041] (X e ) i = μ + σ ⊙ Q i ,
[0042] where ⊙ represents the Hadamard product, that is, element-wise multiplication.
[0043] Furthermore, the hypergraph convolution specifically includes the following steps:
[0044] Use the learned hyperedge features and the constructed hypergraph to update the features of the nodes. The hypergraph convolution formula is:
[0045]
[0046] where W and are both diagonal matrices; removing the regularization term, we get:
[0047]
[0048] Multiply the learned hyperedge features by the weight matrix Θ ∈ R de×d , and then distribute the hyperedge features to each relevant node according to the learned hypergraph structure;
[0049] After introducing self-loops, the formula is modified to:
[0050]
[0051] where, H′ ∈ R n×n is a diagonal event matrix, where each node v i belongs to only one hyperedge e with degree 1 i ;
[0052] Regard the features of the nodes as hyperedge features, that is The final convolution formula is:
[0053]
[0054] This application also includes a PPI prediction system based on a dynamic hypergraph, including: a protein hyperedge sampling module, a protein hyperedge feature update module, a protein hypergraph construction module, and a hypergraph convolution module, which are used to implement a PPI prediction method based on a dynamic hypergraph,
[0055] Protein hyperedge sampling module: According to the protein interaction network, sample the hyperedges using the learned hyperedge feature distribution to obtain an initial hyperedge feature matrix;
[0056] Protein hyperedge feature update module: Map the node features to the hyperedge feature space through a mapping function, calculate the correlation between each node and the hyperedge using an attention mechanism, aggregate the first K n node features, and update the hyperedge features through a multi-layer perceptron and feature splicing, and finally optimize the model through backpropagation;
[0057] Protein hypergraph construction module: Calculate the attention coefficient between the nodes and the hyperedges using an attention mechanism, and assign each node to the first K e hyperedges with the highest attention coefficients;
[0058] Hypergraph convolution module: Combine the Fourier transform and the regularization term, update the node features through the weight matrix, add self-loop node features to enhance the node information representation, and update the node features through trainable parameters.
[0059] This application also includes a computer-readable storage medium, on which program code is stored, which executes a PPI prediction method based on a dynamic hypergraph.
[0060] Compared with the existing technologies, the beneficial effects of the present invention are:
[0061] 1. Dynamic hypergraph structure optimization: This application can adaptively adjust the number of hyperedges and the hypergraph structure according to the learned hyperedge feature distribution, thereby effectively capturing the complex interactions between proteins and improving the expression ability of the model. Compared with the existing static hypergraph methods, this application avoids the information loss problem caused by fixed hyperedges and realizes more flexible hypergraph modeling.
[0062] 2. Adaptive update of hyperedge features: The attention mechanism is used to calculate the correlation between protein nodes and hyperedges, and the hyperedge features are optimized through a multi-layer perceptron (MLP) and feature splicing, enabling the hyperedges to adaptively adjust during the information transmission process at different levels and improving the representation ability of the protein network.
[0063] 3. Enhancing information exchange and avoiding isolated nodes: This application uses the attention mechanism to dynamically allocate hyperedges, enabling each node to be connected to the most relevant hyperedges, avoiding the problem of isolated nodes that may occur in traditional hypergraph methods, ensuring the information flow between proteins, and thus enhancing the effectiveness of PPI prediction.
[0064] 4. Fourier transform combined with regularized hypergraph convolution: This application further enhances the expression ability of node features through hypergraph convolution operations that combine Fourier transform and regularization terms, and ensures the integrity of node information through a self-loop mechanism, improving the robustness of PPI prediction. Brief Description of the Drawings
[0065] Figure 1 It is a schematic diagram of the architecture of this application.
[0066] Figure 2 It is a detailed grouped comparison chart of the number of parameters and training time of HG and AFTGAN under the same hardware resources.
[0067] Figure 3 It is a comparison chart of the F1 scores of HG and GNN-PPI under three partitions using SHS27k as the training set and string as the test set under the condition of a cross-dataset.
[0068] Figure 4 It is a comparison chart of the F1 scores of HG and GNN-PPI under three partitions using SHS148k as the training set and string as the test set under the condition of a cross-dataset. Detailed Implementation Manner
[0069] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0070] The features and performance of the present invention will be further described in detail below in conjunction with embodiments.
[0071] Please refer to Figures 1-4 , a PPI prediction method based on a dynamic hypergraph, as Figure 1 shown, fully considering the time-varying nature and complexity in the protein-protein interaction network, including the following steps:
[0072] Hyperedge sampling: According to the protein-protein interaction network G=(V, E), use the learned hyperedge feature distribution P(X e |X v ) to sample the hyperedges to obtain the initial hyperedge feature matrix X e ;
[0073] The hyperedge sampling specifically includes the following steps:
[0074] According to the protein-protein interaction network G=(V, E), learn the hyperedge feature distribution P(Xe|Xv) to generate hyperedge features;
[0075] Based on the learned hyperedge feature distribution, sample the hyperedges to obtain the initial hyperedge feature matrix X e ;
[0076] Use a trainable Gaussian distribution for hyperedge feature sampling and ensure the trainability of the distribution through the reparameterization trick to improve the flexibility and generalization ability of the model.
[0077] The initial hyperedge feature matrix X e is obtained through the following steps:
[0078] The feature dimension of each hyperedge is independent and follows a trainable Gaussian distribution (X e ) i,j ~N(μ j , diag(σ j ))), where μ ∈ R1×de and σ ∈ R 1×de ;
[0079] Randomly initialize the distribution and obtain the initial hyperedge feature matrix X by sampling m times e ∈ R m×de ;
[0080] Use the reparameterization trick to make μ and σ trainable, that is, sample Q ∈ R from N(0, 1) m×de , and then use the following formula to obtain X e :
[0081] (X e ) i = μ + σ ⊙ Q i ,
[0082] where ⊙ represents the Hadamard product, that is, element-wise multiplication.
[0083] The hyperedge feature matrix X e is obtained by sampling from a trainable Gaussian distribution and trained using the reparameterization trick to ensure the trainability of the distribution.
[0084] Hyperedge feature update: Map the node features to the hyperedge feature space through the mapping function f():
[0085]
[0086] Adopt the attention mechanism to calculate the correlation between each node and the hyperedge to form the hyperedge attention matrix A e , and the formula for the attention coefficient is:[[]]
[0087]
[0088] where A e ∈ R m×n is the attention matrix of the hyperedge, and α ei,vj is the attention coefficient between the hyperedge e i and the node v j ;
[0089] Aggregate the features of the top K n nodes and update the hyperedge features through a multi-layer perceptron and feature concatenation, and finally optimize the model through backpropagation;
[0090] Optimizing the model through backpropagation is:[[]]
[0091]
[0092] where MLP(·) represents the multi-layer perceptron, concate(·) represents feature concatenation, represents for Ae For each row, retain the top k n values and set the rest to zero. By this method, μ and σ are connected to the input feature X v so that backpropagation can be used to update them.
[0093] Hypergraph construction: Use the attention mechanism to calculate the attention coefficients between nodes and hyperedges, and assign each node to the top K e hyperedges with the highest attention coefficients;
[0094] The hypergraph construction specifically includes the following steps:
[0095] Use the hyperedge query q(X e ) to match the keywords of the nodes Differently, first, use the node query to match the keywords of the hyperedges
[0096]
[0097] This reverse operation is performed because it is necessary to ensure that there are no isolated nodes, as isolated nodes cannot exchange information through the constructed hypergraph, which may affect the performance of downstream tasks.
[0098] Put each node into the top k e hyperedges with the highest attention coefficients:
[0099]
[0100] Here, means that for each row of A v , retain the top k e values and set the rest to zero.
[0101] The hypergraph construction is constrained by the supervised constraint and the supervised constraint function, and the loss function in the supervised constraint and the supervised constraint function is:
[0102]
[0103] where C represents the set of node categories in the dataset, c ∈ C represents a category, and d(·) is the distance function, and d(h i , h j ) represents the minimum distance between h i and h j . Use the labeled nodes to obtain the supervised constraint, and use the unsupervised constraint function to improve the accuracy of the learned hypergraph structure and ensure the similarity of node features.
[0104] The hypergraph construction adjusts the number of samples according to the saturation score (SH) of the hypergraph in each iteration, dynamically adjusts the number of hyperedges, making the hypergraph structure more compact and effective; the saturation score S of the hypergraph H :
[0105]
[0106] where E empty ={e∣e∈E,「e = 0} is the set of empty hyperedges in the hypergraph, SH is the saturation score of the hypergraph H, that is, the ratio of non-empty hyperedges to all hyperedges;
[0107] The hyperparameters β∈[0,1] and γ∈[0,1] represent the lower and upper bounds of the saturation score respectively. After each iteration, the number of samples m is adjusted according to the saturation score:
[0108]
[0109] Hypergraph convolution: Combining Fourier transform and regularization terms, updating node features through a weight matrix, adding self-loop node features to enhance node information representation, and updating node features through trainable parameters.
[0110] The hypergraph convolution specifically includes the following steps:
[0111] Using the learned hyperedge features and the constructed hypergraph to update the features of nodes. The hypergraph convolution formula is:
[0112]
[0113] where W and are both diagonal matrices; removing the regularization term, we get:
[0114]
[0115] Multiplying the learned hyperedge features by the weight matrix Θ∈R de×d , and then distributing the hyperedge features to each relevant node according to the learned hypergraph structure; in each layer, this application will reconstruct H according to the hyperedges and node features of the current layer.
[0116] In a hypergraph, a self-loop is a hyperedge that contains only one node. If self-loops are not introduced, the representation of a node will only be affected by the features of neighbor nodes in the previous layer, losing its own features. After introducing self-loops, the formula is modified to:
[0117]
[0118] where H′∈R n×n is a diagonal event matrix, where each node vi Only belongs to a hyperedge e with degree 1 i ;
[0119] Regarding the features of nodes as hyperedge features, that is The final convolution formula is:
[0120]
[0121] Hypergraph convolution uses self-loops to update node features, combines the features of each node with the corresponding hyperedge features, and ensures the integrity of node features in each layer of convolution.
[0122] This application also includes a PPI prediction system based on a dynamic hypergraph, including: a protein hyperedge sampling module, a protein hyperedge feature update module, a protein hypergraph construction module, and a hypergraph convolution module, which are used to implement a PPI prediction method based on a dynamic hypergraph,
[0123] Protein hyperedge sampling module: According to the protein interaction network, sample hyperedges using the learned hyperedge feature distribution to obtain an initial hyperedge feature matrix;
[0124] Protein hyperedge feature update module: Map node features to the hyperedge feature space through a mapping function, calculate the correlation between each node and the hyperedge using an attention mechanism, aggregate the features of the top K n node features, and update the hyperedge features through a multi-layer perceptron and feature splicing, and finally optimize the model through backpropagation;
[0125] Protein hypergraph construction module: Calculate the attention coefficient between nodes and hyperedges using an attention mechanism, and assign each node to the top K e hyperedges with the highest attention coefficients;
[0126] Hypergraph convolution module: Combine Fourier transform and regularization terms, update node features through a weight matrix, add self-loop node features to enhance node information representation, and update node features through trainable parameters.
[0127] [[ID=3,6]]This application also includes a computer-readable storage medium, on which program code is stored to execute the above-mentioned PPI prediction method based on a dynamic hypergraph.
[0128] Experimental verification:
[0129] For the method of predicting PPI based on a dynamic hypergraph proposed in this application, it is abbreviated as "HG" in the experiment.
[0130] Datasets: Extensive experiments were conducted on three public PPI datasets, namely STRING, SHS27k, and SHS148k. The STRING dataset contains human PPI data from the STRING database, with a total of 1,150,830 entries, covering 14,952 proteins and 572,568 interactions. Each protein-protein interaction is annotated as at least one of the following seven types: Activation, Binding, Catalysis, Expression, Inhibition, Post-translational modification (abbreviated as Ptmod), and Reaction. "One interaction + one annotation" is called a PPI data entry. Proteins with more than 50 amino acids and a sequence identity lower than 40% were randomly selected from the human subset of the STRING database to generate two more challenging datasets, SHS27k and SHS148k. The SHS27k dataset contains 16,912 PPI entries, covering 1,663 proteins and 7,401 interactions. The SHS148k dataset contains 99,782 PPI entries, covering 5,082 proteins and 43,397 interactions. In addition, three partitioning algorithms, including Random, Breath-First Search (abbreviated as BFS), and Depth-First Search (abbreviated as DFS), were used to divide each dataset into a training set and a test set at a ratio of 8:2. To better evaluate the generalization ability, the test data was further divided into three subsets according to whether the two proteins have been seen in the training data: (1) BS: both have been seen; (2) ES: one of the proteins has been seen; (3) NS: neither has been seen. In practice, the BFS and DFS partitions are more challenging because their test data only contains the ES and NS subsets, and the test set obtained by BFS is very difficult for the model to predict interactions because it contains a larger proportion of unknown proteins. It is assumed that the dataset split using BFS is more in line with practical applications and may reflect the performance of the model in actual scenarios to a certain extent.
[0131] Evaluation Metrics and Hyperparameters:
[0132] All experiments were conducted using Python v3.8.3 and PyTorch v1.8.1. The experimental results are the average of three repeated experiments and were calculated on an NVIDIA RTX 3090 GPU with 24GB of memory. In the model, the protein node features obtained by Node2vec have a dimension of 256, and the hidden dimension of the classifier is set to 512. The number of layers of the GMB encoder is set to 5. The model was trained on the SHS27k dataset for 600 epochs with a batch size of 2056 and on the SHS148k dataset for 600 epochs with a batch size of 2056. The Adam optimizer was used with a learning rate (lr) of 0.01 and a weight decay of 5e-4. The decay patience value of the learning rate was set to 30, and the decay rate was set to 0.5. The evaluation metric for multi-label PPI prediction is Micro-F1. Micro-F1 is widely used in multi-label classification. Due to the extremely imbalanced different PPI types in the dataset used, micro-average F1 (Micro-F1) may be more suitable than macro-average F1 (macro-F1) for evaluating the performance of multi-label PPI type prediction. Micro-F1 is obtained by calculating the harmonic mean of the total precision and recall of all types.
[0133]
[0134]
[0135] Baseline: Currently, methods that can effectively predict PPI include sequence-based: DL-PPI, DNN-PPI, GNN-PPI, and structure-based: DSSGN-PPI, HIGH-PPI. Furthermore, some scholars have used unlabeled data for pre-training, including AFTGAN, which can predict PPI more accurately. A comparison with these methods was carried out, and the performance comprehensively exceeded that of the non-pre-trained methods and was at least comparable to that of the pre-trained methods. It is worth mentioning that pre-trained models consume a large amount of computing resources and time. The model of this application can achieve the effect of pre-training only by using the most basic PPI network.
[0136] Three important observations can be drawn from Table 1: (1) For models trained only on downstream labeled data, two structure-based models, HIGH-PPI and DSSGNN-PPI, outperform other sequence-based baselines, but these two methods still lag behind the HG method. (2) AFTGAN is a sequence-based model with performance lower than that of structure-based models. However, since AFTGAN adds ESM-1b embeddings containing a large amount of biological information as protein sequence features, the performance of AFTGAN is better than that of other sequence-based models but still lower than that of models using only the PPI network, further demonstrating the efficiency of the model in predicting novel PPI relationships. (3) For the structure-based model: HIGH-PPI, due to the strong relationship between protein interactions and protein structure, its performance shows a significant improvement compared to methods using only sequence information. However, introducing structural information inevitably brings a large computational resource burden. The method of this application uses only sequence information and can outperform methods using structural information in DFS and BFS partitioning.
[0137] Table 1
[0138]
[0139]
[0140] Generalization analysis: To evaluate the generalization ability of the method, a method of training using one dataset and then testing the performance of the model on another larger dataset, i.e., a heterologous training set - test set, was evaluated. For example, a model trained on the SHS27k dataset will encounter proteins with fewer amino acids in the STRING dataset. More importantly, different from the BFS and DFS partitioning on the SHS27k dataset, when migrating to STRING, more than 80% of the PPIs in the test data are in the NS subset. The performance of HG and GNNPPI in two domain migration settings was reported in Figure 3 and Figure 4 It can be seen that HG outperforms GNNPPI in all settings, especially when migrating from SHS148k to STRING.
[0141] Lightweight analysis: As shown in Figure 2 To evaluate the running time of the method compared to other methods, HG and AFTGAN were controlled to use the same hardware resources, and the number of parameters and running time of both were recorded. It can be seen that the model HG of this application is more lightweight and efficient than AFTGAN.
[0142] The above-described embodiments merely represent specific implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application.
Claims
1. A PPI prediction method based on a dynamic hypergraph, characterized in that It includes the following steps: Hyperedge Sampling: According to the protein-protein interaction network G = (V, E), use the learned hyperedge feature distribution P(X e |X v ) to sample hyperedges, obtaining the initial hyperedge feature matrix X e ; Hyperedge feature update: Map node features to the hyperedge feature space through the mapping function f(): Calculate the correlation between each node and the hyperedge using the attention mechanism to form the hyperedge attention matrix A e , and the calculation formula for the attention coefficient is as follows: Among them, A e ∈R m×n is the attention matrix of the hyperedge, and α ei,vj is the attention coefficient between the hyperedge e i and the node v j ; Aggregate the top K n node features, update the hyperedge features through a multi-layer perceptron and feature concatenation, and finally optimize the model through backpropagation; Hypergraph construction: Calculate the attention coefficients between nodes and hyperedges using the attention mechanism, and assign each node to the top K e hyperedges with the highest attention coefficients; Hypergraph convolution: Combine Fourier transform and regularization terms, update node features through the weight matrix, add self-loop node features to enhance node information representation, and update node features through trainable parameters.
2. The PPI prediction method based on a dynamic hypergraph according to claim 1, wherein The optimization of the model through backpropagation is as follows: Among them, MLP(·) represents a multi-layer perceptron, and concate(·) represents feature concatenation. It means for each row of A e keep the first k n values and set the rest to zero.
3. The PPI prediction method based on a dynamic hypergraph according to claim 1, wherein The specific construction of the hypergraph includes the following steps: Use the hyperedge query q(X e ) to match the keywords of the nodes Different. First, use the node query to match the keywords of the hyperedges Put each node into the top-k e hyperedges with the highest attention coefficients: Here means for each row of A v keep the first k e values and set the rest to zero.
4. A PPI prediction method based on a dynamic hypergraph according to claim 1 or 3, characterized in that The construction of the hypergraph is constrained by the supervised constraint and the supervised constraint function, and the loss function in the supervised constraint and the supervised constraint function is: Among them, C represents the set of node categories in the dataset, c ∈ C represents a category, and d(·) is the distance function, d(h i , h j ) represents the minimum distance between h i and h j .
5. The PPI prediction method based on a dynamic hypergraph according to claim 4, characterized in that, The hypergraph construction adjusts the number of hyperedges and optimizes the hypergraph structure according to the saturation score S of the hypergraph; the saturation score S of the hypergraph H , H is as follows: where E empty ={e|e∈E,「e = 0} is the set of empty hyperedges in the hypergraph, and SH is the saturation score of the hypergraph H, that is, the proportion of non-empty hyperedges among all hyperedges; The hyperparameters β∈[0,1] and γ∈[0,1] respectively represent the lower and upper bounds of the saturation score. After each iteration, the sample number m is adjusted according to the saturation score:
6. The PPI prediction method based on a dynamic hypergraph according to claim 1, wherein The specific hyperedge sampling includes the following steps: According to the protein interaction network G=(V,E), learn the hyperedge feature distribution P(Xe∣Xv) to generate hyperedge features; Sample hyperedges based on the learned hyperedge feature distribution to obtain the initial hyperedge feature matrix X e ; Use a trainable Gaussian distribution for hyperedge feature sampling, and ensure the trainability of the distribution through the reparameterization trick to improve the flexibility and generalization ability of the model.
7. A PPI prediction method based on a dynamic hypergraph according to claim 1 or 6, characterized in that The initial hyperedge feature matrix X e is obtained through the following steps: The feature dimension of each hyperedge is independent and follows a trainable Gaussian distribution (X e ) i,j ~N(μ j ,diag(σ j ))), where μ ∈ R 1×de and σ ∈ R 1×de ; Randomly initialize the distribution and obtain the initial hyperedge feature matrix X by sampling m times e ∈R m×de ; Use the reparameterization trick to make μ and σ trainable, that is, sample Q ∈ R from N(0, 1) m×de , and then use the following formula to obtain X e : (X e ) i = μ + σ ⊙ Q i , Where ⊙ represents the Hadamard product, that is, element-wise multiplication.
8. A PPI prediction method based on a dynamic hypergraph according to claim 1, characterized in that The specific hypergraph convolution includes the following steps: Use the learned hyperedge features and the constructed hypergraph to update the features of the nodes. The hypergraph convolution formula is: Among them W and are both diagonal matrices; removing the regularization term, we get: Multiply the learned hyperedge features by the weight matrix Θ ∈ R de×d , and then distribute the hyperedge features to each relevant node according to the learned hypergraph structure; After introducing the self-loop, the formula is modified to: Among them, G′ ∈ R n×n is a diagonal event matrix, where each node v i belongs to only one hyperedge e with degree 1 i ; Regarding the features of nodes as the features of hyperedges, that is The final convolution formula is:
9. A PPI prediction system based on a dynamic hypergraph, characterized in that, It includes: A protein hyperedge sampling module, a protein hyperedge feature update module, a protein hypergraph construction module, and a hypergraph convolution module, which are used to implement the method according to any one of claims 1 to 8. Protein hyperedge sampling module: According to the protein interaction network, sample hyperedges using the learned hyperedge feature distribution to obtain an initial hyperedge feature matrix; Protein hyperedge feature update module: Map node features to the hyperedge feature space through a mapping function, use the attention mechanism to calculate the correlation between each node and the hyperedge, aggregate the features of the top K n node features, and update the hyperedge features through a multi-layer perceptron and feature concatenation. Finally, optimize the model through backpropagation; Protein hypergraph construction module: Using the attention mechanism to calculate the attention coefficients between nodes and hyperedges, and assigning each node to the top K e hyperedges with the highest attention coefficients; Hypergraph convolution module: Combine Fourier transform and regularization terms, update node features through the weight matrix, add self-loop node features to enhance node information representation, and update node features through trainable parameters.
10. A computer-readable storage medium, characterized in that, It stores program code that executes the method according to any one of claims 1 to 8.