Rice secondary metabolite protein association prediction method and device
The heterogeneous map of rice secondary metabolites and proteins is constructed through the heterogeneous map neural network method, and the learning node embedding representation and prediction are solved, which solves the problem of difficulty and cost of detecting the association between secondary metabolites and proteins in traditional methods, achieving high-precision and low-cost prediction effects.
Patent Information
- Application Number
- CN202510125856.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional molecular biology experiments in identifying the association of secondary metabolites and proteins have problems of difficulty and high cost in detection, making it difficult to achieve large-scale low-cost predictions.
Using the heterogeneous graph neural network method, using rice secondary metabolites, proteins and functional annotations as nodes, a heterogeneous graph is constructed, and the node embedding representation is learned, and the learning of the embedded representation is guided through the neighbor comparison learning framework. Finally, the association between secondary metabolites and proteins is predicted based on edge representation.
It significantly improves the accuracy and reliability of the prediction of rice secondary metabolites and protein associations, and reduces time, money and labor costs.
Smart Images

Figure CN119943138A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of bioinformatics, and in particular to a method and device for predicting the association between rice secondary metabolites and proteins. Background Art
[0002] Secondary metabolites refer to small molecules that are synthesized through secondary metabolism after an organism grows to a certain stage, have complex molecular structures, have no obvious physiological functions for the organism, or are not necessary for the growth and reproduction of the organism. This type of metabolite is not directly related to the normal growth, development or reproduction of life. The lack of them will not cause the organism to die immediately, but in the long run, it may affect the survival rate, mating success rate or fertility of the organism. Secondary metabolites play an important role in plant defense against being eaten and resisting environmental pressure, and humans use secondary metabolites as medicines, condiments or recreational drugs. Rice secondary metabolites have an important influence on rice stress resistance, yield, and taste quality. Secondary metabolites-Protein Interaction (SMPI) refers to the interaction or mutual influence between secondary metabolites and proteins. The biosynthesis of secondary metabolites is based on primary metabolites as precursors and is regulated by primary metabolism. The secondary metabolic enzyme system is not very specific in terms of substrate requirements, and its synthesis process is often controlled by multiple genes. Therefore, studying the relationship between secondary metabolites and proteins can explore their biological synthesis pathways and provide a basis for improving rice quality.
[0003] There are three main technical difficulties in the traditional molecular biology experiment to identify the association between secondary metabolites and proteins: first, it is difficult to detect the association between secondary metabolites with low affinity and proteins. Second, it is difficult to detect the affinity of secondary metabolites without destroying the association between related secondary metabolites and proteins. Finally, some proteins are difficult to purify in vitro.
[0004] Therefore, the traditional molecular biology experimental method of identifying the association between secondary metabolites and proteins requires a lot of time, money and manpower, so large-scale and low-cost prediction of rice secondary metabolite protein association is an urgent problem to be solved. Summary of the invention
[0005] Based on this, it is necessary to provide a rice secondary metabolite protein association prediction method and device that can improve prediction accuracy and reduce costs in response to the above technical problems.
[0006] In a first aspect, the present application provides a method for predicting rice secondary metabolite protein association. The method comprises:
[0007] A heterogeneous graph was constructed with rice secondary metabolites, proteins and functional annotations as nodes, rice secondary metabolites and protein interactions, rice secondary metabolites and functional annotation associations, protein and functional annotation associations, rice secondary metabolites and rice secondary metabolites structural similarity, and protein and protein interactions as edges;
[0008] Based on heterogeneous graphs, we aggregate neighbor-level information and relationship-level information to learn node embedding representations, and add a neighbor comparison learning framework to guide the learning of node embedding representations.
[0009] Based on the node embedding representation, the edge representation is constructed, and the association between rice secondary metabolites and proteins is predicted based on the edge representation.
[0010] In one embodiment, learning a node embedding representation based on aggregating neighbor-level information and relationship-level information in a heterogeneous graph includes:
[0011] At the neighbor level, the attention mechanism is used to learn the importance of different neighbor nodes to a given node based on a given relationship in a heterogeneous graph, aggregate the information of different neighbor nodes based on importance, and learn the representation of a given node under a given relationship;
[0012] At the relational level, based on the representation of a given node under a given relation, an attention mechanism is used to learn the importance of different relations.
[0013] The node embedding representation is obtained by combining the representation of a given node under a given relationship and the importance of different relationships.
[0014] In one embodiment, a neighbor contrast learning framework is added to guide the learning of node embedding representations, including:
[0015] Rice secondary metabolites, proteins and functional annotations are used as anchors to obtain the corresponding heterogeneous graph neighbor contrast loss. The heterogeneous graph neighbor contrast loss corresponding to each anchor is combined to determine the total neighbor contrast loss. The total neighbor contrast loss is minimized to enhance the distinguishability of node embedding representation.
[0016] In one embodiment, obtaining the heterogeneous graph neighbor contrast loss includes:
[0017] The anchor point and its heterogeneous neighbor nodes are taken as positive pairs, and the non-heterogeneous neighbor nodes of the anchor point are taken as negative pairs. According to the positive pairs and negative pairs, the heterogeneous graph neighbor contrast loss is represented by the normalized temperature-scaled cross entropy loss.
[0018] In one embodiment, constructing edge representation based on node embedding representation includes:
[0019] The node embedding representations corresponding to rice secondary metabolites and proteins are connected using the average algorithm, L1 norm, L2 norm, Hadamard product and splicing operation methods respectively, and the connection effects are compared and at least one operation method is determined according to the connection effects to connect the node embedding representations corresponding to rice secondary metabolites and proteins to obtain edge representations.
[0020] In one embodiment, predicting the association between rice secondary metabolites and proteins based on edge representation includes:
[0021] The edge representation is input into a multi-layer perceptron to obtain the association prediction probability between rice secondary metabolites and proteins;
[0022] Among them, the multilayer perceptron adopts a cost-sensitive learning strategy to weight the positive and negative edges; the multilayer perceptron constructs the loss function of the multilayer perceptron through weighted binary cross entropy loss and heterogeneous graph neighbor contrast loss.
[0023] In a second aspect, the present application also provides a rice secondary metabolite protein association prediction device. The device comprises:
[0024] A heterogeneous graph construction module is used to construct a heterogeneous graph with rice secondary metabolites, proteins and functional annotations as nodes, rice secondary metabolites and protein interactions, rice secondary metabolites and functional annotation associations, protein and functional annotation associations, rice secondary metabolites and rice secondary metabolites structural similarity, and protein and protein interactions as edges;
[0025] The node representation learning module is used to aggregate neighbor-level information and relationship-level information based on heterogeneous graphs, learn node embedding representations, and add a neighbor comparison learning framework to guide the learning of node embedding representations;
[0026] The association prediction module is used to construct edge representation based on node embedding representation and predict the association between rice secondary metabolites and proteins based on edge representation.
[0027] In a third aspect, the present application further provides a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the rice secondary metabolite protein association prediction method when executing the computer program.
[0028] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above rice secondary metabolite protein association prediction method when executed by a processor.
[0029] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the steps in the above rice secondary metabolite protein association prediction method when executed by a processor.
[0030] The above rice secondary metabolite protein association prediction method and device use rice secondary metabolites, proteins and functional annotations as nodes, rice secondary metabolites and protein interactions, rice secondary metabolites and functional annotation associations, protein and functional annotation associations, rice secondary metabolites and rice secondary metabolites structural similarity, and protein and protein interactions as edges to construct a heterogeneous graph; based on the heterogeneous graph, information at the neighbor level and information at the relationship level are aggregated to learn node embedding representations, and a neighbor comparison learning framework is added to guide the learning of node embedding representations; based on the node embedding representations, edge representations are constructed, and the associations between rice secondary metabolites and proteins are predicted based on the edge representations. By introducing heterogeneous graph neural networks, the present invention can effectively integrate different types of biological data, capture complex multi-level interactions, significantly improve the accuracy and reliability of rice secondary metabolites and protein association predictions, and greatly reduce time, money and labor costs compared to traditional molecular biology experiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic diagram of a heterogeneous graph in one embodiment;
[0032] Figure 2 The figure is a schematic diagram of a method for predicting rice secondary metabolite protein association in one embodiment;
[0033] Figure 3 is a structural block diagram of a rice secondary metabolite protein association prediction device in one embodiment; DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0035] With the accumulation of data and the development of artificial intelligence, deep learning has received a lot of attention recently and has been successfully applied to a wide range of problems in various research fields of computer vision, natural language processing, medical image analysis and bioinformatics. Among them, graph neural networks, existing graph neural networks generally only model homogeneous networks, which contain only one type of node and one type of edge, that is, deep learning technology applied to graph structured data, is one of the hottest research topics in the current deep learning field. The present invention applies heterogeneous graph structures to bioinformatics to study the association between rice secondary metabolites and proteins.
[0036] The present application embodiment provides a method for predicting rice secondary metabolite protein association, comprising the following steps:
[0037] Step 102, constructing a heterogeneous graph with rice secondary metabolites, proteins and functional annotations as nodes, rice secondary metabolites and protein interactions, rice secondary metabolites and functional annotation associations, protein and functional annotation associations, rice secondary metabolites and rice secondary metabolites structural similarity, and protein and protein interactions as edges.
[0038] This example collects relevant data of rice secondary metabolites and proteins, including gene expression data, metabolomics data, protein interaction network data, etc., as well as related functional annotations, to construct a complete rice heterogeneous graph network.
[0039] Among them, the secondary metabolite node can be obtained by downloading the rice secondary metabolite database OryzaCyc 7.0. Specifically, OryzaCyc is a secondary metabolite metabolic pathway database that provides detailed data on rice biochemical reactions, metabolic pathways, enzymes, metabolites, and other information.
[0040] Protein nodes can be obtained by extracting OryzaBase, a comprehensive database covering all aspects of rice bioscience and dedicated to rice genomics and genetics research.
[0041] Functional annotation nodes can be obtained through the Gene Ontology annotation database download, which is an important resource for bioinformatics and biological research, aiming to describe the functions, processes and cellular components of genes and proteins.
[0042] The interaction between secondary metabolites and proteins was also obtained by downloading the OryzaCyc7.0 rice secondary metabolite database. In this database, each secondary metabolite corresponding to rice is assigned to a pathway, which has enzymes (EC ID) and different chemical reactions (Reaction ID) used to identify and classify rice metabolic pathways, and experimentally verified proteins are also assigned to the pathway. Therefore, the same EC ID or Reaction ID can be used as a judgment indicator to obtain the interaction relationship between secondary metabolites and proteins.
[0043] The association of secondary metabolites with functional annotations was obtained through the Gene Ontology database. In the Gene Ontology database, some EC IDs or Reaction IDs have GO annotation information. Through this annotation, the association information between secondary metabolites and functional annotations was obtained.
[0044] The structural similarity between secondary metabolites is determined by calculating the molecular fingerprint of the secondary metabolite small molecule based on the SMILES formula using the small molecule tool RDkit library, and then using the Morgan fingerprint method to calculate the Dace similarity coefficient of the small molecule structure between the secondary metabolites, thereby determining the structural similarity between the secondary metabolites. The SMILES formula of each secondary metabolite is extracted from the OryzaCyc 7.0 database.
[0045] Protein function annotation associations can be obtained from the OryzaBase database. Since the OryzaBase rice comprehensive database provides relevant information such as rice genome sequence, genotype data, genetic map, gene function annotation, etc., protein function annotation associations can be directly extracted from the database.
[0046] Protein-protein interactions are downloaded from the BioGrid protein interaction database. This database collects, archives and provides biologically important protein interaction data, including protein interaction information from a variety of biological systems. These interaction data come from different types of experimental techniques, such as yeast two-hybrid, immunoprecipitation, protein-protein interaction mass spectrometry analysis, etc.
[0047] Based on the above information, it can be used to integrate and construct a complete heterogeneous graph network, such as Figure 1 As shown, it contains three types of nodes and five types of edges. The three types of nodes are 368 secondary metabolite nodes, 12895 protein nodes and 4697 functional annotation nodes. The five types of edges are 6420 secondary metabolites and protein interactions, 1903 secondary metabolites and functional annotation associations, 63546 proteins and functional annotation associations, 86730 secondary metabolites and secondary metabolites structural similarities and 462 proteins and protein interactions.
[0048] Step 104 , based on the heterogeneous graph, aggregates the information at the neighbor level and the information at the relationship level, learns the node embedding representation, and adds a neighbor comparison learning framework to guide the learning of the node embedding representation.
[0049] A heterogeneous graph neural network model is used to integrate the properties of multiple heterogeneous nodes of the rice metabolome and the complex network topology information to learn a unified node embedding representation. In a heterogeneous graph, each node can have different neighbors under different relationships, different neighbors have different importance, and different relationships also have different importance. Two levels of aggregation methods are used here to aggregate information at the neighbor level and information at the relationship level to obtain the final node embedding.
[0050] In heterogeneous graphs, the relationships and attributes between nodes can be diverse, so traditional graph representation learning methods may face challenges. In response to this, this application introduces neighbor contrast learning, which is a self-supervised learning method. First, it finds neighbor nodes of different types through each node; this can be achieved through edges in heterogeneous graphs, and each type of edge connects different types of nodes; secondly, generate contrast samples, where the relevant network in the heterogeneous graph is used as a different contrast view, and the anchor point and its heterogeneous neighbors form a positive pair in the same view, and the anchor point and non-heterogeneous neighbors form a negative pair in different contrast views; finally, the node representation is learned through the neighbor contrast learning loss function, which can encourage heterogeneous neighbor node representations to be close to each other, while heterogeneous non-neighbor node representations are pushed away. The advantage of heterogeneous graph neighbor contrast learning is that it can well preserve the heterogeneous network topology, avoid destroying the relevant information of downstream tasks, and is beneficial to downstream prediction tasks.
[0051] Step 106, constructing edge representation based on the node embedding representation, and predicting the association between rice secondary metabolites and proteins based on the edge representation.
[0052] The methods for constructing edge representation based on node embedding representation mainly include concatenation, weighted summation, message passing, bilinear transformation, and edge type embedding. These methods have their own advantages and disadvantages, and the appropriate method can be selected according to specific task requirements. In heterogeneous graphs, the construction of edge representation needs to pay special attention to the diversification of edge types and node types to capture richer semantic information and improve the performance of downstream tasks.
[0053] Finally, an edge classifier was used to predict the interactions between rice secondary metabolites and proteins based on edge representation.
[0054] This example introduces a heterogeneous graph neural network, which can effectively integrate different types of biological data, capture complex multi-level interactions, and significantly improve the accuracy and reliability of rice secondary metabolites and protein association prediction. In addition, adding annotations as nodes to the heterogeneous graph can more comprehensively capture and utilize the information in the graph, improving the representation ability and application effect of the graph.
[0055] In one embodiment, in step 102, based on the aggregation of neighbor-level information and relationship-level information in the heterogeneous graph, learning the node embedding representation includes: at the neighbor level, using an attention mechanism to learn the importance of different neighbor nodes to a given node based on a given relationship in the heterogeneous graph, aggregating information of different neighbor nodes based on importance, and learning the representation of a given node under a given relationship; at the relationship level, based on the representation of a given node under a given relationship, using an attention mechanism to learn the importance of different relationships; combining the representation of a given node under a given relationship and the importance of different relationships to obtain a node embedding representation.
[0056] Specifically, for the aggregation at the neighbor level, taking the rice secondary metabolite node learning in the heterogeneous graph as an example, firstly given a node i, node j is node i through the given relationship Φ The neighbor nodes of the link are then connected using the attention mechanism to automatically learn the given relationship in the heterogeneous network. Φ The importance of different neighbors to node i is obtained by aggregating the information of different neighbors based on the importance to learn the relationship between node i and Φ The following expression is as follows:
[0057]
[0058] in Indicates the importance of node j to node i; f i ,f j Represents the initial features of nodes i and j; att neighbor Represents the attention mechanism method based on neighbor-level aggregation. After obtaining the importance between node pairs based on a given relationship Φ, it is normalized using the softmax method to obtain the attention weight coefficient The formula can be expressed as:
[0059]
[0060] in, represents all neighbors of node i under a given relationship Φ; || represents the concatenation operation of vectors; is a trainable weight vector based on a given relationship Φ. Next, the secondary metabolite node i is embedded in the given relationship Φ. It is obtained by weighting the neighbor's feature vector and the corresponding attention coefficient, which can be expressed as:
[0061]
[0062] For the aggregation at the relational level, in order to obtain a more comprehensive node representation learning, we need to target different relations {Φ1,Φ2,…,Φ t}Next, learn node representation Next, in order to better understand the importance of each relationship, we can use nonlinear transformation to process the embedding representation of a specific relationship, which can be expressed as follows:
[0063]
[0064] Among them, att relation Attention mechanism representing relation-level aggregation; Indicates the importance of each relationship Φ; V mis all secondary metabolite nodes i; Ψ is a trainable weight vector that can be used to measure the similarity between embedded representations under multiple relationships; W refers to the weight matrix; Represents the secondary metabolite node i in the relationship Φ at the neighbor level k The embedding representation under Φ is obtained by normalizing the results after learning the importance of node representation under different relations. k Weight And perform weighted aggregation to finally obtain a more comprehensive secondary metabolite node representation:
[0065]
[0066] M i represents the final embedding representation of secondary metabolite node i. i By integrating the representations of different relationships of nodes, the nodes can be described more comprehensively. i ) and functional annotation nodes (G i The final node embedding representation of ) will be obtained in the same way, and the formula is as follows:
[0067]
[0068] This embodiment uses a hierarchical attention mechanism to allocate attention weights to node types and edge types respectively, so as to better capture the complex relationships in heterogeneous graphs.
[0069] In one embodiment, a neighbor contrast learning framework is added to guide the learning of node embedding representation, including: using rice secondary metabolites, proteins and functional annotations as anchor points respectively to obtain the corresponding heterogeneous graph neighbor contrast loss; integrating the heterogeneous graph neighbor contrast loss corresponding to each anchor point to determine the total neighbor contrast loss; minimizing the total neighbor contrast loss to enhance the distinguishability of node embedding representation.
[0070] Taking protein nodes as anchors as an example, a protein node is first given as an anchor, where the protein anchor and its heterogeneous neighbors of protein nodes, secondary metabolite nodes, and functional annotation nodes are taken as positive pairs, and its heterogeneous non-neighbors are taken as negative pairs. Based on the given positive and negative pairs, the heterogeneous graph neighbor comparison loss function is represented by the normalized temperature-scale cross entropy loss. Among them, heterogeneous neighbors refer to neighbor nodes that are connected to the target node through different types of edges and have inconsistent node types, and non-heterogeneous neighbors refer to neighbor nodes that are connected to the target node through the same type of edges and have consistent node types.
[0071] Specifically, the heterogeneous graph neighbor contrast loss function formula is expressed as:
[0072]
[0073] in, τ is the temperature coefficient; the positive and negative pairs can be decomposed into formula (8):
[0074]
[0075] Where N represents heterogeneous neighbor nodes.
[0076] Similarly, according to formula (7), the heterogeneous graph neighbor comparison loss function can be calculated when the secondary metabolite node and the functional annotation node are used as anchor points. Finally, the heterogeneous graph neighbor comparison loss functions corresponding to all anchor points are summarized to obtain the total neighbor comparison loss, which is defined as:
[0077]
[0078] Where |V| is the total number of corresponding nodes. Minimizing formula (9) will maximize the consistency between positive pairs and minimize the consistency of negative pairs. It can bring heterogeneous neighbor node representations close to each other, push heterogeneous non-neighbor node representations apart, and map them in different contrasting views in the heterogeneous graph, so that the heterogeneous network topology can be well preserved in the end.
[0079] In one embodiment, based on the node embedding representation, constructing the edge representation includes: respectively connecting the node embedding representations corresponding to the rice secondary metabolites and proteins using an average algorithm, an L1 norm, an L2 norm, a Hadamard product, and a splicing operation method, comparing the connection effects and determining at least one operation method to connect the node embedding representations corresponding to the rice secondary metabolites and proteins according to the connection effects, and obtaining the edge representation.
[0080] Among them, the average algorithm, L1 norm, L2 norm, Hadamard product and splicing operation method are respectively expressed as:
[0081]
[0082] in, represents the edge v connecting secondary metabolite node i and protein node j ij The characteristic representation vector of is, and the symbol ⊙ represents the product of the corresponding elements of two vectors.
[0083] Based on the learned final embedding representations of secondary metabolite nodes and protein nodes, five different ways of constructing edge representations were compared, and the best edge representation method was selected to predict rice secondary metabolite-protein associations.
[0084] In one embodiment, predicting the association between rice secondary metabolites and proteins based on edge representation includes: inputting the edge representation into a multilayer perceptron to obtain the predicted probability of the association between rice secondary metabolites and proteins; wherein the multilayer perceptron adopts a cost-sensitive learning strategy to weight positive edges and negative edges; the multilayer perceptron constructs a loss function of the multilayer perceptron through weighted binary cross entropy loss and heterogeneous graph neighbor contrast loss.
[0085] Based on edge representation, a multilayer perceptron (MPL) can be used as an edge classifier to predict the interaction between rice secondary metabolites and proteins. The edge v connecting node i and node j ij The feature representation vector is input into a multi-layer perceptron MPL, and finally a prediction probability between 0 and 1 is output:
[0086]
[0087] In addition, it should be pointed out that in the rice secondary metabolite and protein association network, the number of associated node pairs (positive edges) is much less than the number of unassociated node pairs (negative edges). For this unbalanced binary classification problem, a cost-sensitive learning strategy will be adopted to give higher weights to positive edges in the loss function, so that the classifier can focus more on accurately identifying positive edges with interactions. Finally, by combining the weighted binary cross entropy loss function L wbce and heterogeneous neighbor contrast loss function L c The loss function L of the model is obtained, which assigns a larger weight to the positive samples and a smaller weight to the negative samples. In the experiment, the ratio of the weight of the positive sample to the weight of the negative sample is set to 10. The λ is used as a weight hyperparameter to balance the heterogeneous contrast loss function:
[0088] L=L wbce +λL c (12)
[0089] In one embodiment, the method also includes evaluating the performance of the prediction model, and it is crucial to select appropriate evaluation indicators. Appropriate indicators can effectively reflect the performance of the model. The level of the evaluation indicator can effectively analyze the model and help to correct and improve the model. In order to better evaluate the performance of the model, AUC_ROC (area under the receiver operating characteristic curve), AUC_PR (area under the precision-recall curve), accuracy, Matthews correlation coefficient (MCC), and F1 value are selected as evaluation indicators of the model.
[0090] In one embodiment, a multi-task learning mechanism can be used in the prediction model to predict multiple related tasks at the same time, further improving the generalization ability and prediction effect of the model.
[0091] In one embodiment, Figure 2 As shown, a rice secondary metabolite protein association prediction method is provided, comprising:
[0092] (1) Data preprocessing
[0093] Collect relevant data on rice secondary metabolites and proteins, including gene expression data, metabolomics data, protein interaction network data, etc., and preprocess the data, including data cleaning, normalization, missing value filling, etc.
[0094] (2) Constructing a heterogeneous graph
[0095] Rice secondary metabolites and proteins are regarded as nodes in the graph, and a heterogeneous graph is constructed, which contains multiple types of nodes (such as secondary metabolites, proteins) and multiple types of edges. Different features and weights are assigned to each node and edge type to reflect its importance in the network.
[0096] (3) Establishing a heterogeneous graph neural network model
[0097] Heterogeneous graph neural networks are used to capture the complex interactive relationships in the graph. Heterogeneous graph neural networks can simultaneously consider the heterogeneity of node types and edge types, aggregate the features of neighbor nodes through a multi-head attention mechanism, and learn node representations.
[0098] (4) Heterogeneous Neighbor Comparative Learning
[0099] In order to better preserve the heterogeneous network topology, a neighbor contrastive learning framework is added to guide node embedding learning.
[0100] (5) Association prediction
[0101] Based on the above feature extraction, a prediction layer is introduced to predict the association between rice secondary metabolites and proteins. Based on edge representation through edge classification loss function, multi-layer perceptron is used as edge classifier to predict the interaction between rice secondary metabolites and proteins.
[0102] (6) Model evaluation
[0103] First, the model was validated using a real dataset to evaluate its prediction accuracy. Secondly, it was compared with the direct application of existing graph neural network methods to demonstrate the advantages of the method of the present invention in large-scale prediction of rice secondary metabolites and protein associations. Finally, its effectiveness was demonstrated in a specific case analysis.
[0104] The present invention studies the prediction performance of the model with different training ratios, that is, setting the model training ratios at 10%, 30%, 50% and 90%. For the selection of these training set ratios, in fact, 10%, 30%, 50% and 90% of the edges in the matrix of secondary metabolites and protein interactions in the five input matrices are randomly extracted, and the other matrices remain unchanged. The four training ratios are randomly split five times, and finally the prediction results of different training set ratios are compared.
[0105] It can be observed from Table 1 that the prediction method of the present invention is significantly higher than 10% and 30% when the training ratio is 50% and 90%. For example, when the training ratio is 90% and the Concatenate method is used to construct the edge representation, the highest value is reached. The AUC_ROC, AUC_PR, Accuracy, F1 and MCC of the model are 98.94±0.06(%), 95.61±0.15(%), 98.29±0.05(%), 90.56±0.25(%), 89.62±0.27(%), which are significantly higher than the prediction results under other training ratios. Compared with the training ratio of 0.3, the five evaluation indicators are increased by 1.53%, 4.75%, 0.89%, 5.3% and 5.73%, respectively, indicating that the higher the training ratio, the better the evaluation indicators. As for the prediction results of the training set with a smaller proportion, a certain degree of accuracy can be achieved because the four other matrix-related edges in the training set (secondary metabolite functional annotation association, secondary metabolite-to-secondary metabolite structural similarity, protein functional annotation association, and protein-to-protein interaction edges) are retained.
[0106] Table 1 Comparison of model results under different training ratios
[0107]
[0108] Table 2 shows the comparison results of five edge representation construction methods. The GCN, HAN and OryHNCGAT models are used to evaluate the effects of the edge representation construction methods from the perspectives of AUC_ROC, AUC_PR, Accuracy, F1 and MCC.
[0109] Table 2 Comparison of experimental results of baseline models
[0110]
[0111]
[0112] In one embodiment, proteins that interact with the secondary metabolite naringenin are used as case studies for this experiment. The present invention is based on an RTXA5000 (24GB) and is implemented in Python. The backend uses the PyTorch framework. The Adam optimizer and the binary cross entropy loss function are used for model training during model training. The loss function automatically processes operations such as the sigmoid function conversion and logarithmic transformation of the probability value. Through back propagation and optimization algorithms, the parameters of the model will be adjusted to minimize the loss function, thereby improving the prediction performance of the model. The initial learning rate (Lr) is set to 0.005 and the weight decay (Lamb) is set to 1e-07; in order to weaken the influence of overfitting on the prediction effect of the model, the dropout rate (Dropout) is 0.5; in order to improve the prediction ability of the model, the mapping dimension (Hid_dim) is set to 64; in addition, in order to fully train the parameters of the model and finally reach a stable state, the training batch (Epoch) is set to 500.
[0113] The optimal hyperparameter combination of the above model was used, and the known associations between rice secondary metabolites and proteins were selected as the training set, and the unknown associations were selected as the candidate set. The model provided the predicted probability of the interaction between rice secondary metabolites and all proteins, and finally ranked these proteins according to the predicted probability; the top ten proteins representing the most likely interaction with the secondary metabolite naringenin were extracted, and the interaction between the secondary metabolites and proteins was verified by literature.
[0114] Table 3 The top ten proteins predicted by the model to be most likely to interact with the secondary metabolite naringenin
[0115]
[0116] As shown in Table 3, it is shown that naringenin may interact with the proteins predicted by the model, such as Os12g0600200, Os10g0579400, Os03g0577000, Os08g0446400, Os03g0116400 and Os05g0156800. Naringenin is a naturally occurring flavanone (flavonoid) that has a variety of biological activities, such as anti-diabetic, anti-atherosclerotic, anti-depressant, immunomodulatory, anti-tumor, anti-inflammatory, DNA protection, hypolipidemic, antioxidant, peroxisome proliferator-activated receptor (PPAR) activator and memory improvement. As a known antioxidant compound, naringenin administration showed neuronal protection and prevented ubiquitin accumulation in hypoxia treatment, and ubiquitinated proteins were significantly downregulated, which means that the model-predicted Os12g0600200 may interact with naringenin; at the same time, WRKY1 in the flavonoid synthesis pathway regulates the expression of related genes, thereby affecting the synthesis of flavonoids, including naringenin; in addition, studies have found that specific methyltransferases such as PfOMT3 can convert naringenin into sakurain, and participate in cellular stress response by expressing specific ribosomal proteins, thereby increasing the strain's tolerance to naringenin and increasing the production of sakurain; literature studies have shown that naringenin can inhibit serine / threonine protein phosphatase activity when used as an active drug for the liver in the medical field. In addition, when mitochondria are hydrolyzed and cleaved by rhomboid-like protein (PARL), naringenin can stabilize the mitochondrial membrane potential and prevent the production of reactive oxygen species, playing an important role in protecting neurons from damage. This indicates that Os03g0116400 predicted by the model may interact with naringenin. Finally, receptor-like protein kinase may affect the synthesis of naringenin and the regulation of anthocyanin biosynthesis pathway by regulating the activity of transcription factors, which also verifies that Os05g0156800 predicted by the model may interact with naringenin.
[0117] OryzaCyc, the most comprehensive rice metabolism database, only includes very few associations between secondary metabolites and proteins. Inspired by the successful application of heterogeneous graph neural networks in various fields, in order to fill the deficiencies of existing databases of rice secondary metabolites and protein associations, and to significantly reduce manpower, material resources and time costs, the present invention proposes a rice secondary metabolite protein association prediction method based on graph neural networks. This method first uses an attention mechanism to learn a unified node representation through a rice heterogeneous graph. Then, in order to better retain the heterogeneous network topology, a neighbor comparison learning framework is added to guide node embedding learning. Through model evaluation testing, this method is superior to the direct application of existing graph neural network methods, and its effectiveness has been successfully demonstrated in a specific case analysis.
[0118] In addition, the heterogeneous graph neural network framework has made improvements on the traditional framework, including:
[0119] (1) Multi-head attention mechanism: The multi-head attention mechanism is introduced into the heterogeneous graph neural network, which enables the model to learn the feature representation of nodes from multiple perspectives and enhance the expressiveness of the model.
[0120] (2) Multi-layer feature extraction: Design a multi-layer heterogeneous graph neural network model to extract high-order features of nodes layer by layer, thereby capturing global and local interactions and improving the model's predictive ability.
[0121] (3) Neighbor contrastive learning on heterogeneous graphs: In order to better preserve the heterogeneous network topology, a neighbor contrastive learning framework is added to guide node embedding learning.
[0122] The heterogeneous graph neural network framework integrates various types of biological data (such as gene expression data, metabolomics data, protein interaction network data) into a unified heterogeneous graph, and takes into account the heterogeneity between different data types; it can simultaneously capture the heterogeneity of node types and edge types, effectively extract high-order features of nodes, and capture complex multi-level interactive relationships.
[0123] On the basis of the above multi-layer feature extraction, the prediction layer is introduced, and the machine learning method is used to make accurate association predictions. At the same time, the multi-task learning mechanism is introduced into the model to support the simultaneous prediction of multiple association tasks, further improving the generalization ability and prediction effect of the model.
[0124] The present invention also proposes a set of efficient data preprocessing methods, including data cleaning, normalization, missing value filling, etc., to ensure the high quality and consistency of data, providing a reliable foundation for subsequent feature extraction and model training.
[0125] The present invention also uses methods such as cross-validation and grid search to optimize the hyperparameters of the model and improve the generalization ability and predictive performance of the model.
[0126] In summary, the beneficial effects of the present invention include:
[0127] (1) Improved prediction accuracy: Through the multi-head attention mechanism and multi-layer feature extraction of heterogeneous graph neural networks, the model can more accurately identify and predict the associations between rice secondary metabolites and proteins.
[0128] (2) Capturing complex interactions: Heterogeneous graph neural networks can simultaneously consider the heterogeneity of nodes and edges, and better capture the complex interactions between secondary metabolites and proteins.
[0129] (3) Model generalization ability: Through cross-validation and grid search optimization, the generalization ability of the model is improved, and it can maintain stable prediction performance on different data sets.
[0130] (4) Broad application prospects: The method of the present invention can be applied not only to rice, but also to other biological systems for predicting the association between secondary metabolites and proteins, and has broad application prospects.
[0131] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0132] Based on the same inventive concept, the present application embodiment also provides a rice secondary metabolite protein association prediction device for implementing the rice secondary metabolite protein association prediction method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations of one or more rice secondary metabolite protein association prediction device embodiments provided below can refer to the limitations of the rice secondary metabolite protein association prediction method above, and will not be repeated here.
[0133] In one embodiment, Figure 3 As shown, a rice secondary metabolite protein association prediction device is provided, comprising:
[0134] A heterogeneous graph construction module is used to construct a heterogeneous graph with rice secondary metabolites, proteins and functional annotations as nodes, rice secondary metabolites and protein interactions, rice secondary metabolites and functional annotation associations, protein and functional annotation associations, rice secondary metabolites and rice secondary metabolites structural similarity, and protein and protein interactions as edges;
[0135] The node representation learning module is used to aggregate neighbor-level information and relationship-level information based on heterogeneous graphs, learn node embedding representations, and add a neighbor comparison learning framework to guide the learning of node embedding representations;
[0136] The association prediction module is used to construct edge representation based on node embedding representation and predict the association between rice secondary metabolites and proteins based on edge representation.
[0137] Each module in the above rice secondary metabolite protein association prediction device can be implemented in whole or in part by software, hardware and a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0138] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in all the above method embodiments when executing the computer program.
[0139] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in all the above method embodiments are implemented.
[0140] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in all the above method embodiments when executed by a processor.
[0141] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0142] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0143] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0144] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for predicting the association between rice secondary metabolites and proteins, characterized in that: The method comprises: A heterogeneous graph was constructed with rice secondary metabolites, proteins and functional annotations as nodes, rice secondary metabolites and protein interactions, rice secondary metabolites and functional annotation associations, protein and functional annotation associations, rice secondary metabolites and rice secondary metabolites structural similarity, and protein and protein interactions as edges; Aggregating neighbor-level information and relationship-level information based on the heterogeneous graph, learning node embedding representation, and adding a neighbor comparison learning framework to guide the learning of the node embedding representation; Based on the node embedding representation, an edge representation is constructed, and based on the edge representation, the association between rice secondary metabolites and proteins is predicted.
2. The method according to claim 1, characterized in that The learning of node embedding representation based on aggregating neighbor-level information and relationship-level information of the heterogeneous graph includes: At the neighbor level, an attention mechanism is used to learn the importance of different neighbor nodes to a given node based on a given relationship in the heterogeneous graph, and information of different neighbor nodes is aggregated based on the importance to learn the representation of the given node under the given relationship; At the relationship level, based on the representation of the given node under the given relationship, an attention mechanism is used to learn the importance of different relationships; The node embedding representation is obtained by combining the representation of the given node under the given relationship and the importance of different relationships.
3. The method according to claim 2, characterized in that The adding of the neighbor comparison learning framework to guide the learning of the node embedding representation includes: Rice secondary metabolites, proteins and functional annotations are respectively used as anchor points to obtain the corresponding heterogeneous graph neighbor contrast losses; the heterogeneous graph neighbor contrast losses corresponding to each anchor point are integrated to determine the total neighbor contrast loss; the total neighbor contrast loss is minimized to enhance the distinguishability of the node embedding representation.
4. The method according to claim 3, characterized in that Obtaining the heterogeneous graph neighbor contrast loss includes: The anchor point and the heterogeneous neighbor nodes of the anchor point are taken as positive pairs, and the non-heterogeneous neighbor nodes of the anchor point are taken as negative pairs; according to the positive pairs and the negative pairs, the heterogeneous graph neighbor contrast loss is represented by a normalized temperature scale cross entropy loss.
5. The method according to claim 1, characterized in that The constructing edge representation based on the node embedding representation comprises: The node embedding representations corresponding to the rice secondary metabolites and proteins are connected by using an average algorithm, an L1 norm, an L2 norm, a Hadamard product and a splicing operation method respectively, and the connection effects are compared and at least one operation method is determined according to the connection effects to connect the node embedding representations corresponding to the rice secondary metabolites and proteins to obtain the edge representation.
6. The method according to claim 1, characterized in that The method for predicting the association between rice secondary metabolites and proteins based on the edge representation includes: Inputting the edge representation into a multi-layer perceptron to obtain the predicted probability of association between the rice secondary metabolites and proteins; The multilayer perceptron adopts a cost-sensitive learning strategy to weight positive edges and negative edges; the multilayer perceptron constructs a loss function of the multilayer perceptron by weighted binary cross entropy loss and the heterogeneous graph neighbor contrast loss.
7. A rice secondary metabolite protein association prediction device, characterized in that: The device comprises: A heterogeneous graph construction module is used to construct a heterogeneous graph with rice secondary metabolites, proteins and functional annotations as nodes, rice secondary metabolites and protein interactions, rice secondary metabolites and functional annotation associations, protein and functional annotation associations, rice secondary metabolites and rice secondary metabolites structural similarity, and protein and protein interactions as edges; A node representation learning module, which is used to learn node embedding representation by aggregating neighbor-level information and relationship-level information based on the heterogeneous graph, and to add a neighbor comparison learning framework to guide the learning of the node embedding representation; The association prediction module is used to construct an edge representation based on the node embedding representation, and predict the association between rice secondary metabolites and proteins based on the edge representation.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.