A method for predicting virus protein-host protein-disease type interactions

By constructing a multi-feature learning logic tensor decomposition model, the problem that the virus-host protein interaction prediction method in the prior art does not consider disease type and the difficulty in mining nonlinear relationships is solved, and the host protein-viral protein interaction prediction under various disease types is achieved.

CN115938471BActive Publication Date: 2025-07-18XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211360854.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-07-18
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

Existing virus-host protein interaction prediction methods focus only on predicting binary relationships, do not consider disease types, and models based on CP decomposition are difficult to mine nonlinear relationships, and fail to effectively utilize feature data.

Method used

A multi-feature learning logical tensor decomposition model based on amino acid data and protein interaction data was constructed, tensor completion was performed using alternating optimization algorithms, and the interaction probability of host protein-viral protein-disease type triplets was modeled, emphasizing the importance of known triplets.

Benefits of technology

It improves the accuracy and scope of viral protein-host protein-disease type interaction prediction, reduces the cost and blindness of biological experiments, and achieves accurate prediction of the potential associations of host protein-viral proteins under various disease types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938471B_ABST
    Figure CN115938471B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting the interaction of virus proteins-host proteins-disease types. The method includes: obtaining the amino acid data of a variety of known virus proteins, the amino acid data of a variety of known host proteins, and the interaction data of virus proteins and host proteins under different disease types; constructing a multi-feature learning logical tensor decomposition model based on the amino acid data and protein interaction data; using an alternating optimization algorithm to perform tensor completion on the constructed multi-feature learning logical tensor decomposition model; and predicting and ranking the interaction probabilities of virus protein-host protein-disease type triples according to the multi-feature learning logical tensor decomposition model after tensor completion. The present invention can integrate various feature data of biological entities, improve the prediction accuracy and prediction scope, and effectively solve the problems of high cost and blindness in biological experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of protein interaction prediction, and particularly relates to a method and device for predicting virus protein-host protein-disease type interaction, and equipment. Background Art

[0002] Viral infection involves a large number of protein-protein interactions between viruses and hosts. These interactions range from the initial binding of viral coat proteins to host membrane receptors to the hijacking of host transcriptional mechanisms by viral proteins. Therefore, knowing the protein-protein interactions between viruses and hosts helps to understand the mechanism of viral infection and design antiviral drugs.

[0003] In recent years, more and more models for predicting host-virus PPIs have been proposed. Eid et al. [7] proposed a sequence-based negative sampling and machine learning framework (DeNovo) to predict virus-host PPIs. This method has strong prediction performance for new viruses sharing the same host. Khorsand et al. [8] extracted various features based on the physicochemical properties of amino acid sequences, combined with different central positions of the host PPI network, and used random forest and majority voting to predict the interactions between human proteins and influenza A virus proteins. Tsukiyama et al. [9] established a host-virus PPIs prediction model combining word2vec and LSTM. This method has effective learning ability for datasets with imbalanced positive and negative sample ratios. Based on amino acid sequences and the human protein interaction network, Dong et al.

[10] developed a multi-task transfer learning method, which showed excellent performance in multiple benchmark datasets and case studies of SARS-CoV-2 virus receptors. Ma et al.

[11] extracted various feature information and similarity information from the amino acid sequences of proteins, and proposed an ensemble learning framework based on sequence projection to predict host-virus PPIs. This method has strong prediction ability for new host (virus) proteins. Huang et al.

[12] first introduced the tensor decomposition method into the prediction of multi-type interaction relationships. Based on CANDECOMP / PARAFAC (CP), this method incorporated biological similarity as a constraint, improving the model's prediction ability regarding types and associations. Through further improvement of CP, Dong et al.

[13] proposed a new computational model (WeightTDAIGN). This method improves the model's prediction ability through a positive sample weighting strategy, and at the same time integrates various auxiliary information of biological entities into the tensor decomposition framework to improve the model's learning ability for low-rank tensor features.

[0004] However, the methods of the prior art have at least the following technical problems:

[0005] First, the existing methods for predicting protein-protein interactions between viruses and hosts only focus on predicting the binary relationship of host-virus PPIs and do not consider the related diseases. Second, the existing CP decomposition-based methods are all linear models and are often difficult to mine complex non-linear relationships. Finally, previous models are designed based on similarity relationships and are not suitable for directly extracting information from feature data. It can be seen that the methods in the prior art have the technical problem of low accuracy in the interaction prediction method.

[0006] References

[0007] [1] Bairoch A. The Universal Protein Resource (UniProt) [J]. Nucleic Acids Research, 2004, 33 (Database issue): D154-D159.

[0008] [2] Cao DS, Xiao N, Xu QS, Chen AF. protr / ProtrWeb: R package and webserver for generating various numerical representation schemes of protein sequences [J]. Bioinformatics, 2014, 31(2): 279-281.

[0009] [3] Choi YK. Emerging and re-emerging fatal viral diseases [J]. Experimental & Molecular Medicine, 2021, 53(5): 711-712.

[0010] [4] Grange ZL, Goldstein T, Johnson CK, et al. Ranking the risk of animal-to-human spillover for newly discovered viruses [J]. Proc Natl Acad Sci U S A, 2021, 118(15).

[0011] [5]Li Z,Li X,Huang Y-Y,et al.Identify potent SARS-CoV-2main proteaseinhibitors via accelerated free energy perturbation-based virtual screeningof existing drugs[J].Proceedings of the National Academy of Sciences,2020,117(44):27381-27387.

[0012] [6]Zhou X,Park B,Choi D,Han K.Ageneralized approach to predictingprotein-protein interactions between virus and host[J].BMCGenomics,2018,19(Suppl 6):568.

[0013] [7]Eid F-E,ElHefnawi M,Heath LS.DeNovo:virus-host sequence-basedprotein–protein interaction prediction[J].Bioinformatics,,2016,32(8):1144–1150.

[0014] [8]Khorsand B,Savadi A,Zahiri J,Naghibzadeh M.Alpha influenza virusinfiltration prediction using virus-human protein-protein interaction network[J].Mathematical Biosciences and Engineering,2020,17(4):3109-3129.

[0015] [9]Tsukiyama S,Hasan MM,Fujii S,Kurata H.LSTM-PHV:prediction ofhuman-virus protein-protein interactions by LSTM with word2vec[J].BriefBioinform,2021,22(6).

[0016]

[10] Dong TN,Brogden G,Gerold G,Khosla M.Amultitask transfer learningframework for the prediction of virus-human protein-protein interactions[J].BMC Bioinformatics,2021,22(1):572.

[11] Ma Y,He T,Tan Y,Jiang X.Seq-BEL:Sequence-Based Ensemble Learning for Predicting Virus-Human Protein-ProteinInteraction[J].IEEE / ACM Trans Comput Biol Bioinform,2022,19(3):1322-1333.

[0017]

[12] Huang F,Yue X,Xiong Z,Yu Z,Liu S,Zhang W.Tensor decompositionwith relational constraints for predicting multiple types of microRNA-diseaseassociations[J].Brief Bioinform,2021,22(3).

[0018]

[13] Ouyang D,Miao R,Wang J,et al.Predicting Multiple Types ofAssociations Between miRNAs and Diseases Based on Graph Regularized WeightedTensor Decomposition[J].Frontiers in Bioengineering and Biotechnology,2022,10:859.

[0019]

[14] Narita A, Kohei Hayashi, Tomioka R, Kashima H. Tensor Factorization Using Auxiliary Information[J]. Data Min Knowl Disc, 2012(25):298–324. Summary of the Invention

[0020] In view of this, the object of the present invention is to provide a method, device and equipment for predicting the interaction of virus proteins-host proteins-disease types, which can integrate various characteristic data of biological entities, greatly improve the prediction accuracy and prediction scope, and effectively solve the problems of high cost and blindness in biological experiments.

[0021] According to one aspect of the present invention, there is provided a method for predicting the interaction of virus proteins-host proteins-disease types, including: obtaining the amino acid data of known various virus proteins, the amino acid data of known various host proteins, and the interaction data of virus proteins and host proteins under different disease types; constructing a multi-feature learning logical tensor decomposition model based on the amino acid data and protein interaction data; using an alternating optimization algorithm to perform tensor completion on the constructed multi-feature learning logical tensor decomposition model; and predicting and ranking the interaction probability of the virus protein-host protein-disease type triple according to the multi-feature learning logical tensor decomposition model after tensor completion.

[0022] According to another aspect of the present invention, there is provided a device for predicting the interaction of virus proteins-host proteins-disease types, including: an acquisition module, a construction module, a training module and a prediction module; the acquisition module is used to obtain the amino acid data of known various virus proteins, the amino acid data of known various host proteins, and the interaction data of virus proteins and host proteins under different disease types; the construction module is used to construct a multi-feature learning logical tensor decomposition model based on the amino acid data and protein interaction data; the training module is used to perform tensor completion on the constructed multi-feature learning logical tensor decomposition model by using an alternating optimization algorithm; and the prediction module is used to predict and rank the interaction probability of the virus protein-host protein-disease type triple according to the multi-feature learning logical tensor decomposition model after tensor completion.

[0023] According to another aspect of the present invention, there is provided a virus protein-host protein-disease type interaction prediction device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the virus protein-host protein-disease type interaction prediction method as described in any one of the above.

[0024] According to still another aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements the virus protein-host protein-disease type interaction prediction method as described in any one of the above.

[0025] It can be found that in the above solutions, the present invention proposes a method for predicting virus-host protein interactions for inferring multiple disease types. The present invention uses three latent factors in a shared low-dimensional space to represent host proteins, virus proteins, and disease types respectively, and further uses a logical function to model the interaction probability of the host protein-virus protein-disease type triple. At the same time, in order to emphasize the importance of known triples, the present invention assigns a higher importance level to them. In addition, the present invention extracts multiple features from the amino acid sequences of human (virus) proteins and incorporates multi-feature learning into the model, further improving the prediction accuracy of the model. The present invention has higher prediction performance and can accurately predict the potential associations between host proteins and virus proteins under multiple disease types. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0027] Figure 1 is a schematic flowchart of an embodiment of the virus protein-host protein-disease type interaction prediction method of the present invention;

[0028] Figure 2 is a schematic diagram of the overall framework of an embodiment of the virus protein-host protein-disease type interaction prediction method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate the present invention, but do not limit the scope of the present invention. Similarly, the following embodiments are only partial embodiments of the present invention rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0030] The present invention provides a method for predicting the interaction of viral proteins-host proteins-disease types, which can integrate various characteristic data of biological entities, greatly improve the prediction accuracy and prediction scope, and effectively solve the problems of high cost and blindness in biological experiments.

[0031] Please refer to Figure 1 、 Figure 2 , Figure 1 which is a schematic flowchart of an embodiment of the method for predicting the interaction of viral proteins-host proteins-disease types of the present invention. Figure 2 which is a schematic diagram of the overall framework of an embodiment of the method for predicting the interaction of viral proteins-host proteins-disease types of the present invention. The following will be described in combination with methods and examples. It should be noted that if there are substantially the same results, the method of the present invention is not limited to Figure 1 the process sequence shown. As Figure 1 shown, the method includes the following steps:

[0032] S101: Obtain the amino acid data of known various viral proteins, the amino acid data of known various host proteins, and the interaction data of viral proteins and host proteins under different disease types.

[0033] In this embodiment, based on a multi-source database, the host protein-viral protein association data under various disease types is extracted, and various characteristic data of the host protein (or viral protein) is calculated. Specifically, the present invention can obtain the above information from existing data and existing methods.

[0034] From the MorCVD database, the interaction relationships between host proteins and viral proteins under various diseases are downloaded one by one, and finally 6 disease types, as well as 8,776 interactions between 1,012 host proteins and 836 viral proteins, are obtained.

[0035] S102: Construct a multi-feature learning logical tensor decomposition model based on the amino acid data and protein interaction data.

[0036] In this embodiment, the protr method is used to calculate the features of the amino acid data of the known host protein and the amino acid data of the known viral protein, extract the protein features of each amino acid data from the calculated amino acid data, and extract the pseudo amino acid composition features and combined triplet descriptor features of the host protein, the pseudo amino acid composition features and combined triplet descriptor features of the viral protein, and the interaction tensor between the host protein and the viral protein under different disease types from the protein features. The multi-feature learning logical tensor decomposition model is constructed by using the above features and interaction tensor as the training input of the multi-feature learning logical tensor decomposition model. The purpose of this setting is to extract multiple features from the amino acid sequence of the host (virus) protein and incorporate multi-feature learning into the model to further improve the prediction accuracy of the model.

[0037] Specifically, use the R package "protr" [2] Extract the amino acid sequences of these proteins from the UniProt database [1] and calculate two features of the host protein respectively: the pseudo amino acid composition feature F1 u and the combined triplet descriptor feature as well as two features of the viral protein: the pseudo amino acid composition feature F1 v and the combined triplet descriptor feature At this time, 8,776 interactions among 6 disease types, 1,012 host proteins and 836 viral proteins, combined with the two features of the host protein and the two features of the viral protein, constitute the benchmark dataset.

[0038] The names of the 6 diseases are as follows: Cardiovascular Infections, Dilated Cardiomyopathy, Endocarditis, Hypereosinophilic Syndrome, Pericarditis, Viral Myocarditis.

[0039] The first 20 of the 1,012 host proteins are as follows: A1L0V1, A5YKK6, F2Z2I2, F8VYE8, O00116, O00165, O00203, O00232, O00299, O00308, O00327, O00391, O00410, O00459, O00487, O00505, O00560, O00571, O00629, O00767.

[0040] The top 20 of 867 viral proteins are as follows: A0A075C673, A0A089NDG7, A0A0C7TA99, A0A0C7TRZ7, A0A0C7TUY8, A0A0F7R9M2, A0A0F7RKT3, A0A0U5B5J2, A0A166MB26, A0A1P7U111, A0A218M2D5, A0A288CFT1, A5A5U1, A5Z256, A6YJF2, A8E1C4, B0FAN4, B2BUF1, B4URE7, B4URF0.

[0041] S103: Tensor completion is performed on the constructed multi-feature learning logical tensor decomposition model using the alternating optimization algorithm.

[0042] In this embodiment, the factor matrix of host proteins, the factor matrix of viral proteins, and the factor matrix of disease types are calculated according to the multi-feature learning logical tensor decomposition model, and the projection matrix of host proteins, the projection matrix of viral proteins are obtained through the multi-feature learning logical tensor decomposition model, and the weight coefficients of the host protein feature projection terms and the weight coefficients of the viral protein feature projection terms are obtained through the multi-feature learning logical tensor decomposition model, and the alternating fixed update optimization of the factor matrix, the projection matrix, and the weight coefficients is performed until the multi-feature learning logical tensor decomposition model converges to the optimal solution.

[0043] The purpose of the setting is to use three latent factors in the shared low-dimensional space to represent host proteins, viral proteins, and disease types respectively, and further use the logical function to model the interaction probability of the host protein-viral protein-disease type triple. To emphasize the importance of known triples, a higher importance level is assigned to them in this embodiment.

[0044] Specifically, according to the pseudo-amino acid composition feature F1 of host proteins u and the combined triple descriptor feature the pseudo-amino acid composition feature F1 of viral proteins v and the combined triple descriptor feature as well as the interaction tensor of I host proteins and J viral proteins under K disease types The factor matrix U ∈ R of host proteins is calculated by the following formula I×D the factor matrix V ∈ R of viral proteins J×D the factor matrix W ∈ R of disease types K×D :

[0045]

[0046] Among them, c≥1 represents the importance level parameter, which controls the importance multiple of known interactions relative to unknown interactions; is a tensor of the (i, j, k)-th element. When there is an interaction relationship between the i-th host protein and the j-th viral protein under disease type k, Otherwise, represents the tensor composed of U, V, and W, that is P i u represents the projection matrix of the i-th host protein feature, represents the projection matrix of the j-th viral protein feature; represents the weight coefficient of the i-th host protein feature projection term, represents the weight coefficient of the j-th viral protein feature projection term, η≥1 is the exponent of α i and β j used to control the importance level of the projection term. μ represents the regularization weight of the projection matrix, and λ represents the weight coefficient of the factor matrix.

[0047] In the specific implementation, the solution of formula (1) can be divided into the following three sub-problems:

[0048] Update the factor matrices {U, V, W}: Fix The partial derivatives with respect to U, V, and W can be obtained as follows:

[0049]

[0050] Among them, the tensor whose (i, j, k)-th element is as shown in formula (2). The matrices and respectively represent the mode-1, mode-2, and mode-3 matrixizations of the tensor . * represents the Hadamard product, that is, the product of the corresponding elements of tensors (or matrices) of the same scale. e represents the Khatri–Rao product of two matrices.

[0051] Update the projection matrix Fix U, V, W, and The iterative formula with respect to can be obtained as follows:

[0052]

[0053] Among them, U + and U - respectively represent the positive part and the negative part of U, that is, U + =(U + |U|) / 2, U- = (|U| - U) / 2, where |·| represents the absolute value. Similarly, V + and V - represent the positive and negative parts of V, respectively.

[0054] Update the weight vector Fix U, V, W, and we can obtain the explicit solution for as follows:

[0055]

[0056] By alternately optimizing the above three sub-problems, it finally converges to the optimal solution.

[0057] S104: According to the multi-feature learning logical tensor decomposition model after tensor completion, predict and rank the interaction probabilities of virus protein-host protein-disease type triples.

[0058] In this embodiment, the protr method is used to extract the pseudo-amino acid composition features and joint triple descriptor features that associate the virus protein and the host protein, as well as the interaction tensors of host proteins and virus proteins under different disease types, from the virus protein amino acid data and the host protein amino acid data. The extracted pseudo-amino acid composition features and joint triple descriptor features that associate the virus protein and the host protein, as well as the interaction tensors of host proteins and virus proteins under different disease types, are input into the multi-feature learning logical tensor decomposition model after tensor completion to predict and rank the interaction probabilities of virus protein-host protein-disease type triples, and the predicted interaction probabilities are sorted in descending order through the multi-feature learning logical tensor decomposition model after tensor completion.. In this embodiment, multiple features are extracted from the amino acid sequences of host (virus) proteins, and multi-feature learning is incorporated into the model, further improving the prediction accuracy of the model. This embodiment has higher prediction performance and can accurately predict the potential associations between host proteins and virus proteins under multiple disease types.

[0059] Specifically, the multi-feature learning logical tensor decomposition model outputs the factor matrix U of host proteins, the factor matrix V of virus proteins, and the factor matrix W of disease types. According to formula (5), the interaction probability between the i-th host protein and the j-th virus protein under the k-th disease type can be calculated as follows:

[0060]

[0061] Take the predicted interaction probability Perform a descending order sorting to obtain the ranking of the association between viral proteins and host proteins under different disease types.

[0062] The interactions of the top 40 host protein-virus protein-disease type with the highest scores are shown in Table 1.

[0063] Table 1 Ranking table of interaction probabilities of 40 host protein-virus protein-disease types

[0064]

[0065]

[0066] To further illustrate the beneficial effects of the method provided by the embodiments of the present invention, the effectiveness is verified through several specific examples below.

[0067] First, regarding CV type and CV triplet For two experimental scenarios, the performance of the embodiments of the present invention is evaluated through a 5-fold cross-validation method.

[0068] CV type Scenario: Evaluate the accuracy of the model in predicting disease types. This method takes host-virus PPIs with at least one disease type and randomly divides them into 5 subsets of equal size. In each experiment, one subset is selected as the test set, and the rest are used as the training set.

[0069] CV triplet Scenario: Evaluate the prediction ability of the model regarding the host protein-virus protein-disease type triples. This method divides the known host-virus PPIs under all disease types into 5 subsets of equal size. In each experiment, one subset is selected as the test set, and the rest are used as the training set.

[0070] Regarding CV type , we are interested in the disease type with the highest score in the test set. Therefore, the average Top-1 precision, average Top-1 recall, and average Top-1 F1 of the predicted disease type are calculated as evaluation metrics. Among them, the formula for precision is as follows:

[0071]

[0072] Among them, TP represents the number of true positive examples and the predicted results are also positive examples, and FP represents the number of true negative examples and the predicted results are positive examples. The Top-1 precision refers to the result calculated according to formula (6) by taking the one with the highest predicted score as the predicted positive example and the others as predicted negative examples.

[0073] The formula for recall is as follows:

[0074]

[0075] Among them, FN represents the number of true positive examples that are predicted as negative examples. The recall rate of Top-1 refers to the result calculated according to formula (7) by taking the example with the highest predicted score as the predicted positive example and the others as predicted negative examples.

[0076] The calculation formula of the F1 value is as follows:

[0077]

[0078] Among them, the definitions of P and R are as shown in formulas (6) and (7). The recall rate of Top-1 refers to the result calculated by combining formulas (6), (7) and (8) by taking the example with the highest predicted score as the predicted positive example and the others as predicted negative examples.

[0079] For CV triplet , we adopt the area under the precision-recall curve (AUPR), the area under the ROC curve (AUC) and the F1 value as evaluation indicators.

[0080] Comparative Example 1

[0081] Selection of comparative methods. CP

[12] : The most basic tensor decomposition method without using any auxiliary information. TFAI

[14] : Tensor factorization based on auxiliary information. TDRC

[12] : Tensor decomposition with relational constraints for predicting multiple types of microRNA-disease associations. WeightTDAIGN

[13] : Predicting multiple associations between miRNAs and diseases based on graph-regularized weighted tensor decomposition. The experimental results of Example 1 and the comparative examples are shown in Table 2.

[0082] Table 2 Comparison of prediction performances of different methods

[0083]

[0084] As shown in Table 2, the method of the present invention achieves the best prediction performance in all indicators. Specifically, for CV type, the top1 accuracy of the method of the present invention reaches 0.9119, which is increased by 25.64%, 16.94%, 16.58% and 21.97% respectively compared with 0.7258 of CP, 0.7798 of TAFI, 0.7822 of TDRC and 0.7476 of WeightTDAIGN. The Top1 recall rate of the method of the present invention is 0.8148, which is higher than 0.6585 of CP, 0.6967 of TAFI, 0.6988 of TDRC and 0.6680 of WeightTDAIGN. The Top1 F1 of the method of the present invention is 0.8711, which is also higher than the prediction results of other methods. Regarding CV triplet , the AUPR value of the method of the present invention is 0.9555, which is increased by 6.95%, 6.81%, 7.64% and 3.41% respectively compared with 0.8934 of CP, 0.8946 of TAFI, 0.8877 of TDRC and 0.9241 of WeightTDAIGN. The AUC value of the method of the present invention is 0.9523 and the F1 value is 0.8890, which are higher than the calculation results of other methods.

[0085] Generally speaking, the present invention proposes a method for predicting virus-host protein interactions of multiple disease types based on multi-feature learning and logical tensor decomposition.

[0086] The present invention can accurately predict the interaction relationships between host proteins and virus proteins of multiple disease types, effectively avoiding the high consumption of manpower and material resources caused by biochemical experiments.

[0087] It can be found that in this embodiment, the present invention proposes a method for predicting virus-host protein interactions for inferring multiple disease types. The present invention uses three latent factors in the shared low-dimensional space to represent host proteins, virus proteins and disease types respectively, and further uses a logical function to model the interaction probability of the host protein-virus protein-disease type triple. At the same time, in order to emphasize the importance of known triples, the present invention assigns a higher importance level to them. In addition, the present invention extracts multiple features from the amino acid sequences of human (virus) proteins and incorporates multi-feature learning into the model, further improving the prediction accuracy of the model. The present invention has higher prediction performance and can accurately predict the potential associations between host proteins and virus proteins under multiple disease types.

[0088] The present invention also provides a device for predicting virus protein-host protein-disease type interactions, which can integrate various feature data of biological entities, greatly improving the prediction accuracy and prediction scope, and effectively solving the problems of high cost and blindness in biological experiments.

[0089] In this embodiment, the virus protein-host protein-disease type interaction prediction device includes an acquisition module, a construction module, a training module, and a prediction module;

[0090] The acquisition module is used to acquire the amino acid data of a variety of known virus proteins, the amino acid data of a variety of known host proteins, and the interaction data between virus proteins and host proteins under different disease types;

[0091] The construction module is used to construct a multi-feature learning logical tensor decomposition model based on the amino acid data and protein interaction data;

[0092] The training module is used to perform tensor completion on the constructed multi-feature learning logical tensor decomposition model by using an alternating optimization algorithm;

[0093] The prediction module is used to predict and rank the interaction probabilities of virus protein-host protein-disease type triples according to the multi-feature learning logical tensor decomposition model after tensor completion.

[0094] Optionally, the construction module can specifically be used for:

[0095] Using the protr method to calculate the features of the amino acid data of the known host proteins and the amino acid data of the known virus proteins, extracting the protein features of each amino acid data from the calculated amino acid data, and extracting the pseudo-amino acid composition features and joint triple descriptor features of the host proteins and the pseudo-amino acid composition features and joint triple descriptor features of the virus proteins and the interaction tensors between host proteins and virus proteins under different disease types from the protein features, and constructing the multi-feature learning logical tensor decomposition model by using the above features and interaction tensors as the training input of the multi-feature learning logical tensor decomposition model.

[0096] Optionally, the training module can specifically be used for:

[0097] Calculating the factor matrix of host proteins, the factor matrix of virus proteins, and the factor matrix of disease types according to the multi-feature learning logical tensor decomposition model, obtaining the projection matrix of host proteins and the projection matrix of virus proteins through the multi-feature learning logical tensor decomposition model, obtaining the weight coefficients of host protein feature projection terms and the weight coefficients of virus protein feature projection terms through the multi-feature learning logical tensor decomposition model, and alternately fixing and updating the optimization of the factor matrix, the projection matrix, and the weight coefficients until the multi-feature learning logical tensor decomposition model converges to the optimal solution.

[0098] The prediction module can specifically be used for:

[0099] The protr method is used to extract the pseudo-amino acid composition features and combined triplet descriptor features that associate the viral protein and the host protein, as well as the interaction tensors of the host protein and the viral protein under different disease types, from the viral protein amino acid data and the host protein amino acid data. The extracted pseudo-amino acid composition features and combined triplet descriptor features that associate the viral protein and the host protein, as well as the interaction tensors of the host protein and the viral protein under different disease types, are input into the multi-feature learning logical tensor decomposition model after tensor completion to predict and rank the interaction probabilities of the viral protein-host protein-disease type triples, and the predicted interaction probabilities obtained by the multi-feature learning logical tensor decomposition model after tensor completion are sorted in descending order.

[0100] Each unit module of the viral protein-host protein-disease type interaction prediction device can respectively execute the corresponding steps in the above method embodiments, so the unit modules will not be elaborated here. For details, please refer to the descriptions of the above corresponding steps.

[0101] The present invention further provides a viral protein-host protein-disease type interaction prediction device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above viral protein-host protein-disease type interaction prediction method.

[0102] Wherein, the memory and the processor are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0103] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.

[0104] The present invention further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0105] It can be found that in the above solution, the present invention proposes a method for predicting virus-host protein interactions for inferring multiple disease types. The present invention uses three latent factors in a shared low-dimensional space to represent host proteins, viral proteins, and disease types respectively, and further uses a logical function to model the interaction probability of the host protein-viral protein-disease type triple. At the same time, in order to emphasize the importance of known triples, the present invention assigns a higher importance level to them. In addition, the present invention extracts multiple features from the amino acid sequences of human (viral) proteins and incorporates multi-feature learning into the model, further improving the prediction accuracy of the model. The present invention has higher prediction performance and can accurately predict the potential associations between host proteins and viral proteins under multiple disease types.

[0106] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.

[0107] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0108] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0109] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods according to various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0110] The above are only some embodiments of the present invention, and thus do not limit the protection scope of the present invention. Any equivalent device or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall equally be included in the patent protection scope of the present invention.

Claims

1. A method for predicting the interaction between viral proteins, host proteins, and disease types, characterized in that, Including: Obtaining the amino acid data of a variety of known viral proteins, the amino acid data of a variety of known host proteins, and the interaction data between viral proteins and host proteins under different disease types; Constructing a multi-feature learning logical tensor decomposition model based on the amino acid data and protein interaction data; Using an alternating optimization algorithm to perform tensor completion on the constructed multi-feature learning logical tensor decomposition model; Predicting and ranking the interaction probabilities of the virus protein-host protein-disease type triples according to the multi-feature learning logical tensor decomposition model after tensor completion; The constructing of the multi-feature learning logical tensor decomposition model based on the amino acid data and protein interaction data includes: Using the protr method to calculate the features of the amino acid data of the known host proteins and the known viral proteins, extracting the protein features of each amino acid data from the calculated amino acid data, and extracting the pseudo amino acid composition features and joint triple descriptor features of the host proteins, as well as the pseudo amino acid composition features and joint triple descriptor features of the viral proteins, and the interaction tensors between host proteins and viral proteins under different disease types from the protein features, and constructing the multi-feature learning logical tensor decomposition model by using the above features and interaction tensors as the training inputs of the multi-feature learning logical tensor decomposition model; The using of the alternating optimization algorithm to perform tensor completion on the constructed multi-feature learning logical tensor decomposition model includes: Calculating the factor matrix of host proteins, the factor matrix of viral proteins, and the factor matrix of disease types according to the multi-feature learning logical tensor decomposition model, obtaining the projection matrix of host proteins and the projection matrix of viral proteins through the multi-feature learning logical tensor decomposition model, obtaining the weight coefficients of the host protein feature projection terms and the weight coefficients of the viral protein feature projection terms through the multi-feature learning logical tensor decomposition model, and alternately fixing and updating the optimization of the factor matrix, the projection matrix, and the weight coefficients until the multi-feature learning logical tensor decomposition model converges to the optimal solution.

2. The method for predicting the interaction between viral proteins-host proteins-disease types according to claim 1, wherein The predicting and ranking the interaction probabilities of the virus protein-host protein-disease type triples according to the multi-feature learning logical tensor decomposition model after tensor completion includes: Using the protr method, pseudo-amino acid composition features, combined triplet descriptor features that associate the viral protein and the host protein, and interaction tensors of host proteins and viral proteins under different disease types are extracted from viral protein amino acid data and host protein amino acid data. The extracted pseudo-amino acid composition features, combined triplet descriptor features that associate the viral protein and the host protein, and interaction tensors of host proteins and viral proteins under different disease types are input into the multi-feature learning logical tensor decomposition model after tensor completion to predict and rank the interaction probabilities of the viral protein-host protein-disease type triples, and the predicted interaction probabilities obtained through the multi-feature learning logical tensor decomposition model after tensor completion are sorted in descending order.

3. A virus protein-host protein-disease type interaction prediction device, characterized in that Including: An acquisition module, a construction module, a training module, and a prediction module; The acquisition module is used to acquire amino acid data of known multiple viral proteins, amino acid data of known multiple host proteins, and interaction data of viral proteins and host proteins under different disease types; The construction module is used to construct a multi-feature learning logical tensor decomposition model based on amino acid data and protein interaction data; The training module is used to perform tensor completion on the constructed multi-feature learning logical tensor decomposition model by using an alternating optimization algorithm; The prediction module is used to predict and rank the interaction probabilities of the viral protein-host protein-disease type triples according to the multi-feature learning logical tensor decomposition model after tensor completion; The construction module is specifically used for: Using the protr method to calculate features of the amino acid data of the known host proteins and the known viral proteins, extracting protein features of each amino acid data from the calculated amino acid data, and extracting pseudo-amino acid composition features and combined triplet descriptor features of the host protein, pseudo-amino acid composition features and combined triplet descriptor features of the viral protein, and interaction tensors of host proteins and viral proteins under different disease types from the protein features, and constructing the multi-feature learning logical tensor decomposition model in a way that uses the above features and interaction tensors as the training input of the multi-feature learning logical tensor decomposition model; The training module is specifically used for: Calculating the factor matrix of the host protein, the factor matrix of the viral protein, and the factor matrix of the disease type according to the multi-feature learning logical tensor decomposition model, obtaining the projection matrix of the host protein and the projection matrix of the viral protein through the multi-feature learning logical tensor decomposition model, obtaining the weight coefficients of the host protein feature projection terms and the weight coefficients of the viral protein feature projection terms through the multi-feature learning logical tensor decomposition model, and alternately fixing and updating and optimizing the factor matrix, the projection matrix, and the weight coefficients until the multi-feature learning logical tensor decomposition model converges to the optimal solution.

4. The virus protein-host protein-disease type interaction prediction device according to claim 3, wherein The prediction module is specifically used for: Using the protr method, pseudo-amino acid composition features, combined triplet descriptor features that associate the viral protein and the host protein, and interaction tensors of host proteins and viral proteins under different disease types are extracted from viral protein amino acid data and host protein amino acid data. The extracted pseudo-amino acid composition features, combined triplet descriptor features that associate the viral protein and the host protein, and interaction tensors of host proteins and viral proteins under different disease types are input into the multi-feature learning logical tensor decomposition model after tensor completion to predict and rank the interaction probabilities of the viral protein-host protein-disease type triples, and the predicted interaction probabilities obtained through the multi-feature learning logical tensor decomposition model after tensor completion are sorted in descending order.

5. A viral protein-host protein-disease type interaction prediction device, characterized in that, Comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the viral protein-host protein-disease type interaction prediction method according to any one of claims 1 to 2.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the viral protein-host protein-disease type interaction prediction method according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Virus-host interaction prediction method based on graph convolutional neural network

    CN112331257A

  • Enzyme prediction method and device and storage medium

    CN115188433A