Method, System and Medium for Predicting Adverse Drug Reactions with Multi-Attribute Feature Filling
By constructing a multi-attribute missing feature filling model for drug-related and unique features, the problem of missing attribute features in the prediction of adverse reactions between drugs is solved, and the prediction accuracy and drug safety are improved.
Patent Information
- Application Number
- CN202211434048.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-11-16
AI Technical Summary
The prior art fails to effectively deal with the situation where the drug has missing attribute characteristics in predicting adverse reactions between drugs, resulting in inaccurate prediction of prediction results and affecting the safety of drug use.
By constructing a multi-attribute missing feature filling model for drug based on common and unique features, fill in the missing features of the drug, and use the filled attribute features to predict adverse reactions between drugs.
It improves the accuracy of predicting adverse reactions between drugs, promotes experimental research on adverse reactions between drugs, and ensures the safety of medication.
Smart Images

Figure CN115831390B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of adverse reaction prediction, and in particular, to a method, system and medium for predicting adverse drug reactions with multi-attribute feature filling. Background Art
[0002] Adverse drug reactions refer to the situation where when two drugs are taken simultaneously, the efficacy or pharmacological effect of one drug is destroyed by the other drug, thereby changing the original in-vivo process of the drug, the response of tissues or organs to the drug, and the physicochemical properties of the drug, resulting in adverse reactions or toxic side effects harmful to the human body.
[0003] Currently, adverse drug reactions have become an important factor delaying disease treatment, aggravating the patient's condition, and affecting the morbidity and mortality of patients. The research on adverse drug reactions has gradually attracted the attention of relevant medical and health institutions and has become the research focus in the current field of medical health. To solve this problem, pharmaceutical companies invest a large amount of funds in clinical adverse drug reaction experiments during the drug R & D stage; currently, the research work on predicting adverse drug reactions is mainly divided into two categories: knowledge-base-based methods and similarity-based methods.
[0004] Knowledge-base-based methods usually detect adverse drug reactions from biomedical texts, electronic medical records, biological heterogeneous databases, and the FDA Adverse Event Reporting System based on technologies such as data mining and natural language processing. This type of method relies on the clinical accumulation of adverse drug reaction data and attempts to discover and extract adverse drug reactions from a large amount of unformatted data; similarity-based methods first extract drug attribute information from a drug database, calculate the attribute similarity score based on the relationship between drug attribute information, and then design a machine learning model to explore the implicit relationship between the similarity score and adverse drug reactions to predict potential adverse drug reactions. This method can achieve the prediction of adverse drug reactions only relying on drug attribute information without the need for a large amount of prior accumulation of adverse drug reaction data.
[0005] However, in the prior art, when predicting adverse drug reactions, in the process of constructing a prediction model for adverse drug reactions based on drug attribute features, the drugs used usually have complete attribute feature information, and drugs with missing attribute features have not been considered: different attribute information of drugs usually comes from different heterogeneous databases, and there are significant differences in the number of drugs included in different databases and the attribute information recorded. For example, comparing the databases DrugBank and SIDER, it can be seen that the number of drugs included in SIDER is much smaller than that included in DrugBank. It can be seen that the vast majority of drugs in DrugBank with molecular structure, target, and enzyme information do not have side effect information, resulting in a serious lack of drug side effect information. In addition, there are also certain degrees of missing for other attributes of drugs due to the differences in the number and type of drugs included in different databases. As the number of attribute factors considered in the model increases, the number of drugs with complete attribute feature information will gradually decrease, and the phenomenon of missing drug attribute feature information is serious.
[0006] Therefore, if this method continues to be used to predict adverse drug reactions, the efficiency of adverse reaction research will be reduced, and the predicted results will be inaccurate. In serious cases, it will lead to the occurrence of medication safety accidents.
[0007] In view of this, the present application is specifically proposed. Summary of the Invention
[0008] The technical problem to be solved by the present invention is that in the prior art, when using the knowledge base method or the similarity method to predict adverse drug reactions, drugs with missing attribute features have not been considered, and the missing attribute features of different drugs often vary greatly. Therefore, if the prior art continues to be used for prediction, the efficiency of adverse reaction research will be reduced, the predicted results will be inaccurate, and further, it will cause the occurrence of medication safety accidents. The purpose is to provide a prediction method, system, and medium for adverse drug reactions with multi-attribute feature filling, which can improve the research rate of adverse drug reactions, improve the accuracy of predicting adverse drug reactions, and ensure medication safety.
[0009] The present invention is achieved through the following technical solutions:
[0010] A prediction method for adverse drug reactions with multi-attribute feature filling, the method steps include:
[0011] Obtain adverse drug reaction data and drug multi-attribute data;
[0012] Based on the drug multi-attribute data, construct a drug multi-attribute missing feature filling model based on common features and unique features;
[0013] The drug multi-attribute missing value filling model is corrected by a cosine similarity regular term, and the Lagrangian function, the alternating direction multiplier method, and the non-negative matrix factorization method are used to solve the corrected drug multi-attribute missing value filling model, so as to obtain the common features and specific features of the drug multi-attributes;
[0014] Based on the common features and specific features of the drug multi-attributes, combined with the adverse reaction data, a prediction model is constructed;
[0015] Obtain the multi-attribute data between any two drugs, calculate the common features and specific features of the drug multi-attributes, and input them into the prediction model to obtain the adverse reaction prediction results between the drugs.
[0016] In traditional adverse reaction prediction between drugs, the knowledge base method or the similarity method is usually used to predict the adverse reactions between drugs. However, when using this method to predict the adverse reactions between drugs, drugs with missing attribute features have not been considered. The missing attribute features of different drugs often vary greatly. Therefore, if drugs with missing attribute features are still used to carry out the adverse reaction prediction between drugs, not only the potential relationship between the drug multi-attributes and the adverse reactions cannot be deeply analyzed, but also the prediction results will be inaccurate, which will further lead to the occurrence of medication safety accidents. The present invention provides an adverse reaction prediction method based on multi-attribute feature filling. By constructing a drug multi-attribute missing feature filling model to fill the missing features of drugs, and then using the filled attribute features to carry out the adverse reaction prediction between drugs, it not only improves the accuracy of the adverse reaction prediction between drugs, but also promotes the experimental research on the adverse reactions between drugs, and ensures the medication safety.
[0017] Preferably, the specific steps for constructing the drug multi-attribute missing feature filling model include:
[0018] Based on the drug multi-attribute data, and based on the relationship between the common features and specific features of the drug attributes and the original feature space, a basic model is constructed;
[0019] The KL divergence is used to measure the distribution difference between the specific features of different attributes, and under the constraint of the distribution difference, the basic model is processed to obtain a multi-attribute missing feature filling model based on the common features and specific features.
[0020] Preferably, the specific method for obtaining the common features and specific features of the drug multi-attributes is:
[0021] The drug multi-attribute missing value filling model is corrected by a cosine similarity regular term to obtain a corrected model;
[0022] Solve the modified model through the Lagrangian function, the alternating direction method of multipliers, and the non - negative matrix factorization method to obtain the common features and specific features of the filled drug attribute feature space and the multi - attribute feature space, as well as the iterative update formula of the reconstruction coefficient matrix of the attribute feature space;
[0023] Iteratively update the modified model variables until the number of iterative updates reaches the maximum value or the model reaches the minimum change threshold to obtain the common features and specific features of the drug multi - attributes.
[0024] Preferably, the multi - attribute data includes molecular structure data, target data, pathway data, side - effect data, phenotype data, and disease data.
[0025] Preferably, the specific expression of the multi - attribute missing feature filling model based on common features and specific features is:
[0026]
[0027]
[0028] represents the Frobenius norm of the matrix, ||·||0 represents the matrix 0 - norm, P is the common feature of the drug multi - attributes, Q m is the specific feature of the m - th attribute of the drug, U m is the original feature space X based on common features and specific features in the m - th attribute m reconstruction coefficient matrix, X m is the m - th attribute feature space of the drug, is the known attribute feature information of the drug, KL is the divergence, represents the marking matrix of the known attribute feature information of the m - th attribute of the drug, represents extracting the drugs with known attribute features in the original feature space X of the attribute m and arranging them in index order to obtain the known attribute feature information of the drug α m represents the sparsity regularization parameter of the m - th attribute reconstruction coefficient matrix, and β represents the regularization parameter of the KL divergence between specific features of different attributes.
[0029] Preferably, the specific expression of the modified model is:
[0030]
[0031]
[0032] S m (d i ,d j)Can be regarded as a vector and The normalized cosine similarity between (P + Q m ) i. and (P + Q m ) j· is the combined representation of the common features and unique features of drug d i and d j γ represents the regularization parameter of the cosine similarity regularization term for the m-th attribute, m represents the l2 norm of the vector.
[0033] Preferably, the specific expression of the prediction model is:
[0034]
[0035] is the contribution of the common features to the adverse reactions between drugs, is the contribution of the unique features to the adverse reactions between drugs, r ij is the adverse reaction relationship between drugs.
[0036] Preferably, the specific expression of the is:
[0037]
[0038] The specific expression of the is:
[0039]
[0040] P i· is the multi-attribute common feature of drug d i P j· is the multi-attribute common feature of drug d j is the unique feature of the m-th attribute of d i is the unique feature of the m-th attribute of d j m w represents the contribution degree of the unique feature of the m-th attribute and to the adverse reactions between d i and d j λ represents the contribution degree of the common features P i· and P j· to the adverse reactions between d i and d j is the common feature - adverse reaction tensor, ε m mThe potential relationship between the unique characteristics of the m-th attribute and adverse reactions.
[0041] The present invention also provides a prediction system for adverse drug reactions among drugs with multi-attribute feature filling, including a data acquisition module, a filling model construction module, an analysis module, a prediction model construction module, and a prediction module;
[0042] The data acquisition module is used to acquire adverse drug reaction data and drug multi-attribute data;
[0043] The filling model construction module is used to construct a filling model for missing features of drug multi-attributes based on the common features and unique features based on the drug multi-attribute data;
[0044] The analysis module is used to correct the filling model for missing drug multi-attributes through a cosine similarity regular term, and use the Lagrangian function, the alternating direction multiplier method, and the non-negative matrix factorization method to solve the corrected filling model for missing drug multi-attributes, and obtain the common features and unique features of the drug multi-attributes;
[0045] The prediction model construction module is used to construct a prediction model based on the common features and unique features of drug multi-attributes in combination with the adverse reaction data;
[0046] The prediction module is used to obtain the multi-attribute data between any two drugs, calculate the common features and unique features of the drug multi-attributes, and input them into the prediction model to obtain the prediction result of adverse drug reactions between drugs.
[0047] The present invention also provides a computer storage medium, on which a computing program is stored. When the computer program is executed by a processor, the method described above is implemented.
[0048] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0049] The present invention provides a method, system and medium for predicting adverse drug reactions among drugs with multi-attribute feature filling. By constructing a filling model for missing features of drug multi-attributes to fill the missing features of drugs, and using the filled attribute features to carry out the prediction of adverse drug reactions among drugs, it not only improves the accuracy of predicting adverse drug reactions among drugs, but also promotes the experimental research on adverse drug reactions among drugs, and ensures the safety of drug use. Description of the Drawings
[0050] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0051] Figure 1 It is a schematic diagram of the prediction method;
[0052] Figure 2 It is a model framework diagram for filling multi-attribute features of drugs;
[0053] Figure 3 It is a model framework diagram for predicting adverse drug reactions between drugs based on common features and specific features;
[0054] Figure 4 It is a partial prediction of adverse drug reactions between drugs;
[0055] Figure 5 It is a partial prediction of adverse drug reactions between drugs. Detailed implementation manners
[0056] To make the purpose, technical solutions and advantages of the present invention more clear and understandable, the following will further elaborate on the present invention in combination with embodiments and drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and do not serve as a limitation to the present invention.
[0057] In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that these specific details do not have to be adopted to implement the present invention. In other embodiments, well-known structures, circuits, materials or methods are not specifically described in order to avoid obscuring the present invention.
[0058] Throughout the specification, the reference to "an embodiment", "embodiment", "an example" or "example" means that the specific features, structures or characteristics described in connection with that embodiment or example are included in at least one embodiment of the present invention. Thus, the phrases "an embodiment", "embodiment", "an example" or "example" that appear throughout the specification do not necessarily all refer to the same embodiment or example. In addition, the specific features, structures or characteristics can be combined in any appropriate combination and / or sub-combination in one or more embodiments or examples. In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0059] In the description of the present invention, the orientation or positional relationship indicated by terms such as "front", "rear", "left", "right", "upper", "lower", "vertical", "horizontal", "high", "low", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the protection scope of the present invention.
[0060] Embodiment 1
[0061] In the traditional prior art, the method of using a knowledge base or the method of similarity is usually adopted to predict the adverse reactions between drugs. However, when using this method to predict the adverse reactions between drugs, drugs with missing attribute features have not been considered. The missing attribute features of different drugs often have huge differences. Therefore, if we continue to use drugs with missing attribute features to carry out the prediction of adverse reactions between drugs, not only can we not deeply analyze the potential relationship between the multi-attributes of drugs and adverse reactions, but also the prediction results will be inaccurate, and further, it will cause the occurrence of medication safety accidents.
[0062] This embodiment provides a method for predicting adverse reactions between drugs with multi-attribute feature filling. Based on the multi-attribute information of drugs, effective filling of missing attribute features is carried out, and a filling model for multi-attribute missing features of drugs based on common features and specific features is established; based on the common features and specific features of drug attributes, an adverse reaction prediction model based on drug multi-attribute information is constructed to explore the contribution degree of different attributes to adverse reactions and realize the prediction of adverse reactions between drugs. This method can provide data support for the experimental research on adverse reactions between drugs, improve the clinical experimental research on drug adverse reactions, and is of great significance for reducing the occurrence of adverse reaction events between drugs, improving the efficiency of adverse reaction research, and promoting medication safety.
[0063] The specific prediction method is as Figure 1 shown, and the method steps include:
[0064] S1: Obtain the adverse reaction data between drugs and the multi-attribute data of drugs;
[0065] The multi-attribute data includes molecular structure data, target data, pathway data, side effect data, phenotype data, and disease data.
[0066] In step S1, adverse drug reaction data among drugs are collected from the TWOSIDES database. The TWOSIDES database records adverse reactions caused by the combined use of two drugs; the molecular structure and target information of drugs are sourced from the DrugBank database; the pathway and disease information of drugs are sourced from the KEGG database; the drug side effect information is sourced from the SIDER database; the drug phenotype information is sourced from the CTD database. Regarding the molecular structure information of drugs, the SMILES molecular formula of drugs is encoded using PubChem substructure fingerprints, and each drug contains 881-dimensional substructure information. Regarding the other attribute information of drugs, the attribute information of drugs is represented by binary vectors, and the vector elements 1 and 0 represent whether the drug contains the characteristic information of the corresponding attribute, respectively. The source databases and characteristic dimensions of the multi-attribute information of drugs are shown in Table 1. Based on the adverse drug reaction data among drugs and the multi-attribute data of drugs, a total of 1,188,258 groups of adverse drug reaction data among drugs are collected, involving 59,377 pairs of adverse reaction drug pairs, N = 567 drugs, K = 258 adverse reactions, covering basically common drugs and adverse reactions. The data collected using this collection method has high reliability. Given a drug set D = {d1, d2,..., d N}, for the adverse reaction between drug d i and d j , a vector r ij ∈ {0, 1} K is constructed to represent the adverse reaction relationship between d i and d j . If the k-th adverse reaction is caused between d i and d j , then Otherwise
[0067] Table 1
[0068]
[0069] As shown in Table 1, the present invention constructs a prediction model for adverse drug reactions among drugs using molecular structure, target, pathway, side effect, phenotype, and disease. M represents the number of attributes. In this embodiment, M = 6, and the matrix represents the feature space of the m-th attribute of the drug, and L m represents the feature dimension of the m-th attribute of the drug. Taking the feature space of the molecular structure of the drug as an example (m = 2), the association relationship between the drug and the target is collected from the DrugBank database, and the feature space of the drug target information is constructed. The feature dimension L2 of the target is 497. Therefore, the target information of drug d i can be represented by a 497-dimensional binary vector. If drug d i is associated with the j-th target, then Otherwise In addition, due to varying degrees of missing information in different attribute information of drugs, that is, in the matrix if the characteristic information of the m-th attribute of drug d has not been recorded in the drug attribute database j then represents a zero vector of dimension L m Therefore, in the m-th attribute feature space of the drug, and respectively represent the known attribute feature information and missing feature information of the drug, and respectively represent the number of drugs with known feature information and the number of drugs with missing feature information in the m-th attribute,
[0070] S2: Based on the multi-attribute data of the drug, construct a drug multi-attribute missing feature filling model based on common features and unique features;
[0071] The specific steps for constructing the drug multi-attribute missing feature filling model include:
[0072] Based on the multi-attribute data of the drug, and based on the relationship between the common features and unique features of drug attributes and the original feature space, construct a basic model;
[0073] Measure the distribution difference between unique features of different attributes through KL divergence, and under the constraint of the distribution difference, process the basic model to obtain a multi-attribute missing feature filling model based on common features and unique features.
[0074] In step S2, construct a drug multi-attribute missing feature filling model. In the feature filling model, explore the common features and unique features of drug multi-attributes. Among them, the common features represent the consistent contribution information for predicting adverse drug reactions among drugs with different attributes, and the unique features represent the unique information of each attribute, which plays a supplementary role in predicting adverse drug reactions. This step constructs a basic model of the common features and unique features of drug attributes and the original feature space, and at the same time introduces the equation constraint between the m-th attribute feature space X m of the drug and its known attribute feature information to ensure that the known attribute feature information remains unchanged during the attribute feature filling process, so as to improve the effectiveness of attribute feature filling. Therefore, the constructed basic model is the multi-attribute missing feature filling objective function, and the multi-attribute missing feature filling objective function based on common features and unique features can be written as:
[0075]
[0076] denotes the Frobenius norm of the matrix; ||·||0 denotes the matrix 0-norm, that is, the number of non-zero elements in the matrix; the matrix represents the common features of the multi-attributes of the drug, and the feature spaces of different attributes of the drug have the same common matrix; the matrix represents the unique features of the m-th attribute of the drug. The feature spaces of different attributes of the drug include their respective unique features, and L represents the dimensions of the common features and the unique features; represents the original feature space X based on the common features and the unique features in the m-th attribute m Reconstruction coefficient matrix. For the m-th attribute, each drug contains a limited number of features, which is much smaller than the feature dimension L of attribute m m , so the attribute feature space of the drug is highly sparse. Therefore, the coefficient matrix U is introduced in formula (1) m with the 0-norm constraint to control the sparsity of the original feature space reconstruction matrix (P + Q m )U m , and α m represents the sparsity regularization parameter of the coefficient matrix U m . In the constraint conditions, represents the known attribute feature information of the drug in the m-th attribute, and the marking matrix is obtained by removing the rows corresponding to the indices of the drugs with missing features in the identity matrix I N×N , represents arranging the drugs that extract the known attribute features in the original feature space X m of the attribute in index order to obtain the drug known attribute feature information matrix The constraint conditions P≥0, Q m ≥0 and U m ≥0 are used to maintain the non-negativity of the matrix.
[0077] As can be seen from formula (1), since there are differences in the feature spaces X m of different attributes of the drug, m = 1,..., M, this step decomposes the feature space of the drug attribute into common features P and unique features Q m , and the M attribute spaces share the same common features P, and different attribute spaces X m have their respective unique features Q m . The common features and the unique features are reconstructed through the sparse coefficient matrix U m . Since the feature spaces X mThe unique features contain the unique information of their respective attribute spaces and are not shared with other attribute spaces. Therefore, this step further restricts the specificity between the unique features of different attributes, provides specific supplementary information between attributes for the prediction of adverse drug reactions based on multiple attributes, and introduces the KL (Kullback-Leible) divergence to measure the distribution difference between the unique features of different attributes:
[0078]
[0079] Measure the difference degree between two unique features Q m and Q n The KL(Q m ||Q n ) ≥ 0. The smaller the difference between Q m and Q n , the smaller the value of the KL divergence. If Q m and Q n are the same, KL(Q m ||Q n ) = 0. Therefore, by measuring the specificity difference between different attribute unique matrices, a drug multi-attribute missing feature filling model is obtained, and the objective function can be written as:
[0080]
[0081] represents the Frobenius norm of the matrix, ||·||0 represents the matrix 0 norm, P is the common feature of the drug multi-attribute, Q m is the unique feature of the m-th attribute of the drug, U m is the original feature space X m reconstruction coefficient matrix based on the common feature and the unique feature in the m-th attribute, X m is the m-th attribute feature space of the drug, is the known attribute feature information of the drug, KL is the divergence, represents the marker matrix of the known attribute feature information matrix of the m-th attribute, represents arranging the drugs that extract the known attribute features in the attribute original feature space X m in the index order to obtain the known attribute feature information matrix α m represents the sparsity regularization parameter of the m-th attribute reconstruction coefficient matrix, β represents the regularization parameter of the KL divergence between different attribute unique features, α m represents the sparsity regularization parameter of the m-th attribute reconstruction coefficient matrix, β represents the regularization parameter of the KL divergence between different attribute unique features.
[0082] S3: Modify the drug multi-attribute missing value filling model with a cosine similarity regularization term, and use the Lagrangian function, the alternating direction multiplier method, and the non-negative matrix factorization method to solve the modified drug multi-attribute missing value filling model to obtain the common features and specific features of the drug multi-attributes;
[0083] The specific method for obtaining the common features and specific features of the drug multi-attributes is as follows:
[0084] Modify the drug multi-attribute missing value filling model with a cosine similarity regularization term to obtain a modified model;
[0085] Use the Lagrangian function, the alternating direction multiplier method, and the non-negative matrix factorization method to solve the modified model to obtain the common features and specific features of the filled drug attribute feature space and multi-attribute feature space, as well as the iterative update formula of the reconstruction coefficient matrix of the attribute feature space;
[0086] Iteratively update the variables of the modified model until the number of iterative updates reaches the maximum value or the model reaches the minimum change threshold to obtain the common features and specific features of the drug multi-attributes.
[0087] In step S2, the original feature space X of the drug attribute m has a feature dimension of L m , and the obtained common feature P and specific feature Q m have a dimension of L; therefore, decomposing the feature space X of the drug attribute m into the common feature P and specific feature Q m can be regarded as mapping the high-dimensional sparse feature space X m to a low-dimensional feature space, which contains the common features and specific features of the attribute space. For drug d i , represents the feature information of the m-th attribute of drug d i , and the feature representation of d i in the low-dimensional feature space can be composed of the common feature representation P i of d in all attributes and the specific feature representation i· in the m-th attribute. Therefore, based on the graph manifold regularization method, the feature representation of the drug in the low-dimensional feature space needs to retain the local geometric structure of the original attribute feature space, that is, in the low-dimensional feature space, the feature representation similarity between drug d and d i is consistent with their similarity in the original attribute feature space. In the feature space of the m-th attribute, the feature representation similarity between drug d j and d i can be expressed as: j
[0088]
[0089] Denote the vector and as the inner product of S m (d i , d j ) can be regarded as the normalized cosine similarity between the vectors and . In the low - dimensional feature space, (P + Q m ) i· and (P + Q m ) j· are the combined representations of the common features and unique features of the drugs d i and d j respectively. Then, the local geometric structure consistency regular term of the drug attribute feature space based on cosine similarity can be expressed as:
[0090]
[0091] The final model framework for filling the missing features of multiple drug attributes is as Figure 2 shown, and the final objective function is as shown in formula (6):
[0092]
[0093] γ m represents the regularization parameter of the cosine similarity regular term for the m - th attribute, represents the l2 - norm of the vector.
[0094] Based on the augmented Lagrangian function and the Alternating Direction Method of Multipliers (ADMM) and the non - negative matrix factorization optimization method, the filled drug attribute feature space X m , the common features P and unique features Q of the multi - attribute feature space m , and the iterative update equations of the reconstruction coefficient matrix U m of the attribute feature space are obtained. By iteratively updating the above model variables, setting the maximum number of iterations or the minimum change threshold of the objective function, the optimal solutions of the above model variables are finally obtained, and the common features and unique features of the multiple drug attributes are obtained.
[0095] S4: Based on the common features and unique features of the multiple drug attributes, combined with the adverse reaction data, construct a prediction model;
[0096] Based on the common features P and unique features Q of the multiple drug attributes obtained in step 3 m, by exploring the influence of different attributes on the prediction of adverse drug reactions, a multi-attribute-based adverse drug reaction prediction model is established to reveal the implicit laws between multi-attributes and adverse reactions.
[0097] For drug d i and d j Regarding the adverse reaction between them, the binary vector r ij ∈{0,1} K represents the adverse reaction relationship between d i and d j . Based on the common features and unique features of drug multi-attributes optimized in step 3, the vectors P i· and P j. respectively represent the common features of d i and d j . The vectors and respectively represent the unique features of the m-th attribute of d i and d j . Since the different attribute feature spaces of drugs share the same common features and contain their respective unique features, the adverse reaction between drug d i and d j can be caused by the joint contribution of common features and unique features. Therefore, the overall objective function of the adverse drug reaction prediction model based on common features and unique features can be written as:
[0098]
[0099] represents the contribution of common features to the adverse reaction between drugs, represents the contribution of unique features to the adverse reaction between drugs. To estimate the vector , the common feature - adverse reaction tensor is introduced. The tensor element represents the potential relationship between the i-th common feature, the j-th common feature and the k-th adverse reaction. Therefore, the vector can be expressed as:
[0100]
[0101] The parameter λ represents the contribution degree of the common features P i· and P j· to the adverse reaction between d i and d j , × k represents the product of the k-th order of the tensor and the vector. On the other hand, since each attribute has its own unique features, the tensor ε m is constructed to represent the potential relationship between the unique features of the m-th attribute and the adverse reaction. Therefore, the vector It may be jointly contributed by the unique characteristics of each of the M attributes:
[0102]
[0103] Parameter w m represents the unique characteristic of the m-th attribute and For d i and d j the contribution degree of the adverse reaction between them. Therefore, the framework of the adverse reaction prediction model between drugs based on common characteristics and unique characteristics is as Figure 3 shown.
[0104] Based on the high-order tensor low-rank CP decomposition method, the tensor and ε m is decomposed into the sum of tensors of rank 1, and the adverse reaction r i between drugs d j and d ij is estimated. Using the stochastic gradient descent method, the implicit parameters of the model are iteratively optimized to update the parameters of the model.
[0105] S5: Obtain the multi-attribute data between any two drugs, calculate the common characteristics and unique characteristics of the drug multi-attributes, and input them into the prediction model to obtain the adverse reaction prediction data between drugs.
[0106] Given the multi-attribute information a of any two drugs d b and d and predict the adverse reaction between drugs d a and d b Based on the model framework filled with drug multi-attribute features, the missing attribute features of drugs d a and d b are filled, and their common characteristics and unique characteristics of each attribute are obtained. Based on the adverse reaction prediction model between drugs with common characteristics and unique characteristics, the adverse reaction prediction between drugs d a and d b can be expressed as:
[0107]
[0108] In this embodiment, using the proposed method for predicting adverse reactions between drugs based on filling drug multi-attribute features, the adverse reactions between drugs can be predicted, and some of the prediction results are supported by relevant literature. The prediction results can provide data support for the research on adverse reactions between drugs and the safety research of new drugs based on biological experimental methods. Figure 4 、 Figure 5 Some of the predicted adverse reactions between drugs are given.Figure 4 The combined use of voriconazole (a triazole antifungal agent for preventing aspergillosis and candidiasis) and dexamethasone (a synthetic corticosteroid used to treat rheumatoid arthritis, cerebral edema, and acute pulmonary edema) will cause adverse reactions such as sepsis, vision attenuation, and osteoporosis. The reason for the adverse reactions is that the two drugs have similar side effects, and some substructures act on the same targets, pathways, and diseases. Figure 5 The combined use of sevelamer (a non-absorbable polyamine used to prevent hyperphosphatemia) and furosemide (a sulfanilic acid derivative used to treat congestive heart failure) will cause adverse reactions such as cardiac arrest, bradycardia, and adynamic ileus.
[0109] The adverse reaction prediction method based on multi-attribute feature filling disclosed in this embodiment fills in the missing features of drugs by constructing a drug multi-attribute missing feature filling model, and uses the filled attribute features to carry out adverse reaction prediction between drugs, which not only improves the accuracy of adverse reaction prediction between drugs, but also promotes the experimental research on adverse reactions between drugs, ensuring the safety of drug use.
[0110] Example Two
[0111] This embodiment discloses an adverse reaction prediction system between drugs with multi-attribute feature filling. This embodiment is to implement the prediction method in Example One, including a data acquisition module, a filling model construction module, an analysis module, a prediction model construction module, and a prediction module;
[0112] The data acquisition module is used to acquire adverse reaction data between drugs and drug multi-attribute data;
[0113] The filling model construction module is used to construct a drug multi-attribute missing feature filling model based on common features and unique features based on the drug multi-attribute data;
[0114] The analysis module is used to correct the drug multi-attribute missing filling model through a cosine similarity regular term, and use the Lagrangian function, the alternating direction multiplier method, and the non-negative matrix factorization method to solve the corrected drug multi-attribute missing filling model to obtain the common features and unique features of the drug multi-attributes;
[0115] The prediction model construction module is used to construct a prediction model based on the common features and unique features of drug multi-attributes, combined with the adverse reaction data;
[0116] The prediction module is used to acquire the multi-attribute data between any two drugs, calculate the common features and unique features of the drug multi-attributes, and input them into the prediction model to obtain the adverse reaction prediction result between drugs.
[0117] Example 3
[0118] This embodiment discloses a computer storage medium, on which a computing program is stored. When the computer program is executed by a processor, the method described in Example 1 is implemented.
[0119] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0120] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be realized by computer program issued instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be realized. These computer program issued instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the issued instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0121] These computer program issued instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing devices to work in a specific manner, so that the issued instructions stored in the computer-readable memory generate a manufactured product including the issued instruction device, and the issued instruction device realizes the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0122] These computer program issued instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the issued instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0123] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting adverse drug reactions among drugs with multi-attribute feature filling, characterized in that, The method steps include: Obtain adverse drug reaction data and multi-attribute drug data; Based on the multi-attribute drug data, construct a filling model for missing multi-attribute drug features based on common features and unique features; The specific expression of the multi-attribute missing feature filling model based on common features and unique features is: denotes the Frobenius norm of the matrix, ||·||0 denotes the matrix 0-norm, P is the common feature of the multi-attributes of the drug, and Q m is the unique feature of the m-th attribute of the drug, and U m is the original feature space X based on the common feature and the unique feature in the m-th attribute m reconstruction coefficient matrix, X m is the feature space of the m-th attribute of the drug, is the known attribute feature information of the drug, KL is the divergence, denotes the known attribute feature information of the drug in the m-th attribute marking matrix, denotes the drug that extracts the known attribute features in the original feature space X of the attribute m arranged in index order to obtain the known attribute feature information of the drug α m denotes the sparsity regularization parameter of the reconstruction coefficient matrix of the m-th attribute, and β denotes the regularization parameter of the KL divergence between the unique features of different attributes; Modify the multi-attribute drug missing filling model through a cosine similarity regularization term, and use the Lagrangian function, alternating direction multiplier method, and non-negative matrix factorization method to solve the modified multi-attribute drug missing filling model to obtain the common features and unique features of the multi-attribute drug; The specific method for obtaining the common features and unique features of the multi-attribute drug is: Modify the multi-attribute drug missing filling model through a cosine similarity regularization term to obtain a modified model; Use the Lagrangian function, alternating direction multiplier method, and non-negative matrix factorization method to solve the modified model to obtain the common features and unique features of the filled drug attribute feature space and multi-attribute feature space, as well as the iterative update formula of the reconstruction coefficient matrix of the attribute feature space; Iteratively update the variables of the modified model until the number of iterative updates reaches the maximum value or the model reaches the minimum change threshold to obtain the common features and unique features of the multi-attribute drug; Based on the common features and unique features of the multi-attribute drug, combine the adverse reaction data to construct a prediction model; The specific expression of the modified model is: S m (d i ,d j ) can be regarded as the normalized cosine similarity between vectors and , (P + Q m ) i· and (P + Q m ) j· are the combined representations of the common and unique features of drug d i and d j , γ m represents the regularization parameter of the m-th attribute cosine similarity regularization term, represents the l2 norm of the vector; Obtain the multi-attribute data between any two drugs, calculate the common features and unique features of the multi-attribute drug, and input them into the prediction model to obtain the adverse reaction prediction result between the drugs.
2. The method for predicting adverse drug reactions among drugs filled with multi-attribute features according to claim 1, wherein The specific steps for constructing a filling model for missing multi-attribute drug features include: Based on the multi-attribute drug data, and based on the relationship between the common features and unique features of drug attributes and the original feature space, construct a basic model; Measure the distribution difference between the unique features of different attributes through KL divergence, and under the constraint of the distribution difference, process the basic model to obtain a filling model for missing multi-attribute drug features based on common features and unique features.
3. The method for predicting adverse drug reactions filled with multi-attribute features according to claim 1, characterized in that, The multi-attribute data includes molecular structure data, target data, pathway data, side effect data, phenotype data, and disease data.
4. The method for predicting adverse drug reactions among drugs filled with multi-attribute features according to claim 1, characterized in that, The specific expression of the prediction model is: For the contribution of common features to adverse drug reactions, For the contribution of unique features to adverse drug reactions, r ij Is the adverse drug reaction relationship between drugs.
5. The method for predicting adverse drug reactions among drugs filled with multi-attribute features according to claim 4, characterized in that, The said The specific expression of which is: The said has the specific expression as follows: P i. is the multi-attribute common feature of drug d i P, the multi-attribute common feature of drug d j. is the multi-attribute common feature of drug d j is the specific feature of the m-th attribute of d i is the specific feature of the m-th attribute of d, w j m represents the specific feature of the m-th attribute and for d i and d j the contribution degree to the adverse reaction between them, λ represents the common feature P i. and P j. for d i and d j the contribution degree to the adverse reaction between them is the common feature - adverse reaction tensor, ε m is the potential relationship between the specific feature of the m-th attribute and the adverse reaction 6. A drug - drug adverse reaction prediction system with multi - attribute feature filling is used to execute the drug - drug adverse reaction prediction method with multi - attribute feature filling as described in any one of claims 1 to 5, characterized in that, It includes a data acquisition module, a filling model construction module, an analysis module, a prediction model construction module, and a prediction module; The data acquisition module is used to obtain adverse drug reaction data and multi-attribute drug data; The filling model construction module is used to construct a filling model for missing multi-attribute drug features based on common features and unique features based on the multi-attribute drug data; The analysis module is used to modify the multi-attribute drug missing filling model through a cosine similarity regularization term, and use the Lagrangian function, alternating direction multiplier method, and non-negative matrix factorization method to solve the modified multi-attribute drug missing filling model to obtain the common features and unique features of the multi-attribute drug; The prediction model construction module is used to construct a prediction model based on the common features and unique features of the multi-attribute drug, combined with the adverse reaction data; The prediction module is used to obtain multi-attribute data between any two drugs, calculate the common features and unique features of the drug multi-attributes, and input them into the prediction model to obtain the adverse reaction prediction result between the drugs.
7. A computer storage medium, on which a computing program is stored, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Propaganda and education pushing method and system based on multi-source information fusion
CN114969557A
Disease auxiliary decision-making system based on personalized state space progress model
CN115019960A