Method for predicting drug performance and efficacy based on artificial intelligence algorithm

By constructing the performance and efficacy matrix of traditional Chinese medicinal materials and using artificial intelligence algorithms to predict the potential performance and efficacy of compounds, the problems of single safety and research paths in drug function relocation are solved, and the safe and efficient expansion of new use of old drugs is achieved.

CN120388646APending Publication Date: 2025-07-29TIANJIN TASLY DIGITAL INTELLIGENCE CHINESE MEDICINE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410083092.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing technology lacks safety guarantee and a single research path in drug function relocation, which limits the development speed of new use of old drugs.

Method used

Based on the digital description of traditional Chinese medicine performance, an artificial intelligence model is built, and the performance and efficacy matrix of traditional Chinese medicine materials is used to predict the potential performance and efficacy of small-molecular compounds through the performance and efficacy matrix of traditional Chinese medicine materials, and an evaluation model is established.

Benefits of technology

Through the multi-dimensional performance and efficacy prediction of traditional Chinese medicinal materials, potential new indications of compounds were discovered, which improved the safety and efficiency of new use of old drugs and expanded the scope of application of drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388646A_ABST
    Figure CN120388646A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting drug performance and efficacy based on an artificial intelligence algorithm. The method comprises the following steps: 1) respectively constructing a performance matrix and an efficacy matrix according to the performance and efficacy of traditional Chinese medicinal materials; 2) generating performance vectors and efficacy vectors of the corresponding traditional Chinese medicinal materials according to the performance matrix and the efficacy matrix; 3) establishing a significant chemical component set and a non-significant chemical component set of efficacy / performance of each medicinal material; 4) taking the Chinese medicinal material molecular structure characteristics of the efficacy / performance of each medicinal material as well as the performance and efficacy of the Chinese medicinal material as a sample training model to obtain a first efficacy / performance evaluation model of the efficacy / performance of the medicinal material; taking molecular structure characteristics in a significant and non-significant chemical component set of the efficacy / performance of each medicinal material and the efficacy and performance of the corresponding medicinal material as a sample training model as a corresponding second efficacy / performance evaluation model; and 5) inputting the molecular structure characteristics of the chemical molecules to be tested into the first and second efficacy / performance evaluation models to obtain the efficacy / performance of the chemical molecules to be tested.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of digital processing of drug properties and efficacy. Specifically, the present invention relates to a method for predicting drug properties and efficacy based on artificial intelligence algorithms. Background Art

[0002] Drug function repositioning, also known as repurposing old drugs, is a strategy for discovering new indications of old drugs or investigational drugs beyond their original approved indications, and expanding their scope of application and uses. Drug function repositioning can not only quickly find effective drugs for diseases lacking treatment means, but also greatly extend the drug life cycle.

[0003] Current methods for repurposing old drugs: One is based on clinical practice, that is, in clinical practice, new uses of drugs are discovered by observing and analyzing the actual treatment rules of clinicians; the other is based on prediction methods, that is, computer algorithms and big data analysis and other technologies are used to predict new uses of drugs. For the former clinical practice, due to the lack of sufficient clinical data, the safety issues of repurposing old drugs cannot be guaranteed. For the latter prediction methods, the principles mostly revolve around the functions of chemical structures or action targets in new indications, and the research paths are single. The deficiencies of these two methods significantly limit the development speed of repurposing old drugs. Therefore, there is an urgent need to find and discover new methods for repurposing old drugs. Summary of the Invention

[0004] In view of this, the present invention provides a method for predicting drug properties and efficacy based on artificial intelligence algorithms, which is applied to the expression of traditional Chinese medicine characteristics for western medicines. Traditional Chinese medicine properties are used to reflect the nature and characteristics of drug actions and are the core of the basic theory of traditional Chinese medicine, mainly including four natures, five flavors, meridian tropism, and toxicity. Traditional Chinese medicine properties are a generalization of the characteristics of drug actions based on efficacy and can be used as a theoretical tool for drug actions. After fully understanding traditional Chinese medicine properties, traditional Chinese medicine experts have accumulated rich and successful clinical practice experience in the treatment of different patients, which fully shows that the theoretical system of traditional Chinese medicine properties can guide the rational use of drugs and is also significantly correlated with the clinical efficacy of drugs. Therefore, this method digitally describes the properties of each dimension corresponding to the core traditional Chinese medicine materials and related compounds and constructs a corresponding artificial intelligence model, aiming to predict the potential properties of each small molecule compound through this model, thereby predicting the potential new indications of the compound, and carrying out corresponding evaluations and verifications, exploring a new method for the successful repurposing of old drugs.

[0005] The technical solution of the present invention is as follows:

[0006] A method for predicting drug properties and efficacy based on artificial intelligence algorithms, the steps of which include:

[0007] 1) Construct a performance matrix of traditional Chinese medicine materials according to the properties of traditional Chinese medicine materials, and construct an efficacy matrix of traditional Chinese medicine materials according to the efficacy of traditional Chinese medicine materials;

[0008] 2) Determine the properties and efficacy of the corresponding Chinese medicinal materials according to the medicinal information of each selected Chinese medicinal material; then, based on the property matrix of the Chinese medicinal materials and the properties of each selected Chinese medicinal material, quantify the selected Chinese medicinal materials to generate the property vectors of the corresponding Chinese medicinal materials; based on the efficacy matrix of the Chinese medicinal materials and the efficacy of each selected Chinese medicinal material, quantify the selected Chinese medicinal materials to generate the efficacy vectors of the corresponding Chinese medicinal materials.

[0009] 3) For each selected Chinese medicinal material with known properties and efficacy, obtain the set of chemical components of the corresponding Chinese medicinal material; based on the set of chemical components of the Chinese medicinal material, establish the set of significant chemical components and non-significant chemical components of the properties of each medicinal material, and the set of significant chemical components and non-significant chemical components of the efficacy of each medicinal material.

[0010] 4) Calculate the average value of the molecular structure characteristics of the set of chemical components of each Chinese medicinal material as the molecular structure characteristic of the corresponding Chinese medicinal material. Take the molecular structure characteristic of each Chinese medicinal material with the i-th medicinal material efficacy and its efficacy vector as a sample to obtain the first efficacy sample set, and divide it into the first efficacy training set and the first efficacy test set; use the first efficacy training set to train and optimize P artificial intelligence algorithm models, and then use the first test set for testing. Select the artificial intelligence algorithm model with the best test effect as the first efficacy evaluation model for the i-th medicinal material efficacy; i = 1 to Q, where Q is the total number of medicinal material efficacies; take the molecular structure characteristic of each Chinese medicinal material with the w-th medicinal material property and its property vector as a sample to obtain the first property sample set, and divide it into the first property training set and the first property test set; use the first property training set to train and optimize P artificial intelligence algorithm models, and then use the first property test set for testing. Select the artificial intelligence algorithm model with the best test effect as the first property evaluation model for the w-th medicinal material property; w = 1 to W, where W is the total number of medicinal material properties.

[0011] 5) Calculate the molecular structure characteristics of each chemical component in the significant chemical component set and the non-significant chemical component set of the efficacy of each medicinal material, take the molecular structure characteristics of each chemical component in the significant chemical component set and the non-significant chemical component set of the efficacy of the i-th medicinal material and the efficacy vector of the i-th medicinal material as a sample, obtain a second efficacy sample set, and divide it into a second efficacy training set and a second efficacy test set; use the second efficacy training set to train and optimize the P artificial intelligence algorithm models, then use the second efficacy test set to test, and select the artificial intelligence algorithm model with the best test effect as the second efficacy evaluation model for the efficacy of the i-th medicinal material; i=1~Q; take the molecular structure characteristics of each chemical component in the significant chemical component set and the non-significant chemical component set of the performance of the w-th medicinal material and the performance vector of the w-th medicinal material as a sample, obtain a second performance sample set, and divide it into a second performance training set and a second performance test set; use the second performance training set to train and optimize the P artificial intelligence algorithm models, then use the second performance test set to test, and select the artificial intelligence algorithm model with the best test effect as the second performance evaluation model for the performance of the w-th medicinal material; w=1~W;

[0012] 6) For each chemical molecule to be tested in the tested Chinese medicinal material, a molecular structure feature of the chemical molecule to be tested is generated and inputted into a first efficacy evaluation model and a second efficacy evaluation model corresponding to the i-th medicinal material efficacy, and a first performance evaluation model and a second performance evaluation model corresponding to the j-th performance, respectively, to determine whether the chemical molecule to be tested has the i-th medicinal material efficacy and the j-th performance; then, the medicinal properties and efficacy of the Chinese medicinal material to be tested are determined based on the detection results of each chemical molecule to be tested in the Chinese medicinal material to be tested.

[0013] Furthermore, the drug information of the Chinese medicinal materials includes information on the four properties, five flavors, meridians, toxicity and efficacy of the Chinese medicinal materials; the performance of the Chinese medicinal materials includes drug characteristics, drug properties and flavors, drug meridians and drug toxicity; the drug characteristics include cold, very cold, slightly cold, hot, very hot, warm, slightly warm, cool, slightly cool, and neutral; the drug properties and flavors include pungent, bitter, sweet, sour, salty, light, and astringent; the drug meridians include lung, large intestine, stomach, spleen, heart, small intestine, bladder, kidney, pericardium, triple burner, gallbladder, and liver. The toxicity of the drugs includes toxic, highly toxic, slightly toxic, and non-toxic; the effects of the Chinese medicinal materials include digestion, wind-dispelling, heat-clearing, detoxification, pain relief, cold-dispelling, expectorant, tranquilizing, purgative, external use, blood regulation, dampness-removing, astringent, antiemetic, anthelmintic, anesthetic, insecticide, softening, diuretic, deficiency-tonifying, tonic, diuretic and dampness-removing, calming the liver and relieving wind, relieving exterior symptoms, removing toxins, removing dead tissue and promoting tissue regeneration, calming the liver, tonifying, hemostasis, attacking toxins, killing insects and relieving itching, removing rheumatism, promoting blood circulation and removing blood stasis, regulating qi, resolving phlegm, relieving cough and relieving asthma, opening the orifices, dispelling insects, drying dampness, dispelling heat, and vomiting.

[0014] Further, the method for obtaining the molecular structure characteristics of traditional Chinese medicine is as follows: Each chemical component in the chemical component set of traditional Chinese medicine is represented by a bit string of length L. The chemical component set includes m chemical components. The bit strings representing these m chemical components are used to form an m*L two-dimensional matrix. The sum of the elements in each column of the m*L two-dimensional matrix is divided by the number of chemical components m to obtain the molecular structure characteristics of the traditional Chinese medicine.

[0015] Further, the length L = 2048.

[0016] Further, the method for establishing the significant chemical component set and non-significant chemical component set of the efficacy / performance of each traditional Chinese medicine is as follows: First, for the efficacy / performance of each traditional Chinese medicine, j traditional Chinese medicines with the efficacy of this traditional Chinese medicine are obtained from the selected traditional Chinese medicines with M known efficacy / performance. For each chemical component in the chemical component set, the number of traditional Chinese medicines N in which this chemical component appears in the corresponding chemical component sets of the M selected traditional Chinese medicines with known performance / efficacy is obtained, and the number of traditional Chinese medicines k with the efficacy of this traditional Chinese medicine is obtained from the N traditional Chinese medicines obtained. Then, the significance value of this chemical component with respect to the efficacy of this traditional Chinese medicine is calculated. If the significance value p(k, M, j, N) of this chemical component with respect to the efficacy / performance of this traditional Chinese medicine is less than the set threshold, then this chemical component is a significant chemical component for the efficacy / performance of this traditional Chinese medicine and is added to the significant chemical component set of the efficacy / performance of this traditional Chinese medicine; otherwise, this chemical component is added to the non-significant chemical component set of the efficacy / performance of this traditional Chinese medicine.

[0017] Further, the P artificial intelligence algorithm models include a deep neural network model, a K-nearest neighbor model, a naive Bayes model, a random forest model, a support vector machine model, and an extreme gradient boosting model.

[0018] Further, the number of decision trees n_estimators in the random forest model is set to (10, 1000), the node splitting uses 'Gini value gini' and 'entropy', the maximum number of features uses'sqrt' and 'log2', and the maximum depth range of the decision tree is set to (1, 3, 5, 7, 9); the kernel function in the support vector machine model is set to RBF, the kernel function coefficient is set to (-3, -15), and the penalty function is set to (-3, -15); the number of decision trees n_estimators in the extreme gradient boosting model is set to (50, 400), the maximum depth range of the decision tree is set to (1, 11), and the learning rate is set to 0.05; the alpha in the naive Bayes model is set in the range of (0.01, 2); the number of nearest neighbors n_neighbors in the K-nearest neighbor model is set to (1, 21); the learning rate in the deep neural network model is set to adaptive, the hidden layer is set in the range of 1 to 4 layers, each hidden layer has 50 neurons, the activation function is set to the identity function identity, the logistic function logistic, the hyperbolic tangent tanh or the rectified linear unit relu, the adaptive moment estimation algorithm adam is used as the optimization algorithm, the loss function is the cross-entropy loss function, and the maximum number of iterations is 500 times.

[0019] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the above method.

[0020] A computer-readable storage medium, on which a computer program is stored, characterized in that the steps of the above method are realized when the computer program is executed by a processor.

[0021] The method flow of the present invention is as Figure 1 shown, and its steps include:

[0022] (A) By sorting out the content of traditional Chinese medicinal materials in the Pharmacopoeia of the People's Republic of China (2020 Edition), including the information of the four natures, five flavors, meridian tropism, toxicity and efficacy of the medicinal materials, a medicinal material property and efficacy database is integrated and formed.

[0023] The properties of the medicinal materials include drug characteristics: cold, severely cold, slightly cold, hot, severely hot, warm, slightly warm, cool, slightly cool, neutral; the nature and flavor of the drugs: pungent, bitter, sweet, sour, salty, bland, astringent; the meridians entered by the drugs: lung, large intestine, stomach, spleen, heart, small intestine, bladder, kidney, pericardium, triple energizer, gallbladder, liver; the toxicity of the drugs: toxic, highly toxic, slightly toxic, non-toxic. Subsequently, digital matching is performed on the properties of the traditional Chinese medicines contained in the traditional Chinese medicine library. The matching form is that if a medicinal material has a certain property, it is marked as 1, and if it does not have this property, it is marked as 0, obtaining the property matrix (616, 33) of the medicinal materials. For example, if a certain drug has bitter, sweet, astringent, slightly warm, enters the liver, heart, and kidney, and is toxic, then the property vector of the medicinal material is (0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 1, 1, 0, 0, 0, 1, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 0, 1, 1, 0, 0, 0).

[0024] The efficacy of the medicinal materials includes that traditional Chinese medicines are divided into 38 categories, including promoting digestion, dispelling wind, clearing heat, detoxifying, relieving pain, dispelling cold, resolving phlegm, calming the mind, purging, external use, regulating blood, removing dampness, astringing, stopping vomiting, expelling worms, anesthetic, killing insects, softening hardness, promoting diuresis, tonifying deficiency, tonifying, promoting diuresis and percolating dampness, calming endogenous wind, inducing sweating, removing toxin and promoting granulation, calming the liver, tonifying deficiency, stopping bleeding, attacking toxin, killing insects and relieving itching, dispelling wind-dampness, promoting blood circulation to remove stasis, regulating qi, resolving phlegm and relieving cough and asthma, resuscitating, expelling insects, drying dampness, dispelling summer heat, inducing vomiting; each traditional Chinese medicine in the traditional Chinese medicine library is classified according to its efficacy, obtaining the efficacy matrix (616, 38) of the medicinal materials. Subsequently, digital matching is performed on the efficacy of the traditional Chinese medicines contained in the traditional Chinese medicine library. The matching form is that if a medicinal material has a certain efficacy, it is marked as 1, and if it does not have this efficacy, it is marked as 0. For example, if a certain medicinal material has the efficacy of attacking toxin, killing insects and relieving itching, promoting blood circulation to remove stasis, promoting diuresis, calming endogenous wind, expelling insects, dispelling wind-dampness, killing insects, and purging, then the efficacy vector of the medicinal material is: (0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 1, 0, 0, 0).

[0025] (B) For traditional Chinese medicines with known properties and efficacy, construct a matrix of traditional Chinese medicines and their properties and efficacy;

[0026] For traditional Chinese medicines with known properties and efficacy, the traditional Chinese medicines include those known to have the said properties and efficacy, and obtain the chemical composition set of the said traditional Chinese medicines;

[0027] According to the chemical composition set of the said traditional Chinese medicines, obtain the significant chemical composition set corresponding to each property and efficacy of the medicinal materials;

[0028] Based on the chemical composition set of the said traditional Chinese medicines and the significant chemical composition set of the properties and efficacy of the medicinal materials, establish an evaluation method for drug properties and efficacy.

[0029] In one embodiment, the traditional Chinese medicine includes traditional Chinese medicine that is known not to have the properties and efficacy of the medicinal materials.

[0030] In one embodiment, the molecular structure characteristics of the traditional Chinese medicine are obtained based on the average value of the molecular structure characteristics of the chemical composition set in the traditional Chinese medicine. The specific operation includes: using a bit string of length 2048 (1, 0, 0, 0, 0, 0, 0, 0, 0,...) to represent the molecular structure of the chemical composition, which is obtained by calculation using RDkit. The molecular structure characteristics of the traditional Chinese medicine are composed of 2048-bit characteristics of m chemical components. The sum of the elements in each column of the m*2048 two-dimensional matrix is divided by the number of chemical components m to obtain the molecular structure characteristics of the traditional Chinese medicine.

[0031] In one embodiment, the molecular structure characteristics of the medicinal materials are corresponded with the corresponding property and efficacy data to construct a correspondence matrix between the molecular structure characteristics of traditional Chinese medicine and the properties and efficacy of the medicinal materials. Each row represents the molecular structure characteristics of a medicinal material plus the properties and efficacy of the medicinal material, and each property and efficacy data of each medicinal material is used as a column for training an artificial intelligence model, such as a deep learning model.

[0032] In one embodiment, a set of significant chemical components for the properties and efficacy of the medicinal materials is established. The specific operation includes: obtaining the total amount of the medicinal materials M, obtaining the number j of medicinal materials with the said properties and efficacy, obtaining the number N of medicinal materials in which each chemical component appears, obtaining the number k of medicinal materials with the said properties and efficacy in which each chemical component appears. The calculation of the significant chemical molecule is as follows: Obtain the significance of each chemical component having the said properties and efficacy, retain the chemical components with a value less than the threshold of 0.05 in the properties and efficacy of the medicinal materials, and obtain the set of significant chemical components for the properties and efficacy of the medicinal materials.

[0033] In one embodiment, the 2048-bit molecular structure characteristics of each significant chemical molecule are obtained through RDkit (https: / / www.rdkit.org / docs / index.html), and a bit string of length 2048 (1, 0, 0, 0, 0, 0, 0, 0, 0,...) is used to represent the molecular structure characteristics of the significant chemical components.

[0034] In one embodiment, a correspondence matrix between the molecular structure characteristics of the significant chemical components and the properties and efficacy of the medicinal materials is established for training an artificial intelligence model, such as a deep learning model.

[0035] (C) A model for evaluating the properties and efficacy of drugs established.

[0036] In one embodiment, the artificial intelligence algorithm includes: Deep Neural Network (DNN) algorithm, K-nearest neighbors (KNN) algorithm, Naive Bayes (NB) algorithm, Random Forest (RF) algorithm, support vector machine (SVM) algorithm, and extreme Gradient Boosting (XGB).

[0037] In one embodiment, 80% of the sample data is randomly selected as the training set, and 20% of the sample data is used as the test set. The training set and the test set are used for training and evaluating the artificial intelligence algorithm model.

[0038] In one embodiment, 12 evaluation models for the properties and efficacy of the medicinal materials are obtained.

[0039] In one embodiment, the number of decision trees n_estimators is set to (10, 1000), the node splitting uses 'Gini value gini' and 'entropy entropy', the maximum number of features uses'square root sqrt' and 'logarithm log2', and the maximum depth range of the decision tree is set to (1, 3, 5, 7, 9). The data in the training set is used to train the random forest model.

[0040] In one embodiment, the kernel function is set to RBF (Radial Basis Function, RBF), the kernel function coefficient is set to (-3, -15), and the penalty function is set to (-3, -15). The data in the training set is used to train the support vector machine model.

[0041] In one embodiment, the number of decision trees n_estimators is set to (50, 400), the maximum depth range of the decision tree is set to (1, 11), and the learning rate is set to 0.05. The data in the training set is used to train the extreme gradient boosting model.

[0042] In one embodiment, the range of alpha is set to (0.01, 2). The data in the training set is used to train the Naive Bayes model.

[0043] In one embodiment, the number of nearest neighbors n_neighbors is set to (1, 21). The data in the training set is used to train the K-nearest neighbor model.

[0044] In one embodiment, the learning rate is set to adaptive, the range of the hidden layer is set from 1 to 4 layers, each hidden layer has 50 neurons, the activation function is set to (identity function, logistic function, tanh function, relu function), the adaptive moment estimation (adam) algorithm is used as the optimization algorithm, the loss function is the cross-entropy loss function, the maximum number of iterations is 500 times, and the deep neural network model is trained using the data in the training set.

[0045] In one embodiment, the test set is used to evaluate the artificial intelligence algorithm model. Based on the evaluation metrics, the model with the highest f1score value of the medicinal material performance and efficacy model is obtained, and the optimal evaluation model of the medicinal material performance and efficacy is obtained using the evaluation method.

[0046] (D) A method for establishing the evaluation of drug performance and efficacy.

[0047] For the chemical molecule to be tested, the 2048-bit molecular structure features of the chemical molecule to be tested are obtained;

[0048] The 2048-bit molecular structure features of the chemical molecule to be tested are used as the input of the artificial intelligence model, and the performance and efficacy of the medicinal material are used as the output.

[0049] When the prediction result of the performance or efficacy model is 1, it is considered that the chemical molecule to be tested has such performance or efficacy, otherwise it is set to 0, indicating that the chemical molecule to be tested does not have such performance or efficacy. Whether the chemical molecule to be tested has the corresponding traditional Chinese medicine performance and efficacy is judged according to the model prediction result, and then the new indications of the molecule to be tested are judged on the basis of traditional Chinese medicine theory.

[0050] The beneficial effects of the present invention are reflected in:

[0051] Based on the underlying logic of the performance and efficacy of traditional Chinese medicine, on the basis of digitizing traditional Chinese medicine information resources, a method for finding new indications for small molecule compounds is constructed through an artificial intelligence multi-model optimization system. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of the method flow of the present invention.

[0053] Figure 2 It is a schematic diagram of the process for predicting and evaluating drug performance and efficacy of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0054] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention. For those skilled in the art, various corresponding changes and deformations can be given according to the above technical solutions and concepts, and all such changes and deformations should be included in the scope of protection of the claims of the present invention.

[0055] Embodiment 1:

[0056] In this embodiment, the efficacy of the drug is predicted based on the method of the present invention, and the calculated prediction results are verified by literature and historical data according to the literature research. As Figure 2 shown, the specific content is as follows:

[0057] Step 1: Establish a Chinese medicinal material library

[0058] Sort out the content of Chinese medicinal materials in the Pharmacopoeia of the People's Republic of China (2020 Edition), including the information of the four natures, five flavors, meridian tropism, toxicity and efficacy of the medicinal materials, and integrate them to form a Chinese medicinal material database.

[0059] Step 2: Construct a Chinese medicinal material efficacy matrix

[0060] According to the above Chinese medicinal material library, summarize and classify the efficacy of Chinese medicinal materials to construct a medicinal material efficacy matrix.

[0061] The efficacy of traditional Chinese medicine is a highly generalized description of the therapeutic and health care effects of drugs under the guidance of traditional Chinese medicine theory. By referring to the list of the efficacy of traditional Chinese medicine, the efficacy of traditional Chinese medicine is divided into 38 categories, including promoting digestion, expelling wind, clearing heat, detoxifying, relieving pain, dispelling cold, resolving phlegm, calming the mind, purging, external use, regulating blood, removing dampness, astringing, stopping vomiting, expelling parasites, anesthetic, insecticidal, softening hardness, promoting diuresis, tonifying deficiency, tonifying, promoting diuresis and removing dampness, suppressing liver wind, inducing sweating, removing toxin and promoting granulation, suppressing liver, tonifying deficiency, stopping bleeding, attacking toxin, insecticidal and relieving itching, dispelling wind and dampness, promoting blood circulation and removing stasis, regulating qi, resolving phlegm, relieving cough and asthma, inducing resuscitation, expelling insects, drying dampness, dispelling summer heat, inducing vomiting; then digitally match the efficacy of traditional Chinese medicine contained in the Chinese medicinal material library to obtain the efficacy matrix (616, 38) of the medicinal materials. For example, cnidium fruit has the efficacy of drying dampness and expelling wind, insecticidal and relieving itching, warming the kidney and strengthening yang, so the efficacy vector of cnidium fruit is (0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 1, 0, 0, 0).

[0062] The present invention digitally matches the traditional Chinese medicine properties or efficacy of the traditional Chinese medicines contained in the traditional Chinese medicine library. The matching form is that if a medicinal material has a certain property, it is marked as 1, and if it does not have the property, it is marked as 0, obtaining the property matrix (616, 33) of the medicinal materials; if it has the efficacy, it is marked as 1, and if it does not have the efficacy, it is marked as 0, obtaining the efficacy matrix (616, 38) of the drugs.

[0063] Step 3: Construct a drug efficacy model

[0064] At the traditional Chinese medicine level, the relevant components (compounds) of each traditional Chinese medicine obtained are respectively represented by the average value of 2048-bit bits; at the compound level, each chemical molecule is respectively represented by 2048-bit bits.

[0065] Taking the average value of the molecular structure characteristics of the chemical component set in the traditional Chinese medicine, the molecular structure characteristics of the traditional Chinese medicine are obtained. Each chemical component is respectively represented by 2048-bit bits (1, 0, 0, 0, 0, 0, 0, 0, 0,......). The molecular structure characteristics of the traditional Chinese medicine are composed of 2048-bit characteristics of m chemical components. The sum of the elements in each column of the m*2048 two-dimensional matrix is divided by the number of chemical components n to obtain the molecular structure characteristics (0.0085, 0.1888, 0.0042, 0.0042, 0.0042, 0.0171,......) of the traditional Chinese medicine.

[0066] The set of traditional Chinese medicinal materials with a certain efficacy (such as "soothing the nerves") is used as the positive data set for training the model, and the set of traditional Chinese medicinal materials without a certain efficacy (such as "soothing the nerves") is used as the negative data set for training the model. Therefore, for 38 kinds of efficacy, the positive data set and the negative data set for each efficacy are respectively constructed and divided into a training set and a test set according to the ratio of 8:2.

[0067] According to the chemical component set in the traditional Chinese medicine, a significant chemical component set of the efficacy of the medicinal material is constructed. j medicinal materials with the efficacy of the medicinal material (such as "soothing the nerves") are obtained from M selected traditional Chinese medicinal materials with known properties and efficacy. For each chemical component in the chemical component set, the number N of medicinal materials in which the chemical component appears in the corresponding chemical component set of the M selected traditional Chinese medicinal materials with known properties and efficacy is obtained, and the number k of medicinal materials with the efficacy of the medicinal material is obtained from the obtained N medicinal materials, and then the significant value of the chemical component with the efficacy of the medicinal material is calculated. If the significant value of the chemical component with the efficacy of the medicinal material is less than the set threshold of 0.05, then the chemical component is a significant chemical component of the efficacy of the medicinal material, and a significant chemical component set of the efficacy of the medicinal material is obtained.

[0068] The samples in the dataset are divided into a training set and a test set according to a ratio of 8:2. Each sample in the training set is represented by the molecular fingerprint of a single medicinal material / compound, and the dependent variable of the task is the performance / effect of the medicinal material / compound. The set of significant chemical components with the said effect (such as "soothing the nerves") is used as the positive dataset for training the model, and the set of significant chemical components without the said effect (such as "soothing the nerves") is used as the negative dataset for training the model. Therefore, for 38 effects, the positive dataset and negative dataset for each effect are constructed respectively, and divided into a training set and a test set according to a ratio of 8:2.

[0069] Construct the artificial intelligence algorithm model, including:

[0070] The number of decision trees n_estimators is set to (10, 1000), the node splitting uses 'Gini value gini' and 'entropy', the maximum number of features uses'square root sqrt' and 'logarithm log2', the range of the maximum depth of the decision tree is set to (1, 3, 5, 7, 9), and the random forest model is trained using the data in the training set.

[0071] The kernel function is set to RBF (Radial Basis Function, RBF), the kernel function coefficient is set to (-3, -15), the penalty function is set to (-3, -15), and the support vector machine model is trained using the data in the training set.

[0072] The number of decision trees n_estimators is set to (50, 400), the range of the maximum depth of the decision tree is set to (1, 11), the learning rate is set to 0.05, and the extreme gradient boosting model is trained using the data in the training set.

[0073] The range of alpha is set to (0.01, 2), and the naive Bayes model is trained using the data in the training set.

[0074] The number of nearest neighbors n_neighbors is set to (1, 21), and the K-nearest neighbor model is trained using the data in the training set.

[0075] The learning rate is set to adaptive, the range of the hidden layer is set from 1 to 4 layers, each hidden layer has 50 neurons, the activation function is set to (identity function identity, logistic function logistic, hyperbolic tangent tanh, rectified linear unit relu), the adaptive moment estimation (adam) algorithm is used as the optimization algorithm, the loss function is the cross-entropy loss function, and the maximum number of iterations is 500 times. The deep neural network model is trained using the data in the training set.

[0076] The prediction model for each efficacy is constructed by six algorithms. Based on these six algorithms, at the traditional Chinese medicine level and the compound level, 12 models are respectively constructed for each efficacy. The model with the highest f1 score value among the 12 models is used as the optimal model for subsequent applications. The grid search technique is used to optimize the parameters of machine learning and deep learning algorithms to obtain the optimal prediction results and parameters. The present invention is based on the collection of traditional Chinese medicinal materials and the collection of significant chemical components, and based on six artificial intelligence algorithms, 12 evaluation models for each efficacy among the 38 medicinal material efficacies are obtained.

[0077] The test set is used to evaluate the artificial intelligence algorithm model. Based on the evaluation index, the model with the highest f1 score value of the medicinal material efficacy model is obtained, and the optimal evaluation model of the medicinal material efficacy is obtained by using the evaluation method.

[0078] For example: Based on the collection of traditional Chinese medicinal materials and the collection of significant chemical components, among the 12 models of "soothing the nerves" in the collection of traditional Chinese medicinal materials and the collection of significant chemical components constructed by six algorithms of DNN, KNN, NB, RF, SVM, and XGB, the f1 score values are 0.627, 0.544, 0.624, 0.538, 0.663, 0.600, 0.722, 0.693, 0.712, 0.728, 0.701, and 0.698 respectively. Then, the f1 score value of the RF algorithm model at the chemical molecular level of "soothing the nerves" is 0.728, and the RF model at the chemical molecular level with the highest f1 score value is selected as the optimal model. Based on the collection of traditional Chinese medicinal materials and the collection of significant chemical components, among the 12 models of "promoting blood circulation to remove blood stasis" in the collection of traditional Chinese medicinal materials and the collection of significant chemical components constructed by six algorithms of DNN, KNN, NB, RF, SVM, and XGB, the f1 score values are 0.562, 0.531, 0.468, 0.496, 0.585, 0.506, 0.821, 0.554, 0.674, 0.641, 0.557, and 0.569 respectively. Then, the f1 score value of the DNN algorithm model at the chemical molecular level of "promoting blood circulation to remove blood stasis" is 0.821, and the DNN algorithm model at the chemical molecular level with the highest f1 score value is selected as the optimal model.

[0079] Step 4: Evaluation of Western medicine on the efficacy model

[0080] The molecular fingerprint of the medicinal material / compound is used as the input vector to obtain the efficacy prediction result of the medicinal material / compound. When the efficacy prediction result is 1, it is considered that the medicinal material / compound has this efficacy, otherwise it is set to 0, indicating that the medicinal material / compound does not have this efficacy. Through literature research, and find literature to verify the calculated prediction results based on literature and historical data.

[0081] The to-be-tested chemical molecules of the sorted Xingdouyun platform (Xingdouyun (tasly.com)) are used to obtain the 2048-bit molecular structure features of each to-be-tested chemical molecule through RDkit, which are used as the input vectors of the optimal model for the medicinal material efficacy. Then, they are input into the optimal models for 38 kinds of efficacies, and the output is the predicted value of the medicinal material efficacy. When the efficacy prediction result is 1, it is considered that the to-be-tested chemical molecule has this efficacy; otherwise, it is set to 0, indicating that the to-be-tested chemical molecule does not have this efficacy. For example, taking the molecular structure features of alprazolam as the input values of 38 optimal models, the model prediction results show that the efficacy of alprazolam is (0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0), that is, alprazolam is predicted to have the efficacy of calming the nerves.

[0082] Step 5: Verification of evaluation results

[0083] Table 1. Prediction results of western medicine compounds in the efficacy model

[0084]

[0085]

[0086] Table 2. Verification of the prediction results of western medicine compounds in the efficacy model

[0087]

[0088]

[0089]

[0090] Whether it has creativity as described in the table means that for an indication not involved in the drug instruction manual, if a new indication is predicted by the algorithm of the present application, it is described as having creativity. From the results, the algorithm of the present application can find new indications, and the retrieved supporting documents verify the accuracy of this algorithm.

[0091] Example Two:

[0092] In this example, the four natures in the performance of the drug are predicted based on the method of the present invention, and the calculation prediction results are verified by literature and historical data according to the literature research. The specific content is as follows:

[0093] Step 1: Establish a traditional Chinese medicine material library

[0094] Organize the content of traditional Chinese medicinal materials in the Pharmacopoeia of the People's Republic of China (2020 Edition), including information on the four natures, five flavors, meridian tropism, toxicity, and efficacy of the medicinal materials, and integrate them to form a database of traditional Chinese medicinal materials.

[0095] Step 2: Construct the four-nature matrix of traditional Chinese medicinal materials

[0096] According to the above-mentioned database of traditional Chinese medicinal materials, summarize and classify the four natures of traditional Chinese medicinal materials, and construct the four-nature matrix of medicinal materials.

[0097] The "four natures" attributes of cold, heat, warmth, coolness, and neutrality of traditional Chinese medicine also come from the summary and accumulation of medical practice. The cold, heat, warmth, coolness, and neutrality attributes of traditional Chinese medicine are not the physical and chemical properties of the drug itself, but are summarized from the perspective of the drug's effect on the human body and the body's reaction. Among them, warmth and coldness belong to two different natures, and there are both commonalities and differences in degree between warmth and heat, cold and coolness. The four natures correspond to the cold and heat properties of the diseases to be treated: drugs that can reduce or eliminate heat syndromes generally belong to cold or cool natures; conversely, drugs that can reduce or eliminate cold syndromes belong to warm or hot natures. Consult the list of the four natures of traditional Chinese medicine and classify the four natures of traditional Chinese medicine into 10 categories, including cold, great cold, slightly cold, heat, great heat, warm, slightly warm, cool, slightly cool, and neutral; then digitally match the four natures of traditional Chinese medicine contained in the database of traditional Chinese medicinal materials to obtain the four-nature matrix (616, 10) of medicinal materials. For example, the four-nature vector of Pinellia ternata is (0, 0, 1, 0, 0, 0, 0, 0, 0, 0), that is, Pinellia ternata belongs to slightly cold in the four natures. The four-nature vector of Fritillaria thunbergii is (1, 0, 0, 0, 0, 0, 0, 0, 0, 0), that is, Fritillaria thunbergii belongs to cold in the four natures.

[0098] Step 3: Construct the four-nature model of drugs

[0099] Obtain the molecular structure characteristics of the traditional Chinese medicine with the average value of the molecular structure characteristics of the chemical composition set in the traditional Chinese medicine. Each chemical composition is represented by a bit vector of length 2048 (1, 0, 0, 0, 0, 0, 0, 0, 0,......). The molecular structure characteristics of the traditional Chinese medicine are composed of 2048-bit characteristics of m chemical compositions. Sum the elements of each column of the m*2048 two-dimensional matrix and divide by the number of chemical compositions m to obtain the molecular structure characteristics of the traditional Chinese medicine (0.0085, 0.1888, 0.0042, 0.0042, 0.0042, 0.0171,......). The four-nature property prediction model is used as a multi-classification model, with a total of 10 types of properties (cold, great cold, slightly cold, heat, great heat, warm, slightly warm, cool, slightly cool, and neutral). The four-nature properties of each traditional Chinese medicinal material are marked respectively. For example, "Pinellia ternata" - "warm", "Fritillaria thunbergii" - "cold". Divide the dataset of 616 traditional Chinese medicinal materials after all markings into a training set and a test set according to a ratio of 8:2.

[0100] According to the set of chemical components in the traditional Chinese medicine, construct the set of significant chemical components of the four natures of the medicinal materials. Obtain j medicinal materials with the efficacy of the medicinal materials (such as "warm") from M selected traditional Chinese medicinal materials with known properties and efficacy. For each chemical component in the set of chemical components, obtain the number N of medicinal materials in which the chemical component appears in the corresponding set of chemical components of the M selected traditional Chinese medicinal materials with known properties and efficacy. Obtain the number k of medicinal materials with the efficacy of the medicinal materials from the obtained N medicinal materials, and then calculate the significant value of the chemical component with the efficacy of the medicinal materials. If the significant value of the chemical component with the efficacy of the medicinal materials is less than the set threshold of 0.05, then the chemical component is a significant chemical component with the efficacy of the medicinal materials, and the set of significant chemical components with the efficacy of the medicinal materials is obtained.

[0101] The set of significant chemical components with the four-nature properties is used as the positive data set for training the model, and the set of significant chemical components without the four-nature properties is used as the negative data set for training the model, and they are divided into a training set and a test set according to a ratio of 8:2.

[0102] Construct the artificial intelligence algorithm model, including:

[0103] Set the number of decision trees n_estimators to (10, 1000), use 'Gini value gini' and 'entropy' for node splitting, use 'root sqrt' and 'logarithm log2' for the maximum number of features, and set the range of the maximum depth of the decision tree to (1, 3, 5, 7, 9), and use the data in the training set to train the random forest model.

[0104] Set the kernel function to RBF (Radial Basis Function, RBF), set the kernel function coefficient to (-3, -15), set the penalty function to (-3, -15), and use the data in the training set to train the support vector machine model.

[0105] Set the number of decision trees n_estimators to (50, 400), set the range of the maximum depth of the decision tree to (1, 11), set the learning rate to 0.05, and use the data in the training set to train the extreme gradient boosting model.

[0106] Set the range of alpha to (0.01, 2), and use the data in the training set to train the naive Bayes model.

[0107] Set the number of nearest neighbors n_neighbors to (1, 21), and use the data in the training set to train the K-nearest neighbor model.

[0108] The learning rate is set to adaptive, the range of the hidden layer is set from 1 to 4 layers, each hidden layer has 50 neurons, the activation function is set to (identity function, logistic function, tanh function, relu function), the adaptive moment estimation (adam) algorithm is used as the optimization algorithm, the loss function is the cross-entropy loss function, the maximum number of iterations is 500 times, and the deep neural network model is trained using the data in the training set.

[0109] Based on the traditional Chinese medicine collection and the significant chemical component collection, 12 evaluation models for the four natures of the traditional Chinese medicine are obtained based on 6 artificial intelligence algorithms.

[0110] The test set is used to evaluate the artificial intelligence algorithm model. Based on the evaluation index, the model with the highest f1score value of the four-nature model of the traditional Chinese medicine is obtained, and the optimal evaluation model of the four natures of the traditional Chinese medicine is obtained using the evaluation method.

[0111] Among the 12 prediction models of the four natures, the f1 score values are 0.517, 0.453, 0.385, 0.448, 0.480, 0.432, 0.467, 0.419, 0.476, 0.485, 0.477, 0.479 respectively. The DNN model at the traditional Chinese medicine level with the highest f1 score value is selected as the optimal model.

[0112] Step 4: Evaluation of the four-nature model by western medicine

[0113] For the 2415 sorted chemical molecules to be tested, the 2048-bit molecular structure features of each chemical molecule to be tested are obtained through RDkit and used as the input vector of the optimal model of the four natures of the traditional Chinese medicine. The input is fed into the optimal model of the four natures, and the four-nature properties of each compound are predicted through the optimal model of the four natures. The output is the predicted value of the four natures of the traditional Chinese medicine. When the prediction result is "cold" among the 10 four-nature properties (cold, severely cold, slightly cold, hot, severely hot, warm, slightly warm, cool, slightly cool, neutral), it is considered that the compound has cold nature. When the prediction result is "hot", it indicates that the compound has hot nature. For example, the molecular fingerprint of Amoxicillin is used as the input of the four-nature model, and the predicted four natures of Amoxicillin is cold.

[0114] Step 5: Verification of the evaluation results

[0115] The four natures of traditional Chinese medicine refer to the four medicinal properties of cold, heat, warmth, and coolness. Among them, coldness and warmth belong to two different natures; while cold and cool, warmth and heat have basically the same nature, only with differences in degree. Warmth is less intense than heat, and coolness is less intense than cold. This is what is meant by "treating cold with heat and heat with cold" in "Plain Questions: Treatise on the Most Essential Questions". Traditional Chinese medicine believes that: Generally, drugs that can alleviate or eliminate heat syndromes belong to cold or cool natures; generally, drugs that can alleviate or eliminate cold syndromes belong to warm or hot natures. In addition, the concept of "neutral nature" is also mentioned in traditional Chinese medicine books throughout history. It refers to drugs with less obvious cold, heat, warmth, or coolness properties and relatively mild effects. Its neutral nature is relative, not an absolute concept. Generally, it is called the four natures.

[0116] Table 3. Prediction results of western medicine compounds in the four-nature model

[0117]

[0118]

[0119] Table 4. Verification of the prediction results of western medicine compounds in the four-nature model

[0120]

[0121]

[0122] Example 3:

[0123] Based on the method of the present invention, this example predicts the five flavors in the properties of drugs, and verifies the calculation prediction results with literature and historical data according to literature research. The specific content is as follows:

[0124] Step 1: Establish a traditional Chinese medicine material database

[0125] Sort out the content of traditional Chinese medicine materials in "Pharmacopoeia of the People's Republic of China (2020 Edition)", including information on the four natures, five flavors, meridian tropism, toxicity, and efficacy of the medicinal materials, and integrate them to form a traditional Chinese medicine material database.

[0126] Step 2: Construct a five-flavor matrix of traditional Chinese medicine materials

[0127] According to the above traditional Chinese medicine material database, summarize and classify the five flavors of traditional Chinese medicine materials to construct a five-flavor matrix of medicinal materials.

[0128] The five flavors in traditional Chinese medicine mainly include pungent, sweet, sour, bitter, salty, and bland. Pungent-flavored herbs have the effects of dispersing, promoting qi movement, and promoting blood circulation; sweet-flavored herbs have the effects of tonifying and regulating the middle energizer, relieving spasm and pain; sour-flavored herbs have the effects of astringing and astringing; bitter-flavored herbs have the effects of clearing heat, drying dampness, purging, and strengthening yin; salty-flavored herbs have the effects of softening hardness and resolving masses, and purging; bland-flavored herbs have the effects of promoting diuresis and removing dampness; the astringent taste is mostly similar to the sour taste and has the effect of astringing and astringing. By referring to the list of the five flavors of traditional Chinese medicine, the five flavors of traditional Chinese medicine are divided into seven categories, including sour, bitter, sweet, pungent, salty, bland, and astringent; then, the five flavors of traditional Chinese medicine contained in the traditional Chinese medicine material library are digitally matched to obtain the five-flavor matrix (616, 7) of the medicinal materials. For example, the five-flavor vector of Fritillaria cirrhosa is (0, 1, 1, 0, 0, 0, 0), that is, Fritillaria cirrhosa belongs to bitter and sweet among the five flavors. The five-flavor vector of Cyathula officinalis is (1, 1, 1, 0, 0, 0, 0), that is, Cyathula officinalis belongs to sour, bitter, and sweet among the five flavors.

[0129] Step 3: Construct a drug five-flavor model

[0130] Using the average value of the molecular structure characteristics of the chemical composition set in the traditional Chinese medicine to obtain the molecular structure characteristics of the traditional Chinese medicine, each chemical composition is represented by a bit string of length 2048 (1, 0, 0, 0, 0, 0, 0, 0, 0,......), and the molecular structure characteristics of the traditional Chinese medicine are composed of 2048-bit characteristics of m chemical compositions. Dividing the sum of the elements in each column of the m*2048 two-dimensional matrix by the number of chemical compositions m to obtain the molecular structure characteristics of the traditional Chinese medicine (0.0085, 0.1888, 0.0042, 0.0042, 0.0042, 0.0171,......).

[0131] The set of traditional Chinese medicinal materials with a certain five-flavor (such as "pungent") is used as the positive data set for training the model, and the set of traditional Chinese medicinal materials without a certain five-flavor (such as "pungent") is used as the negative data set for training the model. Therefore, for the 7 five-flavors, the positive data set and negative data set for each five-flavor are constructed respectively, and they are divided into a training set and a test set according to the ratio of 8:2.

[0132] According to the chemical composition set in the traditional Chinese medicine, construct the significant chemical composition set of the five flavors of the medicinal material. Select j medicinal materials with the efficacy of the medicinal material (such as "pungent") from M selected traditional Chinese medicinal materials with known properties and effects. For each chemical composition in the chemical composition set, obtain the number N of medicinal materials in which the chemical composition appears in the corresponding chemical composition set of the M selected traditional Chinese medicinal materials with known properties and effects, and obtain the number k of medicinal materials with the efficacy of the medicinal material from the obtained N medicinal materials, and then calculate the significant value of the chemical composition with the efficacy of the medicinal material. If the significant value of the chemical component for the efficacy of the medicinal material is less than the set threshold of 0.05, then the chemical component is a significant chemical component for the efficacy of the medicinal material, and a set of significant chemical components for the efficacy of the medicinal material is obtained.

[0133] The set of significant chemical components with the five flavors (e.g., the flavor of "pungent") is used as the positive data set for training the model, and the set of significant chemical components without the five flavors (e.g., the flavor of "pungent") is used as the negative data set for training the model. Therefore, for the 7 kinds of five flavors, a positive data set and a negative data set for each flavor are constructed respectively, and they are divided into a training set and a test set according to the ratio of 8:2.

[0134] Constructing the artificial intelligence algorithm model includes:

[0135] The number of decision trees n_estimators is set to (10, 1000), the node splitting uses 'Gini value gini' and 'entropy', the maximum number of features uses'square root sqrt' and 'logarithm log2', the maximum depth range of the decision tree is set to (1, 3, 5, 7, 9), and the data in the training set is used to train the random forest model.

[0136] The kernel function is set to RBF (Radial Basis Function, RBF), the kernel function coefficient is set to (-3, -15), the penalty function is set to (-3, -15), and the data in the training set is used to train the support vector machine model.

[0137] The number of decision trees n_estimators is set to (50, 400), the maximum depth range of the decision tree is set to (1, 11), the learning rate is set to 0.05, and the data in the training set is used to train the extreme gradient boosting model.

[0138] The range of alpha is set to (0.01, 2), and the data in the training set is used to train the naive Bayes model.

[0139] The number of nearest neighbors n_neighbors is set to (1, 21), and the data in the training set is used to train the K-nearest neighbor model.

[0140] The learning rate is set to adaptive, the range of the hidden layer is set from 1 to 4 layers, each hidden layer has 50 neurons, the activation function is set to (identity function, logistic function, tanh function, relu function), the adaptive moment estimation (adam) algorithm is used as the optimization algorithm, the loss function is the cross-entropy loss function, the maximum number of iterations is 500 times, and the data in the training set is used to train the deep neural network model.

[0141] Based on the traditional Chinese medicine material set and the significant chemical component set, 12 evaluation models for each of the five flavors of the 7 traditional Chinese medicines are obtained based on 6 artificial intelligence algorithms.

[0142] The test set is used to evaluate the artificial intelligence algorithm model. Based on the evaluation index, the model with the highest f1score value of the five-flavor model of the traditional Chinese medicine is obtained, and the optimal evaluation model of the five flavors of the traditional Chinese medicine is obtained by using the evaluation method.

[0143] For example: Based on the traditional Chinese medicine material set and the significant chemical component set, among the 12 models of the "sour" flavor at the traditional Chinese medicine level and the chemical molecule level built by 6 algorithms of DNN, KNN, NB, RF, SVM, and XGB, the f1 score values are 0.718, 0.719, 0.650, 0.657, 0.647, 0.572, 0.683, 0.706, 0.683, 0.717, 0.728, 0.667 respectively. Then the f1 score value of the SVM algorithm model at the chemical molecule level of the "sour" flavor is 0.728, and the SVM model at the chemical molecule level with the highest f1 score value is selected as the optimal model. Based on the traditional Chinese medicine material set and the significant chemical component set, among the 12 models of the "sweet" flavor at the traditional Chinese medicine level and the chemical molecule level built by 6 algorithms of DNN, KNN, NB, RF, SVM, and XGB, the f1 score values are 0.707, 0.684, 0.675, 0.685, 0.667, 0.680, 0.769, 0.743, 0.739, 0.760, 0.763, 0.744 respectively. Then the f1 score value of the DNN algorithm model at the chemical molecule level of the "sweet" flavor is 0.728, and the DNN model at the chemical molecule level with the highest f1 score value is selected as the optimal model.

[0144] Step 4: Evaluation of Western medicines on the five-flavor model

[0145] The 2,415 chemical molecules to be tested are sorted out, and the 2,048-bit molecular structure features of each chemical molecule to be tested are obtained through RDkit. As the input vector of the optimal model of the five flavors of traditional Chinese medicinal materials, it is input into the optimal models of the seven flavors. The output is the predicted value of the five flavors of the traditional Chinese medicinal materials. When the prediction result is 1, it is considered that the compound has this flavor. Otherwise, it is set to 0, indicating that the compound does not have this flavor. For example: Taking the molecular fingerprint of Domperidone as the input value of the seven optimal models, the model prediction result shows that the five flavors of Domperidone are (0, 0, 1, 1, 1, 1, 1), that is, Domperidone is predicted to have astringent, light, sweet, bitter, and sour flavors.

[0146] Step 5: Verification of evaluation results

[0147] Table 5. Prediction results in the five-flavor model of Western medicine compounds

[0148]

[0149] Table 6. Verification of the prediction results of Western medicine compounds in the five-flavor model

[0150]

[0151]

[0152] Example 4:

[0153] Based on the method of the present invention, this example predicts the meridian tropism in the properties of drugs, and verifies the calculation prediction results according to literature research and historical data. The specific content is as follows:

[0154] Step 1: Establish a database of traditional Chinese medicinal materials

[0155] Sort out the content of traditional Chinese medicinal materials in the Pharmacopoeia of the People's Republic of China (2020 Edition), including information on the four natures, five flavors, meridian tropism, toxicity and efficacy of the medicinal materials, and integrate them to form a database of traditional Chinese medicinal materials.

[0156] Step 2: Construct a meridian tropism matrix of traditional Chinese medicinal materials

[0157] According to the above database of traditional Chinese medicinal materials, summarize and classify the meridian tropism of traditional Chinese medicinal materials, and construct a meridian tropism matrix of medicinal materials.

[0158] Meridian tropism in traditional Chinese medicine mainly includes the lung, large intestine, stomach, spleen, heart, small intestine, bladder, kidney, pericardium, triple energizer, gallbladder, and liver. Referring to the list of meridian tropism in traditional Chinese medicine, the meridian tropism of traditional Chinese medicine is divided into twelve categories, including the lung, large intestine, stomach, spleen, heart, small intestine, bladder, kidney, pericardium, triple energizer, gallbladder, and liver; then, the meridian tropism of the traditional Chinese medicine contained in the traditional Chinese medicine material library is digitally matched to obtain the meridian tropism matrix (616, 12) of the medicinal material. For example, the meridian tropism vector of Inula britannica is (1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0), that is, Inula britannica mainly belongs to the lung meridian, large intestine meridian, stomach meridian, and spleen meridian. The meridian tropism vector of Bambusa textilis is (1, 0, 1, 0, 1, 0, 0, 0, 0, 0, 1, 0), that is, Bambusa textilis mainly belongs to the lung meridian, stomach meridian, heart meridian, and gallbladder meridian.

[0159] Step 3: Construct a drug meridian tropism model

[0160] Taking the average value of the molecular structure characteristics of the chemical composition set in the traditional Chinese medicine to obtain the molecular structure characteristics of the traditional Chinese medicine, each chemical composition is represented by a bit string of length 2048 (1, 0, 0, 0, 0, 0, 0, 0, 0,...), and the molecular structure characteristics of the traditional Chinese medicine are composed of 2048-bit characteristics of m chemical compositions. Dividing the sum of the elements in each column of the m * 2048 two-dimensional matrix by the number of chemical compositions m to obtain the molecular structure characteristics of the traditional Chinese medicine (0.0085, 0.1888, 0.0042, 0.0042, 0.0042, 0.0171,...).

[0161] The set of traditional Chinese medicinal materials with a certain meridian tropism (such as the "lung" meridian) is used as the positive data set for training the model, and the set of traditional Chinese medicinal materials without a certain meridian tropism (such as the "lung" meridian) is used as the negative data set for training the model. Therefore, for 12 meridian tropisms, the positive data set and negative data set for each meridian tropism are constructed respectively, and they are divided into a training set and a test set according to a ratio of 8:2.

[0162] According to the chemical composition set in the traditional Chinese medicine, construct the significant chemical composition set of the medicinal material meridian tropism. Obtain j medicinal materials with the efficacy of the medicinal material (such as the "lung" meridian) from M selected traditional Chinese medicinal materials with known properties and effects. For each chemical composition in the chemical composition set, obtain the number of medicinal materials N in which the chemical composition appears in the corresponding chemical composition set of the M selected traditional Chinese medicinal materials with known properties and effects, and obtain the number of medicinal materials k with the efficacy of the medicinal material from the obtained N medicinal materials, and then calculate the significant value of the chemical composition with the efficacy of the medicinal material. If the significant value of the chemical composition with the efficacy of the medicinal material is less than the set threshold of 0.05, then the chemical composition is a significant chemical composition with the efficacy of the medicinal material, and the significant chemical composition set with the efficacy of the medicinal material is obtained.

[0163] The set of significant chemical components with the meridian tropism (e.g., the "lung" meridian) is used as the positive data set for training the model, and the set of significant chemical components without the meridian tropism (e.g., the "lung" meridian) is used as the negative data set for training the model. Therefore, for 12 meridians, the positive data set and negative data set of each meridian are constructed respectively, and divided into a training set and a test set according to the ratio of 8:2.

[0164] Construct the artificial intelligence algorithm model, including:

[0165] The number of decision trees n_estimators is set to (10, 1000), the node splitting uses 'Gini value gini' and 'entropy entropy', the maximum number of features uses'square root sqrt' and 'log2', the maximum depth range of the decision tree is set to (1, 3, 5, 7, 9), and the data in the training set is used to train the random forest model.

[0166] The kernel function is set to RBF (Radial Basis Function, RBF), the kernel function coefficient is set to (-3, -15), the penalty function is set to (-3, -15), and the data in the training set is used to train the support vector machine model.

[0167] The number of decision trees n_estimators is set to (50, 400), the maximum depth range of the decision tree is set to (1, 11), the learning rate is set to 0.05, and the data in the training set is used to train the extreme gradient boosting model.

[0168] The range of alpha is set to (0.01, 2), and the data in the training set is used to train the naive Bayes model.

[0169] The number of nearest neighbors n_neighbors is set to (1, 21), and the data in the training set is used to train the K-nearest neighbor model.

[0170] The learning rate is set to adaptive, the range of the hidden layer is set from 1 to 4 layers, each hidden layer has 50 neurons, the activation function is set to (identity function identity, logistic function logistic, hyperbolic tangent tanh, rectified linear unit relu), the adaptive moment estimation (adam) algorithm is used as the optimization algorithm, the loss function is the cross-entropy loss function, and the maximum number of iterations is 500 times. The data in the training set is used to train the deep neural network model.

[0171] Based on the collection of traditional Chinese medicinal materials and the collection of significant chemical components, 12 evaluation models for each of the 12 meridians are obtained based on 6 artificial intelligence algorithms.

[0172] The test set is used to evaluate the artificial intelligence algorithm model, based on Evaluation index, obtain the model with the highest f1score value of the meridian classification model, and use the evaluation method to obtain the optimal evaluation model of the meridian classification.

[0173] For example, among the 12 models of the "Lung" meridian for the Chinese medicinal materials set and the significant chemical component set built by the six algorithms, namely DNN, KNN, NB, RF, SVM, and XGB, the f1 score values are 0.501, 0.517, 0.417, 0.505, 0.513, 0.516, 0.853, 0.854, 0.847, 0.824, 0.837, and 0.860, respectively. The f1 score value of the SVM algorithm model at the chemical molecular level of the "Lung" meridian is 0.860, and the chemical molecular level SVM model with the highest f1 score value is selected as the optimal model. Among the 12 models of the "Large Intestine" meridian at the traditional Chinese medicine level and chemical molecular level built by the six algorithms of DNN, KNN, NB, RF, SVM, and XGB, the f1score values are 0.734, 0.728, 0.735, 0.718, 0.729, 0.732, 0.842, 0.839, 0.804, 0.833, 0.825, and 0.846, respectively. The f1 score value of the SVM algorithm model at the chemical molecular level of the "Large Intestine" meridian is 0.846. The chemical molecular level SVM model with the highest f1score value is selected as the optimal model.

[0174] Step 4: Evaluation of Western medicine on the meridian model

[0175] RDkit was used to obtain 2048 bits of molecular structural features for each of the 2415 tested chemical molecules. These features were used as input vectors for the optimal meridian models for the medicinal materials. These vectors were then fed into the optimal models for 12 meridians, and the predicted values for the meridians were output. A prediction of 1 indicated that the tested chemical molecule had that meridian; otherwise, it was set to 0, indicating that the tested chemical molecule did not have that meridian. For example, the molecular structural features of phenolphthalein were used as input for the 12 optimal models. The model prediction results showed that the meridians for phenolphthalein were (0, 1, 1, 0, 0, 0, 0, 1, 0, 0, 0), indicating that phenolphthalein was predicted to be associated with the large intestine, stomach, and kidney meridians.

[0176] Step 5: Verification of evaluation results

[0177] Table 7. Prediction results of Western medicine compounds in the meridian model

[0178]

[0179]

[0180] Table 8.6

[0181]

[0182]

[0183]

[0184] Example 5:

[0185] In this example, the toxicity in the performance of the drug was predicted based on the method of the present invention, and the calculation prediction results were verified by literature and historical data according to the literature research. The specific content is as follows:

[0186] Step 1: Establish a traditional Chinese medicine material library

[0187] Sort out the content of traditional Chinese medicine materials in the Pharmacopoeia of the People's Republic of China (2020 Edition), including the information of the four natures, five flavors, meridian tropism, toxicity and efficacy of the medicinal materials, and integrate them to form a traditional Chinese medicine material database.

[0188] Step 2: Construct a meridian tropism matrix of traditional Chinese medicine materials

[0189] According to the above traditional Chinese medicine material library, classify the toxicity of traditional Chinese medicine materials and construct a medicinal material toxicity matrix.

[0190] The toxicity in traditional Chinese medicine mainly includes toxic, highly toxic, slightly toxic, and non-toxic. Consult the list of traditional Chinese medicine toxicity and classify the traditional Chinese medicine toxicity into four categories, including toxic, highly toxic, slightly toxic, and non-toxic; then digitally match the traditional Chinese medicine toxicity of the traditional Chinese medicine materials contained in the traditional Chinese medicine material library to obtain the toxicity matrix (616, 4) of the medicinal materials. For example, the toxicity vector of Polygonum multiflorum is (1, 0, 0, 0), that is, Polygonum multiflorum is toxic. The toxicity vector of Cnidium monnieri is (0, 0, 1, 0), that is, Cnidium monnieri is slightly toxic.

[0191] Step 3: Construct a drug toxicity model

[0192] Taking the average value of the molecular structure characteristics of the chemical composition set in the traditional Chinese medicine, obtain the molecular structure characteristics of the traditional Chinese medicine. Each chemical composition is represented by a bit string of length 2048 (1, 0, 0, 0, 0, 0, 0, 0, 0,...), and the molecular structure characteristics of the traditional Chinese medicine are composed of 2048-bit characteristics of m chemical components. Sum the elements of each column of the m * 2048 two-dimensional matrix and divide by the number of chemical components m to obtain the molecular structure characteristics of the traditional Chinese medicine (0.0085, 0.1888, 0.0042, 0.0042, 0.0042, 0.0171,......).

[0193] As a multi-classification model, the toxicity prediction model is marked with 4 types of toxicity (toxic, highly toxic, slightly toxic, non-toxic). The toxicity of each Chinese medicinal material is marked separately. For example, "Aconitum kusnezoffii" - "highly toxic", "Salvia miltiorrhiza" - "non-toxic". The dataset of 616 marked Chinese medicinal materials is divided into a training set and a test set according to the ratio of 8:2.

[0194] According to the chemical component set in the traditional Chinese medicine, construct a significant chemical component set for the toxicity of the medicinal material. Select j medicinal materials with the efficacy of the medicinal material (such as "highly toxic") from M selected Chinese medicinal materials with known properties and effects. For each chemical component in the chemical component set, obtain the number of medicinal materials N in which this chemical component appears in the corresponding chemical component set of the M selected Chinese medicinal materials with known properties and effects. Then obtain the number of medicinal materials k with the efficacy of this medicinal material from the obtained N medicinal materials, and then calculate the significant value of this chemical component with the efficacy of this medicinal material. If the significant value of this chemical component with the efficacy of this medicinal material is less than the set threshold of 0.05, then this chemical component is a significant chemical component for the efficacy of this medicinal material, and a significant chemical component set for the efficacy of this medicinal material is obtained.

[0195] The set of significant chemical components with the toxicity is used as the positive dataset for training the model, and the set of significant chemical components without the toxicity is used as the negative dataset for training the model, and they are divided into a training set and a test set according to the ratio of 8:2.

[0196] Construct the artificial intelligence algorithm model, including:

[0197] Set the number of decision trees n_estimators to (10, 1000), use 'Gini value gini' and 'entropy' for node division, use 'root sqrt' and 'logarithm log2' for the maximum number of features, and set the maximum depth range of the decision tree to (1, 3, 5, 7, 9). Use the data in the training set to train the random forest model.

[0198] Set the kernel function to RBF (Radial Basis Function, RBF), set the kernel function coefficient to (-3, -15), set the penalty function to (-3, -15), and use the data in the training set to train the support vector machine model.

[0199] The number of decision trees n_estimators is set to (50, 400), the maximum depth range of the decision trees is set to (1, 11), the learning rate is set to 0.05, and the data in the training set is used to train the extreme gradient boosting model.

[0200] The range of alpha is set to (0.01, 2), and the data in the training set is used to train the Naive Bayes model.

[0201] The number of nearest neighbors n_neighbors is set to (1, 21), and the data in the training set is used to train the K-nearest neighbor model.

[0202] The learning rate is set to adaptive, the range of the hidden layer is set from 1 to 4 layers, each hidden layer has 50 neurons, the activation function is set to (identity function, logistic function, tanh, relu), the adaptive moment estimation (adam) algorithm is used as the optimization algorithm, the loss function is the cross-entropy loss function, the maximum number of iterations is 500 times, and the data in the training set is used to train the deep neural network model.

[0203] Based on the traditional Chinese medicine set and the significant chemical composition set, 12 evaluation models for the toxicity are obtained based on 6 artificial intelligence algorithms.

[0204] The test set is used for the evaluation of the artificial intelligence algorithm model. Based on the evaluation index, the model with the highest f1score value for the toxicity of the medicinal material is obtained, and the optimal evaluation model for the toxicity of the medicinal material is obtained by using the evaluation method.

[0205] Among the 12 prediction models for the toxicity, the f1score values are 0.690, 0.729, 0.696, 0.717, 0.712, 0.618, 0.568, 0.616, 0.651, 0.662, 0.643, 0.587 respectively. The KNN model at the traditional Chinese medicine level with the highest f1score value is selected as the optimal model.

[0206] Step 4: Evaluation of Western medicine on the toxicity model

[0207] The 2,415 chemical molecules to be tested are sorted out, and the 2,048-bit molecular structure features of each chemical molecule to be tested are obtained through RDkit. As the input vector of the optimal model for the toxicity of medicinal materials, it is input into the optimal model for toxicity. Through the optimal model for toxicity prediction, the toxicity properties of each compound are obtained, and the output is the predicted value of the toxicity of medicinal materials. When the prediction result is "non-toxic" among the 10 types of toxicity properties (toxic, highly toxic, slightly toxic, non-toxic), it is considered that the compound is not toxic. When the prediction result is "highly toxic", it means that the compound is highly toxic. For example, the molecular fingerprint of paclitaxel is used as the input of the toxicity model, and it is predicted that paclitaxel is toxic.

[0208] Step 5: Verification of evaluation results

[0209] Table 9. Prediction results of Western medicine compounds in the toxicity model

[0210]

[0211] Table 10. Verification of the prediction results of Western medicine compounds in the toxicity model

[0212]

[0213]

[0214] Although specific embodiments of the present invention are disclosed for illustrative purposes, the purpose is to help understand the content of the present invention and implement it accordingly. Those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection required by the present invention is subject to the scope defined by the claims.

Claims

1. A method for predicting the performance and efficacy of drugs based on artificial intelligence algorithms, the steps of which include: 1) Construct a performance matrix of traditional Chinese medicinal materials according to the performance of traditional Chinese medicinal materials, and construct an efficacy matrix of traditional Chinese medicinal materials according to the efficacy of traditional Chinese medicinal materials; 2) According to the drug information of each selected traditional Chinese medicinal material, determine the performance and efficacy of the corresponding traditional Chinese medicinal material; then, according to the performance matrix of traditional Chinese medicinal materials and the performance of each selected traditional Chinese medicinal material, quantify the selected traditional Chinese medicinal materials to generate a performance vector of the corresponding traditional Chinese medicinal material; according to the efficacy matrix of traditional Chinese medicinal materials and the efficacy of each selected traditional Chinese medicinal material, quantify the selected traditional Chinese medicinal materials to generate an efficacy vector of the corresponding traditional Chinese medicinal material; 3) For each selected traditional Chinese medicinal material with known performance and efficacy, obtain the chemical composition set of the corresponding traditional Chinese medicinal material; according to the chemical composition set of the traditional Chinese medicinal material, establish a significant chemical composition set and a non-significant chemical composition set for the performance of each medicinal material, and a significant chemical composition set and a non-significant chemical composition set for the efficacy of each medicinal material; 4) Calculate the average value of the molecular structure characteristics of the chemical composition set of each traditional Chinese medicinal material as the molecular structure characteristic of the corresponding traditional Chinese medicinal material. Take the molecular structure characteristic of each traditional Chinese medicinal material with the i-th medicinal material efficacy and its efficacy vector as a sample to obtain a first efficacy sample set, and divide it into a first efficacy training set and a first efficacy test set; use the first efficacy training set to train and optimize P artificial intelligence algorithm models, and then use the first test set for testing. Select the artificial intelligence algorithm model with the best test effect as the first efficacy evaluation model for the i-th medicinal material efficacy; i = 1 to Q, where Q is the total number of medicinal material efficacies; take the molecular structure characteristic of each traditional Chinese medicinal material with the w-th medicinal material performance and its performance vector as a sample to obtain a first performance sample set, and divide it into a first performance training set and a first performance test set; use the first performance training set to train and optimize P artificial intelligence algorithm models, and then use the first performance test set for testing. Select the artificial intelligence algorithm model with the best test effect as the first performance evaluation model for the w-th medicinal material performance; w = 1 to W, where W is the total number of medicinal material performances; 5) Calculate the molecular structure characteristics of each chemical component in the significant chemical component set and the non-significant chemical component set of the efficacy of each medicinal material, take the molecular structure characteristics of each chemical component in the significant chemical component set and the non-significant chemical component set of the efficacy of the i-th medicinal material and the efficacy vector of the i-th medicinal material as a sample, obtain a second efficacy sample set, and divide it into a second efficacy training set and a second efficacy test set; use the second efficacy training set to train and optimize the P artificial intelligence algorithm models, then use the second efficacy test set to test, and select the artificial intelligence algorithm model with the best test effect as the second efficacy evaluation model for the efficacy of the i-th medicinal material; i=1~Q; take the molecular structure characteristics of each chemical component in the significant chemical component set and the non-significant chemical component set of the performance of the w-th medicinal material and the performance vector of the w-th medicinal material as a sample, obtain a second performance sample set, and divide it into a second performance training set and a second performance test set; use the second performance training set to train and optimize the P artificial intelligence algorithm models, then use the second performance test set to test, and select the artificial intelligence algorithm model with the best test effect as the second performance evaluation model for the performance of the w-th medicinal material; w=1~W; 6) For each chemical molecule to be tested in the tested Chinese medicinal material, a molecular structure feature of the chemical molecule to be tested is generated and inputted into a first efficacy evaluation model and a second efficacy evaluation model corresponding to the i-th medicinal material efficacy, and a first performance evaluation model and a second performance evaluation model corresponding to the j-th performance, respectively, to determine whether the chemical molecule to be tested has the i-th medicinal material efficacy and the j-th performance; then, the medicinal properties and efficacy of the Chinese medicinal material to be tested are determined based on the detection results of each chemical molecule to be tested in the Chinese medicinal material to be tested.

2. The method according to claim 1, wherein The medicinal information of the Chinese medicinal materials includes the information of the four properties, five flavors, meridians, toxicity and efficacy of the Chinese medicinal materials; the performance of the Chinese medicinal materials includes medicinal characteristics, medicinal properties, medicinal meridians and drug toxicity; the medicinal characteristics include cold, very cold, slightly cold, hot, very hot, warm, slightly warm, cool, slightly cool, flat; the medicinal properties include pungent, bitter, sweet, sour, salty, light and astringent; the medicinal meridians include lung, large intestine, stomach, spleen, heart, small intestine, bladder, kidney, pericardium, triple burner, gallbladder and liver; Toxicity includes toxic, highly toxic, slightly toxic, and non-toxic; the effects of the Chinese medicinal materials include digestion, wind-dispelling, heat-clearing, detoxification, pain relief, cold-dispelling, phlegm-resolving, tranquilizing the mind, laxative, external use, blood regulation, dampness-removing, astringent, antiemetic, anthelmintic, anesthetic, insecticide, softening, diuretic, deficiency-tonifying, tonic, diuretic and dampness-removing, calming the liver and relieving wind, relieving exterior symptoms, removing toxins, resolving dead tissue and promoting tissue regeneration, calming the liver, tonifying, hemostasis, attacking toxins, killing parasites and relieving itching, removing rheumatism, promoting blood circulation and removing blood stasis, regulating qi, resolving phlegm, relieving cough and relieving asthma, opening the orifices, dispelling parasites, drying dampness, dispelling heat, and vomiting.

3. The method according to claim 1 or 2, characterized in that, The method for obtaining the molecular structure characteristics of traditional Chinese medicine is as follows: each chemical component in the chemical component set of the traditional Chinese medicine is represented by a bit of length L, the chemical component set includes m chemical components, and the bit representation of the m chemical components is used to form an m*L two-dimensional matrix. The sum of the elements in each column of the m*L two-dimensional matrix is divided by the number of chemical components m to obtain the molecular structure characteristics of the traditional Chinese medicine.

4. The method according to claim 3, wherein The length L = 2048.

5. The method according to claim 1 or 2, characterized in that, The method for establishing the set of significant chemical components and the set of non-significant chemical components for each medicinal material's efficacy / property is as follows: First, for each medicinal material's efficacy / property, j medicinal materials with the efficacy of this medicinal material are obtained from M selected traditional Chinese medicinal materials with known efficacy / property. For each chemical component in the chemical component set, the number of medicinal materials N in which this chemical component appears in the corresponding chemical component set of the M selected traditional Chinese medicinal materials with known properties / efficacy is obtained, and the number of medicinal materials k with the efficacy of this medicinal material is obtained from the obtained N medicinal materials; then the significant value of this chemical component having the efficacy of this medicinal material is calculated. If the significant value p(k, M, j, N) of this chemical component having the efficacy / property of this medicinal material is less than the set threshold, then this chemical component is a significant chemical component of this medicinal material's efficacy / property and is added to the set of significant chemical components of this medicinal material's efficacy / property, otherwise this chemical component is added to the set of non-significant chemical components of this medicinal material's efficacy / property.

6. The method according to claim 1, wherein The P artificial intelligence algorithm models include a deep neural network model, a K-nearest neighbor model, a naive Bayes model, a random forest model, a support vector machine model, and an extreme gradient boosting model.

7. The method according to claim 6, wherein In the random forest model, the number of decision trees n_estimators is set to (10, 1000), the node splitting uses 'Gini value gini' and 'entropy', the maximum number of features uses'sqrt' and 'log2', and the maximum depth range of the decision tree is set to (1, 3, 5, 7, 9); in the support vector machine model, the kernel function is set to RBF, the kernel function coefficient is set to (-3, -15), and the penalty function is set to (-3, -15); in the extreme gradient boosting model, the number of decision trees n_estimators is set to (50, 400), the maximum depth range of the decision tree is set to (1, 11), and the learning rate is set to 0.05; in the naive Bayes model, the range of alpha is set to (0.01, 2); in the K-nearest neighbor model, the number of nearest neighbors n_neighbors is set to (1, 21); in the deep neural network model, the learning rate is set to adaptive, the range of the hidden layer is set to 1 to 4 layers, each hidden layer has 50 neurons, the activation function is set to the identity function identity, the logistic function logistic, the hyperbolic tangent tanh, or the rectified linear unit relu, the adaptive moment estimation algorithm adam is used as the optimization algorithm, the loss function is the cross-entropy loss function, and the maximum number of iterations is 500 times.

8. A server, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in any one of the methods recited in claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of any one of the methods recited in claims 1 to 7 are implemented.