An artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks integrating phenotypic and molecular information

Through deep neural network fusion of phenotype and molecular information, a Chinese medicine prescription evaluation model was constructed, which solved the problem of lack of molecular level information in Chinese medicine prescription evaluation, and achieved high-precision recommendation and rational use of Chinese medicine prescriptions.

CN115376658BActive Publication Date: 2025-08-19TSINGHUA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110736888.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-02
Filing Date
2021-06-30
Publication Date
2025-08-19
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

The lack of molecular-level information in the evaluation of traditional Chinese medicine prescriptions in the prior art has led to a high unreasonable use rate of traditional Chinese medicine prescriptions, making it difficult to achieve accurate recommendations.

Method used

A deep neural network-based method is adopted, combining phenotypic information and molecular information, and a heterogeneous network of Chinese medicine molecules is constructed, and features are extracted using network embedding representations, deep neural network models are constructed to perform vectorized representations and recommendations of Chinese medicine prescriptions.

Benefits of technology

It improves the accuracy and rationality of Chinese medicine prescription evaluation, reduces the unreasonable use rate, and improves the accuracy of Chinese medicine prescription recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115376658B_ABST
    Figure CN115376658B_ABST
Patent Text Reader

Abstract

The present invention provides an artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep learning that integrates phenotypic and molecular information. This method first constructs an intelligent extraction of diagnostic description features based on convolutional neural networks, and an intelligent extraction of drug features based on network embedding, while integrating molecular information into the drug features. Furthermore, based on the extracted diagnostic descriptions and traditional Chinese medicine prescription features, an artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on convolutional neural networks is designed. At the same time, this method also proposes a stratified sampling strategy based on the similarity of traditional Chinese medicine prescriptions for the first time. Experimental results show that our method is superior to the baseline method in the evaluation performance of traditional Chinese medicine prescriptions, and is superior to the model without adding molecular information, and can better learn expert experience. Our method promotes the traditional Chinese medicine based on experience and macro to data-based, macro-micro combined with modern science, helps to reduce the irrational use of traditional Chinese medicine prescriptions, and promotes the precision and intelligence of traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on a deep neural network that integrates phenotypic and molecular information. Background Art

[0002] Traditional Chinese medicine is widely used in clinical practice, but its irrational use is also quite serious. An analysis of non-standard prescriptions of Chinese herbal medicine slices in the outpatient department of Beijing First Hospital of Integrated Traditional Chinese and Western Medicine in 2013 [1] showed that among 2,400 prescriptions, 177 were non-standard prescriptions, accounting for 7.38% of the prescriptions sampled. Another study on irrational prescriptions made by pharmacists before dispensing at Beijing Jishuitan Hospital from January 2011 to May 2013 [2] showed that there were 709 irrational factors in 663 irrational prescriptions, of which the following three situations were the most common: excessive dosage of toxic herbal slices (73.76%), prescriptions with incompatibility taboos not signed (12.13%), and prescription herbal slices input errors (7.76%). These statistical results show that the irrational use of Chinese herbal medicine prescriptions, in addition to common manual operation errors, mainly includes illegal incompatibility taboos, failure to conduct syndrome differentiation and treatment, and failure to consider adverse reactions, which are precisely the more serious irrational use situations. Therefore, accurately recommending Chinese medicine prescriptions and reducing the irrational use rate of Chinese medicine prescriptions is an urgent problem to be solved.

[0003] With the advent of the era of artificial intelligence and big data, increasing research is focusing on using AI to mine the experience of renowned physicians, thereby enabling AI-powered evaluation of Traditional Chinese Medicine (TCM) prescriptions. Indeed, AI technology has already found certain applications in Traditional Chinese Medicine (TCM), such as the standardized collection, processing, and analysis of information from the "Four Diagnoses" of TCM, TCM constitution analysis, and the mining of the experience of renowned TCM physicians. These applications have, to a certain extent, promoted the objectification and standardization of TCM. Therefore, the application of AI technology to the AI-powered evaluation of TCM prescriptions is essential for the precise clinical use of TCM and represents a major trend fostered by the interdisciplinary development of the AI and big data era. The advantages of applying AI technology to the AI-powered evaluation of TCM prescriptions lie in the following: first, the use of AI technology can mine digitized historical TCM case records and clinical diagnosis and treatment records of renowned physicians, thereby better summarizing and inheriting the medication experience of renowned physicians; second, it can promote the rational use of TCM prescriptions in clinical practice, reduce the rate of irrational use of TCM prescriptions, and improve diagnosis and treatment efficiency.

[0004] Currently, artificial intelligence is widely used in the medical field, including medical image processing, disease diagnosis, and drug recommendations. Research on AI technology in the evaluation of traditional Chinese medicine prescriptions can be divided into three categories: the first is data mining based on the extensive prior knowledge accumulated in traditional Chinese medicine; the second is joint recommendation integrating prior knowledge and clinical information; and the third is the identification and matching of herbal active ingredients based on fingerprints of herbal ingredients.

[0005] In the field of data mining of prior knowledge in traditional Chinese medicine, Liang Yao et al. [3] proposed a system for mining prescription relationships from prescription literature, including the relationship between prescription components based on the Trie tree and the relationship between prescription efficacy based on the topic model. Wei Li et al. [4] proposed a decoder with a coverage mechanism and a soft loss function. The study captured 85,166 prescriptions from the traditional Chinese medicine prescription database and obtained 82,044 symptoms. The verification results showed an accuracy of 38.22%, a recall rate of 30.18%, and an F1 value of 33.73%. According to the judgment of traditional Chinese medicine experts, the accuracy of the generated prescriptions was 73%. Jinpeng Chen et al. [5] proposed a symptom-syndrome-traditional Chinese medicine relationship inference method based on a three-part graph. The method first constructed a heterogeneous three-part information network that carries rich information, then systematically extracted path-based topological features from the information network, and finally used an unsupervised method to learn the optimal parameters related to different features to determine the relationship between symptoms and traditional Chinese medicine. In terms of joint recommendation integrating prior knowledge and clinical information, Yang Yun et al. [6] used a ridge regression algorithm based on Gaussian kernel to build a prescription system for traditional Chinese medicine for lung cancer. The prescription system used 2955 cases of traditional Chinese medicine lung cancer outpatient data to complete model training, enabling it to output prescriptions with high accuracy and ultimately be used as a reference for clinical treatment. The accuracy of the prescription was verified by expert evaluation of 108 actual cases. It was found that the accuracy of drugs with a frequency of occurrence greater than 300 times reached 62.9%, the recall rate was 80.2%, and the F1 value was 70.5%. Kuo, Yang et al. [7] proposed a multi-stage analysis method that integrates propensity case matching, complex network analysis, and Chinese medicine enrichment analysis to determine effective prescriptions for specific diseases (such as insomnia). First, propensity case matching was applied to match clinical cases. Then, core network extraction and Chinese medicine enrichment were combined to detect core effective Chinese medicine prescriptions. In terms of the identification and matching of TCM activity based on the fingerprint of TCM components, Chen H et al. [8] started from the perspective of the chemical chromatographic fingerprint of TCM components and established a superimposed multiple linear regression (SMLR) method to predict the biological activity of TCM from the chromatographic fingerprint. Compared with other methods, this method has better versatility and can provide support for the precise use of TCM.

[0006] Most of these methods focus on the phenotypic level, primarily exploring the relationship between symptoms and traditional Chinese medicine (TCM) at the text level. However, they lack molecular-level information. Understanding the molecular mechanisms underlying the therapeutic effects of TCM prescriptions is crucial for accurate TCM prescription recommendations. Therefore, it is crucial to develop a TCM prescription evaluation method that integrates both phenotypic and molecular information. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present invention proposes an artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks that integrates phenotypic and molecular information, thereby rationally and effectively establishing a high-precision artificial intelligence evaluation method for traditional Chinese medicine prescriptions.

[0008] To achieve the above objectives, the present invention provides an artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks that integrates phenotypic and molecular information, which is characterized by comprising the following steps:

[0009] Step 1: Digital and vector representation of diagnostic descriptions is achieved through mathematical modeling, where the diagnostic description includes chief complaint, current medical history, tongue image, and pulse information.

[0010] Step 2: By constructing a heterogeneous network of TCM molecules that integrates molecular information, we use network embedding representation to intelligently extract features from the heterogeneous network, obtain low-dimensional vector features of TCM, and then realize the vectorized representation of TCM prescriptions.

[0011] Step 3: Divide the training set and test set. The division of the training set and test set follows the principle of internal similarity of the disease, ensuring the similarity of both the diagnosis description and the Chinese medicine prescription.

[0012] Step 4: Stratified sampling of TCM prescriptions in the training set. Stratified sampling allows TCM prescriptions with high similarity, especially those with the same top-ranked TCM but slightly different lower-ranked TCM, to be included in the recommendation range. This is used to implement regression prediction of prescription recommendations in the constructed deep neural network model.

[0013] Step 5: Construct and train a neural network model. The neural network model consists of three parts: deep feature extraction of diagnostic description information based on convolutional neural networks, deep feature extraction of traditional Chinese medicine prescription information based on network embedding representation, and artificial intelligence evaluation of traditional Chinese medicine prescriptions based on convolutional neural networks. By training the neural network model to achieve the best, it can be used to intelligently recommend the best traditional Chinese medicine prescriptions in batches for a given diagnostic description.

[0014] Step 6: Neural network model evaluation. Model evaluation includes internal model evaluation and comparison with other baseline methods. Evaluation metrics include, but are not limited to, hit ratio (HR), area under the curve (AUC), and Spearman correlation.

[0015] According to a further embodiment of the present invention, the characteristics of the artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks and integration of phenotypic and molecular information include at least one of the following:

[0016] A) The diagnostic description mainly includes the chief complaint, current medical history, tongue condition, and pulse condition. One-dimensional convolution is performed on the text using convolution kernels of different lengths. The features extracted by the convolution kernels of different lengths are then spliced together and input into the neural network model as the features of a text segment for training.

[0017] B) The method used for deep feature extraction of TCM prescription information mainly includes low-dimensional embedding representation (Network Embedding), including: collecting TCM, compound, and target information from public databases to construct a TCM-compound-target heterogeneous network; using low-dimensional embedding representation methods to perform low-dimensional embedding representation on the heterogeneous network and extract the features of TCM, compounds, and targets; after measuring the features of TCM through low-dimensional embedding representation, further measuring the features of TCM prescriptions, where the features of TCM prescriptions are defined as the mean of the values of each dimension of the TCM contained in the prescription, that is, assuming that the TCM prescription contains m TCMs and the dimension of each TCM feature is d, the feature representation of the TCM prescription is:

[0018]

[0019] C) The division of the training and test sets follows the principle of internal disease similarity. For each disease, the training set accounts for 0.9% and the test set accounts for 0.1%. For each test set data item, it is ensured that at least one item in the training set for the current disease satisfies the diagnosis description similarity greater than or equal to 0.7 and the Traditional Chinese Medicine prescription similarity greater than or equal to 0.7.

[0020] D) Each training data entry contains three main pieces of information: a diagnosis, a disease, and a TCM prescription. We calculate the Jaccard similarity between the current TCM prescription and other TCM prescriptions for the current disease. We then perform stratified sampling based on the Jaccard value. Jaccard values range from 0 to 1, which we divide into 20 equal-length intervals. We sample each interval, and the sample size is proportional to the proportion of the sample size for the current disease to the total sample size: that is,

[0021]

[0022] The specific sampling rule is as follows: K = 50. Let the amount of traditional Chinese medicine prescriptions in this small interval be X. If X ≥ S, then S traditional Chinese medicine prescriptions are randomly sampled without replacement; if 0 < X < S, then all of X are sampled, and new S - X traditional Chinese medicine prescriptions are generated by sequentially deleting the traditional Chinese medicine at the end of the current traditional Chinese medicine prescription in reverse order. If X = 0, then new S traditional Chinese medicine prescriptions are generated by sequentially deleting the traditional Chinese medicine at the end of the current traditional Chinese medicine prescription in reverse order. The training set after sampling is 635120. Through this strategy, not only can the training samples be greatly expanded, but also the "sub - optimal" recommended traditional Chinese medicine prescription information of the same electronic medical record can be captured.

[0023] E) The neural network model includes three parts: deep feature extraction of diagnostic description information based on a convolutional neural network, deep feature extraction of traditional Chinese medicine prescription information based on network embedding representation, and artificial intelligence evaluation of traditional Chinese medicine prescriptions based on a convolutional neural network. Among them, according to an embodiment of the present invention, (1) The deep feature extraction of diagnostic description information based on a convolutional neural network includes: the diagnostic description first passes through an embedding layer with a dimension of 100, and then passes through three one - dimensional convolutional layers with 16 units each. The lengths of the convolutional kernels are 6, 7, and 8 respectively, and the stride is 10. One - dimensional MaxPooling layers are connected behind each convolutional layer, and the features extracted by these three MaxPooling layers are concatenated together as the features of the diagnostic description; (2) The deep feature extraction of traditional Chinese medicine prescription information based on network embedding representation includes: after the features of the traditional Chinese medicine prescription are extracted by the network embedding method, the length is normalized to 256, and then passes through two fully - connected layers with lengths of 128 and 64 in sequence, and the activation function is Relu for both; (3) The artificial intelligence evaluation of traditional Chinese medicine prescriptions based on a convolutional neural network includes: after the features of the diagnostic description and the features of the traditional Chinese medicine prescription are concatenated together, they pass through two one - dimensional convolutional layers and MaxPooling layers with 32 units in sequence, and finally output to two fully - connected layers with 32 and 16 units respectively, and the activation function is Relu for both, and the number of units in the output layer is 1.

[0024] According to one aspect of the present invention, there is provided an artificial intelligence evaluation method for traditional Chinese medicine prescriptions that fuses phenotypic and molecular information based on a deep neural network, which is characterized by including the following steps:

[0025] 1) Extract the features of the diagnostic description, where:

[0026] The diagnostic description includes the chief complaint, current medical history, tongue manifestation, and pulse condition.

[0027] The feature extraction of the diagnostic description is based on TextCNN, including performing one - dimensional convolution on the text using convolutional kernels of different lengths, and then concatenating the features extracted by the convolutional kernels of different lengths together as the features of a piece of text and inputting them into the network for training.

[0028] 2) Extract deep features of traditional Chinese medicine prescription information, including:

[0029] Information on traditional Chinese medicines, compounds, and targets is collected from public databases to construct a heterogeneous network of traditional Chinese medicines, compounds, and targets. A low-dimensional embedding representation method is used to perform a low-dimensional embedding representation on the heterogeneous network and extract the features of traditional Chinese medicines, compounds, and targets. After measuring the features of traditional Chinese medicines through low-dimensional embedding representation, the features of traditional Chinese medicine prescriptions are further measured. The features of traditional Chinese medicine prescriptions are defined as the mean of the values of each dimension of the traditional Chinese medicines contained in the prescription. That is, assuming that a traditional Chinese medicine prescription contains m traditional Chinese medicines and the dimension of each traditional Chinese medicine feature is d, the feature representation of the traditional Chinese medicine prescription is:

[0030]

[0031] 3) Divide the training set and test set, where:

[0032] The division of the training set and the test set follows the principle of internal similarity of the disease to ensure the similarity of both the diagnosis description and the Chinese medicine prescription. This includes: first, using Doc2Vec to train all the diagnosis descriptions, so as to measure the similarity between any two diagnosis descriptions; then, using Jaccard to measure the similarity between any two Chinese medicine prescriptions; finally, setting the training set ratio of each disease to 0.9 and the test set ratio to 0.1. For each test set data, it is guaranteed that in the training set of the current disease, there is at least one number that satisfies the diagnosis description similarity greater than or equal to 0.7 and the Chinese medicine prescription similarity greater than or equal to 0.7.

[0033] 4) Perform stratified sampling on the traditional Chinese medicine prescriptions in the training set, where:

[0034] Each sample contains three pieces of information: diagnosis description, disease, and TCM prescription. This includes: calculating the Jaccard similarity between the current TCM prescription and other TCM prescriptions for the current disease; then performing stratified sampling based on the Jaccard value, where the Jaccard value is distributed between 0 and 1; dividing the 0-1 range into 20 equal-length intervals, sampling in each interval, and sampling in proportion to the proportion of the sample size of the current disease to the total sample size, i.e.:

[0035]

[0036] The specific sampling rule is as follows: K = 50, and let the amount of traditional Chinese medicine prescriptions in this small interval be X. If X ≥ S, then S traditional Chinese medicine prescriptions are randomly sampled without replacement; if 0 < X < S, then all X are sampled, and new S - X traditional Chinese medicine prescriptions are generated by sequentially deleting the traditional Chinese medicine at the end of the current traditional Chinese medicine prescription in reverse order. If X = 0, then new S traditional Chinese medicine prescriptions are generated by sequentially deleting the traditional Chinese medicine at the end of the current traditional Chinese medicine prescription in reverse order. Thus, through this strategy, a large number of training samples can be greatly expanded and the "sub - optimal" traditional Chinese medicine prescription information with the same diagnostic description can be captured.

[0037] 5) Build a neural network model and train it, where:

[0038] The neural network model is divided into three parts:

[0039] Deep feature extraction of diagnostic description information based on a convolutional neural network, where: the diagnostic description first passes through an embedding layer with a dimension of 100, and then passes through three one - dimensional convolutional layers with 16 units respectively. The lengths of the convolutional kernels are 6, 7, and 8 respectively, and the stride is 10. A one - dimensional MaxPooling layer is connected behind each convolutional layer, and the features extracted by the three MaxPooling layers are concatenated together as the features of the diagnostic description.

[0040] Deep feature extraction of traditional Chinese medicine prescription information based on network embedding representation, where: after the features of the traditional Chinese medicine prescription are extracted by the network embedding method, the length is normalized to 256, and then passes through two fully - connected layers with lengths of 128 and 64 respectively, and the activation function is Relu for both.

[0041] Artificial intelligence evaluation of traditional Chinese medicine prescriptions based on a convolutional neural network, where: after the features of the diagnostic description and the features of the traditional Chinese medicine prescription are concatenated together, they pass through two one - dimensional convolutional layers and MaxPooling layers with 32 units respectively, and finally output to two fully - connected layers with 32 and 16 units respectively, and the activation function is Relu for both, and the number of units in the output layer is 1.

[0042] 6) Determine the evaluation index of the neural network model and evaluate it, where:

[0043] The evaluation indexes include the hit rate HR and the area under the curve AUC of the receiver operating characteristic curve ROC, including:

[0044] The hit rate HR is determined according to the following formula:

[0045]

[0046] Among them, the denominator GT is all the test sets, and the numerator NumberOfHits represents the number of samples that are hit.

[0047] The horizontal axis of the receiver operating characteristic curve ROC curve is the false positive rate FPR, and the vertical axis is the true positive rate TPR. The expression formulas are:

[0048]

[0049]

[0050] Among them, FP is the false positive rate, TP is the true positive rate, and TN is the true negative rate.

[0051] The evaluation method is that the higher the hit rate HR and / or AUC, the better the model. The evaluation process includes:

[0052] According to the above HR formula, the HR is directly calculated.

[0053] The calculation process of AUC includes:

[0054] Predict all the TCM prescriptions for each diagnosis description and the disease corresponding to the current diagnosis description, so that each diagnosis description has a known label vector, a predicted score vector, and the Jaccard similarity vector of the current TCM prescription and all the TCM prescriptions for the disease corresponding to the current diagnosis description. Sort the samples in descending order by the predicted scores.

[0055] When the Jaccard threshold is not set, TPR and FPR are calculated directly based on the known label vector and the predicted score vector. When the Jaccard threshold is set, the Jaccard similarity vector is divided from top to bottom according to the Jaccard threshold. Samples with Jaccard similarity greater than the Jaccard threshold are classified as correctly predicted samples, and samples with Jaccard similarity less than the threshold are classified as incorrectly predicted samples. The TPR and FPR at this time are calculated respectively to determine the AUC. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Schematic diagram for evaluation by TCM experts;

[0057] Figure 2 A distribution diagram of the mean scores given by TCM experts to prescriptions obtained using the artificial intelligence evaluation method for TCM prescriptions based on deep neural networks and fusion of phenotypic and molecular information of the present invention;

[0058] Figure 3 This is a hit rate chart of Chinese medicine expert scoring of prescriptions obtained using the artificial intelligence evaluation method for Chinese medicine prescriptions based on deep neural network fusion of phenotypic and molecular information of the present invention. DETAILED DESCRIPTION

[0059] This embodiment of the present invention uses the AI evaluation of TCM prescriptions by National Masters of Traditional Chinese Medicine as an example to design and implement an AI evaluation method for TCM prescriptions that integrates phenotypic and molecular information. We collected over 20,000 electronic medical records from the Yijishan Hospital of Anhui Wannan Medical College from 2013 to March 2020. By defining a series of rules, we selected 6,393 electronic medical records from 10 diseases, each representing a specific category, as the raw data. These 10 diseases include oligomenorrhea, asthenia, internal cancer, epigastric pain, mastitis, rheumatic arthritis, stomach distension, cough, arthritis, and insomnia. Each sample consists of a diagnosis description, a disease, and a corresponding TCM prescription. Most samples have diagnoses of 50 to 200 Chinese characters, and most prescriptions contain 10 to 25 Chinese herbs. The TCM Masters' diagnosis and prescription information primarily consists of three components: diagnosis description, disease, and TCM prescription. The implementation steps mainly include: feature extraction of diagnosis description, feature extraction of traditional Chinese medicine prescription, division of training set and test set, stratified sampling of traditional Chinese medicine prescription in training set, construction and training of neural network model, and evaluation of neural network model. Specific embodiments will explain the present invention in detail.

[0060] Example:

[0061] According to the present invention, an artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks and integrating phenotypic and molecular information comprises the following steps:

[0062] 1. Extracting features of diagnostic description

[0063] The diagnostic description mainly includes the chief complaint, current medical history, tongue condition and pulse condition. The idea of feature extraction of the diagnostic description is based on TextCNN[9]. It mainly uses convolution kernels of different lengths to perform one-dimensional convolution on the text, and then splices the features extracted by the convolution kernels of different lengths together and inputs them into the network as the features of a text for training. Before inputting into the model, it is necessary to extract the features of the diagnostic description. In this embodiment, we use the Tokenizer tool provided by Keras. Tokenizer can convert text into a sequence, that is, a list of word subscripts in the dictionary, thereby realizing the digital representation of the diagnostic description. In addition, Tokenizer also supports padding multiple diagnostic descriptions of unequal lengths to equal length to facilitate the unified use of the model. In this embodiment, the maximum length of the diagnostic description is 411, so all diagnostic results are padded with 0 to 411 at the end.

[0064] 2. Extracting Characteristics of Traditional Chinese Medicine Prescriptions

[0065] The method used for feature extraction of traditional Chinese medicine prescriptions is mainly low-dimensional embedding representation (Network Embedding). Traditional Chinese medicine prescriptions are composed of several traditional Chinese medicines according to the compatibility rule of "monarch, minister, assistant and envoy", and each traditional Chinese medicine has a specific content. This embodiment does not consider the dosage information of traditional Chinese medicines, but only considers which traditional Chinese medicines are included in the traditional Chinese medicine prescription, and focuses on incorporating phenotypic information and molecular information. In this embodiment, the inventors collected information on traditional Chinese medicines, compounds, and targets from the public databases TCMID

[10] , HIT

[11] , SymMap

[12] , and the group's self-built database HerbBioMap

[13] , and constructed traditional Chinese medicine-compound and compound-target networks respectively. The traditional Chinese medicine-traditional Chinese medicine network was constructed using the method independently developed by the research group

[14] . Compound similarity data was extracted from the PubChem database, and the compound similarity threshold was set to be greater than or equal to 90, thereby constructing a compound-compound network. Data were extracted from the protein interaction databases HPRD (Release 9)

[15] , BioGRID (2019 update)

[16] , IntAct

[17] , MINT (2012 update Homo sapiens)

[18] and STRING (V10.5)

[19] to construct a target-target network. Based on the above data collection, a heterogeneous network of Chinese medicine-compound-target was constructed. Node2Vec

[20] was used to perform a low-dimensional embedding representation on the heterogeneous network and extract the features of Chinese medicine, compounds and targets. After measuring the features of Chinese medicine through low-dimensional embedding representation, the features of Chinese medicine prescriptions can be measured. The features of Chinese medicine prescriptions are defined as the mean of the values of each dimension of the Chinese medicine contained in the prescription. That is, assuming that the Chinese medicine prescription contains m Chinese medicines and the dimension of each Chinese medicine feature is d, the feature representation of the Chinese medicine prescription is:

[0066]

[0067] 3. Divide the training set and test set

[0068] The division of the training set and the test set follows the principle of similarity within the disease, ensuring both the similarity of the diagnostic description and the similarity of the TCM prescription. First, all diagnostic descriptions are trained using Doc2Vec

[21] , so that the similarity between any two diagnostic descriptions can be measured. Then, the similarity between any two TCM prescriptions is measured using Jaccard. Finally, the training set of each disease is set to account for 0.9, and the test set accounts for 0.1. For each test set data, it is ensured that there is at least one number in the training set of the current disease that satisfies the diagnostic description similarity greater than or equal to 0.7 and the TCM prescription similarity greater than or equal to 0.7. The total sample size is 6393. After dividing the test set and training set, the training set has 5757 items and the test set has 636 items.

[0069] IV. Stratified Sampling of Traditional Chinese Medicine Prescriptions in the Training Set

[0070] We believe that the traditional Chinese medicine prescriptions written by national master physicians are the best, but other traditional Chinese medicine prescriptions for the current diseases cannot be completely excluded. The writing of traditional Chinese medicine prescriptions follows the compatibility rule of "sovereign, minister, assistant, and guide", and generally speaking, the less important the traditional Chinese medicine is in the traditional Chinese medicine prescription, the lower its importance. Therefore, we hope to design a rule to include traditional Chinese medicine prescriptions with relatively high similarity, especially those with the same top-ranked traditional Chinese medicines but slightly different bottom-ranked traditional Chinese medicines, in the recommended range, and give a specific score to measure such traditional Chinese medicine prescriptions. The rule we designed is called "stratified sampling of traditional Chinese medicine prescriptions", which is mainly carried out for the training set. Each sample mainly includes three aspects of information: diagnosis description, disease, and traditional Chinese medicine prescription. We calculate the Jaccard similarity between the current traditional Chinese medicine prescription and other traditional Chinese medicine prescriptions for the current disease. Then, stratified sampling is carried out according to the Jaccard value. The value of Jaccard ranges from 0 to 1, and we divide 0 - 1 into 20 equal-length small intervals. Sampling is carried out on each small interval, and the sampling quantity is proportional to the proportion of the sample quantity of the current disease in the total sample quantity, that is:

[0071]

[0072] The specific sampling rule is: K = 50, and let the number of traditional Chinese medicine prescriptions in this small interval be X. If X ≥ S, then randomly sample S traditional Chinese medicine prescriptions without replacement; if 0 < X < S, then all X are sampled, and new S - X traditional Chinese medicine prescriptions are generated by successively deleting the traditional Chinese medicines at the end of the current traditional Chinese medicine prescription in reverse order. If X = 0, then new S traditional Chinese medicine prescriptions are generated by successively deleting the traditional Chinese medicines at the end of the current traditional Chinese medicine prescription in reverse order. The training set after sampling is 635120. Through this strategy, not only a large number of training samples are expanded, but also the "sub-optimal" traditional Chinese medicine prescription information of the same diagnosis description can be captured.

[0073] V. Constructing and Training a Neural Network Model

[0074] The neural network model is mainly divided into three parts: deep feature extraction of diagnostic description information based on convolutional neural networks, deep feature extraction of traditional Chinese medicine prescription information based on network embedding representation, and artificial intelligence evaluation of traditional Chinese medicine prescriptions based on convolutional neural networks. (1) Deep feature extraction of diagnostic description information based on convolutional neural networks: The diagnostic description first passes through an embedding layer with a dimension of 100. Then it passes through three one-dimensional convolutional layers with 16 units, the length of the convolution kernel is 6, 7, and 8 respectively, and the step size is 10. Each convolutional layer is connected to a one-dimensional MaxPooling layer. The features extracted by the three MaxPooling layers are spliced together as the features of the diagnostic description. (2) Deep feature extraction of traditional Chinese medicine prescription information based on network embedding representation: After the features of the traditional Chinese medicine prescription are extracted by the network embedding method, the length is normalized to 256, and then passes through two fully connected layers with lengths of 128 and 64 respectively, and the activation function is ReLU. (3) Artificial Intelligence Evaluation of Traditional Chinese Medicine Prescriptions Based on Convolutional Neural Networks: After the features of the diagnosis description and the features of the traditional Chinese medicine prescription are spliced together, they pass through two one-dimensional convolutional layers with 32 units and a MaxPooling layer, and finally output to two fully connected layers with 32 and 16 units respectively. The activation function is ReLU and the number of output layer units is 1. Specifically:

[0075] Each diagnosis description consists of several characters. After the Embedding layer, the dimension of each character is D = 100. Assuming that the number of characters contained in the diagnosis description is N, each diagnosis description is represented by a randomly initialized D-dimensional vector:

[0076]

[0077] S i:j Represents the i-th to j-th characters in the diagnosis description, that is:

[0078]

[0079] The convolution layer includes convolution kernels of different sizes, and each size contains a large number of convolution kernels. The width of the convolution kernel is the same as the width of S, both of which are D = 100. Assuming that the height of the kth convolution kernel is H, the convolution kernel can be expressed as W k =R H×D ,Right now:

[0080]

[0081] The convolution operation is to extract the local features of S. We will give an example to illustrate the process of convolution operation. and s 1,1 Encounter, extracted features for:

[0082]

[0083] In the above formula, s i,j is the value of the j-th dimension of the i-th character of S, It is i,j The weight of is the bias term. Relu is a nonlinear activation function:

[0084] f(x)=max(0,x)

[0085] The convolution operation is W k With a certain step size S c Sliding from the top to the bottom of S, the resulting feature combination is:

[0086]

[0087] The pooling operation is similar to the convolution operation, the only difference is that the pooling operation calculates the mean or max value. We use the max type pooling operation MaxPooling. Assume that the height of the pooling kernel is H p , the step size is S p , then the output of the pooling operation is:

[0088]

[0089] in,

[0090]

[0091]

[0092]

[0093] The diagnostic description undergoes several convolution-pooling operations. After all convolution-pooling operations are completed, all the extracted features are connected in an end-to-end manner to obtain:

[0094]

[0095] in,

[0096] After the features of the traditional Chinese medicine prescription are extracted by the network embedding method, the length is normalized to 256, and then passed through two fully connected layers with lengths of 32 and 64 respectively, and the activation function is ReLU. The extracted features are:

[0097]

[0098] F T and C TThe concatenation is done to get G, which is then fed into the fully connected layer after a series of convolution-pooling operations. The weights of the fully connected layer are defined as W. F , the deviation term is b f , the output of the fully connected layer is:

[0099] y=W F ×G+b f

[0100] After several fully connected layers comes the final output layer, which has 1 unit and a Sigmoid activation function:

[0101]

[0102] The loss function is defined as:

[0103]

[0104] The loss function consists of two parts: the error term and the regularization term. λ is the regularization term coefficient. is the mean square error (MSE) of the sample, defined as:

[0105]

[0106] The above training process is a batch of samples, n is the size of the batch. predicted is the output of the model, i.e., the prediction score of the diagnosis description and the TCM prescription, is the value of the i-th diagnosis description-traditional Chinese medicine prescription combination. real The structure and y predicted The same indicates the strength of the relationship between the diagnosis description and the TCM prescription.

[0107] The training process uses the Adam algorithm, and the weight update rule is:

[0108]

[0109]

[0110]

[0111]

[0112]

[0113] Where t is the number of training steps, η is the learning rate, ∈=10e-8, and β1 and β2 are the forgetting factors for the gradient and second-order gradient, respectively. Dropout for the fully connected layers is set to 0.0005. Training is performed for 1500 epochs, with a learning rate of 1e-4 and a batch size of 256.

[0114] 6. Evaluating Neural Network Models

[0115] Neural network model evaluation consists of three parts: internal model evaluation, comparison with other methods, and expert review. The main evaluation metrics are hit ratio (HR) and the area under the receiver operating characteristic curve (ROC) (AUC).

[0116] Among them, HR is a commonly used indicator to measure recall rate, and the calculation formula is:

[0117]

[0118] The denominator is all the test sets, and the numerator represents the number of test sets.

[0119] The horizontal axis of the ROC curve is the false positive rate FPR, and the vertical axis is the true positive rate TPR.

[0120]

[0121]

[0122] Among them, FP is the false positive rate, TP is the true positive rate, and TN is the true negative rate.

[0123] The evaluation method is that the higher the hit rate HR and / or AUC, the better the model. The evaluation process includes:

[0124] According to the above HR formula, the HR is directly calculated.

[0125] The calculation process of AUC includes:

[0126] Predict all the TCM prescriptions for each diagnosis description and the disease corresponding to the current diagnosis description, so that each diagnosis description has a known label vector, a predicted score vector, and the Jaccard similarity vector of the current TCM prescription and all the TCM prescriptions for the disease corresponding to the current diagnosis description. Sort the samples in descending order by the predicted scores.

[0127] When the Jaccard threshold is not set, TPR and FPR are calculated directly based on the known label vector and the predicted score vector. When the Jaccard threshold is set, the Jaccard similarity vector is divided from top to bottom according to the Jaccard threshold. Samples with Jaccard similarity greater than the Jaccard threshold are classified as correctly predicted samples, and samples with Jaccard similarity less than the threshold are classified as incorrectly predicted samples. The TPR and FPR at this time are calculated respectively to determine the AUC.

[0128] In terms of model internal evaluation, in addition to incorporating molecular information into Node2Vec, we also tried other methods, including LINE

[22] , SDNE

[23] , and no molecular information. Compared with the model without molecular information, the addition of molecular information can significantly improve the prediction effect (Table 1). By comparing FordNet Node2Vec 、FordNet LINE 、FordNet SDNE and FordNet No molecule Hit rate and AUC of four models found by FordNet LINE Adding molecular information can greatly improve the top 1, top 5, top 10 and top 50 recommended Chinese medicine prescriptions. Compared with FordNet No molecule , FordNet LINE The top 1 increased by 24.24%, the top 5 increased by 20.40%, the top 10 increased by 17.28%, and the top 50 increased by 9.24%.

[0129] In comparison with other methods, we compared the baseline methods SVM, Random Forest, and LinearRegression (Table 1). Without setting FST, FordNet LINE Achieved the highest hit rate and FordNet LINE and FordNet No molecule All of them are higher than the baseline methods SVM, Random Forest, and LinearRegression. Without setting FST, FordNet LINE Also achieved the highest AUC (0.813), similarly, FordNet LINE and FordNet No moleculeThe AUCs of the two methods are also higher than those of the baseline methods SVM (AUC=0.563), RandomForest (AUC=0.751), and Linear Regression (AUC=0.513).

[0130] Table 1. Performance comparison of different methods

[0131]

[0132] In terms of expert evaluation, the inventor selected 50 electronic medical records of gastric distension after March 2020, and used the system of the present invention to recommend 10 traditional Chinese medicine prescriptions for each electronic medical record. Three traditional Chinese medicine experts from Yijishan Hospital of Wannan Medical College were invited to evaluate and score the recommended results on a scale of 1-5, with 1 being the least suitable and 5 being the most suitable. Figure 1 The evaluation results show that the scores are mainly concentrated above 4 points ( Figure 2 How to set the threshold to 4.5, the Top5 hit rate is nearly 100% ( Figure 3 The above results show that our model can effectively tap into the experience of national medical masters and accurately recommend traditional Chinese medicine prescriptions.

[0133] References:

[0134] [1]Min L.Analysis of Non-standard Prescription of TCM DecoctionPieces in Outpatient of Beijing First Hospital of Integrated Chinese and Western Medicine in 2013[J].Chinese Journal of Information on TraditionalChinese Medicine,2015,22(6):125-127.

[0135] [2]Jingyan Chen,Jingqi Yang,Fangming He,et al.A Study on UnreasonablePrescriptions in Outpatient Department in Our Hospital[J].Chinese Journal ofInformation on Traditional Chinese Medicine,2015,22(1):122-123.

[0136] [3]Yao L, Zhang Y, Wei B. An Evolution System for Traditional Chinese Medicine Prescription, Berlin, Heidelberg, F, 2014[C]. Springer Berlin Heidelberg.

[0137] [4]Li W, Yang Z. Exploration on Generating Traditional Chinese Medicine Prescriptions from Symptoms with an End-to-End Approach; proceedings of the CCF International Conference on Natural Language Processing and Chinese Computing, F, 2019 [C]. Springer.

[0138] [5]Jinpeng,Chen,Josiah,et al.Mining Symptom-Herb Patterns fromPatient Records Using Tripartite Graph[J].Evid-Based Compl Alt,2015,2015:1-14.

[0139] [6] Yang Yun, Ruan Chunyang, Pei Chaohan, et al. Exploration of building a traditional Chinese medicine prescription system for lung cancer by introducing artificial intelligence [J]. World Science and Technology - Modernization of Traditional Chinese Medicine, 2019, 21(5): 977-982.

[0140] [7] Yang K, Zhang R, He L, et al. Multistage analysis method for detection of effective herb prescription from clinical data [J]. Front Med, 2018, 12(2): 206-217.

[0141] [8] Chen H, Poon J, Poon S K, et al. Ensemble learning for prediction of the bioactivity capacity of herbal medicines from chromatographic fingerprints[J]. Bmc Bioinformatics, 2015, 16(Suppl 12): S4.

[0142] [9] Kim Y. Convolutional neural networks for sentence classification[J]. arXiv preprint arXiv:14085882, 2014:

[0143]

[10] Lin H, Xie D, Yu Y, et al. TCMID 2.0: a comprehensive resource for TCM[J]. Nucleic Acids Research, 2018, (D1): D1117 - D1120.

[0144]

[11] Hao Y, Li Y, Hong K, et al. HIT: linking herbal active ingredients to targets[J]. Nucleic Acids Research, 2011, 39(suppl_1): D1055–D1059.

[0145]

[12] Wu Y, Zhang F, Yang K, et al. SymMap: an integrative database of traditional Chinese medicine enhanced by symptom mapping[J]. Nucleic Acids Research, 2018, 47(D1): D1110–D1117.

[0146]

[13] Ouyang Zibo. Construction and mining of the HerbBioMap2.0 database platform[D]; Tsinghua University.

[0147]

[14] Li S,Zhang B,Jiang D,et al.Herb network construction and co-module analysis for uncovering the combination rule of traditional Chineseherbal formulae[J].BMC Bioinformatics,2010,11(Suppl 11):S6.

[0148]

[15] Keshava Prasad T S,Goel R,Kandasamy K,et al.Human ProteinReference Database--2009update[J].Nucleic Acids Res,2009,37(suppl_1):D767-D772.

[0149]

[16] Rose O,Chris S,Bobby-Joe B,et al.The BioGRID interactiondatabase:2019update[J].Nucleic Acids Res,2018,47(D1):D529–D541.

[0150]

[17] Samuel K,Bruno A,Lionel B,et al.The IntAct molecular interactiondatabase in 2012[J].Nucleic Acids Res,2011,40(D1):D841–D846.

[0151]

[18] Luana L,Leonardo B,Daniele P,et al.MINT,the molecular interactiondatabase:2012update[J].Nucleic Acids Res,2012,40(D1):D857–D861.

[0152]

[19] Damian S,Morris J H,Helen C,et al.The STRING database in 2017:quality-controlled protein–protein association networks,made broadlyaccessible[J].Nucleic Acids Res,2016,45(D1):D362–D368.

[0153]

[20] Grover A,Leskovec J.node2vec:Scalable Feature Learning forNetworks;proceedings of the the 22nd ACM SIGKDD International Conference,F,2016[C].

[0154]

[21] Le Q,Mikolov T.Distributed representations of sentences anddocuments;proceedings of the International conference on machine learning,F,2014[C].

[0155]

[22] Tang J,Qu M,Wang M,et al.Line:Large-scale information networkembedding;proceedings of the Proceedings of the 24th international conferenceon world wide web,F,2015[C].

[0156]

[23] Wang D,Peng C,Zhu W.Structural Deep Network Embedding;proceedingsof the Acm Sigkdd International Conference on Knowledge Discovery&DataMining,F,2016[C].

Claims

1. An artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks that integrates phenotypic and molecular information, characterized by It includes the following steps: 1) Extract the features of the diagnostic description, including the chief complaint, current history of illness, tongue manifestation, and pulse condition, 2) Extract the deep features of the traditional Chinese medicine prescription information, including: Collect information on traditional Chinese medicine, compounds, and targets from public databases to construct a traditional Chinese medicine-compound-target heterogeneous network; use a low-dimensional embedding representation method to perform low-dimensional embedding representation on the heterogeneous network, and extract the features of traditional Chinese medicine, compounds, and targets; further measure the features of the traditional Chinese medicine prescription, 3) Divide the training set and the test set, where: Divide the training set and the test set following the principle of internal similarity of diseases, 4) Perform stratified sampling on the traditional Chinese medicine prescriptions in the training set, where: Each sample contains a diagnostic description, a disease, and a traditional Chinese medicine prescription, including: calculate the Jaccard similarity between the current traditional Chinese medicine prescription and other traditional Chinese medicine prescriptions of the current disease; then perform stratified sampling according to the Jaccard value, where the Jaccard value ranges from 0 to 1; divide the 0-1 into 20 equal-length small intervals, and perform sampling on each small interval, and the sampling quantity is proportional to the proportion of the sample quantity of the current disease in the total sample quantity, that is: The sampling rule is: K = 50, and assume the number of traditional Chinese medicine prescriptions in this small interval is X. If X≥S, randomly sample S traditional Chinese medicine prescriptions without replacement; if 0<X<S, all X are sampled, and new S-X traditional Chinese medicine prescriptions are generated by sequentially deleting the traditional Chinese medicine at the end of the current traditional Chinese medicine prescription in reverse order. If X = 0, new S traditional Chinese medicine prescriptions are generated by sequentially deleting the traditional Chinese medicine at the end of the current traditional Chinese medicine prescription in reverse order; 5) Construct a neural network model and train it, 6) Evaluate, where: The evaluation method is that the higher the hit rate HR and / or AUC, the better the model. Here, AUC is the area under the curve, where the denominator GT is all the test sets, and the numerator NumberOfHits represents the number of samples that are hit, The horizontal axis of the receiver operating characteristic curve ROC curve is the false positive rate FPR, and the vertical axis is the true positive rate TPR. Their expression formulas are respectively: where FP is the false positive rate, TP is the true positive rate, and TN is the true negative rate, The calculation process of AUC includes: Predict each diagnostic description and all the traditional Chinese medicine prescriptions of the disease corresponding to the current diagnostic description, so that each diagnostic description has a known label vector, a predicted score vector, and a Jaccard similarity vector of the current traditional Chinese medicine prescription and all the traditional Chinese medicine prescriptions of the disease corresponding to the current diagnostic description. Sort the samples in descending order according to the predicted scores, For the case where no Jaccard threshold is set, directly calculate TPR and FPR based on the known label vector and the predicted score vector. For the case where a Jaccard threshold is set, divide the Jaccard similarity vector from top to bottom according to the Jaccard threshold, classify the samples with Jaccard similarity greater than the Jaccard threshold as correctly predicted samples, and classify the samples with Jaccard similarity less than the threshold as incorrectly predicted samples, and calculate the TPR and FPR at this time respectively, so as to determine the AUC.

2. The artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks and integration of phenotypic and molecular information according to claim 1, characterized in that: In step 1), the feature extraction of the diagnosis description is based on TextCNN, which involves performing one-dimensional convolution on the text using convolution kernels of different lengths. The features extracted by the convolution kernels of different lengths are then concatenated together and input into the network as the features of a text segment for training. In step 2), the characteristics of the Chinese medicine prescription are defined as the mean of the values of each dimension of the Chinese medicine contained in the prescription. That is, assuming that the Chinese medicine prescription contains m Chinese medicines and the dimension of each Chinese medicine feature is d, the characteristics of the Chinese medicine prescription are expressed as: Step 3) includes: first, using Doc2Vec to train all diagnostic descriptions, so as to measure the similarity between any two diagnostic descriptions; then, using Jaccard to measure the similarity between any two Chinese medicine prescriptions; finally, setting the training set to account for 0.9 of each disease and the test set to account for 0.1, and for each test set data, ensuring that in the training set of the current disease, there is at least one number that satisfies the diagnostic description similarity greater than or equal to 0.7 and the Chinese medicine prescription similarity greater than or equal to 0.7, In step 5), the neural network model is divided into three parts: Deep feature extraction of diagnostic description information based on convolutional neural networks, where the diagnostic description first passes through an embedding layer with a dimension of 100, and then passes through three one-dimensional convolutional layers with 16 units, convolution kernel lengths of 6, 7, and 8, and a stride of 10. Each convolutional layer is followed by a one-dimensional MaxPooling layer, and the features extracted by the three MaxPooling layers are spliced together as the features of the diagnostic description. Deep feature extraction of traditional Chinese medicine prescription information based on network embedding representation, wherein: the features of the traditional Chinese medicine prescription are extracted by the network embedding method, the length is normalized to 256, and then passed through two fully connected layers with lengths of 128 and 64 respectively, and the activation function is ReLU. Artificial intelligence evaluation of traditional Chinese medicine prescriptions based on convolutional neural networks, in which: the features of the diagnosis description and the features of the traditional Chinese medicine prescription are spliced together, and then pass through two one-dimensional convolutional layers with 32 units and a MaxPooling layer, and finally output to two fully connected layers with 32 and 16 units respectively. The activation function is Relu, and the number of units in the output layer is 1.

3. The artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks and integration of phenotypic and molecular information according to claim 2, characterized in that: In the operation of artificial intelligence evaluation of traditional Chinese medicine prescriptions based on convolutional neural networks, Each diagnosis description consists of several characters. After the Embedding layer, the dimension of each character is D = 100. Assuming that the number of characters contained in the diagnosis description is N, each diagnosis description is represented by a randomly initialized D-dimensional vector: S i:j Represents the i-th to j-th characters in the diagnosis description, that is: The convolution layer includes convolution kernels of different sizes. Each size contains a large number of convolution kernels. The width of the convolution kernel is the same as the width of S, which is D=100. Assuming that the height of the kth convolution kernel is H, the convolution kernel can be expressed as W k =R H×D ,Right now: The convolution operation is to extract local features of S, where: when and s 1,1 Encounter, extracted features for: In the above formula, s i,j is the value of the j-th dimension of the i-th character of S, It is i,j The weight of is the deviation term, Relu is a nonlinear activation function: f(x)=max(0,x) The convolution operation is W k With a certain step size S c Sliding from the top to the bottom of S, the resulting feature combination is: The pooling operation is similar to the convolution operation, the only difference is that the pooling operation calculates the mean or max value. Use the max type pooling operation MaxPooling, and set the height of the pooling kernel to H p , the step size is S p , then the output of the pooling operation is: in, The diagnostic description undergoes several convolution-pooling operations. After all convolution-pooling operations are completed, all the extracted features are connected in an end-to-end manner to obtain: in, After the features of the traditional Chinese medicine prescription are extracted by the network embedding method, the length is normalized to 256, and then passed through two fully connected layers with lengths of 32 and 64 respectively. The activation function is ReLU. The extracted features are: F T and C T After a series of convolution-pooling operations, it is input into the fully connected layer. The weight of the fully connected layer is W. F , the deviation term is b f , the output of the fully connected layer is: y=W F ×G+b f After several fully connected layers comes the final output layer, which has 1 unit and a Sigmoid activation function: The loss function is defined as: Loss=H yreal (y predicted )+λ‖W F ‖2 The loss function consists of two parts: error term and regular term, λ is the regular term coefficient, H yreal (y predicted ) is the mean square error (MSE) of the sample, defined as: The above training process is a batch of samples, n is the size of the batch, vector y predicted is the output of the model, i.e., the prediction score of the diagnosis description and the TCM prescription, is the value of the i-th diagnosis description-TCM prescription combination, y real The structure and y predicted The same indicates the strength of the relationship between the diagnosis description and the Chinese medicine prescription. The training process uses the Adam algorithm, and the weight update rule is: Where t is the number of training steps, η is the learning rate, ∈=10e-8, β1 and β2 are the forgetting factors of the gradient and second-order gradient, respectively, the dropout of the fully connected layer is set to 0.0005, the training epoch is 1500, the learning rate is 1e-4, and the batch size is 256.

4. The artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural network integration of phenotypic and molecular information according to claim 2 is characterized in that The step 1) comprises: Use text processing tools and data modeling to achieve mathematical and vectorized representation of diagnostic descriptions, and fill multiple diagnostic descriptions of different lengths to equal lengths to facilitate the unified use of the model. Use convolution kernels of different lengths to perform one-dimensional convolution on the text, then splice the features extracted by convolution kernels of different lengths together and input them into the network as the features of a piece of text for training.

5. The artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural network integration of phenotypic and molecular information according to claim 2, wherein step 2) comprises: Construct a heterogeneous network containing target information. The nodes in the heterogeneous network include traditional Chinese medicine, compounds, and targets. The edges of the heterogeneous network include traditional Chinese medicine-compound, compound-target, traditional Chinese medicine-traditional Chinese medicine, compound-compound, target-target, The features of each node in the heterogeneous network are extracted using low-dimensional embedding representation.

6. The artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural network integration of phenotypic and molecular information according to claim 2, wherein step 3) comprises: The diagnostic description similarity is measured using the text similarity measurement method, and the Chinese medicine prescription similarity is measured using the Jaccard similarity measurement method. The training set of each disease accounts for 0.9, and the test set accounts for 0.

1. For each test set data, it is ensured that there is at least one data in the training set of the current disease that satisfies the diagnostic description similarity greater than or equal to 0.7 and the Chinese medicine prescription similarity greater than or equal to 0.

7.

7. The artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural networks and integration of phenotypic and molecular information according to claim 2, characterized in that: The content of the nodes in the heterogeneous network contains at least one of traditional Chinese medicine, compound, and target. The content of the edge includes at least one of traditional Chinese medicine-traditional Chinese medicine, traditional Chinese medicine-compound, compound-compound, compound-target, and target-target.

8. A storage medium storing a computer program that enables a processor to execute the artificial intelligence evaluation method for traditional Chinese medicine prescriptions based on deep neural network integration of phenotypic and molecular information according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A modeling method of a convolution neural network model for determining a class of sub-image blocks in an image

    CN109410168A

  • Semantic similarity-based personalized traditional Chinese medicine diagnosis and treatment information and traditional Chinese medicine information intelligent matching method

    CN110929511A