A prediction method for TCM syndrome factors based on collaborative filtering
Through a multi-task neural collaborative filtering network based on collaborative filtering, combined with deep representation learning and collaborative filtering, the problem of low accuracy of existing traditional Chinese medicine proof prediction methods is solved, and higher proof prediction accuracy and wider application scope is achieved.
Patent Information
- Application Number
- CN202411382188.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-09-30
AI Technical Summary
The existing traditional Chinese medicine syndrome prediction methods have low accuracy and require manual extraction of symptom variables, which increases the complexity of the model and the coverage is limited to syndrome prediction of a certain patient.
A multi-task neural collaborative filtering network based on collaborative filtering is adopted. Through the deep feature extraction module, shallow feature extraction module, feature fusion module, collaborative filtering module and syndrome prediction module, the training sample set is constructed and data augmented, and synthetic positive samples are generated and synthetic negative samples are used to train the syndrome prediction model.
It improves the accuracy of certima prediction, reduces the complexity of manual feature extraction, expands the application scope of the model, and can more accurately predict multi-label certima.
Smart Images

Figure CN119339917B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical information processing, and in particular to a method for predicting TCM syndrome factors based on collaborative filtering. Background Art
[0002] Syndrome differentiation is a thinking and practice process that comprehensively analyzes the data obtained from the four examinations (observation, auscultation, inquiry, palpation) based on the theory of traditional Chinese medicine, clarifies the nature of the lesion and determines what kind of syndrome it is. According to the theory of traditional Chinese medicine, it analyzes the syndrome (symptoms, signs, etc.) and related data, identifies the syndrome elements such as the location and nature of the disease, and makes a syndrome diagnosis. Discussion and treatment, also known as treatment, is the thinking and practice process of establishing corresponding treatment principles, methods and prescriptions based on the results of syndrome differentiation, and selecting appropriate treatment methods and measures to treat the disease. Syndrome differentiation and discussion and treatment are two inseparable aspects of the diagnosis and treatment of diseases. Syndrome differentiation is to understand the disease and determine the syndrome; discussion and treatment is to establish the treatment method and prescription based on the results of syndrome differentiation. Syndrome differentiation is the premise and basis of discussion and treatment, which is the means and method of treating the disease and the test of whether the syndrome differentiation is correct. Therefore, syndrome differentiation and discussion and treatment are the embodiment of the combination of theory and practice, the specific application of the theoretical system of theory, method, prescription and medicine in clinical practice, and the basic principle of clinical diagnosis and treatment of traditional Chinese medicine.
[0003] With the development of technologies such as AI, deep learning, and reinforcement learning, there are methods to combine deep learning models with TCM dialectics. One method is to manually extract patient variables and then input them into the machine learning model to predict the patient's syndrome. This type mainly focuses on the classification of syndromes and does not involve the classification of syndrome elements. Syndrome elements are the basic elements that constitute the syndrome name, which can reflect the essential meaning of dialectics and can provide a more explanatory AI dialectical method. In addition, all existing methods require manual extraction of symptom variables, which increases the complexity of model use. At the same time, the coverage of the above methods is limited to the prediction of syndromes of a certain type of patient, which further limits the scope of application of the model.
[0004] Another model with a wider range of applications uses the patient's symptom description text as input information, extracts its semantic information through the Bert model, and uses the textCNN network to further optimize the extracted semantic information to predict the patient's syndrome. However, although this method does not require manual feature extraction, it ignores the multi-label prediction problem and label sparsity problem. At the same time, all the above methods only focus on extracting information on the patient side, ignoring the information contained in the syndrome and syndrome, so the prediction results are low in accuracy. Summary of the invention
[0005] In view of the above analysis, the embodiment of the present invention aims to provide a TCM syndrome factor prediction method based on collaborative filtering to solve the problem of low accuracy of existing syndrome factor prediction.
[0006] On the one hand, an embodiment of the present invention provides a method for predicting TCM syndrome factors based on collaborative filtering, comprising the following steps:
[0007] Obtain the patient's symptom description information and the corresponding syndromes and syndrome elements to construct a training sample set;
[0008] Constructing a multi-task neural collaborative filtering network, training the multi-task neural collaborative filtering network based on the training sample set to obtain a trained syndrome factor prediction model; the multi-task includes a syndrome factor prediction task and a syndrome prediction task;
[0009] The symptom description information of the patient to be predicted is input into the trained syndrome factor prediction model to obtain the syndrome factor prediction result of the patient.
[0010] Based on the further improvement of the above method, the multi-task neural collaborative filtering network includes:
[0011] A deep feature extraction module is used to extract the deep features of the sample relative to the syndrome based on the symptom description information of the sample;
[0012] A shallow feature extraction module is used to extract the linear features of the sample relative to the syndrome based on the symptom description information in the sample;
[0013] A feature fusion module, used for fusing the deep features and linear features of the sample relative to the same certificate element to obtain a fusion feature of the sample relative to the certificate element;
[0014] A collaborative filtering module, used for collaboratively filtering the fusion features of the sample relative to each syndrome factor and the features of the syndrome factor relative to the feature domain to which the patient belongs to predict the probability that the sample contains the syndrome factor;
[0015] The syndrome prediction module is used to predict the syndrome of the sample based on the fusion characteristics of the sample relative to the syndrome elements.
[0016] Based on the further improvement of the above method, the feature fusion module uses the following formula to obtain the fusion feature of the sample relative to the certificate element:
[0017]
[0018] in, represents the fusion feature of the i-th sample relative to the j-th certificate element, represents the deep features of the i-th sample relative to the j-th evidence element, It represents the linear characteristic of the i-th sample relative to the j-th certificate element, and β represents the weighting coefficient.
[0019] Based on the further improvement of the above method, the syndrome prediction module uses the following formula to predict the syndrome of the sample:
[0020]
[0021] in, represents the probability that the syndrome prediction module predicts the existence of the jth syndrome in the i-th sample, MLP(·) represents the multi-layer perceptron, and MLP(·) j represents the j-th dimension of the multilayer perceptron output vector, MLP(·) k represents the kth dimension of the multilayer perceptron output vector, N represents the number of elements, represents the fusion feature of the i-th sample relative to the j-th certificate element, represents the average representation information, N represents the number of syndrome elements, and Q represents the number of syndromes.
[0022] Based on the further improvement of the above method, after constructing the training sample set and before training the multi-task neural collaborative filtering network based on the training sample set, it also includes:
[0023] Data enhancement is performed on each original sample in the training sample set based on the symptom corpus to generate a synthetic positive sample and a synthetic negative sample corresponding to each original sample.
[0024] Based on the further improvement of the above method, data enhancement is performed on the training sample set based on the symptom corpus, including:
[0025] For each original sample in the training sample set, extract the original symptom description words of the original sample;
[0026] Generate positive symptom descriptors based on the original symptom descriptors; randomly generate negative symptom descriptors based on the symptom corpus and the negative affix library; concatenate the positive symptom descriptors and the negative symptom descriptors to obtain a synthetic positive sample corresponding to the original sample;
[0027] Randomly extracting symptom description words from the symptom corpus and concatenating them with the original symptom description words of the sample to obtain a synthetic negative sample corresponding to the original sample;
[0028] Get the enhanced training sample set.
[0029] Based on the further improvement of the above method, based on the symptom corpus and the negative affix library, multiple negative symptom description words are generated by multi-step random extraction;
[0030] For each step of random sampling, determining based on the first probability to generate a negative symptom descriptor using the first method or the second method;
[0031] If the first method is adopted to generate negative symptom descriptors, a negative affix is randomly extracted from the negative affix library, a symptom descriptor is randomly extracted from the symptom corpus based on the weight of the symptom descriptor in the symptom corpus, and the extracted negative affix and the symptom descriptor are concatenated to generate a negative symptom descriptor;
[0032] If the second method is used to generate negative symptom descriptors, symptom descriptors are randomly extracted from the symptom corpus, the negative word positions of the symptom descriptors are predicted based on the trained negative word position prediction model, negative affixes are randomly extracted from the negative affix library, and the extracted negative affixes are inserted into the negative word positions to generate negative symptom descriptors.
[0033] Based on the further improvement of the above method, the following formula is used to calculate the training loss of the multi-task neural collaborative filtering network:
[0034]
[0035] Among them, m represents the number of original samples in the current training batch, LE i represents the prediction loss of the certificate element of the i-th original sample, represents the prediction loss of the synthetic positive sample corresponding to the i-th original sample, LS i represents the syndrome prediction loss of the i-th original sample, represents the syndrome prediction loss of the synthetic positive sample corresponding to the i-th original sample, LD i Represents the distance loss between the i-th original sample and the corresponding synthetic positive sample and synthetic negative sample.
[0036] Based on the further improvement of the above method, the deep feature extraction module also includes a decoder, which is used to decode the deep features of the sample relative to the evidence element to obtain a decoding result of the sample; based on the decoding results corresponding to the original sample, the synthetic positive sample and the synthetic negative sample, the distance loss between the original sample and the corresponding synthetic positive sample and the synthetic negative sample is calculated.
[0037] Based on the further improvement of the above method, the following formula is used to calculate the distance loss between the i-th original sample and the corresponding synthetic positive sample and synthetic negative sample:
[0038]
[0039] Among them, D S (E L (x i )) represents the decoding result corresponding to the i-th original sample, D S (E L (x i+ )) represents the decoding result corresponding to the i-th synthetic positive sample, D S (EL (x i- )) represents the decoding result corresponding to the i-th synthetic negative sample, α represents the first probability, a i1 represents the number of negative symptom description words generated by the first method in the i-th synthetic positive sample, a i2 represents the number of negative symptom descriptors generated by the second method in the i-th synthetic positive sample, and ||·|| represents the 2-norm of the vector.
[0040] Compared with the prior art, the present invention constructs a training set by acquiring the patient's symptom description information, syndrome, and syndrome elements, and uses the constructed training set to train a multi-task neural collaborative filtering network, thereby introducing syndrome element information into the scope of model prediction, combining the advantages of collaborative filtering and deep representation learning. Deep representation learning can better mine the linear and nonlinear information contained in the data, and the collaborative filtering module can mine more association information between labels, thereby learning better features, and has a higher parallel efficiency, and can also improve the accuracy of multi-label joint prediction. By combining the syndrome element prediction task with the syndrome prediction task, the accuracy of syndrome element prediction is further improved through multi-task learning, thereby improving the accuracy of dialectical analysis.
[0041] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. In the entire drawings, the same reference symbols represent the same components;
[0043] Figure 1 The present invention is a flowchart of a method for predicting TCM syndrome factors based on collaborative filtering according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0045] A specific embodiment of the present invention discloses a method for predicting TCM syndrome factors based on collaborative filtering, such as Figure 1 As shown, the following steps are included:
[0046] S1. Obtain the patient's symptom description information and the corresponding syndromes and syndrome elements to construct a training sample set;
[0047] S2, constructing a multi-task neural collaborative filtering network, training the multi-task neural collaborative filtering network based on the training sample set to obtain a trained syndrome factor prediction model; the multi-task includes a syndrome factor prediction task and a syndrome prediction task;
[0048] S3. Input the symptom description information of the patient to be predicted into the trained syndrome factor prediction model to obtain the syndrome factor prediction result of the patient.
[0049] Compared with the prior art, the TCM syndrome factor prediction method based on collaborative filtering provided in this embodiment constructs a training set by acquiring the patient's symptom description information, syndrome, and syndrome factor, and uses the constructed training set to train a multi-task neural collaborative filtering network, thereby introducing syndrome factor information into the scope of model prediction, combining the advantages of collaborative filtering and deep representation learning. Deep representation learning can better mine the linear and nonlinear information contained in the data, and the collaborative filtering module can mine more association information between labels, thereby learning better features, and has a higher parallel efficiency, and can also improve the accuracy of multi-label joint prediction. By combining the syndrome factor prediction task with the syndrome prediction task, the accuracy of syndrome factor prediction is further improved through multi-task learning, thereby improving the accuracy of dialectical analysis.
[0050] During implementation, the symptom description information of each patient is obtained and converted into a vector representation using the pre-trained Bert model. The vector corresponding to the symptom description information of the patient in the i-th sample is denoted as x i .
[0051] The traditional collaborative filtering method is matrix factorization (MF), which decomposes patients and syndromes into latent variables of the same size and uses their inner product for prediction. That is, each patient and syndrome is associated with a real-valued latent feature vector, denoted as p i and q j , and then through p i and q j The interaction of factors is used to predict the probability of syndrome.
[0052] During implementation, M and N are assumed to represent the number of patients and the number of syndromes. i ,i∈{1,…,M} represents the text description of the symptom of the i-th patient. Based on the electronic medical record data, we define the patient-symptom interaction matrix Y∈R M×N The element y in i,j for:
[0053]
[0054] Therefore, the multi-label syndrome prediction problem can be formalized as a collaborative filtering problem to accurately predict whether there is a corresponding syndrome based on the patient's symptoms. Specifically, the model-based method can be abstracted as learning in is the probability that the i-th patient has the j-th syndrome, f is the parameterized prediction model, and Θ is the model parameter. However, the traditional method can only establish the correspondence between symptom information and syndromes, but cannot learn the relationship between syndromes and syndromes, as well as the information of the syndromes themselves. The present invention utilizes collaborative filtering methods, uses methods specific to patients and syndromes, and further improves the accuracy of syndrome prediction by adding the learning of syndrome features. From the patient's perspective, different feature vectors can be used for the same patient to predict different syndromes. From the syndrome perspective, when the same syndrome is used as an inner product with different patients, the syndrome feature vector corresponding to the patient can also be taken.
[0055] Therefore, the constructed multi-task neural collaborative filtering network includes:
[0056] A deep feature extraction module is used to extract the deep features of the sample relative to the syndrome based on the symptom description information of the sample;
[0057] A shallow feature extraction module is used to extract the linear features of the sample relative to the syndrome based on the symptom description information of the sample;
[0058] A feature fusion module, used for fusing the deep features and linear features of the sample relative to the same certificate element to obtain a fusion feature of the sample relative to the certificate element;
[0059] A collaborative filtering module, used for collaboratively filtering the fusion features of the sample relative to each syndrome factor and the features of the syndrome factor relative to the feature domain to which the patient belongs to predict the probability that the sample contains the syndrome factor;
[0060] The syndrome prediction module is used to predict the syndrome of the sample based on the fusion characteristics of the sample relative to the syndrome elements.
[0061] During implementation, a neural network combining deep and shallow structures is used to mine the characteristics of the patient's symptom description information. Different deep and shallow networks are used for this mining depending on the syndrome to be predicted.
[0062] Compared with traditional machine learning models, deep learning performs better in multi-label classification prediction problems and can mine more potential patterns in the data. At the same time, the model structure of deep learning can flexibly adjust the model structure according to the actual scenario and data characteristics, so that the model can be perfectly combined with the TCM syndrome prediction scenario. In addition, the deep learning model is more generalizable, that is, it can discover the correlation of rare features that are rare or even never appear in the final outcome. Therefore, the deep feature extraction module uses deep learning to extract the deep features of the samples relative to different syndromes.
[0063] During implementation, the deep feature extraction module has multiple submodules, and each submodule corresponds to a certificate element one by one, that is, a submodule is used to extract the deep features of a sample relative to the corresponding certificate element. The deep features of the i-th sample relative to the j-th certificate element are expressed as
[0064] During implementation, the submodule of the deep feature extraction module can use an existing encoder to convert the vector x corresponding to the symptom description information of the patient of the i-th sample i Input encoder for deep feature extraction.
[0065] Compared with deep learning models, linear models have stronger memory and can remember the historical distribution characteristics of data. For example, if two patients have similar symptom description information, then the two patients may have similar syndrome classification results. The memory-based model provides similar syndrome classification results for patients with similar symptom descriptions. Therefore, a shallow feature extraction module is used to extract the linear features of samples relative to syndromes.
[0066] During implementation, similarly, the shallow feature extraction module has multiple submodules, and the submodules correspond to the evidence elements one by one, that is, a submodule extracts the linear features of the sample relative to the corresponding evidence element. The submodules of the shallow feature extraction module can use a generalized linear model to transform the vector x i Input sub-model to extract linear features. The linear feature of the i-th sample relative to the j-th evidence element is expressed as
[0067] Specifically, the feature fusion module uses the following formula to obtain the fusion feature of the sample relative to the certificate element:
[0068]
[0069] in, represents the fusion feature of the i-th sample relative to the j-th certificate element, represents the deep features of the i-th sample relative to the j-th evidence element, It represents the linear characteristic of the i-th sample relative to the j-th certificate element, and β represents the weighting coefficient.
[0070] The feature fusion module fuses the deep features and linear features to obtain the fused features of the sample equivalent to the evidence elements.
[0071] During implementation, patients are grouped according to their gender, age and other information, and each group is a feature domain. Each syndrome has different feature vectors relative to different feature domains, which can give the model better expressiveness, learn better features, and improve prediction accuracy. It should be noted that the feature vectors of a syndrome relative to different feature domains can be set to the same vector initially. As the model is continuously trained, different feature vectors relative to different feature domains will be learned. The feature vector of the jth syndrome relative to the dth feature domain is expressed as If the ith patient belongs to the dth feature domain, the feature representation of the jth syndrome relative to the feature domain to which the ith patient belongs is:
[0072] In order to better explore the prediction effect of different classifications of syndromes, the embedding vectors q of disease location and disease nature are introduced respectively. p ,q n More specifically, if the syndrome to be predicted is classified as a disease location, q p Add to the feature vector of the syndrome. If the syndrome to be predicted is classified as pathological, q n Added to the feature vector of the syndrome, the feature of the jth syndrome relative to the feature domain of the ith patient is expressed as:
[0073]
[0074] During implementation, the collaborative filtering module performs the inner product of the fusion features of the sample relative to each syndrome element and the features of the syndrome element relative to the feature domain to which the patient belongs, and uses the sigmoid activation function to obtain the probability that the i-th patient has the j-th syndrome element. The process can be expressed as:
[0075]
[0076] Here, σ(·) represents the sigmoid function.
[0077] During implementation, the syndrome prediction module predicts the syndrome of the sample based on the fusion features of the sample relative to the syndrome elements.
[0078] During implementation, the syndrome prediction module can use a multi-layer perceptron. The syndrome prediction module uses the following formula to predict the syndrome of the sample:
[0079]
[0080] in, represents the probability that the syndrome prediction module predicts the existence of the jth syndrome in the i-th sample, MLP(·) represents the multi-layer perceptron, and MLP(·) j represents the j-th dimension of the multilayer perceptron output vector, MLP(·) k represents the kth dimension of the multilayer perceptron output vector, N represents the number of elements, represents the fusion feature of the i-th sample relative to the j-th certificate element, represents the average representation information, N represents the number of syndrome elements, and Q represents the number of syndromes.
[0081] During implementation, for the ith sample, the mean vector of the fusion features of the sample relative to all syndrome elements is calculated to obtain the average representation information; then the information is input into the multi-layer perceptron to predict the syndrome of the sample.
[0082] The accuracy of syndrome factor prediction is improved by complementing each other with syndrome factor prediction. During implementation, the syndrome prediction loss and syndrome factor prediction loss are calculated, and the gradient of the neural network is back-propagated through the mini-batch stochastic gradient descent algorithm to update the network parameters. It should be noted that the feature vector of the syndrome factor relative to the feature domain to which the patient belongs will also be updated.
[0083] When implemented, the loss is calculated using the following formula:
[0084]
[0085] Among them, y i,j Indicates whether the i-th sample has the j-th evidence element, represents the probability that the collaborative filtering module predicts that the i-th sample has the j-th evidence element, N represents the number of evidence elements, m represents the number of samples in the current training batch, and LS i represents the syndrome prediction loss of the i-th sample, LE i represents the prediction loss of the syndrome factor of the i-th sample, K represents the number of syndromes, and t i,j Indicates whether the i-th sample has the j-th syndrome, It indicates the probability that the syndrome prediction module predicts that the i-th sample has the j-th syndrome.
[0086] During implementation, in order to improve the robustness of the multi-task neural collaborative filtering network, after constructing the training sample set and before training the multi-task neural collaborative filtering network based on the training sample set, the method further includes:
[0087] Data enhancement is performed on each original sample in the training sample set based on the symptom corpus to generate a synthetic positive sample and a synthetic negative sample corresponding to each original sample.
[0088] Specifically, performing data enhancement on the training sample set based on the symptom corpus includes:
[0089] For each original sample in the training sample set, extract the original symptom description words of the original sample;
[0090] Generate positive symptom descriptors based on the original symptom descriptors; randomly generate negative symptom descriptors based on the symptom corpus and the negative affix library; concatenate the positive symptom descriptors and the negative symptom descriptors to obtain a synthetic positive sample corresponding to the original sample;
[0091] Randomly extracting symptom description words from the symptom corpus and concatenating them with the original symptom description words of the sample to obtain a synthetic negative sample corresponding to the original sample;
[0092] Get the enhanced training sample set.
[0093] During implementation, for each original sample in the training set, a positive sample and a negative sample are randomly generated, namely, a synthetic positive sample and a synthetic negative sample. The positive sample should have similar core information to the original sample, and an excellent feature extractor can extract the common core information, so that after passing through the same decoder, similar results can be output. The negative sample contains different core information from the original sample, so the features extracted from the negative sample and the original sample should be different, so that new information is added through data enhancement, so that the pre-trained model can extract features more accurately and improve the accuracy of subsequent tasks.
[0094] During implementation, for each original sample, its original symptom descriptor is extracted, and for each original symptom descriptor, a positive affix is randomly selected from the positive affix library and concatenated with the original symptom descriptor to generate a positive symptom descriptor.
[0095] It should be noted that the affirmative affix library stores affirmative affixes, such as "have", "exist", and "happen". Correspondingly, the negative affix library stores negative affixes, such as "no", "did not occur", "have not", and "do not exist".
[0096] The synthetic positive samples include not only positive symptom descriptors, but also negative symptom descriptors. Negative symptom descriptors are randomly generated based on the symptom corpus and the negative affix library.
[0097] Specifically, based on the symptom corpus and the negative affix library, multiple steps of random extraction are used to generate multiple negative symptom description words;
[0098] For each step of random sampling, determining based on the first probability to generate a negative symptom descriptor using the first method or to generate a negative symptom descriptor using the second method;
[0099] If the first method is adopted to generate negative symptom descriptors, a negative affix is randomly extracted from the negative affix library, a symptom descriptor is randomly extracted from the symptom corpus based on the weight of the symptom descriptor in the symptom corpus, and the extracted negative affix and the symptom descriptor are concatenated to generate a negative symptom descriptor;
[0100] If the second method is used to generate negative symptom descriptors, symptom descriptors are randomly extracted from the symptom corpus, the negative word positions of the symptom descriptors are predicted based on the trained negative word position prediction model, negative affixes are randomly extracted from the negative affix library, and the extracted negative affixes are inserted into the negative word positions to generate negative symptom descriptors.
[0101] The first method is simple and direct, but it does not cover the situation where some negative words are in the middle of a phrase. The second method is a little more complicated, but it can better simulate the situation where people put negative words in the middle of a phrase when speaking. The two methods complement each other and can better simulate the situation where negative words appear in reality.
[0102] During implementation, the number of extraction steps using multi-step random extraction can be pre-specified, for example, 3 steps, and can also be used as a hyperparameter of the model, which is specified before training and updated during training.
[0103] For each step of extraction, it is determined whether to generate a negative descriptor in the first manner or the second manner based on the first probability α.
[0104] If the first method is used to generate negative symptom descriptors, an affix is first randomly extracted from the negative affix library, such as "does not exist", and then symptom descriptors are randomly extracted from the symptom corpus based on the weights of the symptom descriptors in the symptom corpus, and the extracted negative affix and symptom descriptors are concatenated to generate negative symptom descriptors.
[0105] Specifically, the weight of each symptom description word in the symptom corpus is calculated in the following way:
[0106] For each symptom description word, a first weight of the symptom description word is obtained based on the number of times the symptom description word appears in all patient symptom description information;
[0107] Clustering all symptom descriptors, and obtaining the second weight of each symptom descriptor in each type based on the number of symptom descriptors in the type;
[0108] A final weight of each symptom descriptor is obtained based on the first weight and the second weight of the symptom descriptor.
[0109] During implementation, for example, if the symptom description appears k1 times in the symptom description information of all patients, then its first weight is k1.
[0110] Then cluster all symptom descriptors in the symptom prediction. For example, a pre-trained language model (such as Bert) can be used to extract features from symptom descriptors, and then cluster the extracted features. The number of clusters can be specified in advance, such as 10, and the features are clustered into 10 classes. The density clustering algorithm can also be used without specifying the number of clusters. For the category with the largest number of symptoms, the highest weight is assigned to each phrase in it. For the category with the second largest number, the second highest weight is assigned to each phrase in it, and so on, to obtain the second weight of each symptom descriptor. The first weight is added to the second weight to obtain the final weight of each symptom descriptor.
[0111] During implementation, for each symptom descriptor, the ratio of its weight to the sum of the weights of all symptom descriptors in the library is used as the extraction probability of the symptom descriptor, and symptom descriptors are randomly extracted from the symptom corpus according to the extraction probability. For example, if the symptom descriptor randomly extracted based on the weight is "facial numbness", the negative descriptor "facial numbness does not exist" is generated by splicing.
[0112] If the second method is used to generate negative descriptors, symptom descriptors are randomly extracted from the symptom corpus, the negative word positions of the symptom descriptors are predicted based on the trained negative word position prediction model, negative affixes are randomly extracted from the negative affix library, and the extracted negative affixes are inserted into the negative word positions to generate negative symptom descriptors.
[0113] Specifically, the trained negative word position prediction model is obtained in the following manner:
[0114] Extract all symptom description words containing negative words in the symptom corpus, use the positions of the negative words in the symptom description words as labels, and use the words after deleting the negative words in the symptom description words as input to construct a second training sample set;
[0115] A semantic model is constructed, and the semantic model is trained based on the second training sample set to obtain a trained negative word position prediction model.
[0116] Specifically, the trained negative word position prediction model is used to predict the position of negative affixes in words so as to insert negative affixes in words. During implementation, firstly, all symptom description words containing negative words in the symptom corpus are screened out, and the position of the negative words in the symptom description words is used as a label. For example, for the symptom description word "no fever", the position 1 of the negative word "none" is used as a label, and the word "fever" after deleting the negative word in the symptom description word is used as a sample input. Thus, a second training sample set is constructed. The constructed semantic model is trained based on the second training sample set to obtain a trained semantic model for predicting the insertion position of negative words.
[0117] After randomly extracting symptom descriptors from the symptom corpus, the corresponding embedding vectors are input into the trained negative word position prediction model to predict the negative word position of the symptom descriptors. Then, the randomly extracted negative affixes are inserted into the predicted negative word positions to generate negative symptom descriptors.
[0118] During implementation, the semantic model may adopt the Bert model.
[0119] The positive symptom descriptors generated based on the original symptom descriptors are combined with the negative symptom descriptors generated in multiple steps to form the symptom description information of the synthetic positive sample.
[0120] At the same time, synthetic negative samples corresponding to the original samples are generated to facilitate subsequent self-supervised learning. For example, multiple symptom description words (all different from the original symptom words) are randomly extracted from the symptom prediction library, and the extracted multiple symptom description words are concatenated with the original symptom words of the original sample to obtain the symptom description information of the synthetic negative sample corresponding to the sample.
[0121] During implementation, the synthesized positive samples and synthesized negative samples can also be further enhanced. This includes but is not limited to random swapping, synonym replacement, random insertion, etc. Random swapping means randomly selecting two phrases and swapping their positions. Synonym replacement means randomly selecting a word and replacing it with its synonym. Synonyms can be obtained from resources such as online dictionaries. Random insertion means randomly selecting a position in the text and randomly inserting a phrase or word.
[0122] The original samples are enhanced by synthesizing positive samples and synthetic negative samples to obtain an enhanced training sample set. The multi-task neural collaborative filtering network is trained based on the enhanced training sample set, which can greatly improve the extraction and distinction of affirmative and negative affixes in traditional Chinese medicine diagnosis and treatment, improve the feature extraction ability of the multi-task neural collaborative filtering network, and thus improve the accuracy of syndrome factor prediction.
[0123] It should be noted that the synthetic positive samples have the same labels as the original samples, the synthetic negative samples have no labels, and the synthetic negative samples do not perform prediction tasks.
[0124] During implementation, the pre-trained Bert model can be used to convert the symptom description information of the synthetic positive sample and the synthetic negative sample corresponding to the i-th original sample into vector representations, which are represented as and
[0125] The vector representations of the synthetic positive samples and the synthetic negative samples are respectively input into the deep feature extraction module to extract the deep features of the synthetic positive samples relative to each certificate element, and the deep features of the synthetic negative samples relative to each certificate element. The vector representations of the synthetic positive samples and the synthetic negative samples are respectively input into the shallow feature extraction module to extract the linear features of the synthetic positive samples relative to each certificate element, and the linear features of the synthetic negative samples relative to each certificate element. The deep features of the synthetic positive samples relative to each certificate element are fused with the linear features through the feature fusion module to obtain the fused features of the synthetic positive samples relative to each certificate element; the deep features of the synthetic negative samples relative to each certificate element are fused with the linear features through the feature fusion module to obtain the fused features of the synthetic negative samples relative to each certificate element. The process of obtaining the fused features of the synthetic positive samples and the synthetic negative samples is the same as the process of obtaining the fused features of the original samples mentioned above, and will not be repeated here.
[0126] Then, the fusion features of the synthetic positive sample relative to each syndrome factor and the features of the syndrome factor relative to the feature domain to which the patient belongs are collaboratively filtered to predict the probability of the presence of syndrome factors in the synthetic positive sample. The syndrome prediction module is used to predict the syndrome of the synthetic positive sample based on the fusion features of the synthetic positive sample relative to all syndrome factors. Since the synthetic positive sample has the same label as the original sample, the syndrome factor and syndrome prediction can be performed through the collaborative filtering module and the syndrome prediction model. The specific process is the same as the original sample prediction process, which will not be repeated here.
[0127] The features of the original sample should be as similar as possible to the synthetic positive sample and as far away from the synthetic negative sample as possible. Therefore, by adding the distance loss between the original sample and the corresponding synthetic positive sample and synthetic negative sample in the loss function, the feature extraction ability of the multi-task neural collaborative filtering network is improved.
[0128] Specifically, the training loss of the multi-task neural collaborative filtering network is calculated using the following formula:
[0129]
[0130] Among them, m represents the number of original samples in the current training batch, LE i represents the prediction loss of the certificate element of the i-th original sample, represents the prediction loss of the synthetic positive sample corresponding to the i-th original sample, LS i represents the syndrome prediction loss of the i-th original sample, represents the syndrome prediction loss of the synthetic positive sample corresponding to the i-th original sample, LD i Represents the distance loss between the i-th original sample and the corresponding synthetic positive sample and synthetic negative sample.
[0131] During implementation, in order to calculate the distance loss between the original sample and the corresponding synthetic positive sample and synthetic negative sample, the deep feature extraction module also includes a decoder, which is used to decode the deep features of the sample relative to the evidence element to obtain a decoding result of the sample; based on the decoding results corresponding to the original sample, the synthetic positive sample and the synthetic negative sample, the distance loss between the original sample and the corresponding synthetic positive sample and the synthetic negative sample is calculated.
[0132] During implementation, the decoder can use an existing decoder network.
[0133] During implementation, for the original sample, firstly, the deep features of each certificate element of the sample extracted by the deep feature extraction module are summed, and the summed result is input into the decoder for decoding to obtain the decoding result of the original sample. Similarly, for the synthetic positive sample, the deep features of each certificate element of the sample extracted by the deep feature extraction module are summed, and the summed result is input into the decoder for decoding to obtain the decoding result of the synthetic positive sample. For the synthetic negative sample, the deep features of each certificate element of the sample extracted by the deep feature extraction module are summed, and the summed result is input into the decoder for decoding to obtain the decoding result of the synthetic negative sample.
[0134] Specifically, the following formula is used to calculate the distance loss between the i-th original sample and the corresponding synthetic positive sample and synthetic negative sample:
[0135]
[0136] Among them, D S (E L (x i )) represents the decoding result corresponding to the i-th original sample, D S (E L (x i+ )) represents the decoding result corresponding to the i-th synthetic positive sample, D S (E L (x i- )) represents the decoding result corresponding to the i-th synthetic negative sample, α represents the first probability, a i1 represents the number of negative symptom description words generated by the first method in the i-th synthetic positive sample, a i2 represents the number of negative symptom descriptors generated by the second method in the i-th synthetic positive sample, and ||·|| represents the 2-norm of the vector.
[0137] It should be noted that the i-th synthetic positive sample is the synthetic positive sample corresponding to the i-th original sample, and the i-th synthetic negative sample is the synthetic negative sample corresponding to the i-th original sample. When minimizing the loss function, the negative sample part can maximize the gap between the representation of the original sample and the negative sample, so that the model can better learn the corresponding representation of the user. When minimizing the loss function, the positive sample part can minimize the gap between the representation of the original sample and the positive sample, so that the model can learn to understand how to deal with negative words. α is a learnable parameter. When minimizing the loss function, the optimal value can be found to optimize the probability of the first and second methods of generating negative words.
[0138] Existing data enhancement methods can all be regarded as disturbances to the original data, and do not increase the amount of additional information. In addition, in the application of traditional Chinese medicine, they cannot solve the problem that the sentences are similar in form but have opposite meanings when there are words such as "no" and "yes" in the text. Therefore, the extracted features are inaccurate, resulting in low accuracy in the prediction of the trained model. The present invention generates synthetic positive samples and synthetic negative samples to add new information, thereby extracting features more accurately and improving the accuracy of the prediction task. Since no other labels are introduced, the work of obtaining label data can be reduced, thereby reducing the workload of clinical diagnosis and treatment.
[0139] During implementation, the TCM syndrome factor prediction method based on collaborative filtering of the present invention can be deployed on an intelligent chip to achieve rapid model training and model reasoning.
[0140] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0141] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for predicting TCM syndrome factors based on collaborative filtering, characterized in that: The following steps are involved: Obtain the patient's symptom description information and the corresponding syndromes and syndrome elements to construct a training sample set; Constructing a multi-task neural collaborative filtering network, training the multi-task neural collaborative filtering network based on the training sample set to obtain a trained syndrome factor prediction model; the multi-task includes a syndrome factor prediction task and a syndrome prediction task; Input the symptom description information of the patient to be predicted into the trained syndrome factor prediction model to obtain the syndrome factor prediction result of the patient; The multi-task neural collaborative filtering network includes: A deep feature extraction module is used to extract the deep features of the sample relative to the syndrome based on the symptom description information of the sample; A shallow feature extraction module is used to extract the linear features of the sample relative to the syndrome based on the symptom description information in the sample; A feature fusion module, used for fusing the deep features and linear features of the sample relative to the same certificate element to obtain a fusion feature of the sample relative to the certificate element; A collaborative filtering module, used for collaboratively filtering the fusion features of the sample relative to each syndrome factor and the features of the syndrome factor relative to the feature domain to which the patient belongs to predict the probability that the sample contains the syndrome factor; The syndrome prediction module is used to predict the syndrome of the sample based on the fusion characteristics of the sample relative to the syndrome elements.
2. The TCM syndrome prediction method based on collaborative filtering according to claim 1, characterized in that: The feature fusion module uses the following formula to obtain the fusion feature of the sample relative to the certificate element: in, represents the fusion feature of the i-th sample relative to the j-th certificate element, represents the deep features of the i-th sample relative to the j-th evidence element, It represents the linear characteristic of the i-th sample relative to the j-th certificate element, and β represents the weighting coefficient.
3. The TCM syndrome factor prediction method based on collaborative filtering according to claim 1, characterized in that: The syndrome prediction module uses the following formula to predict the syndrome of the sample: in, represents the probability that the syndrome prediction module predicts the existence of the jth syndrome in the i-th sample, MLP(·) represents the multi-layer perceptron, and MLP(·) j represents the j-th dimension of the multilayer perceptron output vector, MLP(·) k represents the kth dimension of the multilayer perceptron output vector, N represents the number of elements, represents the fusion feature of the i-th sample relative to the j-th certificate element, represents the average representation information, and Q represents the number of syndromes.
4. The TCM syndrome factor prediction method based on collaborative filtering according to claim 1, characterized in that: After constructing the training sample set and before training the multi-task neural collaborative filtering network based on the training sample set, the method further includes: Data enhancement is performed on each original sample in the training sample set based on the symptom corpus to generate a synthetic positive sample and a synthetic negative sample corresponding to each original sample.
5. The TCM syndrome factor prediction method based on collaborative filtering according to claim 4, characterized in that: Performing data enhancement on the training sample set based on the symptom corpus includes: For each original sample in the training sample set, extract the original symptom description words of the original sample; Generate positive symptom descriptors based on the original symptom descriptors; randomly generate negative symptom descriptors based on the symptom corpus and the negative affix library; concatenate the positive symptom descriptors and the negative symptom descriptors to obtain a synthetic positive sample corresponding to the original sample; Randomly extracting symptom description words from the symptom corpus and concatenating them with the original symptom description words of the sample to obtain a synthetic negative sample corresponding to the original sample; Get the enhanced training sample set.
6. The TCM syndrome factor prediction method based on collaborative filtering according to claim 5, characterized in that: Based on the symptom corpus and the negative affix library, multiple negative symptom description words are generated by multi-step random extraction; For each step of random sampling, determining based on the first probability to generate a negative symptom descriptor using the first method or the second method; If the first method is adopted to generate negative symptom descriptors, a negative affix is randomly extracted from the negative affix library, a symptom descriptor is randomly extracted from the symptom corpus based on the weight of the symptom descriptor in the symptom corpus, and the extracted negative affix and the symptom descriptor are concatenated to generate a negative symptom descriptor; If the second method is used to generate negative symptom descriptors, symptom descriptors are randomly extracted from the symptom corpus, the negative word positions of the symptom descriptors are predicted based on the trained negative word position prediction model, negative affixes are randomly extracted from the negative affix library, and the extracted negative affixes are inserted into the negative word positions to generate negative symptom descriptors.
7. The TCM syndrome factor prediction method based on collaborative filtering according to claim 6, characterized in that: The training loss of the multi-task neural collaborative filtering network is calculated using the following formula: Among them, m represents the number of original samples in the current training batch, LE i represents the prediction loss of the certificate element of the i-th original sample, represents the prediction loss of the synthetic positive sample corresponding to the i-th original sample, LS i represents the syndrome prediction loss of the i-th original sample, represents the syndrome prediction loss of the synthetic positive sample corresponding to the i-th original sample, LD i Represents the distance loss between the i-th original sample and the corresponding synthetic positive sample and synthetic negative sample.
8. The TCM syndrome factor prediction method based on collaborative filtering according to claim 7, characterized in that: The deep feature extraction module also includes a decoder, which is used to decode the deep features of the sample relative to the evidence element to obtain a decoding result of the sample; based on the decoding results corresponding to the original sample, the synthetic positive sample and the synthetic negative sample, the distance loss between the original sample and the corresponding synthetic positive sample and the synthetic negative sample is calculated.
9. The TCM syndrome factor prediction method based on collaborative filtering according to claim 7, characterized in that: The following formula is used to calculate the distance loss between the i-th original sample and the corresponding synthetic positive sample and synthetic negative sample: Among them, D S (E L (x i )) represents the decoding result corresponding to the i-th original sample, D S (E L (x i+ )) represents the decoding result corresponding to the i-th synthetic positive sample, D S (E L (x i- )) represents the decoding result corresponding to the i-th synthetic negative sample, α represents the first probability, a i1 represents the number of negative symptom description words generated by the first method in the i-th synthetic positive sample, a i2 represents the number of negative symptom descriptors generated by the second method in the i-th synthetic positive sample, and ||·|| represents the 2-norm of the vector.
Citation Information
Patent Citations
Traditional Chinese medicine intelligent inquiry tongue diagnosis comprehensive system based on syndrome elements and deep learning
CN112216383A
Traditional Chinese medicine syndrome identification method and device based on intelligent algorithm
CN117766133A