A Traditional Chinese Medicine Syndrome Prediction Method in the Presence of Sample Imbalance

By combining syndromes with sufficient sample size as a separate major category, subdivided syndromes with small sample size into one major category, and building a training sample set and neural network model, the problem of inaccurate prediction of Chinese medicine syndromes caused by sample imbalance in the existing technology is solved, and a high accuracy and interpretability syndrome prediction is achieved.

CN119811690BActive Publication Date: 2025-06-10PEKING UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510293561.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-10
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art cannot accurately predict Chinese medicine syndromes when there is sample imbalance, resulting in the model focusing on types with large samples, affecting the ability of categories with small samples.

Method used

By taking syndromes with sufficient sample size as a separate major category, subdivided syndromes with small sample size are combined into one major category, and each patient marked the major categories of syndromes, subdivided syndromes and syndromes they exist, thereby constructing a training sample set and training a neural network model based on the constructed training sample set.

Benefits of technology

The problem of prediction inaccuracy caused by sample imbalance is solved, overfitting is avoided, the accuracy of syndrome prediction is improved, and the interpretability of the model is increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811690B_ABST
    Figure CN119811690B_ABST
Patent Text Reader

Abstract

The present invention relates to a traditional Chinese medicine syndrome prediction method in the presence of sample imbalance, belonging to the technical field of syndrome prediction, and solves the problem that accurate syndrome prediction cannot be performed in the prior art. The method includes: obtaining the text description of the patient's symptoms and converting it into initial features, and constructing a training sample set based on the patient's initial features, syndromes and syndrome elements; constructing a neural network model, the neural network model including a syndrome prediction task and a syndrome element prediction task; wherein the syndrome prediction task includes the prediction of the major syndromes and the subdivided syndromes of the samples; training the neural network model based on the constructed training sample set to obtain a syndrome prediction model; converting the text description of the symptoms of the patient to be predicted into initial features; predicting the probability of each major syndrome existing in the patient to be predicted based on the trained syndrome prediction model; and determining the syndrome of the patient to be predicted based on the probability of each major syndrome existing in the patient to be predicted. Accurate and interpretable syndrome prediction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of syndrome prediction, and in particular, to a traditional Chinese medicine syndrome prediction method in the presence of sample imbalance. Background Art

[0002] Syndrome differentiation is a thinking and practical process of comprehensively analyzing the data obtained from the four diagnostic methods (inspection, auscultation and olfaction, inquiry, and palpation) based on traditional Chinese medicine theory, clarifying the essence of the disease and determining what syndrome it belongs to. According to traditional Chinese medicine theory, it analyzes syndromes (symptoms, signs, etc.) and relevant data, differentiates syndrome elements such as the location and nature of the disease, and makes a syndrome name diagnosis. Treatment, also known as medical treatment, is a thinking and practical process of establishing corresponding treatment principles, methods, and prescription medications according to the results of syndrome differentiation, and selecting appropriate treatment means and measures to deal with the disease. Syndrome differentiation and treatment are two inseparable aspects that are interrelated in the process of diagnosing and treating diseases. Syndrome differentiation is to recognize the disease and determine the syndrome; treatment is to establish the treatment method and prescribe medications based on the results of syndrome differentiation. Syndrome differentiation is the premise and basis of treatment, and treatment is the means and method of treating diseases, and also a test of whether the syndrome differentiation is correct. Therefore, syndrome differentiation and treatment are the embodiment of the combination of theory and practice, the specific application of the theory system of principle, method, formula, and medicine in clinical practice, and also the basic principle of traditional Chinese medicine clinical diagnosis and treatment.

[0003] With the development of technologies such as deep learning and reinforcement learning, there have begun to be methods of combining deep learning models with traditional Chinese medicine syndrome differentiation. However, existing methods mainly focus on the syndrome classification of certain diseases and cannot perform large-scale syndrome prediction. When the syndrome prediction scope is expanded, it will be found that there is little training data for some rare syndromes, resulting in sample imbalance, which causes the model to focus on the type with a large number of samples and affects the recognition ability of the category with a small number of samples. Summary of the Invention

[0004] In view of the above analysis, the embodiments of the present invention aim to provide a traditional Chinese medicine syndrome prediction method in the presence of sample imbalance to solve the problem that accurate syndrome prediction cannot be performed when there is existing sample imbalance.

[0005] On the one hand, the embodiments of the present invention provide a traditional Chinese medicine syndrome prediction method in the presence of sample imbalance, including the following steps:

[0006] Obtain the text description of the patient's symptoms and convert it into initial features, and construct a training sample set based on the patient's initial features, syndromes, and syndrome elements; wherein, each syndrome with a sample quantity greater than the first threshold is used as a separate major syndrome; the syndromes with a sample quantity less than or equal to the first threshold are minor syndromes, and all minor syndromes are combined into one major category to form a small-sample major syndrome; each sample includes three groups of labels. The first group of labels annotates whether the patient has each major syndrome, the second group of labels is used to express whether the patient has each syndrome in the small-sample major syndrome, and the third group of labels is used to annotate whether the patient has each syndrome element;

[0007] Construct a neural network model, where the neural network model includes a syndrome prediction task and a syndrome element prediction task; among them, the syndrome prediction task includes the prediction of the major syndromes and sub-syndromes of the samples; train the neural network model based on the constructed training sample set to obtain a syndrome prediction model;

[0008] Convert the symptom text description of the patient to be predicted into initial features; based on the trained syndrome prediction model, predict the probability of each major syndrome existing in the patient to be predicted; determine the syndrome of the patient to be predicted based on the probability of each major syndrome existing in the patient to be predicted.

[0009] Based on a further improvement of the above method, the neural network model includes:

[0010] A first feature extraction module for performing deep feature extraction on the samples based on the initial features to obtain the first features of the samples;

[0011] A syndrome element classification module for predicting the syndrome elements of the samples based on the first features or the initial features;

[0012] A major syndrome prediction module for predicting the major syndromes of the samples based on the first features and the syndrome element prediction results;

[0013] A sub-syndrome prediction module for predicting the sub-syndromes in the major syndromes of small samples based on the first features.

[0014] Based on a further improvement of the above method, the sub-syndrome prediction module includes:

[0015] A feature decomposition module for decomposing the first features into invariant features and variable features;

[0016] A third classifier for predicting whether the samples do not have the major syndromes of small samples based on the invariant features;

[0017] A fourth classifier for predicting the probability of the samples having sub-syndromes in the major syndromes of small samples based on the variable features.

[0018] Based on a further improvement of the above method, determining the syndrome of the patient to be predicted based on the probability of each major syndrome existing in the patient to be predicted includes:

[0019] If the probability of the patient to be predicted having the major syndromes of small samples predicted is less than or equal to the second threshold, then the major syndrome with a predicted probability greater than the second threshold is the syndrome of the patient to be predicted;

[0020] Otherwise, based on the similarity between the patient to be predicted and the samples in the training sample library, determine the sub-syndromes of the patient to be predicted, and the sub-syndromes and the major syndromes with a predicted probability greater than the second threshold are the syndromes of the patient to be predicted.

[0021] Based on a further improvement of the above method, determining the syndrome of the patient to be predicted based on the probability of each major syndrome existing in the patient to be predicted, including:

[0022] Determining the probability of each sub-syndrome existing in the patient to be predicted based on the similarity between the patient to be predicted and each sub-syndrome;

[0023] The major syndromes and sub-syndromes with probabilities greater than the second threshold are the syndromes of the patient to be predicted.

[0024] Based on a further improvement of the above method, obtaining the similarity between the patient to be predicted and each sub-syndrome in the following way:

[0025] Extracting the samples with sub-syndromes in the training sample set as the samples to be compared; extracting the variable features and syndrome element prediction results of each sample to be compared based on the trained syndrome prediction model; splicing the variable features and syndrome element prediction results to obtain the features to be compared of each sample to be compared;

[0026] Splicing the variable features and syndrome element prediction results of the patient to be predicted to obtain the features to be compared of the patient to be predicted;

[0027] Calculating the features to be compared of the patient to be predicted and the features to be compared of each sample to be compared to obtain the similarity between the patient to be predicted and each sample to be compared;

[0028] Calculating the similarity between the patient to be predicted and each sub-syndrome based on the similarity between the patient to be predicted and each sample to be compared.

[0029] Based on a further improvement of the above method, calculating the similarity between the patient to be predicted and each sub-syndrome based on the similarity between the patient to be predicted and each sample to be compared, including:

[0030] For each sub-syndrome, taking the mean of the similarities between all the samples to be compared with the sub-syndrome and the patient to be predicted as the similarity between the patient to be predicted and the sub-syndrome.

[0031] Based on a further improvement of the above method, obtaining the probability of each sub-syndrome existing in the patient to be predicted in the following way based on the similarity between the patient to be predicted and each sub-syndrome:

[0032] According to the formula Calculating the initial prediction probability of the j-th sub-syndrome existing in the patient to be predicted ;

[0033] According to the formula Calculating the final prediction probability of the j-th sub-syndrome existing in the patient to be predicted;

[0034] Wherein, denotes the similarity between the patient to be predicted and the j-th sub-syndrome, and P denotes the probability that the patient to be predicted has a small-sample major syndrome predicted. denotes the number of sub-syndromes.

[0035] Based on the further improvement of the above method, the training loss of the neural network model is calculated using the following formula:

[0036] ;

[0037] where denotes the combined loss of syndrome element prediction and major syndrome prediction, denotes the sub-syndrome prediction loss, denotes the weight.

[0038] Based on the further improvement of the above method, the sub-syndrome prediction loss is calculated using the following formula:

[0039] ;

[0040] ; ;

[0041] where denotes the invariant feature of the i-th sample, denotes whether the i-th sample has the j-th sub-syndrome, denotes whether the model predicts that the i-th sample has the j-th sub-syndrome, denotes the third classifier, denotes whether the i-th sample has a small-sample major syndrome, which is 1 if yes and 0 otherwise, denotes the weight, n denotes the number of samples in the current training batch, denotes the number of sub-syndromes.

[0042] Compared with the prior art, the traditional Chinese medicine syndrome prediction method provided in this embodiment when there is sample imbalance constructs a training sample set by taking the syndromes with sufficient sample size as separate major categories, merging the sub-syndromes with small sample sizes as one major category, and annotating each patient with the major syndromes, sub-syndromes and syndrome elements they have. By training a neural network model based on the constructed training sample set, since the major syndromes are the merged syndromes with small sample sizes, there will be no problem of sample imbalance. At the same time, when predicting sub-syndromes, since the sample sizes of these sub-syndromes are the same, there will also be no imbalance problem. Therefore, the trained model can accurately predict major syndromes, sub-syndromes and syndrome elements, solving the problem of inaccurate prediction when there is sample balance, and avoiding overfitting, improving the prediction accuracy. In addition, the existing methods simply establish a direct mapping relationship from symptoms to syndromes, ignoring the analysis process in traditional Chinese medicine diagnosis, so they cannot provide good interpretability for their methods. In order to increase the interpretability of the model, this invention also collects the syndrome element information of patients. By adding the learning of the correspondence between syndrome elements and syndromes and the correspondence between symptom information and syndrome elements, the model can incorporate the information of syndrome elements when predicting syndrome classification, and at the same time give the corresponding syndrome element prediction results, thus achieving the purpose of improving the accuracy of syndrome prediction and providing interpretability for syndrome prediction simultaneously.

[0043] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages can be made obvious from the specification, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained from the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings are only for the purpose of showing specific embodiments, and are not considered as limiting the present invention. Throughout the drawings, the same reference signs represent the same components;

[0045] Figure 1 is a flowchart of the traditional Chinese medicine syndrome prediction method in the embodiment of the present invention when there is sample imbalance. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings, in which the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, not to limit the scope of the present invention.

[0047] A specific embodiment of the present invention discloses a traditional Chinese medicine syndrome prediction method when there is sample imbalance, as Figure 1 shown, including the following steps:

[0048] S1. Obtain the text description of the patient's symptoms and convert it into initial features. Construct a training sample set based on the patient's initial features, syndromes, and syndrome elements. Among them, each syndrome with a sample quantity greater than the first threshold is used as a separate major syndrome category; syndromes with a sample quantity less than or equal to the first threshold are sub-syndromes, and all sub-syndromes are combined into one major category to form a small-sample major syndrome. Each sample includes three groups of labels. The first group of labels indicates whether the patient has each major syndrome. The second group of labels is used to represent whether the patient has each sub-syndrome in the small-sample major syndrome. The third group of labels is used to indicate whether the patient has each syndrome element.

[0049] S2. Construct a neural network model, where the neural network model includes a syndrome prediction task and a syndrome element prediction task. Among them, the syndrome prediction task includes the prediction of the major syndrome and sub-syndrome of the sample. Train the neural network model based on the constructed training sample set to obtain a syndrome prediction model.

[0050] S3. Convert the text description of the symptoms of the patient to be predicted into initial features. Based on the trained syndrome prediction model, predict the probability that the patient to be predicted has each major syndrome. Determine the syndrome of the patient to be predicted based on the probability that the patient to be predicted has each major syndrome.

[0051] During implementation, the text description of the patient's symptoms can be converted into initial features through existing methods, that is, the input recognizable by the model. For example, a pre-trained BERT model can be used to obtain the initial features of each patient.

[0052] If the sample size is less than or equal to the first threshold, the syndrome is considered a small-sample syndrome, and all small-sample syndromes are combined together as a major syndrome category, denoted as a small-sample major syndrome. N is the number of syndrome types. , represents the number of sub-syndromes. represents the number of major syndromes. represents the number of syndromes with a sample size exceeding the first prediction.

[0053] The first threshold is determined according to the actual sample quantities of each syndrome collected. For example, the first threshold can be 10. Sub-syndromes are syndromes with a sample size less than the first threshold.

[0054] It should be noted that each patient may have multiple syndromes and multiple syndrome elements.

[0055] Compared with the prior art, the traditional Chinese medicine syndrome prediction method provided in this embodiment in the presence of sample imbalance constructs a training sample set by taking the syndromes with sufficient sample size as separate major categories, combining the subdivided syndromes with small sample sizes as one major category, and labeling each patient with the major syndromes, subdivided syndromes, and syndrome elements they have. By training a neural network model based on the constructed training sample set, since the major syndromes are combined from the syndromes with small sample sizes, there will be no problem of sample imbalance. At the same time, when predicting the subdivided syndromes, since the sample magnitudes of these subdivided syndromes are the same, there will also be no imbalance problem. Therefore, the trained model can accurately predict the major syndromes, subdivided syndromes, and syndrome elements, solving the problem of inaccurate prediction in the case of sample balance, and avoiding overfitting, improving the prediction accuracy. In addition, the existing methods simply establish a direct mapping relationship from symptoms to syndromes, ignoring the analysis process in traditional Chinese medicine diagnosis, so they cannot provide good interpretability for their methods. In order to increase the interpretability of the model, this invention also collects the syndrome element information of patients. By adding the learning of the correspondence between syndrome elements and syndromes and the correspondence between symptom information and syndrome elements, the model can incorporate the information of syndrome elements when predicting syndrome classification, and at the same time give the corresponding syndrome element prediction results, thereby achieving the purpose of improving the accuracy of syndrome prediction and providing interpretability for syndrome prediction simultaneously.

[0056] Specifically, the neural network model includes:

[0057] The first feature extraction module is used to perform deep feature extraction on the samples based on the initial features to obtain the first features of the samples;

[0058] The syndrome element classification module is used to predict the syndrome elements of the samples based on the first features or the initial features;

[0059] The major syndrome prediction module is used to predict the major syndromes of the samples based on the first features and the syndrome element prediction results;

[0060] The subdivided syndrome prediction module is used to predict the subdivided syndromes in the major syndromes with small samples based on the first features.

[0061] During implementation, the first feature extraction module can adopt an existing NLP model structure.

[0062] The syndrome element classification module, the major syndrome prediction module, and the subdivided syndrome prediction module can adopt existing multi-classifier structures.

[0063] During implementation, the major syndrome prediction module integrates the syndrome element information and predicts the major syndromes based on the first features and the syndrome element prediction results. Specifically, the major syndrome prediction module predicting the major syndromes of the samples based on the first features and the syndrome element prediction results includes:

[0064] After splicing the first feature with the syndrome element prediction structure, the prediction of the large-category syndromes of the samples is carried out.

[0065] During implementation, in order to improve the accuracy of the prediction of the sub-category syndromes, the sub-category syndrome prediction module includes:

[0066] A feature decomposition module, which is used to decompose the first feature into invariant features and variable features;

[0067] A third classifier, which is used to predict whether the sample does not have the large-category syndromes of the small samples based on the invariant features;

[0068] A fourth classifier, which is used to predict the probability that the sample has the sub-category syndromes in the large-category syndromes of the small samples based on the variable features.

[0069] First of all, we hope that the invariant features can better distinguish whether the sample has the large-category syndromes of the small samples, that is, extract the common features of the small-sample syndromes, and use the variable features to assist in distinguishing different sub-category syndromes in the large-category syndromes of the small samples.

[0070] During implementation, the feature decomposition module decomposes the first feature through a masking matrix. Introduce the masking matrix M, that is, the matrix size is the same as the first feature, the invariant feature , the variable feature . The value of the masking matrix M is obtained by model learning, and represents the first feature of the i-th patient.

[0071] During implementation, the third classifier is a binary classifier , and its output represents the probability that the sample does not have the large-category syndromes of the small samples.

[0072] The fourth classifier is a multi-classifier, denoted as , and its output is the probability that the sample has the k-th sub-category syndrome.

[0073] During implementation, the following formula is used to calculate the training loss of the neural network model:

[0074] ;

[0075] Among them, represents the combined loss of syndrome element prediction and large-category syndrome prediction, represents the sub-category syndrome prediction loss, represents the weight.

[0076] During implementation, the combined loss of syndrome element prediction and large-category syndrome prediction is: ;

[0077] ;

[0078] ;

[0079] wherein, represents the prediction loss of the major syndrome types, represents the prediction loss of syndrome elements, represents the label indicating whether the j-th major syndrome type exists in the i-th sample, represents the probability indicating whether the j-th major syndrome type exists in the i-th sample predicted by the model, represents the label indicating whether the k-th syndrome element exists in the i-th sample, represents the probability indicating whether the k-th syndrome element exists in the i-th sample predicted by the model, represents the weight, and n represents the number of samples in the current training batch, represents the number of major syndrome types, the number of types of syndrome elements.

[0080] Specifically, the prediction loss of the sub-syndromes is calculated using the following formula:

[0081] ;

[0082] ;

[0083] ;

[0084] wherein, represents the invariant features of the i-th sample, represents the label indicating whether the j-th sub-syndrome exists in the i-th sample, represents the probability indicating whether the j-th sub-syndrome exists in the i-th sample predicted by the model, represents the third classifier, represents the label indicating whether the i-th sample has a small-sample major syndrome type, which is 1 if yes, otherwise 0, represents the weight, and n represents the number of samples in the current training batch, represents the number of sub-syndromes.

[0085] By classifying small-sample major syndrome types and predicting sub-syndromes based on invariant features and variable features, and adding the classification and sub-syndrome prediction losses, the accuracy of sub-syndrome prediction is improved.

[0086] During implementation, the gradients of the neural network are backpropagated through the mini-batch stochastic gradient descent algorithm, and the parameters of the model are trained by minimizing the loss function. When the model converges, that is, when the preset loss accuracy or the number of iterations is reached, the training is stopped to obtain the syndrome prediction model.

[0087] For a patient to be predicted, convert the text description of their symptoms into initial features; based on the trained syndrome prediction model, predict the probability of the patient to be predicted having each major type of syndrome;

[0088] Then determine the syndrome of the patient to be predicted based on the probability of the patient to be predicted having each major type of syndrome.

[0089] In a specific embodiment, determining the syndrome of the patient to be predicted based on the probability of the patient to be predicted having each major type of syndrome includes:

[0090] If the probability of the patient to be predicted having a small-sample major type of syndrome is less than or equal to the second threshold, then the major type of syndrome with a predicted probability greater than the second threshold is the syndrome of the patient to be predicted;

[0091] Otherwise, determine the sub-syndrome of the patient to be predicted based on the similarity between the patient to be predicted and the samples in the training sample library, and the sub-syndrome and the major type of syndrome with a predicted probability greater than the second threshold are the syndrome of the patient to be predicted.

[0092] If the probability of the patient to be predicted having a small-sample major type of syndrome is less than or equal to the second threshold, it is considered that the patient to be predicted does not have a sub-syndrome, and directly take the major type of syndrome with a predicted probability greater than the second threshold as the syndrome of the patient to be predicted. If the probability of the patient to be predicted having a small-sample major type of syndrome is greater than the second threshold, then it is considered that the patient to be predicted has a sub-syndrome. Therefore, determine the sub-syndrome of the patient to be predicted according to the similarity between the patient to be predicted and the samples in the training sample library. During implementation, the second threshold can be set to 0.5, for example.

[0093] Specifically, determining the sub-syndrome of the patient to be predicted based on the similarity between the patient to be predicted and the samples in the training sample library includes:

[0094] S301. Extract the samples with sub-syndromes in the training sample set as the samples to be compared; based on the trained syndrome prediction model, extract the variable features and syndrome element prediction results of each sample to be compared; splice the variable features and syndrome element prediction results to obtain the features to be compared of each sample to be compared;

[0095] S302. Splice the variable features and syndrome element prediction results of the patient to be predicted to obtain the features to be compared of the patient to be predicted;

[0096] S303. Calculate the features to be compared of the patient to be predicted and the features to be compared of each sample to be compared to obtain the similarity between the patient to be predicted and each sample to be compared;

[0097] S304. Take the sub-syndromes existing in the samples to be compared with a similarity greater than the third threshold as the sub-syndromes of the patient to be predicted.

[0098] During implementation, first, extract the samples with sub-syndromes as the samples to be compared, that is, the patient to be predicted is compared with the samples to be compared.

[0099] Then, based on the syndrome element prediction model obtained through training, obtain the variable features of each sample to be compared and the syndrome element prediction results , and after splicing, obtain the features to be compared of the samples to be compared .

[0100] Similarly, after splicing the variable features and syndrome element prediction results of the patient to be predicted, obtain the feature q to be compared of the patient to be predicted.

[0101] During implementation, the cosine similarity formula can be used to calculate the similarity between the patient to be predicted and the i-th sample to be compared .

[0102] Take the sub-syndromes existing in the samples to be compared with a similarity greater than the third threshold as the sub-syndromes of the patient to be predicted, so as to accurately match the medical record closest to the patient to be predicted. During implementation, the third threshold can be set according to the prediction accuracy requirements.

[0103] In order to accurately obtain the prediction probability of the syndrome, in a specific embodiment, determine the syndrome of the patient to be predicted based on the probability of the patient to be predicted having each major syndrome, including:

[0104] S31. Determine the probability of the patient to be predicted having each sub-syndrome based on the similarity between the patient to be predicted and each sub-syndrome;

[0105] S32. The major syndromes and sub-syndromes with probabilities greater than the second threshold are the syndromes of the patient to be predicted.

[0106] Specifically, the following method is used to obtain the similarity between the patient to be predicted and each sub-syndrome:

[0107] S311. Extract the samples with sub-syndromes in the training sample set as the samples to be compared; based on the trained syndrome prediction model, extract the variable features and syndrome element prediction results of each sample to be compared; splice the variable features and syndrome element prediction results to obtain the features to be compared of each sample to be compared;

[0108] S312. Splice the variable features and syndrome element prediction results of the patient to be predicted to obtain the features to be compared of the patient to be predicted;

[0109] S313. Calculate the similarity between the features to be compared of the patient to be predicted and the features to be compared of each sample to be compared to obtain the similarity between the patient to be predicted and each sample to be compared;

[0110] S314. Calculate the similarity between the patient to be predicted and each sub-syndrome based on the similarity between the patient to be predicted and each sample to be compared.

[0111] Steps S311 - S313 are the same as S301 - S303, and will not be elaborated here.

[0112] Specifically, calculating the similarity between the patient to be predicted and each sub-syndrome based on the similarity between the patient to be predicted and each sample to be compared includes:

[0113] For each sub-syndrome, take the average of the similarities between all samples to be compared with the patient to be predicted that have this sub-syndrome as the similarity between the patient to be predicted and this sub-syndrome.

[0114] Specifically, based on the similarity between the patient to be predicted and each sub-syndrome, the probability that the patient to be predicted has each sub-syndrome is obtained in the following way:

[0115] According to the formula Calculate the initial prediction probability that the patient to be predicted has the j-th sub-syndrome ;

[0116] According to the formula Calculate the final prediction probability that the patient to be predicted has the j-th sub-syndrome;

[0117] Wherein, represents the similarity between the patient to be predicted and the j-th sub-syndrome, represents the probability that the patient to be predicted has the large-category syndrome of the small sample predicted, represents the number of sub-syndromes.

[0118] By fusing the similarity between the patient to be predicted and each sub-syndrome, and the probability that the patient to be predicted has the large-category syndrome of the small sample, the prediction probability that the patient has each sub-syndrome is accurately obtained.

[0119] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disk, a read-only memory, or a random access memory, etc.

[0120] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for predicting TCM syndromes when there is sample imbalance, characterized in that: The following steps are involved: Obtain the patient's symptom text description and convert it into initial features, and construct a training sample set based on the patient's initial features, syndromes and syndrome elements; wherein each syndrome with a sample size greater than a first threshold is regarded as a separate major syndrome; syndromes with a sample size less than or equal to the first threshold are subdivided syndromes, and all subdivided syndromes are merged into a major category to form a small sample major syndrome; each sample includes three groups of labels, the first group of labels marks whether the patient has each major syndrome, the second group of labels is used to express whether the patient has each syndrome in the small sample major syndrome, and the third group of labels is used to mark whether the patient has each syndrome element; Constructing a neural network model, the neural network model includes a syndrome prediction task and a syndrome factor prediction task; wherein the syndrome prediction task includes the prediction of the major syndromes and subdivided syndromes of the sample; training the neural network model based on the constructed training sample set to obtain a syndrome prediction model; The text description of the symptoms of the patient to be predicted is converted into initial features; the probability of each major category of syndrome in the patient to be predicted is predicted based on the trained syndrome prediction model; the syndrome of the patient to be predicted is determined based on the probability of each major category of syndrome in the patient to be predicted.

2. The method for predicting TCM syndromes when there is sample imbalance according to claim 1, characterized in that: The neural network model includes: A first feature extraction module, used for performing deep feature extraction on the sample based on the initial feature to obtain a first feature of the sample; A certificate element classification module, used for predicting the certificate element of the sample based on the first feature or the initial feature; A major syndrome prediction module, used to predict the major syndrome of the sample based on the first feature and the syndrome element prediction result; A subdivided syndrome prediction module is used to predict subdivided syndromes in a small sample of large-category syndromes based on the first feature.

3. The method for predicting TCM syndromes when there is sample imbalance according to claim 2, characterized in that: The subdivided syndrome prediction module comprises: A feature decomposition module, used for decomposing the first feature into an invariant feature and a variable feature; A third classifier is used to predict whether the sample does not have a small sample large category syndrome based on the invariant feature; The fourth classifier is used to predict the probability of the sample containing a subdivided syndrome in a small sample large category syndrome based on the variable feature.

4. The method for predicting TCM syndromes when there is sample imbalance according to claim 3, characterized in that: The syndrome of the patient to be predicted is determined based on the probability of each major syndrome of the patient to be predicted, including: If the predicted probability of the patient to be predicted having a large category of syndromes with a small sample is less than or equal to the second threshold, the large category of syndromes with a predicted probability greater than the second threshold is the syndrome of the patient to be predicted; Otherwise, the subdivided syndrome of the patient to be predicted is determined based on the similarity between the patient to be predicted and the samples in the training sample library, and the subdivided syndrome and the major syndrome whose prediction probability is greater than the second threshold are the syndromes of the patient to be predicted.

5. The method for predicting TCM syndromes when there is sample imbalance according to claim 3, characterized in that: The syndrome of the patient to be predicted is determined based on the probability of each major syndrome of the patient to be predicted, including: Determine the probability that the patient to be predicted has each subdivided syndrome based on the similarity between the patient to be predicted and each subdivided syndrome; The major syndromes and subdivided syndromes with probabilities greater than the second threshold are the syndromes of the patient to be predicted.

6. The method for predicting TCM syndromes when there is sample imbalance according to claim 5, characterized in that: The similarity between the patient to be predicted and each subdivided syndrome is obtained in the following way: Extracting samples with subdivided syndromes from the training sample set as samples to be compared; extracting variable features and syndrome factor prediction results of each sample to be compared based on the trained syndrome prediction model; splicing the variable features and syndrome factor prediction results to obtain the features to be compared of each sample to be compared; The variable features of the patient to be predicted and the syndrome factor prediction results are spliced ​​to obtain the features to be compared of the patient to be predicted; Calculate the features to be compared of the patient to be predicted and the features to be compared of each sample to be compared to obtain the similarity between the patient to be predicted and each sample to be compared; The similarity between the patient to be predicted and each subdivided syndrome is calculated based on the similarity between the patient to be predicted and each sample to be compared.

7. The method for predicting TCM syndromes when there is sample imbalance according to claim 6, characterized in that: The similarity between the patient to be predicted and each subdivided syndrome is calculated based on the similarity between the patient to be predicted and each sample to be compared, including: For each subdivided syndrome, the average of the similarities between all samples to be compared and the patient to be predicted that have the subdivided syndrome is taken as the similarity between the patient to be predicted and the subdivided syndrome.

8. The method for predicting TCM syndromes when there is sample imbalance according to claim 6, characterized in that: Based on the similarity between the patient to be predicted and each subdivided syndrome, the probability of the patient to be predicted having each subdivided syndrome is obtained in the following way: According to the formula Calculate the initial prediction probability that the patient to be predicted has the jth subdivision syndrome ; According to the formula Calculate the final prediction probability of the jth subdivided syndrome in the patient to be predicted; in, represents the similarity between the patient to be predicted and the jth subdivided syndrome, P represents the probability that the patient to be predicted has a small sample of large-category syndromes, Indicates the number of subdivided syndromes.

9. The method for predicting TCM syndromes when there is sample imbalance according to claim 3, characterized in that: The training loss of the neural network model is calculated using the following formula: ; in, It represents the joint loss of syndrome factor prediction and major syndrome prediction. represents the prediction loss of subdivided syndromes, Represents weight.

10. The method for predicting TCM syndromes when there is sample imbalance according to claim 9, characterized in that: The following formula was used to calculate the prediction loss of the subdivided syndromes: ; ; ; in, represents the invariant feature of the i-th sample, Indicates whether the i-th sample has the j-th subdivision syndrome, Indicates whether the i-th sample predicted by the model has the j-th subdivision syndrome. represents the third classifier, Indicates whether the ith sample has a small sample and a large category syndrome. If yes, it is 1, otherwise it is 0. represents the weight, n represents the number of samples in the current training batch, Indicates the number of subdivided syndromes.

Citation Information

Patent Citations

  • Syndrome differentiation method and apparatus for traditional Chinese medicine syndrome elements

    CN107122583A

  • Unbalanced data processing method, device and equipment for medical prediction model

    CN116978571A