Traditional Chinese medicine syndrome differentiation method and device based on difficult text mining
By constructing a Chinese medicine dialectical model, using BERT for text and symptom encoding, and combining difficult text mining and symptom projection, the problem of difficult traditional Chinese medicine dialectical devices being difficult to understand ancient text medical records and confusing similar symptoms is solved, and a more accurate and efficient symptom judgment of Chinese medicine is achieved.
Patent Information
- Application Number
- CN202510571901.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-19
AI Technical Summary
The existing traditional Chinese medicine syndrome differentiation devices are difficult to understand the meaning of ancient Chinese medical records, and are prone to confuse similar traditional Chinese medicine syndrome types, resulting in poor dialectic effect and the inability to provide doctors with accurate and rapid judgment on traditional Chinese medicine syndrome types.
Build a Chinese medicine dialectical model, and use BERT to encode text and proof type through the proof-type coding module, text coding module, proof-type projection module, understanding and migration module and prediction module, and combine difficult text mining and proof-type projection to improve the model's ability to understand ancient text medical records and discriminate similar proof-types.
It improves the ability of traditional Chinese medicine dialectical model to understand ancient medical records and distinguish similar criterion types, reduces the workload of doctors, and improves the accuracy and efficiency of traditional Chinese medicine criterion types.
Smart Images

Figure CN120511029A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method and device for traditional Chinese medicine syndrome differentiation based on difficult text mining. Background Art
[0002] As confidence in traditional culture grows, public confidence in the importance of Traditional Chinese Medicine (TCM) for diagnosis and treatment is growing. TCM prescriptions, in particular, utilize natural ingredients meticulously processed and prepared to maximize their efficacy. This stands in stark contrast to the opaque and often incomprehensible manufacturing processes of Western medicine. The production methods of TCM herbs derived from natural resources not only reduce the use of chemical additives but also make the medicines gentler and safer, further strengthening public confidence in the efficacy and safety of TCM.
[0003] With the release and implementation of the "Several Policy Measures on Accelerating the Development of Traditional Chinese Medicine," the trend of complementary and coordinated development between TCM and Western medicine has been further accelerated. A growing number of researchers are working to optimize and improve various deep learning models in Western medicine for the unique processes of TCM. Among these, TCM syndrome differentiation, as a core task in TCM, has attracted increasing attention. Similar to the automatic coding of ICDs in Western medicine, TCM syndrome differentiation primarily involves analyzing the data, symptoms, and signs collected through the Four Diagnoses (inspection, auscultation, questioning, and palpation) to clarify the cause, nature, and location of a patient's illness, as well as the relationship between cold and heat. The patient's condition is then diagnosed based on the Eight-Principle Syndrome Differentiation, ultimately resulting in a classification of a syndrome. This process places high demands on physicians, who often require extensive clinical experience to accurately determine a patient's TCM syndrome type and tailor treatment accordingly. However, due to the lack of attention paid to traditional Chinese medicine in the past, traditional Chinese medicine hospitals often need an experienced older doctor to teach 5-10 young doctors, and young doctors have low public recognition when they diagnose and treat patients alone. Studies have shown that 80% of the public prefer older traditional Chinese medicine doctors when registering.
[0004] During the process of adapting Western medicine to the Sinicization model, a unique TCM syndrome differentiation device has been designed. This device improves the accuracy of TCM syndrome identification while also assisting young physicians. Physicians simply input clinical text containing the patient's four diagnostic information into the TCM syndrome differentiation device, which quickly analyzes the patient's potential TCM syndrome, thereby assisting physicians with subsequent medication and treatment. Existing TCM syndrome differentiation devices typically treat TCM syndrome differentiation as a single-label classification task, defaulting to analyzing the patient's primary syndrome. This can easily overlook the complexity of some patients' conditions. The emphasis on both primary and secondary syndromes is also reflected in the syndrome differentiation process. Primary syndromes are the core manifestations of the disease and serve as the primary basis for determining the nature of the condition and the direction of treatment, such as typical symptoms of a cold, such as fever and aversion to cold. Secondary syndromes, on the other hand, are other secondary manifestations that accompany the disease, such as sore throat, dry mouth, or fatigue. This combined approach not only reflects the complexity of the condition but also reveals its full picture.
[0005] Due to the unique characteristics of Traditional Chinese Medicine (TCM), physicians often describe patients' conditions in classical Chinese when writing medical records. For example, "Natural expression, rosy complexion, normal body shape, posture, no unusual odor, dark purple tongue with slightly thick and greasy coating, mildly varicose sublingual veins, deep and thready pulse..." These classically written descriptions pose a significant challenge for models to understand their underlying meaning. Furthermore, researchers tend to frame the classification task as a matching problem between TCM syndrome types and clinical text. This matching approach encourages researchers to focus on specific information in clinical text that strongly correlates with specific TCM syndrome types, thereby enhancing the interpretability of model predictions and providing a more intuitive and reliable theoretical basis for clinical practice. However, this matching approach also introduces a new bias: after encoder processing, the vector representations of different TCM syndrome types tend to be overly concentrated in the embedding space, which can easily lead to confusion between TCM syndrome types during model prediction. This phenomenon can weaken the model's discriminative ability and affect the reliability of predictions. In summary, the existing data set construction method is difficult to adapt to actual medical scenarios, and the existing TCM syndrome differentiation devices have difficulty understanding the corresponding meaning of ancient Chinese texts in clinical texts. In addition, TCM syndrome differentiation devices are prone to confusion between similar TCM syndromes when making predictions, making it difficult to achieve satisfactory syndrome differentiation results. Summary of the Invention
[0006] To address the shortcomings of existing TCM syndrome differentiation methods, the present invention proposes a TCM syndrome differentiation method and device based on difficult text mining. On the one hand, it fully considers the varying degrees of difficulty for the model to understand characters of varying difficulty levels during the classification process. On the other hand, it individually maps the representations of similar TCM syndrome types, thereby mapping the more concentrated representations in the vector space to a relatively discrete space. This improves the model's discriminative ability, helping physicians make accurate and reliable TCM syndrome predictions and improving medical efficiency. The present invention proposes a model structure for TCM syndrome differentiation that fully considers the degree of difficulty of characters of varying difficulty levels during the classification process. Through a comprehension transfer module, the representations of more easily understood characters drive the learning of the representations of more difficult characters, thereby helping the model understand the more difficult-to-understand portions of classical clinical text. Furthermore, considering the high similarity of different TCM syndrome types in the vector space, the present invention designs a syndrome projection module that redirects the TCM syndrome type representations in the vector space to a new subspace, thereby increasing the distance between similar TCM syndrome types in the vector space and helping the model improve its discriminative ability. In summary, the present invention not only takes into account the problem of distinguishing between similar TCM syndromes, but also takes into account the difficult-to-understand nature of clinical texts.
[0007] The technical task of the present invention is to implement a TCM syndrome differentiation method based on difficult text mining in the following manner, the details of which are as follows:
[0008] S1. Constructing a training dataset for the TCM syndrome differentiation model: First, the clinical text is processed, and the clinical text and its corresponding TCM syndrome type together constitute a training sample; all training samples are collected to form a training dataset for the TCM syndrome differentiation model; all TCM syndrome types appearing in the TCM syndrome differentiation model training set are summarized to form a TCM syndrome type set;
[0009] S2. Construct a TCM syndrome differentiation model, which mainly includes syndrome type coding module, text coding module, syndrome type projection module, understanding transfer module, and prediction module;
[0010] S3. Training the TCM syndrome differentiation model. The main operations include calculating Drr-LOSS, calculating mask prediction loss, calculating syndrome type prediction loss, constructing the total loss function, and constructing the optimization function.
[0011] Preferably, the construction of a TCM syndrome differentiation model training data set is as follows:
[0012] S101. Processing clinical text: First, screen the patient's original clinical text, remove sensitive information and invalid information, and construct the remaining text into a clinical text;
[0013] S102, assigning TCM syndrome types: assigning corresponding real TCM syndrome types to the clinical texts constructed in S101, and deleting data corresponding to some TCM syndrome types with higher frequencies to alleviate the imbalance problem of the dataset;
[0014] S103, constructing a training data set for a TCM syndrome differentiation model: merging the clinical text in S101 with the TCM syndrome type assigned in S102 to form a training sample for a TCM syndrome differentiation model. All training samples are aggregated to construct a training data set for a TCM syndrome differentiation model.
[0015] S104, constructing a TCM syndrome type set: counting all TCM syndrome types appearing in S102 as a TCM syndrome type set.
[0016] Preferably, the construction of the TCM syndrome differentiation model is as follows:
[0017] Construct a TCM syndrome differentiation model, the main operations of which include syndrome type encoding module, text encoding module, syndrome type projection module, understanding transfer module, and prediction module.
[0018] S201. Construct a text encoding module: The text encoding module is mainly composed of BERT. It receives the clinical text in S101 as input and encodes it using BERT to obtain a clinical text representation. The details are as follows:
[0019] The clinical text D consists of N words, that is, the text length of the clinical text D is N, each word is represented by w, and it is represented as D = w1,…,w i ,…,w N ; Then use BERT to encode it, set the maximum clinical text length, that is, the clinical text length threshold is ζ, where ζ is a manually set hyperparameter; for data with a clinical text length less than ζ, it is padded; for data with a clinical text length greater than ζ, it is truncated; the processed clinical text is sent to BERT to obtain the clinical text representation in represents the set of real numbers, that is, all elements in the matrix H are real numbers, and its dimension is ζ×d e , d e Represents the hidden layer size of BERT;
[0020] S202, constructing a syndrome encoding module: The syndrome encoding module is mainly composed of BERT, which receives the TCM syndrome set obtained in S104; then, a syndrome definition database is constructed, which is constructed with reference to knowledge sources such as "TCM Syndrome Classification and Code" and "Baidu Encyclopedia". Specifically, the database contains relevant medical information such as typical clinical manifestations, etiology, core pathogenesis, syndrome differentiation ideas, related diseases, etc. of TCM syndromes; then, the TCM syndrome definition corresponding to each syndrome in the TCM syndrome definition database is sequentially queried, and all TCM syndrome definitions are aggregated to obtain all TCM syndrome definitions; finally, all TCM syndrome definitions are sent to BERT for encoding to form a complete syndrome representation;
[0021] S20201. Construct a syndrome definition database: Use crawler tools to crawl various types of information on each type of TCM syndrome from websites such as Baidu Encyclopedia and Wikipedia, such as clinical manifestations and core pathogenesis. To ensure the accuracy of this knowledge, refer to the "Classification and Codes of TCM Syndrome Types" for verification, and add the definition of each TCM syndrome type in the book. All of the above information is summarized as additional knowledge of TCM syndrome types; the TCM syndrome type set obtained in S104 is matched with each of the above additional knowledge of TCM syndrome types to finally obtain a syndrome definition database;
[0022] S20202. Obtain all syndrome type representations: Receive the TCM syndrome type set obtained in S104. For each TCM syndrome type in the TCM syndrome type set, search for the corresponding additional knowledge in the syndrome type definition database constructed in S20201. Concatenate the additional knowledge as all TCM syndrome type definitions and pass it into the same BERT as S201 to obtain all syndrome type representations, as follows:
[0023] Assume that the TCM syndrome set Q obtained in S104 = {q 1 ,q 2 ,...,q m}, where each TCM syndrome type is q j ,j∈1,2,...,m, each TCM syndrome can be expressed as q j =w 1q ,…,w Nq , where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104, and Nq represents the number of attributes of the additional knowledge corresponding to the TCM syndrome type. All the additional knowledge is spliced, and the same text length threshold ζ as in S201 is used to fill the data whose text length is less than ζ, and to truncate the data whose text length is greater than ζ. The same BERT as in S201 is used to encode each TCM syndrome type and obtain the syndrome type representation K of the TCM syndrome type. j , j=1,2,..,m, the specific calculation formula is as follows:
[0024] Kj =BERT(q j )
[0025] The syndrome type representation of each TCM syndrome type is spliced together to obtain the total syndrome type representation. The specific calculation formula is as follows:
[0026]
[0027] in, d e Represents the hidden layer size of BERT.
[0028] Preferably, the construction of TCM syndrome differentiation model lies in constructing syndrome type projection module, specifically:
[0029] S203, build the syndrome projection module: the syndrome projection module mainly consists of four sub-modules: projection, masking, shared attention, and mask prediction; the projection sub-module makes full use of all syndrome representations obtained in S20202 The goal is to project all syndrome type representations into a new vector space by introducing trainable weights to obtain mapped syndrome type representations. The masked submodule then adopts a masking strategy to randomly mask the mapped syndrome type representations to obtain masked syndrome type representations. The shared attention submodule then establishes a relationship between the masked syndrome type representations and the clinical text representations, allowing the TCM syndrome differentiation model to capture the deep semantic connection between the two and focus on the key related parts of the two to obtain a shared syndrome type representation. Finally, the masked prediction submodule classifies the shared syndrome type representations through a feedforward neural network to obtain masked predictions.
[0030] S20301. Construct a projection submodule: Multiply all the syndrome type representations obtained in S20202 by the trainable weight vector, map all the syndrome type representations into another vector space, and obtain the mapped syndrome type representations, as follows:
[0031]
[0032] in, is the learnable parameter matrix, d e Represents the hidden layer size of BERT; Represents the total number of certificate types obtained by S20202, where ζ represents the truncation or padding value of the text length, which is defined in S201; the final mapping certificate type representation is obtained
[0033] S20302, constructing a masking submodule: Using a masking strategy for the mapped syndrome type representation obtained in S20301, a one-hot vector is constructed in which only one position is 0 and the rest are 1, so that the mapped syndrome type representation is masked at a random position, i.e., a random TCM syndrome type is masked, thereby obtaining a masked syndrome type representation, as follows:
[0034] First, construct a unique hot vector with a vector length of m Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; randomly select a number from 1 to m and mask the position in the one-hot vector to 0, that is, Then, the masked one-hot vector is multiplied by the mapped pattern representation to obtain the masked pattern representation. The specific calculation formula is as follows:
[0035] R m =R⊙O Mask
[0036] Among them, ⊙ represents the dot product, that is, Each element in O Mask Multiply them to get the masking type representation; d e represents the hidden layer size of BERT; ζ represents the truncation or padding value of the text length, which is defined in S201;
[0037] Constructing a shared attention submodule: The shared attention submodule uses the masked syndrome type representation obtained in S20302 as the query vector and the clinical text representation obtained in S201 as the key-value vector, and performs an attention operation on the two to obtain a syndrome type shared representation. The specific calculation formula is as follows:
[0038]
[0039] Among them, R m is the masked syndrome type representation from S20302, H is the clinical text representation from S201, H share Represents the obtained syndrome type sharing representation, m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; m (H) T Apply Softmax, which converts the similarity score between the two into a probability distribution and assigns a weighting coefficient to each row in H to ensure that the model can weight the contribution of different inputs according to the similarity. The specific calculation formula of the Softmax function is as follows:
[0040]
[0041] In the above process, X is R in S20303 m (H) T, x i Represents the similarity score of the masked syndrome relative to the word at position i in the clinical text;
[0042] S20304. Construct a mask prediction submodule: The mask prediction submodule receives the shared representation of the certificate type obtained in S20303, uses a feedforward neural network to make predictions, and outputs the prediction results, that is, outputs the one-hot vector of the predicted masked position. The specific calculation formula for predicting the masked position is:
[0043] Y mask =MAX(σ(W mask H share +bias))
[0044] Among them, σ represents the activation function, which here represents the Softmax function, that is, the shared expression of the syndrome obtained by S20303 is H share Mapped to all TCM syndrome types obtained in S104; is a one-hot vector of length m, representing the one-hot vector of the predicted masked position, and m represents the number of types of TCM syndromes in the TCM syndrome set constructed in S104; W mask Is a learnable parameter matrix, bias is a learnable parameter, MAX represents the maximum pooling, specifically σ(W mask H share + bias) is a vector of length m. The function of max pooling is to set the maximum value of the m values to 1 and the remaining m-1 positions to 0.
[0045] More optimally, the construction of a TCM syndrome differentiation model involves building an understanding transfer module and a prediction module, specifically:
[0046] S204, constructing an understanding transfer module: The understanding transfer module receives the clinical text representation obtained in S201; the understanding transfer module includes two submodules, one is a shared attention submodule, and the other is a difficult text mining submodule; the shared attention submodule adopts the same shared attention as in S20303, with a custom trainable parameter vector as the query vector of the shared attention, and the clinical text representation as the key-value vector, and obtains the shared attention representation through shared attention; the difficult text mining submodule receives the shared attention representation, and divides the clinical text into two parts, a high-understanding representation and a low-understanding representation, through a difficult character selector, and uses the teacher-student model to perform knowledge transfer learning, and finally obtains the output of the student model as the low-understanding representation, which is sent to the prediction module;
[0047] S20401. Construct a shared attention submodule: The shared attention submodule shares the same parameters as the shared attention in S20303. That is, the two are exactly the same. First, a custom trainable parameter vector is designed as the query vector of the shared attention, so as to transfer the syndrome type representation learned by the shared attention in S20303. Then, the clinical text representation obtained in S201 is received as a key-value vector and the shared attention representation is calculated. The specific calculation formula is as follows:
[0048]
[0049] in, is a custom trainable parameter vector, H is the clinical text representation from S201, Represents the obtained shared attention representation, m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; Softmax function is an activation function, which is introduced in detail in S20303 and will not be repeated here; in the above process, Represents W s Attention score relative to clinical text representation H;
[0050] S20402: Constructing a difficult text mining submodule: First, receive the shared attention representation obtained in S20401 and feed it into the difficult character selector to obtain a high-comprehension representation and a low-comprehension representation. Then, the high-comprehension representation is used as the teacher model, and the low-comprehension representation is used as the student model. Finally, the teacher model's ability to understand text characters is transferred to the student model through KL-LOSS.
[0051] S2040201. Build a difficult character miner: Receive the shared attention representation obtained in S20401 and sort each row of vectors to obtain the ascending order index of the shared attention representation. Then, two strategies are designed to divide the shared attention representation into high-understanding representation and low-understanding representation, as follows:
[0052] Convert the shared attention representation obtained in S20401, and use the Sort function to sort each row vector in the shared attention in ascending order, and then obtain the corresponding descending order index. The formula is as follows:
[0053] H tran_share =[a1,a2,......,a m ]
[0054]
[0055] Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; a k Represents the shared attention representation H obtained in S20401tran_share The k-th row tensor in , e1 represents the index of the highest score position in the k-th row shared attention representation, and Represented as the index of the lowest score position in the k-th row of shared attention representation; the Sort function is a pytorch intrinsic sorting operation that can sort the elements of a tensor of any length in ascending or descending order according to the specified dimension and return the sorted index. In the above formula, the first dimension of the shared attention representation is sorted in descending order;
[0056] In order to divide the shared attention representation into high-understanding representation and low-understanding representation, we first set a length of I k The same all-zero vector ME k and all-one vector MH k , then I k Middle front beta f % position sequence and randomly selected β c % The position sequence is set to 1, and then a new mask matrix MEO is obtained k , and then MEO k Sequentially with a k Dot product, finally get the high-level understanding representation A easy , the formula is as follows:
[0057] A easy =[a1⊙MEO1,a2⊙MEO2,...,a m ⊙MEO m ]
[0058] Similarly, in order to avoid the influence of too low attention on the final result, the middle I k Post-β l % position sequence, front β f % position sequence and randomly selected β c % The position sequence is set to 0, and then a new mask matrix MHO is obtained k , then MHO k Sequentially with a k Dot product, finally get the low-level understanding representation A hard , the formula is as follows:
[0059] A hard =[a1⊙MHO1,a2⊙MHO2,......,a m ⊙MHO m ]
[0060] S2040202. Construct KL-LOSS: Receive the high-understanding representation and low-understanding representation obtained in S2040201, and use Kullback-Leibler divergence, or KL-LOSS, to transfer the model's understanding ability. The specific calculation formula is as follows:
[0061] KL=KL-LOSS(A easy ||A hard )
[0062] Among them, KL-LOSS is a loss function used to measure the difference between two tensor distributions. KL-LOSS can be used to calculate the relative entropy between the two distributions. At the same time, the smaller the KL-LOSS value, the more similar the two distributions are. The specific calculation formula is as follows:
[0063]
[0064] In the above process, P stands for high understanding and A easy ,Q represents low understanding representation;
[0065] S205, constructing a prediction module: receiving the low-level representation obtained in S2040201, using a feedforward neural network to predict the TCM syndrome type and output the result, that is, outputting a multi-hot vector of the TCM syndrome type, and finally outputting the TCM syndrome type through the correspondence between the position of 1 in the vector and the TCM syndrome type. The specific calculation formula is as follows:
[0066] Y label =σ(W label A hard +bias label )
[0067] in, is a one-hot vector of length m, where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; W label Is a learnable parameter matrix, bias label is the bias term; σ represents sigmoid, which converts the activation of the vector into an independent probability, and binarizes the probability into 0 / 1 through the set threshold to obtain a vector of length m; the position of 1 in the vector is mapped to the corresponding TCM syndrome type, and the TCM syndrome type Y is output. label .
[0068] More preferably, after the TCM syndrome differentiation model is constructed, the TCM syndrome differentiation model training data set is trained to optimize the training of the TCM syndrome differentiation model, specifically:
[0069] S3. Training the TCM syndrome differentiation model: Optimize the model by combining KL-LOSS, Drr-LOSS, mask prediction loss, and syndrome type prediction loss according to the weights; first receive the KL-LOSS obtained in S2040202; then calculate Drr-LOSS; then receive the true masked position obtained in S20302 and the predicted masked position obtained in S20304, and calculate the binary cross entropy loss between the two; then receive the TCM syndrome type obtained in S205 and the TCM syndrome type corresponding to the real clinical text to calculate the cross entropy loss; weighted sum of the four losses to obtain the total loss function;
[0070] When the model of this method has not been fully trained, it needs to be trained in the training data set to optimize the model parameters; when the model training is completed, the model can predict the corresponding syndrome type for the input clinical text;
[0071] S401, calculate Drr-LOSS: receive the mapped syndrome type representation obtained in S20301, and use cosine similarity distance to measure the distance between different rows in the mapped syndrome type representation, and maximize the sum of the distances between different rows in the mapped syndrome type representation to enhance the distinction between different TCM syndrome types. The formula is as follows:
[0072]
[0073] Cosine Similarity Distance represents the cosine similarity distance, and the specific calculation formula is as follows: In the above process, X represents R i , Y stands for R j ;
[0074] S402. Calculate the mask prediction loss: Construct the actual mask position obtained in S20302 into the corresponding one-hot vector Y′, then receive the one-hot vector of the predicted mask position obtained in S20304, and calculate the binary cross entropy loss between the two, i.e., the mask prediction loss, as follows:
[0075]
[0076] Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104, Y mask represents the one-hot vector of the predicted occlusion position obtained in S20304, and Y′ represents the one-hot vector of the actual occlusion position obtained in S20302;
[0077] S403, calculating the syndrome prediction loss: receiving the multi-hot vector of the TCM syndrome obtained in S205 and the multi-hot vector of the TCM syndrome corresponding to the real clinical text, and calculating the cross entropy loss. The formula is as follows:
[0078]
[0079] Among them, Y represents the unique heat vector of the TCM syndrome type corresponding to the real clinical text, Y label represents the unique-hot vector corresponding to the predicted TCM syndrome type, and m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104;
[0080] S404. Construct the total loss function: It contains four parts, namely KL-LOSS obtained in S2040202, Drr-LOSS obtained in S401, mask prediction loss obtained in S402, and syndrome prediction loss obtained in S403; and add the four together according to the weights of λ1:λ2:λ3:λ4. The formula is as follows:
[0081]
[0082] λ1:λ2:λ3:λ4 are manually set hyperparameters;
[0083] S405. Construct an optimization function: Use the Adamw algorithm as the optimization function of the model; the learning rate parameter is set to 0.001, and other hyperparameters use the default values in PyTorch; hyperparameters refer to parameters that need to be manually set before starting the training process; these parameters cannot be automatically optimized through training; depending on the actual data set, these parameters need to be manually set by the user.
[0084] A TCM syndrome differentiation device based on difficult text mining, comprising: a TCM syndrome differentiation training data set construction unit, a TCM syndrome differentiation model construction unit, and a TCM syndrome differentiation model training unit;
[0085] The TCM syndrome differentiation model training data set construction unit is used to process the input clinical text. Specifically, each clinical text is assigned its corresponding TCM syndrome type, and finally constructed into a TCM syndrome differentiation model training data set; all syndrome types appearing in the TCM syndrome differentiation model training data set are summarized as all TCM syndrome types.
[0086] The TCM syndrome differentiation model construction unit is responsible for constructing the TCM syndrome differentiation model, and using the teacher-student model to improve the model's ability to understand clinical texts and assign appropriate TCM syndrome types to clinical texts.
[0087] The TCM syndrome differentiation model training unit is used to construct the loss function and optimizer required in the model training process, and finally complete the training and optimization of the model; when the model of this method has not been fully trained, it needs to be trained on the training data set to optimize the model parameters; when the model training is completed, the model can predict the corresponding TCM syndrome type for the input clinical text.
[0088] A storage medium stores a plurality of instructions, which are loaded by a processor to execute the steps of the above-mentioned TCM syndrome differentiation method based on difficult text mining.
[0089] An electronic device comprises: the above-mentioned storage medium; and a processor for executing instructions in the storage medium.
[0090] The TCM syndrome differentiation method and device based on difficult text mining of the present invention have the following advantages:
[0091] (1) This invention proposes a new TCM syndrome differentiation task that is more in line with medical practice and collects a new TCM data set accordingly;
[0092] (2) The present invention can quickly process clinical texts, providing new ideas for subsequent research;
[0093] (3) The present invention is helpful in helping the model to distinguish the representations of similar TCM syndromes, thereby improving the model's ability to distinguish different TCM syndromes;
[0094] (4) The present invention can effectively distinguish the subtle differences in diseases corresponding to TCM syndrome types, thereby making the TCM syndrome differentiation model more accurate;
[0095] (5) The present invention can help the model understand the more difficult-to-understand parts of the clinical text, thereby making it possible to integrate the TCM syndrome differentiation model system into real medical equipment;
[0096] (6) The present invention uses a difficult text mining algorithm to focus on the parts of clinical text that are easier to understand and use the model's ability to understand these parts to help the model explore the parts of clinical text that are more difficult to understand, thereby improving the model's full understanding of clinical texts with classical Chinese style;
[0097] (7) The present invention constructs a TCM syndrome projection method, which projects the TCM syndrome representation into a new vector space and uses the corresponding loss function to improve the specificity of different TCM syndrome representations in the vector space, thereby improving the model's ability to discriminate different TCM syndromes;
[0098] (8) The present invention uses fewer model parameters to enable the TCM syndrome differentiation model to achieve higher syndrome differentiation performance, providing a realistic possibility for deploying the TCM syndrome differentiation system into medical devices;
[0099] (9) The present invention uses natural language processing technology to assign appropriate TCM syndrome types based on the patient's clinical text, which greatly reduces the workload of doctors and also provides the possibility for subsequent patients to self-diagnose at home. It is a major breakthrough in computer artificial intelligence methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] The present invention will be further described below with reference to the accompanying drawings.
[0101] Figure 1 A schematic diagram of the structure of the TCM syndrome differentiation method based on difficult text mining
[0102] Figure 2 Flowchart for building difficult text mining submodule
[0103] Explanation of terms
[0104] Syndrome differentiation in Traditional Chinese Medicine (TCM): Syndrome differentiation in TCM is one of the core concepts of TCM treatment. The basic concept of TCM syndrome differentiation is to comprehensively analyze the patient's symptoms, pulse, tongue coating, and other information to determine the cause, pathogenesis, and course of the disease, thereby clarifying the diagnosis and treatment of the disease.
[0105] TCM Syndrome Types: In Traditional Chinese Medicine (TCM), "Syndrome Types" refers to the classification of disease patterns and is a core component of TCM differentiation. Each disease corresponds to specific syndromes, reflecting different aspects of the pathogenesis, such as cold syndrome, heat syndrome, deficiency syndrome, and excess syndrome. By analyzing different combinations of these syndromes, TCM practitioners can categorize diseases into different syndrome types, enabling more precise treatment plans.
[0106] Main symptom: The main symptom refers to the most prominent and significant symptom in the patient's condition, which is usually the main reason for the patient to seek medical treatment. During the syndrome differentiation process, the doctor will focus on analyzing the main symptom to clarify the main pathological changes of the disease and formulate the core treatment strategy accordingly.
[0107] Concomitant symptoms: Concomitant symptoms refer to other secondary symptoms or syndromes that coexist with the primary symptom. These symptoms may be a natural extension of the primary symptom, while others may reflect other pathological changes related to the primary symptom. By analyzing concomitant symptoms, doctors can further improve their overall understanding of the disease and optimize treatment plans.
[0108] Clinical text: Clinical text can include various textual descriptions related to a patient's health status and is typically used to record the patient's diagnosis, treatment, and monitoring. In the medical database used in this method, clinical text includes descriptions of occupation, admission date, onset solar term, chief complaint, patient medical history, information on the four diagnostic tests, physical examination, specialist examinations, auxiliary examinations, and admission symptoms. DETAILED DESCRIPTION
[0109] The method and device for TCM syndrome differentiation based on difficult text mining of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments.
[0110] Example 1:
[0111] The overall framework of the present invention is as follows Figure 1 As shown. Figure 1 It can be seen that the main framework of the present invention includes the following five modules, namely, syndrome type encoding module, syndrome type projection module, text encoding module, understanding transfer module, and prediction module. Among them, each clinical text and the TCM syndrome type corresponding to the clinical text constitute a training sample; all training samples are collected to form a TCM syndrome differentiation training data set; the clinical text in the TCM syndrome differentiation training data set is used as the input of the text encoding module; at the same time, all TCM syndrome types appearing in the TCM syndrome differentiation training data set, that is, the TCM syndrome type set, are counted as the input of the syndrome type encoding module. The text encoding module takes the clinical text as input, encodes it using BERT, obtains the clinical text representation, and then sends the clinical text representation to the understanding transfer module. The syndrome type encoding module maps all TCM syndrome types to all TCM syndrome type definitions through the syndrome type definition database, and sends all TCM syndrome type definitions to the same BERT as the text encoding module for encoding, thereby obtaining all syndrome type representations, and finally sends all syndrome type representations to the syndrome type projection module. The syndrome projection module consists of four submodules: projection, masking, shared attention, and mask prediction. First, the projection submodule receives all syndrome representations and projects them into another vector space through trainable weights to obtain the mapped syndrome representation. Then, the masking submodule adopts a masking strategy to construct a one-hot vector with a length equal to the hidden dimension of all syndrome representations, and randomly sets one position to 0, and multiplies the one-hot vector with the mapped syndrome representation to obtain the masked syndrome representation. Then, the shared attention submodule uses the masked syndrome representation as the query vector and the clinical text representation as the key-value vector, and sends them into the shared attention to obtain the syndrome shared representation. Finally, the mask prediction submodule performs classification through a feedforward neural network to obtain the mask prediction. The understanding transfer module consists of two submodules: shared attention and difficult text mining. First, the shared attention submodule adopts the same shared attention as that in the syndrome projection module, uses a custom trainable parameter vector as the query vector, and the clinical text representation as the key vector to obtain the shared attention representation, which is then sent to the difficult text mining submodule. Second, the difficult text mining submodule is as follows: Figure 2 As shown in the figure, a difficult character selector is used to divide the clinical text representation into two parts: a high-understanding representation and a low-understanding representation. The high-understanding representation is used as the teacher model, and the low-understanding representation is used as the student model. The understanding ability of the model is transferred from the teacher model to the student model through the KL-LOSS constraint. Finally, the low-understanding representation output by the student model is fed into the prediction module. The prediction module receives the output of the student model and assigns the appropriate TCM syndrome type to the clinical text through a feedforward neural network.
[0112] As attached Figure 1 As shown, the TCM syndrome differentiation method based on difficult text mining of the present invention comprises the following steps:
[0113] S1. Constructing a training data set for the TCM syndrome differentiation model: First, process the obtained clinical texts, associate each clinical text with its corresponding TCM syndrome type to form a training data set, and finally aggregate all the training data to obtain a TCM syndrome differentiation model training data set; summarize all TCM syndrome types appearing in the TCM syndrome differentiation model training data set as a TCM syndrome type set; the specific process is as follows
[0114] S101. Processing clinical text: First, screen the patient's original clinical text, remove sensitive information and invalid information, and construct the remaining text into a clinical text;
[0115] For example, we process clinical text in the MLTCM dataset. The following example shows the process:
[0116] The clinical text of the original dataset records the basic information and clinical records of the patients, as shown in Table 1. The basic information includes attributes such as the patient's hospitalization number and gender. The patient's clinical record consists of multiple attributes such as the patient's medical history, four diagnostic information, physical examination, and specialist examination. Some sensitive information, such as the patient's name and the hospital visited, appears in the medical history. When anonymizing this sensitive information, different regular expressions are designed to delete it according to the unique medical record writing style of different physicians. There are some garbled texts in the anonymized patient clinical records. In order to better preserve the original semantics, the meaningless symbols are identified and the sentences containing the identification symbols are deleted. Finally, the patient's clinical records after anonymization and deletion of invalid information are retained, and the patient's clinical record texts are spliced as clinical texts.
[0117] Table 1
[0118]
[0119] S102, assigning TCM syndrome types: assigning the corresponding real TCM syndrome types to the clinical texts constructed in S101, and deleting the data corresponding to some TCM syndrome types with higher frequencies, thereby alleviating the imbalance problem of the data set;
[0120] For example, in the MLTCM dataset, each clinical case is assigned its corresponding TCM syndrome type. The example is as follows:
[0121] As shown in Table 1 in S101, the TCM syndrome types in the table are used as TCM syndrome types of the TCM syndrome differentiation model training data samples; the frequency of each TCM syndrome type in the TCM syndrome differentiation model training data set is counted, and 1367 and 233 data of the two TCM syndrome types with the highest frequency are deleted respectively to ensure the balance of the TCM syndrome type labels of the data set;
[0122] S103, constructing a training data set for a TCM syndrome differentiation model: merging the clinical text in S101 with the TCM syndrome type assigned in S102 to form a training sample for a TCM syndrome differentiation model. All training samples are aggregated to construct a training data set for a TCM syndrome differentiation model.
[0123] For example, in the MLTCM dataset, each TCM syndrome differentiation training sample is shown in Table 2:
[0124] Table 2
[0125]
[0126] S104, constructing a TCM syndrome type set: counting all TCM syndrome types appearing in S102 as a TCM syndrome type set Q;
[0127] For example, in the MLTCM dataset, the TCM syndrome types corresponding to each TCM syndrome differentiation model training sample are stored in a list. Then, using the unique key value feature of the Python dictionary, all TCM syndrome types that appear in the MLTCM dataset are screened out and named as the TCM syndrome type set, as shown in Table 3 below:
[0128] Table 3
[0129]
[0130] S2. Construct a TCM syndrome differentiation model: The TCM syndrome differentiation model consists of a text encoding module, a syndrome type encoding module, a syndrome type projection module, an understanding transfer module, and a prediction module; first, the text encoding module encodes the clinical text to obtain the clinical text representation; then, a syndrome type encoding module is constructed to map the TCM syndrome type set to all TCM syndrome type definitions through the syndrome type definition database, and encode all TCM syndrome type definitions to obtain all syndrome type representations; then, a syndrome type projection module is constructed to project all syndrome type representations into TCM syndrome types through four parts: projection, masking, shared attention, and mask prediction; then, an understanding transfer module is constructed, and the parameters in the shared attention are used as query vectors to perform attention with the clinical text to obtain shared attention representations, and the difficult text mining module is used to divide the text into two parts: high understanding and low understanding, and the teacher-student model is used to guide the learning of the student model with the understanding ability of the teacher model, and finally the output of the student model is used as the low understanding representation; finally, a prediction module is constructed to obtain the final predicted TCM syndrome type through a feedforward neural network;
[0131] S201. Construct a text encoding module: The text encoding module is mainly composed of BERT. The text encoding module receives the clinical text in S101 as input and encodes it using BERT to obtain a clinical text representation.
[0132] For example, in the MLTCM dataset:
[0133] The clinical text D consists of N words, that is, the text length of the clinical text D is N, each word is represented by w, and it is represented as D = w1,…,w i ,…,w N ; Then use BERT to encode it, set the maximum clinical text length, that is, the clinical text length threshold is ζ, where ζ is a manually set hyperparameter; for data with a clinical text length less than ζ, it is padded; for data with a clinical text length greater than ζ, it is truncated; the processed clinical text is sent to BERT to obtain the clinical text representation in represents the set of real numbers, that is, all elements in the matrix H are real numbers, and its dimension is ζ×d e , d e Represents the hidden layer size of BERT. The process is expressed as follows:
[0134]
[0135] S202, constructing a syndrome encoding module: The syndrome encoding module is mainly composed of BERT, which receives the TCM syndrome set obtained in S104 and then constructs a syndrome definition database. The database is constructed by referring to knowledge sources such as "TCM Syndrome Classification and Code" and "Baidu Encyclopedia". Specifically, the database contains relevant medical information such as typical clinical manifestations, etiology, core pathogenesis, syndrome differentiation ideas, related diseases, etc. of TCM syndromes; then, by sequentially querying the TCM syndrome definition corresponding to each syndrome in the syndrome definition database, all TCM syndrome definitions are aggregated to obtain all TCM syndrome definitions; finally, all TCM syndrome definitions are sent to BERT for encoding to form a complete syndrome representation;
[0136] S20201. Construct a syndrome definition database: Use crawler tools to crawl various types of information on each type of TCM syndrome from websites such as Baidu Encyclopedia and Wikipedia, such as clinical manifestations and core pathogenesis. To ensure the accuracy of this knowledge, refer to the "Classification and Codes of TCM Syndrome Types" for verification, and add the definition of each TCM syndrome type in the book. All of the above information is summarized as additional knowledge of TCM syndrome types; the TCM syndrome type set obtained in S104 is matched with each of the above additional knowledge of TCM syndrome types to finally obtain a syndrome definition database;
[0137] For example, in the syndrome definition database, the information for "heart qi and blood deficiency syndrome" is as follows:
[0138] Table 4
[0139]
[0140]
[0141] S20202. Obtain all syndrome type representations: Receive the TCM syndrome type set obtained in S104. For each TCM syndrome type in the TCM syndrome type set, search for corresponding additional knowledge in the syndrome type definition database constructed in S20201. Concatenate the additional knowledge as all TCM syndrome type definitions and pass it into the same BERT as S201 to obtain all syndrome type representations.
[0142] For example, in the MLTCM dataset:
[0143] Assume that the TCM syndrome set Q obtained in S104 = {q 1 ,q 2 ,...,q m}, where each TCM syndrome type is q j ,j∈1,2,...,m, each TCM syndrome can be expressed as q j =w 1q ,…,w Nq , where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104, and Nq represents the number of attributes of the additional knowledge corresponding to the TCM syndrome type, as shown in Table 4 in S20201. Here, the number of attributes of the additional knowledge is 5. The 5 types of additional knowledge are spliced together, and the same text length threshold ζ as in S201 is used to fill the data with a text length less than ζ, and to truncate the data with a text length greater than ζ. The same BERT as in S201 is used to encode each TCM syndrome type to obtain the syndrome type representation K of the TCM syndrome type. j , j=1,2,..,m, the specific calculation formula is as follows:
[0144] K j =BERT(q j ) (4)
[0145] The syndrome type representation of each TCM syndrome type is spliced together to obtain the total syndrome type representation. The specific calculation formula is as follows:
[0146]
[0147] in, d e Represents the hidden layer size of BERT;
[0148] S203, construct syndrome projection module: The syndrome projection module mainly includes four submodules: projection, masking, shared attention, and mask prediction; the projection submodule makes full use of all syndrome representations Q obtained in S20202, aiming to project all syndrome representations into a new vector space by introducing trainable weights to obtain mapped syndrome representations; then the masking submodule adopts a masking strategy to randomly mask the mapped syndrome representation to obtain a masked syndrome representation; then the shared attention submodule builds a relationship between the masked syndrome representation and the clinical text representation, so that the TCM syndrome differentiation model captures the deep semantic association between the two, and then focuses on the key related parts of the two to obtain a syndrome shared representation; finally, the mask prediction submodule classifies the syndrome shared representation through a feedforward neural network to obtain a mask prediction;
[0149] S20301. Construct a projection submodule: multiply all syndrome type representations obtained in S20202 by a trainable weight vector, map all syndrome type representations into another vector space, and obtain mapped syndrome type representations;
[0150] For example, in the MLTCM dataset:
[0151]
[0152] in, is the learnable parameter matrix, d e Represents the hidden layer size of BERT; Represents the total number of certificate types obtained by S20202, where ζ represents the truncation or padding value of the text length, which is defined in S201; the final mapping certificate type representation is obtained
[0153] S20302, constructing a masking submodule: using a masking strategy on the mapped syndrome type representation obtained in S20301, by constructing a one-hot vector in which only one position is 0 and the rest are 1, so that the mapped syndrome type representation is masked at a random position, i.e., masking a random TCM syndrome type, thereby obtaining a masked syndrome type representation;
[0154] For example, in the MLTCM dataset:
[0155] First, construct a unique hot vector with a vector length of m Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; randomly select a number from 1 to m and mask the position in the one-hot vector to 0, that is, Then, the masked one-hot vector is multiplied by the mapped pattern representation to obtain the masked pattern representation. The specific calculation formula is as follows:
[0156] R m=R⊙O Mask (7)
[0157] Among them, ⊙ represents the dot product, that is, Each element in O Mask Multiply them to get the masking type representation; d e represents the hidden layer size of BERT; ζ represents the truncation or padding value of the text length, which is defined in S201;
[0158] S20303. Construct a shared attention submodule: The shared attention submodule uses the masked syndrome type representation obtained in S20302 as a query vector and the clinical text representation obtained in S201 as a key-value vector, performs an attention operation on the two, and thus obtains a syndrome type shared representation;
[0159] For example, in the MLTCM dataset, the specific calculation formula is as follows:
[0160]
[0161] Among them, R m is the masked syndrome type representation from S20302, H is the clinical text representation from S201, H share Represents the obtained syndrome type sharing representation, m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; m (H) T Apply Softmax, which converts the similarity score between the two into a probability distribution and assigns a weighting coefficient to each row in H to ensure that the model can weight the contribution of different inputs according to the similarity. The specific calculation formula of the Softmax function is as follows:
[0162]
[0163] In the above process, X is R in formula (5) m (H) T , x i Represents the similarity score of the masked syndrome relative to the word at position i in the clinical text;
[0164] S20304. Construct a mask prediction submodule: The mask prediction submodule receives the shared representation of the certificate type obtained in S20303, uses a feedforward neural network to make predictions, and outputs the prediction results, that is, outputs a one-hot vector of the predicted masked position;
[0165] For example, in the MLTCM dataset, the calculation formula for predicting the occlusion position is:
[0166] Y mask =MAX(σ(W mask H share+bias)) (10)
[0167] Among them, σ represents the activation function, which here represents the Softmax function, that is, the shared expression of the syndrome obtained by S20303 is H share Mapped to all TCM syndrome types obtained in S104; is a one-hot vector of length m, representing the one-hot vector of the predicted masked position, and m represents the number of types of TCM syndrome types in the TCM syndrome type set constructed in S104; W mask Is a learnable parameter matrix, bias is a learnable parameter, MAX represents the maximum pooling, specifically σ(W mask H share + bias) is a vector of length m. The function of max pooling is to set the maximum value of the m values to 1 and the remaining m-1 positions to 0.
[0168] S204, constructing an understanding transfer module: The understanding transfer module receives the clinical text representation obtained in S201; the understanding transfer module includes two submodules, one is a shared attention submodule, and the other is a difficult text mining submodule; the shared attention submodule adopts the same shared attention as in S20303, with a custom trainable parameter vector as the query vector of the shared attention, and the clinical text representation as the key-value vector, and obtains the shared attention representation through shared attention; the difficult text mining submodule receives the shared attention representation, and divides the clinical text into two parts, a high-understanding representation and a low-understanding representation, through a difficult character selector, and uses the teacher-student model to perform knowledge transfer learning, and finally obtains the output of the student model as the low-understanding representation, which is sent to the prediction module;
[0169] S20401. Constructing a shared attention submodule: The shared attention submodule shares the same parameters as the shared attention in S20303. That is, the two are identical. First, a custom trainable parameter vector is designed as the query vector for shared attention, thereby migrating the syndrome type representation learned by shared attention in S20303. Then, the clinical text representation obtained in S201 is received as a key-value vector and the shared attention representation is calculated.
[0170] For example, in the MLTCM dataset, the specific calculation formula is as follows:
[0171]
[0172] in, is a custom trainable parameter vector, H is the clinical text representation from S201, Represents the obtained shared attention representation, m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; Softmax function is an activation function, which is introduced in detail in S20303 and will not be repeated here; in the above process, Represents W s Attention score relative to clinical text representation H;
[0173] S20402: Constructing a difficult text mining submodule: First, receive the shared attention representation obtained in S20401 and feed it into the difficult character selector to obtain a high-comprehension representation and a low-comprehension representation. Then, the high-comprehension representation is used as the teacher model, and the low-comprehension representation is used as the student model. Finally, the teacher model's ability to understand text characters is transferred to the student model through KL-LOSS.
[0174] S2040201. Build a difficult character miner: Receive the shared attention representation obtained in S20401 and sort each row of vectors to obtain an ascending index of the shared attention representation. Then, two strategies are designed to divide the shared attention representation into high-understanding representation and low-understanding representation.
[0175] For example, in the MLTCM dataset:
[0176] Convert the shared attention representation obtained in S20401, and use the Sort function to sort each row vector in the shared attention in ascending order, and then obtain the corresponding descending order index. The formula is as follows:
[0177]
[0178] Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; a k Represents the shared attention representation H obtained in S20401 tran_share The k-th row tensor in , e1 represents the index of the highest score position in the k-th row shared attention representation, and Represented as the index of the lowest score position in the k-th row of shared attention representation; the Sort function is a pytorch intrinsic sorting operation that can sort the elements of a tensor of any length in ascending or descending order according to the specified dimension and return the sorted index. In the above formula, the first dimension of the shared attention representation is sorted in descending order;
[0179] In order to divide the shared attention representation into high-understanding representation and low-understanding representation, we first set a length of I k The same all-zero vector ME k and all-one vector MH k , then I k Middle front betaf % position sequence and randomly selected β c % The position sequence is set to 1, and then a new mask matrix MEO is obtained k , and then MEO k Sequentially with a k Dot product, finally get the high-level understanding representation A easy , the formula is as follows:
[0180] A easy =[a1⊙MEO1,a2⊙MEO2,...,a m ⊙MEO m ] (13)
[0181] Similarly, in order to avoid the influence of too low attention on the final result, the middle I k Post-β l % position sequence, front β f % position sequence and randomly selected β c % The position sequence is set to 0, and then a new mask matrix MHO is obtained k , then MHO k Sequentially with a k Dot product, finally get the low-level understanding representation A hard , the formula is as follows:
[0182] A hard =[a1⊙MHO1,a2⊙MHO2,......,a m ⊙MHO m ] (14)
[0183] S2040202. Build KL-LOSS: Receive the high-understanding representation and low-understanding representation obtained in S2040201, and use Kullback-Leibler divergence, namely KL-LOSS, to transfer the model's understanding ability;
[0184] The specific calculation formula is as follows:
[0185] KL=KL-LOSS(A easy ||A hard ) (15)
[0186] Among them, KL-LOSS is a loss function used to measure the difference between two tensor distributions. KL-LOSS can be used to calculate the relative entropy between the two distributions. At the same time, the smaller the KL-LOSS value, the more similar the two distributions are. The specific calculation formula is as follows:
[0187]
[0188] In the above process, P stands for high understanding and A easy , Q represents low understanding, A hard ;
[0189] S205, constructing a prediction module: receiving the low-level understanding representation obtained in S2040201, using a feedforward neural network to predict the TCM syndrome type and output the result, that is, outputting a multi-hot vector of the TCM syndrome type, and finally outputting the TCM syndrome type through the correspondence between the position of 1 in the vector and the TCM syndrome type;
[0190] Calculate TCM syndrome type, the corresponding calculation formula is:
[0191] Y label =σ(W label A hard +bias label ) (17)
[0192] in, is a one-hot vector of length m, where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; W label Is a learnable parameter matrix, bias label is the bias term; σ represents sigmoid, which converts the activation of the vector into an independent probability, and binarizes the probability into 0 / 1 through the set threshold to obtain a vector of length m; the position of 1 in the vector is mapped to the corresponding TCM syndrome type, and the TCM syndrome type Y is output. label ;
[0193] S3. Training the TCM syndrome differentiation model: Optimize the model by combining KL-LOSS, Drr-LOSS, mask prediction loss, and syndrome type prediction loss according to the weights; first receive the KL-LOSS obtained in S2040202; then calculate Drr-LOSS; then receive the true masked position obtained in S20302 and the predicted masked position obtained in S20304, and calculate the binary cross entropy loss between the two; then receive the TCM syndrome type obtained in S205 and the TCM syndrome type corresponding to the real clinical text to calculate the cross entropy loss; weighted sum of the four losses to obtain the total loss function;
[0194] When the model of this method has not been fully trained, it needs to be trained in the training data set to optimize the model parameters; when the model training is completed, the model can predict the corresponding syndrome type for the input clinical text;
[0195] S401, calculate Drr-LOSS: receive the mapped syndrome type representation obtained in S20301, and use cosine similarity distance to measure the distance between different rows in the mapped syndrome type representation, and maximize the sum of the distances between different rows in the mapped syndrome type representation to enhance the distinction between different TCM syndrome types. The formula is as follows:
[0196]
[0197] Cosine Similarity Distance represents the cosine similarity distance, and the specific calculation formula is as follows: In the above process, X represents R i , Y stands for R j ;
[0198] S402. Calculate the mask prediction loss: Construct the actual mask position obtained in S20302 into the corresponding one-hot vector Y′, then receive the one-hot vector of the predicted mask position obtained in S20304, and calculate the binary cross entropy loss between the two, i.e., the mask prediction loss, as follows:
[0199]
[0200] Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104, Y mask represents the one-hot vector of the predicted occlusion position obtained in S20304, and Y′ represents the one-hot vector of the actual occlusion position obtained in S20302;
[0201] S403, calculating the syndrome prediction loss: receiving the multi-hot vector of the TCM syndrome obtained in S205 and the multi-hot vector of the TCM syndrome corresponding to the real clinical text, and calculating the cross entropy loss. The formula is as follows:
[0202]
[0203] Among them, Y represents the unique heat vector of the TCM syndrome type corresponding to the real clinical text, Y label represents the unique-hot vector corresponding to the predicted TCM syndrome type, and m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104;
[0204] S404. Construct the total loss function: It contains four parts, namely KL-LOSS obtained in S2040202, Drr-LOSS obtained in S401, mask prediction loss obtained in S402, and syndrome prediction loss obtained in S403; and add the four together according to the weights of λ1:λ2:λ3:λ4. The formula is as follows:
[0205]
[0206] λ1:λ2:λ3:λ4 are manually set hyperparameters;
[0207] S405. Construct an optimization function: Use the Adamw algorithm as the optimization function of the model; the learning rate parameter is set to 0.001, and other hyperparameters use the default values in PyTorch; hyperparameters are parameters that need to be manually set before starting the training process; these parameters cannot be automatically optimized through training; depending on the actual data set, these parameters need to be manually set by the user;
[0208] For example, in PyTorch, defining the Adam optimization function can be implemented using the following code:
[0209] optim=torch.optim.Adam(lr=0.001)
[0210] The model proposed in this paper achieves better results than other models on the MLTCM dataset. The comparison of experimental results is shown in the following table:
[0211] Model Macro-AUC Micro-AUC Macro-F1 Micro-F1 Precision-P@2 ZY-BERT 49.67 93.94 2.73 27.60 19.83 Chinese-BERT 51.45 93.68 1.85 27.16 19.83 GPT-3.5-ICL 49.97 49.92 0.38 0.54 0.28 The present invention 68.60 91.67 6.34 34.20 25.66
[0212] Example 3:
[0213] As attached Figure 1 As shown, a TCM syndrome differentiation method and device based on difficult text mining includes: a TCM syndrome differentiation training data set construction unit, a TCM syndrome differentiation model construction unit, and a TCM syndrome differentiation model training unit; the functions of steps S1, S2, and S3 in the TCM syndrome differentiation method based on difficult text mining are respectively performed, and the specific functions of each unit are as follows:
[0214] The TCM syndrome differentiation training data set construction unit and the TCM syndrome differentiation model training data set construction unit are used to process the input clinical text. Specifically, each clinical text is assigned its corresponding TCM syndrome type, and finally a TCM syndrome differentiation model training data set is constructed; the TCM syndrome types appearing in the TCM syndrome differentiation model training data set are summarized as all TCM syndrome types;
[0215] The TCM syndrome differentiation model construction unit consists of five parts: a text encoding module, which encodes clinical texts through BERT to obtain clinical text representations; a syndrome type encoding module, which first queries the TCM syndrome types in the syndrome type set in the syndrome type definition database to obtain all TCM syndrome type definitions, and then sends all TCM syndrome type definitions to the same BERT as the text encoding module to obtain all syndrome type representations; a syndrome type projection module consists of four sub-modules: projection, masking, shared attention, and mask prediction. First, all syndrome type representations are projected into another vector space to obtain mapped syndrome type representations, and then a unique vector with a length equal to the number of all TCM syndrome types is constructed. And randomly set a position to 0, multiply the initialized one-hot vector with the mapped syndrome representation to obtain the masked syndrome representation, then obtain the syndrome shared representation through shared attention, and use mask prediction to help the model learn the representation of the masked syndrome; the understanding transfer module includes two sub-modules: shared attention and difficult text mining. In essence, it divides clinical texts into high-understanding texts and low-understanding texts through shared attention, and uses the teacher-student model to transfer the model's ability to high-understanding texts to low-understanding texts, and uses the output of the student model as the low-understanding representation; the prediction module is used to process the above-mentioned low-understanding representation to obtain the TCM syndrome output by the model;
[0216] The TCM syndrome differentiation model training unit is used to construct the loss function and optimizer required in the model training process, and finally complete the training and optimization of the model; when the model of this method has not been fully trained, it needs to be trained on the training data set to optimize the model parameters; when the model training is completed, the model can predict the corresponding TCM syndrome type for the input clinical text.
[0217] Example 4:
[0218] The storage medium based on Example 2 stores a plurality of instructions, which are loaded by a processor to execute the steps of the TCM syndrome differentiation method based on difficult text mining in Example 2.
[0219] Example 5:
[0220] Based on the electronic device of embodiment 4, the electronic device includes: the storage medium of embodiment 4; and a processor for executing instructions in the storage medium of embodiment 4.
[0221] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A TCM syndrome differentiation method based on difficult text mining, characterized by: The method first constructs a TCM syndrome differentiation dataset, and then builds a TCM syndrome differentiation model consisting of a syndrome type encoding module, a text encoding module, a syndrome type projection module, an understanding transfer module, and a prediction module. Finally, the TCM syndrome differentiation model is trained by jointly constraining multiple loss functions. This can help the model understand the meaning of difficult-to-understand characters in clinical texts, alleviate the long-tail problem of the dataset, and achieve the goal of assigning appropriate TCM syndrome types to clinical texts. The details are as follows: S1. Constructing a training dataset for the TCM syndrome differentiation model: First, process the obtained clinical texts, associate each clinical text with its corresponding TCM syndrome type, and form a complete training sample; aggregate all training samples to form a training dataset for the TCM syndrome differentiation model; and summarize all TCM syndrome types appearing in the training dataset for the TCM syndrome differentiation model into a TCM syndrome type set; S2. Constructing a TCM syndrome differentiation model: The TCM syndrome differentiation model consists of a syndrome type encoding module, a text encoding module, a syndrome type projection module, an understanding transfer module, and a prediction module. S3. Training the TCM syndrome differentiation model: Based on the TCM syndrome differentiation model training dataset constructed in S1, four loss functions are used to jointly optimize the model.
2. The TCM syndrome differentiation method based on difficult text mining according to claim 1 is characterized in that The specific process of constructing a TCM syndrome differentiation dataset is as follows: S1. Constructing a training data set for the TCM syndrome differentiation model: First, the obtained clinical texts are processed, and each clinical text is associated with its corresponding TCM syndrome type to form a training data set. Finally, all training data are aggregated to obtain a training data set for the TCM syndrome differentiation model; all TCM syndrome types appearing in the training data set for the TCM syndrome differentiation model are summarized as a TCM syndrome type set; S101. Processing clinical text: First, screen the patient's original clinical text, remove sensitive information and invalid information, and construct the remaining text into a clinical text; S102, assigning TCM syndrome types: assigning corresponding real TCM syndrome types to the clinical texts constructed in S101, and deleting data corresponding to some TCM syndrome types with higher frequencies to alleviate the imbalance problem of the dataset; S103, constructing a training data set for a TCM syndrome differentiation model: merging the clinical text in S101 with the TCM syndrome type assigned in S102 to form a training sample for a TCM syndrome differentiation model. All training samples are aggregated to construct a training data set for a TCM syndrome differentiation model. S104, constructing a TCM syndrome type set: counting all TCM syndrome types appearing in S102 as a TCM syndrome type set.
3. The TCM syndrome differentiation method based on difficult text mining according to claim 1 is characterized in that Construct text encoding module and certificate type encoding module: S201. Construct a text encoding module: The text encoding module is mainly composed of BERT. It receives the clinical text in S101 as input and encodes it using BERT to obtain a clinical text representation. The details are as follows: The clinical text D consists of N words, that is, the text length of the clinical text D is N, each word is represented by w, and it is represented as D = w1,…,w i ,…,w N ; Then use BERT to encode it, set the maximum clinical text length, that is, the clinical text length threshold is ζ, where ζ is a manually set hyperparameter; for data with a clinical text length less than ζ, it is padded; for data with a clinical text length greater than ζ, it is truncated; the processed clinical text is sent to BERT to obtain the clinical text representation in represents the set of real numbers, that is, all elements in the matrix H are real numbers, and its dimension is ζ×d e , d e Represents the hidden layer size of BERT; S202, constructing a syndrome encoding module: The syndrome encoding module is mainly composed of BERT, which receives the TCM syndrome set obtained in S104; then, a syndrome definition database is constructed, which is constructed with reference to knowledge sources such as "TCM Syndrome Classification and Code" and "Baidu Encyclopedia". Specifically, the database contains relevant medical information such as typical clinical manifestations, etiology, core pathogenesis, syndrome differentiation ideas, related diseases, etc. of TCM syndromes; then, the TCM syndrome definition corresponding to each syndrome in the TCM syndrome definition database is sequentially queried, and all TCM syndrome definitions are aggregated to obtain all TCM syndrome definitions; finally, all TCM syndrome definitions are sent to BERT for encoding to form a complete syndrome representation; S20201. Construct a syndrome definition database: Use crawler tools to crawl various types of information on each type of TCM syndrome from websites such as Baidu Encyclopedia and Wikipedia, such as clinical manifestations and core pathogenesis. To ensure the accuracy of this knowledge, refer to the "Classification and Codes of TCM Syndrome Types" for verification, and add the definition of each TCM syndrome type in the book. All of the above information is summarized as additional knowledge of TCM syndrome types; the TCM syndrome type set obtained in S104 is matched with each of the above additional knowledge of TCM syndrome types to finally obtain a syndrome definition database; S20202. Obtain all syndrome type representations: Receive the TCM syndrome type set obtained in S104. For each TCM syndrome type in the TCM syndrome type set, search for the corresponding additional knowledge in the syndrome type definition database constructed in S20201. Concatenate the additional knowledge as all TCM syndrome type definitions and pass it into the same BERT as S201 to obtain all syndrome type representations, as follows: Assume that the TCM syndrome type set Q obtained in S104 = {q 1 ,q 2 ,...,q m }, where each TCM syndrome type is q j ,j∈1,2,...,m, each TCM syndrome can be expressed as Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104, and Nq represents the number of attributes of the additional knowledge corresponding to the TCM syndrome type. All the additional knowledge is spliced, and the same text length threshold ζ as in S201 is used to fill the data whose text length is less than ζ, and to truncate the data whose text length is greater than ζ. The same BERT as in S201 is used to encode each TCM syndrome type to obtain the syndrome type representation K of the TCM syndrome type. j , j=1,2,..,m, the specific calculation formula is as follows: K j =BERT(q j ) The syndrome type representation of each TCM syndrome type is spliced together to obtain the total syndrome type representation. The specific calculation formula is as follows: in, d e Represents the hidden layer size of BERT.
4. The TCM syndrome differentiation method based on difficult text mining according to claim 1 is characterized in that Construct the certificate projection module as follows: S203, build the syndrome projection module: the syndrome projection module mainly consists of four sub-modules: projection, masking, shared attention, and mask prediction; the projection sub-module makes full use of all syndrome representations obtained in S20202 The goal is to project all syndrome type representations into a new vector space by introducing trainable weights to obtain mapped syndrome type representations. The masked submodule then adopts a masking strategy to randomly mask the mapped syndrome type representations to obtain masked syndrome type representations. The shared attention submodule then establishes a relationship between the masked syndrome type representations and the clinical text representations, allowing the TCM syndrome differentiation model to capture the deep semantic connection between the two and focus on the key related parts of the two to obtain a shared syndrome type representation. Finally, the masked prediction submodule classifies the shared syndrome type representations through a feedforward neural network to obtain masked predictions. S20301. Construct a projection submodule: Multiply all the syndrome type representations obtained in S20202 by the trainable weight vector, map all the syndrome type representations into another vector space, and obtain the mapped syndrome type representations, as follows: in, is the learnable parameter matrix, d e Represents the hidden layer size of BERT; Represents the total number of certificate types obtained by S20202, where ζ represents the truncation or padding value of the text length, which is defined in S201; the final mapping certificate type representation is obtained S20302, constructing a masking submodule: Using a masking strategy for the mapped syndrome type representation obtained in S20301, a one-hot vector is constructed in which only one position is 0 and the rest are 1, so that the mapped syndrome type representation is masked at a random position, i.e., a random TCM syndrome type is masked, thereby obtaining a masked syndrome type representation, as follows: First, construct a unique hot vector with a vector length of m Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; randomly select a number from 1 to m and mask the position in the one-hot vector to 0, that is, Then, the masked one-hot vector is multiplied by the mapped pattern representation to obtain the masked pattern representation. The specific calculation formula is as follows: R m =R⊙O Mask Among them, ⊙ represents the dot product, that is, Each element in O Mask Multiply them to get the masking type representation; d e represents the hidden layer size of BERT; ζ represents the truncation or padding value of the text length, which is defined in S201; S20303. Construct a shared attention submodule: The shared attention submodule uses the masked syndrome type representation obtained in S20302 as the query vector and the clinical text representation obtained in S201 as the key-value vector, and performs an attention operation on the two to obtain a syndrome type shared representation. The specific calculation formula is as follows: Among them, R m is the masked syndrome type representation from S20302, H is the clinical text representation from S201, H share Represents the obtained syndrome type sharing representation, m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; m (H) T Apply Softmax, which converts the similarity score between the two into a probability distribution and assigns a weighting coefficient to each row in H to ensure that the model can weight the contribution of different inputs according to the similarity. The specific calculation formula of the Softmax function is as follows: In the above process, X is R in S20303 m (H) T , x i Represents the similarity score of the masked syndrome relative to the word at position i in the clinical text; S20304. Construct a mask prediction submodule: The mask prediction submodule receives the shared representation of the certificate type obtained in S20303, uses a feedforward neural network to make predictions, and outputs the prediction results, that is, outputs the one-hot vector of the predicted masked position. The specific calculation formula for predicting the masked position is: Y mask =MAX(σ(W mask H share +bias)) Among them, σ represents the activation function, which here represents the Softmax function, that is, the shared expression of the syndrome obtained by S20303 is H share Mapped to all TCM syndrome types obtained in S104; is a one-hot vector of length m, representing the one-hot vector of the predicted masked position, and m represents the number of types of TCM syndromes in the TCM syndrome set constructed in S104; W mask Is a learnable parameter matrix, bias is a learnable parameter, MAX represents the maximum pooling, specifically σ(W mask H share + bias) is a vector of length m. The function of max pooling is to set the maximum value of the m values to 1 and the remaining m-1 positions to 0.
5. The TCM syndrome differentiation method based on difficult text mining according to claim 1 is characterized in that Construct the understanding transfer module and the prediction module as follows: S204, constructing an understanding transfer module: The understanding transfer module receives the clinical text representation obtained in S201; the understanding transfer module includes two submodules, one is a shared attention submodule, and the other is a difficult text mining submodule; The shared attention submodule uses the same shared attention as in S20303, with a custom trainable parameter vector as the query vector for shared attention and the clinical text representation as the key-value vector. The shared attention representation is obtained through shared attention. The difficult text mining submodule receives the shared attention representation and divides the clinical text into two parts: high-comprehension representation and low-comprehension representation through a difficult character selector. The teacher-student model is used for knowledge transfer learning, and the output of the student model is finally obtained as the low-comprehension representation, which is sent to the prediction module. S20401. Construct a shared attention submodule: The shared attention submodule shares the same parameters as the shared attention in S20303. That is, the two are exactly the same. First, a custom trainable parameter vector is designed as the query vector for shared attention, so as to transfer the syndrome type representation learned by shared attention in S20303. Then, the clinical text representation obtained in S201 is received as a key-value vector and the shared attention representation is calculated. The specific calculation formula is as follows: in, is a custom trainable parameter vector, H is the clinical text representation from S201, Represents the obtained shared attention representation, m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; Softmax function is an activation function, which is introduced in detail in S20303 and will not be repeated here; in the above process, Represents W s Attention score relative to clinical text representation H; S20402: Constructing a difficult text mining submodule: First, receive the shared attention representation obtained in S20401 and feed it into the difficult character selector to obtain a high-comprehension representation and a low-comprehension representation. Then, the high-comprehension representation is used as the teacher model, and the low-comprehension representation is used as the student model. Finally, the teacher model's ability to understand text characters is transferred to the student model through KL-LOSS. S2040201. Build a difficult character miner: Receive the shared attention representation obtained in S20401 and sort each row of vectors to obtain the ascending order index of the shared attention representation. Then, two strategies are designed to divide the shared attention representation into high-understanding representation and low-understanding representation, as follows: Convert the shared attention representation obtained in S20401, and use the Sort function to sort each row vector in the shared attention in ascending order, and then obtain the corresponding descending order index. The formula is as follows: <h2 style=";text-align:left;direction:ltr">H<h2 style=";text-align:left;direction:ltr"> tran_share <h2 style=";text-align:left;direction:ltr"> (a1,a2,......,a)<h2 style=";text-align:left;direction:ltr"> m <h2 style=";text-align:left;direction:ltr"> ] Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; a k Represents the shared attention representation H obtained in S20401 tran_share The k-th row tensor in , e1 represents the index of the highest score position in the k-th row shared attention representation, and Represented as the index of the lowest score position in the k-th row of shared attention representation; the Sort function is a pytorch intrinsic sorting operation that can sort the elements of a tensor of any length in ascending or descending order according to the specified dimension and return the sorted index. In the above formula, the first dimension of the shared attention representation is sorted in descending order; In order to divide the shared attention representation into high-understanding representation and low-understanding representation, we first set a length of I k The same all-zero vector ME k and all-one vector MH k , then I k Middle front beta f % position sequence and randomly selected β c % The position sequence is set to 1, and then a new mask matrix MEO is obtained k , and then MEO k Sequentially with a k Dot product, finally get the high-level understanding representation A easy , the formula is as follows: A easy =[a1⊙MEO1,a2⊙MEO2,......,a m ⊙MEO m ] Similarly, in order to avoid the influence of too low attention on the final result, the middle I k Post-β l % position sequence, front β f % position sequence and randomly selected β c % The position sequence is set to 0, and then a new mask matrix MHO is obtained k , then MHO k Sequentially with a k Dot product, finally get the low-level understanding representation A hard , the formula is as follows: A hard =[a1⊙MHO1,a2⊙MHO2,......,a m ⊙MHO m ] S2040202. Construct KL-LOSS: Receive the high-understanding representation and low-understanding representation obtained in S2040201, and use Kullback-Leibler divergence, or KL-LOSS, to transfer the model's understanding ability. The specific calculation formula is as follows: KL=KL-LOSS(A easy ||A hard ) Among them, KL-LOSS is a loss function used to measure the difference between two tensor distributions. KL-LOSS can be used to calculate the relative entropy between the two distributions. At the same time, the smaller the KL-LOSS value, the more similar the two distributions are. The specific calculation formula is as follows: In the above process, P stands for high understanding and A easy ,Q represents low understanding representation; S205, constructing a prediction module: receiving the low-level representation obtained in S2040201, using a feedforward neural network to predict the TCM syndrome type and output the result, that is, outputting a multi-hot vector of the TCM syndrome type, and finally outputting the TCM syndrome type through the correspondence between the position of 1 in the vector and the TCM syndrome type. The specific calculation formula is as follows: Y label =σ(W label A hard +bias label ) in, is a one-hot vector of length m, where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; W label Is a learnable parameter matrix, bias label is the bias term; σ represents sigmoid, which converts the activation of the vector into an independent probability, and binarizes the probability into 0 / 1 through the set threshold to obtain a vector of length m; the position of 1 in the vector is mapped to the corresponding TCM syndrome type, and the TCM syndrome type Y is output. label .
6. The TCM syndrome differentiation method based on difficult text mining according to claim 1 is characterized in that The TCM syndrome differentiation model is trained and optimized as follows: S3. Training the TCM syndrome differentiation model: Optimize the model by combining KL-LOSS, Drr-LOSS, mask prediction loss, and syndrome type prediction loss according to the weights; first receive the KL-LOSS obtained in S2040202; then calculate Drr-LOSS; then receive the true masked position obtained in S20302 and the predicted masked position obtained in S20304, and calculate the binary cross entropy loss between the two; then receive the TCM syndrome type obtained in S205 and the TCM syndrome type corresponding to the real clinical text to calculate the cross entropy loss; weighted sum of the four losses to obtain the total loss function; When the model of this method has not been fully trained, it needs to be trained in the training data set to optimize the model parameters; when the model training is completed, the model can predict the corresponding syndrome type for the input clinical text; S401, calculate Drr-LOSS: receive the mapped syndrome type representation obtained in S20301, and use cosine similarity distance to measure the distance between different rows in the mapped syndrome type representation, and maximize the sum of the distances between different rows in the mapped syndrome type representation to enhance the distinction between different TCM syndrome types. The formula is as follows: Cosine Similarity Distance represents the cosine similarity distance, and the specific calculation formula is as follows: In the above process, X represents R i , Y stands for R j ; S402. Calculate the mask prediction loss: Construct the actual mask position obtained in S20302 into the corresponding one-hot vector Y′, then receive the one-hot vector of the predicted mask position obtained in S20304, and calculate the binary cross entropy loss between the two, i.e., the mask prediction loss, as follows: Where m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104, Y mask represents the one-hot vector of the predicted occlusion position obtained in S20304, and Y′ represents the one-hot vector of the actual occlusion position obtained in S20302; S403, calculating the syndrome prediction loss: receiving the multi-hot vector of the TCM syndrome obtained in S205 and the multi-hot vector of the TCM syndrome corresponding to the real clinical text, and calculating the cross entropy loss. The formula is as follows: Among them, Y represents the unique heat vector of the TCM syndrome type corresponding to the real clinical text, Y label represents the unique-hot vector corresponding to the predicted TCM syndrome type, and m represents the number of TCM syndrome types in the TCM syndrome type set constructed in S104; S404. Construct the total loss function: It contains four parts, namely KL-LOSS obtained in S2040202, Drr-LOSS obtained in S401, mask prediction loss obtained in S402, and syndrome prediction loss obtained in S403; and add the four together according to the weights of λ1:λ2:λ3:λ4. The formula is as follows: λ1:λ2:λ3:λ4 are manually set hyperparameters; S405. Construct an optimization function: Use the Adamw algorithm as the optimization function of the model; the learning rate parameter is set to 0.001, and other hyperparameters use the default values in PyTorch; hyperparameters refer to parameters that need to be manually set before starting the training process; these parameters cannot be automatically optimized through training; depending on the actual data set, these parameters need to be manually set by the user.
7. A TCM syndrome differentiation device based on difficult text mining, characterized in that: The model construction unit includes a TCM syndrome differentiation training data set construction unit, a TCM syndrome differentiation model construction unit, and a TCM syndrome differentiation model training unit, which respectively implement the TCM syndrome differentiation method based on difficult text mining described in claims 1-6, as follows: The TCM syndrome differentiation training data set construction unit and the TCM syndrome differentiation model training data set construction unit are used to process the input clinical text. Specifically, each clinical text is assigned its corresponding TCM syndrome type, and finally a TCM syndrome differentiation model training data set is constructed; the TCM syndrome types appearing in the TCM syndrome differentiation model training data set are summarized as all TCM syndrome types; The TCM syndrome differentiation model construction unit consists of five parts: a text encoding module, which encodes clinical texts through BERT to obtain clinical text representations; a syndrome type encoding module, which first queries the TCM syndrome types in the syndrome type set in the syndrome type definition database to obtain all TCM syndrome type definitions, and then sends all TCM syndrome type definitions to the same BERT as the text encoding module to obtain all syndrome type representations; a syndrome type projection module consists of four sub-modules: projection, masking, shared attention, and mask prediction. First, all syndrome type representations are projected into another vector space to obtain mapped syndrome type representations, and then a unique vector with a length equal to the number of all TCM syndrome types is constructed. And randomly set a position to 0, multiply the initialized one-hot vector with the mapped syndrome representation to obtain the masked syndrome representation, then obtain the syndrome shared representation through shared attention, and use mask prediction to help the model learn the representation of the masked syndrome; the understanding transfer module includes two sub-modules: shared attention and difficult text mining. In essence, it divides clinical texts into high-understanding texts and low-understanding texts through shared attention, and uses the teacher-student model to transfer the model's ability to high-understanding texts to low-understanding texts, and uses the output of the student model as the low-understanding representation; the prediction module is used to process the above-mentioned low-understanding representation to obtain the TCM syndrome output by the model; The TCM syndrome differentiation model training unit is used to construct the loss function and optimizer required in the model training process, and finally complete the training and optimization of the model; when the model of this method has not been fully trained, it needs to be trained on the training data set to optimize the model parameters; when the model training is completed, the model can predict the corresponding TCM syndrome type for the input clinical text.
8. The TCM syndrome differentiation device based on difficult text mining according to claim 7 is characterized in that: The model training unit includes: The loss function construction module is used to calculate the distance between different rows in the mapped syndrome type representation using Drr-LOSS, calculate the error between the true masked position and the predicted masked position using mask prediction loss, calculate the difference between high-understanding representation and low-understanding representation using Kl-Loss, and calculate the error between the predicted TCM syndrome type and the true TCM syndrome type using syndrome type prediction loss; The loss function optimization module is used to train and adjust the parameters in model training to reduce the prediction error.
9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the TCM syndrome differentiation method based on difficult text mining according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the TCM syndrome differentiation method based on difficult text mining according to any one of claims 1 to 6.