Case record text data-based cause potential coding method, equipment and medium

Through the causal potential encoding method based on the medical record text data, the BERT sub-model and traditional Chinese medicine coding system are used to quantitatively analyze the medical record text, which solves the problems of subjective differences and low efficiency in traditional Chinese medicine case analysis, and realizes efficient and precise assistance in traditional Chinese medicine diagnosis and treatment.

CN120280118APending Publication Date: 2025-07-08CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145578.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing traditional Chinese medicine case analysis methods rely on manual annotation, have subjective differences and are inefficient, making it difficult to conduct systematic quantitative analysis of the cause, location, disease nature and condition, and cannot effectively support the rapid and accurate diagnosis and treatment of traditional Chinese medicine syndrome differentiation and treatment.

Method used

The causal potential encoding method based on the medical record text data is adopted, and the medical record text is preprocessed and marked using the BERT sub-model, a two-way long and short-term memory network and a conditional random field to generate traditional Chinese medicine codes, and quantitatively assign scores in combination with the 46-bit traditional Chinese medicine coding system to realize quantitative analysis of the cause, location, disease nature and condition.

Benefits of technology

It improves the efficiency and accuracy of traditional Chinese medicine diagnosis and treatment, reduces the subjective deviation of artificial intervention, provides reliable data support and assists in decision-making, and realizes structured processing and quantitative analysis of medical record data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280118A_ABST
    Figure CN120280118A_ABST
Patent Text Reader

Abstract

The invention discloses a cause potential coding method and device based on medical record text data and a medium. The method comprises the steps of obtaining and preprocessing the medical record text data; inputting the preprocessed medical record text data into a coding model, and performing cause potential labeling on the medical record text data to obtain a traditional Chinese medicine code of the medical record text data; the coding model is obtained through training of a training data set, and the training data set comprises preprocessed medical record data and labels; and assigning scores to the traditional Chinese medicine codes according to a traditional Chinese medicine coding system. According to the method, the disease cause, the disease location, the disease nature and the disease condition information in the medical record text are quantitatively analyzed, and data support and auxiliary decision making can be provided for traditional Chinese medicine diagnosis and treatment, so that the efficiency and the accuracy of traditional Chinese medicine clinical diagnosis and treatment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of digitalization of traditional Chinese medicine syndrome differentiation and treatment, and specifically relates to a method, device, and medium for encoding cause, location, nature, and trend based on medical record text data. Background Art

[0002] Traditional Chinese medicine syndrome differentiation and treatment is an important core method in traditional Chinese medicine clinical diagnosis and treatment, which includes two key links: "syndrome differentiation" and "treatment based on syndrome differentiation". "Syndrome differentiation" comprehensively analyzes information such as the patient's symptoms, signs, and medical history to judge the cause, location, nature, and trend of the disease, and summarizes it into specific traditional Chinese medicine syndromes; "treatment based on syndrome differentiation" formulates corresponding personalized treatment plans according to the syndrome differentiation results, including diagnosis and treatment principles, prescription selection, and medication recommendations. Since traditional Chinese medicine diagnosis and treatment has strong individual characteristics and systematicness, its diagnosis and treatment process requires relying on rich traditional Chinese medicine theoretical knowledge and clinical experience.

[0003] With the development of medical digitalization, more and more traditional Chinese medicine medical record data is stored electronically. These medical record texts contain a large amount of traditional Chinese medicine syndrome differentiation and treatment information, including symptom descriptions, syndrome judgments, and treatment records. However, the existing traditional Chinese medicine medical record analysis methods have the following problems:

[0004] Mainly annotate and statistically analyze traditional Chinese medicine medical records manually, relying on the traditional Chinese medicine knowledge level of annotators, which is prone to subjective differences and has low efficiency. The cause, location, nature, and trend in traditional Chinese medicine medical records mostly exist in text form, and it is difficult for existing methods to conduct systematic quantitative analysis on them, lacking a standardized processing process. It is relatively weak in generating diagnosis and treatment suggestions using medical record text data and cannot effectively support doctors in quickly and accurately differentiating syndromes and treating diseases. Summary of the Invention

[0005] To solve the above technical problems, a method, device, and medium for encoding cause, location, nature, and trend based on medical record text data are provided, which quantitatively analyze the cause, location, nature, and trend information in the medical record text, can provide data support and auxiliary decision-making for traditional Chinese medicine diagnosis and treatment, and thus improve the efficiency and accuracy of traditional Chinese medicine clinical diagnosis and treatment.

[0006] To solve the above technical problems, the first aspect of the present invention discloses a method for encoding cause, location, nature, and trend based on medical record text data, including:

[0007] Obtain and preprocess medical record text data;

[0008] Input the preprocessed medical record text data into an encoding model to perform cause, location, nature, and trend annotation on the medical record text data, and obtain the traditional Chinese medicine encoding of the medical record text data; the encoding model is trained by a training data set, and the training data set includes preprocessed medical record data and labels;

[0009] Assign scores to the traditional Chinese medicine (TCM) codes according to the TCM coding system to determine the key codes.

[0010] In some embodiments, the medical record text data includes TCM medical records, symptom descriptions, syndrome type annotations, or prescriptions; performing in-position potential annotations on the medical record text data to obtain the TCM codes of the medical record text data, including:

[0011] Performing syndrome type annotations on the etiology, disease location, disease nature, and disease trend contents in the preprocessed medical record text data to generate a label sequence; the syndrome type annotations include TCM information, diagnosis and treatment methods, used drugs, and TCM symptoms.

[0012] According to the TCM coding system, map the coding sequence corresponding to the label sequence to generate TCM codes.

[0013] In some embodiments, the coding model includes a BERT sub-model, a bidirectional long short-term memory network, and a conditional random field; inputting the preprocessed medical record text data into the coding model, including:

[0014] Using the BERT sub-model to perform word segmentation and feature extraction on the medical record text data to obtain medical record text features.

[0015] Performing context analysis and entity recognition on the medical record text features through the bidirectional long short-term memory network to obtain medical record text features including context relationships.

[0016] Predicting the label sequence of the medical record text features through the conditional random field.

[0017] In some embodiments, the bidirectional long short-term memory network includes a forget gate, an input gate, and an output gate. The forget gate takes the medical record text features at the current time step, the input at the current time step, and the hidden state at the previous time step as inputs, and calculates the forget ratio control signal through the first formula.

[0018] The input gate takes the medical record text features at the current time step and the hidden state at the previous time step as inputs, and generates the control signal of the input gate and the candidate memory cell state through the second set of formulas. The output of the input gate is used to update the memory cell state at the current time step.

[0019] The output gate takes the input at the current time step, the memory cell state at the current time step, and the hidden state at the previous time step as inputs, generates the control signal of the output gate through the third formula, and combines the memory cell state at the current time step to generate the hidden state at the current time step as the output.

[0020] In some embodiments, the first formula is:

[0021] F t =σ(X t Wxf +H t-1 W hf +b f )

[0022] where σ is the Sigmoid activation function, X t is the input vector at the current time step t, W xf is the weight matrix of the input vector X t , representing the weight from the input to the forget gate; H t-1 is the hidden state of the previous time step, representing the memory information of the previous time step; W hf is the weight matrix of the hidden state H t-1 , representing the weight from the hidden state to the forget gate; b f is the bias term of the forget gate, used to offset and adjust the result of the linear transformation;

[0023] The second set of formulas includes:

[0024] I t = σ(X t W xi +H t-1 W hi +b i )

[0025] where, I t is the output of the forget gate, determining how much of the current input vector X t needs to be written into the cell state, with a value range of [0,1]; σ is the Sigmoid activation function, X t is the input vector at the current time step t, W xi is the weight matrix of the input vector X t , H t-1 is the hidden state of the previous time step, representing the memory information of the previous time step; W hi is the weight matrix of the hidden state H t-1 in the input gate; b i is the bias term of the input gate;

[0026]

[0027] where, is the candidate cell state, representing the new candidate memory calculated based on the input at the current time step; tanh is the hyperbolic tangent activation function, used to map the linear result to the interval (-1,1), adding non-linearity; X t is the input vector at the current time step t; W xc is the weight matrix of the input vector X t , used for the candidate cell state; H t-1is the hidden state of the previous time step, representing the memory information of the previous time step; W hc is the weight matrix of the hidden state H t-1 for the candidate cell state; b c is the bias term of the candidate cell state;

[0028]

[0029] where C t is the cell state of the current time step, combining the memory state of the previous time step and the new information of the current time step; F t is the output of the forget gate, C t-1 is the cell state of the previous time step, I t is the output of the forget gate, is the candidate cell state, and ⊙ is element-wise multiplication.

[0030] In some embodiments, the traditional Chinese medicine coding includes an etiology coding factor, a disease location coding factor, a disease nature coding factor, and a disease trend coding factor. The etiology coding factor includes wind, stirring wind, dryness, dampness, water, cold, summer heat, phlegm, fluid deficiency, toxin, food, insect, qi stagnation, blood stasis, collateral obstruction, and bleeding. The disease location coding factor includes the heart, liver, gallbladder, spleen, stomach, intestine, lung, kidney, bladder, uterus, thoroughfare and conception vessels, and sperm chamber. The disease nature coding factor includes heat, deficiency heat, blood heat, cold, accumulation, qi deficiency, blood deficiency, yin deficiency, yang deficiency, insecurity, fu excess, fluid deficiency, and toxin. The disease trend coding factor includes restlessness of the heart, coma, yang hyperactivity, cough and asthma, and qi reversal.

[0031] In some embodiments, scores are assigned to the traditional Chinese medicine coding according to a 46-bit traditional Chinese medicine coding system to determine key coding factors, including:

[0032] Count the occurrence frequencies of each coding factor in the traditional Chinese medicine coding;

[0033] Calculate the syndrome score according to the occurrence frequency of the coding factor and the corresponding weight. The range of the syndrome score is 0 to 9 points;

[0034] Screen key coding factors that meet the preset requirements according to the syndrome score.

[0035] In some embodiments, the preprocessing includes:

[0036] Perform text cleaning, word segmentation processing on the medical record text data, and form structured data.

[0037] According to the second aspect of the present invention, a computer device is disclosed, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of a method for encoding etiology, location, nature, and trend based on medical record text data as described in any one of the above.

[0038] According to the third aspect of the present invention, a computer storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for encoding cause-position potential based on medical record text data as described above are implemented.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] The present invention provides a method, device, and medium for encoding cause-position potential based on medical record text data, which converts complex medical record text data into structured traditional Chinese medicine codes, and uses a 46-bit traditional Chinese medicine coding system to quantitatively score medical record data, providing an intuitive quantitative basis for diagnosis and treatment decisions. This method combines the automatic annotation ability of the coding model, significantly improves the efficiency and accuracy of medical record data processing, reduces the subjective deviation of human intervention, and at the same time accurately reflects the characteristics and correlations of the cause, location, nature, and trend of the disease through the scoring results, providing reliable technical support for traditional Chinese medicine diagnosis and treatment assistance and big data analysis of medical records. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic flowchart of a method for encoding cause-position potential based on medical record text data provided by the present invention;

[0042] Figure 2 It is a schematic flowchart of step S2 of a method for encoding cause-position potential based on medical record text data provided by the present invention;

[0043] Figure 3 It is a schematic flowchart of step S3 of a method for encoding cause-position potential based on medical record text data provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] For better understanding and implementation, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0045] The terms "including" and "having" and any variations thereof in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules does not necessarily have to be limited to those clearly listed steps or modules, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0046] An embodiment of the present invention discloses a location-based potential encoding method based on medical record text data, which quantitatively analyzes the etiology, disease location, disease nature, and disease potential information in the medical record text, can provide data support and auxiliary decision-making for traditional Chinese medicine diagnosis and treatment, and thus improve the efficiency and accuracy of traditional Chinese medicine clinical diagnosis and treatment.

[0047] As Figure 1 shown, this method includes:

[0048] Step S1: Collect and preprocess the medical record text data.

[0049] The medical record text data refers to unstructured or semi-structured text data recording the user's condition information, including the user's basic information, traditional Chinese medicine medical records, symptom descriptions, syndrome type annotations and other data. Exemplarily, it may include the present illness history: "The patient has had excessive thirst, polyuria, weight loss, and fatigue for half a year, which has not been taken seriously. 3 days ago, the weight gain was 10.8 kg / L during an examination in a certain hospital; urine routine: GLU(3+), diagnosed with type 2 diabetes. Currently, the patient feels fatigue, excessive thirst and polyuria, normal diet, normal sleep, normal bowel movement, and shows symptoms of numbness and coldness in hands and feet. Physical examination: red tongue, thin coating, slippery pulse, blood sugar 10.5 mmol / L. Syndrome diagnosis: qi and yin deficiency, internal accumulation of dryness-heat."; traditional Chinese medicine prescription: "Atractylodes macrocephala 12g, Lycium barbarum 12g, Scrophularia ningpoensis 12g, Scutellaria baicalensis 30g, Poria cocos 30g, Pueraria lobata 30g, Angelica sinensis 30g, Rehmannia glutinosa 15g, Chuanxiong rhizome 20g, Chinese yam 9g"; diagnostic etiology: "(September 13th) The patient's symptoms of excessive thirst and polyuria are better. After activity, the fatigue can be attributed to excessive oxygen inhalation. Red tongue, slippery pulse, blood sugar 5.5 mmol / L." etc., which are not limited in this application.

[0050] After obtaining the medical record text data, preprocess the medical record text data to meet the requirements of the encoding model input and improve the accuracy of model encoding. The preprocessing includes text cleaning of the medical record text data, word segmentation processing, and forming structured data. Text cleaning is to remove useless information, symbols, and stop words in the medical record text data, and perform word segmentation processing. The useless information, symbols, and stop words include but are not limited to the user's identity-sensitive information, personal privacy information, or incorrect words, etc.

[0051] The word segmentation processing is to use a traditional Chinese medicine professional vocabulary library for accurate Chinese word segmentation. For example, the original medical record text data is qi stagnation and blood stasis, string-like and slippery pulse. After word segmentation processing, it is [qi stagnation] / [blood stasis] / [pulse] / [string-like and slippery]. Through word segmentation processing, the medical record text data is formed into structured data, so that the encoding model can obtain the required data more quickly and accurately.

[0052] Step S2: Input the preprocessed medical record text data into the encoding model to obtain the traditional Chinese medicine encoding of the medical record text data; the encoding model is obtained by training with a training dataset, and the training dataset includes preprocessed medical record data and labels, and the labels are the traditional Chinese medicine encodings obtained by annotating the in-position potential of the medical record data.

[0053] The encoding model includes a training stage and an application stage. In the training stage, the encoding model is trained with a training dataset and obtained by training using an objective loss function. The training of the encoding model is through a supervised learning method, enabling the encoding model to learn how to extract features from the preprocessed medical record data and correctly predict the labels of multiple medical record data in the training dataset. The labels classify each word or phrase in the medical record data and represent the key elements of traditional Chinese medicine syndrome differentiation and treatment in a structured form, helping the model understand the semantic content of the medical record text and providing support for subsequent diagnosis and treatment assistance and encoding generation.

[0054] The labels in the training dataset are obtained by manual syndrome type annotation. The syndrome type annotation is performed on the etiology, disease location, disease nature, and disease trend content of the medical record data, specifically including aspects such as theory, method, formula, medicine, and symptoms. Exemplarily, theory includes traditional Chinese medicine information such as zang-fu organs and meridians, etiology and pathogenesis, the method is the diagnosis and treatment method adopted according to the corresponding syndrome, such as strengthening healthy qi and eliminating pathogenic factors, regulating yin and yang, tonifying qi and blood, the medicine is the traditional Chinese medicine used, and traditional Chinese medicine symptoms, such as excessive thirst and polydipsia. A label sequence is generated through syndrome type annotation, and the sequence represents the content of the relevant medical record text data; after annotation, the content such as symptoms, theory, method, formula, and medicine is distinguished by different labels. Exemplarily, the label Sym is for symptoms; the label method is for method, and the label drug represents medicine, etc.

[0055] After annotating the unannotated sequence X = {χ1, χ2, Λ, χ n} of length n, the annotated sequence Y = {y1, y2, Λ, y n} is obtained, and the start, inside, and non-entity of the entity are annotated to generate an annotation sequence. For example, numbness of hands and feet is a disease symptom; the hand is the start of this symptom and is annotated as B-sym (Beginning of Symptom, entity start), and the subsequent numbness of feet is I-sym (Inside of Symptom, entity inside), and the unannotated is non-entity 0.

[0056] The training data set is input into the encoding model to be trained. The encoding model includes at least a BERT sub-model, a bidirectional long short-term memory network, and a conditional random field. The BERT sub-model is used to perform word segmentation and feature extraction on the medical record data to obtain medical record training features; the bidirectional long short-term memory network is used to perform context analysis and entity recognition on the medical record training features to obtain medical record training features including context relationships; the conditional random field is used to predict the label sequence of the medical record training features including context relationships.

[0057] The word segmentation and feature extraction by the BERT include:

[0058] Convert the medical record data into an input sample in a preset sample format;

[0059] Perform feature extraction on the input sample, encode it based on the self-attention mechanism, and output the extracted medical record training features.

[0060] These include: Token IDs, Attention Mask, and Segment IDs. Among them, Token IDs are the corresponding word IDs of the medical record in the BERT vocabulary. Generally, BERT uses WordPiece word segmentation, so a word may be decomposed into multiple sub-words. The Attention Mask is a binary mask used to indicate which positions are actual words (1) and padding (0), ensuring that the encoding model only focuses on the actual input during calculation and ignores the padding part. The padding part can be understood as non-entities to focus the attention mechanism on the key Chinese medical texts, that is, entities. Segment IDs are used to distinguish two sentences. If the input is two sentences, such as in a question-answering task, usually the ID of the first sentence is 0 and the ID of the second sentence is 1.

[0061] The embedding layer performs feature embedding, including TokenEmbedding word embedding, SegmentEmbedding segment embedding, and PositionEmbedding position embedding. Among them, Token Embeddings: Each Token ID is mapped into a high-dimensional space to form word embeddings. Segment Embeddings are segment embeddings. If the input sample contains a second sentence, the BERT model will add a segment embedding for each token to distinguish different sentences. The BERT model also introduces Position Embeddings to retain the position information of words and capture sequence information.

[0062] Finally, each token input into the embedding layer will be represented by the sum of the above three embeddings, that is, the representation of each token is:

[0063] Embedding token = Token Embedding + Segment Embedding + Position Embedding

[0064] The BERT sub - model also captures complex dependencies in the input sequence through the Transformer encoder layer. The Transformer encoder is composed of multiple stacked Transformer encoder sub - layers. Each Transformer encoder sub - layer contains a self - attention mechanism and a feed - forward neural network. BERT allows the model to capture complex dependencies in the input sequence.

[0065] The key part of the Transformer encoder layer is the Transformer structure. Transformer is a deep network based on the "self - attention mechanism". It adjusts the weight coefficient matrix according to the degree of association between words in the same sentence to obtain the representation of words:

[0066]

[0067] Among them, Q, K, and V are word vector matrices, and is the Embedding dimension. The multi - head attention mechanism projects Q, K, and V through multiple different linear transformations, and finally concatenates different Attention results.

[0068] Based on the multi - head attention mechanism, the model is allowed to focus on the relationships between different words and other words in each layer. The representation of each word depends not only on its own word vector but also on other words in the context. After each self - attention layer, there is a feed - forward neural network (Feed Forward Neural Network), which independently processes the representation of each position. After each sub - layer, BERT uses layer normalization and residual connection to stabilize the training process and accelerate convergence.

[0069] After the Transformer encoder layer, the output of the BERT sub - model is the last - layer hidden state of each input token. These hidden states are context - aware and are used for the task of medical record named entity recognition and the text classification task of the principles, methods, prescriptions, and herbs in medical records. In the text classification task of medical records, the pooling layer uses the output of the [CLS] token (classification token) as the feature representation of the whole sentence, that is, the medical record text feature. BERT uses the [CLS] token for the medical record classification task during training.

[0070] Input the medical record training features obtained by the BERT sub-model into a bidirectional long short-term memory network to capture context features. The medical record training features contain rich context information and also incorporate the cell state C. t-1 Preserve global information throughout the process to help the model understand each word in the medical record text features. A single LSTM neuron consists of three modules: a forget gate, an input gate, and an output gate.

[0071] The inputs to the forget gate are the input vector X at the current moment t , that is, the medical record training features and the output H of the previous neuron t-1 . After being processed by the sigmoid activation function, the two inputs will be mapped to values between 0 and 1. At this time, the closer the mapped value is to 1, it indicates that the current information is more likely to be forgotten. If the mapped value is closer to 0, it indicates that the information is more likely to be remembered by the neuron. Its role is to determine whether and how much of the previous moment's information should be retained, so as to calculate the forgetting ratio control signal through the first formula. The specific first formula is as follows:

[0072] F t =σ(X t W xf +H t-1 W hf +b f )

[0073] where σ is the Sigmoid activation function, X t is the input vector at the current time step t, W xf is the weight matrix of the input vector X t , representing the weight from the input to the forget gate; H t-1 is the hidden state of the previous time step, representing the memory information of the previous time step; W hf is the weight matrix of the hidden state H t-1 , representing the weight from the hidden state to the forget gate; b f is the bias term of the forget gate, used to offset and adjust the result of the linear transformation.

[0074] The input gate then calculates X t and H t-1 , and determines which information in the two inputs needs to be retained. Among them, the function of the sigmoid activation function is similar to that of the forget gate, which will determine the information to be forgotten and the information to be remembered. Through the tanh function, X t and H t-1 are integrated and mapped to values between -1 and +1 to represent the candidate vector of the current neuron, which is used to update the neuron state C t . The specific second formula is as follows:

[0075] It = σ(X t W xi + H t-1 W hi + b i )

[0076] where I t is the output of the forget gate, determining how much of the current input vector X t needs to be written into the cell state, with a value range of [0, 1]; σ is the Sigmoid activation function, X t is the input vector at the current time step t, W xi is the weight matrix of the input vector X t , H t-1 is the hidden state of the previous time step, representing the memory information of the previous time step; W hi is the weight matrix of the hidden state H t-1 in the input gate; b i is the bias term of the input gate.

[0077]

[0078] where is the candidate cell state, representing the new candidate memory calculated based on the input at the current time step; tanh is the hyperbolic tangent activation function, used to map the linear result to the interval (-1, 1) to add non-linearity; X t is the input vector at the current time step t; W xc is the weight matrix of the input vector X t for the candidate cell state; H t-1 is the hidden state of the previous time step, representing the memory information of the previous time step; W hc is the weight matrix of the hidden state H t-1 for the candidate cell state; b c is the bias term of the candidate cell state.

[0079]

[0080] where C t is the cell state at the current time step, combining the memory state of the previous time step and the new information at the current time step; F t is the output of the forget gate, C t-1 is the cell state of the previous time step, I t is the output of the forget gate, is the candidate cell state, and ⊙ is element-wise multiplication.

[0081] The output gate updates the state of the neuron at the current moment. Taking the input at the current time step, the state of the memory cell at the current time step, and the hidden state at the previous time step as inputs, it generates a control signal for the output gate through the third formula, and combines the state of the memory cell at the current time step to generate the hidden state at the current time step as the output. The third set of formulas is as follows:

[0082] O t = σ(X t W x0 + H t-1 W ho + b0)

[0083] H t = O t ⊙ tanh(C t )

[0084] In the above formulas, I t , F t , O t and C t represent the input gate, forget gate, output gate, and candidate memory element respectively. W x0 , W ho and b0 represent the weight matrices and bias parameters of the forget gate, input gate, and output gate respectively. σ is the Sigmoid activation function, tanh is the hyperbolic tangent activation function. X t is the input of the neuron at the current moment, and H t-1 is the hidden state of the neuron at the previous moment. Through the above calculations, the finally output cell state C t is the weighted sum of the old memory and the new memory, combining long-term and short-term memory information.

[0085] After combining the long-term and short-term bidirectional memory information, the case training features containing context relationships obtained from the above bidirectional long short-term memory network are input into the conditional random field to optimize the prediction of the label sequence of the case data. To obtain the optimal solution of the labeling sequence under the global condition, that is, the sequence with the highest probability, for the given input sequence X and label sequence, the CRF layer will calculate and output a score Y, which reflects the matching degree of the label sequence with the input sequence. The formula is as follows:

[0086]

[0087] Among them, the transition matrix M represents the probability of the label transitioning from y i to y i+1 , and P i,yi represents the probability that the i-th word is labeled as y i , and n is the sequence length.

[0088] The probability of obtaining the predicted sequence Y is given by the following formula:

[0089]

[0090] Among them, Y x is all possible annotation combinations, is the true annotation sequence. In order to maximize the predicted label P(Y|X), the Maximum Likelihood Estimation (MLE) is used to learn the model parameters. The formula is as follows:

[0091]

[0092] The most likely annotation combination is obtained through the conditional random field to ensure that the label predictions of the classifications of "etiology", "disease location", "disease nature", and "disease trend" are consistent with the actual semantic logic, and accurate annotation and classification of the key traditional Chinese medicine information related to "etiology, disease location, disease nature, and disease trend" in the medical record text are realized.

[0093] During the training process, the cross-entropy loss is used to measure the similarity between two probability distributions. In the classification problem, the actual label is represented as a one-hot vector, that is, only one element is 1 and the rest are 0; while the prediction result of the encoding model is usually a probability distribution. The cross-entropy loss evaluates the performance of the model by calculating the cross-entropy between the actual label and the prediction result of the encoding model. The cross-entropy loss function is:

[0094]

[0095] Among them, y i is the probability of the i-th category in the actual label (one-hot encoding), and p i is the probability of the i-th category predicted by the model.

[0096] After training through the above steps, the final encoding model is obtained. In the application stage, the user can directly input the preprocessed medical record text data into the encoding model. As Figure 2 shown, the label sequence is generated by the encoding model and mapped to obtain the traditional Chinese medicine code, which specifically includes:

[0097] Step S21: Perform the above-mentioned syndrome type annotation on the etiology, disease location, disease nature, and disease trend contents in the preprocessed medical record text data to generate a label sequence; the syndrome type annotation includes but is not limited to traditional Chinese medicine information, diagnosis and treatment methods, used drugs, and traditional Chinese medicine symptoms;

[0098] Step S21: According to the 46-bit traditional Chinese medicine coding system, map the coding sequence corresponding to the label sequence to generate the traditional Chinese medicine code.

[0099] When the trained model inputs the preprocessed medical record text, it can generate the label sequence of the medical record text data through the above process, and then map the corresponding coding sequence of the label sequence according to the traditional Chinese medicine coding system to generate the traditional Chinese medicine codes for etiology, disease location, disease nature, and disease trend.

[0100] In some embodiments, the traditional Chinese medicine code includes an etiology coding factor, a disease location coding factor, a disease nature coding factor, and a disease trend coding factor, with a total of 46 bits. The etiology coding factors include wind, stirring wind, dryness, dampness, water, cold, summer heat, phlegm, fluid deficiency, toxin, food, insect, qi stagnation, blood stasis, collateral obstruction, and bleeding. The disease location coding factors include heart, liver, gallbladder, spleen, stomach, intestine, lung, kidney, bladder, uterus, thoroughfare and conception vessels, and sperm chamber. The disease nature coding factors include heat, deficiency heat, blood heat, cold, accumulation, qi deficiency, blood deficiency, yin deficiency, yang deficiency, insecurity, excessive heat in the fu-organs, fluid deficiency, and toxin. The disease trend coding factors include restlessness of the mind, coma, yang hyperactivity, cough and asthma, and qi counterflow.

[0101] Taking diabetes as an example, according to the summary of the relationship between the traditional Chinese medicine coding system and traditional Chinese medicine diagnosis, its etiology includes phlegm, dampness, dryness, liver depression, yang deficiency, yin deficiency, blood stasis, and collateral obstruction. The disease location includes the lung, stomach, and kidney. The disease nature includes deficiency of the root and excess of the tip, with yin deficiency as the root and dryness and heat as the tip. The disease trend includes four stages: heat, depression, deficiency, and impairment. As the screening conditions for the diagnosis and treatment plan of diabetic patients, the discrimination of these four aspects comes from the symptoms, principles of treatment, and formulas and herbs in the medical record text data, which is convenient for corresponding to the content in traditional Chinese medicine texts. Identify and classify the named entities of the data in these four aspects in the medical record text data, match the disease medical records with higher relevance, and recommend medications according to the treatment methods therein.

[0102] The label sequence output by the coding model represents the category of each word in the medical record text. For example, if the medical record text data is "headache accompanied by nausea, liver qi stagnation", the output label sequence is: ["O","O","B-sym","O","I-sym","B-cause"]. Extract the key elements and their corresponding text content according to the label sequence, and use the predefined coding mapping table to map the key elements to the corresponding coding factors. According to the classification of etiology, disease location, disease nature, and disease trend, combine them in sequence to generate a complete traditional Chinese medicine code. For example, the symptom "headache" corresponds to the coding factor: 03, the etiology "liver qi stagnation" corresponds to the coding factor: 04. The disease nature "yin deficiency" corresponds to the coding factor: 02, and the disease trend "initial stage" corresponds to the coding factor: 01, generating the complete traditional Chinese medicine code of 04-02-03-01.

[0103] This application focuses on the refinement and quantification of the four major dimensions of "etiology, disease location, disease nature, and disease trend" in traditional Chinese medicine syndrome differentiation and treatment, providing more precise coding so that it can be directly applied in machine learning models, covering the main elements of traditional Chinese medicine syndrome differentiation and treatment, and being applicable to automated medical record analysis and diagnosis and treatment assistance. The 46-bit traditional Chinese medicine coding system combines traditional Chinese medicine theory and modern data requirements, and through the standardized and quantified description of "etiology, disease location, disease nature, and disease trend", realizes the automated analysis and diagnosis and treatment assistance of traditional Chinese medicine medical records. This system has obvious innovation and practicality in structured design, quantitative analysis, and data compatibility, providing strong support for the digital development of traditional Chinese medicine.

[0104] Step S3: Assign scores to the traditional Chinese medicine codes according to the 46-bit traditional Chinese medicine coding system to determine the key coding factors. As Figure 3 shown, it includes:

[0105] Step S31: Count the occurrence frequencies of each coding factor in the traditional Chinese medicine codes;

[0106] Step S32: Calculate the syndrome score according to the occurrence frequency of the coding factor and the corresponding weight, and the range of the syndrome score is from 0 to 9 points;

[0107] Step S33: Screen the key coding factors that meet the preset requirements according to the syndrome score.

[0108] Assign scores to the traditional Chinese medicine codes according to the traditional Chinese medicine coding system to quantitatively evaluate different types of key elements in the medical record text data, such as etiology, disease location, disease nature, and disease trend, so as to support diagnosis and treatment assistance and medical record analysis. By calculating the frequencies, weights of each coding factor in the traditional Chinese medicine codes and their importance in the diagnosis and treatment decision-making, the quantitative analysis of the medical record data is realized. The score assignment process is based on the coding factors in the four dimensions of etiology, disease location, disease nature, and disease trend, and each coding factor is scored item by item. The range of the syndrome score is usually set as a fixed interval, such as from 0 to 9 points. The syndrome score combines traditional Chinese medicine theory and statistical principles, calculates the initial score through the occurrence frequency of the coding factor, and adjusts the score in combination with the weight to ensure that more important codes account for a higher proportion in the total score.

[0109] For example, for a medical record text where the etiology is "liver qi stagnation" and the etiology coding factor is 04, assuming its frequency is relatively high and it plays a key role in the diagnosis and treatment of a specific disease, its score will be higher than other codes with low occurrence frequency or small influence. In addition, for some combinations such as the simultaneous occurrence of the etiology of "phlegm-dampness obstructing the middle energizer" and the disease location of "spleen and stomach", the combined score can also be calculated through correlation analysis to reflect the complexity and severity of the overall characteristics of the medical record. Finally, the score assignment results can be used for diagnosis and treatment assistance, medical record analysis, as well as disease trend research, matching relevant diagnosis and treatment suggestions, providing quantitative support for traditional Chinese medicine diagnosis and treatment.

[0110] Each disease has its distinct symptoms. The more common symptoms patients with the same disease present, the higher the correlation between the symptom and the disease. Each symptom has a relative corresponding relationship with this traditional Chinese medicine coding. The higher the frequency of a certain coding factor appears, the greater the possibility of the corresponding disease. Using the medium sample size T as the denominator and the occurrence times a as the numerator, determine its occurrence frequency by a / T * 100%. Then, according to the high or low occurrence frequency, correspond to a range of 0 - 9 to determine the syndrome score. A syndrome score above 6 can be regarded as a relatively crucial traditional Chinese medicine coding.

[0111] The present invention provides a method, device, and medium for in-position potential coding based on medical record text data, which transforms complex medical record text data into structured traditional Chinese medicine coding, and uses a 46-bit traditional Chinese medicine coding system to quantitatively score medical record data, providing an intuitive quantitative basis for diagnosis and treatment decisions. This method combines the automatic annotation ability of the coding model, significantly improving the efficiency and accuracy of medical record data processing, reducing the subjective deviation of human intervention. At the same time, through the scoring results, it accurately reflects the characteristics and correlations of the cause, location, nature, and trend of the disease, providing reliable technical support for traditional Chinese medicine diagnosis and treatment assistance and medical record big data analysis.

[0112] The present invention provides a computer device, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the above-mentioned method for in-position potential coding based on medical record text data.

[0113] The present invention also provides an electronic device, which may include: a memory storing executable program code;

[0114] a processor coupled to the memory;

[0115] a transceiver for communicating with other devices or communication networks to receive or send network messages;

[0116] a bus for connecting the memory, the processor, and the transceiver for internal communication.

[0117] The transceiver receives the messages transmitted on the network, passes them to the processor through the bus. The processor calls the executable program code stored in the memory through the bus for processing, and passes the processing results to the transceiver for sending through the bus, thereby implementing the method provided by the embodiments of the present application.

[0118] The embodiments of the present application also provide a non-transitory machine-readable storage medium, on which an executable program is stored. When the executable program is run by the processor, the processor is made to execute the processing method provided by the above-mentioned embodiments.

[0119] An embodiment of the present invention discloses a computer-readable storage medium that stores a computer program for electronic data exchange, wherein the computer program causes a computer to execute the described method.

[0120] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the described method.

[0121] The embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0122] Through the above specific description of the embodiments, those skilled in the art can clearly understand that each implementation mode can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium that can be used to carry or store data.

[0123] Finally, it should be noted that: The embodiments disclosed in the present invention are only the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A coding method for positional potential based on medical record text data, characterized in that, Including: Obtaining and preprocessing medical record text data; Inputting the preprocessed medical record text data into a coding model, performing syndrome location and tendency annotation on the medical record text data, and obtaining the traditional Chinese medicine coding of the medical record text data; The coding model is obtained by training with a training dataset, and the training dataset includes preprocessed medical record data and labels; Assigning scores to the traditional Chinese medicine coding according to the traditional Chinese medicine coding system to determine the key coding.

2. The in-position potential encoding method based on medical record text data according to claim 1, wherein The medical record text data includes traditional Chinese medicine medical records, symptom descriptions, syndrome type annotations or prescriptions; performing syndrome location and tendency annotation on the medical record text data to obtain the traditional Chinese medicine coding of the medical record text data, including: Performing syndrome type annotation on the etiology, disease location, disease nature, and disease tendency content in the preprocessed medical record text data to generate a label sequence; the syndrome type annotation includes traditional Chinese medicine information, diagnosis and treatment methods, used drugs, and traditional Chinese medicine symptoms; According to the traditional Chinese medicine coding system, mapping the coding sequence corresponding to the label sequence to generate the traditional Chinese medicine coding.

3. A positional potential encoding method based on medical record text data according to claim 1, characterized in that, The coding model includes a BERT sub-model, a bidirectional long short-term memory network, and a conditional random field; Inputting the preprocessed medical record text data into the coding model includes: Using the BERT sub-model to perform word segmentation and feature extraction on the medical record text data to obtain medical record text features; Performing context analysis and entity recognition on the medical record text features through a bidirectional long short-term memory network to obtain medical record text features including context relationships; Predicting the label sequence of the medical record text features through a conditional random field.

4. A method for encoding in-position potential based on medical record text data according to claim 3, characterized in that, The bidirectional long short-term memory network includes a forgetting gate, an input gate, and an output gate. The forgetting gate takes the medical record text features at the current time step, the input at the current time step, and the hidden state at the previous time step as inputs, and calculates the forgetting ratio control signal through the first formula; The input gate takes the medical record text features at the current time step and the hidden state at the previous time step as inputs, generates the control signal of the input gate and the candidate memory cell state through the second set of formulas, and the output of the input gate is used to update the memory cell state at the current time step; The output gate takes the input at the current time step, the memory cell state at the current time step, and the hidden state at the previous time step as inputs, generates the control signal of the output gate through the third formula, and combines the memory cell state at the current time step to generate the hidden state at the current time step as the output.

5. A positional potential encoding method based on medical record text data according to claim 4, characterized in that The first formula is: F t = σ(X t W xf + H t-1 W hf + b f ) where σ is the Sigmoid activation function, and X t is the input vector at the current time step t, and W xf is the weight matrix of the input vector X t , representing the weight from the input to the forget gate; H t-1 is the hidden state at the previous time step, representing the memory information at the previous time step; W hf is the weight matrix of the hidden state H t-1 , representing the weight from the hidden state to the forget gate; b f is the bias term of the forget gate, used to offset and adjust the result of the linear transformation; The second set of formulas Including: I t = σ(X t W xi + H t-1 W hi + b i ) Among them, I t is the output of the forget gate, which determines how much of the current input vector X t needs to be written into the cell state, and its value range is [0, 1]; σ is the Sigmoid activation function, X t is the input vector at the current time step t, W xi is the weight matrix of the input vector X t , H t-1 is the hidden state of the previous time step, representing the memory information of the previous time step; W hi is the weight matrix of the hidden state H t-1 in the input gate; b i is the bias term of the input gate; Among them, is the candidate cell state, representing the new candidate memory calculated according to the input at the current time step; tanh is the hyperbolic tangent activation function, which is used to map the linear result to the interval (-1, 1) to add non-linearity; X t is the input vector at the current time step t; W xc is the weight matrix of the input vector X t for the candidate cell state; H t-1 is the hidden state at the previous time step, representing the memory information at the previous time step; W hc is the weight matrix of the hidden state H t-1 for the candidate cell state; b c is the bias term of the candidate cell state; Among them, C t is the cell state at the current time step, combining the memory state of the previous time step and the new information at the current time step; F t is the output of the forget gate, C t-1 is the cell state of the previous time step, I t is the output of the forget gate, is the candidate cell state, and ⊙ is element-wise multiplication.

6. A positional potential encoding method based on medical record text data according to claim 3, characterized in that The traditional Chinese medicine coding includes an etiology coding factor, a disease location coding factor, a disease nature coding factor, and a disease tendency coding factor. The etiology coding factor includes wind, stirring wind, dryness, dampness, water, cold, summer heat, phlegm, fluid deficiency, toxin, food, insect, qi stagnation, blood stasis, collateral obstruction, and bleeding. The disease location coding factor includes the heart, liver, gallbladder, spleen, stomach, intestine, lung, kidney, bladder, uterus, thoroughfare and conception vessels, and seminal chamber. The disease nature coding factor includes heat, deficiency heat, blood heat, cold, accumulation, qi deficiency, blood deficiency, yin deficiency, yang deficiency, insecurity, solidity of the fu-organs, fluid deficiency, and toxin. The disease tendency coding factor includes restlessness of the heart, coma, yang hyperactivity, cough and asthma, and qi reversal.

7. A positional potential encoding method based on medical record text data according to claim 2, characterized in that, Assigning scores to the traditional Chinese medicine coding according to the 46-bit traditional Chinese medicine coding system to determine the key coding factors, including: Counting the occurrence frequencies of each coding factor in the traditional Chinese medicine coding; Calculate the syndrome score according to the occurrence frequency and corresponding weight of the coding factors, and the range of the syndrome score is from 0 to 9 points; Screen the key coding factors that meet the preset requirements according to the syndrome score.

8. A method for encoding in-position potential based on medical record text data according to claim 2, characterized in that The preprocessing includes: Perform text cleaning on the medical record text data, perform word segmentation processing, and form structured data.

9. A computer device, characterized in that, Including: A processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of a method for encoding due-position potential based on medical record text data according to any one of claims 1-8.

10. A computer storage medium, characterized in that, Stored thereon is a computer program, and when the computer program is executed by the processor, the steps of a method for encoding due-position potential based on medical record text data according to any one of claims 1-8 are implemented.