A method for accurate structured representation and semantic comparison of traditional Chinese medicine contextual information

By constructing a semantically enriched network of contextual dimensions, the accuracy problem in TCM contextual information processing is solved, efficient structuring and semantic comparison of TCM contextual information is achieved, and the information processing capability of TCM diagnosis and treatment is improved.

CN119943435BActive Publication Date: 2025-09-26SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510003480.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-09-26
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively process unstructured TCM contextual information in TCM literature and clinical records, resulting in inaccurate information processing and weak retrieval relevance. Traditional methods are unable to capture the multi-dimensional contextual information in TCM diagnosis and treatment.

Method used

A semantic enrichment network of the context dimension is constructed, including a thought chain structuring module, a context information encoding module, an intra-layer context information enrichment module, and an inter-layer information intersection module. The structured information of the TCM context text is extracted through a large language model, and multi-level semantic feature extraction and comparison are performed.

Benefits of technology

It improves the accuracy of structured extraction of TCM context information and the accuracy of semantic comparison, can better capture the semantic characteristics of intuitive symptoms and abstract syndromes in TCM contexts, and improves the accuracy of diagnostic references.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943435B_ABST
    Figure CN119943435B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for accurate structured representation and semantic comparison of traditional Chinese medicine context information, comprising constructing a context dimension semantic enrichment network; extracting key dimension information from the traditional Chinese medicine context text through a thinking chain structuring module and converting it into a structured context text, then extracting multi-level dimensional features through a context information encoding module, and then sequentially capturing intra-layer context information and inter-layer multi-dimensional information through an intra-layer context information enrichment module and an inter-layer information intersection module to obtain semantically rich context semantic features; training the context dimension semantic enrichment network and performing context semantic comparison; the present invention combines a thinking chain prompt tuning method to tap into the traditional Chinese medicine semantic understanding ability of a universal Chinese large language model, accurately extracts traditional Chinese medicine context information, and simultaneously integrates multi-level traditional Chinese medicine context dimension information to obtain semantically rich context semantic features, thereby improving the accuracy of traditional Chinese medicine context semantic comparison and providing valuable reference contexts for traditional Chinese medicine clinicians.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing and artificial intelligence technology, and in particular relates to a method for accurate structured representation and semantic comparison of traditional Chinese medicine context information. Background Art

[0002] Traditional Chinese Medicine (TCM) boasts a long history and has accumulated a vast body of literature documenting theoretical knowledge, clinical diagnostic methods, treatment experiences, and prescriptions. This material not only holds historical value but also offers significant insights for modern TCM practice. In modern TCM practice, physicians often seek insights from similar case histories, symptoms, and diagnoses to better guide diagnosis and treatment. Providing clinicians with similar contextual references through contextual semantic comparison is of great significance. This information comparison can improve diagnostic accuracy and efficiency, help unlock the potential value of TCM literature, and promote the application and development of TCM in modern medicine. However, a large amount of TCM contextual information, such as patient symptom descriptions, physician diagnostic notes, and treatment plans, is typically stored in natural language. These natural language representations are unstructured data and often lack strict formatting and standardization, posing significant challenges in processing. In TCM records, information about the patient's constitution, etiology, progression of the condition, and treatment options is highly dependent on contextual information. For example, a symptom may have different meanings depending on the time of day, climate, and patient's constitution. Traditional information processing methods primarily rely on keyword matching and shallow semantic analysis. Because contextual information is unstructured, traditional information processing technologies struggle to directly understand and process these complex semantics and associations, resulting in inaccurate matching results or weak retrieval relevance. Specifically, in Traditional Chinese Medicine (TCM), terminology is highly context-dependent; the same term can represent different meanings in different contexts. For example, the cause and treatment of a "fever syndrome" can vary depending on the patient's constitution, course of illness, or season. Traditional keyword matching methods struggle to identify these semantic differences, resulting in inaccurate matching results. Techniques that rely solely on shallow semantic analysis also struggle to capture the multidimensional contextual information in TCM diagnosis and treatment. Therefore, existing information processing methods face significant limitations when applied to a highly contextualized and semantically complex field like TCM. Accurate structured representation and semantic matching technologies for TCM contextual text in literature and clinical records are crucial.

[0003] The rapid development of natural language processing (NLP) technology, particularly the emergence of large language models (such as GPT and BERT), has significantly enhanced semantic understanding and processing capabilities. These models, pre-trained on massive amounts of data, are able to capture the rich semantic relationships in language and demonstrate powerful language processing capabilities. However, the application of large language models in Traditional Chinese Medicine (TCM) is still in its infancy. The complexity and context-dependence of TCM terminology increase the challenges of natural language processing in this field. Currently, some studies have attempted to fine-tune large language models using limited TCM data. However, due to the relatively limited amount of TCM literature and clinical data, such fine-tuning has been less than ideal, hindering generalization. On the other hand, large language models are trained on large amounts of data, which contain a significant amount of TCM corpus. This means that even without fine-tuning, large language models possess a certain level of TCM semantic understanding capabilities. The key challenge lies in effectively leveraging these models' capabilities to accurately structure the contextual information in TCM literature and clinical data, thereby obtaining precise semantic representations for comparison and reference across diagnostic contexts. In the context of Traditional Chinese Medicine (TCM), intuitive symptoms and abstract syndrome types are key elements of semantic comparison. Pre-trained large language models have a multi-layered structure, with lower layers capturing intuitive and superficial concepts and higher layers capturing deeper, more abstract patterns or information, each of which is strongly correlated with symptoms and syndrome types. This suggests that we can combine the characteristics of TCM contexts with the capabilities of large language models to develop new semantic structuring and comparison methods to address specific issues in TCM information processing. Summary of the Invention

[0004] In view of this, it is necessary to address the above technical issues. The present invention provides a new method for accurate structured representation and semantic comparison of TCM context information, comprising the following steps:

[0005] Step 1: Construct a context-dimensional semantic enrichment network that can extract structural information and contextual semantic features of TCM context texts, including a thought chain structuring module, a context information encoding module, an intra-layer context information enrichment module, and an inter-layer information intersection module;

[0006] Step 2: Using TCM context text as network input, the thought chain structuring module extracts key dimensional information and converts it into structured context text. The context information encoding module then extracts multi-level dimensional features. The intra-layer context information enrichment module and the inter-layer information intersection module then capture the intra-layer context information and inter-layer multi-dimensional information to obtain semantically rich context semantic features.

[0007] Step 3, using the situational semantic features to train the situational dimension semantic enrichment network;

[0008] Step 4: Use the trained context dimension semantic enrichment network to perform context semantic comparison.

[0009] Furthermore, the thought chain structuring module includes a thought stimulation mechanism and a structured extraction mechanism. The specific steps of extracting key dimension information and converting it into structured context text are as follows:

[0010] Step 20101: Input the TCM context text into a large language model of general Chinese based on the precise structured extraction format requirements, analyze it from multiple preset TCM context dimensions, and provide analysis basis for each dimension to obtain a context element extraction strategy;

[0011] Step 20102: Combining the precise structured extraction format requirements with the context element extraction ideas, construct context element extraction prompts, and obtain the structured context text of the TCM context text through the general Chinese language model.

[0012] Furthermore, the precise structured extraction format of TCM information includes multiple TCM key contextual dimensions developed by TCM experts through grounded theory, including age, symptoms / signs, disease name and syndrome type, etiology and pathogenesis, and treatment principles. The specific format is:

[0013] ;

[0014] Represents 27 key contextual dimensions of TCM.

[0015] Furthermore, the context information encoding module is composed of a multi-layer large language model encoder and a semantic enrichment layer selection mechanism. The specific steps of extracting multi-level dimensional features are as follows:

[0016] In step 2201, all structured contextual texts in the database are encoded by a multi-layer large language model encoder to obtain hidden layer features output by each layer of the encoder;

[0017] Step 2202: Calculate the recall rate of each hidden layer feature in all hidden layer features of the database in the "symptom" dimension search and the "syndrome type" dimension search;

[0018] Step 20203, select the hidden layer feature with the highest recall rate of the "symptom" dimension as the shallow symptom hidden layer feature, and select the hidden layer feature with the highest recall rate of the "syndrome type" dimension as the deep syndrome type hidden layer feature. The shallow symptom hidden layer feature and the deep syndrome type hidden layer feature constitute the multi-level dimensional feature.

[0019] Furthermore, the intra-layer context information enrichment module includes a shallow semantic feedforward network, a deep semantic feedforward network, and a self-attention mechanism network. The specific steps of capturing the intra-layer context information are as follows:

[0020] Step 20301: First, a number of TCM contextual text samples are collected from the database and a batch of shallow symptom hidden layer features and a batch of deep syndrome type hidden layer features are generated through the context information encoding module. The batch of shallow symptom hidden layer features and the batch of deep syndrome type hidden layer features are respectively subjected to dimensionality compression through the shallow semantic feedforward network and the deep semantic feedforward network;

[0021] Step 20302: Sine position coding is added to the compressed batch of shallow symptom hidden layer features and batch of deep syndrome type hidden layer features respectively;

[0022] Step 20303: The encoded batch shallow symptom hidden layer features are processed through a self-attention mechanism network to capture context information, and then through a Gelu nonlinear activation layer and a layer normalization layer to obtain batch full-text shallow symptom features;

[0023] Step 20304: The encoded batch of deep certificate type hidden layer features are passed through the self-attention mechanism network to capture context information to obtain batch full-text deep certificate type features.

[0024] The shallow symptom features of the batch full text and the deep syndrome type features of the batch full text constitute the context information within the layer.

[0025] Furthermore, the inter-layer information intersection module includes a cross-attention mechanism network and a feedforward network. The steps of obtaining semantically rich contextual semantic features are as follows:

[0026] Step 20401: The batch of shallow symptom features of the full text are used as queries, and the batch of deep syndrome type features of the full text are used as keys and values. The deep and shallow semantic information is aggregated through a cross-attention mechanism network and layer normalization to obtain a batch of multi-level contextual semantic representations, where the query, key, and value are the QKV weights in the cross-attention mechanism network.

[0027] In step 20402, the batch multi-level context semantic representation is mapped into a two-dimensional vector in the context semantic representation space through the linear layer in the feedforward network, and then passes through the Gelu activation layer, the layer normalization layer and the mean pooling layer of the text length dimension in sequence to obtain the one-dimensional semantically rich context semantic feature.

[0028] Furthermore, in step 3, the contextual semantic features are used to train the contextual dimension semantic enrichment network, specifically using batch contextual semantic features to calculate model loss and optimize the intra-layer context information enrichment module and the inter-layer information intersection module, and the specific steps are as follows:

[0029] Step 301: Construct a contrast loss function, and the calculation formula is:

[0030]

[0031] in, is the contrast loss function value; is a flag variable. When two situational semantic features are of the same type, is 0, for different classes is 1; Indicates taking the maximum value; is the threshold; , are two situational semantic features; Represents the similarity measure between two contextual semantic features;

[0032] When it is Euclidean distance:

[0033] ;

[0034] When it is cosine similarity:

[0035] ;

[0036] Step 302 , calculating the contrast loss function values ​​of all binary combinations in the batch context semantic features and dividing them by the number of combinations to obtain the total contrast loss of the batch;

[0037] Step 303: freeze all parameters of the large language model in the thought chain structuring module and the context information encoding module, and use the batch total contrast loss to optimize the intra-layer context information enrichment module and the inter-layer information intersection module.

[0038] Furthermore, we use the trained context dimension semantic enrichment network to perform context semantic comparison. The specific steps are as follows:

[0039] Step 401: All TCM context texts in the database are input into the context dimension semantic enrichment network to extract the database context semantic features and save them;

[0040] Step 402: query the TCM context text and input it into the context dimension semantic enrichment network to obtain query context semantic features;

[0041] Step 403 : Calculate the cosine similarity between the query context semantic features and the database context semantic features, and select the top K results in descending order of similarity as output.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] The present invention combines the thought chain tuning method to design a reasonable TCM context text extraction prompt to fully tap the TCM semantic understanding ability of the pre-trained large language model and improve the accuracy of structured extraction of TCM context information;

[0044] The present invention simultaneously integrates different levels of TCM context semantics to learn better context semantic features, designs a context information encoding module to extract multi-level dimensional features of the context from two dimensions of shallow symptom semantics and deep syndrome semantics; the proposed intra-layer context information enrichment module adopts multi-level dimensional features to capture the full-text information of the context; the proposed inter-layer information intersection module adopts the fusion of multi-level semantics to obtain semantically rich context semantic features, thereby improving the accuracy of TCM context semantic comparison. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A schematic flow chart of an implementation method of the present invention is shown;

[0046] Figure 2 A schematic diagram showing the structure of each module working in an embodiment of the present invention is shown; DETAILED DESCRIPTION

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0048] For the purpose of reference and clarity, the technical terms, abbreviations or acronyms used below are summarized and explained as follows:

[0049] LLM: Large Language Model, large language model.

[0050] Prompt: Input the prompt word of the language model to guide the output of the language model.

[0051] GeLU: non-linear activation function.

[0052] Epoch: A complete training of the model using all the data in the training set is called "one generation of training".

[0053] : self-Attention, self-attention mechanism network.

[0054] : cross-Attention, cross attention mechanism network.

[0055] Recall@K: Recall@K is the ratio of the number of samples correctly identified as positive by the model to the total number of positive samples in the database among the top K search results from high to low.

[0056] The present invention discloses a method for accurately structured representation and semantic comparison of traditional Chinese medicine context information to solve many problems existing in the prior art.

[0057] Figure 1 A flow chart of an embodiment of the present invention is shown. A method for accurately structured representation and semantic comparison of TCM context information includes the following steps:

[0058] Step 1: Construct a context-dimensional semantic enrichment network that can extract structural information and contextual semantic features of TCM context texts, including a thought chain structuring module, a context information encoding module, an intra-layer context information enrichment module, and an inter-layer information intersection module;

[0059] Step 2: Using TCM context text as network input, the thought chain structuring module extracts key dimensional information and converts it into structured context text. The context information encoding module then extracts multi-level dimensional features. The intra-layer context information enrichment module and the inter-layer information intersection module then capture the intra-layer context information and inter-layer multi-dimensional information to obtain semantically rich context semantic features.

[0060] Step 3, using the situational semantic features to train the situational dimension semantic enrichment network;

[0061] Step 4: Use the trained context dimension semantic enrichment network to perform context semantic comparison;

[0062] Specifically, if Figure 2 As shown, in step 2, first, all TCM context texts in the database are converted into structured context texts through the thought chain structuring module, and then the context information encoding module extracts and constructs batch multi-level dimensional features from the perspectives of shallow symptom semantics and deep syndrome semantics, and then the batch multi-level dimensional features are input into the intra-layer context information enrichment module and the inter-layer information intersection module to obtain batch semantically enriched context semantic features.

[0063] Furthermore, the specific steps of step 2 in this example are as follows:

[0064] Step 201: Input each TCM context text in the database into the thought chain structuring module to extract key dimension information and convert it into a structured context text. The thought chain structuring module includes a thought stimulation mechanism and a structured extraction mechanism. The specific steps are as follows:

[0065] Step 20101: In accordance with the requirements of the precise structured extraction format, the TCM context text is input into the general Chinese language model and analyzed from multiple preset TCM context dimensions. The analysis basis for each dimension is provided to obtain the context element extraction ideas. In other words, in the thinking stimulation stage, a prompt is designed to allow the large language model to output the extraction ideas of the context elements. The specific steps are as follows:

[0066] Step 2010101: Build the LLM boot prompt. The specific content is as follows:

[0067] Prompt = "You are a TCM expert. Please extract the elements listed below based on the following TCM case. If some elements do not exist in the case, the corresponding content will be left blank. The elements are as follows:"

[0068] Step 2010102: Construct the LLM extraction format prompt. The specific content is as follows:

[0069]

[0070] in, It is a set of 27 key contextual dimensions of Traditional Chinese Medicine evaluated by TCM experts. It represents 27 key contextual dimensions of traditional Chinese medicine, including gender, age, occupation, marriage and childbearing, birth time, population characteristics, menstrual history, symptoms / signs, time and space, luck factors, space, climate and phenology, group knowledge base, doctor's subjective conjecture, disease name and syndrome type, etiology and pathogenesis, auxiliary diagnostic tools, requirements for doctors, treatment principles, internal medication, external treatment, misdiagnosis, conditioning methods, compliance, symptoms, etiology and pathogenesis, examinations and tests, and other comments (notes). In other words, the precise structured extraction format of traditional Chinese medicine context includes multiple key contextual dimensions of traditional Chinese medicine formulated by traditional Chinese medicine experts through grounded theory, including age, symptoms / signs, disease name and syndrome type, etiology and pathogenesis, treatment principles, etc. The specific format is:

[0071] ,

[0072] ;

[0073] Step 2010103, build the LLM example prompt, the specific content is as follows:

[0074] "The following is an example, original text:" + example original text + example extraction results;

[0075] Example text = "Acute exacerbation of chronic bronchitis, with pathogenic heat accumulating in the lungs, resulting in phlegm-heat syndrome. Treatment is to clear heat, promote lung function, resolve phlegm, and calm asthma. Name: XXX, male, 53 years old. Medical record number: XXXX. Initial visit: February 23, 2010. Recurrent cough and asthma attacks for 5 years, worsening with fever for 1 week. For the past 5 years, the patient has frequently experienced coughing, sputum, and wheezing, especially in winter. Despite seeking treatment from various sources, there has been no significant effect. The patient was diagnosed with chronic bronchitis. A week ago, due to a sudden drop in temperature, he failed to add more clothes in time and caught a cold, which caused a relapse. Initial symptoms included chills, fever, headache, cough, sputum, and sore throat. Self-medication with cold medicine and antibiotics was ineffective, and the condition gradually worsened, with the chills disappearing and the fever increasing. Current: Fever, body temperature 38 .8℃, no aversion to cold, paroxysmal cough, choking cough, abundant sputum, thick yellow sputum, wheezing, especially at night, unable to lie flat in severe cases, chest tightness, sweating, thirst, dry and hard stool, red tongue, yellow fur, slippery and rapid pulse. Chest X-ray showed: thickened and disordered markings on both lungs. Chinese medicine diagnosis: wind-heat; the syndrome is pathogenic heat accumulating in the lungs, causing phlegm and heat, and the lungs failing to descend. Western medicine diagnosis: acute exacerbation of chronic bronchitis. Treatment is to clear heat and promote lung function, resolve phlegm and relieve wheezing, and the prescription is Ma Xing Shi Gan Tang with modifications. Prescription: roasted ephedra 6g, almond 10g, gypsum 30g, raw licorice 6g, houttuynia cordata 20g, golden buckwheat 20g, polygonum cuspidatum 20g, trichosanthes seed 12g, light scutellaria 12g, mulberry bark 15g, loquat leaf 12g, plantain 10g. 3 doses. "

[0076] Example Extraction Result = "

[0077] 1. Personal Information

[0078] 1. Gender: Male

[0079] 2. Age: 53

[0080] 3. Occupation:

[0081] 4. Marriage and childbearing:

[0082] 5. Date of Birth:

[0083] 6. Patient Characteristics: For the past five years, the patient had frequently experienced coughing, sputum production, and wheezing, particularly in winter. Despite seeking multiple treatments, they were ineffective and were diagnosed with chronic bronchitis. A week ago, a sudden drop in temperature caused a relapse due to a cold snap, which resulted in a fever, headache, cough, sputum production, and sore throat. Self-medication with cold medicine and antibiotics proved ineffective, leading to a gradual worsening of the condition, with the chills disappearing and the fever becoming severe.

[0084] 2. Symptoms / Signs

[0085] 1. History of menstruation:

[0086] 2. Symptoms / Signs: Recurrent cough and asthma for 5 years, worsening with fever for 1 week. Fever (38.8°C), no aversion to cold, paroxysmal cough, choking cough, copious yellow and thick sputum, wheezing, particularly at night, inability to lie flat in severe cases, chest tightness, sweating, thirst, dry and hard stools, red tongue with a yellow coating, and a slippery and rapid pulse.

[0087] 3. Space and Time

[0088] 1. Date: February 23, 2010

[0089] 2. Luck factor:

[0090] 3. Space:

[0091] 4. Climate and phenology:

[0092] 4. Decision Basis

[0093] 1. Group knowledge base:

[0094] 2. The medical supervisor speculates:

[0095] 5. Dialectical Process

[0096] 1. Disease name and syndrome: Acute exacerbation of chronic bronchitis, pathogenic heat accumulating in the lungs, resulting in phlegm-heat syndrome

[0097] 2. Etiology and pathogenesis:

[0098] 3. Auxiliary Diagnostic Tools: Chest X-ray shows thickened and disordered markings on both lungs. Traditional Chinese Medicine Diagnosis: Wind-heat; syndrome characterized by pathogenic heat accumulating in the lungs, generating phlegm and heat, and the lungs failing to descend and diffuse downward.

[0099] 6. Treatment

[0100] 1. Requirements for doctors:

[0101] 2. Treatment principles: Clear away heat, promote lung function, resolve phlegm and relieve asthma.

[0102] 3. Oral medication: Modified Ma Xing Shi Gan Tang. Prescription: Roasted Ephedra 6g, Apricot Kernel 10g, Gypsum 30g, Raw Licorice 6g, Houttuynia Cordata 20g, Golden Buckwheat 20g, Polygonum Cuspidati 20g, Trichosanthes Fructus 12g, Scutellaria Baicalensis 12g, Morus Albizia Bark 15g, Loquat Leaf 12g, Plantain 10g. 3 doses.

[0103] 4. External treatment:

[0104] 5. Guidance:

[0105] 6. Mistakes and misdiagnosis:

[0106] 7. Maintenance methods:

[0107] 8. Compliance:

[0108] 7. Efficacy Evaluation

[0109] 1. Symptoms: After taking the medicine, the fever gradually subsided and returned to normal (36.8°C). The headache and sore throat disappeared. However, the cough remained, with copious sputum, which turned white but sticky and difficult to cough up. There was shortness of breath, especially at night, chest tightness, dry mouth, red tongue with a greasy yellow coating, and a slippery and rapid pulse.

[0110] 2. Etiology and pathogenesis:

[0111] 3. Examination and examination: Lung heat has been relieved, but phlegm heat is still severe, and the lungs have failed to descend.

[0112] 4. His comments (notes):

[0113]

[0114] Step 2010104: Construct the LLM semantic understanding capability mining prompt. The specific contents are as follows:

[0115] ;

[0116] In step 2010105, the guidance prompt, the format extraction prompt, the example prompt, and the semantic understanding ability mining prompt are sequentially spliced ​​together to obtain the LLM thinking chain prompt.

[0117]

[0118] In step 2010106, construct the LLM input text in the format of [{"role": "system", "content": Thinking Chain Prompt},{"role": "user", "content": Context text to be structured}] and input it into the LLM to obtain the idea of ​​extracting context elements.

[0119] ,

[0120] ;

[0121] Step 2102: Combine the precise structured extraction format requirements and the context element extraction ideas to construct a context element extraction prompt. Use the universal Chinese large language model to obtain the structured context text of the TCM context text, that is, construct the context element extraction prompt input LLM to obtain the TCM context structured text. The specific steps are as follows:

[0122] Step 2010201, the guide prompt, extraction format prompt, example prompt, "The following is the extraction idea", situational element extraction idea, and "Please give the extraction result:" are sequentially spliced ​​to obtain the LLM situational element extraction prompt.

[0123] Step 2010202: Construct LLM input text in the format [{"role": "system", "content": context element extraction prompt},{"role": "user", "content": context text to be structured}] and input it into LLM to obtain 27 dimensions of TCM context structured text.

[0124]

[0125] Step 202: construct a context information encoding module, including a multi-layer large language model encoder and a semantic enrichment layer selection mechanism. The steps of extracting multi-level dimensional features are as follows:

[0126] In step 2201, all structured contextual texts in the database are encoded by a multi-layer large language model encoder to obtain the hidden layer features output by each layer of the encoder; specifically:

[0127] All structured context texts in the database are constructed as input in the format [{"role": "user", "content": TCM context structured text}] and encoded by the LLM encoder to obtain the hidden layer features of each layer output of the LLM encoder. ,in h is the number of LLM encoder layers, The dimension is , is the length of the text, d Output dimension of the LLM hidden layer;

[0128] In this example, we limit the maximum text length to 2048. >2048 samples are removed, for text length 2048 samples, we have each hidden layer feature Padding is done on the first dimension, adding a size to the length of the text (2048- )× d 0 tensor, so that its dimension changes from × d becomes 2048× d , and generate a mask with a corresponding dimension of 2048×1. The value of the first element is 1, and the The value of the element is 0, so that the attention layer in the subsequent steps only considers the previous Valid values;

[0129] Step 2202: Calculate the Recall@1 of each hidden layer feature in the "symptom" dimension retrieval and "syndrome type" dimension retrieval among all hidden layer features in the database.

[0130] Specifically, all TCM context texts in the database are classified into "symptoms" and "syndrome types", and the category labels of "symptoms" and "syndrome types" of each TCM context text are obtained. For a TCM context text, the cosine similarity between its corresponding hidden layer features and all hidden layer features in the database is calculated, and the one with the highest similarity is selected as the retrieval result. If the category labels of the retrieval results are the same, the retrieval is correct.

[0131] For each hidden layer, calculate the average retrieval recall rate Recall@1 of the two category labels of "symptoms" and "syndrome type" of the database hidden layer features, and use them as the "symptom" dimension Recall@1 and the "syndrome type" dimension Recall@1 respectively;

[0132] Step 20203: For each TCM context text, the hidden layer feature with the highest Recall@1 in the “symptom” dimension is selected as the shallow symptom hidden layer feature, which is recorded as ; Select the hidden layer feature with the highest Recall@1 in the “Syndrome Type” dimension as the deep syndrome type hidden layer feature, denoted as , the two together constitute the multi-level dimensional features 、 .

[0133] Step 203: construct an intra-layer context information enrichment module, including a shallow semantic feedforward network, a deep semantic feedforward network, and a self-attention mechanism network; the batch multi-level dimensional features are first passed through the intra-layer context information enrichment module to capture the intra-layer context information, and the intra-layer context information includes a batch of shallow symptom features of the full text and a batch of deep syndrome type features of the full text. The specific steps are as follows:

[0134] Step 20301, in this embodiment, first collect data from the entire database The corresponding shallow symptom hidden layer features and deep syndrome hidden layer features are obtained for each TCM context text sample to form a batch of multi-level dimensional features, which are recorded as 、 , whose dimensions are ×2048× In this embodiment, the number of samples of TCM context texts collected is is 8, that is, the dimension is 8×2048× ;

[0135] Then the batch of shallow symptom hidden layer features and batch of deep syndrome hidden layer features are respectively passed through the shallow semantic feedforward network and the deep semantic feedforward network, and the last dimension is transformed from First expand to 2 Then compress to ; Wherein, the shallow semantic feedforward network and the deep semantic feedforward network are both composed of two fully connected layers;

[0136] Step 20302: Sine position coding is added to the compressed batch of shallow symptom hidden layer features and batch of deep syndrome type hidden layer features respectively;

[0137]

[0138] in, is a shallow semantic feedforward network, is a deep semantic feedforward network, p is a sinusoidal position code, 、 is the batch multi-level dimensional feature before compression, 、 It is a batch of multi-level dimensional features after compression and encoding, with dimensions of ×2048× , batch multi-level dimensional features include batch shallow symptom hidden layer features and batch deep evidence type hidden layer features ;

[0139] In step 20303, the encoded batch of shallow symptom hidden layer features are passed through the self-attention mechanism network to capture context information, and then the batch of full-text shallow symptom features are obtained through the Gelu nonlinear activation layer and the layer normalization layer. The self-attention mechanism network is a bidirectional self-attention layer, and the specific mathematical expression is as follows:

[0140]

[0141] in, LN Representation layer normalization layer, represents a nonlinear activation layer, represents the self-attention mechanism network, represents the QKV weight in the self-attention mechanism network, is the batch shallow symptom hidden layer feature, It is the shallow symptom feature of the batch full text, and its dimensions are ×2048× , Represents a batch mask;

[0142] In step 20304, the encoded batch of deep certificate type hidden layer features are passed through the self-attention mechanism network to capture context information to obtain batch full-text deep certificate type features; wherein the self-attention mechanism network is a bidirectional self-attention layer, and the specific mathematical expression is as follows:

[0143]

[0144] in, is a self-attention mechanism network, represents the QKV weight in the self-attention mechanism network, It is the deep evidence type feature of batch full text, and its dimensions are ×2048× , is the batch deep evidence type hidden layer feature, Represents a batch mask;

[0145] Step 204: construct the inter-layer information intersection module, including a cross-attention mechanism network and a feedforward network; the batch multi-level dimensional features are then passed through the inter-layer information intersection module to capture the inter-layer multi-dimensional information and combine it with the contextual information within the layer to obtain semantically rich contextual semantic features. The specific steps are as follows:

[0146] In step 20401, the shallow symptom features of the batch full text are used as the query, and the deep syndrome type features of the batch full text are used as the key and value. The deep and shallow semantic information are aggregated through a cross-attention mechanism network and layer normalization to obtain a batch of multi-level context semantic representations. The cross-attention mechanism network is a bidirectional cross-attention layer. The specific mathematical expression is as follows;

[0147]

[0148] in, LN Representation layer normalization, is a cross-attention mechanism network, represents the QKV weight in the cross-attention mechanism network, It is a batch multi-level context semantic representation, and its dimension is ×2048× ;

[0149] Step 20402: batch multi-level context semantic representation After being mapped into a two-dimensional vector in the context semantic representation space by the linear layer in the feedforward network, each context semantic representation becomes 2048×1024 dimensions, and then passes through the Gelu activation layer and the layer normalization layer in sequence. LN The mean pooling of the text length dimension is used to obtain the semantically rich contextual semantic features of the batch one dimension. , the specific mathematical expression is as follows:

[0150]

[0151] in, FFNIt is the feedforward network, which consists of two fully connected layers, avepool is the mean pooling layer, and Gelu is the activation layer. LN is the layer normalization layer, is the semantically rich contextual semantic features of the batch, with the dimension ×1024, that is A 1024-dimensional context semantic feature tensor.

[0152] Specifically, in step 3, the contextual semantic features are used to train the contextual dimension semantic enrichment network, specifically using batch contextual semantic features to calculate the model loss and optimize the intra-layer context information enrichment module and the inter-layer information intersection module. The specific steps are as follows:

[0153] Step 301: Construct a contrast loss function, and the calculation formula is:

[0154]

[0155] in, is the contrast loss function value; is a flag variable. When two situational semantic features are of the same type, is 0, for different classes is 1; Indicates taking the maximum value; is the threshold, which is set to 0.5 in this example; , are two situational semantic features; Represents the similarity between two contextual semantic features,

[0156] In this example, Euclidean distance can be used to measure:

[0157] ;

[0158] In this example, cosine similarity can also be used for measurement, specifically:

[0159] ;

[0160] Step 302, calculate The contrast loss function value of all binary combinations in the context semantic features is divided by the number of combinations to obtain the total contrast loss of the batch, that is, the contrast loss is calculated by traversing all binary combinations in the batch. The final mathematical expression of the total contrast loss of the batch is as follows:

[0161]

[0162] in, is the total contrast loss of the batch, is a batch context semantic feature, is the number of batch samples, in this example middle; are two situational semantic features, is the contrast loss function value;

[0163] Step 303: freeze all parameters of the large language model in the thought chain structuring module and the context information encoding module, and use the batch total contrast loss to optimize the intra-layer context information enrichment module and the inter-layer information intersection module.

[0164] Step 4: Use the trained context dimension semantic enrichment network to perform context semantic comparison. The specific steps are as follows:

[0165] Step 401: All TCM context texts in the database are converted into structured context texts through the thought chain structuring module, and input into the context dimension semantic enrichment network to extract the database context semantic features and save them;

[0166] Step 402: The query text of TCM context is structured by the thought chain structuring module and input into the context dimension semantic enrichment network to obtain the query context semantic features;

[0167] Step 403 : Calculate the cosine similarity between the query context semantic features and the database context semantic features, and select the top K results in descending order of similarity as output to obtain a comparison result.

[0168] In this invention, the maximum length of the text is 2048, and the number of TCM context texts collected is That is, the batch size is 8, LLM uses GLM4 and freezes all its parameters, the initial learning rate is 0.00005, 20 epochs are trained on 1 RTX4090 GPU, and the network is optimized using the Adam optimizer.

[0169] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0170] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0171] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0172] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for accurate structured representation and semantic comparison of TCM context information, characterized by: The following steps are involved: Step 1: Construct a context-dimensional semantic enrichment network that can extract structural information and contextual semantic features of TCM context texts, including a thought chain structuring module, a context information encoding module, an intra-layer context information enrichment module, and an inter-layer information intersection module; Step 2: Using TCM context text as network input, the thought chain structuring module extracts key dimensional information and converts it into structured context text. The context information encoding module then extracts multi-level dimensional features. The intra-layer context information enrichment module and the inter-layer information intersection module then capture the intra-layer context information and inter-layer multi-dimensional information to obtain semantically rich context semantic features. Step 3, using the situational semantic features to train the situational dimension semantic enrichment network; Step 4: Use the trained context dimension semantic enrichment network to perform context semantic comparison; The context information encoding module consists of a multi-layer large language model encoder and a semantic enrichment layer selection mechanism. The specific steps of extracting multi-level dimensional features are as follows: Step 2201: All structured contextual texts in the database are encoded by a multi-layer large language model encoder to obtain hidden layer features output by each layer of the encoder; Step 2202: Calculate the recall rate of each hidden layer feature in all hidden layer features in the database for retrieval in the "symptom" dimension and the "syndrome type" dimension. Step 2203: Select the hidden layer features with the highest recall rate in the "symptom" dimension as the shallow symptom hidden layer features, and select the hidden layer features with the highest recall rate in the "syndrome type" dimension as the deep syndrome type hidden layer features. The shallow symptom hidden layer features and the deep syndrome type hidden layer features constitute the multi-level dimensional features. The intra-layer context information enrichment module includes a shallow semantic feedforward network, a deep semantic feedforward network, and a self-attention mechanism network. The specific steps of capturing the intra-layer context information are as follows: Step 20301: First, a number of TCM contextual text samples are collected from the database and a batch of shallow symptom hidden layer features and a batch of deep syndrome type hidden layer features are generated through the context information encoding module. The batch of shallow symptom hidden layer features and the batch of deep syndrome type hidden layer features are respectively subjected to dimensionality compression through the shallow semantic feedforward network and the deep semantic feedforward network; Step 20302: Sine position coding is added to the compressed batch of shallow symptom hidden layer features and batch of deep syndrome type hidden layer features respectively; Step 20303: The encoded batch shallow symptom hidden layer features are processed through a self-attention mechanism network to capture context information, and then through a Gelu nonlinear activation layer and a layer normalization layer to obtain batch full-text shallow symptom features; Step 20304: The encoded batch of deep syndrome type hidden layer features are passed through a self-attention mechanism network to capture context information to obtain batch full-text deep syndrome type features; The shallow symptom features of the batch full text and the deep syndrome type features of the batch full text constitute the context information within the layer.

2. A method for accurate structured representation and semantic comparison of TCM context information according to claim 1, characterized in that: The thought chain structuring module includes a thought stimulation mechanism and a structured extraction mechanism. The specific steps of extracting key dimension information and converting it into structured context text are as follows: Step 20101: Input the TCM context text into a large language model of general Chinese based on the precise structured extraction format requirements, analyze it from multiple preset TCM context dimensions, and provide analysis basis for each dimension to obtain a context element extraction strategy; Step 20102: Combining the precise structured extraction format requirements with the context element extraction ideas, construct context element extraction prompts, and obtain the structured context text of the TCM context text through the general Chinese language model.

3. The method for accurate structured representation and semantic comparison of TCM context information according to claim 2, characterized in that: The precise structured extraction format of TCM information includes multiple key contextual dimensions of TCM developed by TCM experts, including age, symptoms / signs, disease name and syndrome type, etiology and pathogenesis, and treatment principles. The specific format is: ; Represents 27 key contextual dimensions of TCM.

4. The method for accurate structured representation and semantic comparison of TCM context information according to claim 1, characterized in that: The inter-layer information intersection module includes a cross-attention mechanism network and a feedforward network. The steps of obtaining semantically rich contextual semantic features are as follows: Step 20401: The batch of shallow symptom features of the full text are used as queries, and the batch of deep syndrome type features of the full text are used as keys and values. The deep and shallow semantic information is aggregated through a cross-attention mechanism network and layer normalization to obtain a batch of multi-level contextual semantic representations, where the query, key, and value are the QKV weights in the cross-attention mechanism network. In step 20402, the batch multi-level context semantic representation is mapped into a two-dimensional vector in the context semantic representation space through the linear layer in the feedforward network, and then passes through the Gelu activation layer, the layer normalization layer and the mean pooling layer of the text length dimension in sequence to obtain the one-dimensional semantically rich context semantic feature.

5. A method for accurate structured representation and semantic comparison of TCM context information according to claim 4, characterized in that: In step 3, the contextual semantic features are used to train the contextual dimension semantic enrichment network. Specifically, the model loss is calculated using batch contextual semantic features and the intra-layer contextual information enrichment module and the inter-layer information intersection module are optimized. The specific steps are as follows: Step 301: Construct a contrast loss function, and the calculation formula is: ; in, is the contrast loss function value; is a flag variable. When two situational semantic features are of the same type, is 0, for different classes is 1; Indicates taking the maximum value; is the threshold; , are two situational semantic features; Represents the similarity measure between two contextual semantic features; When it is Euclidean distance: ; When it is cosine similarity: ; Step 302 , calculating the contrast loss function values ​​of all binary combinations in the batch context semantic features and dividing them by the number of combinations to obtain the total contrast loss of the batch; Step 303: freeze all parameters of the large language model in the thought chain structuring module and the context information encoding module, and use the batch total contrast loss to optimize the intra-layer context information enrichment module and the inter-layer information intersection module.

6. The method for accurate structured representation and semantic comparison of TCM context information according to claim 1, characterized in that: Use the trained context dimension semantic enrichment network to perform context semantic comparison. The specific steps are as follows: Step 401: All TCM context texts in the database are input into the context dimension semantic enrichment network to extract the database context semantic features and save them; Step 402: query the TCM context text and input it into the context dimension semantic enrichment network to obtain query context semantic features; Step 403 : Calculate the cosine similarity between the query context semantic features and the database context semantic features, and select the top K results in descending order of similarity as output.

Citation Information

Patent Citations

  • Text classification method based on multi-layer feature weighted fusion of mask and causal language model

    CN118332116A

  • Dialectical method for traditional Chinese medicine case text by using natural language model and fuzzy logic

    CN119049627A