Traditional Chinese medicine case text-oriented hierarchical vector processing method and device, equipment and storage medium

Through the hierarchical vector processing method and structured ontology design, combined with the RoBERTa model and graph attention network, the shortcomings of vector representation of traditional Chinese medicine medical texts are solved, efficient and accurate semantic representation is achieved, and the information utilization rate of medical texts is improved.

CN120218013APending Publication Date: 2025-06-27UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510246528.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively deal with traditional Chinese medicine medical texts, especially in the analysis of semantic composite structures and implicit knowledge coverage, which leads to the inaccurate vector representation of traditional Chinese medicine medical texts.

Method used

The hierarchical vector processing method is used to determine the structured ontology design of medical cases based on traditional Chinese medicine theory. Through element-level, field-level and multi-field-level representation methods, combined with the RoBERTa model, Self-Attention network and graph attention network, the semantic representation of traditional Chinese medicine case text is obtained.

Benefits of technology

It realizes efficient, accurate and flexible vector representation of traditional Chinese medicine medical case text, fully captures the complex semantics and implicit knowledge in medical case text, and improves the information utilization rate of subsequent analysis and modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218013A_ABST
    Figure CN120218013A_ABST
Patent Text Reader

Abstract

The invention provides a traditional Chinese medicine case text-oriented hierarchical vector representation method and device, equipment and a storage medium. The method comprises the following steps: determining a structured ontology design of a traditional Chinese medicine case; acquiring professional medical case data in the field of traditional Chinese medicine; preprocessing the professional medical case data; performing structured information extraction on the traditional Chinese medicine case to obtain element-level, field-level and multi-field-level various types of data; constructing an element-level vector representation model based on the RoBERTa model after field enhancement pre-training, and obtaining element-level data semantic representation; for the sequence data, obtaining semantic representation based on an element-level vector representation model; for the element set data, semantics among elements are aggregated based on a Self-Attention network, and semantic representations of the elements are obtained; initializing node features of the multi-field graph; and obtaining vector representation of the traditional Chinese medicine case text based on the graph attention network. According to the scheme, the structural characteristics of the traditional Chinese medicine case text are fully considered, and refined information language representation of the traditional Chinese medicine case text can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer data processing and natural language processing, and focuses on the research of vector representation methods for the complex text structure of traditional Chinese medicine medical records. Based on the observations of its high knowledge density, diverse element types, and complex text structure, the present invention performs vector representation of traditional Chinese medicine medical record texts from the element level, field level to multi-field level, aiming to solve the challenges faced by traditional text representation methods in processing traditional Chinese medicine medical record texts, as well as the deficiencies of existing vector representation methods in traditional Chinese medicine medical record texts. Background Art

[0002] In recent years, with the rapid development of deep learning technology, the vector representation method of texts has been a core issue. However, most of the existing text vector representation methods are designed for general texts and perform poorly in representing domain-specific words and compound structures of professional words, resulting in limited processing effects for traditional Chinese medicine medical record texts. The existing text vector representation methods include the following types:

[0003] 1. Traditional bag-of-words model and TF-IDF: These methods are based on word frequency statistics and ignore the semantic relationships and context information between words, and perform poorly in dealing with the subtle differences of synonyms, near-synonyms, and professional terms that abound in traditional Chinese medicine medical records.

[0004] 2. Word2Vec and GloVe: Although these pre-trained word vector models can capture certain semantic information, they are trained based on large-scale general corpora and have insufficient coverage of professional vocabulary and expression patterns unique to the traditional Chinese medicine field. Professional terms, formula names, symptom descriptions, etc. in traditional Chinese medicine medical records are often not within their training scope, resulting in inaccurate vector representation.

[0005] 3. BERT and domain-specific language models: Although they effectively improve the ability of text understanding and generation through self-attention mechanisms, and existing research has attempted to construct language models in the traditional Chinese medicine field, traditional Chinese medicine medical records often contain complex combinations of medical concepts and logical relationships, such as semantic associations between multiple symptom words and the connections between symptoms and syndrome elements. Existing methods often have difficulty effectively capturing the deep semantics of these compound structures, resulting in information loss or misunderstanding in subsequent analysis and modeling.

[0006] Therefore, how to effectively convert traditional Chinese medicine medical record texts into vector forms for further analysis and modeling has become an important research direction. Summary of the Invention

[0007] In view of this, in order to solve the technical problem that the existing technology lacks semantic composite structure analysis and does not cover various types of high-density implicit Chinese medicine knowledge involved in medical case texts, resulting in poor vector representation of Chinese medicine medical case texts, the embodiments of the present invention provide a hierarchical vector processing solution for Chinese medicine medical case texts.

[0008] Specifically, the present invention provides the following technical solutions:

[0009] On the one hand, the present invention provides a hierarchical vector processing method for Chinese medicine medical case texts. This method is implemented by a hierarchical vector representation model device, and the method includes:

[0010] Based on Chinese medicine theory, determine the structured ontology design of Chinese medicine medical cases;

[0011] Obtain professional medical case data in the field of Chinese medicine; preprocess the professional medical case data; perform structured information extraction on Chinese medicine medical cases based on the prompt engineering technology of large models, and obtain various types of data at the element level, field level, and multi-field level;

[0012] Construct an element-level vector representation model based on the domain-enhanced pre-trained RoBERTa model to obtain the semantic representation of element-level data;

[0013] Field-level data is divided into sequence type and element set type data. Obtain the semantic representation of sequence type data based on the element-level vector representation model; aggregate the semantics between elements based on the Self-Attention network to obtain the semantic representation of element set data;

[0014] The acquisition of the semantic representation of multi-field level data includes the initialization vector encoding representation of element set, sequence, and numerical type data, and the weighted summation of the vector encodings of various types of data based on the graph attention network to represent the semantic representation in Chinese medicine medical case texts.

[0015] Among them, the structured ontology design of Chinese medicine medical cases includes content hierarchy design, data type hierarchy design, and granularity hierarchy design;

[0016] The content hierarchy design of Chinese medicine medical cases includes patient information, clinical information, diagnostic information, treatment information, and other information;

[0017] The data type hierarchy design of Chinese medicine medical cases includes numerical type, sequence type, and element set type;

[0018] The numerical type includes age, date, etc.;

[0019] The sequence type includes patient's chief complaint, current medical history, etc.;

[0020] The element set type includes current symptoms, tongue image sequence, pulse condition sequence, etc.;

[0021] The granularity level design of traditional Chinese medicine (TCM) medical records includes the element level, the field level, and the multi-field level;

[0022] The element level is the entity words in the TCM medical record text, including symptoms, tongue manifestations, pulse conditions, syndromes, syndrome elements, traditional Chinese medicines, etc.;

[0023] The field level is the constituent elements of the medical record, including sequence type data and element set type data;

[0024] The multi-field level is the aggregation of multiple field-level data;

[0025] Among them, the professional medical record data in the field of traditional Chinese medicine is high-quality medical record text data obtained through laboratory accumulation, vertical medical online web crawling, and text data of professional medical books in the medical field; the structural types of medical record data in the field of traditional Chinese medicine include structured data, semi-structured data, and unstructured data;

[0026] Optionally, the preprocessing of the professional medical record data to obtain structured medical record text data in the field of traditional Chinese medicine includes:

[0027] Classifying the professional medical record data based on the data format to obtain XML data, HTML data, PDF data, first TXT data, and standardized first medical record structured data;

[0028] Parsing the XML data and the HTML data to obtain second TXT data;

[0029] Performing OCR text recognition on the PDF data to obtain recognized text data; correcting the recognized text data to obtain third TXT data;

[0030] Cleaning the first TXT data, second TXT data, and third TXT data to obtain fourth TXT data;

[0031] Based on the medical record structured ontology at the content level, designing a theme recognition prompt word template, and the theme scope includes but is not limited to patient information, clinical information, diagnosis information, treatment information, doctor information, and remarks;

[0032] Based on the medical record structured ontology at the granularity level, designing an information extraction prompt word template for each theme content;

[0033] Based on the large model and combined with the above-mentioned prompt word template design, performing information extraction on the fourth TXT data to obtain element-level, field-level, and multi-field-level text data, and merging them with the first medical record structured data to obtain structured medical record data in the field of traditional Chinese medicine;

[0034] Among them, the element-level vector representation model is the RoBERTa model base, which is trained through multiple stages, including enhanced pre-training and contrastive learning in the field of traditional Chinese medicine.

[0035] The RoBERTa model is used to extract the unified vector representation of elements and sequence texts;

[0036] The enhanced pre-training in the field of traditional Chinese medicine is based on the corpus in the field of traditional Chinese medicine and is trained through the masked language modeling objective;

[0037] The contrastive learning introduces the Sentence-BERT siamese network architecture to learn the similarities and differences in vector representations between samples;

[0038] Among them, the field-level vector representation includes the vector representation of sequence data and the vector representation of element set data;

[0039] The vector representation of the sequence data obtains the semantic representation by using the element-level vector representation model;

[0040] The method for the vector representation of the element set data includes element data encoding and set data encoding;

[0041] The element data encoding obtains the semantic representation of each element by using the element-level vector representation model;

[0042] The set data encoding aggregates the semantic encoding representations of each element by using the Self-Attention network;

[0043] The semantic representation of the multi-field-level data includes the initial vector encoding representation of element sets, sequences, and numerical type data, and the weighted sum of the vector encodings of each type of data based on the graph attention network;

[0044] The initial vector encoding representation of each type of data includes the encoding representation of the element set type, the encoding representation of the sequence type, and the encoding representation of the numerical type;

[0045] The encoding representation of the sequence type is obtained by the element-level vector representation model;

[0046] The encoding representation of the element set type is obtained by the field-level encoding representation model;

[0047] The encoding representation of the numerical type is obtained by using One-Hot encoding after discretization processing;

[0048] The graph attention network encoding model characterizes the central node of the traditional Chinese medicine medical record by weighting the features of neighbor nodes through the Attention algorithm;

[0049] On the other hand, a hierarchical vector representation device for traditional Chinese medicine (TCM) medical record texts is provided. This device is applied to the vector representation method of TCM medical record texts and includes:

[0050] A medical record text ontology design module for constructing a structured ontology of medical record texts by combining professional knowledge of TCM theory;

[0051] A medical record text data acquisition module for acquiring professional medical record data in the field of TCM; preprocessing the professional medical record data; and extracting structured information from TCM medical records based on the prompt engineering technology of large models;

[0052] An element-level data semantic representation module for constructing an element-level vector representation model based on the domain-enhanced pre-trained RoBERTa model and obtaining element-level data semantic representations;

[0053] A field-level data semantic representation module. The field-level data is divided into sequence type and element set type data. The sequence type data semantic representation is obtained based on the element-level vector representation model; the semantic between elements is aggregated based on the Self-Attention network to obtain the element set data semantic representation;

[0054] A multi-field-level data semantic representation module for initializing vector encoding representations of various types of data and obtaining the semantic representation of TCM medical record texts by weighting based on the graph attention network.

[0055] Among them, the TCM medical record structured ontology design module includes content hierarchy design, data type hierarchy design, and granularity hierarchy design;

[0056] The content hierarchy design of TCM medical records includes patient information, clinical information, diagnostic information, treatment information, and other information;

[0057] The data type hierarchy design of TCM medical records includes numerical type, sequence type, and element set type;

[0058] The numerical type includes age, date, etc.;

[0059] The sequence type includes patient's chief complaint, current medical history, etc.;

[0060] The element set type includes current symptoms, tongue image sequence, pulse condition sequence, etc.;

[0061] The granularity hierarchy design of TCM medical records includes element level, field level, and multi-field level;

[0062] The element level is the entity words in TCM medical record texts, including symptoms, tongue images, pulse conditions, syndromes, syndrome elements, traditional Chinese medicines, etc.;

[0063] The field level is the constituent elements of medical records, including sequence type data and element set type data;

[0064] The multi-field level is the aggregation of multiple field-level data;

[0065] Among them, the professional medical record data in the field of traditional Chinese medicine is high-quality medical record text data obtained through laboratory accumulation, vertical medical online web crawling, and text data of professional medical books in the medical field; the structural types of the medical record data in the field of traditional Chinese medicine include structured data, semi-structured data, and unstructured data;

[0066] Optionally, the medical record data acquisition module is further configured to:

[0067] Classify the professional medical record data based on the data format to obtain XML data, HTML data, PDF data, first TXT data, and standardized first medical record structured data;

[0068] Parse the XML data and the HTML data to obtain second TXT data;

[0069] Perform OCR text recognition on the PDF data to obtain recognized text data; correct the recognized text data to obtain third TXT data;

[0070] Clean the first TXT data, second TXT data, and third TXT data to obtain fourth TXT data;

[0071] Design a topic recognition prompt word template based on the medical record structured ontology at the content level, and the topic scope includes but is not limited to patient information, clinical information, diagnosis information, treatment information, doctor information, and remarks;

[0072] Design an information extraction prompt word template for each topic content based on the medical record structured ontology at the granularity level;

[0073] Based on the large model and combined with the above-mentioned prompt word template design, perform information extraction on the fourth TXT data to obtain element-level, field-level, and multi-field level text data, and merge them with the first medical record structured data to obtain traditional Chinese medicine medical record structured data;

[0074] Among them, the element-level data semantic representation module is further configured to:

[0075] Use the RoBERTa model trained through multiple stages to extract the unified vector representation of elements and sequence texts, including enhanced pre-training and contrast learning in the field of traditional Chinese medicine.

[0076] The enhanced pre-training in the field of traditional Chinese medicine is based on the corpus in the field of traditional Chinese medicine and is trained through the masked language modeling objective;

[0077] The contrastive learning is to introduce the Sentence-BERT siamese network architecture to learn the similarities and differences in vector representations among samples;

[0078] Among them, the field-level data semantic representation module is further used for:

[0079] Adopt an element-level vector representation model to obtain the vector representation of sequence data;

[0080] The vector representation method of the element set data includes element data encoding and set data encoding;

[0081] The element data encoding uses an element-level vector representation model to obtain the semantic representation of each element;

[0082] The set data encoding uses a Self-Attention network to aggregate the semantic encoding representations of each element;

[0083] Among them, the multi-field-level data semantic representation module is further used for the vector representation of medical record texts or multi-type data, including multi-field-level data feature initialization and a graph attention network encoding model;

[0084] The multi-field-level data feature initialization includes element set type encoding representation, sequence type encoding representation, and numerical type encoding representation;

[0085] The sequence type encoding representation is obtained by an element-level vector representation model;

[0086] The element set type encoding representation is obtained by a field-level encoding representation model;

[0087] The numerical type encoding representation is obtained by using One-Hot encoding after discretization processing;

[0088] The graph attention network encoding model characterizes the central node of traditional Chinese medicine medical records by weighting the features of neighbor nodes through the Attention algorithm;

[0089] On the other hand, a hierarchical vector processing device for traditional Chinese medicine medical record texts is provided. The device includes: a processor; a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, any one of the methods in the above hierarchical traditional Chinese medicine medical record text vector processing method is implemented.

[0090] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above hierarchical traditional Chinese medicine medical record text vector processing method.

[0091] Compared with the prior art, the beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:

[0092] Through a hierarchical representation method from the element level, field level to multi-field level, this solution provides an efficient, accurate and flexible method for representing the text vectors of traditional Chinese medicine case records in natural language processing tasks. Based on traditional Chinese medicine theory, the structured ontology design of traditional Chinese medicine case records is determined. Professional case record data in the field of traditional Chinese medicine is obtained, preprocessed, and structured information extraction is performed on the traditional Chinese medicine case records based on the prompting engineering technology of the large model. For complex case record texts, where the element or sequence data is concerned, feature initialization is performed based on the element-level vector representation model. For element set data, the semantics between elements are aggregated based on the Self-Attention network to obtain its feature initialization. Finally, the fused vector representation of the entire traditional Chinese medicine case record text is obtained based on the graph attention network. The present invention is a refined information semantic representation method that fully considers the structural characteristics of traditional Chinese medicine case record texts and combines the language model embedding representation technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0094] Figure 1 It is a flowchart of a hierarchical vector representation method for traditional Chinese medicine case record texts provided by an embodiment of the present invention;

[0095] Figure 2 It is a schematic diagram of an element-level vector representation model provided by an embodiment of the present invention;

[0096] Figure 3 It is a schematic diagram of an element set data vector representation model provided by an embodiment of the present invention;

[0097] Figure 4 It is a star subgraph of traditional Chinese medicine case record texts provided by an embodiment of the present invention;

[0098] Figure 5 It is a block diagram of a hierarchical vector representation device for traditional Chinese medicine case record texts provided by an embodiment of the present invention;

[0099] Figure 6 It is a schematic structural diagram of a device for hierarchical vector representation of traditional Chinese medicine case record texts provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0100] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.

[0101] Those skilled in the art should be aware that the following specific embodiments or specific implementation manners are a series of optimized setting manners listed by the present invention to further explain the specific inventive content, and these setting manners can be combined with each other or used in association with each other, unless the present invention clearly states that some or a specific embodiment or implementation manner cannot be associated or used jointly with other embodiments or implementation manners. At the same time, the following specific embodiments or implementation manners are only used as the optimized setting manners and are not used to limit the understanding of the protection scope of the present invention.

[0102] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail in conjunction with the accompanying drawings and specific embodiments.

[0103] The embodiment of the present invention provides a hierarchical vector representation method for traditional Chinese medicine medical record texts. This method can be implemented by a vector model representation device, and the vector model representation device can be a terminal or a server. As Figure 1 shown in the flowchart of the hierarchical vector representation method for traditional Chinese medicine medical record texts, the processing flow of this method can include the following steps:

[0104] S1. Based on traditional Chinese medicine theory, determine the structured ontology design of traditional Chinese medicine medical records.

[0105] Among them, the structured ontology design of traditional Chinese medicine medical records includes content level design, data type level design and granularity level design;

[0106] The content level design of traditional Chinese medicine medical records includes patient information, clinical information, diagnosis information, treatment information and other information;

[0107] The data type level design of traditional Chinese medicine medical records includes numerical type, sequence type, and element set type;

[0108] The numerical type includes age, date, etc.;

[0109] The sequence type includes patient's chief complaint, current medical history, etc.;

[0110] The element set type includes current symptoms, tongue image sequence, pulse image sequence, etc.;

[0111] The granularity level design of traditional Chinese medicine medical records includes element level, field level and multi-field level;

[0112] Element level is an entity term in traditional Chinese medicine case texts, including symptoms, tongue manifestations, pulse conditions, syndromes, syndrome elements, traditional Chinese medicines, etc.;

[0113] Field level is a component element of a case, including sequence type data and element set type data;

[0114] Multi-field level is the aggregation of multiple field-level data;

[0115] The designed content hierarchy is used for the subsequent design of the topic recognition Prompt template, indicating how the language model can identify effective information of various topics in the original case text data. The designed granularity hierarchy is used to define various structural data of the case text. The designed data type hierarchy is used to classify various text data of the case, indicating that the model should specifically select a specific vector embedding model to initialize and represent the data.

[0116] Based on the above data patterns and characteristics, the structured ontology of traditional Chinese medicine cases is constructed as shown in Table 1 below.

[0117] Table 1 Design example of structured representation of traditional Chinese medicine cases

[0118]

[0119]

[0120] S2. Obtain professional case data in the field of traditional Chinese medicine; preprocess the professional case data; perform structured information extraction on traditional Chinese medicine cases based on the prompting word engineering technology of the large model to obtain various types of data at the element level, field level, and multi-field level.

[0121] Among them, the professional case data in the field of traditional Chinese medicine is high-quality case text data obtained through laboratory accumulation, vertical medical online web page crawling, and text data of professional medical books in the medical field; the structural types of case data in the field of traditional Chinese medicine include structured data, semi-structured data, and unstructured data;

[0122] Optionally, the preprocessing of the professional case data to obtain structured case text data in the field of traditional Chinese medicine includes:

[0123] Classify the professional case data based on the data format to obtain XML data, HTML data, PDF data, first TXT data, and standardized first structured case data;

[0124] Perform data parsing on the XML data and the HTML data to obtain second TXT data;

[0125] Perform OCR text recognition on the PDF data to obtain recognized text data; correct the recognized text data to obtain third TXT data;

[0126] Perform data cleaning on the first TXT data, the second TXT data, and the third TXT data to obtain the fourth TXT data;

[0127] Based on the medical record structured ontology at the content level, design a topic recognition prompt word template. The topic scope includes but is not limited to patient information, clinical information, diagnostic information, treatment information, doctor information, and remarks. The topic recognition prompt word templates are shown in Tables 2 and 3 below;

[0128] Table 2 Medical Record Text Identification Prompt Word Template

[0129]

[0130]

[0131] Table 3 Topic Recognition Prompt Word Template

[0132]

[0133] Based on the medical record structured ontology at the granularity level, design an information extraction prompt word template for each topic content. The prompt word templates for each topic are shown in Table 4;

[0134] Table 4 Information Extraction Prompt Word Template for Each Topic Content

[0135]

[0136]

[0137]

[0138] Input the above prompt word template and specific instances in the fourth TXT data into the large language model, which can make full use of the powerful information extraction ability of the LLM to obtain the structured information of the specified topic in the medical record text, so as to obtain element-level, field-level, and multi-field-level text data, and merge them with the first medical record structured data to obtain the structured data of traditional Chinese medicine medical records;

[0139] S3. Construct an element-level vector representation model based on the domain-enhanced pre-trained RoBERTa model to obtain the semantic representation of element-level data;

[0140] Among them, the element-level vector representation model is a RoBERTa model trained through multiple stages, including traditional Chinese medicine domain-enhanced pre-training and contrast learning.

[0141] The RoBERTa model is used to extract the unified vector representation of elements and sequence texts;

[0142] The enhanced pre-training in the field of traditional Chinese medicine is based on the corpus in the field of traditional Chinese medicine and trained with a masked language modeling objective;

[0143] Specifically, in order to more accurately capture the element semantics in the medical record text, by performing whole word mask pre-training on the medical record data, the model can more deeply understand the overall meaning of the entity words in the text, rather than just the local information of individual characters or phrases.

[0144] Let the sentence in the medical record dataset be S, and its set of entity words be W. Then the sentence S' after whole word masking is expressed as shown in Equation (1):

[0145]

[0146] where w i represents the entity word in sentence S, and w j represents the non-entity word in sentence S.

[0147] The training objective of the model is to predict the masked entity words based on the context information. Specifically, for each masked entity word position, the model outputs a probability distribution over the vocabulary, indicating the words that may appear at that position. Then, in this embodiment, the cross-entropy loss function is used to measure the accuracy of the model prediction, and the parameters of the model are updated through the backpropagation algorithm.

[0148] Let the probability distribution predicted by the model be P(w∣S′), and the true entity word distribution be Q(w∣s). Then the cross-entropy loss function is expressed as shown in Equation (2):

[0149]

[0150] where W masked represents the set of masked entity words.

[0151] The contrastive learning is to introduce the Sentence-BERT twin network architecture to learn the similarities and differences in the vector representations between samples;

[0152] The Sentence-BERT twin network architecture is as shown in Figure 2 It represents a kind of two-tower architecture, including an encoding module, a pooling module, and a similarity calculation module. Among them, the encoding module is the pre-trained RoBERTa model, the pooling module selects the first token - [CLS] output by the encoding module as the vector representation of the element word, and the similarity calculation module uses the cosine similarity method.

[0153] To improve the language model's ability to represent semantics at the element level, in this embodiment, a positive and negative sample construction and sample balancing processing strategy is designed. Table 5 below illustrates the positive and negative sample dataset construction strategy for the contrastive learning task using symptom entities as an example.

[0154] Table 5 Example of Positive and Negative Sample Dataset Construction for Contrastive Learning Task

[0155]

[0156] As described above, the construction strategy is as follows:

[0157] 1) Introduce conceptual semantics: In this embodiment, positive and negative examples are constructed for candidate words according to the set positive and negative concept label templates.

[0158] 2) Sample balancing processing: To maintain the balance of the number of positive and negative example samples, in this embodiment, similar words that are not query words are randomly selected from the full word list of various entities at the element level as negative examples. This step ensures that the model is not affected by sample imbalance during the training process.

[0159] Based on the dataset constructed by the above method, contrastive learning fine-tuning is performed on the pre-trained RoBERTa language model. Supervised pair learning is an improvement over unsupervised methods. By introducing known class labels to guide the contrastive learning process, its loss function is set as shown in the following formula (3):

[0160]

[0161] where, is the feature representation of the i-th sample; y i is the supervision label of the i-th sample, sim() represents whether it is a positive example, and it is a function that measures the similarity of the features of two samples, such as cosine similarity, etc. p i is the predicted probability of the i-th sample, K is the batch size, is an indicator function, which takes the value of 1 when y k and y i are equal, and 0 otherwise.

[0162] S4. The field-level data is divided into sequence type and element set type data. The semantic representation of the sequence type data is obtained based on the element-level vector representation model; the semantic representation of the element set type data is obtained by aggregating the semantics between elements based on the Self-Attention network;

[0163] wherein, the field-level vector representation includes the vector representation of sequence data and the vector representation of element set data;

[0164] For sequence data, it conforms to the characteristics of language models for numerical representation of unstructured text, and directly adopts element-level vector representation model to obtain semantic representation;

[0165] For element set data, which has the characteristics of high knowledge density and structured text description, the element-level vector representation model has insufficient representation ability. Taking the <Current Symptom> field of clinical information in medical case text as an example, this field contains multiple symptom entity elements, and the element-level vector representation model is difficult to capture the comprehensive vector representation of the entire set. Therefore, the present invention provides a semantic coding representation of each element in the element set data based on the Self-Attention network model to obtain the final vector representation.

[0166] The schematic diagram of the element set representation model architecture is as follows Figure 3 As shown in the figure, it includes input module, set feature extraction module and multi-task training module. The input module uses element-level encoding model to encode the features of a single symptom to obtain its accurate semantic representation; the set feature extraction module uses the Self-Attention layer network after removing the position encoding to capture the association between elements in the element set and generate an overall representation vector to represent the semantic information of the entire element set; the multi-task training module designs two training tasks, including self-supervised contrastive learning and disease name prediction tasks, to enhance the representation ability and generalization performance of the model.

[0167] The following are two training tasks in detail:

[0168] 1. Self-supervised contrastive learning SimCSE Loss——L unsup :This loss function strengthens the effect of overall features in retrieval tasks through self-supervised training tasks. Specifically, L unsup By calculating the similarity between different samples, the model is encouraged to generate similar representation vectors to shorten the distance between similar samples and push the distance between dissimilar samples. unsup As shown in the following formula (4):

[0169]

[0170] here, is with The relevant positive samples are obtained by Dropout sampling in this embodiment with reference to the method of SimCSE. The negative samples in the calculation process adopt the method of In-batch Negatives, that is, other samples different from the current sample are selected as negative samples. The training set is shown in Table 6.

[0171] Table 6. Training set examples of element set data vector representation models

[0172]

[0173] 2. Disease name prediction task: Introduce the disease name label information in the medical records and design a classification task for disease name prediction. This task takes the symptom set as the input and expects the model to predict the corresponding disease name label. The advantage is that it can strengthen the contextual features of the symptom set, so that the implicit representation of the symptom set is constrained by its disease condition to enhance its feature expression ability. At the same time, the disease name prediction task can also provide additional supervision signals to assist the training process of the model. In this part, the cross-entropy loss L CE is used for modeling, as shown in Equation (5):

[0174]

[0175] In this formula, C represents the total number of categories, y i is the true label of the sample, which is a one-hot vector. It is 1 when the sample belongs to the i-th category and 0 otherwise. i is the sample predicted by the model, and p i represents the probability that the sample belongs to the i-th category.

[0176] S5. Initialization of multi-field level data features; Obtaining the multi-field level vector representation in traditional Chinese medicine medical record texts based on the graph attention network;

[0177] 1. Initialization of multi-field level data features

[0178] Define each traditional Chinese medicine medical record as a star-shaped subgraph, with the central node being the original text of the current traditional Chinese medicine medical record, and each field graph node being associated with the central node. As Figure 4 shown, this figure is a simple example graph of an entire medical record, containing important information on various topics in the medical record text. Note that this figure mainly shows the relationship between the central node of the medical record and other field nodes, ignoring the associations between nodes, such as the relationship between the current symptoms and the disease and the prescription. The initial feature representations of each graph node are processed differently according to different data types, as follows:

[0179] (1) Sequence type data is encoded and represented by an element-level embedding model.

[0180] (2) Element set type data is encoded and represented by a field-level embedding model.

[0181] (3) Numeric type data is first discretized, combined with feature engineering to generate enhanced machine learning features, and then One-Hot encoding is used as its feature vector.

[0182] 2. Obtaining the multi-field level vector representation in traditional Chinese medicine medical record texts based on the graph attention network

[0183] In this embodiment, a graph attention neural network is adopted as the basic architecture, and the characteristics of neighbor nodes are weighted by the Attention algorithm to represent the central node of traditional Chinese medicine medical records. At the same time, the present invention proposes to design three denoising auto-encoding tasks based on the auto-encoder structure, including randomly masking nodes, randomly masking label vectors, and graph node vectors, to disrupt and reconstruct the subgraph structure or representation, so as to train the trainable weight parameters of the graph attention network.

[0184] Specifically, for the central traditional Chinese medicine medical record node i, the calculation process of its vector representation is as follows:

[0185] (1) Calculate the coefficient between the central node i and its adjacent nodes. The calculation formula of the correlation coefficient is as follows in formula (6):

[0186] e ij =F([Wh i ||Wh j ),j∈N i (6)

[0187] (2) According to the correlation coefficient, use the softmax function to normalize to obtain the attention coefficient. The calculation formula is as follows (7):

[0188]

[0189] (3) Perform weighted summation on the characteristics of each node. The formula for weighted summation of multi-head attention is as follows (8):

[0190]

[0191] Among them, W is the trainable weight parameter, h is the initialized vector encoding of each node, || represents feature concatenation, F represents the linear layer, which maps high-dimensional features to a real value; e ij represents the correlation coefficient, H represents the number of heads, σ represents the sigmoid activation function, and finally h′ i serves as the graph representation feature of node i.

[0192] In a feasible implementation manner, the features of various types of data in the medical record text are initialized specifically, and then the feature attention weights of each node are calculated through the graph attention network. Finally, the vector representation of the entire traditional Chinese medicine medical record is obtained through weighted summation.

[0193] This solution provides an efficient, accurate, and flexible method for representing traditional Chinese medicine (TCM) medical record texts in natural language processing tasks through a hierarchical representation method from the element level, field level to the multi-field level. Based on TCM theory, a structured ontology design for TCM medical records is determined. Professional medical record data in the TCM field is obtained, preprocessed, and structured information extraction is performed on the TCM medical records based on the prompt engineering technology of large models. For complex medical record texts, where the elements or sequence data, feature initialization is performed based on the element-level vector representation model. For element set data, the semantics between elements are aggregated based on the Self-Attention network to obtain its feature initialization. Finally, the fused vector representation of the entire TCM medical record text is obtained based on the graph attention network. The present invention is a refined information semantic representation method that fully considers the structural characteristics of TCM medical record texts and combines language model embedding representation technology.

[0194] Figure 5 FIG. 4 is a block diagram of a hierarchical vector processing device for TCM medical record texts according to an exemplary embodiment, and this device is used for the hierarchical vector representation method of TCM medical record texts. Refer to Figure 5 , this device includes a TCM medical record text ontology design module 510, a structured medical record text acquisition module 520, an element vector representation acquisition module 530, a field vector representation acquisition module 540, and a multi-field vector representation acquisition module 550. Among them:

[0195] The TCM medical record text ontology design module 510 is used to construct a structured ontology for TCM medical record texts based on TCM theoretical knowledge;

[0196] The structured medical record text acquisition module 520 is used to obtain professional medical record data in the TCM field in combination with the structured ontology of medical record texts; preprocess the professional medical record data; perform structured information extraction on the TCM medical records based on the prompt engineering technology of large models to obtain various types of data at the element level, field level, and multi-field level;

[0197] The element vector representation acquisition module 530 is used to construct an element-level vector representation model based on the domain-enhanced pre-trained RoBERTa model to obtain the semantic representation of element-level data;

[0198] The field vector representation acquisition module 540 is used to obtain the semantic representation of sequence data based on the element-level vector representation model; aggregate the semantics between elements based on the Self-Attention network to obtain the semantic representation of element set data;

[0199] The multi-field vector representation acquisition module 550 is used for multi-field graph node feature initialization; obtain the vector representation of TCM medical record texts based on the graph attention network.

[0200] Among them, the professional medical case data in the field of traditional Chinese medicine is high-quality medical case text data obtained through laboratory accumulation, vertical medical online web crawling, and text data of professional medical books in the medical field; the structural types of the medical case data in the field of traditional Chinese medicine include structured data, semi-structured data, and unstructured data;

[0201] The traditional Chinese medicine medical case text ontology design module 510 is further used for the structured ontology design of traditional Chinese medicine medical cases, including content hierarchy design, data type hierarchy design, and granularity hierarchy design;

[0202] The content hierarchy design of traditional Chinese medicine medical cases includes patient information, clinical information, diagnostic information, treatment information, and other information;

[0203] The data type hierarchy design of traditional Chinese medicine medical cases includes numerical type, sequence type, and element set type;

[0204] The numerical type includes age, date, etc.;

[0205] The sequence type includes patient chief complaints, current medical history, etc.;

[0206] The element set type includes current symptoms, tongue image sequence, pulse condition sequence, etc.;

[0207] The granularity hierarchy design of traditional Chinese medicine medical cases includes element level, field level, and multi-field level;

[0208] The element level is the entity words in the traditional Chinese medicine medical case text, including symptoms, tongue images, pulse conditions, syndromes, syndrome elements, traditional Chinese medicines, etc.;

[0209] The field level is the constituent elements of the medical case, including sequence type data and element set type data;

[0210] The multi-field level is the aggregation of multiple field level data;

[0211] Optionally, the structured medical case text acquisition module 520 is further used for:

[0212] Classify the professional medical case data based on the data format to obtain XML data, HTML data, PDF data, first TXT data, and standardized first medical case structured data;

[0213] Parse the XML data and the HTML data to obtain second TXT data;

[0214] Perform OCR text recognition on the PDF data to obtain recognized text data; correct the recognized text data to obtain third TXT data;

[0215] Clean the first TXT data, second TXT data, and third TXT data to obtain fourth TXT data;

[0216] Based on the medical record structured ontology at the content level, design a topic recognition prompt word template, and the topic scope includes but is not limited to patient information, clinical information, diagnosis information, treatment information, doctor information, and remarks;

[0217] Based on the medical record structured ontology at the granularity level, design an information extraction prompt word template for each topic content;

[0218] Based on the large model and combined with the above prompt word template design, perform information extraction on the fourth TXT data to obtain element-level, field-level, and multi-field-level text data, and merge it with the first medical record structured data to obtain traditional Chinese medicine medical record structured data;

[0219] Optionally, the multi-field vector representation acquisition module 550 is further used for the vector representation of medical record texts or multi-type data, including multi-field data feature initialization and a graph attention network encoding model;

[0220] The multi-field data feature initialization includes element set type encoding representation, sequence type encoding representation, and numerical type encoding representation: the sequence type encoding representation is obtained by an element-level vector representation model; the element set type encoding representation is obtained by a field-level encoding representation model; the numerical type encoding representation is obtained by using One-Hot encoding after discretization processing;

[0221] Based on the graph attention network encoding model, obtain the vector representation of the entire traditional Chinese medicine medical record by weighting the features of neighbor nodes through the Attention algorithm.

[0222] Figure 6 It is a schematic structural diagram of a device for hierarchical vector processing of traditional Chinese medicine medical record texts provided by an embodiment of the present invention. As Figure 6 shown, the device for hierarchical vector representation of traditional Chinese medicine medical record texts may include the above-mentioned Figure 5 shown hierarchical vector representation device for traditional Chinese medicine medical record texts. Optionally, the large language model question answering device 610 may include a first processor 2001.

[0223] Optionally, the device 610 for hierarchical vector representation of traditional Chinese medicine medical record texts may further include a memory 2002 and a transceiver 2003.

[0224] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.

[0225] Next, in combination with Figure 6 specifically introduce each component of the device 610 for hierarchical vector representation of traditional Chinese medicine medical record texts:

[0226] Among them, the first processor 2001 is the control center of the hierarchical vector representation device 610 for traditional Chinese medicine case texts, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0227] Optionally, the first processor 2001 can execute various functions of the hierarchical vector representation device 610 for traditional Chinese medicine case texts by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0228] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 6 the CPU0 and CPU1 shown in

[0229] In a specific implementation, as an embodiment, the hierarchical vector representation device 610 for traditional Chinese medicine case texts may also include multiple processors, such as Figure 6 the first processor 2001 and the second processor 2004 shown. Each of these processors can be a single-CPU or a multi-CPU. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0230] Among them, the memory 2002 is used to store software programs for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.

[0231] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through the interface circuit of the traditional Chinese medicine case text hierarchical vector representation device 610 ( Figure 6 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.

[0232] The transceiver 2003 is used to communicate with a network device or communicate with a terminal device.

[0233] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 6 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0234] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through the interface circuit of the traditional Chinese medicine case text hierarchical vector representation device 610 ( Figure 6 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.

[0235] It should be noted that Figure 6 the structure of the large language model question answering device 610 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0236] In addition, the technical effects of the traditional Chinese medicine case text hierarchical vector representation device 610 may refer to the technical effects of the hierarchical vector representation method for traditional Chinese medicine case texts described in the above method embodiments, and will not be elaborated here.

[0237] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0238] As described above, the specific implementation manners of the present invention are only described, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A hierarchical vector processing method for TCM medical records, characterized in that: The method comprises: Based on the theory of traditional Chinese medicine, determine the structured ontology design of traditional Chinese medicine medical records; Obtain professional medical case data in the field of traditional Chinese medicine; pre-process the professional medical case data; extract structured information from traditional Chinese medicine medical cases based on the prompt word engineering technology of the large model, and obtain various types of data at the element level, field level and multi-field level; Build an element-level vector representation model based on the domain-enhanced pre-trained RoBERTa model to obtain element-level data semantic representation; Field-level data is divided into sequence type and element collection type data. Sequence type data semantic representation is obtained based on element-level vector representation model. Semantic representation of element collection data is obtained based on Self-Attention network aggregation of semantics between elements. The acquisition of multi-field-level data semantic representation includes the initialization vector encoding representation of element sets, sequences and numerical type data and the weighted summation of vector encodings of various types of data based on the graph attention network, so as to represent the semantic representation in the TCM medical case text.

2. The hierarchical vector processing method for TCM medical case text according to claim 1 is characterized in that: The structured ontology design of TCM medical records includes content level design, data type level design and granularity level design; the content level design of TCM medical records includes patient information, clinical information, diagnosis information, treatment information and other information; The hierarchical design of the TCM medical records data type includes a numerical type, a sequence type, and an element set type; The granularity level design of the TCM medical records includes element level, field level and multi-field level.

3. The hierarchical vector processing method for TCM medical case text according to claim 1 is characterized in that: The professional medical case data in the field of traditional Chinese medicine is high-quality medical case text data obtained through laboratory accumulation, vertical medical online web crawling, and professional medical book text data in the field of medicine; the structural types of the medical case data in the field of traditional Chinese medicine include structured data, semi-structured data, and unstructured data; Preprocess professional medical case data to obtain structured medical case text data in the field of traditional Chinese medicine, including: Based on the data format, the professional medical record data is classified to obtain XML data, HTML data, PDF data, first TXT data and standardized first medical record structured data; Parsing the XML data and the HTML data to obtain second TXT data; Performing OCR text recognition on the PDF data to obtain recognized text data; correcting the recognized text data to obtain third TXT data; Cleaning the first TXT data, the second TXT data, and the third TXT data to obtain fourth TXT data; Based on the structured ontology of medical records at the content level, a template for topic identification prompts is designed. The topic range includes patient information, clinical information, diagnosis information, treatment information, doctor information and comments. Based on the structured ontology of medical records at a granular level, information extraction prompt word templates are designed under each topic content; Based on the big model and in combination with the above-mentioned prompt word template, information is extracted from the fourth TXT data to obtain element-level, field-level and multi-field-level text data, which is then merged with the first medical case structured data to obtain TCM medical case structured data.

4. The hierarchical vector processing method for TCM medical case text according to claim 1 is characterized in that: The element-wise vector representation model is the base of the RoBERTa model, which has undergone multi-stage training, including enhanced pre-training and contrastive learning in the field of traditional Chinese medicine: The RoBERTa model is used to extract a unified vector representation of element and sequence text; The enhanced pre-training in the field of traditional Chinese medicine is based on the corpus in the field of traditional Chinese medicine and is trained through the masked language modeling objective; The contrastive learning introduces the Sentence-BERT twin network architecture to learn the similarities and differences between vector representations of samples.

5. The hierarchical vector processing method for TCM medical case text according to claim 1 is characterized in that: The field-level vector representation includes a vector representation of sequence data and a vector representation of element set data; The vector representation of the sequence data uses an element-level vector representation model to obtain semantic representation; The vector representation method of the element set data includes element data encoding and set data encoding; The element data encoding adopts an element-level vector representation model to obtain the semantic representation of each element; The collective data encoding adopts the Self-Attention network to aggregate the semantic encoding representation of each element.

6. The hierarchical vector processing method for TCM medical case text according to claim 1 is characterized in that: The semantic representation of the multi-field-level data includes an initialization vector encoding representation of element sets, sequences and numerical type data and a weighted summation of vector encodings of various types of data based on a graph attention network; The encoding representation of each type of data initialization vector includes an element set type encoding representation, a sequence type encoding representation and a value type encoding representation; The sequence type encoding representation is obtained by an element-level vector representation model; The element set type encoding representation is obtained by a field-level encoding representation model; The numerical type encoding representation is obtained by discretization and One-Hot encoding; The graph attention network coding model characterizes the central node of the TCM medical records by weighting the features of neighboring nodes through the Attention algorithm.

7. A hierarchical vector processing device for TCM medical case texts, the device being used to implement the hierarchical vector processing method for TCM medical case texts as claimed in any one of claims 1 to 6, characterized in that: The device comprises: TCM medical case text ontology design module, which is used to construct a structured ontology of TCM medical case text based on TCM theoretical knowledge; The structured medical case text acquisition module is used to obtain professional medical case data in the field of traditional Chinese medicine by combining the medical case text structured ontology; pre-process the professional medical case data; extract structured information from traditional Chinese medicine medical cases based on the prompt word engineering technology of the large model, and obtain various types of data at the element level, field level and multi-field level; The element vector representation acquisition module is used to build an element-level vector representation model based on the RoBERTa model pre-trained by domain enhancement to obtain the element-level data semantic representation; Field vector representation acquisition module: field-level data is divided into sequence type and element collection type data. Sequence type data semantic representation is acquired based on element-level vector representation model; semantic representation of element collection data is acquired based on Self-Attention network aggregation of inter-element semantics; The multi-field vector representation acquisition module is used to initialize multi-field-level data features; the vector representation of TCM medical case text is obtained based on the graph attention network.

8. A hierarchical vector processing device for TCM medical records, characterized in that: The device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Traditional Chinese medicine ancient and modern situation comparison method based on agency loss

    CN121122776A

  • A traditional chinese medicine ancient and modern situation comparison method based on agent loss

    CN121122776B