A hypergraph-based method for electronic medical record text analysis
By adopting a multi-task learning method based on hypergraphs in electronic medical record text analysis, combining the BERT model and hypergraph perception layer, the problem of negative sample processing difficulties in the existing technology is solved, the performance and robustness of the model are improved, and better text classification effect is achieved.
Patent Information
- Application Number
- CN202411319544.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-09-20
AI Technical Summary
Existing electronic medical record text analysis techniques are difficult to effectively process negative samples, resulting in a degradation in model performance in entity recognition and relationship extraction tasks.
Using a multi-task learning method based on hypergraphs, the entity relationship extraction framework and node classification model are used, and the pre-trained BERT model and hypergraph perception layer are combined to integrate entity recognition and relationship extraction tasks to improve the performance and robustness of the model.
The electronic medical record text processing technology has been improved, the model performance and robustness in entity recognition and relationship extraction tasks have been improved, negative samples can be better processed, and the effectiveness and generalization capabilities of text classification tasks have been improved.
Smart Images

Figure CN119230043B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic medical record texts, and specifically to an electronic medical record text analysis method based on a hypergraph. Background Art
[0002] Electronic medical record text analysis refers to the technology of automatically processing and analyzing electronic medical record texts in the medical field. With the rapid development of medical information technology and the wide application of electronic medical records, electronic medical record text analysis has become increasingly important. Electronic medical record texts contain rich medical information, including patients' medical histories, symptom descriptions, diagnosis results, treatment plans, etc. By analyzing electronic medical record texts, key information can be extracted and mined, providing valuable support for medical research, clinical decision-making, and medical quality improvement. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art, the present invention provides an electronic medical record text analysis method based on a hypergraph, which solves the problem that it is impossible to ensure that its wireless signal can completely cover this area without signal interference.
[0004] To achieve the above objectives, the present invention is realized through the following technical solutions: An electronic medical record text analysis method based on a hypergraph specifically includes the following steps:
[0005] Step 1, Data collection: Determine the source of data collection, communicate with the data source through an interface or API, automatically collect medical conversation records, and securely store the collected data in a database or data warehouse;
[0006] Step 2, Data preprocessing: Perform data cleaning, formatting standardization, deduplication, data segmentation, feature extraction, data conversion, data augmentation, and sensitive information desensitization on the data;
[0007] Step 3, Input the preprocessed data into the entity relationship extraction framework to realize the identification and extraction of entity relationships in the medical record text, and complete the automatic recording of the diagnostic dialogue system and the electronic medical record (EMR);
[0008] Step 4, After the text entity relationship extraction, input the identified entities and the extracted relationships into the node classification model on the hypergraph based on text attributes. According to the characteristics of the medical record text and the task requirements, use the entities as the nodes of the hypergraph and the relationships as the hyperedges of the hypergraph, integrate the results of entity identification and relationship extraction into the hypergraph, where the hyperedges connect multiple nodes, reflecting the association relationships between the nodes, and complete the text classification task.
[0009] Preferably, the entity relationship extraction framework includes vector representation, span representation, entity identification, and relationship extraction, and the entity identification and relationship extraction are integrated to form a multi-task framework;
[0010] Vector representation. The vector representation layer uses fine-tuned pre-trained BERT to extract context information. The input sentence is encoded by byte pair BPE and segmented into a sequence of n tokens. The BPE tokens are propagated through BERT, generating an embedding sequence E of length n + 1: E = (e [CLS] , e1, e2,..., e n ), where e [CLS] represents the embedding representation of the entire sentence context, and each embedding vector where d1 represents the embedding dimension;
[0011] Span representation. The span representation layer uses the embeddings of the tokens generated by BERT to represent spans, maintaining a consistent representation dimension s in different spans: s := (e i , e i+1 ,..., e j ). Apply a fusion process to s, using max pooling in the fusion process:
[0012]
[0013] The model sets the maximum span length L. To reduce complexity, a width embedding matrix is used to incorporate width information w (0 ≤ w ≤ L) into the fused span representation s. Spans of different widths contribute differently to the representation:
[0014]
[0015] where, denotes the concatenation of vectors, and w k is derived from the width embedding matrix obtained by training the model;
[0016] The multi-task framework includes an identification task and a classification task:
[0017] Identification task. The identification task is used to determine whether the current span or span pair is indeed an entity or whether there is a certain relationship. Binary cross-entropy is selected as the loss function, defined as follows:
[0018]
[0019] where q ∈ {0, 1} represents the true class of the span or span pair, where 0 represents a negative sample and 1 represents a positive sample), and p ∈ [0, 1] is the probability estimated by the model when q = 1 is identified as the true class;
[0020] Therefore, the loss of the identification task is calculated as follows:
[0021]
[0022] Classification task. To enhance the diversity of the multi-task learning loss landscape and prevent the two tasks from degenerating into a single task, a pairwise ranking loss function is selected as the loss function for the classification task. This binary cross-entropy loss function has the same task objective as other forms of cross-entropy loss but has a different optimization direction.
[0023] Given a span (pair) representation r and an entity set C or a relation class set C, the score for a class label x ∈ C is calculated using the dot product:
[0024] score c = r T [W classes c ;
[0025] where W classes is the matrix to be learned, the number of columns hit is the number of classes, and [W classes c is the column vector corresponding to class c, whose dimension is equal to the dimension of s;
[0026] For a span (pair), where y + is the correct label and y - is not, and are defined as the scores of y + and y - respectively. Then the formula for the rank loss is:
[0027]
[0028] where γ is the scaling factor, m + and m - are the margins. For non-entity classes or unrelated entity pairs, only L - is calculated to penalize incorrect predictions. In the training step, the ranking loss only selects the highest-scoring incorrect class among all incorrect classes. Then, the pairwise ranking loss is optimized:
[0029] Loss2 = ∑(L + + L - );
[0030] Finally, the loss function of the multi-task framework used is expressed as:
[0031] Loss = αLoss1 + βLoss2;
[0032] Entity recognition. The entity recognition layer directly uses the output of the span representation layer as input and extracts entity classes from the spans using the multi-task framework.
[0033] In the recognition tasks of the entity recognition multi-task framework, the position information between spans is introduced by using the IoU (Intersection over Union) metric; formally, the calculation of IoU is as follows:
[0034]
[0035] where s(i, j) represents the span in the sentence with left and right indices i and j;
[0036] Calculate and utilize the maximum IoU (Intersection / Union) between span t and all entity spans in the sentence, defined as ENIoU(t), and its calculation formula is:
[0037] ENIoU(t) = max(IoU(t, en), en ∈ E entity )
[0038] where E entity represents the set composed of all entity spans in the sentence;
[0039] Subsequently, using ENIoU as the scaling factor, replace the loss function of the recognition task in the multi-task framework with a dynamically scaled cross-entropy loss, and the loss function is:
[0040]
[0041] where δ and ENIoU determine the scaling factor, δ ∈ [0, 1] is the balance factor to solve class imbalance, and it is necessary to set δ < 0.5 to emphasize fewer positive samples; for negative samples, the higher the ENIoU, the greater the loss, so that the model pays more attention to hard negative samples during training;
[0042] For other parts of the multi-task framework, no further modification is required during the entity recognition stage, and the loss of the entity recognition stage is denoted as L1;
[0043] For relation extraction, different from directly using the output of the span representation layer as input during the entity recognition stage, the relation extraction stage makes some adjustments to the input information. For those spans classified as "NA" during the entity recognition multi-task stage, they will be filtered out; we only consider candidate span pairs that have been initially determined as entities, select the local information between the candidate span pairs as their context, and obtain the context representation by fusing their BERT embeddings using maxpooling;
[0044]
[0045] For candidate span pairs of adjacent entities, set v(s1, s2) = 0; in addition, we observe that the Logits in the entity recognition stage may contain useful semantic information;
[0046] Therefore, Logits can be further embedded in the entity recognition stage to enrich the input information; finally, the multi-task input representation of the relation extraction layer is as follows:
[0047]
[0048] Where p(s) represents the Logits in the entity recognition stage; considering that the relationship between s1 and s2 is usually asymmetric, therefore:
[0049]
[0050] For the multi-task framework used in the relation extraction stage, the loss of the relation extraction stage is denoted as L2.
[0051] Preferably, the entity relation extraction framework further includes a joint training loss: for the joint training method of entity recognition and relation extraction, the joint training loss is defined as follows: L = L1 + L2.
[0052] Preferably, the entity relation extraction framework further includes a prediction stage. In the prediction stage, in order to avoid error propagation, only the class scores obtained from the multi-task framework are used, denoted as score c , while the binary recognition task is only used to optimize the network parameters;
[0053] Given a span (pair), the calculation of the predicted class probability P is as follows:
[0054]
[0055] Where θ is a hyperparameter threshold; if the score of each category is lower than θ, the span (pair) is predicted as a non-entity or an irrelevant relationship; otherwise, the span (pair) will be assigned to the class with the highest score.
[0056] Preferably, in step four, the node classification model includes: HyperBERT layer, hypergraph-aware pre-training tasks, and fine-tuning of downstream tasks;
[0057] The HyperBERT layer explicitly mixes the semantic information obtained from each BERT transformer block with the structural information obtained from the Hypergraph Neural Network (HGNN) block;
[0058] The hypergraph-aware pre-training tasks utilize the hypergraph knowledge inherent in TAHG and are applied to semantic and structural representations;
[0059] For fine-tuning of downstream tasks, after pre-training is completed, freeze the parameters of HyperBERT and use it to calculate the representations of each text node. The calculated representations can be further used to fine-tune individual models on different downstream tasks. Train the linear projection from node features to class labels using cross-entropy loss for node classification.
[0060] Preferably, the node classification has several classification methods:
[0061] Disease classification: Classify medical record texts into different disease categories;
[0062] Symptom classification: Classify medical record texts into different symptoms or syndromes;
[0063] Treatment plan classification: Classify medical record texts into different treatment plans or treatment methods.
[0064] The present invention provides a hypergraph-based method for analyzing electronic medical record texts. Compared with the prior art, it has the following beneficial effects:
[0065] 1. Provide an electronic medical record text processing technology, aiming to improve the impact of negative samples on the model performance in entity recognition and relation extraction tasks. By using the multi-task learning method in span-based joint entity relation extraction, solve the problem of performance degradation of the current model when dealing with non-entity spans or irrelevant span pairs.
[0066] 2. Provide a span-based multi-task learning method for dealing with negative samples in span-based joint entity relation extraction tasks. This method alleviates the impact of negative samples on entity and relation classifiers by introducing multi-task learning, and further improves the performance and robustness of the model.
[0067] 3. Provide a node classification technology on hypergraphs based on text attributes, aiming to overcome the difficulties of existing methods in simultaneously capturing hypergraph structure information and rich language attributes in node attributes, thereby improving the effect and generalization ability of node classification. By using a pre-trained BERT model and introducing a dedicated hypergraph-aware layer, the model's understanding ability of high-order structure information on hypergraphs is further enhanced.
[0068] 4. Improve the electronic medical record text processing technology, enabling it to better handle the negative sample problem in entity recognition and relation extraction tasks, and improve the performance and robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0070] Figure 2 It is a framework diagram of entity relation extraction of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0071] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0072] Please refer to Figure 1 , this application provides a method for analyzing electronic medical record texts based on hypergraphs, specifically including the following steps:
[0073] Step 1: First, it is necessary to determine the source of data collection. For electronic medical record text processing technology, data sources may include the hospital's electronic health record system, diagnostic dialogue system, and other relevant medical information systems. By communicating with the interfaces or APIs of the data sources, medical dialogue records are automatically collected. Before processing any personal health information, the consent of the patient must be obtained. This includes a clear description of the purpose and scope of data collection and how the data will be used. To protect patient privacy, the collected data needs to be anonymized by removing all information that may identify an individual's identity, such as name, address, social security number, etc. During the data collection process, quality control measures are implemented to ensure the accuracy and consistency of the data. The collected data is securely stored in a database or data warehouse for subsequent processing and analysis. Ensure that data storage complies with industry standards and regulatory requirements.
[0074] Step 2: Data preprocessing: Perform data cleaning, formatting standardization, duplicate removal, data segmentation, feature extraction, data conversion, data augmentation, and sensitive information desensitization on the data;
[0075] Data cleaning:
[0076] In the data cleaning stage, it is necessary to carefully check the collected medical dialogue data and remove irrelevant information, such as non-medical term dialogue content, advertising information, etc. At the same time, spelling mistakes, grammar mistakes, etc. in the data also need to be corrected. The purpose of this step is to ensure the purity and usability of the data.
[0077] Format standardization:
[0078] Since the collected data may come from different information systems, there may be differences in their formats and encodings. All data needs to be converted into a unified format, including date and time format, medical term standardization, data encoding consistency, etc., which helps subsequent data processing and analysis.
[0079] Duplicate removal operation:
[0080] There may be duplicate records in the dataset, and it is necessary to check and merge or delete these duplicates to ensure the uniqueness of the data. This can not only reduce the size of the dataset but also improve the accuracy of subsequent analysis.
[0081] Data splitting:
[0082] According to requirements, the dataset can be further divided into smaller units, for example, classified according to criteria such as patients, medical conditions, or types of dialogue texts. This helps to conduct in-depth analysis and processing of specific categories of data.
[0083] Missing value handling:
[0084] There will inevitably be some missing values in the dataset, and appropriate methods need to be taken to handle them, such as filling, deleting, or interpolation. This step ensures the integrity of the data and lays a foundation for subsequent analysis.
[0085] Feature extraction:
[0086] From the original medical text data, features helpful for model training need to be extracted, such as medical terms, symptom descriptions, drug names, etc. These features will be used as inputs for training the medical record text processing model.
[0087] Data transformation:
[0088] After feature extraction, the data can also be transformed using some mathematical and statistical methods, such as normalization or standardization, to improve the efficiency and accuracy of model training.
[0089] Data augmentation:
[0090] To improve the generalization ability of the model, some augmentation operations can be performed on the dataset, such as synonym replacement, sentence restructuring, etc. These techniques can artificially create more diverse training samples to make up for the deficiencies of the original data.
[0091] Sensitive information desensitization:
[0092] When processing medical data containing personal privacy information, desensitization processing is required to ensure that patients' sensitive information will not be leaked during subsequent analysis and processing, protecting patients' privacy.
[0093] These preprocessing steps help to improve the quality and consistency of text sentences, facilitating subsequent text entity relationship extraction tasks and classification and diagnosis tasks. At the same time, ensure the protection of patients' privacy and data security throughout the preprocessing process.
[0094] Step 3: Input the preprocessed data into the entity relationship extraction framework to identify and extract the entity relationships in the medical record text, and complete the automatic recording of the diagnostic dialogue system and electronic medical record (EMR);
[0095] The entity relationship extraction framework includes vector representation, span representation, entity recognition, and relationship extraction, and the entity recognition and relationship extraction are integrated to form a multi-task framework;
[0096] The vector representation layer uses fine-tuned pre-trained BERT to extract context information. As Figure 2 shown, the input sentence is encoded by byte pair encoding (BPE) and segmented into a sequence of n tokens. BPE encoding represents uncommon words (such as "pretrain") with common sub-words (such as "pre" and "train"), thus constraining the vocabulary and effectively handling out-of-vocabulary (OOV) and rare words. The BPE tokens are propagated through the interior of BERT, generating an embedding sequence E of length n + 1: E = (e [CLS] , e1, e2,..., e n ), where e [CLS] represents the embedding representation of the entire sentence context, and each embedding vector (where d1 represents the embedding dimension);
[0097] The span representation layer uses the embeddings of the tokens generated by BERT to represent spans. As Figure 2 shown, in order to maintain a consistent representation dimension s in different spans: s := (e i , e i+1 ,..., e j ), a fusion process is applied to s. Max pooling is used in the fusion process:
[0098]
[0099] Based on prior knowledge, longer spans are less likely to effectively represent meaningful entities. Therefore, a maximum span length L is set for the model, mainly to reduce complexity. In addition, a width embedding matrix is needed to incorporate width information w (0 ≤ w ≤ L) into the fused span representation j. The guiding idea of this method is that spans of different widths contribute differently to the representation:
[0100]
[0101] where, denotes the concatenation of vectors, and w k is derived from the width embedding matrix obtained by training the model.
[0102] The category of an entity in a sentence is likely to be related to the overall context information of the sentence. For example, in a medical record, "The patient took antibiotics in the past few days and felt that the symptoms of sore throat and cough had improved." Here, "antibiotics" can strongly indicate the entity class "drug". Therefore, we also need to extend the context embedding information [CLS] to the span representation:
[0103]
[0104] The recognition task is used to determine whether the current entity or entity pair is indeed an entity or whether there is a certain relationship; the binary cross-entropy is selected as the loss function, which is defined as follows:
[0105]
[0106] where q ∈ {0, 1} represents the true category of the span or span pair, where 0 represents a negative sample and 1 represents a positive sample), and p ∈ [0, 1] is the estimated probability of the model when the label q = 1 is identified as the true category;
[0107] Therefore, the loss calculation of the recognition task is as follows:
[0108]
[0109] To enhance the diversity of the multi-task learning loss landscape and prevent the two tasks from degenerating into a single task, the pairwise ranking loss function is selected as the loss function for the classification task. This binary cross-entropy loss function has the same task objective as other forms of cross-entropy loss, but has a different optimization direction.
[0110] Given the span (pair) representation s and the set of entity or relationship classes C, the score of the class label x ∈ C is calculated using the dot product:
[0111] score c = r T [W classes c ;
[0112] where W classes is the matrix to be learned, and the number of columns hit is the number of classes. [W classes c is the column vector corresponding to the class c, and its dimension is equal to the dimension of r.
[0113] For a span (pair), where y + is the correct label and y - is not, and are defined as the scores of y + and y - respectively. Then the calculation formula for the rank loss is:
[0114]
[0115] where γ is the scaling factor, m + and m - are margins. For non-entity classes or unrelated entity pairs, only L - is calculated to penalize incorrect predictions. In the training step, the ranking loss only selects the error class with the highest score among all error classes. Then, the pairwise ranking loss is optimized as follows:
[0116] Loss2 = ∑(L + + L - );
[0117] Finally, the loss function of the multi-task framework used is expressed as:
[0118] Loss = αLoss1 + βLoss2;
[0119] The entity recognition layer directly uses the output of the span representation layer as input and extracts entity classes from the spans using the multi-task framework. In this section, it is adjusted according to the multi-task framework described in Section 3.3.
[0120] In the recognition task of the entity recognition multi-task framework, the position information between spans is introduced by using the IoU (Intersection over Union) metric. Formally, the calculation of IoU is as follows:
[0121]
[0122] where s(i,j) represents the span in the sentence with left and right indices i and j.
[0123] The maximum IoU (Intersection / Union) between the span s and all entity spans in the sentence is calculated and utilized, defined as ENIoU(s). Its calculation formula is:
[0124] ENIoU(t) = max(IoU(t,en), en ∈ E entity );
[0125] where E entity represents the set composed of all entity spans in the sentence.
[0126] Subsequently, using ENIoU as the scaling factor, the loss function of the recognition task in the multi-task framework is replaced with a dynamically scaled cross-entropy loss. The loss function is:
[0127]
[0128] Among them, δ and ENIoU determine the scaling factor. δ ∈ [0, 1] is the balance factor to address class imbalance, and it is necessary to set δ < 0.5 to emphasize the fewer positive samples. For negative samples, the higher the ENIoU, the greater the loss, thus making the model pay more attention to hard negative samples during training.
[0129] For the other parts of the multi-task framework, no further modification is required in the entity recognition stage. Denote the loss in the entity recognition stage as L1.
[0130] For relation extraction, different from directly using the output of the span representation layer as input in the entity recognition stage, some adjustments are made to the input information in the relation extraction stage. For those spans classified as "NA" in the entity recognition multi-task stage, they will be filtered out; we only consider the candidate span pairs that have been preliminarily determined as entities, select the local information between the candidate span pairs as their context, and obtain the context representation by fusing their BERT embeddings using maxpooling.
[0131]
[0132] For adjacent candidate entity pairs, set v(s1, s2) = 0; in addition, we observe that the Logits in the entity recognition stage may contain useful semantic information.
[0133] Therefore, the Logits can be further embedded from the entity recognition stage to enrich the input information; finally, the multi-task input representation of the relation extraction layer is as follows:
[0134]
[0135] where p(s) represents the Logits in the entity recognition stage; considering that the relationship between s1 and s2 is usually asymmetric, thus:
[0136]
[0137] For the multi-task framework used in the relation extraction stage, denote the loss in the relation extraction stage as L2.
[0138] We learned the parameters of the span width embedding w and the parameters of the entity / relation multi-task classifier, and fine-tuned BERT during this process. For the joint training method of entity recognition and relation extraction, the joint training loss is defined as follows: L = L1 + L2.
[0139] In the prediction stage, to avoid error propagation, only the class scores obtained from the multi-task framework are used, denoted as score c , while the binary recognition task is only used to optimize the network parameters.
[0140] Given a span (pair), the calculation of the predicted class probability P is as follows:
[0141]
[0142] where θ is a hyperparameter threshold; if the score for each class is lower than θ, the span (pair) is predicted as a non-entity or an unrelated relationship; otherwise, the span (pair) is assigned to the class with the highest score;
[0143] After the above model, entity relationship extraction from the medical record text is completed, preparing for the text classification task of the subsequent model.
[0144] Step Four. After text entity relationship extraction, the identified entities and the extracted relationships are input into a node classification model on a hypergraph based on text attributes. According to the characteristics of the medical record text and the task requirements, the entities are used as the nodes of the hypergraph, and the relationships are used as the hyperedges of the hypergraph. The results of entity recognition and relationship extraction are integrated into the hypergraph. The hyperedges connect multiple nodes, reflecting the association relationships between the nodes, and the text classification task is completed;
[0145] HyperBERT layer
[0146] The HyperBERT layer explicitly mixes the semantic information obtained by each BERT transformer block with the structural information obtained by the Hypergraph Neural Network (HGNN) block.
[0147] Semantic representation. The semantic representation of the node text attributes is encoded by the BERT encoder. Specifically, BERT consists of alternating multi-head self-attention (MHA) and fully-connected feed-forward network (FFN) blocks, just like the original Transformer architecture. Before each block, layer normalization (LN) is applied, and a residual connection is added after each block. More formally, in the l th -th layer encoder, the hidden state is represented as where N is the maximum length of the input. First, the MHA block maps X to the query matrix the key matrix and the value matrix where m is the number of query vectors and n is the number of key / value vectors. Then, the formula for calculating the attention vector is as follows:
[0148] Attention(Q,K,V)=softmax(A)V,
[0149]
[0150] In practice, the MHA block calculates self-attention through h heads, where each head i is composed of and Independent parameterization maps the input embedding χ into query and key-value pairs. Then, the attention for each head is calculated and concatenated as follows:
[0151]
[0152] where is a trainable parameter matrix. Next, to obtain the semantic hidden state of the input, an FFN block is applied as follows:
[0153]
[0154] where and are linear weight matrices. Finally, layer normalization and residual connection are applied as follows:
[0155]
[0156] Therefore, after passing through L encoder layers, we represent the semantic representation of the text description of the node as
[0157] hypergraph structure representation. Thus, in each HyperBERT layer, a structural representation is generated on the node-centered context hypergraph G i by the hypergraph neural network (HGNN). This hypergraph is defined as a sub-hypergraph that only contains the hyperedges to which node i belongs and has its incidence matrix H i . Formally, we need to formulate the HGNN block in the following form to obtain the hypergraph node embedding:
[0158]
[0159] where D and B are the degree matrices of vertices and hyperedges in the hypergraph respectively. σ(·) is a non-linear activation function such as ReLU. is the learnable weight matrix between the l th th layer and the (l - 1) th th layer. Therefore, after passing through L hypergraph layers, we represent the structural representation of the node as
[0160] text-hypergraph joint representation. After calculating the representations in the semantic space and the hypergraph structure space, the lth HyperBERT layer calculates the mixed text-hypergraph representation as follows:
[0161]
[0162] where It contains a mixture of semantic information and hypergraph structure information to enable information flow between the two modalities.
[0163] Hypergraph-aware pre-training tasks
[0164] To improve the learning ability of HyperBERT without using any manually annotated labels, a novel hypergraph-aware knowledge alignment pre-training algorithm based on contrastive learning loss is required. The hypergraph-aware pre-training tasks used utilize the hypergraph knowledge inherent in TAHG and are applied to semantic and structural representations.
[0165] Semantic contrast loss. Based on the semantic representation of node i For other nodes N(i) belonging to the same hyperedge set as it and the mini-batch instances B(i) excluding node i, the semantic contrast loss objective can be defined as follows:
[0166]
[0167] where τ represents the temperature parameter and e(·) represents the exponential function. Here, for node i, we can regard its semantic representation as the query instance. The positive samples are the representations of the hyperedge members of node i Meanwhile, the negative samples are the representations of other text nodes excluding node i in the same mini-batch
[0168] Structural contrast loss. Similar to the semantic representation, the contrastive learning algorithm also needs to be applied to the hypergraph structure feature space representation of node i, denoted as Formally, the expression of the structural contrast loss is as follows:
[0169]
[0170] where is the query instance. The positive samples are The negative samples are
[0171] Therefore, our semantic and structural contrastive learning losses ensure that nodes belonging to the same hyperedge have similar representations, thus stimulating useful hypergraph knowledge during the learning process.
[0172] Hypergraph-Text Knowledge Alignment. In this work, our goal is to learn expressive representations while encoding informative text semantics in each text node and structural information between hyperedges. However, contrastive learning only in the semantic or structural space is not sufficient because of the lack of information exchange between the two modalities. To better align the knowledge captured by the semantic and hypergraph structure layers, a hypergraph-text knowledge alignment algorithm is used for TAHG (Text-Associated Hypergraph).
[0173] Specifically, for each node \(i\), based on its semantic and structural representations \(\mathbf{h}_i^s\) and \(\mathbf{h}_i^g\) respectively and the objective of the hypergraph-text knowledge alignment loss \(L_{align}\) is defined as follows:
[0174]
[0175] According to the above expression, in the first part \(L_{s2g}\) of the knowledge alignment loss function align1 we need to regard the structural representation \(\mathbf{h}_i^g\) of node \(i\) as the query, and then construct positive and negative samples based on the semantic representation. Specifically, the positive samples include the representation of node \(i\) and the representations of the nodes in the same hyperedge as \(i\) (i.e., \(\mathbf{h}_{j}^s\) where \(j\in N_i^e\)), and the negative samples are the representations of other instances in the same mini-batch (i.e., \(\mathbf{h}_{k}^s\) where \(k\neq i\)). In the second part \(L_{g2s}\) of the knowledge alignment loss, we need to regard the semantic representation \(\mathbf{h}_i^s\) as the query and construct its corresponding positive and negative samples in the same way as the first part. Therefore, the used \(L_{align}\) loss encourages the representations of the same nodes learned in the two independent feature spaces of semantics and structure to be pulled closer together in the embedding space. as the query, and then construct positive and negative samples based on the semantic representation. Specifically, the positive samples include the representation of node \(i\) and the representations of the nodes in the same hyperedge as \(i\) (i.e., )), and the negative samples are the representations of other instances in the same mini-batch (i.e., ). In the second part \(L_{g2s}\) of the knowledge alignment loss, we need to regard the semantic representation align2 as the query and construct its corresponding positive and negative samples in the same way as the first part. Therefore, the used \(L_{align}\) loss encourages the representations of the same nodes learned in the two independent feature spaces of semantics and structure to be pulled closer together in the embedding space. as the query, and construct its corresponding positive and negative samples in the same way as the first part. Therefore, the used \(L_{align}\) loss encourages the representations of the same nodes learned in the two independent feature spaces of semantics and structure to be pulled closer together in the embedding space. align alignment loss encourages the representations of the same nodes learned in the two independent feature spaces of semantics and structure to be pulled closer together in the embedding space.
[0176] Learning Objectives. To better use the hypergraph-aware language model, we need to jointly optimize the proposed semantic, structural, and hypergraph-text knowledge alignment losses. Specifically, the overall training loss objective is defined as follows:
[0177] \(L=\lambda_1L_{s2g}+\lambda_2L_{g2s}+\lambda_3L_{align}\); semantic +\(\lambda_2L_{g2s}\) structural +\(\lambda_3L_{align}\) align ;
[0178] where \(\lambda_1,\lambda_2,\lambda_3\in\mathbb{R}\). In practice, it can be found that setting \(\lambda_1 = \lambda_2=\lambda_3 = 1\) yields good experimental results.
[0179] Fine-tuning for Downstream Tasks
[0180] Once the pre-training is completed, we need to freeze the parameters of HyperBERT and use it to calculate the representations of each text node. Then, these calculated representations can be further used to fine-tune individual models on different downstream tasks. Specifically, we need to train a linear projection from node features to class labels using cross-entropy loss for node classification.
[0181] After the data processing of the above model, the classification task of medical record texts is completed. The model output result is the classification label of the medical record text, indicating the category to which the text belongs. The following are several classification methods.
[0182] Disease classification: Classify medical record texts into different disease categories, such as heart disease, cancer, diabetes, etc.
[0183] Symptom classification: Classify medical record texts into different symptoms or syndromes, such as cough, fever, headache, etc.
[0184] Treatment plan classification: Classify medical record texts into different treatment plans or treatment methods, such as surgery, drug treatment, radiotherapy, etc.
[0185] Here is a further explanation:
[0186] Named entity recognition is an important technique in the field of natural language processing, aiming to identify the named entities that appear in the text and classify them into predefined categories.
[0187] The goal of named entity recognition is to identify entities with specific meanings from the text. These entities can be specific people, places, organizations, etc., or abstract time, dates, currencies, etc. The basic task of named entity recognition is to find the boundaries of entities in the given text and classify them into predefined categories.
[0188] Named entity recognition technology uses a variety of methods and techniques. One commonly used method is the rule-based method, which matches entities in the text through predefined rules and patterns. These rules can be designed based on information such as part of speech, syntactic structure, context, etc. Another commonly used method is the statistical learning-based method, in which machine learning algorithms are used to learn the features and context information of entities from a large amount of labeled training data, and then these models are applied to new texts for named entity recognition.
[0189] In the process of named entity recognition, the following aspects usually need to be considered:
[0190] (1) Entity boundary recognition: Determine the start and end positions of entities in the text. This usually involves processing information such as part of speech, punctuation marks, spaces, etc.
[0191] (2) Entity Classification: Classify the identified entities into predefined categories, such as person names, place names, organization names, etc. This can be done through rule matching, machine learning algorithms, or deep learning models for classification.
[0192] (3) Context Processing: In the entity recognition process, context information is very important for accurately identifying entities. For example, for a given entity, the context can provide clues as to whether it is a person name or a place name.
[0193] Relation extraction is an important technique in the field of natural language processing, aiming to identify the relationships or associations between entities from text. The goal of relation extraction is to extract the semantic relationships between entities from text, and these relationships can be binary relationships or more complex multi - entity relationships.
[0194] The task of relation extraction is to find and identify the relationships between entities in the given text. These entities can be named entities with specific meanings such as people, places, organizations, etc. The goal of relation extraction is to identify the semantic relationships between entities from text, such as "X is the founder of Y", "X is located in Y", etc.
[0195] Relation extraction techniques usually need to solve the following aspects of problems:
[0196] (1) Entity Recognition: Before relation extraction, it is necessary to first identify the entities in the text and determine their boundaries and categories. Entity recognition techniques can be implemented using rules, statistical learning, or deep learning methods.
[0197] (2) Relation Extraction: After identifying the entities, it is necessary to further identify the relationships between the entities. Relation extraction can be implemented based on rules, machine learning, or deep learning methods. Among them, machine learning and deep learning methods usually require a large amount of labeled training data.
[0198] (3) Context Understanding: In the relation extraction process, context information is very important for accurately identifying relationships. For example, the context can provide information such as the part - of - speech, syntactic structure, and semantics of the sentence where the relationship is located, helping to judge the type and meaning of the relationship.
[0199] 3. Hypergraph
[0200] A hypergraph is an extended form in graph theory and is widely used in the fields of mathematics, computer science, and artificial intelligence. Different from the graphs in traditional graph theory, the edges (called hyperedges) in a hypergraph can connect multiple vertices, not just two vertices. The following is a detailed description of the hypergraph:
[0201] A hypergraph consists of a set of vertices and a set of hyperedges. Vertices represent objects or entities, and hyperedges represent the relationships or connections between objects. Different from the edges in traditional graph theory that connect two vertices, a hyperedge can connect multiple vertices, forming a relationship between a hyperedge and multiple vertices. In a hypergraph, a hyperedge is a set that contains multiple vertices. These vertices are called the endpoints or vertex sets of the hyperedge. The size of a hyperedge can be arbitrary, connecting two vertices or multiple vertices (more than two). A hyperedge can be undirected, indicating no directionality between the connected vertices, or directed, indicating directionality between the connected vertices. Hypergraphs can be used to represent and process complex relationships and connection patterns. Compared with traditional binary relation graphs, hypergraphs can more intuitively represent the association relationships between multiple objects. For example, in a knowledge graph, hypergraphs can be used to represent complex relationships between entities, such as multi - element relationships, attribute relationships, etc. Hypergraphs have a wide range of applications in the fields of computer science and artificial intelligence. In tasks such as knowledge graphs, relation extraction, and recommendation systems, hypergraphs can effectively represent and process complex relationships between entities. Hypergraphs also play an important role in the fields of machine learning and data mining. For example, in graph neural networks, hypergraphs can be used as a representation form of input data for learning and inferring complex relationship patterns.
[0202] Some of the data in the above formula are numerically calculated after removing their dimensions, and the content not described in detail in this specification belongs to the prior art well - known to those skilled in the art.
[0203] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A hypergraph-based electronic medical record text analysis method, characterized in that: The specific steps include: Step 1: Data collection: Determine the source of data collection, automatically collect medical conversation records through the interface or API of the data source, and store the collected data securely in a database or data warehouse; Step 2: Data preprocessing: data cleaning, formatting standards, deduplication, data segmentation, feature extraction, data conversion, data enhancement and sensitive information desensitization; Step 3: Input the preprocessed data into the entity relationship extraction framework to realize the entity relationship recognition and extraction of medical record text, and complete the automatic recording of the diagnosis dialogue system and electronic medical records; Step 4: After the text entity relationship extraction, the identified entities and extracted relationships are input into the node classification model on the hypergraph based on text attributes. According to the characteristics of the medical record text and task requirements, the entities are used as nodes of the hypergraph and the relationships are used as hyperedges of the hypergraph. The results of entity recognition and relationship extraction are integrated into the hypergraph. The hyperedges connect multiple nodes to reflect the association relationship between nodes and complete the text classification task. The entity relationship extraction framework includes vector representation, span representation, entity recognition and relationship extraction, and the entity recognition and relationship extraction are integrated to form a multi-task framework; Vector representation: The vector representation layer uses a fine-tuned pre-trained BERT to extract context information. The input sentence is byte-pair encoded by BPE and split into a sequence of n tokens. The BPE tokens are propagated through the internal BERT to generate an embedding sequence E of length n+1: E = (e [CLS] ,e1,e2,...,e n ), where e [CLS] Represents the embedding representation of the entire sentence context, each embedding vector Where d1 represents the embedding dimension; Span representation: The span representation layer uses the embedding of the tokens generated by BERT to represent the span, maintaining a consistent representation dimension s in different spans: = (e i ,e i+1 ,...,e j ), apply a fusion process to s, using max pooling in the fusion process: The model sets the maximum span length L. In order to reduce complexity, the width embedding matrix is used to merge the width information w into the fused span representation j. Spans of different widths contribute differently to the representation: Among them, ° represents the splicing of vectors, w k Derived from the width embedding matrix obtained from the training model; The multi-task framework includes recognition tasks and classification tasks: Recognition task, the recognition task is used to determine whether the current span or span pair is indeed an entity or whether there is a certain relationship; binary cross entropy is selected as the loss function, which is defined as follows: Among them, q∈{0,1} represents the true category of the span or span pair, where 0 represents a negative sample and 1 represents a positive sample, and p∈[0,1] is the estimated probability of the model identifying q=1 as the true category; Therefore, the loss for the recognition task is calculated as follows: Classification task, in order to enhance the diversity of multi-task learning loss pattern and prevent the two tasks from degenerating into a single task, the pairwise ranking loss function is selected as the loss function of the classification task; this pairwise ranking loss function has the same task goal as other forms of cross entropy loss, but has different optimization directions; Given a span representation r, and a set of entity classes C or a set of relationship classes C, the score of the class label x∈C is calculated using the dot product: score c =r T [W classes ] c ; Where W classes is the matrix to be learned, the number of columns is the number of classes, [W classes ] c is the column vector corresponding to class c, whose dimension is equal to the dimension of r; For a span where y + is the correct label, y - no, and They are defined as + and - The score of ; then the calculation formula of rank loss is: where γ is the scaling factor, m + and m - is the margin. For non-entity classes or irrelevant entity pairs, only L is calculated. - to penalize incorrect predictions; in the training step, the ranking loss selects only the error class with the highest score among all error classes; then, the pairwise ranking loss is optimized: Loss2=∑(L + +L - ); Finally, the loss function of the multi-task framework used is expressed as: Loss=αLoss1+βLoss2.
2. The method for analyzing electronic medical record text based on a hypergraph according to claim 1, characterized in that: The entity relationship extraction framework also includes a joint training loss: a joint training method for entity recognition and relationship extraction, and the joint training loss is defined as follows: L = L1 + L2; Entity Recognition,The entity recognition layer directly uses the output of the span representation layer as input and uses a multi-task framework to extract entity categories from spans; In the recognition task of the entity recognition multi-task framework, the position information between spans is introduced by using the IoU metric; formally, the IoU is calculated as follows: Where s(i,j) represents the span space with left and right indexes i and j in the sentence; Calculate and use the maximum IoU between span t and all entity spans in the sentence, defined as ENIoU(t), which is calculated as: ENIoU(t)=max(IoU(t,en),en∈E entity ); Where E entity represents the set of all entity spans in a sentence; Subsequently, using ENIoU as a scale factor, the loss function of the recognition task in the multi-task framework is replaced by a dynamically scaled cross entropy loss, and the loss function is: Among them, δ and ENIoU determine the scaling factor, δ∈[0,1] is the balancing factor to solve the class imbalance, and δ<0.5 needs to be set to emphasize fewer positive samples; for negative samples, the higher the ENIoU, the greater the loss, so that the model pays more attention to hard negative samples during training; For the rest of the multi-task framework, no further modification is done in the entity recognition stage, and the loss of the entity recognition stage is recorded as L1; Relation extraction, unlike the entity recognition stage, which directly uses the span representation layer output as input, the relation extraction stage makes some adjustments to the input information. For those spans that are classified as "NA" in the entity recognition multi-task stage, they will be filtered out; we only consider candidate span pairs that have been preliminarily identified as entities, select the local information between the candidate span pairs as their context, and obtain the contextual representation by fusing their BERT embeddings using maxpooling; For candidate span pairs of adjacent entities, set v(s1,s2) = 0; In addition, we observe that the Logits in the entity recognition stage may contain useful semantic information; Therefore, Logits can be further embedded from the entity recognition stage to enrich the input information; finally, the multi-task input of the relation extraction layer is represented as follows: Where p(s) represents the Logits of the entity recognition stage; considering that the relationship between s1 and s2 is usually asymmetric, therefore: For the multi-task framework used in the relation extraction stage, the loss of the relation extraction stage is L2.
3. The method for analyzing electronic medical record text based on a hypergraph according to claim 1, characterized in that: The entity relationship extraction framework also includes a prediction stage. In the prediction stage, in order to avoid error propagation, only the class scores obtained by the multi-task framework are used, denoted as score c , while the binary recognition task is only used to optimize the network parameters; Given a span or span pair, the predicted class probability P is calculated as follows: where θ is a hyperparameter threshold; if the score of each class is lower than θ, the span is predicted as a non-entity or irrelevant relation; otherwise, the span is assigned to the class with the highest score.
4. The method for analyzing electronic medical record text based on a hypergraph according to claim 1, characterized in that: The node classification model in step 4 includes: HyperBERT layer, hypergraph-aware pre-training tasks, and fine-tuning of downstream tasks; The HyperBERT layer explicitly mixes the semantic information obtained by each BERT transformer block with the structural information obtained by the Hypergraph neural network block; The hypergraph-aware pre-training task leverages the inherent hypergraph knowledge in TAHG and applies it to semantic and structural representations; Fine-tuning of downstream tasks,After pre-training is completed, freeze the parameters of HyperBERT and use it to calculate the representation of each text node. The calculated representation can be further used to fine-tune separate models on different downstream tasks, and use cross entropy loss to train linear projections from node features to category labels for node classification.
5. The method for analyzing electronic medical record text based on a hypergraph according to claim 4, characterized in that: The nodes are classified into several categories: Disease classification: classify medical record text into different disease categories; Symptom classification: classify medical record text into different symptoms or syndromes; Treatment plan classification: Classify medical record text into different treatment plans or treatment methods.
Citation Information
Patent Citations
Text entity generation method, model training method and device
CN113569572A