Entity recognition method for naming Chinese electronic medical text
Through dual-channel neural network and knowledge graph embedding technology, complex semantic relationships and high-dimensional sparsity problems in Chinese electronic medical texts are solved, and the accuracy and stability of naming entity recognition are improved.
Patent Information
- Application Number
- CN202510534078.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing Chinese electronic medical text naming entity recognition method is difficult to deal with the complex semantic relationships and ambiguity problems of medical terms, and the recognition accuracy is low in the case of high-dimensional sparse text features and data sparseness.
A dual-channel neural network structure is adopted, combined with ClinicalBERT and BiLSTM neural networks, and a knowledge graph embedding method is used to extract text features and entity association representations, and label sequence prediction is performed through conditional random field CRF.
Effectively distinguish complex semantic relationships, improve model stability and named entity recognition performance, and overcome the shortcomings of traditional methods in feature modeling and semantic understanding.
Smart Images

Figure CN120068875A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of entity recognition, and particularly to a method for named entity recognition in Chinese electronic medical texts. Background Art
[0002] Currently, the well-known methods for named entity recognition in Chinese electronic medical texts are usually constructed based on neural network models. Common architectures include using models such as BiLSTM-CRF or Transformer. These methods mainly rely on the context information of the text and the features of the labeled data to complete entity recognition. By inputting a text sequence, the neural network captures the context information, and then combines with a sequence labeling algorithm to output the named entity categories. However, these methods have some limitations: on the one hand, traditional methods lack the utilization of professional knowledge in the medical field and are difficult to handle the complex semantic relationships and polysemy problems of medical terms; on the other hand, due to the existence of a large number of synonyms, abbreviations, and non-standard terms in Chinese electronic medical texts, models that solely rely on text information are difficult to accurately capture the semantic associations between entities; in addition, traditional models usually perform poorly when dealing with the problem of decreased entity recognition accuracy caused by data sparsity and high-dimensional sparsity. Summary of the Invention
[0003] In order to overcome the problems that the existing methods for named entity recognition in Chinese electronic medical texts are difficult to handle high-dimensional sparse text features, dense and diverse terms, complex semantic relationships, and limitations of model characteristics, the present invention proposes a method for named entity recognition in Chinese electronic medical texts based on a dual-channel neural network structure and knowledge graph embedding. This method can not only accurately identify the named entities in electronic medical texts, but also effectively distinguish complex semantic relationships, make full use of medical domain knowledge, improve the stability and recognition performance of the model, and overcome the deficiencies of traditional methods in feature modeling and semantic understanding.
[0004] To achieve the above object, the present invention provides the following technical solution: A method for named entity recognition in Chinese electronic medical texts, comprising the following steps: Step 1: Perform preprocessing operations on the original medical text data and retain medical-related terms to form a medical text corpus; Step 2: Convert the medical text corpus into a word list and a knowledge graph entity ID list. The word list contains word encodings and words; use a word segmentation tool to process the word list , align it with the knowledge graph, extract the entity encodings of the medical corpus text and the knowledge graph entity IDs, and form the corresponding word list and entity table. Use a word vector tool to process the word list The training represents low-dimensional dense word vectors. For the entity IDs in the knowledge graph, entity vectors are generated using an embedding method to form a vocabulary that includes word encodings and word vectors. ; Step 3: Using the vocabulary in Step 2 Convert the text of the preprocessed medical text corpus in Step 1 into a word encoding sequence through encoding. Combine with the list of knowledge graph entity IDs described in Step 2 to convert the entities extracted from the medical text corpus into an entity ID sequence. Then use the vocabulary in Step 2 to further convert the preprocessed medical text corpus in Step 1 into a word encoding sequence and map the entity ID sequence to knowledge graph entity embedding vectors; Step 4: Feed the word encoding sequence and knowledge graph entity embedding vectors obtained in Step 3 into the ClinicalBERT neural network and the BiLSTM neural network respectively to learn the text feature vector and entity association representation; Step 5: Concatenate the output features of the ClinicalBERT neural network and the BiLSTM neural network in Step 4, and combine the text feature vector representation and entity representation as the overall output of the neural network; Step 6: Connect one or more fully connected layers to perform feature transformation on the overall output of the neural network concatenated in Step 5, further extract the interaction information between the text and the entity, and use the conditional random field CRF to predict the probability of the entire label sequence, and finally output the optimal predicted label sequence to obtain the named entity recognition result.
[0005] Preferably, in Step 1, the preprocessing operations include word segmentation, stop word removal, and removal of irrelevant characters; in Step 2, for the original medical text data, first use the Jieba word segmentation module to perform word segmentation on the text sequence in the exact mode. After the word segmentation task is completed, traverse the word segmentation results in combination with the stop word list to remove the stop words and form a medical text corpus.
[0006] Preferably, in Step 2, use the word2vec word vector tool to train the vocabulary ; use the skip-gram model to train the words to represent low-dimensional dense word vectors to form a vocabulary , including word numbers and word vectors.
[0007] Preferably, in Step 3, for any sentence in the medical text corpus , in combination with the said vocabulary and the vocabulary including word encodings and word vectors , obtain the sentence under the conversion of the vocabulary to a word encoding sequence , and then in combination with the knowledge graph entity table Under the transformation of , where is a word, is the corresponding knowledge graph entity embedding vector.
[0008] Preferably, in step 4, the ClinicalBERT neural network adopts a multi-head attention mechanism to calculate the input of the word encoding sequence , and calculates the attention weight matrix, which is specifically expressed as: ; where , , respectively refer to the query vector Query, the key vector Key and the value vector Value, refers to the dimensions of the query, key and value vectors, and uses to perform dot product operations on each key vector, and after scaling the operation results, the softmax function is used for normalization to obtain the attention weight matrix.
[0009] Preferably, in step 5, the multi-head attention mechanism is used to weight the features of the input sequence and calculate the attention weight scores; Specifically, the multi-head attention mechanism captures different feature representations simultaneously by using multiple parallel attention heads: ; where represents the output result of the th attention head, is the attention score calculation formula, , , respectively represent the query vector, key vector and value vector of the th attention head.
[0010] Each head calculates self-attention in different subspaces so as to focus on different input information from multiple perspectives: ; where refers to the output result of the multi-head attention mechanism, is the vector concatenation operation, to represent the output results of different attention heads.
[0011] Finally, the heads are concatenated and finally the feature matrix output integrating attention information is obtained after linear transformation: ; Among them, represents the output of the feature matrix of the finally obtained fused attention information, represents the output result of the multi-head attention mechanism, is the output weight matrix; For the knowledge graph entity embedding vector , the TransE method assumes that there is a triple in the knowledge graph, and maps the head entity, relation, and tail entity to a continuous vector space through translation, so that the triple satisfies the following relationship: ; Among them is the vector representation of the head entity, is the vector representation of the relation, is the vector representation of the tail entity; specifically, the TransE method models the relation in the knowledge graph as the displacement between the head entity and the tail entity through translation. For a correct triple , the TransE method aims to minimize the distance function : ; Among them is the norm or norm of the vector, the smaller the value of, the more effective the triple is; The TransE method uses a margin-based ranking loss function for optimization, and its formula is as follows: ; Among them represents the set of positive samples, that is, the correct triples; represents the set of negative samples, and this set is usually generated by randomly replacing or ; is the Margin hyperparameter, which controls the distance between positive and negative samples; is a ReLU operation.
[0012] Preferably, in step 6, the BiLSTM neural network adopts a bidirectional LSTM neural network. For the word encoding sequence and the knowledge graph entity embedding vector , the unidirectional LSTM neural network can obtain the output through text feature training, then the output of the BiLSTM neural network is obtained by splicing the outputs of the bidirectional LSTM neural network: ; Among them and respectively represent the feature outputs extracted by the forward and backward LSTM neural networks.
[0013] The conditional random field CRF is used to predict the label probabilities of each text segment. For the word encoding sequence and the knowledge graph entity embedding vector , for any in the sentence of the medical text corpus , calculate the sentence The calculation formula for the label probability to which it belongs can be expressed as: ; where, where represents the probability that the output sequence appears under the condition of the given input sequence , is the normalization factor, is the exponential function, is to sum over the sequence length , is to sum over the number of feature functions , is the feature function weight parameter, represents the th feature function, where is the current input, and respectively represent the feature functions of the previous label and the current label.
[0014] The present invention has the following advantages: The present invention integrates a medical knowledge graph to enhance semantic information, combines a named entity recognition task, and solves the problems of missing sentence components and complex semantic relationships in electronic medical texts; adopts a dual-channel neural network structure, uses ClinicalBERT to extract local features, combines BiLSTM to extract global context information, and fuses features through an attention mechanism to highlight key feature words; By preprocessing the electronic medical text dataset and training word vectors, the text sequence is converted into a dense low-dimensional word vector, the input feature representation is optimized, the influence of high-dimensional sparsity on the model is reduced, and the stability, robustness, and named entity recognition effect of the model are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a specific implementation flowchart provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0017] The specific implementation process of this embodiment is as follows: As Figure 1 shown, first, collect and construct a dataset of original Chinese electronic medical texts (electronic medical records are used in this embodiment). The experimental dataset comes from the electronic medical text dataset CCKS2019 released by the China Knowledge Graph and Semantic Computing Conference in 2019. After desensitizing the dataset, it is labeled. The dataset is labeled with 6 entity types, namely diseases and diagnoses (Dis), drugs (Drug), operations (Operation), imaging examinations (ImgExam), anatomical parts (Anatomy), and laboratory examinations (LabExam). The CCKS2019 dataset contains 1379 data, of which 1000 data are used for training and 379 data are used for testing.
[0018] For the original electronic medical text dataset, first use the Jieba word segmentation module to segment the text sequence in the exact mode. After the word segmentation task is completed, traverse the word segmentation results in combination with the stop word list to remove the stop words and form a medical text corpus.
[0019] Convert the medical text corpus into a vocabulary , including vocabulary encoding and words, and use the word2vec word vector tool to train the vocabulary , with the default skip-gram model, represent the words as low-dimensional dense word vectors through training, and form a vocabulary , which contains word numbers and word vectors.
[0020] For any sentence in the medical text corpus in this embodiment , in combination with the described vocabulary and the vocabulary containing word encoding and word vectors , obtain under the transformation of the vocabulary as a word encoding sequence , and further obtain the corresponding knowledge graph entity embedding vector under the transformation of combining the knowledge graph entity table , where is a word, is the corresponding knowledge graph entity embedding vector.
[0021] The ClinicalBERT neural network in this embodiment adopts a multi-head attention mechanism to calculate the input of the word encoding sequence and obtains an attention weight matrix, which is specifically expressed as: ; Among them, , , respectively refer to the query vector (Query), key vector (Key), and value vector (Value), refers to the dimensions of the query, key, and value vectors. Use to perform a dot product operation on each key vector, and after scaling the operation result, use the softmax function for normalization to obtain the attention weight matrix.
[0022] This embodiment uses a multi-head attention mechanism to perform feature weighting on the input sequence and calculates the attention weight score. Specifically, the multi-head attention mechanism simultaneously captures different feature representations by using multiple parallel attention heads: ; Among them represents the output result of the th attention head, is the attention score calculation formula, , , respectively represent the query vector, key vector, and value vector of the th attention head.
[0023] Each head calculates self-attention in different subspaces so that the model can focus on different input information from multiple perspectives: ; Among them refers to the output result of the multi-head attention mechanism, is the vector concatenation operation, to represent the output results of different attention heads.
[0024] Finally, the heads are concatenated and finally obtain the feature matrix output that fuses attention information after linear transformation: ; Among them, represents the finally obtained feature matrix output that fuses attention information, represents the output result of the multi-head attention mechanism, is the output weight matrix.
[0025] In this embodiment, the TransE method is used to embed the knowledge graph. For the embedding vectors of the knowledge graph entities , TransE assumes that there is a triple in the knowledge graph . Its core idea is to map the head entity, relation, and tail entity to a continuous vector space through translation, so that the triple satisfies the following relationship: ; where is the vector representation of the head entity, is the vector representation of the relation, is the vector representation of the tail entity. Specifically, TransE models the relation in the knowledge graph as the displacement between the head entity and the tail entity through translation. For a correct triple , the TransE model aims to minimize the distance function : ; where is the norm (sum of absolute values) or norm (Euclidean distance) of the vector. The smaller the value of , the more valid the triple is.
[0026] The core goal of the TransE method adopted in this embodiment is to make the embedding vectors of the correct triples close through optimizing the loss function, while the distance of the wrong triples, that is, the negative samples, is far away, so as to retain the structural information in the knowledge graph. TransE uses a margin-based ranking loss function for optimization, and its formula is as follows: ; where represents the set of positive samples, that is, the correct triples; represents the set of negative samples, and this set is usually generated by randomly replacing or ; is the Margin hyperparameter, which controls the spacing between positive and negative samples; is a ReLU operation, which ensures that the loss is non-negative.
[0027] The BiLSTM neural network in this embodiment uses a bidirectional LSTM neural network. For the word encoding sequence and the knowledge graph entity embedding vectors , the unidirectional LSTM neural network can obtain the output by performing text feature training, then the output of the BiLSTM neural network Obtained by concatenating the outputs of the bidirectional LSTM neural network: ; where and respectively represent the feature outputs extracted by the forward and backward LSTM neural networks.
[0028] In this embodiment, the conditional random field CRF is used to predict the label probabilities of each text segment and output the optimal predicted label sequence; for the word encoding sequence and the knowledge graph entity embedding vector , for any in the sentence in the original corpus , the overall formula for the model to calculate the label probability of the sentence can be expressed as: ; where, represents the probability of the output sequence appearing given the input sequence , is the normalization factor, is the exponential function, is the sum over the sequence length , is the sum over the number of feature functions , is the weight parameter of the feature function , represents the th feature function, where is the current input, and respectively represent the feature functions of the previous label and the current label.
[0029] In this embodiment, the experiment uses python3.8.19, the pytorch2.3.1 framework, and cuda11.8 to train the model; the precision (Precision, P), recall (Recall, R), and F1 value are used as the indicators to evaluate the performance of the named entity recognition model: ; ; ; where, represents the number of positive samples correctly recognized; represents the number of negative samples misrecognized as positive samples; represents the number of positive samples misrecognized as negative samples. represents the proportion of the correctly predicted results in all the predicted results; Represents the proportion of correctly predicted results among all data; is and the harmonic mean of.
[0030] To verify the effectiveness of the method proposed in this embodiment, five groups of comparative experiments were set up: (1) BiLSTM-CRF: The BiLSTM-CRF model combines a bidirectional LSTM network with a conditional random field, organically integrating context information and global features, which helps to reduce the risk of error propagation in sentence processing; (2) Lattice-LSTM: The Lattice-LSTM model introduces a lattice structure to incorporate lexical information into word vectors, enabling it to capture both character-level and word-level features simultaneously, demonstrating multi-granularity feature expression capabilities; (3) FLAT: By transforming the lattice structure into a set of spans and introducing specific position information encoding, the expression ability of the model is enhanced; (4) BERT-BiLSTM-CRF: BERT can perform deep learning in a large-scale corpus to capture rich context semantic information. Different from traditional BiLSTM that relies on the context relationship of local text sequences, BERT provides more comprehensive and accurate semantic expressions, further strengthening the modeling ability of long-term dependencies in sequences; (5) The classification method of this embodiment After multiple rounds of experiments and cross-validation of the experimental results, the model evaluation results of various methods are shown in the following table. Table 1 Named entity recognition results of models of five different methods (unit: %)
[0031] From the experimental results in the above table, it can be concluded that the named entity recognition method of this embodiment has achieved the best results in the evaluation index results, and thus the superiority of the named entity recognition method of this embodiment in the task of named entity recognition in Chinese electronic medical texts can be obtained.
[0032] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A method for entity recognition of named Chinese electronic medical texts, characterized by: The following steps are involved: Step 1: Preprocess the original medical text data and retain medical-related terms to form a medical text corpus; Step 2: Convert the medical text corpus into a vocabulary and knowledge graph entity ID list, vocabulary Contains vocabulary codes and words; uses word segmentation tools to process the vocabulary , align with the knowledge graph, extract the entity encoding of the medical corpus text and the knowledge graph entity ID, and form the corresponding vocabulary and entity table; use the word vector tool to transform the word table The training is represented as a low-dimensional dense word vector, and the entity vector is generated by the embedding method for the knowledge graph entity ID to form a word table containing word encoding and word vector ; Step 3: Using the vocabulary from step 2 The medical text corpus text preprocessed in step 1 is converted into a word encoding sequence through encoding, and the entities extracted from the medical text corpus are converted into entity ID sequences by combining the knowledge graph entity ID list described in step 2. The preprocessed medical text corpus in step 1 is further converted into a word encoding sequence, and the entity ID sequence is mapped into a knowledge graph entity embedding vector; Step 4: Use a dual-channel neural network structure to learn text feature vectors and entity association representations; Step 5: Concatenate the output features in step 4, combine the text feature vector representation and the entity representation, as the overall output of the neural network; Step 6: Connect one or more fully connected layers to perform feature conversion on the overall output of the neural network concatenated in step 5, and predict the probability of the entire label sequence, and finally output the predicted optimal label sequence to obtain the named entity recognition result.
2. The entity recognition method for named Chinese electronic medical text according to claim 1, characterized in that: In step 1, the preprocessing operations include word segmentation, removal of stop words, and removal of irrelevant characters; in step 2, for the original medical text data, the Jieba word segmentation module is first used to perform word segmentation on the text sequence in precise mode. After the word segmentation task is completed, the word segmentation results are traversed in combination with the stop word list, and stop words are removed to form a medical text corpus.
3. The method for entity recognition of Chinese electronic medical text naming according to claim 1, characterized in that: In step 2, use the word2vec word vector tool to train the vocabulary ; Use the skip-gram model to represent word training as a low-dimensional dense word vector to form a vocabulary , containing word numbers and word vectors.
4. The method for entity recognition of Chinese electronic medical text naming according to claim 1, characterized in that: Step 3: For any sentence in the medical text corpus , combined with the vocabulary and a vocabulary containing word encodings and word vectors , get the sentence In the vocabulary The conversion is a word encoding sequence , combined with the knowledge graph entity table Under the transformation, the corresponding knowledge graph entity embedding vector is further obtained ,in, It's a word. is the corresponding knowledge graph entity embedding vector.
5. The method for entity recognition of Chinese electronic medical text naming according to claim 1, characterized in that: Step 4 specifically includes: using the word encoding sequence and knowledge graph entity embedding vector obtained in step 3 to be respectively fed into the ClinicalBERT neural network and the BiLSTM neural network to learn the text feature vector and entity association representation; ClinicalBERT neural network uses a multi-head attention mechanism to encode word sequences The input is calculated to obtain the attention weight matrix, which is specifically expressed as: ; in, , , They refer to the query vector Query, the key vector Key and the value vector Value respectively. Refers to the dimensions of the query, key, and value vectors using Perform a dot product operation on each key vector, scale the operation result, and then normalize it using the softmax function to obtain the attention weight matrix.
6. The method for entity recognition of Chinese electronic medical text naming according to claim 1, characterized in that: In step 5, a multi-head attention mechanism is used to Perform feature weighting and calculate attention weight score; Specifically, the multi-head attention mechanism uses multiple parallel attention heads to simultaneously capture different feature representations: ; in Indicates The output of the attention head is: is the attention score calculation formula, , , Respectively represent The query vector, key vector, and value vector of the attention heads; Each head calculates self-attention in different subspaces to focus on different input information from multiple perspectives: ; in refers to the output of the multi-head attention mechanism, is a vector concatenation operation, to It represents the output results of different attention heads; Finally After linear transformation of the head concatenation, the feature matrix output of the fused attention information is finally obtained: ; in Represents the final feature matrix output of the fused attention information, Represents the output of the multi-head attention mechanism, is the output weight matrix.
7. The method for entity recognition of Chinese electronic medical text naming according to claim 1, characterized in that: In step 6, the BiLSTM neural network uses a bidirectional LSTM neural network to encode the word sequence and knowledge graph entity embedding vector , the unidirectional LSTM neural network is trained with text features to obtain the output , then the output of the BiLSTM neural network is The output of the bidirectional LSTM neural network is concatenated: ;in and Respectively represent the feature outputs extracted by the forward and backward LSTM neural networks; Conditional random field CRF is used to predict the label probability of each text segment. For the word encoding sequence and knowledge graph entity embedding vector , for any Sentences in the medical text corpus , calculate the sentence The calculation formula of the label probability can be expressed as: ; Among them, Indicates that given an input sequence In the case of The probability of occurrence, is the normalization factor, is an exponential function, is the length of the sequence To sum, is the number of characteristic functions To sum, is the characteristic function The weight parameter, Indicates characteristic functions, where For the current input, and Represent the feature functions of the previous label and the current label respectively.
Citation Information
Patent Citations
A knowledge graph embedding method based on a diverse graph attention mechanism
CN109902183A
Knowledge graph and BERT fused Chinese medical named entity recognition method and device
CN112487202A
Auxiliary diagnosis and treatment system and method based on knowledge graph
CN118471486A
Dynamic reasoning method and system based on language model and knowledge graph
CN119150960A
Medical text named entity recognition method and system
CN119514545A