Chinese medicine named entity recognition method and device based on multi-feature fusion
By designing the BERT-wwm-BiLSTM-TextCNN-MHSA-ATVCRF model, combining multi-feature fusion and self-attention mechanism, the problems of complex context dependencies and blurred entity boundaries in Chinese medical texts are solved, efficient naming entity recognition is achieved, accuracy and recall rate are improved, and inference speed is improved.
Patent Information
- Application Number
- CN202411926197.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-23
AI Technical Summary
The context dependencies present in Chinese medical texts are too complex, the entity boundaries are blurred, and the time complexity of existing algorithms lead to inefficient recognition of named entities.
Using a multi-feature fusion method, the BERT-wwm-BiLSTM-TextCNN-MHSA-ATVCRF model is designed, and the Chinese-BERT-wwm word embedding is generated, context features are extracted, TextCNN extracted local features, and feature fusion is performed through multi-head self-attention mechanism. Finally, ATVCRF is used for sequence annotation, reducing time complexity and memory usage.
It effectively solves the problems of complex context dependencies and blurred entity boundaries in Chinese medical texts, improves the accuracy and recall of named entity recognition, and significantly improves the inference speed on large-scale data sets.
Smart Images

Figure CN120031035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Chinese medical natural language processing, and in particular to a Chinese medical named entity recognition method and device based on multi-feature fusion. Background Art
[0002] Named Entity Recognition (NER) is a key component in information extraction tasks. It aims to identify and classify entities with specific meanings from unstructured texts, providing important support for downstream tasks such as text classification, sentiment analysis, and question-answering systems. The application of named entity recognition in the field of Chinese medicine helps to automatically identify and classify key entities in medical literature, research reports, and electronic medical records, such as disease names, drugs, symptoms, etc., thereby improving the efficiency of information retrieval and data analysis. This can not only assist clinicians in quickly obtaining key information and supporting clinical decision-making, but also help researchers analyze large amounts of medical data and discover new medical knowledge and treatments.
[0003] Named entity recognition is essentially a sequence labeling task. Identifying and classifying entities means assigning a label to each element (usually a word or character) in the text to identify whether the element belongs to a certain named entity and determine its role in the entity.
[0004] When using Chinese medical text for named entity recognition tasks, many problems are faced. The meaning of medical terms often depends on the context. For example, "liver disease" and "liver function test" have different interpretations and focuses. If the context dependency information cannot be handled properly, it may affect the accuracy of entity type recognition. The entity boundaries in Chinese medical texts are usually complex and may contain complex combinations of terms, such as "chronic myeloid leukemia", "acute pulmonary thromboembolism", etc. If the start and end boundaries of the entity cannot be accurately identified, the entity may not be fully captured or partially overlap. Therefore, the context dependencies in Chinese medical texts are too complex and the entity boundaries are blurred.
[0005] In addition, in practical applications, Chinese medical named entity recognition tasks usually involve huge amounts of data. The Viterbi algorithm used in the inference stage of the current mainstream sequence labeling model CRF has a high computational complexity. When processing long sequences and large-scale data sets, the inference speed will be greatly affected, and the inference efficiency is low, which further increases the challenge of the task. Summary of the invention
[0006] In order to solve the technical problems in the prior art that the context dependency relationship in Chinese medical texts is too complex, the entity boundary is fuzzy, and the time complexity of the algorithm is high, the embodiment of the present invention provides a Chinese medical named entity recognition method and device based on multi-feature fusion. The technical solution is as follows:
[0007] On the one hand, a Chinese medical named entity recognition method based on multi-feature fusion is provided, the method is implemented by a Chinese medical named entity recognition device based on multi-feature fusion, and the method includes:
[0008] S1. Obtaining a Chinese medical text to be recognized and a Chinese medical named entity recognition model, wherein the Chinese medical named entity recognition model includes an embedding layer, a feature extraction layer, a feature fusion layer, and a labeling layer;
[0009] S2. Based on the Chinese medical text and the embedding layer, obtain a word embedding of the Chinese medical text;
[0010] S3, based on the word embedding of the Chinese medical text and the feature extraction layer, obtaining context features and local features of the Chinese medical text;
[0011] S4, obtaining fused features based on the context features and local features of the Chinese medical text and the feature fusion layer;
[0012] S5. Based on the fused features and the annotation layer, the named entities corresponding to the Chinese medical text are obtained and classified to complete the recognition of the Chinese medical named entities.
[0013] On the other hand, a Chinese medical named entity recognition device based on multi-feature fusion is provided, and the device is applied to a Chinese medical named entity recognition method based on multi-feature fusion, and the device comprises:
[0014] An acquisition unit, used to acquire Chinese medical text to be recognized and a Chinese medical named entity recognition model, wherein the Chinese medical named entity recognition model includes an embedding layer, a feature extraction layer, a feature fusion layer, and a labeling layer;
[0015] An embedding unit, configured to obtain a word embedding of the Chinese medical text based on the Chinese medical text and the embedding layer;
[0016] An extraction unit, configured to obtain context features and local features of the Chinese medical text based on the word embedding of the Chinese medical text and the feature extraction layer;
[0017] A fusion unit, used for obtaining fused features based on the context features and local features of the Chinese medical text and the feature fusion layer;
[0018] An identification unit, configured to obtain the named entities corresponding to the Chinese medical text and classify them based on the fused features and the annotation layer, so as to complete the identification of Chinese medical named entities.
[0019] On the other hand, a Chinese medical named entity recognition device based on multi-feature fusion is provided. The Chinese medical named entity recognition device based on multi-feature fusion includes: a processor; a memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned Chinese medical named entity recognition method based on multi-feature fusion is implemented.
[0020] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned Chinese medical named entity recognition method based on multi-feature fusion.
[0021] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0022] The embodiments of the present invention propose a Chinese medical named entity recognition method based on multi-feature fusion, and design a BERT-wwm-BiLSTM-TextCNN-MHSA-ATVCRF model. First, use Chinese-BERT-wwm to process the input text sequence to generate word embeddings for subsequent processing. Considering the complex context dependencies in the text, use BiLSTM to extract context features from the word embeddings, effectively avoiding the problem of incomplete context modeling; considering that a large number of entities are composed of complex terms with complex boundaries, use TextCNN to extract a large number of local features from the word embeddings through multiple convolutional kernels, which helps to identify and distinguish entities with complex boundaries. Then, fuse the context features and local features through the multi-head self-attention mechanism to obtain a richer and more detailed feature representation, effectively solving the problems of overly complex context dependencies and fuzzy entity boundaries in Chinese medical texts. Use ATVCRF for sequence annotation, and prune the states with lower scores through an adaptive threshold, reducing the time complexity and actual memory occupancy of the algorithm, improving the inference speed of the model, and having higher efficiency when processing large-scale Chinese medical datasets. Finally, output a legal and globally optimal label sequence to complete the Chinese medical named entity recognition task. The method proposed in the embodiments of the present invention can effectively improve the precision and recall rate of the Chinese medical named entity recognition task, and has a higher inference speed on large-scale datasets. Description of the Drawings
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 It is a flow chart of a Chinese medical named entity recognition method based on multi-feature fusion provided by an embodiment of the present invention;
[0025] Figure 2 It is a structural schematic diagram of a Chinese medical named entity recognition model provided by an embodiment of the present invention;
[0026] Figure 3 It is a block diagram of a Chinese medical named entity recognition device based on multi-feature fusion provided by an embodiment of the present invention;
[0027] Figure 4 It is a structural schematic diagram of a Chinese medical named entity recognition device based on multi-feature fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0029] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0030] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0031] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.
[0032] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0033] The embodiment of the present invention provides a Chinese medical named entity recognition method based on multi-feature fusion, which can be implemented by a Chinese medical named entity recognition device based on multi-feature fusion, and the Chinese medical named entity recognition device based on multi-feature fusion can be a terminal or a server. Figure 1 The flowchart of the Chinese medical named entity recognition method based on multi-feature fusion is shown. The processing flow of the method may include the following steps:
[0034] S1. Obtain Chinese medical text to be recognized and a Chinese medical named entity recognition model, where the Chinese medical named entity recognition model includes an embedding layer, a feature extraction layer, a feature fusion layer, and a labeling layer.
[0035] In a feasible implementation, Figure 2 As shown in the model architecture diagram, in the embodiment of the present invention, the text "can also be converted into acute Aspergillus pneumonia" is used as an example of a feasible Chinese medical text to be identified.
[0036] S2. Based on the Chinese medical text and the embedding layer, the word embedding of the Chinese medical text is obtained.
[0037] Among them, the embedding layer can be a Chinese-BERT-wwm model.
[0038] Optionally, the specific operation of S2 may include the following steps S21-S25:
[0039] S21, inputting the Chinese medical text into the embedding layer, and segmenting the Chinese medical text into subwords by using the word segmentation tool in the embedding layer;
[0040] S22, according to a preset vocabulary, convert the subword into an ID sequence, and map the ID sequence to a high-dimensional vector space to obtain a subword vector;
[0041] S23, determining the position embedding and type embedding of the subword according to the ID sequence;
[0042] S24, generating preliminary word embedding information according to the subword vector and the position embedding and type embedding of the subword;
[0043] S25. Input the preliminary word embedding information into the multi-layer Transformer encoder, fuse the context information through the self-attention mechanism, and obtain the word embedding of Chinese medical text.
[0044] In a feasible implementation, the main task of the embedding layer is to map each word in the Chinese medical text into a real vector space of a specific dimension using word embedding technology. With the help of these dense vector representations, the model can capture the semantic features and contextual information of these words, thereby effectively processing and understanding the complex relationships in Chinese medical text. Therefore, the quality of the word embedding generated by the embedding layer directly determines the accuracy of the model in identifying and classifying Chinese medical text entities.
[0045] The embodiment of the present invention uses Chinese-BERT-wwm when generating word embeddings, a pre-trained model based on Whole Word Masking (WWM) technology released by the Harbin Institute of Technology iFlytek Joint Laboratory. In the pre-training process of conventional masked language models (such as BERT), some subwords or word fragments in the sequence will be randomly masked, and the model learns contextual information by predicting these masked parts. However, conventional masked language models cannot guarantee that the entire word is masked at the same time, and may only mask part of a word, which will lead to problems such as insufficient understanding of the model for the complete vocabulary, semantic ambiguity, and decreased learning efficiency. Chinese-BERT-wwm uses WWM technology. When a word is selected for masking, WWM ensures that all subwords of this word are masked at the same time. This strategy can better capture vocabulary-level information and improve the performance of the model in processing complete vocabulary.
[0046] When generating word embeddings for Chinese medical texts, the processing flow of Chinese-BERT-wwm is basically the same as that of BERT. First, the input text sequence is segmented into subwords through a word segmentation tool; then, these subwords are converted into ID sequences according to the vocabulary and mapped to a high-dimensional vector space; next, these vectors are combined with position embeddings (which provide position information of subwords in the sequence) and type embeddings (used to distinguish different sentences or fragments) to generate word embedding information; then, this information is processed through a multi-layer Transformer encoder, and the context information is fused using the self-attention mechanism to obtain the final word embedding. To facilitate the following discussion, the input text sequence is now set to , the word embedding obtained after processing by the embedding layer is .in, represents the tth subword in the input text sequence, Representatives and Corresponding to word embedding, n represents the total number of subwords.
[0047] S3. Based on the word embedding and feature extraction layer of Chinese medical text, the contextual features and local features of Chinese medical text are obtained.
[0048] In a feasible implementation, Chinese medical texts often contain complex contextual dependencies and a large number of entities with fuzzy boundaries. If the word embedding generated by the embedding layer is directly used for subsequent sequence labeling without feature extraction, it may lead to problems such as incomplete context modeling and excessive redundant information, thereby affecting the accuracy of the model in identifying and classifying entities. Therefore, the feature extraction layer of an embodiment of the present invention includes BiLSTM and TextCNN. BiLSTM is used in the feature extraction layer to extract contextual features to effectively capture long-distance contextual features; TextCNN is used to extract local features to effectively identify and distinguish entities with fuzzy boundaries. The embodiment of the present invention helps to obtain a richer feature representation by extracting features from two different angles, thereby effectively improving the performance of the model.
[0049] Optionally, BiLSTM may include multiple LSTM networks, each LSTM network includes an input gate, a forget gate, an output gate and a memory unit; two LSTM networks, namely forward LSTM and backward LSTM, each LSTM includes an input gate, a forget gate, an output gate and a memory unit; TextCNN may include multiple convolutional layers with convolution kernels of different sizes.
[0050] The specific operation of S3 may include the following steps S31-S33:
[0051] S31. According to the word embedding of the Chinese medical text and an adjacent hidden state, the activation value of the input gate, the activation value of the forget gate, the candidate memory unit state and the activation value of the output gate are calculated respectively. According to the activation value of the input gate and the activation value of the forget gate, the candidate memory unit state is updated to obtain the current memory unit state. According to the activation value of the output gate and the current memory unit state, the current hidden state is determined.
[0052] It should be noted that the multiple LSTM networks included in BiLSTM are divided into two types of LSTM networks, namely the forward LSTM network (i.e. Figure 2 The F-LSTM in the Figure 2 LSTM is a variant of recurrent neural network, which aims to solve the gradient vanishing and gradient exploding problems encountered by traditional recurrent neural network when processing long sequence data. Considering that a single direction may not be enough to capture relevant information when LSTM processes Chinese medical texts with complex sequence patterns and dependencies, the embodiment of the present invention uses BiLSTM instead of LSTM to extract contextual features in Chinese medical texts.
[0053] BiLSTM is a model that is further extended from LSTM. In BiLSTM, the word embedding X is passed to two LSTMs at the same time: one from the beginning to the end of the sequence (forward), and the other from the end to the beginning of the sequence (backward). LSTM introduces three gating mechanisms (input gate , Forget Gate , output gate ) and a memory unit To selectively retain or discard information, specifically, the input gate determines how much input information is stored in the memory cell, the forget gate determines how much of the previous memory cell state is retained, the memory cell completes its own update based on the information of the candidate memory cell state, the input gate, and the forget gate, and the output gate determines how much of the memory cell state will be output as the current hidden state, thereby effectively capturing long-distance dependencies. The calculation process of LSTM is shown in the following equations (1) to (12), where, represents the sigmoid activation function, tanh represents the hyperbolic tangent activation function, represents element-level multiplication, W represents the corresponding weight matrix, and b represents the corresponding bias term. and Represent the word embedding and hidden state of the current time step (in this paper, time step t represents the t-th subword unit).
[0054] The following are two examples of LSTM networks:
[0055] In the first aspect, when the LSTM network is a forward LSTM network, the specific operation of S31 may include the following steps S31a1-S31a6:
[0056] S31a1. Calculate the forward activation value of the input gate based on the word embedding of the Chinese medical text, the previous hidden state, and formula (1). :
[0057]
[0058] in, represents the sigmoid activation function, represents the weight matrix of the input gate, represents the bias term of the input gate, represents the word embedding at the current time step t, Represents the previous hidden state of the forward LSTM; You can decide how much input information will be stored in the memory cell.
[0059] S31a2, based on the word embedding of Chinese medical text, the previous hidden state and formula (2), calculate the forward activation value of the forget gate :
[0060]
[0061] in, represents the weight matrix of the forget gate, Represents the bias term of the forget gate; Can determine how many previous memory cell states there are will be retained.
[0062] S31a3, based on the word embedding of Chinese medical text, the previous hidden state and formula (3), calculate the forward candidate memory unit state :
[0063] (3)
[0064] Among them, tanh represents the hyperbolic tangent activation function, represents the weight matrix of the memory unit, Represents the bias term of the memory unit;
[0065] S31a4, according to the forward activation value of the input gate, the forward activation value of the forget gate and formula (4), the forward candidate memory unit state is updated to obtain the forward current memory unit state :
[0066] (4)
[0067] in, represents element-wise multiplication, Represents the state of the forward memory unit at the previous time step;
[0068] S31a5 calculates the forward activation value of the output gate based on the forward word embedding of the Chinese medical text, the previous hidden state, and equation (5) :
[0069] (5)
[0070] in, represents the weight matrix of the output gate, Represents the bias term of the output gate; You can decide how many memory cell states will be output as the current hidden state.
[0071] S31a6, according to the forward activation value of the output gate, the forward current memory unit state and formula (6), determine the forward current hidden state :
[0072] (6).
[0073] Second, when the LSTM network is a backward LSTM, the specific operations of S31 may include S31b1-S31b6:
[0074] S31b1. Calculate the backward activation value of the input gate based on the word embedding of the Chinese medical text, the subsequent hidden state, and formula (7). :
[0075] (7)
[0076] in, represents the sigmoid activation function, represents the weight matrix of the input gate, represents the bias term of the input gate, represents the word embedding at the current time step t, Represents the next hidden state of the backward LSTM;
[0077] S31b2, based on the word embedding of Chinese medical text, the latter hidden state and formula (8), calculate the backward activation value of the forget gate :
[0078] (8)
[0079] in, represents the weight matrix of the forget gate, Represents the bias term of the forget gate;
[0080] S31b3, based on the word embedding of Chinese medical text, the next hidden state and formula (9), calculate the backward candidate memory unit state :
[0081] (9)
[0082] Among them, tanh represents the hyperbolic tangent activation function, represents the weight matrix of the memory unit, Represents the bias term of the memory unit;
[0083] S31b4, according to the backward activation value of the input gate, the backward activation value of the forget gate and formula (10), the backward candidate memory unit state is updated to obtain the backward current memory unit state :
[0084] (10)
[0085] in, represents element-wise multiplication, Represents the state of the backward memory unit at the previous time step;
[0086] S31b5, based on the backward word embedding of the Chinese medical text, the subsequent hidden state and formula (11), calculate the backward activation value of the output gate :
[0087] (11)
[0088] in, represents the weight matrix of the output gate, Represents the bias term of the output gate;
[0089] S31b6, according to the backward activation value of the output gate, the backward current memory unit state and formula (12), determine the backward current hidden state :
[0090] (12).
[0091] S32: Determine context features according to the current hidden state.
[0092] Optionally, the specific operation of S32 may be as follows:
[0093] Pass the current hidden state forward and backward to the current hidden state The spliced features are then concatenated and determined as context features.
[0094] In one possible implementation, the context feature can be used express, The calculation form is as follows:
[0095]
[0096] S33. Input the word embedding of the Chinese medical text into the convolution layer of the convolution kernels of different sizes to obtain a set of feature vectors. Based on the set of feature vectors and the maximum pooling operation, multiple pooling features are obtained. The multiple pooling features are concatenated to obtain the local features of the Chinese medical text.
[0097] In a feasible implementation, the embodiment of the present invention uses TextCNN to extract local features from word embeddings through convolution kernels of different sizes to enhance the ability to capture local information in Chinese medical texts and more effectively identify and classify entities with complex boundaries. Now assume that m convolution kernels are used when extracting local features, and the specific processing process is as follows:
[0098] First, the word embedding is input into the convolutional layer of TextCNN, and different convolution kernels are used according to the following formula (i.e. Figure 2The kernels with different subscripts in the convolution operation are used to perform convolution operations on word embeddings.
[0099]
[0100] in, represents the selected convolution kernel, Represents the size of the convolution kernel, It is the tth subword to the The word embedding of subwords, b is a bias term introduced to enhance the expressiveness and learning flexibility of the model; ReLU is an activation function used to introduce nonlinearity; is the convolution kernel The feature vector generated by a single convolution operation. After each convolution kernel completes the corresponding convolution operation, all the vectors obtained can be stacked into a feature map. That is, the convolution kernel The set of feature vectors obtained after the convolution operation.
[0101] Next, the maximum value in each feature map is selected through the maximum pooling operation, as shown in the following formula, where Represents the pooled features obtained after processing. This process enables the model to focus on the significant parts of local features, thereby increasing sensitivity to important information.
[0102]
[0103] After obtaining the pooling features of convolution kernels of different sizes, they are concatenated according to the following formula to obtain the local features of Chinese medical texts: .
[0104]
[0105] S4. Based on the contextual features and local features of Chinese medical text and the feature fusion layer, the fused features are obtained.
[0106] In a feasible implementation, this algorithm uses a multi-head self-attention mechanism to fuse the extracted context features and local features. First, the context features and local features are spliced, and then the query matrix, key matrix, and value matrix of each head are calculated based on the trainable weight matrix and the spliced features. Next, the attention value of each head is calculated based on these three matrices and spliced together, and the fused features are obtained by transforming them using a linear transformation matrix.
[0107] Among them, the feature fusion layer includes multiple independent self-attention heads.
[0108] Optionally, the specific operation of S4 may include the following steps S41-S43:
[0109] S41, concatenating the context features and local features of the Chinese medical text to obtain concatenated features;
[0110] S42, using each independent self-attention head to process the spliced features respectively to obtain multiple attention features;
[0111] S43. Concatenate multiple attention features, perform linear transformation on the concatenated features, and obtain fused features.
[0112] In a feasible implementation, after obtaining the context features of the Chinese medical text and local features After that, the embodiment of the present invention adopts a multi-head self-attention mechanism in the feature fusion layer to fuse the two. The multi-head self-attention mechanism processes the same set of inputs through multiple independent attention heads. Each head operates in a different subspace and can focus on different aspects and hierarchical structures of the input at the same time, thereby obtaining a richer feature representation, which is better than the single-head self-attention mechanism. It is now assumed that h heads are used in the feature fusion process. The process of the multi-head self-attention mechanism fusing contextual features and local features is as follows:
[0113] First, the context features and local features are concatenated, as shown in the following formula.
[0114]
[0115] For the i-th head, use the trainable weight matrix , and Compute the query matrix , key matrix , value matrix , as shown below.
[0116]
[0117] Using the Query Matrix , key matrix , value matrix The attention value is calculated according to the following formula, where is a scaling factor that aims to stabilize the training process and prevent the dot product value from being too large.
[0118]
[0119] As shown in the following formula, the attention output of each head is concatenated and the linear transformation matrix is used After transformation, the fused features can be obtained.
[0120]
[0121] S5. Based on the fused features and annotation layer, the corresponding named entities of the Chinese medical text are obtained and classified to complete the recognition of Chinese medical named entities.
[0122] In a feasible implementation, this algorithm is based on the idea of dynamic programming to find a score function Maximized label sequence . First, the label of the first subword is determined based on the initial emission score. Then, for each subsequent subword, its label score is recursively calculated. Specifically, the label score is obtained by selecting the maximum value of the previous subword label score and the transfer score after pruned, and adding it to the emission score of the current subword. The pruning operation is based on an adaptive threshold, which is calculated by the mean and standard deviation of the previous subword label score and controlled by an adjustment coefficient. At the same time, the predecessor state of each label is recorded for subsequent backtracking. When processing the last subword, the label with the highest score is selected as the end label of the sequence, and the optimal label sequence is restored by backtracking.
[0123] Optionally, the specific operation of S5 may include the following steps S51-S53:
[0124] S51, calculating the label score of the first subword;
[0125] S52, calculating the sum of the label score of the first subword and each score of the second subword, screening out a sum value that meets the condition from multiple sum values according to a preset condition, determining the sum value that meets the condition as the label score group of the second subword, and pruning the label according to the sum value that meets the condition;
[0126] S53. For each subsequent subword, calculate the maximum value in the label score group of the previous subword and the sum of each score of the current subword, filter out the qualified sum values from multiple sum values according to preset conditions, determine the qualified sum values as the label score group of the current subword, and prune according to the qualified sum values until the last subword is calculated, select the label with the highest score as the end label of the sequence, and backtrack according to the predecessor state to obtain the label sequence that maximizes the label score, and determine the corresponding category of the Chinese medical text according to the label sequence.
[0127] Optionally, the above preset condition can be as follows (13):
[0128] (13)
[0129] in, is the adaptive threshold, Representation Tags , Representation Tags The label set obtained after pruning operation, The value of is the mean calculated from the scores of each label of the t-1th subword and standard deviation get, The calculation formula is as follows (14):
[0130] (14)
[0131] in, and These are adjustment factors. The calculation formula is as follows (15): The calculation formula is as follows (16):
[0132] (15)
[0133] (16).
[0134] In a feasible implementation manner, the feature extraction layer and the feature fusion layer process the obtained It has rich feature representation and contains a lot of key information in Chinese medical texts. The main task of the annotation layer is to Based on the dependency between labels, ATVCRF (Adaptive Threshold Viterbi Conditional Random Field) is used to annotate the text sequence and output a legal and globally optimal label sequence. Now assume that a possible output label sequence is ,in represents the label corresponding to the t-th subword; the entity label set of Chinese medical text is , L is the number of entity tags.
[0135] Score function of ATVCRF The definition is the same as that of CRF, which is used to measure the matching degree of a specific label sequence Y to a given input sequence X. This function consists of two parts: the emission score and the transfer score. The specific calculation method is as follows.
[0136]
[0137] The emission score represents the score of outputting a specific label at each position, That is, it represents the features and labels of the t-th subword of a given input sequence X The initial emission score is determined by the fusion feature It is obtained by mapping it to the label space through linear transformation. For the t-th subword, , Representative subword Marked as a tag score.
[0138] The transfer score represents the score of transferring from one label to the next label. That means from the label Transfer to label The initial transition score is determined by empirical initialization. The score of illegal transitions (such as transitions from label “I-dis” to label “I-dru”) is initialized to a very low value based on prior knowledge, and other legal transitions are initialized randomly.
[0139] use After calculating the scores of each label sequence based on the emission score and the transfer score, the conditional probability that reflects the possibility of a given input sequence X generating a specific label sequence Y is calculated. , as shown below.
[0140]
[0141] During the training process, in order to further improve the accuracy of the model in predicting the label sequence, the negative log-likelihood function is used to quantify the prediction quality of the model for the true label sequence. The calculation method is shown in the following formula.
[0142]
[0143] During the inference process, the traditional CRF uses the Viterbi algorithm to find a score function that Maximized label sequence . The Viterbi algorithm is based on the idea of dynamic programming and recursively calculates the label score corresponding to each subword. The score is calculated by adding the label score of the previous subword to the emission score and transfer score of the label corresponding to the current subword. When calculating each subword, the dynamic programming table and path tracking table are used to record the optimal score and predecessor state for subsequent backtracking. Finally, when calculating the last subword, the label with the highest score is selected as the end label of the sequence, and the optimal label path of the entire sequence is obtained by backtracking through the recorded predecessor state. The Viterbi algorithm needs to check the relevant scores of all labels of the previous subword when calculating the optimal score of all labels for each subword, so each subword involves operations, the total time complexity is Because the Viterbi algorithm needs to store the optimal score and predecessor state of each subword for each tag, the space complexity is .
[0144] In the Viterbi algorithm, each column in the dynamic programming table needs to calculate all possible states. In some cases, the scores of some states are much lower than other states, and they can be pruned in advance (discarding some states that cannot become the final optimal path) to reduce the amount of calculation and memory usage, thereby improving the efficiency of the algorithm. Therefore, the embodiment of the present invention proposes a Viterbi algorithm (Adaptive Threshold Viterbi, ATV) based on an adaptive threshold pruning operation, and its specific calculation steps are as follows:
[0145] The score of each tag for the first subword selection is initialized according to the initial emission score as shown in the following formula.
[0146]
[0147] For each subsequent subword t and possible labels Recursively calculate its score as shown below. That is, calculate the label of the t-th subword When scoring, all tags that meet the conditions for the t-1th subword need to be considered The scores of , and select the maximum value as the score basis of the current score. It represents the set of labels obtained after pruning after processing the t-1th subword.
[0148]
[0149] for , its calculation method can be expressed as above formula (13), where, is the adaptive threshold.
[0150]
[0151] The value of is the mean calculated from the scores of each label of the t-1th subword and standard deviation The calculation methods of the three are shown in the above formula (14), formula (15), and formula (16). Among them, and These are adjustment factors. Mainly affects the mean part of the threshold, controls the looseness of pruning, and the larger Values will reduce pruning; The standard deviation of the threshold affects the strictness of pruning. Values of will result in more pruning. In practice, and The value of needs to be adjusted to balance the pruning effect, computational complexity, and model performance.
[0152]
[0153]
[0154]
[0155] When using the ATV algorithm to recursively calculate the maximum score, in order to facilitate the recovery of the complete optimal label sequence by backtracking, use Record the predecessor state, that is, select the label at the tth subword The optimal label corresponding to the t-1th subword is shown in the following formula.
[0156]
[0157] When calculating the last subword, select the tag with the highest score as the end tag of the sequence, and backtrack based on the predecessor state to obtain the score function Maximized label sequence If we assume that each pruning operation only retains k label scores, the time complexity of the ATV algorithm is , which reduces the time complexity to a certain extent. Although the dynamic programming table and path tracking table are still dimension, the space complexity is still theoretically , but the pruning strategy reduces the amount of data that needs to be stored and processed in each processing, which reduces the actual memory usage. It can be seen that the ATV algorithm reduces the time complexity and actual memory usage to a certain extent, can effectively improve the inference speed of the ATV CRF model, and has higher efficiency when processing long sequences and large-scale data sets.
[0158] In an embodiment of the present invention, a Chinese medical named entity recognition method based on multi-feature fusion is proposed, and a BERT-wwm-BiLSTM-TextCNN-MHSA-ATVCRF model is designed. First, Chinese-BERT-wwm is used to process the input text sequence to generate word embedding for subsequent processing. Considering the complex contextual dependencies in the text, BiLSTM is used to extract contextual features from word embedding, which effectively avoids the problem of incomplete context modeling; considering that a large number of entities are composed of complex terms and have complex boundaries, TextCNN is used to extract a large number of local features from word embedding through multiple convolution kernels, which helps to identify and distinguish entities with complex boundaries. Then, the contextual features and local features are fused through a multi-head self-attention mechanism to obtain a richer and more detailed feature representation, which effectively solves the problem of overly complex contextual dependencies and blurred entity boundaries in Chinese medical texts. ATVCRF is used for sequence annotation, and the state with lower scores is pruned by adaptive thresholds, which reduces the time complexity and actual memory usage of the algorithm, improves the inference speed of the model, and has higher efficiency when processing large-scale Chinese medical data sets. Finally, a legal and globally optimal label sequence is output to complete the Chinese medical named entity recognition task. The method proposed in the embodiment of the present invention can effectively improve the precision and recall rate of Chinese medical named entity recognition tasks, and has a higher reasoning speed on large-scale data sets.
[0159] Figure 3 1 is a block diagram of a Chinese medical named entity recognition device based on multi-feature fusion according to an exemplary embodiment, wherein the device is used in a Chinese medical named entity recognition method based on multi-feature fusion. Figure 3 The device includes an acquisition unit 310, an embedding unit 320, an extraction unit 330, a fusion unit 340 and a recognition unit 350. Among them:
[0160] An acquisition unit 310 is used to acquire a Chinese medical text to be recognized and a Chinese medical named entity recognition model, wherein the Chinese medical named entity recognition model includes an embedding layer, a feature extraction layer, a feature fusion layer, and a labeling layer;
[0161] An embedding unit 320, configured to obtain a word embedding of the Chinese medical text based on the Chinese medical text and the embedding layer;
[0162] An extraction unit 330, configured to obtain context features and local features of the Chinese medical text based on the word embedding of the Chinese medical text and the feature extraction layer;
[0163] A fusion unit 340, configured to obtain fused features based on the context features and local features of the Chinese medical text and the feature fusion layer;
[0164] The recognition unit 350 is used to obtain and classify the named entities corresponding to the Chinese medical text based on the fused features and the annotation layer, so as to complete the recognition of the Chinese medical named entities.
[0165] Optionally, the embedding layer is a Chinese-BERT-wwm model;
[0166] The embedding unit 320 is further used for:
[0167] S21, inputting the Chinese medical text into the embedding layer, and segmenting the Chinese medical text into subwords using a word segmentation tool in the embedding layer;
[0168] S22, according to a preset vocabulary, converting the subword into an ID sequence, and mapping the ID sequence to a high-dimensional vector space to obtain a subword vector;
[0169] S23, determining the position embedding and type embedding of the subword according to the ID sequence;
[0170] S24, generating preliminary word embedding information according to the subword vector and the position embedding and type embedding of the subword;
[0171] S25. Input the preliminary word embedding information into a multi-layer Transformer encoder, fuse the context information through a self-attention mechanism, and obtain the word embedding of the Chinese medical text.
[0172] Optionally, the feature extraction layer includes BiLSTM and TextCNN;
[0173] The BiLSTM includes multiple LSTM networks, each of which includes an input gate, a forget gate, an output gate, and a memory unit; two LSTM networks, namely a forward LSTM and a backward LSTM, each of which includes an input gate, a forget gate, an output gate, and a memory unit;
[0174] The TextCNN includes a plurality of convolutional layers with convolutional kernels of different sizes;
[0175] The extraction unit 330 is further used to:
[0176] S31, according to the word embedding of the Chinese medical text and an adjacent hidden state, respectively calculate the activation value of the input gate, the activation value of the forget gate, the candidate memory unit state and the activation value of the output gate, according to the activation value of the input gate and the activation value of the forget gate, update the candidate memory unit state to obtain the current memory unit state, and determine the current hidden state according to the activation value of the output gate and the current memory unit state;
[0177] S32, determining context features according to the current hidden state;
[0178] S33. Input the word embedding of the Chinese medical text into a convolution layer with convolution kernels of different sizes to obtain a feature vector set, obtain multiple pooling features based on the feature vector set and a maximum pooling operation, and concatenate the multiple pooling features to obtain local features of the Chinese medical text.
[0179] Optionally, the multiple LSTM networks included in the BiLSTM are evenly divided into two types of LSTM networks, namely a forward LSTM network and a backward LSTM network;
[0180] If the LSTM network is a forward LSTM network, the extraction unit 330 is further used to:
[0181] According to the word embedding of the Chinese medical text, the previous hidden state and formula (1), the forward activation value of the input gate is calculated :
[0182]
[0183] in, represents the sigmoid activation function, represents the weight matrix of the input gate, represents the bias term of the input gate, represents the word embedding at the current time step t, Represents the previous hidden state of the forward LSTM;
[0184] According to the word embedding of the Chinese medical text, the previous hidden state and formula (2), the forward activation value of the forget gate is calculated :
[0185]
[0186] in, represents the weight matrix of the forget gate, Represents the bias term of the forget gate;
[0187] According to the word embedding of the Chinese medical text, the previous hidden state and formula (3), the forward candidate memory unit state is calculated :
[0188] (3)
[0189] Among them, tanh represents the hyperbolic tangent activation function, represents the weight matrix of the memory unit, Represents the bias term of the memory unit;
[0190] According to the forward activation value of the input gate, the forward activation value of the forget gate and formula (4), the forward candidate memory unit state is updated to obtain the forward current memory unit state :
[0191] (4)
[0192] in, represents element-wise multiplication, Represents the state of the forward memory unit at the previous time step;
[0193] According to the forward word embedding of the Chinese medical text, the previous hidden state and formula (5), the forward activation value of the output gate is calculated :
[0194] (5)
[0195] in, represents the weight matrix of the output gate, Represents the bias term of the output gate;
[0196] According to the forward activation value of the output gate, the forward current memory unit state and formula (6), the forward current hidden state is determined :
[0197] (6)
[0198] When the LSTM network is a backward LSTM, the S31 calculates the activation value of the input gate, the activation value of the forget gate, the candidate memory unit state and the activation value of the output gate according to the word embedding of the Chinese medical text and the adjacent hidden state, updates the candidate memory unit state according to the activation value of the input gate and the activation value of the forget gate, obtains the current memory unit state, and determines the current hidden state according to the activation value of the output gate and the current memory unit state, including:
[0199] According to the word embedding of Chinese medical text, the latter hidden state and formula (7), the backward activation value of the input gate is calculated :
[0200] (7)
[0201] in, represents the sigmoid activation function, represents the weight matrix of the input gate, represents the bias term of the input gate, represents the word embedding at the current time step t, Represents the next hidden state of the backward LSTM;
[0202] According to the word embedding of Chinese medical text, the latter hidden state and formula (8), the backward activation value of the forget gate is calculated :
[0203] (8)
[0204] in, represents the weight matrix of the forget gate, Represents the bias term of the forget gate;
[0205] According to the word embedding of Chinese medical text, the latter hidden state and formula (9), the backward candidate memory unit state is calculated :
[0206] (9)
[0207] Among them, tanh represents the hyperbolic tangent activation function, represents the weight matrix of the memory unit, Represents the bias term of the memory unit;
[0208] According to the backward activation value of the input gate, the backward activation value of the forget gate and formula (10), the backward candidate memory unit state is updated to obtain the backward current memory unit state :
[0209] (10)
[0210] in, represents element-wise multiplication, Represents the state of the backward memory unit at the previous time step;
[0211] According to the backward word embedding of Chinese medical text, the subsequent hidden state and formula (11), the backward activation value of the output gate is calculated :
[0212] (11)
[0213] in, represents the weight matrix of the output gate, Represents the bias term of the output gate;
[0214] According to the backward activation value of the output gate, the backward current memory unit state and formula (12), the backward current hidden state is determined :
[0215] (12).
[0216] Optionally, the extraction unit 330 is further configured to:
[0217] Pass the current hidden state forward and backward to the current hidden state The spliced features are then concatenated and determined as context features.
[0218] Optionally, the feature fusion layer includes a plurality of independent self-attention heads;
[0219] The fusion unit 340 is further configured to:
[0220] Splicing the context features and local features of the Chinese medical text to obtain splicing features;
[0221] Use each independent self-attention head to process the concatenated features and obtain multiple attention features;
[0222] Multiple attention features are concatenated and linearly transformed to obtain fused features.
[0223] Optionally, the identification unit 350 is further configured to:
[0224] Calculate the tag score of the first subword;
[0225] Calculate the sum of the label score of the first subword and each score of the second subword, select the sum value that meets the condition from multiple sum values according to the preset condition, determine the sum value that meets the condition as the label score group of the second subword, and prune the label according to the sum value that meets the condition;
[0226] For each subsequent subword, the maximum value in the label score group of the previous subword and the sum of each score of the current subword are calculated, and the sum values that meet the conditions are screened out from multiple sum values according to preset conditions, and the sum values that meet the conditions are determined as the label score group of the current subword. Pruning is performed according to the sum values that meet the conditions until the last subword is calculated, and the label with the highest score is selected as the end label of the sequence. The label sequence that maximizes the label score is obtained by backtracking according to the predecessor state, and the corresponding category of the Chinese medical text is determined according to the label sequence.
[0227] Optionally, the preset condition is as follows (13):
[0228] (13)
[0229] in, is the adaptive threshold, Representation Tags , Representation Tags The label set obtained after pruning operation, The value of is the mean calculated from the scores of each label of the t-1th subword and standard deviation get, The calculation formula is as follows (14):
[0230] (14)
[0231] in, and These are adjustment factors. The calculation formula is as follows (15): The calculation formula is as follows (16):
[0232] (15)
[0233] (16).
[0234] In an embodiment of the present invention, a Chinese medical named entity recognition device based on multi-feature fusion is proposed, and a BERT-wwm-BiLSTM-TextCNN-MHSA-ATVCRF model is designed. First, Chinese-BERT-wwm is used to process the input text sequence to generate word embedding for subsequent processing. Considering the complex contextual dependencies in the text, BiLSTM is used to extract contextual features from word embedding, which effectively avoids the problem of incomplete context modeling; considering that a large number of entities are composed of complex terms and have complex boundaries, TextCNN is used to extract a large number of local features from word embedding through multiple convolution kernels, which is helpful to identify and distinguish entities with complex boundaries. Then, the contextual features and local features are fused through a multi-head self-attention mechanism to obtain a richer and more detailed feature representation, which effectively solves the problem of overly complex contextual dependencies and blurred entity boundaries in Chinese medical texts. ATVCRF is used for sequence annotation, and the state with a low score is pruned by an adaptive threshold, which reduces the time complexity and actual memory usage of the algorithm, improves the inference speed of the model, and has higher efficiency when processing large-scale Chinese medical data sets. Finally, a legal and globally optimal label sequence is output to complete the Chinese medical named entity recognition task. The method proposed in the embodiment of the present invention can effectively improve the precision and recall rate of Chinese medical named entity recognition tasks, and has a higher reasoning speed on large-scale data sets.
[0235] Figure 4 is a schematic diagram of the structure of a Chinese medical named entity recognition device based on multi-feature fusion provided by an embodiment of the present invention, such as Figure 4 As shown, the Chinese medical named entity recognition device based on multi-feature fusion may include the above Figure 3The Chinese medical named entity recognition device based on multi-feature fusion is shown. Optionally, the Chinese medical named entity recognition device based on multi-feature fusion 410 may include a first processor 2001 .
[0236] Optionally, the Chinese medical named entity recognition device 410 based on multi-feature fusion may also include a memory 2002 and a transceiver 2003 .
[0237] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0238] Combine the following Figure 4 The components of the Chinese medical named entity recognition device 410 based on multi-feature fusion are specifically introduced:
[0239] The first processor 2001 is the control center of the Chinese medical named entity recognition device 410 based on multi-feature fusion, which can be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiment of the present invention, such as one or more microprocessors (digital signal processor, DSP), or one or more field programmable gate arrays (field programmable gate array, FPGA).
[0240] Optionally, the first processor 2001 can perform various functions of the Chinese medical named entity recognition device 410 based on multi-feature fusion by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.
[0241] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 CPU0 and CPU1 are shown in FIG.
[0242] In a specific implementation, as an embodiment, the Chinese medical named entity recognition device 410 based on multi-feature fusion may also include multiple processors, such as Figure 4The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0243] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.
[0244] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently, and may be accessed through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0245] The transceiver 2003 is used to communicate with a network device or a terminal device.
[0246] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0247] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently, and may be connected to the Chinese medical named entity recognition device 410 based on multi-feature fusion through an interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0248] It should be noted that Figure 4 The structure of the Chinese medical named entity recognition device 410 based on multi-feature fusion shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0249] In addition, the technical effects of the Chinese medical named entity recognition device 410 based on multi-feature fusion can refer to the technical effects of the Chinese medical named entity recognition method based on multi-feature fusion described in the above method embodiment, and will not be repeated here.
[0250] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0251] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0252] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0253] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0254] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0255] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0256] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0257] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0258] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0259] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0260] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0261] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0262] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A Chinese medical named entity recognition method based on multi-feature fusion, characterized in that: The method comprises: S1. Obtaining a Chinese medical text to be recognized and a Chinese medical named entity recognition model, wherein the Chinese medical named entity recognition model includes an embedding layer, a feature extraction layer, a feature fusion layer, and a labeling layer; S2. Based on the Chinese medical text and the embedding layer, obtain a word embedding of the Chinese medical text; S3, based on the word embedding of the Chinese medical text and the feature extraction layer, obtaining context features and local features of the Chinese medical text; S4, obtaining fused features based on the context features and local features of the Chinese medical text and the feature fusion layer; S5. Based on the fused features and the annotation layer, the named entities corresponding to the Chinese medical text are obtained and classified to complete the recognition of the Chinese medical named entities.
2. According to the Chinese medical named entity recognition method based on multi-feature fusion according to claim 1, it is characterized in that: The embedding layer is a Chinese-BERT-wwm model; The step S2 of obtaining a word embedding of the Chinese medical text based on the Chinese medical text and the embedding layer includes: S21, inputting the Chinese medical text into the embedding layer, and segmenting the Chinese medical text into subwords using a word segmentation tool in the embedding layer; S22, according to a preset vocabulary, converting the subword into an ID sequence, and mapping the ID sequence to a high-dimensional vector space to obtain a subword vector; S23, determining the position embedding and type embedding of the subword according to the ID sequence; S24, generating preliminary word embedding information according to the subword vector and the position embedding and type embedding of the subword; S25. Input the preliminary word embedding information into a multi-layer Transformer encoder, fuse the context information through a self-attention mechanism, and obtain the word embedding of the Chinese medical text.
3. The Chinese medical named entity recognition method based on multi-feature fusion according to claim 1 is characterized in that: The feature extraction layer includes BiLSTM and TextCNN; The BiLSTM includes multiple LSTM networks, each of which includes an input gate, a forget gate, an output gate, and a memory unit; two LSTM networks, namely a forward LSTM and a backward LSTM, each of which includes an input gate, a forget gate, an output gate, and a memory unit; The TextCNN includes a plurality of convolutional layers with convolutional kernels of different sizes; The S3 obtains context features and local features of the Chinese medical text based on the word embedding of the Chinese medical text and the feature extraction layer, including: S31, according to the word embedding of the Chinese medical text and an adjacent hidden state, respectively calculate the activation value of the input gate, the activation value of the forget gate, the candidate memory unit state and the activation value of the output gate, according to the activation value of the input gate and the activation value of the forget gate, update the candidate memory unit state to obtain the current memory unit state, and determine the current hidden state according to the activation value of the output gate and the current memory unit state; S32, determining context features according to the current hidden state; S33. Input the word embedding of the Chinese medical text into a convolution layer with convolution kernels of different sizes to obtain a feature vector set, obtain multiple pooling features based on the feature vector set and a maximum pooling operation, and concatenate the multiple pooling features to obtain local features of the Chinese medical text.
4. The Chinese medical named entity recognition method based on multi-feature fusion according to claim 3 is characterized in that: The multiple LSTM networks included in the BiLSTM are evenly divided into two LSTM networks, namely a forward LSTM network and a backward LSTM network; If the LSTM network is a forward LSTM network, the S31 calculates the activation value of the input gate, the activation value of the forget gate, the candidate memory unit state and the activation value of the output gate according to the word embedding of the Chinese medical text and the adjacent hidden state, updates the candidate memory unit state according to the activation value of the input gate and the activation value of the forget gate to obtain the current memory unit state, and determines the current hidden state according to the activation value of the output gate and the current memory unit state, including: According to the word embedding of the Chinese medical text, the previous hidden state and formula (1), the forward activation value of the input gate is calculated : ; in, represents the sigmoid activation function, represents the weight matrix of the input gate, represents the bias term of the input gate, represents the word embedding at the current time step t, Represents the previous hidden state of the forward LSTM; According to the word embedding of the Chinese medical text, the previous hidden state and formula (2), the forward activation value of the forget gate is calculated : ; in, represents the weight matrix of the forget gate, Represents the bias term of the forget gate; According to the word embedding of the Chinese medical text, the previous hidden state and formula (3), the forward candidate memory unit state is calculated : (3) Among them, tanh represents the hyperbolic tangent activation function, represents the weight matrix of the memory unit, Represents the bias term of the memory unit; According to the forward activation value of the input gate, the forward activation value of the forget gate and formula (4), the forward candidate memory unit state is updated to obtain the forward current memory unit state : (4) in, represents element-wise multiplication, Represents the state of the forward memory unit at the previous time step; According to the forward word embedding of the Chinese medical text, the previous hidden state and formula (5), the forward activation value of the output gate is calculated : (5) in, represents the weight matrix of the output gate, Represents the bias term of the output gate; According to the forward activation value of the output gate, the forward current memory unit state and formula (6), the forward current hidden state is determined : (6) When the LSTM network is a backward LSTM, the S31 calculates the activation value of the input gate, the activation value of the forget gate, the candidate memory unit state and the activation value of the output gate according to the word embedding of the Chinese medical text and the adjacent hidden state, updates the candidate memory unit state according to the activation value of the input gate and the activation value of the forget gate, obtains the current memory unit state, and determines the current hidden state according to the activation value of the output gate and the current memory unit state, including: According to the word embedding of Chinese medical text, the latter hidden state and formula (7), the backward activation value of the input gate is calculated : (7) in, represents the sigmoid activation function, represents the weight matrix of the input gate, represents the bias term of the input gate, represents the word embedding at the current time step t, Represents the next hidden state of the backward LSTM; According to the word embedding of Chinese medical text, the latter hidden state and formula (8), the backward activation value of the forget gate is calculated : (8) in, represents the weight matrix of the forget gate, Represents the bias term of the forget gate; According to the word embedding of Chinese medical text, the latter hidden state and formula (9), the backward candidate memory unit state is calculated : (9) Among them, tanh represents the hyperbolic tangent activation function, represents the weight matrix of the memory unit, Represents the bias term of the memory unit; According to the backward activation value of the input gate, the backward activation value of the forget gate and formula (10), the backward candidate memory unit state is updated to obtain the backward current memory unit state : (10) in, represents element-wise multiplication, Represents the state of the backward memory unit at the previous time step; According to the backward word embedding of Chinese medical text, the subsequent hidden state and formula (11), the backward activation value of the output gate is calculated : (11) in, represents the weight matrix of the output gate, Represents the bias term of the output gate; According to the backward activation value of the output gate, the backward current memory unit state and formula (12), the backward current hidden state is determined : (12)。 5. The Chinese medical named entity recognition method based on multi-feature fusion according to claim 4 is characterized in that: The step S32 of determining the context feature according to the current hidden state includes: Pass the current hidden state forward and backward to the current hidden state The spliced features are then concatenated and determined as context features.
6. The Chinese medical named entity recognition method based on multi-feature fusion according to claim 1 is characterized in that: The feature fusion layer includes multiple independent self-attention heads; The S4 obtains fused features based on the context features and local features of the Chinese medical text and the feature fusion layer, including: Splicing the context features and local features of the Chinese medical text to obtain splicing features; Use each independent self-attention head to process the concatenated features and obtain multiple attention features; Multiple attention features are concatenated and linearly transformed to obtain fused features.
7. The Chinese medical named entity recognition method based on multi-feature fusion according to claim 1 is characterized in that: The step S5 obtains and classifies the named entities corresponding to the Chinese medical text based on the fused features and the annotation layer, including: Calculate the tag score of the first subword; Calculate the sum of the label score of the first subword and each score of the second subword, select the sum value that meets the condition from multiple sum values according to the preset condition, determine the sum value that meets the condition as the label score group of the second subword, and prune the label according to the sum value that meets the condition; For each subsequent subword, the maximum value in the label score group of the previous subword and the sum of each score of the current subword are calculated, and the sum values that meet the conditions are screened out from multiple sum values according to preset conditions. The sum values that meet the conditions are determined as the label score group of the current subword, and pruning is performed according to the sum values that meet the conditions until the last subword is calculated. The label with the highest score is selected as the end label of the sequence, and the label sequence that maximizes the label score is obtained by backtracking according to the predecessor state, and the corresponding category of the Chinese medical text is determined according to the label sequence.
8. The Chinese medical named entity recognition method based on multi-feature fusion according to claim 7 is characterized in that: The preset condition is as follows: (13) in, is the adaptive threshold, Representation Tags , Representation Tags The label set obtained after pruning operation, The value of is the mean calculated from the scores of each label of the t-1th subword and standard deviation get, The calculation formula is as follows (14): (14) in, and These are adjustment factors. The calculation formula is as follows (15): The calculation formula is as follows (16): (15) (16)。 9. A Chinese medical named entity recognition device based on multi-feature fusion, the Chinese medical named entity recognition device based on multi-feature fusion is used to implement the Chinese medical named entity recognition method based on multi-feature fusion as described in any one of claims 1-8, characterized in that: The device comprises: An acquisition unit, used to acquire Chinese medical text to be recognized and a Chinese medical named entity recognition model, wherein the Chinese medical named entity recognition model includes an embedding layer, a feature extraction layer, a feature fusion layer, and a labeling layer; An embedding unit, configured to obtain a word embedding of the Chinese medical text based on the Chinese medical text and the embedding layer; An extraction unit, configured to obtain context features and local features of the Chinese medical text based on the word embedding of the Chinese medical text and the feature extraction layer; A fusion unit, used for obtaining fused features based on the context features and local features of the Chinese medical text and the feature fusion layer; The recognition unit is used to obtain and classify the named entities corresponding to the Chinese medical text based on the fused features and the annotation layer, so as to complete the recognition of the Chinese medical named entities.
10. A Chinese medical named entity recognition device based on multi-feature fusion, characterized in that: The Chinese medical named entity recognition device based on multi-feature fusion includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 8 is implemented.