Traditional Chinese medicine nested named entity recognition method based on dynamic span length constraint

By fusion of character-level and word-level embedding and dynamic span length constraint, the complexity and fuzzy boundary problems in nested named entity recognition in traditional Chinese medicine are solved, and the recognition accuracy and efficiency are improved.

CN120764533APending Publication Date: 2025-10-10UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510935267.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Nested entities are a common phenomenon in TCM named entity recognition, and the professional terms in the field of TCM are diverse and have fuzzy boundaries. Existing methods fail to fully explore and utilize the character-level and word-level multidimensional features of the text, resulting in insufficient model understanding ability and poor recognition effect.

Method used

By fusing character-level embedding and word-level embedding, utilizing radical, character, dictionary matching and dependency syntactic features, combined with attention mechanism and bidirectional LSTM encoding, and dynamic span length constraint, the accuracy of entity boundary recognition is improved.

Benefits of technology

The accuracy and efficiency of TCM nested named entity recognition are improved, the number of candidate spans is reduced, and the model's ability to understand TCM domain knowledge is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764533A_ABST
    Figure CN120764533A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing, and particularly relates to a traditional Chinese medicine nested named entity recognition method based on dynamic span length constraint. According to the method, character-level embedding is obtained by fusing character embedding, radical embedding and traditional Chinese medicine dictionary matching word embedding, word-level embedding is obtained by fusing word segmentation embedding, part-of-speech embedding and dependency syntax embedding, and then embedding representation of multi-dimensional feature information is obtained through mask word-word attention fusion character-level embedding and word-level embedding. The understanding ability of the model for traditional Chinese medicine field texts and knowledge is improved; the maximum span length is predicted through the dynamic span length constraint module to constrain generation of the candidate spans, the number of the candidate spans can be reduced, and the entity boundary recognition performance and prediction efficiency of the model can be improved. The accuracy of traditional Chinese medicine named entity recognition can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular is a method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints. Background Art

[0002] The Traditional Chinese Medicine (TCM) knowledge system is vast and complex, involving a vast amount of heterogeneous data from multiple sources, including medical records, literature, and clinical experience. This complexity, dispersion, and unstructured nature pose significant challenges to the integration and utilization of this information. Named Entity Recognition (NER) is the task of extracting meaningful named entities from unstructured text. TCM named entity recognition automatically identifies and extracts entities such as Chinese medicines, diseases, and symptoms from ancient TCM texts and medical records, making it a key step in building a TCM knowledge base.

[0003] Nested named entities, also known as entity overlap, occur when one entity contains one or more other entities within the same text segment. For example, in the sentence "White fur on the tongue caused by plague is due to the presence of the membrane," "The membrane" is a pathogenic entity, while "membrane" is an organismal morphology entity. Nested named entity recognition involves overlapping or intersecting entity boundaries, requiring a multi-level entity structure to be identified compared to traditional named entity recognition, making it more challenging.

[0004] Compared with named entity recognition in general fields, named entity recognition in traditional Chinese medicine faces the problems of complex entity categories and widespread nested phenomena. In addition, there are many professional terms in the field of traditional Chinese medicine, whose expressions are diverse and the boundaries are relatively vague. Therefore, it is very difficult to perform nested named entity recognition in the field of traditional Chinese medicine.

[0005] The patent with publication number CN118095283A discloses a nested named entity recognition method based on word information fusion and boundary detection. This method obtains related word phrases by matching each character in the target sentence with a dictionary, then fuses the words in the phrase according to the weight to obtain word embedding, encodes the characters through the embedding layer to obtain word embedding, fuses the two to obtain word fusion vectors, and then interacts and encodes through the BERT Transformer encoder, and finally performs span boundary prediction and span classification. However, this method does not fully explore and utilize the character-level and word-level multidimensional features of the text. For the field of traditional Chinese medicine, it often contains more professional terms, resulting in insufficient understanding of the model's knowledge in the field of traditional Chinese medicine, which in turn affects the effect of named entity recognition.

[0006] Patent publication number CN118821778A discloses a method for identifying nested named entities in electronic medical records. The method involves encoding the electronic medical record text to obtain character vectors and word vectors; extracting features from the word vectors to obtain contextual feature representation vectors; constructing and extracting sequence graphs on the character vectors to obtain sequence graph feature extraction results; constructing and extracting features from the word vectors to obtain syntactic dependency graph feature extraction results; extracting features from the sequence graph feature extraction results and the syntactic dependency graph feature extraction results to obtain a cascade feature graph, which is then clustered to obtain a cluster extraction result; extracting radical features based on the electronic medical record text; and performing span prediction based on the contextual feature representation vectors, cluster extraction results, and radical feature extraction results to more accurately identify nested entities in electronic medical records. This method deeply exploits sequence graph features, syntactic dependency graph features, and radical features, but it generates a large number of low-quality candidate spans when dealing with nested named entities, resulting in high computational complexity for the model. Summary of the Invention

[0007] To address these issues, the present invention proposes a method for more accurately and efficiently recognizing nested named entities in TCM by fully mining and utilizing the multidimensional character- and word-level features of TCM texts. By fully mining and integrating the character- and word-level features of the text, the model's ability to understand TCM domain knowledge is enhanced. Furthermore, by applying dynamic span length constraints, the number of generated candidate spans is reduced and the accuracy of entity boundary recognition is improved, ultimately improving the performance of nested named entity recognition in TCM.

[0008] The technical solution of the present invention is:

[0009] A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints comprises the following steps:

[0010] S1. The model performs character-level embedding and word-level embedding on the input TCM text, respectively, to obtain character-level embedding and word-level embedding, and defines the input TCM text as the input text sequence ;

[0011] The specific method of the word-level embedding process is to obtain The radical embedding, character embedding and TCM dictionary matching word embedding are obtained, and then the obtained radical embedding, character embedding and TCM dictionary matching word embedding are fused to obtain the word-level embedding; wherein the radical embedding is obtained The radical information of each character in , and then the corresponding radical embedding is obtained by querying the radical embedding matrix, and finally the The corresponding radical embedding sequence, where the initial radical embedding matrix is ​​obtained by random initialization based on the Chinese character radical standard and then updated during the training process; the character embedding is obtained The embedded information of each character in Each character in is input as a separate token into the large language model encoding layer to obtain the corresponding initial character embedding, and then obtain The corresponding initial character embedding sequence; the Chinese medicine dictionary matching word embedding is obtained The TCM dictionary matches the word embedding sequence, specifically to build a TCM dictionary, using the SoftLexicon method to obtain Finally, the word-level embedding is obtained by fusion of radical embedding, character embedding and TCM dictionary matching word embedding through learnable weights;

[0012] The specific method of word-level embedding processing is to obtain The word embedding, part-of-speech embedding and dependency syntactic embedding are obtained, and then the word-level embedding, part-of-speech embedding and dependency syntactic embedding are fused to obtain word-level embedding. Before obtaining the word embedding, Perform word segmentation processing, specifically combining the Chinese natural language processing toolkit HanLP and the constructed Chinese medicine dictionary to obtain The word list is then encoded by a large language model to obtain a word embedding vector sequence; the part-of-speech embedding is based on the word list, and then the HanLP tool is used to perform part-of-speech tagging on each token in the word list based on the PKU part-of-speech tagging set. The initial part-of-speech embedding matrix is ​​obtained by random initialization based on the PKU part-of-speech tagging set. For the word list and the corresponding part-of-speech tagging set, the part-of-speech embedding matrix is ​​obtained. The part-of-speech embedding sequence, where the part-of-speech embedding matrix is ​​updated during the training process; the dependency syntactic embedding is based on the word segmentation embedding vector sequence, combined with the dependency syntactic tree and the dependency relation embedding matrix, and is obtained by encoding based on the combined multi-relation graph convolutional network (CompGCN) The dependency syntax tree is constructed based on the word segmentation list and part-of-speech tags according to the PMT dependency syntax system. The dependency relationship embedding matrix is ​​obtained by randomly initializing the dependency relationship types in the PMT dependency syntax system and updating it during training. Finally, the word-level embedding is obtained by fusion of word segmentation embedding, part-of-speech embedding and dependency syntax embedding through learnable weights.

[0013] S2, the model obtains fused embedding by weighting character-level embedding and word-level embedding through attention;

[0014] S3, the model encodes the fusion embedding through a bidirectional LSTM network to obtain Multi-dimensional feature embedding for each character in ;

[0015] S4. The model sequentially inputs the acquired multidimensional feature embeddings into a boundary classifier to predict whether the character corresponding to the current multidimensional feature embedding is the start or end position of the named entity, where the start position and the end position correspond to the start boundary and the end boundary, respectively. The boundary classifier is implemented using a multi-layer perceptron. After obtaining the start boundary and the end boundary from the input multidimensional feature embeddings, a dynamic span length constraint is used to dynamically define the maximum length of the entity span, thereby generating a candidate entity span.

[0016] S5. The model classifies the generated candidate entity spans to determine whether they belong to the predefined entity type. Specifically, all character embeddings of the candidate entity spans are aggregated through average pooling to obtain the span feature vector, which is then input into the multi-layer perceptron to predict the entity type to which it belongs. Finally, the conditional probability is calculated through softmax.

[0017] S6. Use known TCM texts to train the model described in S1-S5. The training process includes three losses: boundary detection, span estimation, and named entity recognition. Boundary detection corresponds to the sum of the starting boundary cross entropy loss function and the ending boundary cross entropy loss function, span estimation corresponds to the mean square error loss function, and named entity recognition corresponds to the cross entropy loss function. After training, a trained model is obtained.

[0018] S7. Input the TCM text to be recognized into the trained model to obtain the named entity recognition result.

[0019] Furthermore, in S1, the specific method for obtaining character embedding is to define ,in represents the i-th character, , n represents the number of characters in the input text, As a single token, it is fed into the Large Language Model (LLM) encoding layer to obtain an initial embedding. :

[0020] ,

[0021] in Represents the LLM encoding layer; by initially embedding each character, the initial character embedding sequence of the input text is obtained :

[0022] ,

[0023] in, , is the dimension of the vector of the character embedded by LLM;

[0024] The specific method of obtaining radical embedding is to Each character in , by querying the Chinese character dictionary, obtain its radical information, and embed the radical matrix Query to get the corresponding radical embedding, It is obtained by random initialization based on the 201 main radicals specified in the Chinese radical standard GF 0011-2009, expressed as For characters without radicals, use zero vector as radical embedding; define Corresponding radical embedding sequence for:

[0025] ,

[0026] in Represents characters The corresponding radical embedding vector;

[0027] The specific method of obtaining the embedding of matching words in the TCM dictionary is to construct the TCM dictionary by combining manual construction and automatic extraction. , for characters ,exist Query its matching word set and divide all matching words into four sets B, M, E and S according to the position information of the characters in the word, where the set Indicates is the starting vocabulary set, set express A collection of words in the middle, a collection Indicates A collection of words ending with Indicates A single Chinese character is used as a vocabulary set of words, which are represented as follows:

[0028] ,

[0029] ,

[0030] ,

[0031] ,

[0032] in express Chinese-Israeli characters Starts with the character The word that ends, Indicates the character Starts with the character The word that ends, Indicates the character Starts with the character The ending word, if the set is empty, then the special word "None" is added to the set;

[0033] Get Character Vocabulary collection 、 、 and Then, get the vector representation of each set 、 、 and Specifically, we first calculate the word embedding of each word in each set :

[0034] ,

[0035] in It refers to any word in each set; by weighting the word embedding in the word set, the vector representation of the word set is obtained. , vector representation for:

[0036] ,

[0037] in, Expressive words The frequency of occurrence in the text, Represents characters The normalized sum of the frequencies of all words in the matching word set:

[0038] ,

[0039] The same logic applies 、 and , by adding Chinese medicine dictionary matching word embedding :

[0040] ,

[0041] Thus we get The TCM dictionary matches the word embedding sequence:

[0042] ;

[0043] Defining learnable vectors ,in 、 、 They represent the weights corresponding to character embedding, radical embedding, and TCM dictionary matching word embedding, respectively, and are normalized using the softmax function:

[0044] ,

[0045] in Represents the normalized weight value, and the word-level embedding sequence is obtained according to the weight :

[0046] .

[0047] In the process of acquiring word embedding, define the obtained The word list is :

[0048] ,

[0049] in, is the HanLP word segmenter, represents the token of the i-th partition, and s is the number of tokens obtained by partitioning;

[0050] Each token is input into the LLM embedding layer for encoding, and the word segmentation embedding vector sequence is obtained as follows: :

[0051] ,

[0052] In the process of acquiring part-of-speech embedding, the obtained part-of-speech tag sequence is defined as :

[0053] ,

[0054] in represents the HanLP part-of-speech tagger, express The part-of-speech tagging results of

[0055] According to the 43 parts of speech defined in the PKU part-of-speech tagging set, the initial part-of-speech embedding matrix is ​​obtained by random initialization. ,in is the dimension of the LLM embedding vector, for the tag and its part of speech , according to the part-of-speech embedding matrix, obtain its corresponding part-of-speech embedding vector , thus obtaining Part-of-speech embedding sequence

[0056] ;

[0057] In the process of acquiring dependency syntax embedding, the grammatical relationship between tokens is obtained according to the PMT dependency syntax system, and the dependency syntax tree is constructed. , the dimension is There are 32 types of dependency relationships in the PMT dependency syntax system. The initial dependency embedding matrix is ​​obtained by random initialization. , use CompGCN to encode the dependency syntax tree to obtain Dependency syntactic embeddings :

[0058] ,

[0059] in Indicates encoding via CompGCN, is the word embedding vector sequence, is the initial dependency embedding matrix, Indicates the encoding of the i-th token after CompGCN;

[0060] Defining learnable vectors ,in 、 、 They represent the weights corresponding to word segmentation embedding, part-of-speech embedding, and dependency syntax embedding, respectively, and are normalized using the softmax function:

[0061] ,

[0062] in Represents the normalized weight value, and the word-level embedding sequence is obtained according to the weight :

[0063] .

[0064] Furthermore, the specific method of S2 is:

[0065] Constructing the attention mask matrix , where s is The number of tokens divided, n is The number of characters in the definition Represents the mask value of the i-th row and j-th column in the mask matrix:

[0066] ,

[0067] Let the character-level embedding vector be the query, the word-level embedding vector be the key and value, calculate the attention weight by dot product and matrix multiplication, and normalize it by softmax function to get the attention score matrix :

[0068] ,

[0069] The character fusion embedding fusing the word-level embedding information is obtained by attention-weighted word-level embedding and word-level embedding

[0070]

[0071] Further, the specific method of S3 is:

[0072] The fusion embedding is encoded by a bidirectional LSTM network The bidirectional LSTM is composed of two independent LSTM layers, the forward LSTM processes the input sequence from head to tail, and the backward LSTM processes the input sequence from tail to head, and finally the hidden layer sequence of the sentence is obtained by concatenating by position:

[0073]

[0074]

[0075]

[0076] Wherein represents the input feature at time step i, i.e. the character corresponding to the fusion embedding, and respectively represent the forward output feature and the backward output feature of the LSTM at time step i, represents the vector concatenation operation, represents the multi-dimensional feature embedding of the character after fusing the text context feature; thus the multi-dimensional feature embedding sequence of the input sentence is represented as follows:

[0077]

[0078] Further, in S4, the probability of the predicted character being the start boundary and the probability of being the end boundary are calculated as follows:

[0079]

[0080]

[0081] Wherein and respectively represent the multi-layer perceptron module of the start boundary classifier and the end boundary classifier, both of which are composed of two linear layers and ReLU activation function; by setting a confidence threshold , if the predicted probability or ​​​​​​​​If the value is greater than the threshold, the character Considered as the entity's starting or ending boundary;

[0082] The method of using dynamic span length constraint to dynamically define the maximum length of entity span is to If it is determined by the starting boundary classifier to be a potential entity starting character, then in the Chinese medicine dictionary Search by is the set of all matching words starting from , and then introduce the weight function , this function reflects the character The word appears when it is the starting character The probability of , which is calculated as follows:

[0083] ,

[0084] in Expressive words The frequency of occurrence in the training corpus is then calculated by The weighted length of the starting entity :

[0085] ,

[0086] Finally, the length offset is predicted based on the context information :

[0087] ,

[0088] in Represents a multi-layer perceptron for regression prediction; ultimately, the character Maximum length constraint for the starting entity for:

[0089] ,

[0090] For each starting boundary, get its position and distance The ending boundaries within are paired to form candidate spans.

[0091] Furthermore, in S5, a candidate span is defined by arrive The word embedding set of all words in the span is , and obtain its span feature vector by average pooling :

[0092] ,

[0093] Then use multi-layer perceptron and softmax to classify the span feature vector and predict its probability of belonging to each entity type. :

[0094] ,

[0095] in Represents a multi-layer perceptron module for span prediction.

[0096] Furthermore, in S6, the loss function of boundary detection is is the starting boundary cross entropy loss Cross entropy loss with end margin sum:

[0097] ,

[0098] ,

[0099] ,

[0100] in and is the true label of whether the i-th character is the starting boundary or the ending boundary, and For characters is the probability of the starting boundary or the ending boundary, and n is the number of characters;

[0101] Loss function for span estimation The definition is as follows:

[0102] ,

[0103] Where N is the number of training samples, is the actual maximum length constraint value, defined as The end position of the longest entity starting with character , the first entity after this entity starts at ,but The calculation is as follows:

[0104] ,

[0105] in Indicates the rounding operation, i indicates The offset in the input text;

[0106] Loss function for named entity recognition The definition is as follows:

[0107] ,

[0108] in is the number of entity categories, is the probability predicted by the model that the span belongs to the j-th entity, is the true label, which is 1 when the span belongs to the j-th entity and 0 otherwise;

[0109] During model training, minimize the joint training loss :

[0110] ,

[0111] in 、 and is a weight hyperparameter that controls the importance of each task to the model.

[0112] The beneficial effects of the present invention are:

[0113] The proposed technical solution fully exploits and integrates multi-dimensional features at the word and character levels within TCM texts, leveraging both TCM domain knowledge and textual features to effectively improve the accuracy of named entity recognition in TCM. The dynamic span length constraint module reduces the number of candidate spans, improving the model's performance in identifying entity boundaries and prediction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] Figure 1 Flow chart of the method of the present invention.

[0115] Figure 2 This is a model architecture diagram of the present invention.

[0116] Figure 3 Schematic diagram of word-level embedding.

[0117] Figure 4 Schematic diagram of word-level embedding.

[0118] Figure 5 Schematic diagram of fusion embedding. DETAILED DESCRIPTION

[0119] The technical principles and solutions of the present invention are described in detail below with reference to the accompanying drawings:

[0120] The present invention proposes a method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints. The process of the method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints is as follows: Figure 1 As shown in the figure, the model architecture is as follows Figure 2As shown. The method of the application obtains word-level embedding by fusing character embedding, component embedding and traditional Chinese dictionary matching word embedding, obtains word-level embedding by fusing word segmentation embedding, part-of-speech embedding and dependency syntax embedding, then fuses word-level embedding and word-level embedding through mask word-word attention to obtain embedding representation of multi-dimensional feature information, and improves the understanding ability of the model for traditional Chinese field text and knowledge; the maximum span length is predicted through the dynamic span length constraint module to constrain the generation of candidate span, and the performance and prediction efficiency of the model to identify entity boundaries are improved.

[0121] As Figure 1 shown, the overall process of the application includes the following steps:

[0122] Step one: word-level embedding and word-level embedding

[0123] The structural diagram of the word-level embedding module is as shown in Figure 3 , which includes:

[0124] (1) Character embedding based on large language model (LLM)

[0125] LLM is pre-trained on a large-scale text corpus and has strong semantic understanding ability, which can capture rich Chinese character semantic information. For input text , where represents the i-th character, , and n represents the number of characters in the input text. For the i-th character , it is input into the encoding layer of LLM as a separate token to obtain its initial embedding :

[0126]

[0127] where represents the LLM encoding layer. By performing initial embedding on each character, the initial character embedding sequence of the input text is obtained:

[0128]

[0129] where, , is the dimension of the vector of character embedding through LLM.

[0130] (2) Component embedding

[0131] Radicals, as the basic building blocks of Chinese characters, often contain information about the character's meaning or category. For example, in Traditional Chinese Medicine (TCM), characters starting with the radical "疒" (疒) are often associated with diseases, while characters starting with the radical "月" (月) are often associated with body parts. When identifying named entities in TCM, radical information can help the model capture the characteristics of Chinese characters and improve recognition accuracy.

[0132] According to the Chinese radical standard GF 0011-2009 "Chinese character radical table", there are 201 main radicals in total. Therefore, this scheme obtains the initial radical embedding matrix by random initialization. ,in is the dimension of the LLM embedding vector, Each vector in represents the initial embedding corresponding to each radical. During model training, the radical embedding matrix is ​​updated according to the gradient of the loss function, and the optimal embedding of the radical is eventually learned.

[0133] For the input text sequence Each character in , obtain its radical information by querying the Chinese character dictionary, and obtain the corresponding radical embedding by querying the radical embedding matrix. For punctuation marks, special characters, etc. that do not contain radicals, use the zero vector as their radical embedding. Suppose the radical embedding sequence corresponding to the input text is obtained It is expressed as follows:

[0134]

[0135] in Represents characters The corresponding radical embedding vector.

[0136] (3) Chinese medicine dictionary matching word embedding

[0137] This solution uses the SoftLexicon method to introduce the potential domain word embedding corresponding to Chinese characters into the model, and uses the domain word information to enhance named entity recognition, thereby improving the accuracy of entity boundary recognition and entity recognition.

[0138] First, a dictionary in the field of traditional Chinese medicine is constructed by combining manual construction and automatic extraction. .

[0139] For characters , search the dictionary for its matching word set and divide all matching words into four sets: B, M, E and S according to the position information of the word in the word. Indicates The set of domain words that start with express The vocabulary set in the middle of the domain word, the set Indicates A collection of words ending with Indicates A single Chinese character is used as a vocabulary set for a domain word. The four sets are represented as shown in formulas (4)-(7):

[0140]

[0141]

[0142]

[0143]

[0144] in Represents input text Chinese-Israeli characters Starts with the character The word that ends, Indicates the character Starts with the character The word that ends, Indicates the character Starts with the character If the set is empty, the special word "None" is added to the set. Figure 3 As shown, taking the word "epidemic" as an example, Including "plague", "malaria", Contains "None", Contains "anti-plague", S Contains "plague".

[0145] For each word in the word set First, input it as a token into the encoding layer of LLM to obtain its corresponding word embedding :

[0146]

[0147] Then, the frequency of the word in the training text is used as the weight to weight the word embedding in the word set to obtain the embedding representation of the word set. As an example, its embedding representation The calculation is as follows:

[0148]

[0149] in Expressive words The frequency of occurrence in the training text, Represents characters The normalized sum of the frequencies of all words in the set of matching words, which is calculated as follows:

[0150]

[0151] in Indicates taking the set union.

[0152] Similar to equations (8) and (9), the vector representations of the other three sets can be obtained: 、 and The vector representations of the four sets are combined by adding them together to obtain the word Chinese medicine dictionary matching word embedding :

[0153]

[0154] This can be used to obtain the TCM dictionary matching word embedding sequence for the input question T:

[0155]

[0156] (4) Character-level feature fusion

[0157] A weighted fusion approach is adopted to obtain word-level embedding by fusing character embedding, radical embedding and TCM dictionary matching word embedding with learnable weights.

[0158] Defining learnable vectors ,in 、 、 Represents the weights corresponding to character embedding, radical embedding and TCM dictionary matching word embedding, and their initial values ​​are In order to ensure that the sum of the weight vector is 1, the softmax function is used to normalize the weight vector:

[0159]

[0160] in Represents the normalized weight value. Finally, the word-level embedding sequence is obtained according to the weight :

[0161]

[0162] Word-level embeddings

[0163] The structure diagram of the word-level embedding module is as follows Figure 4 Shown, including:

[0164] (1) Word embedding

[0165] Word segmentation is the process of breaking down input text into independent words. Since the segmented words often represent complete expressions and are typically units with independent meaning, good word segmentation results can help the model capture entity boundaries and improve named entity recognition.

[0166] For the field of traditional Chinese medicine, use the constructed dictionary of traditional Chinese medicine Enhance the accuracy of word segmentation in the field of traditional Chinese medicine and avoid entity boundary recognition errors caused by incorrect word segmentation. Use the open source Chinese natural language processing toolkit HanLP for word segmentation, and combine it with the constructed traditional Chinese medicine dictionary to ensure accurate division of the input text, so as to obtain the input text List of word segments :

[0167]

[0168] in, is the HanLP word segmenter, It represents the token of the i-th partition, and s is the number of tokens obtained by the partition.

[0169] Each token is then input into the LLM embedding layer for encoding to obtain a sequence of word embedding vectors. :

[0170]

[0171] in Indicates that LLM embedding layer is used for encoding, is the word embedding of the i-th token.

[0172] (2) Part-of-speech embedding

[0173] Part-of-speech information can help the model distinguish different semantics of the same word. For example, in traditional Chinese medicine, "heat" can be used as both a noun and an adjective. By adding part-of-speech information to word segmentation, the model can be assisted in making correct judgments on entity results.

[0174] Based on the word segmentation results, the HanLP tool is used to perform part-of-speech tagging based on the PKU part-of-speech tagging set, using Enhance the annotation of texts in the field of traditional Chinese medicine to obtain part-of-speech tag sequences :

[0175]

[0176] in represents the HanLP part-of-speech tagger, express The part-of-speech tagging results.

[0177] In the PKU POS tagging set, 43 kinds of POS are defined, and the initial POS embedding matrix is obtained by random initialization , where d is the dimension of the LLM embedding vector, is updated by backpropagation during the model training process. For a token and its POS , the corresponding POS embedding vector can be obtained according to the POS embedding matrix, and the POS embedding sequence of the input text is obtained.

[0178]

[0179] (3) Dependency syntax embedding

[0180] On the basis of tokenization and POS tagging, the dependency syntax analysis tool in the HanLP toolkit is introduced, the syntactic relationship between tokens is obtained according to the PMT dependency syntax system, and the dependency syntax tree is constructed , which is an adjacency matrix with a dimension of . The dependency syntax tree reveals the dependency relationship between different words in the sentence, which can help the model understand the composition and semantics of the sentence. It is represented as a tree structure, and each token in the text is a node of the tree, and the dependency relationship is the edge of the tree. There are 32 types of dependency relationships in the PMT dependency syntax system, and the initial dependency relationship embedding matrix is obtained by random initialization , where d is the dimension of the LLM embedding vector, is updated by backpropagation during the model training process.

[0181] Since the ordinary graph convolutional network is mainly used to process simple undirected graphs, and there are various dependency relationships in the dependency syntax tree and the dependency relationship between different tokens is directed, the CompGCN is used to encode the dependency syntax tree, and the dependency syntax embedding of the input text is obtained:

[0182]

[0183] , where represents encoding by CompGCN, is the token embedding vector sequence, is the initial dependency relationship embedding matrix, represents the encoding of the i-th token after CompGCN.

[0184] (4) Word-level feature fusion ​​

[0185] A weighted fusion approach is adopted to obtain word-level embedding by fusion of word segmentation embedding, part-of-speech embedding and dependency syntax embedding with learnable weights.

[0186] Defining learnable vectors ,in 、 、 Respectively represent the weights corresponding to word segmentation embedding, part-of-speech embedding, and dependency syntax embedding, and their initial values ​​are In order to ensure that the sum of the weight vector is 1, the softmax function is used to normalize it:

[0187]

[0188] in Represents the normalized weight value. Finally, the word-level embedding sequence is obtained according to the weight :

[0189]

[0190] Step 2: Obtain character fusion embedding through masked word-word attention

[0191] Use masked word-word attention to fuse word-level features and word-level features. Its structure is as follows Figure 5 As shown. Since a word segmentation token may correspond to multiple characters, the attention mask matrix is ​​constructed , where s is the input text Divide the number of tokens, n is the input text The number of characters in . Only when the attention value between the token and the corresponding characters that make up the token is non-zero, the other attention values ​​are all 0, ensuring that the attention is focused on the characters that make up the current word. Represents the mask value of the i-th row and j-th column in the mask matrix, then:

[0192]

[0193] Let the character-level embedding vector be the query, the word-level embedding vector be the key and value, calculate the attention weight through dot product and matrix multiplication, and normalize it through the softmax function to get the attention weight of each character and each token. When calculating the attention weight, add the character-word mask matrix , the attention weight is calculated only between the token and the corresponding characters that make up the token. Attention score matrix The calculation is as follows:

[0194]

[0195] Finally, the character fusion embedding that integrates word-level embedding information is obtained through attention-weighted character-level embedding and word-level embedding. :

[0196]

[0197] Step 3: Use bidirectional LSTM encoding

[0198] Fusion embedding via bidirectional LSTM network Encoding further captures contextual information and long-range dependencies, obtaining the final embedding for each character. The bidirectional LSTM consists of two independent LSTM layers. The forward LSTM processes the input sequence from beginning to end, while the backward LSTM processes the input sequence from end to beginning. Finally, the hidden layer sequence of the sentence is obtained by splicing them together according to position. The process is formally expressed as follows:

[0199]

[0200]

[0201]

[0202] in Represents the input features at time step i, that is, character Corresponding fusion embedding, and Represent the forward output features and backward output features of LSTM at time step i, respectively. Represents vector concatenation operation, Represents characters Multi-dimensional feature embedding after integrating text context features. Thus, the multi-dimensional feature embedding sequence of the input sentence can be expressed as follows:

[0203]

[0204] Step 4: Obtain candidate spans using the boundary classifier and dynamic span length constraint module

[0205] (1) Boundary Classifier

[0206] The start boundary classifier and the end boundary classifier are used to predict whether the current character is the start or end position of the named entity, that is, the start boundary and the end boundary. Both the start boundary classifier and the end boundary classifier are implemented based on the multi-layer perceptron (MLP). For each character The corresponding hidden layer output , first encoded by a multi-layer perceptron, and then encoded by a sigmoid activation function to predict the probability of it being the starting boundary or the ending boundary. The probability of the starting boundary and the probability of the end boundary The calculation is as follows:

[0207]

[0208]

[0209] in and The multilayer perceptron modules representing the start boundary and end boundary classifiers are composed of two linear layers and ReLU activation functions. , if the predicted probability or If the value is greater than the threshold, the character Treated as a solid start or end boundary.

[0210] (2) Dynamic span length constraint

[0211] After obtaining the starting and ending boundaries, previous work has used a maximum entity length hyperparameter to generate all possible candidate entity spans based on the starting and ending boundaries within this constraint. The setting of this hyperparameter is directly related to the quality of candidate span generation and the efficiency of the model. If the maximum entity length is set too large, while it ensures that long entities are captured, it will generate too many candidate spans, including a large number of non-entity spans, increasing computational complexity and affecting model training and prediction efficiency. If the maximum entity length is set too small, some entities may be missed. Therefore, it is usually necessary to determine the optimal maximum entity length suitable for a specific domain and specific dataset through experimentation.

[0212] For TCM texts, the entities are often relatively fixed professional terms, especially TCM terminology such as drugs, diseases, prescriptions and acupoints. Therefore, this solution proposes a dynamic span length constraint to dynamically define the maximum length of the entity span. If it is determined by the starting boundary classifier to be a potential entity starting character, then in the Chinese medicine dictionary Search by is the set of all matching words starting from Then introduce the weight function , this function reflects the character The word appears when it is the starting character The probability of , which is calculated as follows:

[0213]

[0214] in Expressive words The frequency of occurrence in the training corpus.

[0215] Then calculate the word The weighted length of the starting entity :

[0216]

[0217] Next, a regression prediction is introduced based on the obtained multidimensional feature embedding that integrates text context features. , predict the length offset based on context information :

[0218]

[0219] in Represents a multi-layer perceptron for regression prediction. Finally, the character Maximum length constraint for the starting entity for:

[0220]

[0221] For each starting boundary, get its position and distance The ending boundaries within are paired to form candidate spans.

[0222] Step 5: Span Classification

[0223] In order to identify nested named entities in text, it is necessary to classify candidate spans and determine whether they belong to predefined entity types. For each candidate span, all its character embeddings are aggregated through average pooling to obtain the span feature vector, which is then input into the multi-layer perceptron to predict its entity type and calculate its conditional probability through softmax. Suppose one of the candidate spans is arrive The word embedding set of all words in the span is , and obtain its span feature vector by average pooling :

[0224]

[0225] Then use multi-layer perceptron and softmax to classify the span feature vector and predict its probability of belonging to each entity type. :

[0226]

[0227] in Represents a multi-layer perceptron module for span prediction.

[0228] Step 6: Model training

[0229] Model training includes three parts of loss: boundary detection, span estimation, and named entity recognition. Boundary detection is to detect whether a character is the start boundary or the end boundary. It can be regarded as a binary classification problem. The loss function is is the starting boundary cross entropy loss Cross entropy loss with end margin sum:

[0230]

[0231]

[0232]

[0233] in and is the true label of whether the i-th character is the starting boundary or the ending boundary, and For characters is the probability of the starting boundary or the ending boundary, and n is the number of characters.

[0234] The span estimation module is a regression task and is trained using the mean squared error loss function. The definition is as follows:

[0235]

[0236] Where N is the number of training samples, is the actual maximum length constraint value, let The end position of the longest entity starting with character , the first entity after this entity starts at ,but The calculation is as follows:

[0237]

[0238] in Indicates the rounding operation, i indicates The offset within the input text.

[0239] The named entity recognition module is a multi-classification problem, and the model is trained by minimizing the cross entropy loss function. The definition is as follows:

[0240]

[0241] in is the number of entity categories, is the probability predicted by the model that the span belongs to the j-th entity, is the true label, which is 1 when the span belongs to the j-th entity and 0 otherwise.

[0242] During model training, minimize the following joint training loss :

[0243]

[0244] in 、 and is a weight hyperparameter that controls the importance of each task to the model.

Claims

1. A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints, characterized in that: The following steps are involved: S1. The model performs character-level embedding and word-level embedding on the input TCM text, respectively, to obtain character-level embedding and word-level embedding, and defines the input TCM text as the input text sequence ; The specific method of the word-level embedding process is to obtain The radical embedding, character embedding and TCM dictionary matching word embedding are then fused to obtain the word-level embedding; The radical embedding is obtained The radical information of each character in , and then the corresponding radical embedding is obtained by querying the radical embedding matrix, and finally the The corresponding radical embedding sequence, where the initial radical embedding matrix is ​​obtained by random initialization based on the Chinese character radical standard and then updated during the training process; the character embedding is obtained The embedded information of each character in Each character in is input as a separate token into the large language model encoding layer to obtain the corresponding initial character embedding, and then obtain The corresponding initial character embedding sequence; the Chinese medicine dictionary matching word embedding is obtained The TCM dictionary matches the word embedding sequence, specifically to build a TCM dictionary, using the SoftLexicon method to obtain Finally, the word-level embedding is obtained by fusion of radical embedding, character embedding and TCM dictionary matching word embedding through learnable weights; The specific method of word-level embedding processing is to obtain The word embedding, part-of-speech embedding and dependency syntactic embedding are obtained, and then the word-level embedding, part-of-speech embedding and dependency syntactic embedding are fused to obtain word-level embedding. Before obtaining the word embedding, Perform word segmentation processing, specifically combining the Chinese natural language processing toolkit HanLP and the constructed Chinese medicine dictionary to obtain The word list is then encoded by a large language model to obtain a word embedding vector sequence; the part-of-speech embedding is based on the word list, and then the HanLP tool is used to perform part-of-speech tagging on each token in the word list based on the PKU part-of-speech tagging set. The initial part-of-speech embedding matrix is ​​obtained by random initialization based on the PKU part-of-speech tagging set. For the word list and the corresponding part-of-speech tagging set, the part-of-speech embedding matrix is ​​obtained. The part-of-speech embedding sequence, where the part-of-speech embedding matrix is ​​updated during the training process; the dependency syntactic embedding is based on the word segmentation embedding vector sequence, combined with the dependency syntactic tree and the dependency relation embedding matrix, and is obtained by encoding based on the combined multi-relation graph convolutional network (CompGCN) The dependency syntax tree is constructed based on the word segmentation list and part-of-speech tags according to the PMT dependency syntax system. The dependency relationship embedding matrix is ​​obtained by randomly initializing the dependency relationship types in the PMT dependency syntax system and updating it during training. Finally, the word-level embedding is obtained by fusion of word segmentation embedding, part-of-speech embedding and dependency syntax embedding through learnable weights. S2, the model obtains fused embedding by weighting character-level embedding and word-level embedding through attention; S3, the model encodes the fusion embedding through a bidirectional LSTM network to obtain Multi-dimensional feature embedding for each character in ; S4. The model sequentially inputs the acquired multidimensional feature embeddings into a boundary classifier to predict whether the character corresponding to the current multidimensional feature embedding is the start or end position of the named entity, where the start position and the end position correspond to the start boundary and the end boundary, respectively. The boundary classifier is implemented using a multi-layer perceptron. After obtaining the start boundary and the end boundary from the input multidimensional feature embeddings, a dynamic span length constraint is used to dynamically define the maximum length of the entity span, thereby generating a candidate entity span. S5. The model classifies the generated candidate entity spans to determine whether they belong to the predefined entity type. Specifically, all character embeddings of the candidate entity spans are aggregated through average pooling to obtain the span feature vector, which is then input into the multi-layer perceptron to predict the entity type to which it belongs. Finally, the conditional probability is calculated through softmax. S6. Use known TCM texts to train the model described in S1-S5. The training process includes three losses: boundary detection, span estimation, and named entity recognition. Boundary detection corresponds to the sum of the starting boundary cross entropy loss function and the ending boundary cross entropy loss function, span estimation corresponds to the mean square error loss function, and named entity recognition corresponds to minimizing the cross entropy loss function. After training, a trained model is obtained. S7. Input the TCM text to be recognized into the trained model to obtain the named entity recognition result.

2. A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints according to claim 1, characterized in that: In S1, the specific method of obtaining character embedding is to define ,in represents the i-th character, , n represents the number of characters in the input text, As a single token, it is fed into the Large Language Model (LLM) encoding layer to obtain an initial embedding. : , in Represents the LLM encoding layer; by initially embedding each character, the initial character embedding sequence of the input text is obtained : , in, , is the dimension of the vector of the character embedded by LLM; The specific method of obtaining radical embedding is to Each character in , by querying the Chinese character dictionary, obtain its radical information, and embed the radical matrix Query to get the corresponding radical embedding, It is obtained by random initialization based on the 201 main radicals specified in the Chinese radical standard GF0011-2009, expressed as For characters without radicals, use zero vector as radical embedding; define Corresponding radical embedding sequence for: , in Represents characters The corresponding radical embedding vector; The specific method of obtaining the embedding of matching words in the TCM dictionary is to construct the TCM dictionary by combining manual construction and automatic extraction. , for characters ,exist Query its matching word set and divide all matching words into four sets B, M, E and S according to the position information of the characters in the word, where the set Indicates is the starting vocabulary set, set express A collection of words in the middle, a collection Indicates A collection of words ending with Indicates A single Chinese character is used as a vocabulary set of words, which are represented as follows: , , , , in express Chinese-Israeli characters Starts with the character The word that ends, Indicates the character Starts with the character The word that ends, Indicates the character Starts with the character The ending word. If the set is empty, the special word "None" is added to the set. Get Character Vocabulary collection 、 、 and Then, get the vector representation of each set 、 、 and Specifically, we first calculate the word embedding of each word in each set : , in It refers to any word in each set; by weighting the word embedding in the word set, the vector representation of the word set is obtained. , vector representation for: , in, Expressive words The frequency of occurrence in the text, Represents characters The normalized sum of the frequencies of all words in the matching word set: , The same logic applies 、 and , by adding Chinese medicine dictionary matching word embedding : , Thus we get The TCM dictionary matches the word embedding sequence: ; Defining learnable vectors ,in 、 、 They represent the weights corresponding to character embedding, radical embedding, and TCM dictionary matching word embedding, respectively, and are normalized using the softmax function: , in Represents the normalized weight value, and the word-level embedding sequence is obtained according to the weight : 。 3. A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints according to claim 2, characterized in that: In S1, during the acquisition of word embedding, the definition of The word list is : , in, is the HanLP word segmenter, represents the token of the i-th partition, and s is the number of tokens obtained by partitioning; Each token is input into the LLM embedding layer for encoding, and the word segmentation embedding vector sequence is obtained as follows: : , In the process of acquiring part-of-speech embedding, the obtained part-of-speech tag sequence is defined as : , in represents the HanLP part-of-speech tagger, express The part-of-speech tagging results of According to the 43 parts of speech defined in the PKU part-of-speech tagging set, the initial part-of-speech embedding matrix is ​​obtained by random initialization. ,in is the dimension of the LLM embedding vector, for the tag and its part of speech , according to the part-of-speech embedding matrix, obtain its corresponding part-of-speech embedding vector , thus obtaining Part-of-speech embedding sequence : ; In the process of acquiring dependency syntax embedding, the grammatical relationship between tokens is obtained according to the PMT dependency syntax system, and the dependency syntax tree is constructed. , the dimension is There are 32 types of dependency relationships in the PMT dependency syntax system. The initial dependency embedding matrix is ​​obtained by random initialization. , use CompGCN to encode the dependency syntax tree to obtain Dependency syntactic embeddings : , in Indicates encoding via CompGCN, is the word embedding vector sequence, is the initial dependency embedding matrix, Indicates the encoding of the i-th token after CompGCN; Defining learnable vectors ,in 、 、 They represent the weights corresponding to word segmentation embedding, part-of-speech embedding, and dependency syntax embedding, respectively, and are normalized using the softmax function: , in Represents the normalized weight value, and the word-level embedding sequence is obtained according to the weight : 。 4. A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints according to claim 3, characterized in that: The specific method of S2 is: Constructing the attention mask matrix , where s is The number of tokens divided, n is The number of characters in the definition Represents the mask value of the i-th row and j-th column in the mask matrix: , Let the character-level embedding vector be the query, the word-level embedding vector be the key and value, calculate the attention weight by dot product and matrix multiplication, and normalize it by softmax function to get the attention score matrix : , By using attention-weighted character-level embedding and word-level embedding, we can obtain character fusion embedding that integrates word-level embedding information. : 。 5. A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints according to claim 4, characterized in that: The specific method for S3 is: Fusion embedding via bidirectional LSTM network For encoding, the bidirectional LSTM consists of two independent LSTM layers. The forward LSTM processes the input sequence from beginning to end, and the backward LSTM processes the input sequence from end to beginning. Finally, the hidden layer sequence of the sentence is obtained by splicing by position: , , , in Represents the input features at time step i, that is, character Corresponding fusion embedding, and Represent the forward output features and backward output features of LSTM at time step i, respectively. Represents vector concatenation operation, Represents characters The multidimensional feature embedding after integrating the text context features; thus the multidimensional feature embedding sequence of the input sentence is expressed as follows: 。 6. A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints according to claim 5, characterized in that: In S4, predict characters The probability of the starting boundary and the probability of the end boundary The calculation is as follows: , , in and The multilayer perceptron modules representing the start boundary and end boundary classifiers are composed of two linear layers and ReLU activation functions; by setting the confidence threshold , if the predicted probability or If the value is greater than the threshold, the character Considered as the entity's starting or ending boundary; The method of using dynamic span length constraint to dynamically define the maximum length of entity span is to If it is determined by the starting boundary classifier to be a potential entity starting character, then in the Chinese medicine dictionary Search by is the set of all matching words starting from , and then introduce the weight function , this function reflects the character The word appears when it is the starting character The probability of , which is calculated as follows: , in Expressive words The frequency of occurrence in the training corpus is then calculated by The weighted length of the starting entity : , Finally, the length offset is predicted based on the context information : , in Represents a multi-layer perceptron for regression prediction; ultimately, the character Maximum length constraint for the starting entity for: , For each starting boundary, get its position and distance The ending boundaries within are paired to form candidate spans.

7. A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints according to claim 6, characterized in that: In S5, a candidate span is defined by arrive The word embedding set of all words in the span is , and obtain its span feature vector by average pooling : , Then use multi-layer perceptron and softmax to classify the span feature vector and predict its probability of belonging to each entity type. : , in Represents a multi-layer perceptron module for span prediction.

8. A method for recognizing nested named entities in traditional Chinese medicine based on dynamic span length constraints according to claim 7, characterized in that: In S6, the loss function of boundary detection is the starting boundary cross entropy loss Cross entropy loss with end margin sum: , , , in and is the true label of whether the i-th character is the starting boundary or the ending boundary, and For characters is the probability of the starting boundary or the ending boundary, and n is the number of characters; Loss function for span estimation The definition is as follows: , Where N is the number of training samples, is the actual maximum length constraint value, defined as The end position of the longest entity starting with character , the first entity after this entity starts at ,but The calculation is as follows: , in Indicates the rounding down operation, i indicates The offset in the input text; Loss function for named entity recognition The definition is as follows: , in is the number of entity categories, is the probability predicted by the model that the span belongs to the j-th entity, is the true label, which is 1 when the span belongs to the j-th entity and 0 otherwise; During model training, minimize the joint training loss : , in 、 and is a weight hyperparameter that controls the importance of each task to the model.

Citation Information

Patent Citations

  • Nested named entity recognition method based on word information fusion and boundary detection

    CN118095283A

  • Electronic medical record nested named entity identification method

    CN118821778A