Chinese named entity recognition method based on local and global character representation enhancement
By integrating local and global character information of Chinese characters, using the self-coding mechanism and interactive gating mechanism, the performance of the Chinese named entity recognition model is enhanced, and the problems of existing methods in entity boundary judgment and semantic ambiguity are solved, and more accurate entity recognition is achieved.
Patent Information
- Application Number
- CN202211273187.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-10-18
AI Technical Summary
The existing Chinese named entity recognition method is difficult to effectively utilize the local and global character information of Chinese characters, resulting in errors in entity boundary judgments and semantic ambiguity, which affects the entity recognition performance of the model.
Using a method based on local and global character representation enhancement, the spatial information and sequence information of the side shape are fused through the autoencoding mechanism, and combined with the interactive gating mechanism, comprehensive character representation is obtained to enhance the semantics and potential boundary information of characters.
It improves the performance of the Chinese named entity recognition model, enhances the ability to represent specific fields and closely related entities, and is better than the existing Chinese NER model based on external information.
Smart Images

Figure CN115455955B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a Chinese named entity recognition method based on local and global character representation enhancement, and belongs to the technical field of natural language processing. Background Art
[0002] Chinese named entity recognition (CNER) is a basic information extraction task that plays a vital role in natural language processing (NLP) applications such as information retrieval, text automatic summarization, question answering, machine translation, knowledge graphs, etc. The goal of CNER is to extract some predefined specific entities from a sentence and identify their correct types, such as person, location, and organization. For CNER, current methods are mainly based on deep learning, which regards it as a sequence labeling task. Unlike statistical-based methods, deep learning-based methods use distributed representations instead of manually designed features to represent characters. Using encoders such as LSTM, CNN, and Transformer, each character in the text is serialized. Then, the automatically labeled sequence can be decoded according to the labeling scheme, and the named entities consisting of multiple characters in the text can be integrated.
[0003] The composition of Chinese named entities is more complex than that of English named entities. Chinese characters can be regarded as a concept between characters and words in English. Chinese characters have more semantics than English characters, but less semantics than words. Some Chinese characters have their own independent meanings, but more Chinese characters need to be combined with other Chinese characters to form a meaningful word. As the basic unit of text, Chinese characters do not have clear word segmentation symbols. The fuzzy word boundaries will cause a lot of boundary ambiguity, which increases the difficulty of defining the boundaries of Chinese named entities. Therefore, word boundary information is indispensable in Chinese. There are many methods to enhance word boundary information by combining dictionary information in CNER tasks. For example, structures such as Lattice and SoftLexicon fuse word embeddings on character embeddings to represent characters to enhance entity boundaries and type information. There are also many methods to introduce external information in CNER tasks. For example, external dictionaries, strokes, pinyin, radicals and glyph features are used as auxiliary information to further enhance the semantic representation ability of embedded vectors. These methods have been proven to be effective and help improve the performance of named entity recognition models. For example, Flat-Lattice, with the help of the powerful Transformer and carefully designed positional encoding, can make full use of Lattice information, has excellent parallelization capabilities, and gives full play to the advantages of the model in capturing remote context dependencies. However, the incorrect introduction of word information will inevitably lead to problems such as incorrect entity boundary judgment and semantic ambiguity, which will affect the entity recognition performance of such models. This leads to a decrease in the accuracy of the entity extraction model. On the other hand, the glyph structure of Chinese characters has independent semantics and represents a specific entity meaning. This glyph structure is called the local information of the character. Specifically, Chinese is a pictographic language consisting of a radical and a phonetic radical. The radical has a strong semantic function, and Chinese characters with the same radical have similar entity meanings. There are relatively few models that use character glyphs for enhancement, and there are still the following shortcomings: (1) The model only extracts features from one aspect of the glyph structure or radical, which limits the model's comprehensive learning of glyph representation. (2) After the model extracts the glyph representation, there is no good method for weighted fusion with its own embedding layer vector, which will affect the results of the NER model.
[0004] In view of the above problems, the present invention proposes a Chinese named entity recognition method based on local and global character representation enhancement. The current mainstream NER method does not consider the comprehensive spatial and sequence character information of Chinese characters. Since the underlying layer of Chinese characters itself carries a large amount of semantic information, it is important to effectively extract it and apply it to the NER task. From this perspective, the present invention uses the radical structure and sequence of the characters to enhance the potential boundaries and semantic information of the characters, and uses the interactive gating mechanism to effectively obtain the local and global information of the comprehensive characters, thereby improving the performance of the character-based NER model. Theoretical and technical verifications were carried out on the Chinese named entity datasets IMCS21 and CMeEE, and the experimental results fully demonstrated the effectiveness of the method. Summary of the invention
[0005] In order to solve the above problems, the present invention provides a Chinese named entity recognition method based on local and global character representation enhancement. The present invention utilizes an autoencoding mechanism to fuse different local information of characters such as spatial information and sequence information of radicals, and utilizes an interactive gating mechanism to control the contribution of local and global character information to character representation, thereby obtaining a comprehensive character representation to enhance character representation, enhancing the semantics and potential boundary information of characters, and enabling the main model to obtain better entity recognition capability. The proposed method is evaluated on two Chinese NER benchmark datasets, and various experimental results not only prove the effectiveness of the method, but also show that the method can improve the representation capability of specific fields and closely related entities.
[0006] The technical solution of the present invention is: a Chinese named entity recognition method based on local and global character representation enhancement, the method comprising the following steps:
[0007] Step 1: Use the character vector trained on the corpus as the initial embedding of the character: map each character to a dense vector representation to obtain the character embedding of each sentence;
[0008] Step 2: Split the character into radicals and other character components, and then use the sequence feature encoder to extract the character's glyph sequence features;
[0009] Step 3: Treat a single character as a two-dimensional image and obtain the glyph structure features of the character through an image feature encoder; the image corresponding to the Chinese character passes through multiple convolutional layers to capture low-level graphic features, and then uses adaptive pooling operations and applies group convolution to map to the final glyph structure features;
[0010] Step 4: Using the self-encoding mechanism, the three vectors of character shape structure features, character shape sequence features and pre-trained character embedding are integrated to obtain the local representation of the character;
[0011] Step 5: First, use the word2vec Skip-Gram model on the domain corpus to train a domain dictionary. Then, query and match each character in the dictionary to obtain several word sets, and then obtain the global representation of the character by weighted distribution and concatenation.
[0012] Step 6. After obtaining the local and global representations of the characters, the interactive gating mechanism is used to filter the features of the two to obtain a comprehensive representation. The comprehensive representation is then sent to the Bi-LSTM for context encoding, and then CRF is used as a decoding layer to obtain the label of the output result.
[0013] As a further solution of the present invention, in Step 1, the input sentence is regarded as a character sequence s={c 1 , c 2 ,···,c n}, and then each character c i are mapped to a dense vector representation Get the character embeddings for each sentence:
[0014]
[0015] where e c Represents a character embedding lookup table.
[0016] As a further solution of the present invention, in Step 2, first, a word-splitting dictionary is used for each word in the data set to construct a query table containing the components of each word; then the character-splitting sequence is sent to a convolutional neural network CNN to extract the character's glyph sequence features, and then a residual network is used to optimize the convolutional layer to alleviate the gradient vanishing problem as the neural network deepens. Finally, the glyph sequence feature embedding is obtained using a maximum pool and a fully connected layer.
[0017] As a further solution of the present invention, Step 2 comprises the following steps:
[0018] Step 2.1, replace the i-th character c i Split into K parts, If the length of a character component is less than K, the empty position is replaced by " <pad>" to fill, and then perform a random embedding operation E on each character component r :
[0019]
[0020] Step 2.2: Randomly embed the obtained characters into a sequence Send it to the convolution operation conv3 with a convolution kernel size of 3 to get the character hidden vector sequence
[0021]
[0022] Step 2.3, perform max-pooling on the vector corresponding to each character component in the character latent vector sequence, and then send it to a fully connected layer f c Perform dimension transformation to get the glyph sequence embedding of the character The glyph sequence embedding dimension of this character is d o ;
[0023]
[0024] As a further solution of the present invention, in Step 3, the glyph structure features can obtain rich pictographic information from the character image to improve the performance of the Chinese named entity recognition model; for images of different font types, they are spliced together to represent the structural image of the character, and the character image is passed through multiple convolutional layers and multiple output channels to capture low-level glyph structure features.
[0025] As a further solution of the present invention, Step 3 comprises the following steps:
[0026] Step 3.1, c i The characters are converted into grayscale images corresponding to 6 different fonts in is an 8-bit grayscale image of the jth font with a size of 12×12. Different image matrices are concatenated to obtain character c i Structural image of
[0027]
[0028] Where concat represents the concatenation operation;
[0029] Step 3.2, then, use the convolution operation conv1 with a convolution kernel size of 5×5 and 384 output channels to capture low-level graphic features and obtain the hidden layer vector
[0030]
[0031] Step 3.3, use the maxpooling operation with a template size of 4×4, The resolution is reduced from 8×8 to 2×2; then a convolution kernel with a size of 1×1 and d s The convolution operation conv2 of the output channels is used to obtain the hidden layer vector
[0032]
[0033] Step 3.4, finally, The convolution kernel size is 2, and the group convolution operation groupconv is performed, and the dimension transformation operation reshape is performed to obtain the glyph structure representation of the character. The glyph structure embedding dimension of this character is d s ;
[0034]
[0035] Reshape represents a dimensional transformation that converts a 2D vector into a 1D vector.
[0036] As a further solution of the present invention, in the Step 4, the character's glyph structure features, glyph sequence features and pre-trained character embedding are first concatenated, and then an automatically fused latent vector is obtained through a transformation layer. Then, an attempt is made to reconstruct the originally concatenated vector from the self-fused latent vector. Finally, the Euclidean distance between the original vector and the reconstructed vector is calculated, and the loss is calculated using the mean square error to obtain information that has been compressed by the intermediate layer but without loss.
[0037] As a further solution of the present invention, in Step 4, an autoencoder network is used to fuse three vectors, namely, the glyph structure features, the glyph sequence features and the pre-trained character embedding, and the model is encouraged to extract multi-granularity features by maximizing the correlation between inputs of different granularities. Specifically, Step 4 includes the following steps:
[0038] Step 4.1, firstly, the glyph structure features Character sequence features and character embedding Perform splicing to obtain the initial splicing vector
[0039]
[0040] Among them, d s Represents the character’s glyph structure embedding dimension size, d o The glyph sequence embedding dimension size of the character, d c Represents the pre-trained embedding dimension size of the character;
[0041] Step 4.2, then Perform the following two linear transformations and activations to obtain the latent vector
[0042]
[0043] Step 4.3, use Reconstruct the original concatenated vector and obtain the reconstructed vector
[0044]
[0045] Step 4.4, use the mean square error loss function to calculate and Loss f :
[0046]
[0047] Step 4.5: Add the loss to the main model sequence labeling model, stimulate the above reconstruction process through the NER downstream task, obtain the information compressed by the middle layer but without loss, and convert the latent vector of the middle layer into As a local representation of fusion.
[0048] As a further solution of the present invention, Step 5 comprises the following steps:
[0049] Step 5.1, character c i Perform query matching in a dictionary D pre-trained using the Skip-Gram model; if a word w in D contains character c i , then according to the different positions of the characters in the word, they are included in four word sets B(c i ),M(c i ),E(c i ),S(c i ); Specifically, if c i appears at the beginning of a word w, the word w is classified into the word set B (c i ); if c i Appears in the middle position of a word w, the word w is classified into the word set M (c i ); if c i If it appears at the end of a word w, the word w is classified into the word set E(c i ); if c i Same as a word w, that is, the character is an independent word, then the word w is classified into the word set S (c i );
[0050] Step 5.2. Count the characters c i The number of times a word w appears in the training data, and the character c i The total number of times all the matched words appear in the training set data is M, then the character c i The frequency of a word w that matches for:
[0051]
[0052] Step 5.3, match the word set B (c i ) multiplies the word vector of each word by its weight and adds them together to get the character c i As a representation of the beginning of a word
[0053]
[0054] Among them, E d (w) represents the embedding vector of word w;
[0055] Step 5.4, loop through the same method in Step 5.3 to get character c i As a representation of the middle character of a word As a representation of the last character of a word and representation as an independent word
[0056] Step 5.5, change the character c i The four representations of are combined to obtain the global representation of each character d g Indicates the size of the global representation dimension of the character;
[0057]
[0058] As a further solution of the present invention, in Step 6, since local representation may have information redundancy compared with global representation, an interactive gating mechanism is used to screen the features of both to obtain a reasonable comprehensive representation. Since the context information in the sentence is helpful for sequence modeling, a bidirectional LSTM network that can capture bidirectional information of the text is used to extract sentence context features. The conditional random field CRF is used as a decoding layer, and the vector encoded by the context encoder is sent to the CRF to find the label sequence with the highest probability by minimizing the negative maximum likelihood function.
[0059] The beneficial effects of the present invention are as follows: the present invention enhances character representation by incorporating local information of Chinese character radicals and global information of domain terms, thereby enhancing the semantics and potential boundary information of characters, and enabling the main model to obtain better entity recognition capabilities. Compared to the Chinese NER model based on external information, the method of the present invention utilizes an autoencoder network in combination with glyph information at the embedding layer, and uses an interactive gating mechanism to filter local information and global information of characters, so that the main model can accurately identify the boundaries and categories of domain entities. Various experimental results not only prove the effectiveness of the model of the present invention, but also show that the main model of the present invention can improve the representation capabilities of specific domains and closely related entities. The performance of the main model of the present invention on two benchmark Chinese data sets is substantially better than that of existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flow chart of the present invention;
[0061] Figure 2 A model diagram for extracting character sequence features proposed by the present invention;
[0062] Figure 3 A model diagram for extracting the structural features of a character shape proposed by the present invention;
[0063] Figure 4 This is a line graph of the hidden layer vector dimension experiment after the self-encoding of the present invention. DETAILED DESCRIPTION
[0064] Example 1: Figure 1-Figure 4 As shown, a Chinese named entity recognition method based on local and global character representation enhancement includes the following steps:
[0065] Step 1. This invention uses two data sets. One is the IMCS21 data set provided by the China Conference on Computational Linguistics (CCL), which includes more than 60,000 sentences. The other is the CMeEE data set, which contains more than 20,000 sentences. The specific data of these two data sets are shown in Table 1:
[0066] Table 1 Dataset statistics
[0067]
[0068] Use the character vectors trained on the corpus as the initial embedding of the characters: map each character to a dense vector representation to obtain the character embedding of each sentence;
[0069] As a further solution of the present invention, in Step 1, the input sentence is regarded as a character sequence s={c 1 , c 2 ,···,c n }, and then each character c i are mapped to a dense vector representation Get the character embeddings for each sentence:
[0070]
[0071] where e c Represents a character embedding lookup table.
[0072] Step 2: Split the character into radicals and other character components, and then use the sequence feature encoder to extract the character's glyph sequence features;
[0073] In the Step 2, a word-splitting dictionary is first used to construct a component query table for each word in the data set; then the character-splitting sequence is sent to a convolutional neural network (CNN) to extract the character shape sequence features, and then a residual network is used to optimize the convolutional layer to alleviate the gradient vanishing problem as the depth of the neural network increases. Finally, the maximum pool and fully connected layer are used to obtain the shape sequence feature embedding.
[0074] As a further solution of the present invention, Step 2 comprises the following steps:
[0075] Step 2.1, replace the i-th character c i Split into K parts, If the length of a character component is less than K, the empty position is replaced by " <pad>" to fill, and then perform a random embedding operation E on each character component r :
[0076]
[0077] Step 2.2: Randomly embed the obtained characters into a sequence Send it to the convolution operation conv3 with a convolution kernel size of 3 to get the character hidden vector sequence
[0078]
[0079] Step 2.3, perform max-pooling on the vector corresponding to each character component in the character latent vector sequence, and then send it to a fully connected layer f c Perform dimension transformation to get the glyph sequence embedding of the character The glyph sequence embedding dimension of this character is d o ;
[0080]
[0081] Step 3: Treat a single character as a two-dimensional image and obtain the glyph structure features of the character through an image feature encoder; the image corresponding to the Chinese character passes through multiple convolutional layers to capture low-level graphic features, and then uses adaptive pooling operations and applies group convolution to map to the final glyph structure features;
[0082] As a further solution of the present invention, in Step 3, the glyph structure features can obtain rich pictographic information from the character image to improve the performance of the Chinese named entity recognition model; for images of different font types, they are spliced together to represent the structural image of the character, and the character image is passed through multiple convolutional layers and multiple output channels to capture low-level glyph structure features.
[0083] As a further solution of the present invention, Step 3 comprises the following steps:
[0084] Step 3.1, c i The characters are converted into grayscale images corresponding to 6 different fonts in is an 8-bit grayscale image of the jth font with a size of 12×12. Different image matrices are concatenated to obtain character c i Structural image of
[0085]
[0086] Where concat represents the concatenation operation;
[0087] Step 3.2, then, use the convolution operation conv1 with a convolution kernel size of 5×5 and 384 output channels to capture low-level graphic features and obtain the hidden layer vector
[0088]
[0089] Step 3.3, use the maxpooling operation with a template size of 4×4, The resolution is reduced from 8×8 to 2×2; then a convolution kernel size of 1×1 and d s The convolution operation conv2 of the output channels is used to obtain the hidden layer vector
[0090]
[0091] Step 3.4, finally, The convolution kernel size is 2, and the group convolution operation groupconv is performed, and the dimension transformation operation reshape is performed to obtain the glyph structure representation of the character. The glyph structure embedding dimension of this character is d s ;
[0092]
[0093] Reshape represents a dimensional transformation that converts a 2D vector into a 1D vector.
[0094] Step 4: Using the self-encoding mechanism, the three vectors of character shape structure features, character shape sequence features and pre-trained character embedding are integrated to obtain the local representation of the character;
[0095] As a further solution of the present invention, in the Step 4, the character's glyph structure features, glyph sequence features and pre-trained character embedding are first concatenated, and then an automatically fused latent vector is obtained through a transformation layer. Then, an attempt is made to reconstruct the originally concatenated vector from the self-fused latent vector. Finally, the Euclidean distance between the original vector and the reconstructed vector is calculated, and the loss is calculated using the mean square error to obtain information that has been compressed by the intermediate layer but without loss.
[0096] As a further solution of the present invention, in Step 4, an autoencoder network is used to fuse three vectors, namely, the glyph structure features, the glyph sequence features and the pre-trained character embedding, and the model is encouraged to extract multi-granularity features by maximizing the correlation between inputs of different granularities. Specifically, Step 4 includes the following steps:
[0097] Step 4.1, firstly, the glyph structure features Character sequence features and character embedding Perform splicing to obtain the initial splicing vector
[0098]
[0099] Among them, d s Represents the character’s glyph structure embedding dimension size, d o The glyph sequence embedding dimension size of the character, d c Represents the pre-trained embedding dimension size of the character;
[0100] Step 4.2, then Perform the following two linear transformations and activations to obtain the latent vector
[0101]
[0102] Step 4.3, use Reconstruct the original concatenated vector and obtain the reconstructed vector
[0103]
[0104] Step 4.4, use the mean square error loss function to calculate and Loss f :
[0105]
[0106] Step 4.5: Add the loss to the main model sequence labeling model, stimulate the above reconstruction process through the NER downstream task, obtain the information compressed by the middle layer but without loss, and convert the latent vector of the middle layer into As a local representation of fusion.
[0107] Step 5: First, use the word2vec Skip-Gram model on the domain corpus to train a domain dictionary. Then, query and match each character in the dictionary to obtain several word sets, and then obtain the global representation of the character by weighted distribution and concatenation.
[0108] As a further solution of the present invention, Step 5 comprises the following steps:
[0109] Step 5.1, character c i Perform query matching in a dictionary D pre-trained using the Skip-Gram model; if a word w in D contains character c i , then according to the different positions of the characters in the word, they are included in four word sets B(c i ),M(c i ),E(c i ),W(c i ); Specifically, if c i appears at the beginning of a word w, the word w is classified into the word set B (c i ); if c i Appears in the middle position of a word w, the word w is classified into the word set M (c i ); if c i If it appears at the end of a word w, the word w is classified into the word set E(c i ); if c i Same as a word w, that is, the character is an independent word, then the word w is classified into the word set S (c i );
[0110] Step 5.2. Count the characters c i The number of times a word w appears in the training data, and the character c i The total number of times all the matched words appear in the training set data is M, then the character c i The frequency of a word w that matches for:
[0111]
[0112] Step 5.3, match the word set B (c i ) multiplies the word vector of each word by its weight and adds them together to get the character c i As a representation of the beginning of a word
[0113]
[0114] Among them, E d (w) represents the embedding vector of word w;
[0115] Step 5.4, loop through the same method in Step 5.3 to get character c i As a representation of the middle character of a word As a representation of the last character of a word and representation as an independent word
[0116] Step 5.5, change the character c i The four representations of are combined to obtain the global representation of each character d g Indicates the size of the global representation dimension of the character;
[0117]
[0118] Step 6. After obtaining the local and global representations of the characters, the interactive gating mechanism is used to filter the features of the two to obtain a comprehensive representation. The comprehensive representation is then sent to the Bi-LSTM for context encoding, and then CRF is used as a decoding layer to obtain the label of the output result.
[0119] As a further solution of the present invention, in Step 6, since local representation may have information redundancy compared with global representation, an interactive gating mechanism is used to screen the features of both to obtain a reasonable comprehensive representation. Since the context information in the sentence is helpful for sequence modeling, a bidirectional LSTM network that can capture bidirectional information of the text is used to extract sentence context features. The conditional random field CRF is used as a decoding layer, and the vector encoded by the context encoder is sent to the CRF to find the label sequence with the highest probability by minimizing the negative maximum likelihood function.
[0120] The Step 6 includes the following steps:
[0121] Character-based NER is a continuous tagging task, and there are strong constraints between adjacent characters. Therefore, the contextual information of the characters in the sentence sequence should also be considered. The sentence sequence is fed into the Bi-LSTM network to extract the sentence sequence representation of the characters. The formula is as follows:
[0122]
[0123]
[0124]
[0125] In the sequence label output stage, CRF is used as a decoder. CRF affects the result of the current label based on the result of the previous label. Specifically, CRF consists of an emission matrix and a transfer matrix. The emission matrix Record the probability of each label, M i,j Represents the probability of the i-th word emitting (predicting) the j-th entity tag. And a transformation matrix T∈R tags×tags , T i,j It represents the probability of the jth label being transferred to the ith label. It is used to simulate the relationship between adjacent labels to be learned in the CRF layer. It is a learnable parameter matrix that can help to explicitly model the transfer relationship between labels and improve the accuracy of named entity recognition. n is the number of characters in the sentence, and tags is the number of entity labels. The characters are encoded by BiLSTM to obtain the latent vector h i , use H to represent the latent vector matrix of the input sequence, and then send it to CRF to find the label sequence with the highest probability by minimizing the negative maximum likelihood function. The formula is as follows:
[0126] M=σ(W t H+b t ) (19)
[0127]
[0128]
[0129] Among them, φ(S,y) is the sum of the emission probability between the observation sequence and the label sequence and the label sequence transfer score, S represents the observation sequence, and y is the true label. and b t ∈R n×tags is the parameter of the linear layer, and Y represents the set of valid label sequences.
[0130] Use the negative log-likelihood function to calculate the loss value for label classification:
[0131] Loss cls = -logp(y|S) (22)
[0132] y is the true sequence label;
[0133] Finally, add the label classification loss and the fusion loss to get the final loss value of the model.
[0134]
[0135] In order to illustrate the effect of the present application, the present invention compares the effects of the traditional NER model Bi-LSTM, the CNER model based on word embedding (SoftLexicon, LGN and FLAT), etc. The model proposed in the present invention can more accurately judge the type and boundary of the entity when performing entity recognition. This is due to the fact that the model of the present invention utilizes the structure of the glyphs and the vector of the sequence to expand the rich information in the dimension of the vector space, so that entities of similar types can be predicted more accurately. The experimental results are shown in Table 2, where Lattice+Glyce is the experimental result of adding glyph structure information to the embedding layer of the Lattice model.
[0136] Table 2 Effects of each model on CMeEE and IMCS21 datasets
[0137]
[0138] It can be observed that: 1. The model of the present invention has achieved the best performance among all models. Compared with MECT, which has the best performance in the base model, the F1 value of the model of the present invention has increased by 1.04% in the CMeEE dataset and by 0.62% in IMCS21. 2. The model of the present invention is better than the above-mentioned models as a whole. Some models have added glyph information on the basis of word information. MECT has incorporated radical information, and Lattice+Glyce has incorporated glyphs. The model of the present invention has both. The latter are models that have integrated word information in different ways, which shows that external glyph information is helpful for understanding Chinese semantics. 3. On the CMeEE dataset, FLAT has the highest recall rate, indicating that it has a strong ability to extract entities in long sentences, but its precision is very low, resulting in an overall performance that is not as good as the model of the present invention. The model of the present invention has achieved the best F1 value on both the CMeEE dataset with more long sentences and the IMCS21 dataset with more short sentences, proving that the model of the present invention has strong robustness.
[0139] In order to prove the effectiveness of the glyph information of the model of the present invention, an ablation experiment was conducted. Among them, the experiment with w / oglobal vector removes the global representation of the characters in the model of the present invention, that is, the model only uses the character representation enhanced by the glyph information. w / o glyph vector only uses character embedding and global representation, and uses a gating mechanism to filter information, and w / o glyph structure vector removes the glyph structure representation when performing local feature fusion. w / o radical sequence vector removes the glyph sequence representation when performing local feature fusion. Experiments were conducted on the CMeEE dataset, and the experimental results are shown in Table 3. It can be seen from the results of all datasets that the use of glyph image information can effectively improve the performance of the model, and is stronger than the improvement effect of using the glyph structure information. After fusing these two glyph features, the improvement effect of the model is most obvious, which proves that using glyph information to enhance the representation of Chinese characters can have a good improvement on the model's entity extraction performance. The present invention further explores the impact of the size of the self-encoding latent vector dimension on the model. The size of the latent vector dimension in the model is set to 50 to 250, and experiments are conducted on the CMeEE dataset. The results are as follows Figure 4 As shown in the figure, it can be found that the performance of the model is better when the dimension is about 200. If the dimension of the latent vector is too low and the representation ability is insufficient, the model performance will deteriorate significantly.
[0140] Table 3 Results of ablation experiments on the CMeEE dataset
[0141]
[0142] In order to prove the effectiveness of the model proposed in the present invention, the number of errors in entity recognition by each model is counted. Table 4 shows the number of entity recognition errors of different models on two data sets, including entity head boundary errors (BE), entity tail boundary errors (EE) and entity type errors (TE). Compared with SoftLexicon's model on CMeEE, our model has reduced the number of entity head boundary errors and entity tail boundary errors by 377 and 394 respectively, and the entity type errors by 68. From the results, it can be seen that the model of the present invention has a significant effect on improving the boundary recognition of entities. There is no doubt that the model of the present invention is very beneficial for the recognition of entity boundaries and entity types.
[0143] Table 4 Statistics of entity recognition error types
[0144]
[0145] In order to prove the effectiveness of the local feature and global feature fusion method proposed in the present invention, the present invention also conducted experiments on three other fusion methods on the CMeEE data set. The fusion method of Filter_1 is to directly add the local and global representations and then send them to the Bi-LSTM encoding. The fusion method of Filter_2 is to directly splice the local and global representations and then send them to the Bi-LSTM encoding. The fusion method of Filter_3 is to use a gating mechanism to process the local and global representations respectively, then add the processed vectors, and then send them to the Bi-LSTM encoding. The experimental results are shown in Table 5. It can be seen that the effect of Filter_1 is not as good as Filter_2, which may be because the latter method can fully preserve local and global information. Filter_3 adds gating and then adds, and the result is worse than the first two. This may be due to the fact that the gating mechanism can well filter out the important parts related to the local and global information and enhance the fitting ability of the model. The model of the present invention uses a gating mechanism to process the local and global representations and then splice the two, so that the local and global information can be fully preserved and the important information of the two can be filtered out, thereby achieving the best model performance.
[0146] Table 5 Ablation experiments combining local and global representations
[0147]
[0148] In order to verify the effectiveness of the local representation self-encoding fusion of the present invention, experiments with two other local feature fusion methods were also carried out on the CMeEE data set. The method of Fusion_1 is to directly splice the character embedding, glyph structure embedding and glyph sequence embedding. The method of Fusion_2 is to embed the character, glyph structure embedding and glyph sequence embedding and then add them after linear transformation. The experimental results are shown in Table 6. It can be seen that the self-encoding fusion method of the present invention has the best effect, with F1 values 0.51 and 1.67 higher than the other two fusion methods. It should be noted that Fusion_1 has the highest recall rate, which may be due to the fact that direct splicing can more comprehensively utilize the three local vectors to identify entities. But on the other hand, the three vectors are located in different vector spaces and have large differences. Direct splicing will introduce redundant information, making its accuracy the lowest. In contrast, the self-encoding method can better fuse the three vectors, thereby taking into account the accuracy and recall rate of entity recognition.
[0149] Table 6 Ablation results of local representation fusion method
[0150]
[0151] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.< / pad> < / pad>
Claims
1. Chinese named entity recognition method based on local and global character representation enhancement, Features: The method comprises the following steps: Step 1: Use the character vector trained on the corpus as the initial embedding of the character: map each character to a dense vector representation to obtain the character embedding of each sentence; Step 2: Split the character into radicals and other character components, and then use the sequence feature encoder to extract the character's glyph sequence features; Step 3: Treat a single character as a two-dimensional image and obtain the character's glyph structure features through an image feature encoder; The image corresponding to the Chinese character passes through multiple convolutional layers to capture low-level graphic features, and then uses adaptive pooling operations and applies group convolution to map to the final glyph structure features; Step 4: Using the self-encoding mechanism, the three vectors of character shape structure features, character shape sequence features and pre-trained character embedding are integrated to obtain the local representation of the character; Step 5: First, use the word2vec Skip-Gram model on the domain corpus to train a domain dictionary. Then, query and match each character in the dictionary to obtain several word sets, and then obtain the global representation of the character by weighted distribution and concatenation. Step 6: After obtaining the local and global representations of the characters, the interactive gating mechanism is used to filter the information of the features of the two to obtain a comprehensive representation; The comprehensive representation is then fed into Bi-LSTM for context encoding, and then CRF is used as a decoding layer to obtain the label of the output result; In the Step 4, the glyph structure features, glyph sequence features and pre-trained character embedding of the character are first concatenated, and then an automatically fused latent vector is obtained through the transformation layer, and then an attempt is made to reconstruct the originally concatenated vector from the self-fused latent vector. Finally, the Euclidean distance between the original vector and the reconstructed vector is calculated, and the loss is calculated using the mean square error to obtain information compressed by the intermediate layer but without loss. The Step 5 includes the following steps: Step 5.1, character c i Perform query matching in a dictionary D pre-trained using the Skip-Gram model; if a word w in D contains character c i , then according to the different positions of the characters in the word, they are included in four word sets B(c i ),M(c i ),E(c i ),S(c i ); Specifically, if c i appears at the beginning of a word w, the word w is classified into the word set B (c i ); if c i Appears in the middle position of a word w, the word w is classified into the word set M (c i ); if c i If it appears at the end of a word w, the word w is classified into the word set E(c i ); if c i Same as a word w, that is, the character is an independent word, then the word w is classified into the word set S (c i ); Step 5.
2. Count the characters c i The number of times a word w appears in the training data, and the character c i The total number of times all the matched words appear in the training set data is M, then the character c i The frequency of a word w that matches for: Step 5.3, match the word set B (c i ) multiplies the word vector of each word by its weight and adds them together to get the character c i As a representation of the beginning of a word Among them, E d (w) represents the embedding vector of word w; Step 5.4, loop through the same method in Step 5.3 to get character c i As a representation of the middle character of a word As a representation of the last character of a word and representation as an independent word Step 5.5, change the character c i The four representations of are combined to obtain the global representation of each character d g Indicates the size of the global representation dimension of the character; 2. The Chinese named entity recognition method based on local and global character representation enhancement according to claim 1, Features: In Step 1, the input sentence is regarded as a character sequence s={c 1 , c 2 ,···,c n }, and then each character c i are mapped to a dense vector representation Get the character embeddings for each sentence: where e c Represents a character embedding lookup table.
3. The Chinese named entity recognition method based on local and global character representation enhancement according to claim 1, Features: In the Step 2, a word-splitting dictionary is first used to construct a component query table for each word in the data set; then the character-splitting sequence is sent to a convolutional neural network (CNN) to extract the character shape sequence features, and then a residual network is used to optimize the convolutional layer to alleviate the gradient vanishing problem as the depth of the neural network increases. Finally, the maximum pool and fully connected layer are used to obtain the shape sequence feature embedding.
4. The Chinese named entity recognition method based on local and global character representation enhancement according to claim 1, Features: The Step 2 includes the following steps: Step 2.1, replace the i-th character c i Split into K parts, If the length of a character component is less than K, the empty position is replaced by " <pad>" to fill, and then perform a random embedding operation E on each character component r :< / pad> Step 2.2: Randomly embed the obtained characters into a sequence Send it to the convolution operation conv3 with a convolution kernel size of 3 to get the character hidden vector sequence Step 2.3, perform max-pooling on the vector corresponding to each character component in the character latent vector sequence, and then send it to a fully connected layer f c Perform dimension transformation to get the glyph sequence embedding of the character The glyph sequence embedding dimension of this character is d o ; 5. The Chinese named entity recognition method based on local and global character representation enhancement according to claim 1, Features: In the Step 3, the glyph structure features can obtain rich pictographic information from the character image to improve the performance of the Chinese named entity recognition model; for images of different font types, they are spliced together to represent the structural image of the character, and the character image is passed through multiple convolutional layers and multiple output channels to capture low-level glyph structure features.
6. The Chinese named entity recognition method based on local and global character representation enhancement according to claim 1, Features: The Step 3 includes the following steps: Step 3.1, c i The characters are converted into grayscale images corresponding to 6 different fonts in is an 8-bit grayscale image of the jth font with a size of 12×12. Different image matrices are concatenated to obtain character c i Structural image of Where concat represents the concatenation operation; Step 3.2, then, use the convolution operation conv1 with a convolution kernel size of 5×5 and 384 output channels to capture low-level graphic features and obtain the hidden layer vector Step 3.3, use the maxpooling operation with a template size of 4×4, The resolution is reduced from 8×8 to 2×2; then a convolution kernel with a size of 1×1 and d s The convolution operation conv2 of the output channels is used to obtain the hidden layer vector Step 3.4, finally, The convolution kernel size is 2, and the group convolution operation groupconv is performed, and the dimension transformation operation reshape is performed to obtain the glyph structure representation of the character. The glyph structure embedding dimension of this character is d s ; Reshape represents a dimensional transformation that converts a 2D vector into a 1D vector.
7. The Chinese named entity recognition method based on local and global character representation enhancement according to claim 1, Features: In the Step 4, an autoencoder network is used to fuse three vectors, namely, the glyph structure features, the glyph sequence features and the pre-trained character embedding, and the model is encouraged to extract multi-granularity features by maximizing the correlation between inputs of different granularities. Specifically, the Step 4 includes the following steps: Step 4.1, firstly, the glyph structure features Character sequence features and character embedding Perform splicing to obtain the initial splicing vector Among them, d s Represents the character’s glyph structure embedding dimension size, d o The glyph sequence embedding dimension size of the character, d c Represents the pre-trained embedding dimension size of the character; Step 4.2, then Perform the following two linear transformations and activations to obtain the latent vector Step 4.3, use Reconstruct the original concatenated vector and obtain the reconstructed vector Step 4.4, use the mean square error loss function to calculate and Loss f : Step 4.5: Add the loss to the main model sequence labeling model, stimulate the above reconstruction process through the NER downstream task, obtain the information compressed by the middle layer but without loss, and convert the latent vector of the middle layer into As a local representation of fusion.
8. The Chinese named entity recognition method based on local and global character representation enhancement according to claim 1, Features: In the Step 6, since the local representation has information redundancy compared to the global representation, the interactive gating mechanism is used to screen the features of the two to obtain a reasonable comprehensive representation. Since the context information in the sentence is helpful for sequence modeling, a bidirectional LSTM network that can capture bidirectional information of the text is used to extract the sentence context features. The conditional random field CRF is used as the decoding layer, and the vector encoded by the context encoder is sent to the CRF to find the label sequence with the highest probability by minimizing the negative maximum likelihood function.