A text named entity recognition model and method

By introducing global context enhancement module and word pair relative position coding attention module in the NER model, the shortcomings of the existing model in capturing global context and relative position relationships are solved, and more accurate recognition of named entities is achieved, especially in complex sentences.

CN119204014BActive Publication Date: 2025-05-30CHENGDU UNIV OF INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411252989.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-05-30
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

The existing NER model has limitations when capturing the global context information of the text, and ignores the relative positional relationship between the named entity and other words in its context, resulting in the inability to accurately judge the boundaries of the named entity in complex sentences, and even recognition errors occur.

Method used

A text named entity recognition model is proposed, including an embedding layer, a multi-layer perceptron, a global context enhancement module, a word pair relative position encoding attention module, a layered feature fusion module and a prediction layer. The model uses the global context enhancement module and the word pair relative position encoding attention module to capture the global context information of the text and the word pair relative position relationship, and generates a score matrix to output the named entity prediction results.

Benefits of technology

Through the global context enhancement and the integration of relative position information, the model's accurate recognition ability of named entities is improved, especially when facing complex sentences, the boundaries of named entities can be judged more accurately, and the accuracy of text named entity recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204014B_ABST
    Figure CN119204014B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of named entity recognition, and discloses a text named entity recognition model and method. The model includes an embedding layer, a multi-layer perceptron, a global context enhancement module, a word pair relative position encoding attention module, a hierarchical feature fusion module, and a prediction layer; the embedding layer is used to generate a text embedding representation according to text data; the multi-layer perceptron is used to convert the text embedding representation into a two-dimensional word pair matrix; the global context enhancement module is used to generate a global context enhancement representation according to the two-dimensional word pair matrix; the word pair relative position encoding attention module is used to generate a word pair relative position encoding attention representation according to the two-dimensional word pair matrix; the hierarchical feature fusion module is used to generate a score matrix according to the global context enhancement representation, the word pair relative position encoding attention representation, and the two-dimensional word pair matrix; the prediction layer is used to output a named entity prediction result according to the score matrix. The present invention realizes the improvement of the accuracy of text named entity recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of named entity recognition, and particularly to a text named entity recognition model and method. Background Art

[0002] Named entity recognition (NER) is a key task in the field of natural language processing, aiming to automatically identify and classify named entities in text, such as person names, place names, organization names, etc. Traditional NER methods mainly rely on manual feature engineering, rule definition, or machine learning methods based on statistical models. However, with the development of deep learning technology, neural network-based models, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformer models based on attention mechanisms, have become the mainstream methods for NER tasks. Although recent deep learning models have made significant progress in NER tasks, there are still some significant deficiencies.

[0003] Firstly, many existing NER models have limitations in capturing the global context information of text. Generally, the models mainly focus on local context, that is, a small number of words around the named entity, but ignore the role of the global semantic background. For example, in long texts or complex semantic environments, simply relying on local context information may not be able to fully capture the true meaning of the named entity.

[0004] Secondly, in the prior art, the models often ignore the relative position relationship between the named entity and other words in its context. Relative position information is crucial for accurately identifying named entities because the boundaries and categories of named entities may be affected by the relative positions of their adjacent words. For example, in the sentence "John met Mary at the park", the relative positions of "John" and "Mary" can provide valuable clues for the model to help distinguish whether these entities exist independently or are related to each other. However, most existing NER models cannot effectively incorporate this relative position information into the entity recognition process. This limitation may cause the model to be unable to accurately judge the boundaries of named entities when facing complex sentences, and even result in recognition errors. Summary of the Invention

[0005] In view of the above deficiencies in the prior art, the present invention provides a text named entity recognition model and method.

[0006] To achieve the above invention object, the technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present application provides a text named entity recognition model, including an embedding layer, a multi-layer perceptron, a global context enhancement module, a word pair relative position encoding attention module, a hierarchical feature fusion module, and a prediction layer;

[0008] The embedding layer is used to generate a text embedding representation according to the text data;

[0009] The multi-layer perceptron is used to convert the text embedding representation into a two-dimensional word pair matrix;

[0010] The global context enhancement module is used to generate a global context enhancement representation according to the two-dimensional word pair matrix;

[0011] The word pair relative position encoding attention module is used to generate a word pair relative position encoding attention representation according to the two-dimensional word pair matrix;

[0012] The hierarchical feature fusion module is used to generate a score matrix according to the global context enhancement representation, the word pair relative position encoding attention representation, and the two-dimensional word pair matrix;

[0013] The prediction layer is used to output a named entity prediction result according to the score matrix.

[0014] In a possible implementation manner of the first aspect, the embedding layer includes a BERT layer, a Word2Vec layer, a bidirectional long short-term memory network layer, and a splicing layer;

[0015] The BERT layer is used to generate a BERT embedding representation of each word according to the text data;

[0016] The Word2Vec layer is used to generate a word embedding representation of each word according to the text data;

[0017] The bidirectional long short-term memory network layer is used to generate a character embedding representation of each word according to the text data;

[0018] The splicing layer is used to sum up the BERT embedding representation of each word, the word embedding representation of each word, and the character embedding representation of each word to form a text embedding representation.

[0019] In a possible implementation manner of the first aspect, the global context enhancement module includes a mask matrix, a sparse convolution layer, and a normalization layer;

[0020] The mask matrix is used to control the calculation area of the sparse convolution layer;

[0021] The sparse convolution layer is used to capture context at different positions of the two-dimensional word pair matrix according to the mask matrix, and obtain the context information between word pairs;

[0022] The normalization layer is used to globally enhance the context information between word pairs with the two-dimensional word pair matrix as a condition, and obtain a global context enhancement representation.

[0023] In a possible implementation of the first aspect, the mask matrix is specifically:

[0024]

[0025] where H is the mask matrix, σ is the sigmoid function, S is the mask parameter matrix, g 1 ,g 2 are both random Gumbel noises, and τ is the temperature parameter.

[0026] In a possible implementation of the first aspect, the sparse convolutional layer is specifically:

[0027] L ij = W sc · M ij · H ij

[0028] where L ij is the context information at the i-th row and j-th column in the context information between word pairs, W sc is the convolutional kernel of the sparse convolution, M ij is the position at the i-th row and j-th column in the two-dimensional word pair matrix, and H ij is the mask element at the i-th row and j-th column in the mask matrix.

[0029] In a possible implementation of the first aspect, the normalization layer is specifically:

[0030]

[0031] where C ij is the global context enhancement information at the i-th row and j-th column in the global context enhancement representation, L ij is the context information at the i-th row and j-th column in the context information between word pairs, mean[M] is the mean of the two-dimensional word pair matrix M, std[M] is the standard deviation of the two-dimensional word pair matrix M, and W and b are both training parameters.

[0032] In a possible implementation of the first aspect, the word pair relative position encoding attention module includes an embedding table and an attention mechanism layer;

[0033] The embedding table is used to save the word pair relative position encoding assigned to the relative position difference between each word pair;

[0034] The attention mechanism layer is used to calculate the relative position difference between word pairs according to the two-dimensional word pair matrix, query the corresponding word pair relative position encoding from the embedding table according to the relative position difference between word pairs, and generate a word pair relative position encoding attention representation according to the word pair relative position encoding.

[0035] In a possible implementation manner of the first aspect, the attention mechanism layer is specifically:

[0036]

[0037] A ij = α ij (W l[i,j] M ij )

[0038] where A ij is the word pair relative position attention information at the i-th row and j-th column in the word pair relative position encoding attention representation, α ij is the attention weight at the i-th row and j-th column in the two-dimensional word pair matrix, W l[i,j] , W a , W h[i,j] are all weight training parameters, M ij is the position at the i-th row and j-th column in the two-dimensional word pair matrix, n is the total number of all candidate entities, k is the traversal serial number for controlling candidate entities, v is the training weight vector, T is the transpose, is the relative position difference between the word pairs at the i-th row and j-th column in the two-dimensional word pair matrix.

[0039] In a possible implementation manner of the first aspect, the hierarchical feature fusion module is specifically:

[0040]

[0041] where S is the score matrix, M is the two-dimensional word pair matrix, C is the global context enhancement representation, σ is the sigmoid function, V_Pool is the vertical pooling operation, H_Pool is the horizontal pooling operation, ⊙ is the Hadamard product, is the concatenation operation, is the outer product operation.

[0042] In a second aspect, the present application provides a text named entity recognition method, which is applied to a text named entity recognition model as described in any item of the first aspect, and includes the following steps:

[0043] Obtain the text data to be recognized;

[0044] Input the text data to be recognized into the text named entity recognition model to obtain the text named entity recognition result.

[0045] The present invention has the following beneficial effects:

[0046] In the embedding layer of the present invention, the text is converted into an embedded one-dimensional vector representation, and this vector is constructed into a two-dimensional word pair matrix, which contains the global word pair information and the relative positions of word pairs in the text. In the global context enhancement module, a sparse convolution and a global context enhancement normalization layer are set. Compared with traditional CNN, sparse convolution only calculates the positions of non-zero weights to improve computational efficiency, saving memory and computational resources. The normalization operation uses the two-dimensional word pair matrix as the conditional input to normalize the context information captured by the sparse convolution to achieve global enhancement of the context. In the word pair relative position attention module, the relative positions of word pairs are introduced into the attention mechanism, so that when calculating the attention weights, the relative position embeddings of each word pair can be combined, and the relative positions of word pairs are used to focus on the entity positions in the text. Thus, the global context understanding ability of the model is improved through global context enhancement, and the relative positions of word pairs in the text are used to help the model focus on the entity positions, realizing the improvement of the accuracy of text named entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic structural diagram of a text named entity recognition model;

[0048] Figure 2 It is a schematic structural diagram of the global context enhancement module;

[0049] Figure 3 It is a schematic structural diagram of the word pair relative position encoding attention module;

[0050] Figure 4 It is a schematic flow diagram of a text named entity recognition method. DETAILED DESCRIPTION OF THE INVENTION

[0051] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

[0052] As Figure 1 shown, an embodiment of the present application provides a text named entity recognition model, including an embedding layer, a multi-layer perceptron, a global context enhancement module, a word pair relative position encoding attention module, a hierarchical feature fusion module, and a prediction layer;

[0053] The embedding layer is used to generate a text embedding representation h according to text data;

[0054] The multi-layer perceptron is used to convert the text embedding representation h into a two-dimensional word pair matrix M;

[0055] The global context enhancement module is used to generate a global context enhanced representation C according to the two-dimensional word pair matrix M;

[0056] The word pair relative position encoding attention module is used to generate a word pair relative position encoding attention representation A according to the two-dimensional word pair matrix M;

[0057] The hierarchical feature fusion module is used to generate a score matrix S according to the global context enhanced representation C, the word pair relative position encoding attention representation A, and the two-dimensional word pair matrix M;

[0058] The prediction layer is used to output a named entity prediction result according to the score matrix S.

[0059] In an alternative embodiment of the present invention, the embedding layer includes a BERT layer, a Word2Vec layer, a bidirectional long short-term memory network layer, and a splicing layer;

[0060] The BERT layer is used to generate a BERT embedding representation of each word according to the text data;

[0061] The Word2Vec layer is used to generate a word embedding representation of each word according to the text data;

[0062] The bidirectional long short-term memory network layer is used to generate a character embedding representation of each word according to the text data;

[0063] The splicing layer is used to sum up the BERT embedding representation of each word, the word embedding representation of each word, and the character embedding representation of each word to form a text embedding representation.

[0064] This embodiment uses BERT to obtain the BERT embedding representation of each word Use pre-trained word vectors (Word2Vec) to generate fixed word embeddings for each word Use a bidirectional long short-term memory network (BiLSTM) to generate character embeddings for each word Then, the above-generated BERT embedding representation Word embedding representation And character embedding representation Are summed up to form a text embedding representation, specifically:

[0065]

[0066] Wherein, Represents the concatenation operation.

[0067] In an alternative embodiment of the present invention, the multi-layer perceptron uses a set of MLPs to take the text embedding representation h obtained by the embedding layer as input and generate a projection representation h mlp . Subsequently, hmlp Perform an outer product with itself to obtain the two-dimensional word pair matrix M of the text, specifically:

[0068] h mlp = MLP(h)

[0069]

[0070] where represents the outer product operation.

[0071] In an alternative embodiment of the present invention, as Figure 2 shown, the global context enhancement module includes a mask matrix, a sparse convolutional layer, and a normalization layer;

[0072] The mask matrix is used to control the calculation area of the sparse convolutional layer;

[0073] The sparse convolutional layer is used to capture context at different positions of the two-dimensional word pair matrix according to the mask matrix, and obtain the context information between word pairs;

[0074] The normalization layer is used to globally enhance the context information between word pairs with the two-dimensional word pair matrix as a condition, and obtain the global context enhancement representation.

[0075] This embodiment uses the mask matrix H ∈ {0, 1} B×1×n×n to control the calculation area of the sparse convolution, specifically:

[0076]

[0077] where H is the mask matrix, σ is the sigmoid function, S is a learnable mask parameter matrix, g 1 , g 2 are both random Gumbel noises, and τ is the corresponding temperature parameter in Gumbel-Softmax.

[0078] This embodiment uses sparse convolution to capture context at different positions of the two-dimensional word pair matrix M obtained by the multi-layer perceptron, and obtains the context information L between word pairs, specifically:

[0079] L ij = W sc · M ij · H ij

[0080] where L ij is the context information of the i-th row and j-th column in the context information between word pairs, W sc is the convolution kernel of the sparse convolution, M ijis the position at the i-th row and j-th column in the two-dimensional word pair matrix, H ij is the masked element at the i-th row and j-th column in the mask matrix.

[0081] In this embodiment, the two-dimensional word pair matrix M obtained by the multi-layer perceptron is used as the conditional input in the normalization operation to globally enhance the context information L, so as to obtain the globally context-enhanced representation C. Specifically:

[0082]

[0083] where C ij is the globally context-enhanced information at the i-th row and j-th column in the globally context-enhanced representation, L ij is the context information at the i-th row and j-th column in the context information between word pairs, mean[M] is the mean of the two-dimensional word pair matrix M, std[M] is the standard deviation of the two-dimensional word pair matrix M, and W and b are both training parameters.

[0084] In an alternative embodiment of the present invention, as Figure 3 shown, the word pair relative position encoding attention module includes an embedding table and an attention mechanism layer;

[0085] The embedding table is used to store the word pair relative position encoding assigned to the relative position difference between each word pair;

[0086] The attention mechanism layer is used to calculate the relative position difference between word pairs according to the two-dimensional word pair matrix, query the corresponding word pair relative position encoding from the embedding table according to the relative position difference between word pairs, and generate a word pair relative position encoding attention representation according to the word pair relative position encoding.

[0087] In this embodiment, a position embedding vector (which is randomly initialized by a normal distribution) is assigned to each relative position difference, and the vector is stored in an embedding table P of size 2m + 1. Here, m is a predefined maximum relative position distance, and relative positions outside the range [-m, +m] are mapped to the same position encoding.

[0088] In this embodiment, for relative position encoding, the relative position difference Δ of each word pair in the input text is calculated through the relative position information between word pairs provided by the word pair matrix; then, in the embedding table P, for each position difference Δ ij , the corresponding encoding vector is looked up Finally, the word pair relative position encoding of each word pair in the obtained two-dimensional word pair matrix is introduced into the attention mechanism to obtain the word pair relative position encoding attention representation A. Specifically:

[0089]

[0090] A ij= α ij (W l[i,j] M ij )

[0091] where A ij is the relative position attention information of the word pair at the i-th row and j-th column in the relative position encoding attention representation of the word pair, α ij is the attention weight at the i-th row and j-th column in the two-dimensional word pair matrix, W l[i,j] , W a , W h[i,j] are all weight training parameters, M ij is the position at the i-th row and j-th column in the two-dimensional word pair matrix, n is the total number of all candidate entities, k is the traversal serial number for controlling the candidate entities, v is a trainable weight vector, T is the transpose, is the relative position difference between the word pairs at the i-th row and j-th column in the two-dimensional word pair matrix.

[0092] In an alternative embodiment of the present invention, this embodiment uses a hierarchical feature fusion module to integrate and fuse the two-dimensional word pair matrix M obtained by the multi-layer perceptron, the global context enhanced representation C obtained by the global context enhancement module, and the relative position attention representation A of the word pair obtained by the relative position encoding attention module of the word pair to obtain a score matrix S, specifically:

[0093]

[0094] where S is the score matrix, M is the two-dimensional word pair matrix, C is the global context enhanced representation, Sig is the sigmoid function, V_Pool is the vertical pooling operation, H_Pool is the horizontal pooling operation, ⊙ is the Hadamard product, is the concatenation operation, is the outer product operation. The score matrix S contains the scores of each word pair.

[0095] In an alternative embodiment of the present invention, this embodiment uses a prediction layer to calculate the probability of each word pair in the score matrix S obtained by the hierarchical feature fusion module as the input of the softmax function, specifically:

[0096]

[0097] where is the probability that the word pair (i, j) belongs to the k-th type.

[0098] As Figure 4 shown, this embodiment of the present application also provides a method for text named entity recognition, including the following steps S1 and S2:

[0099] S1. Obtain the text data to be recognized;

[0100] S2. Input the text data to be recognized into the text named entity recognition model to obtain the text named entity recognition result.

[0101] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0104] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0105] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to these technical revelations disclosed by the present invention, and these deformations and combinations are still within the scope of protection of the present invention.

Claims

1. A text named entity recognition model, characterized in that: It includes embedding layer, multi-layer perceptron, global context enhancement module, word pair relative position encoding attention module, hierarchical feature fusion module and prediction layer; The embedding layer is used to generate a text embedding representation based on the text data; The multi-layer perceptron is used to convert the text embedding representation into a two-dimensional word pair matrix; The global context enhancement module is used to generate a global context enhancement representation according to the two-dimensional word pair matrix; The global context enhancement module includes a mask matrix, a sparse convolutional layer and a normalization layer; The mask matrix is ​​used to control the calculation area of ​​the sparse convolutional layer; The sparse convolution layer is used to capture context at different positions of the two-dimensional word pair matrix according to the mask matrix to obtain context information between word pairs; The normalization layer is used to globally enhance the context information between word pairs using the two-dimensional word pair matrix as a condition to obtain a global context enhanced representation; The word pair relative position encoding attention module is used to generate a word pair relative position encoding attention representation according to a two-dimensional word pair matrix; The hierarchical feature fusion module is used to generate a score matrix based on the global context enhancement representation, the word pair relative position encoding attention representation and the two-dimensional word pair matrix; The prediction layer is used to output the named entity prediction result according to the score matrix.

2. A text named entity recognition model according to claim 1, characterized in that: The embedding layer includes a BERT layer, a Word2Vec layer, a bidirectional long short-term memory network layer, and a concatenation layer; The BERT layer is used to generate a BERT embedding representation of each word based on the text data; The Word2Vec layer is used to generate a word embedding representation of each word based on text data; The bidirectional long short-term memory network layer is used to generate a character embedding representation of each word according to the text data; The concatenation layer is used to add the BERT embedding representation of each word, the word embedding representation of each word, and the character embedding representation of each word to form a text embedding representation.

3. A text named entity recognition model according to claim 1, characterized in that: The mask matrix is ​​specifically: in, is the mask matrix, is the sigmoid function, is the mask parameter matrix, are all random Gumbel noise, is the temperature parameter.

4. A text named entity recognition model according to claim 1, characterized in that: The sparse convolutional layer is specifically: in, is the context information between word pairs. i Row, No. j Column context information, is the convolution kernel of sparse convolution, is the first i Row, No. j The position of the column, is the first i Row, No. j Mask element for the column.

5. A text named entity recognition model according to claim 1, characterized in that: The normalization layer is specifically: in, The global context enhancement representation i Row, No. j The global context enhancement information of the column, is the context information between word pairs. i Row, No. j Column context information, is the mean of the two-dimensional word pair matrix M, is the standard deviation of the two-dimensional word pair matrix M, These are training parameters.

6. A text named entity recognition model according to claim 1, characterized in that: The word pair relative position encoding attention module includes an embedding table and an attention mechanism layer; The embedding table is used to store the relative position codes of the word pairs assigned by the relative position differences between each word pair; The attention mechanism layer is used to calculate the relative position difference between word pairs according to the two-dimensional word pair matrix, and query the corresponding word pair relative position encoding from the embedding table according to the relative position difference between the word pairs, and generate the word pair relative position encoding attention representation according to the word pair relative position encoding.

7. A text named entity recognition model according to claim 6, characterized in that: The attention mechanism layer is specifically: in, The first word in the attention representation encoding the relative position of the word pair i Row, No. j The relative position attention information of the word pairs in the column, is the first i Row, No. j The attention weights of the columns, are weight training parameters, is the first i Row, No. j The position of the column, n is the total number of all candidate entities, k is the traversal number of the control candidate entity, v is the training weight vector, T is the transpose, is the two-dimensional word pair matrix i Row, No. j The relative position difference between word pairs in the column.

8. A text named entity recognition model according to claim 1, characterized in that: The hierarchical feature fusion module is specifically: in, is the score matrix, is a two-dimensional word pair matrix, Enhance representation for global context, is the sigmoid function, For vertical pooling operation, is the horizontal pooling operation, For Hadamard, For serial operation, It is the outer product operation.

9. A method for identifying named entities in text, characterized in that: A text named entity recognition model applied to any one of claims 1 to 7, comprising the following steps: Obtain text data to be recognized; Input the text data to be recognized into the text named entity recognition model to obtain the text named entity recognition results.

Citation Information

Patent Citations

  • Named entity identification method and device, equipment, medium and product

    CN114298049A

  • Named entity recognition method and device based on hybrid lattice self-attention network

    CN114429132A