Named entity recognition method based on two-dimensional semantic expansion

Through the named entity recognition method based on two-dimensional semantic expansion, two-dimensional sentence representation is generated using the coding module, semantic focus module and gated integration module, which solves the problem of poor recognition performance of nested entities and long named entities in the prior art, and achieves a more refined and overall feature extraction effect.

CN119990131AActive Publication Date: 2025-05-13GUIZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510069993.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

Existing named entity recognition methods perform poorly when dealing with nested entities and long named entities, have low feature utilization, and are difficult to make full use of manually labeled data.

Method used

Using a named entity recognition method based on two-dimensional semantic expansion, a model architecture consisting of a coding module, a semantic focus module and a gated integration module is constructed to generate two-dimensional sentence representations, and feature extraction is performed through semantic expansion and lightweight networks.

Benefits of technology

The performance of named entity recognition is improved, especially when dealing with nested entities and long named entities, the fineness and integrity of feature extraction are enhanced, and the problem of entity semantics is dispersed and semantic ambiguity in spans is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990131A_ABST
    Figure CN119990131A_ABST
Patent Text Reader

Abstract

The invention discloses a named entity recognition method based on two-dimensional semantic expansion, which belongs to the technical field of natural language processing, and comprises the following steps: S1, constructing a model architecture consisting of a coding module, a semantic focusing module and a gating integration module; s2, mapping an original sentence into a two-dimensional sentence expression through a coding module; s3, capturing local feature information and global semantic information represented by the two-dimensional sentence through a semantic focusing module; s4, fusing different feature information through a gating integration module to obtain final feature representation, and performing prediction; according to the named entity recognition method based on two-dimensional semantic expansion provided by the invention, aiming at the problems that entity semantics are dispersed in a span and are fuzzy, local semantic information is enhanced through a semantic expansion module, and finer feature extraction is realized; according to the method, a lightweight network architecture is constructed and used for capturing global information of two-dimensional representation, the quality of feature representation is further improved, and therefore the long named entities can be better recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a named entity recognition method based on two-dimensional semantic expansion. Background Art

[0002] Named Entity Recognition (NER) is a basic task in the field of natural language processing, which aims to identify named entities from unstructured text. This task is widely used to support other NLP tasks, such as relationship extraction, entity linking, knowledge graph construction, etc. The tasks of named entity recognition are mainly divided into three categories: flat named entity recognition, nested named entity recognition, and discontinuous named entity recognition. Flat named entity recognition only detects continuous entity spans and their semantic categories in the text. In the nested named entity recognition task, a word can belong to multiple different named entities. Discontinuous named entity recognition deals with entity recognition composed of multiple discontinuous fragments in the text.

[0003] Traditional sequence structures usually represent sentences as a sequential word-unit structure, but this structure is difficult to effectively solve the problem of nested entity recognition with complex structures. For example, "immunoglobulin enhancer" is a biomedical named entity, where "immunoglobulin" is a protein representing the enhancer target. The sequential word-unit structure can only assign one label to a word unit, which can easily lead to a chaotic tag representation of the word unit, thereby deteriorating the performance of named entity recognition.

[0004] Existing named entity recognition methods can be mainly divided into sequence-based methods, hypergraph-based methods and layer-based methods.

[0005] Sequence-based methods solve the named entity recognition task through various tagging schemes. These methods usually use neural network models (such as CNN and Transformer) for representation and then classify through a CRF layer. Hypergraph-based methods convert each sentence into a graph structure to account for the nested semantic structure. In layer-based methods, nested named entities are decomposed into different layers or entity types, and then, different layers or types are extracted through a separate sequence model.

[0006] A sentence can be converted into a two-dimensional representation. This two-dimensional representation method has the ability to expand the sentence semantics into a plane to solve the nested semantic structure. In the flattened sentence representation, all possible entity spans in the sentence can be organized into a matrix. Among them, each element in the matrix corresponds to a possible entity span in the sentence. This representation method helps to deal with nested entities and discontinuous entities. However, the above method still has the following shortcomings:

[0007] (1) Sequence-based methods have low feature utilization in named entity recognition tasks, and cannot fully utilize manually annotated data during training; models that convert original sentence structures into other non-sequence structures all show low performance. The main reason is that this method relies heavily on prior knowledge when converting nested named entities into non-sequence structures, and this conversion method may destroy the internal semantic structure of the sentence, resulting in a decrease in model performance.

[0008] (2) Hierarchical models also show low performance. The main reason is that nested named entities are decomposed into different levels or types and then identified by independent classifiers, which does not fully utilize the information of the annotated corpus.

[0009] (3) In the method of converting a sentence into a two-dimensional representation, adjacent elements in the semantic plane represent overlapping phrases in the sentence. They contain the same word and share the same context features. This dispersed semantics will cause the semantics of the true span to become blurred in the entire context. In addition, since spans are used to represent entities in the two-dimensional representation, when the entity is long, the span contains more words, resulting in more severe semantic weakening, thereby reducing the performance of long named entity recognition. Summary of the invention

[0010] The purpose of the present invention is to provide a named entity recognition method based on two-dimensional semantic expansion to solve the problem of two-dimensional sentence representation existing in the above-mentioned background technology.

[0011] To achieve the above object, the present invention provides a named entity recognition method based on two-dimensional semantic expansion, comprising the following steps:

[0012] S1. Construct a model architecture consisting of an encoding module, a semantic focus module, and a gated integration module.

[0013] S2, mapping an original sentence into a two-dimensional sentence representation through the encoding module;

[0014] S3, capturing the local feature information and global semantic information of the two-dimensional sentence representation through the semantic focus module;

[0015] S4. Different feature information is integrated through the gated integration module to obtain the final feature representation and make predictions.

[0016] Preferably, step S2 specifically comprises:

[0017] Assume that the sentence input to the model is X = {x 1 ,x 2 ,…,x N}, where x iis the i-th word in the sentence, and N is the length of the sentence X. In order to better encode the semantic information and contextual features of the word in the sentence, the pre-trained language model BERT is used to obtain the vector representation H of the sentence, which is defined as:

[0018] H=BERT(X) (1)

[0019] in, d h It is a dimension represented by a token;

[0020] The multi-head biaffine method is used to generate a two-dimensional representation of the vector H, represented by Γ, where the elements in Γ are recorded as Γ i,j , Γ i,j is a span representation of the input sentence X, and the entity boundary is encoded. After encoding the entity boundary, the two-dimensional representation Γ of the sentence is obtained by biaffine the two boundary representations of the potential named entity; two multi-layer perceptrons are designed, and the vectorized sentence representation H is input into the two multi-layer perceptrons respectively, and the start and end boundary representations of the same word but with different semantics are learned; the feature vectors of the encoded start and end boundaries are integrated into the two-dimensional sentence representation through the multi-head biaffine method, which is defined as:

[0021]

[0022] in, h is the hidden layer size, MHBiaffine(·,·) is a multi-head biaffine network, r is the feature size, and the element M in the matrix M i,j A vector representation representing a span of the input sentence X.

[0023] Preferably, the semantic focusing module in step S3 is composed of a lightweight network module for capturing global semantic information, a semantic expansion module for capturing local feature information, and a semantic learning module for feature learning.

[0024] Preferably, the lightweight network module extracts features from the two-dimensional semantic plane through a convolutional neural network. Before the two-dimensional representation Γ of the sentence is input into the convolutional neural network, a Stem block is used for feature smoothing. The Stem block can enhance important features while reducing overfitting. The Stem block consists of a 7×7 convolution with an output channel size of 256, a batch normalization layer and an activation function, which is defined as:

[0025] Γ in =σ(BN(Conv 7×7 (Γ))) (3)

[0026] Among them, Γ inis the output of the Stem block; Conv 7×7 is a convolutional network with a stride of 1 and size of 7×7; BN(·) is a batch normalization layer; σ(·) is a ReLU activation function;

[0027] The lightweight network module consists of a deep convolution module and a channel-based MLP module, where the input of the MLP module is the output of the deep convolution module, and a channel scaling operation and a DropPath operation are used after both modules to improve the generalization and robustness of the extracted features;

[0028] In the deep convolution module, the output feature Γ from the Stem block is in First, a group normalization layer is used to group the feature maps along the channel dimension. The input feature Γ after grouping in Input a deep convolution layer, after passing through the deep convolution layer, input feature Γ in Perform channel scaling and DropPath operations, and implement feature representation Γ after completion in The residual connection of is defined as:

[0029]

[0030] in, is the output of the deep convolution module; GN(·) is a group normalization operation; DConv(·) is a deep convolution with a kernel size of 1×1;

[0031] The input of the MLP module comes from the output features of the deep convolution module Input Features First, the features are processed through a group normalization layer, and then channel MLP is implemented on these features. Compared with spatial MLP, channel MLP can not only effectively reduce the computational complexity, but also meet the requirements of entity recognition tasks. After channel MLP, channel scaling operation, DropPath operation and The residual connection of is defined as:

[0032]

[0033] Among them, CMLP(·) is a channel MLP.

[0034] Preferably, the semantic expansion module adopts the nearest neighbor interpolation algorithm, sets the expansion coefficient of the two-dimensional sentence to C, and the output of each expansion operation is calculated as:

[0035]

[0036] in, is the rounding down operation; Γi,j ∈Γ is an element (or span) in the two-dimensional sentence representation; after the two-dimensional sentence representation Γ is expanded, the output Γ Z It is expressed as follows:

[0037]

[0038] Among them, N Z =N×C is the dimension of the generated sentence representation.

[0039] Preferably, after the two-dimensional sentence representation passes through the lightweight network and semantic expansion respectively, in order to better extract the semantic dependencies between adjacent elements, the semantic learning module uses an independent convolutional layer to learn higher-order semantic features. The convolutional layer consists of a two-dimensional CNN with K convolution kernels, a layer normalization operation, and a ReLU activation function, and is defined as:

[0040]

[0041] in, and represents the convolution output vector; σ represents the GELU activation function, and r is the feature size.

[0042] Preferably, the gated integration module integrates the output of the semantic focus module and Standardize and integrate, specifically:

[0043] In view of the scale change in the semantic expansion module operation, the original size is enlarged C times, and the size to be reduced to the original size is defined as:

[0044]

[0045] in, express An element of is the maximum value among the elements of the C×C window after enlargement; Indicates restoration of the original size;

[0046] The integration step aims to filter out redundant information and select important clues from different representations. After normalization, to generate the final sentence representation, the feature vector h output by BERT is used. CLS , which integrates the semantic information of a sentence and converts h CLS As a global sentence representation, it is input into the gated integration module to achieve the purpose of controlling the weight. The final two-dimensional sentence representation is calculated as follows:

[0047] λ=Sigmod(W 3 LeakeReLU(h CLS W 4)) (10)

[0048]

[0049] Among them, W 3 , W 4 is a learnable parameter; λ is the comprehensive weight of different input features; Sigmod is an activation function; It simultaneously encodes both finer and coarser sentence semantics between different feature representations.

[0050] Preferably, in order to avoid performance degradation due to the increase in network depth, the two-dimensional semantic representation Γ is added to M to reduce the problem of gradient disappearance. Then, the final prediction result is calculated using a multi-layer perceptron, which is defined as:

[0051] P = Sigmod(MLP(Γ+M)) (12)

[0052] in, Represents the predicted probability of all entity types; the MLP consists of a linear layer and a LeakeRELU;

[0053] The loss is calculated using binary cross entropy:

[0054]

[0055] Where n is the number of tokens in the sentence; y i,j is the span [x i ,…,x j ] is a binary vector of the true entity label, i and j are the indices of the word unit, P i,j is the span [x i ,…,x j ] is the predicted probability.

[0056] Therefore, the present invention adopts the above-mentioned named entity recognition method based on two-dimensional semantic expansion, which has the following beneficial effects:

[0057] (1) A local semantic expansion method is proposed to capture local features in the two-dimensional sentence representation and enhance local semantic information, thereby achieving more refined feature extraction and solving the problem of entity semantics being dispersed in the span and semantic ambiguity;

[0058] (2) Construct a lightweight network architecture to further improve the quality of feature representation by capturing the global information of the two-dimensional representation, so as to better realize the recognition of long named entities and solve the problem of semantic weakening when the entities are long;

[0059] (3) A gated integration module is used to fuse different feature information and obtain the final feature representation.

[0060] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 A schematic diagram of a model architecture of a named entity recognition method based on two-dimensional semantic expansion according to the present invention;

[0062] Figure 2 Schematic diagram of entity recognition performance of different lengths according to an embodiment of the present invention. DETAILED DESCRIPTION

[0063] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0064] See also Figure 1 , a named entity recognition method based on two-dimensional semantic expansion includes the following steps:

[0065] S1. Construct a model architecture consisting of an encoding module, a semantic focus module, and a gated integration module.

[0066] S2, the encoding module uses a pre-trained language model and a multi-head bi-affine layer to map an original sentence into a two-dimensional sentence representation;

[0067] S3, the semantic focus module is used to capture the local feature information and global semantic information of the two-dimensional sentence representation to improve the quality of feature representation;

[0068] S4: Different feature information is integrated through the gated integration module to obtain the final feature representation and make predictions. Specifically:

[0069] 1. Encoding module

[0070] Assume that the sentence input to the model is X = {x 1 ,x 2 ,…,x N}, where x i is the i-th word unit of the sentence, and N is the length of the sentence X. In order to better encode the semantic information and contextual features of the word units in the sentence, this embodiment uses the pre-trained language model BERT to obtain the vector representation H of the sentence, which is defined as:

[0071] H=BERT(X) (1)

[0072] in, d hIt is a dimension represented by a token;

[0073] This embodiment uses the planar sentence representation method proposed in the prior art 1, and uses the multi-head biaffine method to generate a two-dimensional representation of the vector H, represented by Γ, where the elements in Γ are recorded as Γ i,j , Γ i,j is a span representation of the input sentence X, which is regarded as a possible named entity representation. Secondly, the second prior art proposes that the entity boundary is clear and distinguishable, and the entity boundary contains the contextual features of the potential named entity. Therefore, the model of this paper encodes the entity boundary to obtain more informative features. After encoding the entity boundary, this paper obtains the two-dimensional representation Γ of the sentence by biaffine the two boundary representations of the potential named entity.

[0074] In this embodiment, two multi-layer perceptrons are designed. The vectorized sentence representation H is input into the two multi-layer perceptrons respectively, and the start and end boundary representations of the same word unit with different semantics are learned; the feature vectors of the encoded start and end boundaries are integrated into the two-dimensional sentence representation through the multi-head biaffine method, which is defined as:

[0075]

[0076] in, h is the hidden layer size, MHBiaffine(·,·) is a multi-head biaffine network, r is the feature size, and the element M in the matrix M i,j A vector representation representing a span of the input sentence X.

[0077] 2. Semantic Focus Module

[0078] The semantic focusing module consists of a lightweight network module for capturing global semantic information, a semantic expansion module for capturing local feature information, and a semantic learning module for feature learning.

[0079] 1. Lightweight network module

[0080] Convolutional neural networks can extract features from two-dimensional semantic planes. Compared with traditional convolutional networks, lightweight networks accelerate and compress convolutional neural networks by fusing deep convolution with channel MLP, thereby improving the quality of feature representation.

[0081] Before the two-dimensional representation Γ of the sentence is input into the convolutional neural network, a Stem block is used for feature smoothing. The Stem block can enhance important features while reducing overfitting. The Stem block consists of a 7×7 convolution with an output channel size of 256, a batch normalization layer and an activation function, which is defined as:

[0082] Γ in =σ(BN(Conv 7×7 (Γ))) (3)

[0083] Among them, Γ in is the output of the Stem block; Conv 7×7 is a convolutional network with a stride of 1 and size of 7×7; BN(·) is a batch normalization layer; σ(·) is a ReLU activation function;

[0084] The lightweight network module consists of a deep convolution module and a channel-based MLP module, where the input of the MLP module is the output of the deep convolution module, and a channel scaling operation and a DropPath operation are used after both modules to improve the generalization and robustness of the extracted features;

[0085] In the deep convolution module, the output feature Γ from the Stem block is in First, a group normalization layer is used to group the feature maps along the channel dimension. The input feature Γ after grouping in Input a deep convolution layer, after passing through the deep convolution layer, input feature Γ in Perform channel scaling and DropPath operations, and implement feature representation Γ after completion in The residual connection of is defined as:

[0086]

[0087] in, is the output of the deep convolution module; GN(·) is a group normalization operation; DConv(·) is a deep convolution with a kernel size of 1×1;

[0088] The input of the MLP module comes from the output features of the deep convolution module Input Features First, the features are processed through a group normalization layer, and then channel MLP is implemented on these features. Compared with spatial MLP, channel MLP can not only effectively reduce the computational complexity, but also meet the requirements of entity recognition tasks. After channel MLP, channel scaling operation, DropPath operation and The residual connection of is defined as:

[0089]

[0090] Wherein, CMLP(·) is a channel MLP. In this embodiment, for the convenience of representation, the channel scaling operation and DropPath operation in formula (5) and formula (6) are omitted.

[0091] 2. Semantic Expansion Module

[0092] The semantic expansion module adopts the nearest neighbor interpolation algorithm proposed in the prior art 3, sets the expansion coefficient of the two-dimensional sentence to C, and the output of each expansion operation is calculated as:

[0093]

[0094] in, is the rounding down operation; Γ i,j ∈Γ is an element (or span) in the two-dimensional sentence representation; after the two-dimensional sentence representation Γ is expanded, the output Γ Z It is expressed as follows:

[0095]

[0096] Among them, N Z =N×C is the dimension of the generated sentence representation.

[0097] 3. Semantic Learning Module

[0098] After the two-dimensional sentence representation passes through the lightweight network and semantic expansion respectively, in order to better extract the semantic dependencies between adjacent elements, the semantic learning module uses an independent convolutional layer to learn higher-order semantic features. The convolutional layer consists of a two-dimensional CNN with K convolution kernels, a layer normalization operation, and a ReLU activation function, which is defined as:

[0099]

[0100] in, and represents the convolution output vector; σ represents the GELU activation function, and r is the feature size.

[0101] 3. Gate control integrated module

[0102] The semantic focus module outputs three matrices and They carry different information. This embodiment uses a gated integration module to filter the redundant information in the above three representations and select valid information for integration. Since the size of the original two-dimensional sentence representation changes during the semantic expansion process, the gated integration module is divided into two steps: normalization and integration. The normalization step converts the changed two-dimensional sentence representation to the same dimension as other representations, so that the subsequent integration process can proceed smoothly. The integration step integrates the representations carrying different information.

[0103] In view of the scale change in the semantic expansion module operation, the original size is enlarged C times, and the size to be reduced to the original size is defined as:

[0104]

[0105] in, express An element of is the maximum value among the elements of the C×C window after enlargement; Indicates restoration of the original size;

[0106] The integration step aims to filter out redundant information and select important clues from different representations. After normalization, to generate the final sentence representation, the feature vector h output by BERT is used. CLS , which integrates the semantic information of a sentence and converts h CLS As a global sentence representation, it is input into the gated integration module to achieve the purpose of controlling the weight. The final two-dimensional sentence representation is calculated as follows:

[0107] λ=Sigmod(W 3 LeakeReLU(h CLS W 4 )) (10)

[0108]

[0109] Among them, W 3 , W 4 is a learnable parameter; λ is the comprehensive weight of different input features; Sigmod is an activation function; It simultaneously encodes both finer and coarser sentence semantics between different feature representations.

[0110] 4. Loss Function

[0111] In order to avoid the performance degradation caused by the increase of network depth, the two-dimensional semantic representation Γ is added to M to reduce the problem of gradient disappearance. Then, the multi-layer perceptron is used to calculate the final prediction result, which is defined as:

[0112] P = Sigmod(MLP(Γ+M)) (12)

[0113] in, Represents the predicted probability of all entity types; the MLP consists of a linear layer and a LeakeRELU;

[0114] The loss is calculated using binary cross entropy:

[0115]

[0116] Where n is the number of tokens in the sentence; y i,j is the span [x i ,…,x j] is a binary vector of the real entity label, i and j are the indices of the token, P i,j is the span [x i ,…,x j ] is the predicted probability.

[0117] Unlike the traditional span-based model, this embodiment organizes all possible entity span representations of a sentence into a two-dimensional sentence plane. On this basis, this embodiment proposes a method for named entity recognition based on two-dimensional semantic expansion. This method captures the semantic information of the two-dimensional plane through semantic expansion operations and lightweight networks, improves the quality of feature representation, and has the ability to distinguish the clue differences between the real span and the background. In addition, compared with the traditional convolutional network, the lightweight network module can achieve convolution compression and acceleration, effectively improve the neural network's ability to identify information, thereby further improving the performance of named entity recognition. As shown in Table 3, when using the same pre-trained language model, the F1 values ​​of the model proposed in this embodiment on the ACE2004 and ACE2005 datasets reached 88.16%, 87.25%, and 81.55%, respectively, which are 0.42%, 0.34%, and 0.15% higher than previous work. After replacing the pre-trained language model, the F1 values ​​of the model in this embodiment on the ACE2004 and ACE2005 datasets were increased to 88.87% and 88.52% respectively. These experimental results are higher than the comparison model, proving the effectiveness of the model in this paper in the nested named entity recognition task.

[0118] In order to verify the scalability of the model in different languages, this embodiment is evaluated on two Chinese data sets, namely, the Resume data set and the Weibo data set. In the experiment on the Chinese data set, this embodiment adopts the division of the training set, validation set, and test set provided by the official. The experimental results are shown in Tables 1 and 2.

[0119] Table 1 Performance of Resume dataset

[0120] Model P R F1 LatticeLSTM 94.81 94.11 94.46 Nflat 95.63 95.52 95.58 LexiconAugmentedNER 96.08 96.13 96.11 This method model 96.05 96.93 96.49

[0121] Table 2 Performance of Weibo dataset

[0122]

[0123]

[0124] As can be seen from Tables 1 and 2, the F1 score of the resume dataset reaches 96.49%, which is 0.38% higher than the F1 score of LexiconAugmentedNER. The performance on the Weibo dataset is 72.84%, which is 0.46% higher than End-to-End NER in F1 score. The above results show that the joint learning of global and local information in two-dimensional sentences is also effective for semantic structures in other planes.

[0125] Table 3 Experimental results on English dataset

[0126]

[0127] Compared with the comparison model, the model of this embodiment has been improved in all aspects. The reason for the improvement is that in the two-dimensional sentence representation, the semantics between adjacent spans are similar and share the same contextual features, which leads to the semantic dispersion and fuzziness of the real entity span. The model of this embodiment extracts global and local information of the two-dimensional sentence plane through local semantic expansion and lightweight network operation, enhances the semantic representation of the sentence, and realizes more refined and overall semantic features. In addition, the model of this embodiment also uses a gated integration module to filter redundant information in different semantic feature representations, retains more valuable information, and improves the quality of feature representation. Therefore, the model of this embodiment has achieved significant improvement.

[0128] This paper aims to solve the problems of entity semantics dispersion in span, semantic ambiguity and semantic weakening in long named entities. This embodiment introduces semantic expansion and lightweight network in the semantic focus module. The module realizes the recognition of named entities of different lengths by jointly capturing the overall and local semantic information in the semantic plane. In order to verify the effectiveness of the model of this embodiment in the recognition of named entities of different lengths, the model in this paper is compared with the comparison model in entities of different lengths, and the experimental data set is the ACE2004 data set. The performance of different named entity lengths is shown in Figure 2. Figure 2 shown.

[0129] As can be seen from the figure, the performance of named entity recognition gradually decreases as the length of the named entity increases. This is mainly due to the gradient vanishing problem, which makes it difficult to encode the semantic dependencies of entities in longer named entities. However, compared with the comparison model, the proposed model shows higher performance. It is particularly noteworthy that when the length of the named entity increases, the performance of the comparison model decreases significantly, while the proposed model maintains stable performance when the length of the named entity is greater than 8.

[0130] Therefore, the present invention adopts the above-mentioned named entity recognition method based on two-dimensional semantic expansion, and proposes a local semantic expansion method to address the problems of entity semantics being dispersed in the span and semantic ambiguity, thereby enhancing local semantic information and achieving more refined feature extraction; by constructing a lightweight network architecture to capture the global information of the two-dimensional representation, the quality of the feature representation is further improved, thereby better identifying long named entities.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A named entity recognition method based on two-dimensional semantic expansion, characterized in that: The following steps are involved: S1. Construct a model architecture consisting of an encoding module, a semantic focus module, and a gated integration module. S2, mapping an original sentence into a two-dimensional sentence representation through the encoding module; S3, capturing the local feature information and global semantic information of the two-dimensional sentence representation through the semantic focus module; S4. Different feature information is integrated through the gated integration module to obtain the final feature representation and make predictions.

2. The method for named entity recognition based on two-dimensional semantic expansion according to claim 1, characterized in that: Step S2 is specifically as follows: Assume that the sentence input to the model is X = {x1, x2, ..., x N }, where x i is the i-th token of the sentence, N is the length of the sentence X; the pre-trained language model BERT is used to obtain the vector representation H of the sentence, which is defined as: H = BERT(X); in, d h It is a dimension represented by a token; The multi-head biaffine method is used to generate a two-dimensional representation of the vector H, represented by Γ, where the elements in Γ are recorded as Γ i,j , Γ i,j is a span representation of the input sentence X, and the entity boundary is encoded. After encoding the entity boundary, the two-dimensional representation Γ of the sentence is obtained by biaffine the two boundary representations of the potential named entity; two multi-layer perceptrons are designed, and the vectorized sentence representation H is input into the two multi-layer perceptrons respectively, and the start and end boundary representations of the same token but with different semantics are learned; the feature vectors of the encoded start and end boundaries are integrated into the two-dimensional sentence representation through the multi-head biaffine method, which is defined as: in, h is the hidden layer size, MHBiaffine(·,·) is a multi-head biaffine network, r is the feature size, and the element M in the matrix M i,j A vector representation representing a span of the input sentence X.

3. The method for named entity recognition based on two-dimensional semantic expansion according to claim 2, characterized in that: The semantic focusing module in step S3 consists of a lightweight network module for capturing global semantic information, a semantic expansion module for capturing local feature information, and a semantic learning module for feature learning.

4. The method for named entity recognition based on two-dimensional semantic expansion according to claim 3, characterized in that: The lightweight network module extracts features from the two-dimensional semantic plane through a convolutional neural network. Before the two-dimensional representation Γ of the sentence is input into the convolutional neural network, a Stem block is used for feature smoothing. The Stem block consists of a 7×7 convolution with an output channel size of 256, a batch normalization layer, and an activation function, which is defined as: C in =σ(BN(Conv 7×7 (C))); Among them, Γ in is the output of the Stem block; Conv 7×7 is a convolutional network with a stride of 1 and size of 7×7; BN(·) is a batch normalization layer; σ(·) is a ReLU activation function; The lightweight network module consists of a deep convolution module and a channel-based MLP module, where the input of the MLP module is the output of the deep convolution module, and a channel scaling operation and a DropPath operation are used after both modules; In the deep convolution module, the output feature Γ from the Stem block is in First, a group normalization layer is used to group the feature maps along the channel dimension. The input feature Γ after grouping in Input a deep convolution layer, after passing through the deep convolution layer, input feature Γ in Perform channel scaling and DropPath operations, and implement feature representation Γ after completion in The residual connection of is defined as: in, is the output of the deep convolution module; GN(·) is a group normalization operation; DConv(·) is a deep convolution with a kernel size of 1×1; The input of the MLP module comes from the output features of the deep convolution module Input Features First, it is processed through a group normalization layer, and then a channel MLP is implemented on these features. After the channel MLP, channel scaling operations, DropPath operations, and The residual connection of is defined as: Among them, CMLP(·) is a channel MLP.

5. The method for named entity recognition based on two-dimensional semantic expansion according to claim 4, characterized in that: The semantic expansion module uses the nearest neighbor interpolation algorithm and sets the expansion coefficient of the two-dimensional sentence to C. The output of each expansion operation is calculated as: in, is the rounding down operation; Γ i,j ∈Γ is an element in the two-dimensional sentence representation; after the two-dimensional sentence representation Γ is expanded, the output Γ Z It is expressed as follows: Among them, N Z =N×C is the dimension of the generated sentence representation.

6. The method for named entity recognition based on two-dimensional semantic expansion according to claim 5, characterized in that: The semantic learning module uses an independent convolutional layer to learn high-order semantic features. The convolutional layer consists of a two-dimensional CNN with K convolution kernels, a layer normalization operation, and a ReLU activation function, which is defined as: in, and represents the convolution output vector; σ represents the GELU activation function, and r is the feature size.

7. The method for named entity recognition based on two-dimensional semantic expansion according to claim 6, characterized in that: The gated integration module integrates the output of the semantic focus module and Standardize and integrate, specifically: In view of the scale change in the semantic expansion module operation, the original size is enlarged C times, and the size to be reduced to the original size is defined as: in, express An element of is the maximum value among the elements of the C×C window after enlargement; Indicates restoration of the original size; After normalization, the feature vector h output by BERT is used CLS , h CLS As a global sentence representation input to the gated integration module, the final two-dimensional sentence representation is calculated as follows: λ=Sigmod(W3LeakeReLU(h CLS W4)); Among them, W3 and W4 are learnable parameters; λ is the comprehensive weight of different input features; Sigmod is an activation function; It simultaneously encodes sentence semantics between different feature representations.

8. The method for named entity recognition based on two-dimensional semantic expansion according to claim 7, characterized in that: The two-dimensional semantic representation Γ is added to M, and then the final prediction result is calculated using a multilayer perceptron, which is defined as: P = Sigmod(MLP(Γ+M)); in, Represents the predicted probability of all entity types; the MLP consists of a linear layer and a LeakeRELU; The loss is calculated using binary cross entropy: Where n is the number of tokens in the sentence; y i,j is the span [x i ,…,x j ] is a binary vector of the real entity label, i and j are the indices of the token, P i,j is the span [x i ,…,x j ] is the predicted probability.

Citation Information

Patent Citations

  • Entity identification method and device, electronic equipment and storage medium

    CN111985239A

  • Named entity recognition method based on mixed scale sentence representation

    CN118966225A