Nested named entity recognition method based on multi-directional gradient feature extraction
By using the extended eight-direction Sobel operator to extract semantic edge features of multi-directional entities in the nested named entity recognition model, the problem of missing in feature extraction is solved, and more efficient nested named entity recognition performance is achieved.
Patent Information
- Application Number
- CN202510069987.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
When extracting entity semantic edge features, the existing nested named entity recognition model only considers edge information in four directions between adjacent spans, and ignores feature information in multiple other directions, resulting in the missing features extracted.
The semantic edge features of multi-directional entity are extracted from planarized sentence representations by using the extended eight-directional Sobel operator, and more complete and distinctive features are obtained through the combination of channel-by-channel convolution and the extended edge gradient operator.
The performance of nested named entity recognition tasks is improved, the extracted features are more complete and distinctive, and the accuracy of entity recognition is improved.
Smart Images

Figure CN119990130A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a nested named entity recognition method for extracting multi-directional gradient features. Background Art
[0002] Information extraction aims to extract structured information from large-scale unstructured or semi-structured natural language texts. The main tasks include entity recognition, relationship extraction, and event extraction. Among them, the named entity recognition (NER) task aims to identify predefined noun phrases with special meanings given a sentence, such as people, locations, and geographic location entities. As a basic task for many applications such as knowledge graphs, question answering, and machine translation, it has received extensive attention and research in the field of natural language processing.
[0003] Early named entity recognition was usually implemented as a sequence labeling task, which assigns a tag to each element in a sentence to indicate its semantic role in the named entity. But the problem with sequence models is that it is difficult to identify nested named entities in a sentence. Nested named entities refer to the phenomenon that an entity contains one or more other entities, and there are multiple labels for the same element. For example, in "Peking University", "Beijing" is labeled as a geographic location entity, and "Peking University" is labeled as an organizational entity, where the entity "Peking University" contains another entity "Beijing".
[0004] For the task of nested named entity recognition, related research models can be roughly divided into hierarchical sequence-based models and span classification-based models. Hierarchical sequence-based models recognize nested entities by dynamically stacking flat sequence recognition layers, but this method brings about the problem of error cascade to a certain extent, and can only guarantee one-way transmission of information without allowing the inner layer to utilize the information of the outer layer. Span-based models regard entity recognition as a span classification task. Enumerating overlapping spans can unfold the nested semantic structure and make full use of the marking characteristics in the span. However, existing span models are mainly based on independent and fragmented spans for classification, which destroys the interaction between them and cannot encode the global semantic features of the sentence.
[0005] In order to make up for the defects of the model based on span classification, some studies have structured the exhaustive span into a two-dimensional sentence representation, which can expand the nested structure in the sentence and encode the interaction between words. However, adjacent spans share the same context, and entities will have semantic overlap in the flat sentence representation. Therefore, combined with the edge detection technology in image processing, the gradient operator is used to enhance and extract the entity semantic edge features in the flat sentence representation, and the Laplace operator is used to capture the semantic gradient between adjacent spans in the four directions of up, down, left, and right. However, the above studies only consider the edge information in the four directions between adjacent spans, ignoring the feature information in multiple other directions, resulting in the lack of extracted entity semantic edge features. Summary of the invention
[0006] The purpose of the present invention is to provide a nested named entity recognition method based on multi-directional gradient feature extraction. Based on the edge detection method in the field of image processing, an extended eight-directional Sobel operator is used in planar sentence representation to solve the problem of insufficient entity semantic edge feature extraction, thereby extracting more complete and discriminative entity semantic edge features in two-dimensional representation.
[0007] To achieve the above object, the present invention provides a nested named entity recognition method based on multi-directional gradient feature extraction, comprising the following steps:
[0008] Step S1, preprocessing the text data set, that is, processing the original data into data suitable for the entity model;
[0009] Step S2: Input the sentence preprocessed in step S1 into the pre-trained BERT model to obtain the context features of the word vector;
[0010] Step S3, performing planar sentence representation on the sentence with context information features obtained in step S2;
[0011] Step S4, extracting multi-directional entity semantic edge features from the sentence represented by the planarization in step S3 by combining channel-by-channel convolution with an extended edge gradient operator; and spatially connecting the sentences after feature extraction by point-by-point convolution to obtain high-order features;
[0012] Step S5: First, the high-order features obtained in step S4 are sent to the multi-layer perceptron, and then residual connections are made with the sentences represented by the flattened sentences obtained in step S3. Finally, Softmax and Argmax are used to predict the classification return index value to complete the screening of candidate entities.
[0013] Preferably, in step S1, the text data set is preprocessed, that is, the original data is processed into data suitable for the entity model; the position of the entity in the sentence is marked, from the beginning to the end, and the entity type is marked with type, and the sentence structure is obtained through pre-training.
[0014] Preferably, in step S2, the sentence preprocessed in step S1 is input into the pre-trained BERT model to obtain the context features of the word vector, and the specific process is as follows:
[0015] First, for the N words x=[x i ],1≤i≤N sentences, each word x i Convert to word fragments and input them into the pre-trained BERT model;
[0016] Then, the word embedding vector is trained using the Bi-LSTM network, running from front to back and back to front, fusing the front and back information together so that the model fully considers the front and back context information at each moment in the input sequence, and obtains the final word representation H:
[0017] H={h1,h2,···,h N};
[0018] Among them, h i ,1≤i≤N represents the concatenation of the bidirectional representation of the i-th word.
[0019] Preferably, in step S3, the sentence with context information features obtained in step S2 is flattened into a sentence representation, and the specific process is as follows:
[0020] The input of multi-head biaffine is two matrices H s ,H e ∈Γ N×h ; The output is R∈Γ N×N×r ; Use the multi-head bi-affine decoder to get the flat sentence representation R as follows:
[0021]
[0022] R=MHBiaffine(H s ,H e );
[0023] Where N, h, and r represent the number of words in a sentence, the size of the hidden vector, and the size of the feature, respectively; W s and W e Represent the word embedding vector of the starting position of each word in the sentence and the word embedding vector of the ending position of each word in the sentence respectively; each cell in the output R represents the feature vector v∈Γ of the span r .
[0024] Preferably, in step S4, the sentence represented by the planarized sentence in step S3 is subjected to channel-by-channel convolution combined with an extended edge gradient operator to extract multi-directional entity semantic edge features; the sentence after feature extraction is spatially connected using point-by-point convolution to obtain high-order features. The specific process is as follows:
[0025] Step S41, using depthwise separable convolution to support multi-directional Sobel gradient operator; wherein the depthwise separable convolution is decomposed into depthwise convolution and 1×1 convolution, i.e., point-by-point convolution;
[0026] The deep convolution performs a separate convolution operation on each channel of the flattened sentence representation R using a gradient operator to extract entity semantic edge features, as shown below:
[0027]
[0028] in, represents the flattened sentence representation after deep convolution operation; w (d) Represents the deep convolution operation function; G∈Γ K×K represents the derivative operator in the form of filter mask, K represents the size of the derivative operator;
[0029] Step S42: Use point-by-point convolution to perform a 1×1 standard convolution operation to learn the interaction between different channels, as shown below:
[0030]
[0031] in, represents the planar sentence representation after point-by-point convolution operation; w (p) represents the point-by-point convolution operation function; C represents the convolution kernel;
[0032] Step S43: Combine the above two convolution formulas into a unified form, as shown below:
[0033] w (s) (R,G,C)=w (p) (w (d) (R,G),C);
[0034] Among them, the specific formulas in eight different directions are expressed as follows:
[0035]
[0036] Among them, w (0) 、w (45) 、w (90) ···w (315)They represent the gradient of the element in the 0° direction, the gradient of the element in the 45° direction, the gradient of the element in the 90° direction, and the gradient of the element in the 135° direction respectively; G0, G 45 , G 90 ···G 315 They respectively represent the derivative operator corresponding to the 0° direction filter mask form, the derivative operator corresponding to the 45° direction filter mask form, the derivative operator corresponding to the 90° direction filter mask form...the derivative operator corresponding to the 135° direction filter mask form;
[0037] Step S44: The gradient of each element is the maximum absolute value of the gradient values in eight directions, as shown below:
[0038] G[w (0,45,···,315) ]≈max{|w0|,|w 45 |,|w 90 |,···,|w 315 |};
[0039] Finally, the sentence after entity semantic edge feature extraction is represented as W.
[0040] Preferably, in step S5, the high-order features obtained in step S4 are first fed into a multi-layer perceptron, and then residually connected with the flattened sentence obtained in step S3, and finally Softmax and Argmax are used to predict the classification return index value to complete the screening of candidate entities. The specific process is as follows:
[0041] Step S51, first, W is directly sent to the MLP layer to further mix the semantic information;
[0042] Step S51, then, W is combined with R to be expressed as A to predict the distribution of named entities, as shown below:
[0043] P θ (C|A) = Softmax(Linear(A));
[0044] Among them, C represents the predicted entity label, which is an N×N entity label matrix; the element C i,j is the corresponding entity span A i,j The entity label of θ; θ is the parameter that the model needs to learn during training;
[0045] The cross entropy function is used as the loss function during training, given the input X and the true label matrix The loss function is calculated as follows:
[0046]
[0047] where z is the number of NER categories + 1, indicating non-entity; y ij Yes A ij Entity label output; N 2 is the normalization factor;
[0048] Step S53, finally, predict the entity label of each span by selecting the category with the highest probability, as shown below:
[0049]
[0050] Among them, P θ Y represents the probability distribution of the model predicting a certain entity label under given parameters θ; i,j Represents the entity tag value corresponding to the span from the i-th word to the j-th word; e represents a specific entity tag category.
[0051] Therefore, the present invention adopts the above-mentioned nested named entity recognition method of multi-directional gradient feature extraction. Compared with the prior art, the present invention makes full use of the complete information of the sentence text and adopts an extended eight-directional Sobel operator in the planar sentence representation to solve the problem of insufficient semantic edge feature extraction, so as to extract more complete and discriminative entity semantic edge features in the planar sentence representation, thereby improving the performance of the nested named entity recognition task and achieving excellent results in entity recognition.
[0052] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a technical route of a nested named entity recognition method for multi-directional gradient feature extraction in the present invention;
[0054] Figure 2 It is an entity recognition model diagram of a nested named entity recognition method for multi-directional gradient feature extraction of the present invention;
[0055] Figure 3 It is a schematic diagram of the directions of the expanded eight-directional Sobel operator of the present invention. DETAILED DESCRIPTION
[0056] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.
[0057] like Figure 1 and Figure 2 As shown, a nested named entity recognition method based on multi-directional gradient feature extraction includes the following steps:
[0058] Step S1, preprocessing the text data set, that is, processing the original data into data suitable for the entity model.
[0059] Mark the position of the entity in the sentence, from the beginning to the end, and use type to mark the entity type, and obtain the sentence structure through pre-training.
[0060] Step S2: Input the sentence preprocessed in step S1 into the model to obtain the contextual features of the word vector.
[0061] For the dataset containing N words x=[x i ],1≤i≤N sentences, each word x i Convert them into word fragments and input them into the pre-trained BERT module. After BERT calculation, each word of the sentence may involve vector representations of several fragments.
[0062] Next, the word embedding vector is trained using the Bi-LSTM network, running from front to back and back to front, fusing the front and back information together so that the model can fully consider the front and back context information at each moment in the input sequence, and obtain the final word representation H:
[0063] H={h1,h2,···,h N};
[0064] Among them, h i ,1≤i≤N represents the concatenation of the bidirectional representation of the i-th word.
[0065] Step S3: Flatten the sentence representation of the sentence with context information features obtained in step S2.
[0066] Input H obtained in step S2 into the multi-headed biaffine to obtain a flattened sentence vector representation. The input of the multi-headed biaffine is two matrices H s ,H e ∈Γ N×h ; The output is R∈Γ N×N×r Therefore, the multi-head bi-affine decoder is used to obtain the flattened sentence representation R as follows:
[0067]
[0068] R=MHBiaffine(H s ,H e );
[0069] Where N, h, and r represent the number of words in a sentence, the size of the hidden vector, and the size of the feature, respectively; W s and W eRepresent the word embedding vector of the starting position of each word in the sentence and the word embedding vector of the ending position of each word in the sentence respectively; each cell (i, j) in the output R can be regarded as the feature vector v∈Γ of the span r .
[0070] Step S4: extract multi-directional entity semantic edge features from the sentence represented by the planarization in step S3 by combining channel-by-channel convolution with an extended edge gradient operator; and spatially connect the sentences after feature extraction using point-by-point convolution to obtain high-order features.
[0071] Step S41, use depthwise separable convolution to support multi-directional Sobel gradient operators, such as Figure 3 shown.
[0072] Depthwise separable convolution can be decomposed into depthwise convolution and 1×1 convolution (also called pointwise convolution). Depthwise convolution is a separate convolution operation using a gradient operator for each channel of the flattened sentence representation R to extract entity semantic edge features, as shown below:
[0073]
[0074] in, represents the flattened sentence representation after deep convolution operation; w (d) It can be regarded as a deep convolution operation function; G∈Γ K×K Represents the derivative operator in the form of a filter mask, and K represents the size of the derivative operator.
[0075] Step S42, use point-by-point convolution to perform a 1×1 standard convolution operation to learn the interaction between different channels, as shown below:
[0076]
[0077] in, represents the planar sentence representation after point-by-point convolution operation; w( p ) can be regarded as a point-by-point convolution operation function; C represents the convolution kernel.
[0078] Step S43: Combine the above two convolution formulas into a unified form, as shown below:
[0079] w (s) (R,G,C)=w (p) (w (d) (R,G),C);
[0080] Among them, the specific formulas in eight different directions are expressed as follows:
[0081]
[0082] Among them, w (0) 、w (45) 、w (90) ···w (315) They represent the gradient of the element in the 0° direction, the gradient of the element in the 45° direction, the gradient of the element in the 90° direction, and the gradient of the element in the 135° direction respectively; G0, G 45 , G 90 ···G 315 They respectively represent the derivative operator corresponding to the 0° direction filter mask form, the derivative operator corresponding to the 45° direction filter mask form, the derivative operator corresponding to the 90° direction filter mask form... and the derivative operator corresponding to the 135° direction filter mask form.
[0083] Step S44: The gradient of each element is the maximum absolute value of the gradient values in eight directions, as shown below:
[0084] G[w (0,45,···,315) ]≈max{|w0|,|w 45 |,|w 90 |,···,|w 315 |};
[0085] Finally, the sentence after entity semantic edge feature extraction is represented as W.
[0086] Step S5: First, the high-order features obtained in step S4 are sent to the multi-layer perceptron, and then residual connection is performed with the flattened sentence obtained in step S3. Finally, Softmax and Argmax are used to predict the classification return index value to complete the screening of candidate entities.
[0087] Step S51: First, W is directly sent to the MLP layer to further mix the semantic information.
[0088] Step S51, then, W is combined with R to be expressed as A to predict the distribution of named entities, as shown below:
[0089] P θ (C|A) = Softmax(Linear(A));
[0090] Among them, C represents the predicted entity label, which is an N×N entity label matrix; the element C i,j is the corresponding entity span A i,j ; θ is the parameter that the model needs to learn during training.
[0091] The cross entropy function is used as the loss function during training, given the input X and the true label matrix The loss function is calculated as follows:
[0092]
[0093] Where z is the number of NER categories + 1 (indicating non-entity); y ij Yes A ij Entity label output; N 2 is the normalization factor.
[0094] Step S53, finally, predict the entity label of each span by selecting the category with the highest probability, as shown below:
[0095]
[0096] Among them, P θ Y represents the probability distribution of the model predicting a certain entity label under given parameters θ; i,j Represents the entity tag value corresponding to the span from the i-th word to the j-th word; e represents a specific entity tag category.
[0097] Example 1
[0098] This embodiment implements the method of the present invention, mainly applied to nested entity data, and the datasets used are the GENIA dataset and the ACE2005 dataset. First, the position of the entity in the sentence is marked, from the beginning to the end, and the entity type is marked with type, and the sentence structure is obtained through pre-training.
[0099] Then, for a given sentence, such as "Mr. A from a certain country unfortunately passed away", the pre-trained model BERT is used to obtain the vector representation of each word or character in the sentence, and then the Bi-LSTM network is used to train the character embedding vector, running from front to back and from back to front to fuse the front and back information together, so that the model comprehensively considers the front and back context information at each moment in the input sequence to obtain the final word representation.
[0100] Next, the word representation H is flattened into a sentence representation through multi-headed biaffine operations to obtain a sentence in matrix form. The sentence after flattening is extracted through a combination of channel-by-channel convolution and extended edge gradient operator to extract multi-directional entity semantic edge features; the sentence after feature extraction is spatially connected using point-by-point convolution to obtain high-order features.
[0101] Finally, the obtained flattened sentence representation is residually connected with the obtained sentence representation after feature extraction to predict the distribution of named entities. The final prediction results are the four entities: "a country", "A", "Mr. A", and "Mr."
[0102] Example 2
[0103] This embodiment is evaluated using the GENIA dataset and the ACE2005 dataset. The ACE2005 dataset is a Chinese nested named entity dataset, which contains documents collected from news agencies, broadcasts, and blogs. Among them, 7 entity types are defined: personnel, organizations, locations, facilities, weapons, vehicles, and geo-entities, with a total of 45,112 entities. In the processing of the dataset, the present invention divides the dataset into a training set, a validation set, and a test set at a ratio of 8:1:1. The GENIA dataset contains 32 nested biomedical entity categories. The present invention focuses on five entity types: DNA, RNA, protein, cell lineage, and cell type, and divides the training set, validation set, and test set into 8.1:0.9:1.0.
[0104] The present invention compares the current mainstream entity recognition models on these two data sets, and uses P value, R value, and F1 value to evaluate the performance of the model method. As shown in Table 1, the experimental results of the present invention are higher than those of other model methods, indicating the effectiveness of the present invention in improving the performance of named entity recognition tasks.
[0105] Table 1 Experimental comparison results
[0106]
[0107] As shown in Table 2, ablation experiments were conducted on the key modules of this civilization on the GENIA dataset. In the absence of residual connections, the performance dropped significantly, proving that removing residual connections would reduce the fault tolerance of the model. After removing the multi-directional gradient operator, the performance dropped by 0.34%, indicating that the multi-directional gradient operator has a positive impact on the nested NER task. The use of multi-directional gradient operators in deep convolution effectively strengthens the semantic representation of the actual corpus, extracts multi-directional semantic edge features between adjacent spans, and improves the accuracy of extracting named entities from sentences. After removing the point-by-point convolution module, the performance dropped by 0.27%, verifying the impact of the 1×1 convolution operation on the model performance, which realizes the fusion of cross-channel information.
[0108] Table 2 Ablation experiment results
[0109]
[0110] Therefore, the present invention adopts the above-mentioned nested named entity recognition method of multi-directional gradient feature extraction. Compared with the prior art, the present invention makes full use of the complete information of the sentence text and adopts an extended eight-directional Sobel operator in the planar sentence representation to solve the problem of insufficient semantic edge feature extraction, so as to extract more complete and discriminative entity semantic edge features in the planar sentence representation, thereby improving the performance of the nested named entity recognition task and achieving excellent results in entity recognition.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A nested named entity recognition method based on multi-directional gradient feature extraction, characterized in that: The following steps are involved: Step S1, preprocessing the text data set, that is, processing the original data into data suitable for the entity model; Step S2: Input the sentence preprocessed in step S1 into the pre-trained BERT model to obtain the context features of the word vector; Step S3, performing planar sentence representation on the sentence with context information features obtained in step S2; Step S4, extracting multi-directional entity semantic edge features from the sentence represented by the planarization in step S3 by combining channel-by-channel convolution with an extended edge gradient operator; and spatially connecting the sentences after feature extraction by point-by-point convolution to obtain high-order features; Step S5: First, the high-order features obtained in step S4 are sent to the multi-layer perceptron, and then residual connections are made with the sentences represented by the flattened sentences obtained in step S3. Finally, Softmax and Argmax are used to predict the classification return index value to complete the screening of candidate entities.
2. The nested named entity recognition method based on multi-directional gradient feature extraction according to claim 1, characterized in that: In step S1, the text dataset is preprocessed, that is, the original data is processed into data suitable for the entity model; the position of the entity in the sentence is marked, from the beginning to the end, and the entity type is marked with type, and the sentence structure is obtained through pre-training.
3. The nested named entity recognition method based on multi-directional gradient feature extraction according to claim 1 is characterized in that: In step S2, the sentence preprocessed in step S1 is input into the pre-trained BERT model to obtain the contextual features of the word vector. The specific process is as follows: First, for the N words x=[x i ],1≤i≤N sentences, each word x i Convert to word fragments and input them into the pre-trained BERT model; Then, the word embedding vector is trained using the Bi-LSTM network, running from front to back and back to front, fusing the front and back information together so that the model fully considers the front and back context information at each moment in the input sequence, and obtains the final word representation H: H={h1,h2,···,h N }; Among them, h i ,1≤i≤N represents the concatenation of the bidirectional representation of the i-th word.
4. The nested named entity recognition method based on multi-directional gradient feature extraction according to claim 1, characterized in that: In step S3, the sentence with context information features obtained in step S2 is flattened into a sentence representation. The specific process is as follows: The input of multi-head biaffine is two matrices H s ,H e ∈Γ N×h ; The output is R∈Γ N×N×r ; Use the multi-head bi-affine decoder to get the flat sentence representation R as follows: R=MHBiaffine(H s ,H e ); Where N, h, and r represent the number of words in a sentence, the size of the hidden vector, and the size of the feature, respectively; W s and W e Represent the word embedding vector of the starting position of each word in the sentence and the word embedding vector of the ending position of each word in the sentence respectively; each cell in the output R represents the feature vector v∈Γ of the span r .
5. The nested named entity recognition method based on multi-directional gradient feature extraction according to claim 1 is characterized in that: In step S4, the sentence represented by the planarized sentence in step S3 is subjected to channel-by-channel convolution combined with an extended edge gradient operator to extract multi-directional entity semantic edge features; the sentence after feature extraction is spatially connected using point-by-point convolution to obtain high-order features. The specific process is as follows: Step S41, using depthwise separable convolution to support multi-directional Sobel gradient operator; wherein the depthwise separable convolution is decomposed into depthwise convolution and 1×1 convolution, i.e., point-by-point convolution; The deep convolution performs a separate convolution operation on each channel of the flattened sentence representation R using a gradient operator to extract entity semantic edge features, as shown below: in, represents the flattened sentence representation after deep convolution operation; w(d) represents the deep convolution operation function; G∈Γ K×K represents the derivative operator in the form of filter mask, K represents the size of the derivative operator; Step S42, use point-by-point convolution to perform a 1×1 standard convolution operation to learn the interaction between different channels, as shown below: in, represents the planar sentence representation after point-by-point convolution operation; w( p ) represents the point-by-point convolution operation function; C represents the convolution kernel; Step S43: Combine the above two convolution formulas into a unified form, as shown below: w ( s ) (R,G,C)=w ( p ) (w ( d ) (R,G),C); Among them, the specific formulas in eight different directions are expressed as follows: Among them, w(0), w( 45 )、w( 90 )···w( 315 ) represent the gradient of the element in the 0° direction, the gradient of the element in the 45° direction, the gradient of the element in the 90° direction...the gradient of the element in the 135° direction; G0, G 45 , G 90 ···G 315 They respectively represent the derivative operator corresponding to the 0° direction filter mask form, the derivative operator corresponding to the 45° direction filter mask form, the derivative operator corresponding to the 90° direction filter mask form...the derivative operator corresponding to the 135° direction filter mask form; Step S44: The gradient of each element is the maximum absolute value of the gradient values in eight directions, as shown below: G[ow (0,45,···,315) ]≈max{|w0|,|w 45 |,|in 90 |,···,|in 315 |}; Finally, the sentence after entity semantic edge feature extraction is represented as W.
6. The nested named entity recognition method based on multi-directional gradient feature extraction according to claim 1, characterized in that: In step S5, the high-order features obtained in step S4 are first fed into a multi-layer perceptron, and then residually connected with the flattened sentence obtained in step S3. Finally, Softmax and Argmax are used to predict the classification return index value to complete the screening of candidate entities. The specific process is as follows: Step S51, first, W is directly sent to the MLP layer to further mix the semantic information; Step S51, then, W is combined with R to be expressed as A to predict the distribution of named entities, as shown below: P θ (C|A)=Softmax(Linear(A)); Among them, C represents the predicted entity label, which is an N×N entity label matrix; the element C i,j is the corresponding entity span A i,j The entity label of θ; θ is the parameter that the model needs to learn during training; The cross entropy function is used as the loss function during training, given the input X and the true label matrix The loss function is calculated as follows: where z is the number of NER categories + 1, indicating non-entity; y ij Yes A ij Entity label output; N 2 is the normalization factor; Step S53, finally, predict the entity label of each span by selecting the category with the highest probability, as shown below: Among them, P θ Y represents the probability distribution of the model predicting a certain entity label under given parameters θ; i,j Represents the entity tag value corresponding to the span from the i-th word to the j-th word; e represents a specific entity tag category.
Citation Information
Patent Citations
Nested named entity semantic enhancement method and system based on edge gradient
CN116227491A
Nested named entity recognition method based on part-of-speech awareness, device and storage medium therefor
US20240111956A1