Method and system for extracting few-sample information
By adopting an iterative optimization method of sentence support sets and multiple loss functions in the recognition of few-sample named entity, the problem of boundary detection is solved, and the accuracy of named entity recognition is significantly improved.
Patent Information
- Application Number
- CN202510367073.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing method of naming entity recognition is prone to accidentally detecting entities that do not belong to predefined categories during the boundary detection stage, resulting in low recognition accuracy.
The sentence support set is used for model training, combined with the entity boundary detector and entity classifier, and iteratively optimizes the model parameters through boundary detection classification loss, boundary perception comparison loss, token comparison loss, regularization loss, prototype classification loss and entity comparison loss to improve the accuracy of entity boundary detection and classification.
By introducing boundary-aware contrast loss and entity comparison loss, the accuracy of entity boundary detection and the effect of entity classification are improved, and the accuracy of naming entity recognition of few samples is significantly improved.
Smart Images

Figure CN120146053A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a few-shot information extraction method and system. Background Art
[0002] Information extraction is one of the core technologies in natural language processing. Named entity recognition, as entity information extraction in information extraction, aims to locate the boundaries of entities in a given sentence and classify them into predefined categories, and its results are usually used in downstream tasks such as entity linking, relation extraction, and syntactic analysis.
[0003] Traditional named entity recognition methods are implemented by deep neural networks trained on large-scale labeled datasets. However, the labeled data is time-consuming and laborious, so it is necessary to explore how to perform few-shot named entity recognition tasks. Currently, many methods regard few-shot named entity recognition as a single-stage sequence labeling task. The core idea is to classify by comparing the semantic distance between the token to be classified and each category prototype, but usually ignores the integrity of the entity and the potential connection between tokens. For this reason, some researchers have proposed a two-stage method, which decomposes the few-shot named entity recognition task into two independent processes: entity boundary detection and entity classification. In the boundary detection stage, the tokens are divided into entity tokens and non-entity tokens to locate the boundaries of named entity recognition, and then in the entity classification stage, the detected entities are classified into predefined categories. However, some entities that do not belong to the predefined categories will be detected in the boundary detection stage, resulting in low accuracy of few-shot named entity recognition. Summary of the Invention
[0004] The present invention provides a few-shot information extraction method and system for improving the entity recognition accuracy of existing few-shot named entity recognition methods.
[0005] A few-shot information extraction method provided in the first aspect of the present invention includes:
[0006] Using the sentence samples of the sentence support set to input an initial named entity recognition model for model training, the initial named entity recognition model includes an entity boundary detector and an entity classifier;
[0007] Performing boundary detection on the sentence samples through the entity boundary detector, outputting entity boundaries, and calculating a boundary detection classification loss and a boundary-aware contrast loss;
[0008] Using the sentence samples and the entity boundaries through the entity classifier for entity classification, outputting a predicted entity label probability distribution of multiple target entity embeddings, and calculating a token contrast loss, a regularization loss, a prototype classification loss, and an entity contrast loss;
[0009] Iteratively optimize the model parameters of the initial named entity recognition model based on the boundary detection classification loss, the boundary-aware contrast loss, the token contrast loss, the regularization loss, the prototype classification loss, and the entity contrast loss, and use the sentence query set for verification to determine the target named entity recognition model;
[0010] Input the sentence to be recognized into the target named entity recognition model for boundary detection and entity classification, and output the target recognition result.
[0011] Further, the boundary detection of the sentence sample by the entity boundary detector, outputting the target position labels of multiple token embeddings and calculating the boundary detection classification loss, includes:
[0012] Convert the sentence sample into a position token embedding sequence based on the boundary BERT embedding layer;
[0013] Use a linear layer to map the position token embedding sequence to the BIOES tag space, and calculate the predicted position label probability distribution of each position token embedding based on the softmax function;
[0014] Through the conditional random field of Viterbi decoding, determine the target position labels of each position token embedding according to each predicted position label probability distribution, and use each target position label to determine the entity boundary;
[0015] Calculate the boundary detection classification loss based on each predicted label probability distribution, and calculate the corresponding boundary-aware contrast loss according to the position entity embeddings, positive sample representations, and negative sample representations of each entity boundary.
[0016] Further, the entity classifier uses the sentence sample and the entity boundary for entity classification, outputs the predicted entity label probability distribution of multiple target entity embeddings, and calculates the token contrast loss, the regularization loss, the prototype classification loss, and the entity contrast loss, including:
[0017] Convert the sentence sample into a semantic token embedding sequence through the semantic BERT embedding layer;
[0018] Determine the semantic entity embeddings corresponding to the semantic token embedding sequence based on the entity boundary;
[0019] Use an entity filtering layer to remove the invalid entity embeddings in each semantic entity embedding to determine the target entity embeddings;
[0020] Through the topic prototype network, calculate the prototype embeddings based on the entity embedding samples of the sentence support set, and determine the predicted entity label probability distribution of each target entity embedding according to each prototype embedding using the temperature-scaled cosine similarity function;
[0021] Calculate the token contrast loss based on the KL divergence from the semantic token embedding sequence;
[0022] Calculate the regularization loss using each of the prototype embeddings, calculate the prototype classification loss based on the predicted entity label probability distribution, and calculate the entity contrast loss using each of the target entity embeddings based on the KL divergence.
[0023] Further, the boundary detection classification loss includes:
[0024] ;
[0025] The boundary-aware contrast loss includes:
[0026] ;
[0027] ;
[0028] ;
[0029] In the formula, is the boundary detection classification loss, is the token index, is the total number of tokens, is the cross entropy, is the position label ground truth, is the th token, is the position label probability distribution, is the boundary-aware contrast loss, is the sigmoid function, is the entity embedding, is the positive sample representation, is the negative sample representation, is the L2 norm.
[0030] Further, the process of determining the target entity embedding includes:
[0031] ;
[0032] ;
[0033] In the formula, is the target entity embedding, is the entity embedding, is the set of semantic entity embeddings, is the th entity set of the entity class label, is the start of the entity, is the end of the entity, For the start of an entity different from For the end of an entity different from For For the end of an entity different from Is the dynamic threshold Is the semantic token embedding with [CLS] added Is the cosine similarity with a scaling factor
[0034] Furthermore, the process of determining the entity embedding includes:
[0035] ;
[0036] In the formula, Is the entity embedding Is the start of the entity Is the end of the entity Is the linear layer Is the entity embedding of the start of the entity Is the entity embedding of the end of the entity Is the concatenation of vectors Is the Row of the learnable matrix
[0037] A few-shot information extraction system provided by the second aspect of the present invention includes:
[0038] A data acquisition module for inputting sentence samples of a sentence support set into an initial named entity recognition model for model training, where the initial named entity recognition model includes an entity boundary detector and an entity classifier;
[0039] A boundary detection module for performing boundary detection on the sentence samples through the entity boundary detector, outputting entity boundaries, and calculating a boundary detection classification loss and a boundary-aware contrast loss;
[0040] An entity classification module for performing entity classification on the sentence samples and the entity boundaries through the entity classifier, outputting a predicted entity label probability distribution of multiple target entity embeddings, and calculating a token contrast loss, a regularization loss, a prototype classification loss, and an entity contrast loss;
[0041] A model optimization module for iteratively optimizing the model parameters of the initial named entity recognition model based on the boundary detection classification loss, the boundary-aware contrast loss, the token contrast loss, the regularization loss, the prototype classification loss, and the entity contrast loss, and verifying with a sentence query set to determine the target named entity recognition model;
[0042] A data recognition module, configured to input a sentence to be recognized into the target named entity recognition model for boundary detection and entity classification, and output a target recognition result.
[0043] A computer device provided in the third aspect of the present invention includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the few-shot information extraction method described in any one of the above.
[0044] A computer-readable storage medium provided in the fourth aspect of the present invention has a computer program stored thereon. When the computer program is executed, the few-shot information extraction method described in any one of the above is implemented.
[0045] A computer program product provided in the fifth aspect of the present invention includes a computer program / instructions. When the computer program / instructions are executed by a processor, the few-shot information extraction method described in any one of the above is implemented.
[0046] It can be seen from the above technical solutions that the present invention has the following advantages:
[0047] The above solution of the present invention provides a few-shot information extraction method, including: using sentence samples of a sentence support set to input an initial named entity recognition model for model training. The initial named entity recognition model includes an entity boundary detector and an entity classifier; performing boundary detection on the sentence samples through the entity boundary detector, outputting entity boundaries, and calculating a boundary detection classification loss and a boundary-aware contrast loss; performing entity classification on the sentence samples and the entity boundaries through the entity classifier, outputting a predicted entity label probability distribution of multiple target entity embeddings, and calculating a token contrast loss, a regularization loss, a prototype classification loss, and an entity contrast loss; iteratively optimizing the model parameters of the initial named entity recognition model based on the boundary detection classification loss, the boundary-aware contrast loss, the token contrast loss, the regularization loss, the prototype classification loss, and the entity contrast loss, and using a sentence query set for verification to determine the target named entity recognition model; inputting the sentence to be recognized into the target named entity recognition model for boundary detection and entity classification, and outputting a target recognition result. Based on the above solution, introducing the boundary-aware contrast loss improves the boundary perception ability of the entity boundary detector, using the token contrast loss helps the designed entity filtering layer filter out invalid entities, and introducing the entity contrast loss effectively mines the implicit information between entities, which can better optimize the performance of the topic prototype network and make full use of limited data, and helps to improve the accuracy of few-shot named entity recognition. Description of the Drawings
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 It is a flowchart of the steps of a few-shot information extraction method provided by an embodiment of the present invention;
[0050] Figure 2 It is an overall architecture diagram of a named entity recognition model provided by an embodiment of the present invention;
[0051] Figure 3 It is a structural block diagram of a few-shot information extraction system provided by an embodiment of the present invention. Detailed implementation manners
[0052] The embodiments of the present invention provide a few-shot information extraction method and system, which are used to improve the technical problem that the entity recognition accuracy of the existing few-shot named entity recognition method is relatively low.
[0053] In order to make the invention purpose, features, and advantages of the present invention more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0054] Please refer to Figure 1 , Figure 1 It is a flowchart of the steps of a few-shot information extraction method provided by an embodiment of the present invention.
[0055] A few-shot information extraction method provided by the present invention includes:
[0056] Step 101: Input sentence samples of a sentence support set into an initial named entity recognition model for model training. The initial named entity recognition model includes an entity boundary detector and an entity classifier.
[0057] It should be noted that, for named entity recognition, an embodiment of the present invention designs a named entity recognition model, as Figure 2 shown, including an entity boundary detector and an entity classifier, and conducts model training to optimize model parameters. During model training, a subset sampled from a sentence training set and the initial model parameters of the model , where represents the sentence support set represents the sentence query set represents the set of entity categories to be classified. The sentence training set contains multiple sentence samples for each entity category. First, the sentence samples in the sentence support set are input into the initial named entity recognition model to facilitate the internal update of the parameters using gradient descent:
[0058]
[0059] In the formula is the learning rate of the internal update is to calculate the gradient is the loss of the model is the forward propagation process of the model is the th training subset in the support set are the model parameters after internal update
[0060] Step 102: Perform boundary detection on the sentence samples through an entity boundary detector, output the entity boundaries, and calculate the boundary detection classification loss and the boundary-aware contrast loss
[0061] Step 102 includes the following sub-steps
[0062] Convert the sentence samples into a position token embedding sequence based on the boundary BERT embedding layer
[0063] Use a linear layer to map the position token embedding sequence to the BIOES tag space, and calculate the predicted position label probability distribution of each position token embedding based on the softmax function
[0064] Through the conditional random field of Viterbi decoding, determine the target position label of each position token embedding according to each predicted position label probability distribution, and use each target position label to determine the entity boundary
[0065] Calculate the boundary detection classification loss based on each predicted position label probability distribution, and calculate the corresponding boundary-aware contrast loss according to the position entity embedding, positive sample representation, and negative sample representation of each entity boundary
[0066] It should be noted that the purpose of the entity boundary detector is to find all named entities in the input sentence by locating the entity boundaries. In this embodiment, boundary detection is regarded as a sequence labeling problem, and the "BIOES" token scheme is used to classify each token in the sequence. "B, I, O, E, S" represent "Begin", "Inside", "Outside", "End", and "Single" of the entity respectively, and the token labels are converted into the boundaries of the entity. The specific implementation process in the boundary detection stage includes:
[0067] (1) Position embedding learning
[0068] Given a sentence sample with a total number of tokens , this embodiment uses boundary BERT as the embedding layer to learn the context embeddings of the tokens in the sentence, which can be expressed by the formula:
[0069]
[0070] Among them, is the position token embedding sequence composed of multiple position token embeddings output by the last hidden layer of the boundary BERT embedding layer, is the boundary BERT embedding layer, is the th position token embedding, is the position token embedding with [CLS] added;
[0071] (2) Position classification
[0072] After obtaining the position token embedding sequence, this embodiment maps it to the BIOES label space through a linear layer Linear and uses a softmax function to calculate the position label probability distribution:
[0073]
[0074] Among them, the position label probability distribution , the position label set ;
[0075] To prevent the predicted position labels from violating the BIOES annotation rules, this embodiment introduces a conditional random field CRF to calculate the emission probability and transition probability of the labels, and uses the Viterbi decoding algorithm when inferring the position labels. Thus, the final target position labels can be expressed as:
[0076]
[0077] Among them, is the token index, Indicates the th token corresponding target position label; for ease of subsequent calculation, the BIOES position label is converted into an entity boundary, and each pair of entity boundaries corresponds to an entity;
[0078] (3) Loss function
[0079] To improve the entity boundary detector's perception ability of entity boundaries, this embodiment proposes to use a boundary-aware contrastive learning method to assist in training the entity detector; assume and are the start and end of the entity respectively, then the positive sample representation corresponding to this entity boundary can be expressed as , and the negative sample representation can be expressed as , and the entity embedding of the position entity embedding can be calculated by the following formula:
[0080]
[0081] where, is the entity embedding, and are the start and end of the entity respectively, is a learnable linear layer, is the concatenation of vectors, is a learnable matrix 's th row, is the dimension;
[0082] Then, the boundary-aware contrastive loss can be expressed as:
[0083]
[0084]
[0085]
[0086] where, is the sigmoid function, is the cosine similarity with a scaling coefficient , is the L2 norm;
[0087] The boundary detection classification loss of the entity boundary detector can be obtained by calculating the cross-entropy of the predicted position label probability distribution and the position label ground truth :
[0088]
[0089] Therefore, the total loss of boundary detection in this stage is:
[0090]
[0091] Step 103: Use an entity classifier to perform entity classification on sentence samples and entity boundaries, output the predicted entity label probability distribution of multiple target entity embeddings, and calculate the token contrast loss, regularization loss, prototype classification loss, and entity contrast loss.
[0092] Step 103 includes the following sub-steps:
[0093] Convert the sentence sample into a semantic token embedding sequence through the semantic BERT embedding layer;
[0094] Determine the semantic entity embedding corresponding to the semantic token embedding sequence based on the entity boundary;
[0095] Use an entity filtering layer to remove invalid entity embeddings in each semantic entity embedding to determine the target entity embedding;
[0096] Through the topic prototype network, calculate the prototype embedding based on the entity embedding samples of the sentence support set, and determine the predicted entity label probability distribution of each target entity embedding according to each prototype embedding using the temperature-scaled cosine similarity function;
[0097] Calculate the token contrast loss based on the KL divergence for the semantic token embedding sequence;
[0098] Calculate the regularization loss using each prototype embedding, calculate the prototype classification loss based on the predicted entity label probability distribution, and calculate the entity contrast loss using each target entity embedding based on the KL divergence.
[0099] It should be noted that the entity classifier is designed to filter and classify the entities obtained in the entity boundary detector. After obtaining the boundary information of the entities through entity boundary detection, it converts the semantic token embeddings learned by the semantic BERT embedding layer into semantic entity embeddings, filters the boundaries of non-entities or entities that do not belong to predefined categories through the entity filtering layer designed in this embodiment, and finally uses the topic prototype network to classify the remaining entities; the specific implementation process in the entity classification stage includes:
[0100] (1) Semantic embedding learning
[0101] Similar to the entity boundary detector, an encoder is first used in the entity classifier to learn the context representation embedding of the tokens. However, since the encoder in this stage pays more attention to the semantic information of the tokens themselves rather than the position information, another BERT, namely the semantic BERT embedding layer, is fine-tuned as the encoder in this stage; the context embedding of the semantic tokens can be expressed as:
[0102]
[0103] Among them, is the semantic token embedding sequence output by the last hidden layer of the semantic BERT embedding layer, is the semantic BERT embedding layer, is the th semantic token embedding, is the semantic token embedding with [CLS] added;
[0104] The entity filtering layer is used to remove the invalid entity embeddings in each semantic entity embedding to determine the target entity embedding;
[0105] (2)Entity filtering
[0106] Combined with the entity boundaries output by the entity boundary detector, the semantic token embedding sequence is transformed into entity embeddings according to the process of Equation 5 to obtain the corresponding semantic entity embeddings;
[0107] In order to remove the invalid entity boundaries, in the entity filtering layer, calculate the entity embedding used as the semantic entity embedding and the similarity between each entity embedding sample in the support set. If the maximum similarity between the semantic entity embedding and the entity embedding samples of each entity category label in the entity category set in the support set is less than the similarity threshold , then this entity will be regarded as an invalid entity and removed:
[0108]
[0109] In the formula, is the remaining semantic entity embedding after filtering, that is, the target entity embedding, is the entity embedding, is the set of semantic entity embeddings, is the th entity set of the entity category label, and are the start and end of entities different from and , is the dynamic threshold, which is calculated by the following formula:
[0110]
[0111] (3)Prototype classification
[0112] After removing the noise entities detected in the first stage, the remaining entities are processed using the topic prototype network Classify; the topic prototype network extends the metric-based few-shot learning principle to the field of topic modeling, establishing a framework that can quickly adapt to new topics while maintaining semantic coherence. It operates in a specialized metric space where topic prototype representations can be effectively compared and classified;
[0113] The prototype representation mechanism implements a non-linear projection function that maps features into a shared embedding space and preserves semantic relationships in this space; for the entity embedding samples given in the sentence support set , the entity embedding projection is calculated through a series of transformations:
[0114]
[0115] where , , and are learnable first weight parameter, second weight parameter, first bias parameter and second bias parameter, is the ReLU function; this projection representation ensures that entities with similar semantic content are clustered together in the topic embedding space; the prototype embedding of each entity category is calculated as the weighted average of the support set examples:
[0116]
[0117] The attention weight is determined by a learned similarity function that considers feature similarity and topic coherence:
[0118]
[0119] where is a balance factor for balancing these two factors, is an evaluation factor for evaluating topic coherence, which can be obtained by calculating the average of the cosine similarities of all entity representations of the same topic (category) in the support set:
[0120]
[0121] Based on the similarity between the entity and the prototype mentioned above, in this embodiment, the final entity label probability distribution is calculated by the temperature-scaled cosine similarity function :
[0122]
[0123] where is the softmax function, is a learnable temperature parameter used to control the sharpness of the probability distribution, ensuring that the model can make confident predictions in appropriate situations while maintaining uncertainty in ambiguous situations;
[0124] It can be understood that during the inference phase, the entity embeddings that are most similar to the semantic entity embeddings The label corresponding to the prototype will be used as the target entity label for prediction :
[0125]
[0126] (4) Loss function
[0127] To help the entity filtering layer filter out misdetected entities, this embodiment considers using token-level contrastive learning:
[0128] Assume that the semantic token embeddings obtained from the semantic BERT embedding layer follow a Gaussian distribution. Consider a token and its semantic embedding , first use two projection layers to generate the parameters of its corresponding Gaussian distribution:
[0129]
[0130] where, and represent the mean and diagonal covariance respectively, and are the mean multi-layer perceptron and the diagonal covariance multi-layer perceptron respectively, represents the exponential linear unit, is a constant used for stable calculation;
[0131] Since the invalid entities detected by the entity boundary detector are usually caused by non-entity tokens, this embodiment uses contrastive learning loss to separate them in the Gaussian embedding space; given two tokens and and their Gaussian embeddings and , use the token KL divergence between these two Gaussian distributions as the distance metric between the two tokens, as follows:
[0132]
[0133] is the trace of the matrix, is the dimension of the Gaussian distribution; since the KL divergence is one-way, the bidirectional token KL divergence in both directions is calculated here:
[0134]
[0135] In the formula, is the one-way token KL divergence of the th token embedding pair with respect to the th token embedding, is the one-way token KL divergence of the th token embedding pair with respect to the th token embedding;
[0136] In the training batch , if the labels of two tokens are the same, they are regarded as a pair of positive samples. On the contrary, two tokens with different labels are regarded as a pair of negative samples; In particular, for the token , its set of token positive samples and set of token negative samples can be expressed as:
[0137]
[0138]
[0139] Then, the token contrast loss of all tokens in this training batch can be calculated by the following formula:
[0140]
[0141] Through the above token contrast loss, the entity filtering layer can increase the distance between entity tokens and non-entity tokens, making it easier to separate non-entity tokens;
[0142] In addition, this embodiment adopts an additional regularization term to prevent prototype collapse and maintain diversity:
[0143]
[0144] Among them, is the regularization loss, is the th entity category label, is the square of the norm;
[0145] The prototype classification loss of the topic prototype network is calculated by the following formula:
[0146]
[0147] Among them, is the true value of the entity label probability distribution;
[0148] Furthermore, to optimize the distribution of semantic entity embeddings and make them closer to the corresponding prototypes, this embodiment uses contrastive loss at the entity level to train the topic prototype network; similar to the token contrastive loss, first calculate the Gaussian embedding of the semantic entity embedding :
[0149]
[0150] where is the mean of the th token embedding, is the diagonal covariance of the th token embedding, is the mean of the entity embedding, is the diagonal covariance of the entity embedding;
[0151] Similar to the method of calculating the KL divergence between token Gaussian embeddings, given two entities and (denoted by and respectively, and are both entity indices) and their Gaussian embeddings and , the bidirectional entity KL divergence between the th entity embedding and the th entity embedding is calculated by the following formula: In the formula,
[0152]
[0153]
[0154] where is the unidirectional entity KL divergence of the th entity embedding with respect to the th entity embedding, is the unidirectional entity KL divergence of the th entity embedding with respect to the th entity embedding;
[0155] Then, the entity positive sample set and the entity negative sample set of the target entity embedding are:
[0156]
[0157]
[0158] Finally, the entity contrastive loss Obtained by the following formula:
[0159]
[0160] The overall loss of entity classification in this stage Can be expressed as:
[0161]
[0162] Step 104: Iteratively optimize the model parameters of the initial named entity recognition model based on the boundary detection classification loss, boundary-aware contrast loss, token contrast loss, regularization loss, prototype classification loss, and entity contrast loss, and use the sentence query set for verification to determine the target named entity recognition model.
[0163] It should be noted that in order to enable the model to learn a good parameter initialization method to help the model quickly adapt to a new domain, this embodiment uses the method of model-agnostic meta-learning to train the models in two stages; according to the losses calculated in the two stages of boundary detection and entity classification, after iteratively optimizing the initial named entity recognition model by gradient descent, use the sentence query set Verify the internally updated model parameters And perform meta-update:
[0164]
[0165] Among them, Is the learning rate of meta-update, Is the sentence query set In the th subset, in Equation (33), the parameter tuning performed through the gradient update step requires calculating the second-order gradient, and here its first-order approximation is used to accelerate the backpropagation process:
[0166]
[0167] Step 105: Input the sentence to be recognized into the target named entity recognition model for boundary detection and entity classification, and output the target recognition result.
[0168] It should be noted that after determining the target named entity recognition model that has completed training, when receiving the sentence to be recognized for named entity recognition, input it into the target named entity recognition model for boundary detection and entity classification, and output the target recognition result; it can be understood that the target recognition result can include the target entity label probability distribution shown in Equation (18), or can further include the target entity label shown in Equation (19).
[0169] To verify the effectiveness of the method proposed in this embodiment, experiments were conducted on two commonly used few-shot named entity recognition datasets, Few-NERD and CrossNER. The experimental results are shown in Tables 1 and 2 as follows:
[0170] Table 1 Experimental Results on Few-NERD Dataset
[0171]
[0172] Table 2 Experimental Results on CrossNER Dataset
[0173]
[0174] As can be seen from Tables 1 and 2, on both datasets, the method of this embodiment has better performance compared with the existing few-shot named entity recognition methods.
[0175] In the embodiment of the present invention, a boundary-aware contrast loss is introduced to improve the boundary awareness ability of the entity boundary detector, a token contrast loss is adopted to help the designed entity filtering layer filter invalid entities, and an entity contrast loss is introduced to effectively mine the implicit information between entities, which can better optimize the performance of the topic prototype network and make full use of limited data. In addition, by combining the interpretability and statistical rigor of traditional topic models with the representation ability of the prototype network, rich prototype representations are constructed, enabling the model to make entity recognition decisions based on local syntactic patterns and global topic contexts, which helps to improve the accuracy of few-shot named entity recognition. The method proposed in this embodiment also has certain reference significance for similar tasks in other natural language processing fields.
[0176] Please refer to Figure 3 , Figure 3 which is a structural block diagram of a few-shot information extraction system provided by an embodiment of the present invention.
[0177] A few-shot information extraction system provided by the present invention includes:
[0178] A data acquisition module 301, configured to input sentence samples of a sentence support set into an initial named entity recognition model for model training. The initial named entity recognition model includes an entity boundary detector and an entity classifier;
[0179] A boundary detection module 302, configured to perform boundary detection on the sentence samples through the entity boundary detector, output entity boundaries, and calculate a boundary detection classification loss and a boundary-aware contrast loss;
[0180] The entity classification module 303 is used to perform entity classification on sentence samples and entity boundaries through an entity classifier, output the predicted entity label probability distribution of multiple target entity embeddings, and calculate the token contrast loss, regularization loss, prototype classification loss, and entity contrast loss;
[0181] The model optimization module 304 is used to iteratively optimize the model parameters of the initial named entity recognition model based on the boundary detection classification loss, boundary-aware contrast loss, token contrast loss, regularization loss, prototype classification loss, and entity contrast loss, and use the sentence query set for verification to determine the target named entity recognition model;
[0182] The data recognition module 305 is used to input the sentence to be recognized into the target named entity recognition model for boundary detection and entity classification, and output the target recognition result.
[0183] Furthermore, the boundary detection module 302 is specifically used for:
[0184] Convert the sentence sample into a position token embedding sequence based on the boundary BERT embedding layer;
[0185] Use a linear layer to map the position token embedding sequence to the BIOES tag space, and calculate the predicted position label probability distribution of each position token embedding based on the softmax function;
[0186] Through the conditional random field of Viterbi decoding, determine the target position label of each position token embedding according to each predicted position label probability distribution, and use each target position label to determine the entity boundary;
[0187] Calculate the boundary detection classification loss based on each predicted label probability distribution, and calculate the corresponding boundary-aware contrast loss according to the position entity embedding, positive sample representation, and negative sample representation of each entity boundary.
[0188] Furthermore, the entity classification module 303 is specifically used for:
[0189] Convert the sentence sample into a semantic token embedding sequence through the semantic BERT embedding layer;
[0190] Determine the semantic entity embedding corresponding to the semantic token embedding sequence based on the entity boundary;
[0191] Use an entity filtering layer to remove the invalid entity embeddings in each semantic entity embedding to determine the target entity embedding;
[0192] Through the topic prototype network, calculate the prototype embedding based on the entity embedding samples of the sentence support set, and determine the predicted entity label probability distribution of each target entity embedding according to each prototype embedding using the temperature-scaled cosine similarity function;
[0193] Calculate the token contrast loss based on the KL divergence for the semantic token embedding sequence;
[0194] Calculate the regularization loss using each prototype embedding, calculate the prototype classification loss based on the predicted entity label probability distribution, and calculate the entity contrast loss using each target entity embedding based on the KL divergence.
[0195] Furthermore, the boundary detection classification loss includes:
[0196] ;
[0197] The boundary-aware contrast loss includes:
[0198] ;
[0199] ;
[0200] ;
[0201] In the formula, is the boundary detection classification loss, is the token index, is the total number of tokens, is the cross entropy, is the position label ground truth, is the th token, is the position label probability distribution, is the boundary-aware contrast loss, is the sigmoid function, is the entity embedding, is the positive sample representation, is the negative sample representation, is the L2 norm.
[0202] Furthermore, the process of determining the target entity embedding includes:
[0203] ;
[0204] ;
[0205] In the formula, is the target entity embedding, is the entity embedding, is the set of semantic entity embeddings, is the th entity set of the entity class label, is the start of the entity, is the end of the entity, is the start of an entity different from the For the end of an entity different from the end of different entities, is a dynamic threshold, is the semantic token embedding with [CLS] added, is the cosine similarity with a scaling factor.
[0206] Furthermore, the process of determining entity embeddings includes:
[0207] ;
[0208] In the formula, is the entity embedding, is the start of the entity, is the end of the entity, is a linear layer, is the entity embedding of the start of the entity, is the entity embedding of the end of the entity, is the concatenation of vectors, is the th row of the learnable matrix.
[0209] An embodiment of the present invention also provides a computer device, including a memory and a processor, and a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the few-shot information extraction method in any of the above embodiments.
[0210] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by the processor, the steps of the few-shot information extraction method in any of the above embodiments are implemented.
[0211] An embodiment of the present invention also provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the few-shot information extraction method in any of the above embodiments are implemented.
[0212] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0213] In several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0214] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0215] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0216] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.
[0217] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for extracting information from a small number of samples, characterized in that: include: Using sentence samples of the sentence support set to input an initial named entity recognition model for model training, the initial named entity recognition model includes an entity boundary detector and an entity classifier; Performing boundary detection on the sentence sample by using the entity boundary detector, outputting entity boundaries, and calculating boundary detection classification loss and boundary perception contrast loss; Performing entity classification using the sentence sample and the entity boundary through the entity classifier, outputting a predicted entity label probability distribution of multiple target entity embeddings, and calculating a token contrast loss, a regularization loss, a prototype classification loss, and an entity contrast loss; Iteratively optimizing model parameters of the initial named entity recognition model based on the boundary detection classification loss, the boundary perception contrast loss, the token contrast loss, the regularization loss, the prototype classification loss, and the entity contrast loss, and verifying it using a sentence query set to determine a target named entity recognition model; The sentence to be recognized is input into the target named entity recognition model for boundary detection and entity classification, and the target recognition result is output.
2. The method for extracting information from a small number of samples according to claim 1, characterized in that: The performing boundary detection on the sentence sample by the entity boundary detector, outputting target position labels for multiple token embeddings and calculating boundary detection classification loss comprises: Converting the sentence sample into a position token embedding sequence based on a boundary BERT embedding layer; A linear layer is used to map the position token embedding sequence to the BIOES tag space, and a softmax function is used to calculate the predicted position label probability distribution of each position token embedding; Determine the target position label for embedding each position token according to the probability distribution of each predicted position label through the conditional random field of Viterbi decoding, and determine the entity boundary using each target position label; The boundary detection classification loss is calculated based on the probability distribution of each predicted label, and the corresponding boundary perception contrast loss is calculated according to the position entity embedding, positive sample representation and negative sample representation of each entity boundary.
3. The method for extracting information from a small number of samples according to claim 1, characterized in that: The entity classifier uses the sentence sample and the entity boundary to perform entity classification, outputs a predicted entity label probability distribution of multiple target entity embeddings, and calculates token contrast loss, regularization loss, prototype classification loss and entity contrast loss, including: Convert the sentence sample into a semantic token embedding sequence through a semantic BERT embedding layer; Determining a semantic entity embedding corresponding to the semantic token embedding sequence based on the entity boundary; Using an entity filtering layer to remove invalid entity embeddings in each of the semantic entity embeddings to determine a target entity embedding; Calculating prototype embeddings based on entity embedding samples of the sentence support set through a topic prototype network, and determining a predicted entity label probability distribution of each target entity embedding using a temperature-scaled cosine similarity function based on each prototype embedding; Calculating a token contrastive loss based on the semantic token embedding sequence according to KL divergence; The regularization loss is calculated using each of the prototype embeddings, the prototype classification loss is calculated based on the predicted entity label probability distribution, and the entity contrast loss is calculated using each of the target entity embeddings based on the KL divergence.
4. The method for extracting information from a small number of samples according to claim 2, characterized in that: Boundary detection classification loss, including: ; Boundary-aware contrast loss, including: ; ; ; In the formula, is the boundary detection classification loss, is the token index, is the total number of tokens, is the cross entropy, is the true value of the position label, For the Tokens, is the probability distribution of position labels, is the boundary-aware contrast loss, is the sigmoid function, is the entity embedding, is the positive sample representation, is the negative sample representation, is the L2 norm.
5. The method for extracting information from a small number of samples according to claim 3, characterized in that: The process of determining the target entity embedding includes: ; ; In the formula, is the target entity embedding, is the entity embedding, is the set of semantic entity embeddings, For the The entity set with entity category labels, is the beginning of the entity, is the end of the entity, For The beginning of different entities, For The end of different entities, is the dynamic threshold, To add semantic token embedding for [CLS], is the cosine similarity with a scaling factor.
6. The method for extracting information from a small number of samples according to claim 2 or 3, characterized in that: The process of determining entity embedding includes: ; In the formula, is the entity embedding, is the beginning of the entity, is the end of the entity, is a linear layer, The entity embedding for the start of the entity, The entity embedding at the end of the entity, is the concatenation of vectors, is the learnable matrix OK.
7. A small sample information extraction system, characterized in that: include: A data acquisition module, used to input sentence samples of a sentence support set into an initial named entity recognition model for model training, wherein the initial named entity recognition model includes an entity boundary detector and an entity classifier; A boundary detection module, used to perform boundary detection on the sentence sample by using the entity boundary detector, output entity boundaries, and calculate boundary detection classification loss and boundary perception contrast loss; An entity classification module, configured to perform entity classification using the sentence sample and the entity boundary through the entity classifier, output a predicted entity label probability distribution of multiple target entity embeddings, and calculate a token contrast loss, a regularization loss, a prototype classification loss, and an entity contrast loss; A model optimization module, configured to iteratively optimize model parameters of the initial named entity recognition model based on the boundary detection classification loss, the boundary perception contrast loss, the token contrast loss, the regularization loss, the prototype classification loss, and the entity contrast loss, and to verify the model using a sentence query set to determine a target named entity recognition model; The data recognition module is used to input the sentence to be recognized into the target named entity recognition model for boundary detection and entity classification, and output the target recognition result.
8. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the few sample information extraction method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method for extracting information from a small number of samples as described in any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method for extracting information from a small number of samples as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Named entity extraction method and device
CN114048746A
Small sample named entity identification method for entity level information enhanced prototype representation
CN116451691A
Entity recognition method, system and equipment based on word vector fusion and medium
CN117669569A
Named entity recognition system and named entity recognition method
US20230205998A1