Small-Sample Entity Recognition Method and System Based on Pseudo-Labels and Cross-Validation
Through the improved K-fold cross-validation and semi-supervised algorithm combined with the R-Drop method of dynamic parameters, the unlabeled data is used to identify small sample entities, which solves the problem of poor performance of small sample annotation data in the existing technology, and achieves higher recognition accuracy and generalization capabilities.
Patent Information
- Application Number
- CN202411855042.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-17
AI Technical Summary
Existing named entity recognition technology is not effective in small sample labeling data, it is difficult to make full use of unlabeled data, and it is impossible to effectively identify small sample entities.
A small sample entity recognition method based on pseudo-label and cross-validation is adopted. Through an improved K-fold cross-validation algorithm and semi-supervised algorithm, combined with the R-Drop method of dynamic parameters, the unlabeled data is used for model training and fine-tuning, and the model generalization ability is improved.
It improves the accuracy and generalization ability of small sample entity recognition, makes full use of labeled data, enhances the semantic understanding ability of the model, and reduces the risk of model overfitting.
Smart Images

Figure CN119740580B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of few-shot entity recognition, and particularly relates to a few-shot entity recognition method and system based on pseudo-labels and cross-validation. Background Art
[0002] Named entity recognition is an important upstream task in the field of natural language processing, and its goal is to identify entities with specific meanings from text, including people, locations, organizations, etc. In recent years, named entity recognition technology based on deep learning models has achieved outstanding results, but there are still problems such as relying on large-scale labeled data and being unable to incrementally learn new types, making most existing models and methods difficult to meet actual needs. The current named entity recognition methods are generally supervised fine-tuning training on labeled data based on open-source models to obtain a model that can solve the task. However, truly high-quality labeled data is often extremely scarce, while unlabeled data is easier to obtain. Such methods are powerless for few-shot labeled data, making full use of labeled data, and using high-quality unlabeled data, and thus cannot effectively obtain high-quality results in few-shot entity recognition. Therefore, the present invention proposes a few-shot entity recognition method and system based on pseudo-labels and cross-validation. Summary of the Invention
[0003] The present invention provides a few-shot entity recognition method and system based on pseudo-labels and cross-validation to solve the above problems.
[0004] The present invention provides a few-shot entity recognition method based on pseudo-labels and cross-validation, including:
[0005] Processing the labeled data based on an improved K-fold cross-validation algorithm to obtain a training set and a test set, training and validating the baseline model UIE to obtain a fine-tuned model;
[0006] Randomly and batch-selecting and processing the unlabeled data in the fine-tuned model based on preset rules to obtain the selected pseudo-data;
[0007] Training the baseline model UIE using a semi-supervised algorithm and an improved K-fold cross-validation algorithm based on the selected pseudo-data and the labeled data to obtain a final model;
[0008] Processing the input text based on the final model to obtain an entity recognition result.
[0009] Preferably, in a few-shot entity recognition method based on pseudo-labels and cross-validation, randomly and batch-selecting and processing the unlabeled data in the fine-tuned model based on preset rules to obtain the selected pseudo-data includes:
[0010] Batch select the unlabeled data according to the preset step size and the sliding window size to obtain a batch of unlabeled data;
[0011] Predict the batch of unlabeled data based on the fine-tuned model to obtain primary pseudo data;
[0012] Mix the primary pseudo data and the labeled data to obtain training data;
[0013] Train the baseline model UIE based on the training data to obtain a test model;
[0014] Test the test data through the test model to obtain test results;
[0015] Obtain the test scores corresponding to each test result, and regard the data with the highest test score as pseudo data.
[0016] Preferably, in a small sample entity recognition method based on pseudo labels and cross-validation, the improved K-fold cross-validation algorithm specifically includes:
[0017] Based on the improved K-fold cross-validation algorithm, set K loops and divide the labeled data into K parts. At the same time, take any one of the subsets as the test set, and the remaining K - 1 subsets as the training set;
[0018] Use K - 1 parts of the K parts as the training set and 1 part as the validation set for training respectively, and finally obtain a model.
[0019] Preferably, in a small sample entity recognition method based on pseudo labels and cross-validation, based on the selected pseudo data and the labeled data, use the semi-supervised algorithm and the improved K-fold cross-validation algorithm to train the baseline model UIE to obtain the final model, including:
[0020] Obtain semi-supervised data based on the selected pseudo data and the real data, and divide the semi-supervised data into K parts based on the improved K-fold cross-validation algorithm to obtain a model training set and a model validation set;
[0021] Based on the model training set and the model validation set, use the semi-supervised algorithm to train the baseline model UIE to obtain the final model.
[0022] Preferably, in a small sample entity recognition method based on pseudo labels and cross-validation, use the semi-supervised algorithm to train the baseline model UIE to obtain the final model, including:
[0023] Introduce a dynamic parameter α into the loss function of the semi-supervised algorithm to adjust the semi-supervised loss of the selected pseudo data and the real data in the data set. The specific formula of the semi-supervised loss function is:
[0024] Losslabel = α * BCELoss(label(Fake)) + (1 - α) * BCELoss(label(True))
[0025] Among them, Loss label represents the semi-supervised loss, α is a dynamic parameter, and its value range is (0, 1); BCEloss(label(Fake)) represents the binary cross-entropy loss of the fake data, and BCE(logits, label(True)) represents the binary cross-entropy loss of the labeled data; label(Fake) represents the fake data label; label(True) represents the labeled data label;
[0026] At the output end of the baseline model UIE, add the dynamic parameter R-Drop algorithm to perform two Dropouts on the model output, and obtain the KL divergence between the two Dropout processes:
[0027]
[0028] Among them, Loss R-Drop represents the KL divergence between the two Dropout processes; P1 w (y|x) represents the sub-model corresponding to the first Dropout process; P2 w (y|x) represents the sub-model corresponding to the second Dropout process; represents the difference distance between the probability distribution of the first Dropout process and the probability distribution of the second Dropout process; represents the difference distance between the probability distribution of the second Dropout process and the probability distribution of the first Dropout process;
[0029] Generate the total loss function Loss for training the semi-supervised algorithm based on the dynamic parameter β, the semi-supervised loss, and the R-Drop loss 总 :
[0030] Loss 总 = Loss label + β(Loss R-Drop,start + Loss R-Drop,end )
[0031] Among them, β is a dynamic parameter, and its value range is (0, 1]; Loss R-Drop,start represents the KL divergence at the beginning of the entity; Loss R-Drop,end represents the KL divergence at the end of the entity; Loss R-Drop,start + Loss R-Drop,end represents the R-Drop loss;
[0032] Based on the total loss function Loss 总 Adjust the model parameters, and combine the model training set and the model validation set to perform semi-supervised training on the baseline model UIE to obtain the final model.
[0033] Preferably, in a small-sample entity recognition method based on pseudo-labels and cross-validation, the input text is processed based on the final model to obtain entity recognition results, including:
[0034] Encode the input text based on the fine-tuning model, and extract the text features of the input text to obtain word vectors;
[0035] Based on the final model, pass through the Linear layer, then pass through the activation function layer Sigmoid to obtain the probability of each word vector, and then through the set threshold comparison to obtain the start and end of the entity segment;
[0036] Extract the entity segment according to the start position, end position and category label, and output it as the entity recognition result.
[0037] Preferably, in a small-sample entity recognition method based on pseudo-labels and cross-validation, it further includes:
[0038] Add a domain recognition model to the input layer of the baseline model UIE, perform domain recognition on the input text based on the domain recognition model, and determine the description domain of the text used;
[0039] And after performing domain annotation on the input text based on the description domain, perform entity recognition processing.
[0040] Preferably, in a small-sample entity recognition method based on pseudo-labels and cross-validation, add a domain recognition model to the input layer of the baseline model UIE, perform domain recognition on the input text based on the domain recognition model, and determine the description domain of the text used, including:
[0041] Based on the domain recognition model, recognize the input text to determine the positions of the text punctuation marks. Based on the text punctuation mark positions, divide the input text into multiple text segments, and perform keyword marking on each text segment;
[0042] Generate corresponding input text vectors for each text segment respectively, compare each text vector with the database text vectors respectively, and obtain the best matching text vector for each text vector based on the comparison results;
[0043] According to the description domain corresponding to the best matching text vector, determine the description domain corresponding to each input text, and obtain the keyword similarity based on the sub-vector corresponding to the keyword marking result of the text segment and the keyword sub-vector of its corresponding best matching text vector;
[0044] When the keyword similarity is greater than or equal to a preset value, determine that the current text segment is a keyword segment, and obtain the description field corresponding to the keyword segment;
[0045] If the description fields corresponding to all the keyword segments are the same, determine that the description field is the description field of the current input text;
[0046] Otherwise, obtain the ratio of the total data volume corresponding to the keyword segment to the total data volume of the input text. If the ratio is greater than or equal to a preset threshold, obtain the description fields corresponding to the adjacent text segments of the input text corresponding to each keyword segment. If the description fields corresponding to the adjacent text segments are the same as the input text segment, or the description fields corresponding to the adjacent text segments are general fields, then use the keyword segment as the target text segment;
[0047] Respectively obtain the description fields corresponding to each target text segment and their corresponding first quantity ratios, and use the description field corresponding to the largest first quantity ratio as the description field of the current input text;
[0048] If the ratio is less than the preset threshold, respectively obtain the description fields corresponding to each valid text segment and their corresponding second quantity ratios, and use the description field corresponding to the largest second quantity ratio as the description field of the current input text.
[0049] Preferably, in a small-sample entity recognition method based on pseudo-labels and cross-validation, it further includes:
[0050] Based on the domain recognition model, determine the keyword marking result and the keyword segment determination result corresponding to the entity segment, and record them;
[0051] Before the entity recognition result is output, perform type detection on the type label corresponding to the entity output result based on the optional label category table corresponding to the description field, and judge whether the label is used correctly. If the label is used correctly, judge whether there is a corresponding type label for the keyword segment and its corresponding keyword. If so, output the entity output result;
[0052] Otherwise, send the entity segment to the final model for secondary entity recognition.
[0053] The present invention provides a small-sample entity recognition system based on pseudo-labels and cross-validation for performing any one of the small-sample entity recognition methods based on pseudo-labels and cross-validation, which is characterized by including:
[0054] A fine-tuning model training module, configured to process the labeled data based on an improved K-fold cross-validation algorithm to obtain a training set and a test set, and train and validate the baseline model UIE to obtain a fine-tuning model;
[0055] A data selection module, which is used to randomly batch select and process unlabeled data in a fine-tuning model based on a preset rule to obtain selected pseudo data;
[0056] A final model training module, which is used to train the baseline model UIE based on the selected pseudo data and labeled data using a semi-supervised algorithm and an improved K-fold cross-validation algorithm to obtain a final model;
[0057] An identification result acquisition module, which is used to process the input text based on the final model to obtain an entity recognition result.
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] In the selection of unlabeled data, the present invention adopts a sliding window search to select locally optimal unlabeled data for the next step of training. In the model training stage, for small-sample entity recognition, in order to make full use of the labeled data, an improved K-fold cross-validation method is designed to make full use of the labeled data. Compared with the traditional training method, the training data is increased. Compared with the traditional K-fold cross-validation, the K-fold is cyclically used to fine-tune on one model, overcoming the problem that the traditional K-fold cross-validation generates K models and trains on different K-1 training sets. For named entity recognition, different entities and entity segment threshold scores may be recognized due to different training sets, and then using their weighted average may be incorrect in the actual scenario and make the calculation of P in the macro-F1 calculation smaller (the predicted entity quantity is larger due to the prediction weighting of each model), thus making the overall index smaller. Compared with traditional semi-supervised learning, the R-Drop method based on dynamic parameters is added to enhance its semantic understanding ability, and jointly with the semi-supervised loss to form the overall loss of semi-supervised learning. The loss of R-Drop is controlled by dynamic parameters. Since the main purpose is semi-supervised learning, the dynamic parameters controlling R-Drop are in the interval (0, 1]. Compared with the traditional R-Drop, its fixed parameters are changed to dynamic parameters, enabling it to learn by itself during the model training process, enhancing the diversity of parameters, and indirectly improving the overall ability of the model.
[0060] Other features and advantages of the present invention will be described in the following description, and, in part, will be obvious from the description or learned by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.
[0061] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0062] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:
[0063] Figure 1 is a flowchart of a small-sample entity recognition method based on pseudo-labels and cross-validation according to the present invention;
[0064] Figure 2 is a flowchart of step 1 of a small-sample entity recognition method based on pseudo-labels and cross-validation according to the present invention;
[0065] Figure 3 is a flowchart of step 2 of a small-sample entity recognition method based on pseudo-labels and cross-validation according to the present invention;
[0066] Figure 4 is a schematic diagram of the dynamic parameter R-Drop algorithm according to the present invention;
[0067] Figure 5 is a flowchart of step 3 of a small-sample entity recognition method based on pseudo-labels and cross-validation according to the present invention;
[0068] Figure 6 is a structural diagram of a small-sample entity recognition system based on pseudo-labels and cross-validation according to the present invention. Detailed Embodiments
[0069] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for explaining and illustrating the present invention and are not used to limit the present invention.
[0070] Embodiment 1:
[0071] The present invention provides a small-sample entity recognition method based on pseudo-labels and cross-validation, as Figure 1 shown, including:
[0072] Step 1: Process the labeled data based on an improved K-fold cross-validation algorithm to obtain a training set and a test set, and train and validate the baseline model UIE to obtain a fine-tuned model;
[0073] Step 2: Randomly batch select and process the unlabeled data in the fine-tuned model based on a preset rule to obtain the selected pseudo-data;
[0074] Step 3: Based on the selected pseudo-data and the labeled data, use a semi-supervised algorithm and an improved K-fold cross-validation algorithm to train the baseline model UIE to obtain a final model;
[0075] Step 4: Process the input text based on the final model to obtain an entity recognition result.
[0076] In this embodiment, the preset rules include the step size and the sliding window size, where the step size is 10% of the total amount of unlabeled data, and the sliding window size is 20% of the total amount of unlabeled data. For example, the amount of unlabeled data can be 5000, and then a sliding block with a step size of 500 and a total length of 1000 is used in the form of a sliding window to select data from randomly selected batch data.
[0077] Both the labeled data and the unlabeled data used in the present invention include multiple description fields.
[0078] In this embodiment, the baseline model UIE refers to a unified model for information extraction.
[0079] Beneficial effects of the above technical solution: The present invention realizes the processing of labeled data based on an improved K-fold cross-validation algorithm, obtains a training set and a test set, trains and validates the baseline model UIE to obtain a fine-tuned model, makes full use of the labeled data, and provides a basis for the prediction of unlabeled data; based on the preset rules, randomly batch-select and process the unlabeled data in the fine-tuned model to obtain the selected pseudo data, realizing the selection of locally optimal unlabeled data, effectively improving the quality of semi-supervised data, and providing a good basis for the training of the final model; based on the selected pseudo data and the labeled data, use the semi-supervised algorithm and the improved K-fold cross-validation algorithm to train the baseline model UIE to obtain the final model, realizing the purpose of training the model with small samples of unlabeled data and at the same time using the improved K-fold cross-validation algorithm to perform fine-tuning on one model, improving the generalization ability and index performance of the small-sample entity recognition model; based on the final model, process the input text to obtain the entity recognition result, and complete the accurate recognition of small-sample entities.
[0080] Embodiment 2:
[0081] On the basis of Embodiment 1, Step 2: Based on the preset rules, randomly batch-select and process the unlabeled data in the fine-tuned model to obtain the selected pseudo data, as Figure 2 shown, including:
[0082] Step 201: Batch-select the unlabeled data according to the preset step size and sliding window size to obtain batch unlabeled data;
[0083] Step 202: Predict the batch unlabeled data based on the fine-tuned model to obtain primary pseudo data;
[0084] Step 203: Mix the primary pseudo data and the labeled data to obtain training data;
[0085] Step 204: Train the baseline model UIE based on the training data to obtain a test model;
[0086] Step 205: Test the test data through the test model until a test result is obtained;
[0087] Step 206: Repeat the above steps until all the test data is tested, and select the data with the highest score in the test set as the selected pseudo-data.
[0088] In this embodiment, the primary pseudo-data refers to the pseudo-label data of the selected pseudo-labeled data predicted based on the fine-tuning model.
[0089] In this embodiment, the score of the total test set data is the macro-F1 index score, and the calculation formula of macro-F1 is as follows:
[0090]
[0091] Among them, macro-F1 represents the macro-F1 index score of the data in the test; N represents the total number of labeled label types, and F1 i refers to the F1 value of the i-th labeled label type;
[0092] Among them, the F1 value of each labeled label type is calculated as follows;
[0093]
[0094] Among them, P i = the number of correctly predicted entity data corresponding to the i-th labeled label type in the test set / the number of entities predicted for the i-th labeled label type in the test set * 100%; R i = the number of correctly predicted entity data corresponding to the i-th labeled label type in the test set / the number of entities in the test set * 100%.
[0095] The beneficial effects of the above technical solutions: The present invention first batch-selects the unlabeled data according to the preset step size and the sliding window size to obtain a batch of unlabeled data, performs existing optimization on the unlabeled data, and predicts the batch of unlabeled data based on the fine-tuning model to obtain primary pseudo-data, realizing the prediction of pseudo-labeled data; and mixes the primary pseudo-data and the labeled data to obtain training data, realizing the full utilization of the labeled data, and then trains the baseline model UIE based on the training data to obtain a test model; tests the test data through the test model until a test result is obtained; repeats the above steps until all the test data is tested, and selects the data with the highest score in the test set as the selected pseudo-data, realizing the automatic selection of unlabeled data.
[0096] Example 3:
[0097] Based on Example 1, the improved K-fold cross-validation algorithm specifically includes:
[0098] Set K loops and divide the labeled data into K parts. At the same time, take any one of the subsets as the test set, and the remaining K - 1 subsets as the training set;
[0099] In each loop, use the training set in each loop for training and the validation set for validation, where the subsets corresponding to the test set selected each time are different.
[0100] Beneficial effects of the above technical solution: The present invention proposes an improved K-fold cross-validation method. Compared with the traditional K-fold cross-validation, it can fine-tune on a model by cycling through the K folds, which can overcome the problem that the traditional K-fold cross-validation generates K models and trains on different K - 1 training sets. For named entity recognition, different entities and entity segment threshold scores may be recognized due to different training sets. Using the weighted average of them may be incorrect in the actual scenario, and in the calculation of macro-F1, the calculation of P is on the small side (the weighted prediction of each model makes the number of predicted entities on the large side), resulting in a decrease in the overall index. It effectively improves the index performance of the model.
[0101] Example 4:
[0102] Based on Example 1, Step 3: Based on the selected pseudo-data and the labeled data, use the semi-supervised algorithm and the improved K-fold cross-validation algorithm to train the baseline model UIE to obtain the final model, as Figure 3 shown, including:
[0103] Step 301: Obtain semi-supervised data based on the selected pseudo-data and real data, and divide the semi-supervised data into K parts based on the improved K-fold cross-validation algorithm to obtain the model training set and the model validation set;
[0104] Step 302: Based on the model training set and the model validation set, use the semi-supervised algorithm to train the baseline model UIE to obtain the final model.
[0105] Beneficial effects of the above technical solution: The present invention mixes the selected pseudo-data and real data to obtain semi-supervised data, divides the semi-supervised data into K parts based on the improved K-fold cross-validation algorithm to obtain the model training set and the model validation set, and uses the semi-supervised algorithm to train the baseline model UIE to obtain the final model, effectively improving the generalization ability of the small-sample entity recognition model and the recognition degree of small-sample entity recognition.
[0106] Example 5:
[0107] Based on Example 4, the baseline model UIE is trained using a semi-supervised algorithm to obtain the final model, including:
[0108] Introduce a dynamic parameter α into the loss function of the semi-supervised algorithm to adjust the semi-supervised loss of the selected pseudo-data and real data in the dataset. The specific formula of the semi-supervised loss function is:
[0109] Loss label = α * BCELoss(label(Fake)) + (1 - α) * BCELoss(label(True))
[0110] where Loss label represents the semi-supervised loss, α is the dynamic parameter, and its value range is (0, 1); BCEloss(label(Fake)) represents the binary cross-entropy loss of the pseudo-data, and BCE(logits, label(True)) represents the binary cross-entropy loss of the labeled data; label(Fake) represents the pseudo-data label; label(True) represents the labeled data label;
[0111] As Figure 4 shown, add the dynamic parameter R-Drop algorithm at the output end of the baseline model UIE to perform two Dropouts on the model output, and obtain the KL divergence between the two Dropout processes:
[0112]
[0113] where Loss R-Drop represents the KL divergence between the two Dropout processes; P1 w (y|x) represents the sub-model corresponding to the first Dropout process; P2 w (y|x) represents the sub-model corresponding to the second Dropout process; represents the difference distance between the probability distribution of the first Dropout process and the probability distribution of the second Dropout process; represents the difference distance between the probability distribution of the second Dropout process and the probability distribution of the first Dropout process;
[0114] Generate the total loss function Loss for training the semi-supervised algorithm based on the dynamic parameter β, semi-supervised loss, and R-Drop loss 总 :
[0115] Loss 总 = Loss label + β(Loss R-Drop,start + Loss R-Drop,end )
[0116] where β is a dynamic parameter with a value range of (0, 1]; Loss R-Drop,start represents the KL divergence at the beginning of the entity; Loss R-Drop,end represents the KL divergence at the end of the entity; Loss R-Drop,start +Loss R-Drop,end represents the R-Drop loss;
[0117] Based on the total loss function Loss 总 the model parameters are adjusted, and combined with the model training set and the model validation set, the baseline model UIE is semi-supervised trained to obtain the final model.
[0118] In this embodiment, both the dynamic parameters α and β are adjustable parameters and can be changed during the model training process.
[0119] In this embodiment, the output probabilities of the two Dropouts are kept consistent.
[0120] In this embodiment, the semi-supervised data (mixed data of pseudo data and labeled data) used for model training is small sample data. For example: 10 entity labels, each entity label is labeled with a quantity of about 50, and the total sample annotation quantity is less than 100.
[0121] Advantages of the above technical solution: In the loss function of the semi-supervised algorithm of the present invention, a dynamic parameter α is introduced to adjust the semi-supervised loss of the selected pseudo data and real data in the data set, which can effectively solve the problem that the confidence of the selected pseudo data in the semi-supervised data is low and there are easy errors, effectively improve the accuracy of the training data. The dynamic parameter R-Drop added at the output end of the baseline model UIE performs two Dropouts on the model output, which can effectively improve the generalization ability and semantic understanding ability of the model, reduce the learning of wrong data and prevent the model from overfitting. Based on the dynamic parameter β, the semi-supervised loss and the R-Drop loss, the total loss function Loss for training the semi-supervised algorithm is generated 总 and based on the total loss function Loss 总 the model parameters are adjusted, and combined with the model training set and the model validation set, the baseline model UIE is semi-supervised trained to obtain the final model. Using the dynamic parameter R-Drop is beneficial for the model to learn during the training process, enhances the diversity of the parameters, and indirectly improves the overall ability of the model.
[0122] Embodiment 6:
[0123] Based on Embodiment 1, in step 4, the input text is processed based on the final model to obtain entity recognition results, as Figure 5 shown, including:
[0124] Step 401: Encode the input text based on the fine-tuned model, and extract the text features of the input text to obtain word vectors;
[0125] Step 402: Then, obtain the probability of each word vector through the activation function layer Sigmoid, and through the set threshold comparison, obtain the beginning and the end of the entity segment;
[0126] Step 403: Decode the model based on the beginning and the end and classify the input text category to obtain a type label;
[0127] Step 404: Extract the entity segment according to the beginning position, the end position, and the category label, and output it as the entity recognition result.
[0128] In this embodiment, the input text can be a paragraph, a sentence, or a document described in natural language.
[0129] In this embodiment, Linear usually refers to a linear layer, which is also called a fully connected layer or a dense layer. Its function is to generate a vector of a fixed size for each word through a linear layer, and this vector can represent the features of the input sequence.
[0130] The Sigmoid function is a commonly used activation function. Its function is to map the value after passing through Linear to the interval (0, 1) for predicting the start position of each token.
[0131] Select the word vectors with a threshold greater than 0.5 as the beginning and the end of the entity segment.
[0132] The beneficial effects of the above technical solution: The present invention first encodes the input text based on the fine-tuned model, and extracts the text features of the input text to obtain word vectors; based on the final model, after passing through the Linear layer and then through the activation function layer Sigmoid, obtain the probability of each word vector, and obtain the beginning and the end of the entity segment; then decode the model based on the beginning and the end and classify the input text category to obtain a type label; finally, extract the entity segment and its category according to the beginning position, the end position, and the category label, and output it as the entity recognition result, which can effectively improve the accuracy of entity recognition and improve the recognition precision.
[0133] Embodiment 7:
[0134] Based on Embodiment 1, a small-sample entity recognition method based on pseudo-labels and cross-validation further includes:
[0135] Add a domain recognition model to the input layer of the baseline model UIE. Based on the domain recognition model, perform domain recognition on the input text to determine the description domain of the text used.
[0136] After performing domain annotation on the input text based on the description domain, then perform entity recognition processing.
[0137] Beneficial effects of the above technical solution: In the present invention, a domain recognition model is added to the input layer of the baseline model UIE. Based on the domain recognition model, perform domain recognition on the input text to determine the description domain of the text used. After performing domain annotation on the input text based on the description domain, then perform entity recognition processing. Identifying and determining the domain of the text to be recognized before entity recognition is beneficial to broaden the range of entity content that the model can recognize and improve the accuracy of the model's entity recognition, thereby achieving an improvement in the generalization ability of the final model.
[0138] Example 8:
[0139] Based on Example 7, add a domain recognition model to the input layer of the baseline model UIE. Based on the domain recognition model, perform domain recognition on the input text to determine the description domain of the text used, including:
[0140] Based on the domain recognition model, recognize the input text to determine the positions of the text punctuation. Based on the text punctuation positions, divide the input text into multiple text segments, and perform keyword marking on each text segment.
[0141] Generate corresponding input text vectors for each text segment respectively. Compare each text vector with the database text vectors respectively. Based on the comparison results, obtain the best matching text vectors for each text vector respectively.
[0142] According to the description domain corresponding to the best matching text vector, determine the description domain corresponding to each input text. Based on the sub-vector corresponding to the keyword marking result of the text segment and the keyword sub-vector of its corresponding best matching text vector, obtain the keyword similarity.
[0143] When the keyword similarity is greater than or equal to the preset value, determine the current text segment as a keyword segment, and obtain the description domain corresponding to the keyword segment.
[0144] If the description domains corresponding to all the keyword segments are the same, determine the description domain as the description domain of the current input text.
[0145] Otherwise, obtain the ratio of the total data volume corresponding to the keyword fields to the total data volume of the input text. If the ratio is greater than or equal to the preset threshold, obtain the description fields corresponding to the adjacent text segments of the input text corresponding to each keyword field. If the description field corresponding to the adjacent text segment is the same as that of the input text segment, or the description field corresponding to the adjacent text segment is a general field, then use the keyword field as the target text segment;
[0146] Respectively obtain the description field corresponding to each target text segment and its corresponding first quantity ratio, and use the description field corresponding to the largest first quantity ratio as the description field of the current input text;
[0147] If the ratio is less than the preset threshold, respectively obtain the description field corresponding to each valid text segment and its corresponding second quantity ratio, and use the description field corresponding to the largest second quantity ratio as the description field of the current input text.
[0148] In this embodiment, the database stores templates of various types of entities in multiple description fields and their corresponding text vectors.
[0149] In this embodiment, the description field refers to the field corresponding to the thing described by the input text content.
[0150] In this embodiment, the general field refers to words or sentences that cannot define the description field without being used with other words in the description. For example, today, tomorrow, the weather is nice today, etc.
[0151] In this embodiment, the first quantity ratio refers to the ratio of the data volume of a certain target text segment to the total data volume of all target text segments.
[0152] In this embodiment, the second quantity ratio refers to the ratio of the data volume of a certain valid text segment to the total data volume of all valid text segments.
[0153] In this embodiment, the valid text segment refers to a text segment whose description field is identified as a non-general field.
[0154] Beneficial effects of the above technical solution: Before entity recognition, the present invention confirms the description field of the text, provides a basis for the model to realize multi-field entity recognition, improves the generalization ability of the final model, and broadens the applicable range of the final model.
[0155] Embodiment 9:
[0156] Based on the embodiment 1, a small-sample entity recognition method based on pseudo-labels and cross-validation further includes:
[0157] Based on the domain recognition model, determine the keyword marking result and keyword field determination result corresponding to the entity segment, and record them;
[0158] Before the entity recognition result is output, based on the optional tag category table corresponding to the description field, perform type detection on the type tag corresponding to the entity output result to determine whether the tag is used correctly. If the tag is used correctly, determine whether there are corresponding type tags for the keyword fields and their corresponding keywords. If so, output the entity output result;
[0159] Otherwise, send the entity segment to the final model for secondary entity recognition.
[0160] In this embodiment, the available tag type table is obtained by processing big data to determine the tag types of the named entity content included in the text description of each description field, and is pre-stored in the database.
[0161] The beneficial effects of the above technical solution: Based on the domain recognition model, the present invention determines the keyword marking result and the keyword field determination result corresponding to the entity segment, and records them; and before the entity recognition result is output, based on the optional tag category table corresponding to the description field, perform type detection on the type tag corresponding to the entity output result to determine whether the tag is used correctly, realizing the verification of the description field in the entity model recognition process; when the tag is used correctly, determine whether there are corresponding type tags for the keyword fields and their corresponding keywords. If so, output the entity output result; otherwise, send the entity segment to the final model for secondary entity recognition, realizing the automatic verification of the entity recognition result, ensuring that the important content of the input text is successfully recognized and marked, which is beneficial to improving the accuracy of the output result and providing a reliable reference for users.
[0162] Embodiment 10:
[0163] The present invention provides a small-sample entity recognition system based on pseudo-labels and cross-validation for performing any one of the methods for small-sample entity recognition based on pseudo-labels and cross-validation described in Embodiments 1-9, as Figure 6 shown, including:
[0164] The fine-tuning model training module is used to process the labeled data based on the improved K-fold cross-validation algorithm to obtain a training set and a test set, and train and validate the baseline model UIE to obtain a fine-tuning model;
[0165] The data selection module is used to randomly batch select and process the unlabeled data in the fine-tuning model based on preset rules to obtain the selected pseudo data;
[0166] The final model training module is used to train the baseline model UIE based on the selected pseudo data and the labeled data using a semi-supervised algorithm and an improved K-fold cross-validation algorithm to obtain a final model;
[0167] The recognition result acquisition module is used to process the input text based on the final model to obtain the entity recognition result.
[0168] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and its equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A small-sample entity recognition method based on pseudo-labels and cross-validation, characterized in that Including: Processing the labeled data based on an improved K-fold cross-validation algorithm to obtain a training set and a test set, training and validating the baseline model UIE to obtain a fine-tuned model; Randomly batch-selecting and processing the unlabeled data in the fine-tuned model based on preset rules to obtain the selected pseudo data; Training the baseline model UIE using the semi-supervised algorithm and the improved K-fold cross-validation algorithm based on the selected pseudo data and the labeled data to obtain the final model; Processing the input text based on the final model to obtain the entity recognition result; Among them, randomly batch-selecting and processing the unlabeled data in the fine-tuned model based on preset rules to obtain the selected pseudo data, including: Batch-selecting the unlabeled data according to the preset step size and the sliding window size to obtain a batch of unlabeled data; Predicting the batch of unlabeled data based on the fine-tuned model to obtain the primary pseudo data; Mixing the primary pseudo data and the labeled data to obtain the training data; Training the baseline model UIE based on the training data to obtain the test model; Testing the test data through the test model to obtain the test result; Obtaining the test score corresponding to each test result, and regarding the data with the highest test score as the pseudo data; Among them, training the baseline model UIE using the semi-supervised algorithm and the improved K-fold cross-validation algorithm based on the selected pseudo data and the labeled data to obtain the final model, including: Obtaining semi-supervised data based on the selected pseudo data and the real data, and dividing the semi-supervised data into K parts based on the improved K-fold cross-validation algorithm for use as the model training set and the model validation set; Training the baseline model UIE using the semi-supervised algorithm based on the model training set and the model validation set to obtain the final model.
2. The small-sample entity recognition method based on pseudo-labels and cross-validation according to claim 1, wherein The improved K-fold cross-validation algorithm specifically includes: Based on the improved K-fold cross-validation algorithm, setting K loops and dividing the labeled data into K parts, and at the same time taking any one subset as the test set and the remaining K-1 subsets as the training set; Using K-1 parts of the K parts as the training set and 1 part as the validation set for training respectively to finally obtain a model.
3. According to the method for small-sample entity recognition based on pseudo labels and cross-validation described in claim 1, training the baseline model UIE using the semi-supervised algorithm to obtain the final model, including: Introducing a dynamic parameter α into the loss function of the semi-supervised algorithm to adjust the semi-supervised loss of the selected pseudo data and the real data in the data set. The specific formula of the semi-supervised loss function is: Among them, Loss label represents the semi-supervised loss, α is a dynamic parameter, and its value range is (0, 1); BCELoss(logits, label(Fake)) represents the binary cross-entropy loss of fake data, and BCELoss(logits, label(True)) represents the binary cross-entropy loss of labeled data; label(Fake) represents the fake data label; label(True) represents the labeled data label; Adding the dynamic parameter R-Drop algorithm to the output end of the baseline model UIE to perform two Dropouts on the model output to obtain the KL divergence between the two Dropout processes; Among them, Loss R-Drop represents the KL divergence between two Dropout processes; P1 w (y|x) represents the sub-model corresponding to the first Dropout process; P2 w (y|x) represents the sub-model corresponding to the second Dropout process; represents the difference distance between the probability distributions of the first Dropout process and the second Dropout process; represents the difference distance between the probability distributions of the second Dropout process and the first Dropout process; Generate the total loss function Loss for training the semi-supervised algorithm based on the dynamic parameter β, semi-supervised loss, and R-Drop loss 总 : Among them, β is a dynamic parameter, and its value range is (0, 1]; represents the KL divergence at the beginning of the entity; represents the KL divergence at the end of the entity; represents the R-Drop loss; Based on the total loss function Loss 总 Adjust the model parameters, and combine the model training set and the model validation set to perform semi-supervised training on the baseline model UIE to obtain the final model.
4. A small sample entity recognition method based on pseudo labels and cross-validation according to claim 1, characterized in that Processing the input text based on the final model to obtain the entity recognition result, including: Encoding the input text based on the fine-tuned model and extracting the text features of the input text to obtain the word vectors; Based on the final model, passing through the Linear layer and then through the activation function layer Sigmoid to obtain the probability of each word vector, and obtaining the start and end of the entity segment; Perform model decoding based on the beginning and the end and classify the input text category to obtain a type label; Extract entity fragments according to the beginning position, the end position, and the category label, and output them as entity recognition results.
5. A small-sample entity recognition method based on pseudo-labels and cross-validation according to claim 1, characterized in that, It also includes: Add a domain recognition model to the input layer of the baseline model UIE, perform domain recognition on the input text based on the domain recognition model, and determine the description domain of the text used; And after performing domain annotation on the input text based on the description domain, perform entity recognition processing.
6. The small sample entity recognition method based on pseudo-labels and cross-validation according to claim 5, characterized in that Add a domain recognition model to the input layer of the baseline model UIE, perform domain recognition on the input text based on the domain recognition model, and determine the description domain of the text used, including: Based on the domain recognition model, recognize the input text to determine the positions of the text punctuation marks. Based on the text punctuation mark positions, divide the input text into multiple text segments, and perform keyword marking on each text segment; Generate corresponding input text vectors for each text segment respectively, compare each text vector with the database text vectors respectively, and obtain the best matching text vector for each text vector based on the comparison results; Determine the description domain corresponding to each input text according to the description domain corresponding to the best matching text vector, and obtain the keyword similarity based on the sub-vector corresponding to the keyword marking result of the text segment and the keyword sub-vector of its corresponding best matching text vector; When the keyword similarity is greater than or equal to the preset value, determine that the current text segment is a keyword segment, and obtain the description domain corresponding to the keyword segment; If the description domains corresponding to all keyword segments are the same, determine that the description domain is the description domain of the current input text; Otherwise, obtain the ratio of the total data volume corresponding to the keyword segments to the total data volume of the input text. If the ratio is greater than or equal to the preset threshold, obtain the description domains corresponding to the adjacent text segments of the input text corresponding to each keyword segment. If the description domains corresponding to the adjacent text segments are the same as the input text segment, or the description domains corresponding to the adjacent text segments are general domains, then use the keyword segments as target text segments; Obtain the description domain corresponding to each target text segment and its corresponding first quantity ratio respectively, and use the description domain corresponding to the largest first quantity ratio as the description domain of the current input text; If the ratio is less than the preset threshold, obtain the description domain corresponding to each effective text segment and its corresponding second quantity ratio respectively, and use the description domain corresponding to the largest second quantity ratio as the description domain of the current input text.
7. A method for small-sample entity recognition based on pseudo-labels and cross-validation according to claim 6, characterized in that It also includes: Based on the domain recognition model, determine the keyword marking result and the keyword segment determination result corresponding to the entity fragment, and record them; Before the entity recognition result is output, perform type detection on the type label corresponding to the entity output result based on the optional label category table corresponding to the description domain to judge whether the label is used correctly. If the label is used correctly, judge whether there is a corresponding type label for the keyword segment and its corresponding keyword. If so, output the entity output result; Otherwise, send the entity fragment to the final model for secondary entity recognition.
8. A small-sample entity recognition system based on pseudo-labels and cross-validation, characterized in that, It includes: The fine-tuning model training module is used to process the labeled data based on the improved K-fold cross-validation algorithm, obtain the training set and the test set, and train and validate the baseline model UIE to obtain the fine-tuning model; The data selection module is used to randomly select a batch of unlabeled data in the fine-tuning model based on preset rules, and after processing, obtain the selected pseudo data; The final model training module is used to train the baseline model UIE based on the selected pseudo data and the labeled data using the semi-supervised algorithm and the improved K-fold cross-validation algorithm to obtain the final model; The recognition result acquisition module is used to process the input text based on the final model to obtain the entity recognition result; Among them, the method for the data selection module to randomly select a batch of unlabeled data in the fine-tuning model based on preset rules and obtain the selected pseudo data after processing includes: Select a batch of unlabeled data according to the preset step size and the sliding window size to obtain a batch of unlabeled data; Predict the batch of unlabeled data based on the fine-tuning model to obtain the primary pseudo data; Mix the primary pseudo data and the labeled data to obtain the training data; Train the baseline model UIE based on the training data to obtain the test model; Test the test data through the test model to obtain the test result; Obtain the test score corresponding to each test result, and regard the data with the highest test score as the pseudo data; Among them, the method for the final model training module to train the baseline model UIE based on the selected pseudo data and the labeled data using the semi-supervised algorithm and the improved K-fold cross-validation algorithm to obtain the final model includes: Obtain semi-supervised data based on the selected pseudo data and the real data, and divide the semi-supervised data into K parts based on the improved K-fold cross-validation algorithm for use as the model training set and the model validation set; Based on the model training set and the model validation set, use the semi-supervised algorithm to train the baseline model UIE to obtain the final model.
Citation Information
Patent Citations
Intention recognition method based on semi-supervised learning, device, equipment and medium
CN113704429A
Semi-supervised named entity recognition method for unlabeled data
CN114266253A