A Natural Language Data Labeling Method and System Based on Reinforcement Learning
The reinforcement learning-based data annotation method integrates human and AI processes to optimize data expansion and quality, addressing inefficiencies in existing natural language processing by enhancing annotation efficiency and model precision.
Patent Information
- Application Number
- CN202111196737.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-14
AI Technical Summary
Natural language data annotation depends on high labor costs and low AI processing accuracy, and low data expansion and labeling efficiency.
A method based on reinforcement learning is adopted, combining four algorithms for data recommendation and generation. By annotating personnel to select data and optimizing the model based on reward values, online training and adjustment of the algorithm output ratio is achieved.
It reduces the workload of manual processing, improves data labeling efficiency and AI processing accuracy, supports the generation of high-quality data on low-quality unlabeled data, and is suitable for a variety of natural language processing scenarios.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method and system for natural language data labeling based on reinforcement learning. Background Art
[0002] Currently, all walks of life are continuously strengthening the application of artificial intelligence technology. However, for current AI applications mainly based on deep learning, their quality highly depends on the quality and quantity of data. The quality and quantity of data determine the launch of the deep network model. In the field of natural language processing, whether it is the development of a dialogue robot or the structured operation of a large amount of text, it is inseparable from a large amount of high-quality labeled text data. How to achieve high-quality data labeling with higher efficiency on the premise of minimizing labor costs is an urgent problem in the industry.
[0003] Most organizations and institutions in the industry often hire professional data labeling teams to manually label their business data according to given standards. This labeling scheme completely relies on manual work, and the labeled data has extremely high quality, but the cost is huge. At the same time, some research institutions choose to train different algorithm models to automatically label data. This method has extremely low cost, but there are several defects: firstly, the algorithm models that can be automatically labeled often also rely on manual labeling to a certain extent; secondly, the accuracy of the model trained with the automatically labeled samples is difficult to exceed the accuracy of the original labeled model, which is not conducive to the iterative optimization of the model prediction accuracy; at the same time, the quality of the automatically labeled text cannot be fully guaranteed. Specifically for natural language data, in some cases, the available datasets for labeling are extremely limited. How to combine data augmentation and data labeling is also an important issue. Summary of the Invention
[0004] Object of the Invention: The present invention provides a method and system for natural language data labeling based on reinforcement learning, which focuses on solving the problems of large manual processing workload for natural language data augmentation and data labeling, and low AI processing accuracy.
[0005] Technical Solution: A method for natural language data labeling based on reinforcement learning includes the following steps:
[0006] (1) Data recommendation and generation are performed according to the standard data manually selected by the labeler.
[0007] (2) The data output by different algorithms in the data recommendation and generation is pushed to the labeler according to a certain ratio, i.e., the algorithm output ratio, and the labeler selects the pushed data.
[0008] (3) Obtain the reward value according to the proportion of the data selected by the labeled personnel in the output data of each algorithm to the total output data of the algorithm. Use the policy gradient method in reinforcement learning to train the model and optimize the algorithm model online.
[0009] The data recommendation and generation adopt four algorithms, specifically including:
[0010] Text matching algorithm based on text character similarity: Compare the similarity between the text in the unlabeled text library and the standard text at the character level through various character similarity metrics such as edit distance and the number of co-occurring words in the text, and output the texts with the top similarity rankings.
[0011] Text matching algorithm based on semantic similarity: Use the simCSE contrastive learning algorithm to train based on the unlabeled database, learn relevant semantic information, and output texts similar to the standard statement from a semantic perspective.
[0012] Text generation algorithm based on the MLM autoencoder model: Split the standard data using a syntactic analysis model for grammar elements, mask a certain grammar element, and use the wobert model to predict and generate synonyms for the masked word.
[0013] Text generation algorithm based on the autoregressive model: Perform transfer learning on the data based on the GPT2 algorithm, and finally generate text content similar to the standard data.
[0014] The specific content of step (2) is that the initial output probabilities of each algorithm are equal. The labeled personnel select the pushed data, and the module automatically collects the proportion of the data selected by the labeled personnel in the output data of each algorithm to the total output data of the algorithm.
[0015] The data output by different algorithms in data recommendation and generation is pushed to the labeled personnel according to a certain proportion, and this proportion is the algorithm output proportion. The proportion of the data selected by the labeled personnel in the output data of each algorithm to the total output data of the algorithm is the reward value. The loss function of the policy gradient method in reinforcement learning is
[0016] loss = -log(p) * reward
[0017] where p is the algorithm output proportion and reward is the reward value. By optimizing the training of this loss function, the algorithm model is optimized online.
[0018] A natural language data labeling system based on reinforcement learning, including:
[0019] A data recommendation and generation module that performs data recommendation and generation according to the standard data manually selected by the labeled personnel.
[0020] The annotation personnel interaction module pushes the data output by different algorithms in data recommendation and generation to the annotation personnel according to a certain ratio, namely the algorithm output ratio, and the annotation personnel select the pushed data;
[0021] The annotation algorithm management and training module obtains the reward value according to the ratio of the data selected by the annotation personnel in the output data of each algorithm to the total output data of the algorithm, and trains the model using the policy gradient method in reinforcement learning to optimize the algorithm model online.
[0022] The data recommendation and generation module uses four algorithms, specifically including:
[0023] Text matching algorithm based on text character similarity: By using various character similarity metrics such as edit distance and the number of co-occurring words in the text, the similarity between the text in the unlabeled text library and the standard text is compared at the character level, and the texts with the top similarity rankings are output;
[0024] Text matching algorithm based on semantic similarity: Using the simCSE contrastive learning algorithm, training is performed based on the unlabeled database, learning relevant semantic information, and outputting texts similar to the standard statement from the semantic perspective;
[0025] Text generation algorithm based on the MLM autoencoder model: The standard data is split into syntactic elements using a syntactic analysis model, a certain syntactic element is masked, and the wobert model is used to predict and generate synonyms for the masked word;
[0026] Text generation algorithm based on the autoregressive model: Based on the GPT2 algorithm, transfer learning is performed on the data, and finally text content similar to the standard data is generated.
[0027] In the annotation personnel interaction module, the initial output probability of each algorithm is equal. The annotation personnel select the pushed data, and the module automatically collects the ratio of the data selected by the annotation personnel in the output data of each algorithm to the total output data of the algorithm.
[0028] The data output by different algorithms in the data recommendation and generation module is pushed to the annotation personnel according to a certain ratio, and this ratio is the algorithm output ratio. The ratio of the data selected by the annotation personnel in the output data of each algorithm to the total output data of the algorithm is the reward value. The loss function of the policy gradient method in reinforcement learning is
[0029] loss=-log(p)*reward
[0030] where p is the algorithm output ratio and reward is the reward value. By optimizing the training of this loss function, the algorithm model is optimized online.
[0031] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:
[0032] The present invention solves the problems of large workload of manual processing for natural language data augmentation and data labeling, and low AI processing accuracy, and combines manual processing with AI processing. A natural language data recommendation and generation module is constructed, which supports searching for data based on a large amount of low-quality unlabeled data, and also supports generating data based on no candidate data, meeting various natural language processing scenarios and conforming to the characteristics of natural language data. A labeling algorithm management and training module is constructed. Through continuous training of the built-in model online, timely fitting for different scenarios of different labelers is achieved. By adjusting the output ratio of each algorithm in the pre-module, the labeling efficiency is continuously improved. Specific implementation manner
[0033] A natural language data labeling method based on reinforcement learning, comprising the following steps:
[0034] (1) Data recommendation and generation are performed according to the standard data manually selected by the labeler;
[0035] (2) The data output by different algorithms in data recommendation and generation is pushed to the labeler according to a certain ratio, i.e., the algorithm output ratio, and the labeler selects the pushed data;
[0036] (3) The reward value is obtained according to the ratio of the data selected by the labeler in the data output by each algorithm to the total output data of the algorithm. The model is trained by using the policy gradient method in reinforcement learning according to the reward value, and the algorithm model is optimized online.
[0037] Four algorithms are adopted for the data recommendation and generation, specifically including:
[0038] A text matching algorithm based on text character similarity: By using various character similarity indexes such as edit distance and the number of co-occurring words in the text, the similarity between the text in the unlabeled text library and the standard text is compared at the character level, and the text with the top similarity ranking is output;
[0039] A text matching algorithm based on semantic similarity: The simCSE contrastive learning algorithm is used to train based on the unlabeled database, learn relevant semantic information, and output text similar to the standard statement from the semantic perspective;
[0040] Text generation algorithm based on the MLM auto-encoding model: The standard data is split into syntactic elements using a syntactic analysis model, such as being split into the subject-predicate-object structure. One of its syntactic elements is masked, and the wobert model is used to predict the masked word. For example, taking "I watched a great fireworks show today. It was very nice" as the standard text, replacing the subject qualifier "nice" with "[MASK]", and using the model to predict the mask. The final output result is "I watched a great fireworks show today. It was very [wonderful || like || nice || excellent || good]", and different words can be used to generalize and enhance the standard data;
[0041] Text generation algorithm based on the autoregressive model: Based on the GPT2 algorithm, transfer learning is performed on the data, and finally text content similar to the standard data is generated.
[0042] Specifically, in step (2), the initial output probabilities of each algorithm are equal. The annotator selects the pushed data, and the module automatically collects the proportion of the data selected by the annotator in the output data of each algorithm to the total output data of that algorithm. If 100 pieces of data are pushed to the annotator in total, and there are 4 algorithms in total, and the initial output probabilities of each algorithm are equal, all being 25%, that is, each algorithm outputs 25 pieces of data. The annotator selects the pushed data, and the module automatically collects the proportion of the selected data in the output data of each algorithm to the total output data of that algorithm. This proportion is the reward of that algorithm. For example, if algorithm A outputs 25 pieces of data in total, and 10 of them are selected by the annotator, then the reward of this algorithm is 0.4.
[0043] In data recommendation and generation, the data output by different algorithms is pushed to the annotator according to a certain proportion, and this proportion is the algorithm output proportion. The proportion of the data selected by the annotator in the output data of each algorithm to the total output data of that algorithm is the reward value. The loss function of the policy gradient method in reinforcement learning is
[0044] loss = -log(p) * reward
[0045] where p is the algorithm output proportion and reward is the reward value. By optimizing the training of this loss function, the algorithm model is optimized online. If the model determines that the output proportion of algorithm A is p and its obtained reward score of algorithm A is reward A then the loss function loss = -log(p) * reward A
[0046] A natural language data marking system based on reinforcement learning, including:
[0047] A data recommendation and generation module that performs data recommendation and generation according to the standard data manually selected by the annotator;
[0048] The annotation personnel interaction module pushes the data output by different algorithms in data recommendation and generation to the annotation personnel according to a certain ratio, i.e., the algorithm output ratio, and the annotation personnel make selections on the pushed data.
[0049] The annotation algorithm management and training module obtains the reward value according to the ratio of the data selected by the annotation personnel in the output data of each algorithm to the total output data of the algorithm, and uses the policy gradient method in reinforcement learning to train the model according to the reward value, and online optimizes the algorithm model.
[0050] The data recommendation and generation module adopts four algorithms, specifically including:
[0051] The text matching algorithm based on text character similarity: Compare the similarity between the text in the unlabeled text library and the standard text at the character level through various character similarity indicators such as edit distance and the number of co-occurring words in the text, and output the text with the top similarity rankings.
[0052] The text matching algorithm based on semantic similarity: Use the simCSE contrastive learning algorithm to train based on the unlabeled database, learn relevant semantic information, and output text similar to the standard sentence from the semantic perspective.
[0053] The text generation algorithm based on the MLM autoencoder model: Split the standard data using a syntactic analysis model into grammatical elements, such as splitting it into a subject-predicate-object structure. Mask one of its grammatical elements and use the wobert model to predict the masked word. For example: Take "I watched a great fireworks show today. It was very nice" as the standard text, replace its subject qualifier "nice" with "[MASK]", and use the model to predict the mask. Finally, the output result is "I watched a great fireworks show today. It was very [wonderful || like || nice || excellent || good]", and different words can promote and enhance the standard data.
[0054] The text generation algorithm based on the autoregressive model: Perform transfer learning on the data based on the GPT2 algorithm, and finally generate text content similar to the standard data.
[0055] In the annotator interaction module, the initial output probabilities of each algorithm are equal. The annotator selects the pushed data, and the module automatically collects the proportion of the data selected by the annotator in the output data of each algorithm to the total output data of that algorithm. If 100 pieces of data are pushed to the annotator in total, and there are 4 algorithms in total, and the initial output probabilities of each algorithm are equal, all being 25%, that is, each algorithm outputs 25 pieces of data. The annotator selects the pushed data, and the module automatically collects the proportion of the selected data in the output data of each algorithm to the total output data of that algorithm. This proportion is the reward of that algorithm. For example, if algorithm A outputs 25 pieces of data in total, and 10 of them are selected by the annotator, then the reward of this algorithm is 0.4.
[0056] In the data recommendation and generation module, the data output by different algorithms is pushed to the annotator according to a certain proportion, and this proportion is the algorithm output proportion. The proportion of the data selected by the annotator in the output data of each algorithm to the total output data of that algorithm is the reward value. The loss function of the policy gradient method in reinforcement learning is
[0057] loss=-log(p)*reward
[0058] where p is the algorithm output proportion and reward is the reward value. By optimizing the training of this loss function, the algorithm model is optimized online. If the model judges that the output proportion of algorithm A is p and it obtains the reward score of algorithm A as reward A , then the loss function loss=-log(p)*reward A .
Claims
1. A text data labeling method based on reinforcement learning, characterized in that, It includes the following steps: (1) Perform data recommendation and generation according to the standard data manually selected by the annotator; (2) Push the data output by different algorithms in data recommendation and generation to the annotator according to a certain proportion, i.e., the algorithm output proportion, and the annotator selects the pushed data; (3) Obtain the reward value according to the proportion of the data selected by the annotator in the data output by each algorithm to the total data output by the algorithm, and train the model using the policy gradient method in reinforcement learning to optimize the model online; The data output by different algorithms in data recommendation and generation is pushed to the annotator according to a certain proportion, which is the algorithm output proportion. The proportion of the data selected by the annotator in the data output by each algorithm to the total data output by the algorithm is the reward value. The loss function of the policy gradient method in reinforcement learning is loss=-log(p)*reward where p is the algorithm output proportion and reward is the reward value. The model is optimized online by optimizing the training of this loss function.
2. The method for text data labeling based on reinforcement learning according to claim 1, wherein The data recommendation and generation adopt four algorithms, specifically including: Text matching algorithm based on text character similarity: Compare the similarity between the text in the unlabeled text library and the standard text at the character level through edit distance and text co-occurrence word number similarity metrics, and output the text with the top similarity rankings; Text matching algorithm based on semantic similarity: Use the simCSE contrastive learning algorithm to train based on the unlabeled database, learn relevant semantic information, and output text similar to the standard statement from the semantic perspective; Text generation algorithm based on the MLM auto-encoding model: Split the standard data using a syntactic analysis model for grammar elements, mask a certain grammar element, and use the wobert model to predict and generate synonyms for the masked word; Text generation algorithm based on the autoregressive model: Perform transfer learning on the data based on the GPT2 algorithm, and finally generate text content similar to the standard data.
3. A method for text data labeling based on reinforcement learning according to claim 1, characterized in that, The specific step (2) is that the initial output probability of each algorithm is equal. The annotator selects the pushed data, and the module automatically collects the proportion of the data selected by the annotator in the data output by each algorithm to the total data output by the algorithm.
4. A text data labeling system based on reinforcement learning, characterized in that, It includes: Data recommendation and generation module, which performs data recommendation and generation according to the standard data manually selected by the annotator; Annotator interaction module, which pushes the data output by different algorithms in data recommendation and generation to the annotator according to a certain proportion, i.e., the algorithm output proportion, and the annotator selects the pushed data; Annotating algorithm management and training module, which obtains the reward value according to the proportion of the data selected by the annotator in the data output by each algorithm to the total data output by the algorithm, and trains the model using the policy gradient method in reinforcement learning to optimize the model online; The data output by different algorithms in the data recommendation and generation module is pushed to the annotator according to a certain proportion, which is the algorithm output proportion. The proportion of the data selected by the annotator in the data output by each algorithm to the total data output by the algorithm is the reward value. The loss function of the policy gradient method in reinforcement learning is loss=-log(p)*reward Where p is the output ratio of the algorithm and reward is the reward value. The model is optimized online by optimizing this loss function for training.
5. A text data tagging system based on reinforcement learning according to claim 4, characterized in that, The data recommendation and generation module adopts four algorithms, specifically including: Text matching algorithm based on text character similarity: By using various character similarity metrics such as edit distance and the number of co-occurring words in the text, the similarity between the text in the unlabeled text library and the standard text is compared at the character level, and the texts with the top similarity rankings are output. Text matching algorithm based on semantic similarity: Using the simCSE contrastive learning algorithm, it is trained based on the unlabeled database to learn relevant semantic information, and texts similar to the standard statement are output from the semantic perspective. Text generation algorithm based on the MLM autoencoding model: The standard data is split into syntactic elements using a syntactic analysis model, a certain syntactic element is masked, and the wobert model is used to predict and generate synonyms for the masked word. Text generation algorithm based on the autoregressive model: Based on the GPT2 algorithm, transfer learning is performed on the data, and finally text content similar to the standard data is generated.
6. The text data tagging system based on reinforcement learning according to claim 4, characterized in that, In the annotator interaction module, the initial output probabilities of each algorithm are equal. The annotator selects the pushed data, and the module automatically collects the proportion of the data selected by the annotator in the output data of each algorithm to the total output data of that algorithm.
Citation Information
Patent Citations
Natural language generation model training method and device
CN113111638A
Remote sensing image text generation and optimization method based on self-reinforcement learning
CN113312925A