Label data inspection method, device, equipment, medium and computer program product
By constructing training data and training reward models, the pre-labeled data is automatically tested, which solves the problem of manual verification work in the existing technology and realizes efficient data labeling inspection.
Patent Information
- Application Number
- CN202510271191.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the accuracy of data labeling requires manual verification, resulting in large workload and low efficiency. How to achieve automatic verification of pre-labeled data has become an urgent problem.
By constructing training data based on historical query data and historical annotation data, generating prompt instructions and inputting the reward model to be trained, obtaining the reward scalar, training the reward model based on scalar loss, and finally entering the trained reward model to be tested data and pre-label data to be tested, obtaining the pre-label test result.
Automatic inspection of pre-labeled data is realized, the workload of manual verification is reduced, and the inspection efficiency of data labeling is improved.
Smart Images

Figure CN120067616A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of labeled data verification, and in particular, to a method, device, equipment, medium and computer program product for verifying labeled data. Background Art
[0002] Data annotation refers to adding annotations to data, and the annotated data is used for model training and evaluation. To ensure the accuracy of data annotation, the existing method for annotating data is generally manual annotation. However, the efficiency of manual annotation is low. To improve the efficiency of data annotation, data pre-annotation can be used to pre-annotate a large number of data. However, the accuracy of the pre-annotated data still needs to be manually verified. To reduce the workload of manually verifying data annotation, how to automatically verify the pre-annotated data has become an urgent technical problem to be solved. Summary of the Invention
[0003] The present invention provides a method, device, equipment, medium and computer program product for verifying labeled data, which is used to solve the problem of automatically verifying labeled data, realize the automatic verification of labeled data, reduce the manual workload, and improve the verification efficiency of data annotation.
[0004] The present invention provides a method for verifying labeled data, including the following steps: Construct training data based on historical query data and historical annotation data; Input the prompt instruction generated based on the training data into the reward model to be trained, and obtain the reward scalar output by the reward model; Train the reward model based on the scalar loss; the scalar loss is determined based on the reward scalar; Input the user query data and the pre-annotated data to be verified into the trained reward model to obtain the pre-annotation verification result.
[0005] According to the method for verifying labeled data provided by the present invention, the historical annotation data includes correct annotation data and historical pre-annotation data; the construction of training data based on historical query data and historical annotation data includes: When the historical pre-annotation data is correctly annotated, convert the historical pre-annotation data into incorrectly annotated data; When the historical pre-annotation data is incorrectly annotated, use the historical pre-annotation data as incorrectly annotated data; Based on the historical query data, the correct annotation data and the incorrect annotation data, construct a data triple; Construct training data based on the data triple.
[0006] A method for verifying labeled data provided by the present invention, before inputting the prompt instruction generated based on the training data into the reward model to be trained to obtain the reward scalar output by the reward model, includes: Generate a correct instruction based on the concatenation result of the correctly labeled data and the historical query data; Generate a prompt instruction based on the correct instruction and the wrong instruction; the wrong instruction is generated based on the concatenation result of the wrongly labeled data and the historical query data.
[0007] A method for verifying labeled data provided by the present invention, the step of inputting the prompt instruction generated based on the training data into the reward model to be trained to obtain the reward scalar output by the reward model includes: Obtain the first hidden state of the minimum segmentation unit corresponding to the correct instruction, and obtain the second hidden state of the minimum segmentation unit corresponding to the wrong instruction; Connect the first hidden state to a linear layer to obtain a correct reward scalar, and connect the second hidden state to a linear layer to obtain a wrong reward scalar.
[0008] A method for verifying labeled data provided by the present invention, the method for verifying labeled data further includes: Determine the scalar difference between the correct reward scalar and the wrong reward scalar; Input the scalar difference into a loss function to obtain the scalar loss output by the loss function.
[0009] A method for verifying labeled data provided by the present invention, the step of inputting the user query data and the pre-labeled data to be verified into the trained reward model to obtain the pre-label verification result includes: Construct a guiding instruction based on the user query data and the pre-labeled data to be verified; Input the guiding instruction into the trained reward model to obtain the normalized score output by the trained reward model; Based on the normalized score, determine the pre-label verification result of the pre-labeled data to be verified.
[0010] The present invention also provides a labeled data verification device, including the following modules: A training data construction module for constructing training data based on historical query data and historical labeled data; A reward scalar obtaining module for inputting the prompt instruction generated based on the training data into the reward model to be trained to obtain the reward scalar output by the reward model; A reward model training module for training the reward model based on the scalar loss; the scalar loss is determined based on the reward scalar; The pre-annotation verification module is used to input the user query data and the pre-annotation data to be verified into the trained reward model to obtain the pre-annotation verification result.
[0011] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the annotation data verification method as described in any one of the above is implemented.
[0012] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the annotation data verification method as described in any one of the above is implemented.
[0013] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the annotation data verification method as described in any one of the above is implemented.
[0014] The annotation data verification method, device, equipment, medium, and computer program product provided by the present invention construct the data to be trained through historical query data and historical annotation data, generate a prompt instruction based on the training data, use the prompt instruction as the input of the reward model to be trained, and obtain the reward scalar output by the reward model to be trained; determine the scalar loss based on the reward scalar, train the reward model through the scalar loss, and finally input the user's query data and the pre-annotation data to be verified into the trained reward model to obtain the verification result of the pre-annotation data output by the trained reward model. The present invention quickly verifies the pre-annotation data to be verified through the trained reward model, reduces the workload of manual verification, and improves the verification efficiency of the annotation data. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is one of the flow diagrams of the annotation data verification method provided by the present invention.
[0017] Figure 2 is the second flow diagram of the annotation data verification method provided by the present invention.
[0018] Figure 3 is the structural diagram of the annotation data verification device provided by the present invention.
[0019] Figure 4It is a schematic structural diagram of the electronic device provided by the present invention. Specific Embodiments
[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0021] The following Figures 1 - 4 describes the method, device, equipment, medium, and computer program product for verifying labeled data of the present invention.
[0022] Figure 1 is one of the flow schematic diagrams of the method for verifying labeled data provided by the present invention. As Figure 1 shown, the method includes the following: Step 100: Construct training data based on historical query data and historical labeled data; Specifically, the specific implementation process of the method for verifying labeled data provided by the present invention includes the following content (the specific steps are only used to illustrate the main content of the technical solution of the present invention and do not represent their sequence).
[0023] Step 1: Construction of training data: Each piece of training data can be a triple, including the user's historical query content, historical correct label, and historical incorrect label. Among them, the historical correct label can be manually labeled, and the historical incorrect label can be a pre-label automatically generated. If the pre-label automatically generated is a correct label, the pre-label automatically generated can be made into an incorrect label through prompt engineering (also known as prompting engineering or instruction engineering: used to guide a large pre-trained language model to generate high-quality, accurate, and targeted outputs), so as to ensure that there are incorrect data labels in the training data set. Similarly, the historical correct label can also be automatically generated through relevant algorithms (or models), and then when the automatically generated label is inaccurate, it can be converted through the above method.
[0024] Step 200: Input the prompt instruction generated based on the training data into the reward model to be trained, and obtain the reward scalar output by the reward model; Specifically, the specific implementation process of the method for verifying labeled data provided by the present invention further includes the following content: Step 2, Construct model input: The prompt instructions input to the reward model to be trained include correct prompt instructions and incorrect prompt instructions. The correct prompt instructions are obtained by concatenating the user's historical query content and historical correct annotations; the incorrect prompt instructions are obtained by concatenating the user's historical query content and historical incorrect annotations.
[0025] Step 3, Output of the reward model to be trained: Input the correct prompt instructions and incorrect prompt instructions into the reward model to be trained to obtain the correct reward scalar and incorrect reward scalar output by the reward model to be trained.
[0026] Step 300, Train the reward model based on the scalar loss; the scalar loss is determined based on the reward scalar; Specifically, the specific implementation process of the annotation data verification method provided by the present invention further includes the following content: Step 4, Loss calculation: In order to enable the reward model to be trained to better learn to distinguish between correct prompt instructions and incorrect prompt instructions and make this distinction as obvious as possible. The present invention constructs a loss function whose output is negatively correlated with the scalar difference. The output of this loss function is the scalar loss in this embodiment. The scalar difference refers to the difference between the above-mentioned correct reward scalar and incorrect reward scalar. For example, the loss function can be a logarithmic loss function constructed using the sigmoid function (activation function).
[0027] Step 400, Input the user query data and the pre-annotation data to be verified into the trained reward model to obtain the pre-annotation verification result.
[0028] Specifically, the specific implementation process of the annotation data verification method provided by the present invention further includes the following content: Step 5, Inference of the trained reward model: Construct a prompt instruction from the user query data and the pre-annotation data to be verified (automatically pre-annotated for the user query data), and obtain a score through the trained reward model; after obtaining this score, obtain a normalized score through logistic regression, with a range between 0 and 1; set a threshold, such as 0.5. If the normalized score is greater than this threshold, it is determined that the pre-annotation data to be verified is correct, otherwise it is determined that the pre-annotation data to be verified is incorrect.
[0029] In this embodiment, training data is constructed using historical query data and historical annotation data. A prompt instruction is generated based on the training data and used as the input to a reward model to be trained, obtaining a reward scalar output by the reward model to be trained. A scalar loss is determined based on the reward scalar, and the reward model is trained using the scalar loss. Finally, the user's query data and pre-annotated data to be verified are input into the trained reward model to obtain the verification result of the pre-annotated data output by the trained reward model. Through the trained reward model, the present invention quickly verifies the pre-annotated data to be verified, reducing the workload of manual verification and improving the verification efficiency of the annotated data.
[0030] Figure 2 is the second flowchart of the method for verifying annotated data provided by the present invention. As Figure 2 shown, the method may further include: Step 110, when the historical pre-annotation data is correctly annotated, convert the historical pre-annotation data into incorrectly annotated data; Step 120, when the historical pre-annotation data is incorrectly annotated, use the historical pre-annotation data as the incorrectly annotated data; Step 130, based on the historical query data, the correctly annotated data, and the incorrectly annotated data, construct a data triple; Step 140, construct training data based on the data triple.
[0031] Specifically, when constructing training data in the present invention, a triple is constructed for each piece of training data. The triple includes the user's historical query content (which can be the text content with query attributes input by the user), the correct annotation, and the incorrect annotation (i.e., the historical annotation data). Among them, the correct annotation and the incorrect annotation can be automatically generated by the model, and then the generated correct annotation is determined to be correct and the generated incorrect annotation is determined to be incorrect through manual verification. In the case where the annotation automatically generated by the model is inaccurate, for example, the correct annotation generated by the model is incorrect and the incorrect annotation generated by the model is correct, in this case, the inaccurate annotation can be converted into an accurate annotation through prompt engineering.
[0032] In this embodiment, training data is constructed using historical query data and historical annotation data. Each piece of training data contains corresponding query content, correct annotation, and incorrect annotation, which enables the model trained based on this training data to better learn accurate data annotation.
[0033] In one embodiment, the method for verifying annotated data provided by the embodiments of the present invention may further include: Step 10, generate a correct instruction based on the concatenation result of the correctly annotated data and the historical query data; Step 20: Generate a prompt instruction based on the correct instruction and the incorrect instruction; the incorrect instruction is generated based on the concatenation result of the incorrect annotation data and the historical query data.
[0034] Specifically, after constructing training data based on historical query data and historical annotation data, on the basis of the training data set, a prompt (i.e., the prompt instruction in this embodiment, which is used to guide the generation of input instructions for an artificial intelligence model to perform a specific task, and the prompt instruction usually appears in the form of natural language text) for input into the reward model to be trained is generated. The prompt instructions proposed by the present invention include correct prompt instructions and incorrect prompt instructions. Among them, the correct prompt instruction is obtained by concatenating the correct annotation data and the historical query data; the incorrect prompt instruction is obtained by concatenating the incorrect annotation data and the historical query data. The prompt instruction includes the correct prompt instruction and the incorrect prompt instruction, and the correct prompt instruction and the incorrect prompt instruction have similar but different processing processes in the reward model to be trained.
[0035] In this embodiment, the correct prompt instruction and the incorrect prompt instruction are respectively generated through the correct annotation data and the incorrect annotation data included in the training data, so that the reward model to be trained can better learn the difference between the correct annotation and the incorrect annotation.
[0036] In one embodiment, the annotation data verification method provided by the embodiments of the present invention may further include: Step 210: Obtain the first hidden state of the smallest segmentation unit corresponding to the correct instruction, and obtain the second hidden state of the smallest segmentation unit corresponding to the incorrect instruction; Step 220: Connect the first hidden state to a linear layer to obtain a correct reward scalar, and connect the second hidden state to a linear layer to obtain an incorrect reward scalar.
[0037] Specifically, input the correct prompt instruction and the incorrect prompt instruction generated above into the reward model to be trained, and obtain the reward scalars output by the reward model to be trained, including the correct reward scalar and the incorrect reward scalar. The process of generating the reward scalar is as follows: Obtain the hidden state (i.e., the hidden state in this embodiment: used to store and transmit internal information when processing the input sequence, which is the first hidden state in this embodiment) of the last token of the correct prompt instruction (i.e., the smallest segmentation unit in this embodiment, which is the smallest unit for the model to understand and process text); obtain the hidden state of the last token of the incorrect prompt instruction (which is the second hidden state in this embodiment). Then, connect a linear layer behind the first hidden state and the second hidden state respectively, and finally output the correct reward scalar and the incorrect reward scalar respectively. The reward scalar can be in the form of a number (score).
[0038] In this embodiment, by inputting a prompt instruction to the reward model to be trained, the reward scalars corresponding to the correctly labeled data and the incorrectly labeled data are obtained, and the ability of the reward model to recognize correct and incorrect labels is improved by enabling the model to clearly distinguish between the correct reward scalar and the incorrect reward scalar.
[0039] In one embodiment, the method for verifying labeled data provided by the embodiments of the present invention may further include: Step 500: Determine the scalar difference between the correct reward scalar and the incorrect reward scalar; Step 600: Input the scalar difference into a loss function to obtain the scalar loss output by the loss function; the loss function is constructed based on the negative correlation between the scalar difference and the scalar loss.
[0040] Specifically, this embodiment mainly introduces the calculation of the loss function of the reward model to be trained. In order to enable the reward model to be trained to better distinguish between correct and incorrect labels and to make the gap between the correct reward scalar and the incorrect reward scalar obvious, the present invention proposes a loss function as shown in Formulas 1 and 2.
[0041] ; (1) ; (2) Wherein, is the loss value calculated by the loss function proposed by the present invention, that is, the scalar loss in this embodiment; is the correct reward scalar; is the incorrect reward scalar; is the scalar difference between the correct reward scalar and the incorrect reward scalar, that is ; represents the natural base of power. It can be seen from the above Formulas 1 and 2 that in the loss function proposed by the present invention, there is a negative correlation between the scalar difference and the scalar loss , and the larger the scalar difference, the smaller the scalar loss. By minimizing the scalar loss of the reward model to be trained, the reward model can better distinguish between correct and incorrect labels.
[0042] In this embodiment, by using a loss function in which the scalar difference and the scalar loss have a negative correlation, the reward model can better distinguish between correct and incorrect labels.
[0043] In one embodiment, the method for verifying labeled data provided by the embodiments of the present invention may further include: Step 410: Construct a guiding instruction based on the user query data and the pre-labeled data to be verified; Step 420: Input the guiding instruction into the trained reward model to obtain the normalized score output by the trained reward model; Step 430: Determine the pre-annotation test result of the pre-annotated data to be tested based on the normalized score.
[0044] Specifically, through the loss function proposed in the present invention, the reward model to be trained is trained so that the trained reward model can better distinguish correct annotations and incorrect annotations. After the reward model meets the convergence requirements or the scalar loss reaches the expectation, the trained reward model is obtained. Then, the user query data and the pre-annotated data to be tested (automatically pre-annotated from the user query data) are input into the trained reward model to obtain a score output by the trained reward model, and this score is normalized to obtain the normalized score in this embodiment. Whether the pre-annotated data to be tested is standard and accurate is judged by a preset score threshold (for example, 0.5). If the normalized score output by the trained reward model is greater than the above score threshold, it is determined that the pre-annotated data to be tested is accurately annotated; otherwise, it is determined that the pre-annotated data to be tested is incorrectly annotated.
[0045] In this embodiment, by inputting the user query data and the corresponding pre-annotated data into the trained reward model and analyzing the normalized score output by the trained reward model, it can be judged whether the pre-annotated data to be tested is standard and accurate, realizing the automatic detection of the pre-annotated data, reducing the workload of manual inspection, and improving the inspection efficiency of the annotated data.
[0046] Next, the annotation data inspection device provided by the present invention will be described. The annotation data inspection device described below can be mutually referred to the annotation data inspection method described above.
[0047] Please refer to Figure 3 , the present invention also provides an annotation data inspection device, including: A training data construction module 301, configured to construct training data based on historical query data and historical annotation data; A reward scalar obtaining module 302, configured to input the prompt instruction generated based on the training data into the reward model to be trained to obtain the reward scalar output by the reward model; A reward model training module 303, configured to train the reward model based on the scalar loss; the scalar loss is determined based on the reward scalar; A pre-annotation inspection module 304, configured to input the user query data and the pre-annotated data to be tested into the trained reward model to obtain the pre-annotation inspection result.
[0048] Optionally, the historical annotation data includes correct annotation data and historical pre-annotation data; the training data construction module includes: A labeled data conversion unit, configured to convert the historical pre-labeled data into mislabeled data when the historical pre-labeled data is correctly labeled; A mislabeled data determination unit, configured to use the historical pre-labeled data as mislabeled data when the historical pre-labeled data is mislabeled; A data triple construction unit, configured to construct data triples based on historical query data, the correctly labeled data, and the mislabeled data; A training data construction unit, configured to construct training data based on the data triples.
[0049] Optionally, the labeled data verification device further includes: A correct instruction generation module, configured to generate a correct instruction based on the concatenation result of the correctly labeled data and the historical query data; A prompt instruction generation module, configured to generate a prompt instruction based on the correct instruction and an incorrect instruction; the incorrect instruction is generated based on the concatenation result of the mislabeled data and the historical query data.
[0050] Optionally, the reward scalar obtaining module includes: A hidden state acquisition unit, configured to acquire a first hidden state of the minimum segmentation unit corresponding to the correct instruction and a second hidden state of the minimum segmentation unit corresponding to the incorrect instruction; A reward scalar obtaining unit, configured to input the first hidden state into a linear layer to obtain a correct reward scalar, and input the second hidden state into a linear layer to obtain an incorrect reward scalar.
[0051] Optionally, the labeled data verification device further includes: A scalar difference determination module, configured to determine the scalar difference between the correct reward scalar and the incorrect reward scalar; A scalar loss obtaining module, configured to input the scalar difference into a loss function to obtain a scalar loss output by the loss function.
[0052] Optionally, the pre-label verification module includes: A guiding instruction construction unit, configured to construct a guiding instruction based on user query data and pre-label data to be verified; A normalized score obtaining unit, configured to input the guiding instruction into a trained reward model to obtain a normalized score output by the trained reward model; A pre-label verification result determination unit, configured to determine a pre-label verification result of the pre-label data to be verified based on the normalized score.
[0053] Figure 4Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 4 shown. The electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute the labeled data verification method, which includes: constructing training data based on historical query data and historical labeled data; inputting the hint instructions generated based on the training data into the reward model to be trained, and obtaining the reward scalar output by the reward model; training the reward model based on the scalar loss; the scalar loss is determined based on the reward scalar; inputting the user query data and the pre-labeled data to be verified into the trained reward model, and obtaining the pre-labeled verification result.
[0054] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs, and other various media that can store program codes.
[0055] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the labeled data verification method provided by the above-mentioned various methods, which includes: constructing training data based on historical query data and historical labeled data; inputting the hint instructions generated based on the training data into the reward model to be trained, and obtaining the reward scalar output by the reward model; training the reward model based on the scalar loss; the scalar loss is determined based on the reward scalar; inputting the user query data and the pre-labeled data to be verified into the trained reward model, and obtaining the pre-labeled verification result.
[0056] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the annotation data verification method provided by the above-mentioned various methods. The method includes: constructing training data based on historical query data and historical annotation data; inputting a prompt instruction generated based on the training data into a reward model to be trained, and obtaining a reward scalar output by the reward model; training the reward model based on a scalar loss; the scalar loss is determined based on the reward scalar; inputting user query data and pre-annotation data to be verified into the trained reward model to obtain a pre-annotation verification result.
[0057] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0058] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for verifying labeled data, characterized in that: include: Build training data based on historical query data and historical annotation data; Inputting the prompt instruction generated based on the training data into the reward model to be trained to obtain a reward scalar output by the reward model; The reward model is trained based on a scalar loss; the scalar loss is determined based on the reward scalar; The user query data and the pre-labeled data to be tested are input into the trained reward model to obtain the pre-labeled test results.
2. The method for verifying labeled data according to claim 1, characterized in that: The historical annotated data includes correct annotated data and historical pre-annotated data; The constructing of training data based on historical query data and historical annotation data includes: In the case where the historical pre-annotated data is correctly annotated, converting the historical pre-annotated data into incorrectly annotated data; In the case where the historical pre-labeled data is incorrectly labeled, treating the historical pre-labeled data as incorrectly labeled data; Constructing data triples based on the historical query data, the correctly labeled data, and the incorrectly labeled data; Training data is constructed based on the data triples.
3. The method for verifying labeled data according to claim 2, characterized in that: The step of inputting the prompt instruction generated based on the training data into the reward model to be trained to obtain the reward scalar output by the reward model comprises: Generate a correct instruction based on the splicing result of the correct labeled data and the historical query data; A prompt instruction is generated based on the correct instruction and the error instruction; the error instruction is generated based on the splicing result of the error marking data and the historical query data.
4. The method for verifying labeled data according to claim 3, characterized in that: The step of inputting the prompt instruction generated based on the training data into the reward model to be trained to obtain the reward scalar output by the reward model includes: Acquire a first hidden state of the minimum segmentation unit corresponding to the correct instruction, and acquire a second hidden state of the minimum segmentation unit corresponding to the incorrect instruction; The first hidden state is connected to the linear layer to obtain a correct reward scalar, and the second hidden state is connected to the linear layer to obtain an incorrect reward scalar.
5. The method for verifying labeled data according to claim 4, characterized in that: The labeled data verification method further includes: determining a scalar difference between the correct reward scalar and the incorrect reward scalar; The scalar difference is input into a loss function to obtain a scalar loss output by the loss function; the loss function is constructed based on the negative correlation between the scalar difference and the scalar loss.
6. The method for verifying labeled data according to claim 1, characterized in that: The step of inputting the user query data and the pre-labeled data to be verified into the trained reward model to obtain the pre-labeled verification result includes: Building guidance instructions based on user query data and pre-annotated data to be inspected; Inputting the guidance instruction into the trained reward model to obtain a normalized score output by the trained reward model; Based on the normalized score, a pre-labeled inspection result of the pre-labeled data to be inspected is determined.
7. A labeling data verification device, characterized in that: include: A training data construction module is used to construct training data based on historical query data and historical annotation data; A reward scalar obtaining module, used for inputting the prompt instruction generated based on the training data into the reward model to be trained, and obtaining the reward scalar output by the reward model; A reward model training module, configured to train the reward model based on a scalar loss; the scalar loss is determined based on the reward scalar; The pre-labeling verification module is used to input the user query data and the pre-labeled data to be verified into the trained reward model to obtain the pre-labeling verification results.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the labeled data verification method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the labeled data verification method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the labeled data verification method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Large language model training method and device and text processing method and device
CN117149989A
Large language model training method and device
CN118036757A
Data labeling method, device and equipment
CN118228047A
Reinforcement learning alignment model training method and system based on AI feedback
CN118735002A
Text data generation method and device, equipment and medium
CN119003694A