Reward Model Training Data Selection for Accurate Dialogue Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy and generalization of reward models in task-oriented dialogue generation technologies are compromised due to poor annotation accuracy in the training data sets, leading to suboptimal reinforcement learning outcomes.
Innovation Solution
A method for determining a training data set of a large reward model involves obtaining candidate question texts and answer requirements, scoring candidate answer texts, selecting target answer texts based on scoring data, and constructing the training data set to improve the accuracy and generalization of the reward model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation is used to obtain training data set for reward model, then the process is simple and fast, but the annotation accuracy is poor leading to low model accuracy and generalization
Solution Approach 1:
The patent introduces a large language model as an intermediary to assist in the annotation process. The LLM generates candidate answer texts and scoring data, which then undergoes multi-dimensional verification including consistency checks, format validation, and quality assessment before being used to train the reward model. This intermediary system bridges the gap between simple automated processes and high accuracy requirements.
Solution Approach 2:
The patent implements preliminary actions by pre-processing candidate answer texts through multiple validation steps before they are used for training. The system performs preliminary consistency verification, format checking, and quality assessment on the annotated data before it enters the reward model training pipeline, ensuring high accuracy is achieved before the actual training begins.
2Reliability
If poor quality training data is used, then the data construction process is fast and simple, but the reward model accuracy and generalization deteriorate causing suboptimal reinforcement learning outcomes
Solution Approach 1:
The system implements self-service through automated consistency verification and quality assessment mechanisms. The annotation system automatically checks for consistency between candidate answers and scoring data, validates formats, and assesses quality without requiring extensive manual intervention. This self-service approach maintains high data quality while reducing the time investment required for data construction.
Solution Approach 2:
The patent incorporates feedback mechanisms where the system evaluates the quality of annotated data through multiple verification steps and uses this feedback to refine the annotation process. The scoring data and candidate answer texts undergo feedback loops including consistency checks and quality assessments, allowing the system to automatically identify and correct issues without extensive manual review, thus maintaining reliability while reducing time loss.
3Measurement precision
If multiple verification steps are added to improve data quality, then the annotation accuracy improves, but the processing time and system complexity increase
Solution Approach 1:
The patent segments the data verification process into distinct modular steps: consistency verification, format validation, and quality assessment. Each segment handles a specific aspect of data quality control independently, allowing the system to process data through multiple verification steps without creating a bottleneck. This segmentation maintains high data quality while preserving processing efficiency by avoiding redundant operations in each step.
Data Source
AI summary
The present disclosure provides a method and an apparatus for determining a training data set of a large reward model, and an electronic device, which relates to the technical field of artificial intelligence, and in particular to the technical fields of deep learning, natural language processing, and large models etc. The specific implementation includes: obtaining a candidate question text, and an answer requirement corresponding to the candidate question text; determining, based on the candidate question text and the answer requirement corresponding to the candidate question text, at least one candidate answer text corresponding to the candidate question text and scoring data of the at least one candidate answer text; selecting, based on the scoring data of the at least one candidate answer text, a target answer text from the at least one candidate answer text; and constructing, based on scoring data of the target answer text and a candidate question text corresponding to the target answer text, the training data set of the large reward model, for training the large reward model. The training data set that is configured for training the large reward model is generated by the electronic device based on the candidate question text and the corresponding answer requirement, resulting in the high accuracy. Thus, the accuracy and generalization of the trained large reward model are improved, and the accuracy of the dialogue model obtained by reinforcement learning based on the large reward model is also improved.


