Reward Model Training Data Selection for Accurate Dialogue Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The accuracy and generalization of reward models in task-oriented dialogue generation technologies are compromised due to poor annotation accuracy in the training data sets, leading to suboptimal reinforcement learning outcomes.

Innovation Solution

A method for determining a training data set of a large reward model involves obtaining candidate question texts and answer requirements, scoring candidate answer texts, selecting target answer texts based on scoring data, and constructing the training data set to improve the accuracy and generalization of the reward model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is used to obtain training data set for reward model, then the process is simple and fast, but the annotation accuracy is poor leading to low model accuracy and generalization

Engineering Contradiction:
Improveannotation accuracyVSAvoiddata construction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a large language model as an intermediary to assist in the annotation process. The LLM generates candidate answer texts and scoring data, which then undergoes multi-dimensional verification including consistency checks, format validation, and quality assessment before being used to train the reward model. This intermediary system bridges the gap between simple automated processes and high accuracy requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements preliminary actions by pre-processing candidate answer texts through multiple validation steps before they are used for training. The system performs preliminary consistency verification, format checking, and quality assessment on the annotated data before it enters the reward model training pipeline, ensuring high accuracy is achieved before the actual training begins.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If poor quality training data is used, then the data construction process is fast and simple, but the reward model accuracy and generalization deteriorate causing suboptimal reinforcement learning outcomes

Engineering Contradiction:
Improvereward model accuracyVSAvoiddata construction time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements self-service through automated consistency verification and quality assessment mechanisms. The annotation system automatically checks for consistency between candidate answers and scoring data, validates formats, and assesses quality without requiring extensive manual intervention. This self-service approach maintains high data quality while reducing the time investment required for data construction.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where the system evaluates the quality of annotated data through multiple verification steps and uses this feedback to refine the annotation process. The scoring data and candidate answer texts undergo feedback loops including consistency checks and quality assessments, allowing the system to automatically identify and correct issues without extensive manual review, thus maintaining reliability while reducing time loss.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If multiple verification steps are added to improve data quality, then the annotation accuracy improves, but the processing time and system complexity increase

Engineering Contradiction:
Improvedata qualityVSAvoiddata processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the data verification process into distinct modular steps: consistency verification, format validation, and quality assessment. Each segment handles a specific aspect of data quality control independently, allowing the system to process data through multiple verification steps without creating a bottleneck. This segmentation maintains high data quality while preserving processing efficiency by avoiding redundant operations in each step.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250371365A1Method for determining training data set of large reward model, and electronic device
Publication Date: 2025.12.04 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20250371365A1 patent drawing
  • US20250371365A1 patent drawing
  • US20250371365A1 patent drawing

AI summary

The present disclosure provides a method and an apparatus for determining a training data set of a large reward model, and an electronic device, which relates to the technical field of artificial intelligence, and in particular to the technical fields of deep learning, natural language processing, and large models etc. The specific implementation includes: obtaining a candidate question text, and an answer requirement corresponding to the candidate question text; determining, based on the candidate question text and the answer requirement corresponding to the candidate question text, at least one candidate answer text corresponding to the candidate question text and scoring data of the at least one candidate answer text; selecting, based on the scoring data of the at least one candidate answer text, a target answer text from the at least one candidate answer text; and constructing, based on scoring data of the target answer text and a candidate question text corresponding to the target answer text, the training data set of the large reward model, for training the large reward model. The training data set that is configured for training the large reward model is generated by the electronic device based on the candidate question text and the corresponding answer requirement, resulting in the high accuracy. Thus, the accuracy and generalization of the trained large reward model are improved, and the accuracy of the dialogue model obtained by reinforcement learning based on the large reward model is also improved.