Conversational Model Training via User Feedback Loops
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current conversational models rely heavily on expert-tagged training data, which is costly and inefficient, and fail to fully utilize user behavioral signals, leading to a mismatch between model performance improvement and user experience enhancement.
Innovation Solution
A method for processing query-response information that generates training samples by acquiring user feedback on initial responses, reducing dependence on expert tagging and aligning model performance with user preferences through conversational feedback, editing feedback, and AI feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If expert-tagged training data is used for conversational model training, then model performance can be improved, but tagging costs are high and efficiency is low
Solution Approach 1:
The system enables self-service by allowing the conversational model to automatically generate initial response information and by utilizing user feedback directly for training sample generation, eliminating the need for expert tagging of all training data
Solution Approach 2:
The system introduces feedback loops where user interactions and preferences are collected and fed back into the training process. User feedback on initial responses is used to generate training samples, creating a continuous improvement cycle that reduces dependency on expert tagging
2Reliability
If expert-tagged training data is used for conversational model training, then model performance can be improved, but tagging costs are high
Solution Approach 1:
The system replaces expensive expert-tagged data with cheaper alternatives: automatically generated initial responses and user feedback data. These substitutes are sufficient for training purposes and significantly reduce tagging costs while maintaining model performance improvement
Solution Approach 2:
The system enables self-service by allowing the conversational model to automatically generate initial response information and by utilizing user feedback directly for training sample generation, eliminating the need for expert tagging of all training data
3Reliability
If traditional training methods are used, then model training can proceed, but user behavioral signals are not fully utilized leading to mismatch between model performance improvement and user experience enhancement
Solution Approach 1:
The system introduces feedback loops where user interactions and preferences are collected and fed back into the training process. User feedback on initial responses is used to generate training samples, creating a continuous improvement cycle that reduces dependency on expert tagging
Solution Approach 2:
The system makes the training process dynamic by continuously incorporating user feedback and behavioral signals. The training data generation adapts to user preferences in real-time, allowing the model to evolve alongside changing user expectations
Data Source
AI summary
A method for processing a query-response information is provided, which relates to a field of artificial intelligence technology, and in particular to fields of deep learning, large models, intelligent query and response, etc. The method for processing a query-response information includes: generating at least one initial response information according to a query information provided by an object; acquiring at least one feedback information corresponding to the at least one initial response information, wherein the feedback information indicates a preference degree of the object for the initial response information; and generating a training sample according to the query information, the at least one initial response information and the at least one feedback information. The present disclosure further provides a method for training a conversational model, an electronic device, and a storage medium.


