Virtual Editors for Virtual Assistant Ground Truth Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for creating ground truth data sets for virtual assistants are inefficient, relying on manual processing by human editors, which limits scalability and raises data protection concerns, and can lead to overfitting due to incomplete or incorrect data.
Innovation Solution
An automated method using multiple virtual editors trained on human editor inputs to analyze and classify user requests, with a filtering process to ensure data quality and relevance, allowing for the creation of a high-quality ground truth data set that is not scaled by the size of the editing team.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human editors manually process user requests to create ground truth datasets, then data quality and accuracy are improved, but processing scalability is limited and data protection risks increase
Solution Approach 1:
The patent creates virtual editors that are copies of human editors' expertise, trained on annotated data from human editors. These virtual editors can process unlimited numbers of requests without the scalability constraints of human teams, while maintaining the quality standards through multiple virtual editors working in parallel
Solution Approach 2:
The patent replaces the mechanical manual processing system with an automated AI-based system. Virtual editors use machine learning models to automatically annotate and classify user requests, eliminating the need for human editors to manually process each request while maintaining annotation quality
2Productivity
If more human editors are hired to process more requests, then processing capacity increases, but data protection risks and costs increase
Solution Approach 1:
Instead of hiring more human editors, the system creates multiple virtual editors that can process requests in parallel. These virtual editors are trained models that can be replicated indefinitely without the data protection and cost implications of employing additional human staff
Solution Approach 2:
The virtual editors process requests autonomously without human intervention. The system self-manages the annotation process, eliminating the need for human editors to handle user data, thereby reducing data protection risks while maintaining processing capacity
3Productivity
If simple automation is used to process requests, then processing speed increases, but data accuracy decreases due to incorrect classifications
Solution Approach 1:
The system implements feedback mechanisms where multiple virtual editors review and validate each other's work. The evaluation unit provides feedback on annotation quality, allowing the system to correct mistakes and maintain high accuracy while processing requests at automated speed
Solution Approach 2:
The virtual editors are trained copies of human editors' expertise, inheriting their accuracy standards. Multiple virtual editors work in parallel, each trained on high-quality annotated data, maintaining the accuracy level of human editors while operating at machine processing speeds
Data Source
Figure 1
Figure 2
AI summary
Techniques for generating a ground-truth dataset for a virtual assistant, wherein the elements of the ground-truth dataset comprise possible user requests to the virtual assistant, wherein the method is at least partially executable on a computing unit and comprises the following steps: • Providing user requests to the computing unit, wherein these are requests intended for, or were intended for, the virtual assistant; • Analyzing the requests with regard to their suitability for inclusion in the ground-truth data; wherein the analysis is performed by at least two virtual editors, each of the virtual editors producing a separate analysis result for each request;• Implementing means for evaluating the analysis results on the computing unit, wherein the evaluation means calculate a value TS representing a probability that the request is suitable for the ground-truth dataset, the calculation being performed based on the analysis results of at least two virtual editors for the respective request; • Comparing the calculated value TS with a predefined threshold TH; • Adding the request to the ground-truth dataset if the calculated value TS is greater than the threshold TH.