Virtual Editors for Virtual Assistant Ground Truth Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for creating ground truth data sets for virtual assistants are inefficient, relying on manual processing by human editors, which limits scalability and raises data protection concerns, and can lead to overfitting due to incomplete or incorrect data.

Innovation Solution

An automated method using multiple virtual editors trained on human editor inputs to analyze and classify user requests, with a filtering process to ensure data quality and relevance, allowing for the creation of a high-quality ground truth data set that is not scaled by the size of the editing team.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human editors manually process user requests to create ground truth datasets, then data quality and accuracy are improved, but processing scalability is limited and data protection risks increase

Engineering Contradiction:
Improvedata qualityVSAvoidprocessing scalability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates virtual editors that are copies of human editors' expertise, trained on annotated data from human editors. These virtual editors can process unlimited numbers of requests without the scalability constraints of human teams, while maintaining the quality standards through multiple virtual editors working in parallel

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical manual processing system with an automated AI-based system. Virtual editors use machine learning models to automatically annotate and classify user requests, eliminating the need for human editors to manually process each request while maintaining annotation quality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If more human editors are hired to process more requests, then processing capacity increases, but data protection risks and costs increase

Engineering Contradiction:
Improveprocessing capacityVSAvoiddata protection risks
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

Instead of hiring more human editors, the system creates multiple virtual editors that can process requests in parallel. These virtual editors are trained models that can be replicated indefinitely without the data protection and cost implications of employing additional human staff

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The virtual editors process requests autonomously without human intervention. The system self-manages the annotation process, eliminating the need for human editors to handle user data, thereby reducing data protection risks while maintaining processing capacity

Inventive Principle:
Principle #25Self-service

3Productivity

If simple automation is used to process requests, then processing speed increases, but data accuracy decreases due to incorrect classifications

Engineering Contradiction:
Improveprocessing speedVSAvoiddata accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements feedback mechanisms where multiple virtual editors review and validate each other's work. The evaluation unit provides feedback on annotation quality, allowing the system to correct mistakes and maintain high accuracy while processing requests at automated speed

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The virtual editors are trained copies of human editors' expertise, inheriting their accuracy standards. Multiple virtual editors work in parallel, each trained on high-quality annotated data, maintaining the accuracy level of human editors while operating at machine processing speeds

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4036909B1Method and data generator for generating a base data set for a virtual assistant
Publication Date: 2024.12.04 DEUTSCHE TELEKOM AG
  • EP4036909B1 patent drawingFigure 1
  • EP4036909B1 patent drawingFigure 2
  • EP4036909B1 patent drawing

AI summary

Techniques for generating a ground-truth dataset for a virtual assistant, wherein the elements of the ground-truth dataset comprise possible user requests to the virtual assistant, wherein the method is at least partially executable on a computing unit and comprises the following steps: • Providing user requests to the computing unit, wherein these are requests intended for, or were intended for, the virtual assistant; • Analyzing the requests with regard to their suitability for inclusion in the ground-truth data; wherein the analysis is performed by at least two virtual editors, each of the virtual editors producing a separate analysis result for each request;• Implementing means for evaluating the analysis results on the computing unit, wherein the evaluation means calculate a value TS representing a probability that the request is suitable for the ground-truth dataset, the calculation being performed based on the analysis results of at least two virtual editors for the respective request; • Comparing the calculated value TS with a predefined threshold TH; • Adding the request to the ground-truth dataset if the calculated value TS is greater than the threshold TH.