Multi-Turn Dialog Dataset Generation Without Human Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing dialog systems face challenges in generating sufficient training data for conversational agents, especially when dealing with multiple agents having different specialties or functions, which typically requires costly and time-consuming human interactions.

Innovation Solution

A system that automatically selects an agent, calculates dialog tree nodes, generates random dialog nodes, and creates multi-turn conversational responses to build datasets without human input, incorporating conversational properties like digressions, disambiguation, and slot filling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If human interactions are used to generate training data for multiple agents, then the quality and diversity of training data improve, but the cost and time required increase significantly

Engineering Contradiction:
Improvetraining dataVSAvoidtime-consuming
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system creates synthetic training data by copying and adapting existing dialog patterns and structures. It generates artificial conversation trajectories that mimic real human interactions without requiring actual human participants, thus producing sufficient training data while eliminating the time-consuming aspect of human-in-the-loop data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The dialog system performs self-service by automatically generating its own training data through simulated conversations. The system uses its own dialog models and policies to create training examples, eliminating the need for external human annotators and significantly reducing the time required for data generation

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If human interactions are used to generate training data for multiple agents, then the quality and diversity of training data improve, but the cost increases

Engineering Contradiction:
Improvetraining dataVSAvoidcostly
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The system replicates diverse dialog patterns and conversation flows through synthetic generation, copying successful interaction patterns from existing data to create varied training examples. This approach achieves data diversity without the high costs associated with recruiting and managing human annotators for multiple agents

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system generates its own training data autonomously using automated dialog simulations, eliminating the need for expensive human labor. The self-service approach allows the system to produce sufficient training data for multiple agents at minimal cost while maintaining data quality through controlled generation processes

Inventive Principle:
Principle #25Self-service

3Productivity

If automated methods are used to generate training data, then the cost and time required decrease, but the conversational properties and realism may be insufficient

Engineering Contradiction:
Improvegeneration efficiencyVSAvoidconversational quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system copies authentic conversational patterns, turn-taking structures, and dialog flow characteristics from real human interactions. By replicating these proven patterns in synthetic data, the system maintains high conversational quality and realism while achieving fast, cost-effective generation through automated processes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses self-service automated generation with sophisticated dialog models that inherently understand conversational nuances. The automated process incorporates conversational properties like digressions, disambiguation, and slot filling through programmatic implementation, ensuring both high productivity and reliable conversational quality

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12451141B2Generating multi-turn dialog datasets
Publication Date: 2025.10.21 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12451141B2 patent drawing
  • US12451141B2 patent drawing
  • US12451141B2 patent drawing

AI summary

An embodiment for generating multi-turn dialog datasets for training of dialog or conversational agents. The embodiment may select an agent from a set of agents. The embodiment may automatically identify sentences from training data of the selected agent that satisfy a first sequential node condition of the selected random dialog node. The embodiment may automatically determine an approach for responding to the first sequential node condition of the selected random dialog node that either satisfies the first sequential dialog node condition, or inserts a multi-turn conversational property, and generate a corresponding response. The embodiment may automatically determine additional approaches for responding to each condition within subsequent sequential child nodes of the selected random dialog node that either satisfy each subsequent sequential child node condition or insert a multi-turn conversational property, and generate corresponding responses. The embodiment may collect and store data relating to the selected agent and the generated responses.