Conversational Training Data Transformation via Language Pattern Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current natural language systems face inefficiencies in processing human requests due to unproductive and irrelevant responses, largely attributed to the presence of duplicates and unrelated data in training datasets derived from methods like deep learning neural networks and web content crawls.

Innovation Solution

A system and method for transforming conversational training data by analyzing language patterns between different knowledge areas using a third machine learning model, allowing for the substitution of segments to create relevant training data for different natural language models, thereby enhancing dialogue consistency and response relevance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If training data is derived from web content crawls and deep learning neural networks, then the quantity of training data is increased, but the relevance and productivity of responses deteriorate due to duplicates and unrelated data

Engineering Contradiction:
Improvequantity of training dataVSAvoidresponse productivity
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent segments the training data transformation process into distinct stages: analyzing language patterns using a third ML model, identifying segments for substitution, and transforming statements. This segmentation allows for targeted processing that maintains data quantity while improving quality through selective transformation of specific language segments across different knowledge areas.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of language representation by transforming statements from one knowledge area to another using a third ML model that recognizes patterns of language variation. This parameter change enables the same training data to be adapted for different domains, increasing relevance without requiring additional data collection.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If training data includes duplicates and unrelated content from web crawls, then the volume of training data is increased, but the consistency of dialogue deteriorates

Engineering Contradiction:
Improvevolume of training dataVSAvoiddialogue consistency
Core Design Contradiction:
Quantity of substanceVSStability of the object's composition

Solution Approach 1:

The patent applies preliminary action by transforming training data statements before they are used to train ML models. The third ML model analyzes language patterns and performs substitutions in advance, ensuring that the training data is already optimized for dialogue consistency before model training begins, preventing inconsistency from propagating through the system.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The third machine learning model serves as an intermediary between the raw training data and the target ML models. It recognizes patterns of language variation and transforms statements accordingly, acting as a mediator that ensures consistency across different knowledge areas while preserving the utility of the training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If general training data is used across different knowledge areas, then the versatility of the training set is increased, but the relevance to specific domains deteriorates

Engineering Contradiction:
Improveversatility of training setVSAvoiddomain relevance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies local quality by transforming specific segments of training data according to the target knowledge area. The third ML model identifies and substitutes only the necessary language segments, preserving domain-specific characteristics while adapting the overall statement. This allows the training data to be locally optimized for each domain while maintaining global versatility.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent achieves universality through the third ML model, which is trained to recognize patterns of language variation across multiple knowledge areas. This single model serves multiple functions by adapting training data for different domains, enabling one training set to be versatile across multiple knowledge areas while maintaining domain-specific relevance through pattern recognition and substitution.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12136043B1Transforming conversational training data for different machine learning models
Publication Date: 2024.11.05 LIKEHUMAN LLC
  • US12136043B1 patent drawing
  • US12136043B1 patent drawing
  • US12136043B1 patent drawing

AI summary

A method includes receiving a command to transform a first statement in a first conversational training set into training data for a second training set, where the first conversational training set trains a first machine learning model in a first knowledge area, and where the second conversational training set trains a second machine learning model in a second knowledge area. The method also includes analyzing language of the first statement using a third machine learning model, where the third machine learning model is trained to recognize patterns of language variation between the first and second knowledge areas. The method also includes, responsive to the analyzing, selecting an original segment of the first statement. The method also includes, responsive to the selecting, transforming the first statement into a second statement, for the second conversational training set, that includes a substitute segment that at least partially replaces the original segment.