Conversational Training Data Transformation via Language Pattern Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current natural language systems face inefficiencies in processing human requests due to unproductive and irrelevant responses, largely attributed to the presence of duplicates and unrelated data in training datasets derived from methods like deep learning neural networks and web content crawls.
Innovation Solution
A system and method for transforming conversational training data by analyzing language patterns between different knowledge areas using a third machine learning model, allowing for the substitution of segments to create relevant training data for different natural language models, thereby enhancing dialogue consistency and response relevance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If training data is derived from web content crawls and deep learning neural networks, then the quantity of training data is increased, but the relevance and productivity of responses deteriorate due to duplicates and unrelated data
Solution Approach 1:
The patent segments the training data transformation process into distinct stages: analyzing language patterns using a third ML model, identifying segments for substitution, and transforming statements. This segmentation allows for targeted processing that maintains data quantity while improving quality through selective transformation of specific language segments across different knowledge areas.
Solution Approach 2:
The patent changes the parameter of language representation by transforming statements from one knowledge area to another using a third ML model that recognizes patterns of language variation. This parameter change enables the same training data to be adapted for different domains, increasing relevance without requiring additional data collection.
2Quantity of substance
If training data includes duplicates and unrelated content from web crawls, then the volume of training data is increased, but the consistency of dialogue deteriorates
Solution Approach 1:
The patent applies preliminary action by transforming training data statements before they are used to train ML models. The third ML model analyzes language patterns and performs substitutions in advance, ensuring that the training data is already optimized for dialogue consistency before model training begins, preventing inconsistency from propagating through the system.
Solution Approach 2:
The third machine learning model serves as an intermediary between the raw training data and the target ML models. It recognizes patterns of language variation and transforms statements accordingly, acting as a mediator that ensures consistency across different knowledge areas while preserving the utility of the training data.
3Adaptability or versatility
If general training data is used across different knowledge areas, then the versatility of the training set is increased, but the relevance to specific domains deteriorates
Solution Approach 1:
The patent applies local quality by transforming specific segments of training data according to the target knowledge area. The third ML model identifies and substitutes only the necessary language segments, preserving domain-specific characteristics while adapting the overall statement. This allows the training data to be locally optimized for each domain while maintaining global versatility.
Solution Approach 2:
The patent achieves universality through the third ML model, which is trained to recognize patterns of language variation across multiple knowledge areas. This single model serves multiple functions by adapting training data for different domains, enabling one training set to be versatile across multiple knowledge areas while maintaining domain-specific relevance through pattern recognition and substitution.
Data Source
AI summary
A method includes receiving a command to transform a first statement in a first conversational training set into training data for a second training set, where the first conversational training set trains a first machine learning model in a first knowledge area, and where the second conversational training set trains a second machine learning model in a second knowledge area. The method also includes analyzing language of the first statement using a third machine learning model, where the third machine learning model is trained to recognize patterns of language variation between the first and second knowledge areas. The method also includes, responsive to the analyzing, selecting an original segment of the first statement. The method also includes, responsive to the selecting, transforming the first statement into a second statement, for the second conversational training set, that includes a substitute segment that at least partially replaces the original segment.


