Large Language Model Training via Synthetic Data Ranking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language processing systems face challenges in efficiently training machine learning models to interpret and resolve ambiguities in text, particularly with open-textured terms, due to the scarcity of large-scale datasets and the need for human-generated data, which increases training costs and limits model performance.

Innovation Solution

The proposed solution involves using large language models to generate and filter responses, iteratively training models based on ranked arguments, and fine-tuning them using synthetic datasets, allowing for the automation of text interpretation and ambiguity resolution through a framework that includes prompting templates and machine learning classifiers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human-generated data is used for training, then model performance improves, but training costs increase and data scarcity limits performance

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent uses synthetic data generated by language models as copies of human-generated training data. These synthetic datasets replicate the structure and characteristics of human-annotated data without requiring actual human annotation, thereby reducing costs while maintaining model performance through iterative generation and filtering processes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The training system becomes self-sufficient by using the language model itself to generate training data. The model generates synthetic responses, self-evaluates them through ranking mechanisms, and iteratively improves without external human intervention for data creation, enabling continuous self-improvement cycles

Inventive Principle:
Principle #25Self-service

2Productivity

If synthetic datasets are used for training, then training efficiency improves, but data quality may decrease

Engineering Contradiction:
Improvetraining efficiencyVSAvoiddata quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent implements feedback loops where generated synthetic data is evaluated through ranking mechanisms, and the results feed back into improving subsequent data generation. The system continuously refines its synthetic data production based on performance metrics and ranking outcomes, ensuring quality improvement over iterations

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary filtering and ranking of synthetic data before actual model training. By pre-processing and quality-assuring the synthetic datasets through automated ranking mechanisms beforehand, the system ensures high data quality is maintained while preserving training efficiency

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If large-scale datasets are required, then model accuracy improves, but data scarcity and processing time increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates synthetic copies of large-scale datasets through automated generation processes. Instead of collecting and processing extensive human-annotated data, the system generates synthetic datasets that replicate the scale and characteristics needed for high-accuracy training, dramatically reducing data collection and preprocessing time

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary generation and organization of synthetic datasets in advance, creating ready-to-use training data structures before actual model training begins. This pre-processing of synthetic data eliminates time-consuming data collection and manual annotation steps while maintaining the scale needed for accurate models

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240403710A1Systems and methods for interpreting language using ai models
Publication Date: 2024.12.05 UNIV OF SOUTH FLORIDA
  • US20240403710A1 patent drawing
  • US20240403710A1 patent drawing
  • US20240403710A1 patent drawing

AI summary

Systems and methods for artificial intelligence (AI)-based systems are described herein. In an aspect, the present disclosure relates to a computer implemented method that includes prompting a first trained large language model (LLM) to generate a plurality of arguments; determine a ranking of the plurality of arguments using a second trained LLM; and training a third LLM based on the ranking of the plurality of arguments and the plurality of arguments.