Locale-Specific Language Model Training via Teacher Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating a language model for a new location are inefficient, as they require manual user input, rely on large datasets that do not account for local nuances, and are unsuitable for low-memory devices, leading to slow training times and incomplete models.

Innovation Solution

A system that uses a language update application with input/output and control circuitry to generate a student language model based on user-provided parameters, comparing data from a teacher language model to the student model to determine necessary updates, reducing training time and improving model relevance to the user's location.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a student language model is trained using all data from teacher language models, then the model completeness improves, but the training time and memory requirements increase significantly

Engineering Contradiction:
Improvemodel completenessVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts only the necessary subset of data from teacher language models that is relevant to the target locale, rather than transferring all data. This is achieved by comparing the teacher model data with existing student model data and identifying gaps specific to the target locale, thereby reducing training time while maintaining model completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by tailoring the data transfer process to the specific needs of each locale. Instead of uniform data transfer, the system identifies and transfers only the locale-specific data portions from teacher models, making the training process efficient and targeted to local requirements.

Inventive Principle:
Principle #3Local quality

2Reliability

If a student language model is trained using all data from teacher language models, then the model completeness improves, but the memory requirements increase significantly

Engineering Contradiction:
Improvemodel completenessVSAvoidmemory requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system extracts and transfers only the essential subset of data from teacher language models that is necessary for the target locale, rather than transferring the entire dataset. This extraction process identifies and eliminates redundant data, reducing memory requirements while preserving model completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the teacher model data into locale-specific and non-locale-specific portions, transferring only the relevant segments to the student model. This segmentation approach reduces the overall data volume that needs to be stored in memory while maintaining the necessary model completeness.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If manual user input is used to train a language model, then the model accuracy for specific use cases improves, but the training time increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training teacher language models on comprehensive datasets before deploying them to guide student model training. This pre-computed knowledge is then efficiently transferred to student models, reducing the time required for manual user input while maintaining or improving model accuracy for specific use cases.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses teacher language models as intermediaries between comprehensive language data and the target student model. Instead of relying solely on manual user input, the teacher models serve as mediators that provide pre-processed, locale-relevant training data, thereby improving accuracy while reducing training time.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If large pools of data are used to train a language model, then the model completeness improves, but the relevance to specific locale nuances decreases

Engineering Contradiction:
Improvemodel completenessVSAvoidlocale relevance
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent implements local quality by making the data transfer process locale-specific. The system identifies and transfers only the portions of teacher model data that are relevant to the target locale's linguistic nuances and characteristics, ensuring both model completeness and locale relevance rather than using uniform large-scale data pools.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system extracts and transfers only the locale-relevant subset of data from teacher language models, removing unnecessary data that does not contribute to specific locale nuances. This extraction process maintains model completeness while enhancing locale-specific relevance.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12039265B2Systems and methods to support a new locale in a language model
Publication Date: 2024.07.16 ADEIA GUIDES INC
  • US12039265B2 patent drawing
  • US12039265B2 patent drawing
  • US12039265B2 patent drawing

AI summary

Systems and methods are presented herein for generating a new language understanding model, based on a user request. A user may input a root language and a locale into an application for generating a student language model. The application may generate the student language model and may identify a teacher language model related to the student language model. The application may compare data from the identified teacher language model to the student language model. The application may determine a subset of data from the teacher language model is not contained in the student language model. If the application determines at least a subset of data from the teacher language model is not in the student language model, the application may add at least the subset of data from the teacher language model to the student language model.