Language Recognition Model Training via Machine Translation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional language recognition models require extensive human labor and time to train, especially when transitioning between languages, due to the need for specialized skills and the generation of training data, making it cumbersome and inefficient.

Innovation Solution

A language processing engine that utilizes historical and machine-generated language data from a reference language to train a target language model, leveraging machine translation and quality checks to reduce human effort and improve accuracy, allowing for the adaptation of models to recognize inputs in multiple languages without requiring additional human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional techniques are used to train language recognition models for different languages, then model accuracy can be achieved, but extensive human labor and time are required

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-translating reference language training data into target language training data before model training. This allows the target language model to be trained using translated utterances and attributes, significantly reducing the time required for human trainers to manually create training data for each language while maintaining model accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating translated copies of reference language training data (utterances and attributes) in the target language. Instead of manually collecting and annotating new training data for each language, the system copies and translates existing high-quality training data, preserving the structure and quality while adapting it to different languages

Inventive Principle:
Principle #26Copying

2Measurement precision

If conventional techniques are used to train language recognition models, then models can recognize features accurately, but extensive human labor and specialized skills are required

Engineering Contradiction:
Improvefeature recognition accuracyVSAvoidhuman intervention level
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The patent implements self-service by enabling the system to automatically translate reference language training data into target language training data using machine translation. This eliminates the need for human trainers with specialized language skills to manually create training data, allowing the system to serve itself in generating multi-language training datasets

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses machine translation as an intermediary to bridge the reference language and target language. The translation component acts as a mediator that converts utterances and attributes from the reference language into the target language, enabling automated training data generation without requiring human trainers to possess specialized skills in multiple languages

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If language recognition models are trained for multiple languages using conventional methods, then model versatility improves, but device complexity increases

Engineering Contradiction:
Improvemulti-language capabilityVSAvoidmodel training complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies universality by creating a unified training framework that can handle multiple languages through a single reference language model and automated translation process. Instead of maintaining separate training pipelines for each language, the system uses one universal approach: translate reference data to target language and train, enabling multi-language capability without proportionally increasing system complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10854189B2Techniques for model training for voice features
Publication Date: 2020.12.01 AMAZON TECH INC
  • US10854189B2 patent drawing
  • US10854189B2 patent drawing
  • US10854189B2 patent drawing

AI summary

Techniques are provided for training a language recognition model. For example, a language recognition model may be maintained and associated with a reference language (e.g., English). The language recognition model may be configured to accept as input an utterance in the reference language and to identify a feature to be executed in response to receiving the utterance. New language data (e.g., other utterances) provided in a different language (e.g., German) may be obtained. This new language data may be translated to English and utilized to retrain the model to recognize reference language data as well as language data translated to the reference language. Subsequent utterances (e.g., English utterances, or German utterances translated to English) may be provided to the updated model and a feature may be identified. One or more instructions may be sent to a user device to execute a set of instructions associated with the feature.