Speech Recognition Training via LLM Prompt Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for expanding intent classification models in speech recognition systems require manual tagging and are time-consuming, costly, and prone to inconsistent results due to human evaluators. Additionally, these systems struggle to detect misclassifications and new intent patterns in changing user behaviors.

Innovation Solution

A method and apparatus for automatically training a speech recognition system by obtaining NLU results, generating prompts for a large-scale language model, determining the appropriateness of these results, and using the determination to train the system. This approach enables the detection of misclassifications and derivation of new intent without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tagging and human evaluation are used to expand intent classification models, then the system can detect misclassifications and new intent patterns, but the process becomes time-consuming and costly

Engineering Contradiction:
Improvedetection accuracy of misclassifications and new intentVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The speech recognition system automatically evaluates its own NLU results using a large-scale language model, eliminating the need for human evaluators. The system self-identifies misclassifications and derives new intent patterns through automated prompt generation and evaluation, making the training process independent of manual intervention while maintaining high detection accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical human evaluation process with an automated computational system. A large-scale language model generates prompts and evaluates NLU results automatically, substituting human cognitive processes with algorithmic operations that are faster and can be performed continuously without fatigue or inconsistency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual tagging and human evaluation are used to expand intent classification models, then the system can detect misclassifications and new intent patterns, but the process becomes costly

Engineering Contradiction:
Improvedetection accuracy of misclassifications and new intentVSAvoidtraining cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The system performs self-evaluation of NLU results using automated prompt generation and large-scale language model assessment, eliminating the need to pay human evaluators. The automated process incurs only computational costs while maintaining high detection accuracy for misclassifications and new intent patterns

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses a large-scale language model to simulate and replicate human evaluation capabilities. The language model generates prompts and evaluates NLU results in a manner that copies human judgment processes, achieving comparable detection accuracy at lower cost through automated computational simulation

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If multiple human evaluators are used to expand intent classification models, then diverse perspectives can be obtained, but inconsistent results occur

Engineering Contradiction:
Improveevaluation perspective diversityVSAvoidevaluation consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The large-scale language model serves as a universal evaluator that can process and evaluate all NLU results consistently. Unlike human evaluators who have individual biases and varying standards, the language model applies uniform evaluation criteria across all cases, ensuring consistent and reliable results while maintaining the ability to detect diverse intent patterns

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system implements automated feedback loops where NLU results are evaluated, misclassifications are identified, and new intent patterns are derived. This continuous feedback process ensures consistent evaluation standards are applied iteratively, improving reliability while the language model's comprehensive understanding maintains adaptability to diverse evaluation perspectives

Inventive Principle:
Principle #23Feedback

4Ease of manufacture

If speech recognition systems are trained with predefined intent, then initial functionality is achieved, but the system cannot adapt to changing user needs over time

Engineering Contradiction:
Improveinitial system setupVSAvoidsystem adaptability to changing user behavior
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent transforms the static predefined intent system into a dynamic adaptive system. Through automated evaluation using large-scale language models, the system continuously identifies new intent patterns and updates its classification model, enabling it to adapt to changing user needs over time while maintaining the simplicity of initial setup

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary automated evaluation and intent derivation in advance, preparing updated intent classifications before they are needed. This allows the system to proactively adapt to changing user behavior patterns, maintaining versatility while keeping the initial setup process simple and straightforward

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250124915A1Method and apparatus for automatically training speech recognition system
Publication Date: 2025.04.17 HYUNDAI MOTOR CO LTD
  • US20250124915A1 patent drawing
  • US20250124915A1 patent drawing
  • US20250124915A1 patent drawing

AI summary

In a method and apparatus for automatically training speech recognition system including one or more NLU engines, the method includes obtaining NLU results output by the one or more NLU engines, generating a prompt for a large-scale language model based on comparing between/among the NLU results, determining whether the NLU results are appropriate by use of a generated prompt for the large-scale language model and training the speech recognition system by use of a determination result