Speech Recognition Training via LLM Prompt Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for expanding intent classification models in speech recognition systems require manual tagging and are time-consuming, costly, and prone to inconsistent results due to human evaluators. Additionally, these systems struggle to detect misclassifications and new intent patterns in changing user behaviors.
Innovation Solution
A method and apparatus for automatically training a speech recognition system by obtaining NLU results, generating prompts for a large-scale language model, determining the appropriateness of these results, and using the determination to train the system. This approach enables the detection of misclassifications and derivation of new intent without human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual tagging and human evaluation are used to expand intent classification models, then the system can detect misclassifications and new intent patterns, but the process becomes time-consuming and costly
Solution Approach 1:
The speech recognition system automatically evaluates its own NLU results using a large-scale language model, eliminating the need for human evaluators. The system self-identifies misclassifications and derives new intent patterns through automated prompt generation and evaluation, making the training process independent of manual intervention while maintaining high detection accuracy
Solution Approach 2:
The patent replaces the mechanical human evaluation process with an automated computational system. A large-scale language model generates prompts and evaluates NLU results automatically, substituting human cognitive processes with algorithmic operations that are faster and can be performed continuously without fatigue or inconsistency
2Measurement precision
If manual tagging and human evaluation are used to expand intent classification models, then the system can detect misclassifications and new intent patterns, but the process becomes costly
Solution Approach 1:
The system performs self-evaluation of NLU results using automated prompt generation and large-scale language model assessment, eliminating the need to pay human evaluators. The automated process incurs only computational costs while maintaining high detection accuracy for misclassifications and new intent patterns
Solution Approach 2:
The patent uses a large-scale language model to simulate and replicate human evaluation capabilities. The language model generates prompts and evaluates NLU results in a manner that copies human judgment processes, achieving comparable detection accuracy at lower cost through automated computational simulation
3Adaptability or versatility
If multiple human evaluators are used to expand intent classification models, then diverse perspectives can be obtained, but inconsistent results occur
Solution Approach 1:
The large-scale language model serves as a universal evaluator that can process and evaluate all NLU results consistently. Unlike human evaluators who have individual biases and varying standards, the language model applies uniform evaluation criteria across all cases, ensuring consistent and reliable results while maintaining the ability to detect diverse intent patterns
Solution Approach 2:
The system implements automated feedback loops where NLU results are evaluated, misclassifications are identified, and new intent patterns are derived. This continuous feedback process ensures consistent evaluation standards are applied iteratively, improving reliability while the language model's comprehensive understanding maintains adaptability to diverse evaluation perspectives
4Ease of manufacture
If speech recognition systems are trained with predefined intent, then initial functionality is achieved, but the system cannot adapt to changing user needs over time
Solution Approach 1:
The patent transforms the static predefined intent system into a dynamic adaptive system. Through automated evaluation using large-scale language models, the system continuously identifies new intent patterns and updates its classification model, enabling it to adapt to changing user needs over time while maintaining the simplicity of initial setup
Solution Approach 2:
The system performs preliminary automated evaluation and intent derivation in advance, preparing updated intent classifications before they are needed. This allows the system to proactively adapt to changing user behavior patterns, maintaining versatility while keeping the initial setup process simple and straightforward
Data Source
AI summary
In a method and apparatus for automatically training speech recognition system including one or more NLU engines, the method includes obtaining NLU results output by the one or more NLU engines, generating a prompt for a large-scale language model based on comparing between/among the NLU results, determining whether the NLU results are appropriate by use of a generated prompt for the large-scale language model and training the speech recognition system by use of a determination result


