Adaptive Intent Estimation for Speech Recognition Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in accurately understanding user utterances due to individual differences in wording, dialects, and non-native language usage, leading to errors in intent recognition and user frustration.
Innovation Solution
An information processing apparatus and method that includes an utterance learning adaptive processing unit, which analyzes user utterances, generates learning data associating entity information with correct intents, and stores this data for improved intent estimation by using similar semantic concepts in new utterances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If natural language understanding (NLU) function is applied to understand user utterance intent, then the system can process speech inputs, but it fails to correctly understand diverse user utterances including dialects and non-native language usage
Solution Approach 1:
The system performs preliminary speech recognition to convert speech to text, then uses NLU to analyze intent before executing actions. This preliminary processing chain allows the system to prepare and analyze user inputs systematically, improving reliability in understanding diverse utterances by breaking down the complex task into manageable stages.
Solution Approach 2:
The patent introduces text data as an intermediary between speech input and intent understanding. The speech recognition unit converts speech to text, which then serves as the input for NLU processing. This intermediary representation helps bridge the gap between diverse speech patterns and standardized intent analysis, enhancing both reliability and adaptability.
2Ease of operation
If speech recognition converts speech to text, then speech input can be processed, but erroneous recognition occurs due to individual differences in articulation and system performance limits
Solution Approach 1:
The system implements a feedback mechanism where the NLU process analyzes the recognized text and can identify errors in speech recognition. When intent understanding fails or produces unexpected results, the system can request clarification or re-recognition, allowing users to correct misrecognized speech. This feedback loop compensates for recognition errors and maintains ease of operation while improving effective accuracy.
Solution Approach 2:
The patent prepares for potential recognition errors by designing a robust NLU process that can handle imperfect text input. The system anticipates erroneous recognition and builds in error tolerance through contextual analysis and intent verification, cushioning against the impact of speech recognition mistakes before they propagate through the system.
3Reliability
If learning data is generated from user utterances with unclear intent, then intent understanding can be improved, but the system complexity increases
Solution Approach 1:
The system performs self-learning by automatically generating learning data from its own operation. When the NLU process encounters uncertain intents, it generates learning data that associates the recognized text with the correct intent, storing this data for future improvements. This self-service approach enhances reliability without requiring external intervention, balancing accuracy improvement with acceptable system complexity.
Data Source
AI summary
Implemented are an apparatus and a method that enable highly accurate intent estimation of a user utterance. An utterance learning adaptive processing unit analyzes a plurality of user utterances input from a user, generates learning data in which entity information included in a user utterance with an unclear intent is associated with a correct intent, and stores the generated learning data is a storage unit. The utterance learning adaptive processing unit generates learning data in which an intent, acquired from a response utterance from the user to an apparatus utterance after input of a first user utterance with an unclear intent, is recorded in association with entity information included in the first user utterance. The learning data is recorded to include superordinate semantic concept information of the entity information. At the time of estimating an intent for a new user utterance, learning data with similar superordinate semantic concept information is used.


