Agent System Speech Recognition with Direction Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing human-machine interface systems struggle to accurately recognize speech contents, especially when unclear wording is used, leading to misinterpretation of user intentions.
Innovation Solution
An agent system that includes a recognizer, an acquirer, and an estimator, which acquires speech and image data to estimate the user's sight direction and objects of interest based on unclear wording, allowing for more accurate recognition of speech contents by correlating unclear information with map and profile data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition is performed using only dictionary-based word matching, then the system is simple to operate, but speech contents cannot be accurately recognized when unclear wording is used
Solution Approach 1:
The speech recognition process is segmented into multiple stages: initial dictionary-based recognition, detection of unclear wording, estimation of sight direction and indicated direction, identification of reference objects, and final recognition using contextual information. This segmentation allows the system to handle unclear wording by breaking down the recognition task into manageable steps.
Solution Approach 2:
The system performs preliminary actions by acquiring image data and estimating the occupant's sight direction and indicated direction before final speech recognition. This preliminary estimation of contextual information (what the occupant is looking at or pointing to) is prepared in advance to resolve unclear wording in the speech.
2Measurement precision
If the system acquires and processes additional image and direction data to resolve unclear wording, then speech recognition accuracy is improved, but the processing time and computational load increase
Solution Approach 1:
The system applies partial action by selectively processing additional information only when unclear wording is detected. Instead of always performing image analysis and direction estimation for every speech input, the system activates these additional processing steps only when the dictionary-based recognition identifies unclear terms, thus reducing unnecessary processing time for clear speech.
Solution Approach 2:
The system uses feedback from the initial speech recognition result to determine whether additional processing is needed. When unclear wording is detected in the initial recognition, the system triggers further analysis of image data and direction estimation, and uses this feedback to refine the final recognition result.
3Reliability
If the system uses only dictionary-based speech recognition, then the device complexity is low, but the reliability of speech content recognition decreases when unclear wording is present
Solution Approach 1:
The system introduces an intermediary process between initial speech recognition and final interpretation. This intermediary step involves estimating the occupant's sight direction and indicated direction from image data, and identifying reference objects based on these directions. This intermediary information acts as a bridge to resolve unclear wording in the speech.
Solution Approach 2:
The system achieves multi-functionality by integrating speech recognition, image processing, direction estimation, and object identification into a unified system. The same system can handle both clear and unclear speech by dynamically activating appropriate processing functions, making it universally applicable to various speech input quality levels.
Data Source
AI summary
An agent system includes: a recognizer configured to recognize speech including speech contents of an occupant in a mobile object; an acquirer configured to acquire an image including the occupant; and an estimator configured to compare wording included in the speech contents of the occupant recognized by the recognizer with unclear information which is stored in a storage and includes wording making the speech contents unclear, to estimate a first direction which is a sight direction of the occupant or a second direction which is indicated by the occupant on the basis of the image acquired by the acquirer when the speech contents of the occupant includes unclear wording, and to estimate an object which is located in the estimated first direction or the estimated second direction. The recognizer is configured to recognize the speech contents of the occupant on the basis of the object estimated by the estimator.


