Agent System Speech Recognition with Direction Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing human-machine interface systems struggle to accurately recognize speech contents, especially when unclear wording is used, leading to misinterpretation of user intentions.

Innovation Solution

An agent system that includes a recognizer, an acquirer, and an estimator, which acquires speech and image data to estimate the user's sight direction and objects of interest based on unclear wording, allowing for more accurate recognition of speech contents by correlating unclear information with map and profile data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition is performed using only dictionary-based word matching, then the system is simple to operate, but speech contents cannot be accurately recognized when unclear wording is used

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech recognition process is segmented into multiple stages: initial dictionary-based recognition, detection of unclear wording, estimation of sight direction and indicated direction, identification of reference objects, and final recognition using contextual information. This segmentation allows the system to handle unclear wording by breaking down the recognition task into manageable steps.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by acquiring image data and estimating the occupant's sight direction and indicated direction before final speech recognition. This preliminary estimation of contextual information (what the occupant is looking at or pointing to) is prepared in advance to resolve unclear wording in the speech.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the system acquires and processes additional image and direction data to resolve unclear wording, then speech recognition accuracy is improved, but the processing time and computational load increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial action by selectively processing additional information only when unclear wording is detected. Instead of always performing image analysis and direction estimation for every speech input, the system activates these additional processing steps only when the dictionary-based recognition identifies unclear terms, thus reducing unnecessary processing time for clear speech.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system uses feedback from the initial speech recognition result to determine whether additional processing is needed. When unclear wording is detected in the initial recognition, the system triggers further analysis of image data and direction estimation, and uses this feedback to refine the final recognition result.

Inventive Principle:
Principle #23Feedback

3Reliability

If the system uses only dictionary-based speech recognition, then the device complexity is low, but the reliability of speech content recognition decreases when unclear wording is present

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system introduces an intermediary process between initial speech recognition and final interpretation. This intermediary step involves estimating the occupant's sight direction and indicated direction from image data, and identifying reference objects based on these directions. This intermediary information acts as a bridge to resolve unclear wording in the speech.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system achieves multi-functionality by integrating speech recognition, image processing, direction estimation, and object identification into a unified system. The same system can handle both clear and unclear speech by dynamically activating appropriate processing functions, making it universally applicable to various speech input quality levels.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11508368B2Agent system, and, information processing method
Publication Date: 2022.11.22 HONDA MOTOR CO LTD
  • US11508368B2 patent drawing
  • US11508368B2 patent drawing
  • US11508368B2 patent drawing

AI summary

An agent system includes: a recognizer configured to recognize speech including speech contents of an occupant in a mobile object; an acquirer configured to acquire an image including the occupant; and an estimator configured to compare wording included in the speech contents of the occupant recognized by the recognizer with unclear information which is stored in a storage and includes wording making the speech contents unclear, to estimate a first direction which is a sight direction of the occupant or a second direction which is indicated by the occupant on the basis of the image acquired by the acquirer when the speech contents of the occupant includes unclear wording, and to estimate an object which is located in the estimated first direction or the estimated second direction. The recognizer is configured to recognize the speech contents of the occupant on the basis of the object estimated by the estimator.