Voice Interaction System Handling Ambiguous User Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice interaction systems face difficulties in responding appropriately to ambiguous user speech, lacking clear intention, which can lead to unpleasant interactions.

Innovation Solution

An information output system comprising a speech acquisition unit, recognition processing unit, and output processing unit that acquires user speech, recognizes its content, outputs questions, and determines guidance information based on the user's positive degree derived from their responses, providing tailored guidance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the system responds to ambiguous user speech using pre-stored interaction scenarios, then the system can provide some level of response, but the response accuracy and user satisfaction deteriorate when user intention is unclear

Engineering Contradiction:
Improveresponse capability to ambiguous speechVSAvoiduser intention recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary actions by storing multiple interaction scenarios in advance and preparing clarification questions before being explicitly asked. When ambiguous speech is detected, the system has pre-prepared guidance information and follow-up questions ready to guide the user toward clarifying their intention, rather than responding directly to the ambiguous input.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback by outputting clarification questions to users when ambiguous speech is detected. The system waits for user responses to these questions, uses the responses to update understanding of user intention, and iteratively refines the response based on this feedback loop until user intention becomes clear.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the system asks clarification questions to resolve ambiguous speech, then user intention recognition accuracy improves, but interaction time and complexity increase

Engineering Contradiction:
Improveuser intention recognition accuracyVSAvoidinteraction time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial action by selectively asking clarification questions only when ambiguous speech is detected, rather than asking questions in all interactions. The system balances between direct response and clarification by evaluating speech ambiguity first, thus avoiding unnecessary interaction delays while maintaining accuracy when needed.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If the system stores multiple interaction scenarios in advance, then response capability to various user inputs improves, but system complexity and memory requirements increase

Engineering Contradiction:
Improveinteraction scenario coverageVSAvoidsystem structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system achieves universality by designing interaction scenarios that can handle multiple types of ambiguous speech through common clarification question templates. Rather than creating entirely separate handling mechanisms for each ambiguous input type, the system uses universal clarification procedures that adapt to different contexts, reducing overall system complexity while maintaining broad adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11657806B2Information output system and information output method
Publication Date: 2023.05.23 TOYOTA JIDOSHA KK
  • US11657806B2 patent drawing
  • US11657806B2 patent drawing
  • US11657806B2 patent drawing

AI summary

An information output system includes a speech acquisition unit configured to acquire a speech of a user, a recognition processing unit configured to recognize the content of the acquired speech of the user, and an output processing unit configured to output a question to the user and to perform processing for outputting a response to the content of the speech of the user who has answered the question. The output processing unit is configured to derive a user's positive degree based on the content of the speech of the user who has answered the question and to determine guidance information to be output to the user based on the derived positive degree.