Adaptive Response Output Control for Speech Recognition Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice UI systems face usability issues in environments where speaking is not preferred or in noisy conditions, as they do not automatically adapt their response output methods to the user's surrounding environment.

Innovation Solution

An information processing device with a response generation unit, decision unit, and output control unit that determines and implements an appropriate response output method based on the current surrounding environment, such as switching to text-based responses in noisy or nighttime conditions, using a combination of voice, display, or external devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the system always outputs responses by voice, then the speech recognition system maintains simple operation, but usability deteriorates in environments where speaking is not preferred or noise is large

Engineering Contradiction:
ImproveusabilityVSAvoidenvironmental adaptability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system dynamically adjusts the response output method based on the detected surrounding environment. The decision unit changes the output mode from voice to text or other methods according to environmental conditions such as noise levels and time of day, making the system adaptive rather than static.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses environmental sensors to detect surrounding conditions and feeds this information back to the decision unit, which then selects the appropriate response output method. This feedback mechanism enables the system to automatically adapt to changing environmental conditions without user intervention.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If the system automatically detects environment and switches output methods, then environmental adaptability improves, but device complexity increases

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system divides the response output function into separate modules: a response generation unit that creates the response content, a decision unit that selects the output method, and an output control unit that executes the output. This segmentation allows each unit to perform its specific function independently, managing complexity through functional decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The decision unit acts as an intermediary between the response generation unit and the output control unit. It receives the generated response and environmental information, then determines the appropriate output method before passing control to the output control unit, thereby managing system complexity through a centralized decision-making layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If voice output is used in noisy environments, then the system maintains simple output control, but information transmission reliability deteriorates

Engineering Contradiction:
Improveoutput control simplicityVSAvoidinformation transmission reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system changes the output parameter from voice to text based on environmental noise levels. When noise is detected to be large, the decision unit switches the output method to text display, which is more reliable in noisy environments where voice output may not be clearly heard or understood.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10776070B2Information processing device, control method, and program
Publication Date: 2020.09.15 SONY GROUP CORP
  • US10776070B2 patent drawing
  • US10776070B2 patent drawing
  • US10776070B2 patent drawing

AI summary

There is provided an information processing device, control method, and program that can improve convenience of a speech recognition system by deciding an appropriate response output method in accordance with a current surrounding environment. A response to a speech from a user is generated, a response output method is decided in accordance with a current surrounding environment, and control is performed such that the generated response is output by using the decided response output method.