Adaptive Audio Directivity for Speech Dialogue Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech dialogue systems struggle to reproduce audio responses optimally, as they often direct sound only towards the speaking person, leading to issues where unintended individuals may hear the response or fail to hear it due to noise, and existing methods do not adapt to the surrounding situation effectively.

Innovation Solution

A sound reproduction method that acquires ambient sound information, separates it into spoken voice and other sounds, compares sound levels, and selects between a reproduction method with or without directivity based on the comparison to ensure the audio response is optimally directed towards the speaking person.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a fixed directivity reproduction method is used to direct sound towards the speaking person, then the speaking person can hear the response clearly, but unintended individuals may also hear the response or the speaking person may fail to hear it due to noise

Engineering Contradiction:
Improveaudio response delivery reliabilityVSAvoidadaptation to surrounding situation
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system dynamically switches between different reproduction methods (first reproduction method with fixed directivity and second reproduction method with variable directivity) based on the detected sound environment. The processor determines whether to use narrow or wide directivity by comparing the sound level of the speaking person with the ambient noise level, allowing the system to adapt to changing situations rather than using a fixed directivity pattern

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the directivity parameter of the audio reproduction based on the sound environment. When the speaking person's voice level is high relative to ambient noise, the system uses narrow directivity to target the speaking person. When ambient noise is high, the system switches to wide directivity to ensure the response is audible despite the noise conditions

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If narrow directivity is used to target the speaking person, then only the intended person hears the response, but the speaking person may fail to hear it in noisy environments

Engineering Contradiction:
Improveprevention of unauthorized hearingVSAvoidaudio response audibility
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system dynamically adjusts the directivity width based on ambient noise conditions. In low-noise environments, narrow directivity is used to prevent unauthorized hearing. In high-noise environments, the system automatically switches to wide directivity to ensure the speaking person can hear the response, thus maintaining reliability while minimizing information loss

Inventive Principle:
Principle #15Dynamics

3Reliability

If wide directivity is used to ensure audibility in noisy environments, then the speaking person can hear the response, but unintended individuals may also hear it

Engineering Contradiction:
Improveaudio response audibilityVSAvoidunauthorized hearing
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system uses wide directivity only when necessary, determined by comparing the speaking person's voice level with ambient noise level. This dynamic approach ensures audibility in noisy environments while minimizing the risk of unauthorized hearing in quiet environments, thus resolving the contradiction between reliability and information loss

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10089980B2Sound reproduction method, speech dialogue device, and recording medium
Publication Date: 2018.10.02 PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
  • US10089980B2 patent drawing
  • US10089980B2 patent drawing
  • US10089980B2 patent drawing

AI summary

A sound reproduction method is provided. The method includes acquiring ambient sound information that includes voice spoken to a speech dialog system and indicates sound around a speaking person who has spoken the voice. The method also includes separating the ambient sound information into first sound information including the spoken voice and second sound information including sound other than the spoken voice. The method further includes comparing the sound level of the first sound information with the sound level of the second sound information, and reproducing an audio response to the spoken voice, by selecting one of a first reproduction method and a second reproduction method that is different in terms of directivity of reproduced sound from the first reproduction method in accordance with a result of the comparison.