Adaptive Audio Directivity for Speech Dialogue Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech dialogue systems struggle to reproduce audio responses optimally, as they often direct sound only towards the speaking person, leading to issues where unintended individuals may hear the response or fail to hear it due to noise, and existing methods do not adapt to the surrounding situation effectively.
Innovation Solution
A sound reproduction method that acquires ambient sound information, separates it into spoken voice and other sounds, compares sound levels, and selects between a reproduction method with or without directivity based on the comparison to ensure the audio response is optimally directed towards the speaking person.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a fixed directivity reproduction method is used to direct sound towards the speaking person, then the speaking person can hear the response clearly, but unintended individuals may also hear the response or the speaking person may fail to hear it due to noise
Solution Approach 1:
The system dynamically switches between different reproduction methods (first reproduction method with fixed directivity and second reproduction method with variable directivity) based on the detected sound environment. The processor determines whether to use narrow or wide directivity by comparing the sound level of the speaking person with the ambient noise level, allowing the system to adapt to changing situations rather than using a fixed directivity pattern
Solution Approach 2:
The system changes the directivity parameter of the audio reproduction based on the sound environment. When the speaking person's voice level is high relative to ambient noise, the system uses narrow directivity to target the speaking person. When ambient noise is high, the system switches to wide directivity to ensure the response is audible despite the noise conditions
2Loss of information
If narrow directivity is used to target the speaking person, then only the intended person hears the response, but the speaking person may fail to hear it in noisy environments
Solution Approach 1:
The system dynamically adjusts the directivity width based on ambient noise conditions. In low-noise environments, narrow directivity is used to prevent unauthorized hearing. In high-noise environments, the system automatically switches to wide directivity to ensure the speaking person can hear the response, thus maintaining reliability while minimizing information loss
3Reliability
If wide directivity is used to ensure audibility in noisy environments, then the speaking person can hear the response, but unintended individuals may also hear it
Solution Approach 1:
The system uses wide directivity only when necessary, determined by comparing the speaking person's voice level with ambient noise level. This dynamic approach ensures audibility in noisy environments while minimizing the risk of unauthorized hearing in quiet environments, thus resolving the contradiction between reliability and information loss
Data Source
AI summary
A sound reproduction method is provided. The method includes acquiring ambient sound information that includes voice spoken to a speech dialog system and indicates sound around a speaking person who has spoken the voice. The method also includes separating the ambient sound information into first sound information including the spoken voice and second sound information including sound other than the spoken voice. The method further includes comparing the sound level of the first sound information with the sound level of the second sound information, and reproducing an audio response to the spoken voice, by selecting one of a first reproduction method and a second reproduction method that is different in terms of directivity of reproduced sound from the first reproduction method in accordance with a result of the comparison.


