HRTF Sound Localization for Clearer Remote Conference Conversations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Participants in remote conferencing systems cannot individually designate specific participants for conversation and focus on their utterances, leading to difficulty in distinguishing which participant is engaged in an action, especially when multiple participants speak simultaneously.
Innovation Solution
An information processing apparatus and terminal that utilize HRTF (Head-Related Transfer Function) data to perform sound image localization processing, allowing sound content to be localized at specific positions based on participant actions, enabling immersive audio experiences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If sound content is output for all participants in a remote conference, then all participants can hear each other, but it becomes difficult to distinguish which specific participant is speaking when multiple participants speak simultaneously
Solution Approach 1:
The patent segments the audio output by participant, assigning unique spatial positions to each participant's sound content. This allows the system to maintain clear audio identification for all participants while distinguishing their individual speech streams through spatial separation, resolving the contradiction between comprehensive audio delivery and participant differentiation.
Solution Approach 2:
The patent introduces a spatial dimension to audio output by localizing sound content at different three-dimensional positions corresponding to each participant's location. This dimensional addition allows multiple participants to be heard simultaneously without confusion, as each participant's voice is positioned in a distinct spatial coordinate, eliminating the need for complex audio mixing while maintaining clear identification.
2Measurement precision
If sound image localization processing is performed for each participant, then clarity of specific participant speech is improved, but processing time and computational resources increase
Solution Approach 1:
The patent pre-calculates and stores HRTF (Head-Related Transfer Function) data for multiple positions before the actual audio processing occurs. This preliminary preparation of spatial audio data allows the system to quickly retrieve and apply appropriate HRTF values during real-time conference operations, achieving high precision sound localization without the time penalty of calculating these transforms in real-time.
Solution Approach 2:
The patent dynamically adjusts audio processing parameters based on participant actions and spatial positions. By changing parameters such as HRTF selection and sound localization intensity according to the current conference state and participant locations, the system maintains high precision localization while optimizing processing time through adaptive rather than fixed processing approaches.
Data Source
AI summary
The present technique relates to an information processing apparatus, an information processing terminal, an information processing method, and a program which enable a sound content in accordance with an action by a participant in a conversation to be output in an immersive state. An information processing apparatus according to an aspect of the present technique includes: a storage unit configured to store HRTF data corresponding to a plurality of positions based on a listening position; and a sound image localization processing unit configured to provide, by performing sound image localization processing using the HRTF data selected in accordance with an action by a specific participant among participants of a conversation having participated via a network, a sound content selected in accordance with the action so that a sound image is localized at a prescribed position. The present technology can be applied to a computer which conducts a remote conference.


