Near-Ear Open Audio Sound Field Expansion With Voice Balancing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The current sound field expansion function in near-ear open audio devices, such as VR and AR, often results in a feeble sound effect for human voice due to the implementation of the Head Related Transfer Function (HRTF) algorithm.
Innovation Solution
A method involving the acquisition of a target transfer function, crosstalk elimination processing, and adjustment of sound intensity ratios between human voice and accompaniment audio to enhance the sound effect while expanding the sound field.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If HRTF algorithm is used to expand the sound field in near-ear open audio devices, then the sound field expansion effect is improved, but the human voice sound quality deteriorates and becomes feeble
Solution Approach 1:
The audio signal is segmented into human voice components and accompaniment components separately. Different processing strategies are applied to each segment: the HRTF algorithm is applied to the accompaniment for sound field expansion, while the human voice is processed with crosstalk elimination and selective intensity enhancement to maintain its clarity and presence.
Solution Approach 2:
Different quality requirements are applied to different parts of the audio signal. The human voice portion receives enhanced processing with crosstalk elimination and intensity boosting to ensure high fidelity, while the accompaniment portion undergoes standard HRTF processing for spatial expansion. This local differentiation resolves the contradiction by optimizing each component according to its specific needs.
2Adaptability or versatility
If sound field expansion is implemented in near-ear open audio devices, then the spatial audio experience is improved, but the clarity and presence of human voice deteriorates
Solution Approach 1:
An audio processing system acts as an intermediary between the input audio signal and the output to the ear. This intermediary performs crosstalk elimination, separates voice from accompaniment, applies selective intensity adjustment, and then applies HRTF processing. The intermediary ensures that the human voice clarity is preserved through careful processing before the spatial expansion is applied.
Solution Approach 2:
Crosstalk elimination and voice-accompaniment separation are performed as preliminary actions before applying the HRTF algorithm. By preparing the audio signal in advance—removing crosstalk and identifying voice portions—the system ensures that when HRTF processing is applied, the human voice clarity is already protected and can be enhanced with selective intensity boosting.
3Adaptability or versatility
If HRTF processing is applied to all audio components, then the sound field expansion is improved, but the sound intensity balance between human voice and accompaniment deteriorates
Solution Approach 1:
The system dynamically adjusts the sound intensity of different audio components based on real-time analysis. After separating human voice and accompaniment, the system applies different intensity adjustments: enhancing the human voice to compensate for HRTF-induced attenuation while maintaining appropriate levels for the accompaniment. This dynamic adjustment resolves the intensity balance issue while preserving sound field expansion.
Solution Approach 2:
The system changes the intensity parameter selectively for different audio components. By identifying human voice portions and applying targeted intensity enhancement, the system compensates for the energy loss that occurs during HRTF processing. This parameter change is applied only where needed (to human voice) while leaving the accompaniment intensity relatively unchanged, thus maintaining overall balance while resolving the specific issue of feeble human voice.
Data Source
AI summary
The disclosure discloses a sound field expansion method, an audio device and a computer readable storage medium, and belongs to a technical field of audio processing. The method comprises: acquiring a target transfer function between a near-ear open audio device and two ears of a user; performing a crosstalk elimination processing on an input audio received by the near-ear open audio device according to the target transfer function to acquire an initial reverberation audio; identifying an actual sound intensity weight ratio between a human voice audio and an accompanying audio in the initial reverberation audio; adjusting a sound intensity of the human voice audio and/or the accompanying audio in the initial reverberation audio according to the actual sound intensity weight ratio to acquire a target reverberation audio; playing the target reverberation audio. The audio device can effectively expand a sound field while ensuring a sound effect of a human voice.


