Audio Signal Processing Apparatus for Distance-Dependent Voice Level Compensation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing systems fail to maintain an appropriate voice level for both distant and near talkers, leading to inadequate voice capture.

Innovation Solution

An audio signal processing method and apparatus that utilize a camera to estimate the position and posture of talkers, generating correction filters to compensate for voice attenuation and maintain stable voice levels and quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the audio processing system uses a fixed gain for all talkers, then the voice of near talkers is captured at an appropriate level, but the voice of distant talkers is attenuated and cannot be obtained at an appropriate level

Engineering Contradiction:
Improvevoice level accuracyVSAvoiddistance adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically adjusts the gain for each talker based on their detected position and distance from the microphone array. Instead of using a fixed gain, the gain is varied in real-time according to the talker's location, allowing appropriate voice levels for both near and distant talkers to be achieved simultaneously.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the gain parameter based on the talker's distance from the microphone array. By calculating the distance and applying distance-dependent gain adjustment, the system compensates for voice attenuation and ensures that voices from different distances are captured at appropriate levels.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If the system applies strong voice enhancement to compensate for distant talker attenuation, then distant talker voice level is improved, but near talker voice quality deteriorates due to over-enhancement

Engineering Contradiction:
Improvedistant talker voice levelVSAvoidvoice quality stability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system applies different processing characteristics to different spatial locations. Near talkers receive minimal or no enhancement while distant talkers receive appropriate gain compensation. This location-dependent processing ensures that each talker's voice is enhanced only to the extent needed, preventing over-enhancement and maintaining voice quality stability.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The enhancement amount is dynamically adjusted based on the talker's distance and position. The system continuously monitors talker location and adapts the enhancement level in real-time, applying stronger enhancement only when and where needed for distant talkers, while maintaining natural levels for near talkers.

Inventive Principle:
Principle #15Dynamics

3Loss of information

If the system uses camera-based position estimation, then distance information is obtained to compensate for voice attenuation, but the system complexity increases due to integration of multiple sensors and processing

Engineering Contradiction:
Improvedistance informationVSAvoidsystem integration complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The camera system serves multiple functions: it captures visual information for position estimation, provides distance information for gain adjustment, and can potentially identify talkers. By making the camera multi-functional, the system avoids adding separate sensors solely for distance measurement, thereby reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses an intermediary processing module that integrates camera and microphone data. This intermediary layer coordinates the position estimation from the camera with the audio processing, managing the complexity of multi-sensor integration through a centralized control mechanism.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3989222B1Audio signal processing method and audio signal processing apparatus
Publication Date: 2025.05.14 YAMAHA CORP
  • EP3989222B1 patent drawingFigure 1
  • EP3989222B1 patent drawingFigure 2
  • EP3989222B1 patent drawingFigure 3

AI summary

An audio signal processing method includes receiving an audio signal corresponding to a voice of a talker (S11), obtaining an image of the talker (S12), estimating position information of the talker using the image of the talker (S13), generating, according to the estimated position information, a correction filter configured to compensate for an attenuation of the voice of the talker (S14), performing filter processing on the audio signal using the generated correction filter (S15), and outputting the audio signal on which the filter processing has been performed (S16).