Spatial Audio-Visual Processing for Residual Audio Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video and audio capture technologies struggle with imperfect audio suppression, leading to residual audio remnants that become distracting when further audio processing is applied, especially in environments where unwanted objects are removed or replaced.

Innovation Solution

An image and sound apparatus that applies audio zoom selectively to spatial audio-visual representations, modifying audio processing based on prior audio-visual manipulations to minimize residual audio remnants by adjusting amplification levels and directions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If directional audio suppression is applied to remove unwanted audio sources, then the primary audio quality is improved, but residual audio remnants remain that become distracting during further audio processing

Engineering Contradiction:
Improveaudio suppression accuracyVSAvoidresidual audio remnants
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The system performs preliminary audio suppression to remove unwanted audio sources before further audio processing. By applying suppression in advance and then detecting residual remnants in subsequent processing stages, the system can identify and handle these remnants before they become distracting to the user.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback by detecting residual audio remnants after suppression and using this information to adjust further audio processing. The detection of remnants provides feedback that allows the system to modify its processing approach to minimize the impact of these remnants during zooming and other audio manipulations.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If audio zoom is applied to enhance specific audio directions, then audio focus is improved, but residual audio remnants become more prominent and distracting

Engineering Contradiction:
Improveaudio focus accuracyVSAvoidamplified residual audio
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary detection of residual audio remnants before applying audio zoom. By identifying the presence and characteristics of remnants in advance, the system can take preliminary actions to suppress or compensate for them before they are amplified by the zoom operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies preliminary anti-action by detecting residual audio remnants and applying counter-measures before audio zoom enhances them. This may involve targeted suppression or filtering of identified remnant frequencies and directions before the zoom operation amplifies the audio signal.

Inventive Principle:
Principle #9Preliminary anti-action

3Measurement precision

If content-aware fill is used to replace unwanted visual objects, then visual quality is improved, but corresponding audio remnants remain that create audio-visual inconsistency

Engineering Contradiction:
Improvevisual editing accuracyVSAvoidaudio-visual synchronization
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system performs preliminary detection of audio remnants in the spatial regions where visual content has been replaced. By identifying these audio-visual inconsistencies in advance, the system can apply targeted audio processing to restore synchronization between the modified visual content and the audio signal.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12418762B2Image and audio apparatus and method
Publication Date: 2025.09.16 NOKIA TECHNOLOGIES OY
  • US12418762B2 patent drawing
  • US12418762B2 patent drawing
  • US12418762B2 patent drawing

AI summary

An apparatus including circuitry configured for causing audio processing to a spatial audio-visual representation of an image and sound apparatus, the spatial audio-visual representation being live or reproduced from recording; and modifying the audio processing applied to an audio-visually manipulated spatial section of the spatial audio-visual representation in response to information a prior audio-visual manipulation with data processing in the audio-visually manipulated spatial section.