Camera-Selected Audio Processing for Consistent Spatial Speech Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing technologies fail to effectively separate and adapt speech and ambient sounds in video calls or recordings based on camera selection, leading to inconsistencies and suboptimal audio quality when camera viewpoints change.

Innovation Solution

Utilizing spatial multi-microphone capture and audio processing techniques that adjust audio modes based on camera selection, separating speech and ambient signals, and maintaining a consistent spatial audio image with the video view, even when camera viewpoints change.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio processing is performed without camera selection awareness, then device complexity is reduced, but audio fidelity and adaptability deteriorate

Engineering Contradiction:
Improveaudio fidelityVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system continuously monitors camera selection state and feeds this information back to the audio processing module, enabling dynamic adjustment of audio parameters. The audio processing apparatus receives camera selection information and uses it to determine appropriate audio processing modes, creating a closed-loop system that adapts audio processing to the current video capture context.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The audio processing system transitions from a static, fixed-mode processing approach to a dynamic, adaptive approach where processing parameters change in real-time based on camera selection. The system can switch between different audio processing modes (e.g., speech enhancement, ambient sound capture, noise reduction) depending on which camera is active and the detected scene context.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If audio processing adapts to camera selection, then audio adaptability improves, but processing time increases

Engineering Contradiction:
Improveaudio adaptabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system pre-configures multiple audio processing modes and algorithms, preparing them in advance for rapid selection. Instead of computing processing parameters in real-time, the system maintains a library of pre-processed audio processing configurations that can be quickly activated based on camera selection, reducing computational overhead and processing delays.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system adjusts audio processing by changing parameters such as gain levels, filter characteristics, and processing intensity based on camera selection state. Rather than performing complex real-time analysis, the system modifies existing processing parameters to match the appropriate audio mode, enabling fast adaptation with minimal computational burden.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If spatial audio image consistency is maintained, then audio quality improves, but device complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidspatial processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system introduces an intermediary mapping layer that connects camera viewpoint information to audio spatial parameters. This intermediary module translates camera selection and scene detection data into appropriate audio panning, spatial positioning, and surround sound configurations, enabling consistent spatial audio images without requiring complex direct analysis of audio sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3238461B1Audio processing based upon camera selection
Publication Date: 2026.03.18 NOKIA TECHNOLOGIES OY
  • EP3238461B1 patent drawingFigure 1~2
  • EP3238461B1 patent drawingFigure 3~4
  • EP3238461B1 patent drawingFigure 5

AI summary

A method including generating respective audio signals from microphones of an apparatus; determining which camera(s) of a plurality of cameras of the apparatus has been selected for use; and based upon the determined camera(s) selected for use, selecting an audio processing mode for at least one of the respective audio signals to be processed, where the audio processing mode at least partially automatically adjusts the at least one respective audio signals based upon the determined camera(s) selected for use.