Audio Signal Generation via Visual Focal Plane Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems fail to effectively integrate audio signals with visual bokeh effects, as audio signals are typically unaffected by visual focus changes, blurring, or other enhancements in photographs and videos.

Innovation Solution

A method and apparatus that dynamically generate audio signals based on the positions of microphones relative to a visual focal plane, incorporating a level of bokeh in the visual content, allowing for the creation of an audio bokeh effect by amplifying or attenuating audio signals from microphones positioned accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If visual bokeh effects are applied to photographs and videos, then visual focus and depth perception are improved, but audio signals remain unaffected and fail to provide corresponding spatial information

Engineering Contradiction:
Improvevisual focus precisionVSAvoidaudio-visual synchronization
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies different audio processing quality to different spatial zones. Audio signals from microphones positioned near the visual focal plane are processed with higher fidelity and emphasis, while audio from other zones is attenuated or blurred, creating local quality differentiation that mirrors the visual bokeh effect

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent extends the two-dimensional visual bokeh concept into the audio dimension by introducing spatial audio processing. It creates an audio depth dimension that corresponds to the visual depth of field, allowing audio signals to be processed according to their spatial relationship with the visual focal plane

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If multiple microphones are used to capture audio from different positions, then audio coverage and spatial information are improved, but the complexity of dynamically generating audio signals based on microphone positions relative to the visual focal plane increases

Engineering Contradiction:
Improvespatial audio informationVSAvoidaudio signal processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs preliminary spatial mapping of microphone positions relative to the visual focal plane before audio processing. By pre-establishing the spatial relationships and processing priorities, the system reduces the computational complexity during real-time audio signal generation while maintaining accurate spatial information

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a virtual copy of the visual focal plane in the audio processing domain. This virtual focal plane serves as a reference model that guides the processing of multiple audio signals, allowing the system to manage complex multi-microphone inputs through a simplified virtual reference framework

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20210368107A1Method, apparatus and computer program product for generating audio signals according to visual content
Publication Date: 2021.11.25 NOKIA TECHNOLOGIES OY
  • US20210368107A1 patent drawing
  • US20210368107A1 patent drawing
  • US20210368107A1 patent drawing

AI summary

A method, apparatus and computer program product are provided for generating audio signals according to visual content. Bokeh refers to a blurring of areas in a photograph or video that are in front of or behind a visual focal plane. Microphones in different positions in the environment of the captured visual content may be mixed according to their positions relative to a visual focal plane of the visual content to generate audio signals. Audio signals may be generated further dependent on a visual effect applied to captured content.