Audio Renderer Joint Audio-Visual Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio systems do not perform joint optimization using audio and visual information, failing to exploit visual features for enhanced audio processing and spatial audio rendering, which limits the accuracy of sound localization and immersion in audiovisual experiences.

Innovation Solution

A machine learning model processes synchronized audio and video streams to infer joint latent representations, enabling spatial audio enhancement by correlating audio and visual cues, and generating output audio channels that are spatially mapped to a target scene, using features such as microphone signals, video signals, and metadata to drive speaker arrangements like binaural or surround sound formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio and visual information are processed separately in traditional audio systems, then the processing complexity is reduced and systems are easier to implement, but the accuracy of sound localization and spatial audio rendering deteriorates

Engineering Contradiction:
Improvesound localization accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges audio and visual processing into a unified joint optimization framework. The system processes audio features (microphone signals, spatial cues) and visual features (video frames, spatial information) simultaneously through integrated algorithms that correlate information between both modalities, thereby improving sound localization accuracy while maintaining manageable system complexity through unified processing architecture

Inventive Principle:
Principle #5Merging (Combining)

2Manufacturing precision

If visual features are not exploited in audio processing, then the audio processing algorithms are simpler and faster, but the quality of spatial audio rendering and immersion deteriorates

Engineering Contradiction:
Improvespatial audio rendering qualityVSAvoidaudio processing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs preliminary extraction of spatial features from both audio and visual streams before the main rendering process. Visual spatial information (object positions, scene geometry) and audio spatial cues are pre-processed and correlated in advance, allowing the main audio rendering algorithm to operate efficiently while benefiting from pre-computed spatial relationships, thus maintaining processing speed while improving rendering quality

Inventive Principle:
Principle #10Preliminary action

3Reliability

If audio information is not considered in video processing algorithms, then the video processing is simpler and more efficient, but the joint audiovisual optimization and immersion deteriorates

Engineering Contradiction:
Improveaudiovisual correlation accuracyVSAvoidvideo processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces spatial temporal correlation as an intermediary mechanism that links audio and video processing. By computing correlation metrics between audio events and visual events in space and time, the system enables bidirectional optimization where video processing can leverage audio information for enhanced spatial accuracy, and audio processing benefits from visual spatial context, without requiring full integration of both processing pipelines

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12010490B1Audio renderer based on audiovisual information
Publication Date: 2024.06.11 APPLE INC
  • US12010490B1 patent drawing
  • US12010490B1 patent drawing
  • US12010490B1 patent drawing

AI summary

An audio renderer can have a machine learning model that jointly processes audio and visual information of an audiovisual recording. The audio renderer can generate output audio channels. Sounds captured in the audiovisual recording and present in the output audio channels are spatially mapped based on the joint processing of the audio and visual information by the machine learning model. Other aspects are described.