3D Audio Visual Representation for Speaker Calibration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Listeners often face difficulty in aurally locating 3D audio objects within a stream due to confusion in directionality, particularly in areas like the Cone of Confusion, and may not understand optimal speaker placement, which can affect the audio experience.

Innovation Solution

A visual representation of 3D audio objects is created on a local playback device using metadata, where audio objects are graphically represented by spheres with size indicating amplitude and color representing instrument type, allowing real-time updates and visual cues for optimal sound field calibration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If a separate video stream is created to provide visual representation of 3D audio objects, then the visual information quality is improved, but the bandwidth requirements increase

Engineering Contradiction:
Improvevisual information qualityVSAvoidbandwidth requirements
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent merges the visual representation function into the existing audio stream by embedding visual cues within the audio metadata structure. The video processor generates visual representations of audio objects using the same metadata that describes audio object positions, eliminating the need for a separate video stream while providing comprehensive visual information about the 3D audio scene

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The audio metadata structure is made multi-functional by using it to convey both audio object information and visual representation data. The same metadata fields that describe audio object location, movement, and characteristics are repurposed to generate visual cues, allowing the audio stream to serve dual audio-visual functions

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If visual representation of 3D audio objects is added to help listeners locate sound objects, then the ease of operation is improved, but the device complexity increases

Engineering Contradiction:
Improveease of locating sound objectsVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system uses the existing audio object metadata to automatically generate visual representations without requiring separate input data or complex processing pipelines. The visual processor leverages the already-present audio object information (position, movement, characteristics) to create visual cues, making the system self-sufficient and reducing overall complexity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces a visual processor as an intermediary component that translates audio metadata into visual representations. This mediator layer simplifies the architecture by handling all visual generation tasks in one dedicated module, rather than distributing complexity across multiple components

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11102606B1Video component in 3D audio
Publication Date: 2021.08.24 SONY GROUP CORP
  • US11102606B1 patent drawing
  • US11102606B1 patent drawing
  • US11102606B1 patent drawing

AI summary

A visual component is added to a 3D audio stream to present on a display at the player side a visual representation of objects in the 3D audio, enabling the user to better understand what is happening in the 3D audio experience. The visual representation may include visual objects with the same location and movement in 3D space as audio objects being played.