Object-Based Audio Format Conversion for Immersive Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio content systems fail to provide a user-customized immersive experience across various playback environments, as they are limited to playing back pre-mixed stereo content without considering the specific capabilities of the playback device.

Innovation Solution

A computer system that processes audio content by converting its format based on the playback environment, using metadata to render immersive experiences on devices with varying capabilities, including smartphones, audio receivers, and home cinemas.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio content is provided in a completed stereo form, then the content can be simply played back on any device, but the user cannot experience a customized immersive audio experience across various playback environments

Engineering Contradiction:
Improveadaptability to playback environmentsVSAvoidcomplexity of audio processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-processing audio content into an object-based format with embedded metadata during the content creation phase. This allows the audio objects and their spatial information to be prepared in advance, enabling flexible adaptation to different playback environments without requiring complex real-time processing during playback.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent utilizes parameter changes by dynamically adjusting audio rendering parameters based on the detected playback environment characteristics. The system modifies spatial audio parameters, channel configurations, and rendering settings according to the specific capabilities of the playback device, enabling the same content to deliver optimized immersive experiences across diverse environments.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If audio content is provided in a completed mixed form, then the playback process is simple, but the content cannot be optimized for specific playback device capabilities

Engineering Contradiction:
Improvesimplicity of playback processVSAvoidoptimization for playback devices
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent applies segmentation by separating the audio content into distinct audio objects with individual metadata rather than providing a pre-mixed stereo signal. This segmentation allows the playback system to independently process and position each audio object according to the specific capabilities of the playback device, achieving both simplified playback operations and device-optimized output.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If the same audio content format is used across all devices, then the system is easy to implement, but the immersive experience cannot be customized for different playback environments

Engineering Contradiction:
Improvecustomized immersive experienceVSAvoidcomplexity of format conversion
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary approach by using a standardized object-based audio format as a universal intermediate representation. This format serves as a mediator between content creation and diverse playback environments, allowing automatic adaptation to different devices through metadata interpretation without requiring complex manual format conversion or reprocessing for each device type.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12604151B2Computer system for processing audio content and method thereof
Publication Date: 2026.04.14 NAVER CORP
  • US12604151B2 patent drawing
  • US12604151B2 patent drawing
  • US12604151B2 patent drawing

AI summary

A computer system for processing audio content may receive content that includes metadata on spatial features about a plurality of objects, convert a format set according to a production environment of the content to a format according to a playback environment in an electronic apparatus, and transmit the content in the converted format to the electronic apparatus. The computer system may support content produced in various production environments and various playback environments.