User-Agnostic Audio Encoding for Personalized Virtual Communication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio rendering techniques, such as ray-tracing, struggle to accurately simulate audio interactions in virtual environments, failing to provide personalized audio communication experiences across multiple users.

Innovation Solution

The implementation of user-agnostic audio encoding and user-specific decoding using artificial neural networks (ANNs) in audio communication nodes, where encoded audio data is generated independently of the user and decoded specifically for each user, ensuring personalized audio reproduction across interconnected devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If user-specific encoding parameters are used for each user, then personalized audio reproduction quality is improved, but system complexity and processing overhead increase

Engineering Contradiction:
Improveaudio reproduction qualityVSAvoidencoding system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system separates the encoding and decoding functions, with encoding being user-agnostic and decoding being user-specific. This segmentation allows the encoding stage to remain simple while the decoding stage handles personalization, resolving the contradiction between personalized quality and system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces encoded audio data as an intermediary representation that is user-agnostic. This intermediary form allows different users to have their own decoding parameters applied to the same encoded data, achieving personalized reproduction without requiring complex user-specific encoding processes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If user-agnostic encoded audio data is generated, then processing efficiency and scalability are improved, but personalized audio experiences for different users are reduced

Engineering Contradiction:
Improveaudio processing efficiencyVSAvoidpersonalized audio reproduction
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system divides the audio processing into two distinct stages: user-agnostic encoding that achieves high efficiency and scalability, and user-specific decoding that restores personalized audio characteristics. This segmentation allows each stage to optimize for its specific function.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The encoding process performs preliminary compression and representation of audio data in a user-agnostic format before transmission or storage. The personalized reproduction is then achieved in the decoding stage using pre-configured user-specific parameters, maintaining both efficiency and personalization.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple user-specific decoding parameters are maintained, then personalized audio communication experiences are improved, but data storage requirements increase

Engineering Contradiction:
Improveaudio communication personalizationVSAvoidstored parameter data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

Instead of storing complete user-specific audio profiles for each user, the system inverts the approach by storing compact decoding parameters that are specific to each user. The encoding remains universal and user-agnostic, reducing the storage burden while maintaining personalization capabilities.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS12142283B2Audio processing
Publication Date: 2024.11.12 SONY INTERACTIVE ENTERTAINMENT LLC
  • US12142283B2 patent drawing
  • US12142283B2 patent drawing
  • US12142283B2 patent drawing

AI summary

Audio communication apparatus comprises a set of two or more audio communication nodes; each audio communication node comprising: an audio encoder controlled by encoding parameters to generate encoded audio data to represent a vocal input generated by a user of that audio communication node, the encoded data being agnostic to which user who generated the vocal input; and an audio decoder controlled by decoding parameters to generate a decoded audio signal as a reproduction of a vocal signal generated by a user of another of the audio communication nodes, the decoding parameters being specific to the user of that other of the audio communication nodes.