Spatial Audio Capture with Dynamic Metadata Updates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for transmitting spatial audio signals struggle to adapt dynamically to changes in capture settings, such as orientation or audio focus, which can affect the quality of the spatial audio experience.

Innovation Solution

An apparatus and method that capture spatial audio signals using multiple microphones and generate metadata-assisted spatial audio (MASA) format data, which includes descriptive metadata indicating capture parameters. This data is encoded and transmitted to remote devices for rendering, allowing for real-time adjustments to capture settings during streaming sessions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If spatial audio signals are transmitted with fixed capture settings, then transmission stability is maintained, but adaptability to changing device orientation and capture parameters deteriorates

Engineering Contradiction:
Improveadaptability to capture setting changesVSAvoidsystem complexity for dynamic metadata updates
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system dynamically updates descriptive metadata during the streaming session to reflect changes in capture settings such as device orientation and audio focus. The audio encoder and transmitter are configured to modify metadata parameters in real-time based on current capture conditions, enabling the spatial audio experience to adapt to changing environments without requiring system reconfiguration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of descriptive metadata to correspond with updated capture settings. When capture parameters such as orientation or audio focus change, the system modifies the metadata parameters accordingly and transmits the updated bitstream to the remote device, allowing the rendering to reflect the new capture conditions while maintaining system stability.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If metadata is updated in real-time during streaming, then spatial audio quality is improved, but data transmission overhead increases

Engineering Contradiction:
Improvespatial audio rendering accuracyVSAvoidmetadata data volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system extracts and transmits only the specific descriptive metadata parameters that are necessary for spatial audio rendering, rather than transmitting all possible metadata. By selectively updating only the changed parameters in the metadata (such as orientation angles or audio focus directions), the system reduces the overall data volume while maintaining rendering accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies partial updates to the metadata by modifying only the specific parameters that have changed during the streaming session, rather than retransmitting the entire metadata set. This partial action approach maintains spatial audio quality by updating necessary parameters while minimizing unnecessary data transmission.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If multiple capture parameters are monitored and updated, then spatial audio adaptability is enhanced, but processing complexity increases

Engineering Contradiction:
Improvecapture setting adaptabilityVSAvoidautomatic parameter monitoring and updating
Core Design Contradiction:
Adaptability or versatilityVSExtent of automation

Solution Approach 1:

The audio encoder is designed with multi-functionality to handle multiple capture parameters simultaneously. It can monitor and update various metadata parameters including orientation, audio focus, and other capture settings within a single unified processing framework, reducing the need for separate processing systems for each parameter type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges the monitoring and updating of multiple capture parameters into a unified metadata update process. By combining the tracking of orientation changes, audio focus adjustments, and other capture settings into a single coordinated system, the patent reduces processing complexity while maintaining comprehensive adaptability to capture setting changes.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250080939A1Spatial audio
Publication Date: 2025.03.06 NOKIA TECHNOLOGIES OY
  • US20250080939A1 patent drawing
  • US20250080939A1 patent drawing
  • US20250080939A1 patent drawing

AI summary

There is herein provided a method comprising capturing spatial audio signals by a plurality of microphones using a first capture setting and generating a set of audio encoder input format data comprising a representation of the spatial audio signals and associated metadata, the metadata including, in part, a set of descriptive metadata indicative of one or more capture parameters associated with the first capture setting. The method further comprises providing the set of audio encoder input format data to an audio encoder for encoding the representation of the spatial audio signals and the associated metadata to a bitstream for a current streaming session and transmitting the bitstream to one or more remote devices for rendering the representation of the spatial audio signals based, at least in part, on the set of descriptive metadata. The method further comprises determining, during the current streaming session, that the first capture setting changes to a second capture setting and in response to the determination that the first capture setting changes to a second capture setting, and while transmitting the bitstream, the means for capturing spatial audio signals is configured to capture spatial audio signals using the second capture setting and the means for generating the audio encoder input format data is configured to change at least one of the one or more capture parameters of the set of descriptive metadata to be associated with the second capture setting.