Screen-Relative Audio Rendering via Dynamic Warping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio content creation and distribution pipelines face challenges in accurately reproducing the location and size of auditory images relative to visual images during playback in environments different from the authoring environment, such as from cinemas to home settings, due to discrepancies in speaker and screen configurations.
Innovation Solution
The method involves 'warping' audio channels by processing audio content to adjust perceived positions of audio elements based on screen-related metadata, allowing for smooth transitions between on-screen and off-screen locations relative to the playback system's display screen, thereby aligning audio with visual elements regardless of the playback environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio content is encoded with fixed speaker positions relative to a reference screen, then the audio program can be efficiently distributed and decoded, but the perceived positions of audio elements will be inaccurate when played back on screens with different sizes and configurations
Solution Approach 1:
The patent applies dynamics by making the audio rendering adaptable to different playback environments. The system dynamically adjusts the perceived positions of audio elements based on the actual screen size and speaker configuration at playback time, rather than using fixed positions. This is achieved through screen-relative rendering that calculates audio element positions as offsets from the screen edges, allowing the same audio program to accurately represent spatial positions across diverse playback configurations.
2Measurement precision
If the audio program is remixed for different playback environments, then the spatial accuracy can be improved, but the complexity of the content creation and distribution pipeline increases
Solution Approach 1:
The patent applies preliminary action by pre-calculating and encoding screen-relative position offsets during the audio production phase. Instead of requiring remixing for different environments, the system pre-processes the audio content to include position information relative to screen edges. This preliminary action enables accurate spatial rendering across different playback configurations without requiring subsequent remixing operations, thereby simplifying the distribution pipeline.
Solution Approach 2:
The patent applies parameter changes by transforming the audio position representation from absolute coordinates relative to a reference screen to relative offsets from screen edges. This parameter transformation allows the same audio data to adapt to different screen sizes and speaker configurations. The encoding process incorporates these relative position parameters, and the playback system uses them to calculate accurate speaker positions based on the actual playback environment.
3Measurement precision
If screen-related metadata is included in the audio program, then screen-relative rendering can be achieved, but the data rate and encoding complexity increase
Solution Approach 1:
The patent applies taking out by extracting only the essential screen-relative position information needed for accurate rendering, rather than including comprehensive screen configuration data. The system encodes position offsets for audio elements relative to screen edges, which are the minimum necessary parameters to achieve accurate audio-visual alignment. This selective extraction of critical metadata minimizes the data rate increase while maintaining the capability for screen-relative rendering across different playback environments.
Data Source
AI summary
In some embodiments, methods for generating an object based audio program including screen-elated metadata indicative of at least one warping degree parameter for at least one audio object, or generating a speaker channel-based program including by warping audio content of an object based audio program to a degree determined at least in part by at least one warping degree parameter, or methods for decoding or rendering any such audio program. Other aspects are systems configured to perform such audio signal generation, decoding, or rendering, and audio processing units (e.g., decoders or encoders) including a buffer memory which stores at least one segment of any such audio program.


