Object-Based Audio Creation from Text via Semantic Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio book and audio play technologies lack the ability to provide an immersive 3D listening experience, as they are limited to channel-based formats that do not allow for the dynamic mixing of audio in a 3D space, missing the benefits of object-based audio content.
Innovation Solution
A method for creating object-based audio content from text input by performing semantic analysis to identify origins of speech and effects, synthesizing speech and effects, and generating metadata to create audio objects, which can be rendered in a channel-based format for an immersive listening experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If channel-based audio formats are used for audio books and audio plays, then the system is simple and widely compatible, but the immersive 3D listening experience cannot be provided
Solution Approach 1:
The patent segments audio content into independent audio objects (speech, effects, music) that can be individually positioned and manipulated in 3D space. Each audio object is treated as a separate entity with its own spatial coordinates, allowing the system to create immersive 3D audio experiences while maintaining compatibility with existing channel-based playback systems through rendering engines that convert object-based audio to channel-based output.
2Adaptability or versatility
If object-based audio content is created for audio books and audio plays, then an immersive 3D listening experience is provided, but the processing complexity and computational requirements increase
Solution Approach 1:
The patent performs semantic analysis and audio object identification in advance during the audio book creation process. Text is analyzed to identify speakers, objects, and actions, and corresponding audio objects are created and positioned in 3D space beforehand. This preliminary processing allows the final playback to simply render the pre-processed audio objects without requiring complex real-time analysis, reducing computational burden during actual use.
3Adaptability or versatility
If traditional text-to-speech processing is used, then the system is simple and fast, but the ability to create emotional and immersive audio experiences is limited
Solution Approach 1:
The patent introduces semantic analysis as an intermediary layer between text input and audio output. This intermediary analyzes the text to extract emotional context, speaker intent, and scene descriptions, then uses this information to guide audio object creation, speech synthesis parameter selection, and spatial positioning. This mediator enables rich emotional expression and immersion without requiring complete redesign of the TTS system.
Data Source
AI summary
Described herein is a method for creating object-based audio content from a text input for use in audio books and/or audio play, the method including the steps of: a) receiving the text input; b) performing a semantic analysis of the received text input; c) synthesizing speech and effects based on one or more results of the semantic analysis to generate one or more audio objects; d) generating metadata for the one or more audio objects; and e) creating the object-based audio content including the one or more audio objects and the metadata. Described herein are further a computer-based system including one or more processors configured to perform said method and a computer program product comprising a computer-readable storage medium with instructions adapted to carry out said method when executed by a device having processing capability.


