Spatial Voice Data Transmission for Wide Viewing Angle Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in providing an efficient method for obtaining a voice output corresponding to a fixed position of a wide viewing angle image, particularly in scene-based audio transmission.
Innovation Solution
A transmission apparatus and method that transmit spatial voice data and information regarding registered viewpoints, using scene-based audio in an HoA format, packaged within an MPEG-H audio stream, allowing for precise voice output synchronization with wide viewing angle images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If spatial voice data and viewpoint information are transmitted using scene-based audio methods, then voice output corresponding to fixed image positions can be obtained, but the system complexity and processing requirements increase
Solution Approach 1:
The patent applies preliminary action by pre-defining and transmitting viewpoint information (azimuth angles, elevation angles) along with spatial voice data before reproduction. This allows the reproduction apparatus to directly use these pre-calculated parameters to generate synchronized voice output without complex real-time calculations, thereby maintaining reliability while reducing processing complexity.
Solution Approach 2:
The patent introduces an intermediary approach by using standardized viewpoint information structures (including azimuth and elevation angles) as a mediator between the audio data and the reproduction process. This intermediary framework enables systematic transmission and processing of spatial audio data, reducing overall system complexity while ensuring reliable synchronization.
2Adaptability or versatility
If viewpoint information is transmitted for each user purpose, then customized voice output can be achieved, but the data transmission volume and processing load increase
Solution Approach 1:
The patent applies segmentation by dividing viewpoint information into multiple groups, each corresponding to different user purposes or scenarios. This allows the transmission apparatus to send only the relevant viewpoint groups needed for specific applications, reducing overall data transmission volume while maintaining adaptability for different user needs.
Solution Approach 2:
The patent implements universality by creating a standardized viewpoint information structure that can serve multiple user purposes. The same basic framework of azimuth and elevation angle data can be applied across different scenarios (e.g., different camera positions, different audio formats), reducing redundant data transmission while maintaining versatility for various applications.
Data Source
AI summary
A voice output corresponding to a fixed position of a wide viewing angle image is easily obtained.A transmission unit configured to transmit spatial voice data and information regarding a predetermined number of registered viewpoints is included. For example, the spatial voice data is data of scene-based audio. Then, for example, the data of the scene-based audio is each component of an HoA format. For example, the information regarding a viewpoint includes information regarding an azimuth angle (azimuth information) and an elevation angle (elevation angle information) that indicate a position of this viewpoint. For example, the data of the scene-based audio and the information regarding the predetermined number of registered viewpoints are transmitted with being included in a packet of object audio.


