Voice Data Position Recognition in MPEG-DASH Streaming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing streaming technologies, such as MPEG-DASH, do not consider the recognition of voice data acquisition positions on video content, limiting the ability to efficiently manage and play back voice data in relation to image data.
Innovation Solution
The proposed system transmits and recognizes information about the acquisition position of voice data on an image by encoding voice metadata with object position information and object IDs, allowing for precise identification and playback of voice data corresponding to specific image regions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If voice data is transmitted without position information, then transmission bandwidth is reduced, but the ability to recognize and selectively playback voice data from specific image regions is lost
Solution Approach 1:
The voice data transmission is segmented by associating it with specific image regions through position information. Each voice data file is linked to a particular region in the image, allowing the playback terminal to selectively acquire only the voice data corresponding to regions of interest, thereby improving playback efficiency without transmitting unnecessary voice data.
Solution Approach 2:
Position information adds a spatial dimension to the voice data transmission system. By incorporating coordinates or region identifiers, the system transitions from transmitting only audio content to transmitting audio content with spatial context, enabling selective playback based on image region analysis.
2Reliability
If all voice data is transmitted, then complete voice coverage is achieved, but transmission bandwidth and processing load increase
Solution Approach 1:
The system extracts and transmits only the necessary voice data files based on the analyzed image content. By identifying which regions in the image contain relevant information, the system extracts and transmits only the corresponding voice data, eliminating the need to transmit all voice data files and thereby reducing bandwidth consumption while maintaining reliability for the relevant content.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present disclosure relates to an information processing device and information processing method capable of recognizing an acquisition position of voice data on an image. A web server transmits image frame size information indicating image frame size of image data and audio position information indicating acquisition position of voice data. The present disclosure is applicable to an information processing system or other like system including file generation device, web server, and video playback terminal to perform tiled streaming using a manner compliant with moving picture experts group phase-dynamic adaptive streaming over HTTP (MPEG-DASH).