Voice Data Position Recognition in MPEG-DASH Streaming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing streaming technologies, such as MPEG-DASH, do not consider the recognition of voice data acquisition positions on video content, limiting the ability to efficiently manage and play back voice data in relation to image data.

Innovation Solution

The proposed system transmits and recognizes information about the acquisition position of voice data on an image by encoding voice metadata with object position information and object IDs, allowing for precise identification and playback of voice data corresponding to specific image regions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If voice data is transmitted without position information, then transmission bandwidth is reduced, but the ability to recognize and selectively playback voice data from specific image regions is lost

Engineering Contradiction:
Improvevoice data acquisition position informationVSAvoidplayback efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The voice data transmission is segmented by associating it with specific image regions through position information. Each voice data file is linked to a particular region in the image, allowing the playback terminal to selectively acquire only the voice data corresponding to regions of interest, thereby improving playback efficiency without transmitting unnecessary voice data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Position information adds a spatial dimension to the voice data transmission system. By incorporating coordinates or region identifiers, the system transitions from transmitting only audio content to transmitting audio content with spatial context, enabling selective playback based on image region analysis.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If all voice data is transmitted, then complete voice coverage is achieved, but transmission bandwidth and processing load increase

Engineering Contradiction:
Improvevoice data completenessVSAvoidtransmission bandwidth
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system extracts and transmits only the necessary voice data files based on the analyzed image content. By identifying which regions in the image contain relevant information, the system extracts and transmits only the corresponding voice data, eliminating the need to transmit all voice data files and thereby reducing bandwidth consumption while maintaining reliability for the relevant content.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP3024249B1Information processing device and information processing method
Publication Date: 2025.05.07 SONY GROUP CORP
  • EP3024249B1 patent drawingFigure 1
  • EP3024249B1 patent drawingFigure 2
  • EP3024249B1 patent drawingFigure 3

AI summary

The present disclosure relates to an information processing device and information processing method capable of recognizing an acquisition position of voice data on an image. A web server transmits image frame size information indicating image frame size of image data and audio position information indicating acquisition position of voice data. The present disclosure is applicable to an information processing system or other like system including file generation device, web server, and video playback terminal to perform tiled streaming using a manner compliant with moving picture experts group phase-dynamic adaptive streaming over HTTP (MPEG-DASH).