Audio Track Segregation for Adaptive Streaming Reproduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies do not effectively facilitate the easy reproduction of audio data from multiple groups, particularly in streaming services that utilize adaptive streaming methods like MPEG-DASH, which primarily focus on video distribution.

Innovation Solution

An information processing apparatus and method that divide audio data into tracks for each kind and arrange related information, enabling efficient reproduction of specific audio tracks by utilizing a file structure that segregates audio streams and metadata into distinct tracks, allowing for adaptive streaming and reduced decoding load.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If audio data of multiple kinds is stored in a single track, then the file structure is simple, but the reproduction efficiency and ease of selecting specific audio tracks deteriorates

Engineering Contradiction:
Improveease of reproductionVSAvoidfile structure complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent divides audio data of multiple kinds into separate tracks within the same container. Each track contains audio data of a specific kind (e.g., voice, music, effect sounds), allowing independent selection and reproduction. This segmentation enables the reproduction terminal to efficiently identify and reproduce only the required audio tracks without processing unnecessary data, thereby improving ease of operation while maintaining manageable file structure through standardized track organization

Inventive Principle:
Principle #1Segmentation

2Productivity

If all audio data is decoded, then complete audio information is available, but the decoding load and processing time increase

Engineering Contradiction:
Improvereproduction efficiencyVSAvoiddecoding load
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent extracts and separates audio data by kind into distinct tracks, allowing the reproduction terminal to selectively decode only the necessary tracks based on user requirements. For example, if only voice audio is needed, the terminal can extract and decode solely the voice track, ignoring music and effect sound tracks. This extraction approach significantly reduces decoding load and processing energy while maintaining complete audio information availability for the selected kinds

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of decoding all audio data completely (excessive action), the system enables partial decoding by allowing the reproduction terminal to process only the necessary portions (required audio kinds). This partial action approach optimizes the balance between processing effort and output quality, avoiding unnecessary decoding operations while ensuring complete reproduction of selected audio tracks

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If audio data is not organized by tracks, then the storage structure is simple, but the adaptability for different reproduction scenarios deteriorates

Engineering Contradiction:
Improveadaptability for reproductionVSAvoiddata organization structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal track-based organization structure that serves multiple reproduction scenarios simultaneously. The same track structure supports various use cases including: reproducing specific audio kinds (voice, music, effects), combining multiple audio tracks, and adapting to different device capabilities. This multi-functional design enhances adaptability without requiring separate storage formats for different scenarios, while the standardized track organization keeps the data structure manageable

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20210326378A1Information processing apparatus and information processing method
Publication Date: 2021.10.21 SONY GROUP CORP
  • US20210326378A1 patent drawing
  • US20210326378A1 patent drawing
  • US20210326378A1 patent drawing

AI summary

The present disclosure relates to an information processing apparatus and an information processing method that enable easy reproduction of audio data of a predetermined kind, of audio data of a plurality of kinds. A file generation device generates an audio file in which audio streams of a plurality of groups is divided into tracks for each one or more of the groups and arranged, and information related to the plurality of groups is arranged. The present disclosure can be applied to an information processing system configured from the file generation device that generates a file, a web server that records the file generated by the file generation device, and a moving image reproduction terminal that reproduces the file, for example.