Audio Tag Insertion for Stream Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio streaming technologies lack efficient methods for segmenting and processing audio streams on the receiving side, particularly in voice-attached distribution services, where precise identification and handling of sound units are necessary for sound output and caption display.
Innovation Solution
A transmitting apparatus and method that inserts tag information into audio frames to indicate the presence of predetermined sound units within the audio stream, allowing for easy segmentation and processing of audio data on the receiving side, including sound unit identification, generation source information, and frame positioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If tag information is inserted into audio frames to identify sound units, then the ease of processing audio streams on the receiving side is improved, but the device complexity increases due to additional information inserting section
Solution Approach 1:
The audio stream is segmented into discrete audio frames, with each frame containing tag information that identifies specific sound units. This segmentation allows the receiving side to efficiently process and identify different sound units without analyzing the entire continuous audio stream, thereby improving ease of processing while the structured segmentation approach keeps the added complexity manageable.
Solution Approach 2:
Tag information identifying sound units is inserted into audio frames during the transmitting side processing before transmission. This preliminary action prepares the audio stream in advance, so that the receiving side can directly utilize the pre-organized tag information for efficient processing without needing to perform complex analysis operations, thus improving ease of operation with minimal additional complexity at the receiving end.
2Measurement precision
If tag information is inserted into each audio frame, then the measurement precision of sound unit identification is improved, but the loss of information increases due to additional metadata
Solution Approach 1:
Specific identifying information about sound units is extracted and placed into tag information within audio frames. This extraction approach allows the receiving side to access only the necessary identification data without processing the entire audio content, improving measurement precision of sound unit identification while minimizing the amount of additional information that needs to be transmitted and stored.
Solution Approach 2:
Tag information is inserted selectively into specific audio frames that contain sound units, rather than uniformly into all frames. This local quality approach ensures that identification information is available precisely where needed (in frames containing sound units), improving measurement precision while reducing the overall amount of tag information that would otherwise be distributed throughout the entire audio stream, thereby minimizing information loss.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A process of an audio stream on a receiving side is facilitated. Encoding processing is performed on audio data and an audio stream in which an audio frame including audio compression data is continuously arranged is generated. Tag information indicating that the audio compression data of a predetermined sound unit is included is inserted into the audio frame including the audio compression data of the predetermined sound unit. A container stream of a predetermined format including the audio stream into which the tag information is inserted is transmitted.