Audio Playback Jitter Buffer Adjustment for Music Frames
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio playback technologies do not distinguish between speech and music frames during network communications, leading to uneven playback and delays, especially under conditions of high network jitter and packet loss, which disrupt the smooth playback of music.
Innovation Solution
An audio playback method and system that identifies the type of audio data frames, adjusts the jitter buffer threshold based on network transmission status, specifically increasing the threshold for music frames during high packet loss or jitter to ensure smoother playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech frames and music frames are treated equally with the same playback duration, then the system is simple to operate, but music playback becomes unsmooth and delays occur during voice communication
Solution Approach 1:
The patent segments audio frames into different types (speech frames and music frames) and applies different playback durations to each type. This segmentation allows the system to handle speech and music differently, improving music playback smoothness while maintaining simple operation for speech communication.
2Device complexity
If the jitter buffer threshold is kept fixed, then the system is simple to manage, but music frame playback becomes unsmooth under high network jitter and packet loss
Solution Approach 1:
The patent implements dynamic adjustment of the jitter buffer threshold based on network conditions and audio frame type. The threshold is increased for music frames under high jitter conditions, allowing more time for packet retransmission and smoothing playback. This dynamic approach improves music playback quality while keeping the system manageable through automated adaptation.
3Speed
If audio frames are transmitted without type identification, then the transmission process is fast and simple, but playback delays occur and music realism cannot be achieved
Solution Approach 1:
The patent performs preliminary identification of audio frame types (speech or music) during the transmission phase. This early classification allows the receiving end to apply appropriate playback durations and jitter buffer adjustments, ensuring music realism is achieved without significantly impacting transmission speed due to the efficient identification process.
Data Source
AI summary
An audio playback method is provided. The method includes identifying a captured audio data frame according to a type of the audio data frame and sending the identified audio data frame to an audio receiving end. The method also includes receiving the audio data frame that is identified according to the type of the audio data frame and determining the type of the audio data frame and evaluating network transmission status based on the identification. Further, the method includes adjusting a threshold value of a jitter buffer that is used to cache the audio data frame when the type of the audio data frame is a music frame and evaluation result of the network transmission status does not meet a preset transmission baseline condition.


