Audio Masker Library for Packet Loss Concealment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice communication over packet-switch networks is plagued by issues such as packet losses and artifacts due to network delays and interference, leading to unnatural sound and poor listener experience, especially when silence is misinterpreted as network failure.
Innovation Solution
An audio processing apparatus and method that utilize audio maskers, such as non-stationary noise, filled pauses, and discourse markers, to conceal defects in audio signals by building a masker library and inserting appropriate maskers into target positions in the audio signal to create a more natural listening experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If packet loss concealment methods such as interpolation or extrapolation are used, then packet losses can be concealed, but artifacts occur in the voice and the heard voice sounds unnatural
Solution Approach 1:
The patent introduces audio maskers as an intermediary element to conceal packet losses and artifacts. Instead of directly interpolating or extrapolating lost audio data, the system inserts pre-processed masker audio segments that mask the defects. These maskers are selected from a library based on contextual matching, providing a natural-sounding concealment that avoids the artificial artifacts produced by traditional methods.
Solution Approach 2:
The patent performs preliminary processing of audio segments to create a masker library before actual packet loss occurs. Audio segments are pre-processed to extract masker candidates, and statistics about contextual information are obtained in advance. When packet loss occurs, the system can quickly select and insert appropriate pre-processed maskers without needing to process audio in real-time, thus avoiding artifacts while maintaining natural sound quality.
2Object-affected harmful factors
If background noise is completely suppressed or empty packets are transmitted, then noise is removed, but talker's silence is misunderstood as network failure
Solution Approach 1:
The patent uses audio maskers as an intermediary to bridge the gap between noise suppression and silence detection. When background noise is suppressed, the system can insert audio maskers during silence periods to indicate that the silence is intentional (talker's pause) rather than a network failure. This intermediary element preserves the benefits of noise suppression while maintaining reliable silence detection.
Solution Approach 2:
The system implements feedback mechanisms by analyzing contextual information of audio segments and using this information to make informed decisions about when to insert audio maskers. The masker selection process considers contextual statistics to determine whether silence should be preserved or masked, providing feedback-based control that prevents misinterpretation of silence as network failure while maintaining effective noise suppression.
3Object-affected harmful factors
If audio maskers are inserted to conceal defects, then perceived quality improves, but device complexity increases due to masker library management
Solution Approach 1:
The patent performs preliminary processing of audio segments to create a masker library in advance. Audio segments are pre-processed to extract masker candidates, and contextual statistics are computed beforehand. This preliminary action reduces the complexity of real-time defect concealment, as the system only needs to select and insert pre-processed maskers rather than processing audio in real-time.
Solution Approach 2:
The system manages masker library complexity by changing parameters such as the size and organization of the masker library. The patent describes building masker libraries with controlled numbers of audio maskers and using contextual matching to select appropriate maskers. By optimizing these parameters, the system balances the quality improvement from masker insertion with the complexity of library management.
Data Source
Figure 1A~1B
Figure 2~3B
Figure 4~5
AI summary
An audio processing apparatus and an audio processing method are described. In one embodiment, the audio processing apparatus include an audio masker separator for separating from a first audio signal an audio material comprising a sound other than stationary noise and utterance meaningful in semantics, as an audio masker candidate. The apparatus also includes a first context analyzer for obtaining statistics regarding contextual information of detected audio masker candidates, and a masker library builder for building a masker library or updating an existing masker library by adding, based on the statistics, at least one audio masker candidate as an audio masker into the masker library, wherein audio maskers in the maker library are used to be inserted into a target position in a second audio signal to conceal defects in the second audio signal.