Audio Masker Library for Packet Loss Concealment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice communication over packet-switch networks is plagued by issues such as packet losses and artifacts due to network delays and interference, leading to unnatural sound and poor listener experience, especially when silence is misinterpreted as network failure.

Innovation Solution

An audio processing apparatus and method that utilize audio maskers, such as non-stationary noise, filled pauses, and discourse markers, to conceal defects in audio signals by building a masker library and inserting appropriate maskers into target positions in the audio signal to create a more natural listening experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If packet loss concealment methods such as interpolation or extrapolation are used, then packet losses can be concealed, but artifacts occur in the voice and the heard voice sounds unnatural

Engineering Contradiction:
Improvepacket loss concealmentVSAvoidartifacts in voice
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces audio maskers as an intermediary element to conceal packet losses and artifacts. Instead of directly interpolating or extrapolating lost audio data, the system inserts pre-processed masker audio segments that mask the defects. These maskers are selected from a library based on contextual matching, providing a natural-sounding concealment that avoids the artificial artifacts produced by traditional methods.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary processing of audio segments to create a masker library before actual packet loss occurs. Audio segments are pre-processed to extract masker candidates, and statistics about contextual information are obtained in advance. When packet loss occurs, the system can quickly select and insert appropriate pre-processed maskers without needing to process audio in real-time, thus avoiding artifacts while maintaining natural sound quality.

Inventive Principle:
Principle #10Preliminary action

2Object-affected harmful factors

If background noise is completely suppressed or empty packets are transmitted, then noise is removed, but talker's silence is misunderstood as network failure

Engineering Contradiction:
Improvebackground noiseVSAvoidsilence detection accuracy
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent uses audio maskers as an intermediary to bridge the gap between noise suppression and silence detection. When background noise is suppressed, the system can insert audio maskers during silence periods to indicate that the silence is intentional (talker's pause) rather than a network failure. This intermediary element preserves the benefits of noise suppression while maintaining reliable silence detection.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback mechanisms by analyzing contextual information of audio segments and using this information to make informed decisions about when to insert audio maskers. The masker selection process considers contextual statistics to determine whether silence should be preserved or masked, providing feedback-based control that prevents misinterpretation of silence as network failure while maintaining effective noise suppression.

Inventive Principle:
Principle #23Feedback

3Object-affected harmful factors

If audio maskers are inserted to conceal defects, then perceived quality improves, but device complexity increases due to masker library management

Engineering Contradiction:
Improveaudio defectsVSAvoidmasker library management
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent performs preliminary processing of audio segments to create a masker library in advance. Audio segments are pre-processed to extract masker candidates, and contextual statistics are computed beforehand. This preliminary action reduces the complexity of real-time defect concealment, as the system only needs to select and insert pre-processed maskers rather than processing audio in real-time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system manages masker library complexity by changing parameters such as the size and organization of the masker library. The patent describes building masker libraries with controlled numbers of audio maskers and using contextual matching to select appropriate maskers. By optimizing these parameters, the system balances the quality improvement from masker insertion with the complexity of library management.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2936489B1Audio processing apparatus and audio processing method
Publication Date: 2017.04.19 DOLBY LABORATORIES LICENSING CORP
  • EP2936489B1 patent drawingFigure 1A~1B
  • EP2936489B1 patent drawingFigure 2~3B
  • EP2936489B1 patent drawingFigure 4~5

AI summary

An audio processing apparatus and an audio processing method are described. In one embodiment, the audio processing apparatus include an audio masker separator for separating from a first audio signal an audio material comprising a sound other than stationary noise and utterance meaningful in semantics, as an audio masker candidate. The apparatus also includes a first context analyzer for obtaining statistics regarding contextual information of detected audio masker candidates, and a masker library builder for building a masker library or updating an existing masker library by adding, based on the statistics, at least one audio masker candidate as an audio masker into the masker library, wherein audio maskers in the maker library are used to be inserted into a target position in a second audio signal to conceal defects in the second audio signal.