Teleconference Audio Packet Reordering for Playback Fidelity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current teleconferencing systems face challenges in recording and playback, where the audio experience during playback differs significantly from the original teleconference due to latency issues and data rate constraints, leading to incomplete or delayed audio data packets, affecting the quality and accuracy of the recorded audio.

Innovation Solution

The system records and processes individual uplink data packet streams, reorders and completes late or missing packets, and applies spatial and temporal adjustments to enhance the playback experience, using a combination of hardware and software components to analyze and render audio data for improved perceptual quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a dummy client records the downlink audio stream during teleconference, then the recording process is simple and easy to implement, but the playback quality is degraded due to latency and data rate constraints

Engineering Contradiction:
Improverecording implementation simplicityVSAvoidplayback audio quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent segments the audio recording process by recording individual uplink streams from each conference participant separately, rather than recording a single mixed downlink stream. This segmentation allows each participant's audio to be processed and reordered independently, improving playback quality while maintaining implementation simplicity through modular recording of separate streams

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by reordering audio packets in advance based on timing information before playback. The system analyzes packet arrival times and reconstructs the original temporal sequence of speech, preparing the audio data in advance to compensate for network latency and ensure accurate playback timing

Inventive Principle:
Principle #10Preliminary action

2Speed

If audio packets are transmitted in real-time during teleconference, then the conference can proceed without delay, but late or missing packets degrade the recording completeness

Engineering Contradiction:
Improveaudio transmission speedVSAvoidaudio data completeness
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The patent applies beforehand cushioning by implementing a buffering and reordering mechanism that anticipates late or missing packets. The system collects audio packets in a buffer, reorders them based on timing information, and waits for potentially late packets before playback, cushioning against the effects of network delays and packet loss to ensure complete audio recording

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

The patent uses feedback by analyzing the timing and sequence of received audio packets to detect late or missing data. The system monitors packet arrival patterns, identifies gaps in the audio stream, and uses this feedback information to request retransmission or interpolate missing data, ensuring complete audio recording

Inventive Principle:
Principle #23Feedback

3Measurement precision

If individual uplink streams are recorded and processed with spatial and temporal adjustments, then playback accuracy is improved, but the system complexity increases

Engineering Contradiction:
Improveplayback accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary processing system that sits between the recorded uplink streams and the final playback. This intermediary component performs spatial and temporal adjustments, packet reordering, and synchronization operations, isolating the complexity from both the recording and playback ends while improving playback accuracy through centralized processing

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3254279B1Conference word cloud
Publication Date: 2018.11.21 DOLBY LABORATORIES LICENSING CORP
  • EP3254279B1 patent drawingFigure 1A
  • EP3254279B1 patent drawingFigure 1B
  • EP3254279B1 patent drawingFigure 1C

AI summary

Various disclosed implementations involve processing and/or playback of a recording of a conference involving a plurality of conference participants. Some implementations disclosed herein involve receiving speech recognition results data, including a plurality of speech recognition lattices and a word recognition confidence score for each of a plurality of hypothesized words of the speech recognition lattices, for a conference recording. A primary word candidate and alternative word hypotheses may be determined for hypothesized words in the speech recognition lattices. A term frequency metric may be calculated for sorting the primary word candidates and the alternative word hypotheses. Hypothesized words may be re-scored according to an alternative hypothesis list.