Conference Audio Playback Spatial Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current teleconferencing systems face challenges in recording and playback, where the audio experience during playback differs significantly from the original conference, due to latency issues and data rate constraints, leading to incomplete or delayed audio packets, and lack of spatial and temporal accuracy.

Innovation Solution

The system records individual uplink data packet streams, reorders and decodes them for playback, incorporating spatial information and conversational dynamics to recreate the original audio experience, allowing for more complete and accurate playback with improved spatial rendering of conference participants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a single monophonic stream containing a mix of all parties is recorded, then the recording system is simple and easy to implement, but the playback experience lacks spatial accuracy and does not faithfully represent the original conference

Engineering Contradiction:
Improverecording system implementationVSAvoidspatial accuracy of playback
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent segments the mixed audio stream into individual uplink data packet streams for each conference participant. Each participant's audio is separated and recorded independently, allowing spatial positioning and individual volume control during playback. This segmentation enables faithful reproduction of the original conference spatial relationships while maintaining manageable system complexity through structured data organization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds spatial dimensionality to the playback by assigning three-dimensional positions to each conference participant based on their original seating arrangement. This transforms the flat monophonic recording into a spatially-aware audio experience, allowing listeners to perceive the original conference geometry and participant locations without complicating the core recording mechanism.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Speed

If audio packets are transmitted in real-time during the conference, then the system operates with low latency during the conference, but data rate constraints cause incomplete or delayed audio packets that degrade playback quality

Engineering Contradiction:
Improvereal-time transmission speedVSAvoidcompleteness of audio packets
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent performs preliminary organization and labeling of audio packets during the conference, assigning unique identifiers and spatial metadata to each uplink stream. This preliminary action ensures that even if packets arrive delayed or out of order during real-time transmission, they can be correctly reassembled and positioned during playback, maintaining both speed and reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms where the teleconferencing bridge acknowledges received packets and tracks transmission status. This feedback loop allows the system to monitor data completeness and request retransmission of missing packets, ensuring reliable playback while maintaining efficient real-time operation through selective retransmission rather than waiting for complete buffer fills.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If individual uplink data packet streams are recorded and processed with spatial information, then the playback experience is more faithful and spatially accurate, but the system complexity and processing requirements increase

Engineering Contradiction:
Improvefaithfulness of playback representationVSAvoidsystem processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal processing framework that handles both simple and complex playback scenarios through the same infrastructure. The system can reproduce individual participant audio, spatial positions, or mixed configurations based on playback requirements, eliminating the need for separate processing pipelines for different playback modes and reducing overall system complexity despite enhanced capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Loss of information

If the recording captures complete audio data from all participants, then the playback is more complete and accurate, but the data rate requirements and storage needs increase

Engineering Contradiction:
Improvecompleteness of conference audioVSAvoiddata rate and storage requirements
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent applies local quality enhancement by allowing selective adjustment of audio parameters for individual participants during playback. Each participant's audio stream can be independently optimized for volume, spatial position, and frequency characteristics, ensuring that important local details are preserved without requiring uniform high-bitrate encoding for all participants, thus reducing overall data requirements while maintaining completeness.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3254454B1Conference searching and playback of search results
Publication Date: 2020.12.30 DOLBY LABORATORIES LICENSING CORP
  • EP3254454B1 patent drawingFigure 1A
  • EP3254454B1 patent drawingFigure 1B
  • EP3254454B1 patent drawingFigure 1C

AI summary

Various disclosed implementations involve processing and/or playback of a recording of a conference involving a plurality of conference participants. Some implementations disclosed herein involve receiving audio data corresponding to a recording of at least one conference involving a plurality of conference participants. The audio data may include conference participant speech data from multiple endpoints, recorded separately and/or conference participant speech data from a single endpoint corresponding to multiple conference participants and including spatial information for each conference participant of the multiple conference participants. A search of the audio data may be based on one or more search parameters. The search may be a concurrent search for multiple features of the audio data. Instances of conference participant speech may be rendered to at least two different virtual conference participant positions of a virtual acoustic space.