Conference Audio Spatial Optimization via Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current teleconferencing systems record audio as a single monophonic stream, lacking spatial information and conversational dynamics, which affects the playback experience and regulatory compliance, especially in industries like banking.

Innovation Solution

The method involves processing audio data from multiple endpoints, analyzing conversational dynamics, and applying spatial optimization techniques to assign virtual positions in a virtual acoustic space, considering factors like doubletalk, voice similarity, and speech frequency, to enhance playback experience and compliance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If audio data is recorded as a single monophonic stream, then the recording system is simple to implement, but the playback experience is degraded and spatial information is lost

Engineering Contradiction:
Improverecording system simplicityVSAvoidplayback experience quality
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent segments the single monophonic audio stream into multiple separate audio channels, each corresponding to a different conference participant or endpoint. This segmentation preserves spatial information and allows for independent processing of each participant's audio, thereby improving playback experience while maintaining implementation feasibility through systematic channel separation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a one-dimensional monophonic recording to a multi-dimensional spatial audio representation by assigning virtual positions to different participants in a virtual acoustic space. This dimensional expansion restores spatial relationships and conversational dynamics that were lost in the single-channel recording.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If spatial optimization is applied to assign virtual positions, then the playback experience is improved, but the processing complexity increases

Engineering Contradiction:
Improveplayback experience qualityVSAvoidprocessing system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service through automatic speaker identification and classification systems that autonomously analyze audio characteristics, determine participant roles, and assign virtual positions without requiring manual configuration. The system automatically detects conversational dynamics and adjusts spatial arrangements in real-time, reducing the need for complex manual processing while maintaining high playback quality.

Inventive Principle:
Principle #25Self-service

3Reliability

If conversational dynamics analysis is performed, then regulatory compliance is enhanced, but the processing time increases

Engineering Contradiction:
Improveregulatory complianceVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing conversational dynamics analysis and generating transcripts during or immediately after the conference recording, rather than delaying processing until playback is required. This advance processing ensures regulatory compliance is met while allowing for efficient playback reproduction of already-analyzed content.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent maintains continuity of useful action by implementing real-time or near-real-time processing of conversational dynamics during the conference itself, allowing for continuous analysis without interrupting the flow of the meeting. This continuous processing approach minimizes post-processing time while ensuring compliance requirements are continuously met.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS11076052B2Selective conference digest
Publication Date: 2021.07.27 DOLBY LABORATORIES LICENSING CORP
  • US11076052B2 patent drawing
  • US11076052B2 patent drawing
  • US11076052B2 patent drawing

AI summary

Various disclosed implementations involve processing and/or playback of a recording of a conference involving a plurality of conference participants. Some implementations disclosed herein involve receiving audio data corresponding to a recording of at least one conference involving a plurality of conference participants. In some examples, only a portion of the received audio data will be selected as playback audio data. The selection process may involve a topic selection process, a talkspurt filtering process and/or an acoustic feature selection process. Some examples involve receiving an indication of a target playback time duration. Selecting the portion of audio data may involve making a time duration of the playback audio data within a threshold time difference of the target playback time duration.