Teleconference Audio Synchronization and Mixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current teleconference systems face challenges in recording high-quality audio due to limitations in capturing all participants' voices simultaneously, leading to noise, static, and unintelligible sections, especially when multiple users speak at once.

Innovation Solution

A computing system that receives local audio files from multiple devices, synchronizes and combines them based on timing information to generate a clear, overlapping-free playback, allowing users to selectively manage audio contributions and enable voice-to-text translation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a single user device records the teleconference from its own perspective, then the recording process is simple, but the audio quality deteriorates due to noise, static, and limited capture of other participants' voices

Engineering Contradiction:
Improverecording process simplicityVSAvoidaudio quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent divides the teleconference audio recording into multiple separate local audio files, each captured by a different participant's device. Each device records only its local participant's voice clearly, avoiding the noise and quality issues of single-device remote recording. These segmented recordings are then combined to create a complete high-quality teleconference record.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges multiple local audio files from different teleconference participants into a single combined audio recording. By synchronizing and combining these separate recordings, the system achieves comprehensive audio coverage of all participants with high quality, eliminating the limitations of single-device recording.

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If multiple users speak simultaneously during a teleconference, then the conversation is more dynamic and collaborative, but the recorded audio becomes unintelligible due to speech overlap

Engineering Contradiction:
Improveconversation dynamicsVSAvoidaudio intelligibility
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent segments the audio recording by participant, creating separate local audio files for each teleconference participant. This segmentation allows the system to clearly identify which participant is speaking at any given time, resolving speech overlap issues while preserving the dynamic nature of multi-user conversation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each participant's local audio device records with high quality for that specific local participant, ensuring clear capture of their voice. This local quality approach means each participant's speech is captured optimally by their own device, and when combined, the overall recording maintains high intelligibility even during simultaneous speech periods.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If local audio devices record only their own user's voice, then the audio clarity for each participant is improved, but the system complexity increases due to need for synchronization and combination of multiple files

Engineering Contradiction:
Improveaudio clarityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

Each participant's local audio device automatically records and captures their own voice with high clarity during the teleconference. This self-service approach allows each device to independently perform the recording function for its user without requiring complex coordination during the actual call, simplifying the user experience while maintaining audio quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary recording actions at each local device during the teleconference, capturing audio with timing information embedded. This preliminary action allows the actual audio capture to happen independently and simultaneously at each device, with the complexity of synchronization and combination deferred to a later processing stage, thereby maintaining audio clarity without compromising real-time call performance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10582063B2Teleconference recording management system
Publication Date: 2020.03.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10582063B2 patent drawing
  • US10582063B2 patent drawing
  • US10582063B2 patent drawing

AI summary

An example operation may include one or more of receiving a plurality of local audio files from a plurality of audio devices that participated in a teleconference, where each local audio file includes a locally captured audio recording of a user of a respective audio device during the teleconference, generating combined audio playback information for the teleconference based on the plurality of local audio files received from the plurality of audio devices, the generating including detecting audio portions within the plurality of local audio files and synchronizing a playing order of the detected audio portions based on timing information included in the plurality of local audio files, and transmitting the combined audio playback information of the teleconference to at least one audio device among the plurality of audio devices.