Audio Signal Synchronization via Watermark Cross-Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Ad-hoc meetings often lack the infrastructure for recording and transcribing conversations, as they typically do not involve pre-set conferencing tools, and existing solutions are not designed for spontaneous gatherings outside formal conference settings.

Innovation Solution

A computer-implemented method that utilizes distributed devices to capture and process audio signals in real-time, synchronizing and attributing speech to generate a transcript, even in unplanned meetings, by designating a reference channel and compensating for time differences across multiple audio channels, and employing audio watermarks and video data for participant identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If distributed devices are used to capture audio in ad-hoc meetings, then the need for pre-set recording devices is reduced, but time synchronization accuracy deteriorates due to varying latency and processing differences across devices

Engineering Contradiction:
Improveadaptability to ad-hoc meeting settingsVSAvoidtime synchronization accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary synchronization mechanism that uses audio watermarks and cross-correlation analysis as a mediator between distributed devices with different clock timings. The watermark signal embedded in audio streams acts as a reference that allows the system to measure and compensate for time differences, enabling synchronized processing despite the varying latency inherent in distributed device architectures.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional mechanical synchronization methods (such as NTP or PTP protocols) with an acoustic-based synchronization approach. By embedding audible or inaudible watermarks in the audio signal itself and using cross-correlation to detect them, the system substitutes network-time-based synchronization with audio-signal-based synchronization, which is more suitable for ad-hoc meeting scenarios where devices join spontaneously.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If audio watermarks and video data are used for participant identification, then speaker attribution accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improvespeaker attribution accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the identification process into distinct stages: first embedding watermarks in audio streams, then detecting and correlating these watermarks to identify speakers, and finally using video data as supplementary information. This segmentation allows the system to handle complex multi-modal data processing in manageable steps, reducing overall processing complexity while maintaining high attribution accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by embedding watermarks in audio streams before the actual speech recognition and attribution processes. This pre-processing step prepares the data in advance, so that when speaker identification is needed, the system can quickly match detected watermarks against known participant profiles without having to analyze raw audio signals from scratch, thereby reducing real-time processing complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10743107B1Synchronization of audio signals from distributed devices
Publication Date: 2020.08.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10743107B1 patent drawing
  • US10743107B1 patent drawing
  • US10743107B1 patent drawing

AI summary

A computer implemented method includes receiving audio signals representative of speech via multiple audio channels transmitted from corresponding multiple distributed devices, designating one of the audio channels as a reference channel, and for each of the remaining audio channels, determine a difference in time from the reference channel, and correcting each remaining audio channel by compensating for the corresponding difference in time from the reference channel.