Audio-Video Speech Enhancement for Distant Microphones

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Capturing high-quality audio in noisy environments is challenging due to background noise, signal-to-noise ratio, reverberation, echoes, and microphone placement, which affects audio clarity and intelligibility, especially when microphones are positioned away from the speaker.

Innovation Solution

Combining facial structure and movement data with large language models to enhance audio quality by processing video and audio data together, using transformer models to filter out noise, fill in missing audio, and correct muffled speech, leveraging multiple cameras and microphones for improved speech clarity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If microphones are positioned away from the speaker, then the system has more flexibility in placement and less intrusion, but audio clarity and signal-to-noise ratio deteriorate

Engineering Contradiction:
Improvemicrophone placement flexibilityVSAvoidaudio clarity
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent uses visual information from video data as an intermediary to compensate for poor audio quality. Facial structure data and mouth movement information serve as mediators that help reconstruct and enhance the original speech signal, allowing the system to maintain microphone placement flexibility while improving audio clarity through cross-modal information integration

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent combines audio data with video data to create enhanced audio output. By merging information from multiple sources (microphone signals, facial structure, mouth movements), the system overcomes the limitations of distant microphone placement and achieves better audio clarity than either modality could provide alone

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If multiple cameras and processing models are used to enhance audio quality, then audio clarity improves, but device complexity increases

Engineering Contradiction:
Improveaudio clarityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the video capture system multi-functional by using it for both visual recording and audio enhancement. The same video data that captures the speaker's image is also used to extract facial structure and mouth movement information for speech enhancement, eliminating the need for separate specialized sensors and reducing overall system complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses its own video capture capability to serve the audio enhancement function. Rather than requiring external specialized hardware, the system leverages its existing video data and processing capabilities to improve audio quality, making the system self-sufficient and reducing complexity

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250336400A1Managing speech using models
Publication Date: 2025.10.30 GOOGLE LLC
  • US20250336400A1 patent drawing
  • US20250336400A1 patent drawing
  • US20250336400A1 patent drawing

AI summary

According to at least one implementation, a method includes obtaining audio data associated with a user, and obtaining video data corresponding to the audio data, the video data from a set of cameras. The method further includes determining features associated with a portion of the user based on the video data and applying a model to the audio data and the features to generate updated audio data, the model configured from second audio data associated with second video data.