Multimodal Audio Processing With Virtual Room Boundaries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conferencing systems struggle to provide high-quality audio signals in open and dynamic meeting spaces, failing to effectively suppress unwanted speech and noise, especially in hybrid meeting setups where physical boundaries are inadequate.

Innovation Solution

Implementing a method that defines virtual room boundaries using video and audio data correlation to filter out unwanted speech and noise, enhancing speech signals by correlating them with participant presence and position within these boundaries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If physical boundaries (closed meeting rooms) are used to suppress unwanted speech and noise, then audio quality is improved, but adaptability to different meeting setups and open spaces deteriorates

Engineering Contradiction:
Improveunwanted speech and noise suppressionVSAvoidadaptability to different meeting setups
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The patent replaces physical acoustic boundaries (mechanical system) with a virtual boundary system based on video-audio correlation and signal processing. The system uses cameras to detect participant positions and microphones to capture audio, then correlates video and audio data to determine which audio signals originate from within the virtual meeting space. This substitution allows the system to achieve noise suppression without requiring physical room boundaries, thereby enabling adaptability to open spaces and different meeting configurations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The virtual room boundaries are dynamic rather than fixed. The system continuously tracks participant positions using video data and adjusts the audio processing accordingly. When participants move within the space, the system dynamically updates which audio signals should be included or excluded. This dynamic approach allows the system to adapt to changing meeting setups, participant arrangements, and space configurations in real-time, resolving the contradiction between noise suppression and adaptability.

Inventive Principle:
Principle #15Dynamics

2Reliability

If fixed acoustic fencing is implemented to limit outside noise, then audio signal quality is improved, but flexibility in space selection and meeting setup deteriorates

Engineering Contradiction:
Improveaudio signal qualityVSAvoidflexibility in space selection
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent replaces fixed acoustic fencing (mechanical system) with a software-based virtual boundary system. Instead of physically enclosing the meeting space to ensure audio quality, the system uses video-audio correlation to dynamically define which audio signals should be captured. This allows the same system to provide high-quality audio in diverse spaces ranging from open offices to conference rooms, eliminating the need for fixed acoustic enclosures and enabling flexible space selection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The virtual boundary system serves multiple functions: it defines meeting space boundaries, tracks participant positions, filters unwanted audio, and adapts to different meeting configurations. This multi-functional approach allows a single system to provide reliable audio quality across various space types and meeting setups, replacing the need for space-specific acoustic treatments or fixed fencing structures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If video-audio correlation with virtual boundaries is used to filter audio signals, then speech intelligibility and signal-to-noise ratio are improved, but system complexity increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges video processing and audio processing into a unified system. The video data from cameras is used to define virtual boundaries and track participant positions, which then directly inform the audio processing to filter and enhance speech signals. This integration allows the system to achieve high speech intelligibility through a coordinated approach rather than separate, complex subsystems. The merging of functions reduces overall system complexity compared to having independent, highly sophisticated audio and video processing systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces video data as an intermediary between the physical environment and audio processing. Instead of directly analyzing audio signals to determine their sources, the system uses video data to identify participant positions and then uses this information to guide audio filtering. This intermediary approach simplifies the audio processing by providing clear spatial context from the video data, making the overall system more manageable despite the added video processing requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12581038B2Audio processing in video conferencing system using multimodal features
Publication Date: 2026.03.17 GN HEARING AS
  • US12581038B2 patent drawing
  • US12581038B2 patent drawing
  • US12581038B2 patent drawing

AI summary

The disclosure relates to a method for processing audio signals in a video conference call. At first, virtual room boundaries defining space relevant for the video conference call are determined. By at least one input transducer, audio data are obtained and by at least one camera, video data are obtained. The video data and audio data are then correlated. On the basis of the correlation and defined virtual room boundaries, the audio data are modified and modified output audio data are generated.