Videoconference Audio Enrichment via Non-Audio Sensor Activity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing videoconferencing systems face challenges in digital accessibility, as users may struggle to follow meetings due to language barriers, skill levels, and digital resource limitations, particularly when video streams include slides or multiple speakers, leading to difficulties in understanding conversations.

Innovation Solution

A method that generates second audio data representative of detected activities using non-audio sensors, which is mixed with captured first audio data to create an enriched audio experience, allowing users to follow meetings more effectively through audio alone, with features like synthetic voice messages and adaptive user parameter settings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If video streams are used to display presentation slides and multiple speakers, then visual information is provided, but users with visual impairments or digital resource limitations cannot fully understand the conversation and activities

Engineering Contradiction:
Improveinformation accessibilityVSAvoidmultimodal delivery requirement
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces an audio intermediary system that translates visual activities (slides, speaker movements, presentations) into descriptive audio signals. This intermediary layer converts information from video streams and sensor data into an accessible audio format, allowing users with visual impairments to understand meeting activities without requiring complex multimodal processing capabilities

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/visual dependency of traditional videoconferencing with an acoustic information delivery system. Instead of relying on users to visually process video streams and sensor inputs, the system substitutes this with an automated audio description channel that conveyS the same information through speech, making the system accessible to users with visual limitations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If audio descriptions of activities are added to the audio stream, then accessibility is improved, but audio data processing complexity increases

Engineering Contradiction:
ImproveaccessibilityVSAvoidaudio data processing
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary processing of sensor data and activity detection before audio generation. By pre-processing sensor inputs, detecting activities, and preparing descriptive audio content in advance of the main audio stream, the system reduces real-time processing complexity while maintaining comprehensive accessibility features

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the audio description generation into distinct functional modules: sensor data acquisition, activity detection, audio synthesis, and stream mixing. This segmentation allows each component to be optimized independently and processed in parallel, reducing overall system complexity while delivering comprehensive audio descriptions

Inventive Principle:
Principle #1Segmentation

3Loss of information

If synchronous mixing of captured audio and generated audio descriptions is performed, then user understanding is improved, but processing time increases

Engineering Contradiction:
Improveactivity information completenessVSAvoidaudio processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary detection and preparation of activity descriptions before they need to be mixed with the audio stream. By pre-detecting activities from sensor data and preparing the corresponding audio descriptions in advance, the system minimizes real-time processing requirements during synchronous mixing, reducing overall processing time while maintaining information completeness

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4287602A1Method for providing audio data, associated device, system and computer program
Publication Date: 2023.12.06 ORANGE SA
  • EP4287602A1 patent drawingFigure 1
  • EP4287602A1 patent drawingFigure 2
  • EP4287602A1 patent drawingFigure 3

AI summary

The present invention relates to a method for providing videoconferencing audio data, an associated device (APP), system (SYS), and computer program. The proposed method comprises a generation (S40) of second audio data (SYN_AUDIO) representative of at least one activity (ACT) detected based on data measured (IN_DATA) by at least one non-audio sensor (SENS), said second generated audio data (SYN_AUDIO) being capable of being mixed with first captured audio data (IN_AUDIO).