Videoconference Audio Enrichment via Non-Audio Sensor Activity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing videoconferencing systems face challenges in digital accessibility, as users may struggle to follow meetings due to language barriers, skill levels, and digital resource limitations, particularly when video streams include slides or multiple speakers, leading to difficulties in understanding conversations.
Innovation Solution
A method that generates second audio data representative of detected activities using non-audio sensors, which is mixed with captured first audio data to create an enriched audio experience, allowing users to follow meetings more effectively through audio alone, with features like synthetic voice messages and adaptive user parameter settings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If video streams are used to display presentation slides and multiple speakers, then visual information is provided, but users with visual impairments or digital resource limitations cannot fully understand the conversation and activities
Solution Approach 1:
The patent introduces an audio intermediary system that translates visual activities (slides, speaker movements, presentations) into descriptive audio signals. This intermediary layer converts information from video streams and sensor data into an accessible audio format, allowing users with visual impairments to understand meeting activities without requiring complex multimodal processing capabilities
Solution Approach 2:
The patent replaces the mechanical/visual dependency of traditional videoconferencing with an acoustic information delivery system. Instead of relying on users to visually process video streams and sensor inputs, the system substitutes this with an automated audio description channel that conveyS the same information through speech, making the system accessible to users with visual limitations
2Adaptability or versatility
If audio descriptions of activities are added to the audio stream, then accessibility is improved, but audio data processing complexity increases
Solution Approach 1:
The patent performs preliminary processing of sensor data and activity detection before audio generation. By pre-processing sensor inputs, detecting activities, and preparing descriptive audio content in advance of the main audio stream, the system reduces real-time processing complexity while maintaining comprehensive accessibility features
Solution Approach 2:
The patent segments the audio description generation into distinct functional modules: sensor data acquisition, activity detection, audio synthesis, and stream mixing. This segmentation allows each component to be optimized independently and processed in parallel, reducing overall system complexity while delivering comprehensive audio descriptions
3Loss of information
If synchronous mixing of captured audio and generated audio descriptions is performed, then user understanding is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary detection and preparation of activity descriptions before they need to be mixed with the audio stream. By pre-detecting activities from sensor data and preparing the corresponding audio descriptions in advance, the system minimizes real-time processing requirements during synchronous mixing, reducing overall processing time while maintaining information completeness
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to a method for providing videoconferencing audio data, an associated device (APP), system (SYS), and computer program. The proposed method comprises a generation (S40) of second audio data (SYN_AUDIO) representative of at least one activity (ACT) detected based on data measured (IN_DATA) by at least one non-audio sensor (SENS), said second generated audio data (SYN_AUDIO) being capable of being mixed with first captured audio data (IN_AUDIO).