Teleconferencing Audio-Visual Cue System for Participant Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In audio teleconferences with multiple attendees, it is difficult to distinguish between speakers due to voice distortion over phone lines and voice-over-IP implementations, making it challenging to identify the current presenter and focus on their slide presentations.
Innovation Solution
A teleconferencing environment that uses both audio and visual cues to identify active participants and presenters by associating stereo-enhanced audio with visual identifiers on a computer screen, promoting and demoting attendees based on their level of participation, and providing visual cues through animated avatars and slide focus indicators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio teleconferencing is used with multiple attendees, then teleconference capability is provided, but participant identification becomes difficult due to voice distortion
Solution Approach 1:
The system segments the audio signal by participant source and assigns each segment to a specific visual identifier (image or icon) on the display. This segmentation allows participants to be visually distinguished despite audio distortion, as each participant's audio is spatially mapped to their corresponding visual representation on the screen.
Solution Approach 2:
Visual identifiers (images or icons) serve as intermediaries between the audio signals and the participants. Instead of relying solely on distorted audio for identification, the system introduces visual elements that mediate the connection between audio sources and participant identities, making identification easier through visual-cued spatial hearing.
2Difficulty of detecting and measuring
If visual identifiers are added to identify participants, then participant identification improves, but system complexity increases
Solution Approach 1:
The visual identifiers serve multiple functions: they represent participant identities, indicate current speakers, and provide spatial cues for audio localization. By making these visual elements multi-functional, the system avoids adding separate components for each function, thereby reducing overall system complexity while improving participant identification.
Solution Approach 2:
The system merges audio signal processing with visual display functions by integrating spatial hearing processing with the presentation of visual identifiers. This consolidation combines what could be separate systems into a unified approach, reducing complexity while achieving both audio-visual synchronization and participant identification.
3Difficulty of detecting and measuring
If stereo-enhanced audio is applied to all participants, then spatial hearing ability is leveraged for identification, but audio processing complexity increases
Solution Approach 1:
The system applies stereo-enhanced audio processing selectively to the audio signal of the current speaker rather than uniformly to all participants. This localized processing approach enhances speaker identification for the active participant while minimizing unnecessary processing for inactive participants, thereby reducing overall audio processing complexity.
Solution Approach 2:
Instead of applying full stereo enhancement to all audio signals continuously, the system applies enhanced processing only when and where needed (for current speakers). This partial action approach provides sufficient speaker identification capability without the excessive processing burden of universal enhancement, optimizing the balance between identification accuracy and processing complexity.
Data Source
AI summary
A teleconferencing environment is provided in which both audio and visual cues are used to identify active participants and presenters. Embodiments provide an artificial environment, configurable by each participant in a teleconference, that directs the attention of a user to an identifier of an active participant or presenter. This direction is provided, in part, by stereo-enhanced audio that is associated with a position of a visual identifier of an active participant or presenter that has been placed on a window of a computer screen. The direction is also provided, in part, by promotion and demotion of attendees between attendee, active participant, and current presenter and automatic placement of an image related to an attendee on the screen in response to such promotion and demotion.


