Speech Fragment Detection in Multipoint Videoconferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current videoconferencing technologies face challenges in maintaining fluid conversations due to latency, sound location issues, and poor fidelity, particularly in multipoint sessions where visual and aural cues from participants may be missed or misinterpreted, leading to confusion and frustration.
Innovation Solution
A conferencing system that detects speech fragments by filtering audio into subbands, determining the energy levels, and generating visual or audio indicia to indicate interruptions or interjections, allowing for separate handling of speech fragments to improve communication flow.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If full-duplex audio is used to allow simultaneous speaking, then participants can be aware of interruptions, but the audio does not effectively indicate which participant spoke or vocalized
Solution Approach 1:
The patent introduces an intermediary indicator (visual or auditory signal) that mediates between the audio input and the participants. This indicator explicitly identifies which participant is speaking or interrupting, resolving the ambiguity of full-duplex audio while maintaining interruption awareness.
Solution Approach 2:
The patent adds a new dimension of information display by introducing visual indicators (such as icons, colors, or highlights) alongside the audio channel. This dimensional addition provides speaker identification without interfering with the full-duplex audio capability.
2Quantity of substance
If composite displays of multiple participants are used in multipoint sessions, then all participants can be seen, but participants may not be able to easily tell which participant is doing what
Solution Approach 1:
The patent applies local quality by making specific regions or elements of the composite display have different visual properties. When a participant speaks or interrupts, their video feed or avatar is highlighted with distinct visual characteristics (such as border coloring, size adjustment, or icon placement), enabling easy identification of active participants within the composite display.
3Measurement precision
If switched multipoint video session is used to show different locations, then participants can be seen individually, but the switching between views takes considerable time and adds confusion
Solution Approach 1:
The patent merges the advantages of individual views and group views by maintaining a composite display that shows all participants simultaneously. By combining multiple video feeds into a single composite view with intelligent highlighting, the system eliminates the need for time-consuming switches while preserving individual participant visibility through selective emphasis.
Data Source
AI summary
A conferencing system and method involves conducting a conference between endpoints. The conference can be a videoconference in which audio data and video data are exchanged or can be an audio-only conference. Audio of the conference is obtained from one of the endpoints, and speech is detected in the obtained audio. The detected speech is analyzed to determine that the detected speech constitutes a speech fragment, and an indicia indicative of the determined speech fragment is generated. For a videoconference, the indicia can be a visual cue to be added to video for the given endpoint when displayed at other endpoints. For an audio-only conference, the indicia can be an audio cue to be added to the audio of the conference at the other end points.


