Video Conference Participant Sorting With MTCNN Cue Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conferencing systems struggle to effectively sort participants based on non-verbal cues such as emotions and gestures, limiting the ability of presenters to gauge audience reaction and communicate effectively.
Innovation Solution
A system that utilizes a multi-cascaded convolutional neural network (MTCNN) to analyze facial expressions, body language, and hand gestures to determine a participant's emotional state and engagement level, then ranks and sorts participants accordingly within the video conference interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If video conferencing systems use traditional participant sorting methods, then the system complexity remains low, but the ability to gauge audience reaction and communication effectiveness deteriorates
Solution Approach 1:
The patent introduces an intermediary component (analysis system) that processes video feeds and extracts non-verbal cues. This intermediary acts as a bridge between the video conferencing system and participant emotional states, enabling the system to gauge audience reaction without requiring direct complex analysis of all participant data
Solution Approach 2:
The patent replaces traditional mechanical sorting methods (based on simple criteria like speaking status) with an automated analysis system that uses machine learning models to detect non-verbal cues. This substitution enables sophisticated emotional and engagement analysis while automating the sorting process, reducing the need for manual intervention
2Measurement precision
If the system analyzes multiple non-verbal cues for each participant, then the accuracy of participant sorting improves, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary analysis by continuously monitoring and pre-processing video feeds in the background. Non-verbal cues are detected and analyzed before sorting is needed, so that when sorting is required, the data is already prepared and ready for quick retrieval and arrangement
Solution Approach 2:
The patent divides the analysis into separate modules, each detecting specific non-verbal cues (facial expressions, body language, gestures) independently. This segmentation allows parallel processing of different cue types, reducing overall processing time while maintaining comprehensive analysis accuracy
3Adaptability or versatility
If the system sorts participants based on real-time emotional states, then the presenter's ability to adjust presentation improves, but the computational load and energy consumption increase
Solution Approach 1:
The patent implements periodic analysis where the system evaluates participant emotional states at regular intervals rather than continuously analyzing every frame. This periodic approach provides timely feedback for presentation adjustment while significantly reducing computational energy consumption compared to continuous real-time analysis
Data Source
AI summary
An example non-transitory machine-readable storage medium comprising instructions executable by a processing resource of a computing device to cause the computing device to: receive a video feed of a participant in a video conference; identify the participants within the video feed; determine a probability that a characteristic is being experienced by the participant; determine a relevancy score of the participant based on the probability that the characteristic is being experienced by the participant; and display the participant relative to other participants in the video conference based on the relevancy score.


