Video Stream Selection for Non-Verbal Activity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video conferencing systems often require users to monitor multiple video feeds to detect non-verbal activities, which can be distracting and inefficient, especially in large groups or on devices with limited display space, and may unnecessarily allocate network bandwidth.

Innovation Solution

A communication system that implements a 'Storied Experience View' by selecting and displaying video streams based on detected non-verbal events, using a central server to manage video streams and allocate bandwidth efficiently, while ensuring that active speakers' audio remains uninterrupted.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If all video streams are displayed to show complete non-verbal activities of all participants, then user awareness of group dynamics is improved, but device display space is overwhelmed and user attention is lost

Engineering Contradiction:
Improvenon-verbal activity informationVSAvoiddisplay space
Core Design Contradiction:
Loss of informationVSArea of stationary object

Solution Approach 1:

The system extracts and displays only the most relevant video streams containing significant non-verbal activities (such as reactions, gestures, or expressions) while filtering out less important feeds. This selective extraction allows the system to maintain awareness of group dynamics without overwhelming the display space with all participant feeds simultaneously.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Different regions or time periods of the video conference are treated with different display priorities. The system dynamically adjusts which video streams receive display priority based on the significance of non-verbal activities detected in local time windows, ensuring that important moments are captured without continuously occupying all display space.

Inventive Principle:
Principle #3Local quality

2Loss of information

If multiple video feeds are continuously monitored to detect all non-verbal activities, then detection completeness is improved, but network bandwidth is wasted

Engineering Contradiction:
Improvenon-verbal event detectionVSAvoidnetwork bandwidth
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The system performs preliminary analysis of video streams to identify potential non-verbal activities before full display or transmission. By pre-detecting significant events using automated video analysis, the system can selectively transmit or display only those streams containing meaningful non-verbal cues, avoiding continuous monitoring and transmission of all feeds.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The video conference system automatically analyzes its own video streams to detect non-verbal activities and makes autonomous decisions about which feeds to display or transmit. This self-service capability eliminates the need for manual monitoring of all feeds while optimizing bandwidth usage based on detected event significance.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If automated detection of non-verbal movements is implemented, then user engagement is improved, but system complexity increases

Engineering Contradiction:
Improveuser engagementVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system replaces manual user monitoring and selection of video feeds with automated computer vision algorithms that detect non-verbal movements and reactions. This substitution of mechanical/manual operations with automated image processing and pattern recognition enhances user engagement through intelligent, dynamic feed selection while managing system complexity through specialized algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP3387826B1Communication event
Publication Date: 2020.09.02 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3387826B1 patent drawingFigure 1
  • EP3387826B1 patent drawingFigure 2
  • EP3387826B1 patent drawingFigure 3

AI summary

In a communication event between a first user and one or more second users via a communication network, a plurality of video streams is received (S502) via the network. Each of the streams carries a moving image of at least one respective user. The moving image of a first of the video streams is displayed (S504) at a user device of the first user for a first time interval. In the moving image of a second of the video streams that is not displayed at the user device in the first time interval, a human feature of the respective user is identified (S506, S508). A movement of the identified human feature during the first time interval that matches one of a plurality of expected movements is detected (S510, S512). In response to the detected movement, at least the moving image of the second video stream is displayed (S516, S518, S520) at the user device for a second time interval.