Dynamic Video View State Control for Communication Sessions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing communication systems often fail to optimize user engagement and resource efficiency by equally displaying video streams of single and multi-person groups, leading to missed social cues and inefficient use of computing resources.

Innovation Solution

Implementing dynamically controlled view states that adjust the size and position of video streams based on the number of individuals depicted, reserving primary areas for multi-person streams and secondary areas for single-person streams, using facial recognition and other technologies to ensure equal representation and improve user interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If video streams are displayed in equal size arrangement, then each stream is given equal visual weight, but multi-person groups do not show sufficient detail for each person

Engineering Contradiction:
Improvedetail representationVSAvoidvisual equality
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent applies local quality by differentiating the display treatment based on the content of each video stream. Single-person streams are displayed at a first size while multi-person streams are displayed at a second size (larger than the first), allowing each type of stream to receive appropriate visual emphasis. This resolves the contradiction by providing detailed representation for multi-person groups through larger display size while maintaining visual equality through consistent application of the sizing rule across all streams.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamic display arrangement that automatically adjusts video stream sizes based on real-time detection of the number of people in each stream. The system continuously monitors video content and dynamically reconfigures the display layout, transitioning between different display states as participants join or leave video streams. This dynamic adaptation ensures optimal detail representation while maintaining ease of operation through automated adjustment.

Inventive Principle:
Principle #15Dynamics

2Productivity

If manual interaction is required to send text messages or emails when social cues are missed, then communication can continue, but workflow is disrupted and productivity is reduced

Engineering Contradiction:
Improveworkflow efficiencyVSAvoidmissed social cues
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent implements feedback mechanisms that automatically detect and respond to social cues in video streams. The system analyzes video content to identify gestures, expressions, and other social signals, then provides immediate feedback through notifications or automatic actions (such as sending text messages or emails). This eliminates the need for manual monitoring and intervention, thereby maintaining productivity while preventing information loss from missed social cues.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-service by automatically detecting social cues and initiating appropriate communication actions without requiring user intervention. The automated detection and response system handles the entire process from cue identification to message delivery, freeing users from manual workflow disruptions while ensuring no social cues are missed.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If video streams are displayed with equal sizing, then display arrangement is simple, but user engagement is reduced due to inability to clearly see important gestures

Engineering Contradiction:
Improvedisplay arrangement simplicityVSAvoiduser engagement
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent applies local quality by assigning different display sizes to different types of video streams based on their content characteristics. Single-person streams use a first size while multi-person streams use a second size (larger than the first), ensuring that streams containing important gestures and social cues are displayed with sufficient detail to maintain user engagement, while keeping the display arrangement rule-based and relatively simple.

Inventive Principle:
Principle #3Local quality

4Loss of information

If follow-up meetings are scheduled to address missed content, then communication completeness is improved, but computing resource usage increases

Engineering Contradiction:
Improvecommunication completenessVSAvoidcomputing resource usage
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The patent implements preliminary action by automatically detecting and responding to social cues in real-time during communication sessions. By proactively identifying missed gestures or social signals and immediately notifying participants or sending follow-up messages, the system prevents information gaps from developing in the first place. This eliminates the need for subsequent follow-up meetings to address missed content, thereby reducing computing resource usage while maintaining communication completeness.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4082164B1Method and system for providing dynamically controlled view states for improved engagement during communication sessions
Publication Date: 2024.09.25 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4082164B1 patent drawingFigure 1A
  • EP4082164B1 patent drawingFigure 1B
  • EP4082164B1 patent drawingFigure 1C

AI summary

The techniques disclosed herein improve user engagement and more efficient use of computing resources by providing dynamically controlled view states for communication sessions based on a number of people depicted in shared video streams. In some configurations, a system can control the size and position of a video rendering based on the number of individuals depicted in a video stream. In some configurations, a stream depicting a threshold number of people can be rendered in the primary display area and other streams can be rendered in a secondary section. The primary area can be sized to scale a video depicting multiple people video to equalize the size of the people with renderings of single-person video streams. This helps a system provide a more granular level of control to equalize the representation of each person displayed within different video streams.