Dynamic Video Orchestration via Probabilistic Transition Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems lack dynamic and immersive video orchestration capabilities, relying on static templates and limited audio event detection, which results in a suboptimal user experience by missing around 70% of useful information.
Innovation Solution
A method and device for generating an output video stream in a video conference using orchestration models with predefined screen templates, transition probabilities, and observation probabilities, dynamically selecting the most suitable display states based on participant actions, such as gestures and audio events, to create a more immersive experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If static templates and limited audio event detection are used for video orchestration, then device complexity is reduced, but loss of information increases (around 70% of useful information is missed)
Solution Approach 1:
The patent transforms static templates into dynamic orchestration models that adapt in real-time. The system uses multiple display states with transition probabilities that dynamically change based on observed participant actions, enabling the video orchestration to respond adaptively to conference dynamics rather than following fixed predetermined patterns.
Solution Approach 2:
The patent changes the parameters of the orchestration system by introducing probability distributions (transition probabilities and observation probabilities) that quantify the likelihood of different display states and observable actions. This probabilistic parameterization enables more nuanced decision-making compared to binary rule-based systems.
2Adaptability or versatility
If fixed number of templates are available for video orchestration, then ease of operation is improved, but adaptability or versatility deteriorates (cannot modify or enhance by single user)
Solution Approach 1:
The patent implements self-service through automated learning where the system observes participant actions and automatically updates the orchestration models without requiring manual reconfiguration. The probabilistic models learn from observed behaviors and adapt to new conference scenarios automatically, eliminating the need for expert intervention while maintaining versatility.
Solution Approach 2:
The patent prepares multiple predefined display states and templates in advance, but unlike fixed systems, these are designed to be selectively combined through probability-based transitions during runtime. This preliminary preparation of modular components enables both ease of operation and adaptability.
3Loss of information
If audio event detection is used for video switching behavior, then device complexity is reduced, but loss of information increases (70 percent of useful information is missing)
Solution Approach 1:
The patent segments participant actions into distinct observable categories (gestures, head motions, face expressions, audio actions, enunciation of keywords, presentation slide actions). This segmentation enables comprehensive capture of different action types while maintaining organized, manageable data structures for processing.
Solution Approach 2:
The patent introduces probabilistic models as intermediaries between raw observation events and video orchestration decisions. The observation probabilities and transition probabilities act as mediators that interpret observed actions and translate them into appropriate display state transitions, enabling sophisticated decision-making without requiring complex direct control logic.
4Ease of manufacture
If expert-created rules templates are used, then manufacturing precision is improved, but ease of manufacture deteriorates (cannot be modified or enhanced by single user)
Solution Approach 1:
The patent enables non-expert users to create and refine orchestration models through automated learning from observed participant actions. The system self-adjusts the probabilistic parameters based on actual conference dynamics, eliminating the need for expert manual tuning while maintaining or improving accuracy through data-driven adaptation.
Solution Approach 2:
The patent implements feedback loops where the system continuously observes participant actions, compares expected versus actual behaviors using the probabilistic models, and updates the orchestration parameters accordingly. This feedback mechanism enables continuous refinement of model accuracy without requiring expert intervention.
Data Source
Figure 1~3
Figure 4~5
Figure 6~7
AI summary
A method for generating an orchestration model of video streams in α video conference comprising α plurality of input video streams (11) and α series of input observation events (33), the method comprising: Providing α user input interface, Displaying the input video streams (11) arranged in accordance with the predefined screen templates, Displaying the current observation events, Recording α sequence of current display states at successive instants in time, in accordance with the current screen templates selected by the user, Determining numbers of transition occurrences that occurred each between two successive display states, Determining the transition probabilities between all the display states from the numbers of transition occurrences, Determining numbers of observation events that occurred for each of the observable actions, Determining the observation probabilities as a function of the numbers of observation events, Storing the orchestration model.