Virtual Director for Automated Camera Selection Based on Actor Gaze

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video production systems that rely on automatic camera switching based on audio activity fail to capture the relevance of silent reactions and interactions, leading to a lack of engagement and dynamic content for viewers.

Innovation Solution

A method and system that automatically select video output data by determining interaction states based on the looking direction of actors, configuring sequences of layouts for each state, and processing video data to detect current and transitioning states, ensuring relevant and dynamic content is displayed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automatic camera switching is based on audio activity only, then the system can automatically select camera shots, but the silent reactions and interactions of actors are not captured

Engineering Contradiction:
Improveautomatic camera switchingVSAvoidsilent reactions and interactions
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The patent introduces an intermediary mechanism (eye gaze detection and interaction state analysis system) that bridges the gap between audio-based automation and visual content understanding. This intermediary layer processes video data to detect looking directions and determines interaction states, enabling the system to capture silent reactions and interactions that audio-only systems miss.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/audio-based switching trigger with a vision-based detection system. Instead of relying on audio signals to trigger camera switches, the system uses computer vision to analyze eye gaze, head pose, and body orientation, substituting the audio detection mechanism with a more sophisticated visual analysis mechanism that can detect silent interactions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If multiple cameras are used to capture different perspectives, then the viewer experience is enhanced, but the complexity of controlling and switching between cameras increases

Engineering Contradiction:
Improvecamera perspectivesVSAvoidcamera control and switching
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a self-service control system where the multi-camera system manages itself through automated interaction state analysis. The system independently processes video data from all cameras, detects actor interaction states, and determines optimal camera selections without requiring external operator intervention. This self-service mechanism reduces control complexity while maintaining versatile multi-camera functionality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the control parameters from manual operator decisions to automated analysis of interaction states, looking directions, and scene dynamics. By transforming the control basis from subjective operator judgment to objective measurable parameters (gaze angles, head pose, body orientation), the system manages multiple camera perspectives more efficiently with reduced complexity.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If a single camera is used with cropping, then the system is simpler, but it cannot actively zoom to multiple people at the same time

Engineering Contradiction:
Improvecamera systemVSAvoidzoom capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent makes the multi-camera system universally applicable to various interaction scenarios by implementing interaction state-based selection. The same system architecture can handle single-person close-ups, multi-person interactions, and dynamic scene changes, providing multi-functionality that surpasses the limited zoom capability of single-camera systems while maintaining reasonable complexity through automated control.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250184597A1Three dimensional virtual director
Publication Date: 2025.06.05 BARCO NV
  • US20250184597A1 patent drawing
  • US20250184597A1 patent drawing
  • US20250184597A1 patent drawing

AI summary

A method for automatically selecting video output data, the video output data including at least a sequence of layouts, the sequence of layouts include a plurality of camera shots provided by a plurality of cameras in different viewpoints, the scene includes at least one actor and an object of interest. The method includes the steps of configuring a plurality of layouts, configuring a plurality of interaction states by determining a sequence of layouts including at least one camera shot for each state. A state depends at least on the looking direction of the at least one actor, capturing the scene with the plurality of cameras to generate video data of the scene, processing the acquired video data of the scene to detect a current interaction state, selecting the video output data to show the sequence of layouts corresponding to the current interaction state.