Autonomous Camera Mimicking Human Operator Behavior

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous camera systems struggle to anticipate and frame dynamic activities like sporting events with 'lead room,' resulting in videos that appear robotic due to their inability to mimic human camera operators.

Innovation Solution

A computer-implemented method that trains a regressor using demonstration data from human camera operators and environmental sensory data to predict pan-tilt-zoom settings, allowing autonomous cameras to mimic human camera behavior by extracting feature vectors and learning to output planned device settings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If hand-coded equations are used to determine camera positioning, then the system can automatically track objects, but the output videos appear robotic and lack anticipatory framing

Engineering Contradiction:
Improveautomatic camera controlVSAvoidaesthetic quality of video
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent copies human camera operator behavior by training a machine learning model on demonstration data from human operators. The system learns to replicate human decision-making patterns for camera positioning, pan-tilt-zoom operations, and anticipatory framing, thereby producing videos that resemble human-operated recordings rather than robotic automated tracking

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces hand-coded mechanical control equations with a data-driven machine learning regressor. Instead of using predetermined algorithms to calculate camera positions, the system uses trained neural networks that process sensory data and output camera settings based on learned patterns from human operator demonstrations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Speed

If object tracking is used to determine camera position, then the camera can follow the subject, but the system cannot anticipate action with sufficient lead room

Engineering Contradiction:
Improvecamera response speedVSAvoidanticipatory time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent implements preliminary action by training the regressor to predict future camera positions based on current and past sensory data. The system learns from human operators how to anticipate upcoming actions and position the camera in advance with appropriate lead room, rather than simply reacting to current object positions

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from demonstration data during training, where the regressor learns from the relationship between sensory inputs and human operator camera settings. This feedback mechanism enables the system to internalize human anticipatory behavior patterns and reproduce them in autonomous operation

Inventive Principle:
Principle #23Feedback

3Stability of the object's composition

If smoothing is applied to object tracking data, then camera motion becomes smoother, but the system still cannot frame shots with adequate lead room

Engineering Contradiction:
Improvecamera motion smoothnessVSAvoidanticipatory framing capability
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The system copies the complete human operator workflow including both smoothing operations and anticipatory adjustments. By learning from demonstrated human behavior, the regressor captures not just the smoothing effect but also the timing and magnitude of anticipatory camera movements that create proper lead room

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10812686B2Method and system for mimicking human camera operation
Publication Date: 2020.10.20 DISNEY ENTERPRISES INC
  • US10812686B2 patent drawing
  • US10812686B2 patent drawing
  • US10812686B2 patent drawing

AI summary

The disclosure provides an approach for mimicking human camera operation with an autonomous camera system. In one embodiment, camera planning is formulated as a supervised regression problem in which an automatic broadcasting application receives one video input captured by a human-operated camera and another video input captured by a stationary camera with a wider field of view. The automatic broadcasting application extracts feature vectors and pan-tilt-zoom states from the stationary camera and the human-operated camera, respectively, and learns a regressor which takes as input such feature vectors and outputs pan-tilt-zoom settings predictive of what the human camera operator would choose. The automatic broadcasting application may then apply the learned regressor on newly captured video to obtain planned pan-tilt-zoom settings and control an autonomous camera to achieve the planned settings to record videos which resemble the work of a human operator in similar situations.