Automated Cinematic Decisions Using 2D Pose Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing systems lack advanced automated cinematic decision-making capabilities, failing to dynamically adjust camera and microphone settings based on real-time environmental and user interactions, leading to suboptimal visual and audio focus during audio-visual communication sessions.

Innovation Solution

An intelligent communication device with a touch-sensitive display, cameras, and microphones that uses a descriptive model to make automated cinematic decisions, including zooming, panning, and audio targeting, by processing 2D Pose data and accessing user privacy settings to enhance the viewing experience without compromising user privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If automated cinematic decision-making is implemented to dynamically adjust camera and microphone settings, then the quality of audio-visual communication is improved, but the device complexity increases

Engineering Contradiction:
Improvequality of audio-visual communicationVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The intelligent communication device performs automated cinematic decisions independently, analyzing 2D Pose data and environmental information without requiring manual control. The device self-adjusts camera zoom, pan, and microphone beamforming based on real-time participant positions and actions, enabling autonomous optimization of audio-visual quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical adjustment of camera and microphone settings is replaced by an intelligent system that uses processed data from 2D Pose estimators and environmental sensors. The system substitutes human operator decisions with automated algorithms that analyze participant positions, actions, and engagement levels to determine optimal cinematic parameters

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If 2D Pose data and environmental information are processed to make cinematic decisions, then the adaptability of the system is improved, but the loss of information due to privacy concerns increases

Engineering Contradiction:
Improveadaptability to different communication scenariosVSAvoiduser privacy
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system extracts only the necessary non-identifying information from 2D Pose data and environmental sensors required for cinematic decisions. It processes participant positions, movements, and actions without capturing or storing personally identifiable information, thereby maintaining adaptability while protecting user privacy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces privacy settings as an intermediary layer between data collection and processing. Users can control what information is accessed through configurable privacy settings, allowing the system to adapt to different communication scenarios while respecting user privacy preferences and preventing unauthorized information loss

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10979669B2Automated cinematic decisions based on descriptive models
Publication Date: 2021.04.13 DISCORD INC
  • US10979669B2 patent drawing
  • US10979669B2 patent drawing
  • US10979669B2 patent drawing

AI summary

In one embodiment, a method includes accessing foreground visual data that comprises a set of coordinate points that correspond to a plurality of surface points of a person in an environment; generating a bounding box for the set of coordinate points, wherein the bounding box comprises every coordinate point in the set of coordinate points; providing instructions to collect background visual data for an area in the environment that is outside of the bounding box; and providing the foreground visual data and the background visual data to an intelligent director associated with the computing device.