Real-Time Director's Cut Using Role-Centric Video Ranking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video systems experience latency and inaccuracies in detecting and switching between speakers during live events, leading to delayed and contextually incorrect camera focus, which affects the quality of the viewing experience.

Innovation Solution

A system that uses multiple cameras and microphones, combined with AI components like role-centric, gaze-centric, and emotion-centric evaluators, to analyze audio and video feeds in real-time, predicting speaker changes and prioritizing camera focus based on participant roles, engagement, and emotional cues, enabling faster and more accurate switching between speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional audio or video detection methods are used to detect speakers, then the system can identify who is speaking, but the detection latency is high (2-5 seconds or more) causing delayed camera switching

Engineering Contradiction:
Improvespeaker detection accuracyVSAvoidcamera switching latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by detecting speaker intent, gaze direction, and speech patterns before the actual speaker change occurs. The gaze-centric detector predicts who will speak next based on eye contact and attention direction, while the speech-centric detector analyzes speech patterns to anticipate speaker transitions, enabling the camera to switch before the traditional 2-5 second latency would occur.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary detection mechanisms (gaze detection, speech pattern analysis, intent detection) that mediate between the actual speaker change and the camera switching action. These intermediaries provide early warning signals about upcoming speaker transitions, allowing the system to prepare and execute camera switches with minimal latency while maintaining accurate speaker identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If the system focuses on the current speaker based on traditional detection, then speaker identification is straightforward, but the context may be incorrect (e.g., showing a speaker who is speaking out of turn)

Engineering Contradiction:
Improvecamera control simplicityVSAvoidcontextual accuracy of speaker focus
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system employs feedback mechanisms where the role-centric evaluator continuously monitors whether the detected speaker is appropriately speaking based on their assigned role. The gaze-centric and speech-centric detectors provide feedback about attention direction and speech patterns, allowing the system to verify contextual appropriateness before directing camera focus, thus ensuring reliability while maintaining operational simplicity through automated evaluation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the parameters used for speaker detection from simple audio/voice activation to multiple dimensions including gaze direction, speech patterns, intent detection, and role appropriateness. By monitoring these additional parameters, the system can determine not just who is speaking but whether they should be the focus, improving contextual accuracy while keeping camera control simple through integrated evaluation.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If manual camera switching is used during live events, then the director can make contextual decisions, but the response time is slow and cannot keep up with rapid speaker changes

Engineering Contradiction:
Improvecontextual decision-making capabilityVSAvoidcamera switching speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system performs self-service by automatically detecting speaker intent, analyzing gaze patterns, evaluating speech contexts, and determining the appropriate next speaker without human intervention. The role-centric, gaze-centric, and speech-centric detectors work together to autonomously make contextual decisions about camera focus, enabling rapid response to speaker changes while maintaining the adaptability of contextual judgment through intelligent algorithms.

Inventive Principle:
Principle #25Self-service

4Area of stationary object

If the system uses wide-shots and slow-pans for camera transitions, then all participants remain visible, but the transitions are slow and reduce engagement

Engineering Contradiction:
Improvevisible area of participantsVSAvoidcamera transition speed
Core Design Contradiction:
Area of stationary objectVSSpeed

Solution Approach 1:

The system performs preliminary detection of speaker intent and gaze direction before the actual camera transition is needed. By anticipating the next speaker through gaze analysis and speech pattern detection, the system can prepare the camera in advance and execute faster transitions without losing contextual information, thereby increasing transition speed while maintaining participant visibility through strategic framing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11924580B2Generating real-time director's cuts of live-streamed events using roles
Publication Date: 2024.03.05 INTEL CORP
  • US11924580B2 patent drawing
  • US11924580B2 patent drawing
  • US11924580B2 patent drawing

AI summary

An example apparatus for generating real-time director's cuts includes a number of cameras to capture videos of a plurality of participants in a scene. The apparatus also includes a number of microphones to capture audio corresponding to each of the number of participants. The apparatus further includes a role-centric evaluator to receive views-of-participants and a role for each of the participants and rank the views-of-participants based on the roles. Each of the views-of-participants are tagged with one of the participants. The apparatus further includes a view broadcaster to display a highest ranking view-of-participant stream.