Key Person Detection in Immersive Video Using Graph Attention Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting key persons in immersive video, such as in sporting events, rely on manual camera operation which is expensive and not scalable, limiting the ability to provide an immersive experience by tracking and focusing on key players in real-time.

Innovation Solution

A system that uses a camera array and graph attention networks to detect key persons by identifying predefined formations and generating feature vectors, which are then used to classify key players through a graph attentional network, allowing for real-time tracking and virtual view generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual camera operation is used to track key players, then the quality of fan engagement and immersive experience is improved, but the cost and scalability deteriorate

Engineering Contradiction:
Improvefan engagement qualityVSAvoidscalability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables automated detection and tracking of key persons through self-service mechanisms. The graph attention network automatically identifies key persons based on formation patterns and visual features without requiring manual camera operation, thereby maintaining engagement quality while eliminating the scalability limitations of manual approaches

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual camera operation with an automated computer vision system. The graph attention network and formation detection algorithms substitute human operators, enabling the system to scale across multiple cameras and scenes while maintaining consistent key person detection and tracking performance

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated detection systems are implemented, then scalability and operational efficiency are improved, but the complexity of the system increases

Engineering Contradiction:
Improveoperational efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex detection task into distinct functional modules: formation detection module that identifies predefined player formations, graph attention network that processes spatial relationships, and key person detection module that identifies specific individuals. This segmentation manages complexity by organizing functions into separate, manageable components that can be processed independently

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate processing stages between raw video input and key person identification. The formation detection acts as an intermediary that filters and structures data before it reaches the graph attention network, which in turn prepares processed features for the final key person detection module, thereby managing system complexity through staged processing

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If real-time tracking of multiple key persons is achieved, then the immersive user experience is improved, but the computational resources and processing time increase

Engineering Contradiction:
Improvetracking accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary detection of formation patterns and spatial relationships before final key person identification. The graph attention network pre-processes visual data to extract meaningful features and relationships, preparing the data structure in advance for faster final detection and tracking decisions, thereby reducing real-time processing delays

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically adjusts detection parameters and processing depth based on scene context and formation types. The system modifies its operational parameters to optimize the balance between tracking accuracy and processing speed, enabling real-time performance by adapting computational resources to the specific requirements of each detection scenario

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230377335A1Key person recognition in immersive video
Publication Date: 2023.11.23 INTEL CORP
  • US20230377335A1 patent drawing
  • US20230377335A1 patent drawing
  • US20230377335A1 patent drawing

AI summary

Techniques related to key person recognition in multi-camera immersive video attained for a scene are discussed. Such techniques include detecting predefined person formations in the scene based on an arrangement of the persons in the scene, generating a feature vector for each person in the detected formation, and applying a classifier to the feature vectors to indicate one or more key persons in the scene.