Game Video Summarization Using Multimodal Metadata Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating an effective summary of computer simulation videos automatically is difficult, and manual summarization is time-consuming.

Innovation Solution

A system utilizing a machine learning engine to process audio-video data, aligning metadata with audio and video modalities, and generating a concise video summary that includes metadata perceptible in the video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual summarization is used to create effective video summaries, then the quality of summary is improved, but the time consumption increases

Engineering Contradiction:
Improvesummary qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the video into multiple clips based on metadata analysis, audio features, and video features. The ML engine processes different segments independently to identify important portions, allowing automated summarization to achieve quality comparable to manual methods while significantly reducing time consumption through parallel processing of segmented content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces metadata as an intermediary element that bridges automated processing and human-perceptible information. Metadata including game event data, emotion, audio, and video features are processed by the ML engine to automatically identify and select important video segments, enabling automated summarization to produce high-quality results without manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated summarization is used to reduce time consumption, then the processing speed is improved, but the summary quality deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidsummary quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent employs dynamic processing where the ML engine continuously analyzes metadata, audio, and video features to adaptively select video segments. The system dynamically adjusts which clips to include based on real-time analysis of multiple feature types, allowing automated processing to maintain high summary quality while achieving fast processing speeds through efficient feature-based selection.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes multiple parameters simultaneously including metadata types (game event data, emotion, audio, video features), audio features (pitch, power, acoustic events), and video features (scene changes, object detection). By analyzing and weighting multiple parameters, the ML engine enables automated summarization to produce high-quality summaries that capture important moments without manual intervention.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If multiple metadata types are integrated into the video summary, then the viewer engagement is improved, but the system complexity increases

Engineering Contradiction:
Improveviewer engagementVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal ML engine that processes multiple types of metadata (game event data, emotion, audio, video features) through a single integrated system. The engine universally handles diverse input types and produces unified video summaries with overlaid metadata, enabling enhanced viewer engagement while managing system complexity through consolidated processing architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12505673B2Multimodal game video summarization with metadata
Publication Date: 2025.12.23 SONY INTERACTIVE ENTERTAINMENT LLC
  • US12505673B2 patent drawing
  • US12505673B2 patent drawing
  • US12505673B2 patent drawing

AI summary

Video and audio from a computer simulation are processed by a machine learning engine to identify candidate segments of the simulation for use in a video summary of the simulation. Text input is then used to reinforce whether a candidate segment should be included in the video summary. Metadata can be added to the summary showing game summary information.