Video Frame Classification for Target-Duration Media Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently adjusting the duration of media content, such as video commercials, while maintaining the message and effectiveness, requiring careful frame selection for different audiences and purposes.

Innovation Solution

A system utilizing machine learning and generative AI to classify frames based on relevance, selecting key frames and regenerating audio signals to create customized media content with a target duration, including techniques for voiceover generation and language adaptation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual frame selection is used to adjust media content duration, then customization precision is improved, but processing time and labor cost increase

Engineering Contradiction:
Improveframe selection precisionVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs automatic frame classification and selection without requiring manual intervention. The machine learning model autonomously analyzes video frames, identifies relevant content, and reconstructs the media content with the target duration, eliminating the need for human editors to manually select frames while maintaining high precision in frame selection.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical frame selection with an automated machine learning-based classification system. The neural network model processes video frames automatically, assigning relevance scores and selecting appropriate frames for reconstruction, thereby substituting human labor with an intelligent automated system that achieves both precision and efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated frame classification is used, then processing speed is improved, but relevance detection accuracy may worsen

Engineering Contradiction:
Improveprocessing speedVSAvoidrelevance detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where the machine learning model continuously refines its frame classification based on the reconstructed media content quality. The relevance scores assigned to frames are adjusted based on the overall coherence and quality of the reconstructed video, ensuring that automated processing achieves both high speed and high accuracy in maintaining content relevance.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If media content is shortened to target duration, then adaptability is improved, but information completeness may worsen

Engineering Contradiction:
Improveduration adaptabilityVSAvoidkey information loss
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system extracts only the most relevant frames from the original media content based on machine learning classification. By identifying and selecting the critical frames that carry the most important information, the system reconstructs the media content at the target duration while preserving key information, effectively taking out only the essential elements needed for the shortened version.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of frame selection based on relevance scores generated by the machine learning model. Instead of using fixed time intervals or random sampling, the system dynamically adjusts frame selection parameters based on the computed relevance of each frame, ensuring that the shortened media content maintains information completeness by prioritizing frames with higher relevance scores.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12541948B2Frame classification to generate target media content
Publication Date: 2026.02.03 ROKU INC
  • US12541948B2 patent drawing
  • US12541948B2 patent drawing
  • US12541948B2 patent drawing

AI summary

Aspects of the disclosed technology provide solutions for processing media content to generate customized media content of a target duration. An example method can include receiving media content of a first duration. The media content may include a plurality of video frames. The method can include steps for receiving one or more parameters, which may include a target duration, classifying each of the plurality of video frames of the media content based on a relevance level of each frame, and generating a target media content of the target duration based on the classification of the plurality of video frames of the media content of the first duration. Systems and machine-readable media are also provided.