Storyboard Generation from Image Frames via Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating graphical summaries of image collections or videos is slow and resource-intensive, making it difficult on devices with limited resources, and often requires extensive user input.
Innovation Solution
A computer-implemented method using machine-learned models to select and crop key image frames based on object detection, arranging them into a storyboard with minimal computational demand, allowing for efficient generation on devices like smartphones.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional methods are used to generate graphical summaries of image collections or videos, then comprehensive visual coverage is achieved, but the process becomes slow and resource-intensive
Solution Approach 1:
The patent extracts only the most important visual elements (key objects) from video frames using object detection, rather than processing or displaying all frames. By identifying and isolating significant objects and their temporal relationships, the system creates concise storyboards that capture essential video content while requiring minimal computational resources and time.
Solution Approach 2:
The system automatically generates storyboards by detecting objects, determining their significance, and arranging them chronologically without requiring manual frame selection or extensive user input. The automated object detection and temporal relationship analysis enable the system to self-organize video content into meaningful summaries, dramatically improving productivity while reducing resource consumption.
2Ease of operation
If traditional summary generation methods are used, then detailed visual information is preserved, but extensive user input is required including image selection and arranging
Solution Approach 1:
The system automatically performs object detection, significance assessment, and chronological arrangement of video frames without requiring user intervention for image selection or positioning. The automated storyboard generator analyzes objects across frames, determines their importance, and organizes them into a coherent narrative sequence, eliminating time-consuming manual operations while preserving essential visual information.
Solution Approach 2:
The patent extracts key objects and their temporal relationships from the video content, automatically identifying which frames and elements are most significant. This extraction process replaces manual image selection and arranging, allowing the system to generate meaningful storyboards quickly without user input while maintaining focus on the most important visual elements.
3Measurement precision
If comprehensive video analysis is performed to ensure accurate object location detection, then detection precision is improved, but computational complexity increases
Solution Approach 1:
The patent extracts and focuses computational resources on detecting and tracking specific objects of interest across video frames, rather than analyzing all visual elements in detail. By identifying key objects and their temporal relationships, the system achieves accurate object location detection with reduced computational complexity, as the model concentrates on significant elements rather than processing the entire video content comprehensively.
Data Source
AI summary
The present disclosure provides systems and methods that generate a summary storyboard from a plurality of image frames. An example computer-implemented method can include inputting a plurality of image frames into a machine-learned model and receiving as an output of the machine-learned model, object data that describes the respective locations of a plurality of objects recognized in the plurality of image frames. The method can include generating a plurality of image crops that respectively include the plurality of objects and arranging two or more of the plurality of image crops to generate a storyboard.


