Representative Image Frame Extraction via Pose Displacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for extracting representative image frames from video data, such as those based solely on feature amount, struggle to accurately summarize and distinguish actions in video images, especially when captured with a fixed angle of view, leading to difficulties in selecting frames that adequately represent the action and differentiate it from other actions.

Innovation Solution

An image frame extraction apparatus and method that acquires video images, extracts features, analyzes these features to identify candidate frames, calculates displacement in a shape space to select frames with significant pose changes, and uses learning models to classify and select frames based on class assignments and probability scores, ensuring the selected frame is representative and distinguishable.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a representative image frame is selected depending solely on the size of the feature amount of each image frame, then the selection process is simple and fast, but it is not necessarily possible to extract an image frame that appropriately represents the video image data

Engineering Contradiction:
Improveselection speedVSAvoidrepresentation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the selection parameter from simple feature amount magnitude to pose displacement in shape space. This transformation allows the system to evaluate frames based on meaningful pose variations rather than arbitrary feature magnitudes, thereby improving representation accuracy while maintaining computational efficiency through mathematical transformation of existing pose data.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Instead of selecting frames with maximum feature amounts, the patent inverts the approach by selecting frames based on maximum displacement from reference poses in shape space. This inversion transforms the selection criterion from absolute magnitude-based to relative difference-based, enabling better capture of action characteristics.

Inventive Principle:
Principle #13The other way round (Inversion)

2Device complexity

If a representative image frame is selected solely by the size of the feature amount of each image frame, then the selection process is straightforward, but it is difficult to ensure that the image frame is sufficiently distinguished from other actions

Engineering Contradiction:
Improveselection complexityVSAvoidaction distinguishability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent transforms the selection parameter from feature amount magnitude to pose displacement in shape space, which inherently provides better action distinguishability. The shape space representation captures essential pose variations that are characteristic of different actions, enabling reliable differentiation without significantly increasing system complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent moves the selection criterion from a one-dimensional feature amount metric to a multi-dimensional pose displacement measurement in shape space. This dimensional transformation enables the system to capture subtle action differences that are not reflected in simple feature magnitudes, thereby improving action distinguishability.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Use of energy by moving object

If a representative image frame is selected solely by the size of the feature amount of each image frame, then the selection process requires minimal computation, but it is difficult to appropriately select an image frame that summarizes the action in a straightforward manner

Engineering Contradiction:
Improvecomputational energyVSAvoidaction summarization quality
Core Design Contradiction:
Use of energy by moving objectVSLoss of information

Solution Approach 1:

The patent changes the evaluation parameter from feature amount to pose displacement in shape space, which provides better action summarization quality. This parameter transformation leverages existing pose estimation results and applies a mathematical transformation, avoiding the need for additional complex computations while significantly improving the quality of action representation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11989943B2Image frame extraction apparatus and image frame extraction method
Publication Date: 2024.05.21 RAKUTEN GROUP INC
  • US11989943B2 patent drawing
  • US11989943B2 patent drawing
  • US11989943B2 patent drawing

AI summary

Disclosed herein is an image frame extraction apparatus that acquires a video image; extracts features of each of a plurality of image frames of the acquired video image; analyzes the extracted features of each of the plurality of image frames, and extracts candidates of a representative frame from the plurality of image frames; and, for each of the extracted candidates of the representative frame, calculates a displacement in a shape space of a pose of an object in the image frame with respect to a reference pose, and select the representative frame from the candidates of the representative frame based on the calculated displacement in the shape space.