Video Cut Grouping Using Temporal Context and Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video processing technologies fail to appropriately group video cuts, leading to an unclear understanding of the video structure, especially when cuts are not similarly grouped.

Innovation Solution

A moving image processing apparatus and method that determines the similarity between video cuts using feature amounts generated from extraction images, grouping similar cuts together and generating feature amounts for dissimilar cuts to enhance understanding of the video structure by prioritizing images with later time codes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If cuts are grouped based on traditional similarity methods, then the grouping process is simple, but the cut structure understanding becomes unclear when cuts are not appropriately grouped

Engineering Contradiction:
Improvecut structure understanding accuracyVSAvoidgrouping process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the parameter used for cut grouping from traditional visual similarity to temporal position-based similarity. By using time code information and comparing temporal positions of cuts, the system achieves more accurate cut structure understanding while maintaining a relatively simple grouping process. The feature amount generation unit generates feature amounts based on temporal positions rather than complex visual analysis.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If all images in each cut group are processed equally, then the processing is straightforward, but temporal context information is lost

Engineering Contradiction:
Improvetemporal context informationVSAvoidimage processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies local quality by differentiating the processing of images based on their temporal positions. The image extraction unit preferentially extracts images with later time codes from each cut group, giving different weights to different images within the same cut group. This ensures that temporal context information is preserved while maintaining efficient processing by not treating all images equally.

Inventive Principle:
Principle #3Local quality

3Reliability

If extraction images are selected without considering time codes, then the selection process is simple, but the temporal sequence and context of video content is not preserved

Engineering Contradiction:
Improvetemporal sequence preservationVSAvoidimage extraction operation
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent applies preliminary action by pre-establishing the time code information for each image before the extraction process. The image extraction unit uses this pre-existing time code information to preferentially select images with later time codes, ensuring temporal sequence preservation without requiring complex real-time analysis during extraction.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8682078B2Moving image processing apparatus, moving image processing method, and program
Publication Date: 2014.03.25 SONY GROUP CORP
  • US8682078B2 patent drawing
  • US8682078B2 patent drawing
  • US8682078B2 patent drawing

AI summary

There is provided a moving image processing apparatus, including a similarity determination unit configured to determine a degree of similarity between a subsequent cut and first and second cut groups based on feature amounts generated from extraction images of the first cut group included in a moving image and feature amounts generated from extraction images of the second cut group included in the moving image, a cut grouping unit configured to group the subsequent cut into a similar cut group similar to the subsequent cut, which is one of the first and second cut groups, when the subsequent cut is similar to the first or second cut group, and group the subsequent cut into a third cut group when the subsequent cut is not similar to either of the first and second cut groups, a feature amount generation unit configured to compare extraction images extracted from the third cut group with the extraction images extracted from the first and second cut groups when the subsequent cut is not similar to either of the first and second cut groups, and generate feature amounts of the third cut group, and an image extraction unit configured to preferentially extract an image with a later time code of the moving image from images included in each cut group, thereby obtaining extraction images of each cut group.