Video Caption Removal via Synchronized Image and Text Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In video editing, users face the inconvenience of having to separately remove captions and images from specified sections of a video, increasing the number of steps required for editing.

Innovation Solution

An information processing apparatus that synchronizes audio, images, and captions, allowing users to specify a playback time section for removal, which automatically removes corresponding partial captions and images within that section.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the apparatus only removes images in a specified section of a video, then the image removal operation is simple, but the caption removal operation becomes complex and requires separate steps

Engineering Contradiction:
Improveimage removal operationVSAvoidcaption removal operation
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent merges the image removal operation and caption removal operation into a single integrated process. When a user specifies a section for removal, the system simultaneously removes both the image data and the corresponding caption data in one operation, eliminating the need for separate caption removal steps and reducing operational complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The video editing apparatus is designed to perform multiple functions through a single operation: it can remove images, remove captions, or remove both simultaneously based on the user's selection. This multi-functional capability allows the same removal interface and process to handle different editing needs without requiring separate specialized operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If the apparatus removes both images and captions simultaneously in a specified section, then the number of editing steps is reduced, but the processing complexity increases

Engineering Contradiction:
Improveediting efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the video data structure into distinct image data portions and caption data portions, each with associated timing information. This segmentation allows the apparatus to independently identify and remove the appropriate portions based on the specified time section, managing processing complexity through structured data organization rather than monolithic processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis to identify the temporal correspondence between image data and caption data before execution. By pre-processing the video structure to map caption timing with image timing, the apparatus prepares the data relationships in advance, enabling simultaneous removal operations to proceed efficiently without complex real-time coordination.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11651167B2Information processing apparatus and non-transitory computer readable medium
Publication Date: 2023.05.16 FUJIFILM BUSINESS INNOVATION CORP
  • US11651167B2 patent drawing
  • US11651167B2 patent drawing
  • US11651167B2 patent drawing

AI summary

An information processing apparatus includes a processor configured to acquire video data that enables playback of a video in which audio, an image, and a caption are chronologically synchronized, receive a section of a playback time of the video, the section being to be removed, and remove a partial caption that corresponds to the audio in the received section and that is at least a portion of the caption from the image in the received section.