Video Captioning Model for Real-Time Motion Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio description techniques for video content are labor-intensive, time-consuming, and limited to pre-recorded content, making them inefficient and costly, and do not effectively provide real-time descriptions for users with visual or cognitive impairments.

Innovation Solution

A data processing apparatus and method that uses a video captioning model to detect predetermined motions in video images and generate description data, including text, audio, and image data, to provide real-time visual descriptions of content, enabling improved accessibility for users with impairments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual creation of descriptive transcripts and human voice actors are used, then audio description can be provided for pre-recorded content, but the process becomes labor-intensive, time-consuming, and costly

Engineering Contradiction:
Improvequality of audio descriptionVSAvoidefficiency of description generation
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces the mechanical system of manual transcript creation and human voice recording with an automated computer vision system that processes video frames, detects motions, and generates captions through algorithmic analysis, thereby eliminating labor-intensive operations while maintaining description quality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the video content itself to provide the necessary information for description generation through automated motion detection and caption creation, eliminating the need for external human operators to create and record descriptions

Inventive Principle:
Principle #25Self-service

2Reliability

If manual audio description techniques are used, then descriptions can be created for pre-recorded content, but real-time description capability is lost

Engineering Contradiction:
Improveaccuracy of visual descriptionVSAvoidtime delay in description provision
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements dynamics by enabling the system to operate in real-time as video frames are processed sequentially, with motion detection and caption generation occurring continuously during video playback rather than requiring pre-processing, thus eliminating time delays while maintaining description accuracy

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary action by pre-training the motion detection algorithms and caption generation models during system setup, enabling rapid real-time processing during actual video description without requiring manual preparation for each video content

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated video captioning is implemented, then real-time description generation is achieved, but system complexity increases

Engineering Contradiction:
Improvespeed of description generationVSAvoidcomplexity of processing system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the complex processing task into distinct modular components: video frame input, motion detection module, caption generation module, and output delivery, allowing each component to be independently optimized and managed while achieving high-speed real-time processing

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4421755A1Apparatus and methods for content description
Publication Date: 2024.08.28 SONY INTERACTIVE ENTERTAINMENT LLC
  • EP4421755A1 patent drawingFigure 1
  • EP4421755A1 patent drawingFigure 2a~3
  • EP4421755A1 patent drawingFigure 4~6

AI summary

A data processing apparatus for determining description data for describing content comprises: a video captioning model to receive an input comprising at least video images associated with the content, wherein the video captioning model is trained to detect one or more predetermined motions of one or more animated objects in the video images and determine one or more captions in dependence on one or more of the predetermined motions, one or more of the captions comprising respective caption data comprising one or more words for describing one or more of the predetermined motions, the respective caption data comprising one or more of audio data, text data and image data; and output circuitry to output description data in dependence on one or more of the captions.