Video Captioning Model for Real-Time Motion Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio description techniques for video content are labor-intensive, time-consuming, and limited to pre-recorded content, making them inefficient and costly, and do not effectively provide real-time descriptions for users with visual or cognitive impairments.
Innovation Solution
A data processing apparatus and method that uses a video captioning model to detect predetermined motions in video images and generate description data, including text, audio, and image data, to provide real-time visual descriptions of content, enabling improved accessibility for users with impairments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual creation of descriptive transcripts and human voice actors are used, then audio description can be provided for pre-recorded content, but the process becomes labor-intensive, time-consuming, and costly
Solution Approach 1:
The patent replaces the mechanical system of manual transcript creation and human voice recording with an automated computer vision system that processes video frames, detects motions, and generates captions through algorithmic analysis, thereby eliminating labor-intensive operations while maintaining description quality
Solution Approach 2:
The system enables self-service by allowing the video content itself to provide the necessary information for description generation through automated motion detection and caption creation, eliminating the need for external human operators to create and record descriptions
2Reliability
If manual audio description techniques are used, then descriptions can be created for pre-recorded content, but real-time description capability is lost
Solution Approach 1:
The patent implements dynamics by enabling the system to operate in real-time as video frames are processed sequentially, with motion detection and caption generation occurring continuously during video playback rather than requiring pre-processing, thus eliminating time delays while maintaining description accuracy
Solution Approach 2:
The system performs preliminary action by pre-training the motion detection algorithms and caption generation models during system setup, enabling rapid real-time processing during actual video description without requiring manual preparation for each video content
3Productivity
If automated video captioning is implemented, then real-time description generation is achieved, but system complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the complex processing task into distinct modular components: video frame input, motion detection module, caption generation module, and output delivery, allowing each component to be independently optimized and managed while achieving high-speed real-time processing
Data Source
Figure 1
Figure 2a~3
Figure 4~6
AI summary
A data processing apparatus for determining description data for describing content comprises: a video captioning model to receive an input comprising at least video images associated with the content, wherein the video captioning model is trained to detect one or more predetermined motions of one or more animated objects in the video images and determine one or more captions in dependence on one or more of the predetermined motions, one or more of the captions comprising respective caption data comprising one or more words for describing one or more of the predetermined motions, the respective caption data comprising one or more of audio data, text data and image data; and output circuitry to output description data in dependence on one or more of the captions.