Dense Video Captioning via Region-Sequence Language Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video description technologies provide inadequate, superficial descriptions that fail to accurately capture the rich content of videos, limiting the effectiveness of video search and evaluation due to reliance on single sentence captions and lack of region-sequence level annotations.

Innovation Solution

A dense video captioning system that generates multiple sentence descriptions for each region of a video using a combination of visual, region-sequence, and language components, including a trained computational model for mapping words to video regions and a sequence-to-sequence learning framework for generating informative and diverse region-sequences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional one-sentence captioning is used for video description, then the system complexity is low, but the description accuracy and information completeness deteriorate

Engineering Contradiction:
Improvedescription accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the video description task into multiple components: region proposal generation, region feature extraction, and sentence generation. Instead of generating a single caption for the entire video, the system divides the video into multiple regions and generates descriptions for each region, thereby improving description accuracy while managing complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-dimension approach (one sentence for entire video) to a multi-dimensional approach by introducing spatial dimension (regions within frames) and temporal dimension (sequences of regions across frames). This allows the system to capture rich video content through region-sequence level annotations, improving description completeness

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If region-sequence level annotations are used for dense video captioning, then the description completeness improves, but the annotation cost and computational expense increase

Engineering Contradiction:
Improveinformation completenessVSAvoidannotation cost
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The patent performs preliminary region proposal generation using trained computational models before detailed annotation. By pre-identifying candidate regions and their sequences, the system reduces the manual annotation workload while maintaining region-sequence level detail, making dense video captioning more feasible

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses trained computational models to generate region proposals and features that copy or replicate expert annotation quality. The models learn from annotated data and can automatically generate region-sequence level descriptions, reducing the need for expensive manual annotation while preserving information completeness

Inventive Principle:
Principle #26Copying

3Measurement precision

If global visual representation is used for video description, then the processing speed is fast, but the mapping accuracy from sentences to visual content deteriorates

Engineering Contradiction:
Improvemapping accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the global visual representation into region-level visual features. Instead of mapping sentences to a single global video representation, the system maps sentences to specific region sequences, improving the accuracy of the sentence-to-visual mapping while maintaining processing efficiency through localized feature extraction

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11790644B2Techniques for dense video descriptions
Publication Date: 2023.10.17 INTEL CORP
  • US11790644B2 patent drawing
  • US11790644B2 patent drawing
  • US11790644B2 patent drawing

AI summary

Techniques and apparatus for generating dense natural language descriptions for video content are described. In one embodiment, for example, an apparatus may include at least one memory and logic, at least a portion of the logic comprised in hardware coupled to the at least one memory, the logic to receive a source video comprising a plurality of frames, determine a plurality of regions for each of the plurality of frames, generate at least one region-sequence connecting the determined plurality of regions, apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video. Other embodiments are described and claimed.