Video Clip Screening for Accurate Description Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video clip description technologies generate inaccurate descriptions due to video clips detected by the video-clip detecting model not having strong correlation with the current video, leading to inaccuracy in final descriptions.

Innovation Solution

A method and apparatus for generating descriptions of video clips that includes a video-clip screening module to screen video proposal clips, selecting only those with strong correlation, and a video-clip describing module to provide descriptions, with both modules being jointly trained to improve accuracy and diversity of descriptions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If all video clips detected by the video-clip detecting model are used for description generation, then the quantity of descriptions is increased, but the accuracy of descriptions deteriorates due to weak correlation between detected clips and current video

Engineering Contradiction:
Improvequantity of video clip descriptionsVSAvoidaccuracy of video clip descriptions
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the video description generation process into two distinct modules: a video-clip detecting model that identifies candidate clips, and a video-clip screening module that filters clips based on correlation strength with the current video. This segmentation allows the system to first generate a comprehensive set of candidate descriptions (maintaining quantity) and then filter for high-quality descriptions (improving accuracy), resolving the contradiction between quantity and accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a video-clip screening module as an intermediary between the video-clip detecting model and the description generation process. This screening module acts as a mediator that evaluates the correlation between detected clips and the current video, selectively passing only highly correlated clips to the description generator. This intermediary mechanism enables the system to maintain high description quantity while ensuring high accuracy through selective filtering.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a video-clip screening module is introduced to filter video clips, then the accuracy of descriptions is improved, but the device complexity increases

Engineering Contradiction:
Improveaccuracy of video clip descriptionsVSAvoidcomplexity of video description model
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the video-clip detecting model and the video-clip screening module into a unified video description model with joint training. By combining these modules and training them together using a unified loss function, the patent reduces the complexity that would arise from separate training processes and multiple independent models. The merged architecture allows gradients to flow through both modules during training, enabling coordinated optimization and simplifying the overall system structure while maintaining high description accuracy.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of manufacture

If video-clip detecting model and video-clip describing model are trained separately in two stages, then the training process is simplified, but the correlation between detected clips and descriptions deteriorates

Engineering Contradiction:
Improveease of model trainingVSAvoidcorrelation between video clips and descriptions
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent combines the previously separate two-stage training process into a single joint training process. The video-clip detecting model and video-clip describing model are trained simultaneously with a unified loss function that optimizes both clip detection accuracy and description quality together. This merging of training processes ensures that the detected clips and generated descriptions are tightly correlated, as the gradient updates during training coordinate both modules to work together effectively, resolving the contradiction between training simplicity and correlation strength.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The joint training process implements a feedback mechanism where the performance of the describing model influences the training of the detecting model and vice versa. The unified loss function provides feedback signals that propagate through both modules, allowing each module to adjust its parameters based on the overall system performance. This feedback loop ensures that the detecting model learns to identify clips that are most beneficial for accurate description generation, thereby maximizing correlation between detected clips and generated descriptions.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11615140B2Method and apparatus for detecting temporal action of video, electronic device and storage medium
Publication Date: 2023.03.28 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11615140B2 patent drawing
  • US11615140B2 patent drawing
  • US11615140B2 patent drawing

AI summary

A method includes screening, by a video-clip screening module in a video description model, a plurality of video proposal clips acquired from a video to be analyzed, to acquire a plurality of video clips suitable for description. The plural video proposal clips acquired from the video to be analyzed may be screened by the video-clip screening module to acquire the plural video clips suitable for description; and then, each video clip is described by a video-clip describing module, thus avoiding description of all the video proposal clips, only describing the screened video clips which have strong correlation with the video and are suitable for description, removing the interference of the description of the video clips which are not suitable for description in the description of the video, guaranteeing the accuracy of the final descriptions of the video clips, and improving the quality of the descriptions of the video clips.