Video Title Generation Using Text Vector Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating video titles are inefficient and labor-intensive, particularly for short videos, as they often require manual annotation and may not accurately represent the content, leading to inconsistent quality and difficulty in scaling with large volumes of data.

Innovation Solution

A method and apparatus for generating video titles by obtaining and analyzing user interaction data to determine central text information with the highest similarity to the video content, using techniques such as text vector conversion and machine learning models like BERT or ERNIE to select the most representative text from various sources like bullet comments, comments, or subtitles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation methods are used to generate video titles, then the titles can be created with human oversight, but the process becomes labor-intensive and inefficient

Engineering Contradiction:
Improvetitle accuracyVSAvoidtitle generation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automatic title generation by having the video content itself provide the necessary information through extracted text features, eliminating the need for manual annotation while maintaining title quality through automated similarity assessment

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical annotation work with an automated computational system that extracts text from video content, converts it to vectors, and automatically selects titles based on similarity metrics

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual title generation is used, then some level of quality control is maintained, but the process cannot scale with large volumes of video data

Engineering Contradiction:
Improvetitle quality consistencyVSAvoidscalability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The automated title generation system serves multiple functions: extracting text from various sources (subtitles, comments, bullet screens), converting text to vectors, calculating similarities, and selecting titles - all within a single scalable framework that handles large volumes of diverse video content

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses parameter-based automated decision making by converting text to vector representations and using similarity threshold parameters to consistently select titles, replacing subjective manual judgment with objective measurable parameters that scale reliably

Inventive Principle:
Principle #35Parameter changes

3Productivity

If existing automated title generation methods are used, then efficiency is improved, but the titles may not accurately represent the video content

Engineering Contradiction:
Improvetitle generation efficiencyVSAvoidcontent representation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system uses similarity calculation as a feedback mechanism to assess how well candidate titles represent video content, continuously selecting and refining title choices based on quantitative similarity measurements between extracted text and video content vectors

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces text vector conversion and similarity calculation as intermediary processes between raw video content and final title selection, using these intermediate representations to objectively measure and ensure accurate content representation

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4209929A1Video title generation method and apparatus, electronic device and storage medium
Publication Date: 2023.07.12 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • EP4209929A1 patent drawingFigure 1
  • EP4209929A1 patent drawingFigure 2
  • EP4209929A1 patent drawingFigure 3~4

AI summary

Provided are a video title generation method and apparatus, an electronic device and a storage medium. The present disclosure relates to a technical field of video, and in particular to a technical field of short video. The method includes: obtaining (S110) a plurality of pieces of optional text information, for a first video file; determining (S120) central text information, from the plurality of pieces of optional text information, the central text information being optional text information with the highest similarity to content of the first video file; and determining (S130) the central text information as a title of the first video file. The present disclosure may determine an interest point in an original video file according to user's interactive behavior data on the original video file, and clip the original video file based on the interest point to obtain a plurality of clipped video files, namely, short videos.