Automatic Video Tagging via Audio Text Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual video tagging is time-consuming and costly, limiting the diversity and quality of training data for video analysis models, particularly in applications like action recognition, where preparing sufficient tagged video clips is necessary for high-performance training.

Innovation Solution

An automatic video tagging method that converts the audio signal from video clips into text sequences, using semantic analysis to select keywords for tagging specific video segments, reducing computational resource consumption and increasing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual video tagging is used, then tagging accuracy can be maintained, but time consumption and cost increase significantly

Engineering Contradiction:
Improvetagging accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces audio signals as an intermediary medium between video content and tagging. Instead of directly analyzing video frames for tagging, the system extracts audio signals from videos, converts them to text via speech-to-text, and uses text-based semantic analysis to generate tags. This intermediary approach reduces the computational complexity of video analysis while maintaining tagging quality through the semantic information preserved in audio transcripts.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If manual video tagging is used, then tag quality can be ensured, but computational resource consumption increases

Engineering Contradiction:
Improvetag qualityVSAvoidcomputational resource consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the audio component from video files for the purpose of tagging, rather than processing the entire video data. By separating and utilizing only the relevant audio information, the system significantly reduces computational resource requirements while still achieving effective video tagging through speech-to-text conversion and natural language processing of the extracted audio content.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If automatic tagging methods are used, then efficiency increases, but tagging quality may deteriorate

Engineering Contradiction:
Improvetagging efficiencyVSAvoidtagging quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent replaces manual mechanical tagging processes with automated computational methods. Specifically, it substitutes human annotators with a pipeline involving speech-to-text conversion and natural language processing algorithms. This automation maintains high tagging efficiency while preserving quality through the semantic accuracy of modern NLP techniques in extracting meaningful keywords from audio transcripts.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11682415B2Automatic video tagging
Publication Date: 2023.06.20 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11682415B2 patent drawing
  • US11682415B2 patent drawing
  • US11682415B2 patent drawing

AI summary

In an approach, a processor extracts an audio signal from a video clip. A processor converts the audio signal into a text sequence. A processor selects a first set of keywords from the text sequence, the first set of keywords corresponding to a first audio segment of the audio signal. A processor tags a target video segment of the video clip with the first set of keywords, the target video segment corresponding to the first audio segment.