Video Tagging Automation via Image Audio Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video processing technologies face challenges in accurately classifying and tagging video content, which affects the precision of video recommendations, as they rely on manual or inefficient methods to extract relevant information from videos.

Innovation Solution

A method and apparatus that acquire target video element information, extract relevant video clips through image or audio recognition, determine keywords based on preset relationships, and match these keywords with tag information to enhance the accuracy and efficiency of video classification and recommendation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual or inefficient methods are used to extract relevant information from videos, then the process is simpler to implement, but the accuracy of video tagging and classification deteriorates

Engineering Contradiction:
Improveaccuracy of video taggingVSAvoidcomplexity of extraction process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual information extraction methods with automated image recognition and audio recognition systems. The extraction unit automatically identifies video clips containing target objects or keywords by analyzing image frames and audio signals, eliminating the need for manual review and significantly improving tagging accuracy while maintaining manageable system complexity through standardized recognition algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The video processing system performs self-service by automatically extracting relevant information from video content without requiring external manual intervention. The extraction unit autonomously analyzes video frames and audio tracks to identify target clips, and the tagging unit automatically generates classification tags based on recognized content, enabling the system to serve itself in the information extraction process.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If automated recognition methods are used to extract video clips, then the accuracy of video tagging improves, but the processing time and computational resources increase

Engineering Contradiction:
Improveaccuracy of video taggingVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts only the essential and relevant portions of video content for analysis. The extraction unit identifies and isolates specific video clips containing target objects or keywords rather than processing the entire video, and further extracts only the critical image frames and audio segments needed for recognition, significantly reducing processing time while maintaining high tagging accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies partial action by performing recognition operations on selected video clips rather than the complete video duration. The extraction unit filters out irrelevant segments and focuses computational resources only on portions of the video that contain target content, achieving accurate tagging with reduced processing time by avoiding excessive analysis of non-essential video portions.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If multiple video element information types are analyzed, then the richness of tagging information improves, but the device complexity increases

Engineering Contradiction:
Improverichness of tagging informationVSAvoidcomplexity of processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements multi-functionality by designing a unified video processing system that handles multiple types of video element information through integrated units. The extraction unit processes both image frames and audio signals, while the tagging unit generates comprehensive tags based on combined analysis of visual and auditory content, enabling rich tagging information with a single versatile system rather than separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges image recognition and audio recognition processes into a unified video analysis framework. The extraction unit combines analysis of visual elements from video frames and auditory elements from audio tracks to identify target clips, and the tagging unit integrates information from both modalities to generate comprehensive classification tags, achieving rich tagging information through combined processing rather than separate systems.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10824874B2Method and apparatus for processing video
Publication Date: 2020.11.03 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US10824874B2 patent drawing
  • US10824874B2 patent drawing
  • US10824874B2 patent drawing

AI summary

Embodiments of the present disclosure disclose a method and apparatus for processing a video. A specific embodiment of the method comprises: acquiring a target video and target video element information of the target video; extracting, based on the target video element information, a target video clip from the target video; obtaining, based on a preset corresponding relationship between video element information and a keyword determining method for a video clip, a keyword representing a category of the target video clip; and matching the keyword and with preset tag information set to obtain tag information of the target video clip, and associating and storing the target video clip and the tag information.