Video Tagging via Audio-Video Correlation Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video content tagging systems require manual intervention and lack wide-ranging functionality, making it difficult to create relevant and comprehensive tags for live-stream or recorded video content without user input.

Innovation Solution

A system and method that utilize natural language understanding (NLU) processing and image recognition to automatically generate tags by analyzing audio and video data, assigning tags based on correlation factors exceeding threshold values, without user intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tagging is used, then tagging accuracy can be ensured, but user engagement is reduced and productivity is lowered

Engineering Contradiction:
Improvetagging accuracyVSAvoiduser engagement
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs automatic tagging without requiring manual user input. The computer device analyzes video data, audio data, and metadata independently to generate tags, allowing the system to serve itself rather than relying on user intervention for each tagging operation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical tagging operations with automated computational analysis. Instead of users manually creating tags, the system uses machine learning models to analyze video frames, audio content, and metadata to automatically generate and assign tags to video portions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automatic tagging is implemented, then productivity is improved, but tagging reliability deteriorates due to lack of verification

Engineering Contradiction:
Improvetagging speedVSAvoidtagging accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system generates multiple candidate tags from different data sources (video analysis, audio analysis, metadata) and uses confidence scores to evaluate each candidate. The tag with the highest confidence score is selected, providing a feedback mechanism that ensures the most reliable tag is chosen while maintaining high productivity through automated processing.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system employs multiple analysis functions working together: video frame analysis, audio content analysis, metadata parsing, and confidence score calculation. This multi-functional approach allows the system to maintain high productivity while improving reliability through cross-verification of tags from multiple sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If comprehensive analysis is performed on both audio and video data, then tagging versatility is improved, but device complexity increases

Engineering Contradiction:
Improvetagging functionalityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system divides the tagging process into separate independent modules: video data analysis, audio data analysis, metadata analysis, and tag generation. Each module processes its specific data type independently, allowing the system to maintain high versatility through comprehensive analysis while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10714144B2Corroborating video data with audio data from video content to create section tagging
Publication Date: 2020.07.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10714144B2 patent drawing
  • US10714144B2 patent drawing
  • US10714144B2 patent drawing

AI summary

Systems and methods for tagging video content are disclosed. A method includes: receiving a video stream from a user computer device, the video stream including audio data and video data; determining a candidate audio tag based on analyzing the audio data; establishing an audio confidence score of the candidate audio tag based on the analyzing of the audio data; determining a candidate video tag based on analyzing the video data; establishing a video confidence score of the candidate video tag based on the analyzing of the video data; determining a correlation factor of the candidate audio tag relative to the candidate video tag; and assigning a tag to a portion in the video stream based on the correlation factor exceeding a correlation threshold value and at least one of the audio confidence score exceeding an audio threshold value and the video confidence score exceeding a video threshold value.