Video Tagging via Audio-Video Correlation Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video content tagging systems require manual intervention and lack wide-ranging functionality, making it difficult to create relevant and comprehensive tags for live-stream or recorded video content without user input.
Innovation Solution
A system and method that utilize natural language understanding (NLU) processing and image recognition to automatically generate tags by analyzing audio and video data, assigning tags based on correlation factors exceeding threshold values, without user intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual tagging is used, then tagging accuracy can be ensured, but user engagement is reduced and productivity is lowered
Solution Approach 1:
The system performs automatic tagging without requiring manual user input. The computer device analyzes video data, audio data, and metadata independently to generate tags, allowing the system to serve itself rather than relying on user intervention for each tagging operation.
Solution Approach 2:
The patent replaces manual mechanical tagging operations with automated computational analysis. Instead of users manually creating tags, the system uses machine learning models to analyze video frames, audio content, and metadata to automatically generate and assign tags to video portions.
2Productivity
If automatic tagging is implemented, then productivity is improved, but tagging reliability deteriorates due to lack of verification
Solution Approach 1:
The system generates multiple candidate tags from different data sources (video analysis, audio analysis, metadata) and uses confidence scores to evaluate each candidate. The tag with the highest confidence score is selected, providing a feedback mechanism that ensures the most reliable tag is chosen while maintaining high productivity through automated processing.
Solution Approach 2:
The system employs multiple analysis functions working together: video frame analysis, audio content analysis, metadata parsing, and confidence score calculation. This multi-functional approach allows the system to maintain high productivity while improving reliability through cross-verification of tags from multiple sources.
3Adaptability or versatility
If comprehensive analysis is performed on both audio and video data, then tagging versatility is improved, but device complexity increases
Solution Approach 1:
The system divides the tagging process into separate independent modules: video data analysis, audio data analysis, metadata analysis, and tag generation. Each module processes its specific data type independently, allowing the system to maintain high versatility through comprehensive analysis while managing complexity through modular architecture.
Data Source
AI summary
Systems and methods for tagging video content are disclosed. A method includes: receiving a video stream from a user computer device, the video stream including audio data and video data; determining a candidate audio tag based on analyzing the audio data; establishing an audio confidence score of the candidate audio tag based on the analyzing of the audio data; determining a candidate video tag based on analyzing the video data; establishing a video confidence score of the candidate video tag based on the analyzing of the video data; determining a correlation factor of the candidate audio tag relative to the candidate video tag; and assigning a tag to a portion in the video stream based on the correlation factor exceeding a correlation threshold value and at least one of the audio confidence score exceeding an audio threshold value and the video confidence score exceeding a video threshold value.


