Video Segment Descriptors for Accurate Content Tagging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in accurately identifying and accessing relevant content due to inaccurate and ambiguous tags associated with videos, leading to the need for improved methods to quickly find content of interest.
Innovation Solution
Content is fragmented into segments based on topical coherence criteria, and descriptors are generated using techniques such as salient tag detection, teaser detection, and optical character recognition (OCR) to accurately represent each segment, enhancing the ability to identify and access relevant content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional keyword tags are used to describe video content, then users can access video content through simple tagging, but the tags are frequently inaccurate and ambiguous making it difficult for users to distinguish and find relevant videos
Solution Approach 1:
The video content is divided into multiple segments based on topical coherence criteria, with each segment receiving its own descriptor. This segmentation allows for more granular and accurate representation of different topics within a single video, improving content identification accuracy while maintaining ease of access through segment-level descriptors.
Solution Approach 2:
The patent introduces descriptors as an intermediary layer between the video content and user queries. These descriptors are generated through multiple techniques including salient text detection, OCR, and teaser analysis, serving as accurate mediators that bridge user search intent with relevant video segments.
2Quantity of substance
If multiple videos are made available to users through social media and content distribution networks, then users have access to innumerable pieces of content, but users have to sift through or watch a large number of videos that are not of interest
Solution Approach 1:
The system performs preliminary analysis of video content to generate descriptors before users search for content. By pre-processing videos to extract salient text, perform OCR, identify teasers, and create segment descriptors in advance, the system enables rapid retrieval and filtering when users search, significantly reducing the time needed to find relevant content among large volumes of available videos.
3Ease of manufacture
If uploaders choose ambiguous terms to describe videos, then videos can be tagged with simple keywords, but it becomes difficult for users to determine the actual subject of the video
Solution Approach 1:
The system enables videos to describe themselves automatically through multiple analysis techniques. Salient text detection extracts meaningful terms directly from the video content, OCR reads text from on-screen graphics and banners, and teaser analysis identifies representative content. This self-service approach eliminates reliance on uploader-provided tags, ensuring accurate content description without adding manual tagging complexity.
Data Source
AI summary
An apparatus, method, system and computer-readable medium are provided for generating one or more descriptors that may potentially be associated with content, such as video or a segment of video. In some embodiments, a teaser for the content may be identified based on contextual similarity between words and/or phrases in the segment and one or more other segments, such as a previous segment. Text/characters may serve as a candidate descriptor(s). In some embodiments, one or more strings of characters or words may be compared with (pre-assigned) tags associated with the content, and if it is determined that the one or more strings or words match the tags within a threshold, the one or more strings or words may serve as a candidate descriptor(s). One or more candidate descriptor identification techniques may be combined.


