Video Cluster Analysis for Duplicate Detection and Storage Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video processing approaches lead to inefficiencies and redundancies due to frequent uploads and storage of the same videos, reducing user experience and causing unnecessary data transmissions and storage inefficiencies.
Innovation Solution
A system and method for defining and analyzing video clusters based on video image frames, where a subset of frames is compared to existing clusters, determining matches within an allowable deviation, and either associating the video with an existing cluster or creating a new one, thereby reducing redundant processing and storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional video processing approaches are used, then videos can be uploaded and shared freely, but redundant uploads and storage of duplicate videos occur frequently
Solution Approach 1:
The system performs preliminary video fingerprinting and cluster matching before allowing video uploads. By pre-establishing video clusters and comparing incoming videos against these clusters, the system proactively identifies duplicates and prevents redundant storage, thereby resolving the contradiction between upload flexibility and data redundancy
Solution Approach 2:
The patent introduces video clusters as an intermediary layer between individual video uploads and storage. Instead of directly storing each uploaded video, the system routes videos through cluster matching, where duplicate videos are identified and linked to existing clusters rather than creating new storage entries, thus eliminating redundancy while preserving upload functionality
2Reliability
If multiple copies of the same video are stored, then all uploads are preserved, but storage efficiency and resource usage decrease
Solution Approach 1:
The system merges duplicate videos into unified video clusters. When a video is identified as a duplicate, it is merged with the existing cluster rather than creating a separate storage entry. This combining approach preserves all video versions through cluster associations while consuming minimal storage space, resolving the contradiction between video preservation and storage efficiency
Solution Approach 2:
The patent uses video fingerprinting to create lightweight digital copies (hashes) of videos for comparison purposes. These fingerprint copies enable efficient duplicate detection without storing full video copies, allowing the system to maintain reliability through accurate identification while minimizing actual storage consumption of duplicate content
3Productivity
If conventional video processing is used, then all videos are processed individually, but processing efficiency and resource utilization are reduced
Solution Approach 1:
The system merges processing operations by handling multiple videos through cluster-based batch processing. Instead of individually processing each video upload, the system groups videos into clusters and processes them collectively, significantly improving processing throughput while reducing the computational energy wasted on redundant operations across multiple independent processing pipelines
4Productivity
If video clusters are defined and analyzed, then redundant videos are identified and grouped, but system complexity increases
Solution Approach 1:
The patent extracts the essential identifying features of videos (fingerprints/hashes) and separates them from the full video content for cluster matching purposes. This extraction approach enables efficient duplicate detection by working with compact feature representations rather than complete videos, improving detection efficiency while avoiding the complexity of analyzing entire video files
Data Source
AI summary
Systems, methods, and non-transitory computer-readable media can identify a first video represented based on a first set of image frames. A first subset of image frames can be extracted from the first set of image frames. The first subset of image frames can be compared to one or more image frames associated with a collection of video clusters. It can be determined that at least a threshold quantity of image frames in the first subset matches, within an allowable deviation, at least some image frames associated with a first video cluster included the collection of video clusters. The first video cluster can be defined to include the first video.


